Gesture classification with magnetomyography

MMG with ADFMR and OPMs addresses the limitations of existing gesture recognition by providing accurate and generalized hand gesture recognition through direct muscle activity detection and machine learning, enhancing performance and reducing variability.

WO2026050775A1PCT designated stage Publication Date: 2026-03-05SONERA INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/044535
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-31
Filing Date
2025-09-02
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Existing gesture recognition technologies face limitations such as visual occlusion, reliance on good lighting, and high variability between individuals, making it challenging to develop a generalized model for hand gesture recognition across diverse populations.

Method used

Utilizing magnetomyography (MMG) with miniaturized magnetic sensors like ADFMR and OPMs to detect muscle activity directly, employing a hierarchical classification approach with machine learning to identify gestures, and using an array of magnetometers on the arm and wrist to capture magnetic signals.

Benefits of technology

Achieves accurate and generalized gesture recognition across participants and sessions, reducing variability and improving performance beyond existing EMG systems, with potential for high-density channel information and reduced environmental noise interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025044535_05032026_PF_FP_ABST
    Figure US2025044535_05032026_PF_FP_ABST
Patent Text Reader

Abstract

Methods and apparatuses (devices, methods, etc.) for determining gestures, e.g., hand gestures, from a wrist and / or forearm worn array of magnetometers. These methods and apparatuses may use a trained machine-learning agent to identify gestures from magnetic signals that have been preprocessed to optimize gesture detection. The trained machine learning agent may apply a hierarchical classification that first identifies a first category of gestures from a second category of gestures and then classifies between the individual gestures within each category. The methods and apparatuses described herein may be of particular use with acoustically driven ferromagnetic resonance (ADFMR) sensors.
Need to check novelty before this filing date? Find Prior Art

Description

GESTURE CLASSIFICATION WITH MAGNETOMYOGRAPHYCLAIM OF PRIORITY

[0001] This patent application claims priority to U.S. provisional patent application no. 63 / 689,700, titled “GESTURE CLASSIFICATION WITH MAGNETOMYOGRAPHY” and filed on August 31, 2024, which is herein incorporated by reference in its entirety.BACKGROUND

[0002] Hand gesture recognition, or the detection of expressive positioning of the hands, has numerous applications including rehabilitation, sign language, and human-computer interfaces (HCIs). A common technology used to implement gesture recognition is computer vision, the method of extracting out joint positions using video streams from infrared and / or visual cameras. However, camera-based gesture recognition has a number of limitations including visual occlusion of the hands or fingers, requirement of good lighting, and reliance on movement to detect activity.

[0003] In contrast, physiological measures, such as surface electromyography (sEMG), can overcome these limitations as they directly detect muscle activity. Though sEMG requires a wrist- or arm mounted device to capture the signals, the hands can be fully obscured and changes in isometric contraction, such as grip strength, can be properly measured. Gesture recognition using sEMG has both research and commercial applications. However, due to physiological variability between individuals, developing a generalized model across a population is incredibly challenging. Despite the difficulties, recent studies have shown notable progress with deep learning approaches: >95% accuracy between sessions of the same subject across 50 gestures, and >92% accuracy between 4800 subjects across 9 gestures. However, the latter study concluded that there is a theoretical upper limit of around 95% accuracy across nine gestures achievable with 16 sensor sEMG on the wrist. Though groundbreaking to demonstrate that it is possible to generalize across so many subjects, 5% error is not practical for commercial applications. There is a clear need to record physiological data using a modality with less variability between participants to enable generalized gesture recognition with out-of-the-box usage.

[0004] Magnetic sensing of muscles, or magnetomyography (MMG), is a novel modality that has been demonstrated to also provide a direct measurement of muscle activity. MMG is understudied due to difficulties in capturing the signals, often requiring extremely sensitive magnetic sensors like superconducting quantum interference devices (SQUIDs) or optically pumped magnetometers (OPMs), and a shielded room to remove earth’s magnetic field and - 1 -SG Docket No.: 14941-711.600various environmental magnetic noise sources.

[0005] In contrast to EMG, MMG does not rely on skin or tissue conductivity, a large contributor of variability in EMG, potentially providing an undistorted view of muscle activity and more high frequency content. In addition, many magnetic sensors for MMG are becoming miniaturized into chip-form factor allowing for higher density of channels rivaling high-density EMG, further increasing the information content available for gesture recognition. Finally, MMG has been shown to directly reflect muscle activity similar to EMG. As a result, MMG may provide less cross-participant and cross-session variability allowing for a more generalized model in gesture recognition.

[0006] Described herein are methods and apparatuses that may provide a more generalized gesture recognition modality. These methods and apparatuses may provide MMG for generalized gesture recognition.SUMMARY OF THE DISCLOSURE

[0007] Described herein are methods and apparatuses (e.g., devices, systems, etc.) for sensing magnetic signals from a muscle, e.g., from outside of the body, and identifying or determining a gesture based on the sensed magnetic signals. As used herein a gesture may refer to a sequence of muscle movements, and may include a complete set of movements, e.g., beginning from a neutral position and ending in the same or a different neutral position, or apportion of a set of movements, e.g., a constricting movement (clenching, pinching, grasping, flexing, etc., and / or a release movement (finger release, e.g., index release, middle release, etc.). Thus the gestures may be compound gestures, including both constricting movements and release movements.

[0008] Any of these methods may include: sensing a plurality of magnetic signals from one or more muscles using an array of magnetometers worn on an arm and / or wrists; dividing the plurality of magnetic signals into windows having a window size of between about 200 and about 1200 ms; applying a trained machine learning agent to identify a gesture from the plurality of windows, wherein the trained machine learning agent is configured to apply a hierarchical classification that first identifies a first category of gestures from a second category of gestures and then classifies between the individual gestures within each category; and outputting the identified gesture.

[0009] In particular, the array of magnetometers may comprise an array of acoustically driven ferromagnetic resonance (ADFMR) sensors. In some cases the array of magnetometers comprises an array of optically pumped magnetometer (0PM) sensors.

[0010] Any of these methods may include comprising removal of bad channels from the- 2 -SG Docket No.: 14941-711.600array of magnetometers. In general, these methods and apparatuses may include 8 or more sensors (e.g., 9 sensors, 10 sensor, 11 sensors, 12 sensors, etc.). The sensors may be arranged in a loop around the forearm and / or wrist. In some examples the method and / or apparatus may be configured to sense from both the arm and the wrist.

[0011] Any of these methods may include filtering the plurality of magnetic signals. In particular, these methods and apparatuses may be configured to filter to allow both high frequency and lower frequency content, between about 30-300 Hz bandwidth.

[0012] Any of these methods or apparatuses may be configured to resample the filtered plurality of magnetic signals. For example, resampling at 1 KHz. Any of these methods may include digitizing the plurality of magnetic signals (e.g., sampling them at a sampling frequency) either before or after filtering.

[0013] As described herein, the first category of gestures may comprise normal gestures. For example, the first category of gestures may be normal gestures comprises contracting gestures such as (but not limited to): clench, flex, Index Pinch, Middle Pinch, swipe right, swipe left, swipe up, and swipe down. The second category of gestures may comprise release gestures, e.g., relaxing gestures, such as but not limited to: extend, index release, middle release.

[0014] In general, any of these methods may include adjusting (e.g., dynamically adjusting) the window size while applying the trained machine learning agent. In some cases the window size may be adjusted and the machine learning agent applied multiple times to the different window sizes. In some cases the window may be a running window (e.g., scanning through the sensed magnetic signals. In some cases the sensed magnetic signals may be divided into windows that are normalized (by stretching or compressing) prior to applying the trained machine learning agent.

[0015] In general, the identified gesture(s) may be output by displaying, storing or transmitting the identified gesture. In some cases the identified gesture(s) may be used as a control input, e.g., to control operation of a device or additional apparatus; e.g., input as a control signal into a computer system, as data into a computing system (e.g., typing or other gestural data, etc.).

[0016] For example, a method as described herein may include: sensing a plurality of magnetic signals from one or more muscles using an array of magnetometers worn on an arm and / or wrists; filtering the plurality of magnetic signals; dividing the plurality of magnetic signals into windows having a window size of between about 200 and about 1200 ms; applying a trained machine learning agent to identify a gesture from the plurality of windows, wherein the trained machine learning agent is configured to apply a hierarchical classification - 3 -SG Docket No.: 14941-711.600that first identifies a first category of gestures from a second category of gestures and then classifies between the individual gestures within each category; repeating the application of the trained machine learning agent while adjusting the window size; and outputting the identified gesture.

[0017] Also described herein are apparatuses configured to perform any of these methods. For example, an apparatus may comprise: an array of magnetometers worn on an arm and / or wrists; one or more processors; and a memory coupled to the one or more processors, the memory storing computer-program instructions, that, when executed by the one or more processors, perform any of the computer-implemented methods described herein.

[0018] The data shown herein includes data collected from multiple sessions of MMG from the wrist and forearm of participants during a hand gesture task. This data demonstrates that gesture recognition as performant as EMG can be performed using single-participant MMG data, and provides a generalized model across participants and sessions. The results suggest that MMG may be a more ideal modality to use for generalized gesture recognition

[0019] Described herein are methods and apparatuses for magnetic sensing of muscles, or magnetomyography (MMG), as a novel approach for gesture recognition. MMG is a modality that has been demonstrated to provide a direct measurement of muscle activity similar to sEMG but without requiring direct skin contact. MMG is understudied due to the inaccessibility of appropriate sensors, typically requiring extremely expensive magnetic sensors like superconducting quantum interference devices (SQUIDs) or optically pumped magnetometers (OPMs). These sensors also need a magnetically shielded room to remove earth’s magnetic field and various environmental noise sources to properly operate. However, recent advancement in sensor technology has improved accessibility of MMG, suggesting its potential applicability beyond research or medical settings.

[0020] MMG has several advantages over EMG in gesture recognition. MMG does not rely on skin or tissue connectivity as the human body is largely transparent to magnetic fields. This in turn suggests that MMG provides an undistorted view of muscle activity as the signal is not distorted by the conductivity of tissue or electrode properties, both large contributors to the variability experienced by sEMG. In addition, many magnetic sensors for MMG can be miniaturized into chip-form factor allowing for higher density of channels rivaling high- density EMG, further increasing the information content available for gesture recognition. As a result, MMG may provide less cross-participant and cross-session variability allowing for a more generalized model in gesture recognition.

[0021] Consequently, we collected multiple sessions of MMG from the wrist and forearm across 31 participants during a hand gesture task. First, we demonstrate gestures can be- 4 -SG Docket No.: 14941-711.600readily classified using MMG with machine learning approaches. Next, we assess crosssession and cross-participant classification accuracies to establish baseline variability. Finally, we present findings that show part of the variability is due to the limitations of the available system rather than the modality of MMG. The results suggest that MMG may be a more ideal modality to use for generalized gesture recognition.

[0022] All of the methods and apparatuses described herein, in any combination, are herein contemplated and can be used to achieve the benefits as described herein.BRIEF DESCRIPTION OF THE DRAWINGS

[0023] A better understanding of the features and advantages of the methods and apparatuses described herein will be obtained by reference to the following detailed description that sets forth illustrative embodiments, and the accompanying drawings of which:

[0024] FIG. l is a graph showing cross-validation accuracy for single-session models. Distributions are over session level cross-validated models trained on the shown channel subsets. Chance-level is 1 / 9.

[0025] FIGS. 2A-2E show examples of confusion matrices across all sessions. FIG. 2A shows a confusion matrix for all channels. FIG. 2B shows a confusion matrix for writs channels. FIG. 2C shows a confusion matrix for forearm channels. FIG. 2D shows a confusion matrix for Y channels, and FIG. 2E shows a confusion matrix for Z channels.

[0026] FIGS. 3A-3I are confusion matrices for 3 participants with low performance (rows). FIGS. 3A-3C correspond to participant 15, FIGS 3D-3F show participant 20, and FIGS. 3G-3I show participant 22.

[0027] FIG. 4 shows an example of Cross-validation accuracy across number of sensors used. Distributions are over session-level cross-validated models trained on the shown number of sensors (both axes). For each session and sensor count we used multiple random samplings of sensors. Chance-level is 1 / 9.

[0028] FIG. 5 is a graph showing cross-validation accuracy across time windows. Distributions are over session-level cross-validated models trained on the shown trial length from movement onset (detected by a time-alignment algorithm). Chance-level is 1 / 9.

[0029] FIGS. 6A-6C show confusion matrices across sessions for 3 time windows. FIG. 6A shows 200 ms window, FIG. 6B shows 500 ms window and FIG. 6C shows 1200 ms window.

[0030] FIGS. 7A-7I show confusion matrices of individual participants (rows). FIGS. 7A-7C show confusion matrices for participant 08 (good), FIGS. 7D-7F show confusion - 5 -SG Docket No.: 14941-711.600matrices for participant 20 (bad), row 3 and FIGS. 7G-7I show confusion matrices for participant 22 (bad).

[0031] FIG. 8 shows a graph of cross-validation accuracy across different frequency bands. Distributions are over session-level cross-validated models trained on the respective frequency band. Chance-level is 1 / 9.

[0032] FIG. 9 is a graph showing cross-session accuracy across test sessions.Distributions are over each session of each subject. In the case of blue plots a single session was used for training, while in the case of orange plots 2 sessions were used. Chance-level is 1 / 9.

[0033] FIGS. 10 A- 10C show confusion matrices for cross-session trainings using 2 training sessions.

[0034] FIGS. 11A-11F show confusion matrices for subject 21 (FIGS. 11A-11C) and subject 33 (FIGS. 11D-11F). 2 training sessions were used in each case, and the confusion matrices are aggregated over all 3 test sessions.

[0035] FIG. 12A-12C shows confusion matrices for the 3 (test) sessions of subject 08. 2 training sessions were used in each case.

[0036] FIG. 13 is a graph showing cross-subject accuracy across test subjects using various channel subsets. Distributions are over the mean (across sessions) accuracies of each test subject. Chance-level is 1 / 9.

[0037] FIGS. 14A-14F show confusion matrices for subject 15 session 01 (FIGS. 14A- 14C) and subject 30 session 01 (FIGS. 14D-14F).

[0038] FIGS. 15A-15C show confusion matrices for cross-subject trainings.

[0039] FIGS. 16A-16I show confusion matrices of 3 bad subjects (FIGS. 16A-16C).FIGS. 16A-16C shows 21 subject 22 (FIGS. 3D-3DF) corresponds to subject 27.

[0040] FIGS. 17A-17C illustrate accuracy over the recording duration. The horizontal axis denotes the ordered trial index within the recording. Note that a sliding window of 20% (about 80, depending on the session) of all trials was used.

[0041] FIG 18 shows the accuracy over the recording duration for 5 bad sessions. The horizontal axis denotes the ordered trial index within the recording point.

[0042] FIG. 19 shows cross-subject accuracy across test subjects using various frequency bands. Distributions are over the mean (across sessions) accuracies of each test subject. Chance-level is 1 / 9.

[0043] FIG. 20 shows cross-subject accuracy across test subjects depending on the size of the training set. Distributions are over the mean (across sessions) accuracies of each test subject. Chance level is 1 / 9.- 6 -SG Docket No.: 14941-711.600

[0044] FIG. 21 A illustrates accuracy over the entire trial length. In FIGS. 21 A-21B, median accuracy and standard deviation across subjects is shown.

[0045] FIG. 2 IB is a graph of the accuracy of 5 example bad sessions.

[0046] FIG. 22A1 shows an example of a layout of an MSR (22A1) and placement of OPMs and sEMG pairs (FIG. 22A2). FIG. 22A3 shows examples of gestures presented to the participant to prompt a trial start (text was also shown with pictures).

[0047] FIG. 22B shows an example of 0PM channel traces (top 8 around wrist, bottom 8 around the forearm) and sEMG pairs during an index pinch and middle pinch trial.

[0048] FIG. 23 A-23B show examples of average amplitude spectral density for a single 0PM (blue) and EMG (red) channel on the dorsal side of the wrist (FIG. 23 A). Solid lines show the average across all Index Pinch trials and dashed lines averages for all Neutral times. SNRs of the two channels (bottom). Note the similarity between the spectral content of the signals.

[0049] FIGS. 24A-24F illustrate examples of a single-session cross-validation results. In all figures, distributions are over session-level cross-validated accuracies. Chance-level is 0.11.

[0050] FIG. 24A show a per-gesture cross-validation accuracy for single-session models using all channels. Release gestures have significantly lower performance than other gestures.

[0051] FIG. 24B show a cross-validation accuracy across various channel subsets. Using either wrist-only or forearm-only channels results in lower performance compared to using both 0PM bands, with forearm being lower than wrist. Only keeping one of the two channels of each sensor also results in lower accuracies, but there is not a difference between the axes.

[0052] FIG. 24C shows an example of cross-validation accuracy scaling with the number of sensors in each band. Performance improves steadily while adding more sensors, and using the wrist 0PM sensors always has higher accuracy than the forearm sensors. Distributions are over session-level cross-validated models trained on the shown number of sensors (both axes). For each session and sensor count we used multiple random samplings of sensors.

[0053] FIG. 24D shows an example of a curve ft to the median cross-validation classification error (100 - accuracy) from (c). The curve ft to the wrist channels is Er = 2.24 + 27.78 / N096, and to the forearm channels is Er = 0 + 37 / N0 8.

[0054] FIG. 24E shows an example of cross-validation accuracy scaling with the window size used since movement onset. Performance improves steadily with larger windows, with statistically significant differences between each subsequent window size. Distributions are over session-level cross-validated models trained on the shown trial length from movement onset (detected by our time-alignment algorithm).

[0055] FIG. 24F shows an example of cross-validation accuracy across various frequency- 7 -SG Docket No.: 14941-711.600bands. The (62.5, 125) Hz band contains the most information content with every other frequency band having less information. Distributions are over session-level cross-validated models trained on the respective frequency band.

[0056] FIGS. 25A-25F show examples of cross-session and cross-participant results. Chance-level is 0.11. FIG. 25 A shows cross-session accuracy across test sessions. Using 2 training sessions instead of 1 substantially improves performance, as well using all channels. Distributions are over each session of each participant. In the case of blue plots a single session was used for training, while in the case of orange plots 2 sessions were used. FIG. 25B shows cross-participant accuracy across test participants depending on the size of the training set. Performance improves steadily when adding more participants to the training set. Distributions are over the mean (across sessions) accuracies of each test participant. FIG. 25C shows curve fit to the median classification errors (100-accuracy) from (b). The following curve was found: Er = 0.0 + 14.33 / N03. Black bars denote the + / - SEM around the median, and the orange band denotes the area between the curve fit to the upper and lower SEM values - giving estimates of the error around the median. FIG. 25D shows per-gesture accuracy for the crossparticipant model using all channels. Swipe down and Swipe left have the worst performance. FIG. 25E shows cross-participant accuracy across test participants using various channel subsets. Using all channels has substantially higher performance than all other combinations. Distributions are over the mean (across sessions) accuracies of each test participant. FIG. 25F shows cross-participant accuracy across test participants using various frequency bands. Best generalization performance is achieved with the 62-125Hz and 187-300Hz ranges, but this is lower than using the whole 30-300Hz range.

[0057] FIGS. 26A-26F shows variability across recording duration, trial timing, and between participants. FIGS. 26A-26C show cross-subject accuracy over the recording duration for 3 channel subsets; all, wrist, and forearm. All cases show a decrease in accuracy over the recording duration except the all-channels good sessions, good refers to the top 14, and bad refers to the bottom 14 sessions according to test accuracy. The horizontal axis denotes the ordered trial index within the recording. Note that a sliding window of 20% (about 80, depending on the session) of all trials was used. FIG. 26D shows cross-subject accuracy over the entire trial length. Accuracy increases rapidly in the beginning of the trial plateauing around 1400ms. Median accuracy and standard deviation across participants. FIGS. 26E-26F shows aggregated confusion matrices of 2 participants with low performance. These were obtained from MPF+logreg models trained on individual sessions (using all channels). The confusion matrices show that there are differences between participants in terms of which gestures are poorly classified.- 8 -SG Docket No.: 14941-711.600DETAILED DESCRIPTION

[0058] Described herein are methods and apparatuses (e.g., devices and systems) for using magnetomyography (MMG) to detect gestures. For example, an apparatus may include an array of MMG sensors in which each sensor includes an ADFMR.

[0059] In general, the methods and apparatuses described herein may be used to effectively classify different gestures, providing enhanced accuracy, particularly with respect to prior art systems.

[0060] In general, these methods and apparatuses may be particularly well suited for use with acoustically driven ferromagnetic resonance (ADFMR) sensor that may be used as the MMG sensor(s). In general, an ADFMR sensor may include a magnetostrictive material (e.g., magnet), a substrate (e.g., a piezoelectric substrate), and an acoustic drive portion. The acoustic drive portion may consist of one of many different types of acoustic resonators, including surface acoustic wave (SAW) resonators, film bulk acoustic resonators (FBAR), and bulk acoustic resonators (e.g., high-tone bulk acoustic resonators, HBAR). The acoustic drive portion may generate an acoustic wave at or near the ferromagnetic resonance of a magnetostrictive element. The acoustic drive portion may include one or more (e.g., a pair of) transducers, such as electrodes, that activate a piezoelectric element to generate acoustic waves. The magnetostrictive element receives the wave as a signal and changes its properties in response to the received wave. Detection circuitry detects the change in the property of the magnetostrictive element and that change is used to determine a result. The detection circuitry may also measure the change in the wave generated by the acoustic drive portion.

[0061] Thus, a single ADFMR sensor and / or array ADFMR sensor unit may include: a piezoelectric substrate comprising the main body of the ADFMR sensor component; an input transducer (e.g., an interdigitated transducer, IDT) that generates a surface acoustic wave (SAW) from an electric signal using the piezoelectric effect which propagates along the substrate; a ferromagnetic film, along the substrate that enables absorption of magnetic fields by the SAW as it propagates thereby modifying the SAW; and an output transducer (e.g., output IDT) that converts the modified SAW to an electric signal. This modified SAW may then be used to determine the strength of the magnetic field. The ADFMR sensor may have many variations that include one, or multi-sensing capabilities.

[0062] The apparatuses described herein may be more accurate and effective than traditional EMG sensor for gesture detection.

[0063] All of the experiments described herein were performed inside a magnetically shielded room (MSR; Magnetic Shield Corporation, MuROOM) with inside dimensions of- 9 -SG Docket No.: 14941-711.6001.3x1.3x2 meters to prevent sensor saturation and contamination from ambient noise sources. The shielded room had around 25,000 fold attenuation of residual fields at DC and up to 8000 fold attenuation at AC. The room was degaussed (demagnetized) by applying a decreasing sinusoidal field using embedded coils prior to each recording to remove magnetization of the walls and guarantee maximum rejection of ambient fields. A projector mounted on the outside of the room displayed a screen on the inside of the MSR through a hole in the wall.

[0064] Optically pumped magnetometer (0PM) sensors were used for MMG measurements. As mentioned above, any of these sensors may instead be ADFMR sensors. Surface electromyography (sEMG) was simultaneously recorded from each subject with single-use, bipolar, MRI safe, pre-gelled sEMG electrodes with non-ferrous contacts (10 mm gelled area diameter) with non-ferrous leads. sEMG was mainly used to validate the MMG data and ensure we were properly recording muscle activity. Finally, IR-based hand tracking camera was used to allow monitoring of the participant’s gestures from outside the MSR and simultaneously collect supplemental spatial hand position data during the gesture task.

[0065] Each 0PM was configured to be sensitive along two orthogonal axes, and were positioned as such that one axis was tangential to the arm, and the other radially into the arm. The OPMs recorded from both axes of each sensor for each recording. The 0PM sensors were housed in a 3-D printed band containing 8 sensors each. Two sensor bands, one for the wrist and one for the forearm, were affixed to the participant’s arm. Multiple size circumference bands were printed to accommodate variations in the circumference of the participants wrists and forearms. The wrist sensor band was placed 4 cm up the arm from the wrist joint, and the forearm sensor band was placed at one third the length of the forearm down from the elbow. The sensors sat in a circular orientation, where one sensor was seated at the mid-line of the anterior face of the forearm and the other 7 sensors sat equidistant from each other around the circumference of the arm.

[0066] 16 sEMG electrodes were also adhered to the participant’s arm. The 16 electrodes were paired into 8 bipolar channels, 4 pairs per wrist and forearm respectively. Each electrode pair was made by adhering two of the sEMG electrodes together with a 2.5cm center-to-center inter-electrode distance. For both the wrist and forearm, two electrode pairs were placed at the mid-line of the anterior and posterior faces of the forearm, and two additional pairs were placed equidistant between those initial pairs, covering the forearm between the two 0PM bands at approximately 90 degrees apart from each other.

[0067] The 0PM and sEMG sensors provided analog signals which were digitized and streamed using 16 or 24-bit NI-DAQ ADC modules (one NI-9202 and two NI-9205 modules in a cDAQ-9185 chassis, National Instruments) at a sampling rate of 2 kHz per channel. The- 10 -SG Docket No.: 14941-711.600data from the Leap Motion Controller was streamed through Unity and directly saved in a separate file at around 100 Hz sampling rate. All data was collected using custom Python and C# code.

[0068] Each participant was seated in a non-ferrous chair inside the MSR after being briefed about the experiment. All sensors were subsequently affixed to their right arm which was then placed on a plastic armrest with foam padding for the duration of each session.

[0069] The task consisted of eleven distinct hand gestures: Clench, Extend, Flex, Index Pinch, Middle Pinch, Swipe Right, Swipe Left, Swipe Up, Swipe Down, and Thumb Tap. The participant was cued to perform a specific gesture and then cued again to return to a "neutral" position. Each cue consisted of the image and name of the gesture as well as the number for the Index and Middle Pinch gestures, the participant was cued to extend all of their fingers before returning to a neutral position to simulate the act of completing the pinch. The order of gestures was pseudo-randomly chosen such that each gesture was performed the same number of times.

[0070] A single session consisted of 50 trials of each gesture. Individual trials lasted 2.5 seconds and participants were instructed to hold their final position until the neutral cue was presented. Neutral positions also lasted 2.5 seconds between each trial. Each session began with 30 seconds of rest in which the participant held the Neutral position to obtain a baseline recording. Additionally, the participant was provided with the option to listen to media through the speaker directed into the room.

[0071] The data was collected from participants, ranging from 21 to 64 years old. There were both female and male participants tested, including right handed and left handed (1 participant), and 1 ambidextrous participant. Recruitment and data collection for human research followed IRB approved protocols (Advarra). Participants were recruited via an online form sent to their company email, and interested individuals were scheduled and consented before their first data collection session. Demographic information (name, email address, phone number) and biometric information (bio-signals, arm length, age, weight, height, etc.) were housed and recorded separately following IRB approved confidentiality protocols. Participants were instructed to undergo three data collection sessions, in which they performed the task over the period of an hour and a half (40 minutes per session). After each session the participant was compensated with a gift card. All participants that consented to research were given a handedness questionnaire to determine their dominant arm, as well as a metal screen to ensure that no metallic objects in or on the participant’s body entered the magnetically shielded room (MSR). Participants were given detailed instructions on how to perform the task, which included a period for them to practice the gestures. Before each- 11 -SG Docket No.: 14941-711.600session, the researcher overseeing the data collection reviewed each gesture with the participant. Using a printout of the same gesture images that were displayed as cues during the task, the overseeing researcher demonstrated each gesture with the respective image, then performed the gestures alongside the participant. During this overview, the researcher guided the participants’ gestures, helping to ensure the individual’s consistency. These tips include real-life analogies and equivalents for each gesture. For example, "the Thumb Tap should be a quick up-and-down, like tapping a button on a touchscreen phone," or “Clench should be a squeeze as hard as a firm handshake." The researcher simultaneously demonstrated the expected gesture with each instruction. Finally, participants were asked to remove their shoes and change into scrubs and swept with a metal detector to fully confirm the absence of ferrous materials before they began each session.Preprocessing

[0072] As the aim of this study was to focus on physiological variability with MMG, not environmental or sensor / equipment variability, sessions which exhibited high noise (likely due to sensors misbehaving, abnormally high environmental noise penetrating the MSR, or unknown metal on the person) were removed. Excluded sessions from all analyses: 03-01 (no chn ordering info) 04-01 (no chn ordering info) 04-02 (no chn ordering info) 04-03 (wrist-only) 04-04 (double-wrist) 04-05 (half data) 11-01 (half data) 13-03 (no metadata) 25- 01 (half data). All sessions that were collected before switching from a 3 to 2.5 second trial length in cross-subject and cross-session analyses were excluded. Excluding these sessions served to remove any variability in task timing, and because these early sessions exhibited more variability in sensor placement and how participants performed gestures due to continuous improvements in the data collection process during the early phases of the study. A data cut-off for all cross-subject analyses was set because training these models takes a long time and should have a consistent set of sessions.

[0073] A common preprocessing pipeline may be applied for classification of gestures. The number and type of preprocessing steps may differ between different analyses, and these are specified in the respective sections. The preprocessing may include the following steps.Filtering

[0074] Raw data may be bandpass filtered between 30-300 Hz, e.g., with a 5th order Butterworth filter. A custom automatic peak finding algorithm may be used to find noise peaks in the Power Spectral Density (PSD) of each session, and these may be removed with a 5th order Butterworth bandstop filter. The algorithm may include a sliding window along the PSD of each channel, where if within a certain frequency bin the z-scored power exceeded a threshold of 5, it was identified as a peak. The statistics for z-scoring were calculated based- 12 -SG Docket No.: 14941-711.600on a surrounding window of 20Hz. The algorithm was run in a 2-pass mode, removing peaks detected in the first pass and then running a second pass, again removing newly detected peaks. Finally, all data was resampled to 1000Hz.Bad channel detection

[0075] After filtering, bad channels may be automatically detected, for example, using the generalized ESD test (ESD), e.g., with a significance level of 0.5, and the maximum number of channels to mark as bad may be set to a threshold (e.g., 25% of all channels). Optionally, these channels might be later removed. The algorithm was applied separately to OPM and EMG channels.Normalization

[0076] Channels may be independently normalized to 0 mean and unit variance.Epoching

[0077] Data may be epoched according to the gesture prompt onset time. Each epoch may last a few (e.g., between 0.5 and 10 seconds, e.g., 2.5 seconds), which may correspond to the gesture offset time. In some cases the Index Pinch and Middle Pinch epochs were immediately followed by Index Release and Middle Release epochs, with no rest period inbetween, due to the nature of these gestures.Bad epoch detection

[0078] The generalized ESD test may be run to automatically remove outlier (bad) epochs. This may be run separately for each gesture. The algorithm may run separately for OPM and EMG channels, and the union of identified epochs may be removed. The significance threshold may be set (e.g., to between 0.05 and 0.25, e.g., 0.1), and the maximum number of epochs to remove may be set, e.g., to between about 1% and 10% (e.g., 5%) of all epochs. The number of epochs (trials) across gestures may be equalized to the gesture with the lowest number of trials.Trial alignment

[0079] A movement onset may be detected based on the time-frequency representation of each trial. After movement onset time was identified, the preceding time period may be removed for each trial. The maximum allowed movement onset time may be set (e.g., to 1 second post gesture prompting), with the minimum being 200 ms. The trial alignment may not be run for Index Release and Middle Release gestures, as these immediately followed gesture offset. The final trial length may be cropped (e.g., to 1.5 seconds) from detected movement onset for all gestures.

[0080] To determine the information content present across each recording, the frequency specific entropy may be calculated using the Hilbert-Huang transform. The Hilbert-Huang- 13 -SG Docket No.: 14941-711.600transform may be calculated with intrinsic mode functions (IMFs) using empirical mode decomposition (EMD). For each IMF awe find the Hilbert transform H:

[0081] Using the Hilbert transform we can then define the analytic signal s,:We can then express the signal in the time frequency domain as:where a, ft) and cotft) are instantaneous amplitude and frequencies calculated with the analytic signal. Finally, we can calculate entropy across frequency of he frequency specific information content H(a>) -.Variability (PCA or VAE)

[0082] To quantify how much variability is present between each session the principal component (PC) space of each session may be calculated, e.g., using all preprocessed MMG channels using the number of PCs that account for 95% of the variance in the data. Every other session’s data may be projected onto the PC projection and calculated the silhouette score to measure how well the different trials can cluster.

[0083] Feature extraction may then be used. The feature extraction technique used may depend on the analysis. In some cases the preprocessed epoched data may be directly used without further feature extraction.Multivariate power frequency

[0084] Multivariate power frequency (MPF) features may be extracted according to a modified version of the technique presented in Kaifosh & Reardon (2024) (A generic noninvasive neuromotor interface for human-computer interaction, CTRL-labs at Reality Labs, David Sussillo, Patrick Kaifosh, Thomas Reardon, bioRxiv 2024.02.23.581779; doi: https: / / doi.org / 10.1101 / 2024.02.23.581779). In one example, shown below, the following frequency bins were used: (0,62.5,125,187.5,300). Two major deviations from the original method are: (1) computing MPF features in two equal-length (750ms) non-overlapping windows, and (2) keeping the full covariance matrix instead of selecting a reduced set of off- diagonal s.

[0085] Classification may be performed with custom Python code using the PyTorch package. Several classification strategies may be used to answer differing questions. In some- 14 -SG Docket No.: 14941-711.600example trainings OPM channels and the same 9 gestures were used: Index Pinch, Middle Pinch, Thumb Tap, Swipe Up, Swipe Down, Swipe Left, Swipe Right, Index Release, and Middle Release. In some analyses we report performance using either the wrist-only OPM channels or the forearm-only channels. Single-participant and generalized modelling approaches may be used.

[0086] Single-session (Session-dependent) models used time-aligned preprocessed data with detected bad channels removed and MPF features were computed on this reduced channel set. The MPF features across all channels and the two time windows were concatenated into a single, large feature vector on which we trained a Logistic Regression classifier. 10-fold cross-validation was employed and we report the average accuracy across the validation folds. For cross-session and cross-subject models no channels were removed to be able to match the number of features across sessions. A modified CNN+LSTM model (Kaifosh & Reardon, 2024) was used. The model was run on preprocessed data downsampled to 1000Hz without any further feature extraction. Cross-session models were run on time- aligned data, while cross-subject models were run on non-time-aligned data. Thus, the full dimensionality of a trial was 32 channels x 1500 timesteps for cross-session models and 32 channels x 2500 timesteps for cross-subject.

[0087] While time-alignment generally improves performance, a more robust method through data augmentation may be used. To improve the generalizability of our model in light of data scarcity, the following data augmentations may be used. TimeWarping may be used to deal with different gesture onset times. The trial may be randomly segmented into 4 segments and applied random interpolation with a factor between 0.5 and 2, effectively shrinking or stretching the timeseries. NoiseAugment may be used by adding random Gaussian noise with a 0.1 standard deviation and random rescaling of each channel with a factor, e.g., of between about 0.5 and 2. These data augmentations may be applied randomly to the training data in each batch.

[0088] The CNN+LSTM model had a single ID convolutional layer with a kernel size of 20 and a stride of 5. This down-sampled the input to 200Hz, which was followed by layer normalization, 3 LSTM layers, and a dense classification layer. The dense layer was applied to the last timestep output of the LSTM. The output channel number of the convolutional layer and the hidden size of the LSTM layers was set to 512. This implementation included dropout on the channel dimension both before and after the convolutional layer with a rate of 0.2. This made the model more robust to sensor positioning differences between sessions and subjects.

[0089] The model was trained with the AdamW optimizer with a learning rate scheduler- 15 -SG Docket No.: 14941-711.600that halved the learning rate every N epochs. Early stopping was used on the validation set by ending training after M epochs (called the patience factor) have passed without improvement in validation accuracy. Batch size, learning rate, number of epochs, learning rate halving, and the patience factor varied depending on the training at hand (e.g. cross-subject vs. crosssession), due to differences in training data amounts. These are reported in the respective sections, below. All results are reported on an independent test set using the model checkpoint at the best validation accuracy in the case of deep learning trainings. Specific data-splitting setups are also described below. In the results below, the inter-quantile range (IQR), which contains the middle 50th percentile of the samples, are shown below for all median values in parentheses, i.e. [QI, Q3 ] .Results

[0090] Multiple sessions of data were collected across 30 participants with a total of 70 sessions with OPMs and sEMG electrode pairs. MMG signals were validated to actively reflect muscle activity by comparing the spectral density between MMG and sEMG. A baseline single-session performance was first established using MPF features with a Logistic Regression model including all 32 channels (FIG. 24A). Across gestures cross-validation accuracy is 95.4% [92.3% - 98.1% IQR], but the Index and Middle Release gestures have much lower accuracy, 87.2% [74.5% - 95.8% IQR] and 87.5% [77.5% - 95.7% IQR], respectively. All other gesture types have either 97.9% or 100% median accuracy across sessions.

[0091] Next, classification was compared for accuracies between subsets of channels to better understand the significance in channel position as well as axis of sensitivity in classification accuracy (FIG. 24B). Each subset contains 16 channels. Using all channels the median cross-validation accuracy is 95.4% [92.3% - 98.1% IQR], while wrist-only channels is slightly lower at 94.1% [89.9% - 97.1% IQR], and forearm-only channels is 93.4% [88.5% - 97.0% IQR]. All comparisons between these 3 channel subsets are statistically significant (p<le-3, Bonferroni corrected for 5 comparisons).

[0092] For example, FIGS. 22A1-22A3 show an example of an experimental setup. The layout of MSR (22A1), the placement of OPMs and sEMG pairs (FIG. 22A2), and images of gestures presented to the participant to prompt a trial start (FIG. 22A3) are shown. Text also accompanied the images. From left to right, top to bottom: Index Pinch, Middle Pinch, Thumb Tap, Swipe Left, Swipe Right, Swipe up, Swipe Down, Neutral. Index and Middle Release were prompted with the Neutral image and adjusted text.

[0093] FIG. 22B shows example traces of 16 0PM channels (top 8 around the wrist, bottom 8 around the forearm) and all sEMG pairs during an Index Pinch and Middle Pinch- 16 -SG Docket No.: 14941-711.600trial. The signals were notched filtered for line noise and bandpass filtered from 30 to 300 Hz. Note the similarity in timing and SNR between the two modalities.

[0094] FIG. 22C shows example average amplitude spectral density for a single 0PM (blue) and EMG (red) channel on the dorsal side of the wrist (top). Solid lines show the average across all Index Pinch trials and dashed lines averages for all Neutral times. SNRs of the two channels (bottom). Note the similarity between the spectral content of the signals.

[0095] FIGS. 24A-24F show single-session cross-validation results. As shown, performance improves steadily with larger windows, with statistically significant differences between each subsequent window size. Distributions are over session-level cross-validated models trained on the shown trial length from movement onset (detected by our timealignment algorithm). The (62.5, 125) Hz band contains the most information content with every other frequency band having less information. Distributions are over session-level cross-validated models trained on the respective frequency ban, showing that while signals from the wrist potentially have more signal related to the gestures compared to the forearm, the forearm provides information in addition to the wrist. Using either tangential or radial channels achieves similar performance, having 93.9% [89.6% - 96.1% IQR] and 93.9% [89.5% - 97.4% IQR] accuracy, respectively. No consistent and significant differences were seen between channel subsets. Additionally, the idea that the number of sensors affects classification performance irrespective of sensor position was challenged. To that end, trainings were run increasing the number of sensors from 1 to 7 for both wrist and forearm bands separately. In each case, 16 random sensor subsets were sampled, except in the case of 1 and 7 sensors, where only 8 total different variations are possible (since each band contains 8 sensors). For each sensor both axes were used, as it was previously determined that there is no large difference between the two. Performance scaling was plotted with number of sensors in FIG. 24C. Across any number of sensors the forearm sensors resulted in statistically lower performance (p<le-3, Bonferroni corrected for 8 comparisons). Though slowly reaching an asymptote, performance continued to increase even when going from 7 to 8 wrist sensors, 93.5% to 94.1% median accuracy respectively (p<le-10). Going from 7 to 8 forearm sensors provides a similar improvement, from 92.7% to 93.4% (p<le-10). This shows that by adding a single sensor to the wrist band the magnitude of improvement (0.6%) is comparable to adding the whole forearm band (1.3% improvement). FIG. 24D shows an exponential curve ft of the form Er = e + AN / NaN to the median classification error (Er) with respect to sensor numbers (N). This shows that performance keeps improving and that forearm performance overtakes wrist-only sensor performance around 14 sensors. In the infinite limit forearm sensors have an error of 0, while wrist-only sensors have an error of 2.24%.- 17 -SG Docket No.: 14941-711.600

[0096] The effect of MPF feature window length on performance. Repeated trainings were run where the MPF feature computation was limited to the following trial lengths after movement onset: 200, 300, 400, 500, 600, 800, 1000, and 1200 ms (FIG. 24E). For calculating the MPF features two non-overlapping windows were used, e.g. two 100 ms windows for the 200 ms training. With only a 500 ms window length close to 90% median (across sessions) accuracy can already be achieved: 89.7% [82.1% - 93.7% IQR]. All consecutive window comparisons were statistically significant (p<le-5, Bonferroni corrected for 8 comparisons). Interestingly, using a window size of 1200 ms yielded 95.5% median accuracy, which is 0.1% higher than using the full 1500ms window (p<le-2). Confusion matrices change depending on the window length. Release gestures are better classified using a smaller window compared to the full trial length. Pinch gestures seem to be most affected by a reduced window length. There are between different subjects with respect to window size.

[0097] Finally, the frequencies that contain most of the gesture-related signals were examined. To this end trainings were repeatedly run separately on each frequency bin used in the MPF computation: (30, 62.5), (62.5, 125, )(125, 187.5), and (187.5, 300) Hz. FIG. 24F shows that the highest accuracy is achieved using the (62.5,125) Hz band (94.7% [90.8% - 97.6% IQR]), almost as high as with the whole frequency range (95.4%). Accuracy drops off (though not substantially) with higher frequencies, 91.5% [84.8% - 96.2% IQR] with the (187.5, 300) Hz band. All consecutive comparisons (i.e. (30, 62.5) with (62.5, 125), etc.) are statistically significant (p<le-5, Bonferroni corrected for 4 comparisons.) Generalization of gesture decoding

[0098] Cross-session classification. 17 participants performed at least 3 sessions. On these participants two cross-session analyses were performed to determine the generalizability of a decoder trained on one session to a different session. The first setup involved one training session, one validation session, and one test session. Mean test session accuracy over each train-test choice (6 folds in total) for each participant is shown. The second setup involved 2 training sessions and 1 test session (3 folds in total). In this case the validation set was sampled from the 2 training sessions randomly (10% of trials). On each fold a CNN+LSTM model was trained on time-aligned data (1.5s trial length) on all channels as well as wrist-only and forearm-only channels. For 1 -train-session trainings the Time- Warping segments were set to 3 and the max scaling for stretching to 1.5. The batch size was set to 128, the total epochs to 5000, and the patience to 800 epochs. Learning rate was set to 2e-5, and it was halved every 800 epochs. For 2-train-session trainings the total epochs were set to 3000, the patience to 500, and used an initial learning rate of 5-e5, halving every 300- 18 -SG Docket No.: 14941-711.600epochs.

[0099] Cross-session performance across all sessions is shown in FIG. 25A. Using 2 training sessions instead of 1 increases performance substantially, e.g. 71.5% [61.8% - 78.6% IQR] vs. 58.6% [47.9% - 72.5% IQR] in the case of all channels. Similar trends can be observed when using wrist-only and forearm-only channels. These comparisons are all significant (p<le-7, Bonferroni corrected for 3 comparisons). Forearm-only cross-session performance is not statistically significantly better than wrist-only in either the 1 -training session or 2-training session case. Using all channels is better than forearm-only in both cases (p<le-3, Bonferroni corrected for 4 comparisons). Confusion matrices for cross-session trainings with 2 training sessions. Release gestures seem to be the most confused, followed by swipe left with swipe down.

[0100] Cross-participant generalization. To determine generalizability across participants cross-participant classification was performed using all available sessions for each participant. This means that the training data was imbalanced across participants, as some had only 1 session, while others had 3. In order to provide robust results in each of the following analyses 30 trainings were run, 1 for each test participant, where the remaining 29 participants were randomly split into train and validation sets. CNN+LSTM was trained on non-time-aligned data (2.5s trial length) on all channels (32).

[0101] Any question to the generalizability of MMG is how performance scales with the number of training participants. To investigate this several cross-participant analyses were performed while varying the number of training participants, 10, 12, 15, 17, 20, and 22, with the rest used for validation. The same leave- 1 -participant-out testing setup was used by sampling the train and validation participants randomly for each test participant. For the 20- training-participant trainings the maximum number of epochs was set to 600, patience epochs were set to 100, and an initial learning rate of le-3 was used, halving every 100 epochs, with a batch size of 512. For the rest of the trainings, the number of epochs was adjusted, patience epochs, and learning-rate halving epochs with respect to the number of participants. I.e. for the 10-training -parti cipant trainings, the number of epochs was set to 1200, and the patience and learning rate halving epochs to 200. This increase is needed due to having less training samples in each epoch.

[0102] FIG. 26B shows that performance scales well with the amount of training participants. Using 10 training participants we found 71.2% [62.3% - 77.9% IQR] median accuracy, and using 15 participants increased accuracy to 73.0% [65.5% - 79.5% IQR]. This increase is not significant, but going from 15 to 20 participants is (p<le-2, Bonferroni corrected for 5 comparisons). The FIG. 25 3: Cross-session and cross-participant results.- 19 -SG Docket No.: 14941-711.600Chance-level is 0.11.

[0103] Cross-session accuracy across test sessions. Using 2 training sessions instead of 1 substantially improves performance, as well using all channels. Distributions are over each session of each participant. In the case of blue plots a single session was used for training, while in the case of orange plots 2 sessions were used.

[0104] Cross-participant accuracy across test participants depending on the size of the training set. Performance improves steadily when adding more participants to the training set. Distributions are over the mean (across sessions) accuracies of each test participant.12- training-participant performance seems to be an outlier as it has higher median accuracy than both the 15- and 17-training-parti cipant analyses. This could be due to the random sampling of the training participants for each test participant, and the effect would probably disappear when running even more samplings.

[0105] An exponential curve of the following form was fit to the median accuracies of the training subsets:Er = e + AN / NaN(1)

[0106] Where Er is the classification error measured as (100 - accuracy percentage), N is the number of participants measured in units of hundreds, and the rest are fitting parameters. The following parameters were found: e = 0.0, AN = 14.33, a» = 0.3. Except e, which is the irreducible error, these parameters are close to those reported in. This curve was plotted in FIG. 25C. In FIG. 25D, the per-gesture differences in performance for the 20-training- participant analysis. The Swipe Down and Swipe Left gestures have the lowest performance, while Index Pinch and Index Release have the best performance. There are significant differences between Index and Middle Pinches, but not between Releases. Swipe Right is significantly better than Swipe Left, but Swipe Up is not significantly better than Swipe Down. All tests were Bonferroni corrected for 4 comparisons.

[0107] Channel subset analysis CNN+LSTM was trained on non-time-aligned data (2.5s trial length) on all channels (32), as well as wrist-only (16), forearm-only (16), tangential- only (16), and radial-only (16) channels. Cross-participant performance for various channel subsets is shown in FIG. 25E. Across these trainings the train-validation splitting for each test participant was matched by using the same random seed. As was in the case of cross-session trainings, wrist-only and forearm-only performed considerably lower than using both (77.7% [65.5% - 84.4% IQR]), and forearm-only (72.3% [62.2% - 76.6% IQR]) performed slightly higher than wrist-only (69.4% [54.4% - 74.5% IQR]), but this was not statistically significant. Using all channels was significantly better than forearm-only (p<le-2, Bonferroni corrected- 20 -SG Docket No.: 14941-711.600for 5 comparisons). Using tangential -only or radial-only channels has similar performance with radial-only (70.6% [57.4% - 74.7% IQR]), slightly higher than tangential -only (69.3% [63.3% - 75.8% IQR]), but not statistically significant. Using all channels was significantly better than tangential -only (p<le-5, Bonferroni corrected for 5 comparisons).

[0108] Out of the 30 total participants 3 participants have higher mean accuracy with wrist-only channels, and 3 participants have higher mean accuracy with forearm-only channels. The confusion matrices across all test participants. Confusion matrices of 14 sessions with highest and lowest test accuracy, respectively, were shown. Low-accuracy sessions seem to have higher confusion rates in the two pinch gestures and the thumb tap.

[0109] Frequency bands. As was done for single-participant trainings, which frequency bands contain the signals with the most information content as well as the best generalizability across participants was investigated. The same setup was run using all channels on data that was bandpass filtered using the following ranges: (30,62),(62,125)(125, 187), (187,300) Hz. The results are plotted in FIG. 25F. Trends are similar to the single-participant case, except that the (187,300) Hz frequency range performs almost as well as the (62,125) Hz range. This suggests that high frequency content, even if containing less overall information, may be more generalizable across participants. Analyzing each pairwise comparison, the only significant difference was between the 62-125Hz and 125-187Hz trainings (p<2e-2). All individual frequency band trainings are below 70% median accuracy and significantly lower than using the entire 30-300 Hz bandwidth (p<le-3). These statistics have been Bonferroni corrected for 10 comparisons.Sources of variability

[0110] A large part of the variability is due to system / sensor position / person doing the gesture. Timing, actual gesture performance, etc. is highly variable even within session. Worse performing participants, both for single participant and cross-participant, typically had particularly inconsistent execution.[OHl] To assess whether we could quantify the variability and connect it to performance we first calculated entropy. Good predictor for single participant performance, not as good for cross participant performance (but better than just correlating between single and cross).

[0112] Latent space projection was attempted to see if certain participants were particularly different compared to others. This was not found to be the case.

[0113] Being able to quantify and find why certain participants perform worse would be very helpful, but more investigation is necessary. Could be very obfuscated due to complex / nonlinear nature of neural networks.

[0114] Accuracy over the duration of each session Possible causes of lower performance - 21 -SG Docket No.: 14941-711.600in certain sessions or participants include the shifting / rotation of the OPM band(s), as well as changes in participant position in the shielded room which changes the bias field the OPMs are subjected to. If this shift increases as the recording goes on, we would expect that crossparticipant performance goes down. We selected the best 14 and worst 14 sessions according to accuracy (for each channel -sub set cross-subject training), and ran a sliding window of 20% of the number of trials over the per-trial accuracies. In FIG. 26 we plot accuracy over consecutive trials within the recording. While the trends are minimal, for example in the allchannels case (FIG. 26A), the low-accuracy sessions do show an overall decrease in performance with respect to recording duration. Interestingly, the high accuracy sessions in the forearm-only training also show such a trend.

[0115] To substantiate these claims quantitatively, we calculated the significance of a linear ft being different from zero for each session partition (i.e. mean accuracy across top- 14 and bottom-14 sessions) and channel subset (Bonferroni corrected for 6 comparisons). While the linear trend slope is very small, we still found significant (p<le-6) downward trends in almost all cases (not only the ones highlighted above). Thus, both high-accuracy and low- accuracy sessions show this trend. The only exception was the high-accuracy sessions for the all-channels training which showed a significant (p<le-2) positive trend. We also correlated each session’s accuracy slope with the overall accuracy value of that session, to see if low- accuracy sessions show a larger downward trend, however we have not found significant results (p>0.1).

[0116] Window size. As with single-session classification, we wanted to assess performance with respect to the window size for cross-subject trainings. To this end we trained a modified CNN+LSTM on non-time-aligned data (2.5 second trials) with an extra data augmentation. This involved randomly cropping the start of each trial up to a maximum of 1000 ms. Testing trials were always 2.5 seconds long. To enable a single CNN+LSTM model to be able to predict across different window lengths, we included the output of each LSTM timestep starting from roughly 200ms in the loss function. This means that output predictions based on the first 200ms were not optimized in the loss. We trained this model on the same splits / folds as our original 20-training-parti cipant all-channel model.

[0117] After training, we generated predictions and computed accuracy at each timestep, resulting in FIG. 26D. Since CNN+LSTM was trained on the entire trial (starting from the prompt) early timesteps have low accuracy, as there is little information present. Performance increases rapidly, plateauing around 1400ms. We ran repeated measures ANOVA across timepoints to assess statistical significance and found p<le-4. Then we ran pairwise comparisons (Wilcoxon signed rank tests, Bonferroni corrected for the number of timepoints) - 22 -SG Docket No.: 14941-711.600between each timepoint and the last timepoint. We found that the first time the p-value goes above 0.05 is at 1410ms, and thus this is the point where accuracy stops increasing significantly. This shows that most information necessary to classify gestures across participants is present in the first 1400ms following the prompt. Taking into account the time to movement onset the actual information content window is likely much smaller.

[0118] Confusion matrices per subject. There is considerable variability in which gestures have better or worse performance between participants. FIGS. 26E and 26F show individual confusion matrices from the single-session MPF+logreg trainings. We selected the two participants with lowest accuracy to illustrate variability. While participant 22 has almost chance-level performance in the release gestures, participant 20 has much lower performance in the swipe gestures compared to others. We investigate the confusion matrices of these two participants for wrist-only and forearm-only trainings.Discussion

[0119] MMG can be used to effectively classify different gestures. The only previous attempt at gesture recognition using MMG relied on 8 channels to classify between three states: index finger flexion, little finger flexion, and neutral. The authors were able to successfully distinguish between the gestures, but were not able to reach >90% accuracy. In addition they showed that sEMG provided more accurate results despite only having 4 channels.

[0120] Here, we expand upon previous work by demonstrating high classification accuracy of various hand gestures using 32 channel MMG. We show a median of >95% accuracy across 9 gestures, rivaling state-of-the-art classification performed with sEMG [cite], despite a high rate of saturated sensors (1.48 saturated sensors per session on average, Table 1). One large difference from the previous study was our much larger bandwidth, ranging from 30 to 300 Hz compared to 25 to 100 Hz. Performing classification with separate frequency bands (FIG. 24F) showed that high frequencies, though attenuated due to the sensor’s frequency response, provided significant information for classification. This suggests that the capability for MMG to capture more high frequency power compared to sEMG due to the lack of attenuation through tissue may provide advantages for gesture recognition applications.

[0121] In addition, we observed differences in the signals obtained from the wrist to those from the forearm. Though intuition says the forearm should provide better signals due to higher muscle mass, sEMG signals have consistently shown higher classification accuracy when recorded at the wrist compared to the forearm, especially with deep learning approaches. Hypotheses as to why focus both on the higher density of muscle groups at the - 23 -SG Docket No.: 14941-711.600wrist, allowing for higher information per channel, as well as the presence of muscles for fine finger motion that is not present in the forearm. We saw similar results with MMG, showing significantly higher accuracy when recorded at the wrist compared to the forearm, demonstrating MMG can also be viable for a wrist wearable device.

[0122] Furthermore, we were able to assess the impact of the axis of sensitivity of MMG on gesture classification. As a magnetic field curls around a current (Biot-Savart law) different vector components are measured by the different axes of sensors. Previous studies have shown the radial and tangential components have different SNRs depending on the muscles measured, with wrist and forearm measurements typically providing larger signals in the radial direction. However, for gesture classification we saw no significant difference in accuracy between radial and tangential components. Additional investigation is necessary to determine the potential benefits of obtaining multiple components of MMG from the same location.

[0123] Coherence shows z and y have lower correlation compared to adjacent or across for low frequencies but not for high frequencies. Further evidence high frequencies are more local. Across shows highest, because same axis of recording. Wrist vs forearm the lowest regardless of frequency showing its different muscle groups. At lower frequencies (<125Hz) z and y is significantly lower than adjacent or across channels. Suggests 3 axis on part of the wrist should be just as accurate. Wrist vs forearm has lowest, further confirming differences in muscle groups captured by the two bands. This does mean a second row of sensors on the wrist may be better than more densely packed single row.

[0124] MMG can be generalized as well as sEMG. We were able to generalize gesture classification using MMG both across sessions and across participants at higher accuracy than previously reported with sEMG. This is likely due to various factors impacting sEMG session variability, such as humidity, skin dryness, skin / fat thickness, and exact placement of the sensors impacting the tissue-sensor impedance being irrelevant to MMG as the human body is largely transparent to magnetic fields. In addition, the non-generalizable crossparticipant error was 0.0% and showed we could enable less than 1% error with N participants, suggesting MMG may be a more ideal modality for enabling out-of-the-box usage of a gesture recognition device with participants to train a model.

[0125] Though the comparison with the sEMG study is not completely equivalent, with differences in number of trials, number of channels, and various other minor factors, we were able to achieve our results with non-ideal sensors. We had more than one saturated channel on average (i.e. no usable signal from that channel). Pure MMG variability, excluding factors introduced by the recording system, is likely even lower. In addition, magnetometers can be - 24 -SG Docket No.: 14941-711.600miniaturized such that multiple channels can be recorded in the same volume as a single pair of bipolar sEMG electrodes. Though we collected data from both the wrist and the forearm, due to cross-talk between the OPMs limiting higher proximity placements, higher density channels around the wrist should provide similar benefits.

[0126] Interestingly, although we found higher single-subject accuracy with wrist sensors we found higher cross-subject accuracy with forearm sensors. However, this is most likely due to the higher chance of wrist sensors being saturated from movement. Frequency bands showed similar results, but with high frequency performing as well as 62.5-125Hz, suggesting that high frequency content is a significant contributor to not only classification accuracy but also for generalization.

[0127] Limitations of the study and future directions. We used OPMs - though in a magnetically shielded room they are prone to saturation as they use biasing coils to remove any DC fields - so movement of the sensor body such that the bias field shifts causes the sensor to saturate. They also have limited bandwidth (200hz cutoff). Additionally, they are prone to introducing crosstalk with their biasing coils which prevented us from placing a higher density of sensors just on the wrist. A small, sensitive sensor that allows for high proximity to the body as well as each other with larger bandwidth and dynamic range would enable much more stable recordings, truly demonstrating the capabilities of MMG.

[0128] Magnetometers that are not band-limited may provide even higher classification accuracy. These methods and apparatuses may be used to detect handwriting, etc., that has been done with sEMG to determine how well MMG performs in complex, smaller SNR signal scenarios. These methods and apparatuses may be used for transfer learning / calibration to enable more generalization. Source space rather than sensor space analysis to see if it performs / generalizes better. These methods and apparatuses may be used for better time alignment / human error detection, and / or better quantification of why certain people perform so much worse in cross-participant.

[0129] Thus, described herein are MMG measured that may be placed / held at the wrist and forearm as a modality for hand gesture recognition. MMG was first seen to be used to accurately classify between 9 distinct gestures. We then demonstrate that MMG can generalize across sessions as well as across participants with fewer sessions and participants compared to sEMG. Finally, we expand upon the factors that contribute to variability observed in the dataset, a large part of which is due the limitations of the sensors used in the study. As a result, MMG may be the ideal modality for use in generalized gesture recognition with the advent of novel, high performing magnetometers.Validation of MMG- 25 -SG Docket No.: 14941-711.600

[0130] In the examples shown herein, the metrics provided may include the number of sensors, number of trials removed, incorrect trials, etc. Data may show correlation / coherence to indicate how much signal spread there is between channels. Other metrics may include entropy, PFI, ablation of features, etc.

[0131] In any of these methods and apparatuses, the number of sensors on just wrist vs forearm may be estimated. In any of these examples the method / apparatus may include sensors from other location to provide more information and / or better performance. The response time (window length) may be varied. In some cases the parameters used may be adjusted based on demographic differences in the user (e.g., handedness, correlation with arm length or circumference, etc.)Accuracy depending on channel subset

[0132] A baseline single-session performance may be established, e.g., using MPF features with a Logistic Regression model including all 0PM channels. This may be compared with results from using various subsets of channels; wrist-only, forearm-only, Y axis-only, and Z axis-only (FIG. 1). Each subset in FIG. 1 contains 16 channels. Using all channels the median cross-validation accuracy is 95.4% [91.7%, 98.3%], while wrist-only channels is slightly lower at 94.1% [89.8%, 97.2%], and forearm-only channels is 93.3% [88.2%, 97.0%]. This shows that signals from the wrist potentially have more signal related to the gestures though there is no statistically significant difference between them (p values, test). Using either Y or Z-axis channels achieves similar performance, having 93.9% [89.1%, 96.7%] and 93.9% [89.1%, 97.4%] accuracy, respectively.

[0133] To determine whether certain gestures were better predicted depending on the subset of selected channels, confusion matrices were averaged across all sessions as shown in FIGS. 2A-2E. The most confused gestures in this example are Index Release and Middle Release, as these are very similar to each other. This could mean that a hierarchical classifier which is first trained to distinguish normal gestures from release gestures and then to classify between the individual gestures could have higher performance. A binary classifier was determined on the two release gestures for comparison but found the same mean 86~ % accuracy as in the multiclass case. However, it seems there is high variability between sessions, with some session-level accuracy ranging between 54% - 99%.Gesture confusion matrices

[0134] To better understand why certain sessions and participants had low accuracy, the confusion matrices across all sessions of participants 15, 20, and 22, which had mean accuracies of 88%, 80%, and 80% using all channels, were examined. Participants 15 and 22 have close to 50% accuracy in the two release gestures, while in participant 20, the swipe - 26 -SG Docket No.: 14941-711.600gestures are confused more. There are also differences depending on whether wrist or forearm channels are used in classification. The release gestures are better classified for participant 20 with forearm-only channels, but swipe gesture accuracy is lower.Effect of number of channels

[0135] To investigate how the number of sensors affects overall performance irrespective of sensor position, trainings were performed increasing the number of sensors from 1 to 7 for both wrist and forearm bands separately. In each case, 16 random sensor subsets were sampled, except in the case of 1 and 7 sensors, where only 8 totally different variations are possible (since the bands contain 8 sensors). For each sensor both axes were used, as there is no large difference between the two. We plot performance scaling with number of sensors in FIG. 4. Across any number of sensors the forearm sensors result in lower performance. Though slowly reaching an asymptote, performance kept increasing even when going from 7 to 8 wrist sensors, 93.3% to 94.1% median accuracy. Increasing the sensor count within one band has a larger effect on overall performance than adding a second band.Effect of window length

[0136] The window length and MPF features may affect performance. Repeated trainings were performed where the MPF feature computation was limited to the following trial lengths following movement onset: 200, 300, 400, 500, 600, 800, 1000, 1200 ms (e.g., FIG. 5). Two non-overlapping windows were used for calculating the MPF features as in the case of 1500ms trial length, i.e. two 100ms windows for the 200ms training, etc. With only a 500ms window length close to 90% median (across sessions), accuracy can already be achieved: 89.8% [81.8%, 93.8%].

[0137] FIG. 6 illustrates how confusion matrices change depending on the window length. Interestingly, release gestures are better classified using a smaller window compared to the full trial length. Pinch gestures seem to be most affected by a reduced window length.

[0138] FIG. 7 shows an example of a plot for a first participant having a high overall accuracy (08), and two participants having lower overall accuracy (20 and 22). For the first participant, the performance is already quite high across most gestures even with a 200ms window (except Swipe Right), while participant 20 has random performance in all swipe gestures until after the window size is over 500ms. In contrast, participant 22 has much worse performance in the pinch gestures at 200ms, and better performance in some swipe gestures. These results could point to differences in the timing of how these participants performed the respective gestures and could also reflect that a fixed time-alignment algorithm may not be optimal for all sessions. Thus, an improved movement onset detection algorithm may improve the low-latency (small window) results.- 27 -SG Docket No.: 14941-711.600Significance of frequency bands

[0139] Certain frequencies may contain most of the gesture-related signal. To this end repeated trainings were performed, separately on each frequency bin, used in the MPF computation, i.e. (0, 62.5), (62.5, 125, )(125, 187.5), (187.5, 300) Hz. FIG. 8 shows that the highest accuracy is achieved using the (62.5, 125) Hz band (94.7% [90.4%, 97.8%]), almost as high as with the whole frequency range (95.4%). Accuracy drops off (though not substantially) with higher frequencies, 91.5% [84.5%, 96.5%] with the (187.5,300) Hz band. Cross-session classification

[0140] 17 participants performed at least 3 sessions. On these subjects, two cross-session analyses were performed to determine the generalizability of a decoder trained on one session to a different session. The first setup involved one training session, one validation session, and one test session. Mean test session accuracy was reported over each train-test choice (6 folds in total) for each subject. The second setup involved 2 training sessions and 1 test session (3 folds in total). In this case the validation set was sampled from the 2 training sessions randomly (10% of trials). On each fold we trained a CNN+LSTM model on time- aligned data (e.g., 1.5s trial length) on all channels as well as wrist-only and forearm-only channels.

[0141] For 1 -train-session trainings the TimeWarping segments were set to 3 and the max scaling for stretching to 1.5. The batch size was set to 128, the total epochs to 5000, and the patience to 800 epochs. Learning rate was set to 2e-5, and it was halved every 800 epochs. For 2-train-session trainings the total epochs were set to 3000, the patience to 500, and used an initial learning rate of 5-e5, halving every 300 epochs.

[0142] There is significant variability depending on which session is used for training and testing. Cross-session performance across all sessions is shown in FIG. 9. Using 2 training sessions instead of 1 increases performance substantially, e.g. 71.5% [61.8%, 78.6%] vs. 58.6% [47.9%, 72.5%] in the case of all channels. Similar trends can be observed when using wrist-only and forearm-only channels. Interestingly, forearm-only cross-session performance is slightly higher, though single-session performance was slightly lower than using wrist-only channels.

[0143] Confusion matrices for cross-session trainings with 2 training sessions across all sessions are shown in FIG. 10. For comparison, the confusion matrices of a subject (21) with low overall cross-session performance, and another subject (33) with consistently high crosssession performance are shown (FIG. 10). It seems that for subject 21 using forearm-only channels gestures were less confused. For subject 33, Swipe Left and Swipe Up were particularly confused, as well as Swipe Right and Swipe Left. Again, there are differences - 28 -SG Docket No.: 14941-711.600depending on whether wrist-only or forearm-only channels are used.

[0144] FIG. 12 shows a graph of the confusion matrices (all channels) of subject 08 separately for each test session. This is interesting because this subject had two high accuracy test sessions (when training on the other sessions), but much lower performance on the 3rd session. This means that the characteristics of this session must be very different from the other two.Cross-participant classification

[0145] To determine generalizability across participants we performed cross-participant classification using all available sessions for each subject. This means that the training data was imbalanced across subjects, as some had only 1 session, while others had 3. In order to provide robust results 30 trainings were performed, 1 for each test subject, where the remaining 29 subjects were randomly split into train (20 subjects) and validation (9 subjects) sets.Channel subsets

[0146] CNN+LSTM was trained on non-time-aligned data (2.5s trial length) on all channels (32), as well as wrist-only (16), forearm-only (16), Y-axis only (16), and Z-axis only (16) channels. For these trainings the maximum number of epochs was set to 600, patience epochs were set to 100, and we used an initial learning rate of le-3, halving every 100 epochs, with a batch size of 512.

[0147] Cross-subject performance for various channel subsets is shown in FIG. 13. Across these trainings we matched the train-validation splitting for each test subject by using the same random seed (42). As was in the case of cross-session trainings, wrist-only and forearm-only performed considerably lower than using both (77.7% [65.5%, 84.4%]), and forearm-only (72.3% [62.2%, 76.6%]) performed slightly higher than wrist-only (69.4% [54.4%, 74.5%]). Using Y-only or Z-only channels has similar performance with Z-only (70.6% [57.4%, 74.7%]), slightly higher than Y-only (69.3% [63.3%, 75.8%]).Session differences

[0148] There are sessions in which using a subset of channels is better than using all channels. For example subject 15 session 01 has 43% accuracy in the all-channel training, 57% in the wrist-only, and 46% in the forearm-only trainings. Subject 30 session 01 has 64% accuracy in the all-channel training, 48% in the wrist-only, and 77% in the forearm-only trainings. Thus, certain subjects may generalize better using wrist-only or forearm-only channels. Confusion matrices for these two sessions are shown in FIGS. 14A-14F. For subject 15 in the all-channel case a lot of gestures are predicted as thumb tap, while this is not the case for the wrist-only training, hence the higher accuracy. For both wrist-only and- 29 -SG Docket No.: 14941-711.600forearm-only trainings however, more gestures are erroneously predicted as swipe up. For subject 30, the forearm-only training has more correct gestures especially thumb taps, swipe ups, and swipe rights.

[0149] The confusion matrices across all test subjects are shown in FIGS. 15A-15C. We also show the confusion matrices of 3 bad subjects (<60% mean accuracy). There is considerable variability in terms of which gestures get misclassified more depending on the subject (FIGS. 16A-16I).Accuracy over the duration of each session

[0150] One possible cause of lower performance in certain sessions or subjects is the shifting / rotation of the OPM band(s). If this shift increases as the recording goes on, crosssubject performance may go down. The best 14 and worst 14 sessions were selected according to accuracy, and a sliding window of 20% of the number of trials over the per-trial accuracies was generated. FIG. 17 shows a graph of accuracy over consecutive trials within the recording. While in this example the trends are minimal, e.g., in the all-channels case, the bad sessions do show an overall decrease in performance with respect to recording duration. Interestingly, the good sessions in the forearm-only training also show such a trend.

[0151] While the wrist channels show some decrease in performance in the beginning of the recording, this stabilizes later (and even increases slightly). However, there may be specific sessions with poor performance which do show a clear downward trend in accuracy during the recording (see FIG. 18).Frequency bands

[0152] As we did for single-participant trainings, we investigated which frequency bands contain the signals with the most information content as well as the best generalizability across subjects. We ran the same setup using all channels on data that was bandpass filtered using the following ranges: (30, 62), (62, 125)(125, 187), (187, 300) Hz.

[0153] The results are plotted in FIG. 19. Trends are similar to the single-participant case, except that the (187,300) Hz frequency range performs almost as well as the (62, 125) Hz range. This suggests that high frequency content, even if containing less overall information, may be more generalizable across subjects. All individual frequency band trainings are below 70% median accuracy, much lower than using the entire 30-300 Hz bandwidth (77.7%).Number of participants in training data

[0154] An important question to the generalizability of MMG is how performance scales with the number of training subjects. To investigate this, two additional cross-participant classifications with 10 training subjects and 15 training subjects were used. The same leave- 1 -subject-out testing setup by sampling the train and validation subjects randomly for each - 30 -SG Docket No.: 14941-711.600test subject, meaning that the validation sets had 19 and 14 subjects respectively was used. FIG. 20 shows that performance scales well with the amount of training subjects. Using 10 training subjects, 71.2% [62.3%, 77.9%] were found with median accuracy, and using 15 subjects increased accuracy to 73.0% [65.5%, 79.5%].Window size

[0155] Finally, as with single-session classification, performance may be assessed with respect to the window size. To this end a modified CNN+LSTM was trained on non-time- aligned data (2.5 second trials) with an extra data augmentation. This involved randomly cropping the start of each trial up to a maximum of 1000 ms. Testing trials were always 2.5 seconds long. To enable a single CNN+LSTM model to be able to predict across different window lengths, the output of each LSTM timestep starting from roughly 200ms in the loss function was included. This means that output predictions based on the first 200ms were not optimized in the loss. We trained this model on the same splits / folds as our original 20- training-subject all-channel model.

[0156] After training, predictions were generated and computed accuracy at each timestep, resulting in FIGS. 21 A-21B. Since CNN+LSTM was trained on the entire trial (starting from the prompt) early timesteps have low accuracy, as there is little information present. Performance increases rapidly, plateauing around 1200 ms. This shows that most information necessary to classify gestures across subjects is present in the first 1200 ms following the prompt. Taking into account the time to movement onset the actual information content window is likely much smaller.

[0157] Five (5) sessions with low overall accuracy are shown in FIG. 2 IB to uncover potential issues. For example for subject 22 session 01 the accuracy increases much sooner, potentially showing that they performed the gestures faster than the average. Others, such as subject 21 session 01 and subject 31 session 01 show a continuous increase in accuracy all the way to the end of the session.

[0158] All publications and patent applications mentioned in this specification are herein incorporated by reference in their entirety to the same extent as if each individual publication or patent application was specifically and individually indicated to be incorporated by reference. Furthermore, it should be appreciated that all combinations of the foregoing concepts and additional concepts discussed in greater detail below (provided such concepts are not mutually inconsistent) are contemplated as being part of the inventive subject matter disclosed herein and may be used to achieve the benefits described herein.

[0159] Any of the methods (including user interfaces) described herein may be implemented as software, hardware or firmware, and may be described as a non-transitory - 31 -SG Docket No.: 14941-711.600computer-readable storage medium storing a set of instructions capable of being executed by a processor (e.g., computer, tablet, smartphone, etc.), that when executed by the processor causes the processor to control perform any of the steps, including but not limited to: displaying, communicating with the user, analyzing, modifying parameters (including timing, frequency, intensity, etc.), determining, alerting, or the like. For example, any of the methods described herein may be performed, at least in part, by an apparatus including one or more processors having a memory storing a non-transitory computer-readable storage medium storing a set of instructions for the processes(s) of the method.

[0160] While various embodiments have been described and / or illustrated herein in the context of fully functional computing systems, one or more of these example embodiments may be distributed as a program product in a variety of forms, regardless of the particular type of computer-readable media used to actually carry out the distribution. The embodiments disclosed herein may also be implemented using software modules that perform certain tasks. These software modules may include script, batch, or other executable files that may be stored on a computer-readable storage medium or in a computing system. In some embodiments, these software modules may configure a computing system to perform one or more of the example embodiments disclosed herein.

[0161] As described herein, the computing devices and systems described and / or illustrated herein broadly represent any type or form of computing device or system capable of executing computer-readable instructions, such as those contained within the modules described herein. In their most basic configuration, these computing device(s) may each comprise at least one memory device and at least one physical processor.

[0162] The term “memory” or “memory device,” as used herein, generally represents any type or form of volatile or non-volatile storage device or medium capable of storing data and / or computer-readable instructions. In one example, a memory device may store, load, and / or maintain one or more of the modules described herein. Examples of memory devices comprise, without limitation, Random Access Memory (RAM), Read Only Memory (ROM), flash memory, Hard Disk Drives (HDDs), Solid-State Drives (SSDs), optical disk drives, caches, variations or combinations of one or more of the same, or any other suitable storage memory.

[0163] In addition, the term “processor” or “physical processor,” as used herein, generally refers to any type or form of hardware-implemented processing unit capable of interpreting and / or executing computer-readable instructions. In one example, a physical processor may access and / or modify one or more modules stored in the above-described memory device. Examples of physical processors comprise, without limitation,- 32 -SG Docket No.: 14941-711.600microprocessors, microcontrollers, Central Processing Units (CPUs), Field-Programmable Gate Arrays (FPGAs) that implement softcore processors, Application-Specific Integrated Circuits (ASICs), portions of one or more of the same, variations or combinations of one or more of the same, or any other suitable physical processor.

[0164] Although illustrated as separate elements, the method steps described and / or illustrated herein may represent portions of a single application. In addition, in some embodiments one or more of these steps may represent or correspond to one or more software applications or programs that, when executed by a computing device, may cause the computing device to perform one or more tasks, such as the method step.

[0165] In addition, one or more of the devices described herein may transform data, physical devices, and / or representations of physical devices from one form to another. Additionally or alternatively, one or more of the modules recited herein may transform a processor, volatile memory, non-volatile memory, and / or any other portion of a physical computing device from one form of computing device to another form of computing device by executing on the computing device, storing data on the computing device, and / or otherwise interacting with the computing device.

[0166] The term “computer-readable medium,” as used herein, generally refers to any form of device, carrier, or medium capable of storing or carrying computer-readable instructions. Examples of computer-readable media comprise, without limitation, transmission-type media, such as carrier waves, and non-transitory-type media, such as magnetic-storage media (e.g., hard disk drives, tape drives, and floppy disks), optical-storage media (e.g., Compact Disks (CDs), Digital Video Disks (DVDs), and BLU-RAY disks), electronic-storage media (e.g., solid-state drives and flash media), and other distribution systems.

[0167] A person of ordinary skill in the art will recognize that any process or method disclosed herein can be modified in many ways. The process parameters and sequence of the steps described and / or illustrated herein are given by way of example only and can be varied as desired. For example, while the steps illustrated and / or described herein may be shown or discussed in a particular order, these steps do not necessarily need to be performed in the order illustrated or discussed.

[0168] The various exemplary methods described and / or illustrated herein may also omit one or more of the steps described or illustrated herein or comprise additional steps in addition to those disclosed. Further, a step of any method as disclosed herein can be combined with any one or more steps of any other method as disclosed herein.

[0169] The processor as described herein can be configured to perform one or more steps- 33 -SG Docket No.: 14941-711.600of any method disclosed herein. Alternatively or in combination, the processor can be configured to combine one or more steps of one or more methods as disclosed herein.

[0170] When a feature or element is herein referred to as being "on" another feature or element, it can be directly on the other feature or element or intervening features and / or elements may also be present. In contrast, when a feature or element is referred to as being "directly on" another feature or element, there are no intervening features or elements present. It will also be understood that, when a feature or element is referred to as being "connected", "attached" or "coupled" to another feature or element, it can be directly connected, attached or coupled to the other feature or element or intervening features or elements may be present. In contrast, when a feature or element is referred to as being "directly connected", "directly attached" or "directly coupled" to another feature or element, there are no intervening features or elements present. Although described or shown with respect to one embodiment, the features and elements so described or shown can apply to other embodiments. It will also be appreciated by those of skill in the art that references to a structure or feature that is disposed "adjacent" another feature may have portions that overlap or underlie the adjacent feature.

[0171] Terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. For example, as used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items and may be abbreviated as " / ".

[0172] Spatially relative terms, such as "under", "below", "lower", "over", "upper" and the like, may be used herein for ease of description to describe one element or feature's relationship to another element(s) or feature(s) as illustrated in the figures. It will be understood that the spatially relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the figures. For example, if a device in the figures is inverted, elements described as "under”, or "beneath" other elements or features would then be oriented "over" the other elements or features. Thus, the exemplary term "under" can encompass both an orientation of over and under. The device may be otherwise oriented (rotated 90 degrees or at other orientations) and the spatially relative descriptors used herein interpreted accordingly. Similarly, the terms "upwardly", "downwardly", "vertical", "horizontal" and the like are used herein for the purpose of- 34 -SG Docket No.: 14941-711.600explanation only unless specifically indicated otherwise.

[0173] Although the terms “first” and “second” may be used herein to describe various features / elements (including steps), these features / elements should not be limited by these terms, unless the context indicates otherwise. These terms may be used to distinguish one feature / element from another feature / element. Thus, a first feature / element discussed below could be termed a second feature / element, and similarly, a second feature / element discussed below could be termed a first feature / element without departing from the teachings of the present invention.

[0174] In general, any of the apparatuses and methods described herein should be understood to be inclusive, but all or a sub-set of the components and / or steps may alternatively be exclusive and may be expressed as “consisting of’ or alternatively “consisting essentially of’ the various components, steps, sub-components or sub-steps.

[0175] As used herein in the specification and claims, including as used in the examples and unless otherwise expressly specified, all numbers may be read as if prefaced by the word "about" or “approximately,” even if the term does not expressly appear. The phrase “about” or “approximately” may be used when describing magnitude and / or position to indicate that the value and / or position described is within a reasonable expected range of values and / or positions. For example, a numeric value may have a value that is + / - 0.1% of the stated value (or range of values), + / - 1% of the stated value (or range of values), + / - 2% of the stated value (or range of values), + / - 5% of the stated value (or range of values), + / - 10% of the stated value (or range of values), etc. Any numerical values given herein should also be understood to include about or approximately that value, unless the context indicates otherwise. For example, if the value " 10" is disclosed, then "about 10" is also disclosed. Any numerical range recited herein is intended to include all sub-ranges subsumed therein. It is also understood that when a value is disclosed that "less than or equal to" the value, "greater than or equal to the value" and possible ranges between values are also disclosed, as appropriately understood by the skilled artisan. For example, if the value "X" is disclosed the "less than or equal to X" as well as "greater than or equal to X" (e.g., where X is a numerical value) is also disclosed. It is also understood that the throughout the application, data is provided in a number of different formats, and that this data, represents endpoints and starting points, and ranges for any combination of the data points. For example, if a particular data point “10” and a particular data point “15” are disclosed, it is understood that greater than, greater than or equal to, less than, less than or equal to, and equal to 10 and 15 are considered disclosed as well as between 10 and 15. It is also understood that each unit between two particular units are also disclosed. For example, if 10 and 15 are disclosed, then 11, 12, 13, and 14 are also- 35 -SG Docket No.: 14941-711.600disclosed.

[0176] Although various illustrative embodiments are described above, any of a number of changes may be made to various embodiments without departing from the scope of the invention as described by the claims. Optional features of various device and system embodiments may be included in some embodiments and not in others. Therefore, the foregoing description is provided primarily for exemplary purposes and should not be interpreted to limit the scope of the invention as it is set forth in the claims.

[0177] The examples and illustrations included herein show, by way of illustration and not of limitation, specific embodiments in which the subject matter may be practiced. As mentioned, other embodiments may be utilized and derived there from, such that structural and logical substitutions and changes may be made without departing from the scope of this disclosure. Such embodiments of the inventive subject matter may be referred to herein individually or collectively by the term “invention” merely for convenience and without intending to voluntarily limit the scope of this application to any single invention or inventive concept, if more than one is, in fact, disclosed. Thus, although specific embodiments have been illustrated and described herein, any arrangement calculated to achieve the same purpose may be substituted for the specific embodiments shown. This disclosure is intended to cover any and all adaptations or variations of various embodiments. Combinations of the above embodiments, and other embodiments not specifically described herein, will be apparent to those of skill in the art upon reviewing the above description.- 36 -SG Docket No.: 14941-711.600

Claims

CLAIMSWhat is claimed is:

1. A method, the method comprising: sensing a plurality of magnetic signals from one or more muscles using an array of magnetometers worn on an arm and / or wrists; dividing the plurality of magnetic signals into windows having a window size of between about 200 and about 1200 ms; applying a trained machine learning agent to identify a gesture from the plurality of windows, wherein the trained machine learning agent is configured to apply a hierarchical classification that first identifies a first category of gestures from a second category of gestures and then classifies between the individual gestures within each category; and outputting the identified gesture.

2. The method of claim 1, wherein the array of magnetometers comprises an array of acoustically driven ferromagnetic resonance (ADFMR) sensors.

3. The method of claim 1, further comprising digitizing the plurality of magnetic signals.

4. The method of claim 1, further comprising remove bad channels from the array of magnetometers.

5. The method of claim 1, wherein the array of magnetometers comprises more than 8 sensors.

6. The method of claim 1, wherein sensing comprises sensing from both the arm and the wrist.

7. The method of claim 1, further comprising resample the filtered plurality of magnetic signals.

8. The method of claim 7, wherein resampling comprising resampling at 1 KHz.

9. The method of claim 1, wherein the first category of gestures comprises normal gestures.

10. The method of claim 9, wherein the normal gestures comprises contracting gestures- 37 -SG Docket No.: 14941-711.600including: clench, flex, Index Pinch, Middle Pinch, swipe right, swipe left, swipe up, and swipe down.

11. The method of claim 1, wherein the second category of gestures comprises release gestures.

12. The method of claim 11, wherein the release gestures comprises relaxing gestures including: extend, index release, middle release.

13. The method of claim 1, further comprising filtering the plurality of magnetic signals.

14. The method of claim 1, further comprising dynamically adjusting the window size while applying the trained machine learning agent.

15. The method of claim 1, wherein outputting comprises: displaying, storing or transmitting the identified gesture.

16. A method, the method comprising: sensing a plurality of magnetic signals from one or more muscles using an array of magnetometers worn on an arm and / or wrists; filtering the plurality of magnetic signals; dividing the plurality of magnetic signals into windows having a window size of between about 200 and about 1200 ms; applying a trained machine learning agent to identify a gesture from the plurality of windows, wherein the trained machine learning agent is configured to apply a hierarchical classification that first identifies a first category of gestures from a second category of gestures and then classifies between the individual gestures within each category; repeating the application of the trained machine learning agent while adjusting the window size; and outputting the identified gesture.

17. A system, the system comprising: an assembly comprising an array of magnetometers configured to be worn on an arm and / or wrist configured to sense a plurality of magnetic signals from one or more muscles; one or more processors; and- 38 -SG Docket No.: 14941-711.600a memory coupled to the one or more processors, the memory storing computerprogram instructions, that, when executed by the one or more processors, perform a computer-implemented method comprising: dividing the plurality of magnetic signals into windows having a window size of between about 200 and about 1200 ms; identifying a gesture from the plurality of windows by applying a hierarchical classification that first identifies a first category of gestures from a second category of gestures and then classifies between the individual gestures within each category; and outputting the identified gesture.

18. The apparatus of claim 17, wherein identifying the gesture comprises applying a trained machine learning agent to identify the gesture from the plurality of windows, further wherein the trained machine learning agent is configured to apply the hierarchical classification.

19. The apparatus of claim 17, wherein the array of magnetometers comprises an array of acoustically driven ferromagnetic resonance (ADFMR) sensors.

20. The apparatus of claim 17, wherein the computer-implemented method further comprises digitizing the plurality of magnetic signals.

21. The apparatus of claim 17, wherein the computer-implemented method further comprises removing bad channels from the array of magnetometers.

22. The apparatus of claim 17, wherein the array of magnetometers comprises more than 8 sensors.

23. The apparatus of claim 17, wherein sensing comprises sensing from both the arm and the wrist.

24. The apparatus of claim 17, wherein the computer-implemented method further comprises resampling the filtered plurality of magnetic signals.

25. The apparatus of claim 24, wherein resampling comprising resampling at 1 KHz.

26. The apparatus of claim 17, wherein the first category of gestures comprises normal gestures.- 39 -SG Docket No.: 14941-711.60027. The apparatus of claim 26, wherein the normal gestures comprises contracting gestures including: clench, flex, Index Pinch, Middle Pinch, swipe right, swipe left, swipe up, and swipe down.

28. The apparatus of claim 17, wherein the second category of gestures comprises release gestures.

29. The apparatus of claim 28, wherein the release gestures comprises relaxing gestures including: extend, index release, middle release.

30. The apparatus of claim 17, wherein the computer-implemented method further comprises filtering the plurality of magnetic signals.

31. The apparatus of claim 17, wherein the computer-implemented method further comprises dynamically adjusting the window size while applying the trained machine learning agent.

32. The apparatus of claim 17, wherein outputting comprises: displaying, storing or transmitting the identified gesture.- 40 -SG Docket No.: 14941-711.600

Citation Information

Patent Citations

  • Recognizing gestures from forearm EMG signals

    US20090327171A1

  • Wearable electromyography sensor array using conductive cloth electrodes for human-robot interactions

    US20170259428A1

  • Gesture based feedback for wearable devices

    US20170364156A1

  • Implant Encoder

    US20230255794A1