Methods and apparatuses for gesture detection and classification

By using unsupervised and self-supervised models of EMG signals, user gestures are detected and classified, and the efficiency and accuracy of gesture detection and classification in the prior art are solved, and the control capability and user experience of artificial reality equipment are improved.

CN113646734BActive Publication Date: 2025-07-08CTRL-LABS CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080026319.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-27
Filing Date
2020-03-30
Publication Date
2025-07-08
Estimated Expiration
2040-03-30

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently and unsupervisedly detect and classify user gestures, especially to accurately control gestures in artificial reality environments, resulting in limited immersive experience.

Method used

Using an unsupervised and self-supervised model based on electromyography (EMG) signals, the user's gestures are detected and classified through principal component analysis (PCA), clustering and machine learning technology, and corresponding control signals are generated to control artificial reality devices.

Benefits of technology

It realizes efficient and accurate detection and classification of gestures in artificial reality environments, and improves the user's control accuracy and immersion of the device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113646734B_ABST
    Figure CN113646734B_ABST
Patent Text Reader

Abstract

An example system may include: a head-mounted device configured to present an artificial reality view to a user; a control device including a plurality of electromyography (EMG) sensors; and at least one physical processor programmed to receive EMG data based on signals detected by the EMG sensors, detect EMG signals within the EMG data corresponding to user gestures, classify the EMG signals to identify gesture types, and provide a control signal based on the gesture types, wherein the control signal triggers the head-mounted device to modify the artificial reality view. Various other methods, systems, and computer-readable media are also disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of U.S. Application No. 62 / 826,478, filed Mar. 29, 2019, and U.S. Application No. 16 / 833,307, filed Mar. 27, 2020. For all purposes, the contents of U.S. Application No. 62 / 826,478 and U.S. Application No. 16 / 833,307 are incorporated herein by reference in their entirety. Brief Description of the Drawings

[0004] The drawings illustrate numerous exemplary embodiments and are part of the specification. Together with the following description, these drawings demonstrate and explain the various principles of the present disclosure.

[0005] Figure 1 An example of the first component extracted from the application of PCA is shown.

[0006] Figure 2 An example of a cluster generated from the detected event is shown.

[0007] Figure 3 An example graph of the first component from PCA performed on the detected discrete events is shown.

[0008] Figures 4A - 4B An epoch corresponding to the discrete event is illustrated, showing aspects of synchronization quality.

[0009] Figures 5A - 5B An aligned epoch corresponding to the detected discrete event is shown.

[0010] Figures 6A - 6B A template corresponding to the PCA analysis performed on the average of two different gestures is shown.

[0011] Figure 7 An example of the detected event on the first PCA component and the corresponding markers generated from two - second data are shown.

[0012] Figure 8 An example of detecting discrete events using a test set is shown.

[0013] Figure 9 An example of discrete events detected in the test dataset is shown.

[0014] Figures 10A - 10B Examples of an index - finger tap event model and a middle - finger tap event model are shown.

[0015] Figures 11A - 11F An example of a user - specific event model for two classes of events is shown.

[0016] Figure 12 Shows example accuracy levels achieved by various single-user event classification models.

[0017] Figure 13A Shows example accuracy levels achieved by two single-user event classification models.

[0018] Figure 13B Shows the relationship between example accuracy levels of two single-user event classification models (single stamp and cumulative window size) and time.

[0019] Figure 14 Shows generalization across time performed to determine the independence of time samples.

[0020] Figure 15 Shows example accuracy levels of a generalized cross-user classification model.

[0021] Figure 16 Shows an example of the transferability of a user-specific classifier based on linear regression.

[0022] Figures 17A - 17Q Shows an example distribution of two classes of gestures.

[0023] Figures 18A - 18B Shows examples of clusters separated using UMAP and PCA.

[0024] Figure 19 Shows examples of accuracy levels achieved using a self-supervised model.

[0025] Figure 20 Shows an example of the relationship between accuracy levels achieved using a supervised user-specific model and a self-supervised user-specific model and the number of training events.

[0026] Figure 21 Shows an example of window size determination for user-specific models and self-supervised models.

[0027] Figures 22A - 22D Shows example models for each event category associated with a first user.

[0028] Figures 23A - 23B Shows an example of an alignment model for each event category associated with a first user and a second user.

[0029] Figures 24A - 24B Exemplary data before and after transformation are shown respectively.

[0030] Figure 25A An exemplary cross - user transfer matrix from all users in a group of users is shown.

[0031] Figure 25B The data size for supervised domain adaptation determined based on the transfer function is shown.

[0032] Figure 26A A wearable system according to some embodiments is shown, which has EMG sensors circumferentially arranged around an elastic band configured to be worn around a user's lower arm or wrist.

[0033] Figure 26B is a cross - sectional view through Figure 26A one of the EMG sensors shown in

[0034] Figure 27A and Figure 27B Schematically shows components of a computer - based system on which some embodiments are implemented. Figure 27A Shows a schematic diagram of a control device of a computer - based system, and Figure 27B shows an exemplary dongle part that can be connected to a computer, where the dongle part is configured to communicate with the control device (and a similar configuration can be used within a head - mounted device that communicates with the control device).

[0035] Figure 28 Shows an exemplary implementation where a wearable device interfaces with a head - mounted wearable display.

[0036] Figure 29 and Figure 30 Shows an exemplary method.

[0037] Figure 31 Is an illustration of an exemplary augmented reality glasses that can be used in conjunction with embodiments of the present disclosure.

[0038] Figure 32 Is an illustration of an exemplary virtual reality headset that can be used in conjunction with embodiments of the present disclosure.

[0039] In all the figures, the same reference symbols and descriptions indicate similar but not necessarily identical elements. While the exemplary embodiments described herein are amenable to various modifications and alternative forms, specific embodiments have been shown by way of example in the figures and are described in detail herein. However, the exemplary embodiments described herein are not intended to be limited to the particular forms disclosed. Rather, the disclosure covers all modifications, equivalents, and alternatives falling within the scope of the appended claims.

[0040] Detailed Description of Exemplary Embodiments

[0041] Examples of the present disclosure are directed to the detection of signals from a user and the control of an artificial reality device based on the detected signals. As explained in more detail below, embodiments of the present disclosure may include a system having a head-mounted device configured to present an artificial reality view to a user and a control device including a plurality of electromyography (EMG) sensors. One or more processors, which may be located in any system component, may be programmed to detect EMG signals corresponding to user gestures associated with EMG data received from the sensors and classify the EMG signals to identify gesture types. Control signals may trigger the head-mounted device to modify the artificial reality view (e.g., based on the gesture type).

[0042] Accurate control of (real or virtual) objects within an artificial reality environment may be useful for maintaining an immersive experience. Gestures may be a useful way to control objects and do not require interaction with any real physical object. For example, actions such as pressing a key on a keyboard, turning a dial, pressing a button, selecting an item from a menu (and many other actions) may be simulated by user gestures. A tap gesture may simulate a key press. Additionally, identifying which body part (e.g., which finger) has been used to perform a gesture allows for further control of the artificial reality environment.

[0043] In accordance with the general principles described herein, the features of any of the embodiments described herein may be used in combination with one another. These and other embodiments, features, and advantages will be more fully understood when the following detailed description is read in conjunction with the accompanying figures and claims.

[0044] Reference Figures 1 - 30 , a detailed description of gesture recognition models (including unsupervised models and self-supervised models) is provided below. Figures 1 - 13B Event detection and classification are shown, where the term "event" may include gestures such as finger taps. Figures 14 - 25B The temporal dependence, clustering, training, and accuracy of various models are further shown. Figures 26A - 26B An example control device is shown. Figures 27A - 27BShows a schematic diagram of a control device. Figure 28 Shows an example system including a head-mounted device. Figures 29 - 30 Shows an example computerized method, and Figure 31 and Figure 32 Shows an example AR / VR application.

[0045] The present disclosure is directed to an event detector model that can be used to detect user gestures. Such a detector model may involve recording a series of EMG signals (a dataset) when one or more users perform different gestures. In some examples, example gestures may include finger taps (e.g., simulated button presses), but other types of gestures can similarly be used to implement the example event detector model.

[0046] Gestures can include discrete events spanning a limited time period and, in some embodiments, can be characterized by one or more electromyogram signals (including electromyogram wavelets) representing muscle activation. Configuring a system to detect and classify such gestures using machine learning techniques may involve a large number of labeled training samples. Thus, a system that can quickly learn gestures from few samples and capture and interpret meaningful features from human gestures in an unsupervised or self-supervised manner is highly desirable. The examples described herein provide such unsupervised models and / or self-supervised models.

[0047] Figure 1 Shows a first component (PCA, vertical line) and detected peaks (dots) that can be extracted from the application of principal component analysis. Multiple events are shown, which are divided into two groups separated by rest periods. The illustrated events can be detected using a peak detection process that can also detect peaks recorded during the rest period, which correspond to local maxima during rest.

[0048] The dataset can include EMG signals corresponding to index finger taps and middle finger taps. The dataset can be divided into a training set and a test set, where the training set includes 50 consecutive finger taps per finger recorded at approximately 2 Hz, and the test set includes 20 consecutive finger taps per finger recorded at approximately 2 Hz. The above dataset may represent less than 2 minutes of recorded data. Any other suitable dataset can also be used as the training set.

[0049] The covariance mapped to the tangent space can be chosen as a feature. A short time window (30 ms) and a stride of 5 samples (corresponding to a data rate of 400 Hz) can be used for feature extraction. By applying principal component analysis (PCA) to 5 components, the dimensionality of the feature space can be reduced to find events in the dataset. Thereafter, the data can be centered (e.g., by removing the median), and finally, local maxima (peaks) can be identified on the first component.

[0050] Figure 2 Clusters that can be generated from the detected events (including the detected events recorded during the rest period) are shown. Three clusters are shown, respectively for each type of finger tap (data groups 100 and 102, which correspond to index finger tap and middle finger tap). For those events recorded during the rest period (which may not be considered useful events), an additional cluster may appear. This additional cluster 104 can be located in the lower left corner, which indicates a cluster with low energy samples. This cluster can be removed by discarding all corresponding events below, for example, a predetermined energy level threshold.

[0051] The data around each event can be sliced in epochs to prepare for clustering analysis. In one example, a 150 - ms window can be centered on each event to slice the data, and any other suitable window size can be used in a similar manner. Thereafter, each epoch can be vectorized, and a K - Means clustering process can be performed to extract three clusters. For visualization purposes, a dimensionality reduction process based on Uniform Manifold Approximation and Projection (UMAP) can be applied to plot Figure 2 the shown clusters (including approximately fifty events for each class of events).

[0052] Figure 3 A plot of the first component from principal component analysis (PCA) is shown, which can be performed on the detected discrete events. The data can be plotted with respect to the first component generated by the principal component analysis. In this example, the index finger events are shown first (on the left), followed by the rest period, and then the middle finger events on the right.

[0053] In some examples, temporal adjustment can be performed on the recorded events. The timing of each event can be associated with the local maxima on the first component identified by the execution of PCA analysis. Then, ground truth can be generated from the acquired samples to train an event detection model.

[0054] Figure 4A and Figure 4BAn epoch corresponding to a discrete event is illustrated, showing aspects of synchronization quality. There may be some jitter and misalignment in different epochs.

[0055] In some examples, by analyzing the autocorrelation between an epoch and the average of all events, the best offset can be found for each epoch, which can reduce or eliminate jitter and misalignment. Thus, different offsets (-10 to 10 samples) can be tested, and then the timing that maximizes the correlation can be selected. The testing process can be iteratively performed until all epochs are correctly aligned.

[0056] Figure 5A and Figure 5B An aligned epoch corresponding to the detected discrete event is shown.

[0057] Figure 6A and Figure 6B A graph corresponding to two templates of PCA analysis is shown, which can be performed on the averages of two different gestures. Figure 6A corresponds to index finger tap data, while Figure 6B corresponds to middle finger tap data. The templates can be based on the average energy of each event's epoch obtained after synchronization. The first PCA component (among the five components of PCA) may have a significant difference in amplitude between the taps of two fingers (index finger and middle finger), while the other components may have different signal forms.

[0058] When an event is detected (the event occurs), the binary time series can be marked with a value of 1, and when no event is detected (e.g., the event may not have occurred yet), the binary time series can be marked with a value of 0. A model for predicting such a time series can be trained based on the marked samples. Then the output of the model can be compared with a predetermined energy threshold, and debounced to configure the event detector.

[0059] Exemplary parameters can be configured for the ground truth of the model. After resynchronization, the event can be centered on the peak of the first PCA component. The model can rely on the complete event time process, and once the user finishes its execution, the model can predict the event. Thus, the marking can be shifted or deviated based on the event timing. This parameter can be called "offset".

[0060] In some examples, the model may not perfectly predict the correct single time sample corresponding to an event. Thus, the model can be configured to predict values (e.g., 1) on several consecutive time samples centered around the event. This parameter can be called "pulse width".

[0061] In some examples, the offset can be set at 75 ms after the event peak (about 30 samples after the peak of the event), and the pulse width can be set to 25 ms. These examples and other examples are non - restrictive, and depending on the particularity of the signals used during the training of the event detector model, other parameter values can be used.

[0062] Figure 7 Shows events with corresponding labels that can be detected using the first PCA component. These labels can be generated for 2 - second data. The event detector model can be implemented as a multi - layer perceptron (MLP) model or other suitable machine - learning model. Features can be collected from a 150 - ms (about 60 samples) sliding window over the PCA features (e.g., for each time sample, a vectorized vector of the first 60 time samples (i.e., 300 dimensions) with five PCA components can be generated).

[0063] The model can be trained to predict the labels used. The model can be applied to a test set, and the inferred output can be compared with a predetermined threshold and de - jittered to elicit the identification of discrete events.

[0064] Figure 8 Shows the detection of discrete events on a test set, including two outputs (solid lines) from the model and the discrete events (dashed lines) that can be generated from the test set.

[0065] Figure 9 Shows discrete events that can be detected in a test dataset, including, for example, five components generated using PCA analysis performed on the test set and events that can be detected in the same set. The model can detect all possible events, and there may be a clear disambiguation between two types of discrete events.

[0066] In some examples, events can be classified based on snapshots taken from EMG signals. The event detector can detect or record snapshots taken around the time event. The event classifier model can be trained to distinguish different types or classes of events. This classification is possible in part because each event is associated with a class or type of event that corresponds to a characteristic or stereotypical signal associated with a particular muscle activation synchronous with the event occurrence. 18 datasets can be used, and each dataset can be collected from a different user. These datasets include recordings of EMG signals captured from key - down, key - up, and tap events. The total number of events used by each user can be approximately 160 (80 for each index finger and middle finger).

[0067] The covariance can be estimated using a time window of 40 ms and a step size of 2.5 ms (which is generated by a feature sampling frequency of 400 Hz). Then, the covariance can be projected into the tangent space, and the dimensionality can be reduced by selecting the diagonal and two adjacent channels (represented by the values above and below the diagonal in the matrix). Applying the above operations results in a feature space with a dimensionality of 48.

[0068] Signal windows ranging from -100 ms to +125 ms around each key press event can be extracted (e.g., sliced and buffered). Such a window can include approximately 90 EMG sample values. At the end of the aforementioned operations, a dataset of size 160×90×48 (N_events × N_time_samples × N_features) can be obtained for each user.

[0069] Figure 10A and Figure 10B Examples of the index finger tap event model and the middle finger tap event model are shown respectively. A model for each event can be generated by taking the average of the EMG values for each event category (e.g., index finger tap and middle finger tap) of all such events that occur. Examples of tap events are shown in Figure 10A and Figure 10B are shown.

[0070] In Figure 10A and Figure 10B In the event models shown, two signals can be identified, one signal corresponding to a key press and one signal corresponding to a key release. The same features may appear to be active in both the index finger key press category and the middle finger key press category, but their respective amplitudes are significantly different and provide a good basis for discrimination.

[0071] Figures 11A - 11F Examples of user-specific event models for two classes of events are shown. Figure 11A 、 Figure 11C and Figure 11E correspond to index finger key presses, while Figure 11B 、 Figure 11D and Figure 11F correspond to middle finger key presses. Each user can display different patterns for each event category. Although the timing is usually the same, significant amplitude differences can be observed between the signals.

[0072] Several classification models can be used to implement a single-user event classification model. In some examples, each trial can be vectorized into a large vector (whose dimensionality corresponds to the number of time points × the number of features). Once such a large vector is generated, a classifier can be generated based on logistic regression, random forest, or a multi-layer perceptron, and the classifier can be implemented in the gesture classification model.

[0073] In some examples, the dimensionality of the data (in the feature dimension) can be reduced by applying a spatial filter, then vectorizing the result, and using a classifier. Examples of spatial filters can be based on the extraction of, for example, common spatial patterns (CSP), or the xDawn enhancement of evoked potentials integrated with linear discriminant analysis (LDA). By applying CSP, a subspace can be determined that maximizes the variance difference of the sources. In the xDawn method, the spatial filter can be estimated based on class means rather than the raw data (which can improve the signal-to-noise ratio (SNR)).

[0074] In some examples, a model can be developed by a method including one or more of the following: concatenating the event models of each class (e.g., middle finger key press and index finger key press) to each trial; estimating the covariance matrix; performing tangent space mapping, and applying LDA. Such an approach can produce a compact representation of the signal and may be effective in cases of low SNR.

[0075] A stratified random split with 90% training and 10% testing can be used in part to maintain class balance. A random split can also be used. Using a linear regression classifier can achieve an average accuracy of 99% for users, where for the worst user, an accuracy of 95% can be achieved.

[0076] Figure 12 The accuracy levels achieved by each test model for single-user event classification are shown. Each point in the figure represents a single user. The classifier can generally perform at an accuracy level similar to the Figure 12 accuracy levels shown.

[0077] The training set size can be modified. The size of the training set can be changed from 5% to 90% in the split. The amount of test data can be kept fixed at 10%. Two classifiers (LR and XDCov+LDA) can be used. Ten stratified random splits with 10% testing and variable training size can be used for cross-validation.

[0078] At around 80 events, the accuracy may reach a stable level. In the case of a classifier based on logistic regression, 20 events can be used to achieve 95% accuracy. The classifier based on XDCov+LDA may require more events to converge.

[0079] Figure 13AShows an example accuracy level as a function of the number of training events, which can be achieved by two different implementations of a single-user event classification model. The results of the LR (solid line) and XDCov+LDA (dashed line) methods are shown. The remaining dashed and dotted lines give a qualitative indication of the possible uncertainty of the LR results (the dotted line above and the dashed line roughly in the lower-middle) and the XDCov+LDA results (the remaining dashed lines and the dotted line below).

[0080] The window size can also be adjusted. The size of the window used to classify events may affect the latency of event detection. Therefore, the performance of the model may vary according to the window size parameter, which can be adjusted accordingly.

[0081] In some implementations, the single time point of classification can be used to reveal which time point contains information. Alternatively, an increasing window size (including all past time points) from -100 ms to +125 ms after a key press event, for example, can be used. For each time point or window size, a user-specific model can be trained, and then the performance of the resulting classifier or model can be evaluated. As discussed above, a logistic regression model or other suitable model can be used to implement the classifier. Cross-validation can be implemented using 10 stratified random splits (where 10% is reserved for testing purposes and 90% for training purposes). These values, as well as other values discussed in this article, are exemplary and not restrictive.

[0082] Figure 13B Shows an example accuracy level that can be achieved by a single timestamp and a cumulative window size. The results show that most time points in the window can contain information that allows the model to classify them above the chance level (with an accuracy of approximately 50%, for example). For a key press, the maximum accuracy can be achieved at -25 ms, while for a key release, the maximum accuracy can be achieved around +70 ms. Using a cumulative window that includes all past time samples, the maximum accuracy level can be achieved at the end of the window. An average accuracy level of 95% can be achieved using all timestamps before the key press event. Waiting for the release wave can improve the accuracy by providing supplementary information. The remaining dashed and dotted lines represent a qualitative indication of the possible uncertainty.

[0083] Cross-time generalization can be used to determine the degree of independence of time samples. As part of cross-time generalization, a classifier can be trained at a single time point and then tested at another time point. This method can determine whether the different processes involved in an event are stationary. If the same combination of sources is similarly active at two different time points, it may mean that the single-user model can be migrated or used to classify events generated by other users.

[0084] For each user and each time point, a classifier based on logistic regression can be trained. Then, the accuracy of each classifier can be evaluated at every other time point (for the same user). Then, the accuracy across all users and the structure of the accuracy matrix can be averaged.

[0085] Figure 14 Cross - time generalization that can be performed to determine the independence of time samples is shown. Two clusters can be observed in the accuracy matrix, one cluster corresponding to key presses and the other to key releases. From the transitions observed within each cluster, this can mean that each time sample does not carry much complementary information, and that using a carefully selected subset of samples may be sufficient to achieve optimal accuracy (alternatively, compressing the feature space with singular value decomposition SVD may be useful).

[0086] In some examples, a generalized cross - user classification model can be used. A classifier can be trained with data collected from several users, and the obtained trained classifier can be tested to obtain its performance on test users. As discussed above, several types of classifiers can be implemented to determine the best type of classifier. For cross - validation purposes, data extracted from one user may be left out. On average, the accuracy achieved on the implemented models may be around 82%. Large differences between users can also be observed.

[0087] Figure 15 The accuracy level of the generalized cross - user classification model is shown, and it is shown that some classifiers can achieve 100% accuracy while other classifiers may only achieve less than 60% accuracy. Figure 16 It is also shown that a classifier based on linear regression can be used to achieve a reasonable level of accuracy.

[0088] In some examples, model transfer across user pairs can be used. A classifier model can be trained based on data extracted from one user, and then the accuracy of the model can be evaluated with respect to the data of each other user. The classifier model can be based on logistic regression.

[0089] Figure 16 The transferability of user - specific classifiers based on linear regression is shown, indicating that a large variability in transfer accuracy can be observed. Some user - specific models can be transferred well to some other users. Some user - specific models seem to be good receivers (e.g., Figure 16 the user model for "Alex" shown in [reference]), which have good transferability for most other users, while other user - specific models (e.g., the user model for "Rob") do not seem to match well with other users.

[0090] In some examples, user adaptation can also be used. Based on an investigation of the single-user event classification model, it is even possible to separate the categories derived from a single user, and a reasonably accurate single-user event classification model can be obtained using a relatively small amount of labeled training data.

[0091] From the perspective of the generalized cross-user classification model results, it can be inferred that some user-specific classification models are sufficiently transferable to other users. Based on these preliminary results, the following examples are as follows. In some examples, models from other (different) users can be used to obtain a good estimate of the labels for the current user. Additionally, using this estimate of the labels, a user-specific model can be trained to obtain performance close to that of a single-user model trained with labeled data.

[0092] User embeddings can also be used. An embedding space can be generated in which two event categories can be clustered. The user transfer matrix indicates that for each test user, generally some (e.g., two) single-user models can be sufficiently transferred. A user embedding space can be constructed that includes the output of a collection of single-user models. Specifically, a simple nearest centroid classifier for covariance features (XDCov+MDM) can be constructed. The advantage of the XDCov+MDM method over linear regression or other alternative probability models is that even if the model may be improperly calibrated, the events may still contribute to cluster separability.

[0093] The output of the XDCov+MDM model can be a function of softmax, which is applied to the distances to the centroids of each event category. In some examples (e.g., binary classification), one dimension can be used for each user-specific pattern. However, the number of dimensions can be extended according to the classification type (e.g., classification that can be performed from a pool of more than two possible classes, e.g., greater than binary classification).

[0094] The embedding associated with a user can be trained using samples from all users (minus one user) in a group of users. Thereafter, samples associated with the user not used in the training of the embedding can be projected into the trained embedding. Thus, a space of X-1 dimensions can be produced, where X is the number of users from the group of users.

[0095] Figures 17A - 17Q Examples of the distribution of two classes (index finger tap and middle finger tap) for each dimension are shown. In some models, the separation of these two classes can be distinguished, while other models show approximately the same distribution. In some examples, when the model is not optimally calibrated (i.e., the optimal separation between classes may not be at 0.5), the model can still effectively separate the two classes.

[0096] After generating the embeddings as discussed above, a clustering process can be performed to separate clusters corresponding to different types of event categories (e.g., index finger taps and middle finger taps or pinches or snaps or other gesture types to be separated). For example, a K-means process can be run on the set of data points generated using the embeddings.

[0097] Figure 18A and Figure 18B An example of clusters separated using UMAP and PCA is shown, indicating that such clusters can be plotted using Uniform Manifold Approximation and Projection (UMAP) as in Figure 18A or Principal Component Analysis (PCA) as shown in Figure 18B Multiple clusters (e.g., two clusters) can be seen, and each of these clusters can correspond to different event categories (e.g., gesture types) and different labels. Since the embedding space conveys a meaning (which can be called "proba"), each cluster can be associated with its corresponding category.

[0098] A self-supervised user model can also be developed. After a set of labels can be generated using, for example, clustering techniques, these labels can be used to train a user-specific model from the original data set. For example, if it is known that the selected classification model will not substantially overfit the model and may be insensitive to the noise contained in the labeled data, then XDCov and linear displacement analysis, or other suitable classification models, can be implemented.

[0099] Figure 19 An example accuracy level achieved using a self-supervised model is shown, indicating that after training a self-supervised model, an accuracy of approximately 99% for label or classification estimation can be achieved. In this example, two training iterations may be sufficient.

[0100] An accuracy of 98% can be achieved using the full training set, which can include all the data points from a group of users.

[0101] Figure 20 An accuracy level achieved using a supervised user-specific model and a self-supervised user-specific model is shown, indicating that the self-supervised model performs better than the user-specific model trained with labeled data. The remaining dash and dotted lines give a qualitative indication of possible uncertainties.

[0102] The window size can be adjusted to improve the performance of the self-supervised model. As the window size increases, observing the accuracy of the self-supervised model can be used to determine the optimal window size. For cross-validation of the model, data from one user can be omitted. For clustering and user-specific models, a 10-fold random split with 10% test data and 90% training data can be used. In this case, it can be determined that the self-supervised model performs better at the full window size. This can be explained by observing that in this instance, a small window size does not produce separable clusters. Therefore, a large window size can be used to obtain labeled data, and then a relatively small window size (e.g., using the labels) can be used to train the user-specific model.

[0103] Figure 21 Shows the window size determination for a user-specific model (solid line) and a self-supervised model (lower short dash). The remaining short dashes and dotted lines give a qualitative indication of possible uncertainties.

[0104] A similar method can be used to study the effect of data size. An ensemble of single-user models can be used to evaluate performance. Cross-validation can include leaving one user out for alignment and then using the same 10-fold random split with 10% test data and a training size that increases from 5% to 90%. The ensemble method may reach 96% accuracy after 30 events, and then the accuracy may level off after that for a larger number of events.

[0105] Supervised domain adaptation can use a Canonical Partial Least Squares (CPLS) model. In some examples, a domain adaptation-based method can be used instead of building a user-specific model (e.g., by determining a data transformation that can lead to sufficient cross-user transfer). The CPLS model can be used to perform domain adaptation. A transformation function can be determined to align the models of each event category (e.g., different gesture types such as index finger tap, middle finger tap, index finger to thumb pinch, middle finger to thumb pinch, finger snap, etc.) of one user with the models of each event category of another user.

[0106] Figures 22A - 22B Shows the models of each event category associated with the first user.

[0107] Figures 22C - 22D Shows the models of each event category associated with the second user.

[0108] Figures 23A - 23BShows the alignment of models of event categories associated with a first user and a second user, indicating that the model of an event category of one user can be aligned with the corresponding model of an event category of another user. The vertical dashes correspond to button presses. The alignment may be effective, in part because the original models of each event category for the two users may be very different, but they may become almost identical after alignment.

[0109] The aligned data distribution can be studied by considering the UMAP embeddings of the data before and after transformation.

[0110] Figures 24A - 24B Shows example data before and after transformation. Figure 24A Shows that the original data can be clearly separated, and the largest variations can be seen between the two users. After transformation, the two event categories of the events can be matched with a high degree of accuracy (e.g., as Figure 24B shown).

[0111] The transformation process can be studied for each pair of users in a group of users. After performing alignment, the user-to-user migration matrix can be reproduced. A single-user model can be trained, and then for each test user, the data can be aligned, and the accuracy of the model can be tested on the transformed data. For test users, cross-validation can include estimating the event category model for the first 40 events (or other number of events), then performing domain adaptation, and finally testing the accuracy of the model on the remaining events (e.g., 120 events). The numerical values used in these (and other) examples are exemplary and not restrictive.

[0112] Figure 25A Shows the cross-user migration from all users in a group of users, indicating that the process can enhance the migration of a single-user model to any other user.

[0113] The amount of data required to achieve optimal adaptation can be determined. An ensemble of single-user models can be used for performance evaluation, in part because it is possible to adapt the data between user pairs. Cross-validation can include leaving out one user for alignment and then using a 10-fold random split (which has 10% test data and increases the training size from 5% to 90%). The numerical values are exemplary and not restrictive.

[0114] Figure 25B Shows the data size for supervised domain adaptation determined based on the migration function, indicating the relationship between accuracy and the number of training events. The results show that the ensemble can reach 96% accuracy after 30 events and may then level off. The remaining dashes and dotted lines give a qualitative indication of possible uncertainties.

[0115] Figures 26A - 26BAn example device is shown, which may include one or more of the following: a human-machine interface, an interface device, a control device, and / or a control interface. In some examples, the device may include a control device 2600, in which example (as Figure 26A shown), the control device 2600 may include a plurality (e.g., 16) of neuromuscular sensors 2610 (e.g., EMG sensors), which are circumferentially arranged around an elastic band 2620 configured to be worn around a user's lower arm or wrist. In some examples, the EMG sensors 2610 may be circumferentially arranged around the elastic band 2620. The band may include flexible electrical connections 2640 (shown in Figure 26B ), which may interconnect the discrete sensors and electronic circuits, which in some examples may be encapsulated in one or more sensor housings 2660. Each sensor 2610 may have a skin contact portion 2650, which may include one or more electrodes. Any suitable number of neuromuscular sensors 2610 may be used. The number and arrangement of the neuromuscular sensors may depend on the particular application for which the control device is used. For example, a wearable control device configured as an armband, wristband, or chest band may be used to generate control information to control an augmented reality system, control a robot, control a vehicle, scroll through text, control an avatar, or for any other suitable control task. As shown, the sensors may be coupled together using flexible electronics incorporated into a wireless device.

[0116] Figure 26B A cross-sectional view through one of the sensors 2610 of the control device 2600 shown in Figure 26A is shown. The sensor 2610 may include a plurality of electrodes located within the skin contact surface 2650. The elastic band 2620 may include an outer flexible layer 2622 and an inner flexible layer 2630, which may at least partially enclose the flexible electrical connector 2640.

[0117] In some embodiments, a hardware-based signal processing circuit may optionally be used to process the output of one or more sensing components (e.g., to perform amplification, filtering, rectification, and / or other suitable signal processing functions). In some embodiments, at least some of the signal processing of the output of the sensing components may be performed in software. Thus, the signal processing of the signals sampled by the sensors may be performed in hardware, in software, or by any suitable combination of hardware and software, as aspects of the techniques described herein are not limited in this regard. Non-limiting examples of analog circuits for processing signal data from the sensors 2610 are discussed in more detail below with reference to Figure 27A and Figure 27B .

[0118] Figure 27A and Figure 27B shows a schematic diagram of internal components of a device that may include one or more EMG sensors, such as 16 EMG sensors. The device may include a wearable device (such as the control device 2710 schematically shown in Figure 27A ), and a dongle portion 2750 (schematically shown in Figure 27B ) that can communicate with the control device 2710 (e.g., using Bluetooth or another suitable short-range wireless communication technology). In some examples, the functionality of the dongle portion (e.g., a circuit similar to the circuit shown in Figure 27B ) may be included within the head-mounted device, thereby allowing the control device to communicate with the head-mounted device.

[0119] Figure 27A shows that the control device 2710 may include one or more sensors 2712 (such as the sensors 2610 described above in connection with Figure 26A and Figure 26B ). Each sensor may include one or more electrodes. Sensor signals from the sensors 2712 may be provided to an analog front end 2714, which may be configured to perform analog processing of the sensor signals (e.g., noise reduction, filtering, etc.). The processed analog signal may then be provided to an analog-to-digital converter (ADC) 2716, which may convert the processed analog signal into a digital signal, which may then be further processed by one or more computer processors. Example computer processors that may be used according to some embodiments may include a microcontroller (MCU) 2722. The MCU 2722 may also receive signals from other sensors (e.g., inertial sensors such as an inertial measurement unit (IMU) sensor 2718, or other suitable sensors). The control device 2710 may also include a power supply 2720 or receive power from the power supply 2720, which may include a battery module or other power source. The output of the processing performed by the MCU 2722 may be provided to an antenna 2730 for transmission to Figure 27B the dongle portion 2750 shown in

[0120] Figure 27BThe dongle portion 2750 can include an antenna 2752, which can be configured to communicate with an antenna 2730 associated with the control device 2710. Communication between the antennas 2730 and 2752 can be performed using any suitable wireless technology and protocol, non-limiting examples of which include radio frequency signaling and Bluetooth. As shown, signals received by the antenna 2752 of the dongle portion 2750 can be received via a Bluetooth radio (or other receiver circuitry) and provided to the host computer via an output 2756 (e.g., a USB output) for further processing, display, and / or for effecting control over one or more specific physical or virtual objects.

[0121] In some examples, the dongle can be inserted into a separate computing device that can be located within the same environment as the user but is not carried by the user. The separate computing device can receive control signals from the control device and further process those signals to provide further control signals to the head-mounted device. The control signals can trigger the head-mounted device to modify the artificial reality view. In some examples, the dongle (or an equivalent circuit in the head-mounted device or other device) can be network-enabled, allowing communication with a remote computer via a network, and the remote computer can provide control signals to the head-mounted device to trigger the head-mounted device to modify the artificial reality view. In some examples, the dongle can be inserted into the head-mounted device to provide improved communication capabilities, and the head-mounted device can perform further processing (e.g., modify the AR image) based on control signals received from the control device 2710.

[0122] In some examples, the configuration of the dongle portion can be included within a head-mounted device (e.g., an artificial reality head-mounted device). In some examples, the circuitry described above in Figure 27B can be provided by components of the head-mounted device (i.e., integrated within the components of the head-mounted device). In some examples, the control device can communicate with the head-mounted device using the wireless communication described and / or similar illustrative circuitry or circuitry having similar functionality.

[0123] The head-mounted device can include something similar to that described above with respect to Figure 27BThe antenna of the described antenna 2752. The antenna of the head-mounted device can be configured to communicate with the antenna associated with the control device. Communication between the antennas of the control device and the head-mounted device can be carried out using any suitable wireless technology and protocol, non-limiting examples of which include radio frequency signaling and Bluetooth. Signals received by the antenna of the head-mounted device (such as control signals) can be received by a Bluetooth radio (or other receiver circuitry) and provided to a processor within the head-mounted device, which can be programmed to modify the artificial reality view for the user in response to the control signal. For example, in response to the detected type of gesture, the control signal can trigger the head-mounted device to modify the artificial reality view presented to the user.

[0124] Example devices can include a control device and one or more devices (such as one or more dongle portions, head-mounted devices, remote computer devices, etc.) that communicate with the control device (e.g., via Bluetooth or another suitable short-range wireless communication technology). The control device can include one or more sensors, which can include electrical sensors (including one or more electrodes). The electrical output from the electrodes (which can be referred to as sensor signals) can be provided to an analog circuit that is configured to perform analog processing of the sensor signals (such as filtering, etc.). The processed sensor signals can then be provided to an analog-to-digital converter (ADC) that can be configured to convert the analog signals into digital signals that can be processed by one or more computer processors. Example computer processors can include one or more microcontrollers (MCUs), such as the nRF52840 (manufactured by NORDIC SEMICONDUCOTR). The MCU can also receive inputs from one or more other sensors. The device can include one or more other sensors (such as an orientation sensor), which can be an absolute orientation sensor and can include an inertial measurement unit. Example orientation sensors can include the BNO055 inertial measurement unit (manufactured by BOSCH SENSORTEC). The device can also include a dedicated power supply (such as a power supply and battery module). The output of the processing performed by the MCU can be provided to the antenna for transmission to the dongle portion or another device. Other sensors can include myogram (MMG) sensors, surface electromyogram (SMG) sensors, electrical impedance tomography (EIT) sensors, and other suitable types of sensors.

[0125] The dongle portion or other devices such as a head-mounted device may include one or more antennas configured to communicate with a control device and / or other devices. Communication between system components may use any suitable wireless protocol (e.g., radio frequency signaling and Bluetooth). Signals received by the antenna of the dongle portion (or other device) may be provided to a computer via an output (e.g., a USB output) for further processing, display, and / or to effect control of one or more specific physical or virtual objects.

[0126] Although the examples provided Figure 26A 、 26B and Figure 27A 、 27B are discussed in the context of an interface having an EMG sensor, the examples may also be implemented in a control device (e.g., a wearable interface used with other types of sensors, which other types of sensors include but are not limited to mechanomyogram (MMG) sensors, surface electromyogram (SMG) sensors, and electrical impedance tomography (EIT) sensors). The methods described herein may also be implemented in a wearable interface that communicates with a computer host via wires and cables (e.g., a USB cable, an optical fiber cable).

[0127] Figure 28 An example system 2800 is shown, which may include a head-mounted device 2810 and a control device 2820 (which may represent a wearable control device). In some examples, system 2800 may include a magnetic tracker. In these examples, a transmitter for the magnetic tracker may be mounted on the control device 2820, and a receiver for the magnetic tracker may be mounted on the head-mounted device 2810. In other examples, the transmitter for the magnetic tracker may be mounted on the head-mounted device or otherwise located within the environment. In some embodiments, system 2800 may also include one or more optional control gloves 2830. In some examples, many or all of the functions of the control glove may be provided by the control device 2820. In some examples, the system may be an augmented reality and / or virtual reality system. In some examples, the control glove 2830 may include multiple magnetic tracker receivers that may be used to determine the orientation and / or position of various parts of the user's hand. In some examples, the control device 2820 may be similar to Figure 26A and Figure 26B the control device shown. In some examples, the control device may include an electronic circuit similar to Figure 27A (and / or Figure 27B ) the electronic circuit shown.

[0128] In some examples, the control glove 2830 (which may be more simply referred to as a glove) may include one or more magnetic tracker receivers. For example, the fingers of the glove may include at least one receiver coil, and detection of tracker signals induced by a magnetic tracker transmitter from at least one receiver coil may be used to determine the position and / or orientation of at least a portion of the finger. One or more receiver coils may be associated with each part of the hand (such as fingers (e.g., thumb), palm, etc.). The glove may also include other sensors (such as electroactive sensors) that provide sensor signals indicative of the position and / or configuration of the hand. The sensor signals (such as magnetic tracker receiver signals) may be transmitted to a control device (such as a wearable control device). In some examples, a control device (such as a wrist-worn control device) may communicate with the control glove and receive sensor data from the control glove using wired and / or wireless communication. For example, a flexible electrical connector may extend between the control device (such as, a wrist-worn control device) and the glove. In some examples, the control device may include the glove, and / or may include a wristband.

[0129] In some examples, the control device 2820 may include an EMG control interface that is similar to Figure 26A and Figure 26B the device shown. Positioning the magnetic tracker transmitter on or near the control device 2820 may cause noise to be introduced into the signals recorded by the control device 2820 due to induced current and / or voltage. In some embodiments, electromagnetic interference caused by the magnetic tracker transmitter may be reduced by positioning the transmitter at a greater distance from the control device 2820. For example, the transmitter may be mounted on the head-mounted device 2810, while the magnetic tracker receiver may be mounted on the control device 2820. For example, this configuration works well when the user keeps their arm away from their head, but may not work well if the user moves their arm to a position close to the head-mounted device. However, many applications do not require extensive proximity between the user's head and hand.

[0130] A control device (such as the wearable control device 2820) may include an analog circuit and an analog-to-digital converter. The analog circuit includes at least one amplifier configured to amplify analog electrical signals originating from the user's body (such as from electrodes in contact with the skin, and / or one or more other sensors), and the analog-to-digital converter is configured to convert the amplified analog electrical signals into digital signals that may be used to control the system (such as a virtual reality (VR) and / or augmented reality (AR) system).

[0131] In some examples, an augmented reality system can include a magnetic tracker. The magnetic tracker can include a transmitter located in the head-mounted device or elsewhere and one or more receivers that can be associated with an object to be tracked or a body part of a user (such as a user's hand, or other limb or part thereof, or a joint).

[0132] Figure 29 An example method (2900) for classifying an event is shown, which includes: obtaining electromyogram (EMG) data (2910) from a user, the EMG data including an EMG signal corresponding to the event; detecting the EMG signal corresponding to the event (2920); classifying the EMG signal as from an event type (2930); and generating a control signal (2940) based on the event type. The example method may further include triggering a head-mounted device based on the control signal to modify an artificial reality view, thereby controlling an artificial reality environment.

[0133] Figure 30 An example method (3000) for classifying an event is shown, which includes: detecting an EMG signal corresponding to the event (3010); using a trained model to classify the EMG signal as corresponding to an event type (3020); and generating a control signal for an artificial reality environment (3030) based on the event type. The control signal can trigger a head-mounted device to modify an artificial reality view.

[0134] In some examples, Figure 29 and Figure 30 can represent flowcharts of example computer-implemented methods for detecting at least one gesture and using at least one detected gesture type to control an artificial reality system (such as an augmented reality system or a virtual reality system). One or more of the steps shown in the figures can be performed by any suitable computer-executable code and / or computer device (such as a control device, a head-mounted device, other computer devices communicating with the control device, or computer devices communicating with a device providing sensor signals). In some examples, Figure 29 and / or Figure 30 one or more of the steps shown in can represent an algorithm using a method (such as those described herein), the structure of the algorithm including multiple sub-steps and / or being represented by multiple sub-steps. In some examples, the steps of a particular example method can be performed by different components of the system (including, for example, a control device and a head-mounted device).

[0135] In some examples, event detection and classification can be performed by unsupervised or self-supervised models, and these methods can be used to detect user gestures. The model can be trained for a specific user, or in some examples, the model can be trained on different users, and the training data of these different users is suitable for use with the current user. In an example training method, EMG data can be detected and optionally recorded for analysis. The EMG data can be used to train the model, and the EMG data can be obtained when one or more users perform one or more gestures. Example gestures include finger tapping (e.g., simulated button presses), other finger movements (e.g., finger curling, swiping, pointing gestures, etc.) or other types of gestures, and / or sensor data can be similarly used to train an example event detector model.

[0136] In some embodiments, by constructing an embedding space including a single-user model, clearly separable event clusters can be obtained. Clustering techniques can be implemented to determine the labels of each event, and then the labeled data can be used to train a user-specific model. By using at least one of these techniques, a very high accuracy (e.g., 98% accuracy) can be achieved in a completely unsupervised manner. For example, using a relatively small number of samples (e.g., fewer than 40 event samples), a relatively high accuracy (e.g., 95% accuracy) can be achieved.

[0137] In addition, in some embodiments, the single-user event template can be adapted to other users, thereby further reducing the additional amount of data that may be required to use the model with the adapted users. For example, PLS can be used to adapt the domain by aligning the datasets across users. For example, PLS can be trained to align event templates across users. The integration of the aligned user templates can result in a high accuracy (e.g., 96% accuracy), thus requiring very little event data to be collected (e.g., fewer than 10 events).

[0138] A pose can be defined as a body position that is stationary over time and can theoretically be maintained indefinitely. In contrast, in some examples, a gesture can be defined as including a dynamic body position, which can have a start time and an end time each time it occurs. Thus, a gesture can be defined as a discrete event of a specific gesture type. Representative examples of gesture types include finger snapping, finger tapping, finger curling or bending, pointing, swiping, rotating, grasping, or other finger movements. In some examples, a gesture can include the movement of at least a portion of the arm, wrist, or hand or other muscle activations. In some examples, visually perceptible movement of the user may not be required, and a gesture can be defined by a muscle activation pattern (independent of any visually perceptible movement of a part of the user's body).

[0139] When a gesture event is detected (e.g., in a continuous electromyography (EMG) data stream), a generic event detector can generate an output signal. A control signal for a computer device (e.g., an artificial reality system) can be based on the output signal of the generic event detector. The generic event detector can generate an output signal each time a user performs a gesture. In some examples, the output signal can be generated independent of the type of gesture performed. In some examples, when the event detector detects an event (e.g., a gesture), an event classifier can be executed. The event classifier can determine information related to the gesture performed by the user (e.g., gesture type). The gesture type can include one or more of the following: the physical action performed, the body part used to perform the physical action (e.g., a finger or other body part), the user's intended action, and other physical actions performed at the same time or approximately the same time. The control signal can also be based on a combination of sensor data from one or more sensor types. The corresponding control signal can be sent to an augmented reality (AR) system, and the control signal can be at least partially based on the gesture type. The control signal can modify the artificial reality display by one or more of the following: selection of an item, execution of a task, movement of an object a certain degree and / or direction that can be at least partially determined by the gesture type, interaction with a user interface of an object (e.g., a real or virtual object), or other actions. In some embodiments, a gesture can be classified as a specific gesture type based on one or more electromyography signals (e.g., EMG wavelets).

[0140] In some examples, a method for detecting an event (e.g., a gesture) can include: obtaining a first set of electromyography (EMG) data including EMG signals corresponding to a gesture of a first user; training a first classifier by clustering event data determined from the obtained first set of EMG signals; using the first classifier to label a second set of obtained EMG data; and using the labeled second set of EMG data to train an event detector.

[0141] In some examples, a method for classifying an event (e.g., a gesture) can include one or more of the following steps: generating a plurality of single-user event classifiers; using the plurality of single-user classifiers to generate a multi-user event classifier; using the generated multi-user classifier to label electromyography (EMG) data; generating a data transformation corresponding to a plurality of users; generating a single-user classifier related to a first user among the plurality of users; using the data transformation for a second user among the plurality of users and the single-user classifier for the first user to label received EMG data of the second user; and using the labeled EMG data to train an event detector.

[0142] In some examples, a method for training an event detector (such as a gesture detector) is provided. The method may include one or more of the following steps: obtaining electromyogram (EMG) data including EMG signals corresponding to a gesture; generating feature data from the EMG data; detecting an event in the feature data; generating epochs using the feature data, where each epoch may be centered around one of the detected events; clustering the epochs into types, where at least one type may correspond to a gesture; aligning the epochs by type to generate aligned epochs; training a tagging model using the aligned epochs; tagging the feature data using the tagging model to generate tagged feature data; and training the event detector using the tagged feature data.

[0143] In some examples, a method for training an event classifier may include one or more of the following steps: obtaining electromyogram (EMG) data including EMG signals corresponding to a plurality of gestures; generating feature data from the EMG data; detecting an event in the feature data using an event detector; generating epochs using the feature data, each epoch centered around one of the detected events; generating a single-user event classification model using the epochs; tagging the EMG data using the single-user event classification model; and training the event classifier using the tagged EMG data.

[0144] In some examples, a method for generating a single-user event classification model using epochs may include one or more of the following steps: generating vectorized epochs using the epochs; and generating a single-user event classification model by training one or more of a logistic regression, a random forest, or a multi-layer perceptron classifier using the vectorized epochs. In some examples, generating a single-user event classification model using epochs includes: generating spatially filtered and dimension-reduced epochs using the epochs; generating vectorized epochs using the spatially filtered and dimension-reduced epochs; and generating a single-user event classification model by training one or more of a logistic regression, a random forest, or a multi-layer perceptron classifier using the vectorized epochs. In some examples, generating a single-user event classification model using epochs includes: generating one or more event models using the epochs, each event model corresponding to a gesture; generating combined epochs by combining each epoch with one or more event models; and generating a single-user event classification model by training one or more of a logistic regression, a random forest, or a multi-layer perceptron classifier using the combined epochs.

[0145] In some examples, a method for training an event classifier is provided. The method may include one or more of the following steps: obtaining electromyogram (EMG) data including EMG signals corresponding to multiple gestures of multiple users; generating feature data from the EMG data; detecting an event in the feature data using an event detector; generating an epoch using the feature data, each epoch centered on one of the detected events; generating a cross-user event classification model using the epochs; labeling the EMG data using the cross-user event classification model; and training the event classifier using the labeled EMG data.

[0146] In some examples, a method for training an event classifier is provided. The method may include one or more of the following steps: generating an embedding model using multiple single-user event classification models; generating an embedded event using the embedding model and electromyogram (EMG) data including EMG signals corresponding to multiple gestures of a user; clustering the embedded events into clusters corresponding to multiple gestures; associating labels with the EMG data based on the clustered embedded events; and training an event classifier for the user using the EMG data and the associated labels.

[0147] In some examples, a method for training an event classifier is provided. The method may include one or more of the following steps: generating an event template for each of multiple events for each of multiple users; determining an alignment transformation between the event templates for each of the multiple events across the multiple users; transforming the EMG data of a first user using the alignment transformation determined for a second user; associating labels with the EMG data using the transformed EMG data and a single-user event classification model of the second user; and training an event classifier for the user using the EMG data and the associated labels.

[0148] In some examples, a system for gesture detection is provided. The system may include at least one processor and at least one non-transitory memory including instructions that, when executed by the at least one processor, cause the system for gesture detection to perform operations including: associating an event label with a portion of electromyogram data using an event detector; in response to associating the event label with the portion of the electromyogram data, associating a gesture label with the portion of the electromyogram data using an event classifier; and outputting an indication of at least one of the event label or the gesture label.

[0149] Examples described herein may include various suitable combinations of example aspects, provided that the aspects are not incompatible.

[0150] Example systems and methods can include user-based models for detecting gestures in an accurate and unsupervised manner. Event detector models are provided—these event detector models can be trained on a limited user dataset of a particular user and using labeling and clustering methods, which can improve the accuracy of the event detector while limiting the number of event data instances.

[0151] By constructing an embedding space that includes single-user models, clearly separable event clusters can be obtained. Clustering techniques can be implemented to determine the labels for each event, and then the labeled data can be used to train user-specific models. In some examples, by applying the process in a completely unsupervised manner, an accuracy of 98% can be achieved. Additionally, an accuracy of 95% can be achieved using a limited number (e.g., 40) of event samples.

[0152] Domain adaptation using PLS can include one or more of the following. The datasets of cross-user pairs can be aligned by training PLS to align event templates. The integration of the aligned single users can result in an accuracy of 96%. The alignment requires very little data to be executed (e.g., less than 10 events).

[0153] When a gesture event is detected in a continuous electromyogram (EMG) data stream, a general event detector can emit an output signal. Example general event detectors can generate an output signal each time a user performs a gesture, and can generate this output signal independently of the type of gesture performed.

[0154] When the event detector recognizes a gesture event, an event classifier can be executed. The event classifier can then determine the type of gesture performed by the user.

[0155] In some examples, a method for detecting an event can include one or more of the following: obtaining a first set of electromyogram (EMG) data including EMG signals corresponding to gestures of a first user; training a first classifier by clustering event data determined from the obtained first set of EMG signals; and using the first classifier to label a second obtained set of EMG data; and using the labeled second set of EMG data to train an event detector. Example methods can include providing a general event detector.

[0156] In some examples, a method for classifying events may include one or more of the following: generating multiple single-user event classifiers; using the multiple single-user classifiers to generate a multi-user event classifier; using the generated multi-user classifier to label electromyography (EMG) data; generating data transforms corresponding to multiple users; generating a single-user classifier related to a first user among the multiple users; using a data transform for a second user among the multiple users and the single-user classifier for the first user to label received EMG data for the second user; and using the labeled EMG data to train an event detector. The example method may include providing a general event classifier.

[0157] In some examples, a method for training an event detector may include one or more of the following: obtaining electromyography (EMG) data including EMG signals corresponding to gestures; generating feature data from the EMG data; detecting events in the feature data; using the feature data to generate epochs, each epoch centered on one of the detected events; clustering the epochs into types, with at least one type corresponding to a gesture; aligning the epochs by type to generate aligned epochs; using the aligned epochs to train a labeling model; using the labeling model to label the feature data to generate labeled feature data; and using the labeled feature data to train an event detector. The example method may include generating a classifier to label unlabeled data and then using the labeled data to generate an event detector.

[0158] In some examples, a method for training an event classifier may include one or more of the following: obtaining electromyography (EMG) data including EMG signals corresponding to multiple gestures; generating feature data from the EMG data; using an event detector to detect events in the feature data; using the feature data to generate epochs, each epoch centered on one of the detected events; using the epochs to generate a single-user event classification model; using the single-user event classification model to label the EMG data; and using the labeled EMG data to train the event classifier. The example method may include generating a single-user event classification model to label unlabeled data and then using the labeled data to generate the event classifier.

[0159] In some examples, using epochs to generate a single-user event classification model may include one or more of the following: using the epochs to generate vectorized epochs; and generating a single-user event classification model by training one or more of a logistic regression, random forest, or multi-layer perceptron classifier using the vectorized epochs. The example method may include generating a single-user event classification model from vectorized trials.

[0160] In some examples, generating a single-user event classification model using epochs may include one or more of the following: generating spatially filtered and dimension-reduced epochs using epochs; generating vectorized epochs using the spatially filtered and dimension-reduced epochs; and generating a single-user event classification model by training one or more of a logistic regression, a random forest, or a multi-layer perceptron classifier using the vectorized epochs. The method can be used to generate a single-user event classification model based on dimension-reduced data generated by spatially filtering a trial.

[0161] In some examples, generating a single-user event classification model using epochs may include one or more of the following: generating one or more event models using epochs, each event model corresponding to a gesture; generating combined epochs by combining each epoch with one or more event models; and generating a single-user event classification model by training one or more of a logistic regression, a random forest, or a multi-layer perceptron classifier using the combined epochs. An example method may include generating a single-user event classification model by generating an event template and connecting the event template with a trial.

[0162] In some examples, a method for training an event classifier includes one or more of the following: obtaining electromyography (EMG) data including EMG signals corresponding to multiple gestures of multiple users; generating feature data from the EMG data; detecting events in the feature data using an event detector; generating epochs using the feature data, each epoch centered on one of the detected events; generating a cross-user event classification model using the epochs; and using the cross-user event classification model to label the EMG data; and training an event classifier using the labeled EMG data. An example method may include generating a cross-user event classification model to label unlabeled data and then using the labeled data to generate an event classifier.

[0163] In some examples, a method for training an event classifier may include one or more of the following: generating an embedding model using multiple single-user event classification models; generating embedded events using the embedding model and electromyography (EMG) data including EMG signals corresponding to multiple gestures of a user; clustering the embedded events into clusters corresponding to multiple gestures; associating labels with the EMG data based on the clustered embedded events; and training an event classifier for the user using the EMG data and the associated labels. An example method may include generating an event classification model independent of the user to label unlabeled data based on an ensemble of single-user event classification models and then using the labeled data to generate an event classifier.

[0164] In some examples, a method for training an event classifier can include one or more of the following: generating an event template for each of a plurality of events for each of a plurality of users; determining an alignment transformation between event templates for each of the plurality of events across the plurality of users; transforming EMG data of a first user using at least one of the alignment transformations determined for a second user; associating a label with the EMG data using the transformed EMG data and a single-user event classification model of the second user; and training an event classifier for the user using the EMG data and the associated label. The example method can include transforming data using an alignment transformation between users in order to be labeled by a single-user specific event classification model and then using the labeled data to generate an event classifier.

[0165] In some examples, a system for gesture detection can be configured to identify gestures using an event detector and classify gestures using an event classifier, where the event detector can be trained using a training method (such as the training methods described herein). In some examples, a system for gesture detection can include: at least one processor; and at least one non-transitory memory including instructions that, when executed by the at least one processor, cause the system for gesture detection to perform operations including: associating an event label with a portion of electromyography data using the event detector; in response to associating the event label with the portion of electromyography data, associating a gesture label with the portion of electromyography data using the event classifier; and outputting an indication of at least one of the event label or the gesture label.

[0166] Exemplary computer-implemented methods can be performed by any suitable computer-executable code and / or computing system, where one or more steps of the method can represent an algorithm, the structure of which can include multiple sub-steps and / or can be represented by multiple sub-steps.

[0167] In some examples, a system includes at least one physical processor and a physical memory including computer-executable instructions that, when executed by the physical processor, cause the physical processor to perform one or more methods or method steps as described herein. In some examples, a computer-implemented method can include the detection and classification of gestures and the control of an artificial reality system using the detected gesture type.

[0168] In some examples, a non-transitory computer-readable medium includes one or more computer-executable instructions that, when executed by at least one processor of a computing device, cause the computing device to perform one or more method steps as described herein. In some examples, a computer-implemented method may include detecting and classifying a gesture and using the detected gesture type to control an artificial reality system.

[0169] Examples include a control device that includes a plurality of electromyography (EMG) sensors and / or other sensors and at least one physical processor programmed to receive sensor data, detect sensor signals within the sensor data corresponding to a user gesture, classify the sensor signals to identify a gesture type, and provide a control signal based on the gesture type. The control signal may trigger a head-mounted device to modify an artificial reality view.

[0170] Exemplary Embodiment

[0171] Example 1. An example system includes: a head-mounted device configured to present an artificial reality view to a user; a control device including a plurality of electromyography (EMG) sensors, the electromyography (EMG) sensors including electrodes that contact a user's skin when the user wears the control device; at least one physical processor; and a physical memory including computer-executable instructions that, when executed by the physical processor, cause the physical processor to: process one or more EMG signals detected by the EMG sensors; classify the processed one or more EMG signals into one or more gesture types; provide a control signal based on the gesture type, wherein the control signal triggers the head-mounted device to modify at least one aspect of the artificial reality view.

[0172] 2. The system of Example 1, wherein the at least one physical processor is located within the control device.

[0173] Example 3. The system of any one of Examples 1-2, wherein the at least one physical processor is located within the head-mounted device or within an external computer device in communication with the control device.

[0174] Example 4. The system of any one of Examples 1-3, wherein, when executed by the physical processor, the computer-executable instructions cause the physical processor to classify the processed EMG signals into one or more gesture types using a classifier model.

[0175] Example 5. The system of any one of Examples 1-4, wherein the classifier model is trained using training data including a plurality of EMG training signals for the gesture type.

[0176] Example 6. The system according to any one of Examples 1-5, wherein the training data is obtained from a plurality of users.

[0177] Example 7. The system according to any one of Examples 1-6, wherein the head-mounted device includes a virtual reality head-mounted device or an augmented reality device.

[0178] Example 8. An example method includes: obtaining one or more electromyogram (EMG) signals from a user; processing the one or more EMG signals to generate associated feature data; classifying the associated feature data into one or more gesture types using a classifier model; and providing a control signal to an artificial reality (AR) device based on the one or more gesture types, wherein the classifier model is trained using training data including a plurality of EMG training signals for the one or more gesture types.

[0179] Example 9. The method according to Example 8, wherein the classifier model is trained by clustering the feature data determined from the EMG training signals.

[0180] Example 10. The method according to any one of Examples 8-9, wherein the classifier model is trained using EMG training signals obtained from a plurality of users.

[0181] Example 11. The method according to any one of Examples 8-10, wherein the plurality of users does not include the user.

[0182] Example 12. The method according to any one of Examples 8-11, further comprising training the classifier model by: obtaining EMG training signals corresponding to gesture types; training the classifier model by clustering the EMG training data obtained from the EMG training signals.

[0183] Example 13. The method according to any one of Examples 8-12, wherein the classifier model is further trained by: determining the time dependence of the EMG training signals with respect to the time of the maximum value of the corresponding EMG training signals; aligning the time dependence of the plurality of EMG training signals by adding a time offset to at least one of the plurality of EMG training signals; obtaining signal features from the aligned plurality of EMG training signals; and training the classifier model to detect EMG signals having the signal features.

[0184] Example 14. The method according to any one of Examples 8 - 13, wherein the classifier model is further trained by the following steps: obtaining training data including EMG training signals corresponding to gesture types; and averaging the EMG training signals for each occurrence corresponding to the gesture type to obtain a gesture model for the gesture type, wherein the classifier model classifies EMG signals using the gesture model.

[0185] Example 15. The method according to any one of Examples 8 - 14, wherein the gesture model is a user - specific gesture model for the gesture type.

[0186] Example 16. The method according to any one of Examples 8 - 14, wherein the gesture model is a multi - user gesture model based on EMG training data obtained from multiple users, and the multi - user gesture model is a combination of multiple user - specific gesture models.

[0187] Example 17. The method according to any one of Examples 8 - 16, wherein the artificial reality device includes a head - mounted device configured to present artificial reality images to a user, and the method further includes modifying the artificial reality images based on a control signal.

[0188] Example 18. The method according to any one of Examples 8 - 17, wherein modifying the artificial reality images includes selecting or controlling an object in the artificial reality image based on the gesture type.

[0189] Example 19. A non - transitory computer - readable medium including one or more computer - executable instructions that, when executed by at least one processor of a computing device, cause the computing device to: receive one or more electromyogram (EMG) signals detected by an EMG sensor; process the one or more EMG signals to identify one or more features corresponding to a user gesture type; use the one or more features to classify the one or more EMG signals as a gesture type; provide a control signal based on the gesture type; and transmit the control signal to a head - mounted device to trigger a modification of an artificial reality view in response to the control signal.

[0190] Example 20. The non - transitory computer - readable medium according to Example 19, wherein the computing device is configured to classify EMG signals to identify a gesture type based on a gesture model determined according to training data obtained from multiple users.

[0191] Embodiments of the present disclosure may include or be implemented in conjunction with various types of artificial reality systems. Artificial reality is a form of reality that has been adjusted in some manner before being presented to a user, and may include, for example, virtual reality, augmented reality, mixed reality, hybrid reality, or some combination and / or derivative thereof. Artificial reality content may include entirely computer-generated content or computer-generated content combined with captured (e.g., real-world) content. Artificial reality content may include video, audio, haptic feedback, or some combination thereof, any of which may be presented in a single channel or multiple channels (e.g., stereoscopic video that produces a three-dimensional (3D) effect for a viewer). Additionally, in some embodiments, artificial reality may also be associated with applications, products, accessories, services, or some combination thereof, which are used, for example, to create content in artificial reality and / or otherwise use in artificial reality (e.g., perform activities in artificial reality).

[0192] Artificial reality systems may be implemented in a variety of different form factors and configurations. Some artificial reality systems may be designed to operate without a near-eye display (NED). Other artificial reality systems may include an NED that also provides visibility of the real world (e.g., Figure 31 the augmented reality system 3100 in Figure 32 ), or that immerses a user visually in the artificial reality (e.g.,

[0193] Turning to Figure 31 , the augmented reality system 3100 may include a glasses device 3102 having a frame 3110 configured to hold a left display device 3115(A) and a right display device 3115(B) in front of a user's eyes. The display devices 3115(A) and 3115(B) may act together or independently to present an image or a series of images to the user. While the augmented reality system 3100 includes two displays, embodiments of the present disclosure may be implemented in augmented reality systems having a single NED or more than two NEDs.

[0194] In some embodiments, the augmented reality system 3100 may include one or more sensors, such as sensor 3140. Sensor 3140 may generate measurement signals in response to the movement of the augmented reality system 3100 and may be located on substantially any part of the frame 3110. Sensor 3140 may represent a position sensor, an inertial measurement unit (IMU), a depth camera assembly, a structured light emitter and / or detector, or any combination thereof. In some embodiments, the augmented reality system 3100 may include or may not include sensor 3140, or may include more than one sensor. In embodiments where sensor 3140 includes an IMU, the IMU may generate calibration data based on the measurement signals from sensor 3140. Examples of sensor 3140 may include, but are not limited to, an accelerometer, a gyroscope, a magnetometer, other suitable types of sensors for detecting movement, sensors for error correction of the IMU, or some combination thereof.

[0195] The augmented reality system 3100 may also include a microphone array having a plurality of acoustic transducers 3120(A)-1320(J) (collectively referred to as acoustic transducers 3120). The acoustic transducers 3120 may be transducers that detect changes in air pressure caused by sound waves. Each acoustic transducer 3120 may be configured to detect sound and convert the detected sound into an electronic format (e.g., analog or digital format). Figure 2 The microphone array in may include, for example, ten acoustic transducers: 3120(A) and 3120(B), which may be designed to be placed within the respective ears of the user; acoustic transducers 3120(C), 3120(D), 3120(E), 3120(F), 3120(G), and 3120(H), which may be located at various positions on the frame 3110; and / or acoustic transducers 3120(I) and 3120(J), which may be located on the respective neckbands 3105.

[0196] In some embodiments, one or more of the acoustic transducers 3120(A)-3120(F) may be used as output transducers (e.g., speakers). For example, acoustic transducers 3120(A) and / or 3120(B) may be earbuds or any other suitable type of headphones or speakers.

[0197] The configuration of the acoustic transducers 3120 in the microphone array may vary. Although the augmented reality system 3100 is in Figure 31It is shown as having ten acoustic transducers 3120, but the number of acoustic transducers 3120 can be greater than or less than ten. In some embodiments, using a greater number of acoustic transducers 3120 can increase the amount of audio information collected and / or the sensitivity and accuracy of the audio information. Conversely, using a smaller number of acoustic transducers 3120 can reduce the computational power required for the associated controller 3150 to process the collected audio information. Additionally, the position of each acoustic transducer 3120 in the microphone array can vary. For example, the position of the acoustic transducer 3120 can include defined positions on the user, defined coordinates on the frame 3110, the orientation associated with each acoustic transducer 3120, or some combination thereof.

[0198] The acoustic transducers 3120(A) and 3120(B) can be located at different parts of the user's ear, such as behind the pinna, behind the tragus, and / or within the auricle or fossa. Alternatively, in addition to the acoustic transducers 3120 within the ear canal, there can be additional acoustic transducers 3120 on or around the ear. Positioning the acoustic transducers 3120 near the user's ear canal can enable the microphone array to collect information about how sound reaches the ear canal. By positioning at least two acoustic transducers 3120 on opposite sides of the user's head (e.g., as a binaural microphone), the augmented reality device 3100 can simulate binaural hearing and capture the 3D stereo sound field around the user's head. In some embodiments, the acoustic transducers 3120(A) and 3120(B) can be connected to the augmented reality system 3100 using a wired connection 3130, and in other embodiments, the acoustic transducers 3120(A) and 3120(B) can be connected to the augmented reality system 3100 using a wireless connection (e.g., a Bluetooth connection). In other embodiments, the acoustic transducers 3120(A) and 3120(B) may not be used in combination with the augmented reality system 3100 at all.

[0199] The acoustic transducers 3120 on the frame 3110 can be positioned along the length of the temple, across the bridge, above or below the display devices 3115(A) and 3115(B), or some combination thereof. The acoustic transducers 3120 can be oriented such that the microphone array can detect sounds in a wide range of directions around the user wearing the augmented reality system 3100. In some embodiments, an optimization process can be performed during the manufacture of the augmented reality system 3100 to determine the relative positioning of each acoustic transducer 3120 in the microphone array.

[0200] In some examples, the augmented reality system 3100 can include or be connected to an external device (e.g., a paired device), such as a neckband 3105. The neckband 3105 generally represents any type or form of paired device. Thus, the following discussion of the neckband 3105 can also apply to various other paired devices, such as a charging case, a smartwatch, a smartphone, a wristband, other wearable devices, a handheld controller, a tablet computer, a laptop computer, other external computing devices, etc.

[0201] As shown, the neckband 3105 can be coupled to the glasses device 3102 using one or more connectors. The connectors can be wired or wireless and can include electrical and / or non-electrical (e.g., structural) components. In some cases, the glasses device 3102 and the neckband 3105 can operate independently without any wired or wireless connection between them. While Figure 31 components of the glasses device 3102 and the neckband 3105 are shown in example locations on the glasses device 3102 and the neckband 3105, these components can be located elsewhere on the glasses device 3102 and / or the neckband 3105 and / or be distributed differently on the glasses device 3102 and / or the neckband 3105. In some embodiments, components of the glasses device 3102 and the neckband 3105 can be located on one or more additional peripheral devices paired with the glasses device 3102, the neckband 3105, or some combination thereof.

[0202] Pairing an external device, such as the neckband 3105, with the augmented reality glasses device enables the glasses device to achieve the form factor of a pair of glasses while still providing sufficient battery and computing power for extended capabilities. Some or all of the battery power, computing resources, and / or additional features of the augmented reality system 3100 can be provided by the paired device or shared between the paired device and the glasses device, thus generally reducing the weight, heat distribution, and form factor of the glasses device while still maintaining the desired functionality. For example, the neckband 3105 can allow components that would otherwise be included on the glasses device to be included in the neckband 3105 because users can tolerate a heavier weight load on their shoulders than they would on their heads. The neckband 3105 can also have a larger surface area on which to spread and dissipate heat into the surrounding environment. Thus, the neckband 3105 can allow for a larger battery and computing capacity than might otherwise be possible on a stand-alone glasses device. Since the weight borne in the neckband 3105 is less invasive to the user than the weight borne in the glasses device 3102, the user can tolerate wearing a lighter glasses device and can tolerate carrying or wearing the paired device for a longer period of time than the user would tolerate wearing a heavy stand-alone glasses device, enabling the user to more fully incorporate the artificial reality environment into their daily activities.

[0203] The neckband 3105 can be communicatively coupled to the glasses device 3102 and / or other devices. These other devices can provide certain functions (e.g., tracking, positioning, depth mapping, processing, storage, etc.) to the augmented reality system 3100. In Figure 31 an embodiment, the neckband 3105 can include two acoustic transducers (e.g., 3120(I) and 3120(J)), which are part of a microphone array (or potentially form their own microphone sub-array). The neckband 3105 can also include a controller 3125 and a power supply 3135.

[0204] The acoustic transducers 3120(I) and 3120(J) of the neckband 3105 can be configured to detect sound and convert the detected sound into an electronic format (analog or digital). In Figure 31 an embodiment, the acoustic transducers 3120(I) and 3120(J) can be positioned on the neckband 3105 to increase the distance between the neckband acoustic transducers 3120(I) and 3120(J) and other acoustic transducers 3120 located on the glasses device 3102. In some cases, increasing the distance between the transducers 3120 of the microphone array can improve the accuracy of beamforming performed using the microphone array. For example, if the acoustic transducers 3120(C) and 3120(D) detect sound, and the distance between the acoustic transducers 3120(C) and 3120(D) is greater than, for example, the distance between the acoustic transducers 3120(D) and 3120(E), the determined source location of the detected sound may be more accurate than the case where the sound is detected by the acoustic transducers 3120(D) and 3120(E).

[0205] The controller 3125 of the neckband 3105 can process information generated by sensors on the neckband 3105 and / or the augmented reality system 3100. For example, the controller 3125 can process information from the microphone array describing the sound detected by the microphone array. For each detected sound, the controller 3125 can perform a direction of arrival (DOA) estimation to estimate the direction in which the detected sound arrives at the microphone array. When the microphone array detects sound, the controller 3125 can populate an audio data set with this information. In an embodiment where the augmented reality system 3100 includes an inertial measurement unit, the controller 3125 can compute all inertial and spatial calculations from the IMU located on the glasses device 3102. Connectors can transfer information between the augmented reality system 3100 and the neckband 3105 and between the augmented reality system 3100 and the controller 3125. The information can be in the form of optical data, electrical data, wireless data, or any other transmittable data form. Moving the processing of information generated by the augmented reality system 3100 to the neckband 3105 can reduce the weight and heat in the glasses device 3102, making it more comfortable for the user.

[0206] The power source 3135 in the neckband 3105 can supply power to the glasses device 3102 and / or the neckband 3105. The power source 3135 can include, but is not limited to, a lithium-ion battery, a lithium polymer battery, a primary lithium battery, an alkaline battery, or any other form of electrical energy storage device. In some cases, the power source 3135 can be a wired power source. Including the power source 3135 on the neckband 3105 rather than on the glasses device 3102 can help better distribute the weight and heat generated by the power source 3135.

[0207] As mentioned, some artificial reality systems can substantially replace one or more of a user's sensory perceptions of the real world with a virtual experience, rather than mixing artificial reality with actual reality. An example of this type of system is a head-mounted display system (e.g., Figure 32 the virtual reality system 3200 in), which mainly or completely covers the user's field of view. The virtual reality system 3200 can include a front rigid body 3202 and a band 3204 shaped to surround the user's head. The virtual reality system 3200 can also include output audio transducers 3206(A) and 3206(B). Additionally, although not shown in Figure 32 , the front rigid body 3202 can include one or more electronic components, which include one or more electronic displays, one or more inertial measurement units (IMUs), one or more tracking transmitters or detectors, and / or any other suitable device or system for creating an artificial reality experience.

[0208] An artificial reality system can include various types of visual feedback mechanisms. For example, the display device in an augmented reality system 3100 and / or a virtual reality system 3200 can include one or more liquid crystal displays (LCDs), light-emitting diode (LED) displays, organic LED (OLED) displays, digital light projection (DLP) microdisplays, liquid crystal on silicon (LCoS) microdisplays, and / or any other suitable type of display screen. The artificial reality system can include a single display screen for both eyes, or can provide a display screen for each eye, which can provide additional flexibility for zoom adjustment or for correcting the user's refractive error. Some artificial reality systems can also include an optical subsystem having one or more lenses (e.g., conventional concave or convex lenses, Fresnel lenses, adjustable liquid lenses, etc.) through which the user can view the display screen. These optical subsystems can be used for various purposes, including collimation (e.g., making an object appear to be at a greater distance than its physical distance), magnification (e.g., making an object appear larger than its actual size), and / or transmitting light (e.g., transmitting light to the viewer's eyes). These optical subsystems can be used in non-pupil-forming architectures (e.g., a single-lens configuration that directly collimates light but can cause so-called pincushion distortion) and / or pupil-forming architectures (e.g., a multi-lens configuration that can produce barrel distortion to counteract pincushion distortion).

[0209] In addition to or instead of using a display screen, some artificial reality systems can include one or more projection systems. For example, the display device in an augmented reality system 3100 and / or a virtual reality system 3200 can include a micro-LED projector that projects light (using, for example, a waveguide) into the display device, such as a transparent combination lens that allows ambient light to pass through. The display device can refract the projected light towards the user's pupil and can enable the user to view artificial reality content and the real world simultaneously. The display device can achieve this using any one of a variety of different optical components, including waveguide components (e.g., holographic, planar, diffractive, polarization, and / or reflective waveguide elements), light manipulation surfaces and elements (e.g., diffractive, reflective, and refractive elements and gratings), coupling elements, etc. The artificial reality system can also be configured with any other suitable type or form of image projection system (e.g., a retinal projector used in a virtual retinal display).

[0210] An artificial reality system may also include various types of computer vision components and subsystems. For example, the augmented reality system 3100 and / or the virtual reality system 3200 may include one or more optical sensors, such as two-dimensional (2D) or 3D cameras, structured light emitters and detectors, time-of-flight depth sensors, single-beam or sweeping lidar sensors, 3D LiDAR sensors, and / or any other suitable type or form of optical sensor. The artificial reality system may process data from one or more of these sensors to identify the user's location, map the real world, provide the user with context about the surrounding real-world environment, and / or perform various other functions.

[0211] The artificial reality system may also include one or more input and / or output audio transducers. For example, elements 3206(A) and 3206(B) may include voice coil speakers, ribbon speakers, electrostatic speakers, piezoelectric speakers, bone conduction transducers, cartilage conduction transducers, tragus vibration transducers, and / or any other suitable type or form of audio transducer. Similarly, the input audio transducer may include a capacitive microphone, a dynamic microphone, a ribbon microphone, and / or any other type or form of input transducer. In some embodiments, a single transducer may be used for both audio input and audio output.

[0212] In some examples, the artificial reality system may include a haptic (i.e., tactile) feedback system, which may be incorporated into a headset, gloves, a bodysuit, a handheld controller, environmental devices (such as chairs, floor mats, etc.), and / or any other type of device or system. The haptic feedback system may provide various types of skin feedback, including vibration, force, traction, texture, and / or temperature. The haptic feedback system may also provide various types of kinesthetic feedback, such as motion and compliance. Motors, piezoelectric actuators, jet systems, and / or various other types of feedback mechanisms may be used to implement the haptic feedback. The haptic feedback system may be implemented independently of other artificial reality devices, within other artificial reality devices, and / or in combination with other artificial reality devices.

[0213] By providing tactile sensations, auditory content, and / or visual content, an artificial reality system can create a complete virtual experience or enhance a user's real-world experience in various contexts and environments. For example, an artificial reality system can assist or extend a user's perception, memory, or cognition within a particular environment. Some systems can enhance a user's interaction with others in the real world, or can enable a more immersive interaction of a user with others in a virtual world. Artificial reality systems can also be used for educational purposes (e.g., for teaching or training in schools, hospitals, government organizations, military organizations, commercial enterprises, etc.), entertainment purposes (e.g., for playing video games, listening to music, watching video content, etc.), and / or for accessibility purposes (e.g., as hearing aids, visual aids, etc.). The embodiments disclosed herein can implement or enhance a user's artificial reality experience in one or more of these contexts and environments and / or in other contexts and environments.

[0214] As detailed above, the computing devices and systems described and / or illustrated herein broadly represent any type or form of computing device or system capable of executing computer-readable instructions (such as those included within the modules described herein). In their most basic configuration, these computing devices can each include at least one memory device and at least one physical processor.

[0215] In some examples, the term "memory device" generally refers to any type or form of volatile or non-volatile storage device or medium capable of storing data and / or computer-readable instructions. In one example, a memory device can store, load, and / or maintain one or more of the modules described herein. Examples of memory devices include, but are not limited to, random access memory (RAM), read-only memory (ROM), flash memory, hard disk drive (HDD), solid state drive (SSD), optical disc drive, cache, variations or combinations of one or more of these components, or any other suitable storage memory.

[0216] In some examples, the term "physical processor" generally refers to any type or form of hardware-implemented processing unit capable of parsing and / or executing computer-readable instructions. In one example, a physical processor can access and / or modify one or more of the modules stored in the memory device described above. Examples of physical processors include, but are not limited to, microprocessors, microcontrollers, central processing units (CPUs), field programmable gate arrays (FPGAs) implementing soft-core processors, application specific integrated circuits (ASICs), portions of one or more of these components, variations or combinations of one or more of these components, or any other suitable physical processor.

[0217] Although shown as separate elements, the modules described and / or shown herein may represent portions of a single module or application. Additionally, in some embodiments, one or more of these modules may represent one or more software applications or programs that, when executed by a computing device, may cause the computing device to perform one or more tasks. For example, one or more of the modules described and / or shown herein may represent modules stored and configured to run on one or more of the computing devices or systems described and / or shown herein. One or more of these modules may also represent all or part of one or more special-purpose computers configured to perform one or more tasks.

[0218] Additionally, one or more of the modules described herein may transform data, physical devices, and / or representations of physical devices from one form to another. For example, one or more of the modules described herein may receive data to be transformed (e.g., data based on detected signals from a user, such as EMG data), transform the data, output the result of the transformation to perform a function (e.g., output control data, control an AR system, or other functions) or otherwise use the result of the transformation to perform a function, and store the result of the transformation to perform a function. Additionally or alternatively, one or more of the modules described herein may transform a processor, volatile memory, non-volatile memory, and / or any other part of a physical computing device from one form to another by executing on the computing device, storing data on the computing device, and / or otherwise interacting with the computing device.

[0219] In some embodiments, the term "computer-readable medium" generally refers to any form of device, carrier, or medium capable of storing or carrying computer-readable instructions. Examples of computer-readable media include, but are not limited to, transmission-type media (such as, a carrier wave) and non-transitory types of media such as magnetic storage media (e.g., hard disk drives, tape drives, and floppy disks), optical storage media (e.g., compact discs (CDs), digital video discs (DVDs), and BLU-RAY discs), electronic storage media (e.g., solid state drives and flash media), and other distribution systems.

[0220] The order of process parameters and steps described and / or shown herein is given only as an example and may vary as needed. For example, although the steps shown and / or described herein may be shown or discussed in a particular order, these steps do not necessarily need to be performed in the order shown or discussed. The various exemplary methods described and / or shown herein may also omit one or more steps described or shown herein, or include additional steps in addition to those disclosed.

[0221] The use of ordinal terms (e.g., "first", "second", "third") does not in itself imply any priority, precedence, or order of one claim element with respect to another or the temporal order in which acts of a method are performed. These terms may be used only as labels to distinguish one element from another having the same or similar name.

[0222] The foregoing description has been provided to enable others skilled in the art to best utilize various aspects of the exemplary embodiments disclosed herein. This exemplary description is not intended to be exhaustive or to be limited to any precise form disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments disclosed herein should be considered in all respects to be illustrative rather than restrictive. When determining the scope of the present disclosure, reference should be made to the appended claims and their equivalents.

[0223] Unless otherwise stated, terms such as "connected to" and "coupled to" (and their derivatives) as used in the specification and claims shall be construed to allow both direct and indirect (i.e., through other elements or components) connection. Further, the terms "a" or "an" as used in the specification and claims shall be construed to mean "at least one of". Finally, for convenience of use, the terms "including" and "having" (and their derivatives) as used in the specification and claims may be interchanged with the word "comprising" and have the same meaning as the word "comprising".

Claims

1. A system for gesture detection, comprising: A head-mounted device configured to present an artificial reality view to a user; A control device including a plurality of electromyography sensors, the electromyography sensors including electrodes that contact the skin of the user when the user wears the control device; At least one physical processor; And A physical memory including computer-executable instructions that, when executed by the physical processor, cause the physical processor to: Process one or more electromyography signals detected by the electromyography sensors to generate processed one or more electromyography signals, the electromyography signals being detected within a specific time window around an event; Classify the processed one or more electromyography signals into one or more gesture types using a user-specific classifier model, the user-specific classifier model having an associated accuracy level; Determine an optimal time window size by observing a change in the accuracy of the user-specific classifier model as the size of the time window changes; And Provide a control signal based on the classified one or more gesture types, wherein the control signal triggers the head-mounted device to modify at least one aspect of the artificial reality view, wherein: The user-specific classifier model is trained using training data including event markers determined by a multi-user classifier model; and The multi-user classifier model is trained using multi-user data obtained from multiple users.

2. The system according to claim 1, wherein The at least one physical processor is located within the control device.

3. The system according to claim 1, wherein, The at least one physical processor is located within the head-mounted device or within an external computer device in communication with the control device.

4. The system according to claim 1, wherein, The computer-executable instructions, when executed by the physical processor, cause the physical processor to classify the processed electromyography signals into one or more gesture types using a classifier model.

5. The system according to claim 4, wherein The classifier model is trained using training data including a plurality of electromyography training signals for the gesture types.

6. The system according to claim 5, wherein, The training data is obtained from multiple users.

7. The system according to claim 1, wherein, The head-mounted device includes a virtual reality headset or an augmented reality device.

8. A method for gesture detection, comprising: Obtaining one or more electromyography signals from a user; Processing the one or more electromyography signals to generate associated feature data, the electromyography signals being detected within a specific time window around an event; Classifying the associated feature data into one or more gesture types using a user-specific classifier model, the user-specific classifier model having an associated accuracy level; Determining an optimal time window size by observing a change in the accuracy of the user-specific classifier model as the size of the time window changes; and Providing a control signal to an artificial reality device based on the classified one or more gesture types Wherein, the user-specific classifier model is trained using training data including event labels determined by a multi-user classifier model; and the multi-user classifier model is trained using multi-user data obtained from multiple users.

9. The method according to claim 8, wherein, The classifier model is trained by clustering feature data determined from electromyogram training signals.

10. The method according to claim 8, wherein, The multiple users do not include the user.

11. The method according to claim 8, further comprising training the classifier model by the following steps: Obtaining electromyogram training signals corresponding to gesture types; Training the classifier model by clustering electromyogram training data obtained from the electromyogram training signals.

12. The method according to claim 8, wherein The classifier model is further trained by the following steps: Determining the time dependence of the electromyogram training signal with respect to the time of the maximum value of the corresponding electromyogram training signal; Aligning the time dependence of the multiple electromyogram training signals by adding a time offset to at least one of the multiple electromyogram training signals; Obtaining signal features from the aligned multiple electromyogram training signals; and Training the classifier model to detect electromyogram signals having the signal features.

13. The method according to claim 8, wherein, The classifier model is further trained by the following steps: Obtaining training data including electromyogram training signals corresponding to gesture types; and Averaging the electromyogram training signals corresponding to each occurrence of the gesture type to obtain a gesture model for the gesture type, wherein the classifier model uses the gesture model to classify electromyogram signals.

14. The method according to claim 13, wherein, The gesture model is a user-specific gesture model for the gesture type.

15. The method according to claim 13, wherein, The gesture model is a multi-user gesture model based on electromyogram training data obtained from multiple users, The multi-user gesture model is a combination of multiple user-specific gesture models.

16. The method according to claim 8, wherein, The artificial reality device includes a head-mounted device configured to present artificial reality images to a user, and the method further includes modifying the artificial reality images based on the control signal.

17. The method according to claim 16, wherein Modifying the artificial reality images includes selecting or controlling an object in the artificial reality images based on the gesture type.

18. A non-transitory computer-readable medium comprising one or more computer-executable instructions that, when executed by at least one processor of a computing device, cause the computing device to: Receive one or more electromyogram signals detected by an electromyogram sensor; Process the one or more electromyogram signals to generate processed one or more electromyogram signals, the electromyogram signals being detected within a specific time window around an event and identifying one or more features corresponding to one or more user gesture types; Use a user-specific classifier model to classify the one or more electromyogram signals into the one or more gesture types using the one or more features, the user-specific classifier model having an associated accuracy level; By observing the change in the accuracy of the user-specific classifier model as the size of the time window changes, change the size of the time window to determine the optimal time window size; and Provide a control signal based on the one or more gesture types classified; Transmit the control signal to the head-mounted device to trigger a modification of the artificial reality view in response to the control signal, wherein: The user-specific classifier model is trained using training data including event markers determined by a multi-user classifier model; and The multi-user classifier model is trained using multi-user data obtained from multiple users.

19. The non-transitory computer-readable medium according to claim 18, wherein, The computer device is configured to classify the electromyogram signal to identify the one or more gesture types based on a gesture model determined according to training data obtained from multiple users.

Citation Information

Patent Citations

  • Method and apparatus for a gesture controlled interface for wearable devices

    US20160313801A1

  • Prosthetic Virtual Reality Training Interface and Related Methods

    US20180301057A1