System and method for recognizing user gestures or other movement or states

WO2025186543A8PCT designated stage Publication Date: 2025-10-02THE UNIV COURT OF THE UNIV OF EDINBURGH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/GB2025/050395
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-05
Filing Date
2025-02-27
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Myoelectric control systems face performance degradation due to EMG pattern variability over time, caused by factors like user behavior variation, electrode shift, muscle fatigue, and noise, leading to degraded model performance in long-term use.

Method used

A system and method that utilizes sensors to obtain muscle, nerve, and brain activity data, applies a classifier to generate pseudo-labels, and updates a gesture recognition model by balancing sample data in storage, using techniques like K-means clustering and feature reduction to improve model calibration.

Benefits of technology

Enhances the stability and accuracy of gesture recognition models by maintaining a balanced dataset, reducing variability, and adapting to user-specific changes, thereby improving long-term performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GB2025050395_02102025_PF_FP_ABST
    Figure GB2025050395_02102025_PF_FP_ABST
Patent Text Reader

Abstract

A system comprising: at least one sensor configured to obtain sensor data representing at least one of muscle, nerve and / or brain activity of a user; a processing resource configured to: obtain unlabelled sample data from said sensor data for the user for a plurality of samples; apply a classifier to the sample data and / or data derived from said sample data to obtain a pseudo-label for each sample, wherein the pseudo-label represents similarity of the sample to a selected one of a plurality of gestures, movements or states or to each of the plurality of gestures, movements or states; wherein the system further comprises storage for storing sample data for a set of samples together with the obtained pseudo-label, wherein the processing resource is further configured to: update the sample data stored in the storage based on the obtained pseudo-labels of the samples; calibrate a trained or partially trained gesture recognition model or other machine learning derived model for the user using at least the sample data of the updated sample data stored in the storage.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] System and method for recognizing user gestures or other movement or states

[0002] Field

[0003] The present disclosure relates to a system and method for recognising user gestures or other movement, and more particularly, the present disclosure relates to a system and method for recognising a gesture, for example, from a set of possible gestures by processing signals from sensors that detect muscular or other activity.

[0004] Background

[0005] The electrical activity of muscles, that is the myoelectric signals, are measured conventionally from the surface of the skin. Myoelectric control systems may translate electromyographic (EMG) signals to control commands of human-machine interfaces. Such systems enable users to interact with diverse devices, for example, exoskeleton and prosthesis in clinical applications or virtual keyboard and menu navigation in mixed reality and metaverse applications.

[0006] The variability of EMG patterns over time (both intra-day and inter-day) is a problem that may limit the practical application of myoelectric control systems. Diverse known or unknown factors, such as the behaviour variation of users, noises, electrode shift, muscle fatigue, limb position and other physiological factors jointly lead to the variability of EMG patterns. In real world applications, user motor learning / adaptation (even in open-loop myoelectric control) may also lead to a gradual change in EMG patterns. The EMG variability leads to substantially degraded performance of a model in long-term use.

[0007] The performance of myoelectric control models may degrade with varied nuisance factors between training and testing stages. For example, models pre-trained on data from other users may show a degraded performance when applied on a new user. Even the model applied to the same user may show degraded performance over long term usage (for example, over multiple days) due to nuisance factors such as electrode shift, different noise sources and varied EMG characteristics. It is known to train a model for learning nuisance-invariant EMG patterns. For example, US20230162023A1 proposes to train a neural network with nuisance factors disentangled.

[0008] Summary In accordance with a first aspect, there is provided a system comprising: at least one sensor configured to obtain sensor data representing at least one of muscle, nerve and / or brain activity of a user; a processing resource configured to: obtain unlabelled sample data from said sensor data for the user for a plurality of samples; apply a classifier to the sample data and / or data derived from said sample data to obtain a pseudo-label, or other classification label, for each sample, wherein the pseudo-label represents similarity of the sample to a selected one of a plurality of gestures, movements or states or to each of the plurality of gestures, movements or states; wherein the system further comprises storage for storing sample data for a set of samples together with the obtained pseudo-label, wherein the processing resource is further configured to: update the sample data stored in the storage based on the obtained pseudo-labels of the samples so that the storage stores a balanced and / or a more balanced set of samples; calibrating a trained or partially trained gesture recognition model or other machine learning derived model for the user using at least the sample data of the updated sample data stored in the storage.

[0009] The storage may comprise a data buffer, a data store, a storage device or a storage resource.

[0010] The updating of the sample data stored in the storage based on the obtained pseudolabels of the samples may provide a balanced and / or at least a more balanced set of samples stored in the storage.

[0011] Updating the storage may comprise selectively removing samples from the storage based on the pseudo-labels. Updating the storage may comprise counting samples based on their pseudo-labels or otherwise determining the occurrence of samples for each pseudo-label and selecting a sample for removal based on the occurrence of samples. Updating the storage may comprise selectively removing samples from the storage based on the pseudo-labels.

[0012] Updating the storage may comprise determining a pseudo-label having the highest occurrence in the set of stored samples and selectively removing one or more, for example, the oldest, of the oldest samples in the storage having the highest occurrence pseudo-label. Updating the storage may form a balanced and / or more balanced set of samples. Each pseudo-label may correspond to a classification. The set of samples may comprises a number of classes, wherein the set of samples are divided between the classes. Each pseudo-label and / or class may correspond to one of a plurality of gestures, movements or states.

[0013] The balanced set of samples may comprise a set of samples in which the count of samples for each class is within a range, for example, a range defined relative to an average count and / or other derived central number. The range may be based on the counts for each class and / or a value for an equally divided count. The balanced set of samples may comprise a set of samples in which the count of samples are within a predefined spread, for example, from each other and / or from an average count across classed.

[0014] The balanced set of samples may comprise a set of samples in which the count of sample for each class is within a pre-determined spread and / or variance. The balanced set of sample may comprise a set of samples in which the count of samples for each class is within a pre-determined count from the average number or other predetermined value. The balanced set of samples may comprises a set of samples in which the count for each class is within 20%, optionally 15%, optionally 10%, optionally 5% of the average count and / or from each other.

[0015] The balanced set of samples may comprise a set of samples having a measure of imbalance below a pre-determined threshold value.

[0016] The updated set of samples may be more balanced than the set of samples before the updating. The updated set of samples may comprise a set of samples in which a spread and / or variance of the count of samples for each class is smaller than the set of samples before the updating. The updated set of samples may comprise a set of samples in which a measure of imbalance is smaller than the set of samples before the updating.

[0017] The gesture recognition model may comprises a random forest model. The classifier may be trained using a clustering procedure. The sensor data may comprise electrical, magnetic, capacitive, and / or mechanical data, of muscle, nerve and / or brain activity of a plurality of subjects. The sensor data may correspond to a measurement corresponding to a known gesture, movement or state of a set of gestures, movements or states by one of the subjects.

[0018] The trained or partially trained gesture recognition model or other machine learning derived model may be configured to assign a class from a set of classes to a sample. The classifier may be further configured to assign a pseudo-label from the same set of classes.

[0019] The gesture recognition model or other machine learning derived model may be configured to output predicted labels of gesture, movement or other states. The class may be one of a pre-determined set of gestures. The clustering procedure may comprise a K-means clustering procedure.

[0020] The model may be configured to output a classification and / or a control signal based on gesture, movement or state recognition.

[0021] The sensor data may comprise electromyographic (EMG) data. Each sensor of the at least one sensor may be configured to obtain sensor data comprises a plurality of EMG data channels.

[0022] The processing resource may be further configured to apply a feature reduction process on the sample data to obtain lower dimensional data, wherein the further gesture recognition model is performed on the lower dimensional data

[0023] The classifier may be trained on lower dimensional data obtained from a feature reduction process performed on sample and / or sensor data.

[0024] The classifier may be configured to classify lower dimensional data, wherein the method comprises applying a feature reduction process to the sample data prior to the classification process.

[0025] The feature reduction process may substantially preserve the local distribution structure and / or other desired parameter of the higher dimensional data. The feature reduction process may comprise a manifold learning method, optionally at least one of t-distributed stochastic neighbour embedding t-SNE, uniform manifold approximation and projection (LIMAP).

[0026] The feature reduction process may comprise flattening a curved high-dimensional manifold on to a lower dimensional sub-space.

[0027] The updating and / or calibration may be performed periodically and / or in response to the storage being above a threshold value.

[0028] The calibration process may comprise an auto-calibration and / or self-calibration process.

[0029] The at least one sensor may be configured to perform measurements, for example, at least one of electrical, magnetic, capacitive, and / or mechanical measurements.

[0030] The sensor data may comprise unlabelled data collected during normal or routine use of the system.

[0031] The storage may comprise or form part of a random access memory (RAM). The storage may comprise part of a temporary storage resource.

[0032] Calibrating the model may comprises comprise updating and / or refining at least one trained model parameter or weight based on a balanced and / or more balanced set of samples

[0033] Each sample of the sample data or other data derived from the sample data may be represented in a latent space for gestures, movements or state and wherein the pseudolabels represent or are determined based on distance of a sample from a centroid or other position in a latent space for the gestures, movements or states. The centroid may be obtained from a clustering process, for example, a K-means clustering process.

[0034] The processing resource may be configured to obtain a representation of the sample, for example, a vector in a latent space, and determining a distance between the representation of the sample and respective representations, for example, vectors in a latent space, of each of the plurality of gestures, movements or states. The respective representations may correspond to centroid in the latent space. The processing resource may be configured to determine said centroids based on a plurality of further samples. The processing resource may be configured to assign a pseudo-label or other classification label based on the gesture, movement or state that is closest to the representation of the sample.

[0035] The obtained pseudo-label may represent the similarity of the sample to a selected one of a plurality of gestures, movements or states or to each of the plurality of gestures. The similarity between the sample and a gesture, movement or state may be determined or represented by a measure of distance between a representation, for example, a vector representation, of the sample and a corresponding representation of the gesture, movement or state. The pseudo-label may be assigned based on the gesture, movement or state having a representation at the smallest measure of distance from the representation of the sample. The pseudo-label may be assigned based on most similar gesture, movement or state or closest representation.

[0036] The at least one sensor may comprise a plurality of sensors, for example electrodes, located or configured to be located about a forearm or other limb.

[0037] Obtaining the sample data may comprise performing a sampling process and / or a feature extraction process on the sensor data. Obtaining the sample data may comprise applying a sliding window to the sensor data to obtain windowed samples of the sensor data.

[0038] Obtaining the sample data may comprise processing sensor data to determine values for a plurality of features as inputs to the gesture recognition model and / or classifier, wherein the features comprise: energy descriptors, distribution descriptors and spectrum descriptors.

[0039] The energy descriptors may comprise: absolute value (MAV), root mean square (RMS) waveform length (WL), slope sign changes (SSC),and zero crossings (ZC). The distribution descriptors may comprise skewness of the EMG signals. The spectrum descriptors may comprise mean frequency (MNF), median frequency (MDF), peak frequency (PKF), and variance of central frequency (VCF) The system in various embodiments may be applied, for example, in improving user movement, metaverse interfaces, robotics control, exoskeleton control, prosthetics, fatigue detection, monitoring of the body during exercise, rehabilitation and the like, and in training user movement for such activities.

[0040] The system may further comprises a memory resource for storing the trained gesture recognition model or its parameters.

[0041] The output of the gesture recognition model or other machine learning derived model may comprise a plurality of time-varying outputs, each time-varying output representing similarity to a respective one of the plurality of gestures or other movements or states. The gesture recognition model may be configured to receive sensor data and / or data derived from said sensor data and obtain an output representing or associated with a gesture, other movement or state.

[0042] Updating the sample data in the storage may comprise performing one or more read or write operations data stored in the storage. Updating the sample data in the storage may comprise at least deleting or overwriting sample data. Updating the sample data may comprises deleting sample data representing one or more sample with the highest occurrence in the storage and writing sample data representing one or more further sample. Updating the sample data may comprises overwriting sample data representing one or more samples with the highest occurrence in the storage with sample data representing one or more further samples.

[0043] In accordance with a second aspect, there is provided a method comprising: obtaining sensor data representing at least one of muscle, nerve and / or brain activity of a user; obtaining unlabelled sample data from said sensor data for the user for a plurality of samples; applying a classifier to the sample data and / or data derived from said sample data to obtain a pseudo-label for each sample, wherein the pseudo-label represents similarity of the sample to a selected one of a plurality of gestures, movements or states or to each of the plurality of gestures, movements or states; storing sample data for a set of samples together with the obtained pseudo-label, for example in storage, wherein the processing resource is further configured to: update the stored sample data based on the obtained pseudo-labels of the samples; calibrating a trained or partially trained gesture recognition model or other machine learning derived model for the user using at least the updated sample data stored in the storage.

[0044] The system may be for monitoring user gestures or other movement or states.

[0045] In accordance with a third aspect, there is provided a computer program product comprising computer-readable instructions that are executable to perform a method according to the second aspect or as described herein.

[0046] In accordance with a fourth aspect, there is provided a system comprising at least one sensor configured to obtain sensor data representing at least one of muscle, nerve and / or brain activity of a user; a processing resource configured to: obtain a pre-trained or at least partially trained model for gesture recognition or other machine learning derive model based on sensor data and / or sample data, wherein the model is trained using sensor data and / or sample data for a plurality users, wherein the model comprises a random forest model; perform a calibration process for calibrating the random forest model for a further user, wherein calibrating the random forest comprises: obtaining sample data from sensor data for a plurality of samples for the further user; applying a pruning and / or grafting process on one or more decision trees of the random forest using the sample data for the further user to obtain one or more further decision trees; appending the pre-trained random forest model with the one or more further decision trees.

[0047] The one or more decision trees may comprises a plurality of nodes. The nodes may comprise root nodes; decision nodes and leaf nodes. Root nodes may represent the highest node in the decision tree. Each decision nodes may represent a learnt rule, activity and / or decision having two or more outcomes. The decision nodes may be trained using training data. Decision nodes may have a single incoming edge and a plurality of outgoing edges. Leaf nodes may represent a node where further division is not possible and / or represent a final classification. Leaf nodes may have no outgoing edges. Each trained decision tree may be used to predict a classification representing a gesture, other movement or state, using new data.

[0048] The pruning and / or grafting process may comprise a bottom-up pruning and / or grafting strategy. The pruning and / or grafting process may be an iterative process performed at two or more nodes, optionally all each node of the one or more decision trees. The pruning and / or grafting process may be performed recursively at each node, starting at the plurality of leaf nodes.

[0049] The calibration process comprises applying a one-shot calibration process.

[0050] The calibration process may comprise evaluating performance of one or more decision trees before pruning and / or grafting and after pruning and / or grafting. The pruned and / or grated decision trees may be retained based on said evaluation. The performance may be evaluated using a performance

[0051] The appending of the pre-trained random forest model may comprise a part trained using data for a plurality of users and a part trained using data for the further user. The appended of the pre-trained random forest model may provide a random forest model that is calibrated and / or fine-tuned for the further user.

[0052] The calibration process may be performed at more than one time. The calibration process may be performed at periodic intervals.

[0053] The pruning process may comprises adding one or more further nodes to the one or more decision trees and retaining and / or discarding the further nodes based on the performance of the tree with the additional nodes. The performance may be determined using further validation data. The retaining and / or discarding may be based on a comparison of the performance of the tree before and / or after the adding of the further nodes.

[0054] The pruning process may comprises determining an error rate of the one or more decision trees with and without further nodes and replacing and / or removing the node based on the measure of relevance.

[0055] The measure of relevance and / or error rate may be obtained using a validation data set The pruning process may comprise comparing an error rate of a decision tree without a further node with the error rate of the decision tree with a further node and retaining the additional node if the error rate is increased and discarding the additional node if the error rate is decreased The grafting process may comprise identifying a new branch comprising one or more decision nodes that improves predictive performance of the tree and adding the new branch to the one or more decision trees of the random forest.

[0056] The grafting process comprises fine tuning a decision rule of a node by adding one or more further nodes to the node, optionally a further decision tree, and / or by changing the predicted label of the node to another predicted label.

[0057] The output of the gesture recognition model includes a plurality of time-varying outputs, each time-varying output representing similarity to a respective one of the plurality of gestures or other movements or states. The gesture recognition model may be configured to receive sensor data and / or data derived from said sensor data and obtain an output representing or associated with a gesture, other movement or state.

[0058] In accordance with a fifth aspect, there is provided a method comprising: obtaining sensor data representing at least one of muscle, nerve and / or brain activity of a user; obtaining a pre-trained or at least partially trained model for gesture recognition or other machine learning derive model based on sensor data and / or sample data, wherein the model is trained using sensor data and / or sample data for a plurality users, wherein the model comprises a random forest model; performing a calibration process for calibrating the random forest model for a further user, wherein calibrating the random forest comprises: obtaining sample data from sensor data for a plurality of samples for the further user; applying a pruning and / or grafting process on one or more decision trees of the random forest; training one or more further decision trees using the sample data for the user; appending the pruned random forest with the one or more further decision trees.

[0059] According to a sixth aspect, there is provided a computer program product comprising computer-readable instructions that are executable to perform a method according to the fourth aspect.

[0060] Features of one aspect may be provided as features of any other aspects. For example, features of the system may be provided as features of the method and vice versa. In addition, features of the first to third aspects may be provided as features of the fourth to sixth aspects.

[0061] Brief Description of Drawings

[0062] Various embodiments will now be described by way of example only, and with reference to the accompanying drawings, of which:

[0063] Figure 1 is a schematic diagram of an apparatus according to an embodiment;

[0064] Figure 2 is a schematic diagram showing placement of sensors on a subject and classification of hand gestures;

[0065] Figure 3 is a schematic diagram showing a training and self-calibration process, in accordance with an embodiment;

[0066] Figure 4 is a set of graphs that illustrate accuracy of the self-calibration model;

[0067] Figure 5 illustrates a decision tree pruning and / or grafting process, and Figures 6 to 8 depicts results obtained from experimental analyses;

[0068] Figures 9(a) and 9(b) represent tree pruning and grafting algorithms;

[0069] Figure 10 is a table of results from further experimental analysis; Figures 11 to 15 depict results from further experimental analysis.

[0070] Detailed Description

[0071] An apparatus or system 10 according to an embodiment is illustrated schematically in Figure 1. The system 10 comprises a computing apparatus 12, in this case a personal computer (PC) or workstation, which is connected to sensor(s) 14.

[0072] In the present embodiment, the sensor 14 is a surface electromyographic (sEMG) sensor, but in other embodiments, it can be any device that can detect the state of and / or changes in muscular tissue or muscle activity, or nerve activity, for example peripheral nerve activity, or brain activity, or spinal cord activity or any other physiological state or process of interest. The sensor may comprise a plurality of sensor devices. Figure 2 depicts a plurality of sensor positioned on a subject’s forearm, in accordance with an embodiment. The sensor(s) 14 is configured to obtain sensor data.

[0073] Some examples of sensor technologies that can be used to detect user movements or muscle activity or other activity state or process include electrical, magnetic, capacitive and mechanical sensors. The sensor(s) 14 is configured to generate data that is representative of at least one anatomical region of a patient or other subject. In some embodiments, there may be more than one sensor. In some embodiments, the sensors are used to recognize gestures made by the user and can monitor user movement.

[0074] The gestures may comprise muscular movement, position and / or relative position of musculature, optionally at least one of hand gestures and facial gestures. Any suitable muscles or sets of muscles may be the subject of the sensor measurements. While electromyographic signals, in particular surface electromyographic signals, are discussed, these can be replaced with capacitive signals, magnetomyograhic signals, mechanomyograms, ultrasound myographic signals, infrared myographic signals, optical myographic signals and the like in other embodiments. The above signals may, for example, be obtained from muscles, the spinal cord, nerves or the brain or a combination of these. The signals may include non-muscular signals measured from the anatomy. The sensor data typically represents data collected over a time interval and includes time varying data. The gesture may correspond to a hand grip.

[0075] In the present embodiment, data obtained by the sensor(s) 14 is provided to computing apparatus 12. Computing apparatus 12 comprises a processing resource 18 for processing of data. The processing resource 18 of Figure 1 comprises a central processing unit (CPU) and Graphical Processing Unit (GPU), and may further comprise a Tensor Processing Unit (TPU). The processing resource 18 provides for automatically or semi-automatically processing data from the sensor(s) 14. In other embodiments, the data to be processed may comprise any other suitable data obtained from measurements performed with respect to a patient or subject, which may be other than measurements based on muscular changes. The processing resource 18 is configured to receive measured signals represented by sensor data from the sensor(s) 14.

[0076] The processing resource 18 is further configured to provide a gesture recognition model or other machine learning derived model, for example, a random forest based model, configured to recognise or identify a gesture from data provided to it. In some embodiments, the model is trained and therefore configured to identify one or more of a gesture, a movement or other state of a user.

[0077] The processing resource 18 has sampling circuitry 21 configured to sample sensor data obtained by the sensor(s) 14. The sensor data is processed by the sampling circuity to obtain sample data for a plurality of samples. The apparatus 10 also has a data storage resource 16 for storing data, for example, a training data set 30 containing sensor data or data derived from sensor data, for example, the sample data. As described with reference to Figure 2, the training data comprises sample data representing a plurality of samples collected for a plurality of users.

[0078] The system 10 also has a data buffer 32 for storing data. The data buffer 32 is any suitable memory device or part of a suitable memory device for storage, for example, temporary storage of data. The data buffer may form part of a storage resource or device or data store. The data buffer 32 may form part of a RAM or hard drive of the system. The processing resource 18 is configured to perform read and write operations on the data stored on data buffer 32. In the following embodiments, the sensor data is processed by the sampling circuity to obtain sample data for a plurality of samples and the sample data is stored in the data buffer 32. In some embodiments, the sampling circuitry 21 and / or data buffer 32 are provided together with the sensor 14 as a sensing device. In some embodiments, the data buffer forms part of the data storage resource 16 or a further data storage resource or device.

[0079] The processing resource 18 includes feature selection circuitry 22 configured to extract values for a plurality of features for a plurality of samples from the sensor data either obtained by the sensor(s); model training circuitry 24 configured to train the gesture recognition model or other machine learning derived model, for example, a random forest based model; pseudo-labelling circuitry 26 configured to assign pseudo-labels to samples based on a classification process.

[0080] In the following embodiments, pseudo-labels are assigned to samples. In embodiments, the pseudo-labels may be examples of a type of classification label for classifying the samples as one or more of a plurality of gestures. The assignment of the pseudo-label or other classification label is based on a similarity of the sample to a selected one of a plurality of gestures, movements or states or to each of the plurality of gestures, movements or states. In further detail,

[0081] The model training circuitry 24 is configured to perform an initial training process for the gesture recognition model, based on initial training data 30 for a plurality of users and also to perform a further training or fine-tuning process of the model for a specific or target user, for example, based on further sensor data obtained for that user. In some embodiments, the model training circuitry comprises first model training circuitry for the initial training and further fine-tuning circuitry for the subsequent fine-tuning.

[0082] The processing resource 18 further comprises prediction circuitry for predicting a gesture, movement or other state from the sensor data using the trained model. In the present embodiment, the model is configured to receive new sensor data or data derived from sensor data and output a prediction of one or more gestures selected from a set of gestures. The model is trained such that the output of the model is representative or at least indicative of the one or more gestures that the new sensor data or the data derived therefrom represents. In some embodiments, the output is a label or labels corresponding to the gesture or gestures. In some embodiments, the output also includes one or more scores representing, for example, a likelihood or probability that the new sensor data corresponds to one or more, optionally all of the gestures.

[0083] In embodiments where the model is trained to recognise other movements or states, the output will accordingly be a prediction and / or associated score for said other movements or states, according to the training of the mode.

[0084] The system 10 also has a display screen 33, for example, for outputting performance data during use of a model and / or for presenting data to a user during model training. The system 10 also has an input device 34, for example, a keyboard or other device for receiving user input.

[0085] In the present embodiment, the circuitries 21 , 22, 24, 26, 28 are each implemented in the CPU and / or GPU and / or TPU by means of a computer program having computer- readable instructions that are executable to perform the method of the embodiment. In other embodiments, the circuitries may be implemented as one or more ASICs (application specific integrated circuits) or FPGAs (field programmable gate arrays). Any suitable processing resource may be used, not limited only to a CPU and / or GPU and / or TPU.

[0086] The computing apparatus 12 also includes a hard drive and other components of a PC including RAM, ROM, a data bus, an operating system including various device drivers, and hardware devices including a graphics card. Such components are not shown in Figure 1 for clarity.

[0087] As described in the following, the system of Figure 1 provides a machine learning model with inputs comprising the conventional features of multi-channel myoelectric signals. In the present embodiments, the machine learning model is a random-forest based model, however, it will be understood that other machine learning models may be used. The machine learning model described in the following is trained to recognise or identify gestures, other movements or states of a user.

[0088] Figure 2 (right) depicts an electrode configuration correspond to eight sensors (202a, 202b, 202c, 202d) arranged on forearm, in this case 2cm from below the elbow. Figure 2 (left) also depicts six possible gestures. These are labelled include power, lateral, tripod, pointer, open and rest. In the present embodiment, the sensors 14 are eight Trigno Quattro sensors (Delsys, USA) with a 2000 Hz sampling rate and 10-500 Hz passband. Eight channels of EMG signals were recorded with the sensors. Reference electrodes may be placed on the wrist or on other body parts, or may be integrated in the EMG sensor. Sensors were placed on the forearm, in this embodiment, 2cm below the elbow. Training data was obtained from a plurality of users using the sensors. Training data was obtained for each participant to perform the six possible gestures. For each gesture a number of repetitions were performed. In the present embodiment, for each hand gesture, 10 repetitions (6s each) were performed in 10 trials. Data recorded in the first 2 seconds reaction and transition period of each trial were removed, with the last 4s retained. A 5s inter-trial resting period was provided. While eight channels were recorded in the present embodiment, it will be understood that the method can be extended to any number of channels.

[0089] While the present embodiment describes using sensor data from eight sensors to detect six possible gestures, it will be understood that other numbers or types of sensors may be used. In addition, other numbers and types of gestures may be detected. In particular, the number of sensors used may be dependent on the number of gestures that are being recognised. In some embodiments, for example, a lower number of sensors may be used to detect a lower number of gestures. By way of a non-limiting example, sensor data from four sensors may be used to detect three gestures. Embodiments can cover the case where at least one sensor targets at least one muscle and decodes at least one gesture. In some embodiments, using more sensors may allow for more muscles to be targeted and more gestures may be decoded.

[0090] Any number and type (electric, magnetic, capacitive, mechanical etc) of sensors may be used in other embodiments. A sensor based on any other technology that senses the activity of muscles or nerves may also be used. At least one sensor was configured to output measured signals on at least one measurement channel. The at least one sensor may be configured to sense the activity of at least ten muscles of interest. In other embodiments, the at least one sensor may sense the activity of between 5 and 50 muscles of interest. In other embodiments, the at least one sensor may sense the activity of, for example, between 15 and 35 muscles of interest.

[0091] While the following embodiments describe detecting muscle activity of the arm or forearm, the same method can be used in sensor for any part of the body, for example, a neck, face or leg or any portion of the body where muscle activity is detectable.

[0092] In the following embodiments, collection of sensor data and sample data is described. The collection of data is described in the following. The sensor data is EMG data. Sample data for a plurality of samples can be obtained from sensor data as described in the following follows. A sampling and feature extraction process is performed on the raw EMG data to obtain values for a plurality of features for a plurality of samples. In the present embodiment, the raw EMG data is continuous, time varying data. In the present embodiment, eight sensors are attached to a user’s forearm and each sensor provides a channel of continuous sensor data, specifically EMG data. A sampling process is applied to the EMG data of each channel to obtain a plurality of samples. In the present embodiment, the sampling process is a sliding window applied to the EMG data of each channel to obtain windowed samples. For each windowed sample, a feature extraction process is performed to obtain a plurality of features. As a result of the sampling and feature extraction process, sample data for a plurality of samples are obtained, the sample data representing values for the features of interest. In the present embodiments, samples were extracted via a sliding window. In each window and each EMG channel, ten features were extracted to represent EMG signals from the three complementary aspects. The obtained sample data may be considered as belonging to a high dimensional vector space, specifically having dimensions of Number of EMG channels x Number of Features. In this embodiment, the high dimensional vector space is an 80 dimensional vector space. In the present embodiment, for the assignment of pseudo-labels, a dimensional reduction process is performed on the sample space to map each sample onto a lower dimensional space as part of a manifold learning process, for example, a t- SNE process. The lower dimensional space may alternatively be referred to as a latent space.

[0093] In the present embodiment, the feature extraction process includes obtaining values for features that are grouped as: energy descriptors, distribution descriptors and spectrum descriptors. In further detail, the energy descriptors comprise: absolute value (MAV), root mean square (RMS) waveform length (WL), slope sign changes (SSC),and zero crossings (ZC). The distribution descriptors comprise a measure of skewness of the EMG signals. The spectrum descriptors comprise mean frequency (MNF), median frequency (MDF), peak frequency (PKF), and variance of central frequency (VCF). In other embodiments, a subset of these features may be used. In other embodiments, alternative features may be extracted and used.

[0094] The apparatus 10 of Figure 1 is configured to perform a method that is now described with reference to Figure 3. Figure 3 is a flow-chart depicting a calibration method 300 for a machine learning model, specifically a gesture recognition model. The method 300 can be considered as composed of an initial training and calibration process for a user (including, for example, pre-training steps 302, 304 and a fine-tuning stage of steps 306 and 308). In addition, the method 300 includes a further self-calibration that is performed following the initial calibration stage.

[0095] At step 302 a pre-training dataset is obtained. The pre-training dataset is obtained during a pre-training stage in which each participant performs each of the six gestures and sensor data is recorded. The sensor data is EMG data. The sensor data is processed to obtain sample data for a plurality of samples for each user. Collectively, the plurality of samples for the plurality of users make up the pre-training dataset. For each hand gesture, a ground truth label is recorded and stored together with the sample such that the pre-training dataset comprises labelled samples. At step 304 a random forest model for each user is pre-trained using data collected from a plurality of other users. The training data is obtained for the plurality of users using the sensors as described above. The training data is used to partially train an RF model having 200 decisions trees. Considering the pre-training dataset is relatively large, 7% samples were randomly drawn via bootstrap to train each decision tree, to avoid extremely complex decision rules and encourage diversity among decision trees. Each feature from each subject was normalised separately. The partially trained model is referred to as a pre-trained random forest 304, the random forest comprising a plurality of pre-trained decision trees.

[0096] At step 306, a fine-tuning or initial calibration dataset is obtained. The collection of the initial calibration dataset corresponds to the sample collection process performed to obtain the training dataset but, in contrast to the pre-training dataset which represents sensed samples from other users, the initial calibration dataset is obtained by performing a sample collection process on the user. The result is a fine-tuning or initial calibration dataset that includes labelled sample data for a plurality of samples for the user. The samples of the calibration data are referred to as calibration samples.

[0097] At step 308, the pre-trained random forest is fine-tuned using the calibration dataset obtained at step 306. The fine tuning process of the pre-trained random forest comprises a pruning process followed by a further tree training and appending process. The pruning process uses a bottom-up pruning strategy is applied. Such a strategy includes performing a number of pruning operations. Each pruning operation involves pruning branches of the pre-trained decision trees using the calibration data. Specifically, a pruning operation involves providing sample data of a calibration sample to a decision tree of the pre-trained random forest model and pruning leaf nodes where the calibration data would fall according to the original decision nodes. The pruning operation continues to remove parent node until the root node is reached. The result is pre-trained tree with one or more branches pruned, referred to as a pruned tree. The calibration data with ground truth labels from the user, obtained in step 306, are used to validate the effect of each pruning operation. The pruning operation was performed subject to the condition that validation accuracy was improved. Further detail on the pruning process is provided with reference to Figure 5. Following the pruning process, a plurality of further decision trees are trained, from scratch, using only the calibration data from the user obtained at step 306. The new trained trees are appended to the pre-trained, pruned random forest model to form a fine-tuned random forest. The new trained trees are referred to as appended trees. In the present embodiment, the pre-trained and fine-tuned forest both consist of 400 decision trees. The fine-tuned random forest obtained at step 308 consists of 200 pretrained and pruned trees and 200 appended trees. As a result of the initial calibration stage, a random forest model calibrated for the user is obtained.

[0098] Following the initial calibration process for the user, the method includes further self or auto-calibration process, as described in the following. The self-calibration process uses data stored in a data buffer. In the following two data collection blocks are described, however, it will be understood that the data collection may be continuous performed, in some embodiments or at regular time intervals. In some embodiments, the further sample data for self-calibration is continually obtained as a user uses the system and the self-calibration steps are regularly performed.

[0099] At step 310, samples from the fine-tuning dataset are stored in the data buffer. As described in the following, the data buffer is used to save the latest testing samples. In the present embodiment, the data buffer has a size set to store 1500 windowed samples (which is about 500 kB using float32 precision). Initially calibration samples are stored in the data buffer.

[0100] Figure 3 depicts two testing blocks: first testing block 312 and second testing block 314. While two testing blocks are depicted in Figure 3, it will be understood that the method may include more than two testing blocks.

[0101] At the first testing block, a plurality of further samples are obtained for a user. The further samples are unlabelled. The further samples may be referred to as self-calibration samples. These samples are not initially labelled. These samples are collected, for example, during normal use of the system rather than through a separate calibration and ground truth data collection stage. As such, in contrast to the pre-training dataset that was collected using a data collection protocol, the collected samples can be considered to be an unbalanced dataset in that the dataset includes unequal numbers of samples. The result of the collection of the further sample is the unbalanced testing dataset (block 1) 316 in the data buffer.

[0102] Following collection of the unbalanced samples, a pseudo-labelling process is performed to assign pseudo-labels to the samples. To assign pseudo-labels to the high-dimensional EMG features, a dimensionality reduction process is first performed on the sample data. The dimensionality reduction process maps the high dimensional EMG data to a lower dimensional sub-space. The lower dimensional space may alternatively be referred to as a latent space. To self-calibrate a model, manifold learning and clustering were employed to assign pseudo-labels on buffered samples, which were then used to update the model parameters.

[0103] To better embed the complex distribution structure of EMG features, a dimensionality reduction enabling manifold learning was applied. In the present embodiment, a t- Distributed Stochastic Neighbor Embedding (t-SNE) is performed, as described in Maaten & Hinton, 2008. This embedding comprises mapping the original 80-dimensional (8 EMG channels * 10 types of features) EMG feature space into a lower dimensional subspace, in this example, a 3-dimensional subspace. The mapping is configured to preserve a local distribution structure of the original space. After dimensionality reduction via t SNE, the samples are represented by three dimensional data, referred to as reduced dimensional sample data. The initialised labels of all testing samples before clustering were assigned as the predictions directly given by the current model.

[0104] While t-SNE embedding is described, other suitable manifold learning methods, for example, LIMAP (Uniform Manifold Approximation and Projection) method may be used. An importation part of the manifold learning is the flattening of a high dimensional manifold (on which the high dimensional EMG data is distributed) to a low-dimensional subspace.

[0105] Each sample has an initial pseudo-label assigned to it correspond to the output of the random forest for that sample.

[0106] A clustering process is then performed on the samples in the reduced dimension space (the latent space). In the present embodiment, the clustering process is a K-means clustering process. The initialised labels of all testing samples before clustering were assigned as the predictions directly given by the current model. The pseudo-labels can be understood as approximate labels. Each sample is represented as, for example, a point or vector in the latent space.

[0107] The clustering method, in this embodiment, a K-means clustering method, consists of two key steps. The first step is to find the centroid location of all samples belonging to each class in the latent space. As a result of the first step, a first centroid location in the latent space for each class is obtained. The second step is to update the pseudo-labels of samples according to the class label of their nearest centroid in the latent space. As a result of the second step, each sample has an updated class label.

[0108] The first step of finding the centroid location of all samples in a class is then performed again using the updated class labels. As a result a revised centroid location in the latent space is obtained. In the present embodiment, the pseudo-labels represent or are determined from a distance between the sample in the latent space from the centroid or other position.

[0109] The two steps of obtaining an updated centroid location and updating class labels based on the updated centroid location are iterated until a convergence condition is met. Although the initial pseudo-labels can be given randomly, a more accurate assignment of initial pseudo-labels can speed up the convergence process and contribute to more accurate final pseudo-labels. Following convergence, the class of each label corresponds to the assigned final pseudo-label used to balance the buffer.

[0110] The clustering process allows a sample to be classified according to a measure of distance between the sample and the centroid locations. The pseudo-label is assigned based on the most similar class, in this embodiment, that is determined by the class having a centroid closest to the sample representation.

[0111] While centroids are describes, other reference points in latent space of each clusters / groups of samples may be suitable to determine the most similar class to a sample by a measure of distance.

[0112] It is noted that a point (e.g. user state) in the latent space, may be defined by a set of numbers in a vector. If the space is higher-dimensional then the point is defined by a larger number of points and hence the vector is longer. In a 3D space, each point is defined by three numbers x, y, and z. Through the pseudo-labelling procedure, each point is assigned a pseudo-label.

[0113] The pseudo-labels are assigned by integrating manifold learning via t-SNE and clustering via K-Means. EMG features are normally with a high dimensionality, distributed on a curved manifold in the feature space. The motivation to apply manifold learning is to simplify the data distribution in the high-dimensional feature space. Through manifold learning, the curved manifold can be flattened in a low-dimensional space. The motivation for clustering is to jointly consider (1) the knowledge learned by the current model (for the initialisation of pseudo-labels before clustering) and (2) the statistical information and distribution structure of a batch of saved data. Without these components (i.e. directly assigning pseudo-labels using the predictions of the current model), the model would be more likely to fall into a loop to learn biased knowledge given by itself. The ablation experiment also demonstrated the necessity of the manifold learning and the clustering step. As a result of the clustering process, the samples in the buffer are clustered into a number of distinct groups.

[0114] Following the clustering process, a sample balancing process is performed at stage 318. In the present embodiment, when the data buffer reaches its maximum size, the oldest sample corresponding to the gesture label (previously determined as pseudo-labels) with the most samples is deleted. Accordingly, samples in the data buffer tend to be balanced across different hand gestures.

[0115] The result of the reduction, clustering and buffer update is that the updated buffer is sufficiently balanced, at step 318, or at least balanced to a greater degree than prior to the updating. A balanced buffer may include a substantially equal number of samples belonging to each class, indicated by their assigned pseudo-labels. In some embodiments, the updated buffer is more balanced in that the number of samples in each class are closer together or are less spread apart in the updated buffer than prior to the buffer update. In some embodiments, the balancing of the number of pseudolabels may be represented as a distribution across gesture classes, and the balancing of the samples can be considered to flatten the distribution. For N classes, the more even (or closer the count of) the number of samples in different classes is, the more balanced the dataset is. Ideally there are 1 / N of all samples for each class. A number of suitable measures may be used to quantify a degree of balance of a dataset. For example, a GINI coefficient or entropy can be a good measure. For example, a higher GINI measure calculated on the number of samples in all classes, will correspond to a more balanced dataset.

[0116] In some embodiments, a balanced set of samples may comprise a set of samples in which the number (count) of samples for each class is within a pre-defined range or spread, for example, as defined from an average count or other central number. The permitted range may be defined relative to an equal divided number, for example, for N samples and Y classed, the equally divided number is Y / N and the permitted range may be defined relative to this value.

[0117] The balanced set of samples may comprise a set of samples in which the count of samples for each class is within a predetermined variance from an average value. The permitted spread or variance may comprise a percentage range, for example, 30%, optionally 20%, optionally 10%, further optionally 5% range of the average value or other centrally defined number.

[0118] It will be understood that the permitted range may be determined using a different reference or different techniques. Other methods of determining a balanced dataset may be used, for example, entropy based measurements of balance may be used.

[0119] In example embodiments, a balanced set of samples may comprise a set of samples in which a measure of imbalance is below a threshold. The measure of imbalance may be determined using, for example, the counts of samples in each class. The measure of imbalance may be determined by calculating the ratio of a count for a class relative to one or more other classes. A measure of imbalance may be defined as a ratio of a count of samples in a first class to samples in another class. For example, a 1 :1 ratio would correspond to two equal counts.

[0120] In some embodiments, the set of samples is more balanced after the updating process. A more balanced dataset may comprise a dataset in which the count of samples in each set has a smaller spread or variance and / or a measure of imbalance, for example, as described above, is smaller. The updated set of samples may be more balanced than the set of samples before the updating. The updated set of samples may comprise a set of samples in which a spread and / or variance of the count of samples for each class is smaller than the set of samples before the updating. The updated set of samples may comprise a set of samples in which a measure of imbalance is smaller than the set of samples before the updating. In some embodiments, updating the set of samples comprises identifying one or more classes having the lowest count(s) and increasing their number.

[0121] After pseudo-label assignment, both sample data in the data buffer with the pseudolabels (the balanced samples) and the calibration data 306 (previously used in the fine- tuning stage) with ground-truth labels are combined together to form a self-calibration data set. The self-calibration data set was used to train new decision trees to replace the original ones. It will be understood that it is the higher dimensional sample data that is combined with the calibration data for the refining / training, rather than the lower dimensional data used for pseudo-labelling. The pseudo-labels are used as labels for the refining / training step. By applying a pseudo-labelling and updating the buffer, challenges associated with an unbalanced set of samples may be addressed.

[0122] In the present embodiment, only the assignment of pseudo-labels is performed on the lower dimensional data. The applied t-SNE based dimensionality reduction and the clustering operations are applied to a batch of samples, which is only possible in the process of pseudo-label assignment. The training / re-calibration of random forest is performed on original high-dimensional data, because the model need to give predicted labels on each high-dimensional testing sample.

[0123] The self-calibration data is used to calibrated and / or retrain the model. In the present embodiment, the self-calibration data set is used to train new decision trees and replace the original ones to provide a self-calibrated random forest 320.

[0124] In further detail, it is understood that the pruned decision trees were pre-trained on data from other users which could provide different information with the data from the target user, the pruned decision trees are retained and the appended decision trees, as appended at stage 308, are replaced with the new trained decision trees. To prevent the model from forgetting previous knowledge, each time a random selected sub-set of the appended trees are replaced. In the present embodiment, 80 randomly selected appended decision trees are replaced. The choice of parameter was justified using an offline analysis and directly used in the real time experiment without any further parameter tuning. Following block 312, a first self-calibrated random forest model 320 is trained and ready for use to predict gestures for new data. The calibration of the machine learning model includes the step of refining or updating model parameters.

[0125] A further testing block 314 is performed. At step 324, further testing samples are collected in the storage buffer, substantially as described with reference to step 316. At step 326, a dimensional reduction and clustering process is performed to assign pseudolabels to the further testing samples in the buffer, followed by a further buffer balancing process, substantially as described with reference to step 318. The result is a further balanced storage buffer.

[0126] The further balanced storage buffer is then combined with the fine-tuning dataset, from step 306 and used to refine the random forest model, at step 328. The result of the further block 322, is a further self-calibrated random forest model.

[0127] It will be understood that Figure 3 depicts two self-calibration blocks, this process is repeated to allow the model to be refined and self-calibrated over a period of time. Figure 4 depicts results over 10 testing blocks comparing a self-calibrated model to a model that is not self-calibrated. Five self-calibration blocks are performed per day and measurements made of the model at each self-calibration block. As can be observed from Figure 4, self-calibration improves accuracy of the model. The testing data are unbalanced which can simulate a more realistic scenario, because in real world applications, the samples within a short period of time are not necessarily balanced. However, the samples in the data buffer tend to be balanced. Figure 4 depicts results for self-calibration 402 and results without self-calibration 404.

[0128] Figure 5 depicts a pruning process in further detail. Before the process of Figure 5, training data for a plurality of users was used for pre-training an RF model. The model had 200 decision trees. A subject-wise feature normalization was performed on the pretraining dataset. Specifically, the extracted features from each subject were normalized separately to a mean value of zero and a standard deviation of one to align the basic statistical metrics of feature distributions from different subjects.

[0129] Given a pre-trained RF model, the follow-up calibration procedure involved three strategies, namely, pruning, grafting, and appending decision trees, as presented in Figure 5. These steps correspond to steps 302, 304 306 and 308 of Figure 3.

[0130] Pruning is a known strategy to simplify a model and at the same time improve its generalizability. Grafting relates to growing new branches on a leaf node to further specify a decision rule. For both pruning and grafting, calibration data from a new subject was first collected and used to modify the pre-trained decision tree. The calibration data from the new user includes ground truth labels. The pre-trained decision tree is trained using training data for a plurality of other users. The pre-trained decision tree comprises a plurality of nodes including root nodes, decision nodes and leaf nodes and connections between nodes. Leaf nodes are understood as nodes that have no additional nodes stemming from them in that they do not split the data any further and provide a final classification. Figure 9(a) depicts a tree pruning algorithm in further detail.

[0131] For pruning a bottom-up strategy was adopted. Such a bottom up strategy includes determining the relevance of each individual node in a decision tree by recursively working upwards from leaf nodes. If a node is determined to be not relevant the node is either removed or replaced by a lead node.

[0132] In further detail, starting with a pre-trained decision tree (Tree) with a node u, at depth d where the depth of the root node is zero and the depth of a leaf node with the longest decision path is dmax. Using samples from a target new user (Starget) as the validation dataset, the error rate of a decision tree was estimated and denoted as Error (Tree, Starget). The Error may be classified as, for example, a number of misclassified samples given by a tree divided by the total number of all samples in Starget. When performing the pruning operation on a node, all its children nodes (and any children of children, if any) are removed from the tree. All samples falling into the removed children nodes now instead fall into their parent nodes, which are, following the pruning operation, the new leaf nodes. The prediction label of the new leaf node after decision tree pruning was defined as the mode of ground truth labels of all validation samples falling into this node. In the bottom- up strategy, parent nodes of the deepest leaf nodes was initially found (denoted as parentNodes = {ui, ...,Uj , ur,}) with their depth equal to dmax. If the estimated error rate decreased after pruning these deepest nodes, the pruning operation was performed. The same inspection procedure was repeated on nodes at lower depths until the root node was inspected. In the bottom-up strategy, parent nodes of the deepest leaf nodes are initially found (denoted as parentNodes = {ui, ...,Uj , .., ur,}) with their depth equal to dmax. If the estimated error rate decreased after pruning these deepest nodes, then the pruning operation is performed. The same inspection procedure was repeated on nodes at lower depths until the root node was inspected.

[0133] With reference to Figure 5, which is an illustrative example, the pre-trained decision tree represented by 500a has two first generation child nodes (504, 506) and a plurality of second generation child nodes wherein each second generation child node originates from a first generation child node (i.e. the parent of each second generation child node is a first generation child node). Figure 5 represents two second generation child nodes 510 and 508 having as their parent the first generation child node 504. The other second generation child nodes are not presented.

[0134] New calibration data for a user representing a plurality of labelled samples is provided (502a) to the pre-trained decision tree. A new sample of the data 502a initially falls onto the first, second generation node 510 and / or the second, second generation node 508. As part of the pruning operation, these nodes are removed so that the new sample data will now fall onto parent node 504. In this example, a new decision tree without nodes 510, 508 is formed that instead has parent node 504 as a new leaf node. The pre-trained tree and the new, pruned tree are subject to a comparison process, and as a result of the comparison, the pruning step is either maintained or discarded. In the present embodiment, the comparison process includes comparing the error rate of the pretrained decision tree to the error rate of the new decision tree. If the error rate decreases then the change is retained. If the error rate does not decrease then the change is discarded. This process is then repeated, recursively, for all nodes to find a pruned decision tree. It will be understood that alternative pruning operations may be used, in other embodiments. In addition, alternative comparison processes may be used. Turning to grafting, as depicted in Figure 5, a pre-trained decision tree is provided. The grafting process identifies any new branches that can be added to the decision tree to improve performance, for example, to improve predictive accuracy. Figure 9(b) depicts a tree grafting algorithm in further detail.

[0135] In further detail, given a set of leaf nodes, leafNodes = {ui, ...,Uj , .., ur,}, a set of validation samples from the new target subject falling into each leaf node denoted as Si = {Sj1, ..., Sj1, ..., SiK}. The set of involved ground truth labels of the sample set s, was defined as: C.i = {7‘, . . . , f‘}. ffi « JVdflss= 6.

[0136] The predicted label of a leaf node Uiwas defined as ljpredict. If m > 1 (that is, validation samples with more than one label fell into a leaf node Ui) or m=1 but i #= ivredlctthen the decision rule related to the leaf node u, is not reliable and therefore should be finetuned. Concretely, the decision rule related to a leaf node was fine-tuned either by grafting a new decision tree to the leaf node to further purify the sample labels (in the case of m > 1), or by directly flipping the predicted label (in the case of m = 1 but i #= j-predict |n tjS contextflipping the predicted label means that that the predicted label is changed to another value. Specifically, if only one calibration sample falls into a leaf node, and the ground truth label of this sample is different with the predicted label of the leaf node, then the predicted label is changed to the ground truth label.

[0137] With reference to Figure 5, which is an illustrative example, the pre-trained decision tree represented by 500b has two first generation child nodes (512, 514) and a plurality of second generation child nodes wherein each second generation child node originates from a first generation child node (i.e. the parent of each second generation child node is a first generation child node). Figure 5 represents two second generation child nodes 516 and 518 having as their parent the first generation child node 512. The other second generation child nodes are not depicted in Figure 5.

[0138] The new sample data 502b representing a plurality of samples are provided to the decision tree. If more than one sample of the new samples falls on a leaf and that leaf is determined to be unreliable, then a new branch (referred to as grafted branch) is grafted to the leaf. In Figure 5, the new branch includes new leaves 520 and 522. The new branch represents one or more further decision rule that correctly classifies the new sample. After performing either pruning or grafting on the pre-trained model, an additional 200 decision trees trained on data from the new target user are appended. This constructed an RF model with 400 trees. By appending new decision trees, the decision rules of the calibrated RF model consider both the generalized data distribution from a large number of pre-training users and the specific data distribution from the new user. Decision tree appending is simply planting a new decision trees in the forest model. These new decision trees are trained only on the calibration data from the new user, and are independent to the pre-trained decision trees.

[0139] As described above, a self-calibrating random forest (RF) common model, which can first, be pre-trained on other users and easily adapt to a new user via one-shot calibration, and second, keep self-calibrating itself once in a while based on pseudolabels of testing samples saved in a data buffer is descried. The pseudo-labels were assigned jointly by unsupervised manifold learning and clustering. The effectiveness of the model has been validated in both offline and real time experiments, in both openloop and closed-loop settings, and in intra-day, inter-day (2 consecutive days) and longterm (up to 5 weeks) applications. Instead of performance degradation, it was observed that the self-calibrating RF model gradually improves its performance and keep in a stable level in long-term applications. For the dynamic self-calibrating model, the effects of bidirectional user-model co-adaption in closed-loop experiments. With visual feedback, users can also adapt to the dynamic self-calibrating model and learn to perform the target hand gestures with significantly lower EMG amplitudes which may help a user to save effort.

[0140] The following embodiments relate to a system in which calibration data is gathered automatically for different gestures by clustering and pseudo-labelling measurements obtained during normal use by the participant. The selective deletion from the buffer provides a balanced set of samples for use in a self-calibration auto-calibration step which can, for example, be repeated periodically. As the model is first pre-trained using data from other users in the database then only a small amount of data is therefore needed for a new user, in some embodiments, only one second signals per gesture / class may be required. In the above embodiments, a standard random forest is trained and then calibrated the pre-trained random forest using a few unlabelled data from a new user to adapt the decision rules. The data buffer is used to allow the decision trees selfrecalibrate autonomously. The parameters of our model can be updated once in a while automatically. During the usage period, the parameters of our random forest can be updated automatically. The new idea of our method is not using random forest, but enabling a random forest to recalibrate itself automatically. The self-calibration may be performed in response to a trigger event, and may be, for example, periodic.

[0141] In contrast to methods intended to disentangle nuisance factors, the embodiments are build on a random forest, which is a more computationally efficient, lightweight, and explainable model compared with neural networks. A further difference, is that, the method may provide a flexible model that can be updated automatically during ongoing usage (instead of being pre-determined and fixed during applications). The model can therefore adapt to the time-varying nuisance factors, and even nuisance factors that are unknown at the initial training stage. In addition, EMG pattern recognition using simple, easily parallelizable, computationally efficient, and explainable RF, which is a practical alternative solution in mobile applications with limited computation resources. RF-based models may also achieve better across-day generalizability compared with efficient standard LDA. The proposed RF calibration method is source-free, which may also satisfying data privacy regulations in practice.

[0142] Each of the feature selection, model training, pseudo-labelling, can be implemented in any suitable manner in other embodiments, and are not limited only to the implementations discussed in relation to Figures 2, 3, 4 and 5.

[0143] The following non-limiting comments on experiments performed for are provided. In experiment 1 , 20 participants were recruited (aged 22-43 years, 12 males, 8 females). Eight electrodes were placed across the circumference of the forearm as shown in Figure 2). Delsys Trigno sensors with a 2000 Hz sampling rate and 10-500 Hz passband were used for data collection. Each participant then performed six hand gestures listed in Figure 1. For each hand gesture, 10 repetitions (6s each) were performed in 10 trials. Data recorded in the first 2s reaction and transition period of each trial were removed, with the last 4s retained. A 5s inter-trial resting period was provided.

[0144] In experiment 2, 18 new participants were recruited (aged 22-28 years, 11 males, 7 females). Data were collected on two successive days. Experiment on day 1 consisted of two sections, the calibration and the testing sections. In the calibration section, participants performed only one repetition per gesture in a 2s trial. In each trial, participant could react to the experiment instruction program and shape their hand in one second, and hold the gesture in the next second. Only signals during the latter period were retained (the same for all our following experiments) and used in our analyses. The testing section comprised 5 testing blocks. In each testing block, participants performed five repetitions per gesture (30 repetitions totally), with a pseudo-randomised order. Participants had 2 seconds and 5 minutes for inter-trial and inter-block rest, respectively. On day 2, all participants started the testing section (the same as day 1) directly without the calibration section. Datasets collected in experiments 1 and 2 were used in offline analyses.

[0145] Experiment 3 was a real time myoelectric control experiment. 28 new participants (aged 21-42 years, 13 males, 15 females) were recruited. Experiment 3 was also conducted on two successive days. On day 1 , the experiment consisted of a calibration section and a testing section, similar to experiment 2, except that, in each of the 5 testing blocks, data were unbalanced across different hand gestures. Specifically, each testing block did not necessarily involve exactly 5 repetitions per gesture. In this way, a more realistic application scenario was simulated where data in a short period of time may be unbalanced, which is expected to shift the overall data distribution. Data was kept in 5 successive blocks on each day balanced across different hand gestures, so that the model performance can be evaluated.

[0146] On day 1 , models were calibrated immediately after the calibration section, and run in real time in the testing section. Two models are run in the testing section, one with and the other without self-calibration (details will be described later). The 5 testing blocks on day 1 were conducted in an open-loop mode, that is, participants did not see the output of both models. Therefore, any differences between the two models is completely due to their decoding capacity.

[0147] On day 2, the testing section directly started without a calibration section. The testing section on day 2 consisted of two parts, 5 testing blocks each part. In part 1 , the experiment is exactly the same as the testing section on day 1. After part 1 , all 28 participants were divided into two groups (Group A and Group B). Participants in both groups are with a similar average accuracy in part 1 . After grouping each participant, the part 2 started. The experiment in part 2 for Group A was exactly the same as part 1 (open-loop mode), while the experiment for Group B was closed-loop with visual feedback. Specifically, participants in Group B were provided with the real time outcome of the self-calibrating model, together with the overall accuracy score of the model in each trial (shown at the end of a trial). 2 participants in Group B were randomly selected to take our long-term experiment. The 2 participants came back each week (up to 5 weeks) to repeat our closed-loop experiment (exactly the same as the part 2 on day 2).

[0148] Two baseline machine learning models were implemented for comparison. First, a userspecific LDA (linear discriminant analysis) model was trained from scratch using data from only the new target user. LDA was selected here because it is generally considered by myoelectric community as a gold standard for the state-of-the-art models. Second, a user-specific RF model (also with 400 decision trees) was trained from scratch using data from only the new target user. By comparing the performance of the pre-trained and fine-tuned RF model with the two baseline models, the superiority of the model in achieving a higher initial accuracy without self-calibration is demonstrated. By further comparing model performances with and without self-calibration, the effectiveness of the self-calibrating RF model is demonstrated.

[0149] Both offline and real time validations were performed. The offline validation was performed using the two datasets collected in experiments 1 and 2. All data from the 20 participants in dataset 1 were allocated in the pre-training dataset. For data from the 18 participants in dataset 2, each participant was in turn viewed as the target testing participant, with data from the other 17 participants used as additional pre-training data. Accordingly, for each of the 18 testing participant, data from 37 participants were available for model pre-training. Given the pre-trained RF model, data collected from the testing participant in the calibration section (1 repetition per gesture with Is holding duration) were used to fine-tune the RF model via pruning and appending decision trees. Once-pre-trained and fine-tuned, the model was duplicated into two copies, one with and the other without self-calibration. For the self-calibrating model, self-calibration was performed after each testing block. The baseline user-specific RF and LDA models were also trained (from scratch) using the same data collected from the calibration section. All these models will be compared in our offline analyses. Considering the pre-trained and fine-tuned RF model outperformed user-specific RF and LDA models in offline analyses in the real time experiment, only two models, i.e. the finetuned RF with and without self-calibration, were ran in an open-loop mode (experiments on day 1 and the part 1 on day 2). In experiments of the part 2 on day 2, participants were divided into two groups. For participants in Group B, visual feedback corresponding to the self-calibrating model was presented.

[0150] To explore the differences between Groups A and B (with and without visual feedback, respectively), the average classification accuracy and the EMG amplitude in our real time experiments were compared. To compare the effect of visual feedback on EMG amplitude, the EMG amplitude (quantified by RMS) of all participants in part 1 on day 2 (before grouping, all in open-loop mode) was normalized to the same baseline level. Then the average amplitudes of participants with and without feedback in experiments of the part 2 on day 2 was compared.

[0151] The factor of noises was simulated by injecting Gaussian noises to the collected EMG data. In real world applications using sparse electrodes (e.g. 8 electrodes in our work), one corrupted electrode due to strong noises would substantially degrade the performance of most models. The noises were injected to only testing data, because the collection of pre- training and fine-tuning dataset can be supervised to avoid EMG signals in extremely low quality. For the testing data collected on day 1 and day 2, one corrupted electrode was randomly selected (not necessarily the same on two days). Signals in the corrupted channel were injected additive white Gaussian noises with signal-to-noise ratio (SNR) of 20 dB, 15 dB, 10 dB and 5 dB.

[0152] Results in offline analyses are presented in Figure 6. According to Figure 6(a) and Figure 6(b), the accuracy of pruned and appended RF model achieved a higher accuracy compared with standard RF and LDA models on both days. With self-calibration, the accuracy of pruned and appended RF could be further improved. Specifically, as presented in Figure 6(a), the accuracy with self-calibration showed a slowly increased trend on each day, demonstrating that the self-calibrating model can progressively adapt to the data distribution in the testing process. Results in the ablation experiment was presented in Figure 6(c). With either clustering or t-SNE removed, the self-calibrating model achieved a relatively lower accuracy, demonstrating the necessity of each component in the whole framework. Figure 6(a) depicts results for pruned and appended with self-calibration 602, pruned and appended without self-calibration 604, standard RF 606 and standard LDA 608.

[0153] Figure 7 presents the accuracy variation with decreasing SNR. RF models showed a higher robustness against in-creasing levels of noises, compared with LDA. The pruned and appended RF model with self-calibration achieved the highest accuracy in all cases, although the improvement tended to decrease with stronger noises. Figure 7(a) depicts results for pruned and appended with self-calibration 702a, pruned and appended without self-calibration 704a, standard RF 706a and standard LDA 708a. Figure 7(b) depicts results for pruned and appended with self-calibration 702b, pruned and appended without self-calibration 704b, standard RF 706b and standard LDA 708b.

[0154] Results in the real time experiment are presented in Figure 8. According to Figure 8(a) and Figure 8(b), pruned and appended RF with self-calibration significantly outperformed the model without self-calibration on both day 1 and the part 1 of day 2. As shown in Figure 6(a), the accuracy of model without self-calibration progressively dropped from 84.3% (the first block on day 1) to 78.6% (the last block in part 1 on day 2). By contrast, the accuracy of self-calibrating model progressively increased on day 1 (from 84.3% to 86.2%), then dropped at the beginning of day 2 (from 86.2% to 82.8%), and finally slowly increased by the end of part 1 on day (from 82.8% to 86.5%). In the part 2 on day 2, the average accuracy of Group A 87.7%, without visual feedback) and Group B (88.9%, with visual feedback) did not show a significant difference. However, the EMG amplitude of participants in Group B with visual feedback was significantly lower than that of Group A without visual feedback, demonstrating that participants tended to learn the most effortsaving way to perform hand gestures with the help of feedback.

[0155] From Group B, two participants are randomly selected to take our 5-week long-term experiment. According to the curve in Figure 8(a), the accuracy can be maintained at a very high level. From week 1 (the part 2 on day 2) to week 5, the average accuracy was 91.9%, 95.3%, 95.7%, 96.1 % and 97.9%, even gradually increasing in a long-term use. In the last week, the extremely high accuracy of 97.9% demonstrated the exciting prospects of our model in real world applications.

[0156] In this work, an aim is to develop a plug-and-play myoelectric control system. First, a common RF model is pre-trained on data from other participants and then fine-tune the model using only 1 -repetition per gesture from the target testing participant. However, the large variability of EMG characteristics which may change over time degrade the performance of the model, due to extrinsic factors such as noises, electrode shift etc. Additionally, even without the extrinsic factors, the human motor system is inherently complex. Even when a myoelectric user repeat the same hand gesture, muscle activity can differ vastly, which is an inherent factor leading to degraded model performance. In addition, myoelectric users may gradually forget the details (e.g. the force level, the finger angles or shape of their hand) of each hand gesture they have performed in the calibration session. Overall, using data from only 1 -repetition per gesture is quite challenging to capture the general EMG characteristics of a new user. As presented in Figure 6, the accuracy of model without self-calibration gradually decreases even on the same day (blocks 1-5 on day 1), demonstrating the progressively increased level of distribution shift during the experiment. By contrast, in Figure 6, the accuracy of the selfcalibrating model shows an increasing trend on both day 1 and day 2. The self-calibrating model could adapt to the gradually shifted data distribution, and even learn a more generalised data distribution from the shifted data to further improve the model generalisability.

[0157] In our method, the pseudo-labels were assigned by integrating manifold learning via t- SNE and clustering via K-Means. EMG features are normally with a high dimensionality, distributed on a curved manifold in the feature space. The motivation to apply manifold learning is to simply the data distribution in the high-dimensional feature space. Through manifold learning, the curved manifold can be flattened in a low-dimensional space. The motivation to applying clustering is to jointly consider (1) the knowledge learned by the current model (for the initialisation of pseudo-labels before clustering) and (2) the statistical information and distribution structure of a batch of saved data. Without these components (i.e. directly assigning pseudo-labels using the predictions of the current model), the model would be more likely to fall into a loop to learn biased knowledge given by itself. The ablation experiment also demonstrated the necessity of the manifold learning and the clustering steps.

[0158] With simulated noises injected to an EMG channel, the model with self-calibration achieved the highest accuracy in all cases. However, it is noteworthy that with extremely strong noises, the performance improvement due to self-calibration tends to be less obvious. As discussed before, self-calibration can adapt the model to a slowly shifted data distribution within a short period of time. Although the data distribution may substantially shifted in long-term applications, the dynamic and progressively shifted data distribution during the application process serves as a bridge to connect the original and the final EMG patterns, so that the difficulty of each calibration could be reduced, contributing to a largely improved performance. Accordingly, it is foreseeable that there are extremes of the self-calibration mechanism. When the data distribution suddenly and vastly shifted, the benefits of the self-calibration mechanism would be reduced. That is why with extremely strong noises, the performance improvement due to self-calibration reduced.

[0159] The effect of feedback on user adaptation in closed-loop myoelectric control has been investigated in previous studies. Visual feedback has been proved effective in facilitating motor learning of users, so that users can learn useful myoelectric skills and adapt to a fixed model, contributing to improved performance compared with that in an open-loop mode. The mechanism of bidirectional user-model co-adaptation when users interacting with a dynamic self-calibrating model (rather than a fixed model). It is found that, compared with user adaptation, the model adaptation via self-calibration can make the greatest contribution from the aspect of improving accuracy. With self-calibration, the model performance was significantly improved even in an open-loop mode. Further proving visual feedback can only slightly improve the performance (see Group B in Figure 8). In addition to the aspect of accuracy, another interesting finding is that visual feedback can reduce the amplitude of users' EMG. With visual feedback, users can learn the most efficient manner to perform hand gestures and save efforts in long-term applications, which is another benefits of feedback for myoelectric control.

[0160] The system in various embodiments may be applied, for example, in improving user movement, metaverse interfaces, robotics control, exoskeleton control, prosthetics, fatigue detection, monitoring of the body during exercise, rehabilitation and the like, and in training user movement for such activities.

[0161] The following non-limiting comments on experiments performed for tree pruning and tree grafting are provided. Two experiments were performed, referred to as experiment B1 and B2. Experiment B1 was offline. EMG data was collected from 20 able-bodied subjects (12 males, 8 females, age range: 22-43). Eight electrodes were equally spaced across the circumference of the forearm. The EMG data were collected using Delsys Trigno sensors (Delsys, Inc.), with a sampling rate of 2000 Hz and bandpass filter of IQ- 500 Hz. Each subject performed six hand gestures, as presented in Figure 2. For each gesture, 10 repetitions were performed in 10 successive trials, each six seconds long. In each trial, subjects were required to mimic an instructed target hand gesture and then hold the gesture until the trial ended. Signals recorded in the first two seconds of each trial were removed to get rid of the transient period. The EMG signals within the last four seconds were retained. There was a five-second resting period between trials. Data from Experiment 1 was used to pre-train the RF model.

[0162] Experiment B2 was a real-time myoelectric control experiment. 18 new able-bodied subjects were recruited (11 males, 7 females, age range: 22-28). Experiment

[0163] 2 was conducted on two successive days. On the first day, the experiment consisted of two sections, namely, calibration and testing. In the calibration section, subjects were asked to perform each hand gesture only once in a 2s trial. Each two-second trial afforded the participants one second to shape their hand so that they could hold the gesture reliably during the second one second period. Only the data during the latter period was used for model tuning, or training a user-specific model from scratch (as a benchmark). The testing section comprised five blocks. In each block, five repetitions of each gesture were instructed in a pseudo-randomised way. Subject had two seconds and five minutes for rest between trials and between blocks, respectively. On the second day, subjects started the testing section without any further tuning. Electrodes were replaced on day 2 according to the electrode positions marked on day 1 .

[0164] In both of the tree pruning and tree grafting experiments, a window of 200 ms was used, sliding at 100 ms, for feature extraction. Ten features were extracted, which described the EMG data from three complementary aspects. First, the mean absolute value (MAV), waveform length (WL), zero crossings (ZC) and slope sign changes (SSC) of the EMG signals, combined with root mean square (RMS), were used as energy descriptors. Second, the skewness of the EMG signals was extracted as a distribution descriptor. Third, the mean frequency (MNF), median frequency (MDF), peak frequency (PKF), and variance of central frequency (VCF) were extracted as spectrum descriptors.

[0165] For the data collected from 20 subjects in experiment B1 , an offline leave-one-subject- out cross- validation was performed. For each subject, data from the other 19 subjects were allocated into the pre-training dataset to pre-train an RF model. For the 6 gestures x 10 trials / gesture = 60 trials from the testing subject, only data in the first trial of each gesture were used to calibrate the pre-trained model. To simulate the most challenging validation using minimal calibration data from the testing subject, only data within a 1s duration (randomly segmented out of the total 4s valid duration) of a trial were allocated into the calibration dataset to fine-tune the pre-trained RF. The choice of using only 1s duration of a trial in the calibration dataset is also the same as the counterpart in our subsequent online experiment (experiment B2) where only data within a 1s holding period were available in each calibration trial. For the other two baseline user-specific RF and LDA models, the same calibration dataset was used to train a brand-new userspecific model, and all other configurations were the same. Overall, in the offline validations, the performances of four models were compared, namely, RF calibrated via grafting and appending decision trees, RF calibrated via pruning and appending decision trees, standard RF, and standard LDA.

[0166] In experiment B2, an RF model was first pre-trained using all offline data from all 20 subjects of experiment 1 . When a new user started the experiment on the first day, data in the 1s holding duration of one trial for each gesture was collected in the calibration session, and used to calibrate the pretrained RF model. Standard user-specific RF and LDA models were also trained using the same calibration dataset (1 trial per gesture from the new user). In the testing part of experiment B2, the three models, namely RF calibrated via pruning and appending decision trees, the user-specific RF and the userspecific LDA, were implemented and gave their three predictions in each sliding window. Accuracy was evaluated on all predictions in 10 windows within the 1s holding period of a testing trial. On the second day of experiment 2, no model calibration was performed, and the testing section directly started using the same three models trained / calibrated on the first day.

[0167] A summary of data allocation strategies in both experiments for all models is presented in Table of Figure 10. Experiment B2 was conducted in an open-loop mode, that is the participants did not see the outcome of the decoders. As such any observed differences between the three decoders is wholly because of the decoding capacity of the model. The experiment was run on a laptop (CPU: 11th Gen Intel(R) Core(TM) i5-1145G7 @ 2.60GHz). An ablation experiment was performed. Considering many key components were included in the method, an ablation experiment with each component removed from the processing pipeline was performed, to show the necessity and contribution of each individual component. These key components are: (1) energy descriptors, (2) distribution descriptors, (3) spectrum descriptors, (4) subject-wise feature normalization, (5) pruned decision trees, and (6) appended decision trees. Note that when removing pruned decision trees or appended decision trees, the same number of total decision trees, namely 400 trees, was maintained by doubling the other types of decision trees.

[0168] To compare the performances of different models (> 3 groups), the Friedman test was first performed. The Nemenyi post-hoc test, a multi-comparison test to identify pair-wise group differences, was then applied. Significant differences were claimed if p < 0.05 was obtained.

[0169] Results of the experiment B1 are presented in Figure 11. RF with pruning and appending strategies and RF with grafting and appending strategies yielded an accuracy of 88.3% and 88.2%, respectively, both significantly higher than the outcome of LDA. By definition, pruning leads to smaller models with fewer parameters. Therefore, decision tree pruning rather than grafting was selected in the real-time implementation. Additionally, considering the final fine-tuned RF model involves many components, an offline ablation experiment was run to examine the effect of removing each individual component. Results were presented in Figure 12. Removing each individual component from the overall framework would lead to lower accuracy, indicating the necessity of all these components in the whole framework. Significance was observed when removing the energy features, pruned decision trees, or appended decision trees, demonstrating the important roles of these components.

[0170] Results of our real-time implementations of experiment B2 are presented in Figure 13. On the first day, RF calibrated with pruning and appending decision trees achieved the highest average accuracy of 81.5% but the accuracy improvement over the standard LDA (75.5%) and standard RF (77.6%) is not significant. On the second day, RF calibrated via pruning and appending decision trees significantly outperformed both the standard LDA and the standard RF. This results verifies the superiority of our method in an inter-day application. The final pre-trained and then fine-tuned RF model involves two groups of decision trees, i.e. , the pruned decision trees and the appended decision trees. For the pruned decision trees, because they were pre-trained on a large dataset from a large number of subjects, each decision tree consists of ~ 925 nodes. For the appended decision trees trained on a small dataset from a new subject, each decision tree comprised ~ 15 nodes. Accordingly, the whole model consists of 925x200+15x200 = 188, 000 nodes. Four parameters were saved in each node, i.e., the index of the left child (2 bytes), the index of the right child (2 bytes), the index of the selected feature in a node (1 byte) and the threshold of the selected feature (2 bytes), with a size of 7 bytes each node. Therefore, the size of the whole model is 188,000 nodes x 7 bytes / node = 1.26 MB. Other memory usage in the whole processing is 190.1 ± 19.2 KB (peak memory).

[0171] The average accuracy of each individual decision tree and the distribution of their predictions was evaluated and are presented in Figures 14 and Figure 15, respectively. First, as shown in Figure 14, by pruning each pre-trained decision tree, the average accuracy of the pruned decision trees was significantly higher than the pre-trained decision trees, demonstrating that the decision tree pruning operation mainly improves the performance of RF from the above listed aspect 1. Second, as shown in Figure 15, the predictions of the pre-trained trees and pruned trees were similar so their predictions were distributed near each other. This is due to the fact that, the pruning operation was performed on top of each pre-trained tree, so the predictions of pruned trees are more accurate but also quite similar as pre-trained trees. However, the predictions of the appended decision trees were far from the other trees in Figure 15, demonstrating that the appended decision trees could provide different complementary information with the other types of decision trees. The appended trees improve the ambiguity between all decision trees, improving the overall performance. A high average accuracy and a high ambiguity of decision trees together improve the overall performance of the RF model. Figure 15 depicts a visualization of predictions of different decision trees in a 2- dimensional space via t-Distributed Stochastic Neighbor Embedding (t-SNE). Predictions in this figure were drawn from a representative subject on the second day. Predictions on testing samples were used as visualization variables and all decision trees were used as visualization data points. Data from other subjects follow a similar pattern. A single data point represents a decision tree.

[0172] The effectiveness of the grafting and pruning operations on decision trees was also evaluated. Decision tree grafting could further specify the decision rules of a pre-trained decision tree based on the feature distribution of a new user; so that the new decision rules can be better adapted to the new user. By using data from a new user as the validation dataset, pruning could simplify the decision rules. Both grafting and pruning were effective and did not show significantly different performance. Decision tree pruning in the real-time implementation was selected because it led to smaller models.

[0173] Although description of particular embodiments has been provided above, it should be understood that these embodiments are illustrative only and that the claims are not limited to those embodiments. Those skilled in the art will be able to make modifications and alternatives to the described embodiments which are contemplated as falling within the scope of the appended claims. Each feature disclosed or illustrated in the present specification may be incorporated in any embodiment, whether alone or in any appropriate combination with any other feature disclosed or illustrated herein. In particular, one of ordinary skill in the art will understand that one or more of the features of the embodiments of the present disclosure described above with reference to the drawings may produce effects or provide advantages when used in isolation from one or more of the other features of the embodiments of the present disclosure and that different combinations of the features are possible other than the specific combinations of the features of the embodiments of the present disclosure described above.

Claims

CLAIMS1 . A system comprising: at least one sensor configured to obtain sensor data representing at least one of muscle, nerve and / or brain activity of a user; a processing resource configured to: obtain unlabelled sample data from said sensor data for the user for a plurality of samples; apply a classifier to the sample data and / or data derived from said sample data to obtain a pseudo-label for each sample, wherein the pseudo-label represents similarity of the sample to a selected one of a plurality of gestures, movements or states or to each of the plurality of gestures, movements or states; wherein the system further comprises storage for storing sample data for a set of samples together with the obtained pseudo-label, wherein the processing resource is further configured to: update the sample data stored in the storage based on the obtained pseudo-labels of the samples; calibrate a trained or partially trained gesture recognition model or other machine learning derived model for the user using at least the sample data of the updated sample data stored in the storage.

2. The system of claim 1 , wherein the updating the sample data provides a balanced and / or at least a more balanced set of samples and / or wherein the updated sample data comprises a balanced and / or at least more balanced set of samples.

3. The system of any preceding claim, wherein updating the sample data stored in the storage comprises selectively removing samples from the storage based on the pseudo-labels4. The system of any preceding claim, wherein updating the storage comprises determining a pseudo-label having the highest occurrence in the set of samples stored in the storage and selectively removing one or more, for example, the oldest, of the oldest samples in the storage having the highest occurrence pseudo-label5. The system of any preceding claim, wherein the gesture recognition model comprises a random forest model and / or wherein the classifier is trained using a clustering procedure.

6. The system of any preceding claim, wherein the trained or partially trained gesture recognition model or other machine learning derived model is configured to assign a class from a set of classes to a sample and wherein the classifier is further configured to assign a pseudo-label from the same set of classes.

7. The system of any preceding claim, wherein the model is configured to output based on gesture, movement or state recognition, a classification and / or a control signal8. The system of any preceding claim, wherein the sensor data comprises electromyographic (EMG) data, optionally wherein each at least one sensor is configured to obtain sensor data comprises a plurality of EMG data channels.

9. The system of any preceding claim, wherein the processing resource is further configured to apply a feature reduction process on the sample data to obtain lower dimensional data, wherein the further gesture recognition model is performed on the lower dimensional data10. The system of any preceding claim, wherein the classifier is trained on lower dimensional data obtained from a feature reduction process performed on sample and / or sensor data.11 . The system of any preceding claim wherein the classifier is configured to classify lower dimensional data, wherein the method comprises applying a feature reduction process to the sample data prior to the classification process.

12. The system of claim 10 or 11 wherein the feature reduction process substantially preserves the local distribution structure and / or other desired parameter of the higher dimensional data and / or wherein the feature reduction process comprises a manifold learning method, optionally at least one of t-distributed stochastic neighbour embedding t-SNE, uniform manifold approximation and projection (LIMAP).

13. The system of any preceding claim, wherein the updating and / or calibrating is performed periodically and / or in response to the storage being above a threshold value.

14. The system of any preceding claim, wherein the at least one sensor is configured to perform measurements, for example, at least one of electrical, magnetic, capacitive, and / or mechanical measurements.

15. The system of any preceding claims, wherein the sensor data comprises unlabelled data collected during normal or routine use of the system.

16. The system of any preceding claim wherein the storage comprises or forms part of a RAM or temporary storage resource.

17. The system of any preceding claim, wherein calibrating the model comprises comprise updating and / or refining at least one trained model parameter or weight based on balanced set of samples18. The system of any preceding claim, wherein each sample of the sample data or other data derived from the sample data is represented in a latent space for gestures, movements or state and wherein the pseudo-labels represent or are determined based on distance of a sample from a centroid or other position in a latent space for the gestures, movements or states.

19. The system of any preceding claim, wherein the at least one sensor comprises a plurality of sensors, for example electrodes, located about a forearm or other limb20. The system of any preceding claim, wherein obtaining the sample data comprises performing a sampling process and / or a feature extraction process on the sensor data.21 . The system of any preceding claim, wherein obtaining the sample data comprises processing sensor data to determine values for a plurality of features as inputs to the gesture recognition model and / or classifier, wherein the features comprise: energy descriptors, distribution descriptors and spectrum descriptors.

22. The system of claim 21 , wherein at least one of:a) the energy descriptors comprise: absolute value (MAV), root mean square (RMS) waveform length (WL), slope sign changes (SSC),and zero crossings (ZC); b) the distribution descriptors comprise skewness of the EMG signals; c) the spectrum descriptors comprise mean frequency (MNF), median frequency (MDF), peak frequency (PKF), and variance of central frequency (VCF)23. A method comprising: obtaining sensor data representing at least one of muscle, nerve and / or brain activity of a user; obtaining unlabelled sample data from said sensor data for the user for a plurality of samples; applying a classifier to the sample data and / or data derived from said sample data to obtain a pseudo-label for each sample, wherein the pseudo-label represents similarity of the sample to a selected one of a plurality of gestures, movements or states or to each of the plurality of gestures, movements or states; storing sample data for a set of samples together with the obtained pseudo-label, for example in storage, wherein the processing resource is further configured to: updating the stored sample data based on the obtained pseudo-labels of the samples so that the storage stores a balanced and / or a more balanced set of samples; calibrating a trained or partially trained gesture recognition model or other machine learning derived model for the user using at least the sample data of the balanced set of samples stored in the storage.

24. A computer program product comprising computer-readable instructions that are executable to perform a method according to claim 23.