Multi-channel gesture recognition method and device
Through the multi-channel gesture recognition method, electromyography and pressure signals are collected, combined with deep learning technology, accurate recognition of complex gestures and gestures of different velocities is achieved, and the problem of difficulty in recognition in the existing technology is solved.
Patent Information
- Application Number
- CN202510291766.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-03-12
AI Technical Summary
Existing hand recognition methods are difficult to accurately identify complex gestures and gestures of different velocities, especially in methods based on human myoelectric signal.
The multi-channel gesture recognition method is used to collect the electromyography signal during hand movements through electrodes and the pressure sensor to collect real-time pressure signals between the fingertips and the grasping object. Then, signal preprocessing, feature extraction and model construction are carried out, and gestures and force recognition are used using a bidirectional LSTM network and a graph convolutional network.
It realizes accurate classification and recognition of complex gestures and gestures of different velocities, improving the accuracy and reliability of gesture recognition.
Smart Images

Figure CN119970010A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of rehabilitation, and in particular to a multi-channel gesture recognition method and device. Background Art
[0002] Stroke is a disease that causes damage to the neuromuscular pathways due to obstruction of blood supply to the brain. It can lead to functional disorders such as movement, sensation, and language. Hand function rehabilitation training is an important part of the patient's rehabilitation process.
[0003] Existing hand recognition methods are mainly divided into vision-based gesture recognition and human electromyography-based gesture recognition. Vision-based gesture recognition uses a camera to collect visual information of gestures, and users do not need to wear hardware devices. This non-contact information collection greatly improves the comfort of interaction. However, vision-based gesture recognition is easily affected by the lighting environment. The nature and intensity of the light source, the position of the light source and the camera have a direct impact on the imaging effect. In addition, the diversity of gestures, human gestures are three-dimensional and multi-degree-of-freedom, and it is difficult to accurately judge complex and changeable gestures on a two-dimensional projection plane.
[0004] Existing gesture recognition based on human electromyographic signals focuses on the decoding and motion recognition of electromyographic signals, but it is difficult to accurately classify complex gestures or gestures of different strengths, especially in the absence of strength recognition function. Summary of the invention
[0005] In order to overcome the deficiencies of the prior art, one of the objectives of the present invention is to provide a multi-channel gesture recognition method based on human electromyographic signals and capable of accurately classifying complex gestures and gestures of different strengths.
[0006] In order to overcome the shortcomings of the prior art, a second object of the present invention is to provide a multi-channel gesture recognition device based on human electromyographic signals and capable of accurately classifying complex gestures and gestures of different strengths.
[0007] One of the purposes of the present invention is achieved by the following technical solution:
[0008] A multi-channel gesture recognition method comprises the following steps:
[0009] Signal acquisition: The electrodes are set on the forearm to collect N-channel electromyographic signals during hand movements, where N is an integer greater than 1; the pressure sensor is set on the fingertips to collect real-time pressure signals between the fingertips and the grasped object;
[0010] Signal preprocessing: The dynamic cutoff frequency range of the filter is calculated through the real-time pressure signal, and the electromyographic signal of each channel is adaptively band-pass filtered according to the frequency range; the electromyographic signal of each channel after filtering is subjected to wavelet threshold denoising, and the electromyographic signal of each channel after denoising is subjected to multimodal normalization based on pressure value parameter compensation; the normalized N-channel electromyographic signals are integrated into a signal matrix:
[0011] Feature extraction: Extract the features of time domain, frequency domain, time-frequency domain and spatial correlation from the signal matrix in turn, and concatenate all the features of each time frame in chronological order to form a high-dimensional feature matrix;
[0012] Model building: Use a bidirectional LSTM network to learn the temporal evolution of gestures, build an anatomically constrained muscle connection map, and capture multi-channel spatial correlations through graph convolution; design channel-temporal joint attention weights to form a dynamic spatiotemporal graph convolution network model;
[0013] Result output: The high-dimensional feature matrix is input into the dynamic spatiotemporal graph convolutional network model, which performs gesture and force recognition to form a gesture classification probability distribution and a force classification probability distribution. When the probability is greater than or equal to the first threshold, the result is directly output; when the probability is greater than or equal to the second threshold and less than the first threshold, the sliding window weighted voting is triggered; when the probability is less than the second threshold, resampling is triggered.
[0014] Furthermore, in the signal preprocessing step, the lowest frequency in the dynamic cutoff frequency range is Maximum frequency Where f 0 is the reference center frequency, α is the dynamic adjustment coefficient, P(t) is the current pressure value, P max is the maximum pressure value.
[0015] Furthermore, in the signal preprocessing step, the wavelet threshold denoising is specifically performed by performing multi-layer decomposition based on the wavelet basis function and calculating the threshold of each layer of wavelet decomposition. tanh(0.5·SNR j ), where σ j is the noise standard deviation of the jth layer, N j is the coefficient length, SNR j It is the dynamic weighting factor of the signal-to-noise ratio. After wavelet decomposition, the high-frequency components of each layer are threshold processed.
[0016] Furthermore, the high-frequency components of each layer are thresholded as follows: the absolute value of the high-frequency component is less than λ j The absolute value of the high-frequency component is greater than or equal to λ j The signal is shrunk according to the threshold value and the valid signal is retained.
[0017] Furthermore, in the signal preprocessing step, the multimodal normalization based on pressure value parameter compensation is specifically: the standardized signal μ EMG is the mean value of the electromyographic signal, σ EMG is the standard deviation of the electromyographic signal, and β is the pressure compensation factor.
[0018] Furthermore, in the signal preprocessing step, the signal matrix is T is the length of the time series, N is the number of myoelectric signal channels, each row of the signal matrix is the preprocessed signal of N channels at each time point at that moment, and each column of the signal matrix is the complete time series of a channel.
[0019] Furthermore, in the feature extraction step, the time domain feature extraction is specifically to calculate the mean, variance, and zero-crossing rate of the signal matrix using a sliding window; the frequency domain feature extraction is specifically to extract the power spectrum density and divide the energy integral into multiple sub-bands; the time-frequency domain feature extraction is specifically to add the EMG signal spectrum slope compensation factor to improve the MFCC coefficient extraction; the spatial correlation feature extraction is specifically to calculate the mutual information entropy between channels In the formula, I(X i ,X j ) represents the mutual information entropy between the EMG signals of the i-th channel and the j-th channel, p(x i ,x j ) represents the joint probability distribution of the electromyographic signals of the i-th channel and the j-th channel.
[0020] Furthermore, in the model building step, the channel-time joint attention weight Where h t is the LSTM hidden state, x c is the feature vector of channel c, W h ,W x ,v is a trainable parameter, c′ and t′ represent the combination of other channels and time steps respectively.
[0021] Furthermore, the multi-channel gesture recognition method also includes an identity recognition step, which specifically includes: identifying the user identity through a user identity identifier or a biometric feature, and when the user is a new user, freezing the underlying network and fine-tuning the top-level classifier to achieve personalized adaptation.
[0022] The second object of the present invention is achieved by adopting the following technical solution:
[0023] A multi-channel gesture recognition device, used to implement any one of the multi-channel gesture recognition methods described above, the multi-channel gesture recognition device comprising
[0024] An electrode, wherein the electrode comprises a release film, a hydrogel, a medical tape, an insulating oil layer, a carbon paste circuit, a silver chloride circuit and a PET layer, wherein the silver chloride circuit is arranged on the PET layer, the carbon paste circuit, the insulating oil layer and the hydrogel are arranged on the silver chloride circuit, the insulating oil layer protects the carbon paste circuit and the silver chloride circuit, the hydrogel contacts the skin surface to conduct electricity, the medical tape is located on the insulating oil layer to provide mechanical support, and the electrode is arranged on the forearm to collect myoelectric signals during hand movements;
[0025] A pressure sensor, which is a thin-film pressure sensor used to collect real-time pressure signals between the fingertips and the grasped object;
[0026] A processor is used to identify a gesture and a force corresponding to the gesture according to the electromyographic signal and the pressure signal.
[0027] Compared with the prior art, the multi-channel gesture recognition method of the present invention collects N-channel electromyographic signals during hand movements and real-time pressure signals between the fingertips and the grasped object; calculates the dynamic cutoff frequency range of the filter through the real-time pressure signal, and performs adaptive bandpass filtering on the electromyographic signals of each channel according to the frequency range; performs wavelet threshold denoising on the electromyographic signals of each channel after filtering, and performs multimodal normalization based on pressure value parameter compensation on the electromyographic signals of each channel after denoising; integrates the normalized N-channel electromyographic signals into a signal matrix: extracts the features of time domain, frequency domain, time-frequency domain and spatial correlation from the signal matrix in turn, and splices all the features of each time frame in chronological order to form a high-dimensional feature matrix; adopts a bidirectional The LSTM network learns the temporal evolution of gesture movements, constructs an anatomically constrained muscle connection map, and captures multi-channel spatial correlation through graph convolution; designs channel-temporal joint attention weights to form a dynamic spatiotemporal graph convolution network model; inputs the high-dimensional feature matrix into the dynamic spatiotemporal graph convolution network model, and the dynamic spatiotemporal graph convolution network model performs gesture and strength recognition to form a gesture classification probability distribution and a strength classification probability distribution. When the probability is greater than or equal to the first threshold, the result is directly output; when the probability is greater than or equal to the second threshold and less than the first threshold, the sliding window weighted voting is triggered; when the probability is less than the second threshold, resampling is triggered. Through the above steps, complex gestures and gestures of different strengths can be accurately classified and recognized. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 is a flow chart of the multi-channel gesture recognition method of the present invention;
[0029] Figure 2 is a gesture classification diagram of the multi-channel gesture recognition method of the present invention;
[0030] Figure 3A classification diagram of different strengths of gestures in the multi-channel gesture recognition method of the present invention;
[0031] Figure 4 It is an electromyographic signal diagram of the multi-channel gesture recognition method of the present invention;
[0032] Figure 5 It is an algorithm framework diagram of the multi-channel gesture recognition method of the present invention;
[0033] Figure 6 A model architecture diagram of the multi-channel gesture recognition method of the present invention;
[0034] Figure 7 FIG. 4 is a diagram of the electrode structure of the multi-channel gesture recognition device of the present invention. DETAILED DESCRIPTION
[0035] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0036] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more related listed items.
[0037] See also Figure 1 as well as Figure 5 The present invention provides a multi-channel gesture recognition method, comprising the following steps:
[0038] Signal acquisition: The electrodes are set on the forearm to collect N-channel electromyographic signals during hand movements, where N is an integer greater than 1; the pressure sensor is set on the fingertips to collect real-time pressure signals between the fingertips and the grasped object;
[0039] Signal preprocessing: The dynamic cutoff frequency range of the filter is calculated through the real-time pressure signal, and the electromyographic signal of each channel is adaptively band-pass filtered according to the frequency range; the electromyographic signal of each channel after filtering is subjected to wavelet threshold denoising, and the electromyographic signal of each channel after denoising is subjected to multimodal normalization based on pressure value parameter compensation; the normalized N-channel electromyographic signals are integrated into a signal matrix:
[0040] Feature extraction: Extract the features of time domain, frequency domain, time-frequency domain and spatial correlation from the signal matrix in turn, and concatenate all the features of each time frame in chronological order to form a high-dimensional feature matrix;
[0041] Model building: Use a bidirectional LSTM network to learn the temporal evolution of gestures, build an anatomically constrained muscle connection map, and capture multi-channel spatial correlations through graph convolution; design channel-temporal joint attention weights to form a dynamic spatiotemporal graph convolution network model;
[0042] Result output: The high-dimensional feature matrix is input into the dynamic spatiotemporal graph convolutional network model, which performs gesture and force recognition to form a gesture classification probability distribution and a force classification probability distribution. When the probability is greater than or equal to the first threshold, the result is directly output; when the probability is greater than or equal to the second threshold and less than the first threshold, the sliding window weighted voting is triggered; when the probability is less than the second threshold, resampling is triggered.
[0043] Specifically, Figure 2 as well as Figure 3 As shown in the figure, gestures are divided into simple gestures and complex gestures. There are 10 simple gestures, namely: Index Flex (IF): index finger flexion, Index Ext (IE): index finger extension, Middle Flex (MF): middle finger flexion, Middle Ext (ME): middle finger extension, Ring Flex (RF): ring finger flexion, Ring Ext (RE): ring finger extension, Pinky Flex (PF): little finger flexion, Pinky Ext (PE): little finger extension, Thumb Flex (TF): thumb flexion, Thumb Ext (TE): thumb abduction. There are 10 kinds of complex gestures, namely: One (O): the gesture represents the number 1 (index finger extended, other fingers bent), Two (T): the gesture represents the number 2 (index finger and middle finger extended, other fingers bent), Three (TH): the gesture represents the number 3 (index finger, middle finger and ring finger extended, other fingers bent), Four (FO): the gesture represents the number 4 (all fingers except the thumb extended), Five (FI): the gesture represents the number 5 (palm flat, all fingers extended), Fist (FT): clenched fist, Flat (FL): flat palm, Thumb Up (TU): thumb up, Tip pinch (TP): pinch the index finger and thumb, Tripod pinch (TRP): triangle pinch, index finger, middle finger and thumb pinch.
[0044] Among them, each of the three complex gestures of index finger and thumb pinching, middle finger, index finger, thumb triangle pinching and fisting includes three different strengths. They are:
[0045] Pinch 1: pinch the index finger and thumb gently, pinch normally, and pinch hard (light, medium, and heavy strength).
[0046] Pinch 2: Gentle pinch, normal pinch, and strong pinch (light, medium, and heavy strength).
[0047] Fist: Light fist, normal fist, strong fist (light, medium, heavy strength).
[0048] The user's recovery level is determined by identifying the type and strength of the user's gestures.
[0049] The signal acquisition steps are as follows: the electrodes are arranged in a transverse "T" shape, which can accurately cover the main muscle groups of the human forearm and ensure the accuracy and stability of signal acquisition. The electrodes include multiple electrode points, which cover the main muscle areas of the forearm, including key muscle groups such as elbow flexion, pronation and supination, wrist flexion and finger flexion. The overall size of the electrode patch is 194mm x 126mm, and the electrode diameter is 9mm, which ensures that the electrode can be effectively attached to the forearm to avoid signal loss. In this embodiment, the number of electrode points is 8, so an 8-channel electromyographic signal is generated, such as Figure 4 The electrodes are connected to the data acquisition device, and the data are transmitted to the computer in real time with a sampling frequency of 500HZ.
[0050] The pressure sensor is a thin-film pressure sensor, which is set between the fingertips and the grasped object. The specific model is Likan Technology RP-C18.3-LT. The thin-film pressure sensor collects the pressure parameters between the fingertips and the grasped object, and reflects the pressure through the change of resistance value.
[0051] The signal preprocessing steps are as follows:
[0052] The filter cutoff frequency range is dynamically adjusted according to the real-time pressure value to suppress motion artifacts (such as sweat or skin sliding interference) caused by changes in muscle contraction strength and improve signal quality. The dynamic cutoff frequency range of the filter is adjusted in real time by the pressure value, and the lowest frequency Maximum frequency Where f 0 is the reference center frequency, α is the dynamic adjustment coefficient, P(t) is the current pressure value, P max is the maximum pressure value. In this embodiment, f 0 =250Hz, α=200Hz.
[0053] Wavelet threshold denoising is specifically as follows:
[0054] Based on the Daubechies-4 wavelet basis function, multi-layer decomposition is performed, and the improved SUREShrink threshold algorithm is used for the detail coefficient, and the noise suppression ratio reaches 32.7dB. Specifically: calculate the threshold of each layer of wavelet decomposition Where σ j is the noise standard deviation of the jth layer, N j is the coefficient length, SNR j is the dynamic weighting factor of the signal-to-noise ratio. After wavelet decomposition, the high-frequency components of each layer are threshold processed. The absolute value of the high-frequency component (detail coefficient) is less than λ j The absolute value of the high-frequency component (detail coefficient) is greater than or equal to λ j The signal is shrunk according to the threshold value and the valid signal is retained.
[0055] During the multimodal normalization process, pressure value parameter compensation is introduced to perform Z-score normalization. Specifically, the standardized signal μ EMG is the mean value of the electromyographic signal, σ EMG is the standard deviation of the electromyographic signal, and β is the pressure compensation factor, which is obtained through the real-time value of the pressure sensor and is used to suppress the signal amplitude fluctuation caused by pressure changes, ensure signal stability, and reduce motion interference. In this embodiment, β=0.05.
[0056] The signal matrix is T is the length of the time series, T = total time × 500. N is the number of EMG signal channels. Each row of the signal matrix is the preprocessed signal of N channels at each time point at that moment, that is, the preprocessed signal of 8 channels at that moment. Each column of the signal matrix corresponds to the complete time series of a channel.
[0057] The feature extraction step adopts a multi-domain joint feature extraction strategy, as shown in Table 1, which specifically includes:
[0058] Table 1
[0059]
[0060] Specifically: extracting time domain features is to calculate the mean, variance, and zero-crossing rate of the signal matrix using a sliding window; extracting frequency domain features is to extract power spectrum density and divide multiple sub-band energy integrals; extracting time-frequency domain features is to add a compensation factor for the tilt of the EMG signal spectrum to improve MFCC coefficient extraction; extracting spatial correlation features is to calculate the mutual information entropy between channels. In the formula, I(X i ,X j ) represents the mutual information entropy between the EMG signals of the i-th channel and the j-th channel, p(x i ,x j) represents the joint probability distribution of the electromyographic signals of the i-th channel and the j-th channel. The total dimension of the feature set is 316, which reduces 47% of redundant features compared with the traditional method. The muscle synergy is quantified by mutual information entropy to improve the separability of complex gestures.
[0061] All features of each time frame are concatenated in chronological order to form a high-dimensional feature matrix. Specifically, 316-dimensional features (24+160+104+28) are extracted for each time frame (such as a 200ms window) to form a T′×316 matrix.
[0062] Please continue reading Figure 6 In the model building step, a dynamic spatiotemporal graph convolutional network model (DST-GCN model) is built, and dual-branch spatiotemporal feature extraction is adopted. The dual-branch spatiotemporal branch is a temporal branch and a spatial branch. The temporal branch uses a bidirectional LSTM network (hidden layer 128 nodes) to learn the temporal evolution law of gesture movements. The spatial branch is to build an anatomically constrained muscle connection diagram (adjacency matrix d ij is the muscle spacing), and the multi-channel spatial correlation is captured by graph convolution. Design channel-time joint attention weights, channel-time joint attention weights Where h t is the LSTM hidden state, x c is the feature vector of channel c, W h ,W x ,v is a trainable parameter, c′ and t′ represent the combination of other channels and time steps respectively. In the step of building the model, it is necessary to collect training data and train the model. Specifically, when collecting training data, the subjects are between 18 and 40 years old, have no history of muscle or nervous system diseases, no history of skin allergies, good skin condition, suitable for wearing electrodes, have basic hand functions, and can complete gestures. Exclude the existence of upper limb motor dysfunction, recent upper limb trauma or surgery (in the past three months), and allergic reactions to electrode equipment. When collecting training data, the subjects rest for 3s, move for 5s, and rest for 5s. The number of repetitions: each gesture is performed 10 times, and 3 sets of data are collected repeatedly for each action group. Total time = (5 seconds action + 5 seconds rest) x 10 times x 3 groups. The experimental site is a quiet, well-lit laboratory environment. The subjects sit in a comfortable chair with their arms placed on the table to ensure stability.
[0063] The multi-channel gesture recognition method also includes an identity recognition step, which is implemented using an incremental learning mechanism. The incremental learning mechanism identifies new users through user identity identifiers (such as login ID) or biometric features (such as electromyography baseline patterns), and adopts the method of freezing the weights of the underlying network (DST-GCN) and only fine-tuning the top-level classifier to achieve personalized adaptation. 50 groups of samples from new users are used to update the classifier parameters through back propagation to achieve personalized adaptation. The whole process takes only 3 minutes, which is efficient and accurate.
[0064] The specific steps of the result output are: input the high-dimensional feature matrix into the dynamic spatiotemporal graph convolutional network model, the dynamic spatiotemporal graph convolutional network model performs potential and force recognition, and outputs: simple gesture classification probability distribution (10 types of gestures); Probability distribution of complex gesture classification (10 types of gestures); Probability distribution of strength classification (3 levels: light / medium / severe). When the output probability of the main classifier is greater than or equal to the first threshold, the result is output directly; when the probability is greater than or equal to the second threshold and less than the first threshold, the sliding window weighted voting (window length 500ms) is triggered; when the probability is less than the second threshold, resampling is triggered (new data is collected within 200ms). In this embodiment, the first threshold is 0.85 and the second threshold is 0.6. The results are output using a real-time visual interface, which specifically displays the electromyographic signal waveform, feature heat map, and classification confidence curve.
[0065] The present application also discloses a multi-channel gesture recognition device for implementing the above multi-channel gesture recognition method. The multi-channel gesture recognition device includes electrodes, pressure sensors and a processor. The electrodes collect electromyographic signals during hand movements, and the pressure sensors collect real-time pressure signals between the fingertips and the grasped objects; the processor recognizes gestures and the strength corresponding to the gestures based on the electromyographic signals and the pressure signals.
[0066] Please continue reading Figure 7 The electrode includes a release film, a hydrogel, a medical tape, an insulating oil layer, a carbon paste circuit, a silver chloride circuit and a PET layer.
[0067] The release film is the bottom layer, which is used to protect the adhesive part of the electrode to prevent contamination or adhesion before use. The release film is peeled off during use to expose the hydrogel or adhesive underneath.
[0068] The hydrogel is used to ensure that the electrode can make good contact with the skin surface and play a conductive role, transmitting the electromyographic signal to the sensor. The hydrogel is sticky, allowing the electrode to fit on the skin and maintain stable conductivity at the electrode-skin interface.
[0069] Medical tape provides mechanical support to ensure that the overall structure of the electrode can be firmly attached to the skin. Medical tape is usually waterproof and breathable to ensure comfort during long-term wear and prevent sweat from affecting electrode performance.
[0070] The insulating oil layer is used to protect the conductive material inside the electrode, prevent the external environment from interfering with the electrode performance, and ensure that the working area of the electrode can remain dry and stable.
[0071] Carbon paste circuit is one of the conductive layers of the electrode, and is used to conduct signals. Carbon paste has good conductivity, corrosion resistance and flexibility, and is a common conductive material in electrodes.
[0072] As one of the main conductive materials, silver chloride circuit can efficiently transmit electromyographic signals and has good signal stability and low noise characteristics. Silver chloride is a commonly used material for electromyographic electrodes and has good biocompatibility.
[0073] The PET layer is the topmost, transparent layer. The PET layer provides physical support for the electrode and maintains its flexibility. The PET material is transparent, flexible, and durable, ensuring that the electrode can maintain stable performance in different usage scenarios.
[0074] Motor manufacturing process First, the PET substrate is annealed to eliminate internal stress and provide a stable foundation for the conductive circuit. Specifically, the thickness of the PET substrate is 35-40μm, preferably 38μm. Next, a 6-12μm thick silver chloride layer is printed on the PET substrate as the main conductive material, and then a conductive carbon paste layer of the same thickness is printed to enhance conductivity and protect the silver paste circuit. Then apply a 10-20μm transparent insulating oil layer to protect the conductive part from the external environment. After that, double-sided tape and a blue reinforcement sheet with a thickness of 120-130μm are used to enhance the structural strength of the electrode to ensure its durability and stability during use. Preferably 125μm. A hydrogel layer with a diameter of 7-12mm is applied to the conductive area of the electrode to ensure good contact with the skin and signal conduction. Preferably 9mm. Finally, the electrode is cut to the designed size (194mm x 126mm) by a precision die-cutting process to ensure that it fits the user's forearm muscle area precisely. The release film is the bottom layer, which is used to protect the hydrogel and medical tape parts to prevent adhesion or contamination when the electrode is not in use. The entire production process focuses on the precise processing and combination of materials, ensuring the efficiency and reliability of the electrode in the acquisition of electromyographic signals.
[0075] Compared with the prior art, the multi-channel gesture recognition method of the present invention collects N-channel electromyographic signals during hand movements and real-time pressure signals between the fingertips and the grasped object; calculates the dynamic cutoff frequency range of the filter through the real-time pressure signal, and performs adaptive bandpass filtering on the electromyographic signals of each channel according to the frequency range; performs wavelet threshold denoising on the electromyographic signals of each channel after filtering, and performs multimodal normalization based on pressure value parameter compensation on the electromyographic signals of each channel after denoising; integrates the normalized N-channel electromyographic signals into a signal matrix: extracts the features of time domain, frequency domain, time-frequency domain and spatial correlation from the signal matrix in turn, and splices all the features of each time frame in chronological order to form a high-dimensional feature matrix; adopts a bidirectional The LSTM network learns the temporal evolution of gesture movements, constructs an anatomically constrained muscle connection map, and captures multi-channel spatial correlation through graph convolution; designs channel-temporal joint attention weights to form a dynamic spatiotemporal graph convolution network model; inputs the high-dimensional feature matrix into the dynamic spatiotemporal graph convolution network model, and the dynamic spatiotemporal graph convolution network model performs potential and strength recognition to form a gesture classification probability distribution and a strength classification probability distribution. When the probability is greater than or equal to the first threshold, the result is directly output; when the probability is greater than or equal to the second threshold and less than the first threshold, the sliding window weighted voting is triggered; when the probability is less than the second threshold, resampling is triggered. Through the above steps, complex gestures and gestures of different strengths can be accurately classified and recognized.
[0076] The above embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in this field, several modifications and improvements can be made without departing from the concept of the present invention, which are equivalent modifications and improvements made to the above embodiments based on the essential technology of the present invention, and all of them belong to the protection scope of the present invention.
Claims
1. A multi-channel gesture recognition method, characterized in that: The following steps are involved: Signal acquisition: The electrodes are set on the forearm to collect N-channel electromyographic signals during hand movements, where N is an integer greater than 1; the pressure sensor is set on the fingertips to collect real-time pressure signals between the fingertips and the grasped object; Signal preprocessing: Calculate the dynamic cutoff frequency range of the filter through the real-time pressure signal, and perform adaptive bandpass filtering on the electromyographic signal of each channel according to the frequency range; Perform wavelet threshold denoising on the electromyographic signals of each channel after filtering, perform multimodal normalization based on pressure value parameter compensation on the electromyographic signals of each channel after denoising; integrate the normalized N-channel electromyographic signals into a signal matrix; Feature extraction: Extract the features of time domain, frequency domain, time-frequency domain and spatial correlation from the signal matrix in turn, and concatenate all the features of each time frame in chronological order to form a high-dimensional feature matrix; Model building: Use a bidirectional LSTM network to learn the temporal evolution of gestures, build an anatomically constrained muscle connection map, and capture multi-channel spatial correlations through graph convolution; design channel-temporal joint attention weights to form a dynamic spatiotemporal graph convolution network model; Result output: The high-dimensional feature matrix is input into the dynamic spatiotemporal graph convolutional network model, which performs gesture and force recognition to form a gesture classification probability distribution and a force classification probability distribution. When the probability is greater than or equal to the first threshold, the result is directly output; when the probability is greater than or equal to the second threshold and less than the first threshold, the sliding window weighted voting is triggered; Resampling is triggered when the probability is less than a second threshold.
2. The multi-channel gesture recognition method according to claim 1, characterized in that: In the signal preprocessing step, the lowest frequency in the dynamic cutoff frequency range is Maximum frequency Where f0 is the reference center frequency, α is the dynamic adjustment coefficient, P(t) is the current pressure value, P max is the maximum pressure value.
3. The multi-channel gesture recognition method according to claim 1, characterized in that: In the signal preprocessing step, the wavelet threshold denoising is specifically: performing multi-layer decomposition based on wavelet basis functions, calculating the threshold of each layer of wavelet decomposition Where σ j is the noise standard deviation of the jth layer, N j is the coefficient length, SNR j It is the dynamic weighting factor of the signal-to-noise ratio. After wavelet decomposition, the high-frequency components of each layer are threshold processed.
4. The multi-channel gesture recognition method according to claim 3, characterized in that: The threshold processing of the high-frequency components of each layer is as follows: the absolute value of the high-frequency component is less than λ j The absolute value of the high-frequency component is greater than or equal to λ j The signal is shrunk according to the threshold value and the valid signal is retained.
5. The multi-channel gesture recognition method according to claim 1, characterized in that: In the signal preprocessing step, the multimodal normalization based on pressure value parameter compensation is specifically: the standardized signal μ EMG is the mean value of the electromyographic signal, σ EMG is the standard deviation of the electromyographic signal, and β is the pressure compensation factor.
6. The multi-channel gesture recognition method according to claim 1, characterized in that: In the signal preprocessing step, the signal matrix is T is the length of the time series, N is the number of myoelectric signal channels, each row of the signal matrix is the preprocessed signal of N channels at each time point at that moment, and each column of the signal matrix is the complete time series of a channel.
7. The multi-channel gesture recognition method according to claim 1, characterized in that: In the feature extraction step, the time domain feature extraction is specifically to calculate the mean, variance, and zero-crossing rate of the signal matrix using a sliding window; the frequency domain feature extraction is specifically to extract the power spectrum density and divide the energy integral into multiple sub-bands; the time-frequency domain feature extraction is specifically to add the EMG signal spectrum slope compensation factor to improve the MFCC coefficient extraction; the spatial correlation feature extraction is specifically to calculate the mutual information entropy between channels In the formula, I(X i ,X j ) represents the mutual information entropy between the EMG signals of the i-th channel and the j-th channel, p(x i ,x j ) represents the joint probability distribution of the electromyographic signals of the i-th channel and the j-th channel.
8. The multi-channel gesture recognition method according to claim 1, characterized in that: In the model building step, the channel-time joint attention weight Where h t is the LSTM hidden state, x c is the feature vector of channel c, W h ,W x ,v is a trainable parameter, c′ and t′ represent the combination of other channels and time steps respectively.
9. The multi-channel gesture recognition method according to claim 1, characterized in that: The multi-channel gesture recognition method also includes an identity recognition step, which specifically includes: identifying the user identity through a user identity identifier or a biometric feature, and when the user is a new user, freezing the bottom layer network and fine-tuning the top layer classifier to achieve personalized adaptation.
10. A multi-channel gesture recognition device, used to implement the multi-channel gesture recognition method according to any one of claims 1 to 9, characterized in that: The multi-channel gesture recognition device comprises An electrode, wherein the electrode comprises a release film, a hydrogel, a medical tape, an insulating oil layer, a carbon paste circuit, a silver chloride circuit and a PET layer, wherein the silver chloride circuit is arranged on the PET layer, the carbon paste circuit, the insulating oil layer and the hydrogel are arranged on the silver chloride circuit, the insulating oil layer protects the carbon paste circuit and the silver chloride circuit, the hydrogel contacts the skin surface to conduct electricity, the medical tape is located on the insulating oil layer to provide mechanical support, and the electrode is arranged on the forearm to collect myoelectric signals during hand movements; A pressure sensor, which is a thin-film pressure sensor used to collect real-time pressure signals between the fingertips and the grasped object; A processor is used to identify a gesture and a force corresponding to the gesture according to the electromyographic signal and the pressure signal.
Citation Information
Patent Citations
Gesture recognition method based on BP neural network
CN106293057A
On-body sensor system and method for automatic interpretation of visual body signals
US20230358848A1
Ai enabled multisensor connected telehealth system
US20250000361A1
Cited By
Head action recognition method and device, electronic equipment and storage medium
CN121542583A