CNNAttention-BiLSTM parallel model-based hand joint angle continuous motion estimation method
By adopting the CNN_Attention-BiLSTM parallel model and multi-head attention mechanism in the continuous motion estimation model, the problems of poor generalization and slow training speed of existing models are solved, and the motion estimation effect with higher accuracy and faster inference time is achieved.
Patent Information
- Application Number
- CN202411939589.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-30
AI Technical Summary
The existing continuous motion estimation model has poor generalization, slow training speed and few compatible actions, resulting in low prediction accuracy.
The continuous motion estimation method of hand joint angle based on the CNN_Attention-BiLSTM parallel model is adopted. Through the parallel combination of convolutional neural network and recurrent neural network, a multi-head attention mechanism is added to extract and fuse multiple features to improve the accuracy and robustness of the model.
Improve the accuracy and efficiency of continuous motion estimation, and achieve prediction results with higher accuracy and faster inference time on ARM architecture devices.
Smart Images

Figure CN120071385A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of biological signal detection and pattern recognition, and in particular to a method for estimating continuous motion of hand joint angles based on a CNN_Attention-BiLSTM parallel model. Background Art
[0002] Electromyography (EMG) is used to record the electrical signals generated when muscles contract. EMG reflects the subject's movement intention and contains a large amount of temporal and spatial information. Due to its low acquisition cost and convenient acquisition, surface EMG is widely used in gesture recognition, robot control, clinical diagnosis, sports science and other fields. Surface EMG is an important direction, which provides a more natural method for human-computer collaboration. With the development of artificial intelligence technology, deep learning algorithms have been widely used in image processing, anomaly detection, text prediction, human-computer collaboration and other fields, greatly improving life production efficiency. At present, simple motion classification can no longer meet the requirements of human-computer interaction. Accurate continuous motion estimation plays a vital role in achieving more natural human-computer interaction. In recent years, deep learning technology has been increasingly used in the field of biomedical signal pattern recognition. Due to the excellent performance of this method in prediction accuracy, its future application scenarios are very broad.
[0003] There are many issues that affect the prediction accuracy of continuous motion estimation. In terms of signals, the original sEMG signal has a low signal-to-noise ratio and is unstable, and the amount of data for a single individual is small, which makes it difficult to fully train the deep learning model. In terms of models, existing continuous motion estimation models often only consider the spatial or temporal information of sEMG signals, and it is difficult to extract two features at the same time, resulting in low prediction accuracy. In addition, the recurrent neural network (RNN), which is the most widely used in this field, still has the problem of difficult model convergence. Summary of the invention
[0004] In view of the problems of poor generalization, slow training speed, and few compatible actions of existing models that affect prediction accuracy, this application proposes a method for estimating continuous motion of hand joint angles based on the CNN_Attention-BiLSTM parallel model, combining convolutional neural networks and recurrent neural networks in a parallel manner, and adding a multi-head attention mechanism to the hybrid network, making full use of the advantages of a single network, thereby improving the accuracy and robustness of the model while ensuring efficiency. The present invention extracts multiple features from the original sEMG signal and then fuses them. The proposed network cleverly combines convolutional neural networks, recurrent neural networks, and attention mechanisms, and can perform high-precision continuous motion estimation of hand joint angles.
[0005] The technical means adopted by the present invention are as follows:
[0006] A method for continuous motion estimation of hand joint angles based on a CNN_Attention-BiLSTM parallel model, comprising the following steps:
[0007] Obtain surface electromyogram signal data and real-time corresponding finger joint angle data during continuous motion of multiple finger joints;
[0008] Send the surface electromyogram signal data into four parallel feature extraction branches respectively, fuse the four extracted features, and make the sampling frequency of the fused features consistent with the sampling frequency of the finger joint angles; the four parallel feature extraction branches are respectively used to extract the root mean square, peak stress, vibration expectation value, and unbiased standard deviation of the surface electromyogram signal data;
[0009] Use the fused features and the corresponding finger joint angles as the input data and output data of the CNN_Attention-BiLSTM parallel model respectively, and train the CNN_Attention-BiLSTM parallel model;
[0010] Obtain the surface electromyogram signal data to be estimated, send the surface electromyogram signal data into four parallel feature extraction branches respectively, fuse the four extracted features to generate the fused features to be estimated, input the fused features to be estimated into the trained CNN_Attention-BiLSTM parallel model, and obtain the finger joint angle data output by the CNN_Attention-BiLSTM parallel model.
[0011] Furthermore, the CNN_Attention-BiLSTM parallel model includes a CNN_Attention branch and a BiLSTM branch, where:
[0012] The CNN_Attention branch includes a multi-scale convolution module and a multi-head attention module, and the calculation formula of the multi-scale convolution module is:
[0013]
[0014] Among them, the input signal is X, and the output of the multi-scale convolution module is X′, represents the convolution operation, C is the number of convolution kernels, D is the number of channels of the input signal, n i is the convolution kernel size, i = 1, 2, 3, and σ is the ELU activation function;
[0015] The calculation formula of the multi-head attention module is:
[0016] Q = W q x + b q
[0017] K = Wk x + b k
[0018] V = W v x + b v
[0019]
[0020] head i = Attention(Q i , K i , V i )
[0021] MultiHead = Concat(head 1 , …, head n )W + b
[0022] Where Q is the query vector, K is the key vector, V is the weight value vector, Attention(Q, K, V) is the attention operation, head i is the single - head attention operation, MultiHead is the multi - head attention operation, W q , W k and W v are in R C×C , b q , b k and b v are in R C , C is the channel dimension of the single - head attention mechanism layer, d k is the dimension of Q and K, W is in R D×D , b is in R D , Concat means concatenation by dimension, softmax is the activation function, n is the number of multi - head attention heads, and x is the input signal;
[0023] The calculation formula of the BiLSTM module is as follows:
[0024] f t = σ(W f ·[h t-1 , x t + b f )
[0025] i t = σ(W i ·[h t-1 , x t + b i )
[0026]
[0027] o t= σ(W o [h t-1 , x t + b o )
[0028] h t = o t * tanh(C t )
[0029] Among them, i t is the input gate, W i is its weight matrix, b i is its bias term; f t is the forget gate, W f is its weight matrix, b f is its bias term; o t is the output gate, W o is its weight matrix, b o is its bias term; C t-1 is the cell state at the previous moment, C t is the cell state at the current moment, is the candidate cell state, W C is the weight matrix, b C is the bias term; x t is the input at the current moment, h t-1 is the hidden state at the previous moment, h t is the hidden state at the current moment; σ is the sigmoid function, and tanh is the hyperbolic tangent function;
[0030] Then, it passes through a max-pooling layer to obtain the output of the BiLSTM module.
[0031] Furthermore, finally, the outputs of the CNN_Attention branch and the BiLSTM branch are concatenated according to the channel dimension, and the concatenated features are mapped to 10 finger joint angles through two fully connected layers.
[0032] Furthermore, obtaining the surface electromyogram signal data during continuous movement of multiple finger joints further includes performing noise reduction preprocessing on the surface electromyogram signal data based on a Butterworth filter.
[0033] Furthermore, fusing the four extracted features includes:
[0034] Step 1: Segment several 100 - ms sliding windows from the surface electromyogram signal and the resampled hand movement joint angle signal, with a step size of 0.5 ms;
[0035] Step 2: Extract four features from the segmented data intervals, namely root mean square error (RMS), peak stress (PS), expected value of amplitude change rate (SE), and unbiased standard deviation (USTD). The mathematical expressions of the four features are as follows;
[0036]
[0037] Among them, N represents the window size, and x i represents the i-th sEMG sample in the window, and m 0 represents the intensity of muscle contraction, and m 2 describes the change of surface electromyogram signal, and m 4 describes the change of m 2 , and Δ 2 represents the second derivative, and x represents the expected value of a single sliding window of electromyogram samples;
[0038] Step 3: Finally, fuse them in a dimension-wise concatenation manner. Since the number of signal channels after extracting a single feature is 12, the number of signal channels after fusing the four features is 48.
[0039] Furthermore, before training the CNN_Attention-BiLSTM parallel model, it also includes normalizing the fused features based on the following formula:
[0040]
[0041] Among them is the fused feature after normalization, x max is the maximum value of the fused feature, and x min is the minimum value of the fused feature.
[0042] Furthermore, the method also includes:
[0043] Deploy the trained CNN_Attention-BiLSTM parallel model to an ARM device for offline testing.
[0044] Compared with the prior art, the present invention has the following advantages:
[0045] The present invention extracts and fuses four features from the original electromyogram signal, and then uses the CNN_Attention-BiLSTM parallel network model for continuous estimation of hand joint angles. The model consists of a multi-scale convolution module and a multi-head attention mechanism serial connection module, and is then composed by parallel connection with the BiLSTM module. While improving the continuous estimation accuracy, it ensures efficiency, thereby obtaining a prediction result with higher accuracy and faster inference time on a device with an ARM architecture. Description of the Drawings
[0046] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0047] Figure 1 It is a flowchart of a method for continuous motion estimation of hand joint angles based on a CNN_Attention-BiLSTM parallel model of the present invention.
[0048] Figure 2 It is the execution process of the motion estimation method based on the method of the present invention in the embodiments of the present invention.
[0049] Figure 3 It is the training process of the CNN_Attention-BiLSTM parallel model in the embodiments of the present invention.
[0050] Figure 4 It is the prediction result diagram of continuous motion estimation of hand joint angles in the embodiments of the present invention. Detailed implementation manners
[0051] In order to enable those skilled in the art to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0052] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0053] Such as Figure 1As shown in the figure, the present invention provides a method for continuously estimating the hand joint angle based on a CNN_Attention-BiLSTM parallel model, which mainly includes the following steps:
[0054] S1. Obtain the surface electromyogram signal data during continuous movement of multiple finger joints and the corresponding finger joint angle data in real time.
[0055] Specifically, use an electromyogram device to collect the surface electromyogram signal data during continuous movement of multiple hand joints, use a data glove to collect the real-time joint angle data information of the hand, synchronize the joint angle data and the surface electromyogram signal data in real time, and align the data. Preprocess the collected surface electromyogram signal, and use a Butterworth filter to perform noise reduction processing on the electromyogram signal. The calculation expression is:
[0056]
[0057] where n is the filter order, ω c is the cut-off frequency, and ω represents the electromyogram signal frequency.
[0058] S2. Send the surface electromyogram signal data into four parallel feature extraction branches respectively, fuse the four extracted features, and make the sampling frequency of the fused features consistent with the sampling frequency of the finger joint angle. The four parallel feature extraction branches are respectively used to extract the root mean square, peak stress, vibration expectation value, and unbiased standard deviation of the surface electromyogram signal data.
[0059] Furthermore, perform normalization processing on the surface electromyogram signal data after feature fusion to improve the convergence speed and prediction accuracy of the model:
[0060]
[0061] where is the value of the surface electromyogram signal data after normalization, x max is the maximum value of the surface electromyogram signal data, and x min is the minimum value of the surface electromyogram signal data.
[0062] Furthermore, fuse the four extracted features, including:
[0063] Step 1: Segment several 100-ms sliding windows from the surface electromyogram signal and the resampled hand movement joint angle signal, with a step size of 0.5 ms;
[0064] Step 2: Extract 4 features from the segmented data interval, namely root mean square error RMS, peak stress PS, expectation of amplitude change speed SE, and unbiased standard deviation USTD. The mathematical expressions of the 4 features are as follows;
[0065]
[0066] Among them, N represents the window size, and x i represents the i-th sEMG sample in the window, and m 0 represents the intensity of muscle contraction, and m 2 describes the change of surface electromyogram signal, and m 4 describes m 2 change, and Δ 2 represents the second derivative, and x represents the expected value of a single sliding window of electromyogram samples;
[0067] Step 3: Finally, fuse them in a dimension-wise concatenation manner. Since the number of signal channels after extracting a single feature is 12, the number of signal channels after fusing 4 features is 48.
[0068] S3. Respectively use the fused features and the corresponding finger joint angles as the input data and output data of the CNN_Attention-BiLSTM parallel model, and train the CNN_Attention-BiLSTM parallel model.
[0069] Use the normalized surface electromyogram signal data and finger joint angle data as the input-output pair of the CNN_Attention-BiLSTM parallel network to train the model. And finally obtain the prediction result on the test set, and continuously deploy the saved model to the ARM device Raspberry Pi to test its inference time to ensure the unity of accuracy and latency. The training process of the CNN_Attention-BiLSTM parallel network model is as Figure 3 shown. Specifically, input the normalized surface electromyogram signal data into the CNN_Attention network module and the BiLSTM module respectively to obtain the outputs of the two modules.
[0070] The calculation expression of the multi-scale convolution module (CNN) is as follows:
[0071]
[0072] Among them, the input signal is X, and the output of the multi-scale convolution module is X′, represents the convolution operation, C is the number of convolution kernels, D is the number of channels of the input signal, and n i is the convolution kernel size, i = 1, 2, 3, and σ is the ELU activation function.
[0073] Then, use the output of the multi-scale convolution module (CNN) as the input and inject it into the multi-head attention module (Attention). The calculation expression is as follows:
[0074] Q = W qx + b q
[0075] K = W k x + b k
[0076] V = W v x + b v
[0077]
[0078] head i = Attention(Q i , K i , V i )
[0079] MultiHead = Concat(head 1 , …, head n )W + b
[0080] Among them, Q is the query vector, K is the key vector, V is the weight value vector, Attention(Q, K, V) is the attention operation, head i is the single - head attention operation, MultiHead is the multi - head attention operation, W q , W k and W v are in R C×C , b q , b k and b v are in R C . C is the channel dimension of the single - head attention mechanism layer, d k is the dimension of Q and K, W is in R D×D , b is in R D . Concat means concatenation by dimension, softmax is the activation function, n is the number of multi - head attention heads, and x is the input signal.
[0081] This application establishes a multi - head attention mechanism with 8 heads (n = 8) in parallel. This module divides the input sequence into n groups by dimension. Then, it concatenates the outputs of multiple single - head attention layers according to the corresponding dimension. Finally, a fully - connected layer provides the output of the CNN_Attention network module.
[0082] The calculation expression of the BiLSTM module is as follows:
[0083] f t = σ(W f · [h t-1 , x t + b f )
[0084] i t = σ(W i · [h t-1 , x t + b i )
[0085]
[0086] o t = σ(W o [h t-1 , x t + b o )
[0087] h t = o t * tanh(C t )
[0088] Among them, i t is the input gate, W i is its weight matrix, b i is its bias term; f t is the forget gate, W f is its weight matrix, b f is its bias term; o t is the output gate, W o is its weight matrix, b o is its bias term; C t-1 is the cell state at the previous moment, C t is the cell state at the current moment, is the candidate cell state, W C is the weight matrix, b C is the bias term; x t is the input at the current moment, h t-1 is the hidden state at the previous moment, h t is the hidden state at the current moment; σ is the sigmoid function, and tanh is the hyperbolic tangent function.
[0089] Then, it passes through a max - pooling layer to obtain the output of the BiLSTM module.
[0090] Finally, the outputs of the two branch modules are concatenated according to the channel dimension, and the features are mapped to 10 target angles through two fully - connected layers.
[0091] Train the model and save the trained model.
[0092] S4. Obtain the surface electromyogram signal data to be estimated, send the surface electromyogram signal data into four parallel feature extraction branches respectively, fuse the four extracted features to generate the fused feature to be estimated, and input the fused feature to be estimated into the trained CNN_Attention-BiLSTM parallel model to obtain the finger joint angle data output by the CNN_Attention-BiLSTM parallel model.
[0093] As a preferred embodiment of the present invention, the estimation method further includes:
[0094] S5. Deploy the trained CNN_Attention-BiLSTM parallel model to the ARM device for offline testing.
[0095] Use the Pytorch tool to perform offline processing on the model, perform offline deployment on the trained model, deploy the model on the ARM device (Raspberry Pi 4B model), and measure the inference time of one window operation. Obtain the prediction result of the real-time joint angle.
[0096] As Figure 4 shown, in the embodiment of the present invention, the angle signal predicted by the model is compared with the real joint angle signal, and the performance of the model is intuitively reflected by plotting the predicted joint angle curve and the actual joint angle curve. Among them, the orange represents the real joint angle curve, and the blue represents the predicted joint angle curve. It can be seen from the figure that the joint angle signal predicted by the method proposed in this application has a high fitting degree with the real signal, small fluctuations, and is closest to the joint angle signal of human movement.
[0097] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for estimating continuous motion of hand joint angles based on CNN_Attention-BiLSTM parallel model, characterized in that: The following steps are involved: Acquire surface electromyographic signal data and real-time corresponding finger joint angle data during continuous motion of multiple finger joints; The surface electromyography signal data are respectively sent to four parallel feature extraction branches, the four extracted features are fused, and the sampling frequency of the fused features is kept consistent with the sampling frequency of the finger joint angle; the four parallel feature extraction branches are respectively used to extract the root mean square, peak stress, vibration expectation value and unbiased standard deviation of the surface electromyography signal data; The fusion feature and the corresponding finger joint angle are used as input data and output data of the CNN_Attention-BiLSTM parallel model respectively, and the CNN_Attention-BiLSTM parallel model is trained; The surface electromyography signal data to be estimated is obtained, and the surface electromyography signal data is respectively sent to four parallel feature extraction branches, the four extracted features are fused to generate fused features to be estimated, and the fused features to be estimated are input into the trained CNN_Attention-BiLSTM parallel model to obtain the finger joint angle data output by the CNN_Attention-BiLSTM parallel model.
2. According to claim 1, a method for estimating continuous motion of hand joint angles based on a CNN_Attention-BiLSTM parallel model, characterized in that: The CNN_Attention-BiLSTM parallel model includes the CNN_Attention branch and the BiLSTM branch, where: The CNN_Attention branch includes a multi-scale convolution module and a multi-head attention module. The calculation formula of the multi-scale convolution module is: Among them, the input signal is X, and the output of the multi-scale convolution module is X ′ , represents the convolution operation, C is the number of convolution kernels, D is the number of channels of the input signal, n i is the convolution kernel size, i=1,2,3, σ is the ELU activation function; The calculation formula of the multi-head attention module is: Q=W q x+b q K=W k x+b k V=W v x+b v head i =Attention(Q i ,K i ,V i ) MultiHead=Concat(head1,…,head n )W+b Among them, Q is the query vector, K is the key vector, V is the weight value vector, Attention(Q,K,V) is the attention operation, and head i is a single-head attention operation, MultiHead is a multi-head attention operation, W q , W k and W v In R C×C Middle, b q , b k and b v In R C In the above figure, C is the channel dimension of the single-head attention mechanism layer, and d k is the dimension of Q and K, W in R D×D In R D In the above equation, Concat means concatenation by dimension, softmax is the activation function, n is the number of multi-head attention heads, and x is the input signal; The calculation formula of the BiLSTM module is as follows: f t =σ(W f ·[h t-1 ,x t ]+b f ) i t =σ(W i ·[h t-1 ,x t ]+b i ) the t =σ(W o [h t-1 ,x t ]+b o ) h t =o t *tanh(C t ) Among them, i t is the input gate, W i Its weight matrix, b i is its bias term; f t is the forget gate, W f Its weight matrix, b f is its bias term; o t is the output gate, W o Its weight matrix, b o is its bias term; C t-1 is the cell state at the previous moment, C t is the cell state at the current moment, is the candidate cell state, W C is the weight matrix, b C is the bias term; x t Input for the current time, h t-1 is the hidden state of the previous moment, h t The hidden state at the current moment; σ is the sigmoid function, and tanh is the hyperbolic tangent function; Then a maximum pooling layer is used to obtain the output of the BiLSTM module.
3. The method for estimating continuous motion of hand joint angles based on the CNN_Attention-BiLSTM parallel model according to claim 2, characterized in that: Finally, the outputs of the CNN_Attention branch and the BiLSTM branch are concatenated according to the channel dimension, and the concatenated features are mapped to 10 finger joint angles through two fully connected layers.
4. The method for estimating continuous motion of hand joint angles based on the CNN_Attention-BiLSTM parallel model according to claim 1, characterized in that: Acquiring surface electromyographic signal data when multiple joints of the fingers move continuously also includes performing noise reduction preprocessing based on a Butterworth filter on the surface electromyographic signal data.
5. The method for estimating continuous motion of hand joint angles based on the CNN_Attention-BiLSTM parallel model according to claim 1, characterized in that: The four extracted features are fused, including: Step 1: Segment several 100ms sliding windows from the surface electromyography signal and the resampled hand motion joint angle signal, with a step size of 0.5ms; Step 2: Extract 4 features from the segmented data interval, namely, root mean square error RMS, peak stress PS, expected SE of amplitude change rate, and unbiased standard deviation USTD. The mathematical expressions of the 4 features are as follows; Among them, N represents the window size, x i represents the i-th sEMG sample in the window, m0 represents the strength of muscle contraction, m2 describes the change of surface EMG signal, m4 describes the change of m2, Δ 2 represents the second-order derivative, represents the expected value of a single sliding window of EMG samples; Step 3: Finally, they are fused by splicing them by dimension. Since the number of signal channels after extracting a single feature is 12, the number of signal channels after the fusion of the four features is 48.
6. The method for estimating continuous motion of hand joint angles based on the CNN_Attention-BiLSTM parallel model according to claim 1, characterized in that: Before training the CNN_Attention-BiLSTM parallel model, the fusion feature is normalized based on the following formula: in is the fusion feature after normalization, x max is the maximum value of the fusion feature, x min is the minimum value of the fused feature.
7. The method for estimating continuous motion of hand joint angles based on the CNN_Attention-BiLSTM parallel model according to claim 1, characterized in that: The method further comprises: Deploy the trained CNN_Attention-BiLSTM parallel model to the ARM device for offline testing.
Citation Information
Cited By
Priori feature assisted directional coordinate attention remote sensing road extraction method
CN120976762A
Neural state recognition method and device based on EMG signal, storage medium and program product
CN121242601A