Electromyographic signal limb movement intention recognition method
By introducing multi-scale position encoding and cross-layer interaction mechanism into the electromyographic signal recognition method combining CNN and Transformer, the problem of time scale and channel correlation not being captured in the existing technology is solved, and high-precision recognition of limb movement intention from electromyographic signals is achieved.
Patent Information
- Application Number
- CN202510944768.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-10-17
AI Technical Summary
In the existing technology, the electromyographic signal limb movement intention recognition method based on the combination of CNN and Transformer fails to fully capture the signal dependence at different time scales and the signal correlation within and outside the channel, resulting in inaccurate recognition.
A method for recognizing limb movement intention using electromyographic signals is adopted. By introducing multi-scale position encoding and cross-layer interaction mechanism based on convolutional neural networks and Transformer, the features of electromyographic signals at different scales and depths are extracted, and fused and position encoded. Combined with attention extraction and residual features, accurate classification of electromyographic signals is achieved.
It significantly improves the recognition accuracy of limb movement intentions from electromyographic signals, enhances the characterization capability of time-dependent features and cross-channel feature correlations, and overcomes the limitations of traditional methods in motion detail depiction and temporal dynamic modeling.
Smart Images

Figure CN120805054A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of signal recognition, in particular to a myoelectric signal limb movement intention recognition method. BACKGROUND
[0002] At present, recognizing limb movement intention is crucial for helping stroke patients recover functional movement and improve the quality of life. Especially in the recovery of hand function, the fine control of hand movement is directly related to the daily life ability of patients. Since the hand is the main part of the human body for fine operation and interaction with the outside world, hand rehabilitation has become a key link in post-stroke rehabilitation treatment. Related research shows that functional electrical stimulation (FES) technology, especially FES based on myoelectric signal control, can effectively solve this problem. In this process, machine learning and deep learning technology plays a crucial role in the classification and recognition of limb movement intention based on surface electromyography (sEMG) signals.
[0003] In the prior art, CNN is usually combined with Transformer for feature extraction, but since sEMG is time series data reflecting changes in electrical signals in muscle activity, there may be dependencies at different time scales. For example, some movements may involve rapid muscle activation, while others may require slower muscle coordination. The current research combining CNN with Transformer fails to fully capture the sEMG signal dependency at different time scales and is difficult to capture both intra-channel and cross-channel signal correlations, resulting in inaccurate recognition of limb movement intention. SUMMARY
[0004] Therefore, it is necessary to provide a myoelectric signal limb movement intention recognition method to improve the recognition accuracy of limb movement intention.
[0005] The present application adopts the following technical solutions: The present application provides a myoelectric signal limb movement intention recognition method, comprising: Collecting myoelectric signals of human movement; Extracting features of myoelectric signals at different scales and depths, and fusing features at different scales and depths to obtain fused features; position encoding the fused features at multiple time scales and channels to obtain position encoding features, extracting attention features from the position encoding features, and adding the attention features to the fused features to obtain residual features; classifying the residual features to obtain probability values of each movement category; Determining the movement category with the largest probability value as the movement intention of human movement.
[0006] Optionally, a probability value of each motion category is obtained through a motion intention recognition model; the motion intention recognition model is constructed by introducing a multi-scale position encoding and a cross-layer interaction mechanism on the basis of a convolutional neural network and a Transformer; the motion intention recognition model comprises a feature extraction module, a Transformer encoder and a classifier which are connected in series; The feature extraction module comprises two branches; one branch comprises a first convolutional layer, a first max-pooling layer and a first batch normalization layer which are connected in series; the other branch comprises a second convolutional layer, a second max-pooling layer, a second batch normalization layer, a third convolutional layer, a third max-pooling layer, a third batch normalization layer and a fourth convolutional layer which are connected in series; The Transformer encoder comprises a first reshaping layer, a Transformer encoding module and a second reshaping layer which are connected in series; the Transformer encoding module comprises a plurality of Transformer encoding units which are connected in series; each Transformer encoding unit comprises a position encoding unit, a multi-head attention unit and a feedforward unit; The classifier comprises a global average pooling layer, a first full connection layer, a first regularization layer, a second full connection layer, a second regularization layer, a third full connection layer and an activation function layer which are connected in series.
[0007] Optionally, features of the electromyographic signal at different scales and depths are extracted, and the features at different scales and depths are fused to obtain fused features, comprising: The electromyographic signal is input into the two branches of the feature extraction module; the electromyographic signal is controlled to pass through the first convolutional layer, the first max-pooling layer and the first batch normalization layer in sequence to obtain first features; and the electromyographic signal is controlled to pass through the second convolutional layer, the second max-pooling layer, the second batch normalization layer, the third convolutional layer, the third max-pooling layer, the third batch normalization layer and the fourth convolutional layer in sequence to obtain second features; The first features and the second features are added to obtain the fused features.
[0008] Optionally, the fused features are position-encoded at a plurality of time scales and channels to obtain position-encoded features; the position-encoded features are attention-extracted to obtain attention features; and the attention features are added to the fused features to obtain residual features, comprising: The fused features are input into the Transformer encoder; the fused features are reshaped through the first reshaping layer to obtain first reshaped features; The first reshaped feature is input into a Transformer encoding module, for any layer of the Transformer encoding unit, a position encoding matrix is generated by a position encoding unit; the position encoding matrix is added to the first reshaped feature to obtain a position encoding feature; the position encoding feature is extracted by a multi-head attention unit to obtain a candidate attention feature; the candidate attention feature is processed by a feedforward unit to obtain an attention feature; wherein the attention feature output by each layer of the Transformer encoding unit is the input of the next layer of the Transformer encoding unit; The attention feature is reshaped by a second reshaping layer to obtain a second reshaped feature; The fusion feature and the second reshaped feature are added to obtain a residual feature.
[0009] Optionally, the first reshaped feature is encoded by a position encoding unit to obtain a position encoding matrix, comprising: According to the index of the layer of the Transformer encoding unit, a scaling factor is calculated; According to the position encoding matrix, a position term and a divisor term are calculated; According to the scaling factor, the position term and the divisor term, a sine value and a cosine value are determined; According to the sine value and the cosine value, a two-dimensional position encoding matrix is generated; A dimension is added to the two-dimensional position encoding matrix to obtain a three-dimensional position encoding matrix.
[0010] Optionally, the position encoding feature is extracted by a multi-head attention unit to obtain a candidate attention feature, comprising: The initial attention feature of the position encoding feature is extracted by a multi-head attention mechanism; The initial attention feature is added to the position encoding matrix to obtain a fusion encoding feature; The fusion encoding feature is processed by layer normalization to obtain a candidate attention feature.
[0011] Optionally, the candidate attention feature is processed by a feedforward unit to obtain an attention feature, comprising: The candidate attention feature is processed by a feedforward neural network, and the candidate attention feature and the processed candidate attention feature are added to obtain a fusion attention feature; The fusion attention feature is processed by layer normalization to obtain an attention feature.
[0012] Optionally, the residual feature is classified to obtain a probability value of each motion category, comprising: The residual feature is input into a classifier, and the residual feature is processed by a global average pooling layer to obtain a pooling feature; The pooled features are sequentially input into a first fully connected layer, a first regularization layer, a second fully connected layer, a second regularization layer and a third fully connected layer to obtain fully connected features; The fully connected features are input into an activation function layer to determine the probability value of each motion category through the activation function.
[0013] The application provides a myoelectric signal limb motion intention recognition device. The acquisition module is configured to acquire myoelectric signals of human motion. The recognition module is configured to extract features of the myoelectric signals at different scales and depths, fuse the features at different scales and depths to obtain fused features, perform position encoding on the fused features at multiple time scales and channels to obtain position encoded features, perform attention extraction on the position encoded features to obtain attention features, add the attention features to the fused features to obtain residual features, and classify the residual features to obtain probability values of each motion category. The determination module is configured to determine a motion category with the largest probability value as a motion intention of the human motion.
[0014] The application provides a computer readable storage medium, which stores a computer program.
[0015] The application provides a computer device, which comprises a memory, a processor and a computer program stored in the memory and executable on the processor.
[0016] The above at least one technical solution adopted by the application can achieve the following beneficial effects: In the application, the fused features at different scales and depths are position encoded at multiple time scales, which effectively enhances the representation ability of the time-dependent features of the myoelectric signals and the correlation of the cross-channel features, overcomes the limitations of traditional position encoding in motion detail description and time dynamic modeling, and adds the attention features to the fused features to realize feature information transmission and fusion between different levels, improve the feature extraction depth and representation ability of the method, and thus significantly improve the accuracy of myoelectric signal limb motion intention recognition. BRIEF DESCRIPTION OF DRAWINGS
[0017] The accompanying drawings, which are included to provide a further understanding of the application, constitute a part of the application and serve to explain the application together with the specification. The illustrative embodiments of the application and their description serve to explain the application and do not limit the application in any manner. In the drawings:
[0018] Figure 1A myoelectric signal limb movement intention recognition method provided by the application is shown in the flowchart; Figure 2 A structural diagram of the feature extraction module is shown in the flowchart; Figure 3 A structural diagram of the Transformer encoder is shown in the flowchart; Figure 4 A structural diagram of the classifier is shown in the flowchart; Figure 5 16 gestures of the SIA_delsys_16_movements dataset are shown in the flowchart; Figure 6 49 gestures of the Ninapro DB2 dataset are shown in the flowchart; Figure 7 A myoelectric signal limb movement intention recognition method provided by the application is shown in the flowchart; Figure 8 Classification accuracy of different window sizes is shown in the flowchart; Figure 9 Training and verification accuracy and loss curves of the model proposed by the application on different datasets are shown in the flowchart; Figure 10 The mean and variance of the Precision, Recall and F1 Score of the different gesture recognition results of the model proposed by the application in the test set are shown in the flowchart based on the statistical results of 10 training results; Figure 11 A computer device for implementing the myoelectric signal limb movement intention recognition method provided by the application is shown in the flowchart. DETAILED DESCRIPTION
[0019] To make the purpose, technical solutions and advantages of the application clearer, the technical solutions of the application will be described clearly and completely below by combining the specific embodiments of the application and corresponding drawings. Obviously, the described embodiments are only some of the embodiments of the application, not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the application.
[0020] Term explanation: Convolutional Neural Network (CNN): CNN is a deep learning model designed specifically for processing grid-like data, with the core idea of efficiently extracting features through local perception and parameter sharing. The workflow of CNN typically involves an alternating stack of convolutional layers, activation functions, and pooling layers: convolutional layers use learnable filters to slide across the input, capturing local features, and enhance expressiveness through non-linear activation functions; pooling layers compress the feature map size, reducing computational load and enhancing translation invariance. The fully connected layers at the end of the network integrate global information and output the prediction results. During training, CNN optimizes filter parameters through backpropagation, with the unique advantages of parameter sharing and hierarchical feature learning. Transformer: Transformer is a deep learning model based on self-attention mechanism, originally used for processing sequence data and now extended to multiple fields such as speech and signal. It abandons the traditional recurrent structure (RNN) and achieves efficient parallel computation through an encoder-decoder architecture: the encoder converts the input sequence into hidden representations, and the decoder generates the target sequence accordingly. The core mechanism of self-attention calculates the association weight between any two positions in the sequence (Query-Key-Value interaction), dynamically capturing global dependency relationships and solving the problem of long-distance information decay; multi-head attention runs multiple sets of self-attention modules in parallel to extract diverse features from different subspaces. To supplement the sequence order information, the model introduces position encoding. Each layer also contains a feedforward neural network and residual connection + layer normalization to accelerate training convergence. Multi-scale position encoding: Multi-scale position encoding is a technique that integrates position information at different scales in deep learning models, aiming to enhance the model's understanding of data space or temporal structure. By generating position representations at multiple scales, the model can capture both short-range precise position relationships and long-range feature associations. This technique is particularly suitable for scenarios such as electromyography signal classification, effectively addressing the problem of insufficient feature capture in complex time dynamics with single-scale encoding, and improving the model's performance in tasks such as electromyography signal classification.
[0021] Electromyography (EMG) signals directly reveal the underlying motor intention behind human limb movement by reflecting the electrical activity of skeletal muscles. For stroke patients, the recognition of limb movement intention is crucial for restoring functional movement and improving quality of life, especially in the recovery of hand function. Since hand movement involves fine motor operations, hand rehabilitation has become a key link in post-stroke rehabilitation therapy.
[0022] sEMG, as an electrophysiological representation of neuromuscular activity, has important application value in fields such as motor function analysis, rehabilitation robot control, and human-computer interaction. However, when using sparse electrode configurations for signal acquisition, the inherent physical limitations and signal characteristics will significantly increase the difficulty of data processing, a challenge that has been confirmed in multiple studies.
[0023] First, sEMG is highly sensitive to noise interference. Due to the limited coverage of electrodes on muscle regions, the signal is easily affected by motion artifacts, environmental electromagnetic interference, and cross-talk from adjacent muscle groups. Second, the spatial sparsity of the signal leads to incomplete representation of muscle activation patterns. Traditional high-density sEMG can reconstruct the muscle activation area through spatial sampling of electrode arrays, while the signal under sparse configuration only reflects the recruitment state of local muscle fibers.
[0024] To address the technical bottlenecks of sEMG signal processing, early researchers often use traditional machine learning algorithms for sEMG signal classification, and the features are extracted through program calculation after setting the sliding window size and step length manually. The effect of motion recognition task depends on the quality of the signal feature set obtained. Support vector machine (SVM), linear discriminant analysis (LDA), and random forest (RF) classifiers have been proven to achieve high recognition rates when classifying a limited number of action types. However, when applied to large datasets containing a large number of action types (such as NinaPro DB2), traditional learning algorithms cannot meet the actual accuracy and stability requirements. Therefore, limb motion intention recognition based on machine learning faces significant challenges in developing fine motor function rehabilitation training for more motion types in wearable myoelectric control FES rehabilitation training systems.
[0025] Deep learning-based feature extraction methods have shown significant advantages in recent years. In particular, convolutional neural networks (CNN) can model the cross-channel signal correlation adaptively through hierarchical convolution kernels, and use spatial pooling operations to enhance the robust representation of local muscle activation patterns. However, a single CNN model often cannot fully capture the temporal features of dynamic limb movements, and it is difficult to fully consider the features of individual channel signals and the correlation of cross-channel signals, resulting in poor performance of final motion intention recognition. Therefore, many researchers often combine CNN with other technologies to enhance their respective abilities and further improve the accuracy and robustness of action recognition.
[0026] Transformers excel in capturing global features of sEMG, especially when analyzing complex temporal dependencies and long-term signals. The self-attention mechanism enables Transformers to capture global features of input sequences as a whole, rather than being limited to local contextual information, which allows them to better identify movement features and reduce reliance on feature extraction techniques when processing EMG signals. Additionally, Transformers support parallel processing, significantly improving efficiency in handling large-scale data compared to traditional recurrent neural networks (RNNs). However, Transformers have limitations in local feature extraction compared to convolutional neural networks (CNNs). CNNs can effectively capture detailed and local structural features of input data through local receptive fields and convolution operations, which are crucial for handling complex spatiotemporal features like EMG signals. Although the self-attention mechanism of Transformers can model global dependencies, it lacks the inherent mechanism of CNNs to impose spatial feature constraints to directly capture local details.
[0027] Therefore, to compensate for the limitations of Transformer models in local feature extraction, researchers often combine CNNs with Transformers. This hybrid architecture achieves complementary advantages in local and global feature modeling: CNNs are used first to extract local features of data, encoding local patterns in the signal, while Transformers are used for subsequent processing to capture long-range global dependencies. In this way, CNNs ensure that key details are effectively extracted, while Transformers provide flexibility and a global perspective for modeling as a whole, resulting in higher accuracy and robustness in sEMG movement intention recognition. This combination not only enhances the model's ability to recognize dynamic movements, but also effectively reduces the reliance on the weaknesses of a single model in certain specific aspects.
[0028] Although previous research combining CNNs with Transformers has achieved complementary advantages in local and global feature modeling, sEMG is time series data reflecting changes in electrical signals during muscle activity, which may exhibit dependencies at different time scales. For example, some movements may involve rapid muscle activation, while others may require slower muscle coordination. Previous research combining CNNs with Transformers has failed to adequately capture sEMG signal dependencies at different time scales and has difficulty capturing both intra-channel and cross-channel signal correlations simultaneously, resulting in suboptimal model performance.
[0029] Therefore, gesture action recognition methods based on deep learning have been widely proven to have more significant advantages in the degree of automation of feature extraction, the ability to model complex actions, and cross-scene adaptability compared to traditional machine learning methods (such as SVM, KNN, decision tree, etc.). According to the differences in network structure design and information fusion methods, existing gesture recognition methods based on deep learning can be divided into two categories: single deep learning technology and hybrid deep learning technology. Among them, single technology usually refers to feature extraction and classification based on a single deep neural network structure (such as CNN, RNN, LSTM, or Transformer), which has the advantages of simple structure, high training efficiency, etc., but has certain limitations in processing complex temporal dynamics and multi-dimensional feature fusion; while hybrid technology can improve the model's expression ability and generalization performance by fusing multiple neural network structures (such as CNN and LSTM, CNN and Transformer, etc.), or combining auxiliary modules (such as attention mechanism, position encoding mechanism, residual connection, etc.), which can more effectively capture local details and global dependencies in gesture signals, and achieve higher recognition accuracy, but still has some defects.
[0030] In summary, the existing technology in this field has the following defects: (1) In the existing technology, deep neural networks usually rely on multi-layer convolution and down-sampling operations for high-level semantic modeling during feature extraction, but in this process, low-level local detail information is easily attenuated or lost layer by layer, especially when dealing with input data such as electromyographic signals with high-frequency fluctuations and significant physiological noise interference, the feature expression has problems such as excessive abstraction and information deficiency, which leads to a decrease in recognition accuracy and insufficient robustness when the model faces complex gesture patterns or continuous action sequences. To solve the above problems, the present application proposes a cross-level feature interaction mechanism, which realizes the effective fusion of local temporal features extracted in the shallow convolutional neural network (CNN) and global context information modeled in the deep Transformer structure by introducing a residual connection structure; this mechanism preserves the local detail information of the original input while enhancing the expression integrity of the deep network, improving the generalization ability and cross-channel consistency of the model in complex action recognition scenarios, and effectively alleviating the information loss problem caused by multi-layer stacking.
[0031] (2) In the prior art, the standard position encoding mechanism adopted by the Transformer is usually static, single-scale fixed wavelength encoding, which cannot effectively adapt to the multi-time scale dynamic change characteristics existing in the electromyographic signal, especially in modeling the local and global time sequence dependence relationship in the action starting, transition state and continuous action, which easily leads to slow response of the model to short transient signals, weak long-time dependence modeling ability, thereby affecting the recognition accuracy and the precision of action boundary judgment. To solve the above problems, the present application proposes a multi-scale position encoding mechanism, which realizes flexible modeling of position encoding under different time scales by introducing a dynamic sine and cosine wavelength adjustment strategy based on hierarchical depth control; at the same time, combined with the expandable cross-channel phase coupling design, the joint modeling ability of local phase information and global time dependence is enhanced, and the position encoding method is embedded in the multi-head attention structure of the Transformer, so that the model has dynamic time scale perception ability, thereby improving the recognition accuracy of the action starting and change process, and has higher generalization performance and robustness in complex gesture sequence and cross-individual application scenarios.
[0032] (3) In the prior art, the convolutional neural network (CNN) has strong expression ability in local feature extraction and can effectively capture short-time dependence and local pattern changes, but its receptive field is limited and it is difficult to model global context relationships; relatively, the Transformer structure has advantages in processing long-distance dependence and sequence modeling, but it is not sensitive to the large amount of high-frequency, non-stationary local change characteristics existing in the electromyographic signal, resulting in limitations of the single structure in feature extraction integrity and discrimination ability. To overcome the above problems, the present application proposes a feature extraction architecture based on the fusion of CNN and Transformer structure, which fully utilizes the high stability of CNN in local time sequence feature modeling and the modeling advantage of Transformer in global dependence capture; through the preposed CNN module to extract stable local patterns, the intermediate features are transmitted into the Transformer for global context encoding, realizing the joint modeling of features in spatial and temporal dimensions, effectively improving the expression ability and context perception ability of the model to complex dynamic gestures, while ensuring the preservation of spatial structure features, enhancing the overall classification precision and model generalization performance.
[0033] To address the aforementioned issues, the present invention provides a method for recognizing limb movement intentions using myoelectric signals. Specifically, a movement intention recognition model (CNN-MSTINet) is constructed. Multi-scale position encoding and a cross-layer interaction mechanism are introduced into the CNN-Transformer backbone model. This approach aims to enhance the ability to preserve and express myoelectric signal feature information during dimensionality reduction, thereby better extracting complex features from sEMG signals. Multi-scale position encoding encodes signal positions at different time scales and channels, enabling the model to simultaneously capture short-term and long-term temporal dependencies and model the complex dynamics of sEMG signals at multiple scales. This effectively enhances the model's ability to represent temporal dependencies and cross-channel feature correlations in surface electromyographic (sEMG) signals, overcoming the limitations of traditional position encoding in depicting movement details and modeling temporal dynamics. This encoding method ensures that the unique signal features of each channel are effectively preserved and understood during processing, while also enhancing feature correlation and information sharing between channels. By capturing global temporal and channel dependencies, the accuracy and robustness of sEMG signal recognition are significantly improved, thereby facilitating a more comprehensive understanding and utilization of the complex features of sEMG signals. In addition, by building a cross-level information interaction mechanism, the feature information transmission and fusion between different levels of the neural network can be realized, further improving the feature extraction depth and representation ability of the model, thereby significantly improving the accuracy and robustness of gesture recognition.
[0034] The technical solutions provided by various embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0035] Figure 1 The flowchart of the method for recognizing limb movement intention by electromyographic signal in the present invention specifically includes the following steps: S101, collecting electromyographic signals of human body movements.
[0036] Collecting electromyographic signals of human motion includes: collecting surface electromyographic signals of human motion, preprocessing the surface electromyographic signals of human motion to obtain electromyographic signals, and the preprocessing operations include: noise removal and standardization.
[0037] S102, extracting features of the electromyographic signal at different scales and depths, and fusing the features at different scales and depths to obtain fused features; performing position encoding on the fused features at multiple time scales and channels to obtain position encoding features, performing attention extraction on the position encoding features to obtain attention features, and adding the attention features to the fused features to obtain residual features; classifying the residual features to obtain a probability value for each motion category.
[0038] Optionally, the probability value of each motion category is obtained through a motion intention recognition model; the motion intention recognition model is constructed based on convolutional neural network and Transformer by introducing multi-scale position encoding and cross-layer interaction mechanism; the motion intention recognition model comprises a feature extraction module, a Transformer encoder and a classifier connected in sequence.
[0039] The feature extraction module combines the concepts of Inception network and residual learning, and comprises two branches. One branch comprises a first convolutional layer, a first max-pooling layer and a first batch normalization layer connected in sequence. The first convolutional layer has 64 filters, a convolution kernel size of 5 and uses a ReLU activation function. The first max-pooling layer has a pooling size of 10. The other branch comprises a second convolutional layer, a second max-pooling layer, a second batch normalization layer, a third convolutional layer, a third max-pooling layer, a third batch normalization layer and a fourth convolutional layer connected in sequence. The second convolutional layer has 512 filters, a convolution kernel size of 5 and uses a ReLU activation function, followed by a second max-pooling layer with a pooling size of 5. The third convolutional layer has 256 filters, a convolution kernel size of 2 and uses a ReLU activation function, followed by a third max-pooling layer with a pooling size of 2. The fourth convolutional layer has 64 filters, a convolution kernel size of 5 and uses a ReLU activation function.
[0040] The features of the electromyographic signals at different scales and depths are extracted, and the features at different scales and depths are fused to obtain fused features, including: inputting the electromyographic signals into two branches of the feature extraction module, controlling the electromyographic signals to pass through the first convolutional layer, the first max-pooling layer and the first batch normalization layer in sequence to obtain first features, and controlling the electromyographic signals to pass through the second convolutional layer, the second max-pooling layer, the second batch normalization layer, the third convolutional layer, the third max-pooling layer, the third batch normalization layer and the fourth convolutional layer in sequence to obtain second features; and adding the first features and the second features to obtain the fused features.
[0041] After the fourth convolutional layer, the first features and the second features extracted from the two branches are fused together to obtain the fused features. As shown in Figure 2 Figure 2 This is a schematic diagram of the feature extraction module, which consists of two branches: one branch contains a convolutional layer (Conv1D), a maximum pooling layer (MaxPooling1D), and a batch normalization layer. The convolutional layer has 64 filters, a kernel size k of 5, and a ReLU activation function. The maximum pooling layer has a pooling size p of 10. The other branch contains three convolutional layers (Conv1D), two maximum pooling layers (MaxPooling1D), and two batch normalization layers. Each convolutional layer is followed by a maximum pooling layer and a batch normalization layer. The first convolutional layer has 512 filters, a kernel size of 5, and a ReLU activation function, followed by a maximum pooling layer with a pooling size of 5. The second convolutional layer has 256 filters, a kernel size of 2, and a ReLU activation function, followed by a maximum pooling layer with a pooling size of 2. The third convolutional layer has 64 filters, a kernel size of 5, and a ReLU activation function. Extracting features at different scales and depths through two branches helps enhance the model's representation and generalization capabilities. In addition, the residual structure helps alleviate the vanishing gradient problem and improves model training performance. The convolution and maximum pooling processes can be expressed as formulas (1) and (2), respectively.
[0042] (1) in, Represents the output feature map produced by the convolution operation; and Represents the number of layers and feature maps respectively; It is j The bias of the feature map; It is the convolution operation of the convolution layer; Represents the feature set of the input convolutional layer; Represents the activation function.
[0043] (2) In this equation, and The above processing steps ensure that the data features after convolution are scale-invariant, shift-invariant, and locally dependent.
[0044] Optionally, the Transformer encoder includes a first reshaping layer, a Transformer encoding module, and a second reshaping layer connected in series; the Transformer encoding module includes multiple layers of Transformer encoding units connected in series, and each layer of Transformer encoding units includes a position encoding unit, a multi-head attention unit, and a feedforward unit.
[0045] The fusion feature is positionally encoded on multiple time scales and channels to obtain a positionally encoded feature, the positionally encoded feature is attention extracted to obtain an attention feature, and the attention feature is added to the fusion feature to obtain a residual feature, including: inputting the fusion feature into a Transformer encoder, reshaping the fusion feature through a first reshaping layer to obtain a first reshaped feature; inputting the first reshaped feature into a Transformer encoding module, and generating a position encoding matrix through a position encoding unit for any layer of Transformer encoding unit; adding the position encoding matrix and the first reshaped feature to obtain a positionally encoded feature; attention extracting the positionally encoded feature through a multi-head attention unit to obtain a candidate attention feature; processing the candidate attention feature through a feedforward unit to obtain an attention feature; reshaping the attention feature through a second reshaping layer to obtain a second reshaped feature; adding the fusion feature and the second reshaped feature to obtain a residual feature.
[0046] The attention feature is a single but multi-scale information fused encoding feature tensor, the attention feature output by each layer of Transformer encoding unit is the input of the next layer of Transformer encoding unit, and the attention feature output by the last layer of Transformer encoding unit is the input feature of the second reshaping layer.
[0047] It should be noted that the process of signal encoding on different time scales and channels is realized through two designs. On the one hand, a position encoding unit is introduced in each layer of Transformer encoding unit, and the position encoding is scaled according to the number of layers. The time granularity of the first layer encoding is relatively fine, and the subsequent layers focus on the structure of a longer time period. The "time length" of each layer is different, which is equivalent to using different "time scales" to process the same signal. This hierarchical and layer-by-layer deepening processing method enables the entire Transformer encoding module to have multi-scale time modeling capability.
[0048] On the other hand, during the stage of input signal feature extraction through a convolutional neural network (feature extraction module), the model uses multiple groups of different convolution operations, which extract channel information of different frequencies and different feature dimensions. Subsequently, in the Transformer encoding module, the multi-head attention mechanism further learns on different channels respectively, which is equivalent to each "attention head" focusing on a certain type of signal feature.
[0049] The position encoding matrix is generated by a position encoding unit, including: calculating a scaling factor according to the index of the layer of the Transformer encoding unit; calculating a position term and a divisor term according to a sequence vector; determining a sine value and a cosine value according to the scaling factor, the position term and the divisor term; generating a two-dimensional position encoding matrix according to the sine value and the cosine value; and adding a dimension to the two-dimensional position encoding matrix to obtain a three-dimensional position encoding matrix.
[0050] Specifically, the multi-scale position encoding is introduced into the Transformer to form a position encoding unit, and the process of generating the position encoding matrix by the position encoding unit includes the following steps: First, the scaling factor is calculated according to the defined number of encoder layers, and the formula of the scaling factor is shown as formula (3).
[0051] (3) wherein, layer_index represents the index of the layer, specifically the index of the current Transformer encoder layer.
[0052] Then, the position term and the divisor term are generated, which are used to calculate the sin and cos values together with the scaling factor. The generation of the position contains two steps. First, as shown in formula (4), the sequence vector is generated, and each element in the sequence vector represents the number of the corresponding time step, which can be generated by simply incrementing the natural number, starting from 0 and adding 1 one by one until the end of the sequence. Subsequently, the sequence vector is converted into a column vector, as shown in formula (5). Finally, the formula of the divisor term is shown as formula (6).
[0053] (4) (5) wherein, represents the length of the input sequence.
[0054] (6) wherein, embed_dim represents the embedding dimension. range(0, embed_dim , 2) represents a sequence generated with a step of 2 from 0 to embed_dim - 1; exp refers to the exponential function.
[0055] The sine and cosine values are calculated using the position term and the divisor term, and are then used to generate the position encoding. The calculation formulas of the sine and cosine values are shown as formulas (7) and (8) respectively.
[0056] (7) (8) where, position denotes the position index, div_term denotes the divisor term, scale denotes the scaling factor.
[0057] The sine and cosine values are integrated to generate the final position encoding tensor PE, as shown in equation (9).
[0058] (9) where, i is the position index, denoting the i th position in the input sequence, k is the column index of the final position encoding matrix.
[0059] The position encoding matrix PE is then reshaped to the format of (sequence length, embedding dimension). An additional dimension is added in the batch dimension, resulting in a three-dimensional position encoding matrix PE.
[0060] The three-dimensional position encoding matrix is added to the input signal (first reshaped feature), and then processed through the multi-head attention unit, as shown in equation (10). Specifically, the position encoding feature is extracted through the multi-head attention unit to obtain the candidate attention feature, including: extracting the initial attention feature of the position encoding matrix through the multi-head attention mechanism; adding the initial attention feature to the position encoding matrix to obtain the fusion encoding feature; performing layer normalization processing on the fusion encoding feature to obtain the candidate attention feature.
[0061] (10) where, , , where , and are the learnable weight matrices of the Query, Key, and Value of the i th head, respectively. Each weight matrix maps the input X to a new space, where, i the value range of h is 1 to . The Softmax function is used to normalize the attention scores, denotes the dimension of the key vector, is used for scaling to prevent the dot product value from being too large.
[0062] The output of this process is subjected to dropout and layer normalization operations to maintain the stability of the model.
[0063] Optionally, the candidate attention features are processed by a feedforward unit to obtain an attention feature, including: processing the candidate attention features by a feedforward neural network, and adding the candidate attention features and the processed candidate attention features to obtain a fused attention feature; performing layer normalization on the fused attention features to obtain an attention feature.
[0064] The normalized signals (candidate attention features) are further processed by a feedforward neural network, followed by dropout and layer normalization steps to enhance the generalization and robustness of the model. The mathematical formula of the feedforward neural network is shown in Equation (11).
[0065] (11) in, is the normalized input signal, and are the weight matrices of the feedforward neural network, and are the bias vectors of the feedforward neural network, and ReLU is the rectified linear unit activation function.
[0066] like Figure 3 As shown, Figure 3 Figure 1 is a schematic diagram of the Transformer encoder architecture. The fused features output by the feature extraction module pass through five layers of Transformer encoding units. Each layer uses different positional encoding, and the encoding granularity gradually decreases to capture information at different scales. In the final stage, the processed signal is reshaped through a reshape layer and added to the lower-level features to form a residual connection. The residual connection helps preserve the information of the lower-level features, enabling the model to better understand and capture the complex patterns in the sEMG signal. The formula for the residual connection is shown in Equation (12).
[0067] (12) in, is the residual feature, It is the high-level feature tensor signal after the reshape layer (the second reshaped feature). It is the saved low-level feature tensor signal (fusion feature).
[0068] Optionally, the classifier comprises, in sequence, a global average pooling layer, a first fully connected layer, a first regularization layer, a second fully connected layer, a second regularization layer, a third fully connected layer, and an activation function layer; the residual feature is classified to obtain a probability value of each motion category, comprising: inputting the residual feature into the classifier, processing the residual feature through the global average pooling layer to obtain a pooled feature; sequentially passing the pooled feature through the first fully connected layer, the first regularization layer, the second fully connected layer, the second regularization layer, and the third fully connected layer to obtain a fully connected feature; inputting the fully connected feature into the activation function layer to determine the probability value of each motion category through the activation function.
[0069] Specifically, the residual feature is input into the classifier for classification, as shown in Figure 4 Figure 4 is a structural diagram of the classifier. The classifier first converts the residual feature into a fixed-size vector through a global average pooling layer, and then classifies using three fully connected layers, which respectively contain 512, 256, and 128 neurons, each layer using a ReLU activation function, and applying a 30% dropout (regularization layer) after each layer. The goal of this part is to map the previously extracted and processed feature maps to the final classification result, and to achieve a multi-class classification task through a Softmax activation function, as shown in equation (13).
[0070] (13) In the current experimental data set, the number of classifications is 16, represented by the variable s . The input data is represented by the variable , and is subjected to probability calculation to determine the likelihood of each action occurring. The final output result identifies the action with the highest probability of occurrence.
[0071] S103, determining the motion category with the largest probability value as the running intention of the human body.
[0072] In one embodiment, the present application provides an electromyographic signal limb motion intention recognition method, which combines the advantages of Transformer and CNN (convolutional neural network), effectively extracts the features of sparse electromyographic signals through multi-scale position encoding and cross-layer interaction mechanism, and improves the accuracy and robustness of the classifier, especially suitable for real-time motion recognition on edge devices. The method comprises the following steps:
[0073] 1. Signal preprocessing: removing noise and standardizing the collected surface electromyographic (sEMG) signals to enhance signal quality.
[0074] 2、Feature extraction: Local features of electromyographic signals are extracted using CNN, and local structure information is enhanced through convolutional and pooling layers to fully exploit local muscle activity patterns.
[0075] 3、Global feature modeling: The self-attention mechanism of the Transformer network is combined to capture the global dependence and long-term temporal features of the signal, improving the accuracy of motion intention recognition.
[0076] 4、Multi-scale position encoding: Signal encoding is performed on different time scales and channels, allowing the model to capture both short-term and long-term temporal dependencies and model the complex dynamics of EMG signals.
[0077] 5、Cross-layer interaction mechanism: A cross-layer interaction mechanism is introduced to enhance the fusion of local and global features and improve the model's generalization ability in various motion intention recognition tasks.
[0078] The invention points of the present application include: 1、A Transformer-CNN electromyographic signal gesture recognition algorithm combining multi-scale position encoding and cross-layer interaction mechanism is designed for effective local and global feature modeling and capturing global dependencies across channels and time. CNN can effectively capture the details and local structure features of the input data through local receptive fields and convolution operations, while ensuring a good balance between recognition accuracy and parameter quantity. The Transformer structure captures long-distance global dependencies and introduces a multi-scale position encoding mechanism to enhance the modeling ability of dynamic changes at different time granularities. Combined with the cross-layer feature interaction mechanism, it effectively alleviates the information decay problem in deep networks and achieves accurate recognition of action initiation, transition, and continuous gestures in electromyographic signals.
[0079] 2、Multi-scale position encoding encodes signals at different time scales and channels, allowing the model to capture both short-term and long-term temporal dependencies while modeling the complex dynamics of EMG signals at multiple scales. This encoding ensures that the unique signal features of each channel are effectively preserved and understood during processing, while enhancing feature correlation and information sharing between channels. By capturing global dependencies in time and channels, the accuracy and robustness of EMG signal recognition are greatly improved.
[0080] 3、A cross-layer feature interaction mechanism is proposed, which realizes the effective fusion of local temporal features extracted in shallow convolutional neural networks (CNN) and global context information modeled in deep Transformer structures through the introduction of residual connection structures. This mechanism preserves the local detail information of the original input while enhancing the expression integrity of deep networks, improving the model's generalization ability and cross-channel consistency in complex action recognition scenarios, and effectively alleviating the information loss problem caused by multi-layer stacking.
[0081] In one embodiment, in order to verify the effectiveness of the method provided by the present application, an experiment can be conducted according to the method of the present application. In the experiment, two publicly accessible data sets, SIA_delsys_16_movements and Ninapro DB2, are selected.
[0082] 1. SIA_delsys_16_movements: This dataset was collected offline using a Delsys device with a sampling frequency of 2000 Hz. Figure 5 As shown, Figure 5 The following is a schematic diagram of 16 gestures in the SIA_delsys_16_movements dataset. The dataset records 16 different hand movement signals from four healthy participants. Each movement is repeated for 6 seconds, maintained for 6 seconds, and rested for 4 seconds between each repetition. Figure 5 As shown, six electrodes were precisely placed on specific muscles of the forearm, including the extensor carpi radialis longus, flexor carpi radialis, brachioradialis, extensor carpi ulnaris, extensor digitorum, and superficial flexor digitorum, to capture surface electromyography (sEMG) signals associated with hand movements.
[0083] 2. Ninapro DB2: This dataset contains surface electromyography (sEMG) signals collected from 40 subjects during finger movements and object grasping. The 49 types of movements are divided into three experimental groups: Exercise B, Exercise C, and Exercise D. Figure 6 As shown, Figure 6 Schematic diagram of the 49 gestures in the NinaPro DB2 dataset. Exercises B and C maintain the same settings as in the NinaPro DB1 dataset, while Exercise D adds nine new force pattern data. These data are obtained by collecting the force distribution patterns when the subject presses their finger against the force sensor.
[0084] Since surface electromyography (sEMG) signals are often contaminated by unnecessary components such as power supply noise and receiver noise during the acquisition process, in order to extract reliable electromyography features and ensure the accuracy of subsequent analysis, the sEMG signals were first preprocessed. Bandpass filtering was implemented using MATLAB R2023b, and the frequency band was set to 20 Hz to 450 Hz to retain the main electromyography signal frequency band while effectively removing low-frequency and high-frequency interference noise. Then, the 8th-order wavelet denoiser of the Symlet mother wavelet was used to remove the noise. In order to simulate the real-time processing and classification of sEMG signals, this application decomposes the sEMG signal into individual segments of window length × channel, and the sliding window scheme is as follows: Figure 7 As shown, Figure 7A flowchart of a myoelectric signal limb movement intention recognition method provided by the present application. Finally, according to relevant research findings, the absolute maximum window size of the myoelectric signal should not exceed 300 milliseconds.
[0085] Since different window sizes have a certain impact on the classification accuracy of the model, the present application divides the data into window sizes ranging from 100 milliseconds to 300 milliseconds to explore the impact of different window sizes on classification accuracy. As shown in Figure 8 , Figure 8 The classification accuracy for different window sizes is shown in the figure. The experimental results show that when the window size is 300 milliseconds, the classification accuracy of the model is the highest. Therefore, the present application selects a window length of 300 milliseconds and a sliding step size of 100 milliseconds.
[0086] After preprocessing the data, the present application takes the 1st, 3rd, 4th and 6th repetitions of each gesture as the training set, and the 2nd and 5th repetitions as the test set. The ratio of test set to validation set is about 7:3. In this way, it can be ensured that the EMG data in the training set and the test set are completely non-overlapping. This segmentation method can effectively evaluate the performance of the model on unseen data, thereby verifying its generalization ability and reliability.
[0087] 3、Experimental environment In the present application, Windows 11 operating system and NVIDIA RTX 4060 GPU are used as the computing platform, and TensorFlow 2.6 framework is used to realize and evaluate various models including CNN-MSTINet model and comparative models. In order to ensure the rigor and repeatability of the experiment, all models are trained, validated and tested under the same hardware and software environment.
[0088] In order to optimize the performance of the model, the Adam optimization algorithm is selected, and the learning rate is set to 0.001. In addition, the classification cross-entropy is used as the loss function to quantify the difference between the model prediction and the actual label. During the training process, 32 samples are processed per batch, and the entire training process lasts for 100 epochs.
[0089] 4、Experimental results In the experiment, the proposed model will be comprehensively evaluated on the specified dataset using accuracy, precision, recall, and F1 score. In addition, in order to highlight the effectiveness and advantages of the proposed model, the present invention directly compares the four indicators with several state-of-the-art network models. In the sEMG-based gesture recognition task, accuracy can effectively measure the overall recognition ability of the model in all gesture categories, providing a quick and representative overview of model prediction performance. F1 score is calculated based on precision and recall, where precision represents the proportion of correctly predicted positive samples, and recall represents the proportion of correctly identified actual positive samples. F1 score can provide a more comprehensive evaluation of the model's performance on unbalanced datasets, avoiding misleading results that may be caused by relying solely on accuracy.
[0090] As shown in Figure 9 , Figure 9 The training and validation accuracy and loss curves of the proposed model of the present invention on different datasets are as follows: Figure 9 The (a) figure in the (a) figure represents the SIA_delsys_16_movements dataset, Figure 9 The (b) figure in the (b) figure represents the NinaproDB2 dataset. Figure 9 The accuracy and loss curves of the proposed model of the present invention on the training and test data of the SIA_delsys_16_movements and Ninapro DB2 datasets during the training process, as well as the experimental results of classifying each gesture, are shown. The results are presented in the form of line graphs, which provide a detailed overview of the performance of each gesture in terms of precision, recall, and F1 score. From Figure 9 It can be seen that the accuracy of the training data and validation data shows a relatively stable upward trend. Although the training accuracy has reached a high level, and there is a certain overfitting phenomenon, the accuracy gap between the training set and the validation set gradually narrows. This indicates that the model still has good generalization ability on the validation data.
[0091] Figure 10 The performance of the CNN-MSTINet model in each gesture classification task on the two datasets is shown, Figure 10 The mean and variance of the Precision, Recall, and F1 Score of the proposed model of the present invention for different gesture recognition results in the test set are shown. Figure 10The (a) in the table 1 shows the classification results on the SIA_delsys_16_movements dataset, and the results show that among the 16 gesture classes, the model performs well in most gesture classification tasks except for gesture 7 and gesture 8. Specifically, the accuracy of gesture 7 and gesture 8 is 71.99% and 73.64% respectively, and the corresponding recall rate and F1 score are low, indicating that these two gestures are more difficult in the classification process. However, the accuracy, recall rate and F1 score of other gestures are close to or close to full score, indicating that the model has high classification performance on most gesture tasks.
[0092] Figure 10 The (b) in the table 1 shows the classification results on the Ninapro DB2 dataset. Although there are a certain number of gestures that do not perform well among the 49 gesture classes, the accuracy of most gestures remains above 80%, indicating that the model still maintains good classification ability when dealing with complex multi-class tasks. In addition, although the accuracy and recall rate of some classes fluctuate, the model as a whole exhibits strong robustness and stability, and can effectively cope with the diversity of the dataset and the complexity of the task. This result further verifies the generalization ability of the CNN-MSTINet model on different datasets.
[0093] In the present application, the proposed CNN-MSTINet model is compared with mainstream gesture classification methods, mainly from the perspective of classification accuracy to evaluate the performance.
[0094] As shown in Table 1, two datasets, SIA_delsys_16 and Ninapro DB2, are used for different gesture classification tasks. Among them, in the SIA_delsys_16 dataset, the CNN-MSTINet model achieves a classification accuracy of 96.08% under a 300 millisecond window length, significantly surpassing existing gesture classification methods such as LeNet (71.67%), LCNN (83.08%), Two-stream CNN (71.54%) and CNN-VIT (93.47%), showing stronger classification ability.
[0095] The performance of the CNN-MSTINet model is further verified on the Ninapro DB2 dataset. With a 200-millisecond window length, the method provided by the present application achieves an accuracy of 83.44%, and as the window length increases, the accuracy improves to 85.77%, surpassing other methods including SVM (77.44%), RF (75.27%), CNN (78.71%), and Vision Transformer (80.02%). At the same time, the CNN-MSTINet model still shows higher accuracy compared to other advanced deep learning methods (such as MLP-Mixer with CNN), indicating that the method provided by the present application has stronger generalization ability and superior classification performance.
[0096] For multiple sub-tasks of Ninapro DB2 (such as Exercise B and Exercise D), the CNN-MSTINet model also performs well under 200-millisecond and 300-millisecond window lengths, with accuracies of 85.04% / 86.86% (Exercise B) and 96.00% / 96.85% (Exercise D), respectively. These results show that the CNN-MSTINet model provided by the present application has strong robustness and stability when facing various gesture classification tasks.
[0097] Overall, the experimental results of the CNN-MSTINet model on the two datasets fully demonstrate its superiority in gesture classification tasks. Compared with other traditional machine learning methods and deep learning models, it has higher classification accuracy and shorter inference time, providing a solid foundation for the application of real-time gesture recognition systems.
[0098] Table 1 Table 2 shows the accuracy results of the single-user recognition experiment on the SIA_delsys_16 dataset. The experiment shows that the model proposed by the present application performs very well on different users, with an accuracy of 96.70% for user 1, 96.51% for user 2, 96.83% for user 3, and 96.78% for user 4. These results show that the model can effectively adapt to individual differences, including gender, body type, muscle memory, and movement characteristics, thereby maintaining high recognition accuracy among different users. These excellent performances verify the good generalization ability and robustness of the proposed method in processing sEMG data, especially when facing individual differences, it still maintains consistency and high efficiency.
[0099] Table 2 In the application, the myoelectric signal limb movement intention recognition method can be executed without considering the order of each step. Figure 1 The execution order of each step can be determined according to requirements, and the application does not limit this.
[0100] The myoelectric signal limb movement intention recognition method provided by one or more embodiments of the application is based on the same idea, and the application further provides a corresponding myoelectric signal limb movement intention recognition device, which comprises: The acquisition module is configured to acquire myoelectric signals of human movement.
[0101] The recognition module is configured to input the myoelectric signals into a movement intention recognition model, extract features of the myoelectric signals at different scales and depths through a feature extraction module, fuse the features at different scales and depths to obtain fused features, perform position coding on the fused features at multiple time scales through a Transformer encoder, add the position-coded features to the fused features to obtain residual features, and classify the residual features through a classifier to obtain probability values of each movement category. The movement intention recognition model is constructed based on the introduction of multi-scale position coding and cross-layer interaction mechanism into a convolutional neural network and a Transformer.
[0102] The determination module is configured to determine the movement category with the largest probability value as the movement intention of the human movement.
[0103] The specific limitations of the myoelectric signal limb movement intention recognition device can be referred to the limitations of the myoelectric signal limb movement intention recognition method described above, and will not be repeated here. Each module in the myoelectric signal limb movement intention recognition device described above can be realized by software, hardware and a combination thereof in whole or in part. The above modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so as to be called and executed by the processor to perform the operations corresponding to the above modules.
[0104] The application further provides a computer readable storage medium, which stores a computer program, and the computer program can be used to execute the myoelectric signal limb movement intention recognition method provided above. Figure 1 The application further provides a computer readable storage medium, which stores a computer program, and the computer program can be used to execute the myoelectric signal limb movement intention recognition method provided above.
[0105] The application further provides a computer readable storage medium, which stores a computer program, and the computer program can be used to execute the myoelectric signal limb movement intention recognition method provided above. Figure 11 The application further provides a computer readable storage medium, which stores a computer program, and the computer program can be used to execute the myoelectric signal limb movement intention recognition method provided above. Figure 11 As shown in FIG. 6, at the hardware level, the computer device comprises a processor, an internal bus, a network interface, a memory and a non-volatile memory, and can further comprise other hardware required by business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs to realize the aboveFigure 1 The provided myoelectric signal limb movement intention recognition method.
[0106] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiments. Any reference to memory, storage, database or other medium used in the embodiments of the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0107] Any combination of the technical features of the above embodiments can be made. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist, it should be considered as the range disclosed by the present application.
Claims
1. A method for identifying limb movement intention using electromyographic signals, characterized in that: include: Collect electromyographic signals of human body movements; Extract the features of the electromyographic signal at different scales and depths, and fuse the features at different scales and depths to obtain fused features; Perform position encoding on the fusion features at multiple time scales and channels to obtain position encoding features, perform attention extraction on the position encoding features to obtain attention features, and add the attention features to the fusion features to obtain residual features; The residual features are classified to obtain the probability value of each motion category, and the motion category with the largest probability value is determined as the movement intention of the human body.
2. The method according to claim 1, characterized in that The probability value of each motion category is obtained through the motion intention recognition model. The motion intention recognition model is constructed by introducing multi-scale position encoding and cross-layer interaction mechanisms based on convolutional neural networks and Transformers. The motion intention recognition model includes a feature extraction module, a Transformer encoder, and a classifier connected in series. The feature extraction module includes two branches, one branch includes a first convolutional layer, a first maximum pooling layer, and a first batch normalization layer connected in series, and the other branch includes a second convolutional layer, a second maximum pooling layer, a second batch normalization layer, a third convolutional layer, a third maximum pooling layer, a third batch normalization layer, and a fourth convolutional layer connected in series; The Transformer encoder includes a first reshaping layer, a Transformer encoding module, and a second reshaping layer connected in series. The Transformer encoding module includes multiple layers of Transformer encoding units connected in series, and each layer of Transformer encoding units includes a position encoding unit, a multi-head attention unit, and a feedforward unit. The classifier includes a global average pooling layer, a first fully connected layer, a first regularization layer, a second fully connected layer, a second regularization layer, a third fully connected layer and an activation function layer connected in series.
3. The method according to claim 2, characterized in that Extract the features of the electromyographic signal at different scales and depths, and fuse the features at different scales and depths to obtain fused features, including: Inputting the electromyographic signal into the two branches of the feature extraction module respectively, controlling the electromyographic signal to pass through the first convolution layer, the first maximum pooling layer, and the first batch normalization layer in sequence to obtain the first feature, and controlling the electromyographic signal to pass through the second convolution layer, the second maximum pooling layer, the second batch normalization layer, the third convolution layer, the third maximum pooling layer, the third batch normalization layer, and the fourth convolution layer in sequence to obtain the second feature; Add the first feature and the second feature to obtain the fusion feature.
4. The method according to claim 2, characterized in that Perform position encoding on the fusion features at multiple time scales and channels to obtain position encoding features, perform attention extraction on the position encoding features to obtain attention features, and add the attention features to the fusion features to obtain residual features, including: The fused features are input into the Transformer encoder, and the fused features are reshaped through the first reshaping layer to obtain the first reshaped features; The first reshaped feature is input into the Transformer encoding module. For any layer of Transformer encoding unit, a position encoding matrix is generated through the position encoding unit. The position encoding matrix is added to the first reshaped feature to obtain the position encoding feature. The position encoding feature is extracted by the multi-head attention unit to obtain the candidate attention feature. The candidate attention feature is processed by the feedforward unit to obtain the attention feature. The attention feature output by each layer of Transformer encoding unit is the input of the next layer of Transformer encoding unit. The attention feature is reshaped through the second reshaping layer to obtain the second reshaped feature; The fusion feature and the second reshaped feature are added to obtain the residual feature.
5. The method according to claim 4, characterized in that Generate a position encoding matrix through the position encoding unit, including: Calculate the scaling factor based on the index of the layer of the Transformer coding unit; Calculate the position term and divisor term based on the sequence vector; Determine the sine and cosine values based on the scaling factor, the position term, and the divisor term; Generate a two-dimensional position coding matrix according to the sine value and cosine value; Adding a dimension to the two-dimensional position encoding matrix results in a three-dimensional position encoding matrix.
6. The method according to claim 4, characterized in that The position encoding features are extracted through multi-head attention units to obtain candidate attention features, including: Extract the initial attention features of the position encoding features through the multi-head attention mechanism; Add the initial attention feature to the position encoding matrix to obtain the fused encoding feature; The fused encoding features are layer-normalized to obtain candidate attention features.
7. The method according to claim 4, characterized in that The candidate attention features are processed by the feedforward unit to obtain the attention features, including: The candidate attention features are processed by a feedforward neural network, and the candidate attention features and the processed candidate attention features are added together to obtain a fused attention feature; The fused attention features are layer-normalized to obtain the attention features.
8. The method according to claim 2, characterized in that Classify the residual features and obtain the probability value of each motion category, including: The residual features are input into the classifier and processed through the global average pooling layer to obtain the pooled features; The pooled features are sequentially passed through the first fully connected layer, the first regularization layer, the second fully connected layer, the second regularization layer, and the third fully connected layer to obtain the fully connected features; The fully connected features are input into the activation function layer, and the probability value of each motion category is determined by the activation function.