Mechanical dog dual-mode control method based on myoelectric gestures and SSVEP
By adopting the dual-mode control method of electromyography and SSVEP in the mechanical dog control system, the problems of single control method and poor signal stability are solved, and flexible and efficient control in complex environments are achieved.
Patent Information
- Application Number
- CN202510638915.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-05-19
AI Technical Summary
The existing mechanical dog control system has a single control method and insufficient flexibility. The SSVEP signal has poor stability in complex environments, and the electromyography signal is affected by user action and environmental factors.
The mechanical dog dual-mode control method based on electromyography and SSVEP is adopted to identify the target by collecting environmental data and entering the detection mode or alert mode. In detection mode, use the electromyography signal for gesture recognition, and in alert mode, use the SSVEP signal for gesture recognition, and convert the recognition results into control commands.
It realizes flexible control in different scenarios, improves operational safety and efficiency, and enhances the system's anti-interference ability and user comfort.
Smart Images

Figure CN120155928A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of mechanical dog control, and specifically relates to a dual-mode control method of a mechanical dog based on electromyographic gestures and SSVEP. Background Art
[0002] Robot dogs have shown great application potential in various environments. However, traditional control methods, such as remote controls and voice commands, often fail to provide sufficient accuracy and response speed in complex environments, and are difficult to meet rapidly changing needs. In recent years, bioelectric signal control has gradually become a research hotspot in the field of human-computer interaction. This technology collects human physiological signals, analyzes and decodes these signals, and ultimately controls the device. It has the advantages of fast response speed and high concealment. Among them, steady-state visual evoked potential (SSVEP) induces specific EEG signals through light source stimulation of different frequencies, providing a control method that does not require complex training; while electromyographic signals (EMG) can quickly identify gestures and achieve real-time control of the device by capturing electrical signals generated by muscle activity. These two technologies have played an important role in improving the convenience of operation of the robot dog.
[0003] However, existing physiological electrical signal control technology still faces many challenges in practical applications. SSVEP signals are susceptible to environmental noise and electromagnetic interference in complex environments, and long-term gaze may cause visual fatigue in users; electromyographic signals may become unstable due to muscle fatigue. Existing signal processing algorithms and stimulation paradigms are still insufficient in terms of real-time and anti-interference capabilities. In addition, a single control method is also difficult to meet the needs of complex environments. Summary of the invention
[0004] The purpose of this application is to provide a dual-mode control method for a mechanical dog based on electromyographic gestures and SSVEP, so as to solve the problems of the existing mechanical dog control system, such as the single control mode and insufficient flexibility, the poor stability of SSVEP signals in complex environments, and the influence of electromyographic signals on user actions and environmental factors.
[0005] In order to achieve the above purpose, the technical solution of this application is as follows: A dual-mode control method for a mechanical dog based on electromyographic gesture and SSVEP, comprising: Collect environmental data and perform target recognition. When the predetermined target is recognized, it enters the alert mode, otherwise it enters the detection mode; In the detection mode, the electromyographic signals are collected for gesture recognition, and the gesture recognition results are converted into control instructions to control the robot dog to perform corresponding operations; In the alert mode, the SSVEP signal is collected for identification and the identification result is converted into a control instruction to control the robot dog to perform the corresponding operation.
[0006] Further, the collecting environmental data for target recognition includes: Collecting environmental images and using the YOLOv5 target detection model to recognize targets.
[0007] Further, in the detection mode, collecting EMG signals for gesture recognition includes: Preprocessing the collected EMG signals; Inputting the preprocessed EMG signals into a deep learning neural network model for gesture recognition.
[0008] Further, the deep learning neural network model is a spatio-temporal feature fusion learning network, including a recurrent multi-scale convolution module, a spatio-temporal graph convolution module, a cross-attention module, and a classification module; Among them, the recurrent multi-scale convolution module includes a multi-scale convolution sub-module and a convolutional bidirectional long short-term memory neural network sub-module, which perform the following operations: The multi-scale convolution sub-module uses different convolutional kernels to perform multi-scale operations on the preprocessed EMG signals, generating three short-term time features; Then the three short-term time features pass through the convolutional bidirectional long short-term memory neural network sub-module to obtain long-term time features; Concatenating the short-term time features and the long-term time features into long and short features; The spatio-temporal graph convolution module processes the preprocessed EMG signals, including a plurality of spatio-temporal graph convolution layers connected in sequence. Each layer adds a time convolution block and a residual connection on the basis of the traditional graph convolutional network GCN; the spatio-temporal graph convolution layer performs the following operations: The input features extract spatial features through the graph convolutional network GCN; Then perform time convolution on the spatial features through the time convolution block and perform residual connection with the input features.
[0009] Further, in the warning mode, a stimulus encoding method of mutual modulation of luminance modulation and radial scaling frequency modulation is used to induce SSVEP signals.
[0010] Further, the luminance modulation is encoded based on a sine wave function, and its calculation formula is:
[0011] Among them, is the luminance modulation frequency, is the current frame index, is the screen refresh rate, is the frame luminance value; The radial scaling frequency modulation follows the following function:
[0012] Among them, is the motion modulation frequency, is the scaling radius of the When the luminance modulation frequency and the motion modulation frequency act together, an intermodulation frequency IMF is generated based on the combination of the two:
[0013] Among them, is the luminance modulation frequency, is the motion modulation frequency, and are the integers 1 and 2 respectively.
[0014] Furthermore, the collected SSVEP signals are identified by using an attention-enhanced spatio-temporal network, and the attention-enhanced spatio-temporal network includes a spatial attention module, a temporal attention module, a temporal processing module, and a classification module.
[0015] Furthermore, the spatial attention module includes a convolutional layer, a batch normalization layer, an activation layer, an ECA module, and a dropout layer; the temporal attention module includes a convolutional layer, a batch normalization layer, an activation layer, an ECA module, and a dropout layer.
[0016] A dual-mode control method for a robotic dog based on EMG gestures and SSVEP proposed in this application has the following beneficial effects: 1. Two control modes adapted to different scenarios are provided. In the detection mode, using the EMG signal-driven control technology, the operator can command the robotic dog silently, thereby reducing the exposure risk and improving the safety of the detection task. At the same time, precise control ensures that the robotic dog efficiently executes the task. In the alert mode, the SSVEP signal control method is concealed, and the operator only needs to gaze at the visual stimulus, reducing the risk of being discovered, improving the safety of the alert task, and being able to quickly and accurately control the actions of the robotic dog, improving the task efficiency. The operator can flexibly select according to the actual task scenario without complex switching operations, greatly improving the operation flexibility.
[0017] 2. Signal processing algorithms for each of the two modes are proposed. The EMG signals in the detection mode are processed through a fusion network model, which integrates the long-term and short-term temporal and spatial feature information of the signals to achieve high-precision gesture recognition and significantly improve the control performance. The SSVEP signals in the alert mode are processed through a network model based on the attention mechanism, which can efficiently identify the visual attention state of the operator and ensure the stability and response speed of the robotic dog's actions. At the same time, reliable signal acquisition devices, such as high-precision EMG bracelets in the detection scenario and high-sensitivity EEG devices in the alert scenario, provide a solid guarantee for the stable operation of the system.
[0018] 3. A stimulus coding method based on the intermodulation of luminance and radial scaling frequency is adopted, which enhances the robustness and user comfort in outdoor environments and promotes technological innovation in this field.
[0019] This application combines multiple advantages and has outstanding performances in terms of operation flexibility, precision, task execution safety and efficiency, and technological innovation in bio-signal processing. These advantages make it have important application value in tasks such as detection and alert, lay a solid foundation for the development of future human-machine collaborative control robotic dog technology, and are expected to be widely applied in more complex environments and task scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is a flow chart of the dual-mode control method for a robotic dog based on EMG gestures and SSVEP of this application.
[0021] Figure 2 It is a schematic diagram of the spatio-temporal feature fusion learning network structure for EMG signal recognition.
[0022] Figure 3 It is a schematic diagram of the attention-enhanced spatio-temporal network structure for SSVEP signal recognition. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] In order to make the objectives, technical solutions and advantages of this application clearer, the following further details this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0024] This application provides a dual-mode control method and system for a robotic dog based on EMG gestures and SSVEP, adopts a stimulus coding method of synchronous modulation of luminance and radial scaling to improve the stability and anti-interference ability of SSVEP signals, and designs a fusion network model combining cross-attention to effectively extract signal features and achieve gesture recognition. At the same time, the dual-mode control method can significantly improve the operation efficiency and response ability of the robotic dog under changing conditions, reduce the control cost, enhance the system flexibility, reduce the computational amount, and ensure its ability to effectively execute various tasks.
[0025] An embodiment of the present application, such as Figure 1 As shown, a dual-mode control method for a mechanical dog based on electromyographic gestures and SSVEP is provided, including: Collect environmental data and perform target recognition. When the predetermined target is recognized, it enters the alert mode, otherwise it enters the detection mode; In the detection mode, the electromyographic signals are collected for gesture recognition, and the gesture recognition results are converted into control instructions to control the robot dog to perform corresponding operations; In the alert mode, the SSVEP signal is collected for identification and the identification result is converted into a control instruction to control the robot dog to perform the corresponding operation.
[0026] In this embodiment, depending on whether the predetermined target is recognized, different modes are adopted to control the robot dog to perform corresponding operations, including forward, backward, stop, turn and the like, to ensure the reliable execution of the task.
[0027] For the recognition of targets in the environment, a visual recognition method is usually used, that is, a camera is used to capture video images, and then the target in the image is identified to determine whether the predetermined target is recognized. An auditory recognition method can also be used, that is, the sound in the surrounding environment is collected to identify the target.
[0028] After identifying the target, it is necessary to determine whether it is the intended target. If it is not the intended target, it means that the current situation requires continued detection, that is, the detection mode is executed; if it is the intended target, it means that the current situation requires alert, and the alert mode is executed.
[0029] It should be noted that the predetermined target can be a specific person, a specific object, etc. As long as the predetermined target is identified, it is considered that an emergency is detected and it is necessary to switch to the alert mode.
[0030] This embodiment controls the robot dog in two modes. The first mode is the detection mode, that is, gesture recognition and control are performed through electromyographic signals; the second mode is the alert mode, that is, SSVEP signal recognition and control are performed.
[0031] A specific embodiment of the present application is described by taking visual recognition as an example, and acoustic recognition is not described here. For visual recognition, it is necessary to use an image acquisition device such as a camera to capture an environmental image, and then perform target recognition based on the environmental image. Target recognition can use the YOLOv5 target detection model to identify the target, or use other target detection models, which are not described here.
[0032] The YOLOv5 object detection model takes an input image Divide into A grid is used to predict whether each grid contains a target, as well as the location and category of the target. is the number of grids into which the image is divided in each dimension. Each grid cell will output a prediction value containing confidence, bounding box coordinates, and class probabilities. The prediction result of YOLO v5 for each grid cell can be expressed as: (1) where is the confidence, representing the probability that the target is contained within the grid; is the offset of the center of the bounding box relative to the grid, is the width and height of the bounding box; is the probability of each class, representing the likelihood that the target within the grid belongs to different classes. To ensure that the bounding box coordinates adapt to the image size, YOLOv5 normalizes these outputs and uses the activation function to ensure that the prediction values are within the range [0,1]. The position and size of the bounding box are calculated using the following formulas: (2) (3) (4) (5) where and are the width and height of the input image. represents the confidence of the bounding box, represents the conditional probability distribution of the target belonging to each class when the given bounding box contains the target. The function ensures that the output value of is within the range [0,1], while the function normalizes the class probabilities to ensure that the sum of the probabilities of all classes is 1.
[0033] In this way, YOLO v5 can effectively localize and classify the targets in the image. The identified target information, including the bounding box coordinates and class probabilities of the targets, will be transmitted to the visualization interface in real time to help the operator provide accurate environmental perception information and ensure that the task can be efficiently completed even in complex environments.
[0034] In a specific embodiment of the present application, in the detection mode, electromyographic signals are collected for gesture recognition, including: Step 1.1, preprocess the collected electromyographic signals.
[0035] In the detection mode, in this embodiment, an electromyography bracelet is adopted, which is built-in with an 8-channel high-sensitivity electromyography sensor and a 9-axis motion sensor, and is equipped with modules such as Bluetooth BLE4.2, and can capture the bioelectric changes generated by the muscle movement of the user's arm and the arm movement data. Regarding the capture of the electromyography signal of the operator by the electromyography bracelet, it is a relatively mature technology in this field and will not be elaborated here.
[0036] When preprocessing the collected electromyography signal, a first-order Butterworth low-pass filter with a cut-off frequency of 1 Hz is used to filter the original electromyography signal to eliminate high-frequency noise and extract the signal envelope. Then, the moving average method is used to smooth the signal to reduce the influence of random noise in the signal, and the smoothness of the moving average method is set to 10. Finally, according to the model input, a sliding window technique with a length of 200 milliseconds and a step size of 25 milliseconds is adopted to cut the electromyography signal.
[0037] Step 1.2: Input the preprocessed electromyography signal into the deep learning neural network model for gesture recognition.
[0038] In this embodiment, for gesture recognition, a deep learning neural network is used for gesture recognition. At present, there are already some specific network structures of deep learning neural networks for gesture recognition, which will not be elaborated one by one here.
[0039] In a specific embodiment, the deep learning neural network model adopts a spatio-temporal feature fusion learning network, such as Figure 2 shown, including a recurrent multi-scale convolution module (RMCM), a spatio-temporal graph convolution module with a spatial partitioning strategy (STGCM-SPS), a cross-attention module (CAFM), and a classification module.
[0040] The recurrent multi-scale convolution module in this embodiment includes a multi-scale convolution sub-module and a convolutional bidirectional long short-term memory neural network sub-module (CNN-BiLSTM).
[0041] In the recurrent multi-scale convolution module, the preprocessed signal is subjected to primary feature extraction through the multi-scale convolution sub-module, where represents the number of sampling samples, is the time length, is the number of channels. The multi-scale convolution sub-module performs multi-scale operations on the preprocessed electromyography signal using different convolutional kernels to generate three short-term time features , where represents the convolutional output dimension.
[0042] The specific feature extraction process of the features is as follows: (6) Among them, represents a convolution operation, respectively representing the convolution kernel size.
[0043] The innovation of the multi-scale convolution sub-module lies in that through multi-scale convolution operations, it can capture both the short-term changes and long-term dependence features of the signal, which is crucial for the dynamic changes of EMG signals.
[0044] Subsequently, the short-term time features are input into the convolutional bidirectional long short-term memory neural network sub-module to capture the long-term time features of the signal . The convolutional bidirectional long short-term memory neural network CNN-BiLSTM is a multi-input single-output network model, and its core lies in the ingenious combination of the local feature extraction ability of CNN and the sequence modeling ability of BiLSTM, which will not be elaborated here.
[0045] Then, through the concatenation operation, the short-term time features and the long-term time features are fused into long and short-term time features : (7) where represents the concatenation operation.
[0046] The spatio-temporal graph convolution module in this embodiment processes the preprocessed EMG signal, including a plurality of spatio-temporal graph convolution layers connected in sequence. Each layer adds a temporal convolutional block (Temporal Convolutional Network, TCN) and a residual connection on the basis of the traditional graph convolutional network GCN. The spatio-temporal graph convolution layer performs the following operations: The input features extract spatial features through the graph convolutional network GCN; Then, the spatial features are temporally convolved through the temporal convolutional block, and then a residual connection is made with the input features.
[0047] Specifically, an undirected weighted graph based on the inter-channel functional connectivity of the preprocessed EMG signal is constructed , where is the node set, is the adjacency matrix. The edge weights in the undirected weighted graph are defined by the Pearson correlation coefficient (Pearson correlation coefficient, PCC), and the calculation formula is: (8) where is the covariance of nodes and , and the nodes are the respective channels for EMG signal acquisition, and respectively represent nodes and Standard deviation of The value range of is 0 to 1, where indicates that there is no correlation between nodes,
[0048] In addition, according to the set threshold determine the elements in the adjacency matrix : (9) Based on the adjacency matrix , calculate the normalized Laplacian matrix , the formula is: (10) where, is the degree matrix, and its diagonal elements are the degrees of nodes, defined as: .
[0049] Then, through the Graph Convolutional Network (GCN), process the Laplacian matrix and the input signal to extract the spatial features , and its update formula is: (11) where, is the input feature, is the weight matrix, is the activation function, is the time dimension, is the number of nodes, represents the output dimension. It not only extracts the mutual relationship between nodes in the spatial dimension but also ensures that the model can pay attention to the deeper spatial feature relationship.
[0050] Then perform time convolution and residual connection to obtain the spatial features output by the spatio-temporal graph convolutional layer: (12) where the input is, is the ReLU activation function, represents the output dimension, ensuring the efficiency and accuracy of the model in gesture recognition.
[0051] In one instance, the spatio-temporal graph convolutional module includes 9 spatio-temporal graph convolutional layers connected in sequence, obtaining spatial features from the preprocessed electromyogram signal , and the output of each layer is input to the next layer for further processing.
[0052] In this embodiment, the cross-attention module reshapes the long-term and short-term temporal features and the spatial features through a linear projection layer to and to meet the input requirements of the cross-attention mechanism. The purpose of this step is to map features of different dimensions to the same feature space for better feature fusion and information interaction. And according to and generate queries ( ), keys ( ), and values ( ): (13) where are the linear projection weight matrices for queries, keys, and values respectively, represents the feature dimension. They are adjusted according to the dimension of the features and the dimension of the embedding space to ensure the effectiveness of data transformation.
[0053] Next, the obtained queries, keys, and values are used to calculate the cross-attention matrix : (14) where function is used to normalize the attention weights to ensure their sum is 1, is the dimension of the embedding space. Normalizing the attention weights ensures the uniformity and accuracy of the distribution. It enables the model to highlight the feature information that is most important for the prediction task while maintaining the complexity of the input features.
[0054] After calculating the cross-attention matrix, set the number of layers of CAFM to , and the final cross-attention fusion features are obtained through the following formula: (15) (16) where represents layer normalization, is a feed-forward network implemented through linear layers, with an output dimension of . At this stage, the output of each layer is normalized to eliminate the problem of internal covariate shift that may occur during training and enhance the training stability of the model.
[0055] In this embodiment, the classification module inputs into the classification module. Through convolution and pooling operations, the final high-level features are obtained, and finally passed to the Softmax layer to obtain the predicted labels : (17) in, is the convolution operation, is the average pooling operation.
[0056] This embodiment also converts the gesture recognition results into control instructions, captures the operator's electromyographic signals through the gForcePro+ electromyographic bracelet, and directly outputs the digital code of the corresponding gesture category after decoding by STCFF-Net (such as "01" for "clenched fist", "02" for "open palm", etc.). These digital codes are transmitted to the control system of the robot dog in real time through the Wi-Fi wireless communication link. After receiving the digital code, the robot dog parses and executes the corresponding action instructions (such as "01" for moving forward, "02" for stopping, etc.), thereby realizing efficient and real-time human-machine collaborative control, and driving the robot dog to perform operations such as moving forward, moving backward, stopping, and turning.
[0057] It should be noted that the spatiotemporal feature fusion learning network of this embodiment integrates short-term temporal features (through a cyclic multi-scale convolution module), long-term temporal features (through a CNN-BiLSTM module), and spatial features (through a spatiotemporal graph convolution module), and introduces a cross-attention mechanism (CAFM) to optimize feature fusion. Through multi-scale spatiotemporal feature fusion, the accuracy of electromyographic gesture recognition is significantly improved (especially in complex environments). The cross-attention mechanism effectively suppresses noise interference, ensures the stability of gesture control, and enhances anti-interference. The optimized network structure reduces computational complexity and achieves low-latency real-time control.
[0058] In the alert mode, the SSVEP signal is used to control the action of the robot dog. Specifically, a display is used as a visual stimulus presentation device, and a flashing letter or number is displayed in the center of the display as a stimulus target. The brightness and size of the stimulus target change at a specific frequency. Brightness modulation causes the brightness of the stimulus target to change periodically between 0 and 1, while radial scaling modulation causes the size of the stimulus target to simulate near and far motion, producing a visual "breathing" effect. When the operator looks at the flashing stimulus target in the center of the display, his visual system is stimulated, thereby inducing SSVEP signals.
[0059] In order to efficiently induce a stable signal, the present application proposes a stimulation coding method of intermodulation of brightness modulation and radial scaling frequency modulation, wherein the brightness modulation part is encoded based on a sine wave function, and its calculation formula is: (18) in, is the brightness modulation frequency, is the current frame index, is the screen refresh rate. The brightness range is 0-1. is the luminance value of the frame.
[0060] The radial scaling frequency modulation range of the stimulation target is from 0.2 Hz to 3.4 Hz, with a step size of 0.2 Hz. The frequency modulation follows the function: (19) where is the motion modulation frequency, is the scaling radius of the frame. The near and far motion of the target is simulated through radial scaling modulation, creating a visual "breathing" effect, thereby stimulating specific EEG signal frequencies.
[0061] When the luminance modulation frequency and the motion modulation frequency act together, an intermodulation frequency (IMF) is generated based on their combination: (20) where is the luminance modulation frequency, is the motion modulation frequency, and are the integers 1 and 2 respectively. The generation of the intermodulation frequency modulates the visual stimulation signal in multiple frequency dimensions, enhancing the intensity and stability of the SSVEP signal.
[0062] In this embodiment, the SSVEP stimulation coding method of luminance and radial scaling frequency intermodulation, the intermodulation frequency stimulates the SSVEP signal in multiple frequency bands, enhancing the detectability of the signal. The combined modulation method reduces the interference of outdoor light and motion noise on the signal. The breathing visual stimulation reduces the fatigue of long-term fixation.
[0063] Then, the SSVEP signal is preprocessed, including three key steps: filtering, amplification, and denoising and smoothing, to ensure the stability and quality of the signal and improve the accuracy of subsequent control. First, the SSVEP is band-pass filtered at 8 - 30 Hz to extract the effective frequency band information in the signal: (21) where represents the input original EEG signal, is the impulse response function of the band-pass filter, is the output signal after filtering. This filter retains the important frequency band information in the signal through convolution operation while reducing the interference of other frequency components.
[0064] Next, the filtered signal is subjected to gain adjustment to further enhance the signal strength. The calculation formula for gain adjustment is: (22) where is the gain coefficient of the amplifier, is the filtered signal, is the amplified signal.
[0065] Finally, the sliding window averaging method is used to denoise and smooth the signal, reducing local fluctuations and noise interference, and improving the stability and clarity of the signal. The calculation formula for sliding window averaging is: (23) where is the window length of the sliding window, is the sampling interval, is the denoised signal. By using the sliding window averaging method to smooth local fluctuations and reduce the influence of noise, the signal stability is improved.
[0066] In this embodiment, for the recognition of SSVEP signals, the attention-enhanced spatiotemporal network (AESTN) is used to extract the spatiotemporal features of SSVEP signals, and the spatial and temporal attention mechanisms are combined to optimize the feature representation, and finally the accurate recognition of specific frequency visual stimuli is achieved.
[0067] The model architecture of AESTN is as Figure 3 shown, including: a spatial attention module, a temporal attention module, a temporal processing module, and a classification module. The input signal of the network is , where represents the number of channels, represents the length of the time window.
[0068] In a specific embodiment, the spatial attention module includes a convolutional layer (Conv2d), a batch normalization layer (Batch Normal), an activation layer (PRelu), an ECA module, and a dropout layer.
[0069] In the spatial attention module, the input signal performs spatial feature screening through two-dimensional convolution operation, and the convolution kernel size is , and the output is calculated as follows: (24) where represents the spatial convolution kernel weight. The extracted spatial feature Subsequently, enhancement is performed through batch normalization and the PReLU activation function to obtain enhanced features , and the formula is: (25) where, represents the activation function, represents batch normalization.
[0070] To enhance the model's perception ability of discriminative features in the input signal, a lightweight channel attention (ECA) module is introduced into the spatial attention module. By calculating the channel attention weights to highlight important features and suppress invalid information, the feature selection ability of the spatial attention module in the channel dimension is effectively enhanced: (26) where, represents the element-wise multiplication operation, is the spatial feature weighted by channel attention. The output of the ECA module, after passing through the dropout layer, is output to the temporal attention module.
[0071] The temporal attention module in this embodiment includes a convolutional layer (Conv2d), a batch normalization layer (Batch Normal), an activation layer (PRelu), an ECA module, and a dropout layer. The temporal attention module uses the temporal attention mechanism to model the dynamic changes of the signal.
[0072] The temporal features are calculated by the following formula: (27) where, represents the temporal convolution kernel weights. Similarly, batch normalization and the PReLU activation function are applied to obtain , and the ECA module is introduced to extract more prominent features: (28) where, represents the output of the temporal attention module.
[0073] The output of the ECA module, after passing through the dropout layer, is output to the temporal processing module.
[0074] The temporal processing module in this embodiment uses a BiLSTM network to input the temporally enhanced features into the BiLSTM network to capture the global spatio-temporal features in the EEG signal: (29) where, represents the weight matrix of the LSTM cell, represents the bias vector, and Represent the forward and backward outputs of BiLSTM respectively.
[0075] Finally, the output of the model is classified through the flattening layer and the fully connected layer to accurately judge the operator's visual attention state: (30) in, represents flattening the output of the BiLSTM layer into a one-dimensional vector, and represents the weights and biases of the fully connected layer, through Output classification probabilities.
[0076] Finally, the SSVEP EEG recognition results are converted into control instructions and transmitted to the robot dog control system in real time. The instructions trigger the robot dog to perform operations including moving forward, backward, stopping, and turning, ensuring the reliable execution of the task.
[0077] The system uses the signal processing algorithm on the host computer platform to decode the SSVEP signal and convert the recognition result into a control command. The generated control command is transmitted to the control system of the robot dog through wireless communication to control the robot dog to perform operations such as forward, backward, stop, and turn, thereby achieving efficient execution of the task. In order to verify the performance of the system in a complex environment, outdoor tests were carried out. The test results show that the SSVEP control system has high stability and accuracy in practical applications.
[0078] This embodiment of the attention-enhanced spatiotemporal network uses the spatial attention module to filter key channel information, the temporal attention module to capture dynamic timing features, and the ECA module model to capture potential useful signals in the data. The attention mechanism adaptively focuses on effective features and suppresses noise interference. The spatiotemporal attention mechanism combined with the ECA module to enhance the key frequency band enables the model to quickly capture the key features of the SSVEP signal and improve recognition accuracy.
[0079] The technical solution of this application adopts a dual-mode dynamic switching mechanism to switch the detection mode (electromyography control) and the alert mode (SSVEP control) according to the environmental target detection results (YOLOv5). It can flexibly adapt to different task scenarios (covert detection / rapid alert). The electromyography mode reduces the risk of exposure, and the SSVEP mode realizes contactless covert control. Mode switching ensures the consistency and efficiency of task execution.
[0080] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A dual-mode control method for a mechanical dog based on electromyographic gestures and SSVEP, characterized in that: The dual-mode control method of the mechanical dog based on electromyographic gesture and SSVEP includes: Collect environmental data and perform target recognition. When the predetermined target is recognized, it enters the alert mode, otherwise it enters the detection mode; In the detection mode, the electromyographic signals are collected for gesture recognition, and the gesture recognition results are converted into control instructions to control the robot dog to perform corresponding operations; In the alert mode, the SSVEP signal is collected for identification and the identification result is converted into a control instruction to control the robot dog to perform the corresponding operation.
2. The dual-mode control method of a mechanical dog based on electromyographic gesture and SSVEP according to claim 1 is characterized in that: The collecting of environmental data and performing target identification includes: Collect environmental images and use the YOLOv5 target detection model to identify targets.
3. The dual-mode control method of a mechanical dog based on electromyographic gesture and SSVEP according to claim 1 is characterized in that: In the detection mode, collecting electromyographic signals for gesture recognition includes: Preprocessing the collected electromyographic signals; The preprocessed EMG signals are input into the deep learning neural network model for gesture recognition.
4. The dual-mode control method of a mechanical dog based on electromyographic gesture and SSVEP according to claim 3 is characterized in that: The deep learning neural network model is a spatiotemporal feature fusion learning network, including a cyclic multi-scale convolution module, a spatiotemporal graph convolution module, a cross attention module and a classification module; The cyclic multi-scale convolution module includes a multi-scale convolution submodule and a convolutional bidirectional long short-term memory neural network submodule, and performs the following operations: The multi-scale convolution submodule uses different convolution kernels to perform multi-scale operations on the preprocessed EMG signals to generate three short-term temporal features; Then the three short-term temporal features are passed through the convolutional bidirectional long short-term memory neural network submodule to obtain the long-term temporal features; Concatenate short-term time features and long-term time features into long-term and short-term features; The spatiotemporal graph convolution module processes the preprocessed electromyographic signal, including a plurality of spatiotemporal graph convolution layers connected in sequence, each layer adding a temporal convolution block and a residual connection on the basis of the traditional graph convolution network GCN; the spatiotemporal graph convolution layer performs the following operations: The input features are used to extract spatial features through the graph convolutional network GCN; The spatial features are then temporally convolved through the temporal convolution block and then residually connected with the input features.
5. The dual-mode control method of a mechanical dog based on electromyographic gesture and SSVEP according to claim 1, characterized in that: In the vigilance mode, a stimulus coding method of brightness modulation and radial scaling frequency modulation intermodulation was used to induce SSVEP signals.
6. The dual-mode control method of a mechanical dog based on electromyographic gesture and SSVEP according to claim 5 is characterized in that: The brightness modulation is encoded based on a sine wave function, and its calculation formula is: ; in, is the brightness modulation frequency, is the current frame index, is the screen refresh rate; The radial scaling frequency modulation follows the following function: ; in, is the motion modulation frequency, It is The scaling radius of the frame; When the brightness modulation frequency and motion modulation frequency When working together, the intermodulation frequency IMF is generated based on the combination of these two: ; in, is the brightness modulation frequency, is the motion modulation frequency, and The integers are 1 and 2 respectively.
7. The dual-mode control method of a mechanical dog based on electromyographic gesture and SSVEP according to claim 1, characterized in that: The SSVEP signal is collected for recognition and an attention-enhanced spatiotemporal network is used for recognition. The attention-enhanced spatiotemporal network includes a spatial attention module, a temporal attention module, a temporal processing module and a classification module.
8. The dual-mode control method of a mechanical dog based on electromyographic gesture and SSVEP according to claim 7, characterized in that: The spatial attention module includes a convolution layer, a batch normalization layer, an activation layer, an ECA module and a random inactivation layer; the temporal attention module includes a convolution layer, a batch normalization layer, an activation layer, an ECA module and a random inactivation layer.
Citation Information
Patent Citations
Embedded type system of outer skeleton robot
CN103722550A
Brain-myoelectricity artificial limb control device and method based on scene steady-state visual evoking
CN104398325A
Multi-mode intelligent control system and method based on electroencephalogram and myoelectricity information
CN107957783A
Service robot control method based on brain-machine interaction
CN109015635A
Facilitating user-proficiency in using radar gestures to interact with an electronic device
CN110908516A