Robot control method and device based on electroencephalogram signals, computer readable storage medium and computer program product
Patent Information
- Application Number
- CN202611081722.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-21
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-07-21
AI Technical Summary
但是,上述方法在频域特征建模以及跨受试泛化能力方面仍存在一定不足,导致导致对脑电信号的分类识别稳定性和准确性较差,最终造成协作机器人的控制准确度相对较低
[0018]采用上述技术方案后,本发明实施例至少具有如下有益效果:本发明实施例通过先对实时采集到的原始脑电信号进行带通滤波等预处理,可有效滤除无关噪声、剔除异常信号,保留与控制者意图相关的有效脑电成分,为后续特征提取和类别识别奠定高质量的数据基础,减少噪声对控制精度的干扰;还通过多频带分解将目标脑电信号拆解为多个不同频率范围的子带,提取不同频带下的脑电特征,来有效体现控制者不同的神经活动状态和控制意图;通过频域变换将提取到的特征转换至频域,有效增强了对频率特征及其谐波信息的表达能力;而利用预先训练获得的基于深度学习的Transformer网络对多频带频域特征进行处理,Transformer网络的自注意力机制可自适应挖掘各子带频域特征之间的内在相关性,实现多频带特征的高效融合,同时通过深层网络结构能进一步提取特征中的深层语义信息,有效解决了多频带特征融合不充分、深层特征挖掘不足的问题,显著提升目标类别识别的准确性和鲁棒性,从而实现对脑电信号的分类识别,最后,基于所述目标类别生成并输出用于控制机器人对应动作的控制指令,实现基于脑机接口的机器人控制。
Smart Images

Figure CN122593634B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of data processing technology, and in particular to a robot control method, device, computer-readable storage medium, and computer program product based on electroencephalogram (EEG) signals. Background Technology
[0002] Currently, Brain-Computer Interface (BCI), as a system that can convert brain activity information into control commands for external devices, can more conveniently help humans enhance their ability to control external devices. When controlling external devices, the operator does not need to move their limbs; they only generate brain signals through external stimuli or spontaneous imagination. After online analysis of these brain signals, the generated brain activity information can be converted into control commands, thereby controlling the movement of an external collaborative robot. BCI technology is of great significance for patients with movement disorders (limb disabilities, stroke, amyotrophic lateral sclerosis, cerebral palsy), as enabling these patients to communicate and interact with the outside world through BCI is currently an important approach. Currently, most input signals used in BCI systems are the operator's brain signals.
[0003] One existing method for robot control based on electroencephalogram (EEG) signals involves first acquiring EEG signals generated by the controller observing stimuli at predetermined frequencies. Then, models such as Convolutional Neural Networks (CNNs) and Transformers are used to decode these signals. By automatically learning deep features within the EEG signals, the method classifies and identifies the type of stimulus received by the controller, thereby controlling the robot's movement according to the corresponding control commands. However, this method still has limitations in frequency domain feature modeling and cross-subject generalization, leading to poor stability and accuracy in EEG signal classification and ultimately resulting in relatively low control accuracy for collaborative robots. Summary of the Invention
[0004] The technical problem to be solved by the embodiments of the present invention is to provide a robot control method based on electroencephalogram (EEG) signals, which can improve the accuracy and stability of robot control.
[0005] A further technical problem to be solved by the embodiments of the present invention is to provide a robot control device based on electroencephalogram (EEG) signals, which can improve the accuracy and stability of robot control.
[0006] A further technical problem to be solved by the embodiments of the present invention is to provide a computer-readable storage medium for storing a computer program that can improve the accuracy and stability of robot control.
[0007] A further technical problem to be solved by the embodiments of the present invention is to provide a computer program product that can improve the accuracy and stability of robot control.
[0008] To address the aforementioned technical problems, this invention first provides the following technical solution: a robot control method based on electroencephalogram (EEG) signals, comprising the following steps: The system controls the EEG acquisition device to acquire raw EEG signals generated when the controller observes a stimulus image flashing at an inherent stimulus frequency in real time. The raw EEG signals are then preprocessed to obtain the target EEG signal. The preprocessing includes at least bandpass filtering. The target EEG signal is decomposed into multiple sub-bands with different frequency ranges using a multi-band filter bank, and frequency domain transformation is performed on each sub-band to obtain multi-band frequency domain features. The multi-band frequency domain features are input into a pre-trained deep learning-based Transformer network for processing to obtain the target category. The Transformer network uses a self-attention mechanism to model the correlation between different sub-bands to achieve feature fusion and deep feature extraction of the multi-band frequency domain features. Based on the target category, control commands are generated and output to control the corresponding actions of the robot.
[0009] Furthermore, the multi-band filter bank includes several sub-filters with different bandpass ranges. The total bandpass range of each sub-filter covers the fundamental and harmonic ranges and can improve the signal-to-noise ratio of harmonics.
[0010] Furthermore, the step of performing frequency domain transformation on each of the sub-bands to obtain multi-band frequency domain features specifically includes: A random sliding window technique is used to extract data from each sub-band with a fixed-length time window, and a fast Fourier transform is used to transform the extracted data from the time domain to the frequency domain. The real and imaginary parts of the data transformed from the time domain to the frequency domain are concatenated along the frequency dimension to obtain the spectral characteristics of each sub-band signal; and The spectral features of each sub-band are stacked along the sub-band dimension to form the multi-band frequency domain features.
[0011] Furthermore, the step of inputting the multi-band frequency domain features into a pre-trained deep learning-based Transformer network for processing to obtain the target category specifically includes: The spatial filter obtained through pre-training is used to convolve each channel of the multi-band frequency domain features to obtain the corresponding spatial feature tensor. The spatial feature tensor is divided into several original spectral words by block embedding, and each original spectral word is positionally encoded to obtain a target spectral word with positional information. The target spectral words in each subband are sequence-modeled using a pre-trained word-level encoder, and the result of the sequence modeling is subjected to mean pooling to obtain the corresponding global representation vector. The global representation vectors are stacked along the sub-band dimension to form a sub-band word sequence. A pre-trained frequency-level encoder is used to explicitly model the sub-band word sequence, and the result of the explicit modeling is subjected to mean pooling to obtain frequency band context features; and The frequency band context features are input into a preset classification head for identification and classification to output the target category.
[0012] Furthermore, the step of using a pre-trained spatial filter to convolve each channel of the multi-band frequency domain features to obtain the corresponding spatial feature tensor specifically includes: The multi-band frequency domain features are convolved using a convolution kernel of a predetermined size to output convolutional features; and The convolutional features are sequentially processed through batch normalization, ELU activation function, and Dropout layer to obtain the spatial feature tensor.
[0013] Furthermore, a sine-cosine positional coding method is used to positionally encode each of the original spectral words to obtain target spectral words with positional information.
[0014] Furthermore, the word-level encoder and the frequency band-level encoder have the same structure, both including at most two sub-modules. Each sub-module is a standard Transformer encoder structure, including a multi-head self-attention mechanism and a feedforward neural network. In addition, each residual connection path is equipped with LayerScale and DropPath improvement mechanisms.
[0015] On the other hand, in order to solve the above-mentioned further technical problems, the present invention provides the following technical solution: a robot control device based on electroencephalogram (EEG) signals, which is connected to an EEG acquisition device and a robot, respectively, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the robot control method based on EEG signals as described in any one of the above.
[0016] Furthermore, in order to solve the aforementioned technical problems, the present invention provides the following technical solution: a computer-readable storage medium, the computer-readable storage medium including a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to execute the robot control method based on EEG signals as described in any one of the above.
[0017] On another front, in order to solve the aforementioned further technical problems, the present invention provides the following technical solution: a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, it implements the robot control method based on EEG signals as described in any of the above claims.
[0018] After adopting the above technical solution, the embodiments of the present invention have at least the following beneficial effects: By performing preprocessing such as bandpass filtering on the raw EEG signals acquired in real time, the embodiments of the present invention can effectively filter out irrelevant noise and eliminate abnormal signals, retaining the effective EEG components related to the controller's intention, laying a high-quality data foundation for subsequent feature extraction and category recognition, and reducing the interference of noise on control accuracy; furthermore, by decomposing the target EEG signal into multiple sub-bands of different frequency ranges through multi-band decomposition, the EEG features under different frequency bands are extracted to effectively reflect the different neural activity states and control intentions of the controller; by transforming the extracted features to the frequency domain through frequency domain transformation, the expressive power of frequency features and their harmonic information is effectively enhanced. The system utilizes a pre-trained deep learning-based Transformer network to process multi-band frequency domain features. The Transformer network's self-attention mechanism can adaptively mine the intrinsic correlation between the features of each sub-band frequency domain, achieving efficient fusion of multi-band features. At the same time, the deep network structure can further extract deep semantic information from the features, effectively solving the problems of insufficient fusion of multi-band features and insufficient mining of deep features, significantly improving the accuracy and robustness of target category recognition, thereby realizing the classification and recognition of EEG signals. Finally, based on the target category, control commands for controlling the corresponding actions of the robot are generated and output, realizing robot control based on brain-computer interface. Attached Figure Description
[0019] Figure 1 This is a flowchart illustrating the steps of an optional embodiment of the robot control method based on electroencephalogram (EEG) signals of the present invention.
[0020] Figure 2 The flowchart below shows a specific step S3 of an optional embodiment of the robot control method based on electroencephalogram (EEG) signals of the present invention.
[0021] Figure 3This is a network architecture diagram of a multi-band filter bank and a Transformer network, representing an optional embodiment of the robot control method based on electroencephalogram (EEG) signals of the present invention.
[0022] Figure 4 This is a network architecture diagram of each sub-module in an optional embodiment of the robot control method based on electroencephalogram (EEG) signals of the present invention.
[0023] Figure 5 This is a visual stimulus phase and frequency map of a stimulus image, representing an optional embodiment of the robot control method based on electroencephalogram signals of the present invention.
[0024] Figure 6 This is a diagram of the robot control system architecture of an optional embodiment of the robot control device based on electroencephalogram (EEG) signals of the present invention.
[0025] Figure 7 This is a schematic diagram of an optional embodiment of the robot control device based on electroencephalogram (EEG) signals of the present invention.
[0026] Figure 8 This is a functional block diagram of an optional embodiment of the robot control device based on electroencephalogram (EEG) signals of the present invention. Detailed Implementation
[0027] The present application will now be described in further detail with reference to the accompanying drawings and specific embodiments. It should be understood that the following illustrative embodiments and descriptions are only used to explain the present invention and are not intended to limit the present invention. Moreover, the embodiments and features in the embodiments of the present application can be combined with each other unless otherwise specified.
[0028] like Figure 1 As shown, an optional embodiment of the present invention provides a robot control method based on electroencephalogram (EEG) signals, comprising the following steps: S1: Control the EEG acquisition device 3 to acquire the raw EEG signal generated when the controller observes the stimulus image flashing at the inherent stimulus frequency in real time, and preprocess the raw EEG signal to obtain the target EEG signal. The preprocessing includes at least bandpass filtering. S2: The target EEG signal is decomposed into multiple sub-bands with different frequency ranges using a multi-band filter bank, and frequency domain transformation is performed on each sub-band to obtain multi-band frequency domain features. S3: The multi-band frequency domain features are input into a pre-trained deep learning-based Transformer network for processing to obtain the target category. The Transformer network uses a self-attention mechanism to model the correlation between different sub-bands to achieve feature fusion and deep feature extraction of the multi-band frequency domain features. S4: Generate and output control instructions for controlling the corresponding actions of robot 5 based on the target category.
[0029] This invention preprocesses the raw EEG signals acquired in real time using bandpass filtering and other methods to effectively filter out irrelevant noise and abnormal signals, retaining only the valid EEG components related to the controller's intentions. This lays a high-quality data foundation for subsequent feature extraction and category recognition, reducing the interference of noise on control accuracy. Furthermore, it decomposes the target EEG signal into multiple sub-bands with different frequency ranges through multi-band decomposition, extracting EEG features from different frequency bands to effectively reflect the controller's different neural activity states and control intentions. Frequency domain transformation converts the extracted features to the frequency domain, effectively enhancing the ability to express frequency features and their harmonic information. Finally, it utilizes pre-trained data based on... The Transformer network of deep learning processes multi-band frequency domain features. The self-attention mechanism of the Transformer network can adaptively mine the intrinsic correlation between the features of each sub-band frequency domain, realizing efficient fusion of multi-band features. At the same time, the deep network structure can further extract deep semantic information from the features, effectively solving the problems of insufficient fusion of multi-band features and insufficient mining of deep features, significantly improving the accuracy and robustness of target category recognition, thereby realizing the classification and recognition of EEG signals. Finally, based on the target category, control commands for controlling the corresponding actions of robot 5 are generated and output, realizing robot control based on brain-computer interface.
[0030] In practical implementation, to improve subsequent processing efficiency and reduce interference, the preprocessing may include mean removal, notch filtering, resampling, and data truncation based on a decision time window. Mean removal can eliminate DC offset; the bandpass filter can use 3-64Hz; the notch filter uses 50Hz to suppress power frequency interference. Since the sampling rate of the EEG acquisition device 3 is typically high (e.g., 1024 Hz), to reduce computational complexity and improve online inference efficiency, resampling reduces the signal frequency to 250 Hz. To ensure the continuity and real-time performance of online processing, the system can construct a fixed-length circular buffer for dynamic storage of the data stream, overwriting old data to avoid system latency caused by frequent memory allocation. At the start of each trial, the current buffer position is recorded. After each trial, data for the corresponding time period is truncated based on the number of sampling points as the valid EEG segment for the current trial, thereby achieving accurate data extraction and time alignment, i.e., data truncation based on a decision time window. In practical implementation, a 1-second decision time window can be used.
[0031] In an optional embodiment of the present invention, the multi-band filter bank includes several sub-filters with different bandpass ranges. The total bandpass range of each sub-filter covers the fundamental and harmonic ranges and can improve the signal-to-noise ratio of the harmonics. In this embodiment, since the response of SVEP is mainly generated by the fundamental and harmonics, the fundamental frequency refers to the brain response triggered by the stimulation frequency itself. For example, if the flicker frequency is 12Hz, a significant 12Hz component will also appear in the SSVEP signal. Harmonics are frequency components that are integer multiples of the fundamental frequency. Although harmonics can provide additional frequency information, too many higher-order harmonics may introduce noise and interference, reducing the signal-to-noise ratio (SNR) and thus affecting the classification accuracy. Therefore, in the design, it should be ensured that the total bandpass range of each sub-filter covers the fundamental and harmonic ranges and can improve the signal-to-noise ratio of the harmonics as much as possible.
[0032] In practical implementation, the multi-band filter bank of this invention includes four sub-filters with different passband ranges. During the training phase of the Transformer network, two offline datasets with different passband ranges are used. The passband ranges of the training data in the first offline dataset are 3-8Hz, 8-15Hz, 15-30Hz, and 30-50Hz. The sub-filters should fully contain the information of the specific harmonics of all stimuli and discard low SNR data above 50Hz. The passband ranges of the training data in the second offline dataset are 6-15Hz, 9-30Hz, 15-45Hz, and 30-64Hz, and low SNR data above 64Hz are discarded. In actual implementation, the passband ranges of the four sub-filters of the filter bank are set to 3-15Hz, 9-30Hz, 15-45Hz, and 30-64Hz. The specific filter bank can be implemented using a filter bank-based convolutional neural network (FB-tCNN).
[0033] In an optional embodiment of the present invention, the step of performing frequency domain transformation on each of the sub-bands to obtain multi-band frequency domain features specifically includes: A random sliding window technique is used to extract data from each sub-band with a fixed-length time window, and a Fast Fourier Transform (FFT) is used to transform the extracted data from the time domain to the frequency domain. The real and imaginary parts of the data transformed from the time domain to the frequency domain are concatenated along the frequency dimension to obtain the spectral characteristics of each sub-band signal; and The spectral features of each sub-band are stacked along the sub-band dimension to form the multi-band frequency domain features.
[0034] In this embodiment, data (EEG segments) are extracted from each sub-band at the start of the stimulus using a fixed-length time window. To ensure the randomness of the training data, a random sliding window technique is used to extract the data. After the extracted data is transformed from the time domain to the frequency domain using a Fast Fourier Transform (FFT), to preserve complete spectral information, this embodiment does not perform power spectrum compression (such as modulus or squaring) on the transformation result. Instead, the real and imaginary parts of the complex FFT are extracted separately and concatenated in the frequency dimension. This method can preserve the amplitude and phase information of the frequency components, which helps the network to identify the phase-locked features of the stimulus frequency. The spectral features of all sub-bands are further stacked in the sub-band dimension to ultimately form multi-band frequency domain features.
[0035] In an optional embodiment of the present invention, such as Figure 2 As shown, step S3 specifically includes: S31: The pre-trained spatial filter is used to convolve each channel of the multi-band frequency domain feature to obtain the corresponding spatial feature tensor; S32: The spatial feature tensor is divided into several original spectral words using patch embedding, and each of the original spectral words is positionally encoded to obtain target spectral words with positional information; S33: The pre-trained word-level encoder is used to perform sequence modeling on the target spectral words of each sub-band, and mean pooling is performed on the result of the sequence modeling to obtain the corresponding global representation vector; S34: Stack the global representation vectors along the sub-band dimension to form a sub-band word sequence. Explicitly model the sub-band word sequence using a pre-trained frequency-level encoder, and perform mean pooling on the explicit modeling result to obtain frequency band context features; and S35: Input the frequency band context features into a preset classification head for identification and classification to output the target category.
[0036] In this embodiment, since EEG signals are multi-channel data, their spatial dimension includes spatial distribution information of the response of different cortical regions of the brain to stimuli. To effectively utilize the interrelationship information between channels and effectively model the spatial distribution pattern of SSVEP signals, a spatial filter is first used to convolve each channel of the multi-band frequency domain features. Discriminative spatial projection is then learned on all EEG channels through two-dimensional convolution, effectively improving the model's ability to model spatial patterns in multi-channel signals and enhancing its generalization ability in short time windows and cross-subject scenarios. Furthermore, due to the Transformer... The architecture itself does not have the ability to model the order of elements in the sequence. Therefore, before inputting the spectral word sequence into the encoder, the frequency position of each word must be explicitly introduced. Thus, after dividing the spatial feature tensor into several original spectral words using block embedding, positional encoding is first performed on each of the original spectral words to obtain target spectral words with positional information, assisting the Transformer network in better capturing the relative and absolute positional information in the sequence structure. Furthermore, a pre-trained word-level encoder can be used to perform sequence modeling on the target spectral words in each sub-band to capture discriminative dependencies in the local and global spectral structures. Moreover, after sequence modeling is completed, a simple mean expression is used... The pooling strategy averages the embedding representations of all words in the sequence dimension to obtain the overall semantic vector of the spectral segment in each sub-band. This effectively aggregates contextual information, preserves the overall structural features across words, and avoids introducing additional learnable parameters, exhibiting good stability and efficiency. Furthermore, to fuse the spectral feature information extracted from each sub-band and learn the relationships between them, the global representation vectors are first stacked in the sub-band dimension to form a sub-band word sequence. Then, a pre-trained frequency band-level encoder is used to explicitly model the sub-band word sequence, effectively learning the potential correlations and collaborative patterns between sub-bands. To balance training stability, model capacity control, and small-sample generalization ability, the results after explicit modeling are also subjected to mean pooling to obtain frequency band contextual features, effectively aggregating contextual information. Finally, the frequency band contextual features are input into a preset classification head for recognition and classification to output the target category, achieving target classification. In specific implementations, the classification head can be a fully connected layer structure, a convolutional layer structure, or a lightweight attention structure, etc.
[0037] In an optional embodiment of the present invention, the step of convolving each channel of the multi-band frequency domain features with a pre-trained spatial filter to obtain the corresponding spatial feature tensor specifically includes: The multi-band frequency domain features are convolved using a convolution kernel of a predetermined size to output convolutional features; and The convolutional features are sequentially processed through batch normalization, ELU activation function, and Dropout layer to obtain the spatial feature tensor.
[0038] In this embodiment, a convolution kernel is first used to perform convolution operations on each channel of the multi-band frequency domain features. Then, batch normalization, ELU activation function and Dropout layer are used to improve the convergence stability and generalization ability of the model, and finally the spatial feature tensor is obtained.
[0039] In specific implementation, the multi-band filter bank and Transformer network are as follows: Figure 3 As shown, the size of the convolution kernel is set to (C, 1), where C is the number of channels, the stride is 1, there is no padding, and the number of output channels is 4C, so that each frequency point of the output contains the spatial information of C electrodes of the EEG acquisition device 3; the shape of the input tensor is [B, 1, C, 2F], where B represents the batch size, C is the number of channels, 2F is the dimension after concatenating the real and imaginary parts of the frequency domain, and the shape of the final output spatial feature tensor is [B, 4C, 2F].
[0040] In an optional embodiment of the present invention, a sine-cosine positional encoding method is used to positionally encode each of the original spectral words to obtain target spectral words with positional information. In this embodiment, the core idea of the sine-cosine positional encoding method is to use sine functions of different frequencies to encode the positional information of the original spectral words into a continuous vector, and then sum and fuse it with their spectral representation. This method has two major advantages: firstly, the encoding method is fixed, without introducing additional parameters, thus avoiding the risk of overfitting in small sample scenarios; secondly, its continuously differentiable form can maintain stable propagation in deep networks.
[0041] In practice, if the token sequence length is set to L and the embedding dimension to D, then the position encoding of the pos-th token is defined by the following formula: (Formula 1) (Formula 2) Where i represents the position in the embedding dimension, the final generated [L, D] position encoding matrix will be added element by element to the block embedding output to form the target spectral lexical with position information.
[0042] In an optional embodiment of the present invention, such as Figure 3As shown, the word-level encoder and the band-level encoder have the same structure, both including at most two layers of sub-modules. Each sub-module is a standard Transformer encoder structure, including a multi-head self-attention mechanism and a feedforward neural network. Furthermore, each residual connection path also incorporates LayerScale and DropPath improvement mechanisms. In this embodiment, both the word-level encoder and the band-level encoder are composed of sub-modules with at most two layers of a standard Transformer encoder structure. Each sub-module includes a multi-head self-attention mechanism and a feedforward neural network. Each residual connection path also incorporates LayerScale and DropPath improvement mechanisms. LayerScale is a strategy that adds a learnable diagonal matrix after the residual connection output; while DropPath is a regularization technique that randomly discards the entire residual path during training, forcing the model to learn on different paths, thereby enhancing the model's robustness and generalization ability.
[0043] In practice, the LayerScale mechanism is implemented by introducing scaling factors γ1 and γ2 to scale the output of each residual branch, effectively alleviating the gradient amplification or vanishing problem during deep network training. The input feature sequence is set as follows: ; Where B is the batch size, L is the word sequence length, and D is the word embedding dimension; First, project the input feature sequence X into Query, Key, and Value respectively: (Formula 3) in, , , It is a linear mapping matrix that can be learned and trained, d k It is the dimension of each attention head; The output of single-head self-attention is: (Formula 4) The multi-head self-attention mechanism divides the input into h heads, calculates the attention process described above for each head separately, and then concatenates the results before performing a linear transformation: (Formula 5) (Formula 6) In this embodiment of the invention, the number of attention heads is set to 4, and a 64-dimensional spectral token embedding dimension is adopted, so that the dimension of each attention head is 16. The output after multi-head self-attention processing is first dropped, then multiplied by a learnable scaling factor γ1 to control the residual signal strength of this sub-layer. The scaled output then undergoes a DropPath operation, randomly discarding some residual paths during training to guide the model to learn from multiple paths. The result after residual connections is fed into LayerNorm1 for normalization to alleviate distribution offset and enhance training stability. Subsequently, the feedforward neural network (FFN) consists of two linear transformations and a GELU activation function, specifically as follows: (Formula 7) The feedforward neural network is used to independently perform nonlinear projection transformations on each word, enhancing the model's ability to express complex features. The output of the feedforward neural network is then multiplied by a learnable scaling factor γ2, and after passing through DropPath and residual connections, it is fed into LayerNorm2 for normalization, completing the entire sub-layer structure. The module structure, taking the word encoder as an example, is as follows: Figure 4 As shown in the figure. In practice, experiments have shown that when the number of layers in a submodule exceeds two, the model training convergence slows down, especially under small sample conditions, which can easily lead to gradient oscillations and overfitting, affecting generalization ability.
[0044] Furthermore, when controlling the robot using the embodiments of the present invention, the stimulus image adopts a 4×3 stimulus layout, containing a total of 12 periodically flashing visual stimulus targets, specifically as follows: Figure 5 As shown, each stimulus target flashes at a different stimulation frequency to induce an SSVEP EEG response at the corresponding frequency when the controller gazes at it. In robot control applications, each stimulus target corresponds to a specific robot control command, thereby enabling the user to control the robot to perform corresponding operations by gazing at different stimulus targets.
[0045] The stimulus frequency and initial phase follow the Joint Frequency-Phase Modulation (JFPM) scheme from the Tsinghua University 40-objective benchmark dataset, which will not be elaborated here. Furthermore, for the above scheme, a subset containing 12 objects was selected from the original 40 classification objects, with corresponding indices {0, 3, 6, 9, 12, 15, 18, 21, 24, 27, 30, 33}, and its frequency-phase assignment was directly used to ensure consistency with offline training. Figure 6 As shown, visual stimuli are generated using sinusoidal luminance modulation. The normalized luminance is defined as: (Formula 8) Where RR is the screen refresh rate, i is the frame index, s∈[0,1], and each trial begins with a cue phase indicating the target, followed by all 12 stimuli flashing simultaneously during the stimulation period.
[0046] In terms of stimulus design, the system employs a 12-target frequency-phase coding paradigm. It generates visual flicker stimuli with different frequency and phase combinations by modulating screen brightness using a sine function, thereby inducing stable frequency response signals in the occipital-parietal lobe brain region for each target. The stimulus interface is implemented using a graphical user interface library, with the main thread controlling the Block loop, Trial timing, and cue phase switching, thus ensuring the accuracy of the experimental procedure and the precision of time control.
[0047] The EEG acquisition device 3 uses an EEG acquisition cap, which receives raw data in real time through a network data stream interface with a sampling rate of 1024 Hz. Six electrodes related to the visual cortex, namely P3, Pz, P4, O1, Oz and O2, are selected for decoding analysis.
[0048] Finally, to verify the classification performance of the proposed network model, comparative experiments were conducted on two public datasets with existing models such as tCNN, FBtCNN, SSVEPNet, and SSVEPformer. An ablation experiment was designed to optimize the network structure, and subjects were recruited offline for online classification experiments. The proposed network model was then applied to robot control, successfully completing the set grasping task. The table below summarizes the online classification accuracy of six subjects.
[0049]
[0050] Overall, the control method proposed in this embodiment (FB-FDFormer) achieved the highest average accuracy (78.3%) in online evaluation, outperforming all other comparative methods. Despite the relatively limited number of subjects in the online experiment, the control method of this embodiment still exhibited a relatively stable performance fluctuation range across different individuals, without significant performance degradation, demonstrating good cross-individual robustness and generalization ability, indicating that the model has strong adaptability in real-world online application scenarios.
[0051] From the individual results, the embodiment of the present invention achieved a classification accuracy of 68.33% on subject S01, slightly lower than FBtCNN, but better than FBtCNN on the other five subjects (p = 0.186). Although the difference with FBtCNN did not reach statistical significance, the overall trend still shows performance improvement. In contrast, the differences with other comparison methods all reached statistical significance (p < 0.05), indicating that FB-FDFormer can consistently achieve better classification performance in online scenarios.
[0052] In terms of Information Transfer Rate (ITR), the embodiments of the present invention also achieved the best performance, with an ITR of 131.74 bits / min, while the ITRs of tCNN, FBtCNN, SSVEPNet, and SSVEPformer are 100.86, 106.23, 118.94, and 115.02 bits / min, respectively. Since ITR considers both classification accuracy and decision time, this result further demonstrates that the embodiments of the present invention not only improve recognition accuracy but also achieve higher information output efficiency per unit time. This is particularly important for practical online brain-computer interface systems, because system performance depends not only on accuracy but also on interaction speed.
[0053] In summary, the embodiments of this invention achieved both high classification accuracy and superior information transmission efficiency in online experiments, indicating that the learned time-frequency feature representation possesses strong discriminative power and stability. More importantly, the features learned by the model under offline training conditions can be effectively transferred to the online deployment environment without significant performance degradation or overfitting. These results further validate the application potential and engineering feasibility of the embodiments of this invention in the actual SSVEP brain-computer interface system.
[0054] In addition, regarding robot control, the system adopts a client-server architecture, where the client is responsible for the acquisition and processing of EEG signals, and the server receives control commands and executes robot control via TCP / IP, such as... Figure 6 As shown; combined Figure 5 As shown, the system is designed with 12 visual stimuli, each corresponding to a discrete control command. Unlike traditional continuous control, this system adopts a discrete incremental control strategy, meaning that each recognition triggers only a small-amplitude motion step, rather than directly planning the complete trajectory. The classification output is first converted into discrete control signals, and then published by the ROS2 node to the corresponding motion control topic to achieve modular control. The control commands are divided into the following four categories: Category 1: Robot's mobile base control: forward, left turn, right turn, chassis control via linear velocity v posted to the ROS2 / cmd_vel topic. x and angular velocity ω z accomplish: (Formula 9) To avoid continuous false triggering, the system adopts a pulse-type time window control mechanism, that is, within a fixed time T... d After the applied velocity is applied, it automatically returns to zero according to the following formula, thereby enhancing the safety and predictability of the control: (Formula 10).
[0055] The second category: End-effector control of robots: This involves incremental control in the six degrees of freedom (±X, ±Y, ±Z) directions of Cartesian space, based on the task space. Let the current end-effector pose be: (Formula 11) The first three terms represent position, and the last three represent attitude. The system applies a fixed step increment as follows: (Formula 12) Here, Δp represents the corresponding change in the X / Y / Z directions, and the step size of the change is set to 0.02m. This small step increment strategy can effectively avoid instability near the singularity point of inverse kinematics, while enhancing the human-computer interaction controllability of the system.
[0056] Category 3: Gripper Opening and Closing: The opening and closing of the grippers are independently controlled using the Modbus industrial communication protocol. Specifically, the ROS2 control node acts as the master station, establishing a master-slave communication connection with the gripper controller via a serial bus (RS-485). After the SSVEP classification result is mapped to gripper control commands, the system sends the target position or action trigger signal to the gripper's internal control register via Modbus write register instructions, thereby realizing the opening or closing operation.
[0057] The fourth type: Automatic grasping trigger: When the SSVEP classification result is mapped to the "automatic grasping" command, the system no longer executes single-step incremental motion, but instead calls a predefined grasping process. This process includes multiple stages such as target pose determination, motion planning generation, and gripper closure control. First, the grasping pose is calculated based on the relative positional relationship between the current end effector pose and the target object; then, the motion planning module is called to generate a collision-free trajectory and drive the robotic arm to move to the target position; finally, the gripper closure command is triggered to complete the grasping action. The entire process is uniformly scheduled and executed by the control node, realizing automated closed-loop control from brain signal triggering to complete grasping behavior.
[0058] In addition, to ensure the accuracy of robot motion control, the system establishes the geometric model of the robotic arm based on the Denavit–Hartenberg (D–H) parameter method; to ensure the safety of system operation, the system is also equipped with an emergency stop function to prevent unexpected robot movement and potential collision risks.
[0059] On the other hand, such as Figure 7As shown, this embodiment of the invention further provides a robot control device 1 based on electroencephalogram (EEG) signals, which is connected to an EEG acquisition device 3 and a robot 5, respectively. It includes a processor 10, a memory 12, and a computer program stored in the memory 12 and configured to be executed by the processor 10. When the processor 10 executes the computer program, it implements the robot control method based on EEG signals as described in any of the above embodiments.
[0060] For example, the computer program can be divided into one or more modules / units, which are stored in the memory 12 and executed by the processor 10 to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the EEG-based robot control device 1. For example, the computer program can be divided into... Figure 8 The functional modules in the robot control device 1 based on EEG signals include the device control and data preprocessing module 41, the multi-band decomposition and frequency domain transformation module 42, the target recognition module 43, and the command output module 44, which respectively perform the above steps S1-S4.
[0061] The brainwave signal-based robot control device 1 can be a desktop computer, laptop, handheld computer, or cloud server, etc. The brainwave signal-based robot control device 1 may include, but is not limited to, a processor 10 and a memory 12. Those skilled in the art will understand that the schematic diagram is merely an example of the brainwave signal-based robot control device 1 and does not constitute a limitation on the brainwave signal-based robot control device 1. It may include more or fewer components than shown, or combine certain components, or use different components. For example, the brainwave signal-based robot control device 1 may also include input / output devices, network access devices, buses, etc.
[0062] The processor 10 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor. The processor 10 is the control center of the EEG-based robot control device 1, connecting all parts of the EEG-based robot control device 1 via various interfaces and lines.
[0063] The memory 12 can be used to store the computer programs and / or modules. The processor 10 implements various functions of the EEG-based robot control device 1 by running or executing the computer programs and / or modules stored in the memory 12 and calling the data stored in the memory 12. The memory 12 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as image recognition function, image overlay function, etc.), etc.; the data storage area may store data created based on the use of the EEG-based robot control device 1 (such as image data, etc.). In addition, the memory 12 may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0064] If the functions described in the embodiments of the present invention are implemented in the form of software functional modules or units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the embodiments of the present invention can implement all or part of the processes in the methods described above, or they can be accomplished by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by the processor 10, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.
[0065] In another aspect, embodiments of the present invention also provide a computer-readable storage medium, the computer-readable storage medium including a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform the robot control method based on EEG signals as described in any of the above.
[0066] In another aspect, embodiments of the present invention also provide a computer program product, including a computer program that, when executed by a processor, implements the robot control method based on electroencephalogram (EEG) signals as described in any of the above embodiments.
[0067] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0068] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the scope of protection of the present invention.
Claims
1. A robot control method based on electroencephalogram (EEG) signals, characterized in that, The method includes the following steps: The system controls the EEG acquisition device to acquire raw EEG signals generated when the controller observes a stimulus image flashing at an inherent stimulus frequency in real time. The raw EEG signals are then preprocessed to obtain the target EEG signal. The preprocessing includes at least bandpass filtering. The target EEG signal is decomposed into multiple sub-bands with different frequency ranges using a multi-band filter bank, and frequency domain transformation is performed on each sub-band to obtain multi-band frequency domain features. The multi-band frequency domain features are input into a pre-trained deep learning-based Transformer network for processing to obtain the target category. The Transformer network uses a self-attention mechanism to model the correlation between different sub-bands to achieve feature fusion and deep feature extraction of the multi-band frequency domain features. Based on the target category, generate and output control commands for controlling the corresponding actions of the robot; Specifically, the step of inputting the multi-band frequency domain features into a pre-trained deep learning-based Transformer network for processing to obtain the target category includes: The spatial filter obtained through pre-training is used to convolve each channel of the multi-band frequency domain features to obtain the corresponding spatial feature tensor. The spatial feature tensor is divided into several original spectral words by block embedding, and each original spectral word is positionally encoded to obtain a target spectral word with positional information. The target spectral words in each subband are sequence-modeled using a pre-trained word-level encoder, and the result of the sequence modeling is subjected to mean pooling to obtain the corresponding global representation vector. The global representation vectors are stacked along the sub-band dimension to form a sub-band word sequence. A pre-trained frequency-level encoder is used to explicitly model the sub-band word sequence, and the result of the explicit modeling is subjected to mean pooling to obtain frequency band context features; and The frequency band context features are input into a preset classification head for identification and classification to output the target category; The control commands include a first type of control command for controlling the movement of the robot's mobile base. By responding to the first type of control command, a linear velocity v is generated according to the following formula. x and angular velocity ω z velocity vector v: Furthermore, a pulse-type time window control mechanism is adopted, and the control time is set according to the following formula at a fixed time T. d After applying the velocity vector v to the movable base, the velocity vector v is then returned to zero. Wherein, v0 is the velocity vector v generated in response to each of the first type of control commands.
2. The robot control method based on electroencephalogram (EEG) signals as described in claim 1, characterized in that, The multi-band filter bank includes several sub-filters with different bandpass ranges. The total bandpass range of each sub-filter covers the fundamental and harmonic ranges and can improve the signal-to-noise ratio of harmonics.
3. The robot control method based on electroencephalogram (EEG) signals as described in claim 1, characterized in that, The step of performing frequency domain transformation on each of the sub-bands to obtain multi-band frequency domain features specifically includes: A random sliding window technique is used to extract data from each sub-band with a fixed-length time window, and a fast Fourier transform is used to transform the extracted data from the time domain to the frequency domain. The real and imaginary parts of the data transformed from the time domain to the frequency domain are concatenated along the frequency dimension to obtain the spectral characteristics of each sub-band signal; and The spectral features of each sub-band are stacked along the sub-band dimension to form the multi-band frequency domain features.
4. The robot control method based on electroencephalogram (EEG) signals as described in claim 1, characterized in that, The step of using a pre-trained spatial filter to convolve each channel of the multi-band frequency domain features to obtain the corresponding spatial feature tensor specifically includes: The multi-band frequency domain features are convolved using a convolution kernel of a predetermined size to output convolutional features; and The convolutional features are sequentially processed through batch normalization, ELU activation function, and Dropout layer to obtain the spatial feature tensor.
5. The robot control method based on electroencephalogram (EEG) signals as described in claim 1, characterized in that, The sine-cosine positional encoding method is used to positionally encode each of the original spectral words to obtain target spectral words with positional information.
6. The robot control method based on electroencephalogram (EEG) signals as described in claim 1, characterized in that, The term-level encoder and the frequency band-level encoder have the same structure, both including at most two sub-modules. Each sub-module is a standard Transformer encoder structure, specifically including a multi-head self-attention mechanism and a feedforward neural network. Furthermore, each residual connection path is equipped with LayerScale and DropPath improvement mechanisms.
7. A robot control device based on electroencephalogram (EEG) signals, connected to an EEG acquisition device and a robot respectively, characterized in that, The device includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the EEG-based robot control method as described in any one of claims 1 to 6.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device containing the computer-readable storage medium to perform the robot control method based on electroencephalogram (EEG) signals as described in any one of claims 1 to 6.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the robot control method based on electroencephalogram (EEG) signals as described in any one of claims 1-6.
Citation Information
Patent Citations
Multi-agent cooperative control system based on multi-mode brain-computer interface-visual tracking
CN119620861A
SSVEP (Steady-State Visual Evoked Potential) classification method based on time-frequency collaborative channel attention and multistage fusion
CN121542822A