Motor imagery classification and identification method
The cross-temporal-spatial-frequency feature fusion strategy enhances EEG-based motion imagination classification by using a dual Mamba model to extract and fuse features, addressing noise interference and non-linear characteristics in EEG signals.
Patent Information
- Application Number
- CN202510803370.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-07-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
During the collection process, EEG signals are susceptible to physiological noise such as electromyography and ophthalmic electrophysiology, and their nonlinear characteristics make it difficult for traditional methods to effectively capture complex timing time-frequency characteristics, affecting the accuracy of motion imagination classification.
The cross-space-time-frequency feature fusion strategy is adopted, and the time-time-time-frequency features and timing-frequency features of motion imagination EEG data are extracted through the multi-branch space-time convolution module and the Mamba module, and the feature fusion is carried out through the fusion module, and finally classified identification is performed.
It improves the accuracy of motion imagination classification, effectively captures global characteristics, and improves the accuracy and reliability of classification tasks.
Smart Images

Figure CN120316591A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of electroencephalogram (EEG) signal monitoring and processing, and particularly to a method for classifying and recognizing motor imagery. Background Art
[0002] In the field of Brain-Computer Interface (BCI), realizing human-computer interaction by decoding EEG signals has become an important direction of research and application. Motor Imagery (MI) is an important branch of it, mainly referring to the activity of activating related brain regions by an individual imagining certain movement behaviors without actual limb movement. Electroencephalogram (EEG) signals, as bioelectrical signals directly reflecting brain electrical activities, are recorded. By collecting EEGs generated during the motor imagery process, they can be used to distinguish and recognize an individual's movement intention. This technology has been widely applied in fields such as stroke motor function rehabilitation and external mechanical control, and its application prospects continue to expand with the development of technology.
[0003] However, during the collection of EEG, it is extremely vulnerable to interference from physiological noises such as electromyogram and electrooculogram. The inclusion of these noises will not only reduce the signal quality of EEG, but may also cause a significant reduction in the classification accuracy in motor imagery tasks. In addition, EEG signals have strong non-linear characteristics, and their temporal and time-frequency characteristics show large differences due to different inter-individual neural activities. Traditional linear methods often have difficulty effectively capturing these complex characteristics, which may affect the classification performance of the model. Therefore, how to effectively remove noises and improve the classification accuracy of motor imagery based on EEG has become an important technical problem in the BCI field. Summary of the Invention
[0004] An object of this application is to propose a method for classifying and recognizing motor imagery. The cross-temporal and spatio-temporal frequency feature fusion strategy provided by this application can effectively improve the classification accuracy of MI tasks.
[0005] According to an embodiment of the first aspect of this application, a method for classifying and recognizing motor imagery is provided. The method includes: Obtaining motor imagery EEG data; Performing temporal feature extraction on the motor imagery EEG data based on a spatio-temporal feature extraction module to obtain temporal features, where the spatio-temporal feature extraction module includes a multi-branch spatio-temporal convolution module and a Mamba module; Performing temporal and time-frequency feature extraction on the motor imagery EEG data based on a time-frequency feature extraction module to obtain temporal and time-frequency features, where the time-frequency feature extraction module includes a multi-branch time-frequency convolution module and a Mamba module; Fuse the temporal-spatial features and the temporal-frequency features to obtain fused features, and classify and identify the motor imagery EEG data based on the fused features.
[0006] In some embodiments, the spatio-temporal feature extraction module and the time-frequency feature extraction module are two parallel feature extraction modules. The two parallel feature extraction modules are respectively and independently used to obtain the motor imagery EEG data, and respectively extract features from the motor imagery EEG data to obtain the temporal-spatial features and the temporal-frequency features corresponding to the motor imagery EEG data.
[0007] In some embodiments, the temporal-spatial feature extraction module extracts temporal-spatial features from the motor imagery EEG data. The spatio-temporal feature extraction module includes a multi-branch spatio-temporal convolution module and a Mamba module, including: Extract local features of different scales from the motor imagery EEG data respectively based on the multi-branch spatio-temporal convolution module in the spatio-temporal feature extraction module to obtain multi-scale local temporal-spatial features; Input the multi-scale local temporal-spatial features into the Mamba module in the spatio-temporal feature extraction module, perform linear mapping expansion on the multi-scale local temporal-spatial features based on the Mamba module, and obtain global temporal-spatial features through an activation function and discrete state space processing.
[0008] In some embodiments, the time-frequency feature extraction module extracts temporal-frequency features from the motor imagery EEG data. The time-frequency feature extraction module includes a multi-branch time-frequency convolution module and a Mamba module, including: Obtain the frequency domain features corresponding to the motor imagery EEG data; Extract local features of different scales from the frequency domain features of the motor imagery EEG data respectively based on the spatio-temporal convolution module with multi-branch different-sized convolutional kernels in the time-frequency feature extraction module to obtain multi-scale local temporal-frequency features; Input the multi-scale local temporal-frequency features into the Mamba module in the time-frequency feature extraction module, perform linear mapping expansion on the multi-scale local temporal-frequency features based on the Mamba module, and obtain global temporal-frequency features through an activation function and discrete state space processing.
[0009] In some embodiments, the step of fusing the temporal-spatial features and the temporal-frequency features to obtain fused features, and classifying and identifying the motor imagery EEG data based on the fused features includes: Convert the temporal-frequency features to temporal-spatial features, and fuse the data of the two temporal-spatial features through a convolutional layer to obtain fused features; Classify and recognize the motor imagery EEG data based on the fused features and the fully connected layer.
[0010] In some embodiments, the method further includes enhancing the motor imagery EEG data, and extracting temporal features and temporal time-frequency features based on the enhanced motor imagery EEG data; The enhancing of the motor imagery EEG data includes: Perform sequence decomposition on all channels of the original motor imagery EEG data to obtain a plurality of decomposition terms; Randomly recombine the decomposition terms to generate the enhanced motor imagery EEG data.
[0011] In some embodiments, the multi-branch spatio-temporal convolution module includes three convolutional layers with different-sized convolutional kernels and different strides, wherein the different convolutional kernels are used to extract spatio-temporal local features in different-scale spaces; The multi-branch time-frequency convolution module includes two convolutional layers with different-sized convolutional kernels and different strides, wherein the different convolutional kernels are used to extract time-frequency local features in different-scale spaces.
[0012] In some embodiments, the spatio-temporal feature extraction module extracts temporal features from the motor imagery EEG data to obtain temporal spatio-temporal features, wherein the spatio-temporal feature extraction module includes a multi-branch spatio-temporal convolution module and a Mamba module, and includes: Expand the hidden dimension of the temporal feature vector output by the multi-branch spatio-temporal convolution module through linear mapping to obtain extended data; Process the extended data based on the convolutional function and the activation function to obtain intermediate data; Determine the state data based on the intermediate data and the discrete state space model; Combine the state data and the intermediate data based on linear mapping to obtain combined data; Combine the combined data with the input residual to obtain the temporal features.
[0013] In some embodiments, the time-frequency feature extraction module extracts time-frequency features from the motor imagery EEG data to obtain time-frequency features, wherein the time-frequency feature extraction module includes a multi-branch time-frequency convolution module and a Mamba module, and includes: Expand the hidden dimension of the time-frequency feature vector output by the multi-branch time-frequency convolution module through linear mapping to obtain extended data; Process the extended data based on the convolutional function and the activation function to obtain intermediate data; Determine the state data based on the intermediate data and the discrete state space model; Combining the state data and the intermediate data based on a linear mapping to obtain combined data; Combining the combined data with an input residual to obtain temporal and frequency features.
[0014] In some embodiments, the method further includes: Obtaining motor imagery electroencephalogram data corresponding to multiple users, and aligning the motor imagery electroencephalogram data corresponding to the multiple users by means of domain alignment; Performing temporal feature extraction on the motor imagery electroencephalogram data after domain alignment of the multiple users based on a spatio-temporal feature extraction module to obtain temporal features, where the spatio-temporal feature extraction module includes a multi-branch spatio-temporal convolution module and a Mamba module; Performing temporal and frequency feature extraction on the motor imagery electroencephalogram data after domain alignment of the multiple users based on a time-frequency feature extraction module to obtain temporal and frequency features, where the time-frequency feature extraction module includes a multi-branch time-frequency convolution module and a Mamba module; Fusing the temporal features and the temporal and frequency features corresponding to the multiple users to obtain fused features, and classifying and recognizing the motor imagery electroencephalogram data based on the fused features.
[0015] This application also provides a motor imagery classification and recognition device, where the device includes: A data acquisition module, configured to acquire motor imagery electroencephalogram data; A spatio-temporal feature extraction module, configured to perform temporal feature extraction on the motor imagery electroencephalogram data based on the spatio-temporal feature extraction module to obtain temporal features, where the spatio-temporal feature extraction module includes a multi-branch spatio-temporal convolution module and a Mamba module; A time-frequency feature extraction module, configured to perform temporal and frequency feature extraction on the motor imagery electroencephalogram data based on the time-frequency feature extraction module to obtain temporal and frequency features, where the time-frequency feature extraction module includes a multi-branch time-frequency convolution module and a Mamba module; A fusion module, configured to fuse the temporal features and the temporal and frequency features to obtain fused features, and classify and recognize the motor imagery electroencephalogram data based on the fused features.
[0016] This application also provides a computer-readable storage medium, storing a computer program, where when the computer program is executed by a processor, the steps of a motor imagery classification and recognition method provided in any one of the above embodiments are implemented.
[0017] This application also provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of a motor imagery classification and recognition method provided in any one of the above embodiments are implemented.
[0018] The present application also provides a computer program product, on which a program or instructions are stored, and when the program or instructions are executed by a processor, the steps of a motor imagery classification and recognition method provided in any of the above embodiments are implemented.
[0019] The present application provides a neural network method for spatio-temporal and time-frequency feature fusion. The method mainly includes: using spatio-temporal convolution to extract the temporal features of the electroencephalogram (EEG) signal time series; converting the time series to the frequency domain by means of the Fourier transform (FFT), and then extracting the temporal spatio-temporal features. The obtained temporal spatio-temporal features and temporal domain features respectively fully mine the global features through the Mamba model, and then perform feature fusion through the fusion module proposed in the present application, and finally form a motor imagery classification and recognition method with high performance.
[0020] To solve the problems of the prior art, the present application provides a motor imagery recognition solution, which increases the feature dimension by separately extracting temporal spatio-temporal features and temporal time-frequency features from the time series, uses a dual-Mamba structure to fully extract the global features in the time series and frequency domain, and then uses a fusion module to fuse the dual features. This cross spatio-temporal and time-frequency feature fusion strategy can effectively improve the classification accuracy of the MI task.
[0021] The motor imagery recognition method based on the dual-Mamba spatio-temporal and time-frequency feature fusion network provided in the above embodiments has the following specific beneficial effects: Adopt a multi-layer spatio-temporal convolution structure, and the convolution kernel sizes and strides of each layer are different, so as to efficiently extract local features. On this basis, the Mamba model is introduced to accurately capture the global time features from the multi-layer local features, and improve the ability to mine the global information of the time dimension in the time series data.
[0022] Use the fast Fourier transform (FFT) to convert the time series data to the frequency domain, and use a convolutional neural network (CNN) to extract the local features in the frequency domain. Subsequently, the global temporal spatio-temporal features are obtained from the temporal time-frequency features by means of the Mamba model, and finally the data is restored from the frequency domain to the time domain by the inverse Fourier transform, realizing the in-depth mining and integration of the time-frequency domain information.
[0023] Construct a fusion module to organically fuse the global features captured by the above two methods, and retain the key information in the two types of features to the greatest extent. The fused features are input into a classifier for classification, effectively improving the accuracy and reliability of the classification task.
[0024] In order to extract global features, combined with the Mamba model, it can effectively capture global time features and is more efficient than the traditional Transformer when dealing with sequence problems.
[0025] This application constructs a motor imagery classification model based on the Mamba model. The Mamba model integrates the advantageous features of CNN and RNN. It can not only comprehensively and fully extract the global features in the data, but also has the ability of concurrent training, thus significantly improving the training efficiency. This innovative method provides a new way to solve the inherent problems of traditional models and is expected to achieve more excellent application results in related fields.
[0026] Additional aspects and advantages of this application will be given in part in the following description, will become apparent in part from the following description, or will be understood through the practice of this application. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The above and / or additional aspects and advantages of this application will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, where: Figure 1 is a schematic flowchart of a motor imagery recognition method provided in an embodiment of this application; Figure 2 is a schematic diagram of a motor imagery experimental paradigm provided in an embodiment of this application; Figure 3 is a schematic flowchart of a data augmentation method provided in an embodiment of this application; Figure 4 is a structural diagram of a Mamba module provided in an embodiment of this application; Figure 5 is a schematic flowchart of a motor imagery recognition method provided in another embodiment of this application; Figure 6 is a structural diagram of a motor imagery recognition device provided in an embodiment of this application; Figure 7 is a schematic diagram of a dual Mamba structure provided in an embodiment of this application; Figure 8 is a schematic diagram of the structure of a computer device provided in an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] In order to make the objectives, technical solutions and advantages of this application more clear and understandable, the following further details this application in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0029] The technical solutions between the various embodiments of the present invention can be combined with each other, but it must be based on the premise that those of ordinary skill in the art can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.
[0030] To make the above objects, features, and advantages of the present application more obvious and understandable, the following detailed description of specific embodiments of the present application will be provided in conjunction with the accompanying drawings.
[0031] Refer to Figure 1 , a method for motor imagery recognition provided by the present application, the method comprising: S101, acquiring motor imagery EEG data.
[0032] In some embodiments, the motor imagery EEG data of the target user can be acquired through an EEG acquisition device.
[0033] For example, in some embodiments, the paradigm for guiding the user to perform the motor imagery task combines mirror therapy. The user sits upright in front of the mirror table, ensuring that the reflected image of their healthy limb in the mirror can be clearly visually captured. This setting aims to fully utilize the intrinsic mechanism of brain nerve plasticity, activate the key nerve circuits in the brain responsible for motor control and motor perception through the visual feedback pathway, and thus create favorable conditions for the repair and reconstruction of nerve function.
[0034] Among them, the schematic diagram of the motor imagery experimental paradigm can be seen in Figure 2 .
[0035] In a specific example, in order to accurately monitor brain nerve activities, the user (patient) can wear a 16-channel EEG acquisition device, and strictly check the standardization and accuracy of the device wearing before the experiment. During the experiment, the patient performs the motor imagery task in the mirror environment according to the voice and visual prompts provided by the system.
[0036] In some embodiments, to ensure the reliability and effectiveness of the experimental data, this task can be repeated in multiple rounds to enhance the cumulative effect of nerve stimulation and promote the gradual recovery of nerve function.
[0037] In some embodiments, each patient can also be required to perform the experiment twice on two days, with 90 rounds of training in each experiment.
[0038] In some embodiments, it may also include further data preprocessing of the initially acquired motor imagery EEG data.
[0039] Specifically, considering that the EEG data collected by the EEG device contains noise interference, a band-pass filter can be used for filtering, and most of the noise interference can be effectively removed through the filtering algorithm.
[0040] In some embodiments, an 8 - 30 Hz band - pass filter can be selected. Through the filtering algorithm, low - frequency and high - frequency noise interferences can be effectively removed while retaining the frequency band in the EEG signal that has a strong correlation with motor imagery. Among them, the band - pass filter can be a Butterworth filter or other filters, which is not limited here.
[0041] In some embodiments, the filtered EEG signal can also be normalized to eliminate the scale deviation of the data. For example, through the min - max normalization method, the signal can be scaled to a fixed range (such as the range of [0, 1]), while retaining the distribution shape of the original data.
[0042] In some embodiments, the normalization operation is as follows. The multi - channel EEG time - series signal is represented as the following two - dimensional matrix:
[0043] Among them, is the two - dimensional matrix representation of the acquired EEG time - series signal (motor imagery EEG data), C is the number of channels, and L is the sequence length.
[0044] First, normalize each channel. For example, the min - max normalization process can be used to obtain the normalized EEG signal matrix, as shown in the following formula.
[0045] Among them, taking the EEG signal of the channel as an example, the normalization process is as follows:
[0046] Through the above formula, the data is filtered and normalized to ensure the cleanliness of the data and the standardization of the data scale.
[0047] In some embodiments, the method further includes performing enhancement processing on the motor imagery EEG data, and extracting temporal features and temporal time - frequency features based on the enhanced motor imagery EEG data.
[0048] In some embodiments, the enhancement processing of the motor imagery EEG data includes: performing sequence decomposition on all channels of the original motor imagery EEG data to obtain multiple decomposition terms; randomly recombining the decomposition terms to generate the enhanced motor imagery EEG data.
[0049] In some embodiments, the data enhancement method can be a temporal enhancement method based on time - series decomposition and random recombination. The implementation process includes: performing sequence decomposition on all channels of the original EEG data of the same category to obtain multiple decomposition terms, and then randomly recombining these decomposition terms to generate a new EEG sequence. It can be seen that Figure 3 , Figure 3Schematic diagram of the data augmentation method provided in one embodiment of this application. Taking the decomposition of k items as an example:
[0050] Among them, is the same as and represents the EEG data for one training session. ={ }, where i ∈ [1, N] is a set of N preprocessed EEG trials available for training a given class. The signal of each training trial is segmented into k consecutive and non-overlapping segments (A, B, C). Then, from these segments, a new artificial trial can be generated. , is an integer randomly selected from [1, N], representing the segment selected from the th trial. In the example, the parameter is set, which represents the multiple of the increased new data volume to the original data volume, and this value can be set differently according to experimental conditions.
[0051] S102. Based on the spatio-temporal feature extraction module, temporal spatio-temporal features are extracted from the motor imagery EEG data, and the obtained temporal spatio-temporal features, where the spatio-temporal feature extraction module includes a multi-branch spatio-temporal convolution module and a Mamba module.
[0052] In some embodiments, the motor imagery EEG data is input into the spatio-temporal feature extraction module, and the temporal spatio-temporal features of the motor imagery EEG data are obtained through the spatio-temporal feature extraction module.
[0053] In some embodiments, the spatio-temporal feature extraction module includes two sub-modules. The first sub-module includes a multi-branch spatio-temporal convolution module, and the second sub-module includes a Mamba module. Specifically, the obtained motor imagery EEG data can be first input into the first sub-module, and then the feature data output by the first sub-module is input into the second sub-module again, and the temporal spatio-temporal features of the motor imagery EEG data are extracted through the two sub-modules.
[0054] In some embodiments, all channel data of the original EEG data and the data-augmented EEG data can be input into the spatio-temporal feature extraction module for the purpose of expanding the dataset.
[0055] In some embodiments, the temporal spatio-temporal features are extracted from the motor imagery EEG data based on the spatio-temporal feature extraction module, and the obtained temporal spatio-temporal features, where the spatio-temporal feature extraction module includes a multi-branch spatio-temporal convolution module and a Mamba module, include: Based on the multi-branch spatio-temporal convolution module in the spatio-temporal feature extraction module, local features of different scales are extracted from the motor imagery EEG data respectively to obtain multi-scale local temporal spatio-temporal features; Input the multi-scale local temporal-spatial features into the Mamba module in the spatio-temporal feature extraction module. Based on the Mamba module, linearly map and expand the multi-scale local temporal-spatial features, and obtain the global temporal-spatial features through an activation function and discrete state space processing.
[0056] Among them, the multi-branch spatio-temporal convolution module is used to extract local features of the motor imagery data, and the Mamba module is used to extract global features of the motor imagery EEG data. The Mamba model is based on a selective state space model, achieves linear time complexity, can efficiently process time series, and can selectively capture global time dependencies.
[0057] In this way, the local and global features of the motor imagery EEG data can be extracted through the multi-branch spatio-temporal convolution module and the Mamba module in the spatio-temporal feature extraction module.
[0058] In some embodiments, the multi-branch spatio-temporal convolution module includes three convolutional layers with different-sized convolutional kernels and different strides. Among them, different convolutional kernels are used to extract spatio-temporal local features in different-scale spaces.
[0059] In the above embodiments, through different-sized convolutional kernels and variable strides, multi-scale parallel capture of local features is achieved, avoiding feature loss that may be caused by a single kernel size, and improving the integrity of local features.
[0060] In the above embodiments, refer to Figure 7 As shown in the schematic diagram of the dual Mamba structure provided by this application, in the spatio-temporal feature extraction module, all channel data of the original EEG data and the augmented EEG data are input into the spatio-temporal feature extraction module to expand the dataset size. The spatio-temporal feature extraction module consists of two sub-modules: the first sub-module is the multi-branch spatio-temporal convolution module, and different kernel sizes are set in the convolutional layers of each branch to capture local features of different scales, while keeping the input and output data dimensions the same. In a specific example, a 3-branch spatio-temporal convolution can be designed, with kernel sizes of (1, 11), (3, 21), and (5, 41) respectively, and the padding method is set to same padding.
[0061] In some embodiments, the setting of the kernel size is obtained through a large number of experimental comparative analyses. Small convolutional kernels can extract local fine-grained features, and large convolutional kernels can extract local coarse-grained features. The 16×1000 single-trial EEG data The two-dimensional matrix is horizontally concatenated after passing through the 3-branch spatio-temporal convolution kernel to obtain a 48×1000 two-dimensional matrix .
[0062] The second sub-module is the Mamba module, which is a new architecture based on the state space model. It introduces a selective mechanism to construct a structured state space sequence (S4) model, which can identify information like the attention mechanism and is particularly remarkable in time series prediction. The structured state space sequence (S4) model is inspired by continuous systems and maps a one-dimensional function or sequence through a hidden state to .
[0063]
[0064] Specifically, the S4 model uses A as the state matrix, and B and C as the input and output matrices. The state equation represented by the S4 model is as follows:
[0065] The model captures global features through the above expressions, where is the hidden state, containing information about the entire sequence. For the discrete case, the parameters A , B are converted to discrete parameters , through the step size Δ. The discretized form of the state equation is expressed as:
[0066] where , and .
[0067] Finally, the model calculates the output through global convolution for efficient parallelizable training. The structured convolution kernel can be expressed as:
[0068] where W represents the length of the input x sequence. The Mamba module introduces a data-dependent selection mechanism, which extends S4 to S6. It is calculated through a scanning process and integrates a hardware-aware parallel algorithm into its loop pattern. In S6, the parameters can be dynamically adjusted according to the input elements currently being processed, enabling the model to respond differently to different inputs.
[0069] As the second sub-module, the overall process of the Mamba module is as follows:
[0070] In some embodiments, the spatio-temporal feature extraction module extracts temporal features from the motor imagery EEG data to obtain temporal spatio-temporal features, where the spatio-temporal feature extraction module includes a multi-branch spatio-temporal convolution module and the Mamba module, including: The temporal and spatial feature vectors output by the multi-branch spatio-temporal convolution module are extended in hidden dimension through a linear mapping to obtain extended data x and res ; Based on a convolution function and an activation function, the extended data is processed to obtain intermediate data , res and the intermediate data z is obtained through the SiLU activation function; Based on a discrete state space model and the intermediate data the state data y is determined; Based on a linear mapping, the state data y and the intermediate data z are combined to obtain combined data ; The combined data is combined with the input residual to obtain the temporal and spatial features .
[0071] In the example, the structural diagram of the Mamba module is as shown in Figure 4 . The output of the first sub-module is extended in hidden dimension through a linear mapping to obtain x and res , and is obtained through a convolution function and the SiLU activation function . The state y is generated by a discrete state space model (SSM). res z is obtained through the SiLU activation function. Then y and z are combined through a linear mapping to obtain . Finally, is combined with the input residual to obtain the output .
[0072] S103. Based on the time-frequency feature extraction module, temporal and time-frequency features are extracted from the motor imagery EEG data, and the temporal and time-frequency features are obtained, where the time-frequency feature extraction module includes a multi-branch time-frequency convolution module and a Mamba module.
[0073] In some embodiments, the motor imagery EEG data is input into the time-frequency feature extraction module, and the temporal and time-frequency features of the motor imagery EEG data are obtained through the time-frequency feature extraction module.
[0074] In some embodiments, the time-frequency feature extraction module includes two sub-modules. The first sub-module includes a multi-branch time-frequency convolution module, and the second sub-module includes a Mamba module. Specifically, the obtained motor imagery EEG data can be first input into the first sub-module, and then the feature data output by the first sub-module is input into the second sub-module again, and the temporal and time-frequency features of the motor imagery EEG data are extracted through the two sub-modules.
[0075] In some embodiments, all channel data of the original EEG data and the data-augmented EEG data can be input into the time-frequency feature extraction module, aiming to expand the dataset.
[0076] In some embodiments, the time-frequency feature extraction module extracts temporal time-frequency features from the motor imagery EEG data to obtain temporal time-frequency features, where the time-frequency feature extraction module includes a multi-branch time-frequency convolution module and a Mamba module, including: Obtain the frequency-domain features corresponding to the motor imagery EEG data; Based on the spatio-temporal convolution modules with multiple branches of different-sized convolutional kernels in the time-frequency feature extraction module, extract different-scale local features from the frequency-domain features of the motor imagery EEG data to obtain multi-scale local temporal time-frequency features; Input the multi-scale local temporal time-frequency features into the Mamba module in the time-frequency feature extraction module, and based on the Mamba module, perform linear mapping expansion on the multi-scale local temporal time-frequency features, and obtain global temporal time-frequency features through an activation function and discrete state space processing.
[0077] Among them, the multi-branch time-frequency convolution module is used to extract local features of the motor imagery data, The Mamba module is used to extract global features of the motor imagery EEG data. The Mamba model is based on a selective state space model, achieves linear time complexity, can efficiently process time series, and can selectively capture global time dependencies.
[0078] In this way, the local and global features of the motor imagery EEG data can be extracted through the multi-branch time-frequency convolution module and the Mamba module in the time-frequency feature extraction module.
[0079] Compared with the problem of insufficient utilization of frequency-domain information when only focusing on the temporal spatio-temporal features of EEG signals, in the embodiments of the present application, the temporal data is converted to the frequency domain through FFT, the CNN extracts local frequency-domain features, and the Mamba models the long-term dependencies of different frequency components, which can solve the problem of difficult global modeling of the frequency domain of non-stationary signals.
[0080] In some embodiments, the multi-branch time-frequency convolution module includes two convolutional layers with different-sized convolutional kernels and different strides, where the different convolutional kernels are used to extract time-frequency local features in different-scale spaces.
[0081] In the above embodiments, through different-sized convolutional kernels and variable strides, multi-scale parallel capture of local features is achieved, avoiding feature loss that may be caused by a single kernel size, and improving the integrity of local features.
[0082] In the above embodiment, the obtained motor imagery EEG data is, for example, single-trial EEG data of 16×1000 The two-dimensional matrix is first transformed to the frequency domain through Fourier transform (FFT). The FFT transform can be expressed as:
[0083] In this example, only the unilateral spectrum is taken, retaining the zero frequency and positive frequencies. At this time The effective length is 501. Then it enters the time-frequency feature extraction module, which consists of two sub-modules: In the above embodiment, referring to Figure 7 which is the schematic diagram of the double Mamba structure provided by this application. In the time-frequency feature extraction module, the first sub-module consists of two convolutional layers with different kernel sizes, which can extract the temporal time-frequency features in a small range; the second sub-module consists of Mamba modules, whose function is to improve the model's dependence on global time-frequency.
[0084] In the example, the size of the first convolutional kernel in the first sub-module is (1, 1), which can independently achieve information fusion between channels; the size of the second convolutional kernel is (3, 3) to achieve local feature fusion across channels. To maintain a longer sequence length, the same padding method is selected for both. After being processed by the first sub-module, the output two-dimensional matrix has a dimension of 16×501. Then it is sent into the Mamba module, and using its characteristic of selectively extracting important features, global features are captured. To facilitate the fusion of temporal features, an inverse Fourier transform needs to be performed on the temporal time-frequency features:
[0085] Before performing the inverse Fourier transform, a zero-padding operation needs to be performed on the two-dimensional matrix to expand the sequence length to its original length, obtaining the temporal time-frequency features .
[0086] In some embodiments, the time-frequency feature extraction module extracts temporal time-frequency features from the motor imagery EEG data to obtain temporal time-frequency features, where the time-frequency feature extraction module includes a multi-branch time-frequency convolution module and Mamba modules, including: Expanding the hidden dimension of the time-frequency feature vector output by the multi-branch time-frequency convolution module through a linear mapping to obtain extended data; Processing the extended data based on a convolution function and an activation function to obtain intermediate data; Based on the discrete state space model and the intermediate data Determining state data; Combining the state data and the intermediate data based on a linear mapping to obtain combined data; Combining the combined data with the input residual to obtain temporal time-frequency features.
[0087] The schematic structural diagram of the Mamba module is shown in Figure 4 .
[0088] In some embodiments, the spatio-temporal feature extraction module and the time-frequency feature extraction module are two parallel feature extraction modules. The two parallel feature extraction modules are respectively and independently used to obtain the motor imagery EEG data, and respectively perform feature extraction on the motor imagery EEG data to obtain the sequential spatio-temporal features and sequential time-frequency features corresponding to the motor imagery EEG data.
[0089] The spatio-temporal feature extraction module and the time-frequency feature extraction module are two parallel feature extraction modules. By inputting the motor imagery EEG data in parallel to the two feature extraction modules, the sequential spatio-temporal features and sequential time-frequency features of the motor imagery EEG data can be respectively extracted.
[0090] S104, fuse the sequential spatio-temporal features and the sequential time-frequency features to obtain fused features, and classify and identify the motor imagery EEG data based on the fused features.
[0091] In some embodiments, the fusing the sequential spatio-temporal features and the sequential time-frequency features to obtain fused features, and classifying and identifying the motor imagery EEG data based on the fused features includes: Convert the sequential time-frequency features to sequential spatio-temporal features, and fuse the data of the two sequential spatio-temporal features through a convolutional layer to obtain fused features; classify and identify the motor imagery EEG data based on the fused features and a fully connected layer.
[0092] In the above embodiments, the fusion module integrates the global dynamic features extracted by spatio-temporal convolution and the global frequency features extracted by frequency-domain CNN to form a cross-dimensional fusion, which can cover all the information of the signal and improve the generalization ability of the model.
[0093] Specifically, in the above embodiments, a convolutional layer is selected to automatically fuse and classify the sequential spatio-temporal features and the sequential time-frequency features. The sequential spatio-temporal features have a two-dimensional matrix size of 16×1000, and the sequential time-frequency features have a two-dimensional matrix size of 16×1000. The two features are first linearly convolved in a convolutional kernel of (1, 1), and then matrix multiplication is performed to obtain a 16×16 two-dimensional matrix, which fuses the spatio-temporal and time-frequency features on the channels. After that, linear convolution is performed through a convolutional layer of (1, 1). Finally, the extracted features are classified and predicted through a fully connected layer.
[0094] In some embodiments, in order to ensure the robustness and generality of the feature extraction process, a domain alignment method is adopted in this example to construct domain-invariant representations of the training data and the test data.
[0095] In some embodiments, the method further includes: Obtaining the motor imagery EEG data corresponding to multiple users, and aligning the motor imagery EEG data corresponding to multiple users through a domain alignment method; Performing temporal-spatial feature extraction on the motor imagery EEG data after domain alignment of multiple users based on a spatio-temporal feature extraction module to obtain temporal-spatial features, wherein the spatio-temporal feature extraction module includes a multi-branch spatio-temporal convolution module and a Mamba module; Performing temporal-frequency feature extraction on the motor imagery EEG data after domain alignment of multiple users based on a time-frequency feature extraction module to obtain temporal-frequency features, wherein the time-frequency feature extraction module includes a multi-branch time-frequency convolution module and a Mamba module; Fusing the temporal-spatial features and the temporal-frequency features corresponding to multiple users to obtain fused features, and performing classification and recognition on the motor imagery EEG data based on the fused features.
[0096] Because the EEG data of each person is specific, training a model for each patient will increase a lot of workload. A general model can be trained first and aligned with the general model through a domain alignment method.
[0097] Through the above embodiments, a general motor imagery recognition model can be constructed based on the data of multiple users, thereby improving the generality of the model. To avoid the inconsistency of data between different users, it also includes constructing a test set and a training set through a domain alignment method.
[0098] In some embodiments, it may further include obtaining the motor imagery EEG data corresponding to a single target user, aligning the motor imagery EEG data of the target user with the data in the general motor imagery recognition model through a domain alignment method, and then respectively identifying the data of the target user based on the trained general motor imagery recognition model.
[0099] The implementation steps of domain alignment are as follows:
[0100] Wherein, represents the EEG data of a single trial, R represents the mean covariance matrix of N trials, is the EEG data after domain alignment. Through domain alignment, the distribution difference between the training data set and the test data set can be effectively reduced, which helps to achieve more accurate and reliable feature extraction.
[0101] In the above embodiments, the present application provides a motor imagery recognition solution. By separately extracting temporal-spatial features and temporal-frequency features from the time series, the feature dimension is increased, and a dual Mamba structure is used to fully extract the global features in the time series and frequency domain. Then, a fusion module is used to fuse the dual features. This cross spatio-temporal-frequency feature fusion strategy can effectively improve the classification accuracy of the MI task.
[0102] In some embodiments, the present embodiment provides a motor imagery recognition method based on a dual Mamba spatio-temporal and frequency feature fusion network. Refer to Figure 5 , the method includes the following steps: S501, Obtain the original electroencephalogram (EEG) data of motor imagery; refer to Figure 2 for the motor imagery experimental paradigm. Obtain the original EEG signal of motor imagery, incorporate mirror therapy to activate mirror neurons and improve the patient's motor imagery.
[0103] S502, Perform data augmentation on the original EEG data to expand the training samples required by the deep network model. Perform data augmentation on the original EEG data and use a temporal augmentation method based on the decomposition and random recombination of the time series for data expansion.
[0104] S503, Temporal feature extraction. Input the augmented EEG data into the spatio-temporal feature extraction module to extract temporal-spatial features.
[0105] S504, Temporal-frequency feature extraction. Input the augmented EEG data into the temporal-frequency feature extraction module to extract temporal-frequency features.
[0106] S505, Fuse the temporal-spatial features and temporal-frequency features to obtain the required multi-dimensional EEG features; finally, send the multi-dimensional features into the classifier to obtain the motor imagery classification result. Input the temporal-spatial features and temporal-frequency features output after parallel processing by the spatio-temporal feature extraction module and the temporal-frequency feature extraction module into the fusion module for effective feature fusion and classification.
[0107] During the collection of EEG, it is extremely vulnerable to interference from physiological noises such as electromyogram (EMG) and electrooculogram (EOG). The inclusion of these noises will not only reduce the signal quality of EEG, but may also lead to a significant reduction in the classification accuracy in the motor imagery task. In addition, EEG signals have strong non-linear characteristics, and their temporal-frequency features show large differences due to different inter-individual neural activities. Traditional linear methods often have difficulty effectively capturing these complex characteristics, which may affect the classification performance of the model. Therefore, how to effectively remove noise and improve the classification accuracy of EEG-based motor imagery has become an important technical problem in the BCI field.
[0108] To solve the above problems, a method based on a deep neural network can be adopted to automatically extract features and classify EEG signals. Compared with traditional signal processing and pattern recognition methods, deep learning models can extract high-level features from the original EEG signals. For example, models such as convolutional neural networks (CNNs) perform well in extracting spatio-temporal features and can effectively capture the temporal and frequency features in EEG signals. However, due to the limitation of the receptive field size, CNNs can only effectively capture local features in the input data, have a weak ability to model global semantic relationships, and cannot accurately capture global features. If the receptive field is enlarged, problems such as gradient disappearance or explosion may occur, further weakening the transmission efficiency of global information. The emergence of recurrent neural networks (RNNs) makes up for the defect that CNNs are difficult to capture global features to a certain extent. By using the context dependence of time series to model long-distance dependencies, but RNNs have problems such as forgetting historical information and slow training.
[0109] The motion imagination recognition method based on a dual-Mamba spatio-temporal and time-frequency feature fusion network provided in the above embodiments has the following specific beneficial effects: Adopt a multi-layer spatio-temporal convolution structure, where the convolution kernel sizes and strides of each layer are different to efficiently extract local features. On this basis, introduce the Mamba model to accurately capture global time features from multi-layer local features and improve the ability to mine global information in the time dimension of time series data.
[0110] Use the fast Fourier transform (FFT) to convert time series data to the frequency domain, and use a convolutional neural network (CNN) to extract local features in the frequency domain. Subsequently, with the help of the Mamba model, obtain global spatio-temporal and time-frequency features from the spatio-temporal and time-frequency features, and finally restore the data from the frequency domain to the time domain through the inverse Fourier transform to achieve in-depth mining and integration of time-frequency domain information.
[0111] Construct a fusion module to organically fuse the global features captured by the above two methods, and retain the key information in the two types of features to the greatest extent. The fused features are input into a classifier for classification, effectively improving the accuracy and reliability of the classification task.
[0112] To extract global features, combined with the Mamba model, it can effectively capture global time features and is more efficient than traditional Transformers when dealing with sequence problems.
[0113] The present invention constructs a motor imagery classification model based on the Mamba model. The Mamba model integrates the advantageous features of CNN and RNN. It can not only comprehensively and fully extract the global features in the data, but also has the ability of concurrent training, thus significantly improving the training efficiency. This innovative method provides a new way to solve the inherent problems of traditional models and is expected to achieve more excellent application results in related fields.
[0114] This application also provides a motor imagery classification and recognition device 600, and the device 600 includes: A data acquisition module 601, configured to acquire motor imagery EEG data; A spatio-temporal feature extraction module 602, configured to perform spatio-temporal feature extraction on the motor imagery EEG data based on the spatio-temporal feature extraction module to obtain spatio-temporal features, where the spatio-temporal feature extraction module includes a multi-branch spatio-temporal convolution module and a Mamba module; A time-frequency feature extraction module 603, configured to perform sequential time-frequency feature extraction on the motor imagery EEG data based on the time-frequency feature extraction module to obtain sequential time-frequency features, where the time-frequency feature extraction includes a multi-branch time-frequency convolution module and a Mamba module; A fusion module 604, configured to fuse the sequential spatio-temporal features and the sequential time-frequency features to obtain fusion features, and perform classification and recognition on the motor imagery EEG data based on the fusion features.
[0115] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the motor imagery classification and recognition method provided in any one of the above embodiments are implemented.
[0116] A computer program product has a program or instruction stored thereon. When the program or instruction is executed by a processor, the steps of the motor imagery classification and recognition method provided in any one of the above embodiments are implemented.
[0117] In some embodiments, the method further includes monitoring the EEG data of a user through a non-contact EEG device. For example, the target user can be an infant, a vegetative person who cannot control their body, etc. The non-contact EEG acquisition device is not directly worn on the head of the target user, which can make the data acquisition more flexible. And when the head pose of the target user changes, the pose of the corresponding EEG acquisition device can also be adjusted adaptively.
[0118] In this way, for infants who need to perform motor imagery classification and recognition, the real-time acquisition of EEG data can also be achieved through a non-contact EEG acquisition device, and then the motor imagery classification and recognition can be performed on the acquired EEG data.
[0119] Adopt a non-contact electroencephalogram (EEG) signal acquisition method. During the EEG signal acquisition process, the EEG electrodes are not directly attached to the scalp, and the EEG signals of the monitored brain region are accurately monitored by adjusting the pose of the EEG electrodes.
[0120] In some embodiments, the steps of monitoring the EEG data of a user through a non-contact EEG device include: Step 1, obtaining a head image corresponding to the target user at the current moment.
[0121] In some embodiments, the target user may include infants or other users who need EEG monitoring.
[0122] In some embodiments, the head image at the current moment may be obtained by real-time acquisition through a depth sensor or other image acquisition devices.
[0123] In some embodiments, the image acquisition device can be fixed on a bracket. The position of the bracket is immovable, and the pose of the image acquisition device can be adjusted.
[0124] Step 2, determining the head pose of the target user corresponding to the current moment based on the head image corresponding to the current moment.
[0125] Among them, the pose includes position and direction.
[0126] In some embodiments, after obtaining the head image of the target user corresponding to the current moment through the image acquisition device, it further includes performing pose analysis on the head image to determine the head position and head direction of the target user at the current moment.
[0127] Step 3, determining the pose of the EEG cap based on the head pose at the current moment and the head pose corresponding to the previous moment, where the EEG cap includes a plurality of EEG electrodes, and the plurality of EEG electrodes monitor the EEG data of the target user in the pose corresponding to the EEG cap.
[0128] Among them, the EEG cap is similar to a helmet, and the EEG electrodes can be arranged inside the helmet, and each EEG electrode can move horizontally and vertically.
[0129] In some embodiments, the previous moment is earlier than the current moment, and the head pose corresponding to the previous moment is the historical head position and direction calculated based on the head image acquired at the previous moment.
[0130] In some embodiments, the pose of the EEG cap can be determined based on the head pose at the current moment and the previous N historical head poses. The EEG cap includes a plurality of EEG electrodes, and the plurality of EEG electrodes are used to monitor the EEG data of the target user. N is an integer greater than or equal to 1.
[0131] In some embodiments, the target user may include infants and young children. When collecting brain data of infants and young children, the heads of infants and young children are prone to uncontrolled movement, making it easier for the relative position between the brain electrodes and the scalp of infants and young children to shift, resulting in unstable signal acquisition positions. Therefore, during the EEG monitoring process, it is necessary to continuously obtain the head images of infants and young children, determine the head position and head direction of infants and young children at the current moment based on the obtained head images corresponding to the current moment, compare the head pose at the current moment with the historical poses at the previous one or more moments, so as to evaluate the amplitude of the change in the head pose of infants and young children, and further determine whether it is necessary to adjust the position and direction of the brain electrodes to ensure the accuracy of monitoring the brain data of infants and young children.
[0132] In the above embodiments, when the target user is a newborn or an infant, in order to avoid irritation and damage to the skin of newborns and infants, a non-contact EEG monitoring method can be used. The non-contact EEG monitoring method makes the EEG cap not directly contact the head of the infant. However, when the head of the infant moves, it is easy to cause the brain electrodes to deviate from the EEG monitoring area, resulting in inaccurate EEG data monitoring.
[0133] The EEG monitoring method provided in this application can calculate the similarity between the current pose and the previous N historical poses based on the head pose at the current moment and the previous N historical head poses, so as to determine whether it is necessary to adjust the EEG cap and how to adjust the pose of the EEG cap to ensure that the brain electrodes can continuously and accurately monitor the monitoring brain area, thereby continuously and stably monitoring the EEG signals of the target user.
[0134] In some embodiments, determining the pose of the EEG cap based on the head pose at the current moment and the head pose corresponding to the previous moment includes: Step 1, determining the distance between the head pose at the current moment and the head pose at the previous moment.
[0135] Step 2, in response to the distance being greater than a preset threshold, adjusting the current pose of the EEG cap based on the head pose at the current moment, and the brain electrodes in the EEG cap continue to monitor the EEG data of the target user in the adjusted pose.
[0136] Step 3, in response to the distance not being greater than the preset threshold, the brain electrodes in the EEG cap continue to monitor the EEG data of the target user in the current pose.
[0137] In some embodiments, the pose of the EEG cap can be determined based on the head pose at the current moment and the previous N historical head poses, including: determining the distance between the head pose at the current moment and the previous N head poses. Based on the distance being greater than a preset threshold, it indicates that the head pose of the target user has changed significantly, and it is necessary to synchronously adjust the pose of the EEG cap. Therefore, the pose of the EEG cap can be adjusted based on the head pose at the current moment, and the head EEG data of the target user can be continuously monitored based on the adjusted pose of the EEG cap, so that the pose of the EEG cap is always consistent with the head pose of the target user.
[0138] In the above embodiments, the EEG of a specific area of the target user's head is monitored through the EEG electrodes in the EEG cap. A preset deviation threshold of the EEG electrodes is set, and the distance between the head pose at the current moment and the previous N head poses at the previous moment is calculated in real time. If the distance is greater than the preset threshold, the pose of the EEG cap is adjusted according to the previous pose, so that the EEG electrodes in the EEG cap can continuously and accurately monitor a specific area of the target user's head. If the distance is less than or equal to the preset threshold, it indicates that the pose of the target user is the same or basically the same between the current moment and the previous moment. At the current moment, there is no need to adjust the pose of the EEG cap or the EEG electrodes. Therefore, at the current moment, the EEG monitoring can be continued based on the pose of the EEG cap at the previous moment. At the next moment, the head pose of the target user can be obtained through the sensor again, and the distance between the head pose at the current moment and the previous N head poses is repeatedly determined, and the size of the distance at this time and the preset threshold are judged, and the determination result decides whether to adjust the pose of the EEG.
[0139] It can be understood that the EEG electrodes are arranged in the EEG cap, and adjusting the pose of the EEG cap also realizes the adjustment of the pose of the EEG electrodes.
[0140] In some embodiments, the adjusting the current pose of the EEG cap based on the head pose at the current moment includes: taking the head pose of the target user at the current moment as the control target of the robotic arm, and the robotic arm controls and adjusts the current position and current direction of the EEG cap, so that the EEG cap moves to a pose matching the head position and head direction of the target user at the current moment, and the EEG electrodes in the EEG cap collect the EEG data of the target user at the adjusted pose.
[0141] Specifically, the EEG cap can be arranged at the end of the robotic arm, and the pose of the EEG cap is controlled by the robotic arm. When it is detected that the head pose of the target user changes significantly at the current moment, the pose of the EEG cap is synchronously adjusted by the robotic arm, so that the pose of the EEG cap can be adapted to and matched with the head pose of the target user at the current moment, and the head EEG data of the target user can be continuously and accurately detected.
[0142] In some embodiments, the step of determining the head pose of the target user includes: determining the set of head edge point coordinates of the target user based on the head image; determining the head pose of the target user based on the set of head edge point coordinates.
[0143] In some embodiments, the acquired head image can be segmented based on an image segmentation algorithm to remove irrelevant information and obtain a target image. Then, edge detection is performed on the target image through an edge detection algorithm to obtain the set of edge point coordinates of the head image, and the head pose of the target user is determined based on the set of head edge point coordinates.
[0144] It can be understood that the edge detection algorithm and the image segmentation algorithm can be existing mature algorithms, which are not limited here.
[0145] In some embodiments, the method further includes: determining the head size of the target user based on the head image of the target user; determining the scaling coefficient of the EEG cap and the scaling coefficient of the arrangement of the brain electrodes in the EEG cap based on the head size of the target user; adjusting the size of the EEG cap based on the scaling coefficient, and adjusting the coordinate position layout of the brain electrodes in the EEG cap based on the arrangement scaling coefficient.
[0146] It can be understood that the head sizes of different target users are different, so the requirements for the size of the EEG cap by different target users are also different. It is very important to set an EEG cap that matches the head size of the corresponding target user, which is the basic condition for obtaining accurate monitoring data.
[0147] In some embodiments, it further includes determining the head size of the target user based on the head image of the target user, and then adjusting the size of the EEG cap adaptively based on the head size of the target user, so that the size of the EEG cap matches the head size of the target user.
[0148] The brain electrodes are arranged in the EEG cap, and the positions of the brain electrodes in the EEG cap can change. When the size of the EEG cap changes, the positions of the brain electrodes in the EEG cap will also change adaptively.
[0149] In some embodiments, the relative positional relationship of the brain electrodes in the EEG cap is fixed. The initial layout coordinates of the brain electrodes in the EEG cap can be obtained first, and the spacing between the brain electrodes is scaled based on the scaling coefficient of the EEG cap to obtain the target coordinates of the brain electrodes.
[0150] In this way, the head size of the target user is determined based on the head image of the target user, and then the size of the EEG cap and the positions of the brain electrodes in the EEG cap are adjusted adaptively, so that the finally adjusted EEG cap and brain electrodes match the target user, in order to monitor more accurate brain data.
[0151] In some embodiments, the head size of the target user includes a head transverse size and a head longitudinal size; The step of determining the head size of the target user includes: determining the ear key points, nose key points, and head key points of the target user based on the head image, where the ear key points include a left ear key point and a right ear key point; determining the head transverse size of the target user based on the ear key points and the head key points; determining the head longitudinal size of the target user based on the nose key points and the head key points.
[0152] In some embodiments, the head size can be determined based on the head key points. For example, after obtaining a head image based on a sensor, a trained key point detection model is used to output the coordinates of the head key points, nose key points, left ear key points, and right ear key points.
[0153] By fitting the top-of-head curve through the coordinates of the left ear key point concave and the right ear key point and the coordinates of the head edge point set, and calculating the curve length as the head transverse size.
[0154] By calculating the distance between two points along the head surface through the nose key point and the midpoint of the top-of-head curve, and multiplying by 2, the head longitudinal size can be obtained.
[0155] Determine the head size of the target user based on the head transverse size and head longitudinal size of the target user, so as to determine the size of the EEG cap and the coordinate arrangement of the brain electrodes according to the head size.
[0156] In some embodiments, the method further includes obtaining a head image of the target user based on a depth sensor; The step of determining the shooting pose of the depth sensor includes: obtaining a working plane image corresponding to the working plane and the plane normal vector corresponding to the working plane based on the depth sensor, where the working plane includes the plane corresponding to the working platform for monitoring the EEG data of the target user; determining the shooting position and shooting direction of the depth sensor based on the on-site observation distance, the working distance of the depth sensor, and the plane normal vector.
[0157] In some embodiments, a head image of the target user can be captured by a depth sensor, and then the head pose and head size of the target user can be calculated from the head image.
[0158] In some embodiments, the pose setting of the sensor is very important. If the pose setting of the sensor is unreasonable, then accurate head pose and head size data cannot be obtained from the captured head image, which may lead to unreasonable brain electrode monitoring poses and inaccurate brain data.
[0159] In some embodiments, it further includes setting the shooting pose of the sensor. First, the working plane image corresponding to the working plane is collected by the sensor, and the shooting pose of the sensor is determined based on the working plane image.
[0160] In some embodiments, after obtaining the working plane image corresponding to the working plane, the method further includes: extracting a plurality of internal coordinate point sets from the working plane image based on a preset marker template; converting the coordinate systems of the plurality of internal coordinate point sets into a world coordinate system, and determining the plane normal vector corresponding to the working plane based on the internal coordinate point sets in the world coordinate system.
[0161] Specifically, the target user to be examined lies flat on the working platform, and the working platform is photographed by an RGBD sensor to obtain a working plane image. Using a 2D edge detection / image segmentation algorithm and combining a marker template with a known shape and size, the internal point coordinates of n markers on the working plane and the corresponding depth information are obtained. In some embodiments, the coordinates of these points can be recorded as C_POINTS (camera_set_1, camera_set_2,…, camera_set_n). Then, the 3D coordinates in the camera coordinate system are converted to the world coordinate system W_POINTS (world_set_1, world_set_2,…, world_set_n) of the bracket base using the external parameters.
[0162] Using the coordinate information of all points in W_POINTS, plane fitting is performed to obtain the plane normal vector N = [nx, ny, nz]. According to the working distance and field of view of the sensor, the optimal observation distance d is set, and the center position P_c = [x_c, y_c, d_c] of the working plane is calculated based on the marker point set. Then the sensor shooting position can be P1 = P_c + N * d, and the direction is perpendicular to the working plane where the working platform is located. Furthermore, the target pose of the sensor is determined based on the determined position and direction, and the sensor is adjusted to this target pose by the robotic arm to obtain the head data of the target user.
[0163] The present application also provides a non-contact electroencephalogram monitoring system, which includes: a bracket module, a robotic arm module, an image acquisition module, a work platform module, an electroencephalogram cap module, and a control module; the image acquisition module is installed on the bracket module and is used to acquire the head image of the target user and / or the work platform image corresponding to the work platform module; the electroencephalogram cap module is installed at the end of the robotic arm, the robotic arm is used to adjust the position and pose of the electroencephalogram cap, and the brain electrodes in the electroencephalogram cap are used to monitor the electroencephalogram data of the target user; the work platform module is used to carry the target user; the control module is used to control the robotic arm to adjust the position and pose of the electroencephalogram cap based on the head position and pose of the target user, and the brain electrodes in the electroencephalogram cap are used to monitor the electroencephalogram data of the target user.
[0164] It can be understood that the computer device provided by the present application can be a server, and its internal structure diagram can be as Figure 8 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store relevant data. The network interface of the computer device is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements the method provided by the present application.
[0165] Those skilled in the art can understand that Figure 8The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements. Those of ordinary skill in the art can understand that all or part of the process of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application may include at least one of non-volatile and volatile memories. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0166] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0167] The above-described embodiments only represent several implementation manners of this application, and their descriptions are relatively specific and detailed, but they should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of the patent of this application should be subject to the appended claims.
Claims
1. A method for classifying and recognizing motor imagery, characterized in that, The method includes: Obtaining motor imagery EEG data; Performing temporal and spatial feature extraction on the motor imagery EEG data based on a spatio-temporal feature extraction module to obtain temporal and spatial features, where the spatio-temporal feature extraction module includes a multi-branch spatio-temporal convolution module and a Mamba module; Performing temporal and frequency feature extraction on the motor imagery EEG data based on a time-frequency feature extraction module to obtain temporal and frequency features, where the time-frequency feature extraction module includes a multi-branch time-frequency convolution module and a Mamba module; Fusing the temporal and spatial features and the temporal and frequency features to obtain fused features, and classifying and recognizing the motor imagery EEG data based on the fused features.
2. The method according to claim 1, wherein The spatio-temporal feature extraction module and the time-frequency feature extraction module are two parallel feature extraction modules. The two parallel feature extraction modules are respectively and independently used to obtain the motor imagery EEG data, and respectively perform feature extraction on the motor imagery EEG data to obtain the corresponding temporal and spatial features and temporal and frequency features of the motor imagery EEG data.
3. The method according to claim 1, wherein The performing temporal feature extraction on the motor imagery EEG data based on the spatio-temporal feature extraction module to obtain temporal and spatial features, where the spatio-temporal feature extraction module includes a multi-branch spatio-temporal convolution module and a Mamba module, includes: Performing extraction of local features of different scales on the motor imagery EEG data respectively based on the multi-branch spatio-temporal convolution module in the spatio-temporal feature extraction module to obtain multi-scale local temporal and spatial features; Inputting the multi-scale local temporal and spatial features into the Mamba module in the spatio-temporal feature extraction module, and performing linear mapping expansion on the multi-scale local temporal and spatial features based on the Mamba module, and obtaining global temporal and spatial features through an activation function and discrete state space processing.
4. The method according to claim 1, wherein The performing temporal and frequency feature extraction on the motor imagery EEG data based on the time-frequency feature extraction module to obtain temporal and frequency features, where the time-frequency feature extraction module includes a multi-branch time-frequency convolution module and a Mamba module, includes: Obtaining the frequency domain features corresponding to the motor imagery EEG data; Performing extraction of local features of different scales on the frequency domain features of the motor imagery EEG data respectively based on the spatio-temporal convolution module with multi-branch different-sized convolutional kernels in the time-frequency feature extraction module to obtain multi-scale local temporal and frequency features; Inputting the multi-scale local temporal and frequency features into the Mamba module in the time-frequency feature extraction module, and performing linear mapping expansion on the multi-scale local temporal and frequency features based on the Mamba module, and obtaining global temporal and frequency features through an activation function and discrete state space processing.
5. The method according to claim 1, wherein The fusing the temporal and spatial features and the temporal and frequency features to obtain fused features, and classifying and recognizing the motor imagery EEG data based on the fused features, includes: Converting the temporal and frequency features to temporal and spatial features, and fusing the data of the two temporal and spatial features through a convolutional layer to obtain fused features; Classifying and recognizing the motor imagery EEG data based on the fused features and a fully connected layer.
6. The method according to claim 1, characterized in that, The method further includes enhancing the motor imagery EEG data, and extracting temporal-spatial features and temporal-frequency features based on the enhanced motor imagery EEG data; The enhancing of the motor imagery EEG data includes: Performing sequence decomposition on all channels of the original motor imagery EEG data to obtain a plurality of decomposition terms; Randomly recombining the decomposition terms to generate the enhanced motor imagery EEG data.
7. The method according to claim 1, characterized in that, The multi-branch spatio-temporal convolution module includes three convolutional kernels with different sizes and convolutional layers with different strides, wherein different convolutional kernels are used to extract spatio-temporal local features in different scale spaces; The multi-branch time-frequency convolution module includes two convolutional kernels with different sizes and convolutional layers with different strides, wherein different convolutional kernels are used to extract time-frequency local features in different scale spaces.
8. The method according to claim 1, characterized in that, The extracting of temporal features from the motor imagery EEG data based on the spatio-temporal feature extraction module to obtain temporal-spatial features, wherein the spatio-temporal feature extraction module includes a multi-branch spatio-temporal convolution module and a Mamba module, includes: Expanding the hidden dimension of the temporal-spatial feature vector output by the multi-branch spatio-temporal convolution module through linear mapping to obtain extended data; Processing the extended data based on a convolution function and an activation function to obtain intermediate data; Determining state data based on the intermediate data and a discrete state space model; Combining the state data and the intermediate data based on linear mapping to obtain combined data; Combining the combined data with the input residual to obtain temporal-spatial features.
9. The method according to claim 1, wherein The extracting of temporal-frequency features from the motor imagery EEG data based on the time-frequency feature extraction module to obtain temporal-frequency features, wherein the time-frequency feature extraction module includes a multi-branch time-frequency convolution module and a Mamba module, includes: Expanding the hidden dimension of the time-frequency feature vector output by the multi-branch time-frequency convolution module through linear mapping to obtain extended data; Processing the extended data based on a convolution function and an activation function to obtain intermediate data; Determining state data based on the intermediate data and a discrete state space model; Combining the state data and the intermediate data based on linear mapping to obtain combined data; Combining the combined data with the input residual to obtain temporal-frequency features.
10. The method according to claim 1, wherein The method further includes: Obtaining motor imagery EEG data corresponding to multiple users, and aligning the motor imagery EEG data corresponding to multiple users through a domain alignment method; Extracting temporal features from the motor imagery EEG data aligned in multiple user domains based on the spatio-temporal feature extraction module to obtain temporal features, wherein the spatio-temporal feature extraction module includes a multi-branch spatio-temporal convolution module and a Mamba module; Extracting temporal-frequency features from the motor imagery EEG data aligned in multiple user domains based on the time-frequency feature extraction module to obtain temporal-frequency features, wherein the time-frequency feature extraction module includes a multi-branch time-frequency convolution module and a Mamba module; Fusing the temporal-spatial features and the temporal-frequency features corresponding to multiple users to obtain fused features, and classifying and recognizing the motor imagery EEG data based on the fused features.
Citation Information
Patent Citations
Motor imagery electroencephalogram decoding method based on multi-band dual-stage feature extraction network
CN117743942A
Sleep staging method and system based on self-supervised learning and Mama network
CN118490179A
SSVEP (Steady-State Visual Evoked Potential) signal classification method based on Mama model
CN118760946A
Physiological signal fragment analysis method based on adaptive bidirectional selection state space
CN119302668A
Motor imagery electroencephalogram signal classification and identification method and system based on data enhancement and multi-modal feature fusion
CN119475109A
Cited By
Complex action-based space-time frequency domain feature fusion action recognition method and system
CN121327782A