A complex action-based spatio-temporal frequency domain feature fusion action recognition method and system thereof
By fusing the spatiotemporal frequency network TSF-Net with the action recognition method, and utilizing complex action paradigms to activate more brain regions, the problem of insufficient EEG signal features in existing technologies is solved, and high-precision multi-category action recognition is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-12
- Publication Date
- 2026-03-27
AI Technical Summary
Existing motor imagery paradigms are simple in design, resulting in insufficient richness of EEG signal features and poor discrimination, making it difficult to meet the needs of high-precision multi-category action recognition.
The spatiotemporal frequency network TSF-Net is adopted, which integrates time, space and frequency features through the spatiotemporal partitioning branch network TS-Net and the frequency branch network F-Net. It activates more brain regions by using a complex action experiment paradigm and combines the partition weighted splicing method for action recognition.
It improves the quality of EEG signals and the accuracy of action recognition, enhances the model's generalization ability and classification accuracy, and solves the problem of difficult classification of complex actions.
Smart Images

Figure CN121327782B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of electroencephalogram signal processing, and particularly relates to a complex action-based spatio-temporal-frequency feature fusion action recognition method and system. BACKGROUND
[0002] A brain computer interface (BCI) is a new type of human-computer interface that converts invisible brain activity into actual control instructions to control external devices and realizes the communication between the brain and the device by establishing a direct relationship between neural activity information and actual behavior. The technology has been widely applied in medical rehabilitation, education, military and entertainment fields, and has shown great development potential and broad application prospects.
[0003] Motor imagery, as a classic paradigm of brain-computer interface, has shown important clinical application value in the field of neural rehabilitation. By activating brain neuron activity without actually moving the body, it helps people with severe neuromuscular diseases (such as stroke, spinal cord injury, etc.) to undergo neural rehabilitation treatment. The current motor imagery paradigm design is usually simple, mainly involving simple actions such as left hand clenching and right hand clenching. Due to the single imagination content and low cognitive load, it is difficult for subjects to induce significant and stable electroencephalogram (EEG) signal features when performing tasks, resulting in insufficient representation of EEG signals in time, space and frequency domains, and poor discrimination between different action categories. In addition, the range of brain regions activated by simple action imagination is limited, and the feature extraction is easily affected by individual differences and noise interference, thereby affecting the generalization ability and recognition accuracy of the classification model.
[0004] Therefore, there is an urgent need for a motor imagery recognition method that can effectively fuse time, space and frequency multidimensional features, and design a reasonable complex action experiment paradigm to improve the quality of electroencephalogram signals and the performance of the model recognition, and meet the demand for high-precision and multi-class action recognition in practical applications. SUMMARY
[0005] The purpose of the present application is to improve the quality of complex motor imagery signals and improve the recognition accuracy, and a complex action-based spatio-temporal-frequency feature fusion action recognition method and system are proposed. The present application proposes an experimental paradigm for comparing simple actions with complex actions, adopts a time-space-frequency network TSF-Net (including a dual-branch of a time-space partition branch network TS-Net and a frequency branch network F-Net), fuses the three-dimensional features of time, space and frequency, and adopts a partitioned and weighted splicing method for the spatial dimension to highlight important channels. The data set collected through the experimental paradigm is subjected to action recognition, and has important application value and practicality.
[0006] In a first aspect, the present application provides a complex action-based spatio-temporal frequency domain feature fusion action recognition method, which comprises the following steps:
[0007] Collecting electroencephalogram signals of complex action motor imagery paradigm;
[0008] Pretreating the collected electroencephalogram signals to obtain processed signals ;
[0009] Pretreating the collected electroencephalogram signals to obtain processed signals Extracting global spatio-temporal features using a time-space partition branch network TS-Net;
[0010] Pretreating the collected electroencephalogram signals to obtain processed signals Calculating power spectral density and extracting frequency domain features using a frequency branch network F-Net;
[0011] Flattening and splicing the frequency domain features and the global spatio-temporal domain features, and sending them into a classifier for action recognition.
[0012] Preferably, the complex action motor imagery paradigm is that the subject needs to imagine corresponding actions after being prompted by a screen, and the actions include left hand clenched, right hand clenched, left hand drawing 8, right hand drawing 8 and resting state.
[0013] Preferably, in the complex action motor imagery paradigm, one imagination trial lasts for 8 seconds, the signal prompt period accounts for the first 2 seconds, a video prompt of a real person making actions is used, the experimental development period accounts for the middle 4 seconds, the subject imagines, the screen is blank during imagination, the rest period accounts for the last 2 seconds, the subject rests quietly in the seat, and the above total of 8 seconds ends the signal prompt period of the next trial.
[0014] Each session during imagination contains 75 trials, and a total of 6 sessions, with 2 minutes of rest time between sessions, and the subject is allowed to move the body on the seat; wherein, first, 3 sessions including left hand clenched and right hand clenched are performed, each session being as follows: random combination of left hand clenched, right hand clenched and resting state, a total of 30 left hand clenched + 30 right hand clenched + 15 resting state; then, 3 sessions of left hand drawing 8 and right hand drawing 8 are performed, each session being as follows: random combination of left hand drawing 8, right hand drawing 8 and resting state, a total of 30 left hand drawing 8 + 30 right hand drawing 8 + 15 resting state.
[0015] Preferably, the pretreatment includes re-reference, filtering, down-sampling and artifact removal.
[0016] Preferably, the time-space partition branch network TS-Net comprises a time domain feature learning module and a local space domain feature learning module connected in series.
[0017] The time domain feature learning module adopts a time convolution layer, a number F1 of convolution kernels of the time convolution layer is set to 2, which corresponds to two core time domain modes respectively associated with typical time features of the movement preparation stage ERD and the stage after the movement ends ERS;
[0018] The local spatial domain feature learning module adopts a plurality of parallel convolution layers and a fusion module, each convolution layer corresponding to a brain partition; the number of convolution kernels of each convolution layer is F1 x D, wherein D is a dilation factor, ELU is selected as an activation function, and local spatial information is learned by respectively performing convolution on channel dimensions of different brain partitions; the fusion module is used for splicing partition features output by the above convolution layers in the channel dimension.
[0019] Preferably, the size of the convolution kernel of the time convolution layer in the time domain feature learning module is 1 x 100, the step is 1, the Padding mode is set to Same, and a BatchNorm layer and a Drop layer are further set after the time convolution layer;
[0020] The size of the convolution kernel of each convolution layer in the local spatial domain feature learning module is set to x 1, the step is 1, Ci represents the number of channels of the i-th brain partition, a BatchNorm layer, an AvgPool layer and a Drop layer are set.
[0021] Preferably, the brain partition includes six regions:
[0022] Region 1: {P7, P5, P3, P1, PZ, PO7, PO5, PO3, POZ, O1, OZ, CB1}
[0023] Region 2: {P8, P6, P4, P2, PZ, PO8, PO6, PO4, POZ, O2, OZ, CB2}
[0024] Region 3: {T7, C5, C3, C1, CZ, TP7, CP5, CP3, CP1, CPZ}
[0025] Region 4: {T8, C6, C4, C2, CZ, TP8, CP6, CP4, CP2, CPZ}
[0026] Region 5: {FP1, FPZ, AF3, F7, F5, F3, F1, FZ, FT7, FC5, FC3, FC1, FCZ}
[0027] Region 6: {FP2, FPZ, AF4, F8, F6, F4, F2, FZ, FT8, FC6, FC4, FC2, FCZ}.
[0028] Preferably, the frequency branch network F-Net comprises two layers of convolutional layers connected in series;
[0029] The number of convolutional kernels of the first layer of convolutional layers is set to 8, corresponding to 8 kinds of sub-band and combination of motor imagery related electroencephalogram: low alpha wave (8-10Hz), high alpha wave (10-13Hz), low beta wave (13-20Hz), medium beta wave (20-25Hz), high beta wave (25-30Hz), low alpha and low beta energy ratio, high alpha and high beta phase synchronization, and full alpha and full beta band energy difference;
[0030] The number of convolutional kernels of the second layer of convolutional layers is FNet_F1 x FNet_D, wherein FNet_D is an expansion factor.
[0031] Preferably, in the frequency branch network F-Net, the size of the convolutional kernel of the first layer of convolutional layers is 1 x 16, the step is 1, and the Padding mode is set to Same; ELU is selected as the activation function, and a BatchNorm layer and a Dropout layer are further set after the activation function.
[0032] The size of the convolutional kernel of the second layer of convolutional layers is set to 4 x 4, the sliding step is set to 3 x 3, ELU is selected as the activation function, and a BatchNorm layer, an AvgPool layer, and a Dropout layer are set after the activation function.
[0033] In a second aspect, the present application provides a motion recognition system, comprising:
[0034] A data acquisition module is responsible for acquiring electroencephalogram signals of complex motion motor imagery paradigm;
[0035] A data preprocessing module is responsible for preprocessing the acquired electroencephalogram signals to obtain processed signals ;
[0036] A global spatio-temporal feature extraction module is responsible for extracting global spatio-temporal features from the preprocessed signals using a spatio-temporal partition branch network TS-Net;
[0037] A frequency domain feature extraction module is responsible for extracting frequency domain features from the preprocessed signals by calculating power spectral density and using a frequency branch network F-Net;
[0038] A motion recognition module is responsible for flattening and splicing the frequency domain features and the global spatio-temporal domain features, and sending them to a classifier for motion recognition.
[0039] In a third aspect, the present application provides a computing device comprising a memory and a processor, wherein the memory stores executable code, and the processor executes the executable code to implement the method.
[0040] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed in a computer, causes the computer to perform the method.
[0041] Compared with the prior art, the present application has the following advantages:
[0042] 1. The present application constructs a time-space-frequency domain double-branch network (TSF-Net), deeply fuses and splices the global time-space features extracted by a time-space partition branch network (TS-Net) and the multi-sub-band power spectrum features extracted by a frequency branch network (F-Net), breaks through the limitation of traditional methods which are limited in time-space domain, realizes the complementary feature extraction of motor imagery electroencephalogram signals in time, space and frequency dimensions, and adopts the function brain region partition weighting splicing method in the space dimension to highlight the important channel motion recognition method, obtains comprehensive features, highlights the more important features in the space dimension, solves the problem of complex motion classification difficulty, and improves the classification precision.
[0043] 2. The experimental paradigm of the present application combines complex motion and simple motion, the complex motion can increase the difficulty of the imagination task of the test personnel, which helps to build more complex motion scenarios in the brain, and further more fully activates the corresponding regions of the brain, increases the cognitive load and imagination difficulty of the task, helps to more fully activate the corresponding functional areas of the brain, and induces more discriminative electroencephalogram patterns. In addition, the 3D picture is fed back to the test personnel in real time, realizes a neural feedback system, achieves the purpose of improving the correctness of the imagination content, and also helps the test personnel to correct the imagination content. BRIEF DESCRIPTION OF DRAWINGS
[0044] In order to more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings needed in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0045] Figure 1 is a schematic diagram of the TSF-Net network structure of the present application.
[0046] Figure 2 is a schematic diagram of brain partition.
[0047] Figure 3 is a motion description diagram of the complex motion motor imagery paradigm.
[0048] Figure 4 is a flowchart of a complex action motor imagery paradigm. DETAILED DESCRIPTION
[0049] The following will be combined with the accompanying Figure 1 , 2 , 3, 4, the complex action and simple action provided by the application are further described in the motor imagery paradigm and the complex action-based spatiotemporal frequency domain feature fusion action recognition method.
[0050] As Figure 1 shown, the application provides a complex action-based spatiotemporal frequency domain feature fusion action recognition method, which fuses the three-dimensional features of time, space and frequency, and highlights important channels by adopting the partition weighted splicing method of the space dimension, so as to improve the action recognition rate of motor imagery. The specific implementation steps are as follows:
[0051] Step 1: Collecting the electroencephalogram signal of the complex action motor imagery paradigm
[0052] 1.1 Experimental paradigm: After being prompted by the screen, the subject needs to imagine the corresponding action, and the complex action motor imagery paradigm task is as shown in Figure 3 , including left hand clenched fist, right hand clenched fist, left hand drawing 8, right hand drawing 8 and resting state.
[0053] As Figure 4 shown, one imagination test lasts 8 seconds, the signal prompt period accounts for the first 2 seconds, and the video prompt of a real person making action is used, the experimental development period accounts for the middle 4 seconds, the subject imagines, and the screen is blank during imagination. The rest period accounts for the last 2 seconds, and the subject rests quietly in the seat. After the above total of 8 seconds, enter the signal prompt period of the next test, each session during the imagination process contains 75 tests, a total of 6 sessions, and there are 2 minutes of rest time between sessions. The subject can move his body on the seat. Among them, first perform 3 sessions including left hand clenched fist and right hand clenched fist, each session as follows: random combination of left hand clenched fist, right hand clenched fist and resting state, a total of 30 left hand clenched fist+30 right hand clenched fist+15 resting state; then perform 3 sessions of left hand drawing 8 and right hand drawing 8, each session as follows: random combination of left hand drawing 8, right hand drawing 8 and resting state, a total of 30 left hand drawing 8+30 right hand drawing 8+15 resting state.
[0054] 1.2 Experimental equipment: The Neuroscan 64-channel conductive cap (HEO and VEO eye electrode channels are removed during actual operation) is used to collect electroencephalogram (EEG) data. The conductive cap conforms to the international 10-20 standard lead system and is connected with the SynAmps2 amplifier for collection. The sampling rate is set to 1000 Hz, and the impedance value of each electrode should be kept below 30 kilo-ohms. The E-Prime software is used to design the labeling program, which is connected with the collection device through the serial port to realize the functions of paradigm presentation and automatic labeling.
[0055] Collecting N-channel EEG signals:
[0056]
[0057] Wherein, N is the number of channels, T is the number of samples, L is the number of sampling points, is the data of N channels at time t;
[0058] Step 2: Preprocessing the collected EEG signals to obtain the processed signals ;
[0059] 2.1 Re-reference: By eliminating the common mode interference (such as environmental noise) between electrodes, the spatial consistency of EEG signals is enhanced. The common average reference (Common Average Reference) is used in the present application, the signal values of all electrodes are averaged, and then this average value is taken as the reference electrode to reduce the artifacts in the signal. In the original EEG signal, the signal value of the i-th electrode at time t is , the total number of electrodes is N (after removing the HEO and VEO eye electrode, N = 62), then the signal after common average reference is :
[0060]
[0061] 2.2 Filtering: Retaining the EEG rhythms related to motor imagery (alpha wave 8-13 Hz, beta wave 13-30 Hz), and filtering out high-frequency electromyogram, low-frequency drift and other noises. The signal is subjected to Butterworth band-pass filtering in the present application, zero-phase filtering is realized through signal processing tools (butter and filtfilt functions of MATLAB), and signal waveform distortion is avoided, finally the collected EEG signal is filtered to 8-30 Hz.
[0062] 2.3 Down-sampling: Reducing the data sampling rate under the premise of retaining key information, reducing the complexity of subsequent calculation. The decimate signal processing function in MATLAB is used for integer multiple down-sampling, and the data is down-sampled from 1000 Hz to 200 Hz.
[0063] 2.4 Artifact removal: Eliminate non-cerebral noise (such as blinking, frowning, swallowing, etc. Action-induced signal) to avoid interference with motor imagery-related features. The present application uses the Automatic Artifact Removal toolbox to automatically remove eye and muscle artifacts in EEG.
[0064] Step 3: Preprocessing of the signal The spatio-temporal partition branch network TS-Net is used to extract time domain features, local spatial domain features in turn, and construct global spatio-temporal features.
[0065] The spatio-temporal partition branch network TS-Net includes a time domain feature learning module and a local spatial domain feature learning module connected in turn.
[0066] The time domain feature learning module adopts a time convolution layer, which is specifically as follows: the number of convolution kernels of the time convolution layer is F1=2, corresponding to two core time domain modes, respectively associated with the typical time features of ERD (motor preparation stage) and ERS (after motor ends). The convolution kernel size is 1 × 100, the step is 1, the padding mode is set to Same, the BatchNorm layer and the Drop layer (DropoutRate=0.5) are set.
[0067] The local spatial domain feature learning module adopts a plurality of parallel convolution layers and a fusion module, each convolution layer corresponding to a brain partition; wherein the brain partition includes 6 regions, such as Figure 2 :
[0068] Region 1: {P7, P5, P3, P1, PZ, PO7, PO5, PO3, POZ, O1, OZ, CB1}
[0069] Region 2: {P8, P6, P4, P2, PZ, PO8, PO6, PO4, POZ, O2, OZ, CB2}
[0070] Region 3: {T7, C5, C3, C1, CZ, TP7, CP5, CP3, CP1, CPZ}
[0071] Region 4: {T8, C6, C4, C2, CZ, TP8, CP6, CP4, CP2, CPZ}
[0072] Region 5: {FP1, FPZ, AF3, F7, F5, F3, F1, FZ, FT7, FC5, FC3, FC1, FCZ}
[0073] Region 6: {FP2, FPZ, AF4, F8, F6, F4, F2, FZ, FT8, FC6, FC4, FC2, FCZ}
[0074] The number of convolution kernels of each convolution layer in the local spatial feature learning module is F1 x D, where D = 8 is the expansion factor. The convolution kernel size is set to x 1, where is the number of channels of the i-th partition. The step is 1, and ELU is selected as the activation function. The local spatial information is learned by performing convolution on the channel dimension of different partitions respectively. Finally, the BatchNorm layer, the AvgPool layer (size 1 x 16), and the Drop layer (DropoutRate = 0.5) are used to process the data. For an input of size 1 x x T, the final output shape is , where T is the number of time points.
[0075] The fusion module is used to concatenate the partition features output by each convolution layer in the channel dimension, where the weights of brain partitions 1-6 are 0.1, 0.1, 0.25, 0.25, 0.15, and 0.15, respectively.
[0076] Step 4: Calculate the power spectral density of the preprocessed signal and use the frequency branch network F-Net to extract frequency domain features.
[0077] The power spectral density is the Welch power spectral density, which is specifically:
[0078]
[0079] where, is the Fourier transform result of the i-th signal segment, and L is the total number of short data segments into which the original long signal is divided.
[0080] The frequency branch network F-Net includes two layers of convolution layers connected in series.
[0081] The first convolutional layer has 8 kernels (FNet_F1=8), corresponding to 8 EEG signal sub-bands and combinations related to motor imagery: low alpha waves (8-10Hz), high alpha waves (10-13Hz), low beta waves (13-20Hz), medium beta waves (20-25Hz), high beta waves (25-30Hz), low alpha to low beta energy ratio, high alpha to high beta phase synchronization, and energy difference between all alpha and all beta bands. The kernel size is 1 × 16, the stride is 1, and the padding mode is set to Same. ELU is selected as the activation function, and a BatchNorm layer and a Dropout layer (DropoutRate=0.5) are set to process the output data. This layer is for learning. Local feature information in the frequency domain dimension.
[0082] The second convolutional layer has FNet_F1 × FNet_D kernels, where FNet_D=4 is the dilation factor, expanding the feature dimension and increasing the diversity of frequency domain feature combinations the model can learn. The kernel size is set to 4 × 4, the stride is set to 3 × 3, and ELU is selected as the activation function. A BatchNorm layer, an AvgPool layer (size 1 × 4), and a Dropout layer (DropoutRate=0.5) are used to process the output data, compressing the final output data dimension, reducing the amount of data, and lowering computational complexity. For an input of size 1 × C × T, the final output shape is... Where C is the number of channels, T is the number of time points, and F is the frequency band range.
[0083] Step 5: Flatten the frequency domain features and global spatiotemporal domain features, then concatenate them and feed them into the classifier for action recognition. In this embodiment, the classifier uses a Softmax layer. The Softmax layer normalizes the raw scores output by the fully connected layer, converting them into probabilities for each category (the sum of probabilities is 1), as shown in the formula:
[0084]
[0085] in This is the original score of the i-th category, and K is the total number of categories. In this invention, K=5. It is the predicted probability of the i-th category. When determining the category, the category with the highest predicted probability is selected as the final action recognition result.
[0086] Step 6: Design a loss function, train the time-frequency-space domain network constructed in steps 4, 5, and 6 to obtain a set of network parameters, select PyTorch's CrossEntropyLoss (cross-entropy loss) as the loss function, use the Adam optimizer, set the learning rate to 0.001, set the BatchSize to 300, and set the iteration number epoch to 100. The cross-entropy loss of all samples is:
[0087]
[0088] where N is the total number of samples, K is the total number of categories, K = 5 in the present application, is the predicted probability of the i-th category in the n-th sample, is the label encoding of the n-th sample, only the true category corresponds to = 1, and the rest are 0.
[0089] Step 7: Use the trained model to recognize the actions of new subjects and display the corresponding 3D animation on the screen to provide feedback to the subjects. To further verify the effectiveness and superior detection performance of the time-frequency-space domain feature fusion action recognition method constructed in the present application, a comparative experiment is performed.
[0090] Experiments were conducted on the self-collected data set after preprocessing in step 2, and the five-classification accuracy under different methods was obtained for comparison. Two algorithms with the best performance in traditional machine learning and deep learning models, namely FBCSP+SVM and EEGNet, were selected as the comparison methods of the present model, wherein EEGNet adopts the EEGNet-8,2 model (8 time filters, 2 space filters). The experimental results are shown in Table 1.
[0091] Table 1 Comparison of different methods under self-collected data set
[0092] Subject No. FBCSP+SVM (%) EEGNet (%) Invention TSF-Net (%) 1 41.18 37.73 47.59 2 27.41 29.18 34.76 3 43.54 52.39 54.27 4 34.26 51.24 46.63 5 32.68 35.63 41.12 6 37.94 38.62 43.93 7 43.28 47.21 44.28 8 43.84 36.61 38.89 9 35.56 41.24 48.62 Average accuracy rate 37.74 41.09 44.45
[0093] Compared with other methods, the following conclusions can be drawn: the TSF-Net proposed in the present application has advantages compared with the two classic methods of FBCSP+SVM and EEGNet. Among the 9 subjects, 6 subjects achieved the highest accuracy on TSF-Net, and the average accuracy of TSF-Net was the highest, which improved by 6.71% compared with FBCSP+SVM and by 3.36% compared with EEGNet, reflecting the advantages brought by increasing frequency dimension features and not being limited to single brain region spatial division.
[0094] The above is the preferred embodiment of the present application, it should be pointed out that, for those skilled in the art, without departing from the principles of the present application, can also make a number of improvements and refinements, these improvements and refinements are also considered to be within the scope of the present application.
Claims
1. A complex action-based spatiotemporal frequency domain feature fusion action recognition method, characterized in that, The method comprises: Collecting electroencephalogram signals of a complex action motor imagery paradigm; The collected electroencephalogram signal is preprocessed to obtain a processed signal ; on the pre-processed signal extract global spatio-temporal features using a spatio-temporal partitioned branch network, TS-Net; The preprocessed signal The power spectral density is calculated, and the frequency domain features are extracted using a frequency branch network (F-Net). The flattened frequency domain features and global space-time domain features are spliced and sent to a classifier for action recognition. The space-time partition branch network TS-Net comprises time domain feature learning modules and local space domain feature learning modules connected in series. The time domain feature learning module adopts a time convolution layer, the number of convolution kernels F1 of the time convolution layer is set to 2, which corresponds to two core time domain modes, and is respectively associated with the typical time features of the motor preparation stage ERD and the stage after the movement ends ERS; The local space domain feature learning module adopts a plurality of parallel convolution layers and a fusion module, each convolution layer corresponds to a brain partition; the number of convolution kernels of each convolution layer is F1 x D, wherein D is a dilation factor, ELU is selected as an activation function, and local space information is learned by convolution in the channel dimension of different brain partitions; the fusion module is used for splicing the partition features output by the above convolution layers in the channel dimension.
2. The method of claim 1, wherein, The complex action motor imagery paradigm is that the subjects need to imagine corresponding actions after being prompted by a screen, and the actions include left hand clenched, right hand clenched, left hand drawing 8, right hand drawing 8 and resting state.
3. The method of claim 1, wherein, The convolution kernel size of the time convolution layer in the time domain feature learning module is 1 x 100, and the step is 1; The convolution kernel size of each convolution layer in the local spatial feature learning module is set as 1x1, and the step is 1, represents the number of channels of the i-th brain partition.
4. The method of claim 1, wherein, The brain partition comprises six regions: Region 1: {P7, P5, P3, P1, PZ, PO7, PO5, PO3, POZ, O1, OZ, CB1}; Region 2: {P8, P6, P4, P2, PZ, PO8, PO6, PO4, POZ, O2, OZ, CB2}; Region 3: {T7, C5, C3, C1, CZ, TP7, CP5, CP3, CP1, CPZ}; Region 4: {T8, C6, C4, C2, CZ, TP8, CP6, CP4, CP2, CPZ}; Region 5: {FP1, FPZ, AF3, F7, F5, F3, F1, FZ, FT7, FC5, FC3, FC1, FCZ}; Region 6: {FP2, FPZ, AF4, F8, F6, F4, F2, FZ, FT8, FC6, FC4, FC2, FCZ}.
5. The method of claim 1, wherein, The frequency branch network F-Net comprises two layers of convolution layers connected in series; The number of convolution kernels FNet_F1 of the first layer of convolution layers is set to 8, corresponding to 8 kinds of brain electrical signal sub-frequencies and combinations related to motor imagery: low alpha wave, high alpha wave, low beta wave, medium beta wave, high beta wave, low alpha and low beta energy ratio, high alpha and high beta phase synchronization, and full alpha and full beta frequency band energy difference; The number of convolution kernels of the second layer of convolution layers is FNet_F1 x FNet_D, wherein FNet_D is a dilation factor.
6. The method of claim 5, wherein, In the frequency branch network F-Net, the convolution kernel size of the first layer of convolution layers is 1 x 16, and the step is 1; The convolution kernel size of the second layer of convolution layers is set to 4 x 4, and the sliding step size is set to 3 x 3.
7. Action recognition system implementing the method of any one of claims 1 to 6, characterized in that, The method comprises: A data acquisition module is responsible for collecting electroencephalogram signals of complex action motor imagery paradigm; A data preprocessing module is responsible for preprocessing the collected electroencephalogram signals to obtain processed signals ; A global space-time feature extraction module is responsible for extracting global space-time features from the preprocessed signals extract global space-time features using a space-time partition branch network TS-Net; The frequency domain feature extraction module is responsible for extracting frequency domain features from the preprocessed signal calculating power spectral density and extracting frequency domain features using a frequency branch network (F-Net); An action recognition module is responsible for splicing the flattened frequency domain features and global space-time domain features and sending them to a classifier for action recognition.
8. A computing device comprising a memory and a processor, wherein, The memory stores executable code, and the processor executes the executable code to implement the method in any one of claims 1-6.
9. A computer readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed in the computer, the computer executes the method in any one of claims 1-6.
Citation Information
Patent Citations
EEG (electroencephalogram) classification method based on multi-domain feature fusion
CN120892917A