EEG-fNIRS-based early fusion decoding method for motor imagery

By employing an early fusion decoding method for EEG-fNIRS, deep coupling and complementarity of EEG and fNIRS signals are achieved through convolution and pooling operations, cross-attention mechanisms, and a Transformer encoder. This solves the problem of insufficient spatiotemporal resolution in existing technologies and improves the decoding accuracy and robustness of motion visualization tasks.

CN121302171BActive Publication Date: 2026-07-31SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SOUTH CHINA UNIV OF TECH
Filing Date
2025-10-29
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing EEG and fNIRS single-modal motion imagery decoding systems are insufficient in terms of spatiotemporal resolution and anti-interference, making it difficult to meet the requirements of adaptability to complex scenarios and long-term stability. Early fusion strategies did not fully utilize the spatiotemporal complementary characteristics.

Method used

We design an early fusion decoding method based on EEG-fNIRS. We extract signal features through convolution and pooling operations, combine cross-attention mechanism and Transformer encoder to achieve deep coupling and complementarity of EEG and fNIRS signals, and use attention-weighted pooling module to adaptively learn time step features.

Benefits of technology

It significantly improves the decoding accuracy and robustness of motion imagery tasks, achieves efficient feature fusion and temporal alignment of multi-channel signals, captures full sequence dependencies, and enhances the adaptability and stability of the decoding system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121302171B_ABST
    Figure CN121302171B_ABST
Patent Text Reader

Abstract

This invention discloses an early fusion decoding method for motion imagery based on EEG-fNIRS, comprising: acquiring and preprocessing EEG and fNIRS signals from a motion imagery task; extracting and aligning EEG and fNIRS feature information using an EEG and fNIRS feature extraction and alignment module, respectively; performing deep fusion of the time-dimensional aligned features of EEG and fNIRS using a bidirectional cross-attention module to obtain early fusion features of EEG-fNIRS; inputting the early fusion features of EEG-fNIRS into a Transformer encoder, and adaptively fusing information from different time steps through an attention-weighted pooling module to obtain fusion features of EEG-fNIRS; and inputting the fusion features of EEG-fNIRS into a multilayer perceptron to output the motion imagery task category. This invention can fully utilize the spatiotemporal coupling characteristics of EEG and fNIRS signals to achieve deep fusion of cross-modal features, significantly improving the decoding performance of motion imagery tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of brain-computer interfaces and neural engineering, and in particular to an early fusion decoding method for motor imagery based on EEG-fNIRS. Background Technology

[0002] Motor imagery, a crucial component of brain-computer interface technology, refers to the mental simulation of specific limb movements without actual execution. This simulation shares neural similarities with actual voluntary movement and plays a vital role in fields such as medicine, communication, and entertainment. Current motor imagery decoding largely relies on single-modality methods like electroencephalography (EEG) and functional near-infrared spectroscopy (fNIRS). While EEG offers millisecond-level temporal resolution, providing key temporal features for real-time decoding of motor intentions, its spatial resolution is limited and it is susceptible to electromyographic artifacts. In contrast, fNIRS, by monitoring changes in hemoglobin concentration to reflect brain region activation, offers advantages such as strong resistance to electromagnetic interference and high spatial resolution. However, it suffers from inherent hemodynamic response delays and insufficient temporal resolution, making it difficult to meet the real-time requirements of motor imagery decoding. These inherent limitations of single-modality methods pose significant challenges to the adaptability to complex scenarios and long-term stability of existing decoding systems. In recent years, many researchers have discovered that fNIRS, by integrating EEG and fNIRS and fully utilizing their spatiotemporal coupling characteristics, can effectively improve the performance of motor imagery decoding.

[0003] Currently, the stages of EEG and fNIRS fusion strategies are mainly divided into early fusion, mid-stage fusion, and late-stage fusion. Compared to mid-stage and late-stage fusion, early fusion has received more attention because it can more fully exploit the spatiotemporal complementarity between EEG and fNIRS, especially in tasks such as motion imagery that require high spatiotemporal resolution coordination. However, existing EEG-fNIRS early fusion strategies have not fully utilized the spatiotemporal complementarity between EEG and fNIRS, and still face problems such as difficulties in matching spatiotemporal heterogeneity, feature redundancy, and high dimensionality, which restrict further performance improvement. Therefore, this paper designs a novel early fusion decoding algorithm for motion imagery that effectively utilizes the spatiotemporal coupling characteristics of EEG and fNIRS signals, which can better improve the decoding performance of motion imagery tasks. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings and deficiencies of the prior art and propose an early fusion decoding method for motion imagery based on EEG-fNIRS. This method fully utilizes the spatiotemporal coupling characteristics of EEG signals and fNIRS signals, breaks through the limitations of single-modal representation capabilities, and significantly improves the decoding accuracy and robustness of motion imagery tasks.

[0005] To achieve the above objectives, the technical solution provided by this invention is: an early motion imagery fusion decoding method based on EEG-fNIRS, comprising the following steps:

[0006] S1: Acquire the EEG signal and fNIRS signal when performing the motion imagery task, and perform corresponding preprocessing on the EEG signal and fNIRS signal to obtain signal segments of uniform size;

[0007] S2: An EEG feature extraction and alignment module based on convolution and pooling operations was designed for the preprocessed EEG signal to capture local detail features of the EEG signal in the time dimension and perform dimensionality reduction to obtain EEG time dimension alignment features; an fNIRS feature extraction and alignment module based on convolution operations was designed for the preprocessed fNIRS signal to capture local detail features of the fNIRS signal in the time dimension and obtain fNIRS time dimension alignment features.

[0008] S3: A bidirectional cross-attention module based on the cross-attention mechanism was designed, which uses EEG time dimension alignment features and fNIRS time dimension alignment features as query and key values ​​respectively. Through cross-attention mechanism, weighted average, splicing and linear projection operations, deep coupling and complementarity between EEG time dimension alignment features and fNIRS time dimension alignment features are achieved, and early fusion features of EEG-fNIRS are obtained.

[0009] S4: Input the early fusion features of EEG-fNIRS into the Transformer encoder, output the features at different time steps, then use the attention weighted pooling module to adaptively learn the weights of the features at different time steps, and perform weighted summation on the features at different time steps to obtain the final EEG-fNIRS fusion features.

[0010] S5: Input the EEG-fNIRS fused features into a multilayer perceptron for classification, and output a probability distribution vector representing the task category. The category corresponding to the highest probability is the classification result of the motion imagery task.

[0011] Further, in step S1, the preprocessing of the EEG signal includes: firstly, using bandpass filtering and trap filtering to remove noise and interference, then downsampling the EEG signal to a preset frequency, then using independent component analysis to remove artifact signals, and using the average reference method to rereference the original EEG signal, and finally performing baseline correction and signal segmentation to obtain EEG signal segments for a fixed time period.

[0012] The preprocessing of the fNIRS signal includes: first, using bandpass filtering to remove noise; then, downsampling the fNIRS signal to a preset frequency; next, using the modified Beer-Lambert law to convert the fNIRS signal; and finally, performing baseline correction and signal segmentation to obtain the fNIRS signal segment with the same time period as the EEG signal segment.

[0013] Furthermore, the specific operation steps of step S2 are as follows:

[0014] S21: For the input EEG signal An EEG feature extraction and alignment module was designed. First, temporal convolution was used to extract... Local features in the time dimension are extracted, then batch normalized to accelerate convergence. Next, depthwise convolution is used to fuse multi-channel features, and batch normalization is performed again to obtain EEG spatiotemporal features with fused channel information. The formula is expressed as follows:

[0015] ;

[0016] In the formula, Indicates batch normalization, Represents depthwise convolution. Represents temporal convolution. Represents the number of convolution kernels at the EEG time. The channel multiplier representing the depthwise convolution of EEG. The number of channels representing the EEG signal. This represents the stride of the first temporal convolution of the EEG;

[0017] S22: To align the temporal dimension features of EEG and fNIRS signals, the spatiotemporal features of EEG are analyzed. Perform a first time-dimensional pooling, then apply time convolution and batch normalization again to obtain the long-term dependency features of EEG. The formula is expressed as follows:

[0018] ;

[0019] In the formula, This represents the stride of the second temporal convolution of the EEG. This indicates the first time-dimension pooling;

[0020] S23: Long-term dependence on EEG characteristics A second time-dimensional pooling is performed, followed by dimensional transformation to obtain EEG time-dimensional aligned features. The formula is expressed as follows:

[0021] ;

[0022] In the formula, Indicates dimensional transformation. This indicates the second time-dimension pooling;

[0023] S24: For the input fNIRS signal A module for fNIRS feature extraction and alignment was designed, which sequentially uses temporal convolution, batch normalization, depthwise convolution, and batch normalization to obtain the spatiotemporal features of fNIRS that incorporate channel information. The formula is expressed as follows:

[0024] ;

[0025] In the formula, Represents the number of convolutional kernels in the fNIRS time signature. The channel multiplier represents the depthwise convolution of fNIRS. The number of channels representing the fNIRS signal. This represents the stride of the first temporal convolution of fNIRS;

[0026] S25: Spatiotemporal characteristics of fNIRS Temporal convolution and batch normalization are applied again to obtain the long-term dependency features of fNIRS. The formula is expressed as follows:

[0027] ;

[0028] In the formula, This represents the stride of the second temporal convolution of fNIRS;

[0029] S26: Long-term dependence on fNIRS features After dimensional transformation, the time dimension alignment feature of fNIRS is obtained. The formula is expressed as follows:

[0030] .

[0031] Furthermore, in step S3, the bidirectional cross-attention module performs the following operations:

[0032] S31: Alignment features for the time dimension of the input EEG Alignment features with the time dimension of fNIRS First, select features that align the EEG time dimension. As query data, align the fNIRS time dimension with features. As key-value data, they are then linearly mapped to obtain features aligned to the EEG time dimension. Query vector derived from mapping and features aligned by the time dimension of fNIRS Mapped key vector Sum value vector The formula is expressed as follows:

[0033] ;

[0034] In the formula, represent The query vector mapping matrix, and Represent The key vector mapping matrix and the value vector mapping matrix;

[0035] S32: By query vector and key vector The dot product operation is used to construct the correlation matrix, and a scaling factor is introduced to suppress the gradient vanishing problem. Then, the Softmax function is used for normalization calculation, and the result is compared with the value vector. Weighted summation is performed to obtain the EEG enhancement features. The formula is expressed as follows:

[0036] ;

[0037] In the formula, The spatial dimension after linear mapping;

[0038] S33: Align features with the time dimension of fNIRS As query data, align features with the EEG time dimension. As key-value data, they are then linearly mapped to obtain features aligned to the time dimension of fNIRS. Query vector derived from mapping and features aligned by the EEG time dimension Mapped key vector Sum value vector The formula is expressed as follows:

[0039] ;

[0040] In the formula, represent The query vector mapping matrix, and Represent The key vector mapping matrix and the value vector mapping matrix;

[0041] S34: By query vector and key vector The dot product operation is used to construct the correlation matrix, and a scaling factor is introduced to suppress the gradient vanishing problem. Then, the Softmax function is used for normalization calculation, and the result is compared with the value vector. We perform weighted summation to obtain the fNIRS enhanced features. The formula is expressed as follows:

[0042] ;

[0043] S35: Enhanced EEG Features and fNIRS enhanced features A weighted average is performed to obtain the bidirectional fusion features. The formula is expressed as follows:

[0044] ;

[0045] In the formula, and These represent EEG enhancement features. and fNIRS enhanced features Confidence level in classification of motion imagery tasks;

[0046] S36: To preserve fine-grained cross-modal interaction information, EEG enhancement features are used. fNIRS Enhancement Features and bidirectional fusion features By splicing the data, multi-granularity fused features are obtained. The formula is expressed as follows:

[0047] ;

[0048] In the formula, This represents the concatenation function. This indicates that the tensor is spliced ​​along its last dimension.

[0049] S37: Using a linear projection layer to fuse multi-granularity features Dimensionality reduction was performed to obtain early fusion features of EEG-fNIRS. The formula is expressed as follows:

[0050] ;

[0051] In the formula, The weights represent the linear projection layer. This represents the bias vector of the linear projection layer.

[0052] Furthermore, the specific operation steps of step S4 are as follows:

[0053] S41: Early fusion features of EEG-fNIRS The input is fed into an L-layer Transformer encoder. Each layer of the encoder includes a multi-head self-attention operation and a feedforward network function, all of which adopt a layer normalization structure. After L-layer cascading processing, the features H at different time steps are obtained, as expressed by the following formula:

[0054] ;

[0055] ;

[0056] In the formula, The features represent those after multi-head self-attention operation and layer normalization. Representation layer normalization, This represents the multi-head self-attention transformation function. Represents the feedforward network function;

[0057] S42: For features H at different time steps, the attention-weighted pooling module adaptively learns the weights of features at different time steps. First, the features are compressed through the first linear layer, then the ReLU activation function is used to selectively preserve the features. Next, attention scores are mapped through the second linear layer, and finally, the Softmax function is used for normalization calculation to obtain the weights a of features at different time steps, as shown in the following formula:

[0058] ;

[0059] In the formula, Represents the weights of the first linear layer. This represents the bias of the first linear layer. Represents the weights of the second linear layer. This represents the bias of the second linear layer. Indicates the activation function;

[0060] The weights 'a' of features at different time steps and the features 'H' at different time steps are summed using matrix multiplication to obtain the final EEG-fNIRS fused features. The formula is expressed as follows:

[0061] .

[0062] Furthermore, in step S5, the final EEG-fNIRS fusion features are... Input the data into a multilayer perceptron for task classification, and output a predicted label representing the task category. This yields the classification results for the motor imagery task, expressed by the following formula:

[0063] ;

[0064] In the formula, This represents a multilayer perceptron.

[0065] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0066] 1. This invention designs an EEG feature extraction and alignment module based on convolution and pooling operations, which realizes efficient feature fusion and accurate alignment in the time dimension of multi-channel EEG signals.

[0067] 2. This invention designs an fNIRS feature extraction and alignment module based on convolution operation, which realizes efficient feature fusion and accurate alignment in the time dimension of multi-channel fNIRS signals.

[0068] 3. This invention designs a bidirectional cross-attention module based on a cross-attention mechanism, which can fully capture the temporal coupling information between EEG signals and fNIRS signals, and realize deep coupling and complementarity between cross-modal features.

[0069] 4. This invention utilizes a Transformer encoder to encode early fusion features of EEG-fNIRS, and captures dependencies across the entire sequence range through its self-attention mechanism to obtain features at different time steps.

[0070] 5. This invention utilizes an attention-weighted pooling module to replace the traditional global average pooling, which can adaptively fuse information from different time steps to obtain the final EEG-fNIRS fused features. Attached Figure Description

[0071] Figure 1 This is a framework diagram of the method of the present invention.

[0072] Figure 2 This is a schematic diagram of the EEG feature extraction and alignment module.

[0073] Figure 3 This is a schematic diagram of the fNIRS feature extraction and alignment module.

[0074] Figure 4 This is a schematic diagram of a bidirectional cross-attention module.

[0075] Figure 5 This is a schematic diagram of the attention-weighted pooling module. Detailed Implementation

[0076] The present invention will be further described in detail below with reference to the embodiments and accompanying drawings, but the embodiments of the present invention are not limited thereto.

[0077] like Figure 1As shown, this embodiment discloses an early motion imagery fusion decoding method based on EEG-fNIRS, the specific details of which are as follows:

[0078] 1) Acquire the EEG signal and fNIRS signal when performing the motion imagery task, and perform corresponding preprocessing on the EEG signal and fNIRS signal to obtain signal segments of uniform size.

[0079] The motor imagery task used in this embodiment is a binary motor imagery task published in Berlin, Germany, involving 29 subjects. The EEG signal preprocessing operations included: first, bandpass filtering and trap filtering were used to remove noise and interference; then, the EEG signal was downsampled to 200Hz; next, independent component analysis was used to remove artifacts; and the original EEG signal was rereferenced using an average reference method. Finally, baseline correction and signal segmentation were performed to obtain a 3s EEG signal segment. The fNIRS signal preprocessing operations included: first, bandpass filtering was used to remove noise; then, the fNIRS signal was downsampled to 10Hz; next, the fNIRS signal was converted using a modified Beer-Lambert law; and finally, baseline correction and signal segmentation were performed to obtain a 3s fNIRS signal segment.

[0080] 2) Input the preprocessed EEG signal and fNIRS signal into the EEG feature extraction and alignment module and the fNIRS feature extraction and alignment module, respectively.

[0081] like Figure 2 As shown, an EEG feature extraction and alignment module based on convolution and pooling operations was designed for the preprocessed EEG signal. This module captures local detail features of the EEG signal in the time dimension and performs dimensionality reduction to obtain the EEG time dimension aligned features. Figure 3 As shown, a convolution-based fNIRS feature extraction and alignment module was designed for the preprocessed fNIRS signal to capture local detail features in the time dimension and obtain the fNIRS time-dimensional alignment features. The specific operation steps are as follows:

[0082] 2.1) For the input EEG signal An EEG feature extraction and alignment module was designed. First, temporal convolution was used to extract... Local features in the time dimension are extracted, then batch normalized to accelerate convergence. Next, depthwise convolution is used to fuse multi-channel features, and batch normalization is performed again to obtain EEG spatiotemporal features with fused channel information. The formula is expressed as follows:

[0083] ;

[0084] In the formula, Indicates batch normalization, Represents depthwise convolution. Represents temporal convolution. Represents the number of convolution kernels at the EEG time. The channel multiplier representing the depthwise convolution of EEG. The number of channels representing the EEG signal. This represents the stride of the first temporal convolution of the EEG;

[0085] 2.2) To align the temporal features of EEG and fNIRS signals, the spatiotemporal features of EEG were analyzed. Perform a first time-dimensional pooling, then apply time convolution and batch normalization again to obtain the long-term dependency features of EEG. The formula is expressed as follows:

[0086] ;

[0087] In the formula, Indicates batch normalization, Represents temporal convolution. Represents the number of convolution kernels at the EEG time. This represents the stride of the second temporal convolution of the EEG. This indicates the first time-dimension pooling;

[0088] 2.3) Long-term dependence on EEG A second time-dimensional pooling is performed, followed by dimensional transformation to obtain EEG time-dimensional aligned features. The formula is expressed as follows:

[0089] ;

[0090] In the formula, Indicates dimensional transformation. This indicates the second time-dimension pooling;

[0091] 2.4) For the input fNIRS signal A module for fNIRS feature extraction and alignment was designed, which sequentially uses temporal convolution, batch normalization, depthwise convolution, and batch normalization to obtain the spatiotemporal features of fNIRS that incorporate channel information. The formula is expressed as follows:

[0092] ;

[0093] In the formula, Indicates batch normalization, Represents depthwise convolution. Represents temporal convolution. Represents the number of convolutional kernels in the fNIRS time signature. The channel multiplier represents the depthwise convolution of fNIRS. The number of channels representing the fNIRS signal. This represents the stride of the first temporal convolution of fNIRS;

[0094] 2.5) Spatiotemporal characteristics of fNIRS Temporal convolution and batch normalization are applied again to obtain the long-term dependency features of fNIRS. The formula is expressed as follows:

[0095] ;

[0096] In the formula, Indicates batch normalization, Represents temporal convolution. Indicates the number of convolution kernels in the fNIRS time interval. This represents the stride of the second temporal convolution of fNIRS;

[0097] 2.6) Long-term dependence features of fNIRS After dimensional transformation, the time dimension alignment feature of fNIRS is obtained. The formula is expressed as follows:

[0098] ;

[0099] In the formula, This indicates a dimensional transformation.

[0100] 3) A bidirectional cross-attention module based on the cross-attention mechanism was designed, such as... Figure 4 As shown, the EEG time-dimensional alignment features and the fNIRS time-dimensional alignment features are used as query and key-value pairs, respectively. Through cross-attention mechanism, weighted averaging, concatenation, and linear projection operations, deep coupling and complementarity between the EEG time-dimensional alignment features and the fNIRS time-dimensional alignment features are achieved, resulting in early EEG-fNIRS fusion features. The specific operation steps are as follows:

[0101] 3.1) Alignment features of the input EEG time dimension Alignment features with the time dimension of fNIRS First, select features that align the EEG time dimension. As query data, align the fNIRS time dimension with features. As key-value data, they are then linearly mapped to obtain features aligned to the EEG time dimension. Query vector derived from mapping and features aligned by the time dimension of fNIRS Mapped key vector Sum value vector The formula is expressed as follows:

[0102] ;

[0103] In the formula, represent The query vector mapping matrix, and Represent The key vector mapping matrix and the value vector mapping matrix;

[0104] 3.2) By query vector and key vector The dot product operation is used to construct the correlation matrix, and a scaling factor is introduced to suppress the gradient vanishing problem. Then, the Softmax function is used for normalization calculation, and the result is compared with the value vector. Weighted summation is performed to obtain the EEG enhancement features. The formula is expressed as follows:

[0105] ;

[0106] In the formula, The dimension of the space after linear mapping.

[0107] 3.3) Align features with the time dimension of fNIRS As query data, align features with the EEG time dimension. As key-value data, they are then linearly mapped to obtain features aligned to the time dimension of fNIRS. Query vector derived from mapping and features aligned by the EEG time dimension Mapped key vector Sum value vector The formula is expressed as follows:

[0108] ;

[0109] In the formula, represent The query vector mapping matrix, and Represent The key vector mapping matrix and the value vector mapping matrix;

[0110] 3.4) By query vector and key vector The dot product operation is used to construct the correlation matrix, and a scaling factor is introduced to suppress the gradient vanishing problem. Then, the Softmax function is used for normalization calculation, and the result is compared with the value vector. We perform weighted summation to obtain the fNIRS enhanced features. The formula is expressed as follows:

[0111] ;

[0112] In the formula, The dimension of the space after linear mapping.

[0113] 3.5) Enhance EEG features and fNIRS enhanced features A weighted average is performed to obtain the bidirectional fusion features. The formula is expressed as follows:

[0114] ;

[0115] In the formula, and These represent EEG enhancement features. and fNIRS enhanced features Confidence level in the classification of motion imagery tasks.

[0116] 3.6) To preserve fine-grained cross-modal interaction information, EEG enhancement features are used. fNIRS Enhancement Features and bidirectional fusion features By splicing the data, multi-granularity fused features are obtained. The formula is expressed as follows:

[0117] ;

[0118] In the formula, This represents the concatenation function. This indicates that the tensor is spliced ​​along its last dimension.

[0119] 3.7) Using a linear projection layer to fuse multi-granularity features Dimensionality reduction was performed to obtain early fusion features of EEG-fNIRS. The formula is expressed as follows:

[0120] ;

[0121] In the formula, The weights represent the linear projection layer. This represents the bias vector of the linear projection layer.

[0122] 4) Input the early EEG-fNIRS fusion features into the Transformer encoder to obtain features at different time steps. Then, use the attention-weighted pooling module to adaptively learn the weights of the features at different time steps, and perform weighted summation on the features at different time steps to obtain the final EEG-fNIRS fusion features. The specific operation steps are as follows:

[0123] 4.1) Early fusion features of EEG-fNIRS The input is fed into an L-layer Transformer encoder. Each layer of the encoder includes a multi-head self-attention operation and a feedforward network function, all of which adopt a layer normalization structure. After L-layer cascading processing, the features H at different time steps are obtained, as expressed by the following formula:

[0124] ;

[0125] ;

[0126] In the formula, The features represent those after multi-head self-attention operation and layer normalization. Representation layer normalization, This represents the multi-head self-attention transformation function. Represents the feedforward network function;

[0127] 4.2) For features H at different time steps, an attention-weighted pooling module is used to adaptively learn the weights of the features at different time steps, such as... Figure 5 As shown, feature compression is first performed through the first linear layer, followed by selective feature preservation using the ReLU activation function. Attention scores are then mapped through the second linear layer, and finally, the Softmax function is used for normalization to obtain the weights 'a' of the features at different time steps. The formula is as follows:

[0128] ;

[0129] In the formula, Represents the weights of the first linear layer. This represents the bias of the first linear layer. Represents the weights of the second linear layer. This represents the bias of the second linear layer. Indicates the activation function;

[0130] The weights 'a' of features at different time steps and the features 'H' at different time steps are summed using matrix multiplication to obtain the final EEG-fNIRS fused features. The formula is expressed as follows:

[0131] ;

[0132] 5) Fuse the final EEG-fNIRS features Input the data into a multilayer perceptron for task classification, and output a predicted label representing the task category. This yields the classification results for the motor imagery task, expressed by the following formula:

[0133] ;

[0134] In the formula, This represents a multilayer perceptron.

[0135] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An EEG-fNIRS-based early fusion decoding method for motor imagery, characterized in that, Includes the following steps: S1: Acquire the EEG signal and fNIRS signal when performing the motion imagery task, and perform corresponding preprocessing on the EEG signal and fNIRS signal to obtain signal segments of uniform size; S2: An EEG feature extraction and alignment module based on convolution and pooling operations was designed for the preprocessed EEG signal to capture local detail features of the EEG signal in the time dimension and perform dimensionality reduction to obtain EEG time dimension aligned features. A convolution-based fNIRS feature extraction and alignment module was designed for the preprocessed fNIRS signal to capture local detail features of the fNIRS signal in the time dimension and obtain the fNIRS time dimension alignment features. S3: A bidirectional cross-attention module based on the cross-attention mechanism was designed, which uses EEG time dimension alignment features and fNIRS time dimension alignment features as query and key values ​​respectively. Through cross-attention mechanism, weighted average, splicing and linear projection operations, deep coupling and complementarity between EEG time dimension alignment features and fNIRS time dimension alignment features are achieved, and early fusion features of EEG-fNIRS are obtained. S4: Input the early fusion features of EEG-fNIRS into the Transformer encoder, output the features at different time steps, then use the attention weighted pooling module to adaptively learn the weights of the features at different time steps, and perform weighted summation on the features at different time steps to obtain the final EEG-fNIRS fusion features. S5: Input the EEG-fNIRS fused features into a multilayer perceptron for classification, and output a probability distribution vector representing the task category. The category corresponding to the highest probability is the classification result of the motion imagery task.

2. The EEG-fNIRS-based motor imagery early fusion decoding method according to claim 1, characterized in that, In step S1, the preprocessing of the EEG signal includes: first, using bandpass filtering and trap filtering to remove noise and interference; then, downsampling the EEG signal to a preset frequency; next, using independent component analysis to remove artifact signals; and then using the average reference method to rereference the original EEG signal; finally, performing baseline correction and signal segmentation to obtain EEG signal segments for a fixed time period. The preprocessing of the fNIRS signal includes: first, using bandpass filtering to remove noise; then, downsampling the fNIRS signal to a preset frequency; next, using the modified Beer-Lambert law to convert the fNIRS signal; and finally, performing baseline correction and signal segmentation to obtain the fNIRS signal segment with the same time period as the EEG signal segment.

3. The motion imagery early fusion decoding method based on EEG-fNIRS according to claim 2, characterized in that, The specific steps of step S2 are as follows: S21: For the input EEG signal An EEG feature extraction and alignment module was designed. First, temporal convolution was used to extract... Local features in the time dimension are extracted, then batch normalized to accelerate convergence. Next, depthwise convolution is used to fuse multi-channel features, and batch normalization is performed again to obtain EEG spatiotemporal features with fused channel information. The formula is expressed as follows: ; In the formula, Indicates batch normalization, Represents depthwise convolution. Represents temporal convolution. Represents the number of convolutional kernels at the EEG time. The channel multiplier representing the depthwise convolution of EEG. The number of channels representing the EEG signal. This represents the stride of the first temporal convolution of the EEG; S22: To align the temporal dimension features of EEG and fNIRS signals, the spatiotemporal features of EEG are analyzed. Perform a first time-dimensional pooling, then apply time convolution and batch normalization again to obtain the long-term dependency features of EEG. The formula is expressed as follows: ; In the formula, This represents the stride of the second temporal convolution of the EEG. This indicates the first time-dimension pooling; S23: Long-term dependence on EEG characteristics A second time-dimensional pooling is performed, followed by dimensional transformation to obtain EEG time-dimensional aligned features. The formula is expressed as follows: ; In the formula, Indicates dimensional transformation. This indicates the second time-dimension pooling; S24: For the input fNIRS signal A module for fNIRS feature extraction and alignment was designed, which sequentially uses temporal convolution, batch normalization, depthwise convolution, and batch normalization to obtain the spatiotemporal features of fNIRS that incorporate channel information. The formula is expressed as follows: ; In the formula, Represents the number of convolutional kernels in the fNIRS time signature. The channel multiplier represents the depthwise convolution of fNIRS. The number of channels representing the fNIRS signal. This represents the stride of the first temporal convolution of fNIRS; S25: Spatiotemporal characteristics of fNIRS Temporal convolution and batch normalization are applied again to obtain the long-term dependency features of fNIRS. The formula is expressed as follows: ; In the formula, represents the step size of the second time convolution of fNIRS. S26: fNIRS long-time dependence features After dimension transformation, fNIRS time dimension alignment features are obtained The formula is as follows: 。 4. The EEG-fNIRS-based motor imagery early fusion decoding method according to claim 3, characterized in that, In step S3, the bidirectional cross-attention module performs the following operations: S31: Alignment features for the input EEG time dimension Alignment features with the time dimension of fNIRS First, select features that align the EEG time dimension. As query data, align the fNIRS time dimension with features. As key-value data, they are then linearly mapped to obtain features aligned to the EEG time dimension. Query vector derived from mapping and features aligned by the time dimension of fNIRS Mapped key vector Sum value vector The formula is expressed as follows: ; In the formula, represent The query vector mapping matrix, and Represent The key vector mapping matrix and the value vector mapping matrix; S32: By query vector and key vector The dot product operation is used to construct the correlation matrix, and a scaling factor is introduced to suppress the gradient vanishing problem. Then, the Softmax function is used for normalization calculation, and the result is compared with the value vector. Weighted summation is performed to obtain the EEG enhancement features. The formula is expressed as follows: ; In the formula, is the dimension of the linear mapping post-space; S33: Align features with the time dimension of fNIRS As query data, align features with the EEG time dimension. As key-value data, they are then linearly mapped to obtain features aligned to the time dimension of fNIRS. Query vector derived from mapping and features aligned by the EEG time dimension Mapped key vector Sum value vector The formula is expressed as follows: ; In the formula, represent The query vector mapping matrix, and Represent The key vector mapping matrix and the value vector mapping matrix; S34: By query vector and key vector The dot product operation is used to construct the correlation matrix, and a scaling factor is introduced to suppress the gradient vanishing problem. Then, the Softmax function is used for normalization calculation, and the result is compared with the value vector. We perform weighted summation to obtain the fNIRS enhanced features. The formula is expressed as follows: ; S35: EEG enhanced features and fNIRS enhanced features weighted average to obtain bidirectional fusion features The formula is as follows: ; In the formula, and These represent EEG enhancement features. and fNIRS enhanced features Confidence level in classification of motion imagery tasks; S36: To reserve the cross-modal interaction information of fine granularity, the EEG enhanced features , the fNIRS enhanced features and the bidirectional fusion features are spliced to obtain the multi-granularity fusion features , which are expressed as follows: ; wherein denotes a concatenation function, denotes a concatenation along the last dimension of the tensor; S37: using a linear projection layer to the multi-granularity fusion features Dimensionality reduction is performed to obtain the EEG-fNIRS early fusion features , which is expressed by the following formula: ; In the formula, a weight representing a linear projection layer, a bias vector representing a linear projection layer.

5. The motion imagery early fusion decoding method based on EEG-fNIRS according to claim 4, characterized in that, The specific steps of step S4 are as follows: S41: EEG-fNIRS early fusion features Input into the L-layer Transformer encoder, each layer of the encoder includes a multi-head self-attention operation and a feedforward network function, both of which use layer normalization structure, after L-layer cascade processing, different time step features H are obtained, which is expressed as follows: ; ; wherein, represents the features after the multi-head self-attention operation and layer normalization, denotes layer normalization, denotes a multi-head self-attention transformation function, denotes a feed-forward network function; S42: For features H at different time steps, the attention-weighted pooling module adaptively learns the weights of features at different time steps. First, the features are compressed through the first linear layer, then the ReLU activation function is used to selectively preserve the features. Next, attention scores are mapped through the second linear layer, and finally, the Softmax function is used for normalization calculation to obtain the weights a of features at different time steps, as shown in the following formula: ; wherein, weights representing the first linear layer, bias representing the first linear layer, weights representing the second linear layer, bias representing the second linear layer, denotes an activation function; The weight a of the feature of different time steps and the feature H of different time steps are weighted and summed through matrix multiplication operation to obtain the final EEG-fNIRS fusion feature The formula is as follows: 。 6. The motion imagery early fusion decoding method based on EEG-fNIRS according to claim 5, characterized in that, In step S5, the final EEG-fNIRS fusion features are... Input the data into a multilayer perceptron for task classification, and output a predicted label representing the task category. This yields the classification results for the motor imagery task, expressed by the following formula: ; In the formula, This represents a multilayer perceptron.