Brain activity signal recognition method and related apparatus
By using a deentangled multimodal spatiotemporal learning model, the problem of capturing spatiotemporal correlations in hybrid EEG-fNIRS systems is solved, achieving higher signal recognition accuracy and generating more refined multimodal representations.
Patent Information
- Application Number
- CN202510204350.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-02-24
AI Technical Summary
Existing hybrid EEG-fNIRS BCI systems struggle to effectively capture the spatiotemporal correlation between EEG and fNIRS signals in brain activity signal recognition, resulting in limited classification performance and low recognition accuracy.
A disentangled multimodal spatiotemporal learning model is adopted, including a channel reconstruction module, a multimodal attention module, a multi-branch graph convolution module, and a classification module. Through multimodal spatiotemporal coupling and disentangled representation learning, the spatiotemporal coupling characteristics and correlations of EEG and fNIRS signals are captured.
It significantly improves the accuracy of brain activity signal recognition, generates more refined multimodal representations, and enhances the ability to recognize the brain activity signals of subjects.
Smart Images

Figure CN120217074B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of brain-computer interface technology, and in particular to a method and related apparatus for recognizing brain activity signals. Background Technology
[0002] Brain-computer interfaces (BCIs) have greatly advanced human-computer interaction by establishing direct communication between the brain and external devices. Currently, BCI technology is widely used in fields such as robot control, workload monitoring, and medical rehabilitation, playing a crucial role, especially in assisting patients with limb disabilities or other neuromuscular degeneration. BCI systems can be implemented using various neuroimaging modalities, including stereoscopic electroencephalography (SEEG), functional magnetic resonance imaging (fMRI), electroencephalography (EEG), and functional near-infrared spectroscopy (fNIRS). Among these modalities, EEG and fNIRS have attracted significant attention due to their non-invasiveness and ease of operation.
[0003] Electroencephalography (EEG) records the brain's electrophysiological activity via scalp electrodes, offering good temporal resolution but relatively limited spatial resolution and susceptibility to artifacts and noise. In contrast, fNIRS reflects brain activity by measuring cerebral blood flow and metabolic changes, providing better spatial resolution and less susceptibility to noise, but with relatively lower temporal resolution. Combining these two modalities provides more comprehensive information on brain activity, effectively compensating for the limitations of a single modality. Since near-infrared light does not interfere with electrical signals, simultaneous measurement of EEG and fNIRS is possible, leading to widespread interest in hybrid EEG-fNIRS BCI systems in both basic research and practical applications.
[0004] EEG and fNIRS signals originate from multiple brain regions and fluctuate over time, exhibiting complex spatiotemporal characteristics. In light of this, researchers have developed various methods to utilize their spatiotemporal information. In hybrid EEG-fNIRS BCI systems, traditional machine learning methods such as Linear Discriminant Analysis (LDA), Support Vector Machines (SVM), and k-Nearest Neighbors (K-NN) have been widely applied. While these traditional methods have achieved some success, they largely rely on manual feature engineering, which limits their ability to effectively capture the complex spatiotemporal dynamics of EEG and fNIRS signals.
[0005] In recent years, the rise of deep learning technology has brought significant progress to hybrid EEG-fNIRS systems. Current research often employs a late fusion strategy, that is, performing feature fusion after extracting EEG and fNIRS features separately. However, this strategy often fails to capture the potential spatiotemporal correlations between EEG and fNIRS signals, resulting in limited classification performance and low recognition accuracy. Therefore, improving the accuracy of brain activity signal recognition in subjects has become an urgent technical problem to be solved. Summary of the Invention
[0006] This invention provides a method and related device for identifying brain activity signals, which can effectively capture the spatiotemporal coupling characteristics and correlations of EEG and fNIRS signals, generate more refined multimodal representations, and thus significantly improve the accuracy of identifying brain activity signals in subjects.
[0007] In a first aspect, embodiments of the present invention provide a method for recognizing brain activity signals, comprising:
[0008] Acquire the subject's EEG and fNIRS signals;
[0009] The EEG signal and the fNIRS signal are input into a disentangled multimodal spatiotemporal learning model to obtain the brain activity signal recognition results of the subject. The disentangled multimodal spatiotemporal learning model is constructed based on multimodal spatiotemporal coupling and disentangled representation learning. The disentangled multimodal spatiotemporal learning model includes a channel reconstruction module, a multimodal attention module, a multi-branch graph convolution module, and a classification module. The channel reconstruction module is used to generate features with rich spatiotemporal patterns for each modality. The multimodal attention module is used to comprehensively capture intermodal correlations and enhance the representation of each modality. The multi-branch graph convolution module is used for disentangled representation learning and capturing effective spatiotemporal coupling features. The classification module is used to adaptively fuse the disentangled representations and perform task prediction.
[0010] In some embodiments, inputting the EEG signal and the fNIRS signal into the disentangled multimodal spatiotemporal learning model to obtain the brain activity signal recognition result of the subject includes:
[0011] The EEG signal and the fNIRS signal are input to the channel reconstruction module to obtain the corresponding preliminary features;
[0012] The preliminary features are input into the multimodal attention module to obtain the corresponding enhanced features;
[0013] The enhanced features are input into the multi-branch graph convolutional module to obtain corresponding public and private features;
[0014] The public and private features are input into the classification module to obtain the brain activity signal recognition result.
[0015] In some embodiments, the channel reconstruction module includes an EEG embedding block and a near-infrared embedding block. The step of inputting the EEG signal and the fNIRS signal into the channel reconstruction module to obtain corresponding preliminary features includes:
[0016] The EEG signal E (0) The input is fed into the EEG embedding block, and the first preliminary feature E is extracted. (1) The EEG embedding block includes a temporal convolutional layer, a spatial convolutional layer, an average pooling layer, and a rearrangement layer.
[0017] The fNIRS signal F (0) The input is fed into the near-infrared embedding block to extract the second preliminary feature F. (1) The near-infrared embedding block includes a temporal convolutional layer, a spatial convolutional layer, and a rearrangement layer.
[0018] In some embodiments, the multimodal attention module includes a feature connectivity layer and a modal attention interaction block, and notations are introduced in the multimodal attention module. To express the first preliminary feature E (1) and the second preliminary feature F (1) The first preliminary feature E (1) and the second preliminary feature F (1) The input is fed into the feature connection layer to obtain hybrid features. The hybrid feature P is embedded into two spaces, denoted as K = LN(P) and V = LN(P), respectively, and the query for each layer is defined as Q. (l) =LN(U (l) The attention interaction of the l-th layer is defined as follows:
[0019]
[0020] Where LN represents layer normalization, l = 1, ..., N, and N is the number of layers; the attention score matrix is calculated using formula (1). To measure the attention level from the i-th time step in a single modality to the j-th time step in a mixed modality, where i = 1, ..., M, j = 1, ..., 2M; and to further enhance the diversity of representations using a multi-head attention strategy;
[0021] The forward propagation process of layer l is represented as follows:
[0022]
[0023] Wherein, FC is the feedforward layer. Subsequently, the enhanced features H∈{H} of the two modal outputs are... e H f} is represented as {H e H f}=(E (N+1) ) T ,F (N+1) ) T By extracting complementary features from the mixed modalities, the enhanced features H learned by the modal attention interaction block have stronger representational capabilities.
[0024] In some embodiments, in the multi-branch graph convolution module, given H e and H f The common features of the two modalities are obtained through the same common encoding function. and
[0025] Z ce =C (E,F) (H e ;θ c (4)
[0026] Z cf =C (E,F) (H f ;θ c ), (5)
[0027] Among them, C (E,F) (·) represents the common encoding function based on a graph convolutional network, and these two modes are in C (E,F) (·) share the same parameter θ c ;
[0028] Given H e and H f Different private encoding functions are used to learn the private features of the two modalities. and
[0029]
[0030] Among them, P E (·) and P F Both (·) are implemented using graph convolutional networks, with two private encoding functions assigning separate parameters to the corresponding modes. and
[0031] The specific implementation of the graph convolutional network is as follows: For enhanced features A graph structure was dynamically constructed for each sample. The associations between learning channels are studied, where the adjacency matrix A is defined as follows:
[0032]
[0033] Specifically, the ReLU activation function is deployed to ensure the non-negativity of the adjacency matrix. This represents element-wise product; it assumes that the connections between nodes are undirected and that the adjacency matrix of the graph structure is... It is symmetrical, A base The elements are calculated based on the dot product between the node feature vectors, and MASK is a trainable mask with the same dimension as the adjacency matrix.
[0034] The adjacency matrix is normalized, denoted as:
[0035]
[0036] Where D = diag(d1, d2, ..., d K Let ) be the degree matrix of A, where d m =∑ n A(m,n), where I is the identity matrix;
[0037] Using the enhanced feature matrix H and the normalized adjacency matrix The GCN layer is represented as:
[0038]
[0039] in, Let b be the weight matrix and b be the deviation vector.
[0040] In some embodiments, in the classification module, for two private features and and two common features and Transform the two private features and two public features into a form of length d = K × M. out The vectors are used, and an attention mechanism is employed to learn their importance, as follows:
[0041] (α e ,α ce ) = att(v e ,v ce ), (11)
[0042] (α f ,α cf ) = att(v f ,v cf ), (12)
[0043] Among them, v e , This indicates the flattened Z e and Z ce The normalized embedding obtained by applying L2-normalization, α e and α ce This represents the attention value between them; similarly, α f and α cf It also follows the same principle;
[0044] Using vectors Obtain attention value ω e =v e Similarly, we can obtain vector v. ce Attention value ω ce Then, the softmax function is used to adjust the attention value ω. e and ω ce Normalization is performed to obtain the final weight α. e =softmax(ω e ), where α e The larger the value, the more important the corresponding embedding; similarly, we can obtain α. ce =softmax(ω ce Then, combining these two embeddings, we obtain the final embedding v. Eout =α e ·v e +α ce ·v ce ;
[0045] In obtaining Then, a linear transformation is performed to obtain the class prediction:
[0046] y Eout =softmax(v Eout W e +b e ), (13)
[0047] in, C is the number of categories; similarly, given v f and v cf ,get
[0048] Then, these two predictions are added together to obtain the final output:
[0049]
[0050] in, This indicates element-wise addition. use This represents the probability of belonging to class c.
[0051] In some embodiments, the training method of the disentangled multimodal spatiotemporal learning model includes:
[0052] Construct a target loss function, which is determined based on the task loss function, consistency loss function, and difference loss function;
[0053] The unentangled multimodal spatiotemporal learning model is trained based on the target loss function to obtain the trained unentangled multimodal spatiotemporal learning model;
[0054] Specifically, for the task loss function, label smoothing is used to prevent model overfitting and improve generalization ability; assuming that each batch contains B samples during training, the model's predicted output is... By calculating the cross-entropy loss for each batch of samples and taking the average within each batch:
[0055]
[0056] in, It is the predicted probability of the b-th sample in the batch in class c. It is the smoothed true probability of the same sample and class;
[0057] For the aforementioned consistency loss function, consistency constraints are used to further enhance the generality of the two outputs of the common encoding function for the two modalities; let... and For each row, v ce and v cf Let B be a matrix, and let B represent the batch size. Define the following constraints:
[0058]
[0059] in, It is the square of the F-norm;
[0060] For the aforementioned differential loss function, orthogonal constraints are applied to ensure that private and public coding functions of the same modality can effectively capture different features of the input; let... and Its row corresponds to v e and v f The matrix is given by B, where B is the batch size; the orthogonality constraint between the outputs of the private encoding function and the public encoding function is calculated by the following formula:
[0061]
[0062] Combined with the task loss function The consistency loss function and the difference loss function The target loss function is calculated as follows:
[0063]
[0064] Where α and β are trade-off parameters, with values ranging from {0.01, 0.1, 1, 3, 10}. The objective loss function simultaneously optimizes the task loss function. The consistency loss function and the difference loss function
[0065] Secondly, embodiments of the present invention also provide a brain activity signal recognition device, the device comprising:
[0066] The acquisition module is used to acquire the subject's EEG and fNIRS signals;
[0067] The recognition module is used to input the EEG signal and the fNIRS signal into the unentangled multimodal spatiotemporal learning model to obtain the brain activity signal recognition result of the subject. The unentangled multimodal spatiotemporal learning model is constructed based on multimodal spatiotemporal coupling and unentangled representation learning. The unentangled multimodal spatiotemporal learning model includes a channel reconstruction module, a multimodal attention module, a multi-branch graph convolution module, and a classification module. The channel reconstruction module is used to generate features with rich spatiotemporal patterns for each modality. The multimodal attention module is used to comprehensively capture the correlation between modalities and enhance the representation of each modality. The multi-branch graph convolution module is used for unentangled representation learning and capturing effective spatiotemporal coupling features. The classification module is used to adaptively fuse the unentangled representations and perform task prediction.
[0068] Thirdly, embodiments of the present invention also provide an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the brain activity signal recognition method as described in the first aspect.
[0069] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing computer-executable instructions for performing the brain activity signal recognition method as described in the first aspect.
[0070] According to embodiments of the present invention, a method and related apparatus for identifying brain activity signals are provided. The method includes: acquiring the EEG signal and fNIRS signal of a subject; inputting the EEG signal and fNIRS signal into a disentangled multimodal spatiotemporal learning model to obtain the subject's brain activity signal identification result. The disentangled multimodal spatiotemporal learning model is constructed based on multimodal spatiotemporal coupling and disentangled representation learning. The model includes a channel reconstruction module, a multimodal attention module, a multi-branch graph convolution module, and a classification module. The channel reconstruction module generates features with rich spatiotemporal patterns for each modality. The multimodal attention module comprehensively captures intermodal correlations and enhances the representation of each modality. The multi-branch graph convolution module learns disentangled representations and captures effective spatiotemporal coupling features. The classification module adaptively fuses the disentangled representations and performs task prediction. Based on this, embodiments of the present invention can effectively capture the spatiotemporal coupling features and correlations of EEG and fNIRS signals, thereby generating more refined multimodal representations and significantly improving the accuracy of identifying the subject's brain activity signals. Attached Figure Description
[0071] Figure 1 This is a flowchart of a brain activity signal recognition method provided in one embodiment of the present invention;
[0072] Figure 2 This is a schematic diagram of the experimental paradigm and channel positions of a dataset provided in one embodiment of the present invention;
[0073] Figure 3 This is a schematic diagram of the overall framework of a disentangled multimodal spatiotemporal learning model provided in one embodiment of the present invention;
[0074] Figure 4 This is a flowchart of step S102 provided in one embodiment of the present invention;
[0075] Figure 5 This is a flowchart of step S401 provided in one embodiment of the present invention;
[0076] Figure 6 This is a flowchart of a training method for a disentangled multimodal spatiotemporal learning model provided in one embodiment of the present invention;
[0077] Figure 7 This is a scatter plot of the classification accuracy (x-axis) of EEG signals and fNIRS signals and the classification accuracy (y-axis) of mixed modes provided in an embodiment of the present invention;
[0078] Figure 8 This is a schematic diagram of a brain activity signal recognition device provided in one embodiment of the present invention;
[0079] Figure 9This is a schematic diagram of an electronic device provided in one embodiment of the present invention. Detailed Implementation
[0080] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0081] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the following drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0082] In this embodiment of the invention, the terms "furthermore," "exemplarily," or "optionally" are used as examples, illustrations, or descriptions and should not be construed as being more preferred or advantageous than other embodiments or designs. The use of the terms "furthermore," "exemplarily," or "optionally" is intended to present the relevant concepts in a specific manner.
[0083] To facilitate a more convenient description of the working principle of the embodiments of the present invention, the following introduction of relevant technical scenarios is given first.
[0084] In recent years, the rise of deep learning technology has brought significant progress to hybrid EEG-fNIRS systems. Current research often employs a late fusion strategy, that is, performing feature fusion after extracting EEG and fNIRS features separately. However, this strategy often fails to capture the potential spatiotemporal correlations between EEG and fNIRS signals, resulting in limited classification performance and low recognition accuracy. Therefore, improving the accuracy of brain activity signal recognition in subjects has become an urgent technical problem to be solved.
[0085] Based on this, the present invention provides a method and related apparatus for identifying brain activity signals. The method includes: acquiring the EEG signal and fNIRS signal of a subject; inputting the EEG signal and fNIRS signal into a disentangled multimodal spatiotemporal learning model to obtain the subject's brain activity signal identification result. The disentangled multimodal spatiotemporal learning model is constructed based on multimodal spatiotemporal coupling and disentangled representation learning. The disentangled multimodal spatiotemporal learning model includes a channel reconstruction module, a multimodal attention module, a multi-branch graph convolution module, and a classification module. The channel reconstruction module generates features with rich spatiotemporal patterns for each modality; the multimodal attention module comprehensively captures intermodal correlations and enhances the representation of each modality; the multi-branch graph convolution module learns disentangled representations and captures effective spatiotemporal coupling features; and the classification module adaptively fuses the disentangled representations and performs task prediction. Based on this, the embodiments of the present invention can effectively capture the spatiotemporal coupling features and correlations of EEG and fNIRS signals, thereby generating a more refined multimodal representation and significantly improving the accuracy of identifying the subject's brain activity signals.
[0086] It should be noted that in various specific embodiments of the present invention, when processing data related to the identity or characteristics of a subject, such as subject information, subject behavior data, subject historical data, and subject location information, is required, the subject's permission or consent will be obtained first. Furthermore, the collection, use, and processing of this data will comply with relevant laws, regulations, and standards. In addition, when embodiments of the present invention require obtaining the subject's personal information, the subject's individual permission or consent will be obtained through methods such as collecting information from forms or completing informed consent forms. Only after obtaining the subject's individual permission or consent will the necessary subject-related data for the normal operation of the embodiments of the present invention be obtained.
[0087] The embodiments of the present invention will be further described below with reference to the accompanying drawings.
[0088] like Figure 1 As shown, Figure 1 This is a flowchart of a brain activity signal recognition method provided in an embodiment of the present invention. The brain activity signal recognition method may include, but is not limited to, steps S101 to S102.
[0089] Step S101: Acquire the subject's EEG signal and fNIRS signal;
[0090] Step S102: Input the EEG signal and fNIRS signal into the unentangled multimodal spatiotemporal learning model to obtain the brain activity signal recognition results of the subject. The unentangled multimodal spatiotemporal learning model is constructed based on multimodal spatiotemporal coupling and unentangled representation learning. The unentangled multimodal spatiotemporal learning model includes a channel reconstruction module, a multimodal attention module, a multi-branch graph convolution module, and a classification module. The channel reconstruction module is used to generate features with rich spatiotemporal patterns for each modality. The multimodal attention module is used to comprehensively capture the correlation between modalities and enhance the representation of each modality. The multi-branch graph convolution module is used for unentangled representation learning and capturing effective spatiotemporal coupling features. The classification module is used to adaptively fuse the unentangled representations and perform task prediction.
[0091] Understandably, during the model training phase, this invention uses the 2017 Berlin Public Dataset HBCI as experimental data. The dataset includes 29 subjects (14 men and 15 women). The experimental paradigm is as follows: Figure 2 As shown in (a), subjects were required to perform 30 trials for each task. Each trial began with a 2-second cue, followed by a 10-second task time, and ended with a randomly assigned rest period of 15 to 17 seconds. The tasks included: baseline (resting state), mental arithmetic, left-hand motor imagery, and right-hand motor imagery. This invention divides the data into two datasets based on the tasks: the MA dataset (baseline and mental arithmetic) and the MI dataset (left-hand and right-hand motor imagery). EEG and fNIRS data were collected simultaneously, with channel positions as shown... Figure 2 As shown in Figure (b), 30 electrode channels (dots) record EEG signals, and 14 light sources (solid squares) and 16 detectors (shaded squares) constitute 36 fNIRS channels (solid black lines) to record fNIRS signals. This invention further divides the EEG and fNIRS signals into 3-second segments with a step size of 1 second to evaluate real-time BCI performance. Therefore, the sizes of the EEG and fNIRS signals are 30 × 600 (channel × time) and 72 × 30 (channel × time), respectively. The fNIRS data is stacked with HbO and HbR in the channel dimension, with a total of 600 samples per subject (10 segments × 30 trials × 2 tasks).
[0092] It is understandable that, such as Figure 3As shown, the Disentangled Multimodal Spatiotemporal Learning (DMSL) model of this invention consists of a Channel reconstruction module, a Multi-modal attention module, a Multi-branch Graph Convolutional Module (MCCN) with consistency and difference constraints, and a Classification module. The DMSL model simultaneously performs multimodal spatiotemporal coupling and disentangled representation learning within a unified architecture. Specifically, the Channel reconstruction module generates features with rich spatiotemporal patterns for each modality; the Multi-modal attention module comprehensively captures intermodal correlations and enhances the representation of each modality, thereby improving the representational power and robustness of single modalities; the MCN module learns disentangled representations and captures effective spatiotemporal coupling features to capture the inherent complex spatiotemporal relationships in the signal; and the Classification module adaptively fuses the disentangled representations and performs task prediction.
[0093] It is understandable that, such as Figure 4 As shown, step S102 may include, but is not limited to, the following steps:
[0094] Step S401: Input the EEG signal and fNIRS signal to the channel reconstruction module to obtain the corresponding preliminary features;
[0095] Step S402: Input the preliminary features into the multimodal attention module to obtain the corresponding enhanced features;
[0096] Step S403: Input the enhanced features into the multi-branch graph convolution module to obtain the corresponding public and private features;
[0097] Step S404: Input the public and private features into the classification module to obtain the brain activity signal recognition results.
[0098] The EEG and fNIRS signals are input into the deentangled multimodal spatiotemporal learning model. Specifically, the EEG and fNIRS signals are input into the channel reconstruction module to obtain corresponding preliminary features. The preliminary features are then input into the multimodal attention module to obtain corresponding enhanced features. The enhanced features are then input into the multi-branch graph convolution module to obtain corresponding public and private features. Finally, the public and private features are input into the classification module to obtain the brain activity signal recognition result.
[0099] It is understandable that, such as Figure 5As shown, the channel reconstruction module includes an EEG embedding block and a near-infrared embedding block. Step S401 may include, but is not limited to, the following steps:
[0100] Step S501, EEG signal E (0) Input to the EEG embedding block, and the first preliminary feature E is extracted. (1) The EEG embedding block includes a temporal convolutional layer, a spatial convolutional layer, an average pooling layer, and a rearrangement layer.
[0101] Step S502, the fNIRS signal F (0) The input is fed into the near-infrared embedding block to extract the second preliminary feature F. (1) The near-infrared embedding block includes a temporal convolutional layer, a spatial convolutional layer, and a rearrangement layer.
[0102] Changes in cognitive processes in the brain are reflected at different time points and activation levels in different brain regions. Hypothesis (E) (0) ,F (0) ) are inputs from EEG and fNIRS, where, Here, C (·) It is the number of channels, S (·) This refers to the number of sampling points. This invention reconstructs channels by adding a dimension, with each reconstructed channel obtaining different information from the original input.
[0103] The structure of the channel reconstruction module is shown in Table 1. For the EEG embedding block, the first two layers focus on the temporal dimension and the interaction between electrode channels, respectively, followed by batch normalization and ELU activation. The third layer performs average pooling to reduce overfitting and complexity. Finally, the features in the convolutional module are reordered by compressing and transposing the dimensions. For the near-infrared embedding block, this invention specifically designs the kernel size to ensure that the feature dimension of fNIRS is aligned with the feature dimension of EEG, thereby simplifying the computation.
[0104] Table 1 Structure of the Channel Reconstruction Module
[0105]
[0106] By using EEG embedding blocks and near-infrared embedding blocks as feature extractors, this invention obtains a pair of features for two modalities (E... (1) ,F (1) ).here, and The first and second preliminary features extracted are input into the subsequent multimodal attention module, where M represents the number of time points and K represents the number of convolutional kernels.
[0107] It is understood that the multimodal attention module of this invention includes two stages: a Feature Connectivity (CAT) layer and a Modality Attention Interaction (MAI) block, such as... Figure 3 As shown.
[0108] After using the aforementioned channel reconstruction module, the present invention obtains the preliminary features E of these two modes. (1) and F (1) To simplify the expression, this invention introduces notations. This invention obtains the hybrid features of these two modes through the CAT layer. The present invention then embeds P into two spaces, denoted as K = LN(P) and V = LN(P) respectively, and the query for each layer is defined as Q. (l) =LN(U (l) The attention interaction of the l-th layer is defined as follows:
[0109]
[0110] Where LN represents layer normalization, l = 1, ..., N, and N is the layer number. Specifically, the attention score matrix calculated using formula (1) is used. The degree of attention from the i-th time step within a single modality to the j-th time step in a mixed modality is measured, where i = 1, ..., M, j = 1, ..., 2M. This invention also employs a multi-head attention strategy to further enhance the diversity of representations.
[0111] The forward propagation process of layer l is represented as follows:
[0112]
[0113] Wherein, FC is the feedforward layer. Subsequently, the outputs H∈{H} of the two modes are... e H f} is represented as {H e H f}=(E (N+1) ) T ,F (N+1) ) T}, where H e H is the first enhancement feature. f This is the second enhancement feature. By extracting complementary features from the mixed modalities, the enhancement feature H learned by the MAI block has stronger representational power.
[0114] Understandably, this invention further utilizes a multi-branch graph convolution module to explicitly extract the complex features contained in H, enabling the invention to balance these characteristics. Given H e and Hf The common features of the two modalities are represented as the first common feature. Second common features It can be obtained through the same common encoding function:
[0115] Z ce =C (E,F) (H e ;θ c (4)
[0116] Z cf =C (E,F) (H f ;θ c ), (5)
[0117] Among them, C (E,F) (·) is a common encoding function based on a graph convolutional network. It's important to note that these two modes in C... (E,F) (·) share the same parameter θ c .
[0118] Similarly, given H e and H f This invention uses different private encoding functions to learn private features, namely the first private feature. Second private feature It can be obtained through different private encoding functions:
[0119]
[0120] Among them, the first private encoding function P E (·) and the second private encoding function P F (·) is also implemented using a graph convolutional network. Two private encoding functions assign separate parameters to the corresponding modes. and
[0121] The specific implementation of the graph convolutional network is as follows: for the enhanced features obtained after the MAI block... This invention dynamically constructs a graph structure for each sample. The associations between learning channels are studied, where the adjacency matrix A is defined as follows:
[0122]
[0123] Specifically, the ReLU activation function is deployed to ensure the non-negativity of the adjacency matrix. This represents element-wise product. This invention assumes that the connections between nodes are undirected, and that the adjacency matrix of the graph structure... It is symmetrical. A baseThe elements are calculated based on the dot product between the node feature vectors, and MASK is a trainable mask with the same dimension as the adjacency matrix.
[0124] Furthermore, the adjacency matrix is normalized, denoted as:
[0125]
[0126] Where D = diag(d1, d2, ..., d K Let ) be the degree matrix of A, where d m =∑ n A(m,n), where I is the identity matrix.
[0127] Using the enhanced feature matrix H and the normalized adjacency matrix The GCN layer is represented as:
[0128]
[0129] in, Let b be the weight matrix and b be the deviation vector.
[0130] It is understandable that for two private features and and two common features and This invention converts these features into a length of d = K × M out The vectors are used, and an attention mechanism is employed to learn their importance, as follows:
[0131] (α e ,α ce ) = att(v e ,v ce ), (11)
[0132] (α f ,α cf ) = att(v f ,v cf ), (12)
[0133] Among them, v e , This indicates the flattened Z e and Z ce The normalized embedding obtained by applying L2-normalization, α e and α ce This represents the attention value between them. Similarly, α f and α cf It also follows the same principle.
[0134] This invention uses vectors Obtain attention value ω e =v e q. Similarly, this invention can also obtain vector v. ce Attention value ω ce Then, this invention uses the softmax function to adjust the attention value ω. e and ω ce Normalization is performed to obtain the final weight α. e =softmax(ω e ). α e The larger the value, the more important the corresponding embedding. Similarly, this invention can also obtain α. ce =softmax(ω ce Then, the present invention combines these two embeddings to obtain the final embedding v. Eout =α e ·v e +α ce ·v ce .
[0135] In obtaining Subsequently, the present invention performs a linear transformation to obtain class prediction:
[0136] y Eout =softmax(v Eout W e +b e ), (13)
[0137] in, C is the number of categories. Similarly, given v f and v cf The present invention can obtain
[0138] The present invention then combines these two predictions to obtain the final output:
[0139]
[0140] in, This indicates element-wise addition. Here, use This represents the probability of belonging to class c.
[0141] It is understandable that, such as Figure 6 As shown, the training method for the disentangled multimodal spatiotemporal learning model may include, but is not limited to, the following steps:
[0142] Step S601: Construct the target loss function, which is determined based on the task loss function, consistency loss function, and difference loss function;
[0143] Step S602: Train the unentangled multimodal spatiotemporal learning model based on the target loss function to obtain the trained unentangled multimodal spatiotemporal learning model.
[0144] For the training and optimization process of the disentangled multimodal spatiotemporal learning model, a target loss function is constructed. The target loss function is determined based on the task loss function, consistency loss function, and difference loss function. The disentangled multimodal spatiotemporal learning model is trained based on the target loss function to obtain the trained disentangled multimodal spatiotemporal learning model.
[0145] 1) Mission loss
[0146] This invention uses label smoothing to prevent model overfitting and improve generalization ability. Assuming each batch contains B samples during training, the model's predicted output is... This invention calculates the cross-entropy loss for each batch of samples and takes the average value within the batch:
[0147]
[0148] in, It is the predicted probability of the b-th sample in the batch in class c. It is the smoothed true probability of the same sample and class.
[0149] 2) Consistency loss
[0150] This invention uses consistency constraints to further enhance the generality of the two outputs of the common encoding function. Let... and For each row, v ce and v cf The matrix is given by B, where B represents the batch size. This invention defines the following constraints:
[0151]
[0152] in, It is the square of the F-norm.
[0153] 3) Loss of difference
[0154] This invention applies orthogonal constraints to ensure that private and public coding functions of the same modality capture different features of the input. Let... and Its row corresponds to v e and v f The matrix is given by B, where B is the batch size. The orthogonality constraint between the outputs of the private and public coding functions is calculated using the following formula:
[0155]
[0156] 4) Overall objective function
[0157] Combined with mission losses Consistency loss and difference loss The final target loss function is calculated as follows:
[0158]
[0159] Here, α and β are trade-off parameters, with values ranging from {0.01, 0.1, 1, 3, 10}. The final objective loss function simultaneously optimizes all three loss functions mentioned above.
[0160] The brain activity signal recognition method of the present invention will be further described below with reference to specific embodiments.
[0161] Step 1: In terms of form, let This indicates input from EEG and fNIRS. These are the corresponding real tags, here. It is a one-hot output vector, where B is the batch size and C is the number of classes. The goal of DMSL is to learn a robust and efficient neural network to predict... tags
[0162] Step 2: Construct a disentangled multimodal spatiotemporal learning (DMSL) model, such as Figure 3 As shown, in the DMSL model, the Channel Reconstruction Module generates features with rich spatiotemporal patterns for each modality, the Multi-modal Attention Module comprehensively captures intermodal correlations and enhances the representation of each modality, the Multi-branch Graph Convolution Module with modal consistency and dissimilarity constraints is used for deentangled representation learning and capturing effective spatiotemporal coupling features, and the Classification Module adaptively fuses the deentangled representations and performs task prediction.
[0163] Step 3: In the channel reconstruction module, this invention constructs a single-modal feature extractor. First, this invention reconstructs the channels by adding a depth dimension, with each reconstructed channel obtaining different information from the original input. As shown in Table 1, the first two layers focus on the temporal dimension and electrode channel interactions, respectively, followed by batch normalization and ELU activation. The third layer performs average pooling, and finally, the channels are reordered. A similar operation is performed for the fNIRS sub-model. This invention designs the kernel size to ensure that the feature dimensions of fNIRS are aligned with those of EEG, thereby simplifying computation.
[0164] By using feature extractors from two sub-models, this invention obtains a pair of features (E) for two modalities. (1) ,F (1) ).here, and As the initial features extracted, they are input into the subsequent multimodal attention module, where M represents the number of time points and K represents the number of convolutional kernels.
[0165] Step 4: The multimodal attention module consists of two stages: a Feature Connectivity (CAT) layer and a Modality Attention Interaction (MAI) module. This invention introduces notation. And obtain hybrid features through the CAT layer. The present invention then embeds P into two spaces, denoted as K = LN(P) and V = LN(P) respectively, and the query for each layer is defined as Q. (l) =LN(U (l) The attention interaction of the l-th layer is defined as follows:
[0166]
[0167] Where LN represents layer normalization, l = 1, ..., N, and N is the layer number. This invention also employs a multi-head attention strategy to further enhance the diversity of representations.
[0168] The forward propagation process of layer l is represented as follows:
[0169]
[0170] Wherein, FC is the feedforward layer. Subsequently, the outputs H∈{H} of the two modes are... e H f} is represented as {H e H f}=(E (N+1) ) T ,F (n+1) ) tBy extracting complementary features from the mixed modalities, the enhanced features H learned by the MAI module have stronger representational capabilities.
[0171] Step 5: In the multi-branch graph convolutional block, enhance the feature H∈{H e H f The feature is mapped to two different GCN branches: a public GCN branch and a private GCN branch, to further explicitly extract the complex features it contains. Given the first augmented feature H... e Second enhancement feature H f Common representation of the two modalities and It can be obtained through the same common encoding function:
[0172] Z ce =C (E,F) (H e ;θ c (4)
[0173] Z cf =C (E,F) (H f ;θ c ), (5)
[0174] Among them, C (E,F) (·) is a common encoding function based on a graph convolutional network. It's important to note that these two modes in C... (E,F) (·) share the same parameter θ c .
[0175] Similarly, given H e and H f This invention uses different private encoding functions to learn private representations. and It can be obtained through different private encoding functions:
[0176]
[0177] Among them, the first private encoding function P E (·) and the second private encoding function P F (·) is also implemented using a graph convolutional network. Two private encoding functions assign separate parameters to the corresponding modes. and
[0178] The specific implementation of the graph convolutional network is as follows. The enhanced feature representation obtained after the MAI block... This invention dynamically constructs a graph structure for each sample. The associations between learning channels are studied, where the adjacency matrix A is defined as follows:
[0179]
[0180] Specifically, the ReLU activation function is deployed to ensure the non-negativity of the adjacency matrix. This represents element-wise product. This invention assumes that the connections between nodes are undirected, and that the adjacency matrix of the graph structure... It is symmetrical. A base The elements are calculated based on the dot product between the node feature vectors, and MASK is a trainable mask with the same dimension as the adjacency matrix.
[0181] Furthermore, the adjacency matrix is normalized, denoted as:
[0182]
[0183] Where D = diag(d1, d2, ..., d K Let ) be the degree matrix of A, where d m =∑ n A(m,n), where I is the identity matrix.
[0184] Using the enhanced feature matrix H and the normalized adjacency matrix The GCN layer is represented as:
[0185]
[0186] in, Let b be the weight matrix and b be the deviation vector.
[0187] Step 6: In the task classification block, this invention further combines two private features. and and two common features and These features are converted to a length of d = K × M out The vectors are then used, and an attention mechanism is employed to learn their respective importance, as follows:
[0188] (α e ,α ce ) = att(v e ,v ce ), (11)
[0189] (α f ,α cf ) = att(v f ,v cf ), (12)
[0190] Among them, v e , This indicates the flattened Ze and Z ce The normalized embedding obtained by applying L2-normalization, α e and α ce This represents the attention value between them. Similarly, α f and α cf It also follows the same principle.
[0191] This invention uses vectors Obtain attention value ω e =v e q. Similarly, this invention can also obtain vector v. ce Attention value ω ce Then, this invention uses the softmax function to adjust the attention value ω. e and ω ce Normalization is performed to obtain the final weight α. e =softmax(ω e ). α e The larger the value, the more important the corresponding embedding. Similarly, this invention can also obtain α. ce =softmax(ω ce Then, the present invention combines these two embeddings to obtain the final embedding v. Eout =α e ·v e +α ce ·v ce .
[0192] In obtaining Subsequently, the present invention performs a linear transformation to obtain class prediction:
[0193] y Eout =softmax(v Eout W e +b e ), (13) in, C is the number of categories. Similarly, given v f and v cf The present invention can obtain
[0194] The present invention then combines these two predictions to obtain the final output:
[0195]
[0196] in, This indicates element-wise addition. Here, use This represents the probability of belonging to class c.
[0197] Step 7: During the training process, this invention obtains four different feature representations of multi-branch graph convolution, as well as the final output after task classification.
[0198] For the output feature Z of private and public coding functions e and Z ce , and Z f and Z cf First, normalize the feature matrix to v. e and v ce , and v f and v cf .set up Its row corresponds to v e The matrix is given by B, where B represents the batch size. Similarly, V can be obtained. ce V f and V cf For the two output features Z of the common coding function ce and Z cf This invention uses consistency constraints to further enhance their universality:
[0199]
[0200] in It is the square of the F-norm.
[0201] Meanwhile, the output features Z of private and public coding functions of the same modality e and Z ce , and Z f and Z cf This invention applies orthogonal constraints to calculate the difference loss:
[0202]
[0203] For the task prediction loss, assuming each batch contains B samples during training, the model's prediction output is... This invention calculates the cross-entropy loss for each batch of samples and takes the average value within the batch:
[0204]
[0205] in It is the predicted probability of the b-th sample in the batch in class c. It is the smoothed true probability of the same sample and class.
[0206] The final overall objective function of DMSL combines task loss. Consistency loss and difference loss The calculation is as follows:
[0207]
[0208] Where α and β are the trade-off parameters, and their values are set to the range {0.01, 0.1, 1, 3, 10}.
[0209] It should be noted that this invention verifies the effectiveness of each modality and different components in DMSL, and ablation experiments were conducted on the MI and MA datasets. The results are shown in Table 2.
[0210] Table 2. Ablation experimental results of the leave-one-out-of-subjects method on the MI and MA datasets.
[0211]
[0212] Note: "w / o" indicates the removal of the above factors.
[0213] 1) Ablation experiments of different modes
[0214] First, this invention explores the performance of DMSL by removing the EEG and fNIRS modalities separately. When the EEG modalities are removed, the model's performance significantly decreases, indicating that the EEG modalities play a dominant role in multimodal tasks. Furthermore, compared to multimodal DMSL, unimodal DMSL consistently performs worse, further demonstrating that the model of this invention can effectively extract complementary features between EEG and fNIRS.
[0215] 2) Ablation experiments of different components
[0216] This invention designs three different DMSL variants, as follows:
[0217] DMSL w / o Phase 1: Multi-modal Attention (Phase 1) was removed from DMSL to verify the effectiveness of the multi-modal attention block.
[0218] DMSL w / o Phase 2: Multi-branch GCN (Phase 2) was removed from DMSL to verify the effect of multi-branch graph convolutional blocks.
[0219] DMSL w / o CAT: To verify the effectiveness of mixed-modal guidance, CAT was removed from DMSL.
[0220] The key improvement of the DMSL method is the addition of an improved multimodal attention module to learn fused features and the introduction of a representation learning mechanism for disentanglement to capture the complex spatiotemporal relationships inherent in the signal. Therefore, this invention conducts an ablation study on the dataset, removing the multimodal attention block (Phase 1) and the multi-branch graph convolutional module (Phase 2) respectively. It can be seen that removing Phase 1 significantly reduces the model's classification performance, further validating the effectiveness of the multi-head attention mechanism. Removing Phase 2 also results in a decrease in performance, indicating that applying disentanglement representation learning can improve feature extraction performance. Furthermore, when this invention removes the concatenation (CAT) operation of Phase 1, the performance decrease further confirms that the improved strategy of enhancing single-modal branches with mixed modalities has better guiding significance than the single-modal enhancement strategy.
[0221] To demonstrate the statistical improvement of the multimodal method over the unimodal method, this invention performs a paired t-test on DMSL.
[0222] like Figure 7 As shown in (a) and (b), in the MA dataset, the proportion of subjects whose performance was improved by the mixed-modality model compared to the single-modality EEG and fNIRS models was 55.17% (p = 0.193) and 86.21% (p < 0.001), respectively. Similarly, Figure 7 Figures (c) and (d) show that, in the MI dataset, the mixed-modality model significantly outperformed both the EEG and fNIRS models, with improved performance observed in 58.62% (p = 0.599) and 93.10% (p < 0.001), respectively. These results demonstrate that DMSL can achieve superior performance in the EEG-fNIRS mixed-modality BCI system, outperforming the single-modality BCI system. It should be noted that... Figure 7 The dots above the red diagonal in (a) to (d) show that performance was improved by fusing multimodal approaches. The percentage represents the proportion of subjects whose performance was improved by the mixed modality approach to the total number of subjects. The p-value of the paired t-test is marked in the lower right corner.
[0223] The proposed method was compared with feature fusion methods such as EF-Net, Linear Fusion, Tensor Fusion, and pth-PF. The results are shown in Table 3.
[0224] Table 3 compares the classification accuracy ± standard deviation (%) of the algorithm on the MA and MI datasets.
[0225]
[0226]
[0227] As can be seen from the data comparison in Table 3, the DMSL model proposed in this invention is significantly superior to other models.
[0228] Based on this, the present invention proposes a brain activity signal recognition method. Its proposed DMSL model, under a unified architecture, can simultaneously achieve synergistic optimization of multimodal spatiotemporal coupling and disentangled representation learning. This invention not only effectively overcomes the limitations of traditional methods in failing to fully capture the spatiotemporal coupling features and correlations between EEG and fNIRS signals, but also extracts key features of each modality through disentangled representation learning, thereby generating more refined and effective multimodal representations and significantly improving the model's expressive power. The DMSL model demonstrates significant performance advantages in mental arithmetic and motor imagery tasks. Furthermore, ablation experiments further confirm that the various components of DMSL are both indispensable and complementary, fully validating the model's effectiveness and robustness.
[0229] It should be noted that the DMSL model proposed in this invention is not only applicable to the current experimental task, but can also be extended to brain-computer interface application scenarios such as emotion state recognition and driving fatigue detection, and has broad potential and high technical value in practical applications.
[0230] In addition, such as Figure 8 As shown, one embodiment of the present invention also discloses a brain activity signal recognition device, the device comprising:
[0231] Acquisition module 110 is used to acquire the subject's EEG signal and fNIRS signal;
[0232] The recognition module 120 is used to input EEG signals and fNIRS signals into the unentangled multimodal spatiotemporal learning model to obtain the recognition results of the subject's brain activity signals. The unentangled multimodal spatiotemporal learning model is constructed based on multimodal spatiotemporal coupling and unentangled representation learning. The unentangled multimodal spatiotemporal learning model includes a channel reconstruction module, a multimodal attention module, a multi-branch graph convolution module, and a classification module. The channel reconstruction module is used to generate features with rich spatiotemporal patterns for each modality. The multimodal attention module is used to comprehensively capture the correlation between modalities and enhance the representation of each modality. The multi-branch graph convolution module is used for unentangled representation learning and capturing effective spatiotemporal coupling features. The classification module is used to adaptively fuse the unentangled representations and perform task prediction.
[0233] The brain activity signal recognition device of this invention is used to execute the brain activity signal recognition method in the above embodiments. Its specific processing procedure is the same as that of the brain activity signal recognition method in the above embodiments, and will not be described in detail here.
[0234] In addition, such as Figure 9As shown, one embodiment of the present invention also discloses an electronic device, including: at least one processor 210; at least one memory 220 for storing at least one program; and when the at least one program is executed by the at least one processor 210, implementing the brain activity signal recognition method as in any of the preceding embodiments.
[0235] In addition, one embodiment of the present invention discloses a computer-readable storage medium storing computer-executable instructions for performing the brain activity signal recognition method as described in any of the preceding embodiments.
[0236] The system architecture and application scenarios described in the embodiments of this invention are for the purpose of more clearly illustrating the technical solutions of the embodiments of this invention, and do not constitute a limitation on the technical solutions provided by the embodiments of this invention. As those skilled in the art will know, with the evolution of system architecture and the emergence of new application scenarios, the technical solutions provided by the embodiments of this invention are also applicable to similar technical problems.
[0237] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0238] In hardware implementations, the division between functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, as is known to those skilled in the art, communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0239] The terms “component,” “module,” “system,” etc., used in this specification are used to refer to computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, or a computer. As illustrated, applications running on computing devices and computing devices can both be components. One or more components may reside in a process or execution thread, and components may be located on a single computer or distributed among two or more computers. Furthermore, these components can be executed from various computer-readable media on which various data structures are stored. Components can communicate, for example, via local or remote processes based on signals having one or more data packets (e.g., data from two components interacting with another component between a local system, a distributed system, or a network, such as the Internet interacting with other systems via signals).
Claims
1. A method for recognizing brain activity signals, characterized in that, include: Acquire the subject's EEG and fNIRS signals; The EEG signal and the fNIRS signal are input into a disentangled multimodal spatiotemporal learning model to obtain the brain activity signal recognition results of the subject. The disentangled multimodal spatiotemporal learning model is constructed based on multimodal spatiotemporal coupling and disentangled representation learning. The disentangled multimodal spatiotemporal learning model includes a channel reconstruction module, a multimodal attention module, a multi-branch graph convolution module, and a classification module. The channel reconstruction module is used to generate features with rich spatiotemporal patterns for each modality. The multimodal attention module is used to comprehensively capture intermodal correlations and enhance the representation of each modality. The multi-branch graph convolution module is used for disentangled representation learning and capturing effective spatiotemporal coupling features. The classification module is used to adaptively fuse the disentangled representations and perform task prediction.
2. The brain activity signal recognition method according to claim 1, characterized in that, The step of inputting the EEG signal and the fNIRS signal into the unentangled multimodal spatiotemporal learning model to obtain the brain activity signal recognition results of the subject includes: The EEG signal and the fNIRS signal are input to the channel reconstruction module to obtain the corresponding preliminary features; The preliminary features are input into the multimodal attention module to obtain the corresponding enhanced features; The enhanced features are input into the multi-branch graph convolutional module to obtain corresponding public and private features; The public and private features are input into the classification module to obtain the brain activity signal recognition result.
3. The brain activity signal recognition method according to claim 2, characterized in that, The channel reconstruction module includes an EEG embedding block and a near-infrared embedding block. The EEG signal and the fNIRS signal are input into the channel reconstruction module to obtain corresponding preliminary features, including: The EEG signal E (0) The input is fed into the EEG embedding block, and the first preliminary feature E is extracted. (1) The EEG embedding block includes a temporal convolutional layer, a spatial convolutional layer, an average pooling layer, and a rearrangement layer. The fNIRS signal F (0) The input is fed into the near-infrared embedding block to extract the second preliminary feature F. (1) The near-infrared embedding block includes a temporal convolutional layer, a spatial convolutional layer, and a rearrangement layer.
4. The brain activity signal recognition method according to claim 3, characterized in that, The multimodal attention module includes a feature connectivity layer and a modal attention interaction block. Notation is introduced into the multimodal attention module. To express the first preliminary feature E (1) and the second preliminary feature F (1) The first preliminary feature E (1) and the second preliminary feature F (1) The input is fed into the feature connection layer to obtain hybrid features. The hybrid feature P is embedded into two spaces, denoted as K = LN(P) and V = LN(P), respectively, and the query for each layer is defined as Q. (l) =LN(U (l) The attention interaction of the l-th layer is defined as follows: Where LN represents layer normalization, l = 1, ..., N, and N is the number of layers; the attention score matrix is calculated using formula (1). To measure the degree of attention from the i-th time step in a single modality to the j-th time step in a mixed modality, where i = 1, ..., M, j = 1, ..., 2M; and to further enhance the diversity of representations by employing a multi-head attention strategy; The forward propagation process of layer l is represented as follows: Wherein, FC is the feedforward layer. Subsequently, the enhanced features H∈{H} of the two modal outputs are... e H f } is represented as {H e H f }={(E (N+1) ) T ,(F (N+1) ) T By extracting complementary features from the mixed modalities, the enhanced features H learned by the modal attention interaction block have stronger representational capabilities.
5. The brain activity signal recognition method according to claim 4, characterized in that, In the multi-branch graph convolution module, given H e and H f The common features of the two modalities are obtained through the same common encoding function. and Z ce =C (E,F) (H e ;θ c ), (4)Z cf =C (E,F) (H f ;θ c ), (5) Among them, C (E,F) (·) represents the common encoding function based on a graph convolutional network, and these two modes are in C (E,F) (·) share the same parameter θ c ; Given H e and H f Different private encoding functions are used to learn the private features of the two modalities. and Among them, P E (·) and P F Both (·) are implemented using graph convolutional networks, with two private encoding functions assigning separate parameters to the corresponding modes. and The specific implementation of the graph convolutional network is as follows: For enhanced features A graph structure was dynamically constructed for each sample. The associations between learning channels are studied, where the adjacency matrix A is defined as follows: Specifically, the ReLU activation function is deployed to ensure the non-negativity of the adjacency matrix. This represents element-wise product; it assumes that the connections between nodes are undirected and that the adjacency matrix of the graph structure is... It is symmetrical, A base The elements are calculated based on the dot product between the node feature vectors, and MASK is a trainable mask with the same dimension as the adjacency matrix. The adjacency matrix is normalized, denoted as: Where D = diag(d1, d2, ..., d K Let ) be the degree matrix of A, where d m =∑ n A(m,n), where I is the identity matrix; Using the enhanced feature matrix H and the normalized adjacency matrix The GCN layer is represented as: in, Let b be the weight matrix and b be the deviation vector.
6. The brain activity signal recognition method according to claim 5, characterized in that, In the classification module, for two private features and and two common features and Transform the two private features and two public features into a form of length d = K × M. out The vectors are used, and an attention mechanism is employed to learn their importance, as follows: (a e ,a ce )=att(v e ,v ce ), (11) (a f ,a cf )=att(v f ,v cf ), (12) in, This indicates the flattened Z e and Z ce The normalized embedding obtained by applying L2-normalization, α e and α ce This represents the attention value between them; similarly, α f and α cf It also follows the same principle; Using vectors Obtain attention value ω e =v e Similarly, we can obtain vector v. ce Attention value ω ce Then, the softmax function is used to adjust the attention value ω. e and ω ce Normalization is performed to obtain the final weight α. e =softmax(ω e ), where α e The larger the value, the more important the corresponding embedding; similarly, we can obtain α. ce =softmax(ω ce Then, combining these two embeddings, we obtain the final embedding v. Eout =α e ·v e +α ce ·v ce ; In obtaining Then, a linear transformation is performed to obtain the class prediction: y Eout =softmax(v Eout W e +b e ), (13) in, C is the number of categories; similarly, given v f and v cf ,get Then, these two predictions are added together to obtain the final output: in, This indicates element-wise addition. use This represents the probability of belonging to class c.
7. The brain activity signal recognition method according to claim 6, characterized in that, The training method for the disentangled multimodal spatiotemporal learning model includes: Construct a target loss function, which is determined based on the task loss function, consistency loss function, and difference loss function; The unentangled multimodal spatiotemporal learning model is trained based on the target loss function to obtain the trained unentangled multimodal spatiotemporal learning model; Specifically, for the task loss function, label smoothing is used to prevent model overfitting and improve generalization ability; assuming that each batch contains B samples during training, the model's predicted output is... By calculating the cross-entropy loss for each batch of samples and taking the average within each batch: in, It is the predicted probability of the b-th sample in the batch in class c. It is the smoothed true probability of the same sample and class; For the aforementioned consistency loss function, consistency constraints are used to further enhance the generality of the common encoding function output of the two modalities; let... and For each row, v ce and v cf Let B be a matrix, and let B represent the batch size. Define the following constraints: in, It is the square of the F-norm; For the aforementioned differential loss function, orthogonal constraints are applied to ensure that private and public coding functions of the same modality can effectively capture different features of the input; let... and Its row corresponds to v e and v f The matrix is given by B, where B is the batch size; the orthogonality constraint between the outputs of the private encoding function and the public encoding function is calculated by the following formula: Combined with the task loss function The consistency loss function and the difference loss function The target loss function is calculated as follows: Where α and β are trade-off parameters, and the objective loss function simultaneously optimizes the task loss function. The consistency loss function and the difference loss function 8. A brain activity signal recognition device, characterized in that, The device includes: The acquisition module is used to acquire the subject's EEG and fNIRS signals; The recognition module is used to input the EEG signal and the fNIRS signal into the unentangled multimodal spatiotemporal learning model to obtain the brain activity signal recognition result of the subject. The unentangled multimodal spatiotemporal learning model is constructed based on multimodal spatiotemporal coupling and unentangled representation learning. The unentangled multimodal spatiotemporal learning model includes a channel reconstruction module, a multimodal attention module, a multi-branch graph convolution module, and a classification module. The channel reconstruction module is used to generate features with rich spatiotemporal patterns for each modality. The multimodal attention module is used to comprehensively capture the correlation between modalities and enhance the representation of each modality. The multi-branch graph convolution module is used for unentangled representation learning and capturing effective spatiotemporal coupling features. The classification module is used to adaptively fuse the unentangled representations and perform task prediction.
9. An electronic device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, when the processor executes the computer program, it implements the brain activity signal recognition method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing computer-executable instructions for performing the brain activity signal recognition method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Behavior intention recognition method and system based on cross-modal hypergraph, and terminal
CN117828281A
EEG-fNIRS motor imagery recognition method and device based on heterogeneous graph network
CN118013352A