Brain activity signal identification method and related device

By using the de-entangled multimodal spatiotemporal learning model in the hybrid EEG-fNIRS BCI system, the space-time coupling characteristics of the signal is captured, and the problem of low recognition accuracy in existing systems is solved, and higher signal recognition accuracy is achieved.

CN120217074AActive Publication Date: 2025-06-27WUYI UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510204350.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-27
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

The existing hybrid EEG-fNIRS BCI system is difficult to effectively capture the spatial and temporal correlation between signals, resulting in low recognition accuracy.

Method used

The de-entangled multimodal spatiotemporal learning model is adopted, which includes a channel reconstruction module, a multimodal attention module, a multi-branch graph convolution module and a classification module to capture the spatiotemporal coupling characteristics and their correlations of EEG and fNIRS signals.

Benefits of technology

By generating more refined multimodal representations, the accuracy of identification of subjects' brain activity signals is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217074A_ABST
    Figure CN120217074A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a brain activity signal identification method and a related device. The method comprises the following steps: acquiring an EEG signal and an fNI RS signal of a subject; the EEG signals and the fNI RS signals are input into the deentanglement multi-mode space-time learning model, the brain activity signal recognition result of the subject is obtained, the channel reconstruction module is used for generating features with rich space-time modes for each mode, the multi-mode attention module is used for comprehensively capturing correlation between modes and enhancing representation of each mode, and the brain activity signal recognition result of the subject is obtained through the deentanglement multi-mode space-time learning model. The multi-branch graph convolution module is used for deentanglement representation learning and capturing effective space-time coupling features, and the classification module is used for adaptively fusing deentanglement representations and executing task prediction. Based on this, the embodiment of the invention can effectively capture the space-time coupling characteristics and correlation thereof of the EEG and fNI RS signals, and generate finer multi-modal representation, thereby significantly improving the accuracy of brain activity signal recognition of the subject.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the technical field of brain-computer interfaces, and in particular, to a method for identifying brain activity signals and related devices. Background Art

[0002] Brain-computer interfaces (BCIs) have greatly promoted the development of human-computer interaction by establishing direct communication between the brain and external devices. Currently, BCI technology has been widely applied in fields such as robot control, workload detection, and medical rehabilitation, and plays a crucial role especially in assisting patients with limb disabilities or other neuromuscular degenerative diseases. BCI systems can be implemented by means of various neuroimaging modalities, including stereoelectroencephalogram (SEEG), functional magnetic resonance imaging (fMRI), electroencephalogram (EEG), and functional near-infrared spectroscopy (fNIRS). Among these modalities, EEG and fNIRS have received much attention due to their non-invasiveness and ease of operation.

[0003] Electroencephalogram (EEG) records the electrophysiological activities of the brain through scalp electrodes, having good temporal resolution, but its spatial resolution is relatively limited and it is vulnerable to artifacts and noise interference. In contrast, fNIRS reflects brain activity by measuring cerebral blood flow and metabolic changes, having better spatial resolution and being less affected by noise, but its temporal resolution is relatively low. Combining these two modalities can provide more comprehensive brain activity information and effectively make up for the limitations of a single modality. Since near-infrared light does not interfere with electrical signals, this makes it possible to synchronously measure EEG and fNIRS, so hybrid EEG-fNIRS BCI systems have received extensive attention in both basic research and practical applications.

[0004] EEG and fNIRS signals originate from multiple brain regions and fluctuate over time, having complex spatio-temporal characteristics. In view of this, researchers have developed various methods to utilize their spatio-temporal information. In hybrid EEG-fNIRS BCI systems, traditional machine learning methods such as linear discriminant analysis (LDA), support vector machine (SVM), and k-nearest neighbor (K-NN) have been widely applied. Although these traditional methods have achieved certain results, they rely to a large extent on manual feature engineering, which limits their ability to effectively capture the complex spatio-temporal dynamics of EEG and fNIRS signals.

[0005] In recent years, the rise of deep learning technology has brought significant progress to hybrid EEG-fNIRS systems. Current research mostly adopts a late fusion strategy, that is, after separately extracting EEG and fNIRS features, feature fusion is carried out. However, this strategy often has difficulty capturing the potential spatio-temporal correlation between EEG and fNIRS signals, thereby resulting in limited classification performance and low recognition accuracy. Therefore, how to improve the accuracy of brain activity signal recognition for subjects has become a technical problem to be solved urgently. Summary of the Invention

[0006] The embodiments of the present invention provide a method and related device for brain activity signal recognition, which can effectively capture the spatio-temporal coupling features and their correlations of EEG and fNIRS signals, generate a more refined multi-modal representation, and thus significantly improve the accuracy of brain activity signal recognition for subjects.

[0007] In a first aspect, the embodiments of the present invention provide a method for brain activity signal recognition, including:

[0008] Obtain the EEG signal and fNIRS signal of a subject;

[0009] Input the EEG signal and the fNIRS signal into a disentangled multi-modal spatio-temporal learning model to obtain the recognition result of the brain activity signal of the subject, wherein the disentangled multi-modal spatio-temporal learning model is constructed based on multi-modal spatio-temporal coupling and disentangled representation learning, and the disentangled multi-modal spatio-temporal learning model includes a channel reconstruction module, a multi-modal attention module, a multi-branch graph convolutional module, and a classification module. The channel reconstruction module is used to generate features with rich spatio-temporal patterns for each modality, the multi-modal attention module is used to comprehensively capture the correlation between modalities and enhance the representation of each modality, the multi-branch graph convolutional module is used for disentangled representation learning and capturing effective spatio-temporal coupling features, and the classification module is used to adaptively fuse the disentangled representations and perform task prediction.

[0010] In some embodiments, the inputting the EEG signal and the fNIRS signal into the disentangled multi-modal spatio-temporal learning model to obtain the recognition result of the brain activity signal of the subject includes:

[0011] Input the EEG signal and the fNIRS signal into the channel reconstruction module to obtain corresponding preliminary features;

[0012] Input the preliminary features into the multi-modal attention module to obtain corresponding enhanced features;

[0013] Input the enhanced features into the multi-branch graph convolutional module to obtain corresponding common features and private features;

[0014] Input the common features and the private features into the classification module to obtain the brain activity signal recognition result.

[0015] In some embodiments, the channel reconstruction module includes an EEG embedding block and an fNIRS embedding block. Inputting the EEG signal and the fNIRS signal into the channel reconstruction module to obtain corresponding preliminary features includes:

[0016] Input the EEG signal E (0) into the EEG embedding block to extract the first preliminary feature E (1) , where the EEG embedding block includes a temporal convolutional layer, a spatial convolutional layer, an average pooling layer, and a rearrangement layer;

[0017] Input the fNIRS signal F (0) into the fNIRS embedding block to extract the second preliminary feature F (1) , where the fNIRS embedding block includes a temporal convolutional layer, a spatial convolutional layer, and a rearrangement layer.

[0018] In some embodiments, the multimodal attention module includes a feature connection layer and a modal attention interaction block. In the multimodal attention module, introduce a notation to represent the first preliminary feature E (1) and the second preliminary feature F (1) . Input the first preliminary feature E (1) and the second preliminary feature F (1) into the feature connection layer to obtain a mixed feature Embed the mixed feature P into two spaces, denoted as K = LN(P) and V = LN(P) respectively, and the query for each layer is defined as Q (l) = LN(U (l) ). The definition of the attention interaction at the l-th layer is as follows:

[0019]

[0020] where LN represents layer normalization, l = 1, …, N, and N is the number of layers; use the attention score matrix calculated in formula (1) to measure the attention degree of the i-th time step within a single modality pointing to the j-th time step of the mixed modality, where i = 1, …, M, j = 1, …, 2M; and adopt a multi-head attention strategy to further enhance the diversity of the representation;

[0021] The forward propagation process of the l-th layer is expressed as:

[0022]

[0023] Among them, FC is the feed-forward layer, Subsequently, the enhanced features H output by the two modalities, where H ∈ {H e , H f}, are represented as {H e , H f} = (E (N+1) ) T , F (N+1) ) T}. By extracting complementary features from the mixed modalities, the enhanced features H learned by the modality attention interaction block have stronger representation capabilities.

[0024] In some embodiments, in the multi-branch graph convolutional module, given H e and H f , the common features of the two modalities are obtained through the following same common encoding function and

[0025] Z ce = C (E,F) (H e ; θ c ), (4)

[0026] Z cf = C (E,F) (H f ; θ c ), (5)

[0027] Among them, C (E,F) (·) is a common encoding function based on a graph convolutional network, and the two modalities share the same parameter θ (E,F) in C c ;

[0028] Given H e and H f , different private encoding functions are used to learn the private features of the two modalities and

[0029]

[0030] Among them, P E (·) and P F (·) are both implemented through a graph convolutional network, and the two private encoding functions assign separate parameters and

[0031] Among them, the specific implementation of the graph convolutional network is as follows: For the enhanced feature a graph structure is dynamically constructed for each sample To learn the correlation between learning channels, where the adjacency matrix A is defined as:

[0032]

[0033] Among them, the ReLU activation function is deployed to ensure the non - negative property of the adjacency matrix. denotes the element - wise product; assuming that the connection between nodes is undirected and the adjacency matrix of the graph structure is symmetric, the elements of A base are calculated based on the dot product between node feature vectors, and MASK is a trainable mask with the same dimension as the adjacency matrix.

[0034] Normalize the adjacency matrix, denoted as:

[0035]

[0036] Among them, D = diag(d1, d2, …, d K ) is the degree matrix of A, where d m = ∑ n A(m,n), and I is the identity matrix.

[0037] Use the enhanced feature matrix H and the normalized adjacency matrix to represent the GCN layer as:

[0038]

[0039] Among them, is the weight matrix, and b is the bias vector.

[0040] In some embodiments, in the classification module, for two private features and and two public features and Convert the two private features and two public features into vectors of length d = K × M out and use the attention mechanism to learn their importance as follows:

[0041] (α e , α ce ) = att(v e , v ce ), (11)

[0042] (α f , α cf ) = att(v f , v cf ), (12)

[0043] Among them, v e , Denote the flattened Z e and Z ce as the normalized embeddings obtained by applying L2 - normalization, α e and α ce represent the attention values between them; similarly, α f and α cf also follow the same principle;

[0044] Use the vector to obtain the attention value ω e = v e q, and similarly obtain the attention value ω ce of the vector v ce ; then, use the softmax function to normalize the attention values ω e and ω ce to obtain the final weight α e = softmax(ω e ), where the larger α e , the more important the corresponding embedding is. Similarly, obtain α ce = softmax(ω ce ); then, combine these two embeddings to obtain the final embedding v Eout = α e ·v e + α ce ·v ce ;

[0045] After obtaining , further perform a linear transformation to obtain the class prediction:

[0046] y Eout = softmax(v Eout W e + b e ), (13)

[0047] where, C is the number of classes; similarly, given v f and v cf , obtain

[0048] Subsequently, add these two predictions together to obtain the final output:

[0049]

[0050] where, denotes element - wise addition, is used to represent the probability of belonging to class c.

[0051] In some embodiments, the training method of the disentangled multi-modal spatio-temporal learning model includes:

[0052] Construct a target loss function, which is determined according to a task loss function, a consistency loss function, and a difference loss function;

[0053] Train the disentangled multi-modal spatio-temporal learning model based on the target loss function to obtain the trained disentangled multi-modal spatio-temporal learning model;

[0054] Among them, for the task loss function, label smoothing is used to prevent the model from overfitting and improve the generalization ability; assuming that each batch contains B samples during the training process, the predicted output of the model is By calculating the cross-entropy loss of each batch of samples and taking the average within the batch:

[0055]

[0056] where is the predicted probability of the b-th sample in the batch on the category c, is the smoothed true probability of the same sample and category;

[0057] For the consistency loss function, a consistency constraint is used to further enhance the generality of the two outputs of the common encoding functions of the two modalities; let and be matrices whose rows correspond to v ce and v cf respectively, where B represents the batch size, and the following constraint is defined:

[0058]

[0059] where is the square of the F-norm;

[0060] For the difference loss function, an orthogonality constraint is applied to ensure that the private encoding function and the common encoding function of the same modality can effectively capture different features of the input; let and be matrices whose rows correspond to v e and v f respectively, where B is the batch size; the orthogonality constraint between the output of the private encoding function and the output of the common encoding function is calculated by the following formula:

[0061]

[0062] Combining the task loss function the consistency loss function and the difference loss function The target loss function is calculated as follows:

[0063]

[0064] where α, β are trade-off parameters, and the value range is set to {0.01, 0.1, 1, 3, 10}, and the target loss function optimizes the task loss function simultaneously. The consistency loss function and the difference loss function

[0065] In a second aspect, an embodiment of the present invention further provides a brain activity signal recognition device, including:

[0066] An acquisition module, configured to acquire the EEG signal and fNIRS signal of a subject;

[0067] An identification module, configured to input the EEG signal and the fNIRS signal into a disentangled multi-modal spatio-temporal learning model to obtain a recognition result of the brain activity signal of the subject, where the disentangled multi-modal spatio-temporal learning model is constructed based on multi-modal spatio-temporal coupling and disentangled representation learning, and the disentangled multi-modal spatio-temporal learning model includes a channel reconstruction module, a multi-modal attention module, a multi-branch graph convolution module, and a classification module. The channel reconstruction module is used to generate features with rich spatio-temporal patterns for each modality. The multi-modal attention module is used to comprehensively capture the correlation between modalities and enhance the representation of each modality. The multi-branch graph convolution module is used for disentangled representation learning and capturing effective spatio-temporal coupling features. The classification module is used to adaptively fuse the disentangled representations and perform task prediction.

[0068] In a third aspect, an embodiment of the present invention further provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that when the processor executes the computer program, the brain activity signal recognition method described in the first aspect is implemented.

[0069] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, storing computer-executable instructions, and the computer-executable instructions are used to execute the brain activity signal recognition method described in the first aspect.

[0070] According to the brain activity signal recognition method and related devices provided by the embodiments of the present invention, the brain activity signal recognition method includes: obtaining the EEG signal and fNIRS signal of a subject; inputting the EEG signal and fNIRS signal into a disentangled multi-modal spatio-temporal learning model to obtain the brain activity signal recognition result of the subject, where the disentangled multi-modal spatio-temporal learning model is constructed based on multi-modal spatio-temporal coupling and disentangled representation learning, and the disentangled multi-modal spatio-temporal learning model includes a channel reconstruction module, a multi-modal attention module, a multi-branch graph convolution module, and a classification module. The channel reconstruction module is used to generate features with rich spatio-temporal patterns for each modality, the multi-modal attention module is used to comprehensively capture the correlations between modalities and enhance the representation of each modality, the multi-branch graph convolution module is used for disentangled representation learning and capturing effective spatio-temporal coupling features, and the classification module is used to adaptively fuse the disentangled representations and perform task prediction. Based on this, the embodiments of the present invention can effectively capture the spatio-temporal coupling features and their correlations of EEG and fNIRS signals, thereby generating a more refined multi-modal representation, and thus significantly improving the accuracy of brain activity signal recognition for subjects. Description of the Drawings

[0071] Figure 1 is a flowchart of the brain activity signal recognition method provided by an embodiment of the present invention;

[0072] Figure 2 is a schematic diagram of the experimental paradigm and channel positions of the data set provided by an embodiment of the present invention;

[0073] Figure 3 is a schematic diagram of the overall framework of the disentangled multi-modal spatio-temporal learning model provided by an embodiment of the present invention;

[0074] Figure 4 is a flowchart of step S102 provided by an embodiment of the present invention;

[0075] Figure 5 is a flowchart of step S401 provided by an embodiment of the present invention;

[0076] Figure 6 is a flowchart of the training method of the disentangled multi-modal spatio-temporal learning model provided by an embodiment of the present invention;

[0077] Figure 7 is a scatter plot of the classification accuracy (x-axis) of EEG signals and fNIRS signals and the classification accuracy of the mixed modality (y-axis) provided by an embodiment of the present invention;

[0078] Figure 8 is a schematic diagram of the brain activity signal recognition device provided by an embodiment of the present invention;

[0079] Figure 9It is a schematic diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0080] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, but not to limit the present invention.

[0081] It should be noted that although functional module division is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first", "second", etc. in the specification, claims and the following drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence.

[0082] In the embodiments of the present invention, words such as "furthermore", "exemplarily" or "optionally" are used to represent examples, illustrations or explanations, and should not be construed as being more preferred or having more advantages than other embodiments or design solutions. The use of words such as "furthermore", "exemplarily" or "optionally" is intended to present relevant concepts in a specific manner.

[0083] In order to more conveniently describe the working principle of the embodiments of the present invention hereinafter, an introduction to the related technical scenarios will be given first.

[0084] In recent years, the rise of deep learning technology has brought significant progress to hybrid EEG-fNIRS systems. Current research mostly adopts a late fusion strategy, that is, after separately completing the feature extraction of EEG and fNIRS, feature fusion is then performed. However, this strategy often has difficulty capturing the potential spatio-temporal correlation between EEG and fNIRS signals, thereby resulting in limited classification performance and low recognition accuracy. Therefore, how to improve the accuracy of brain activity signal recognition for subjects has become an urgent technical problem to be solved.

[0085] Based on this, the present invention provides a method and related device for identifying brain activity signals. The method for identifying brain activity signals includes: obtaining the EEG signal and fNIRS signal of a subject; inputting the EEG signal and fNIRS signal into a disentangled multimodal spatio-temporal learning model to obtain the identification result of the subject's brain activity signal. The disentangled multimodal spatio-temporal learning model is constructed based on multimodal spatio-temporal coupling and disentangled representation learning. The disentangled multimodal spatio-temporal learning model includes a channel reconstruction module, a multimodal attention module, a multi-branch graph convolutional module, and a classification module. The channel reconstruction module is used to generate features with rich spatio-temporal patterns for each modality. The multimodal attention module is used to comprehensively capture the correlation between modalities and enhance the representation of each modality. The multi-branch graph convolutional module is used for disentangled representation learning and capturing effective spatio-temporal coupling features. The classification module is used to adaptively fuse the disentangled representations and perform task prediction. Based on this, the embodiments of the present invention can effectively capture the spatio-temporal coupling features and their correlations of EEG and fNIRS signals, thereby generating a more refined multimodal representation, and thus significantly improving the accuracy of identifying brain activity signals of the subject.

[0086] It should be noted that in each specific embodiment of the present invention, when it comes to performing relevant processing based on data related to the subject's identity or characteristics, such as subject information, subject behavior data, subject historical data, and subject location information, the permission or consent of the subject will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present invention need to obtain the personal information of the subject, the separate permission or separate consent of the subject will be obtained through methods such as collecting in a form or filling out an informed consent form. After clearly obtaining the separate permission or separate consent of the subject, the necessary subject-related data for the normal operation of the embodiments of the present invention will be obtained.

[0087] The following further elaborates on the embodiments of the present invention with reference to the accompanying drawings.

[0088] As Figure 1 shown, Figure 1 is a flowchart of a method for identifying brain activity signals provided by an embodiment of the present invention. The method for identifying brain activity signals may include but is not limited to steps S101 to S102.

[0089] Step S101: Obtain the EEG signal and fNIRS signal of the subject;

[0090] Step S102: Input the EEG signal and the fNIRS signal into the disentangled multi-modal spatio-temporal learning model to obtain the recognition result of the subject's brain activity signal. The disentangled multi-modal spatio-temporal learning model is constructed based on multi-modal spatio-temporal coupling and disentangled representation learning. The disentangled multi-modal spatio-temporal learning model includes a channel reconstruction module, a multi-modal attention module, a multi-branch graph convolutional module, and a classification module. The channel reconstruction module is used to generate features with rich spatio-temporal patterns for each modality. The multi-modal attention module is used to comprehensively capture the correlations between modalities and enhance the representation of each modality. The multi-branch graph convolutional module is used for disentangled representation learning and capturing effective spatio-temporal coupling features. The classification module is used to adaptively fuse the disentangled representations and perform task prediction.

[0091] It can be understood that in the model training stage, the present invention uses the publicly available dataset HBCI in Berlin in 2017 as the experimental data. The dataset includes 29 subjects (14 males and 15 females). The experimental paradigm is as Figure 2 shown in (a) below. The subject needs to perform 30 trials for each task. Each trial starts with a 2-second cue, followed by a 10-second task time, and ends with a randomly assigned 15- to 17-second rest period. The tasks include: baseline (rest state), mental arithmetic, left hand motor imagery, and right hand motor imagery. The present invention divides the data into two datasets according to the tasks: the MA dataset (baseline and mental arithmetic) and the MI dataset (left hand motor imagery and right hand motor imagery). EEG and fNIRS data are collected simultaneously. The channel positions are as Figure 2 shown in (b) below. 30 electrode channels (dots) record the EEG signal, and 14 light sources (solid squares) and 16 detectors (shaded squares) form 36 fNIRS channels (black solid lines) to record the fNIRS signal. The present invention further divides the EEG and fNIRS signals into 3s segments with a step size of 1s to evaluate the real-time BCI performance. Therefore, the sizes of the EEG and fNIRS signals are 30×600 (channel×time) and 72×30 (channel×time), respectively. Among them, the fNIRS data is stacked by HbO and HbR in the channel dimension. Each subject has a total of 600 samples (10 segments×30 trials×2 tasks).

[0092] It can be understood that as Figure 3As shown in the figure, the Disentangled Multimodal Spatiotemporal Learning (DMSL) model of the present invention is composed of a channel reconstruction module, a multi-modal attention module, a multi-branch graph convolutional module with consistency constraints and difference constraints, and a classification module. The disentangled multimodal spatiotemporal learning model simultaneously performs multimodal spatiotemporal coupling and disentangled representation learning within a unified architecture. Among them, the channel reconstruction module is used to generate features with rich spatiotemporal patterns for each modality; the multi-modal attention module is used to comprehensively capture the correlation between modalities and enhance the representation of each modality, enhancing the representation ability and robustness of the single modality; the multi-branch graph convolutional module is used for disentangled representation learning and capturing effective spatiotemporal coupling features to capture the complex spatiotemporal relationships inherent in the signal; the classification module is used to adaptively fuse the disentangled representations and perform task prediction.

[0093] It can be understood that, as Figure 4 shown in the figure, step S102 may include but is not limited to the following steps:

[0094] Step S401, input the EEG signal and the fNIRS signal into the channel reconstruction module to obtain corresponding preliminary features;

[0095] Step S402, input the preliminary features into the multi-modal attention module to obtain corresponding enhanced features;

[0096] Step S403, input the enhanced features into the multi-branch graph convolutional module to obtain corresponding common features and private features;

[0097] Step S404, input the common features and private features into the classification module to obtain the recognition result of the brain activity signal.

[0098] Input the EEG signal and the fNIRS signal into the disentangled multimodal spatiotemporal learning model. Specifically, by inputting the EEG signal and the fNIRS signal into the channel reconstruction module to obtain corresponding preliminary features, inputting the preliminary features into the multi-modal attention module to obtain corresponding enhanced features, inputting the enhanced features into the multi-branch graph convolutional module to obtain corresponding common features and private features, and inputting the common features and private features into the classification module to obtain the recognition result of the brain activity signal.

[0099] It can be understood that, as Figure 5As shown, the channel reconstruction module includes an EEG embedding block and an fNIRS embedding block. Step S401 may include but is not limited to the following steps:

[0100] Step S501, input the EEG signal E (0) into the EEG embedding block to extract the first preliminary feature E (1) , where the EEG embedding block includes a temporal convolutional layer, a spatial convolutional layer, an average pooling layer, and a rearrangement layer;

[0101] Step S502, input the fNIRS signal F (0) into the fNIRS embedding block to extract the second preliminary feature F (1) , where the fNIRS embedding block includes a temporal convolutional layer, a spatial convolutional layer, and a rearrangement layer.

[0102] Changes in cognitive processes in the brain are reflected in the activation levels at different timestamps and brain regions. Assume (E (0) , F (0) ) is the input from EEG and fNIRS, where, here, C (·) is the number of channels, and S (·) is the number of sampling points. The present invention reconstructs the channels by adding a dimension, and each reconstructed channel obtains different information from the original input.

[0103] The structure of the channel reconstruction module is shown in Table 1. For the EEG embedding block, the first two layers respectively focus on the temporal dimension and the interaction between electrode channels, and then batch normalization and ELUs activation are performed. The third layer performs average pooling to reduce overfitting and complexity. Finally, the features in the convolutional module are reordered by compressing and transposing the dimensions. For the fNIRS embedding block, the present invention specifically designs the size of the kernel to ensure that the feature dimension of fNIRS is aligned with that of EEG, thereby simplifying the calculation.

[0104] Table 1 Structure of the channel reconstruction module

[0105]

[0106] By using the EEG embedding block and the fNIRS embedding block as feature extractors, the present invention obtains a pair of features (E (1) , F (1) ) of two modalities. Here, and as the first and second preliminary features extracted are input into the subsequent multi-modal attention module, M represents the number of time points, and K represents the number of convolutional kernels.

[0107] It can be understood that the multi-modal attention module of the present invention includes two stages: a feature connection (CAT) layer and a modality attention interaction (MAI) block, as Figure 3 shown.

[0108] After passing through the above-mentioned channel reconstruction module, the present invention obtains the preliminary features E (1) and F (1) . To simplify the expression, the present invention introduces the notation The present invention obtains the mixed features of these two modalities through the CAT layer. Then the present invention embeds P into two spaces, denoted as K = LN(P) and V = LN(P) respectively, and the query of each layer is defined as Q (l) = LN(U (l) ). The definition of the attention interaction of the l-th layer is as follows:

[0109]

[0110] where LN represents layer normalization, l = 1,..., N, and N is the number of layers. Specifically, the attention score matrix calculated by formula (1) is used to measure the attention degree of the i-th time step in the single modality pointing to the j-th time step in the mixed modality, where i = 1,..., M and j = 1,..., 2M. The present invention also adopts a multi-head attention strategy to further enhance the diversity of the representation.

[0111] The forward propagation process of the l-th layer is expressed as:

[0112]

[0113] where FC is the feed-forward layer. Subsequently, the outputs H ∈ {H e , H f} of the two modalities are expressed as {H e , H f} = (E (N+1) ) T , F (N+1) ) T}, where H e is the first enhanced feature and H f is the second enhanced feature. By extracting complementary features from the mixed modality, the enhanced feature H learned by the MAI block has stronger representation ability.

[0114] It can be understood that the present invention further uses a multi-branch graph convolution module to explicitly extract the complex features contained in H, which enables the present invention to balance these characteristics. Given H e and Hf , the common feature representations of the two modalities are the first common feature and the second common feature which can be obtained through the same common encoding function:

[0115] Z ce = C (E,F) (H e ; θ c ), (4)

[0116] Z cf = C (E,F) (H f ; θ c ), (5)

[0117] where C (E,F) (·) is a common encoding function based on a graph convolutional network. It should be noted that the two modalities share the same parameter θ in C (E,F) . c .

[0118] Similarly, given H e and H f , the present invention uses different private encoding functions to learn private features, which are the first private feature and the second private feature which can be obtained through different private encoding functions:

[0119]

[0120] where the first private encoding function P E (·) and the second private encoding function P F (·) are also implemented through a graph convolutional network. The two private encoding functions assign separate parameters to the corresponding modalities and

[0121] The specific implementation of the graph convolutional network is as follows: For the enhanced features obtained after the MAI block the present invention dynamically constructs a graph structure for each sample to learn the associations between channels, where the adjacency matrix A is defined as:

[0122]

[0123] where the ReLU activation function is deployed to ensure the non - negative property of the adjacency matrix, denotes the element - wise product. The present invention assumes that the connections between nodes are undirected and the adjacency matrix of the graph structure is symmetric. A baseThe elements are calculated based on the dot product between the node feature vectors, and MASK is a trainable mask with the same dimension as the adjacency matrix.

[0124] Furthermore, the adjacency matrix is normalized, denoted as:

[0125]

[0126] where D = diag(d1, d2, …, d K ) is the degree matrix of A, where d m = ∑ n A(m,n), and I is the identity matrix.

[0127] Using the enhanced feature matrix H and the normalized adjacency matrix The GCN layer is represented as:

[0128]

[0129] where is the weight matrix and b is the bias vector.

[0130] It can be understood that for two private features and as well as two public features and The present invention transforms these features into vectors of length d = K × M out and uses the attention mechanism to learn their importance, as follows:

[0131] (α e , α ce ) = att(v e , v ce ), (11)

[0132] (α f , α cf ) = att(v f , v cf ), (12)

[0133] where v e , represents the normalized embeddings obtained by applying L2 - normalization to the flattened Z e and Z ce , and α e and α ce represent the attention values between them. Similarly, α f and α cf also follow the same principle.

[0134] The present invention uses the vector Obtain the attention value ω e = v e q. Similarly, the present invention can also obtain the attention value ω ce of the vector v ce . Then, the present invention uses the softmax function to normalize the attention values ω e and ω ce to obtain the final weight α e = softmax(ω e ). The larger α e , the more important the corresponding embedding is. Similarly, the present invention can also obtain α ce = softmax(ω ce ). Then, the present invention combines these two embeddings to obtain the final embedding v Eout = α e ·v e + α ce ·v ce .

[0135] After obtaining , the present invention further performs a linear transformation to obtain a class prediction:

[0136] y Eout = softmax(v Eout W e + b e ), (13)

[0137] where C is the number of classes. Similarly, given v f and v cf , the present invention can obtain

[0138] Subsequently, the present invention adds these two predictions together to obtain the final output:

[0139]

[0140] where denotes element-wise addition. Here, is used to represent the probability of belonging to class c.

[0141] It can be understood that, as Figure 6 shown, the training method of the disentangled multi-modal spatio-temporal learning model may include but is not limited to the following steps:

[0142] Step S601, construct an objective loss function, and the objective loss function is determined according to a task loss function, a consistency loss function, and a difference loss function;

[0143] Step S602: Train the disentangled multi-modal spatio-temporal learning model based on the target loss function to obtain a trained disentangled multi-modal spatio-temporal learning model.

[0144] For the training and optimization process of the disentangled multi-modal spatio-temporal learning model, by constructing a target loss function, which is determined according to the task loss function, the consistency loss function, and the difference loss function, train the disentangled multi-modal spatio-temporal learning model based on the target loss function to obtain a trained disentangled multi-modal spatio-temporal learning model.

[0145] 1) Task loss

[0146] The present invention uses label smoothing to prevent the model from overfitting and improve the generalization ability. Assume that each batch contains B samples during the training process, and the predicted output of the model is The present invention calculates the cross-entropy loss of each batch of samples and takes the average within the batch:

[0147]

[0148] where is the predicted probability of the b-th sample in the batch on the category c, is the smoothed true probability of the same sample and category.

[0149] 2) Consistency loss

[0150] The present invention uses a consistency constraint to further enhance the generality of the two outputs of the common encoding function. Let and be matrices whose rows correspond to v ce and v cf respectively, where B represents the batch size. The present invention defines the following constraint:

[0151]

[0152] where is the square of the F-norm.

[0153] 3) Difference loss

[0154] The present invention applies an orthogonality constraint to ensure that the private encoding function and the common encoding function of the same modality capture different features of the input. Let and be matrices whose rows correspond to v e and v f respectively, where B is the batch size. The orthogonality constraint between the output of the private encoding function and the output of the common encoding function is calculated by the following formula:

[0155]

[0156] 4) Overall objective function

[0157] Combined with the task loss Consistency loss and the difference loss The final objective loss function is calculated as follows:

[0158]

[0159] where α, β are trade-off parameters, and the value range is set to {0.01, 0.1, 1, 3, 10}. The final objective loss function optimizes the above three loss functions simultaneously.

[0160] The brain activity signal recognition method of the present invention will be further described below in conjunction with specific embodiments.

[0161] Step1: Formally, let represent the inputs from EEG and fNIRS, be the corresponding true labels, where is a one-hot output vector. Here, B is the batch size and C is the number of classes. The goal of DMSL is to learn a robust and effective neural network to predict the label of

[0162] Step2: Construct a disentangled multi-modal spatio-temporal learning (DMSL) model as shown in Figure 3 . In the DMSL model, the Channel reconstruction Module generates features with rich spatio-temporal patterns for each modality. The Multi-modal attention Module comprehensively captures the correlations between modalities and enhances the representation of each modality. The Multi-branch GCNModule with modality consistency and difference constraints is used for disentangled representation learning and capturing effective spatio-temporal coupling features. And the ClassificationModule adaptively fuses the disentangled representations and performs task prediction.

[0163] ​Step 3: In the channel reconstruction module, the present invention constructs a unimodal feature extractor. First, the present invention reconstructs the channels by adding a depth dimension, and each reconstructed channel obtains different information from the original input. As shown in Table 1, the first two layers respectively focus on the time dimension and the electrode channel interaction, then perform batch normalization and use ELUs for activation, the third layer performs average pooling, and finally reordering. For the fNIRS sub-model, the present invention performs similar operations. The present invention designs the kernel size to ensure that the feature dimension of fNIRS is aligned with that of EEG, thus simplifying the calculation.

[0164] Through the feature extractors of the two sub-models, the present invention obtains a pair of features of the two modalities (E (1) , F (1) ). Here, and , as the preliminary features extracted, are input into the subsequent multi-modal attention module, where M represents the number of time points and K represents the number of convolutional kernels.

[0165] Step 4: The multi-modal attention module includes two stages: one is the feature concatenation (CAT) layer and the modality attention interaction (MAI) module. The present invention introduces the notation and obtains the mixed feature through the CAT layer. Then the present invention embeds P into two spaces, denoted as K = LN(P) and V = LN(P) respectively, and the query of each layer is defined as Q (l) = LN(U (l) ). The definition of the attention interaction of the l-th layer is as follows:

[0166]

[0167] where LN represents layer normalization, l = 1, …, N, and N is the number of layers. The present invention also adopts the multi-head attention strategy to further enhance the diversity of the representation.

[0168] The forward propagation process of the l-th layer is expressed as:

[0169]

[0170] where FC is the feed-forward layer, Subsequently, the outputs of the two modalities H ∈ {H e , H f} are expressed as {H e , H f} = (E (N+1) ) T , F (n+1) ) t}。By extracting complementary features from the mixed modality, the enhanced feature H learned by the MAI module has stronger representational ability.

[0171] Step5: In the multi-branch graph convolutional block, the enhanced feature H ∈ {H e , H f} is mapped to two different GCN branches: the common GCN branch and the private GCN branch to further explicitly extract the complex features contained therein. Given the first enhanced feature H e and the second enhanced feature H f , the common representations and of the two modalities can be obtained through the same common encoding function:

[0172] Z ce = C (E,F) (H e ; θ c ), (4)

[0173] Z cf = C (E,F) (H f ; θ c ), (5)

[0174] where C (E,F) (·) is the common encoding function based on a graph convolutional network. It should be noted that the two modalities share the same parameter θ (E,F) in C c .

[0175] Similarly, given H e and H f , the present invention uses different private encoding functions to learn the private representations, and which can be obtained through different private encoding functions:

[0176]

[0177] where the first private encoding function P E (·) and the second private encoding function P F (·) are also implemented through a graph convolutional network. The two private encoding functions assign separate parameters and

[0178] The specific implementation of the graph convolutional network is as follows. For the enhanced feature representation obtained after the MAI block The present invention dynamically constructs a graph structure for each sample to learn the association between channels, where the adjacency matrix A is defined as:

[0179]

[0180] Among them, the ReLU activation function is deployed to ensure the non - negative property of the adjacency matrix. denotes the element - wise product. The present invention assumes that the connections between nodes are undirected, and the adjacency matrix of the graph structure is symmetric. The elements of A base are calculated based on the dot product between node feature vectors, and MASK is a trainable mask with the same dimension as the adjacency matrix.

[0181] Furthermore, the adjacency matrix is normalized, denoted as:

[0182]

[0183] where D = diag(d1, d2, …, d K ) is the degree matrix of A, where d m = ∑ n A(m,n), and I is the identity matrix.

[0184] The GCN layer is represented using the enhanced feature matrix H and the normalized adjacency matrix as:

[0185]

[0186] where is the weight matrix and b is the bias vector.

[0187] Step6: In the task classification block, the present invention further combines two private features and as well as two public features and to transform these features into vectors of length d = K × M out and uses the attention mechanism to learn their corresponding importance, as follows:

[0188] (α e , α ce ) = att(v e , v ce ), (11)

[0189] (α f , α cf ) = att(v f , v cf ), (12)

[0190] where v e , denotes the flattened Ze and Z ce The normalized embedding obtained by applying L2 - normalization, α e and α ce represent the attention values between them. Similarly, α f and α cf also follow the same principle.

[0191] The present invention uses the vector to obtain the attention value ω e = v e q. Similarly, the present invention can also obtain the attention value ω ce of the vector v ce . Then, the present invention uses the softmax function to normalize the attention values ω e and ω ce to obtain the final weight α e = softmax(ω e ). The larger α e , the more important the corresponding embedding is. Similarly, the present invention can also obtain α ce = softmax(ω ce ). Then, the present invention combines these two embeddings to obtain the final embedding v Eout = α e ·v e + α ce ·v ce .

[0192] After obtaining , the present invention further performs a linear transformation to obtain the class prediction:

[0193] y Eout = softmax(v Eout W e + b e ), (13) where, C is the number of classes. Similarly, given v f and v cf , the present invention can obtain

[0194] Subsequently, the present invention adds these two predictions together to obtain the final output:

[0195]

[0196] where, represents element - wise addition. Here, is used to represent the probability of belonging to class c.

[0197] Step 7: During the training process, the present invention obtains four different feature representations of the multi-branch graph convolution and the final output after task classification.

[0198] For the output features Z of the private and public encoding functions e and Z ce , and Z f and Z cf , first normalize the feature matrices to v e and v ce , and v f and v cf . Let be the matrix whose rows correspond to v e , where B represents the batch size. Similarly, we can obtain V ce , V f and V cf . For the two output features Z ce and Z cf of the public encoding function, the present invention uses a consistency constraint to further enhance their generality:

[0199]

[0200] where is the square of the F-norm.

[0201] Meanwhile, for the output features Z e and Z ce of the private and public encoding functions for the same modality, as well as Z f and Z cf , the present invention applies an orthogonality constraint to calculate the difference loss:

[0202]

[0203] For the task prediction loss, assume that each batch during the training process contains B samples, and the predicted output of the model is The present invention calculates the cross-entropy loss for each batch of samples and takes the average within the batch:

[0204]

[0205] where is the predicted probability of the b-th sample in the batch for class c, is the smoothed true probability for the same sample and class.

[0206] Finally, the overall objective function of DMSL combines the task loss the consistency loss and the difference loss and is calculated as:

[0207]

[0208] Where α and β are trade-off parameters, and their value ranges are set to {0.01, 0.1, 1, 3, 10}.

[0209] It should be noted that the present invention verifies the effectiveness of each modality and different components in DMSL, and ablation experiments are carried out on the MI and MA datasets, and the results are shown in Table 2.

[0210] Table 2 Ablation experiment results of the leave-one-subject verification method on the MI and MA datasets

[0211]

[0212] Note: "w / o" means removing the above factors.

[0213] 1) Ablation experiments of different modalities

[0214] First of all, the present invention removes the EEG and fNIRS modalities respectively to explore the performance of DMSL. When the EEG modality is removed, the performance of the model drops significantly, indicating that the EEG modality plays a leading role in the multi-modal task. In addition, compared with the multi-modal DMSL, the performance of the single-modal DMSL is always worse, further indicating that the model of the present invention can effectively extract complementary features between EEG and fNIRS.

[0215] 2) Ablation experiments of different components

[0216] The present invention designs three different DMSL variants as follows:

[0217] DMSL w / o Phase 1: Remove the Multi-modal Attention (Phase 1) from DMSL to verify the effect of the multi-modal attention block.

[0218] DMSL w / o Phase 2: Remove the Multi-branch GCN (Phase 2) from DMSL to verify the effect of the multi-branch graph convolution block.

[0219] DMSL w / o CAT: In order to verify the effect of guiding with mixed modalities, remove CAT from DMSL.

[0220] The key improvement of the DMSL method is to add an improved multi-modal attention module to learn fused features and introduce a representation learning for disentanglement to capture the complex spatio-temporal relationships inherent in the signals. Therefore, the present invention conducts an ablation study on the dataset, removing the multi-modal attention block (Phase 1) and the multi-branch graph convolution module (Phase 2) respectively. It can be seen that when Phase 1 is removed, the classification performance of the model drops significantly, further verifying the effectiveness of the multi-head attention mechanism. When Phase 2 is removed, the experimental results also decline, indicating that the application of disentangled representation learning can improve the performance of feature extraction. In addition, when the concatenation (CAT) operation of Phase 1 is removed respectively in the present invention, the further decline in performance further confirms that the strategy of the present invention to enhance the single-modal branch with mixed modalities has better guidance than the single-modal enhancement strategy.

[0221] To prove the statistical improvement of the multi-modal method over the single-modal method, the present invention conducts a paired t-test on DMSL.

[0222] As Figure 7 shown in (a) and (b) therein, in the MA dataset, compared with the single-modal EEG and fNIRS models, the proportions of subjects with improved performance of the mixed-modal models are 55.17% (p = 0.193) and 86.21% (p < 0.001) respectively. Similarly, Figure 7 shown in (c) and (d) therein, in the MI dataset, the performance of the mixed-modal models is significantly better than that of the EEG and fNIRS models, and the proportions of subjects with improved performance are 58.62% (p = 0.599) and 93.10% (p < 0.001) respectively. These results indicate that DMSL can achieve superior performance in the EEG-fNIRS hybrid BCI system, outperforming the single-modal BCI system. It should be noted that Figure 7 the dots above the red diagonal line in (a) to (d) therein improve the performance by fusing multi-modalities, and the percentage represents the ratio of the number of subjects improved by the mixed-modal method to the total number of subjects, and the p-values of the paired t-test are marked in the lower right corner.

[0223] The present invention compares the performance of the proposed method with feature fusion methods such as EF-Net, Linear Fusion, Tensor Fusion, pth-PF, etc., and the results are shown in Table 3.

[0224] Table 3 Classification accuracies ± standard deviations (%) of the comparison algorithms in the MA and MI datasets

[0225]

[0226]

[0227] As can be seen from the data comparison in Table 3, the DMSL model proposed by the present invention is significantly superior to other models.

[0228] Based on this, a method for identifying brain activity signals proposed by the present invention, the proposed DMSL model can simultaneously achieve the collaborative optimization of multi-modal spatio-temporal coupling and disentangled representation learning under a unified architecture. The present invention not only effectively overcomes the limitation that traditional methods fail to fully capture the spatio-temporal coupling characteristics and correlations between EEG and fNIRS signals, but also extracts the key features of each modality through disentangled representation learning, thereby generating a more refined and effective multi-modal representation, significantly enhancing the expressive power of the model. The DMSL model shows significant performance advantages in mental arithmetic tasks and motor imagery tasks. In addition, ablation experiments further confirm that each component of DMSL is both indispensable and complementary, fully verifying the effectiveness and robustness of the model.

[0229] It should be noted that the DMSL model proposed by the present invention is not only applicable to the current experimental tasks, but also can be extended to brain-computer interface application scenarios such as emotion state recognition and driving fatigue detection, and has broad potential and high technical value in practical applications.

[0230] In addition, as Figure 8 shown, an embodiment of the present invention also discloses a device for identifying brain activity signals, which includes:

[0231] An acquisition module 110, configured to acquire EEG signals and fNIRS signals of a subject;

[0232] An identification module 120, configured to input the EEG signals and fNIRS signals into a disentangled multi-modal spatio-temporal learning model to obtain an identification result of the brain activity signals of the subject, wherein the disentangled multi-modal spatio-temporal learning model is constructed based on multi-modal spatio-temporal coupling and disentangled representation learning, and the disentangled multi-modal spatio-temporal learning model includes a channel reconstruction module, a multi-modal attention module, a multi-branch graph convolution module, and a classification module. The channel reconstruction module is used to generate features with rich spatio-temporal patterns for each modality, the multi-modal attention module is used to comprehensively capture the correlations between modalities and enhance the representation of each modality, the multi-branch graph convolution module is used for disentangled representation learning and capturing effective spatio-temporal coupling features, and the classification module is used to adaptively fuse the disentangled representations and perform task prediction.

[0233] The device for identifying brain activity signals in the embodiment of the present invention is used to execute the method for identifying brain activity signals in the above embodiment, and its specific processing process is the same as that of the method for identifying brain activity signals in the above embodiment, and will not be elaborated here one by one.

[0234] In addition, as Figure 9As shown in the figure, an embodiment of the present invention also discloses an electronic device, including: at least one processor 210; at least one memory 220 for storing at least one program; when the at least one program is executed by the at least one processor 210, the brain activity signal recognition method in any of the previous embodiments is implemented.

[0235] In addition, an embodiment of the present invention also discloses a computer-readable storage medium, in which computer-executable instructions are stored, and the computer-executable instructions are used to execute the brain activity signal recognition method in any of the previous embodiments.

[0236] The system architecture and application scenarios described in the embodiments of the present invention are for more clearly explaining the technical solutions of the embodiments of the present invention, and do not constitute a limitation on the technical solutions provided by the embodiments of the present invention. Those skilled in the art know that with the evolution of the system architecture and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present invention are equally applicable to similar technical problems.

[0237] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.

[0238] In the hardware implementation, the division between the functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component can have multiple functions, or a function or step can be executed by several physical components in cooperation. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or can be implemented as hardware, or can be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory, or other memory technologies, CD-ROM, digital versatile disk (DVD), or other optical disk storage, magnetic cassette, tape, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.

[0239] As used in this specification, the terms "component", "module", "system", etc. are used to refer to computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable, an execution thread, a program, or a computer. By way of illustration, both an application running on a computing device and the computing device can be components. One or more components can reside within a process or execution thread, and a component can be located on one computer or distributed between two or more computers. Further, these components can execute from various computer-readable media having various data structures stored thereon. A component can, for example, communicate by way of signals with other systems, such as according to one or more data packets (e.g., data from two components interacting with another component from a local system, a distributed system, or a network, such as the Internet).

Claims

1. A method for identifying brain activity signals, characterized in that: include: Acquire the subject's EEG signal and fNIRS signal; The EEG signal and the fNIRS signal are input into a disentangled multimodal spatiotemporal learning model to obtain a brain activity signal recognition result of the subject, wherein the disentangled multimodal spatiotemporal learning model is constructed based on multimodal spatiotemporal coupling and disentangled representation learning, and the disentangled multimodal spatiotemporal learning model includes a channel reconstruction module, a multimodal attention module, a multi-branch graph convolution module and a classification module, the channel reconstruction module is used to generate features with rich spatiotemporal patterns for each modality, the multimodal attention module is used to comprehensively capture the correlation between modalities and enhance the representation of each modality, the multi-branch graph convolution module is used for disentangled representation learning and capturing effective spatiotemporal coupling features, and the classification module is used to adaptively fuse disentangled representations and perform task prediction.

2. The brain activity signal recognition method according to claim 1, characterized in that: The step of inputting the EEG signal and the fNIRS signal into the disentangled multimodal spatiotemporal learning model to obtain a brain activity signal recognition result of the subject includes: Inputting the EEG signal and the fNIRS signal into the channel reconstruction module to obtain corresponding preliminary features; Inputting the preliminary features into the multimodal attention module to obtain corresponding enhanced features; Inputting the enhanced features into the multi-branch graph convolution module to obtain corresponding public features and private features; The public features and the private features are input into the classification module to obtain the brain activity signal recognition result.

3. The brain activity signal recognition method according to claim 2, characterized in that: The channel reconstruction module includes an EEG embedding block and a near infrared embedding block. The EEG signal and the fNIRS signal are input into the channel reconstruction module to obtain corresponding preliminary features, including: The EEG signal E (0) Input to the EEG embedding block and extract the first preliminary feature E (1) , wherein the EEG embedding block includes a temporal convolution layer, a spatial convolution layer, an average pooling layer and a rearrangement layer; The fNIRS signal F (0) Input to the near infrared embedding block and extract the second preliminary feature F (1) , wherein the near-infrared embedding block includes a temporal convolution layer, a spatial convolution layer and a rearrangement layer.

4. The method for identifying brain activity signals according to claim 3, characterized in that: The multimodal attention module includes a feature connection layer and a modal attention interaction block. In the multimodal attention module, the notation To express the first preliminary feature E (1) and the second preliminary feature F (1) , the first preliminary feature E (1) and the second preliminary feature F (1) The input is sent to the feature connection layer to obtain the mixed feature The hybrid feature P is embedded into two spaces, denoted as K = LN(P) and V = LN(P), and the query of each layer is defined as Q (l) =LN(U (l) ), the attention interaction of layer l is defined as follows: Where LN represents layer normalization, l = 1, ..., N, N is the number of layers; the attention score matrix calculated using formula (1) To measure the degree of attention from the i-th time step in a single modality to the j-th time step in a mixed modality, where i = 1, ..., M, j = 1, ..., 2M; and a multi-head attention strategy is used to further enhance the diversity of representation; The forward propagation process of the lth layer is expressed as: Among them, FC is the feed-forward layer, Then, the enhanced features H∈{H e ,H f } is represented by {H e ,H f }={(E (N+1) ) T ,(F (N+1) ) T }, by extracting complementary features from the mixed modalities, the enhanced feature H learned by the modal attention interaction block has stronger representation ability.

5. The brain activity signal recognition method according to claim 4, characterized in that: In the multi-branch graph convolution module, given H e and H f , the common features of the two modalities are obtained through the same common encoding function as follows and Z ce =C (E,F) (H e ;θ c ), (4)Z cf =C (E,F) (H f ;θ c ), (5) Among them, C (E,F) (·) is a common encoding function based on a graph convolutional network, and these two modes are in C (E,F) (·) share the same parameter θ c ; Given H e and H f , using different private encoding functions to learn private features of the two modalities and Among them, P E (·) and P F (·) are both implemented through graph convolutional networks, and the two private encoding functions assign separate parameters to the corresponding modes. and Among them, the specific implementation of the graph convolutional network is as follows: For enhanced features A graph structure is dynamically constructed for each sample To learn the association between channels, the adjacency matrix A is defined as: Among them, the ReLU activation function is deployed to ensure the non-negative property of the adjacency matrix. represents the element-wise product; it is assumed that the connections between nodes are undirected and the adjacency matrix of the graph structure is It is symmetrical. base The elements of are calculated based on the dot product between the node feature vectors, and MASK is a trainable mask with the same dimension as the adjacency matrix; Normalize the adjacency matrix and record it as: Where D = diag(d1, d2, ..., d K ) is the degree matrix of A, where d m =∑ n A(m,n), I is the identity matrix; Using the enhanced feature matrix H and the normalized adjacency matrix The GCN layer is represented as: in, is the weight matrix and b is the bias vector.

6. The brain activity signal recognition method according to claim 5, characterized in that: In the classification module, for two private features and And two common features and Convert two private features and two public features into a length of d = K × M out vectors and use the attention mechanism to learn their importance as follows: (a e ,a ce )=att(v e ,v ce ), (11) (a f ,a cf )=att(v f ,v cf ), (12) in, Represents the flattened Z e and Z ce The normalized embedding obtained by applying L2-normalization, α e and α ce represents the attention value between them; similarly, α f and α cf The same principles apply; Using Vectors Get the attention value ω e =v e q, and similarly we get vector v ce The attention value ω ce ; Then, use the softmax function to adjust the attention value ω e and ω ce Normalize and get the final weight α e =softmax(ω e ), where α e The larger the value, the more important the corresponding embedding is. Similarly, we can get α ce =softmax(ω ce ); Then, combine these two embeddings to get the final embedding v Eout =α e ·v e +α ce ·v ce ; In obtaining After that, further linear transformation is performed to obtain class prediction: y Eout =softmax(v Eout W e +b e ), (13) in, C is the number of categories; similarly, given v f and v cf ,get These two predictions are then added together to get the final output: in, represents element-by-element addition, use represents the probability of belonging to class c.

7. The method for identifying brain activity signals according to claim 6, characterized in that: The training method of the disentangled multimodal spatiotemporal learning model includes: Constructing a target loss function, wherein the target loss function is determined according to a task loss function, a consistency loss function, and a difference loss function; Training the disentangled multimodal spatiotemporal learning model based on the target loss function to obtain the trained disentangled multimodal spatiotemporal learning model; Among them, for the task loss function, label smoothing is used to prevent model overfitting and improve generalization ability; assuming that each batch contains B samples during training, the predicted output of the model is By calculating the cross entropy loss for each batch of samples and taking the average within the batch: in, is the predicted probability of the bth sample in the batch for category c, is the smoothed true probability of the same sample and category; For the consistency loss function, a consistency constraint is used to further enhance the universality of the common encoding function output of the two modalities; let and Its rows correspond to v ce and v cf A matrix of , where B represents the batch size, defines the following constraints: in, is the square of the F-norm; For the difference loss function, an orthogonal constraint is applied to ensure that the private encoding function and the public encoding function of the same modality can effectively capture different features of the input; let and For its row corresponding to v e and v f where B is the batch size; the orthogonality constraint between the private encoding function output and the public encoding function output is calculated as follows: Combined with the task loss function The consistency loss function And the difference loss function The objective loss function is calculated as: Among them, α, β are trade-off parameters, and the target loss function also optimizes the task loss function The consistency loss function And the difference loss function 8. A brain activity signal recognition device, characterized in that: The device comprises: An acquisition module, used to acquire the EEG signal and fNIRS signal of the subject; A recognition module is used to input the EEG signal and the fNIRS signal into a disentangled multimodal spatiotemporal learning model to obtain a recognition result of the subject's brain activity signal, wherein the disentangled multimodal spatiotemporal learning model is constructed based on multimodal spatiotemporal coupling and disentangled representation learning, and the disentangled multimodal spatiotemporal learning model includes a channel reconstruction module, a multimodal attention module, a multi-branch graph convolution module and a classification module, the channel reconstruction module is used to generate features with rich spatiotemporal patterns for each modality, the multimodal attention module is used to comprehensively capture the correlation between modalities and enhance the representation of each modality, the multi-branch graph convolution module is used for disentangled representation learning and capturing effective spatiotemporal coupling features, and the classification module is used to adaptively fuse disentangled representations and perform task prediction.

9. An electronic device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the brain activity signal recognition method as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to execute the brain activity signal recognition method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Behavior intention recognition method and system based on cross-modal hypergraph, and terminal

    CN117828281A

  • EEG-fNIRS motor imagery recognition method and device based on heterogeneous graph network

    CN118013352A

  • Multi-modal sentiment analysis method and system based on fine-grained semantic decomposition network

    CN118378212A

  • Neural biometric authentication system and method

    WO2025003473A1