A deep learning embedding network for decoding brain signals
By adopting a deep learning embedding network based on graph random walk and physical position embedding in brain signal decoding, combined with adaptive feature fusion and multi-task Transformer classification model, the problems of insufficient model depth and insufficient data enhancement are solved, significantly improving the characterization and generalization of the model, and achieving more accurate brain signal decoding.
Patent Information
- Application Number
- CN202410595776.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-05-14
AI Technical Summary
The lack of model depth in brain signal decoding results in poor generalization ability, which is unable to fully capture the complex relationships and linear and nonlinear relationships of brain networks. The data augmentation method fails to make full use of the deep information of brain region location and mutual relationships, limiting the representation ability of the model and the sharing of cross-site learning.
The physical enhancement model based on construction graph random walk and physical position embedding is adopted, combined with adaptive feature fusion and multi-task Transformer classification model, node embedding is obtained through unsupervised learning, simulate the diffusion process of brain neural information, and improve the characterization and generalization of the model through innovative position coding and soft parameter sharing mechanisms.
It significantly improves the model's decoding ability of brain signals, enhances characterization and generalization, can more accurately capture the complex relationships and multi-task information of brain networks, and improves the applicability and robustness of the model on different data sets.
Smart Images

Figure CN118468135B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of brain signal processing, and in particular to a deep learning embedding network for decoding brain signals. Background Art
[0002] In terms of brain signal decoding, the complexity of brain networks and the scarcity of task-specific data pose great challenges, leading to two major problems: 1. Insufficient model depth leads to poor generalization ability, and 2. Inability to fully capture linear and nonlinear relationships, weakening representation capabilities. The current popular strategy is to use data augmentation to enrich the information content of limited data sets, and then use neural networks to learn from these augmented data. This approach aims to address the limitations imposed by data scarcity and network complexity, with the goal of improving model performance and understanding of auditory processing tasks.
[0003] In the prior art, in the Transformer-Encoder model using large-scale multi-site resting-state fMRI data, a large number of samples from multiple sites are used to build an end-to-end model to distinguish patients with major depression. The model is mainly divided into three parts: processing and feature extraction of multi-site heterogeneous data, an end-to-end transformer-encoder model, and a pre-training module using reconstruction loss. First, since the data structure of multi-site data may be different, a mask mechanism is adopted, and the mask is used to fill the tail elements with arbitrary values, and the functional connection matrix is calculated as a feature through Pearson correlation; secondly, the encoder part of the transformer technology is adopted, and multi-head attention is used to obtain long and short sequence dependencies, and sinusoidal position encoding is used to allow the model to capture position information. In addition, the model removes the decoder part to reduce the complexity of the model; finally, the pre-training module initializes the transformer-encoder model under unsupervised conditions. First, the functional connection matrix X is randomly added with a mask in the ratio of r, r∈(0,1) to obtain Then by linear projection: As the input of the model, the reconstruction loss is finally calculated on the masked elements: To capture valuable features and patterns in high-dimensional representations, thereby accelerating the model training process and improving model performance.
[0004] However, the prior art mainly has the following deficiencies:
[0005] 1. In the process of data augmentation, the model mainly relies on the strategy of repeated slicing. Although this approach increases the amount of data, it does not actually introduce new information, and does not fully explore and utilize the deep information of the location and relationship of brain regions, resulting in relatively limited knowledge captured from the data.
[0006] 2. Although the model utilizes data resources from multiple sites, it does not share learning parameters across sites, which limits its ability to represent learning, thereby affecting the generalization performance of the model and limiting its applicability and robustness on different datasets. Summary of the invention
[0007] The purpose of the present invention is to provide a deep learning embedding network for decoding brain signals, which solves the technical problem of capturing the complex relationship of brain networks in brain signal decoding and improving the representation and generalization of the model.
[0008] To achieve the above object, the present invention provides the following technical solutions:
[0009] A deep learning embedding network for decoding brain signals, including a physically enhanced model based on building graph random walks and physical location embedding, and a classification model based on adaptive feature fusion and multi-task transformer;
[0010] The physical enhancement model is used to simulate the diffusion process of brain neural information: by constructing a graph random walk and unsupervised learning, the node embedding X of all brain regions in the graph structure is obtained; on the other hand, two position codes are defined, the two position codes are combined into three-dimensional coordinates to form a final position code X′, and the final position code X′ is merged into the node embedding X as the output of the physical enhancement model;
[0011] On the one hand, the classification model uses adaptive feature fusion, that is, by adaptively adjusting the weight of each node (based on the degree of the node), to fuse node features in multiple rounds of propagation and achieve dimensionality upgrade; on the other hand, through the multi-task Transformer method, the node features after dimensionality upgrade are fully used to learn the activity relationship between tasks, and then the node representations of neighbors are fused through the attention mechanism. Finally, a probabilistic representation is obtained for judging categories, which significantly improves the quality of node representation.
[0012] Preferably, the graph structure includes nodes, node features and edge connections, and the nodes are the brain regions; the graph structure is obtained by the following method:
[0013] The BOLD time series of each brain region of the subjects under the corresponding task obtained by functional magnetic resonance imaging are regarded as node features;
[0014] Based on the BOLD sequence, edge connections were constructed between two brain regions (nodes) using the Pearson correlation coefficient.
[0015] Preferably, the step of obtaining the node embedding X of all brain regions in the graph structure by constructing a graph random walk and unsupervised learning comprises the following steps:
[0016] First, the feature learning framework is defined. The purpose is to learn the mapping function by maximizing the co-occurrence probability of nodes in the sequence. By adjusting the parameters p and q, a balance is made between new nodes and returning old nodes according to the tightness of the network itself. The objective function is defined as follows:
[0017]
[0018] Among them, the conditional probability is defined as:
[0019]
[0020] Node u j Move to the next node v j The probability is:
[0021]
[0022] Finally, embed the nodes of all brain regions into X∈R Num(V)×d Carry out the next step of graph analysis task;
[0023] In the formula, f j :V j →R d , f j is the mapping function that needs to be learned, through which nodes can be mapped to node embeddings; V j is the node of the jth person; is the set of nodes sampled by the Alias algorithm; n is the nth node sampled by the Alias algorithm; α pq is a node; For node u j and v j The connection weight between them; For node u j And the connection weight of the node sampled by the nth Alias algorithm.
[0024] Preferably, the two position codes include a first position code and a second position code, wherein the first position code is a position code based on node centrality:
[0025]
[0026] Where V is the set of all nodes; N s (u) is the set of adjacent nodes of the node under the sampling of the Alias algorithm; Num is the number of nodes in the set.
[0027] Preferably, the second position encoding utilizes the attraction and repulsion between two nodes to update the position coordinates of the nodes through iteration, comprising the following steps:
[0028] First, input the weight matrix Randomly initialize the position of each node and define the coefficient
[0029] Secondly, for each pair of nodes (u, v), iterate as follows:
[0030] The position shift under the repulsive force is recorded as:
[0031]
[0032] The position displacement under the action of attraction is recorded as:
[0033]
[0034] The resulting position update is recorded as:
[0035]
[0036] The repulsive force is defined as:
[0037]
[0038] Attraction is defined as:
[0039]
[0040] Finally, when the number of iterations reaches the maximum value, the coordinates of each node at this time are returned and combined with the first position code to form the final position code:
[0041] X′=X||{PE u} u∈V ;
[0042] Where k is a predefined parameter; w uv is the correlation between two nodes; area is the size of the canvas for physical coordinate layout, V is the set of all nodes; pos(u) is the position of the u node; pos(v) is the position of the v node; t is a parameter that simulates "temperature" and is used to control the moving step of the node in each iteration; PE u Encode the position of the obtained node.
[0043] Preferably, the method for achieving node dimension increase through adaptive feature fusion is:
[0044] The expression for upgrading the final position code X′ is:
[0045] X′ (p+1) =W×(X′ (p) ⊙W′)+X′ (p) ;
[0046] in,
[0047]
[0048] Each time the new embedding is obtained, it will be arranged in rows with the original embedding, and finally the dimensionality-enhanced embedding is obtained:
[0049] S,S∈M N,p+1,d+3 ;
[0050] Where p is the number of rounds, W i ' i is the weight matrix defined based on the node degree, D ii is the degree of node i; ε is a constant; M N,p+1,d+3 It is a three-dimensional matrix of (N, p+1, d+3); N is the number of summary points; d is the embedding length of the final position code X′.
[0051] Preferably, the attention mechanism fuses the node representations of neighbors and finally obtains a probability representation for class determination, comprising the following steps:
[0052] First, an encoder layer is constructed, which includes a parameter transformation layer, a multi-head self-attention mechanism, a feedforward neural network, etc.
[0053] Secondly, a soft parameter sharing mechanism is introduced through the parameter transformation layer to fully learn the activities between tasks and share parameter representations;
[0054] Finally, through the multi-head self-attention mechanism, the complex relationships between complex features are captured according to the attention scores.
[0055] Preferably, the constructing the encoder layer comprises the following steps:
[0056] First, build and apply the multi-head attention mechanism:
[0057] MultiHead(Q,K,V)=Caoncat(head 1 ,…,head n )H O ;
[0058] in,
[0059]
[0060] In the formula, Q=SH Q ,K=SH K ,V=SH V ,H Q ,H K ,H Vis a mapping matrix, which can be obtained through model learning and training; Q, K, V are query vectors, key vectors, and value vectors respectively; H O is S, which is the first layer input of the encoder; S is the embedding after dimensionality increase;
[0061] Second, construct the parameter transformation layer; set the parameter used by the encoder to Θ, then: Θ′ i =W i Θ+b i ;
[0062] Θ′ i is the parameter after the transformation of parameter Θ, which is shared by all the multi-head attention in the encoder layer, so the parameters shared by each layer can be calculated:
[0063]
[0064] Where i is the number of encoder layers; W i is the learned weight matrix; b i is the learned bias; Q i , K i 、V i are the query vector, key vector, and value vector of each layer respectively; are the mapping matrices corresponding to each layer.
[0065] Third, construct a feedforward neural network; the encoders all contain a feedforward neural network, which consists of two linear transformation layers and a nonlinear activation function, combined after a multi-head attention mechanism:
[0066] FFN(x)=max(0,xW 1 +b 1 )W 2 +b 2 ;
[0067] Where x is the output of multi-head attention; W 1 ,b 1 and W 2 ,b 2 are the weight matrix and bias of the two-layer linear transformation respectively; max{·} is a nonlinear transformation that only takes the positive output representation (if it is negative, the output is zero); FFN(x) is the feedforward neural network expression.
[0068] Compared with the prior art, the present invention has the following beneficial effects:
[0069] The present invention innovatively uses the random walk method to simulate the neural signal transmission process in the brain, obtains the embedded representation of the node through an unsupervised method, and then innovatively proposes the second position encoding (NeuroSyncFruchterman-Reingold) method to simulate the physical coordinates of the brain area based on correlation constraints; finally, a soft parameter sharing mechanism is introduced to calculate attention by sharing the parameters used in each coding layer, thereby improving the model's utilization of multiple task information. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 It is a schematic diagram of the processing flow of the present invention;
[0071] Figure 2 This is a schematic diagram of feature embedding after enhancement according to an embodiment of the present invention;
[0072] Figure 3 Embed the schematic diagram for the original features; DETAILED DESCRIPTION
[0073] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0074] See also Figure 1-Figure 3 , a deep learning embedding network for decoding brain signals, including a physically enhanced model based on building graph random walks and physical location embedding, and a classification model based on adaptive feature fusion and multi-task transformer;
[0075] The brain signal comes from a graph structure, which is mainly composed of nodes, node features, and edge connections. The nodes are the brain regions. The graph structure is obtained by the following method:
[0076] The BOLD time series of each brain region of the subjects under the corresponding task obtained by functional magnetic resonance imaging technology are regarded as node features;
[0077] Based on the BOLD sequence, edge connections are constructed between two brain regions (nodes) through the Pearson correlation coefficient. Here, this embodiment only retains edges with correlation coefficients greater than or equal to 0.25. All attributes of the entire graph structure will be used later.
[0078] The physical enhancement model is used to simulate the diffusion process of brain nerve information:
[0079] The node embedding X of all brain regions in the graph structure is obtained by constructing graph random walks and unsupervised learning.
[0080] The physical enhancement model is a graph neural network embedding learning scheme based on node sequences, which is used to simulate the diffusion process of neural information. Based on the node2vec model, the weighted random walk mechanism and the Skip-gram model are introduced to learn the vector representation of nodes in the graph.
[0081] First, we define the feature learning framework, which aims to learn the mapping function f by maximizing the co-occurrence probability of nodes in the sequence. j :V j →R d By adjusting the parameters p and q, a balance is struck between new nodes and returning to old nodes. p controls the tendency of the random walk to return to the previous node, and q controls the tendency of the random walk to move away from the previous direction. The specific value depends on whether the network itself is densely connected or relatively sparse.
[0082] The objective function is defined as follows:
[0083]
[0084] Among them, the conditional probability is defined as:
[0085]
[0086] Node u j Move to the next node v j The probability is:
[0087]
[0088] Finally, embed the nodes of all brain regions into X∈R Num(V)×d Carry out the next step of graph analysis task;
[0089] In the formula, f j :V j →R d , f j is the mapping function that needs to be learned, through which nodes can be mapped to node embeddings; V j is the node of the jth person; is the set of nodes sampled by the Alias algorithm; n is the nth node sampled by the Alias algorithm; α pq is a node; For node u j and v j The connection weight between them; For node u j And the connection weight of the node sampled by the nth Alias algorithm.
[0090] On the other hand, two position codes are defined, and the two position codes are combined into three-dimensional coordinates to form the final position code X′, and the final position code is merged into the node embedding X as the output of the physical enhancement model;
[0091] The two position codes include a first position code and a second position code, wherein the first position code is a position code based on node centrality:
[0092]
[0093] Where V is the set of all nodes; N s (u) is the set of adjacent nodes of the node under the sampling of the Alias algorithm; Num is the number of nodes in the set.
[0094] The second position coding is named NeuroSyncFruchterman-Reingold, which provides a more sophisticated and advanced spatial coding strategy for the structural and functional analysis of brain networks based on the correlation constraint mechanism. This embodiment uses the attraction and repulsion between two nodes to update the position coordinates of the nodes through iteration. It is a heuristic algorithm that does not require labeled training.
[0095] The repulsive force is defined as:
[0096]
[0097] Attraction is defined as:
[0098]
[0099] Where k is a freely selectable parameter; w uv is the correlation between two nodes.
[0100] Iteratively updating the node position coordinates includes the following steps:
[0101] First, input the weight matrix Randomly initialize the position of each node and define the coefficient
[0102] Secondly, for each pair of nodes (u, v), iterate as follows:
[0103] The position shift under the repulsive force is recorded as:
[0104]
[0105] The position displacement under the action of attraction is recorded as:
[0106]
[0107] The resulting position update is recorded as:
[0108]
[0109] Finally, when the number of iterations reaches the maximum value, which is set to 2000 in this embodiment, the coordinates of each node at this time are returned and combined with the first position code to form the final position code:
[0110] X′=X||{PE u} u∈V ;
[0111] Wherein, area is the canvas size of the physical coordinate layout, which is 480000 (800×600) in this embodiment; V is the set of all nodes; pos(u) is the position of the u node; pos(v) is the position of the v node; t is a parameter simulating "temperature" to control the moving step of the node in each iteration; PE u Encode the position of the obtained node.
[0112] On the one hand, the classification model achieves node dimension increase through adaptive feature fusion.
[0113] The embedding method after dimensionality increase of the final position code X′ is:
[0114] The expression for upgrading the final position code X′ is:
[0115] X′ (p+1) =W×(X′ (p) ⊙W′)+X′ (p) ;
[0116] in,
[0117]
[0118] Each time the new embedding is obtained, it will be arranged in rows with the original embedding, and finally the dimensionality-enhanced embedding is obtained:
[0119] S,S∈M N,p+1,d+3 ;
[0120] Where p is the number of rounds, W i ' i is the weight matrix defined based on the node degree, D ii is the degree of node i; ε is a constant; M N,p+1,d+3 is a three-dimensional matrix of (N, p+1, d+3); N is the number of summary points; d is the embedding length of the final position code X′; ⊙ is the Hadamard product.
[0121] On the other hand, through the multi-task Transformer method, the node features after dimensionality increase are used to fully learn the activity relationship between tasks, and then the node representations of neighbors are fused through the attention mechanism. Finally, the probabilistic representation is obtained for category judgment, which significantly improves the quality of node representation.
[0122] The following steps are involved:
[0123] First, an encoder layer is constructed, which includes a parameter transformation layer, a multi-head self-attention mechanism, a feedforward neural network, etc.
[0124] Secondly, a soft parameter sharing mechanism is introduced through the parameter transformation layer to fully learn the activities between tasks and share parameter representations;
[0125] Finally, through the multi-head self-attention mechanism, the complex relationships between complex features are captured according to the attention scores.
[0126] The construction of the encoder layer comprises the following steps:
[0127] First, we build and apply a multi-head attention mechanism. The multi-head attention mechanism means that different types of features of data information can be captured through different "heads", which significantly improves the generalization of model learning. The multi-head attention mechanism can be expressed as:
[0128] MultiHead(Q,K,V)=Caoncat(head 1 ,…,head n )H O ;
[0129] in,
[0130]
[0131] In the formula, Q=SH Q ,K=SH K ,V=SH V ,H Q ,H K ,H V is a mapping matrix, which can be obtained through model learning and training; Q, K, V are query vectors, key vectors, and value vectors respectively; H O is S, which is the first layer input of the encoder; S is the embedding after dimensionality increase;
[0132] Second, construct a parameter transformation layer; the parameter transformation layer is represented by a fully connected linear transformation before the multi-head attention mechanism. Set the parameters used by the encoder to Θ, and the linear transformation corresponds to a weight matrix W i and a bias vector b i , where i is the number of encoder layers. Then the parameter transformation can be expressed as:
[0133] Θ′ i =W i Θ+b i ;
[0134] Θ′ i is the parameter after the transformation of parameter Θ, which is shared by all the multi-head attention in the encoder layer, so the parameters shared by each layer can be calculated:
[0135]
[0136] Where i is the number of encoder layers; W i is the learned weight matrix; b i is the learned bias; Q i , K i 、V i are the query vector, key vector, and value vector of each layer respectively; are the mapping matrices corresponding to each layer.
[0137] Third, construct a feedforward neural network; each encoder contains a feedforward neural network, which consists of two linear transformation layers and a nonlinear activation function, combined with a multi-head attention mechanism. The role is to introduce nonlinear transformation, capture high-level features in the data, and increase the total number of model parameters and improve the model's expressiveness:
[0138] FFN(x)=max(0,xW 1 +b 1 )W 2 +b 2 ;
[0139] Where x is the output of multi-head attention; W 1 ,b 1 and W 2 ,b 2 are the weight matrix and bias of the two-layer linear transformation respectively; max{·} is a nonlinear transformation that only takes the positive output representation (if it is negative, the output is zero); FFN(x) is the feedforward neural network expression.
[0140] In order to reduce the unnecessary parameters of the model and reduce the complexity of the model, this embodiment does not construct a decoder layer, but normalizes the output of the encoder layer, reduces the dimensionality through a linear layer, and then uses a simple attention mechanism to obtain the representation of each node, and outputs it through another linear layer to obtain the representation in the category space, thereby achieving fine-grained classification. (Existing technology, no further description).
[0141] The present invention has achieved the best performance in brain signal, especially auditory signal tasks, that is, it can judge what the subject is thinking with excellent accuracy by scanning the subject's brain signals. The present invention has the following application scenarios:
[0142] 1. Brain-computer interface field: The present invention can utilize information about brain functional connections to achieve more accurate and efficient human-computer interaction control. It can analyze brain signals and convert them into control commands to control external devices such as prostheses, robots, or smart home devices, thereby improving the convenience and flexibility of control. In the field of rehabilitation therapy, it can improve the speed and effect of recovery of patients' motor functions.
[0143] 2. Educational improvement: It can analyze students’ brain activity when they are performing learning tasks, and even help students with special learning needs (such as autistic patients, etc.) to develop personalized education plans based on their brain activity during learning and improve learning outcomes.
[0144] This example evaluates the proposed model on an fMRI dataset collected from healthy subjects. In the fMRI experiment, the subjects were asked to imagine and listen to four categories of auditory information (human, animal, machine, nature), so we obtained a total of eight categories of auditory neural activity and classified them to test the performance of the model. When preprocessing the fMRI data, the first five volumes of images in each run were discarded. All images were realigned to remove motion artifacts, and then the voxel size was resampled to 3×3×3mm using the T1 image, and kernel registration and normalization to the MNI space were performed. The normalized images were smoothed with a 6mm Gaussian kernel. The time series were then extracted to form functional connections.
[0145] In the self-built data set, the classification accuracy of the model PEMT-Net of the present invention reached 95.41%, and the ablation experiment (Multi-Transformer removing the physical location encoding: 82.52%, PET-Net removing the soft parameter sharing layer: 86.74%) and the comparative experiment (Node2vec: 79.31%, etc.) both achieved the best results.
[0146] The comparison results are shown in the following table.
[0147] method Accuracy Accuracy Recall F1 score GraphSAGE 67.48±0.53 67.56±0.54 67.52±0.53 67.47±0.53 DeepWalk 71.69±0.43 71.68±0.30 71.70±0.53 71.63±0.37 Node2vec 79.31±0.22 79.10±0.19 79.31±0.55 79.23±0.38 Multi-Transformer 82.52±0.59 85.12±0.52 83.28±0.45 82.74±0.27 PET-Net 86.74±0.64 89.07±0.55 86.64±0.49 86.17±0.29 PEMT-Net 95.41±0.38 95.32±0.42 95.28±0.36 95.26±0.23
[0148] like Figure 2 - Figure 3As shown, the present invention uses a neural network and a heuristic algorithm based on physical interaction to simulate the transmission path of neural information in the brain and the spatial coordinates of brain regions, providing a new idea for brain signal decoding, and experiments have proved that the data distribution after enhancement by this method is much better than the initial data. 3. The linear and nonlinear information is captured simultaneously through a multi-round adaptive fusion embedding method; the parameter transformation layer is used to fully learn the activities between tasks and share parameter representation, which significantly improves the quality of node representation.
Claims
1. A method for constructing a deep learning embedding network for decoding brain signals, characterized in that: Including physical enhancement models based on building graph random walks and physical location embedding, and classification models based on adaptive feature fusion and multi-task transformer; The physical enhancement model is used to simulate the diffusion process of brain neural information: by constructing a graph random walk and unsupervised learning, the node embedding X of all brain regions in the graph structure is obtained; on the other hand, two position codes are defined, the two position codes are combined into three-dimensional coordinates to form a final position code X′, and the final position code X′ is merged into the node embedding X as the output of the physical enhancement model; The two position codes include a first position code and a second position code, wherein the first position code is a position code based on node centrality: Where V is the set of all nodes; N s (u) is the set of adjacent nodes of the node under the sampling of the Alias algorithm; Num is the number of nodes in the set; The second position encoding utilizes the attraction and repulsion between two nodes to update the position coordinates of the nodes through iteration, including the following steps: First, input the weight matrix Randomly initialize the position of each node and define the coefficient Secondly, for each pair of nodes (u, v), iterate as follows: The position shift under the repulsive force is recorded as: The position displacement under the action of attraction is recorded as: The resulting position update is recorded as: The repulsive force is defined as: Attraction is defined as: Finally, when the number of iterations reaches the maximum value, the coordinates of each node at this time are returned and combined with the first position code to form the final position code: X′=X||{PE u } u∈V ; Where k is a predefined parameter; w uv is the correlation between two nodes; area is the size of the canvas for physical coordinate layout, V is the set of all nodes; pos(u) is the position of the u node; pos(v) is the position of the v node; t is a parameter that simulates "temperature" and is used to control the moving step of the node in each iteration; PE u Encode the position of the obtained node; On the one hand, the classification model realizes node dimensionality upgrade through adaptive feature fusion; on the other hand, through the multi-task Transformer method, the node features after dimensionality upgrade are fully learned to learn the activity relationship between tasks, and then the node representations of neighbors are fused through the attention mechanism, and finally a probabilistic representation is obtained for category judgment.
2. A method for constructing a deep learning embedding network for decoding brain signals according to claim 1, characterized in that: The graph structure includes nodes, node features and edge connections; the graph structure is obtained by the following method: The BOLD time series of each brain region of the subjects under the corresponding task obtained by functional magnetic resonance imaging are regarded as node features; Based on the BOLD sequence, edge connections were constructed between two brain regions using the Pearson correlation coefficient.
3. The method for constructing a deep learning embedding network for decoding brain signals according to claim 1, characterized in that: The method of obtaining the node embedding X of all brain regions in the graph structure by constructing a graph random walk and unsupervised learning includes the following steps: First, we define the feature learning framework. The goal is to learn the mapping function by maximizing the co-occurrence probability of nodes in the sequence. By adjusting the parameters p and q, we balance between adding new nodes and returning old nodes. The objective function is defined as follows: Among them, the conditional probability is defined as: Node u j Move to the next node v j The probability is: Finally, embed the nodes of all brain regions into X∈R Num(V)×d Carry out the next step of graph analysis task; In the formula, f j :V j →R d , f j is the mapping function that needs to be learned, through which nodes can be mapped to node embeddings; V j is the node of the jth person; is the set of nodes sampled by the Alias algorithm; n is the nth node sampled by the Alias algorithm; α pq is a node; For node u j and v j The connection weight between them; For node u j And the connection weight of the node sampled by the nth Alias algorithm.
4. The method for constructing a deep learning embedding network for decoding brain signals according to claim 1, characterized in that: The method of achieving node dimension increase through adaptive feature fusion is: The expression for upgrading the final position code X′ is: X′ (p+1) =W×(X′ (p) ⊙W′)+X′ (p) ; in, Each time the new embedding is obtained, it will be arranged in rows with the original embedding, and finally the dimensionality-enhanced embedding S is obtained: S∈M N,p+1,d+3 ; Where p is the number of rounds, W′ ii is the weight matrix defined based on the node degree, D ii is the degree of node i; ε is a constant; M N,p+1,d+3 It is a three-dimensional matrix of (N, p+1, d+3); N is the number of summary points; d is the embedding length of the final position code X′.
5. The method for constructing a deep learning embedding network for decoding brain signals according to claim 1, characterized in that: The attention mechanism fuses the node representations of neighbors and finally obtains a probability representation for class determination, including the following steps: First, an encoder layer is constructed, which includes a parameter transformation layer, a multi-head self-attention mechanism, and a feedforward neural network; Secondly, a soft parameter sharing mechanism is introduced through the parameter transformation layer to fully learn the activities between tasks and share parameter representations; Finally, through the multi-head self-attention mechanism, the complex relationships between complex features are captured according to the attention scores.
6. A method for constructing a deep learning embedding network for decoding brain signals according to claim 5, characterized in that: The construction of the encoder layer includes the following steps: First, build and apply the multi-head attention mechanism: MultiHead(Q,K,V)=Caoncat(head1,…,head n )H O ; in, In the formula, Q=SH Q ,K=SH K ,V=SH V ,H Q ,H K ,H V is a mapping matrix, which can be obtained through model learning and training; Q, K, V are query vectors, key vectors, and value vectors respectively; H O is S, which is the first layer input of the encoder; S is the embedding after dimensionality increase; Second, construct the parameter transformation layer; set the parameter used by the encoder to θ, then: I' i =W i I+b i ; Θ′ i is the parameter after the transformation of parameter Θ, which is shared by all encoder layers for multi-head attention, so the parameters shared by each layer are calculated: Where i is the number of encoder layers; W i is the learned weight matrix; b i is the learned bias; Q i , K i 、V i are the query vector, key vector, and value vector of each layer respectively; These are the mapping matrices corresponding to each layer; Third, construct a feedforward neural network; the encoders all contain a feedforward neural network, which consists of two linear transformation layers and a nonlinear activation function, combined after a multi-head attention mechanism: FFN(x)=max(0,xW1+b1)W2+b2; Where x is the output of the multi-head attention; W1,b1 and W2,b2 are the weight matrix and bias of the two-layer linear transformation respectively; max{·} is a nonlinear transformation that only takes the representation of positive output; FFN(x) is the feedforward neural network expression.
Citation Information
Patent Citations
Functional brain network classification method based on pre-training and graph neural network
CN113313232A
Rich semantic dialogue generation method fusing visual situation
CN115964467A