A depression state detection system based on a multi-modal nuclear magnetic image graph neural network

By constructing an LSTM-DD-GAT model and using the stimulation range of a TMS coil to divide fMRI data into grids, combined with graph neural networks and self-attention mechanisms, the problem of insufficient consideration of individual differences in existing technologies is solved, and individualized detection and accurate diagnosis of depressive states are achieved.

CN119867753BActive Publication Date: 2025-11-07江淮前沿技术协同创新中心 +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411887199.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-19
Publication Date
2025-11-07
Estimated Expiration
2044-12-19

AI Technical Summary

Technical Problem

Current technologies for diagnosing depression rely on standard brain region methods and subjective experience, failing to effectively account for individual differences, resulting in insufficient accuracy in diagnosis and treatment.

Method used

An LSTM-DD-GAT model based on multimodal MRI image graph neural network was constructed. The fMRI data grid was divided by the stimulation range of TMS coil. A personalized brain functional connectivity map was constructed using graph neural network and self-attention mechanism. The robustness and generalization ability of the model were improved by combining a domain classifier.

Benefits of technology

It enables individualized detection of depressive states, improves the accuracy and adaptability of diagnosis, and enhances the applicability of the model to different individuals and the targeted nature of treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119867753B_ABST
    Figure CN119867753B_ABST
Patent Text Reader

Abstract

The application discloses a kind of depression state detection systems based on multi-modal nuclear magnetic image graph neural network.The system includes brain data processing module, feature extraction module and classifier module;Brain data processing module is according to the input fMRI data, and multiple graph nodes are obtained by dividing grid aggregation to brain, the graph node and edge value are calculated, and then the construction of brain function connection graph is completed;Feature extraction module analyzes brain function connection graph, and extracts feature vector H by network layer and pooling layer, and represents the characteristics of brain function connection;Classifier module trains depression state classifier, and finally the trained depression state classifier outputs the probability of depressive patient or healthy individual according to the input feature vector H.The application uses graph neural network and self-attention mechanism to comprehensively explore the depression features in fMRI data, providing support for depression state detection.Furthermore, by using domain classifier, the robustness and generalization ability of the depression state classifier are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of medical rehabilitation treatment, and particularly relates to a depression state detection system based on a multi-modal nuclear magnetic image graph neural network. BACKGROUND

[0002] Depression, as a common mental disorder, is mainly manifested in persistent low mood, difficulty in concentrating, and other symptoms. This disease has a long incubation period, repeated episodes, and is difficult to completely cure, which seriously endangers the psychological health of patients. With the acceleration of modern social rhythm, the increase of life pressure and the influence of other factors, the age distribution of depression patients is becoming more and more extensive, and the total number of patients is showing a significant upward trend. Depression not only affects the mood and cognitive function of patients, but also may lead to social barriers, decline in professional ability, tension in family relationships and a series of problems. These problems further exacerbate the psychological pressure of patients, making the treatment and rehabilitation of depression more difficult.

[0003] The current clinical use of neuromodulation treatment provides an effective non-drug treatment option for patients with depression, especially for patients who are ineffective or unsuitable for drug treatment. Neuromodulation treatment has the advantages of non-invasiveness or minimally invasive, reducing the risk of surgery, and using electricity, magnetism, light or chemical means to regulate the activity of specific brain regions, thereby targeting the treatment of certain brain dysfunction diseases. Transcranial magnetic stimulation (TMS) is a kind of neuromodulation technology, which has the characteristics of non-invasiveness and reversibility, uses the magnetic field generated by the coil to produce local induced current in the cerebral cortex, regulates the excitability of cerebral cortex neurons, and thus affects the function of the human brain. At present, transcranial magnetic stimulation has been widely used in the research of brain function and the treatment of certain neuropsychiatric diseases, such as depression, schizophrenia, Parkinson's disease, epilepsy and migraine, and has good clinical application value and prospect.

[0004] With the continuous development of medical imaging technology, doctors and researchers can observe and study the structure and function of the brain without causing harm to the human body, improve the efficiency of diagnosis and the accuracy of treatment, and cope with the problem of individual differences, and make personalized targeted treatment. Today, magnetic resonance imaging (MRI) technology has been widely used in the diagnosis and treatment of brain dysfunction. Structural magnetic resonance imaging (sMRI) can clearly present detailed brain structure information of individual patients, can reduce the differences between individuals in the anatomical level, and realize the stereoscopic geometric positioning of the stimulation target of the patient through three-dimensional modeling of the brain. Compared with the traditional stimulation target positioning based on the standard brain region, it can be more accurately positioned and quantified, and based on sMRI, the maximum current generated by the TMS coil can be controlled in the specified area. Functional magnetic resonance imaging (fMRI) is an imaging technique used to study brain function. It indirectly measures neural activity by detecting changes in blood flow in the brain when it is active. fMRI can be used to study the functional interactions and cooperative work between different regions of the brain. Through fMRI, the neural activity of other parts functionally connected to the stimulation site can be affected, such as deep subcortical regions that are not easily stimulated. fMRI has important value in studying brain functional connectivity, which can help scientists and doctors understand how the brain processes information and the relationship between brain dysfunction and disease.

[0005] The analysis of fMRI data is often based on standard brain partitioning methods to divide the brain into regions for individuals, and the regions divided with subjective factors are used as seed points or regions of interest (ROI) for further connection analysis, which is highly dependent on the professional knowledge and subjective experience of researchers, or uses standard space registration, but this process does not well consider the differences between individuals. The brain is scanned using an MRI scanner to obtain a series of time-series three-dimensional images. Each image at a time point contains a snapshot of the brain. In fMRI data analysis, voxels can be regarded as small regions of the brain, and each voxel contains signal intensity at a specific time. TMS is a non-invasive brain stimulation technique that generates a short magnetic field pulse by placing an electromagnetic coil on the scalp to generate a small current in a specific area of the brain, affecting the activity of neurons. The size of the brain area affected by the TMS coil is about 5mm x 5mm x 5mm. To elucidate the functional mechanism of the brain, researchers use various methods to explore and understand the working principle of the brain using fMRI. Many of these methods are based on pre-defined brain atlas node construction techniques for graph neural networks. This method usually relies on brain region division standards such as automatic anatomical markers (AAL), Brodmann areas, etc. to determine the nodes in the graph using standardized templates. However, this method may have limitations in explaining individual differences, dynamic changes, or high-resolution data (Depression state detection system based on physiological signal synchronization and multi-feature fusion (CN202310576947.6) Depression state recognition system based on multi-level feature fusion (CN202310761351.3)). SUMMARY

[0006] The present application is based on the problem of depression state detection based on multi-modal nuclear magnetic resonance image graph neural network. Through the individual level graph neural network LSTM-DD-GAT based on fMRI, the diagnosis of depression state of individuals is realized, thereby solving the problems in the above background art.

[0007] Graph Neural Networks (GNN) can identify functionally similar or closely interacting region sets in the brain by analyzing fMRI data, which helps to reveal the modular organization of the brain. The advantage of GNN in processing fMRI data is that it can learn directly on graph structure without converting graph data into traditional grid or vector form. This direct processing method allows GNN to retain the original topological structure information of graph data, thus more accurately capturing the complexity and dynamics of brain functional networks. And it can adapt to the differences in brain structure of different individuals, because they can be trained for each individual's brain map, thus providing personalized analysis results. GNN can effectively capture complex patterns and relationships in such graph structures.

[0008] The application provides a depression state detection system based on a multi-modal magnetic resonance imaging graph neural network, and a new network model architecture, namely LSTM-DD-GAT (LSTM-Domain Discriminator-Graph Attention Network), is constructed. In the LSTM-DD-GAT model, the fMRI data is first processed by a graph construction module, which determines the positions and quantities of graph nodes according to the stimulation range of the TMS coil, and the connection relationship between them is determined based on correlation analysis. The constructed graph is analyzed by a feature extraction module to reveal the potential neural mechanisms related to depression. The extracted features are input into a domain classifier to identify and eliminate the feature differences between the source domain and the target domain. Finally, the label classifier is used to predict the depression state of the individual.

[0009] The purpose of the application is to divide the fMRI data into a grid based on the stimulation range of the TMS coil, and to construct a graph network in this way. The graph neural network and the self-attention mechanism are used to comprehensively explore the depression features in the fMRI data, providing support for depression state detection. The domain classifier is used to improve the robustness and generalization ability of the depression state classifier.

[0010] The purpose of the application is achieved at least by one of the following technical solutions.

[0011] A depression state detection system based on a multi-modal magnetic resonance imaging graph neural network, comprising a brain data processing module, a feature extraction module and a classifier module;

[0012] The brain data processing module divides the brain into a grid according to the input fMRI data, aggregates multiple graph nodes, calculates the graph node and edge values, and then completes the construction of the brain functional connection graph.

[0013] The feature extraction module analyzes the brain functional connectivity graph, and extracts a feature vector H through a network layer and a pooling layer, which represents the characteristics of brain functional connectivity.

[0014] The classifier module trains a depression state classifier, and finally outputs the probability of a depression patient or a healthy individual according to the input feature vector H by the trained depression state classifier.

[0015] Further, in the brain data processing module, the fMRI data is first input, and the brain is divided into an equidistant three-dimensional grid matching the TMS coil stimulation range.

[0016] The voxels in each grid are aggregated into a graph node, and a long short-term memory network (LSTM) is used to extract time series information therefrom and capture dynamic changes between the graph nodes. Then, the functional connectivity between the graph nodes is calculated, and the correlation between the graph nodes is calculated based on the BOLD signal time series of the voxels between different graph nodes using correlation analysis methods such as Pearson correlation coefficient or Spearman rank correlation coefficient or partial correlation coefficient or phase synchronization. When the time series of the voxels of two graph nodes show significant correlation in statistics, an edge between the graph nodes is established for the graph node pair with a correlation exceeding the threshold value τ, and a functional connectivity graph is formed. The functional connectivity graph data finally output includes the graph node and correlation edge connection information, thereby providing a basis for subsequent LSTM-DD-GAT model analysis of depression-related neural activity.

[0017] Further, the size of each grid is determined according to the specific purpose of the study and the stimulation characteristics of the TMS coil.

[0018] The size of the grid should match the stimulation range of the TMS coil to ensure that the details of the neural activity related to depression are captured, thereby better assisting subsequent analysis of treatment. For example, assuming that the stimulation range of the TMS coil is 10 mm in diameter and it is desired to capture the neural activity within the stimulation area, each grid can be divided into a 5x5x5 mm cube. This size can ensure sufficient resolution within the stimulation area to detect the activity of brain regions, while not being overly subdivided, resulting in excessive computational load.

[0019] After the grid is divided, the voxels (i.e., volume pixels) within each grid are aggregated into a graph node; this graph node represents the collective characteristics of the voxels within the grid region, including local neural activity patterns. Since the TMS stimulation area and treatment conditions are related to the stimulation range of the TMS coil, by analyzing the stimulation area under appropriate grid division, the neural activity in the corresponding stimulation area after stimulation can be better analyzed and detected, avoiding the situation of analyzing too many unnecessary areas or ignoring necessary area information.

[0020] Further, the feature extraction module comprises a double-layer graph attention network layer GATs, a graph pooling layer, and a plurality of groups of sequentially connected groups of graph attention network GAT layers and graph pooling layers, the double-layer graph attention network layer GATs comprising two layers of graph attention network GAT layers connected in sequence;

[0021] The double-layer graph attention network layer GATs receives functional connection graph data including feature vectors of graph nodes and edge weights between graph nodes, extracts local neighborhood information, captures local features and part of global features of the graph nodes, and outputs updated functional connection graph data to the graph pooling layer, which reduces the sampling of the graph and outputs the reduced complexity graph representation data to the readout layer and the graph attention network GAT layer in the first group of the plurality of groups of graph attention network GAT layers and graph pooling layers, extracts local neighborhood information, and outputs updated functional connection graph data to the graph pooling layer in the first group, which further processes and refines the features in an aggregated feature manner to help restore or reconstruct important local features, maintains a balance between local and global information, and allows the graph pooling layer to filter out graph nodes with higher importance for information transmission. The graph pooling layer in the first group reduces the sampling of the graph and outputs the reduced complexity graph representation data to the readout layer and the graph attention network GAT layer in the second group, and so on. Each group of graph attention network GAT layers and graph pooling layers extracts local neighborhood information of the nodes and refines the node features one group at a time. The graph pooling layer in the last group outputs the reduced complexity graph representation data to the readout layer, which finally integrates the input data and outputs a fixed-length feature vector representing the global information of the entire graph.

[0022] Further, in the graph attention network GAT layer, a self-attention mechanism is introduced to assign different importance weights to each graph node in the brain functional connection graph. The mechanism calculates the importance scores of each graph neighbor node to the target node, which reflect the correlation between the neighbor nodes and the target node. According to these scores, the graph attention network GAT layer can aggregate the features of the neighbor nodes with weights, thereby better capturing complex node relationships in the information integration process. This allows the graph attention network GAT layer to flexibly focus on more important neighbor nodes for the target node, improving the accuracy of data processing.

[0023] In the graph attention network GAT layer, a self-attention mechanism is introduced, which is as follows:

[0024] For a graph node v and its graph neighbor node u, the attention score e vu which can be calculated by the following formula:

[0025]

[0026] where h v and h u are the feature vectors of the graph node v and the graph neighbor node u, respectively, W1is a learnable weight matrix, a is another set of learnable parameters, || denotes the concatenation of vectors, N(v) denotes the set of all graph neighbor nodes of the graph node v; LeakyReLU is an activation function that allows gradient to pass in the negative region by providing a non-zero small slope for negative input values.

[0027] After calculating the attention score, the feature vector of the graph neighbor node u of each graph node v is updated by the following formula:

[0028] h′ u = σ(∑ k∈N(v) α vu W2h k );

[0029] where h′ u is the updated feature vector of the graph neighbor node u, α vu is the normalized attention score e vu , σ is a nonlinear activation function, and W2is a learnable weight matrix.

[0030] Further, in the graph pooling layer, a graph pooling method based on the self-attention mechanism is used to calculate the attention score of the node, so as to distinguish which nodes should be retained and which nodes should be discarded; a graph pooling technology combining multiple self-attention mechanisms is used to evaluate the importance of the node, and the relationship between the full graph information and the node features is determined through the graph convolution network to determine the attention score;

[0031] In the multiple sets of graph attention network GAT layers and the graph pooling layer, the graph pooling layer first processes through graph convolution and activation function to obtain the self-attention score of each graph node; secondly, according to the ranking of the self-attention score, the top k graph nodes with the highest score are retained to form a mask; finally, the features of the specified graph nodes and the topological information between these graph nodes are retained according to the mask, as follows:

[0032] X is the node representation of the graph node after the graph attention network GAT layer, and the first attention score Z1is calculated using graph convolution:

[0033]

[0034] where A is the adjacency matrix, is the adjacency matrix with self-connection, is the degree matrix of , I is the identity matrix, and Θ1is a t x 1 parameter.

[0035] The second attention score Z2 is calculated using the full graph information, and the feature representation of the full graph information is first obtained using the average method:

[0036]

[0037] where n is the number of graph nodes, u i is the graph node representation of the ith graph node after processing by the graph attention network GAT layer, Θ2 is a t x t parameter matrix, and the dot product operation result of the full graph information and the graph node information is used to determine the attention weight of the graph node, and after obtaining the attention weight of each graph node, the full graph information can be obtained by weighted average node features:

[0038]

[0039]

[0040] Therefore, the total attention score Z of each graph node is:

[0041]

[0042] According to the total attention score, a part of the nodes of the input graph is retained:

[0043]

[0044] where k is the pooling ratio, which determines the number of graph nodes that need to be retained, Top-Rank returns the index of the topk nodes based on the total attention score Z, and Zmask is the mask of the retained nodes.

[0045] Further, the readout layer is responsible for aggregating the features of all graph nodes to form a global graph representation, converting high-dimensional node features into lower-dimensional global feature representation for classification tasks; in essence, it is actually doing a pooling function, but the readout layer aggregates node embeddings into graph embeddings, and its output can be directly used for graph-level classification tasks; the addition pooling operation has superior expression ability compared to the mean pooling and maximum pooling, so the addition pooling operation is used to mine depression features in fMRI data, as follows:

[0046]

[0047] where and are the graph embedding and node embedding matrices of the readout layer corresponding to the lth group of GAT layers and graph pooling layers, respectively, and the feature H input to the classifier is the feature splicing of all readout layers, i.e.

[0048] The main task of the readout layer is to aggregate the features of each node to generate a global graph representation for subsequent classification tasks. The specific operation is to pool all node embedding features by addition to summarize a global feature vector. This addition pooling method has been proven to be superior to mean pooling and max pooling in capturing the expression characteristics of different graph structures. The output of the readout layer is a low-dimensional global feature representation that can effectively reflect the features in the fMRI data.

[0049] Further, in the classifier module, a domain classifier and a depression state classifier are constructed; due to the difference in fMRI sample collection conditions, there is inconsistency in the distribution between different data sets. In the transfer learning framework, the data set with labels is called the source domain, and the data set lacking labels is the target domain; therefore, the labels of the source domain and the target domain are:

[0050]

[0051] The domain classifier performs linear transformation and dimensionality reduction mapping of the input feature H to a two-dimensional space through a fully connected layer, and the data points in this space are used to distinguish samples from the source domain and the target domain, and then a Sigmoid function is used for binary classification, i.e., to identify whether the sample comes from the source domain or the target domain:

[0052]

[0053] In the domain classifier, 1 represents the target domain, and 0 represents the source domain, is the parameter matrix of the fully connected layer of the domain classifier; during backpropagation, the gradient of the domain classification loss of the domain classifier is automatically negated before being backpropagated to the parameters of the feature extractor, thereby realizing the adversarial loss.

[0054] The purpose of the depression state classifier is to analyze the input feature H to identify whether the feature data belongs to a depression patient or a healthy individual; the input of the depression state classifier is the data obtained by the domain classifier through linear transformation and dimensionality reduction mapping of the input feature H to a two-dimensional space through a fully connected layer, and the output is the probability of each class (depression state or normal state); the labels corresponding to depression and health are:

[0055]

[0056] The depression state classifier outputs the probability of each class through a fully connected layer and a Sigmoid function:

[0057]

[0058] In the depression state classifier, 1 represents the depression state, and 0 represents the normal state, is the parameter matrix of the full connection layer of the depression state classifier; the result output by the depression state classifier is the probability of the depression state (marked as 1) and the normal state (marked as 0); the probabilities reflect the confidence degree of the model on the input data belonging to a certain state.

[0059] Further, in the classifier module, an existing overall data set is input, the overall data set includes source domain (training set) data and target domain (test set) data, wherein the source domain data is represented as X S , the target domain data is represented as X T , in the classifier module, the core target of training is to minimize the depression state classifier loss function L C and maximize the domain classifier loss function L D The gradient descent method is used to optimize the depression state classifier and the domain classifier, and the parameters of the depression state classifier and the domain classifier are adjusted through back propagation to reduce the overall loss function, as follows:

[0060]

[0061] The optimization process is as follows:

[0062]

[0063] Wherein, W is the parameter of the feature extraction module; by minimizing the loss function of the depression state classifier, the depression state classifier can make the prediction result as close to the true label as possible; at the same time, by maximizing the loss function of the domain classifier, the target is to make the source domain of the sample unable to be effectively identified.

[0064] The target of the domain classifier is to maximize the domain classification loss, and to confuse the target domain data and the source domain data, but the target of the depression state classifier is to minimize the depression state classification loss, and to realize the accurate classification of the depression state. The depression state classifier and the domain classifier are mutually opposed in the training process to finally realize the mutual balance between the depression state classification loss and the domain classification loss. The purpose of the domain classifier and the depression state classifier is to enhance the generalization ability of the depression state classifier and ensure that the depression state classifier can accurately predict the depression state.

[0065] Compared with the prior art, the advantages of the technical scheme of the present application are as follows:

[0066] The clinical diagnosis of depression mainly depends on the patient's self-report and the doctor's clinical experience, and the target selection of TMS also depends on the doctor's professional experience. And the fMRI data is divided into brain regions according to the standard brain partition method, and the regions divided with subjective factors are connected for analysis, which depends on the professional knowledge and subjective experience of the researchers.

[0067] In order to solve the above problems, the application divides the fMRI data into a grid based on the stimulation range of the TMS coil, thereby constructing a graph network. The graph neural network and the self-attention mechanism are used to comprehensively explore the depression features in the fMRI data, thereby providing support for depression state detection. Then, the domain classifier is used to improve the robustness and generalization ability of the model. The application proposes a depression state detection system based on a multi-modal magnetic image graph neural network, constructs an LSTM-DD-GAT network model architecture (including four modules), realizes the judgment of the individual depression state, and constructs a brain function graph network structure based on the voxel level by analyzing the fMRI data. The structure can capture the functional connection relationship between different regions in the brain. The GAT can effectively process the graph structure data, learn the local neighborhood information of the nodes, and extract the global features related to the depression state. Through the training and verification of the LSTM-DD-GAT model, the evaluation of the individual depression state is realized.

[0068] 1. Compared with the traditional diagnosis and treatment scheme, the brain data processing module of the application utilizes the characteristics of the TMS coil stimulation to divide the fMRI data into a grid, so that the analysis and diagnosis results are more suitable for the treatment of TMS stimulation.

[0069] 2. The LSTM-DD-GAT network model architecture in the depression state detection system of the application improves the processing ability of the model for time series signals through the LSTM layer, reveals the potential time sequence relationship in the fMRI data, reduces the difference between the source domain and the target domain data distribution through the introduction of the domain classifier, enhances the generalization ability of the model for different domains, and improves the adaptability.

[0070] 3. In the specific processing of feature extraction, the application proposes a feature extraction module combining a multi-layer GAT, a Graph Pooling layer and a readout layer, which can highlight the local topological structure and node features of the graph, determine the attention score in combination with the relationship between the full graph information and the node features, and make the feature extraction module fully extract the representation ability of the full graph. BRIEF DESCRIPTION OF DRAWINGS

[0071] Figure 1 FIG. 1 is a structural schematic diagram of the brain data processing module in the embodiment of the application;

[0072] Figure 2 FIG. 2 is a structural schematic diagram of the feature extraction module in the embodiment of the application;

[0073] Figure 3 FIG. 3 is a structural schematic diagram of the classifier module in the embodiment of the application. DETAILED DESCRIPTION

[0074] In order to make the objectives, technical solutions and advantages of the present application clearer, the specific embodiments of the present application are described in detail below with reference to the drawings and examples.

[0075] Embodiments

[0076] A depression state detection system based on a multi-modal nuclear magnetic image graph neural network, comprising a brain data processing module, a feature extraction module and a classifier module;

[0077] The brain data processing module divides the brain into a plurality of graph nodes by aggregating the grid according to the input fMRI data, calculates the graph node and edge values, and further completes the construction of the brain functional connectivity graph;

[0078] The feature extraction module analyzes the brain functional connectivity graph, extracts a feature vector H through a network layer and a pooling layer, and represents the characteristics of the brain functional connectivity;

[0079] The classifier module trains a depression state classifier, and finally outputs the probability of a depression patient or a healthy individual according to the input feature vector H by the trained depression state classifier.

[0080] As shown in Figure 1 In the brain data processing module, the fMRI data is first input, and the brain is divided into an equal-interval three-dimensional grid matching the stimulation range of the TMS coil;

[0081] The voxels of each grid are aggregated into graph nodes, and the Long Short-Term Memory (LSTM) is used to extract time series information therefrom to capture the dynamic changes between the graph nodes; then the functional connectivity between the graph nodes is calculated, the correlation between the graph nodes is calculated based on the BOLD signal time series of the voxels between different graph nodes, and when the time series of the voxels of two graph nodes show significant correlation in statistics, the edges between the graph nodes are established only for the graph node pairs whose correlation exceeds the threshold value τ, thereby forming a functional connectivity graph; the functional connectivity graph data finally output includes the graph node and correlation edge connection information, thereby providing a basis for the subsequent LSTM-DD-GAT model to analyze depression-related neural activity.

[0082] In one embodiment, the correlation analysis method is a Pearson correlation coefficient, a Spearman rank correlation coefficient, a partial correlation coefficient or a phase synchronization method;

[0083] Further, the size of each grid is determined according to the specific purpose of the study and the stimulation characteristics of the TMS coil;

[0084] The size of the grid should match the stimulation range of the TMS coil to ensure that the details of neural activity related to depression are captured, thereby better assisting subsequent analysis and treatment; for example, assuming that the stimulation range of the TMS coil is 10 mm in diameter and it is desired to capture neural activity within the stimulation area, each grid can be divided into 5x5x5 mm cubes. This size can ensure sufficient resolution within the stimulation area to detect activity in brain regions, while not being too finely divided, resulting in excessive computation.

[0085] After dividing the grid, the voxels (i.e., volume pixels) within each grid are aggregated into a graph node; this graph node represents the collective characteristics of the voxels within the grid area, including local neural activity patterns. Since the TMS stimulation area and treatment situation are related to the stimulation range of the TMS coil, by analyzing the stimulation area under appropriate grid division, the neural activity in the corresponding stimulation area after stimulation can be better analyzed and detected, avoiding the situation of analyzing too many unnecessary areas or ignoring necessary area information.

[0086] As shown in Figure 2 , the feature extraction module includes a double-layer graph attention network layer GATs, a graph pooling layer, and multiple sets of sequentially connected groups of graph attention network GAT layers and graph pooling layers, the double-layer graph attention network layer GATs including two layers of graph attention network GAT layers connected in sequence;

[0087] The double-layer graph attention network layer GATs receives functional connection graph data, including feature vectors of graph nodes and edge weights between graph nodes, extracts local neighborhood information, captures local features and part of global features of graph nodes, and outputs updated functional connection graph data to a graph pooling layer, which reduces the sampling of the graph and outputs reduced complexity graph representation data to a readout layer and a graph attention network GAT layer in the first group of multiple sets of graph attention network GAT layers and graph pooling layers, extracts local neighborhood information, and outputs updated functional connection graph data to a graph pooling layer in the first group, which further processes and refines the features in an aggregated feature manner to help restore or reconstruct important local features, maintains a balance between local and global information, and allows the graph pooling layer to filter out graph nodes with higher importance for information transmission. The graph pooling layer in the first group reduces the sampling of the graph and outputs reduced complexity graph representation data to the readout layer and the graph attention network GAT layer in the second group, and so on. Each group of graph attention network GAT layers and graph pooling layers extracts the local neighborhood information of the nodes and refines the node features in each group. The graph pooling layer in the last group outputs reduced complexity graph representation data to the readout layer, which finally integrates the input data and outputs a fixed-length feature vector representing the global information of the entire graph.

[0088] Further, in the graph attention network GAT layer, a self-attention mechanism is introduced to assign different importance weights to each graph node in the brain functional connectivity graph; the mechanism calculates the importance scores of each graph neighbor node to the target node, which reflect the correlation between the neighbor node and the target node; according to the scores, the graph attention network GAT layer can weightedly aggregate the features of the neighbor nodes, so as to better capture the complex node relationship in the information integration process; so that the graph attention network GAT layer can flexibly focus on the neighbor nodes more important to the target node, and improve the accuracy of data processing;

[0089] In the graph attention network GAT layer, a self-attention mechanism is introduced, which is as follows:

[0090] For a graph node v and its graph neighbor node u, the attention score e vu can be calculated by the following formula:

[0091]

[0092] where h v and h u represent the feature vectors of the graph node v and the graph neighbor node u respectively, W1 is a learnable weight matrix, a is another set of learnable parameters, || represents the concatenation of vectors, N(v) represents the set of all graph neighbor nodes of the graph node v; LeakyReLU is an activation function that provides a non-zero small slope for negative input values, thereby allowing gradient to pass in the negative region.

[0093] After calculating the attention score, the feature vector of each graph neighbor node u of the graph node v is updated by the following formula:

[0094] h′ u = σ(∑ k∈N(v) α vu W2h k );

[0095] where h′ u is the updated feature vector of the graph neighbor node u, α vu is the normalized attention score e vu , σ is a nonlinear activation function, and W2 is a learnable weight matrix.

[0096] Further, in the graph pooling layer, a graph pooling method based on the self-attention mechanism is used to calculate the attention score of the node, so as to distinguish which nodes should be retained and which nodes should be discarded; a graph pooling technology combining multiple self-attention mechanisms is used to evaluate the importance of the nodes, and the relationship between the full graph information and the node features is determined through the graph convolution network to determine the attention score;

[0097] In the multi-group graph attention network GAT layer and the graph pooling layer, the graph pooling layer first processes through graph convolution and activation function to obtain the self-attention score of each graph node; secondly, according to the ranking of the self-attention score, the top k nodes with the highest score are reserved to form a mask; finally, the features of the specified graph nodes and the topological information between these graph nodes are reserved according to the mask, as follows:

[0098] X is the node representation of the graph node after the graph attention network GAT layer, and the first attention score Z1 is calculated using graph convolution:

[0099]

[0100] Where A is the adjacency matrix, is the adjacency matrix with self-connection, is the degree matrix of , I is the identity matrix, and Θ1 is a t x 1 parameter;

[0101] The second attention score Z2 is calculated using the full graph information. First, the average method is used to obtain the feature representation of the full graph information:

[0102]

[0103] Where n is the number of graph nodes, u i is the graph node representation of the i-th graph node after the graph attention network GAT layer, Θ2 is a t x t parameter matrix, and the dot product of the full graph information and the graph node information is used to determine the attention weight of the graph node. After obtaining the attention weight of each graph node, the full graph information can be obtained by weighted average node feature:

[0104]

[0105]

[0106] Therefore, the total attention score Z of each graph node is:

[0107]

[0108] According to the total attention score, a part of the nodes of the input graph is reserved:

[0109]

[0110] Where k is the pooling ratio, which determines the number of graph nodes to be reserved, Top-Rank returns the index of the top k nodes based on Z, and Zmask is the mask for the reserved nodes.

[0111] Further, the readout layer is responsible for aggregating the features of all graph nodes to form a global graph representation, converting high-dimensional node features into a lower-dimensional global feature representation for classification tasks; in essence, it is actually a pooling function, but the readout layer aggregates node embeddings into graph embeddings, and its output can be directly used for graph-level classification tasks; the addition pooling operation has superior expressive power compared to mean pooling and max pooling, so it is used to mine depression features in fMRI data, as follows:

[0112]

[0113] wherein, and are the graph embedding and node embedding matrices of the readout layer corresponding to the lth group of GAT layers and graph pooling layers, respectively, and the feature H input to the classifier is the concatenation of the features of all readout layers, i.e.

[0114] The main task of the readout layer is to aggregate the features of each node to generate a global graph representation for subsequent classification tasks. The specific operation is to perform addition pooling on the embedding features of all nodes to summarize a global feature vector. This addition pooling method has been proven to be superior to mean pooling and max pooling in capturing the expressive features of different graph structures. The output of the readout layer is a low-dimensional global feature representation that can effectively reflect the features in the fMRI data.

[0115] In one embodiment, as shown in Figure 3 , in the classifier module, a domain classifier and a depression state classifier are constructed; due to the differences in fMRI sample collection conditions, there is inconsistency in the distribution between different data sets. In the transfer learning framework, the data set with labels is called the source domain, and the data set lacking labels is the target domain. Therefore, the labels of the source domain and the target domain are:

[0116]

[0117] The domain classifier performs linear transformation and dimensionality reduction mapping of the input feature H to a two-dimensional space through a fully connected layer, and the data points in this space are used to distinguish samples from the source domain and the target domain, and then a Sigmoid function is used for binary classification, i.e., to identify whether the sample comes from the source domain or the target domain:

[0118]

[0119] In the domain classifier, 1 represents the target domain, and 0 represents the source domain, is the parameter matrix of the domain classifier's fully connected layer; during the backpropagation process, the gradient of the domain classifier's domain classification loss is automatically negated before being backpropagated to the parameters of the feature extractor, thereby achieving the adversarial loss.

[0120] The purpose of the depression state classifier is to analyze the input feature H to identify whether the feature data belongs to a depression patient or a healthy individual; the input of the depression state classifier is the data obtained by the domain classifier through linear transformation dimensionality reduction mapping of the input feature H to a two-dimensional space by the fully connected layer, and the output is the probability of each class (depression state or normal state); the labels corresponding to depression and health are:

[0121]

[0122] The depression state classifier outputs the probability of each class through the fully connected layer and the Sigmoid function:

[0123]

[0124] In the depression state classifier, 1 represents the depression state, and 0 represents the normal state, is the parameter matrix of the depression state classifier's fully connected layer; the output of the depression state classifier is the probability of the depression state (labeled as 1) and the normal state (labeled as 0); these probabilities reflect the model's confidence level in determining whether the input data belongs to a certain state.

[0125] Further, in the classifier module, the existing overall data set is input, which includes source domain (training set) data and target domain (test set) data, where the source domain data is represented as X S , and the target domain data is represented as X T In the classifier module, the core objective of training is to minimize the depression state classifier loss function L C and maximize the domain classifier loss function L D The gradient descent method is used to optimize the depression state classifier and the domain classifier, and the parameters of the depression state classifier and the domain classifier are adjusted through backpropagation to reduce the overall loss function, as follows:

[0126]

[0127] The optimization process is as follows:

[0128]

[0129] where W is the parameter of the feature extraction module; by minimizing the loss function of the depression state classifier, the depression state classifier can make the prediction result as close to the true label as possible; at the same time, by maximizing the loss function of the domain classifier, the goal is to make the source domain of the sample unable to be effectively identified.

[0130] The goal of the domain classifier is to maximize the domain classification loss, confusing the target domain data with the source domain data, but the goal of the depression state classifier is to minimize the depression state classification loss, achieving accurate classification of the depression state. The depression state classifier and the domain classifier are mutually antagonistic in the training process, and finally achieve mutual balance between the depression state classification loss and the domain classification loss. The goal of the domain classifier and the depression state classifier is to enhance the generalization ability of the depression state classifier and ensure that the depression state classifier can accurately predict the depression state.

[0131] The preferred embodiments of the application disclosed above are only used to help understand the application and core idea. For those skilled in the field, according to the idea of the application, there will be changes in specific application scenarios and implementation operations, and the description should not be understood as limiting the application. The application is limited by the claims and their entire scope and equivalents.

Claims

1. A depression state detection system based on a multi-modal magnetic image graph neural network, characterized in that, The brain data processing module, the feature extraction module and the classifier module are included. The brain data processing module divides the brain into multiple graph nodes according to the input fMRI data, calculates the graph node and edge values, and then completes the construction of the brain functional connectivity graph. The feature extraction module analyzes the brain functional connectivity graph, extracts the feature vector H through the network layer and the pooling layer, and represents the characteristics of the brain functional connectivity. The classifier module trains the depression state classifier, and finally outputs the probability of a depression patient or a healthy individual according to the input feature vector H. In the brain data processing module, the fMRI data is input first, and the brain is divided into an equal-interval three-dimensional grid matching the TMS coil stimulation range. The voxels of each grid are aggregated into graph nodes, and the long short-term memory network is used to extract time series information therefrom to capture the dynamic changes between the graph nodes. Then, the functional connectivity between the graph nodes is calculated, the correlation between the graph nodes is calculated based on the BOLD signal time series of the voxels between different graph nodes, and when the voxel time series of two graph nodes shows significant correlation in statistics, the edge between the graph nodes is established only for the graph node pairs with correlation exceeding the threshold value τ, thereby forming a functional connectivity graph. The size of each grid is determined according to the specific purpose of the study and the stimulation characteristics of the TMS coil. The size of the grid should match the stimulation range of the TMS coil to ensure that the details of the neural activity related to depression are captured. After the grid is divided, the voxels in each grid are aggregated into a graph node; this graph node represents the collective characteristics of the voxels in the grid region, including the local neural activity pattern.

2. The depression state detection system based on the multi-modal magnetic image graph neural network according to claim 1, wherein, The feature extraction module includes a double-layer graph attention network layer GATs, a graph pooling layer, and multiple groups of sequentially connected graph attention network GAT layers and graph pooling layers. The double-layer graph attention network layer GATs includes two layers of graph attention network GAT layers connected in sequence. The double-layer graph attention network layer GAT receives functional connection graph data including feature vectors of graph nodes and edge weights between graph nodes, extracts local neighborhood information, and outputs updated functional connection graph data to a graph pooling layer in a layer, the graph pooling layer performs down-sampling on the graph, and outputs a reduced complexity graph representation data to a readout layer and a graph attention network GAT layer in a first group of the graph pooling layer and the graph attention network GAT layer, extracts local neighborhood information, and outputs updated functional connection graph data to a graph pooling layer in the first group, the graph pooling layer in the first group performs down-sampling on the graph, and outputs a reduced complexity graph representation data to a readout layer and a graph attention network GAT layer in a second group, and so on, each group of graph attention network GAT layers and graph pooling layers extracts local neighborhood information of nodes and refines node features in groups, and the graph pooling layer in the last group outputs a reduced complexity graph representation data to the readout layer, and the readout layer finally integrates the input data and outputs a fixed-length feature vector representing global information of the entire graph.

3. The depression state detection system based on the multi-modal magnetic image graph neural network according to claim 2, wherein, In the graph attention network GAT layer, a self-attention mechanism is introduced, specifically as follows: For a graph node v and its graph neighbor node u, the attention score e vu is computed by the following equation: where h v and h u denote the feature vectors of the graph node v and the graph neighbor node u, respectively, Wi is a learnable weight matrix, a is another set of learnable parameters, || denotes the concatenation of vectors, N(v) denotes the set of all graph neighbor nodes of the graph node v; LeakyReLU is an activation function that allows gradients to pass in the negative region by providing a small non-zero slope for negative input values. After calculating the attention score, the feature vector of the graph neighbor node u of each graph node v is updated by the following formula: h′ u = σ(∑ k∈N(v) α vu W2h k ); where h' = h + h u is the updated feature vector of the neighbor node u, and vu is the attention score e vu is the normalized attention score, and σ is a nonlinear activation function, and W2is a learnable weight matrix.

4. The depression state detection system based on the multi-modal nuclear magnetic image graph neural network according to claim 3, characterized in that, In the multiple groups of graph attention network GAT layers and graph pooling layers, the graph pooling layer first obtains the self-attention score of each graph node through graph convolution and an activation function; Secondly, according to the ranking of the self-attention score, the top k nodes with the highest score are retained to form a mask; finally, the features of the specified graph nodes and the topological information between these graph nodes are retained according to the mask, specifically as follows: X is the node representation of the graph node after the graph attention network GAT layer, and the first attention score Z1 is calculated using graph convolution: Where A is the adjacency matrix. It is an adjacency matrix with self-connections. yes The degree matrix, I is the identity matrix, and Θ1 is the parameter of t×1; The second attention score Z2 is calculated using the whole graph information, and the feature expression of the whole graph information is first obtained by using the average method: where n is the number of graph nodes, u i is the graph node representation of the i-th graph node after processing by the graph attention network GAT layer, Θ2 is a t x t parameter matrix, and the result of dot product operation using the global graph information and the graph node information is used to determine the attention weight of the graph node, and after obtaining the attention weight of each graph node, the global graph information is obtained by weighted average node features: Therefore, the total attention score Z of each graph node is: According to the total attention score, a part of the nodes of the input graph are retained: Z mask = Z ids ; Wherein, k is the pooling ratio, which determines the number of graph nodes to be retained, Top-Rank returns the index of the topk nodes based on the total attention score Z, and Zmask is the mask of the retained nodes.

5. The depression state detection system based on the multi-modal magnetic image graph neural network according to claim 4, wherein, The readout layer is responsible for aggregating the features of all graph nodes to form a global graph representation, converting high-dimensional node features into lower-dimensional global feature representations for classification tasks; in essence, it is actually a pooling operation, but the readout layer aggregates node embedding into graph embedding, and its output can be directly used for graph-level classification tasks; an addition pooling operation is adopted, specifically as follows: wherein, and are the graph embedding and node embedding matrices of the corresponding readout layers of the l-th group of GAT layers and graph pooling layers, respectively, and the features H of the final input to the classifier are the concatenation of the features of all readout layers, i.e.

6. The depression state detection system based on the multi-modal magnetic image graph neural network according to claim 1, wherein, In the classifier module, a domain classifier and a depression state classifier are constructed; the dataset with labels is called the source domain, and the dataset lacking labels is the target domain, and the labels of the source domain and the target domain are respectively: The domain classifier maps the input feature H to a two-dimensional space by linear transformation and dimension reduction through a fully connected layer, and the data points in this space are used to distinguish the samples from the source domain and the target domain, and then a sigmoid function is used for binary classification, i.e., to identify whether the sample comes from the source domain or the target domain: In the domain classifier, 1 represents the target domain, and 0 represents the source domain, is the parameter matrix of the full connection layer of the domain classifier; in the process of back propagation, the gradient of the domain classification loss of the domain classifier is automatically negated before being back propagated to the parameters of the feature extractor, thereby realizing the adversarial loss.

7. The depression state detection system based on the multi-modal magnetic image graph neural network according to claim 6, wherein, The purpose of the depression state classifier is to analyze the input feature H to identify whether the feature data belongs to a depression patient or a healthy individual; the input of the depression state classifier is the data obtained by the domain classifier by linear transformation and dimension reduction of the input feature H through a fully connected layer, and the output is the probability of each category (depression state or normal state); the labels corresponding to depression and health are: The depression state classifier outputs the probability of each category through a fully connected layer and a sigmoid function: In the depression state classifier, 1 represents a depression state, and 0 represents a normal state, is a parameter matrix of a full connection layer of the depression state classifier; and a result output by the depression state classifier is a probability of a depression state and a normal state.

8. The depression state detection system based on the multi-modal magnetic image graph neural network according to claim 7, wherein, In the classifier module, input the existing overall dataset, the overall dataset includes source domain (training set) data and target domain (test set) data, wherein the source domain data is represented as X S , and the target domain data is represented as X T In the classifier module, the core target of training is to minimize the depression state classifier loss function L C and maximize the domain classifier loss function L D The gradient descent method is used to optimize the depression state classifier and the domain classifier, and the parameters of the depression state classifier and the domain classifier are adjusted through back propagation to reduce the overall loss function, as follows:

9. The depression state detection system based on the multi-modal magnetic image graph neural network according to claim 8, wherein, In the classifier module, the optimization process is as follows: Where W is the parameter of the feature extraction module; by minimizing the loss function of the depression state classifier, the depression state classifier can make the prediction result as close to the true label as possible; at the same time, by maximizing the loss function of the domain classifier, the goal is to make the source domain of the sample unable to be effectively identified.

Citation Information

Patent Citations

  • Depression state detection system based on physiological signal synchronization and multi-feature fusion

    CN116725535A

  • Depression state recognition system based on multilevel feature fusion

    CN116763311A

  • Depression identification system based on multi-scale residual image attention network

    CN117898730A

  • Devices and methods for identifying brain stimulation targets

    WO2024186459A2