Brain Network Classification Method Based on Spatiotemporal Graph Convolution

Through the brain network classification method based on spatiotemporal graph convolution, the problem of the inability to simultaneously extract the temporal and spatial characteristics of resting state fMR image data in the prior art is solved, and higher classification accuracy and interpretability are achieved.

CN115496953BActive Publication Date: 2025-05-27HEBEI UNIV OF TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202211345300.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-31
Publication Date
2025-05-27
Estimated Expiration
2042-10-31

AI Technical Summary

Technical Problem

The existing brain network classification methods cannot simultaneously and effectively extract the temporal and spatial characteristics of resting fMR image data, resulting in the defects of high error rate and poor interpretation in the classification prediction task.

Method used

The brain network classification method based on spatiotemporal graph convolution is adopted to extract temporal features and spatial features through the split spatiotemporal graph convolution layer, and combine triple loss function and spatiotemporal graph attention module to improve the training effect and classification accuracy of the model.

Benefits of technology

Effectively extracting the spatial and temporal features of brain regions improves the accuracy and interpretability of brain network classification and reduces the impact of noise on the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115496953B_ABST
    Figure CN115496953B_ABST
Patent Text Reader

Abstract

The present invention is a brain network classification method based on spatio-temporal graph convolution. For the time series of resting-state functional magnetic resonance imaging (fMRI) data, split spatio-temporal graph convolution is used to simultaneously extract temporal features and spatial features. The method first analyzes the resting-state fMRI using voxel-based morphometry, divides the cerebral cortex into multiple regions of interest using a brain template, models the brain as a graph-structured brain network, and uses multiple spatio-temporal graph convolution layers to extract temporal features and spatial features respectively, reducing the complexity of the convolution operation. To further improve the training effect, the method uses a triplet loss function to enhance the model training effect and a spatio-temporal graph attention module to further extract spatial and temporal features, thereby improving the classification accuracy based on the brain network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technical solution of the present invention relates to a method for identifying graphics, specifically a brain network classification method based on spatiotemporal graph convolution. Background Art

[0002] In recent years, with the development and innovation of science and technology, countries around the world have successively launched corresponding brain planning projects, aiming to model and study the human brain through imaging technologies such as nuclear magnetic resonance and positron emission tomography, and use sophisticated molecular technologies such as single nucleotide polymorphisms to achieve early diagnosis of brain diseases (such as Alzheimer's disease, schizophrenia and autism) and timely intervention and treatment. Recent studies have shown that some phenotypic characteristics (such as gender, age, smoking and drinking, etc.) have an important influence on some neurodegenerative diseases. Therefore, researchers also use brain imaging technology to study the differences in phenotypic factors in brain structure and function, analyze their effects on brain diseases, and achieve targeted treatment and prevention. For example, the 2022 World Alzheimer's Disease Prevention and Control Association reported that the lifetime prevalence of Alzheimer's disease in people over 65 years old is about 21.1% (female) and 11.6% (male), so it is very necessary to study the impact of gender differences on brain diseases.

[0003] The brain has very complex functions and structures, and is constantly changing. Resting-state functional MRI records the changes in the blood oxygen concentration-dependent contrast signal of the brain at rest over a period of time through long-term scanning, and depicts the spontaneous activity and periodic changes of the brain. Therefore, the temporal characteristics of resting-state functional MRI data can be used to study the process of brain function changes, and to study the manifestation of gender differences at the level of brain function, so as to analyze the effects of these differences on brain diseases. With the popularization and promotion of nuclear magnetic resonance imaging technology, more and more institutions and hospitals around the world have successively collected a large amount of high-quality imaging data and conducted long-term follow-up diagnosis and treatment of patients, providing an important data source for studying brain structure and function.

[0004] CN114287908A discloses a brain connection classification method of multi-passband graph convolution, which uses the functional imaging features of the whole brain to extract the functional connection network from the white matter and gray matter partitions of the brain, and uses multi-passband feature learning on the basis of the existing graph convolution network to fuse low-pass filtering and band-pass filtering through graph scattering convolution. This method does not consider the temporal dynamics of functional magnetic resonance imaging data, and performs correlation analysis on the blood oxygen concentration dependent contrast signals at all time points to construct the brain network of the whole brain, which cannot fully utilize the temporal dynamics. CN114863181A discloses a gender classification method based on predicted probability knowledge distillation, which uses the teacher-student knowledge distillation model and uses ResNet18 and CNN networks for training to achieve gender prediction. This method uses a deep convolution model for gender classification tasks, which improves the classification accuracy, but ignores the interpretability of the brain network model and cannot explain the prediction results from the brain structure and function level. CN110477909A discloses a gender classification method based on resting-state EEG data. The method collects resting-state EEG data, pre-processes the raw EEG data and removes artifacts before inputting it into a convolutional neural network for training and testing. Although the method removes the influence of artifacts, it only uses the most primitive CNN network for gender classification tasks, and does not fully consider the influence of raw data noise on the training process.

[0005] In summary, among the existing brain network classification methods, the current feature extraction methods cannot effectively extract the temporal and spatial features of resting-state functional magnetic resonance imaging data at the same time. They have the defects of high error rate and poor interpretability in classification prediction tasks, and their accuracy needs to be improved. Summary of the invention

[0006] The technical task of the present invention is to address the above shortcomings and provide a brain network classification method based on spatiotemporal graph convolution. For the time series of resting-state functional magnetic resonance imaging data, the split spatiotemporal graph convolution is used to simultaneously extract temporal features and spatial features. The method first uses a voxel morphology-based analysis method to analyze resting-state functional magnetic resonance images, divides the cerebral cortex into multiple regions of interest using a brain template, models the brain as a graph-structured brain network, and uses multiple spatiotemporal graph convolution layers to extract temporal features and spatial features respectively, reducing the complexity of the convolution operation. In order to further improve the training effect, the method uses a triplet loss function to improve the model training effect, and uses a spatiotemporal graph attention module to further extract spatial features and temporal features, thereby improving the classification accuracy based on the brain network. The method disclosed in the present invention can effectively extract the spatial features of brain regions for time series data and make full use of their temporal dynamics.

[0007] In the above, “Spatial-Temporal Graph Convolution” is called “Spatial-Temporal Graph Convolution”, abbreviated as S-TGConv, “Spatial-Temporal Graph Attention Block” is abbreviated as S-TGAtten, “Triplet Loss” is called “Triplet Loss”, “Region Of Interest” is called “ROI”, and “Voxel-Based Morphometry” is called “Voxel-Based Morphometry”, abbreviated as VBM.

[0008] The technical solution adopted by the present invention to solve the technical problem is: a brain network classification method based on spatiotemporal graph convolution, characterized in that the brain network classification method based on spatiotemporal graph convolution, characterized in that the classification method includes the following contents:

[0009] After acquiring resting-state functional magnetic resonance imaging data, multiple regions of interest were extracted using the Human Connectome Project Multimodal Parcellation Atlas (HCP-MMP) using a voxel-based analysis method;

[0010] Construct a graph-structured brain network based on the time series features of the region of interest, and obtain the adjacency matrix A shared by all samples;

[0011] Build a classification prediction model:

[0012] The entire network of the classification prediction model includes a spatiotemporal graph convolution S-TGConv module with 64 channels, a temporal kernel size of 11 and a spatial kernel size of 5, a spatiotemporal graph convolution S-TGConv module with 128 channels, a temporal kernel size of 11 and a spatial kernel size of 5, a spatiotemporal graph attention S-TGAtten module with 128 channels, a normalized BN layer, a Dropout layer, two fully connected layers, and a classification function; the two S-TGConv modules use different numbers of channels and kernel sizes, and the number of channels is ranked from small to large. After the series connection, the S-TGAtten module is connected, and the features processed and spliced ​​by the spatiotemporal graph convolution in the second S-TGConv module are summed with the features directly output by the S-TGAtten module, and then the output after the LeakyReLU activation function is sequentially connected to the normalized BN layer, a Dropout layer, and two fully connected layers, and then the output of the classification prediction model is obtained through the classification function; the number of channels of the S-TGAtten module is consistent with the number of channels of the second S-TGConv module, which is used to extract weight information of different time points and different ROIs;

[0013] The S-TGConv module includes a spatial graph convolution layer and a temporal graph convolution layer in parallel. The outputs of the spatial graph convolution layer and the temporal graph convolution layer are connected to the LeakyReLU activation function through a splicing operation to obtain the output of the spatiotemporal graph convolution S-TGConv module;

[0014] The process of the spatiotemporal graph attention S-TGAtten module is as follows: for the features output by the spatiotemporal graph convolution S-TGConv module, first use a shared linear transformation weight matrix W on each node, and then use a layer with parameters W f The linear change layer extracts high-frequency features and low-frequency features, and concatenates the two features through a layer with a parameter of W t The linear change layer is finally transformed nonlinearly through the LeakyReLU activation function to obtain the output of the spatiotemporal graph attention S-TGAtten module.

[0015] The brain network classification method based on spatiotemporal graph convolution, the specific steps of the method are:

[0016] The first step is to build a resting-state functional magnetic resonance imaging dataset:

[0017] Step 1.1, resting state functional magnetic resonance imaging data preprocessing:

[0018] A resting-state functional magnetic resonance imaging classification dataset was obtained, and the original image data was preprocessed by VBM using fMRISurface software, including skull removal, segmentation, registration, and spatial smoothing. After removing the influence of the cerebellum, the brain was divided into multiple regions of interest (ROIs) using the corresponding templates. The standardized values ​​of the blood oxygen concentration-dependent contrast (BOLD) signals were calculated as the sequence features of the ROIs. Each sample obtained time series signal data containing N ROIs. The regions of interest were used as nodes, so the number of nodes was N and the sequence length was M. For the nth ROI, its features can be represented as an M-dimensional time series data. Therefore, the characteristics of a sample can be expressed as X = [x 1 ,…,x n ,…,x N ] T ∈R N×M , whose initial number of channels is 1; the data set is divided into a training set and a test set;

[0019] Step 1.2, set a fixed time window size, randomly sample M time points in the time window and add them to the training set, expand the training set to obtain the expanded training set; use the test set and the expanded training set as the resting-state functional magnetic resonance imaging data set;

[0020] The second step is to build a graph structure brain network:

[0021] Step 2.1, calculate the adjacency matrix A shared by all samples:

[0022] First, after z-score normalization of the features of each sample in step 1.1, the Pearson correlation coefficient (PCC) shared by all samples is calculated to construct a graph-structured brain network. Under standardized data, feature relationship metrics such as PCC, Euclidean distance, Cosine similarity, and square of Euclidean distance can be considered equivalent; the output range of PCC is [-1, 1], which can measure the similarity of features. For random variables A and B, the PCC calculation method is shown in formula (1):

[0023]

[0024] In formula (1), μ A , μ B , σA, σB are the sample mean and standard deviation of random variables A and B respectively, and E(·) represents the expected value of the random variable.

[0025] In order to further reduce the complexity of the model and prevent the graph convolutional network from being over-smoothed, the K-nearest neighbor algorithm is used to calculate the KNN graph to construct the graph structure brain network. For the rth ROI and the pth ROI, their time series features are and (x r ,x p ∈R M ), define its characteristic distribution distance D(x r ,x p ) is the Euclidean distance between time series, as shown in formula (2):

[0026]

[0027] For the current ROI, after calculating the distances to all other ROIs, its connection with the nearest 50% of the ROIs is retained to form a KNN graph.

[0028] In order to calculate the adjacency matrix A shared by all samples, we first use the feature concatenation function concat(·) to concatenate the ROI sequences of all samples, and calculate the Pearson correlation coefficient according to formula (1), and calculate its KNN graph at the same time. Finally, we calculate the adjacency matrix A shared by all samples according to formula (3).

[0029]

[0030] Among them, U 0 is the number of samples, Xu =[x 1 ,…,x n ,…,x N ] T ∈R N×M is the time series feature matrix of the u-th sample, KNN and concat represent the calculation of KNN graph and feature concatenation operation respectively;

[0031] Step 2.2: Construct a graph-structured brain network: Use the adjacency matrix A shared by all samples as the edge weight, and use the time series feature X of the sample as the feature of each node to construct a graph-structured brain network;

[0032] The third step is to use the classification prediction model for classification:

[0033] In step 3.1, a layer of S-TGConv module with 64 channels is used to extract spatial and temporal features from the input features. The size of the temporal kernel and spatial kernel used in the S-TGConv module are both 11, the step size is 1, and the features processed by the spatial graph convolution and the temporal graph convolution are concatenated and connected to a LeakyReLU activation function.

[0034] To further simplify the graph convolution operation, consider the 2D convolution operation in the case of only one frame, and the convolution kernel size is μ×σ, then a simplified graph convolution can be defined as:

[0035]

[0036] In formula (4), p(·) is a sampling function, which is used to limit the neighbor node range (neighborhood size) of the current node x (i.e., ROI). Note that this is only a spatial convolution operation on one time node; w is a filter weight function, which is independent of the position of the current node, so it can be shared on all input graphs; in order to perform spatial and temporal convolution operations at the same time, the sampling function p(·) needs to be redefined; for the i-th node v in the t-th frame, ti , first define its neighborhood range B(v ti )for:

[0037]

[0038] In formula (5), γ is the spatial kernel width, and λ is the temporal kernel width; Among them A tj represents the element in the tth row and jth column of the adjacency matrix A. BFS(i,j) calculates the node v based on the adjacency matrix A shared by all samples. tj To node v ti The minimum number of hops (Hop), so Represents node v tj and vti The path length; the parameters q and j are used to limit the range of the spatial and temporal convolution kernels, respectively, v qj represents the jth node at time point q, v tj represents the jth node at time point t; by defining the spatial kernel width γ and the temporal kernel width λ, the range of spatial convolution and temporal convolution in formula (5) is limited, thereby capturing the spatial and temporal features at the same time and avoiding the over-smoothness of the graph convolution operation; through formula (4)-formula (5), the spatial and temporal convolution kernel sizes are also limited, and then the node v is defined ti The sampling function p(·) is the neighborhood range B(v ti ), the graph convolution operation can be redefined as shown in formula (6):

[0039]

[0040] In formula (6), Z ti (v tj )=|{v tk = l tj (v tj )}| is used to balance the contribution between different parameters, l ti (v tj )Compute node v tj The feature representation, f in (v tj ) is the node x at time t j Features

[0041] After a layer of spatiotemporal graph convolution with 64 channels and 11 temporal and spatial convolution kernels, the LeakyReLU activation function is used to increase the nonlinear ability of the network. The formula of LeakyReLU is shown in formula (7).

[0042] f(z)=max(σz,z) (7)

[0043] Among them, z is a random variable, σz represents the standard deviation of the random variable z, and f(z) is the LeakyReLU activation function.

[0044] For each sample, input its feature X and the graph structure network composed of the adjacency matrix A, and after passing through an S-TGConv module, the output feature is F 1 ;

[0045] Step 3.2: Output the feature F in 3.1 1The output is to an S-TGConv module with 128 channels, 11 temporal convolution kernels, and 5 spatial convolution kernels. Considering that the brain network features have begun to be over-smoothed after two layers of S-TGConv modules, the spatial kernel size is reduced to reduce the number of spatial neighbor nodes sampled in the spatial convolution operation, and the output feature F is obtained. 2 ;

[0046] Step 3.3, feature F 2 The input is sent to the S-TGAtten module with a channel number of 128. The processing process of the S-TGAtten module is: first, a shared linear transformation weight matrix W is used on each node, and then a layer of parameters W is used. f The linear change layer is then split to obtain high-frequency features and low-frequency features, and these two features are concatenated through a layer with a parameter of W. t The linear change layer is finally transformed nonlinearly through the LeakyReLU activation function to obtain the feature F 3 , the formula for this step is:

[0047] F 3 =LeakyReLU(Concat((WF 2 )W f )W t ) (8)

[0048] Step 3.4, Feature F 3 After a normalized BN layer, then a Dropout layer, and two fully connected layers φ fc1 ,φ fc2 After a classify classification function, the final output classification result P is obtained c , as shown in formula (9):

[0049] P c =classify(φ fc2 (φ fc1 (φ dropout (φ BN (F 3 ))))) (9)

[0050] This completes feature learning based on spatiotemporal graph convolution;

[0051] Step 3.5: Use triplet graph convolution to enhance the training effect; given a sample X a As the original sample, randomly sample a positive sample X p and a negative sample X n (X p With X a The labels are consistent, and Xp With X a The labels are inconsistent); based on l 2 Regularization function, defining the triple loss function L of the current sample td for:

[0052]

[0053] Where O represents the number of triplets sampled in a mini-batch training, and f t (·) is the sampling function, Ф is the interval parameter, which can control the collected negative samples To the current sample The feature distance of , so the training difficulty of the model can be adjusted by adjusting the value of Ф; the present invention uses the Hard-mining strategy to select the smallest possible λ t value to ensure that the collection of more difficult classification Sample, improve the training effect of the model and reduce the degree of overfitting. The loss function outputs the current sample The probability values ​​of positive samples and negative samples;

[0054] Calculate the classification result P c The predicted label Use the loss function with regularization term and the triple loss function to enhance the training effect to calculate the classification result of step 3.4; the total loss function Loss is:

[0055]

[0056] in, is the category prediction value, y is the true value corresponding to each category, ε is the label smoothing parameter, and C is the number of categories;

[0057] The classification prediction model is trained using the expanded training set, and the optimal classification prediction model is saved for brain network classification to complete the classification prediction task;

[0058] In the fourth step, the collected functional magnetic resonance imaging data set is preprocessed according to step 1.1, and the time series features X of the processed samples and the adjacency matrix A shared by all samples in step 2.1 are used to form a graph structure brain network, which is input into the optimal classification prediction model in the third step to perform the classification task on the generated brain network;

[0059] Step 5: For any sample, for the output feature F in step 3.3 3 ,use Calculate the weight of the i-th ROI, e ij For feature F 3The element in the i-th row and j-th column; calculate the weights of N ROIs, and indicate the importance of the ROI according to the weights. Later, the actual number of the ROI is used to determine the part of the brain where it is located.

[0060] The brain network classification method is used in brain network classification of gender classification, Alzheimer's disease symptom classification, age classification, smoking and drinking, etc.

[0061] The present invention also protects a computer-readable storage medium, in which a computer program is stored. The computer program is suitable for executing the above-mentioned brain network classification method based on spatiotemporal graph convolution when loaded by a computer.

[0062] A brain network classification platform based on spatiotemporal graph convolution, which executes the brain network classification method based on spatiotemporal graph convolution, comprises:

[0063] A data acquisition module, used for acquiring resting-state functional magnetic resonance imaging data;

[0064] A region of interest extraction module, for extracting multiple regions of interest of a sample using a voxel-based analysis method through the Human Connectome Project Multimodal Parcellation Atlas (HCP-MMP);

[0065] The graph structure network construction module is used to construct a graph structure brain network based on the time series features of the region of interest and obtain the adjacency matrix A shared by all samples;

[0066] Classification prediction model, used to classify and predict graph structured brain networks;

[0067] The localization module uses the spatiotemporal graph attention module to obtain important spatial nodes after the spatiotemporal graph convolution module. For time nodes and ROIs that may be significantly different from the classification task, the spatiotemporal graph attention module tends to assign higher weights, and can more accurately locate key brain areas according to the weight size;

[0068] Data output module, used to display and output classification results.

[0069] Compared with the prior art, the present invention adopts the above technical solution, and the outstanding substantive features and significant improvements of the present invention are as follows:

[0070] (1) The brain network classification method based on spatiotemporal graph convolution is to divide the brain into multiple regions of interest using a voxel-based morphological analysis method, and perform a detailed analysis of resting-state functional magnetic resonance imaging data. In order to extract the time series characteristics of the brain, including the characteristics of the regions of interest and their correlation and temporal dynamics, the brain is modeled as a graph-structured brain network, and the spatiotemporal graph convolution is used to simultaneously extract spatial and temporal features. The method of the present invention proposes an effective spatiotemporal convolutional brain network classification method based on graph structure, especially for graph structure models containing time series, which can simultaneously consider the temporal dynamics of the data (spontaneous activity and periodic changes of the brain) and spatial structural relationships (correlation between ROIs and their characteristics); the spatiotemporal graph convolution module is used to simultaneously extract the spatial and temporal features of the model, and the spatiotemporal graph attention module is used to obtain important spatial nodes after the spatiotemporal graph convolution module. For time nodes and ROIs that may have important differences with the classification task, the spatiotemporal graph attention mechanism tends to give higher weights, so that the key brain areas can be located more accurately. During training, the Hard-mining strategy is used to randomly sample positive and negative samples, and a triple loss function is added to improve the training effect of the model.

[0071] (2) Compared with other sliding window and time series interception methods, the present invention adopts a window random sampling method. Through multiple random sampling, it can fully utilize the time information of resting-state functional magnetic resonance imaging data, while reducing the impact of noise on the model, which is beneficial to enhancing the training effect of the model and improving the classification accuracy based on the brain network.

[0072] (3) CN114287908A discloses a brain connection classification method of multi-passband graph convolution, which uses the functional imaging features of the whole brain to extract the functional connection network from the white matter and gray matter partitions of the brain, and uses multi-passband feature learning on the basis of the existing graph convolution network to fuse low-pass filtering and band-pass filtering through graph scattering convolution. This method does not consider the temporal dynamics of functional magnetic resonance imaging data, and performs correlation analysis on the blood oxygen concentration dependent contrast signal at all time points to construct the brain network of the whole brain, which cannot fully utilize the temporal dynamics. Compared with CN114287908A, the method of the present invention uses a spatiotemporal graph convolution structure to simultaneously extract the spatial features and temporal dynamics in the graph network constructed based on the time series, including the correlation between ROIs, the features of ROIs and the spontaneous activity of the brain during the scanning period of the resting state functional magnetic resonance imaging, fully extracts the discriminative information of the data, enhances the classification effect of the model, and improves the interpretability of the model.

[0073] (4) CN114863181A discloses a gender classification method based on prediction probability knowledge distillation. This method uses a teacher-student knowledge distillation model to train ResNet18 and CNN networks to achieve gender prediction. This method uses a deep convolutional model to perform gender classification tasks, which improves the classification accuracy, but ignores the interpretability of the brain network model and cannot explain the prediction results from the perspective of brain structure and function. Compared with CN114863181A, the method of the present invention uses a spatiotemporal graph attention module, assigns different weights to the features extracted by the spatiotemporal graph convolution module, improves the classification effect of the model, and can locate the key areas and key time points of the brain network.

[0074] (5) CN110477909A discloses a gender classification method based on resting-state EEG data. After preprocessing and removing artifacts from the original EEG data, the data is input into a convolutional neural network for training and testing. Although this method removes the influence of artifacts, it only uses the most primitive CNN network for gender classification tasks and does not fully consider the influence of the original data noise on the training process. Compared with CN110477909A, this method increases the training effect of the model by repeated random sampling and using a triplet loss function, thereby reducing the influence of noise on the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] The present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0076] Figure 1 This is an overall flow chart of the brain network classification method based on spatiotemporal graph convolution of the present invention.

[0077] Figure 2 This is a processing flow chart of feature extraction of the S-TGConv module and the S-TGAtten module in the method of the present invention.

[0078] Figure 3 Schematic diagram of the specific structure of the S-TGAtten module in the method of the present invention. DETAILED DESCRIPTION

[0079] Figure 1 As shown in the embodiment, the overall prediction and classification process of the brain network classification method based on spatiotemporal graph convolution of the present invention is as follows:

[0080] Input brain imaging data → preprocess resting-state functional magnetic resonance imaging data → build a classification prediction model to extract data spatial and temporal features, calculate the prediction results → calculate the loss value between the prediction results and the true value, optimize the classification prediction model, use the optimal classification prediction model for classification prediction → output the classification results.

[0081] Figure 2 As shown in the embodiment, the processing steps of feature extraction and feature learning of the S-TGConv module and the S-TGAtten module in the classification prediction model used in the present invention are as follows:

[0082] Input data → Calculate the Pearson correlation coefficient and the group-shared adjacency matrix A to build a graph-structured brain network → Use an S-TGConv module with 64 channels and 11 spatial and temporal kernel sizes to extract spatial and temporal features → Use a spatial graph convolution with 128 channels and a spatial kernel size of 5 to extract spatial features, and use a temporal graph convolution with a temporal kernel size of 11 to extract temporal features. The spatial graph convolution and the temporal graph convolution constitute the spatiotemporal graph convolution → Use a spatiotemporal attention module with 128 channels to obtain the temporal and spatial attention weights → Sum the features processed and spliced ​​by a spatiotemporal graph convolution with 128 channels, a spatial kernel size of 5, and a temporal kernel size of 11 with the features directly output by the spatiotemporal attention module → Use the LeakyReLU activation function to enhance the nonlinear ability → Output features.

[0083] Figure 3 As shown in the embodiment, the specific structure of the S-TGAtten module proposed in the present invention is as follows: first, input the feature F 2 , using a shared linearly varying weight matrix W, and then using a layer with parameters W f The linear change layer is then used to extract high-frequency features and low-frequency features, and these two features are concatenated in the form of high frequency first and low frequency later. t The linear change layer is finally activated by the LeakyReLU function to perform nonlinear changes and output the feature F 3 .

[0084] The entire network of the classification prediction model includes 1 layer of spatiotemporal graph convolution S-TGConv module with 64 channels, time kernel and spatial kernel size of 11, 1 layer of spatiotemporal graph convolution S-TGConv module with 128 channels, time kernel size of 11 and spatial kernel size of 5, 1 layer of spatiotemporal graph attention S-TGAtten module with 128 channels, a normalized BN layer, a Dropout layer, two fully connected layers, and a classification function. Two of the S-TGConv modules use different numbers of channels and kernel sizes. The S-TGAtten modules are connected in series according to the number of channels from small to large. The features processed and concatenated by the spatiotemporal graph convolution in the second S-TGConv module are summed with the features directly output by the S-TGAtten module. The output connection after the LeakyReLU activation function is connected to a normalized BN layer, a Dropout layer, two fully connected layers, and a classification function in sequence. The number of channels in the S-TGAtten module is consistent with that of the last S-TGConv module, which is used to extract weight information at different time points and different ROIs.

[0085] In the first step, the input resting-state functional magnetic resonance imaging data were preprocessed by performing skull removal, registration, spatial smoothing and other operations. Then, the blood oxygen concentration-dependent contrast signal in the voxel was mapped to 22 main ROIs using the HCP-MMP template, and data with a dimension of 22×1200×1 was obtained, including 22 ROIs, 1200 time nodes, and an initial channel number of 1.

[0086] In the second step, the first S-TGConv module is used to extract temporal and spatial features, with 64 channels, 11 temporal and spatial kernel sizes, and a step size of 1. The resulting feature map size is 22×128×64.

[0087] In the third step, the second S-TGConv module with 128 channels, 5 spatial kernels, 11 temporal kernels, and a stride of 1 is used, and the resulting feature map size is 22×128×128.

[0088] In the fourth step, a spatiotemporal graph attention S-TGAtten module with 128 channels is used to process the outputs of the two series-connected S-TGConv modules. The features after spatiotemporal graph convolution processing and splicing in the second S-TGConv module are summed with the features directly output by the spatiotemporal attention module, and then processed by the LeakyReLU activation function to obtain an output feature map size of 22×128×128.

[0089] In the fifth step, a normalized BN layer, a Dropout layer, a fully connected layer with an output dimension of 256, a fully connected layer and a classification function are used on the feature map output in the fourth step. The second fully connected layer is consistent with the number of categories of the classified task to obtain the final classification result.

[0090] In the sixth step, the model is trained using the loss function between the classification result in the fifth step and the true value, the best performing parameters on the validation set are saved during the model training process, and the best model is used for classification to implement a brain network classification method based on spatiotemporal graph convolution.

[0091] In the seventh step, use the 22×128×128 feature map output by the S-TGAtten module to obtain the weights of 22 ROIs, and define the importance of ROI in the classification task according to the weight size.

[0092] After preprocessing, the resting-state functional magnetic resonance imaging data to be tested is input into the optimal classification prediction model for classification, and the classification results are output. The attention mechanism of the S-TGAtten module is used to learn the most critical brain areas in the brain connection for the classification task, thereby improving the classification accuracy based on the brain network and locating the relevant brain parts at the same time.

[0093] Example 1

[0094] This embodiment uses the model structure of the brain network classification method based on spatiotemporal graph convolution to perform gender classification tasks. The specific steps are as follows:

[0095] The first step is to build a gender-classified resting-state functional magnetic resonance imaging dataset:

[0096] Step 1.1, resting state functional magnetic resonance imaging data preprocessing:

[0097] The resting-state functional magnetic resonance imaging data of the Human Genome Project HCP S1200 project were obtained, and the original image data were preprocessed by VBM using fMRISurface software, and operations such as skull removal, segmentation, alignment and spatial smoothing were performed respectively. After removing the influence of the cerebellum, the cerebral cortex was divided into N = 22 main ROIs using the HCP-MMP template, and the final features were obtained after calculating the standardized values ​​of the blood oxygen concentration-dependent contrast signal; among them, the data set contains a total of 1112 samples, but some samples have the phenomenon of too short time series in sampling, and some data have label errors, data errors, excessive noise, and incomplete follow-up data. Finally, 1096 adolescent samples (age: 22-25, male: 598, female: 498) were retained, with a sampling time point M = 1200 frames, a sampling time window TR = 0.72 seconds, and a total sampling time length of 15 minutes. The time series feature of each sample is X = [x1 ,…,x n ,…,x N ] T ∈R N×M ; Divide the data set into training set and test set according to five-fold cross validation;

[0098] Step 1.2: Set a fixed time window size, randomly sample M = 1200 time points in the time window and add them to the training set, expand the training set to obtain the expanded training set; use the test set and the expanded training set as the resting-state functional magnetic resonance imaging dataset; set the time window size to 128, randomly sample 1200 time points in the time window and add them to the training set, then the initial feature dimension of the samples in the training set is 22×128×1, which is represented by 22 ROIs, 128 randomly sampled time points, and an initial channel number of 1; thus, the construction of the gender-classified resting-state functional magnetic resonance imaging dataset is completed;

[0099] The second step is to build a graph-structured brain network:

[0100] Step 2.1, calculate the adjacency matrix A shared by all samples:

[0101] First, the features of each sample are z-score standardized, and the Pearson correlation coefficient (PCC) shared by all samples is calculated to construct a graph-structured brain network. Under standardized data, feature relationship metrics such as PCC, Euclidean distance, Cosine similarity, and square of Euclidean distance can be considered equivalent; the output range of PCC is [-1, 1], which can measure the similarity of features. For random variables A and B, the PCC calculation method is shown in formula (1):

[0102]

[0103] In formula (1), μ A , μ B , σA, σB are the sample mean and standard deviation of random variables A and B respectively, and E(·) represents the expected value of the random variable.

[0104] In order to further reduce the complexity of the model and prevent the graph convolutional network from being over-smoothed, the K-nearest neighbor algorithm is used to calculate the KNN graph to construct the graph structure brain network. For the rth ROI and the pth ROI, their time series features are and (x r ,x p ∈R M ), define its characteristic distribution distance D(x r ,x p) is the Euclidean distance between time series, as shown in formula (2):

[0105]

[0106] For the current ROI, after calculating the distances to all other ROIs, its connection with the nearest 50% of the ROIs is retained to form a KNN graph.

[0107] In order to calculate the adjacency matrix A shared by all samples, we first use the feature concatenation function concat(·) to concatenate the ROI sequences of all samples, and calculate the Pearson correlation coefficient according to formula (1). At the same time, we calculate the KNN graph, and multiply the Pearson correlation coefficient and the KNN graph to obtain the adjacency matrix A shared by all samples. The specific implementation is shown in formula (3):

[0108]

[0109] Among them, U 0 is the number of samples, X u =[x 1 ,…,x n ,…,x N ] T ∈R N×M is the time series feature matrix of the u-th sample. KNN, concat, and PCC represent the calculation of KNN graph, feature concatenation operation, and PCC operation, respectively.

[0110] Step 2.2: Construct a graph-structured brain network: Take the adjacency matrix A shared by all samples as the edge weight, and the time series feature X of the sample as the feature of each node to construct a graph-structured brain network;

[0111] The third step is to use the classification prediction model for classification:

[0112] In step 3.1, a S-TGConv module with a channel number of 64 is used to extract spatial and temporal features from the input features. The size of the temporal kernel and spatial kernel used here are both 11, with a step size of 1. After concatenation, a LeakyReLU activation function is used to implement the feature extraction process using the S-TGConv module.

[0113] To further simplify the graph convolution operation, consider the 2D convolution operation in the case of only one frame, and the convolution kernel size is μ×σ, then a simplified graph convolution can be defined as:

[0114]

[0115] In formula (4), p(·) is a sampling function, which is used to limit the neighbor node range (neighborhood size) of the current node x (i.e., ROI). Note that this is only a spatial convolution operation on one time node; w is a filter weight function, which is independent of the position of the current node, so it can be shared on all input graphs; in order to perform spatial and temporal convolution operations at the same time, the sampling function p(·) needs to be redefined; for the i-th node v in the t-th frame, ti , first define its neighborhood range B(v ti )for:

[0116]

[0117] In formula (5), γ is the spatial kernel width, and λ is the temporal kernel width; Among them A tj Represents the element in the tth row and jth column of the adjacency matrix A. BFS(i,j) calculates the node v according to the adjacency matrix A. tj To node v ti The minimum number of hops (Hop), so Represents node v tj and v ti The path length; the parameters q and j are used to limit the range of the spatial and temporal convolution kernels, respectively, v qj represents the jth node at time point q, v tj represents the jth node at time point t; by defining the spatial kernel width γ and the temporal kernel width λ, the range of spatial convolution and temporal convolution in formula (5) is limited, thereby capturing the spatial and temporal features at the same time and avoiding the over-smoothness of the graph convolution operation; through formula (4)-formula (5), the spatial and temporal convolution kernel sizes are also limited, and then the node v is defined ti The sampling function p(·) is the neighborhood range B(v ti ), the graph convolution operation can be redefined as shown in formula (6):

[0118]

[0119] In formula (6), Z ti (v tj )=|{v tk = l tj (v ti )}| is used to balance the contribution between different parameters, l tj (v tj )Compute node v tj The feature representation, f in (v tj ) is the node x at time t j Features After passing through a convolution layer with 64 channels and 11 temporal and spatial convolution kernels, the LeakyReLU activation function is used to increase the nonlinear ability of the network. The formula of LeakyReLU is shown in formula (7).

[0120] f(z)=max(σz,z) (7)

[0121] Where z is a random variable, σz represents the standard deviation of the random variable z, and f(z) is the LeakyReLU activation function;

[0122] For each sample, we input its feature X and the graph structure brain network constructed by the adjacency matrix A, and after passing through an S-TGConv module, we get the output feature F 1 , the dimensions are 22×128×64;

[0123] Step 3.2: Output the feature F in 3.1 1 Output to an S-TGConv module with 128 channels, 11 temporal convolution kernels, and 5 spatial convolution kernels to obtain the output feature F 2 , F 2 The dimensions are 22×128×128;

[0124] Step 3.3, feature F 2 Input to the S-TGAtten module with 128 channels. First, a shared linear transformation weight matrix W is used on each node, and then a layer with parameters W is used. f The linear change layer is then split to obtain high-frequency features and low-frequency features, and these two features are concatenated through a layer with a parameter of W. t The linear change layer is finally transformed nonlinearly through the LeakyReLU activation function to obtain the feature F 3 , the formula of this step is shown in formula (8), feature F 3 The dimensions are 22×128×128;

[0125] F 3 =LeakyReLU(concat((WF 2 )W f )W t ) (8)

[0126] Step 3.4, Feature F 3 After a normalized BN layer, then a Dropout layer, and two fully connected layers φ fc1 ,φ fc2 After a sigmoid activation function, the final output classification result P is obtained. c , as shown in formula (9):

[0127] P c =sigmoid(φ fc2 (φ fc1 (φ dropout (φ BN (F 3 ))))) (9)

[0128] Specifically, the Dropout layer parameter p drop = 0.5, the number of connections of the first fully connected layer is 256, and the number of connections of the second layer is 2. Then the sigmoid activation function is used to obtain the final classification scores P for the two genders. c ;

[0129] At this point, feature learning based on spatiotemporal graph convolution is completed and used for brain network classification to complete the gender classification prediction task.

[0130] Step 3.5: Use triplet graph convolution to enhance the training effect; given a sample X a As the original sample, randomly sample a positive sample X p and a negative sample X n (X p With X a The labels are consistent, and X p With X a The labels are inconsistent); based on l 2 Regularization function, defining the triple loss function L of the current sample td for:

[0131]

[0132] Calculate the classification result P c The predicted label Use the loss function with regularization term and the triple loss function to enhance the training effect to calculate the classification result of step 3.4; the total loss function Loss is:

[0133]

[0134] in, is the category prediction value, ε is the label smoothing parameter, y 0 Represents the label of the negative sample, y 1 Represents the label of the positive sample. Specifically, in this embodiment, y 0 =0 means male, y 1 =1 means female;

[0135] The classification prediction model is trained using the expanded training set, and the optimal classification prediction model is saved for brain network classification to complete the classification prediction task;

[0136] In the fourth step, the collected functional magnetic resonance imaging data set is preprocessed according to step 1.1, and the time series features X of the processed samples and the adjacency matrix A shared by all samples in step 2.1 are used to form a graph structure brain network, which is input into the optimal classification prediction model in the third step to perform classification tasks on the generated brain network;

[0137] Step 5: For any sample, for the output feature F in step 3.3 3 ,use Calculate the weight of the i-th ROI, e ij For feature F 3 The element in the i-th row and j-th column in ; calculate the weights of N ROIs, and use the weights to represent the importance of the ROI in the gender classification task. Later, the actual number of the ROI is used to determine the part of the brain where it is located.

[0138] Example 2

[0139] This embodiment uses the model structure of the brain network classification method based on spatiotemporal graph convolution to perform the classification task of Alzheimer's disease (AD). The specific steps are as follows:

[0140] The first step is to build a resting-state functional magnetic resonance imaging dataset for Alzheimer's disease:

[0141] Step 1.1, resting state functional magnetic resonance imaging data preprocessing:

[0142] The resting-state functional magnetic resonance imaging data of the ADNI dataset were obtained, and the original image data were preprocessed by VBM using the fMRISurface software, and operations such as skull removal, segmentation, registration, and spatial smoothing were performed. After removing the influence of the cerebellum, the cerebral cortex was divided into N = 22 main ROIs using the HCP-MMP template, and the final features were obtained by calculating the standardized values ​​of the blood oxygen concentration-dependent contrast signal. The time series length M was 1200 frames, and the time series feature of a sample can be expressed as X = [x 1 ,…,x n ,…,x N ] T ∈R N×M ; The dataset contains three types of samples, namely Alzheimer's disease (AD) patients, mild cognitive impairment (MCI) patients, and normal people (HC), among which MCI is an early symptom of AD;

[0143] In step 1.2, a fixed time window size is set, M time points are randomly sampled in the time window and added to the training set, and the training set is expanded to obtain the expanded training set; the test set and the expanded training set are used as the resting state functional magnetic resonance imaging data set; the time window size is set to 128, 1200 time points are randomly sampled in the time window and added to the training set, then the initial feature dimension of the samples in the training set is 22×128×1, which is represented by 22 ROIs, 128 randomly sampled time points, and the initial number of channels is 1;

[0144] This completes the construction of the resting-state functional magnetic resonance imaging dataset for Alzheimer's disease symptom classification;

[0145] The second step is to build a graph structure brain network:

[0146] Step 2.1, calculate the adjacency matrix A shared by all samples:

[0147] First, the features of each sample are z-score standardized, and the Pearson correlation coefficient (PCC) shared by all samples is calculated to construct a graph-structured brain network. Under standardized data, feature relationship metrics such as PCC, Euclidean distance, Cosine similarity, and square of Euclidean distance can be considered equivalent. The output range of PCC is [-1, 1], which can measure the similarity of features. For random variables A and B, the PCC calculation method is shown in formula (1):

[0148]

[0149] In formula (1), μ A , μ B , σA, σB are the sample mean and standard deviation of random variables A and B respectively, and E(·) represents the expected value of the random variable. In order to further reduce the complexity of the model and prevent the graph convolution network from being over-smoothed, the K-nearest neighbor algorithm is used to calculate the KNN graph to construct the graph structure brain network. For the rth ROI and the pth ROI, their time series features are and (x r ,x p ∈R M ), define its characteristic distribution distance D(x r ,x p ) is the Euclidean distance between time series, as shown in formula (2):

[0150]

[0151] For the current ROI, after calculating the distance between it and all other ROIs, the connection relationship with the 50% ROIs closest to it is retained to form a KNN graph. In order to calculate the adjacency matrix A shared by all samples, the feature concatenation function concat(·) is first used to concatenate the ROI sequences of all samples, and the Pearson correlation coefficient is calculated according to formula (1). At the same time, its KNN graph is calculated, and the two are multiplied to obtain the adjacency matrix A shared by all samples. The specific implementation is shown in formula (3):

[0152]

[0153] Among them, U 0 is the number of samples, X u =[x 1 ,…,x n ,…,x N ] T ∈R N×M is the time series feature matrix of the u-th sample. KNN and concat represent the calculation of KNN graph and feature concatenation operations respectively.

[0154] Step 2.2: Construct a graph-structured brain network: Take the adjacency matrix A shared by all samples as the edge weight, and the time series feature X of the sample as the feature of each node to construct a graph-structured brain network;

[0155] The third step is to use the classification prediction model for classification:

[0156] In step 3.1, a S-TGConv module with 64 channels is used to extract spatial and temporal features from the input features. The size of the temporal kernel and spatial kernel used here are both 11, with a step size of 1, and a LeakyReLU activation function after concatenation.

[0157] To further simplify the graph convolution operation, consider the 2D convolution operation in the case of only one frame, and the convolution kernel size is μ×σ, then a simplified graph convolution can be defined as:

[0158]

[0159] In formula (4), p(·) is a sampling function, which is used to limit the neighbor node range (neighborhood size) of the current node x (i.e., ROI). Note that this is only a spatial convolution operation at one time point; w is a filter weight function, which is independent of the position of the current node, so the weight function can be shared on all input graphs; in order to perform spatial and temporal convolution operations at the same time, the sampling function p(·) needs to be redefined; for the i-th node v in the t-th frame ti , first define its neighborhood range B(v ti )for:

[0160]

[0161] In formula (5), γ is the spatial kernel width, and λ is the temporal kernel width; Among them A tj Represents the element in the tth row and jth column of the adjacency matrix A. BFS(i,j) calculates the node v according to the adjacency matrix A. tj To node v ti The minimum number of hops (Hop), so Represents node v tj and v ti The path length; the parameters q and j are used to limit the range of the spatial and temporal convolution kernels, respectively, v qj represents the jth node at time point q, v tj represents the jth node at time point t; by defining the spatial kernel width γ and the temporal kernel width λ, the range of spatial convolution and temporal convolution in formula (5) is limited, thereby capturing the spatial and temporal features at the same time and avoiding the over-smoothness of the graph convolution operation; through formula (4)-formula (5), the spatial and temporal convolution kernel sizes are also limited, and then the node v is defined ti The sampling function p(·) is the neighborhood range B(v ti ), the graph convolution operation can be redefined as shown in formula (6):

[0162]

[0163] In formula (6), Z ti (v tj )=|{v tk = l tj (v tj )}| is used to balance the contribution between different parameters, l ti (v tj )Compute node v tj The feature representation, f in (v tj ) is the node x at time t j Features After passing through a convolution layer with 64 channels and 11 temporal and spatial convolution kernels, the LeakyReLU activation function is used to increase the nonlinear ability of the network. The formula of LeakyReLU is shown in formula (7).

[0164] f(z)=max(σz,z) (7)

[0165] Where z is a random variable, σz represents the standard deviation of the random variable z, and f(z) is the LeakyReLU activation function. For each sample, its feature X and adjacency matrix A are input, and feature extraction is performed according to the rules of formula (4); after passing through an S-TGConv module, the output feature is F 1 , the dimensions are 22×128×64;

[0166] Step 3.2: Output the feature F in 3.1 1 Output to an S-TGConv module with 128 channels, 11 temporal convolution kernels, and 5 spatial convolution kernels, and the output feature is F 2 , F 2 The dimensions are 22×128×128;

[0167] Step 3.3, feature F 2 Input to the S-TGAtten module with 128 channels; first use a shared linear transformation weight matrix W on each node, then use a layer with parameters W f The linear change layer extracts high-frequency features and low-frequency features, and concatenates these two features through a layer with a parameter of W t The linear change layer is finally transformed nonlinearly through the LeakyReLU activation function to obtain the feature F 3 , the formula for this step is, F 3 The dimensions are 22×128×128:

[0168] F 3 =LeakyReLU(conncat((WF 2 )W f )W t ) (8)

[0169] Step 3.4, Feature F 3 After a normalized BN layer, a Dropout layer, an adaptive average pooling layer, two fully connected layers and a softmax activation function, the final output classification result P is obtained. c , as shown in formula (9):

[0170] P c =softmax(φ fc2 (φ fc1 (φ dropout (φ BN (F 3 ))))) (9)

[0171] Specifically, the Dropout layer parameter p drop= 0.5, the number of connections in the first fully connected layer is 256, and the number in the second layer is 3. Then the classification score P of the Alzheimer's disease dataset is obtained through the softmax activation function. c ;

[0172] At this point, feature learning based on spatiotemporal graph convolution has been completed and used for brain network classification to complete the Alzheimer's disease classification task.

[0173] Step 3.5: Use triplet graph convolution to enhance the training effect; given a sample X a As the original sample, randomly sample a positive sample X p and a negative sample X n (X p With X a The labels are consistent, and X p With X a The labels are inconsistent); based on l 2 Regularization function, defining the triple loss function L of the current sample td for:

[0174]

[0175] Calculate the classification result P c The predicted label The classification results of step 2.5 are calculated using the loss function with regularization term and the triple loss function to enhance the training effect. The total loss function is:

[0176]

[0177] in, is the category prediction value, ε is the label smoothing parameter, y 0 Represents the label of the normal sample, y 1 represents the label of MCI patients, y 2 Represents the label of AD patients, specifically y 0 =0,y 1 =1,y 2 =2;

[0178] The classification prediction model is trained using the expanded training set, and the optimal classification prediction model is saved for brain network classification to complete the classification prediction task;

[0179] In the fourth step, the collected functional magnetic resonance imaging data set is preprocessed according to step 1.1, and the time series features X of the processed samples and the adjacency matrix A shared by all samples in step 2.1 are used to form a graph structure brain network, which is input into the optimal classification prediction model in the third step to perform the classification task on the generated brain network;

[0180] Step 5: For any sample, for the output feature F in step 3.3 3 ,use Calculate the weight of the i-th ROI, e ij For feature F 3 The element in the i-th row and j-th column in ; the weights of N ROIs are calculated, and the importance of ROI in the Alzheimer's disease symptom classification prediction task is represented by the weight size. Later, the actual number of ROI is used to determine the part of the brain.

[0181] After acquiring resting-state functional magnetic resonance imaging data, the classification method of the present invention uses a voxel-based analysis method to extract multiple regions of interest through the Human Connectome Project Multimodal Segmentation Atlas (HCP-MMP); using the split spatiotemporal graph convolution layer, spatial convolution and temporal convolution operations are performed on the extracted time series data respectively to extract spatial features and temporal features, and at the same time, the triplet loss function is used to improve the training performance and increase the classification accuracy of the model. Finally, an improved spatiotemporal graph attention module is used to further locate important brain areas.

[0182] Any matters not described in the present invention are applicable to the prior art.

Claims

1. A brain network classification method based on spatiotemporal graph convolution, It is characterized in that The classification method includes the following: After acquiring resting-state functional magnetic resonance imaging data, multiple regions of interest were extracted using the human connectome project multimodal parcellation atlas HCP-MMP using a voxel-based analysis method; A graph-structured brain network is constructed based on the time series characteristics of the region of interest and the adjacency matrix A shared by all samples; Construct a classification prediction model, input the graph structure brain network into the classification prediction model, and obtain the brain network classification result: The entire network of the classification prediction model includes 1 spatiotemporal graph convolution S-TGConv module with 64 channels, time kernel and space kernel size of 11, 1 spatiotemporal graph convolution S-TGConv module with 128 channels, time kernel size of 11 and space kernel size of 5, 1 spatiotemporal graph attention S-TGAtten module with 128 channels, a normalization layer BN layer, a Dropout layer, two fully connected layers, and a classification function; the two S-TGConv modules use different numbers of channels and kernel sizes, and the numbers are arranged from small to large according to the number of channels. The S-TGAtten module is connected in series, and the features processed and spliced ​​by the spatiotemporal graph convolution in the second S-TGConv module are summed with the features directly output by the S-TGAtten module. The output after the LeakyReLU layer is sequentially connected to a normalized BN layer, a Dropout layer, and two fully connected layers, and then the output of the classification prediction model is obtained through the classification function; the number of channels of the S-TGAtten module is consistent with the number of channels of the second S-TGConv module, which is used to extract weight information of different time points and different ROIs; The S-TGConv module includes a spatial graph convolution layer and a temporal graph convolution layer in parallel. The outputs of the spatial graph convolution layer and the temporal graph convolution layer are connected to the LeakyReLU activation function through a splicing operation to obtain the output of the spatiotemporal graph convolution S-TGConv module; The process of the spatiotemporal graph attention S-TGAtten module is as follows: for the features output by the spatiotemporal graph convolution S-TGConv module, first use a shared linear transformation weight matrix W on each node, and then use a layer with parameters W f The linear change layer extracts high-frequency features and low-frequency features, and concatenates the two features through a layer with a parameter of W t The linear change layer is finally transformed nonlinearly through the LeakyReLU activation function to obtain the output of the spatiotemporal graph attention S-TGAtten module.

2. A brain network classification method based on spatiotemporal graph convolution, the specific steps of the method are: The first step is to build a resting-state functional magnetic resonance imaging dataset: Step 1.1, resting state functional magnetic resonance imaging data preprocessing: A resting-state functional magnetic resonance imaging classification dataset was obtained, and the original image data was preprocessed with VBM using fMRISurface software, including skull removal, segmentation, registration, and spatial smoothing. After removing the influence of the cerebellum, the brain was divided into multiple regions of interest (ROIs) using the corresponding templates. The normalized values ​​of the blood oxygen concentration-dependent contrast BOLD signals were calculated as the time series features of the ROIs. Each sample obtained time series signal data containing N ROIs. The regions of interest were used as nodes, the number of nodes was N, and the sequence length was M. The time series features of the nth ROI were represented as an M-dimensional time series data. Then the time series feature of a sample is X = [x 1 ,…,x n ,…,x N ] T ∈R N×M , whose initial number of channels is 1, divides the data set into training set and test set; Step 1.2, set a fixed time window size, perform random sampling of the time window on M time points and add them to the training set, expand the training set to obtain the expanded training set; use the test set and the expanded training set as the resting-state functional magnetic resonance imaging data set; The second step is to build a graph structure brain network: Step 2.1, calculate the adjacency matrix A shared by all samples: Use the feature concatenation function concat(·) to concatenate the time series features of all samples according to the sequence number of ROIs, and calculate the Pearson correlation coefficient. Each element represents the correlation between brain regions. Use the feature concatenation function concat(·) to concatenate the time series features of all samples according to the sequence number of ROI, and then use the K-nearest neighbor algorithm to calculate the KNN graph. Multiply the Pearson correlation coefficient and the KNN graph to obtain the adjacency matrix A shared by all samples; in, U 0 is the number of samples, X u =[x 1 ,…,x n ,…,x N ] T ∈R N×M is the time series feature matrix of the u-th sample, KNN, concat, and PCC represent the calculation of KNN graph, feature concatenation operation, and PCC operation respectively; Step 2.2: Construct a graph-structured brain network: Use the adjacency matrix A shared by all samples as the edge weight, and use the time series feature X of the sample as the feature of each node to construct a graph-structured brain network; The third step is to use the classification prediction model for classification: Step 3.1: Input the graph structured brain network constructed in the second step into a layer of S-TGConv module with 64 channels to extract features. The size of the temporal kernel and spatial kernel used here are both 11, and the step size is 1; Step 3.2: Output the feature F in 3.1 1 Output to a layer of S-TGConv module with 128 channels, 11 temporal convolution kernels, and 5 spatial convolution kernels, and get the output feature F 2 ; Step 3.3, feature F 2 Input to the S-TGAtten module with 128 channels, randomly given a shared linear transformation weight matrix W, in F 2 Use a shared linear transformation weight matrix W on each ROI feature, and then use a layer with parameters W f The linear change layer extracts high-frequency features and low-frequency features, and concatenates the two features through a layer with a parameter of W t The linear change layer is finally transformed nonlinearly through the LeakyReLU activation function to obtain the feature F 3 , the formula for this step is: F 3 =LeakyReLU(conncat((WF 2 )W f )W t ) (8) Step 3.4, Feature F 3 After a normalized BN layer, a Dropout layer, and two fully connected layers φ fc1 ,φ fc2 After a classification function classify, the final output classification result P is obtained c , as shown in formula (6): P c =classify(φ fc2 (f fc1 (f dropout (f BN (F 3 ))))) (9) Step 3.5: Use triplet graph convolution to enhance training effect: Given a sample X a As the original sample, randomly sample a positive sample X p and a negative sample X n , X p With X a The labels are consistent, and X p With X a The labels are inconsistent; based on l 2 Regularization function, defining the triple loss function L of the current sample td for: Where O represents the number of triplets sampled in a mini-batch training, and f t (·) is the sampling function, Ф is the interval parameter; Calculate the classification result P c The predicted label Use the loss function with regularization term and the triple loss function to enhance the training effect to calculate the classification result of step 3.4; the total loss function Loss is: in, is the category prediction value, y is the true value corresponding to each category, ε is the label smoothing parameter, and C is the number of categories; The expanded training set is used to train the classification prediction model, and the optimal classification prediction model is saved for brain network classification to complete the classification prediction task; In the fourth step, the collected functional magnetic resonance imaging data set is preprocessed according to step 1.1, and the time series features X of the processed samples and the adjacency matrix A shared by all samples in step 2.1 are used to form a graph structure brain network, which is input into the optimal classification prediction model in the third step to perform the classification task on the generated brain network; Step 5: For any sample, for the output feature F in step 3.3 3 ,use Calculate the weight of the i-th ROI, e ij For feature F 3 The element in the i-th row and j-th column; calculate the weights of N ROIs, and indicate the importance of the ROI according to the weights. Later, the actual number of the ROI is used to determine the part of the brain where it is located.

3. The brain network classification method based on spatiotemporal graph convolution according to claim 2, It is characterized in that The calculation process of the KNN graph is: For the rth ROI and the pth ROI, their time series features are and Define its characteristic distribution distance D(x r ,x p ) is the Euclidean distance between time series, as shown in formula (1): For the current ROI, after calculating the distances to all other ROIs, the connection relationships with the 50% of ROIs closest to it are retained to form a KNN graph.

4. The brain network classification method based on spatiotemporal graph convolution according to claim 2, It is characterized in that The spatiotemporal graph convolution processing process in the spatiotemporal graph convolution module is: Define the sampling function p(·); for the i-th node v in the t-th frame ti , first define its neighborhood range B(v ti ) as: In formula (3), γ is the spatial kernel width, and λ is the temporal kernel width; Among them A tj represents the t-th row and j-th column element of the adjacency matrix A shared by all samples. BFS(i,j) calculates the node v according to the adjacency matrix A shared by all samples. tj To node v ti The minimum number of hops (Hop), so Represents node v tj and v ti The path length; the parameters q and j are used to limit the range of the spatial and temporal convolution kernels, respectively, v qj represents the jth node at time point q, v tj Represents the jth node at time point t; By defining the spatial kernel width Υ and the temporal kernel width λ, the range of spatial convolution and temporal convolution is limited, thereby capturing spatial and temporal features at the same time and avoiding the over-smoothness of graph convolution operations; Define node v ti The sampling function p(·) is the neighborhood range B(v ti ), the graph convolution operation is redefined as shown in formula (4): In formula (4), Z ti (v ij )=|{v tk = l tj (v tj )}| is used to balance the contribution between different parameters, l ti (v tj )Compute node v tj The feature representation, f in (v tj ) is the node x at time t j Features w is the filter weight function, which is independent of the position of the current node and is shared across all input graphs and learned during training; For each sample, its feature X and the adjacency matrix A shared by all samples are input, and feature extraction is performed according to the rules of formula (4) to achieve the purpose of using the S-TGConv module to perform spatiotemporal graph convolution processing.

5. The brain network classification method based on spatiotemporal graph convolution according to any one of claims 1 to 4, It is characterized in that The method is used in gender classification, Alzheimer's disease classification, age classification, and brain network classification of smoking and drinking.

6. A computer-readable storage medium, wherein a computer program is stored in the storage medium, and the computer program is suitable for executing the brain network classification method based on spatiotemporal graph convolution according to any one of claims 1 to 4 when loaded by a computer.

7. A brain network classification platform based on spatiotemporal graph convolution, It is characterized in that The brain network classification method based on spatiotemporal graph convolution according to any one of claims 1 to 4 is implemented, and the platform comprises: A data acquisition module, used for acquiring resting-state functional magnetic resonance imaging data; A region of interest extraction module is used to extract multiple regions of interest of a sample using a voxel-based analysis method through the Human Connectome Project Multimodal Segmentation Atlas HCP-MMP; The graph structure network construction module is used to construct a graph structure brain network based on the time series features of the region of interest and obtain the adjacency matrix A shared by all samples; Classification prediction model, used to classify and predict graph structured brain networks; The localization module uses the spatiotemporal graph attention module to obtain important spatial nodes after the spatiotemporal graph convolution module. For time nodes and ROIs that may be significantly different from the classification task, the spatiotemporal graph attention module tends to assign higher weights, and can more accurately locate key brain areas according to the weight size; Data output module, used to display and output classification results.

Citation Information

Patent Citations

  • Gender classification method based on resting EEG data

    CN110477909A

  • Multi-passband graph convolution fusion brain connection classification method

    CN114287908A

  • Gender classification method and system based on prediction probability knowledge distillation

    CN114863181A

  • Real-time sign language intelligent recognition method, device and system

    CN113221663A

  • Process and system for image classification

    WO2022041222A1