An extensible multi-level graph neural network model based on multi-modal image data

By constructing a multi-level graph neural network model of multi-scale residual network and brain hierarchical graph neural network, the fusion problem of multi-modal medical imaging data in disease recognition is solved, and efficient disease recognition and diagnosis is achieved.

CN115393269BActive Publication Date: 2025-07-29UNIV OF CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210822138.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-13
Publication Date
2025-07-29
Estimated Expiration
2042-07-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively utilize multimodal medical image data, especially image information and non-image information, and cannot systematically reflect brain structure atrophy, brain interval structure and functional association, resulting in poor disease recognition effect.

Method used

A scalable multi-level graph neural network model based on multi-scale residual network and brain hierarchical graph neural network is constructed. Image features are extracted through multi-scale residual network, combined with brain hierarchical graph neural network and population hierarchical graph neural network, and multi-modal images and non-image features are fused for disease recognition.

Benefits of technology

It improves the accuracy and efficiency of disease identification, can quickly identify diseases, provide reliable auxiliary diagnostic indicators, has low network complexity and high computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115393269B_ABST
    Figure CN115393269B_ABST
Patent Text Reader

Abstract

The present invention specifically relates to an extensible multi-level graph neural network model based on multi-modal brain imaging data. The steps include: preprocessing the multi-modal imaging data to be processed; in the training stage, constructing a multi-scale residual network for the preprocessed multi-modal imaging data to extract features of each modality; for each sample, constructing a brain-level graph neural network model to extract its corresponding brain-level features and attention coefficients between brain regions; for the sample population, constructing an extensible multi-level graph neural network model, and using a weighted inductive spatial-domain graph convolution operator to learn the sample population to achieve learning and classification of the sample data and complete model training; in the testing stage, after preprocessing and feature extraction of each sample data in the test set respectively, inputting it into the trained network model to complete the test. The present invention provides an extensible and adjustable learning framework for the continuous improvement of multi-modal imaging data analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a learning classification method for multi-modal image data, and particularly to an extensible multi-level graph neural network based on multi-modal brain image data. Background Art

[0002] In recent years, deep learning technology has been widely applied to medical image analysis and auxiliary diagnosis of diseases. Multi-modal medical images can complement each other and reflect the structural, functional, and metabolic changes of the brain from different perspectives. How to effectively utilize multi-modal medical image data, improve the efficiency and accuracy of disease diagnosis, and provide effective auxiliary diagnosis indicators for clinicians is an important problem to be solved in computer-aided diagnosis. The present invention mainly relates to an extensible multi-level graph neural network based on multi-modal brain image data.

[0003] Magnetic Resonance Imaging (MRI) has been widely used in the clinical diagnosis of brain diseases. Among them, structural magnetic resonance imaging mainly includes T1-weighted images and Diffusion Tensor Imaging (DTI), which can effectively present changes at the structural level such as gray matter atrophy and nerve fiber damage in the brain; Functional MRI (fMRI) can effectively reflect the degree of brain function degradation; Quantitative Susceptibility Mapping (QSM) can effectively reflect the iron deposition characteristics in brain tissue. Positron Emission Computed Tomography (PET), as an imaging method reflecting metabolism, has important significance for the early detection of neurodegenerative diseases. Multi-modal medical images can reflect the development and changes of neurodegenerative diseases from different aspects. The computer-aided diagnosis method based on multi-modal medical images can greatly improve the efficiency and accuracy of quantitative analysis of image data, making the diagnosis results more accurate and achieving the purpose of early detection, early diagnosis, and early treatment of diseases. In traditional machine learning methods, features related to diseases are extracted from multi-modal data, and classifiers such as support vector machines, logistic regression, and AdaBoost are combined to classify and identify diseases; in deep learning methods, the image data of the whole brain or a certain brain region (such as the hippocampal region) is regarded as an ordinary 3D image or processed into 2D slice images, and then used as the input of the deep learning model to complete classification, which can deeply mine and effectively learn brain image information. However, traditional machine learning methods or other deep learning methods have problems such as being unable to systematically reflect brain structure atrophy, structural associations and functional association relationships between brain regions, and being unable to effectively combine non-image information of multi-modal data, and the recognition effects for different disease courses are poor. The human brain has a complex network structure. The graph neural network method can comprehensively analyze the image features and connection relationship features of brain regions by modeling the human brain as a graph, and effectively utilize non-image information to establish similarity relationships between subjects by modeling the subjects as graphs.

[0004] Therefore, in the research on using graph neural networks for the identification of neurodegenerative diseases based on multi-modal medical image data, it will be able to effectively learn and analyze the characteristics of brain structure atrophy, structural connection relationships and functional connection relationships between brain regions, and provide more reliable auxiliary diagnostic indicators for disease identification and diagnosis under different disease courses. Summary of the Invention

[0005] The object of the present invention is to construct an extensible multi-level graph neural network model based on multi-modal magnetic resonance imaging data, realize the learning, extraction and analysis of multi-modal image features of sample data, and further complete the rapid identification of diseases, so as to improve the classification performance of computer-aided diagnosis of multi-modal image data.

[0006] The present invention is realized through the following technical solutions, including the following steps:

[0007] (1) For the multi-modal image data to be processed, perform preprocessing operations using general medical image preprocessing software to obtain various parameter maps of the multi-modal image data and the correlation coefficients between brain regions.

[0008] (2) In the training stage, for the preprocessed multi-modal image data, construct a feature extraction network based on a multi-scale residual network to extract features from each modal data.

[0009] (3) In the training stage, for each sample, based on the image features already extracted by the feature extraction network, construct a brain-level graph neural network model to extract its corresponding brain-level features and the attention coefficients between brain regions.

[0010] (4) In the training stage, for the sample population, based on the features extracted by the feature extraction network and the brain-level graph neural network, construct an extensible multi-level graph neural network model to realize the learning and classification of sample data and complete the model training.

[0011] (5) In the testing stage, after preprocessing and feature extraction of each sample data in the test set respectively, input it into the trained network model to complete the test.

[0012] Further, in the above technical solution, the construction of the feature extraction network based on the multi-scale residual network in step (2) is specifically as follows:

[0013] (1) Along the three directions of the coronal, sagittal, and transverse planes of the three-dimensional image, accumulate all the slices in each direction and convert them into two-dimensional images of the human brain in the three directions, which are used as the three-channel input of the feature extraction network.

[0014] (2) The feature extraction network uses the ResNet-18 network as the backbone network, including three input channels, 17 convolutional layers, and 1 fully connected layer.

[0015] (3) Replace the 3×3 convolutional kernel in the original ResNet-18 network with 1×1, 3×3, 5×5 convolutional kernels and a 3×3 pooling layer to form a multi-scale residual module.

[0016] (4) Embed an attention mechanism module in the network. The attention mechanism mask is generated by performing a 3×3 convolution on the features of all input attention modules. This mask is used to re-weight the features to enhance the network's attention to key features, thus completing the construction of the "feature extraction network".

[0017] Furthermore, in the above technical solution, the construction of the brain hierarchical graph neural network model in step (3) is specifically as follows:

[0018] (1) In the brain hierarchical graph neural network model, extract brain hierarchical image features based on the correlation relationship between brain regions. Among them, the brain of each sample is represented as a graph, with different brain regions as the vertices of the graph, and the extracted single-modal image features as the vertex features. There are two construction methods for the edge weights connecting the vertices: the correlation coefficient between brain regions and the graph attention coefficient.

[0019] (2) The graph network model sequentially includes a GCN convolutional layer, a Topk pooling layer, a GCN convolutional layer, and a Readout layer, as shown in the following formula:

[0020]

[0021] Among them, Z1 is the result of the first graph convolution, σ is the linear activation function σ(·) = max(0, ·), X is the original feature matrix, and W (0) is the parameter matrix of the first graph convolution. is the normalized adjacency matrix, where A is the adjacency matrix of the graph, I is the identity matrix, represents the degree of node i. After applying the graph pooling operation on the graph after the first convolution, a new graph is obtained, and its corresponding feature matrix and adjacency matrix are and A p , and the normalized adjacency matrix of the new graph is obtained from A p in the same normalization manner. Finally, after another graph convolution operation and Readout operation, the final output result representing the graph can be obtained:

[0022]

[0023] Among them, W (1) is the parameter matrix of the second graph convolution, and Z2 is the output result of the second graph convolution.

[0024] (3) According to the characteristics of different modality data, the correlation coefficient between brain regions can be extracted through the preprocessing process or obtained by calculating the correlation coefficient of the extracted feature vectors between brain regions.

[0025] (4) The graph attention coefficient can be obtained by learning through the graph network model of the graph attention mechanism. This model includes an embedding layer, an attention propagation layer, a Topk pooling layer, an embedding layer, and a Readout layer in sequence of connection. The formula is as follows:

[0026] H1 = σ(XV (0) );

[0027] where H1 is the result of the first embedding layer, σ is the linear activation function, X is the original feature matrix, and V (0) is the weight matrix of the first embedding layer; after the first embedding layer is the attention propagation layer, and its output H2 is:

[0028] H2 = PH1;

[0029] where H2 is the result of the first embedding layer, and the propagation matrix is composed of the attention coefficients between nodes. N is the number of nodes, and its calculation formula is as follows:

[0030]

[0031] where P ij is the attention coefficient of vertex j to i, β is the shared weight of the propagation layer, h i and h j represent the feature vectors of vertices i and j, exp represents the exponential operation, represents all the neighborhood nodes of vertex i, then vertex k can be taken as all the neighborhood nodes of vertex i and vertex i itself. cosh i , h j ) represents the cosine value of the angle between the feature vectors h i and h j ; ||.|| represents the L2 norm; after applying the TopK pooling operation on H2, a new feature matrix is obtained. Finally, after another graph convolution operation and Readout operation, the output result of the graph model using the graph attention mechanism can be obtained:

[0032]

[0033] where V (1) is the weight matrix of the first embedding layer, and Z is the final output result.

[0034] Furthermore, in the above technical solution, the construction of the expandable multi-level graph neural network model in step (4) is as follows:

[0035] (1) Based on the feature extraction network and the brain hierarchical graph neural network, the image features have been extracted and input into the population hierarchical graph neural network to construct an extensible multi-level graph neural network model; during the training process, the training set samples of the model are represented as a graph, each sample is regarded as a vertex of the graph, and the non-image information of the sample (age, gender, clinical score, education level, etc.) is used as the edge weight information of the connection between vertices, and the multi-modal image combination features of each sample are used as vertex features; among them, the multi-modal image combination features can be the brain hierarchical image features extracted by different modal data through the brain region graph neural network respectively, or the features obtained by cascading the image features extracted by some modal data through the feature extraction network and the brain hierarchical image features extracted by some modal data through the brain region graph model; the edge weights between vertices in the graph network are encoded according to the similarity of samples, and the similarity between two samples m and n can be represented by A mn denote:

[0036]

[0037] where K represents the types of non-image information, S m and S n represent the feature vectors of samples m and n, and the feature similarity between samples m and n can be obtained using distance-related methods, and the formula is as follows:

[0038]

[0039] where and are the averages of vectors S n and S m respectively, ⊙ represents the dot product operation, and ||.||2 represents the L2 norm; represents the measure of some non-image information M:

[0040]

[0041] where θ represents the threshold, and M k (.) represents a certain non-image information;

[0042] (2) The population hierarchical graph neural network model includes two weighted inductive W-GraphSAGE graph convolutional layers with optimized aggregation methods and one Softmax layer; among them, the W-GraphSAGE graph convolution formula is as follows:

[0043]

[0044] where represents the feature representation of node u in the (k - 1)th layer, W is the weight matrix of this layer, is the feature representation of node v at the k-th layer; σ is the linear activation function, a u are all the neighbors of node v of the normalized weight coefficients:

[0045]

[0046] wherein, represents the edge weight between the neighbor node u and node v, represents that vertex k can be taken as all the neighborhood nodes of vertex v and vertex v itself, exp represents the exponential operation, softmax v (e u ) represents the normalization operation.

[0047] Furthermore, in the above technical solution, after preprocessing and feature extraction are respectively performed on each sample data in the test set, it is input into the trained network model to complete the test, specifically as follows:

[0048] (1) Preprocess the different modality data of the test set data respectively;

[0049] (2) Input the preprocessed multi-modal data into the feature extraction network respectively to extract the image features corresponding to different brain regions of different modality data;

[0050] (3) For some modality data, use the brain hierarchical graph neural network to extract the brain hierarchical image features and the inter-brain region attention coefficients;

[0051] (4) Input the extracted image features and non-image features into the multi-level graph neural network model to complete the test.

[0052] Compared with the prior art, the beneficial effects of the present invention are:

[0053] (1) The present invention proposes an extensible multi-level graph neural network model based on multi-modal image data, which is constructed by embedding a multi-scale residual network and a brain hierarchical graph neural network model into a population hierarchical graph neural network model. This model has good scalability, effectively fuses clinical multi-modal image features and non-image features, and provides an extensible and adjustable learning framework for the continuous improvement of clinical data analysis.

[0054] (2) Use a multi-scale residual network (ResNet) as the feature extraction network to extract the features of different modality data, fully mine the hidden features contained in different voxels of each brain region, and at the same time have a lower network complexity and higher computational efficiency.

[0055] (3) The brain hierarchical graph neural network model models each sample brain, completing the comprehensive learning of brain region image features and the connection relationship features between brain regions; meanwhile, combining the graph attention mechanism to extract the attention coefficients between brain regions, providing auxiliary references for disease course recognition and the extraction of imaging biomarkers.

[0056] (4) The population hierarchical graph neural network model models the sample population, weights the inductive spatial domain graph convolution operator, and uses the sample similarity relationship established by non-image information to achieve rapid disease recognition, with high computational efficiency and accuracy. The prediction time for a single new sample is 0.15% of the conventional frequency domain graph convolution method, and it can quickly predict diseases for subjects. Description of the Drawings

[0057] Figure 1 This is the algorithm flow chart of the scalable multi-level graph neural network model for AD recognition of the present invention.

[0058] Figure 2 This is the structural diagram of the feature extraction network in the model of this aspect.

[0059] Figure 3 This is the schematic diagram of the brain hierarchical graph neural network model in the model of the present invention.

[0060] Figure 4 This is the graph network model based on the graph attention mechanism in the model of the present invention.

[0061] Figure 5 This is the schematic diagram of the population hierarchical graph neural network model in the model of the present invention.

[0062] Figure 6 This is the relationship graph between the target brain region and other highly relevant brain regions in the experimental results of the present invention. Detailed Embodiments

[0063] The following details the embodiments of the present invention with reference to the accompanying drawings. These embodiments are implemented on the premise of the technical solution of the present invention, and detailed implementation manners and processes are given. However, the protection scope of the present invention is not limited to the following embodiments.

[0064] The present invention takes Magnetic Resonance Imaging (MRI) data as the application object, and learns and classifies MRI data of samples with Alzheimer's Disease (AD), Mild Cognitive Impairment (MCI), and Normal Control (NC). In this experiment, three modalities of data, namely T1-weighted data, DTI data, and fMRI, are selected. The data used in the experiment comes from the publicly available dataset Alzheimer's Disease Neuroimaging Initiative (ADNI) and Peking University Third Hospital. The data from Peking University Third Hospital includes 149 AD patients, 95 MCI patients, and 102 NCs, with their average ages being 72.57, 73.97, and 69.16 respectively, and their average MMSE scores being 17.97, 27.16, and 28.54 respectively; the data from ADNI includes 42 AD patients, 65 MCI patients, and 65 NCs, with their average ages being 73.36, 78.35, and 81.25 respectively. The MMSE scores of the AD data are all missing, and the average MMSE scores of the MCI and NC data are 27.14 and 28.94 respectively. In this experiment, a five-fold cross-random experiment is conducted on all the training data and test data of the two datasets, that is, 80% of the individual data is selected as training each time, and the rest is used as testing, and the average value is taken as the final result after repeating five times.

[0065] The implementation process of the scalable multi-level graph neural network model based on multi-modal image data in the present invention is as follows:

[0066] In the first step, preprocessing operations are performed on the multi-modal MRI image data to be processed. In this embodiment, the specific operations are as follows:

[0067] First of all, in this embodiment, the Spm8 tool running on the Matlab platform is used to process the T1-weighted data of each sample, including removing the skull, retaining the brain tissue, segmenting the brain tissue image into gray matter, white matter, and cerebrospinal fluid, and aligning the image to the standard space with the T1 standard image in the Montreal Neurological Institute (MNI) brain template as the template.

[0068] Then, the FSL tool running on the Linux system was used to process the DTI data of each sample. First, the b0 images were extracted and skull-stripped. A synthetic b0 image was generated using a synthetic algorithm, and Topup correction was performed using the original b0 image and the synthetic b0 image. The Eddy tool in FSL was used to align all other original DTI images with the b0 image to correct the head movement and eddy current distortion of the subjects during data acquisition. The Dtifit module in FSL was used to perform diffusion tensor fitting at each voxel in the subject's brain, and parametric maps of fractional anisotropy (FA), mean diffusivity (MD), axial diffusivity (AxD), and radial diffusivity (RD) were calculated. All individual FA images were nonlinearly aligned with the FMRIB58_FA template in the standard space through the TBSS (Trac-Based Spatial Statistics) algorithm implemented in FSL. Subsequently, the aligned transformation matrix was applied to the MD, AxD, and RD parametric maps to achieve the purpose of alignment and registration. The ICBM-DTI-81 template was used to divide each parametric map into 50 white matter regions of interest. The PROBTRACKX tool in FSL was used to perform probabilistic fiber tractography to obtain the white matter fiber tract counts between brain regions.

[0069] Finally, the DPARSF (Data Processing Assistant for Resting-State fMRI), provided in the toolbox for brain imaging data processing and analysis (DPABI) running on Matlab software, was used to process the fMRI data of each sample, including removing the first 10 imaging time points, slice time correction, head movement correction, spatial smoothing, band-pass filtering (0.01 - 0.08 Hz), MNI space normalization, calculating the amplitude of low frequency fluctuation (ALFF) and regional homogeneity (ReHo) parametric maps. Using the AAL template in the MNI space, the average BOLD time series of 90 brain regions were obtained from the fMRI data of each subject, and the Pearson correlation coefficient between the average BOLD time series of pairwise brain regions was calculated to obtain the functional connectivity between brain regions.

[0070] In the second step, in the training stage, for the preprocessed multi-modal image data, a feature extraction network based on a multi-scale residual network was constructed to extract features from each modal data. In this embodiment, the specific operations are as follows:

[0071] First, the feature extraction network uses a ResNet-18 network, which contains 17 convolutional layers and 1 fully connected layer. Its network structure diagram is as shown in Figure 1 . In this model, it includes a multi-scale residual module and an attention mechanism module composed of 1×1, 3×3, 5×5 convolutional kernels and a 3×3 pooling layer. The attention mechanism mask is generated by performing a 3×3 convolution on the features of all input attention modules. Each input three-dimensional MRI data is accumulated along the coronal, sagittal, and transverse directions, and all slices in each direction are accumulated to convert into two-dimensional images of the human brain in three directions, which are used as the three-channel input of the feature extraction network. Adam is selected as the optimizer, and it is trained for 50 epochs with a batch size of 32, an L2 regularization rate of 0.001, and a learning rate of 0.001.

[0072] Then, the gray matter images segmented from the T1-weighted data of each training set sample are input into the feature extraction network to complete training and learning, and 16-dimensional feature vectors can be extracted in each brain region; the FA, MD, AxD, and RD parameter maps obtained from the DTI data of each sample are input into the feature extraction network to complete training and learning, and 64-dimensional feature vectors can be extracted from each parameter map in each brain region; the ALFF and ReHo obtained from the fMRI data of each sample are input into the feature extraction network to complete training and learning, and 64-dimensional feature vectors can be extracted from each parameter map in each brain region.

[0073] In the third step, in the training stage, for each sample, based on the extracted image features, a brain hierarchical graph neural network model is constructed to extract its corresponding brain hierarchical features and the attention coefficients between brain regions. In this embodiment, the specific operations are as follows:

[0074] First, in the brain hierarchical graph neural network model as shown in Figure 2 , the brain of the subject is represented as a graph, where different brain regions are regarded as the vertices of the graph. The graph network model includes one layer of GCN convolutional layer, one layer of Topk pooling layer, one layer of GCN convolutional layer, and one layer of Readout layer, and the formula is as follows:

[0075] z1 = σ(A XW (0) ).

[0076] Among them, Z1 is the result of the first layer of graph convolution, σ is the linear activation function σ(·) = max(0, ·), X is the original feature matrix, W (0) is the parameter matrix of the first layer of graph convolution, is the normalized adjacency matrix, where A is the adjacency matrix of the graph, I is the identity matrix, Denote the degree of node i; After applying the graph pooling operation on the graph after the first convolution, a new graph is obtained, and its corresponding feature matrix and adjacency matrix are and A p , and from A p obtain the normalized adjacency matrix of the new graph in the same normalization manner Finally, after another graph convolution operation and Readout operation, the final output result representing the graph can be obtained:

[0077]

[0078] W (1) is the parameter matrix of the second-layer graph convolution, and Z2 is the output result of the second-layer graph convolution. Among them, Adam is selected as the optimizer, and it is trained for 100 epochs with a dropout rate of 0.1, an L2 regularization rate of 0.00005, and a learning rate of 0.001.

[0079] Then, for DTI data and fMRI data, construct brain hierarchical graph neural networks respectively. For DTI data, the features extracted from the four parameter maps of FA, MD, AxD, and RD from the feature extraction network are used as the vertex features of the graph, and the fiber tract count matrix obtained from fiber imaging is used as the edge weight of the graph. After learning through the brain hierarchical graph neural network, the DTI brain hierarchical image features are extracted; for fMRI data, the features extracted from the two parameter maps of ALFF and ReHo from the feature extraction network are used as the vertex features of the graph, and the calculated functional connectivity matrix between brain regions is used as the edge weight of the graph. After learning through the brain hierarchical graph neural network, the fMRI brain hierarchical features are extracted.

[0080] Finally, adopt the graph network model based on the graph attention mechanism as shown in Figure 3 to extract the attention coefficients between brain regions for T1-weighted data, DTI data, and fMRI data respectively, and use them as the edge weights of the graph to analyze their relationship with the correlation parameters between brain regions. The graph network model of the graph attention mechanism includes one embedding layer, one attention propagation layer, one Topk pooling layer, one embedding layer, and one Readout layer. The formula is as follows:

[0081] H1 = σ(XV (0) ).

[0082] Among them, H1 is the result of the first embedding layer, σ is the linear activation function, X is the original feature matrix, and V (0) is the weight matrix of the first embedding layer. After the first embedding layer is the attention propagation layer, and its output is:

[0083] H2 = PH1.

[0084] Among them, H2 is the result of the first embedding layer, and the propagation matrix It is composed of the attention coefficients between nodes. N is the number of nodes, and its calculation formula is as follows:

[0085]

[0086] Among them, P ij is the attention coefficient of vertex j to i, β is the shared weight of the propagation layer, h i and h j represent the feature vectors of vertices i and j. exp represents the exponential operation, represents all the neighborhood nodes of vertex i, then it means that vertex k can be taken as all the neighborhood nodes of vertex i and vertex i itself. cosh i , h j ) represents the cosine value of the angle between the feature vectors h i and h j The ||.|| represents the L2 norm. After applying the TopK pooling operation on H2, a new feature matrix is obtained Finally, after another graph convolution operation and Readout operation, the output result of the graph model using the graph attention mechanism can be obtained:

[0087]

[0088] Among them, V (1) is the weight matrix of the first-layer embedding layer, and Z is the final output result.

[0089] Step 4: For the extracted image features and non-image features in the training set, construct an extensible multi-level graph neural network model to classify and identify the sample population and complete the model training. In this embodiment, the specific operations are as follows:

[0090] First, based on the image features extracted by the feature extraction network and the brain-level graph neural network, input them into the population-level graph neural network to construct an extensible multi-level graph neural network model. During the training process, the entire population of the training set of this model is represented as a graph, each sample is regarded as a vertex of this graph, and the non-image information of the sample is used as the edge weight information for the connection between vertices. The population-level graph neural network model includes two weighted inductive W-GraphSAGE graph convolution layers with an optimized aggregation method and one Softmax layer, as Figure 4 shown.

[0091] Among them, the W-GraphSAGE graph convolution formula is as follows:

[0092]

[0093] Among them, represents the feature representation of node u at the (k-1)-th layer, W is the weight matrix of this layer, is the feature representation of node v at the k-th layer. σ is the linear activation function, a u is all the neighbors of node v of the normalized weight coefficients:

[0094]

[0095] Among them, represents the edge weight between the neighbor node u and node u, represents that vertex k can be taken as all the neighborhood nodes of vertex v and vertex v itself, exp represents the exponential operation, softmax v (e u ) represents the normalization operation. Adam is selected as the optimizer for the model, and it is trained for 100 rounds with a dropout rate of 0.1, an L2 regularization rate of 0.0,0005, and a learning rate of 0.001.

[0096] Then, the DTI brain hierarchical features, fMRI brain hierarchical features obtained from the brain hierarchical graph neural network model, and the T1 image features obtained from the feature extraction network are concatenated to obtain the combined features.

[0097] Finally, the combined features of the three modalities of all training samples are used as the vertex features of the graph, and the non-image information (age and gender) of each individual is used to establish the association between groups as the edge weights of the graph. The edge weights between vertices in the graph network are encoded according to the similarity of the samples, and the similarity between two samples m and n can be represented by A mn is expressed as:

[0098]

[0099] Among them, K represents the type of non-image information, S m and S n represent the feature vectors of samples m and n, and the feature similarity between samples m and n can be obtained using distance-related methods, and the formula is as follows:

[0100]

[0101] Among them, and are the averages of the vectors S n and S m respectively, ⊙ represents the dot product operation, and ||.||2 represents the L2 norm. represents the measure of some non-image information M, which is expressed as:

[0102]

[0103] Among them, θ represents the threshold, and M k (.) represents a certain non-image information. For the gender factor, θ is 0, that is, when the genders of two subjects are the same, the metric is 1, otherwise it is 0; for the age factor, θ is the standard deviation of the ages of all subjects in the figure, that is, when the age difference between two subjects is less than this standard deviation, the metric is 1, otherwise it is 0. Finally, the learning and training of the model are completed through the group-level graph neural network model.

[0104] Step 5, in the test stage, based on the network model trained on the training set, after preprocessing and feature extraction for each sample data in the test set respectively, it is input into the network model to complete the test. In this embodiment, the specific operations are as follows:

[0105] First, the gray matter images obtained by splitting the weighted data of the test set T1, the FA, MD, AxD, and RD parameter maps obtained from DTI data, and the ALFF and ReHo maps obtained from fMRI data are respectively input into the feature extraction network to extract image features;

[0106] Secondly, the image features extracted from the FA, MD, AxD, and RD parameter maps obtained from the test set DTI data from the feature extraction network and the fiber bundle count matrix obtained from fiber imaging are used as the input of the brain-level graph neural network model to complete the DTI brain-level feature extraction. At the same time, the image features extracted from the ALFF and ReHo parameter maps obtained from fMRI data from the feature extraction network and the calculated functional connectivity matrix between brain regions are used as the input of the brain-level graph neural network model to complete the fMRI brain-level feature extraction.

[0107] Then, the graph network model with graph attention mechanism is used to extract the attention coefficients between brain regions for T1-weighted data, DTI data, and fMRI data respectively.

[0108] Finally, the T1 features extracted by the feature extraction network, the DTI and fMRI brain-level features extracted by the brain-level graph neural network model are feature-cascaded, and the combined features obtained are used as the vertex features of the graph; the non-image information (age and gender) of each sample is used as the edge weight of the graph, and the group-level graph neural network model is used to complete the classification test.

[0109] In this embodiment, the system used is Ubuntu 18.04 system, the GPU configuration is NVIDIA GeForce 1080Ti 11G×8, and the development software environment is a server with python 3.7. In all the experiments of the present invention, two datasets, ADNI and the Third Hospital of Peking University, are selected for a five-fold cross-validation experiment to analyze the test effect of the method proposed in the present invention on multi-modal MRI data. At the same time, in order to analyze the robustness of the method proposed in the present invention, in the robustness verification experiment, the dataset of the Third Hospital of Peking University is selected as the training set, and the ADNI dataset is selected as the validation set for 10 Monte Carlo random experiments.

[0110] First of all, the verification of the effectiveness and robustness of the method in this aspect in this embodiment is shown in Table 1. Based on three commonly used MRI modal data, the method of the present invention shows excellent classification performance and strong robustness in the three binary classification test tasks of AD, MCI, and NC.

[0111] Table 1 Test effect and robustness verification of the scalable multi-level graph neural network model (%)

[0112]

[0113] Secondly, the functions of each algorithm module in the present invention are shown in the distillation experiment in Table 2. It can be seen from Table 2 that better classification performance can be achieved with multi-modal image data than with single-modal data. In the MCI vs NC classification task, removing the fMRI modal data has a greater impact on the model performance than removing the T1 modal or DTI modal data. The brain function of MCI patients in the early stage of AD is abnormal earlier than gray matter atrophy and white matter fiber damage, which is consistent with the clinical manifestations. After removing the W-GraphSAGE module (i.e., using the conventional GraphSAGE module), the model performance decreases slightly, indicating that W-GraphSAGE can effectively utilize the correlation relationships established by the non-image information of the subjects in the population-level graph neural network. After removing the brain-level graph neural network model, the model performance drops significantly (the accuracy and F1 in the three classification tasks decrease by 2.8 - 4.5% and 2.7 - 4.3% respectively), indicating that the brain-level graph neural network model can effectively extract brain-level image features.

[0114] Table 2 Results of the distillation experiment of the scalable multi-level graph neural network model (%)

[0115]

[0116]

[0117] Then, the method proposed in the present invention was compared and analyzed with other existing methods, as shown in Table 3. All comparison methods used the same experimental environment, training, and test data. Among them, for the SVM method, the gray matter volume features extracted from the T1-weighted images and the structural connection matrix features extracted from the DTI images were respectively subjected to Z-score normalization and mRMR-FCD feature selection, and finally SVM was used to complete the classification. For the 3D-CNN method, the whole-brain T1-weighted images were used as the input of the 3D-CNN model, and the subject age was encoded into a vector, which was combined with the output of the convolutional layer to obtain the final classification result. Based on the image features and non-image information extracted from DTI data and fMRI data, the MSE-GCN method completed the classification by establishing a multi-scale enhanced group hierarchical graph convolutional network (Multi-Scale Enhanced GCN, MSE-GCN). Based on rs-fMRI data, the ChebyNet method established a brain hierarchical graph neural network to achieve classification. As can be seen from Table 3, the method of the present invention can simultaneously take advantage of the brain hierarchical graph neural network and the group hierarchical graph neural network, comprehensively utilize the internal information of the brain and the non-image information of the samples to improve the accuracy, classification consistency, and robustness of disease detection. At the same time, only three commonly used MRI modality data were adopted in this embodiment, but the model has good scalability and can effectively fuse more clinical multi-modal imaging features and non-imaging features, providing an extensible and adjustable learning framework for the continuous improvement of clinical data analysis.

[0118] Table 3 Comparison results of the detection performance of the method of the present invention and other methods (%)

[0119]

[0120]

[0121] Subsequently, the classification performance and calculation time of the three graph convolutional operators of W-GraphSAGE, conventional GraphSAGE, and GCN after optimizing the weights were compared. For the AD vs NC task, the training time, test time, and the time required for predicting a single sample in a single-fold experiment of the three graph convolutional operators were calculated. As can be seen from Table 4 and Table 5, the W-GraphSAGE convolutional operator is higher than GraphSAGE in terms of classification accuracy and consistency, and is almost the same as GCN, but the training time-consuming is not high. Based on the already trained model, it can quickly predict the disease for new samples, effectively saving the prediction time in actual clinical applications.

[0122] Table 4 Comparison of the classification performance of W-GraphSAGE and other methods (%)

[0123]

[0124] Table 5 Comparison of Computation Time between W-GraphSAGE and Other Methods

[0125] Method Training time consumption Testing time consumption Prediction time consumption for a single new sample GCN 23.004s 0.019s 23.007s GraphSAGE 9.710s 1.627s 0.004s W-GraphSAGE 13.048s 3.624s 0.036s

[0126] Finally, for the three-modal data, the structural similarity (SSIM) of the attention matrix and the correlation matrix between the target brain region and other brain regions, the cosine similarity and the distance correlation coefficient of the attention coefficients and correlation parameters between the target brain region and other brain regions are used to measure their similarity. The results are shown in Table 6. It can be seen that the structural similarity (SSIM) index of the brain region attention matrix and the brain region correlation matrix is relatively large; from the perspective of brain regions, the cosine similarity and the distance correlation coefficient between the attention coefficients and the correlations between a single brain region and other brain regions are also relatively large. For the three-modal data, the 37th left hippocampus region (Hippocampus_L) of T1 data, the 38th right hippocampus region (Hippocampus_R) of fMRI data, and the 37th left cingulum hippocampus region (Cingulum(hippocampus)left, CGH-L) of DTI data are respectively selected as the target brain regions, and the top 30 brain regions with the largest to smallest correlation coefficients or attention coefficients between other brain regions and the target brain regions are analyzed. From Figure 5 It can be seen that there is a high overlap among the brain regions with a large relationship between the attention coefficients and the brain region correlations, indicating that there is a large similarity between the brain region attention matrix and the brain region correlation matrix. In practical applications, the preprocessing time of the correlation parameters between some brain regions is relatively long. For example, for the DTI data of a sample, the calculation of the fiber bundle count between brain regions usually takes 2 to 3 hours. In contrast, the calculation of the attention coefficients between brain regions only takes less than 10 seconds. Therefore, in the absence of correlation parameters, the attention coefficients can provide auxiliary reference for the real-time identification of the clinical AD course and the extraction of imaging biomarkers.

[0127] Table 6 Calculation of the Overall and Brain Region Similarities of Attention Coefficients and Connection Relationships

[0128]

[0129] Generally speaking, in the AD recognition application based on multi-modal MRI image data in this embodiment, the method of the present invention makes full use of the image features extracted from multi-modal MRI data and the connection features between brain regions, comprehensively analyzes the roles of gray matter atrophy, damaged nerve fibers in white matter, and abnormal brain functional connections in different disease courses, and at the same time combines the non-imaging data of the subjects to accurately identify different disease courses. This method has good scalability, allowing further addition of image features and non-image information from multi-modal data, providing an extensible and adjustable learning framework for the continuous improvement of clinical data analysis.

[0130] The above description has fully introduced the specific implementation manners of the present invention. It should be noted that any modifications made by those skilled in the art to the specific implementation manners of the present invention do not depart from the scope of the claims of the present invention. Correspondingly, the scope of the claims of the present invention is not limited solely to the foregoing specific implementation manners.

Claims

1. A scalable multi - level graph neural network model based on multi - modal image data, characterized in that, It includes the following steps: (1) For the multi-modal image data to be processed, use general medical image preprocessing software to perform preprocessing operations to obtain various parameter maps of the multi-modal image data and the correlation coefficients between brain regions; (2) In the training stage, for the preprocessed multi-modal image data, construct a feature extraction network based on a multi-scale residual network to extract features from each modality data; (3) In the training stage, for each sample, based on the image features already extracted by the feature extraction network, construct a brain hierarchical graph neural network model to extract its corresponding brain hierarchical features and the attention coefficients between brain regions; (4) In the training stage, for the sample population, based on the features extracted by the feature extraction network and the brain hierarchical graph neural network, construct an extensible multi-level graph neural network model to realize the learning and classification of sample data and complete the model training; (5) In the testing stage, after preprocessing and feature extraction of each sample data in the test set respectively, input it into the trained network model to complete the test.

2. The scalable multi-level graph neural network model based on multi-modal image data according to claim 1, wherein The construction of the feature extraction network based on the multi-scale residual network in step (2) is as follows: (1) Along the three directions of the coronal, sagittal, and transverse planes of the three-dimensional image, accumulate all the slices in each direction and convert them into two-dimensional images of the human brain in three directions, which are used as the three-channel input of the feature extraction network; (2) The feature extraction network uses the ResNet-18 network as the backbone network, including three input channels, 17 convolutional layers, and 1 fully connected layer; (3) Replace the 3×3 convolutional kernel in the original ResNet-18 network with 1×1, 3×3, 5×5 convolutional kernels and a 3×3 pooling layer to form a multi-scale residual module; (4) Embed an attention mechanism module in the network. The attention mechanism mask is generated by performing 3×3 convolution on the features of all input attention modules, and use this mask to re-weight the features to improve the network's attention to key features and complete the construction of the "feature extraction network".

3. The scalable multi-level graph neural network model based on multi-modal image data according to claim 1, wherein The construction of the brain hierarchical graph neural network model in step (3) is as follows: (1) In the brain hierarchical graph neural network model, extract brain hierarchical image features based on the correlation relationship between brain regions; among them, the brain of each sample is represented as a graph, different brain regions are used as the vertices of the graph, and the single-modal image features already extracted are used as vertex features. The edge weights connecting the vertices include two construction methods: the correlation coefficient between brain regions and the graph attention coefficient; (2) The graph network model includes a GCN convolutional layer, a Topk pooling layer, a GCN convolutional layer, and a Readout layer in the connection order, and the formula is as follows: ; Among them, is the result of the first-layer graph convolution, is the linear activation function , is the original feature matrix, is the parameter matrix of the first-layer graph convolution, is the normalized adjacency matrix, where , is the adjacency matrix of the graph, is the identity matrix, is the negative half-power operation on the diagonal matrix , that is, replacing each diagonal element in with the reciprocal of its square root, the i-th diagonal element in represents the degree of node 𝑖; after applying the graph pooling operation on the graph after the first convolution, a new graph is obtained, and its corresponding feature matrix and adjacency matrix are and , and obtains the normalized adjacency matrix of the new graph in the same normalization manner, is the negative half-power operation on the diagonal matrix after the graph pooling operation, and finally, after another graph convolution operation and 𝑅𝑒𝑎d𝑜𝑢𝑡 operation, the final output result representing the graph can be obtained: ; Among them, is the parameter matrix of the second-layer graph convolution, is the output result of the second-layer graph convolution; (3) According to the characteristics of different modality data, the correlation coefficient between brain regions can be obtained by extracting or calculating the correlation coefficient of the feature vectors already extracted between brain regions during the preprocessing process; The graph attention coefficient can be obtained by learning through a graph network model of the graph attention mechanism. This model includes, in the connection order, an embedding layer, an attention propagation layer, a Topk pooling layer, an embedding layer, and a Readout layer. The formula is as follows: ; Among them, is the result of the first-layer embedding layer, 𝜎 is the linear activation function, and 𝑋 is the original feature matrix. is the weight matrix of the first-layer embedding layer; after the first-layer embedding layer is the attention propagation layer, and its output is: ; Among them, is the result of the first-layer embedding layer, and the propagation matrix is composed of the attention coefficients between nodes, is the number of nodes, and its calculation formula is as follows: , Among them, is the attention coefficient of vertex 𝑗 to 𝑖, 𝛽 is the shared weight of the propagation layer, and represent the feature vectors of vertices 𝑖 and 𝑗, represents the exponential operation, represents all neighborhood nodes of vertex 𝑖, then represents vertex can be taken as all neighborhood nodes of vertex 𝑖 and vertex 𝑖 itself, represents the feature vector and the cosine value of the angle between them , represents the 𝐿2 norm; after applying the TopK pooling operation on , a new feature matrix is obtained; finally, after another graph convolution operation and a Readout operation, the output result of the graph model using the graph attention mechanism can be obtained: ; Among them, is the weight matrix of the first embedding layer, is the final output result.

4. The scalable multi-level graph neural network model based on multi-modal image data according to claim 1, wherein The construction of the scalable multi-level graph neural network model in step (4) is as follows: 1) Based on the feature extraction network and the brain hierarchical graph neural network, the image features have been extracted and input into the population hierarchical graph neural network to construct an extensible multi-level graph neural network model; during the training process, the training set sample population of this model is represented as a graph, each sample is regarded as a vertex of this graph, the non-image information of the sample, such as age, gender, clinical score, and education level, is used as the edge weight information for the connection between vertices, and the multi-modal image combined features of each sample are used as vertex features; among them, the multi-modal image combined features can be the brain hierarchical image features extracted by different modal data through the brain region graph neural network respectively, or the features obtained by cascading the image features extracted by part of the modal data through the feature extraction network and the brain hierarchical image features extracted by part of the modal data through the brain region graph model; the edge weights between vertices in the graph network are encoded according to the similarity of the samples, and the similarity between two samples 𝑚 and 𝑛 can be represented by denoted as: , Among them, K represents the type of non-image information, and represent the feature vectors of samples 𝑚 and 𝑛, and the feature similarity between samples 𝑚 and 𝑛 can be obtained using a distance-related method, and the formula is as follows: , Among them, and are the average values of vectors and respectively. ⊙ represents the dot product operation, represents the 𝐿2 norm; represents a measure of some non-image information 𝑀: , where 𝜃 represents a threshold value, represents a certain non-image information; (2) The population-level graph neural network model includes two weighted inductive W-GraphSAGE graph convolutional layers with optimized aggregation methods and a Softmax layer. Among them, the W-GraphSAGE graph convolution formula is as follows: ; Among them, represents a node in the layer's feature representation, is the weight matrix of this layer, is the node in the layer's feature representation; 𝜎 is a linear activation function, are the normalized weight coefficients of all neighbors of node 𝑣, where, : ; Among them, represents the edge weight between neighbor node 𝑢 and node 𝑣, , represents the vertex can be taken as all neighborhood nodes of vertex v and vertex v itself, represents the exponential operation, represents the normalization operation.

5. The scalable multi-level graph neural network model based on multi-modal image data according to claim 1, wherein After preprocessing and feature extraction are respectively performed on each sample data in the test set in step (5), it is input into the trained network model to complete the test. The specific steps are as follows: (1) Preprocess the different modality data of the test set data separately; (2) Input the preprocessed multi-modal data into the feature extraction network respectively to extract the image features corresponding to different brain regions of different modality data; (3) For some modality data, use the brain-level graph neural network to extract the brain-level image features and the attention coefficients between brain regions; (4) Input the extracted image features and non-image features into the multi-level graph neural network model to complete the test.

Citation Information

Patent Citations

  • Brain network modeling and individual prediction method based on multi-modal magnetic resonance image

    CN113616184A

  • Systems and methods for quantitatively characterizing alzheimer's disease risk events based on multimodal biomarker data

    US20210212629A1