AD classification prediction method and system based on multi-mode multi-layer graph neural network model

Through the multimodal multi-layer graph neural network model, combined with GAT and GCN, the edge weights are dynamically adjusted, and the graph construction and training process is optimized, the existing AD diagnostic model has been solved, and more efficient AD classification performance is achieved.

CN120387100APending Publication Date: 2025-07-29HUAZHONG AGRI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510443447.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing early AD diagnostic model based on graph neural networks has high computational complexity when building brain area networks, and fails to fully explore brain interval relationships, resulting in limited model performance improvement.

Method used

A multimodal multi-layer graph neural network model is used to extract ROI features through Gaussian kernel functions, and a two-layer GNN architecture is built with GAT and GCN, which integrates image and non-image information, dynamically adjusts edge weights, and optimizes graph construction and training processes.

Benefits of technology

It improves AD classification performance, simplifies brain area feature extraction, fully explores the relationship information between brain ROI and subjects, and improves the overall performance and diagnostic accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387100A_ABST
    Figure CN120387100A_ABST
Patent Text Reader

Abstract

According to the AD classification prediction method and system based on the multi-mode multi-layer graph neural network model, the AD classification performance is improved by combining imaging and non-image information, and the AD classification prediction method and system have wide application potential in the field of cranial nerve diseases. According to the method, a multilayer graph neural network of a double-layer GNN architecture based on GAT and GCN is constructed based on the relationship between brain ROIs and the relationship between subjects, hierarchical training is performed on the brain ROI relationship and the subject relationship, and the performance of the model in AD classification is improved. According to the method, the Gaussian kernel function is adopted to construct the brain network between the brain ROIs as the ROI features, the process of extracting the brain region features is simplified, and the training requirement of the GNN is better met. According to the method, the multi-source non-image information matrix is constructed, the graph construction and training process of the second layer GNN is optimized, and the influence of the non-image information on the edge weight is adjusted through back propagation, so that the relationship among subjects is considered more comprehensively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning and neuroimaging processing, and particularly relates to an AD classification prediction method and system based on a multi-modal multi-layer graph neural network model. Background Art

[0002] Early identification of the disease signs of AD (Alzheimer's Disease) is of inestimable value for implementing effective management and intervention strategies. At present, sMRI (structural Magnetic Resonance Imaging) and PET (Positron Emission Tomography) are usually combined for application to enhance the prediction accuracy of the AD diagnosis model. With the development of deep learning technology, the combination of neuroimaging technology and deep learning technology has significantly improved the diagnostic performance of AD and other degenerative diseases, and has gradually become the core method for AD classification problems.

[0003] In the early deep learning research for AD classification, CNNs (Convolutional Neural Networks) have achieved good classification performance in AD diagnosis combining images by virtue of their excellent image feature extraction ability and the advantages of convolutional layers and pooling layers in automatically learning local patterns. However, CNNs have limitations in processing non-image information and modeling global relationships, and it is difficult to fully exploit the potential value of multi-modal data, thus restricting the further improvement of performance.

[0004] To overcome this limitation, GNNs (Graph Neural Networks) are used to extend CNNs to non-Euclidean spaces through graph topology and feature propagation between nodes, effectively making up for the defects of CNNs in processing relational information and improving the AD diagnosis accuracy by fusing image and non-image information. However, there are two main problems in the existing AD early diagnosis models based on graph neural networks (GNN): one is that the constructed brain region (Regions of Interest, ROIs) network is relatively complex, increasing the computational complexity of the model; the other is that the mutual relationships between brain regions are ignored, resulting in the model failing to fully exploit the potential effective information. These problems limit the efficiency and information mining ability of the model. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide an AD classification prediction method and system based on a multimodal multi-layer graph neural network model (Multimodal Multi-layer Graph Neural Network, MMGNN) to improve the AD classification performance.

[0006] The technical solution adopted by the present invention to solve the above technical problem is: an AD classification prediction method based on a multimodal multi-layer graph neural network model, including the following steps: S0: Obtain the image information and non-image information of the subject; S1: Use a Gaussian kernel function to extract the feature vectors of the ROIs from the image information; S2: Construct an ROI graph according to the feature vectors of the ROIs; introduce an attention mechanism through GAT to obtain the hidden layer features of the subject from the ROI graph; S3: Extract the relationships between subjects by calculating the differences in the non-image information of the subjects; S4: Construct a subject graph according to the hidden layer features of the subjects and the relationships between the subjects; output the subject category prediction value through GCN, and apply a threshold to judge the category of the subject.

[0007] According to the above solution, in step S0, the specific steps are: The image information of the subject includes sMRI and PET; the non-image information of the subject includes category, gender, age, and MMSE.

[0008] Further, in step S1, the specific steps are: S11: Perform standardized preprocessing on the data of sMRI and PET; S12: Extract the corresponding original features of the ROIs defined based on the neuromorphic brain atlas; S13: Use a Gaussian kernel function to map the original features of each ROI, extract the structural feature vectors from sMRI, and extract the functional feature vectors from PET.

[0009] Further, in step S14, the specific steps are: After obtaining the original features of all ROIs by comparing the brain atlas to extract images, calculate the feature coefficients of a certain ROI with other ROIs through a Gaussian kernel, and combine these feature coefficients to obtain the feature vector of the ROI.

[0010] According to the above solution, in step S2, the specific steps are: S21: Represent the ROI with nodes, the node features represent the feature vectors of the ROI, the edges represent the relationships between ROIs, the adjacency matrix of the edges represents the relationship strength between nodes, and construct an ROI graph consisting of nodes, node features, edges, and the adjacency matrix of the edges; S22: Calculate the Spearman correlation coefficient between two ROIs according to the number of ROIs and the Euclidean distance between the two ROIs; S23: Under the constraint of the Spearman correlation coefficient, obtain the attention score between two ROIs by concatenating, transforming, weighting, and non-linearizing the feature vectors of the two ROIs; S24: Normalize the attention scores between nodes to obtain attention coefficients and form an attention matrix; Dynamically allocate node feature weights according to the magnitudes of the normalized attention scores; S25: Obtain the hidden layer features of the subject through GAT according to the attention scores.

[0011] According to the above solution, in step S4, the specific steps are: S41: Represent the subject with nodes, the node features represent the hidden layer features obtained from GAT, the edges represent the relationships between subjects, the adjacency matrix of the edges represents the relationship strength between nodes, and construct a subject graph consisting of nodes, node features, edges, and the adjacency matrix of the edges; S42: GCN processes the adjacency matrix of the edges in the subject graph, the matrix composed of the hidden layer features of the subject, and the weight matrix with function and function and outputs the predicted value of the subject category.

[0012] Further, in step S41, the cosine similarity is used to calculate the similarity of the hidden layer features, and the adjacency matrix of the edges is calculated by combining the hidden layer features and non-image information.

[0013] According to the above solution, it further includes the following steps: S5: Adopt a cross-entropy loss function with class weights to solve the class imbalance problem by imposing a higher penalty on the misclassification of the subject categories.

[0014] AD classification prediction system based on a multi-modal multi-layer graph neural network model, Data acquisition sub-module, used to acquire the image information and non-image information of the subject; Image information extraction sub-module, used to extract the feature vectors of the ROI from the image information by using a Gaussian kernel function; ROI graph sub-module, used to construct an ROI graph according to the feature vectors of the ROI; Introduce an attention mechanism through GAT to obtain the hidden layer features of the subject from the ROI graph; A non-image information extraction sub-module, used to extract the relationship between subjects by calculating the differences in non-image information of the subjects; A subject graph sub-module, used to construct a subject graph based on the hidden layer features of the subjects and the relationship between the subjects; output the subject category prediction value through GCN, and apply a threshold to judge the category of the subjects.

[0015] A computer memory, which stores a computer program executable by a computer processor, and the computer program executes an AD classification prediction method based on a multi-modal multi-layer graph neural network model.

[0016] The beneficial effects of the present invention are as follows: 1. The AD classification prediction method and system based on the multi-modal multi-layer graph neural network model of the present invention realize the improvement of AD classification performance by combining imaging and non-image information, and have broad application potential in the field of brain nerve diseases.

[0017] 2. The present invention constructs a multi-layer graph neural network with a double-layer GNN architecture based on GAT and GCN respectively based on the relationship between brain ROIs and the relationship between subjects, and performs hierarchical training on the relationship between brain ROIs and the relationship between subjects, improving the performance of the model in AD classification.

[0018] 3. The present invention uses a Gaussian kernel function to construct a brain network between brain ROIs as ROI features, simplifies the process of brain region feature extraction, and better adapts to the training requirements of GNN.

[0019] 4. The present invention constructs a multi-source non-image information matrix, optimizes the graph construction and training process of the second-layer GNN, and adjusts the influence of non-image information on edge weights through backpropagation, so as to more comprehensively consider the relationship between subjects.

[0020] Of course, it is not necessary for any product implementing the present invention to achieve all the above-mentioned advantages simultaneously. Description of the Drawings

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0022] Figure 1 is the flowchart of the embodiment of the present invention.

[0023] Figure 2 is the architecture diagram of the MMGNN model of the embodiment of the present invention.

[0024] Figure 3 It is a process diagram for extracting brain ROI features in an embodiment of the present invention.

[0025] Figure 4 It is a multi-modal early fusion strategy diagram in an embodiment of the present invention.

[0026] Figure 5 It is a multi-modal mid-term fusion strategy diagram in an embodiment of the present invention.

[0027] Figure 6 It is a multi-modal late fusion strategy diagram in an embodiment of the present invention.

[0028] Figure 7 It is a comparison diagram of ACC values at different numbers of GAT attention heads in an embodiment of the present invention.

[0029] Figure 8 It is a comparison diagram of ACC values at different hidden layer dimensions in an embodiment of the present invention.

[0030] Figure 9 It is a diagram showing the influence of non-image information on the ACC values of four tasks in an embodiment of the present invention. Detailed implementation manners

[0031] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0032] Embodiment 1 In brain science research, two relationships are usually studied: one is the relationship between subjects, mainly studying the correlation between the differences in non-image information such as age and gender of different subjects and disease symptoms; the other is the relationship between brain ROIs, mainly studying the correlation between image information such as brain region anatomical structure and functional connectivity and disease symptoms. Different from the existing GNN models that perform multi-layer training based on a single relationship, in this embodiment, a two-layer GNN architecture is designed for the two different relationships between brain ROIs and subjects, so as to more reasonably utilize the two natural relationships in brain science: that is, introducing GAT (Graph Attention Network), using the self-attention mechanism to accurately capture these complex interaction patterns; at the same time, to model the relationship between subjects, GCN (Graph Convolutional Network) is adopted, and through the convolutional operation of the normalized adjacency matrix, the weighted aggregation of neighborhood features is realized, so as to efficiently integrate node features; giving full play to the complementary advantages of the two models: GAT can handle complex heterogeneous relationships, while GCN is good at efficient information aggregation, so as to jointly improve the overall performance of the model.

[0033] See Figure 1 , the specific steps of the AD classification prediction method based on the multi-modal multi-layer graph neural network model are as follows: S0: Obtain the image information and non-image information of the subjects; S1: Use the Gaussian kernel function to extract the feature vectors of the ROIs from the image information; S2: Construct an ROI graph based on the feature vectors of the ROIs; Introduce an attention mechanism through GAT to obtain the hidden layer features of the subjects from the ROI graph; S3: Extract the relationships between the subjects by calculating the differences in the non-image information of the subjects; S4: Construct a subject graph based on the hidden layer features of the subjects and the relationships between the subjects; Output the subject category prediction values through GCN, and apply a threshold to judge the category of the subjects.

[0034] Furthermore, in step S0, the specific steps are: The image information of the subjects includes sMRI and PET; The non-image information of the subjects includes category, gender, age, and MMSE.

[0035] In step S1, the specific steps are: S11: Perform standardized preprocessing on the data of sMRI and PET; S12: Extract the corresponding original features based on the ROIs defined by the neuromorphic brain atlas; S13: Use the Gaussian kernel function to map the original features of each ROI, extract the structural feature vectors from sMRI, and extract the functional feature vectors from PET.

[0036] Furthermore, in step S14, the specific steps are: After extracting the original features of all ROIs by comparing the brain atlas extraction images, calculate the feature coefficients of a certain ROI and other ROIs through the Gaussian kernel, and combine these feature coefficients to obtain the feature vector of the ROI.

[0037] In step S2, the specific steps are: S21: Use nodes to represent the ROIs, node features to represent the feature vectors of the ROIs, edges to represent the relationships between the ROIs, and the adjacency matrix of the edges to represent the relationship strength between the nodes, and construct an ROI graph composed of nodes, node features, edges, and the adjacency matrix of the edges; S22: Calculate the Spearman relationship coefficient between two ROIs according to the number of ROIs and the Euclidean distance between the two ROIs; S23: Under the constraint of the Spearman relationship coefficient, obtain the attention scores between two ROIs by concatenating, transforming, weighting, and non-linearizing the feature vectors of the two ROIs; S24: Normalize the attention scores between nodes to obtain attention coefficients and form an attention matrix; dynamically allocate node feature weights according to the magnitudes of the normalized attention scores. S25: Obtain the hidden layer features of the subjects through GAT according to the attention scores.

[0038] In step S4, the specific steps are as follows: S41: Use nodes to represent the subjects, node features to represent the hidden layer features obtained from GAT, edges to represent the relationships between the subjects, and the adjacency matrix of the edges to represent the relationship strength between nodes, and construct a subject graph composed of nodes, node features, edges, and the adjacency matrix of the edges. S42: GCN processes the adjacency matrix of the edges in the subject graph, the matrix composed of the hidden layer features of the subjects, and the weight matrix with function and function and outputs the predicted values of the subject categories.

[0039] Furthermore, in step S41, calculate the similarity of the hidden layer features using cosine similarity, and combine the hidden layer features and non-image information to calculate the adjacency matrix of the edges.

[0040] It further includes the following steps: S5: Adopt a cross-entropy loss function with class weights to solve the class imbalance problem by imposing higher penalties on the misclassification of subject categories.

[0041] Embodiment 2 The steps of this embodiment are the same as those of Embodiment 1, except that each step is applied to a specific instance. This embodiment specifically includes the following steps: S0: Obtain the image information and non-image information of the subjects. The dataset used in this embodiment is from the publicly accessible Alzheimer's Disease Neuroimaging Initiative (ADNI) database, which contains data from 890 subjects in total. Among them, 195 are AD patients, 379 are normal controls (NC), 161 are early mild cognitive impairment (EMCI), and 155 are late mild cognitive impairment (LMCI). In this embodiment, two types of image information, namely sMRI and PET, as well as four types of non-image information, namely the subject category, gender, age, and Mini-Mental State Examination (MMSE), of the subjects in the above four categories are collected. Among them, the subject category refers to which one of the four categories of AD, NC, EMCI, and LMCI the subject belongs to. MMSE is a standardized test used to evaluate cognitive function, widely used for screening dementia and cognitive impairment, as well as tracking the progression of the disease. The scoring range of MMSE is from 0 to 30 points, and the lower the score, the higher the degree of impaired cognitive function. The specific collection situation of the non-image information is shown in Table 1.

[0042] Table 1 Alzheimer's Disease Data Collection

[0043] S1: Extract the feature vectors of ROIs from the image information using the Gaussian kernel function; S11: Before extracting features from the sMRI and PET data, use the FSL toolbox 6.0 to perform standardized preprocessing on the sMRI and PET data, including skull stripping, affine registration, noise filtering and artifact correction, rigid alignment, and non-linear alignment.

[0044] S12: Extract the corresponding original features based on N0 brain ROIs defined by the neuromorphological brain atlas; the brain atlas used here comes from the version of 136 brain ROIs defined by Neuromorphometrics, Inc. These original features include the white matter volume, standard deviation, thickness, etc. of each ROI.

[0045] S13: Use the Gaussian kernel function to map the features of each ROI to generate the structural and functional features of the brain. As a similarity measurement method, the Gaussian kernel is widely used in the field of link prediction. Its non-linear mapping ability and local approximation characteristics not only reduce the complexity of modeling the relationships of brain ROIs, but also enable the data to better adapt to the training requirements of GNN. The specific process is as Figure 3 shown.

[0046] Figure 3In this embodiment, the sMRI and PET data are first preprocessed, and then the original features are extracted through the brain atlas, and finally the Gaussian kernel function is used to extract the structural and functional features of the brain. Among them, the features extracted from the sMRI images are structural features, and the features extracted from the PET images are functional features. Specifically, after the original features of N0 ROIs are obtained by comparing the brain atlas extraction images, the feature coefficients of the i-th and j-th ROIs of the subject are calculated through the Gaussian kernel , as shown in formula (1).

[0047] (1) where and respectively represent the original features of the i-th and j-th ROIs of the subject, and the ranges of i and j are from 1 to N0. is the parameter of the Gaussian kernel, which is used to control the width of the kernel function.

[0048] In formula (1), with i unchanged, j ranges from 1 to , after calculating all , the features of the i-th ROI can be obtained.

[0049] S2: Construct an ROI map according to the ROI feature vector; introduce an attention mechanism through GAT to obtain the hidden layer features of the subject from the ROI map; the specific steps are as follows: S21: In the GNN model, it is usually necessary to construct a graph, which consists of nodes, edges, node features, and the adjacency matrix of the edges that measures the strength of the relationship between nodes. GAT is a variant of GNN. By introducing an attention mechanism, it weights the node features in the graph, so as to more flexibly capture the important relationships between nodes. This mechanism can not only handle heterogeneous relationships, but also enhance the model's ability to express complex network structures, especially suitable for modeling graph data with non-uniform connection features. Since there are heterogeneous relationships between the brain ROIs, GAT can more effectively capture these complex interactions between nodes through its dynamic weighting mechanism, thus improving the modeling ability of the complex relationships between the brain ROIs and the representation ability of the brain ROI features. In the GAT layer network architecture, the construction of the brain ROI graph is crucial, because the quality of the graph directly determines the training effect and the final performance of the GAT model. Therefore, in this embodiment, the relationships between ROIs are constructed. In the ROI graph, the nodes represent the various ROIs of the brain, the node features correspond to the features of each ROI, and the edges represent the relationships between these ROIs.

[0050] S22: The Spearman correlation coefficient is used to describe the non - linear relationship between ROIs. The Spearman coefficient is a non - parametric statistical method, widely used to evaluate the correlation between variables, especially suitable for revealing complex relationships between variables. The Spearman relationship coefficient between the i - th brain ROI and the j - th brain ROI of a subject is shown in formula (2). The calculation method is as follows: (2) Where, represents the Euclidean distance between the features of the i - th ROI and the j - th ROI of a subject. is the number of brain ROIs.

[0051] S23: The Spearman relationship coefficient is mainly used to initialize the adjacency matrix of edges , as shown in formula (3). These initialized weights will be further adjusted through the attention mechanism during the training process to more accurately reflect the relationship strength between ROIs.

[0052] (3) Where, represents the attention score between the i - th ROI and the j - th ROI, representing the relationship strength between them; represents the concatenation operation; is a matrix used to transform the dimensionality of the ROI feature vector; is the feature vector of the i - th ROI of a subject calculated through the Gaussian kernel; is a trainable weight vector; is a non - linear activation function.

[0053] S24: When calculating the attention scores of all neighbors of the i - th ROI to this ROI, the values may be significantly unbalanced. If these values are directly used for the i - th ROI to aggregate neighbor information, it will cause some nodes to have too much influence while other nodes have too little influence. To avoid this problem, the attention matrix is further calculated. By normalizing the attention scores between nodes, the node feature weights are dynamically allocated according to the size of the normalized scores, so as to more effectively aggregate neighbor information. Formula (4) describes the calculation method of the attention coefficient , representing the normalized attention score of the j - th ROI to the i - th ROI, representing its relative relationship strength to the i - th ROI among all neighbors , as follows: (4) Note that formula (4) only calculates the nodes of , which is the neighborhood of the i-th ROI in the figure and represents the set of neighbors of this ROI.

[0054] S25: Apply the algorithm of the GAT layer to obtain the hidden layer features to enhance the representation ability of the brain ROI features and use them as the input of the GCN layer. The algorithm of the GAT layer is shown in formula (5): (5) where represents the number of attention heads in the GAT, is the trainable weight matrix, represents the hidden layer features of the subject, represents the number of brain ROIs. represents the concatenation operation, represents the non-linear activation function.

[0055] S3: Extract the relationships between subjects from the multi-source non-image information (category, gender, age, and MMSE score) of the subjects by calculating the differences, as shown in formulas (6) to (9): (6) (7) (8) (9) where and represent the categories of the i-th and j-th subjects respectively, and represent the genders of the i-th and j-th subjects respectively, and represent the ages of the i-th and j-th subjects respectively, and represent the MMSE scores of the i-th and j-th subjects respectively.

[0056] For two subjects, the values of the parameters (category), (gender), (age), and (MMSE) are set to 1 and 0 respectively. When they are equal, take 1; when they are not equal, take 0. The value ranges of i and j are both from 1 to , represents the number of subject data processed by GCN in one training.

[0057] ​S4: Construct a subject graph based on the relationships between subjects; output the predicted subject category values through GCN, and apply a threshold to determine the category of the subject. The specific steps are as follows: S41: The feature of GCN lies in that its convolution operation is based on a normalized adjacency matrix. For structured graph data, this method can efficiently perform weighted aggregation of neighborhood features, which not only improves the efficiency of information aggregation but also has low computational complexity and strong generalization ability. In this embodiment, the relationships between subjects exhibit relatively structured graph data characteristics, enabling GCN to fully capture the mutual associations between subjects, thereby improving the classification performance of the model. Before using GCN, a graph needs to be constructed, and the quality of graph construction is directly related to the model training effect. In the subject graph, the nodes in the graph are subjects, and the node features are the hidden layer features obtained from GAT; the edges represent the relationships between subjects, and the adjacency matrix of the edges is calculated by combining the hidden layer features and non-image information to measure the relationship strength between nodes. In this embodiment, the adjacency matrix of the edges is denoted as A, and its calculation formula is shown in Equation (10): (10) Where represents the relationship strength between subject i and subject j; and are the hidden layer features of subject i and subject j, calculated from Equation (5); is the similarity matrix. Here, cosine similarity is used as the calculation method to determine the similarity of the hidden layer features because cosine similarity has good robustness in dealing with complex relationships. 、 、 and respectively represent the importance coefficients of non-image information (such as status, gender, age, and MMSE score) in the matrix and reflect the contribution degree of this information to the matrix . During the training process, these parameters will adjust the weights through backpropagation to achieve better results.

[0058] After constructing the GCN graph, in this embodiment, the algorithm of the GCN layer is used to obtain the preliminary result, and a threshold is applied to determine the category of the subject. This algorithm consists of a GCN layer and an output layer activated by the and functions, as shown in Equation (11): (11) Where represents the predicted output, the matrix is the adjacency matrix of the edges in the GCN graph, is a matrix composed of the hidden layer features of N subjects. is a trainable weight matrix used for linearly transforming node features and is used here to adjust the dimension of the matrix. S5: Set the loss function After obtaining the prediction results of the GCN layer, a cross-entropy loss function with class weights is adopted to solve the class imbalance problem by imposing a higher penalty on the misclassification of the minority class. Equation (12) is the calculation formula for the weighted cross-entropy loss function as follows: (12) In Equation (12), represents the predicted output, represents the true label of the i-th object, represents one training process of the data of subjects, which is used to adjust the contributions of positive and negative samples in the loss function to solve the class imbalance problem. In the classification task with classes NC (normal) and AD (Alzheimer's disease), it is assumed that the number of samples in NC is half of that in AD. To balance the calculation of the loss function, a higher weight will be assigned to the NC class, thereby reducing the impact brought by class imbalance.

[0059] Experiments and Results This embodiment more efficiently integrates image and non-image data, simplifies the feature extraction process of the regions of interest (ROIs) of the brain, and further improves the accuracy of AD diagnosis. To verify the effectiveness of the model, a comprehensive and systematic evaluation work was carried out relying on the authoritative database of the Alzheimer's Disease Neuroimaging Initiative (ADNI). The results show that compared with traditional methods, the MMGNN model exhibits higher prediction performance in multiple AD classification tasks, providing more reliable technical support for the early diagnosis of AD.

[0060] The model implementation in this example is based on the Python language and the PyTorch framework. The experimental code was run on a computer equipped with an NVIDIA RTX 3090 GPU and 24GB of memory. To ensure the stability of the model results, this example uses a five-fold cross-validation method. The core concept of this method is to divide the original dataset into five equal subsets, four of which are used for model training and one for testing. This entire process is repeated five times, each time selecting a different subset as the test set and the remaining subset for training, thereby comprehensively evaluating the model's performance. Finally, the five performance evaluation results are averaged to obtain the final evaluation result. The model parameters are set as follows: To prevent overfitting, a dropout rate of 0.5 is used, combined with a weight decay of 0.001 for regularization to achieve a balance between the model's convergence speed and generalization ability. The initial learning rate is set to 0.1, and the training process is carried out for 600 epochs. The learning rate is adjusted by a decay factor of 0.1 every 150 epochs to further improve the fitting effect in the later stages of training. In order to evaluate the accuracy of the model, this embodiment is tested on four binary classification tasks: AD vs. NC, NC vs. EMCI, EMCI vs. LMCI, and LMCI vs. AD, and accuracy (ACC), sensitivity (SEN), and specificity (SPE) are selected as key performance indicators.

[0061] Performance comparison of MMGNN based on different fusion methods In order to achieve the best multimodal fusion effect for the proposed MMGNN model, this example compares the results of three multimodal fusion methods: early fusion (Fusion a), mid-term fusion (Fusion b), and late fusion (Fusion c) in the MMGNN model. These three different multimodal fusion strategies are as follows: Figure 4 、 5 , as shown in 6.

[0062] Figure 4 It is an early fusion strategy that concatenates structural features and functional features as fusion features, and then uses a layer of GAT and a layer of GCN to obtain the classification results. Figure 5 It is a mid-term fusion strategy, in which the structural features and functional features are trained through a layer of GAT respectively and then concatenated, and then passed through a layer of GCN to obtain the classification results. Figure 6For the late fusion strategy, the structural features and functional features are trained using one layer of GAT and one layer of GCN respectively, and then the output results are weighted and summed to obtain the classification result. For the image data of the two most important modalities, sMRI and PET, in the dataset, the above multi-modal fusion strategy is adopted in this embodiment to conduct four binary classification experiments of NC vs. AD, NC vs. EMCI, EMCI vs. LMCI, and LMCI vs. AD respectively, and at the same time, binary classification experiments under the sMRI and PET single-modal image data are conducted. The results are shown in Table 2.

[0063] Table 2 Experimental results of different modality fusion strategies

[0064] As can be seen from Table 2, Fusion a based on early fusion achieved excellent performance in the classification task, with an average accuracy (ACC) of 95.36%, which is at least 4.07% higher than other strategies. The average ACCs of the other two fusion strategies are 91.29% and 88.59% respectively. In addition, compared with using only sMRI single-modal data (accuracy rate of 92.84%), it is increased by 2.52%; compared with using only PET single-modal data (accuracy rate of 90.89%), it is increased by 4.47%. Therefore, in the subsequent experiments, Fusion a of early fusion is selected as the fusion strategy in this embodiment. Performance comparison of combined double-layer graph neural networks In this embodiment, a two-layer graph neural network architecture is set up. In order to obtain the best combined result of the double-layer network, this embodiment compares the combinations of various popular frameworks, and all experimental results are based on the early fusion strategy. The experimental results are shown in Table 3.

[0065] Table 3 Comparison results of combinations of various graph neural networks and MLP

[0066] In Table 3, the parameter configuration of the MLP is the same as that of other graph neural networks. As can be seen from Table 3, when the first layer is the MLP and the second layer is the GCN, the average ACC is improved by 25.20% compared with the combination where the second layer is the MLP. This result strongly verifies the effectiveness of the GCN graph neural network. It can also be seen from Table 3 that the GAT+GCN combination achieves the best result, with an average ACC of 95.36%, which is 1.26% higher than the MLP+GCN combination, further proving the effectiveness of the double-layer GNN framework. In addition, further analysis of Table 3 reveals that when the second layer uses the Graph Isomorphism Network (GIN) or GAT, the combination with the MLP in the first layer shows higher performance than the combination with the GAT in the first layer. This result indicates that an inappropriate combination of graph neural networks may lead to a decline in performance, highlighting the importance of the adaptation between graph neural network combinations.

[0067] Performance Comparison between this Embodiment and Related Works To verify the advancement of the proposed MMGNN model, this embodiment conducts a performance comparison between it and related research on the AD classification problem in recent years. The specific results are shown in Table 4.

[0068] Table 4 Performance Comparison with Related Works

[0069] As can be seen from Table 4, in the NC vs. AD classification task, the model proposed by Y. Zhang performed the best among previous related works, while the model in this embodiment exceeded Y. Zhang's model in terms of ACC by 1.9% and improved by 4.6% in terms of SPE. Similarly, in the NC vs. EMCI classification task, the model proposed by X. Song et al. performed the best among previous related works, while the model in this embodiment improved by 2.6% and 2.5% in terms of ACC and SPE respectively. The model in this embodiment obtained better results through five-fold cross-validation, demonstrating its good stability.

[0070] Hyperparameter Experiments The two key hyperparameters in this embodiment are the number of attention heads in the GAT and the dimension of the output of the GAT hidden layer. Experiments were conducted on four binary classification tasks for verification, and the results are as Figure 7 and Figure 8 shown.

[0071] The number of attention heads was set to 2, 4, 8, and 16 respectively, and it was found that when the number was 8, the model had the best average ACC performance in the four classification tasks. Similarly, when the dimension of the output of the hidden layer was set to 2, 4, 8, and 16 respectively, the results showed that when it was set to 4, the model had the best average ACC performance in the four classification tasks.

[0072] Ablation experiments based on non-image information In this embodiment, ablation experiments were conducted to analyze the influence of non-image information, and the results are as Figure 9 shown below. In four classification tasks, different scenarios were set, including the scenario without any non-image information (e1) and scenarios lacking a certain specific information, such as category (e2), gender (e3), age (e4), MMSE score (e5), and normal (e6). In scenario e1, the ACC was 84.56%. While in the normal scenario e6, the ACC reached 95.36%, which was 10.8% higher than that in e1, highlighting the importance of non-image information in improving the model performance. Especially in scenario e3, when the gender information was missing, the ACC was the lowest, indicating that gender is the most influential among these factors. This study emphasizes the important role of gender in affecting cognition and brain reserve, especially for female patients with Alzheimer's disease.

[0073] In this embodiment, Alzheimer's disease (AD) classification was performed by combining imaging and non-image information. The innovation lies in using a non-linear function to simplify the process of brain region feature extraction, and constructing multi-layer graph neural networks based on the relationships between brain ROIs and the relationships between subjects respectively. In addition, this embodiment also analyzed the influence of different fusion strategies and combinations of graph neural networks on the model performance, as well as the influence of non-image information on the model. Compared with the existing related work models, the method proposed in this embodiment achieved better performance. In future work, the ADNI dataset will be mainly used, and combined with other brain imaging datasets to verify the generalization ability of the model.

[0074] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution is prior or posterior. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0075] Embodiment 3 This embodiment is used to implement the principle of the above method embodiment to construct an AD classification prediction system based on a multi-modal multi-layer graph neural network model, including a data acquisition sub-module, an image information extraction sub-module, an ROI graph sub-module, a non-image information extraction sub-module, and a subject graph sub-module; the data acquisition sub-module is used to acquire the image information and non-image information of the subject; the image information extraction sub-module is used to extract the feature vectors of the ROI from the image information by using a Gaussian kernel function; the ROI graph sub-module is used to construct an ROI graph according to the feature vectors of the ROI, and introduce an attention mechanism through GAT to obtain the hidden layer features of the subject from the ROI graph; the non-image information extraction sub-module is used to extract the relationships between subjects by calculating the differences in the non-image information of the subjects; the subject graph sub-module is used to construct a subject graph according to the hidden layer features of the subject and the relationships between subjects, output the subject category prediction value through GCN, and apply a threshold to judge the category of the subject.

[0076] The multi-modal multi-layer graph neural network model (MMGNN) consists of one layer of graph attention network (GAT) and one layer of graph convolutional network (GCN). The input of the GAT layer is the brain ROI features extracted from sMRI and PET data; by introducing the self-attention mechanism, GAT can dynamically adjust the weights of information propagation between nodes, so as to flexibly capture the complex interaction relationships between brain ROIs. Subsequently, the hidden layer features generated by the GAT layer are passed to the GCN layer as input. The GCN layer performs convolutional operations on the normalized adjacency matrix to achieve weighted aggregation of neighborhood features, efficiently aggregates node features, and finally generates preliminary results. Then, the scores of each subject are normalized through the Softmax layer, and the results are obtained by classifying according to the defined threshold. The structure of this model is as Figure 2 shown.

[0077] Figure 2 In it, ① Collect data including image information and non-image information, extract the brain ROI features of the image information, and construct a non-image information matrix. ② The GAT network architecture uses the brain ROI features to construct a graph and trains to output hidden layer features. ③ Use the non-image information matrix to perform weighted summation to construct a non-image similarity matrix, calculate the similarity using the hidden layer features to obtain an image similarity matrix, and multiply these two matrices to construct the adjacency matrix of the subject. ④ The GCN network architecture uses the hidden layer features and the adjacency matrix of the subject to construct a graph and trains to output the prediction result, and sets a threshold to obtain the classification result.

[0078] Each sub-module is mainly used to implement each step of the method embodiment, which will not be elaborated here.

[0079] It should be noted that according to the needs of implementation, each step / component described in this application can be split into more steps / components, or two or more steps / components or partial operations of steps / components can be combined into new steps / components to achieve the purpose of the present invention.

[0080] This embodiment further includes a processor, a communication interface, a memory, and a communication bus; wherein the processor, the communication interface, and the memory complete communication with each other through the communication bus; a computer program is stored in the memory, and when the program is executed by the processor, the processor executes the steps of the AD classification prediction method based on the multi-modal multi-layer graph neural network model.

[0081] This embodiment also provides a computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by the processor, the processor implements the AD classification prediction method based on the multi-modal multi-layer graph neural network model.

[0082] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects.

[0083] Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0084] The present application is described with reference to the flowcharts of the method and computer program product according to Embodiment 1 of the present application and the block diagrams of the device (system) according to Embodiment 3. It should be understood that each process or block in the flowchart or block diagram can be implemented by computer program instructions, and the combination of processes or blocks in the flowchart or block diagram can also be implemented by computer program instructions.

[0085] These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a multi-modal multi-layer graph neural network model-based AD classification prediction system for implementing the functions specified in one Figure 1 one process or multiple processes or blocks Figure 1 one block or multiple blocks.

[0086] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions in the process Figure 1One process or multiple processes or boxes Figure 1 The functions specified in one box or multiple boxes.

[0087] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the steps of the method for AD classification prediction based on a multi-modal multi-layer graph neural network model specified in one process or multiple processes or boxes Figure 1 One process or multiple processes or boxes Figure 1 The steps of the method for AD classification prediction based on a multi-modal multi-layer graph neural network model specified in one box or multiple boxes.

[0088] The above embodiments are only used to illustrate the design concept and features of the present invention, and the purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, all equivalent changes or modifications made according to the principles and design concepts disclosed in the present invention are within the protection scope of the present invention.

Claims

1. An AD classification prediction method based on a multi-modal multi-layer graph neural network model, characterized in that: It includes the following steps: S0: Obtain the image information and non-image information of the subject; S1: Use the Gaussian kernel function to extract the feature vectors of the ROI from the image information; S2: Construct an ROI graph according to the feature vectors of the ROI; Introduce an attention mechanism through GAT to obtain the hidden layer features of the subject from the ROI graph; S3: Extract the relationships between subjects by calculating the differences in the non-image information of the subjects; S4: Construct a subject graph according to the hidden layer features of the subjects and the relationships between the subjects; Output the subject category prediction value through GCN, and apply a threshold to judge the category of the subject.

2. The AD classification prediction method based on the multi-modal multi-layer graph neural network model according to claim 1, characterized in that: In the step S0 described above, the specific steps are: The image information of the subject includes sMRI and PET; The non-image information of the subject includes category, gender, age, and MMSE.

3. The AD classification prediction method based on the multi-modal multi-layer graph neural network model according to claim 2, characterized in that: In the step S1 described above, the specific steps are: S11: Perform standardized preprocessing on the data of sMRI and PET; S12: Extract the corresponding original features based on the ROI defined by the neuromorphic brain atlas; S13: Use the Gaussian kernel function to map the original features of each ROI, extract the structural feature vectors from sMRI, and extract the functional feature vectors from PET.

4. The AD classification prediction method based on the multi-modal multi-layer graph neural network model according to claim 3, characterized in that: In the step S14 described above, the specific steps are: After extracting the original features of all ROIs by comparing the brain atlas extraction images, calculate the feature coefficients of a certain ROI and other ROIs through the Gaussian kernel, and combine these feature coefficients to obtain the feature vector of the ROI.

5. The AD classification prediction method based on the multi-modal multi-layer graph neural network model according to claim 1, characterized in that: In the step S2 described above, the specific steps are: S21: Use nodes to represent ROIs, node features to represent the feature vectors of ROIs, edges to represent the relationships between ROIs, and the adjacency matrix of the edges to represent the relationship strength between nodes, and construct an ROI graph composed of nodes, node features, edges, and the adjacency matrix of the edges; S22: Calculate the Spearman relationship coefficient of two ROIs according to the number of ROIs and the Euclidean distance between the two ROIs; S23: Under the constraint of the Spearman relationship coefficient, obtain the attention score between two ROIs by concatenating, transforming, weighting, and non-linearizing the feature vectors of the two ROIs; S24: Normalize the attention scores between nodes to obtain attention coefficients and form an attention matrix; Dynamically allocate node feature weights according to the magnitude of the normalized attention scores; S25: Obtain the hidden layer features of the subject through GAT according to the attention scores.

6. The AD classification prediction method based on the multi-modal multi-layer graph neural network model according to claim 1, characterized in that: In the step S4 described above, the specific steps are: S41: Use nodes to represent subjects, node features to represent the hidden layer features obtained from GAT, edges to represent the relationships between subjects, and the adjacency matrix of the edges to represent the relationship strength between nodes, and construct a subject graph composed of nodes, node features, edges, and the adjacency matrix of the edges; S42: The GCN outputs the predicted value of the subject category after processing the adjacency matrix of the edges in the subject graph, the matrix composed of the subject hidden layer features, and the weight matrix using the function and the function.

7. The AD classification prediction method based on the multi-modal multi-layer graph neural network model according to claim 6, wherein: In the step S41 described above, use the cosine similarity to calculate the similarity of the hidden layer features, and combine the hidden layer features and non-image information to calculate the adjacency matrix of the edges.

8. The AD classification prediction method based on the multi-modal multi-layer graph neural network model according to claim 1, wherein: It also includes the following steps: S5: Adopt a cross-entropy loss function with class weights to solve the class imbalance problem by imposing a higher penalty on the misclassification of the subject categories.

9. An AD classification prediction system based on a multi-modal multi-layer graph neural network model, characterized in that: A data acquisition sub-module for acquiring the image information and non-image information of the subject; An image information extraction sub-module for extracting the feature vectors of the ROI from the image information by using a Gaussian kernel function; An ROI graph sub-module for constructing an ROI graph according to the feature vectors of the ROI, and introducing an attention mechanism through GAT to obtain the hidden layer features of the subject from the ROI graph; A non-image information extraction sub-module for extracting the relationships between subjects by calculating the differences of the subjects in the non-image information; A subject graph sub-module for constructing a subject graph according to the hidden layer features of the subject and the relationships between subjects, outputting the subject category prediction value through GCN, and applying a threshold to judge the category of the subject.

10. A computer memory, characterized in that: It stores a computer program executable by a computer processor, and this computer program executes the AD classification prediction method based on the multi-modal multi-layer graph neural network model according to any one of claims 1 to 8.