A data classification method and system based on deep multi-path attention adaptive graph convolutional network

Through the Deep Multi-path Attention Adaptive Graph Convolutional Network (DMAGCN) model, the problems of uncertainty and inconsistent dataset distribution in the diagnosis of autism spectrum disorder are solved, and high-accuracy and generalization data classification are achieved, especially in the diagnosis of autism spectrum disorder.

CN119478550BActive Publication Date: 2025-09-30WENZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411914971.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-09-30
Estimated Expiration
2044-12-24

AI Technical Summary

Technical Problem

Existing technologies face problems such as high uncertainty, inconsistent dataset distribution, and small dataset size in diagnosing autism spectrum disorder (ASD), resulting in poor robustness of classification models.

Method used

A method based on the Deep Multi-path Attention Adaptive Graph Convolutional Network (DMAGCN) is adopted. Through a model composed of a backbone network and multiple branch networks, common features and unique features are extracted respectively. The graph network is combined with non-imaging data for multimodal information extraction, and the softmax function is used for classification.

Benefits of technology

It has improved the classification accuracy of autism spectrum disorder diagnosis and the generalization ability of the model, especially when processing uncertain data, and promoted the development of related disease research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119478550B_ABST
    Figure CN119478550B_ABST
Patent Text Reader

Abstract

The present invention discloses a data classification method and system based on a deep multi-path attention adaptive graph convolutional network, which relates to the field of medical image processing technology. rs-fMRI data is acquired and pre-processed to obtain a BOLD sequence, a functional connectivity feature vector is constructed and input into a DMAGCN model, and finally the optimal model is obtained through five-fold cross-validation for classification. The present invention ensures data quality through data preprocessing, laying the foundation for accurate analysis. The functional connectivity feature vector effectively represents the data. After input into the model, the Transformer backbone network and the MLP branch network can extract multi-source domain features. In combination with the graph network, non-imaging data is utilized, so that the model can learn rich features and enhance generalization and adaptability. The present invention improves classification accuracy, especially when processing uncertain data classification such as autism spectrum disorder, promotes research on related diseases, and provides an efficient and reliable method for medical data classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical image processing technology, and more particularly to a data classification method and system based on a deep multi-path attention adaptive graph convolutional network. Background Art

[0002] Autism Spectrum Disorder (ASD) is a childhood neurodevelopmental disorder characterized by stereotyped behaviors, narrow interests, and communication difficulties. Early intervention is more effective, but some ASD patients go undetected in childhood and symptoms are observed in adolescence. Therefore, early prevention is crucial for autism.

[0003] Over the past few decades, and particularly in the past decade, advances in neuroimaging technology have provided a critical step forward, enabling the measurement of functional and structural changes associated with ASD. Several methods based on resting-state functional magnetic resonance imaging (rs-fMRI) have been proposed for the diagnosis of various brain disorders and have garnered widespread attention in the discovery and classification of ASD biomarkers. It is a key tool that can reveal brain dysfunction based on the blood oxygenation level-dependent (BOLD) signal when the subject is at rest. Functional connectivity is calculated by using the average BOLD signal between two brain regions in rs-fMRI. Due to its effectiveness in identifying brain functional organization and biomarkers for neuropsychiatric disorders, it has been widely used in computer-assisted ASD diagnosis. Most rs-fMRI-based ASD classification methods are developed using functional connectivity as a feature.

[0004] Deep learning has been successfully applied to disease diagnosis and has achieved remarkable results. However, human neural activity is characterized by uncertainty. This is because the process of clinically acquiring rs-fMRI data is subject to interference caused by the human body's own noise and equipment, making it difficult to reflect true neural activity. Secondly, ASD is highly heterogeneous, and there are obvious individual differences in the neural activity of different ASD patients. These factors make rs-fMRI data contain significant uncertainty, which is not conducive to building a robust ASD auxiliary diagnosis model. At the same time, clinical neuroimaging datasets often face the problem of small dataset size due to their expensive acquisition and time-consuming labeling. Therefore, in some studies, such as ASD diagnosis, multi-site rs-fMRI data are often merged to expand the dataset, which leads to the second problem: in most cases, samples from different scanners or acquisition protocols do not follow the same distribution. Sun et al. minimized domain shift by aligning the second-order statistics of the source and target domain distributions; Liu et al. mitigated marginal distribution differences between domains by adjusting the global structure of predicted multi-site data; Wang et al. aimed to reduce data distribution differences by determining a common low-rank representation of data from multiple sites; Eslami et al. used autoencoders and single-layer perceptrons (SLPs) to improve the quality of extracted features and optimize model parameters. These methods are effective in addressing the problem of inconsistent feature distribution in datasets, but they are not detailed enough in processing features, which can lead to misdiagnosis, which is very serious in the medical field.

[0005] Therefore, how to handle the classification of uncertain data is an urgent problem that those skilled in the art need to solve. Summary of the Invention

[0006] In view of this, the present invention provides a data classification method and system based on a deep multi-path attention adaptive graph convolutional network, and proposes a new deep neural network - DMAGCN, which can handle uncertain data classification; it includes a backbone network and multiple branch networks. The backbone network is constructed by Transformer, and common features of different source domains are extracted. The branch networks are constructed by MLP and GCN, and the number is the same as the number of source domains. This part extracts the unique features of the source domain. GCN combines image data and non-image data by constructing a population map, and extracts multimodal information for model training; in the feature classification stage, the features obtained from each branch are connected, and the softmax function is used to obtain the final classification result, thereby realizing the classification of uncertain data.

[0007] In order to achieve the above object, the present invention adopts the following technical solutions:

[0008] A data classification method based on deep multi-path attention adaptive graph convolutional network, including:

[0009] Resting-state functional magnetic resonance imaging data were obtained from the Autism Brain Imaging Data Exchange database and preprocessed using a configurable pipeline for connectome analysis to obtain BOLD sequences.

[0010] The functional feature matrix is ​​constructed using the BOLD sequence, and the upper triangular part of the functional feature matrix is ​​straightened to obtain the functional connectivity feature vector as the input of the DMAGCN model;

[0011] Construct a DMAGCN model, which consists of a backbone network and multiple branch networks. The backbone network is built with Transformer, and the branch networks are composed of MLP and graph networks respectively. The backbone network is used to extract common features of all source domains, and the branch networks are used to extract unique features of a single source domain.

[0012] The DMAGCN model was verified based on the five-fold cross-validation method to obtain the optimal DMAGCN model;

[0013] The data to be classified is input into the optimal DMAGCN model to obtain the data classification results.

[0014] Optionally, the preprocessing involves skull delineation, slice timing correction, motion correction, global mean intensity normalization, noise signal regression, bandpass filtering 0.01–0.1 Hz, and registration of resting-state functional magnetic resonance imaging data to a standard anatomical space.

[0015] Optionally, the calculation formula of the functional connectivity feature vector is as follows:

[0016]

[0017] Among them, X i and X j is the average time series of the i-th ROI region and the j-th ROI region; X i,t ,X j,t It's X i and X j BOLD intensity at time moment t; and denotes the mean of the average BOLD time series of the i-th brain region and the j-th brain region, respectively; T denotes the total number of time points of the average BOLD series; and corr denotes the Pearson correlation coefficient.

[0018] Optionally, the backbone network built with MLP and Transformer extracts features from imaging data. In addition to imaging features, the graph network uses non-imaging data as edges between nodes to provide multimodal information to train the DMAGCN model. After training, the graph network will be discarded, and other networks will jointly make decisions on the target domain.

[0019] Optionally, the backbone network adopts Transformer as the backbone network, and the overall structure of Transformer includes a multi-head self-attention module, a feedforward network, a residual connectivity layer and a normalization layer;

[0020] The self-attention mechanism is the core of the Transformer-encoder and is calculated from the Query, Key, and Value matrices:

[0021]

[0022] in N and M represent the length of Query and Key, D k and D v Represents the dimensions of Key and Value; Softmax is an activation function that converts attention scores into probabilities; the multi-head attention mechanism is applied in Transformer:

[0023]

[0024] in, are the parameter matrices corresponding to Q, K, and V respectively; W 0 is the parameter matrix for multi-head attention calculation; the feedforward network then applies two linear transformations with Gelu activation function to the output of the multi-head self-attention:

[0025] X=FFN(x)=Gelu(xW1+b1)W2+b2;

[0026] Where x is the output of the previous layer, W1, W2, b1, and b2 represent the training parameter matrix and bias values.

[0027] Optionally, in order to complete the feature extraction of multiple tasks, multiple branch networks are set up to extract the unique features of each source domain. Represents the input features, which consists of fully connected layers. Each layer has multiple nodes. Each node receives the output of the node in the previous layer as input. The output of the k-th node in the l-th layer is:

[0028]

[0029] Among them, k (l)is the output of the lth layer, l=1,2,...,L;k=1,2,...,K;j=1,2,...,J;k≠j; is the feature of the kth node in the l-1th layer; θ jk (l) represents the connection weight between the kth nodes in the lth layer; α l represents the activation function of the lth layer;

[0030] Use the backbone network to propose common features of all source domains:

[0031]

[0032] Use the branch network to extract the unique features of the i-th source domain:

[0033]

[0034] Where i=1,2,3...I represents the i-th source domain sample or target domain sample; s is the source domain dataset; f c,i (·) is the common feature extractor; f s,i (·) is the extractor of unique features; Θ s,i Parameters of the i-th source domain-specific feature extractor; Θ c,i represents the parameters of the i-th common feature extractor; represents the common features of the i-th source domain sample in the source domain; represents the unique features of the i-th source domain sample.

[0035] Optionally, the maximum mean difference is used to measure the distribution distance between the common features and unique features related to the source domain, specifically:

[0036]

[0037] Where n is the number of samples in the source domain; i and j represent the sample numbers, i≠j; represents the Gaussian kernel Hilbert space.

[0038] The domain alignment loss is:

[0039] L domain =L com +L s +L cs .

[0040] Optionally, the graph network is specifically an edge variational graph convolutional network, which uses the spatial perception of the brain network and the population relationship of the dataset to train and optimize the model;

[0041] Given N subjects’ data consisting of imaging and non-imaging data, we construct a population graph: G = (V, E, W), where |V| = N represents the vertex set. is a set of edges, where the weight of the edge is W; define the node characteristics As the C-dimensional feature vector extracted from the imaging data of the i-th subject; (x i ,x j ) between the weight w i,j ∈W is defined as a learnable function representing the non-imaging data information: φ:(x i ,x j ), which is modeled and trained by the Pairwise Associative Encoder PAE:

[0042] h i =φ(x i ,Ω);

[0043] h j =φ(x j ,Ω);

[0044]

[0045] Among them, x% i is the normalized input; τ is the ReLU function; h i and h j is the input feature x i and x j Mapping in the same feature space; Ω represents the parameters trained in PAE;

[0046] The uncertainty-aware prediction loss function is:

[0047]

[0048] Among them, P(x i ) represents the predicted value of the i-th sample, P%(x i ) is the true value of the i-th sample;

[0049] Therefore, the total loss function of the DMAGCN model is:

[0050]

[0051] L=λL domain +L ev ;

[0052] Where λ changes from 0 to 1 over time, γ is a hyperparameter, and ρ represents the number of iterations.

[0053] Optionally, the convolution layer of the graph convolution is composed of Chebyshev convolution, and the recurrence relationship of Chebyshev polynomials is:

[0054] T0(L)=1,T1(L)=L;

[0055] T k (L)=2LT k-1 (L)-T k-2 (L);

[0056] The formula for Chebyshev convolution is:

[0057]

[0058] Among them, T k (L) represents the topological structure L of the graph G after the Chebyshev polynomial calculation in the kth term; H l Represents the node feature vector of the lth layer; θ k l is the convolution kernel parameter.

[0059] A data classification system based on deep multi-path attention adaptive graph convolutional network, including:

[0060] The data acquisition and processing module obtains resting-state functional magnetic resonance imaging data from the Autism Brain Imaging Data Exchange database and preprocesses them using a configurable pipeline for connectome analysis to obtain BOLD sequences.

[0061] The feature extraction module uses the BOLD sequence to construct a functional feature matrix, and straightens the upper triangular part of the functional feature matrix to obtain the functional connectivity feature vector as the input of the DMAGCN model;

[0062] The model construction module builds the DMAGCN model, which consists of a backbone network and multiple branch networks. The backbone network is built using Transformer, and the branch networks are composed of MLP and graph networks respectively. The backbone network is used to extract common features of all source domains, and the branch networks are used to extract unique features of a single source domain.

[0063] The model validation module verifies the DMAGCN model based on the five-fold cross-validation method to obtain the optimal DMAGCN model;

[0064] The result output module inputs the data to be classified into the optimal DMAGCN model to obtain the data classification results.

[0065] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses a data classification method and system based on a deep multi-path attention adaptive graph convolutional network. First, the BOLD sequence is obtained by obtaining and preprocessing the rs-fMRI data from the ABIDE database, and then a functional connection feature vector is constructed to input into the DMAGCN model. The model is composed of a specific structural network, and finally the optimal model is obtained through five-fold cross validation for classification. The present invention ensures data quality through data preprocessing and lays the foundation for accurate analysis. The functional connection feature vector effectively represents the data. After inputting the model, the Transformer backbone network and the MLP branch network can extract multi-source domain features. In combination with the graph network, the non-imaging data is utilized, so that the model can learn rich features and enhance generalization and adaptability. Five-fold cross validation ensures model reliability and improves classification accuracy. It performs well in classifying uncertain data such as autism spectrum disorder, promotes research on related diseases, and provides an efficient and reliable method for medical data classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0067] Figure 1 A schematic diagram of the process of constructing the functional connectivity feature vector provided by the present invention;

[0068] Figure 2 The overall structure diagram of the DMAGCN model provided by the present invention;

[0069] Figure 3 This is a diagram of the trunk and branch network structure provided by the present invention. DETAILED DESCRIPTION

[0070] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0071] The embodiment of the present invention discloses a data classification method based on a deep multi-path attention adaptive graph convolutional network, comprising:

[0072] Resting-state functional magnetic resonance imaging data were obtained from the Autism Brain Imaging Data Exchange database and preprocessed using a configurable pipeline for connectome analysis to obtain BOLD sequences.

[0073] The functional feature matrix is ​​constructed using the BOLD sequence, and the upper triangular part of the functional feature matrix is ​​straightened to obtain the functional connectivity feature vector as the input of the DMAGCN model;

[0074] Construct a DMAGCN model, which consists of a backbone network and multiple branch networks. The backbone network is built with Transformer, and the branch networks are composed of MLP and graph networks respectively. The backbone network is used to extract common features of all source domains, and the branch networks are used to extract unique features of a single source domain.

[0075] The DMAGCN model was verified based on the five-fold cross-validation method to obtain the optimal DMAGCN model;

[0076] The data to be classified is input into the optimal DMAGCN model to obtain the data classification results.

[0077] In a specific embodiment, the dataset is provided by the Autism Brain Imaging Data Exchange (ABIDE) and has been preprocessed using a configurable pipeline for connectome analysis (C-PAC). The acquisition device is Siemens, and the process involves skull striping, slice timing correction, motion correction, global mean intensity normalization, interference signal regression, bandpass filtering (0.01-0.1 Hz), and registering fMRI images to a standard anatomical space (MNI152). The dataset used comes from six acquisition sites and includes N=1735 subjects. Specific statistical information is shown in the following table:

[0078] Table 1 Dataset information

[0079]

[0080]

[0081] Explanation of parameters in the table:

[0082] Repetition time (TR): The time interval between successive pulse trains applied to the same slice.

[0083] Echo Time (TE): is the time between the peak of the transmitted and received pulses.

[0084] Short TR and short TE enhance T1-weighted contrast, making tissues with short T1 delay times (such as fat) appear bright, while tissues with longer T1 times (such as fluid) appear dark.

[0085] Longer TR and TE values ​​produce T2-weighted images in which tissues with longer T2 relaxation times (e.g., fluids) appear bright. Longer TR ensures complete longitudinal relaxation, while longer TE allows for adequate lateral dephasing of T2 contrast.

[0086] Voxel size: affects image resolution. The smaller the voxel, the higher the resolution. It also affects the results of data analysis. Too large a voxel may not be able to capture subtle differences between brain regions.

[0087] In a specific embodiment, the original feature extraction is as follows Figure 1 As shown in the figure, for each sample, the Pearson correlation coefficient between any pair of ROI regions in each brain is calculated to obtain the final functional connectivity feature matrix. Due to the symmetry of the matrix, only the upper triangle of the matrix is ​​selected and this part is stretched into a one-dimensional vector as the functional connectivity feature vector. The calculation formula of the functional connectivity feature vector is as follows:

[0088] The calculation formula of functional connectivity feature vector is as follows:

[0089]

[0090] Among them, X i and X j is the average time series of the i-th ROI region and the j-th ROI region; X i,t ,X j,t It's X i and X j BOLD intensity at time moment t; and denotes the mean of the average BOLD time series of the i-th brain region and the j-th brain region, respectively; T denotes the total number of time points of the average BOLD series; and corr denotes the Pearson correlation coefficient.

[0091] In a specific embodiment, the whole process flow is as follows Figure 2As shown in the figure, it includes four parts: data preprocessing, functional connectivity feature extraction, extraction of mixed features, and classification. Specifically, the rs-fMRI data is first preprocessed to obtain the BOLD sequence. The BOLD sequence is then used to construct a functional feature matrix, and the upper triangular part of the matrix is ​​straightened to obtain the functional connectivity feature vector. Because the data sets involved come from different imaging centers, this is a multi-source domain problem. The model consists of a backbone network plus multiple branch networks. The backbone network is built using Transformer, and the branch networks are composed of MLP and graph networks respectively. The network built by MLP and Transformer mainly extracts features from imaging data. In addition to imaging features, the graph network uses non-imaging information as edges between nodes to provide multimodal information for model training. The backbone network is used to extract common features of all source domains, and the branch network is used to extract unique features of a single source domain. After training, the graph network will be discarded, and the other networks will jointly make decisions on the target domain.

[0092] Branch network structure such as Figure 3 As shown in the figure, it consists of a fully connected layer. The input features are the output of the Transformer structure. The data will be reduced in dimension in the deep neural network. The neural network will learn the feature information from it and extract key features for subsequent model optimization and classification.

[0093] In a specific embodiment, the backbone network uses Transformer as the backbone network, and the overall structure of Transformer includes a multi-head self-attention module, a feedforward network, a residual connectivity layer, and a normalization layer;

[0094] The self-attention mechanism is the core of the Transformer-encoder and is calculated from the Query, Key, and Value matrices:

[0095]

[0096] in N and M represent the length of Query and Key, D k and D v Represents the dimensions of Key and Value; Softmax is an activation function that converts attention scores into probabilities; the multi-head attention mechanism is applied in Transformer:

[0097]

[0098] in, are the parameter matrices corresponding to Q, K, and V respectively; W 0is the parameter matrix for multi-head attention calculation; the feedforward network then applies two linear transformations with Gelu activation function to the output of the multi-head self-attention:

[0099] X=FFN(x)=Gelu(xW1+b1)W2+b2;

[0100] Where x is the output of the previous layer, W1, W2, b1, and b2 represent the training parameter matrix and bias values.

[0101] In order to complete the multi-task feature extraction, multiple branch networks are set up to extract the unique features of each source domain. Represents the input features, which consists of fully connected layers. Each layer has multiple nodes. Each node receives the output of the node in the previous layer as input. The output of the k-th node in the l-th layer is:

[0102]

[0103] Among them, k (l) is the output of the lth layer, l=1,2,...,L;k=1,2,...,K;j=1,2,...,J;k≠j;oj ( l- 1) is the feature of the kth node in the l-1th layer; θ jk (l) represents the connection weight between the kth nodes in the lth layer; α l represents the activation function of the lth layer;

[0104] Use the backbone network to propose common features of all source domains:

[0105]

[0106] Use the branch network to extract the unique features of the i-th source domain:

[0107]

[0108] Where i=1,2,3...I represents the i-th source domain sample or target domain sample; s is the source domain dataset; f c,i (·) is the common feature extractor; f s,i (·) is the extractor of unique features; Θ s,i Parameters of the i-th source domain-specific feature extractor; Θ c,i represents the parameters of the i-th common feature extractor; represents the common features of the i-th source domain sample in the source domain; represents the unique features of the i-th source domain sample.

[0109] Maximum mean difference (MMD) is the most widely used loss function in transfer learning, especially in domain adaptation. It is mainly used to measure the distance between the distributions of two different but related random variables. Here, MMD is used to measure the distance between the distributions of common and unique features related to the source domain. Specifically:

[0110]

[0111] Where n is the number of samples in the source domain; i and j represent the sample numbers, i≠j; represents the Gaussian kernel Hilbert space.

[0112] The domain alignment loss is:

[0113] L domain =L com +L s +L cs .

[0114] In a specific embodiment, uncertainty-aware population graph construction, specifically an edge-variational graph convolutional network, provides a method for constructing an adaptive population graph with partially labeled nodes and variational edges. This method integrates imaging and non-imaging data from a population for uncertainty-aware disease prediction, and its effectiveness has been demonstrated. This model uses this as one of its branches, aiming to leverage the spatial perception of brain networks and the population relationships of the dataset to train and optimize the model.

[0115] Given N subjects’ data consisting of imaging and non-imaging data, we construct a population graph: G = (V, E, W), where |V| = N represents the vertex set. is a set of edges, where the weight of the edge is W; define the node characteristics As the C-dimensional feature vector extracted from the imaging data of the i-th subject; (x i ,x j ) between the weight w i,j ∈W is defined as a learnable function representing the non-imaging data information: φ:(x i ,x j ), which is modeled and trained by the Pairwise Associative Encoder PAE:

[0116] h i =φ(x i ,Ω);

[0117] h j =φ(x j ,Ω);

[0118]

[0119] Among them, x% i is the normalized input; τ is the ReLU function; h i and h j is the input feature x i and x j Mapping in the same feature space; Ω represents the parameters of the mapping network φ(·);

[0120] The convolution layer of graph convolution is composed of Chebyshev convolution, and the recurrence relationship of Chebyshev polynomials is:

[0121] T0(L)=1,T1(L)=L;

[0122] T k (L)=2LT k-1 (L)-T k-2 (L);

[0123] The formula for Chebyshev convolution is:

[0124]

[0125] Among them, T k (L) represents the topological structure L of the graph G after the Chebyshev polynomial calculation in the kth term; H l Represents the node feature vector of the lth layer; is the convolution kernel parameter.

[0126] The uncertainty-aware prediction loss function is:

[0127]

[0128] Among them, P(x i ) represents the predicted value of the i-th sample, is the true value of the i-th sample;

[0129] Therefore, the total loss function of the DMAGCN model is:

[0130]

[0131] Where λ changes from 0 to 1 over time, γ is a hyperparameter, and ρ represents the number of iterations.

[0132] In a specific embodiment, experimental results and analysis are also included;

[0133] Experimental setup

[0134] The experimental data comes from the ABIDE database. A new dataset was constructed using data from six sites, including NYU, USM, UM, Leuven, YALE, and UCLA. The experiment was based on five-fold cross-validation, a validation scheme commonly used in many studies. In the experiment, one of the six imaging centers was selected as the target domain, and the rest were selected as the source domain. Each source domain was divided into five subsets (the number of samples in each subset was similar, and the number of ASD patients and healthy controls in the subset was also roughly the same). In each fold of the cross-validation, four subsets were selected from each independent source domain as labeled source domain training sets, and the target domain data of all five subsets were selected as unlabeled target domain training sets. The target domain labels were not used in training, and were only used in testing to evaluate the classification performance of the model on the target domain.

[0135] The experiment ensured that each comparison algorithm maintained the same data partitioning as the proposed algorithm. The above process was repeated ten times, and the average was taken to evaluate each method. ACC, SEN, SPE, and AUC were used as evaluation metrics to quantitatively assess the classification performance of all methods.

[0136] Performance indicator evaluation

[0137] Accuracy (ACC), sensitivity (SEN), specificity (SPE), and AUC were used to measure the classification performance of all relevant methods.

[0138] ACC represents the proportion of correctly classified samples to the total number of samples, SEN represents the proportion of correctly classified samples among samples with ASD as the true sample, and SPE represents the proportion of correctly classified samples among samples with healthy controls as the true sample. The higher the values ​​of these three indicators, the better the classification performance of the model. The calculation method of ACC, SEN and SPE is as follows:

[0139]

[0140] In the above formula, TP is the number of true positive examples, i.e., samples with the true label ASD, that are predicted correctly; FN is the number of false negative examples, i.e., samples with the true label ASD, that are predicted incorrectly; TN is the number of true negative examples, i.e., samples with the true label healthy controls, that are predicted correctly; and FP is the number of false positive examples, i.e., samples with the true label healthy controls, that are predicted incorrectly.

[0141] Experimental Summary

[0142] Comparison with state-of-the-art methods and baselines

[0143] In order to fully verify the effectiveness of the method proposed in this invention, its results are compared with those of the latest methods.

[0144] ST-Transformer: A linear spatiotemporal multi-head attention unit is proposed to obtain spatial and temporal representations of fMRI data. In addition, a data balancing method based on Gaussian GAN is introduced to address the data imbalance problem in real-world ASD datasets for ASD subtype diagnosis.

[0145] ST-ASDNet: Two modules are proposed: Bidirectional Long Short-Term Memory Transformer (BLSTM-Transformer) and Fully Convolutional Network Transformer (FCN-Transformer), which respectively obtain the spatial and temporal features of fMRI data.

[0146] BrainNETTF: proposes an orthogonal clustering readout operation based on self-supervised soft clustering and orthogonal projection. This design considers the underlying functional modules that determine the similar behaviors between ROI groups, resulting in distinguishable cluster-aware node embeddings and information graph embeddings.

[0147] RGTNet: We propose a residual graph transformer network (RGTNet) for FC learning. We design a graph encoder to extract temporally relevant features with long-range dependencies, from which we can model an interpretable FC matrix. Furthermore, we introduce a residual technique to deepen the GCN architecture, thereby learning higher-level information.

[0148] AIMAFE: A multi-graph deep feature representation method based on stacked denoising autoencoders (SDA) is proposed. A multi-layer perceptron (MLP) and ensemble learning method are proposed to perform the final ASD recognition task.

[0149] MDANN: Captures the relationships in multimodal data (functional neuroimaging data and PC data) by integrating multi-layer neural networks, attention mechanisms, and feature fusion.

[0150] PLSNet: It designs a time series encoder for context-rich feature extraction, followed by a functional connection generator to model correlations with long-range dependencies. It uses positional embedding to uniquely identify each graph region. It also embeds a sparse method to filter significant nodes during message diffusion, which also helps reduce dimensionality.

[0151] MVS-GCN: A graph-structured learning algorithm that adaptively constructs clean brain networks via a supervised learning scheme. Compared to whole-brain networks, coarse graph representations facilitate brain network embedding learning and disease diagnosis. Furthermore, graph-structured learning considers group-level consistency across subjects from multiple sites by highlighting indicative edges, thereby eliminating noisy correlations in brain networks.

[0152] LRCDRl: Adopting a low-rank representation, it alleviates the marginal distribution differences between domains by aligning the global structure of projected multi-site data. To reduce the conditional distribution differences of all site data, LRCDR learns class-discriminative representations of data from multiple source and target domains to enhance the intra-class compactness and inter-class separability of the projected data.

[0153] This embodiment uses support vector machine (SVM), random forest (RF), and naive Bayes classifier (NB). In neuroimaging research, RF, SVM, and NB are often used as baselines.

[0154] Support Vector Machine: A linear classifier that uses the largest margin in feature space as its learning strategy. By introducing kernel methods to map the original features into a high-dimensional space, it can effectively become a nonlinear classifier. The penalty parameter C = 1, the kernel function kernel = "linear", and gamma = 1.

[0155] Random Forest: Using the decision tree as its basic unit, Random Forest integrates multiple decision trees using the principle of ensemble learning. Each decision tree randomly extracts data with replacement from the training set and then randomly selects a subset of features as training data. This ensures independence between different trees and improves the random forest's noise resistance. The number of decision trees in the random forest is n = 100.

[0156] Naive Bayes classifier: Based on Bayesian theory, it assumes that all features have a conditionally independent Gaussian distribution. The Naive Bayes algorithm learns the joint probability distribution of input and output data, and then uses Bayesian theorem to infer the label with the largest posterior probability as the prediction.

[0157] Table 2 Comparison of experimental results with various methods on ABIDE

[0158]

[0159]

[0160] Table 3 Comparison with machine learning methods

[0161]

[0162] Tables 2 and 3 show the average performance of 10 repetitions of 5-fold cross-validation. The proposed method outperforms all compared methods in terms of accuracy, achieving 75.6% on the AAL atlas, 73.6% on CC200, and 79.4% on dosenbatch160. Compared with Transformer-related models, the accuracy on AAL is 2.2% higher than that of the RGTNet model. Compared with GCN-related models, the accuracy on both AAL and CC200 is improved, reaching up to 6.7%. Compared with domain adaptation methods, the accuracy is improved by 2.5%. Comprehensive analysis shows that although other methods have demonstrated good performance in classification, they are inferior to the proposed method in addressing the inconsistent distribution of rs-fMRI feature space. The reasons for this are: 1. Feature extraction is not detailed enough. The proposed method divides features into common features and unique features, taking a broader perspective. 2. This invention uses a multi-network approach, setting up a network for each sample, reducing the risk of forgetting knowledge from the same network due to excessive knowledge. 3. This invention uses the construction of an uncertainty population graph and integrates it into a multi-branch network. Compared to other methods, this method considers more non-imaging data and the relationship between samples, making the model more generalizable.

[0163] Ablation experiments

[0164] This example verifies the effectiveness of each part of the model, and the results are shown in Table 4:

[0165] Table 4 Validity verification results

[0166]

[0167]

[0168] Ablation experiments were conducted at six sites: UM, NYU, USM, Leuven, UCLA, and YALE. The results are analyzed as follows:

[0169] 1. Removing the DANN part will result in a 4.5% drop in accuracy compared to the whole. DANN integrates the feature extractor and classifier, achieves domain alignment and classification in a simple way, and achieves domain alignment in an adversarial way. At the same time, removing the branch network alignment L s The accuracy dropped by 2.3% after removing the alignment between the common features and the unique features. cs After that, the accuracy dropped by 4.5%, proving the effectiveness of feature refinement.

[0170] 2. Remove the population map L evAfter that, the accuracy dropped by 7.8%, the highest among all losses. This proves that using EV-GCN as a branch network, introducing the relationship between samples, and non-imaging data is very effective for the diagnosis and classification of ASD, which enhances the generalization of the model.

[0171] The present invention proposes a new multi-center domain adaptive neural network, which is developed based on Transformer and population graph. Transformer helps to model the long-term dependencies between extracted data on time series data. The GCN network uses the form of population graph, combines imaging data and non-imaging data to optimize the model, and enhances the generalization of the model. In addition, this paper sets up a backbone network and multiple branch networks in the process of feature learning to extract common features and unique features of the source domain. This method attempts to introduce population graphs in multi-task ASD. The present invention is verified on the datasets of six imaging centers UM, NYU, USM, Leuven, YALE and UCLA in the three atlases of ABIDEI, and the accuracy can be increased to 79.4% at most.

[0172] A data classification system based on deep multi-path attention adaptive graph convolutional network, including:

[0173] The data acquisition and processing module obtains resting-state functional magnetic resonance imaging data from the Autism Brain Imaging Data Exchange database and preprocesses them using a configurable pipeline for connectome analysis to obtain BOLD sequences.

[0174] The feature extraction module uses the BOLD sequence to construct a functional feature matrix, and straightens the upper triangular part of the functional feature matrix to obtain the functional connectivity feature vector as the input of the DMAGCN model;

[0175] The model construction module builds the DMAGCN model, which consists of a backbone network and multiple branch networks. The backbone network is built using Transformer, and the branch networks are composed of MLP and graph networks respectively. The backbone network is used to extract common features of all source domains, and the branch networks are used to extract unique features of a single source domain.

[0176] The model validation module verifies the DMAGCN model based on the five-fold cross-validation method to obtain the optimal DMAGCN model;

[0177] The result output module inputs the data to be classified into the optimal DMAGCN model to obtain the data classification results.

[0178] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0179] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A data classification method based on deep multi-path attention adaptive graph convolutional network, characterized by: include: Resting-state functional magnetic resonance imaging data were obtained from the Autism Brain Imaging Data Exchange database and preprocessed using a configurable pipeline for connectome analysis to obtain BOLD sequences. The functional feature matrix is ​​constructed using the BOLD sequence, and the upper triangular part of the functional feature matrix is ​​straightened to obtain the functional connectivity feature vector as the input of the DMAGCN model; Construct a DMAGCN model, which consists of a backbone network and multiple branch networks. The backbone network is built with Transformer, and the branch networks are composed of MLP and graph networks respectively. The backbone network is used to extract common features of all source domains, and the branch networks are used to extract unique features of a single source domain. The DMAGCN model was verified based on the five-fold cross-validation method to obtain the optimal DMAGCN model; Input the data to be classified into the optimal DMAGCN model to obtain the data classification results; The backbone network built with MLP and Transformer extracts features from imaging data. In addition to imaging features, the graph network uses non-imaging data as edges between nodes to provide multimodal information to train the DMAGCN model. After training, the graph network will be discarded, and other networks will jointly make decisions on the target domain.

2. A data classification method based on deep multi-path attention adaptive graph convolutional network according to claim 1, characterized in that: The preprocessing involved skull striping, slice timing correction, motion correction, global mean intensity normalization, noise signal regression, bandpass filtering 0.01–0.1 Hz, and registration of resting-state functional magnetic resonance imaging data to a standard anatomical space.

3. The data classification method based on deep multi-path attention adaptive graph convolutional network according to claim 1 is characterized in that The calculation formula of the functional connectivity feature vector is as follows: ; in, and It is i ROI area and j The average time series of the ROI area; , yes and In the t BOLD intensity at a certain time moment; and Respectively represent i brain regions and j The mean of the average BOLD time series of the brain regions, T represents the total number of time points of the average BOLD series, and corr represents the Pearson correlation coefficient.

4. The data classification method based on deep multi-path attention adaptive graph convolutional network according to claim 1 is characterized in that The backbone network uses Transformer as the backbone network. The overall structure of Transformer includes a multi-head self-attention module, a feedforward network, a residual connectivity layer and a normalization layer; The self-attention mechanism is the core of the Transformer-encoder and is calculated from the Query, Key, and Value matrices: ; in , , , N and M Indicates the length of Query and Key, and Represents the dimensions of Key and Value; Softmax is an activation function that converts attention scores into probabilities; the multi-head attention mechanism is applied in Transformer: ; in, are respectively corresponding to The parameter matrix of is the parameter matrix for multi-head attention calculation; the feedforward network then applies two linear transformations with Gelu activation function to the output of the multi-head self-attention: ; in, x is the output of the previous layer, 、 、 and Represents the parameter matrix and bias values ​​for training.

5. The data classification method based on deep multi-path attention adaptive graph convolutional network according to claim 1 is characterized in that In order to complete the multi-task feature extraction, multiple branch networks are set up to extract the unique features of each source domain. Represents the input features, which is composed of fully connected layers. Each layer has multiple nodes, and each node receives the output of the previous layer node as input. l Tier k The output of each node is: ; in, is the output of layer l, ; ; ;k≠j; For the -1st floor The characteristics of each node; Indicates the Tier Node and The connection weights between nodes; represents the activation function of the lth layer; Use the backbone network to propose common features of all source domains: ; Extract the first i Unique characteristics of the source domain: ; in, Indicates the source domain samples or target domain samples; is the source domain dataset; is an extractor of common features; Extractor for unique features; No. Parameters of a source domain-specific feature extractor; Indicates the Parameters of the common feature extractor; Indicates the The common features of source domain samples in the source domain; Indicates the The unique features of the source domain samples.

6. A data classification method based on deep multi-path attention adaptive graph convolutional network according to claim 5, characterized in that: The maximum mean difference is used to measure the distribution distance between the common features and unique features related to the source domain, specifically: ; ; ; in, is the number of samples in the source domain; and Indicates the sample number, ; represents the Gaussian kernel Hilbert space; The domain alignment loss is: 。 7. A data classification method based on deep multi-path attention adaptive graph convolutional network according to claim 6, characterized in that: The graph network is specifically an edge variational graph convolutional network, which uses the spatial perception of the brain network and the population relationship of the dataset to train and optimize the model; Given data for N subjects consisting of imaging and non-imaging data, we construct a population graph: ,in represents a set of vertices, is a set of edges, where the weight of the edge is ; Define node characteristics As from the i The C-dimensional feature vector extracted from the imaging data of the subjects; The weight between Defined as a learnable function representing non-imaging data information , which is modeled and trained by the pairwise association encoder PAE: ; ; ; in, and The input features and Mapping in the same feature space; Represents the mapping network Parameters; The uncertainty-aware prediction loss function is: ; in, represents the predicted value of the i-th sample, is the true value of the i-th sample; Therefore, the total loss function of the DMAGCN model is: ; ; in changes from 0 to 1 over time, is a hyperparameter, Represents the number of iterations.

8. The data classification method based on deep multi-path attention adaptive graph convolutional network according to claim 7 is characterized in that: The convolution layer of graph convolution is composed of Chebyshev convolution, and the recurrence relationship of Chebyshev polynomials is: ; ; The formula for Chebyshev convolution is: ; in, Represents the topological structure of graph G After the Chebyshev polynomial calculation, Expression of terms; Indicates the The node feature vector of the layer; is the convolution kernel parameter.

9. A data classification system based on deep multi-path attention adaptive graph convolutional network, characterized by: A data classification method based on a deep multi-path attention adaptive graph convolutional network according to any one of claims 1 to 8 is applied, comprising: The data acquisition and processing module obtains resting-state functional magnetic resonance imaging data from the Autism Brain Imaging Data Exchange database and preprocesses them using a configurable pipeline for connectome analysis to obtain BOLD sequences. The feature extraction module uses the BOLD sequence to construct a functional feature matrix, and straightens the upper triangular part of the functional feature matrix to obtain the functional connectivity feature vector as the input of the DMAGCN model; The model construction module builds the DMAGCN model, which consists of a backbone network and multiple branch networks. The backbone network is built using Transformer, and the branch networks are composed of MLP and graph networks respectively. The backbone network is used to extract common features of all source domains, and the branch networks are used to extract unique features of a single source domain. The model validation module verifies the DMAGCN model based on the five-fold cross-validation method to obtain the optimal DMAGCN model; The result output module inputs the data to be classified into the optimal DMAGCN model to obtain the data classification results.