LSTM effective connectivity method for identifying MCI progress and reversal based on resting state fMRI

By combining LSTM with large-scale Granger causality analysis, an effective connectivity analysis framework for resting-state fMRI modalities was constructed, which solved the problems of unrevealed multimodal data dependence and causal effects, and enabled accurate identification and early intervention of MCI progression and reversal.

CN121237381APending Publication Date: 2025-12-30SHANGHAI SECOND POLYTECHNIC UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511394993.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

Existing technologies for the early diagnosis of Alzheimer's disease rely on multimodal data fusion, which is costly, fails to reveal the causal mechanisms between brain regions, ignores the MCI reversal phenomenon, and lacks systematic research.

Method used

By combining Long Short-Term Memory Network (LSTM) and Large-Scale Granger Causal Analysis (LSTM-lsGC), an effective connectivity analysis framework based on a single resting-state fMRI modality is constructed. By replacing the traditional vector autoregression model with LSTM, nonlinear causal interactions between brain regions are captured, and graph theory features are used to extract MCI progression and reversal.

Benefits of technology

This study achieved accurate identification of MCI progression and reversal in a single rs-fMRI modality, revealed the brain network reorganization mechanism, provided a new method for early intervention of Alzheimer's disease, and improved classification accuracy and identification ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BSA0000301590790000011
    Figure BSA0000301590790000011
  • Figure BSA0000301590790000012
    Figure BSA0000301590790000012
  • Figure BSA0000301590790000016
    Figure BSA0000301590790000016
Patent Text Reader

Abstract

The invention discloses an effective connectivity (EC) framework based on a Granger causality method enhanced by a long short-term memory (LSTM) network. The EC framework is used for classifying MCI subtypes and identifying key brain regions related to disease progression. According to the method, resting state functional magnetic resonance imaging (rs-fMRI) data of 52 MCI patients (including 7 rMCIs, 29 sMCIs and 16 pMCIs) are collected, and a framework fusing a healthy control-AD difference template (HAD) and a novel effective connectivity algorithm-LSTM-based large-scale Granger causality analysis (LSTM-lsGC) is constructed. According to the method, dimension reduction is carried out through principal component analysis, LSTM is used for replacing a traditional vector autoregression model to describe time dynamic features, then an effective connection matrix is constructed through Granger causal analysis, subtype discrimination is carried out through a random forest classifier after graph theory features are extracted, and Dunn inspection corrected through an error discovery rate is adopted for statistical analysis. According to the method, the LSTM enhanced Granger causal method is integrated to effectively realize the conjoint analysis of the MCI progress and reversal process, the proposed framework not only shows higher classification precision, but also reveals the key brain network change, and provides a potential approach for the early diagnosis and dynamic monitoring of MCI.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to a brain effective connection calculation method based on a long short-term memory network and large-scale Granger causality analysis (LSTM-lsGC), and belongs to the technical field of early diagnosis of Alzheimer's disease assisted by neural images and artificial intelligence. BACKGROUND

[0002] Current technologies are mostly focused on predicting the conversion of MCI to Alzheimer's disease (AD) by using multi-modal neural images (such as sMRI, PET) or functional connectivity (FC), which has improved the classification accuracy to a certain extent, but still has three limitations: first, it relies on multi-modal data fusion, which has high clinical implementation cost and low feasibility; second, traditional functional connectivity cannot reveal the causal mechanism between brain regions; and third, the phenomenon of MCI reversal (rMCI) is not systematically studied, and the potential role of brain network plasticity in cognitive recovery is ignored.

[0003] To break through these limitations, the application proposes to combine a long short-term memory network (LSTM) with large-scale Granger causality analysis (LSTM-lsGC) to construct an effective connection analysis framework based on a single resting-state fMRI modality. By establishing a healthy control and AD difference template (HAD), significant difference brain regions with low fluctuation amplitude (ALFF) are extracted, and then LSTM is used to replace the traditional vector autoregressive model to capture the nonlinear causal interaction between brain regions, and finally the graph theory features are extracted from the effective connection matrix to realize the accurate identification of MCI progression and reversal. This method first systematically integrates effective connection analysis and graph theory model, and is committed to revealing the brain network reorganization mechanism in the bidirectional conversion of MCI, and providing new direction of image biomarkers for early intervention. SUMMARY

[0004] The purpose of the application is to propose an LSTM effective connection method based on resting-state fMRI for identifying MCI progression and reversal, aiming to classify MCI subtypes and identify key brain regions related to disease progression.

[0005] The purpose of the application can be achieved by the following technical method: introducing principal component analysis in Granger causality analysis based on a long short-term memory network (LSTM), thereby proposing an LSTM-lsGC model. The model is based on the large-scale Granger causality (lsGC) method, which is an extension of the traditional multivariate Granger causality (mvGC), to construct an overall processing framework. Assuming that x is a stationary multivariate time series, the traditional mvGC estimates by fitting a multivariate vector autoregressive model (MVAR) with a time lag order P:

[0006]

[0007] where X(t) denotes the vector of all time series, A is the coefficient matrix, and E(t) is the uncorrelated noise process. For a particular variable x i (i.e., one component of X), the model excluding x i can be expressed as:

[0008]

[0009] where denotes the time series vector excluding x i , A* is the corresponding coefficient matrix, denotes the noise vector excluding x i . Based on this, the Granger causality from x i to x j can be defined as:

[0010]

[0011] denotes the noise covariance matrix excluding the variable , and Cov(E) denotes the noise covariance matrix of the complete model. This simplified formula is used to test whether the residual covariance significantly increases when the variable x i is excluded from the model, thereby determining whether there is a Granger causal effect of x i on other variables in the time series vector X.

[0012] The traditional multivariate autoregressive (MVAR) method has a limitation: for high-dimensional data, the estimation of model parameters is difficult to achieve. To overcome this problem, according to the claims of the invention, principal component analysis (PCA) is combined with Granger causality analysis. Specifically, the original high-dimensional data is divided into two parts: one part is the target region of interest (ROIs) (x i , x j ) for which we want to explore the causal relationship, and the other part is the remaining ROIs. For the remaining ROIs, we use PCA for dimensionality reduction and convert them into low-dimensional form, and then combine them with the extracted target ROIs. Subsequently, the LSTM model is used for prediction, and the Granger causality relationship is calculated based on this.

[0013] Specifically, the time series obtained from the extracted region of interest (ROIs) are divided into two parts: x i ∈R 1 ×T , x j ∈R 1×T denotes the target ROIs, denotes the remaining ROIs. We perform dimensionality reduction on the remaining ROIs, which can be mathematically expressed as:

[0014]

[0015] where W is a transformation matrix of dimension L x (H-2), and the reduced time series X LD is subsequently merged with x i and x j .

[0016] The long short-term memory network (LSTM) is particularly suitable for brain connectivity estimation due to its ability to effectively model time series with different transmission delays. Its input includes the current time step x i , the previous hidden state h t-1 , and the previous cell state C t-1 . The forget gate and the memory gate control the discard and retention of information, respectively. The new information is calculated and passed to the next cell after being adjusted by the gates. The sigmoid function (σ) compresses the value to the interval [0, 1], and the tanh function maps the value to the range [-1, 1]. To replace the traditional multivariate autoregressive (MVAR), we use the LSTM model to predict the time series:

[0017] X(t) = LSTM(X(t-1)) + e(t)

[0018]

[0019] Finally, the Granger causality value is stored in the (i, j) position of the Granger matrix G. The matrix has a dimension of H x H, and the diagonal elements are zero, so it contains H(H-1) non-zero LSTM-lsGC values.

[0020] Here, LSTM(·) represents the LSTM model, to predict the time series, and e(t) is the estimation error. To calculate the Granger causality value from x i to x j , we construct two groups of inputs: the first group contains x j , x i , and X LD , and the second group contains only x j and X LD , omitting x i . Using these two groups of inputs, the LSTM model generates prediction results and then calculates the residual variance containing x i and x j , and the residual variance not containing x i .

[0021] Finally, the Granger causality value The value stored in the (i, j) position of the Granger matrix G. The matrix dimension is H x H (H is the total number of brain regions), and the diagonal elements are always zero, so it contains H(H-1) nonzero LSTM-lsGC values.

[0022] The network is composed of nodes and edges, and the brain network proposed in this paper also follows this structure: the nodes correspond to the regions of interest (ROIs), and the edges are defined by the connectivity measure. We constructed the brain network for each subject according to the method described above. Before extracting the graph metrics, it is usually necessary to set the edges below the threshold to zero through threshold processing. In this study, we used the sparsity threshold to determine the retention degree of the edges in the individual brain network, and the threshold was defined as the ratio of the number of connections above the threshold to the maximum possible number of connections in the network, to measure the network connection cost. It is worth noting that, due to the lack of a unified standard to set a single sparsity threshold, and different thresholds will lead to different results

[16] , we set the threshold matrix to the range of 0.05 to 0.95 (step 0.05), so that the network is neither too sparse nor too dense. Then, the functional connectivity matrix under each threshold is binarized (non-zero elements are set to 1, and zero elements remain 0). To eliminate the randomness of the threshold, we calculated the area under the curve of each subject's graph metrics as the threshold changes as a feature [38, 39], including global graph metrics (small-worldness, network efficiency (global / local efficiency), assortativity, hierarchy) and local graph metrics (clustering coefficient, shortest path, efficiency, local efficiency, degree centrality, betweenness centrality). A total of 1019 features were obtained in this part, including 5 global features and 1014 local features (6 local metrics x 169 ROIs), which were used for subsequent feature selection.

[0023] Feature selection reduces the number and dimension of features by removing redundant features, and can improve classification performance and reduce overfitting. Feature selection methods are usually divided into three categories: filter, wrapper, and embedded. The minimum redundancy maximum relevance (mRMR) algorithm is a filter-based feature selection method, which selects features based on mutual information, correlation or similarity scores, while considering the correlation between features F and labels L and the redundancy between features. i The correlation between the feature set F and the label L is defined by the average value of the mutual information value between a single feature f

[0024]

[0025] The redundancy of all features in the set F is defined by the average value of the mutual information value between a single feature f i and a feature f j .

[0026]

[0027] The feature selection algorithm used in this study can effectively maintain the strong correlation between features and categories while reducing the number of features.

[0028] The beneficial effects of this invention are:

[0029] This invention effectively overcomes three major limitations of traditional methods: over-reliance on multimodal data, inability to analyze causal interactions in brain regions, and neglect of MCI reversal mechanisms. By pioneering a framework based on LSTM-enhanced Granger causality analysis (LSTM-lsGC) and efficient connectivity computation, it achieves for the first time bidirectional and accurate identification of MCI progression and reversal in a single rs-fMRI modality. This technology fully considers the dynamic nonlinear characteristics and threshold randomness of brain networks, utilizing dimensionality reduction integration and area under the curve (AUC) feature construction methods to ultimately achieve high-dimensional causal feature extraction and the discovery of interpretable brain network biomarkers, providing a new approach for early intervention and cognitive recovery mechanism research in Alzheimer's disease. Detailed Implementation

[0030] The implementation of the present invention will be further described in detail below. It should be understood that the following specific embodiments are only used to illustrate the present invention and do not limit the scope of the present invention.

[0031] All data presented were selected from the publicly available Alzheimer's Disease Neuroimaging Initiative (ADNI) database. Studies involving human subjects complied with ADNI ethical guidelines: the studies followed Good Clinical Practice Guidelines, Section 50 (Protection of Human Subjects) and Section 56 (Institutional Review Board / Research Ethics Committee) of the U.S. Federal Code, and complied with state and federal regulations. All participants and / or their authorized representatives signed written informed consent forms and HIPAA authorizations prior to the start of the study. The primary objective of ADNI was to validate the combined use of serial MRI, positron emission tomography (PET), other biomarkers, and clinical neuropsychological assessments to monitor the progression of MCI and AD. All patients used in this study had complete resting-state functional magnetic resonance imaging (rs-fMRI) data. The diagnostic criteria for MCI in the ADNI project are as follows: (1) Mini-Mental State Examination (MMSE) score between 24 and 30; (2) Clinical Dementia Rating Scale (CDR) score of 0.5; (3) presence of memory complaints and objective memory decline confirmed by the Wechsler Memory Scale Logical Memory II test after education level adjustment; (4) no significant functional impairment in other cognitive domains and normal daily living abilities (not meeting the criteria for dementia). This study included 28 HC, 23 AD, 29 sMCI, and 23 MCI conversions (including 7 rMCI and 16 pMCI). Detailed information is shown in Table 1. MCI conversions occurred within 6-36 months, and sMCI patients did not convert to AD or were normal during the 36-month follow-up period. Subjects who did not undergo MMSE, CDR, or FAQ assessments were excluded. All MCI patients had a CDR score of 0.5, an MMSE score between 24 and 30, an FAQ score between 0 and 20, no significant abnormalities in other cognitive domains, retained daily living abilities, and did not meet the criteria for dementia.

[0032] In accordance with the ADNI protocol, participants underwent rs-fMRI scans using a 3T Philips scanner. During data acquisition, participants were required to remain awake, with their eyes open, and to minimize head movement. Functional MRI images were acquired using an echo-planar imaging (EPI) sequence with the following parameters: repetition time (TR) 3000.0 ms, echo time (TE) 30.0 ms, and flip angle (FA) 80.0°, for a total of 48 slices. The pixel pitch in the X and Y dimensions was 3.3 mm, and the slice thickness was 3.3 mm.

[0033] The fMRI data of patients were preprocessed using DPABI software (http: / / rfmri.org / dpabi) based on the MATLAB 2018b platform

[31] . The preprocessing steps included: removing the first 10 time points of all participants to eliminate the magnetization balance effect; then slice time correction, head motion correction, spatial normalization based on EPI template (resampling to 3mm×3mm×3mm voxels), Gaussian kernel smoothing with a full width at half maximum (FWHM) of 4mm, delinearization, and covariate regression were performed in sequence.

[0034] Template construction is based on low-frequency (0.01-0.08Hz) filtered signals to calculate ALFF values ​​for HC and AD groups without additional filtering. To simplify cross-subject comparison, various standardization methods are used for rs-fMRI indicators: ALFF is converted into z-scores (called zALFF) by subtracting the mean within the gray matter mask and dividing by the standard deviation within the mask. Two-sample t-tests are used to calculate the voxel differences in zALFF between HC and AD groups, and the false discovery rate (FDR) is used for correction. Voxels with significant inter-group differences are defined as differential voxels, and these voxels are then matched with AAL maps

[32] . If a differential voxel overlaps with the brain region corresponding to AAL, it is defined as a new region of interest (ROI), and finally a new template containing 53 ROIs is obtained (named differentialmask). In addition, voxels that overlap with differentialmask are removed from the AAL template to obtain the AAldel template (still retaining 116 ROIs, but the number of voxels is reduced). Finally, by integrating the differential mask and AALdel templates, a differential HC-AD template (HAD) containing 169 ROIs was constructed for this study. Unlike most studies, this study did not exclude cerebellar ROIs—because changes in cerebellar regions under different MCI states are also valuable for research; therefore, all ROIs were explored and analyzed using the HAD template across the entire brain.

[0035] Based on the aforementioned LSTM-based research method, this study divided the three groups of participants into training and test sets in an 8:2 ratio. The feature subsets obtained after feature selection were ranked according to the importance score of the random forest, with higher scores ranking higher. Features were then added to the classification feature set sequentially according to this ranking, and the feature subset corresponding to the highest accuracy was selected. We used a grid search algorithm to determine the optimal parameters and calculated the accuracy of the random forest classifier corresponding to these feature subsets based on the five-fold cross-validation method. Notably, to reduce feature redundancy and computational cost, we implemented an interruption condition: if the accuracy remained unchanged for ten consecutive iterations, the process was terminated early, and the current result was retained. Since an imbalance in the sample size of the groups to be classified may lead to classification bias, we overcame this problem by updating the class weights based on the number of samples in each class. The highest accuracy and its corresponding optimal feature subset in the above classification process were determined through 50 repeated experiments.

[0036] Accuracy, sensitivity, and specificity are used as evaluation metrics to measure classification performance and test its stability and robustness. The calculation formulas are as follows:

[0037]

[0038] The above metrics are calculated separately for each category label (e.g., label 0, label 1, label 2), and finally the average sensitivity and average specificity of all labels are obtained by summing them up. AUC (Area Under the ROC Curve) is calculated as follows: First, both the predicted and true labels are converted into one-hot encoded forms; then, the true positive rate (TPR) and false positive rate (FPR) are calculated for each class at different thresholds, and a multi-class ROC curve is plotted accordingly; finally, the area under the curve is calculated by integration to evaluate the overall classification performance of the model.

[0039] This invention, for the first time, explores effective connectivity and changes in brain network characteristics based on resting-state functional magnetic resonance imaging (rs-fMRI) and achieves the identification of patients with reversible microvascular injury (rMCI), stable MCI (sMCI), and progressive MCI (pMCI). We constructed dementia-specific brain templates using Alzheimer's disease (AD) patients to more accurately extract regions of interest (ROI) signals closely related to MCI. In brain network construction, a novel effective connectivity method—LSTM-lsGC—was proposed and validated. In the classification framework, graph theory and machine learning were combined for feature selection, and the data imbalance problem was addressed by updating class weights, ultimately achieving effective classification. The proposed framework can distinguish the three states of MCI well, with an overall accuracy of 84.92% and an AUC of 0.84. The results show that the LSTM-lsGC and graph theory fusion framework based on rs-fMRI can effectively predict the transformation trend of MCI. The discovery of key brain regions such as the precentral gyrus and hippocampus may serve as diagnostic biomarkers for predicting future changes in MCI patients, providing a new approach for the early detection and differential diagnosis of AD.

Claims

1. An effective connectivity (EC) framework based on long short-term memory network (LSTM) enhanced Granger causality method for classifying MCI subtypes and identifying key brain regions related to disease progression, which mainly comprises three core links of template construction, effective connectivity calculation based on LSTM-lsGC model, feature selection and classification, and the specific process is as follows: 1) HAD template: by comparing the ALFF indicators of the healthy control group (HC) and the Alzheimer's disease group (AD), the voxels with significant differences are obtained and the HAD template is constructed; 2) effective connectivity matrix: the Granger causality is calculated by applying the LSTM-lsGC model; 3) features and classification: the indicators extracted from the global and local features are used to identify the three states of MCI. In addition, this study also respectively adopts the automatic anatomical marker template (AAL) and the pair-wise conditional Granger causality model (PGC) in the MVGC toolbox to construct the brain effective connectivity network for comparative experiments, so as to verify the effectiveness and superiority of the proposed method. All the data are selected from the public Alzheimer's Disease Neuroimaging Initiative (ADNI) database, all the data are selected from the public Alzheimer's Disease Neuroimaging Initiative (ADNI) database, and the present application includes 28 HC, 23 AD, 29 sMCI and 23 MCI converters (including 7 rMCI and 16 pMCI) in total. The MCI converters have state transition within 6-36 months, and the sMCI patients have not been converted into AD or normal within 36 months of follow-up period. The subjects who have not been evaluated by MMSE, CDR or FAQ are excluded, all the MCI patients have CDR score of 0.5, MMSE score of 24-30, FAQ score of 0-20, no significant abnormality in other cognitive fields, daily life ability reservation and no dementia standard. The participants in the present application use a 3T Philips scanner for rs-fMRI scanning. The participants are required to keep awake and open-eyed state and reduce head motion as much as possible during data acquisition. The functional magnetic resonance image is acquired by using the echo planar imaging (EPI) sequence, and the parameters are as follows: repetition time (TR) 3000.0 ms, echo time (TE) 30.0 ms, flip angle (FA) 80.0°, and a total of 48 layers are acquired. The pixel spacing in X and Y dimensions is 3.3 mm, and the layer thickness is 3.3 mm. The template construction is based on the calculation of the ALFF values of the HC and AD groups from the low-frequency band (0.01-0.08 Hz) filtered signal, without additional filtering. The main contribution of the present application is to enhance the Granger causality analysis method for high-dimensional feature utilization: we introduce principal component analysis in the long short-term memory network (LSTM) based Granger causality analysis, thereby proposing an LSTM-lsGC model. The model is based on the large-scale Granger causality (lsGC) as an extension method of the traditional multivariate Granger causality (mvGC) to construct the overall processing framework. Assuming that x is a stationary multivariate time series, the traditional mvGC estimates by fitting a multivariate vector autoregressive model (MVAR) with a time lag order P: wherein, X(t) denotes the vector of all time series, A is the matrix of coefficients, and E(t) is the uncorrelated noise process. For a particular variable x i (i.e. one component of X), the model excluding this variable x i can be written as: A limitation of the traditional multivariate autoregressive (MVAR) approach is that for high-dimensional data, the model parameter estimation is difficult to achieve. To overcome this problem, principal component analysis (PCA) is combined with Granger causality analysis according to claim 1. Specifically, the original high-dimensional data is divided into two parts: one part is the region of interest (ROIs) (x i , x j ) for which we want to explore the causal relationship, and the other part is the rest of the ROIs. For the rest of the ROIs, we use PCA for dimension reduction, transforming them into a low-dimensional form, and then merging them with the extracted target ROIs. Subsequently, the LSTM model is used for prediction, and on this basis, the Granger causality relationship is calculated. The sued long short-term memory network (LSTM) is particularly suitable for brain connectivity estimation due to its ability to effectively model time series with different transmission delays. Its input includes the current time step x t , the previous hidden state h t-1 , and the previous cell state C t-1 . The forget gate and the memory gate control the discard and retention of information, calculate the new information after gating adjustment and transmit it to the next cell. The sigmoid function (sigma) compresses the value to the interval [0, 1], and the tanh function maps the value to the range [-1, 1]. To replace the traditional multivariate autoregressive (MVAR), we use the LSTM model to predict the time series: X(t) = LSTM(X(t-1)) + e(t) Finally, the Granger causality values are stored in the (i, j) position of the Granger matrix G. This matrix has dimension H x H, with diagonal elements being zero, thus containing H(H - 1) non-zero LSTM-lsGC values. According to the effective connectivity (EC) framework of a long short-term memory network (LSTM) enhanced Granger causality method according to claim 1, an individual brain network is constructed based on nodes (ROIs) and edges (connectivity indicators), a sparsity threshold value in the range of 0.05-0.95 (step length 0.05) is adopted to perform binary processing on the connection matrix, and the area under the curve of global and local graph indicators is extracted as features, a total of 1019 features are obtained. Subsequently, the mRMR algorithm is used for feature selection, the correlation between features and labels and the redundancy between features are measured to reduce the dimension and improve the classification performance.