Method for predicting AD conversion probability through deep learning method based on attention

By applying a deep learning method that imparts weights to features and time points in AD prediction, the problem of ignoring the importance of cross-modal interaction and modality in the prior art is solved, and more efficient early AD diagnostic prediction is achieved.

CN119993535AActive Publication Date: 2025-05-13CHONGQING UNIV OF POSTS & TELECOMM
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510086087.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-13
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

When using multimodal timing data to predict Alzheimer's disease (AD) progression, the prior art ignores the importance of cross-modal interactions and the modality in different tasks, resulting in poor model effectiveness.

Method used

The attention-based deep learning method is adopted to give weights to different features and time points through the attention mechanism, capture the relationship between features and time points, determine their importance, and then predict the AD conversion probability.

Benefits of technology

It improves the accuracy of early diagnostic prediction of AD and the generalization ability of the model, reduces the computational complexity and cost, and provides better application prospects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993535A_ABST
    Figure CN119993535A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data mining, and particularly relates to a method for predicting AD conversion probability by an attention-based deep learning method, which comprises the following steps: acquiring various examination data of a patient to be predicted, and dividing the data into follow-up data and baseline data according to follow-up features and baseline features; pre-processing follow-up data, and inputting the pre-processed follow-up data into the constructed base model to obtain follow-up data representation; and inputting the follow-up data representation and the baseline data into a final decision maker to obtain the probability that the next access of the to-be-predicted patient is converted from MCI to AD. On the basis of the principle of an attention mechanism, an attention variant form is designed for a prediction scene of current time sequence data, and weights are given to different features and time points, so that the relationship between the features and the time points is captured, and the importance of the different features and the time points is determined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data mining, and in particular relates to a method for predicting AD conversion probability based on an attention-based deep learning method. Background Art

[0002] Alzheimer's disease (AD) is a neurodegenerative disease that is closely related to age. It is estimated that the prevalence rate is about 26.4% in people aged 65-74, 38.6% in people aged 75-84, and 35.4% in people aged 85 and above. Scientific research points out that the main pathogenic factors of AD are senile plaques formed by the deposition of β-amyloid protein (Aβ) in the brain and neurofibrillary tangles caused by hyperphosphorylation of Tau protein. The accumulation of these abnormal proteins interferes with the normal communication of the nervous system, thereby causing a gradual decline in cognitive function. Mild cognitive impairment (MCI) is considered to be a potential early manifestation of AD. Although only about 10%-12% of MCI patients eventually develop AD each year, it affects about 22% of the elderly over 65 years old, of which about half of MCI cases are caused by AD. In addition, about 22% of people aged 65 or above are in a preclinical AD state, that is, their cognitive functions appear normal, but there are obvious signs of AD pathology. Especially among the elderly population aged 90 and above, this proportion is as high as 50%.

[0003] Given the irreversibility of AD and the high cost of care, it is of great significance to use deep learning methods to accurately predict the progression of Alzheimer's disease in the early stages.

[0004] In recent years, many studies have used artificial intelligence methods to study AD, but a single modality can only provide information on a specific aspect, and some modalities may not be sensitive enough for early AD detection, and baseline non-time series data cannot predict progression. Therefore, using multimodal time series data to predict the progression of Alzheimer's disease has many advantages and can make up for the problems of single modality and non-time series data. Multimodal fusion methods can be roughly divided into three categories: early fusion, late fusion, and mid-term fusion.

[0005] Early fusion, also known as feature-level fusion, refers to the fusion of data from different modalities in the early stages of the model before putting them into the model for training. Its advantage is that it can capture low-level correlation information between different modalities, but its disadvantage is also obvious. The fusion features generated after the transformation and scaling of each modality feature usually have higher dimensions, which will increase the complexity and computational cost of the model. In 2022, Xu L et al. proposed an early fusion method to predict the progression of Alzheimer's disease with multimodal time series data, processed different modal data into potential learning representations, then filled the data, and finally predicted, and obtained better prediction results. In the same year, El-Sappagh S et al. proposed a two-stage deep learning model to predict the progression of Alzheimer's disease, using neuroimaging data, CS, CSF biomarkers, neuropsychological battery markers and demographic data to explore the impact of different time series on the results. In the first stage, multi-classification predictions were performed for CN, MCI and AD, and in the second stage, the exact conversion of MCI to AD was predicted as a regression task, which obtained better results than the comparison model. Deep learning models generally lack interpretability. The interpretability of models in medical scenarios is a relatively important issue. This study uses the attention mechanism to increase the interpretability of the model.

[0006] Late fusion, also known as decision-level fusion, is the fusion of prediction results of different modalities in the late stage of the model. Its advantages are that each modality is processed independently, model training is simple, and it is easy to integrate. It can simply handle the asynchrony of data, and the entire system can be expanded as the number of modalities increases. The exclusive prediction model of each modality can better model the modality, and predictions can also be made when the model input lacks certain modalities. The disadvantage is that it may not be able to capture the interactive information between different modalities. Huang SC et al. found that when using CT scans and EHR data to predict pulmonary embolism detection, the average-based late fusion with regularized DNN as a submodel is better than early, mid-term and other late fusion strategies. In the field of Alzheimer's disease prediction, Qiu S et al. found that MRI data and other non-imaging data can be effectively combined to predict AD progression. There is still limited work in this area. Using the NACC dataset, MRI scans, MMSE (Mini-Mental State Examination) and LM (Logical Memory) data were used to predict Alzheimer's disease progression using different algorithms, and then the results were integrated using the majority voting method. The results are better than individual models.

[0007] Mid-term fusion fuses features of different modalities at the intermediate level. Usually, the method of feature interaction and fusion is adopted in the intermediate layer of the model, such as feature combination through attention mechanism or shared network layer. Its advantage is that it has advantages in capturing intermediate correlation information between different modalities, and can better balance the advantages and disadvantages of early fusion and late fusion. The disadvantage is that the implementation is more complicated and requires the design of a reasonable fusion mechanism. Lee G et al. proposed a mid-term fusion method to predict Alzheimer's disease. GRU was used for different modalities to predict the conversion of AD, and then logistic regression was used to integrate the results, which was better than the baseline.

[0008] In summary, the current limitation of multimodal fusion is that it ignores cross-modal interactions and the importance of various modalities in different tasks. In addition, time series data will also ignore the importance of different time points, leading to suboptimal model results. Summary of the invention

[0009] In order to solve the above technical problems, the present invention provides a method for predicting AD conversion probability based on a deep learning method of attention, comprising:

[0010] S1: Obtain various examination data of the patient to be predicted, and divide the data into follow-up data and baseline data according to follow-up characteristics and baseline characteristics;

[0011] S2: After preprocessing the follow-up data, it is input into the constructed base model to obtain the follow-up data representation;

[0012] S3: The follow-up data is combined with the baseline data and input into the final decision maker to obtain the probability of the patient to be predicted to convert from MCI to AD at the next visit.

[0013] Beneficial effects of the present invention:

[0014] The present invention is based on the principle of attention mechanism. Aiming at the current prediction scenario of time series data, we have designed a variant form of attention, which can be regarded as a kind of local attention. Weights are assigned to different features and time points to capture the relationship between features and time points, and determine the importance of different features and time points. For small sample data, such simplified operations can reduce the complexity of calculations, and there is no need to calculate the attention score at each forward propagation, saving computing costs. The present invention is implemented on the ADNI dataset, which proves that the present invention is superior to the traditional single machine learning model in terms of the accuracy of the prediction results of early diagnosis of AD and the generalization ability of the model. The present invention has high prediction accuracy and high practicality. Through the early diagnosis of Alzheimer's patients, it helps to make relevant preparations and preventive measures in time to slow down the deterioration of the disease and increase the probability of patient survival. It has good application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 A schematic diagram of the structure of the AD conversion probability model predicted by the attention-based deep learning method in the present invention;

[0016] Figure 2 This is a schematic diagram of data preprocessing in the present invention;

[0017] Figure 3 It is a schematic diagram of the accuracy of different scenarios in the invention;

[0018] Figure 4 Schematic diagram of F2 scores in different scenarios in the present invention. DETAILED DESCRIPTION

[0019] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0020] The present invention proposes a method for predicting AD conversion probability based on a deep learning method based on attention, such as Figure 1 As shown, the method includes: obtaining various examination data of the patient to be predicted, dividing the data according to follow-up characteristics and baseline characteristics, performing a series of preprocessing on the follow-up data, and inputting the data into the constructed base model to obtain the follow-up data representation, and then combining the baseline data with the final decision maker to obtain the probability of the patient to be predicted converting from MCI to AD in the next visit;

[0021] The process of building a model for predicting AD conversion probability using an attention-based deep learning approach includes:

[0022] S1: Obtain the patient's examination data and divide them according to follow-up characteristics and baseline characteristics.

[0023] Select a data set that meets the actual production environment to complete the AD risk probability prediction. We selected the ADNI database. The Alzheimer's Disease Neuroimaging Program (ADNI) is a longitudinal multicenter study that aims to provide a large amount of clinical, biomarker and imaging data across multiple institutions and research organizations to promote research on early diagnosis, tracking and treatment of Alzheimer's disease. The database includes structural and functional magnetic resonance imaging (MRI and fMRI), positron emission tomography (PET), cerebrospinal fluid biomarkers, demographic data, and clinical and cognitive assessments collected from patients with Alzheimer's disease (AD), mild cognitive impairment (MCI), normal elderly (CN) and other related diseases.

[0024] The present invention has collected demographic data, cognitive scores, cerebrospinal fluid and MRI samples of 2419 subjects from three subgroups of ADNI-1, ADNI-2 and ADNI-3, where the characteristics of the demographic data are baseline characteristics (the longitudinal changes are meaningless), and the rest are follow-up characteristics. According to the label division of ADNI, the clinical status of the subjects at a certain time point can be divided into three categories, namely normal cognition (CN), mild cognitive impairment (MCI) and Alzheimer's disease (AD). The ADNI data set is sampled once every six months, and the time of the initial examination is set as the first time, the second time after half a year, the third time after one year, and so on. A total of the first six data are used, and Table 1 shows the statistics of the selected data at baseline.

[0025] Table 1 Baseline statistics

[0026]

[0027] S2: Preprocess the data set to obtain a structured and complete data set.

[0028] The follow-up characteristics and baseline characteristics are divided into two tables and saved. The data table of follow-up characteristics is processed by the following series (such as Figure 2 ), there were 11,096 visits for 1,597 patients remaining. Multiple imputation was used to fill in the missing data. The baseline characteristics of the data table were retained for patients with the same RID as the processed patient data. The data were randomly divided into a training set and a test set with a ratio of 7:3 to predict the conversion of MCI to AD.

[0029] S3: Input the processed follow-up data into the constructed follow-up data model to obtain data representation.

[0030] Based on the principle of the attention mechanism, we designed a variant of attention for the current time series data prediction scenario, which can be regarded as a local attention, giving weights to different features and time points, thereby capturing the relationship between features and time points, and determining the importance of different features and time points. For small sample data, such a simplified operation can reduce the complexity of the calculation, and there is no need to calculate the attention score in each forward propagation, saving computing costs.

[0031] In the custom attention layer, as trainable parameters, they are learned through backpropagation during training. These weights are initialized to 1, which means that in the initial state, they have no scaling effect on the input data, so that the model starts training from an unbiased state and all inputs are treated equally. After forward propagation and backpropagation, these weights will be scaled to varying degrees according to the learned values ​​as training progresses. The optimizer (Adam) is used to update the loss gradients, which indicate the amount and direction that each weight needs to change to minimize the loss. The resulting weighted input is then fed into the subsequent model for training.

[0032]

[0033] in represents the weighted input, x t represents input, α represents feature weight, and β represents time weight.

[0034] Recurrent Neural Network (RNN) is designed to process time series data. It effectively captures the relationship characteristics between sequences through the internal structural design of the network. It has a simple internal structure and performs well on short sequence tasks. However, it performs poorly on long sequence tasks, and there will be gradient vanishing or explosion problems. In order to solve this problem, Long Short-Term Memory Network (LSTM) and Gated Recurrent Unit (GRU) were designed and found to be effective in many places. There are many advantages to choosing GRU as the main component. First, it has fewer parameters and can speed up training. LSTM requires too many parameters, the training complexity is high, and there is also a risk of overfitting. The amount of data input to the model is small, and GRU is easier to train with a small number of samples. After testing, it was found that GRU did have better performance, so CRU was selected as the RNN training unit at the beginning.

[0035] GRU consists of two parts, the update gate and the reset gate. The calculation method is to use x(t) and h(t-1) to concatenate for linear transformation, and then activate it through the sigmoid function. After that, the update gate value acts on h(t-1), which controls how much information from the previous time step can be used. Then use this updated h(t-1) to perform basic RNN operations, that is, concatenate it with x(t) for linear transformation, and activate it through tanh to get a new h(t). Finally, the gate value of the reset gate will act on the new h(t), and (1-gate value) will act on h(t-1). Then the results of the two calculations are added together to get the final implicit state output h(t). This process means that the reset gate has the ability to reset all previous calculations.

[0036] z t =σ(W z ·[ht-1 ,x t ])

[0037] r t =σ(W r ·[h t-1 ,x t ])

[0038]

[0039] In this process, we use time series learning methods such as GRU in combination with the attention mechanism to learn the feature representation of the model based on the attention layer, and use the obtained feature representation combined with the baseline data as the input of the final decision maker.

[0040] S4: Use multilayer perceptron as the final decision maker to predict transition probabilities

[0041] For the final prediction, a multilayer perceptron was used to predict the conversion from MCI to AD. The information learned from the GRU combined with attention was concatenated with the cross-sectional (baseline) data and input into a multilayer perceptron to predict the probability of conversion from MCI to AD at the next visit.

[0042]

[0043] Where y′ represents the predicted output, W1 and W2 are trainable linear transformation matrices, and c t represents the context vector of the RNN unit after the attention layer, D represents the baseline data, and b1 and b2 represent the bias vectors.

[0044] S5: Update the parameters of each module according to the loss function to obtain the AD conversion probability model predicted by the attention-based deep learning method.

[0045] The trainable parameters of the RNN unit and the multilayer perceptron of the present invention are learned by using the binary cross entropy function, and the loss function is used to reduce the false negative medical records as much as possible and improve the sensitivity of the model.

[0046]

[0047] In Equation 7, α is a weight set from 0 to 1 for prediction, y is the true quasi-segment, and y′ is the predicted diagnostic. Based on the proposed custom loss function, all trainable parameters of RNN, RNN autoencoder, and MLP are updated when training the model in backpropagation. All model weights are trained using the Adaptive Moment Estimation (Adam) optimizer, a commonly used optimization algorithm that can adaptively adjust the learning rate of each parameter and usually converge faster than the classic stochastic gradient descent algorithm. The learning rate is set to 0.001. RNN units, batch size, dropout rate, and L2 regularization are tuned hyperparameters.

[0048] S8: Use the attention-based deep learning method to predict the AD conversion probability model to predict the patient data to obtain the probability of the patient converting from MCI to AD during the next visit.

[0049] In order to verify the effectiveness of the present invention, a comparative experiment was conducted on other traditional machine learning algorithms. The present invention selected prediction accuracy and F2 as evaluation parameters:

[0050] 1) Accuracy — ACC

[0051] The accuracy rate indicates the proportion of all correctly predicted samples to the total samples. The calculation formula for the binary classification problem is:

[0052]

[0053] Among them, TP represents the number of positive samples predicted correctly, TN represents the number of negative samples predicted correctly, FP represents the number of positive samples predicted incorrectly, and FN represents the number of negative samples predicted incorrectly.

[0054] 2) F2 score

[0055] The F2 score is an indicator that balances precision and recall. However, unlike the F1 score, the F2 score places more emphasis on recall. The specific calculation formula is as follows:

[0056]

[0057] Where β is a weight parameter. For the F2 score, β = 2, which means that recall is twice as important as precision.

[0058] In the experiment, we used longitudinal multimodal data from the ADNI dataset and baseline demographic data (these features are meaningless when accumulated over time) to train and test the proposed model. The training set and test set accounted for 70% and 30% respectively. When setting up the experiment, the probability of whether the next side will switch was predicted for the first two, three, four, five and six times. The comparative experiment was set up with support vector machine (SVM), random forest (RF) and decision tree (DT). Since these three algorithms cannot process longitudinal data, they were trained using the same data as the RNN unit to obtain the prediction results.

[0059] Figure 3 , Figure 4 The accuracy of the experiment and the comprehensive index F2 score are displayed respectively. It can be seen that these two indicators of the present invention are better than the selected baseline model. The average accuracy of five predictions can reach 92.2%, and the average F2 index is 84.8%, which is the best performance among all models and the results are highly competitive.

[0060] In view of the shortcomings of existing early diagnosis and prediction methods for Alzheimer's disease, this paper predicts the AD conversion probability model based on the deep learning method of attention. The main contributions are as follows:

[0061] 1. Using the idea of ​​attention mechanism, weights are assigned to features and time points respectively. The weights obtained through model training are used in training to weight the input data and obtain the importance of different features and time points. Experiments also prove the effectiveness and superiority of the proposed attention layer.

[0062] 2. A method for predicting the probability of AD conversion based on multimodal data fusion based on attention is proposed. It combines baseline characteristics and follow-up characteristics, considers the complexity of data in the actual production environment, and combines data of multiple modalities to provide decision support for doctors. The experimental results prove the effectiveness and superiority of the present invention.

[0063] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for predicting AD conversion probability based on a deep learning method based on attention, characterized in that: include: S1: Obtain various examination data of the patient to be predicted, and divide the data into follow-up data and baseline data according to follow-up characteristics and baseline characteristics; S2: After preprocessing the follow-up data, it is input into the constructed base model to obtain the follow-up data representation; S3: The follow-up data is combined with the baseline data and input into the final decision maker to obtain the probability of the patient to be predicted to convert from MCI to AD at the next visit.

2. The method for predicting AD conversion probability based on an attention-based deep learning method according to claim 1, characterized in that: Obtain various examination data of the patient to be predicted, and divide the data into follow-up data and baseline data according to follow-up characteristics and baseline characteristics, including: Obtain various examination data of the patients to be predicted from the ADNI database, including demographic data, cognitive scores, cerebrospinal fluid and MRI samples from three subgroups, ADNI-1, ADNI-2 and ADNI-3; the characteristics of the demographic data are baseline characteristics, and their longitudinal changes are meaningless, and the rest are follow-up characteristics; The ADNI database includes structural and functional magnetic resonance imaging (MRI) and fMRI, positron emission tomography (PET), cerebrospinal fluid biomarkers, demographic data, and clinical and cognitive assessment data collected from patients with Alzheimer's disease (AD), mild cognitive impairment (MCI), normal elderly people (CN), and patients with other related diseases.

3. The method for predicting AD conversion probability based on an attention-based deep learning method according to claim 1, characterized in that: The base model includes: a dual attention layer and a temporal model; The dual attention layer is used to obtain the feature importance and time step importance of the follow-up data, and express them with importance weights to obtain weighted output; The time series model adopts a GRU network, and the weighted output is passed through the GRU network to obtain the feature representation of the follow-up data.

4. The method for predicting AD conversion probability based on an attention-based deep learning method according to claim 3, characterized in that: The dual attention layer is used to obtain the feature importance and time step importance of the follow-up data, and express them with importance weights to obtain weighted outputs, including: in, represents the weighted output, x t represents follow-up data, α represents feature weight, and β represents time weight.

5. The method for predicting AD conversion probability based on an attention-based deep learning method according to claim 1, characterized in that: The follow-up data were preprocessed, including: deleting features that were not related to the follow-up data, excluding data of patients who had one visit or were diagnosed with CN, deleting data of patients with APOE4 deletion, and deleting features with a missing rate of >60%.

6. The method for predicting AD conversion probability based on an attention-based deep learning method according to claim 1, characterized in that: The follow-up data is combined with the baseline data and input into the final decision maker to obtain the probability of the patient to be predicted to convert from MCI to AD at the next visit, including: Among them, y′ represents the predicted output, W1 and W2 are trainable linear transformation matrices, and c t represents the feature representation of the follow-up data, D represents the baseline data, and b1 and b2 represent the deviation vectors.

Citation Information

Patent Citations

  • Depression identification method based on facial video, depression identification device based on facial video, and storage medium

    CA3166026A1

  • ICU patient death risk prediction method based on comparative learning

    CN115148359A

  • Brain magnetic resonance image classification system based on attention guidance 3D neural network

    CN115908923A

  • Intercity travel OD demand prediction model training method, prediction method and system

    CN117076922A

  • Cancer risk prediction method based on deep learning

    CN118507048A