A brain region time series position encoding method for predicting alzheimer's disease type
Patent Information
- Application Number
- CN202410083919.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-19
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-01-19
AI Technical Summary
[0008]现有的方法很少有效利用脑区时间序列中的位置信息,同时也很少有方法充分考虑fMRI数据复杂的长程相互作用
[0049]本发明的有益效果是:本发明利用每个样本的fMRI数据,然后从fMRI数据中提取多个脑区的时间序列;接下来通过位置编码技术挖掘脑区时间序列中的时间位置信息得到新的时间序列;然后将新的时间序列输入到由Transformer组成的多头注意力模型中进行数据长程信息挖掘;最后将更新的脑区特征输入到全连接分类器来实现样本疾病类型的预测。本发明的实验结果与传统方法相比,本发明提出的方法能够提高样本疾病类型预测的性能,证明提取脑区时间序列中的时间位置信息并挖掘数据中的长程关联能提升预测样本疾病类型的准确性。
Smart Images

Figure CN117912676B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for encoding the temporal location of brain regions for predicting Alzheimer's disease types, belonging to the field of systems biology. Background Technology
[0002] Alzheimer's disease (AD) is a chronic neurodegenerative disease and a form of dementia. It causes severe impairment in memory, thinking ability, motor skills, and language. This disease not only has a huge impact on patients but also places a burden on their families and society. Since there is currently no effective cure, accurate diagnosis of Alzheimer's disease and identification of the corresponding brain regions and genes are crucial for the prevention and treatment of AD.
[0003] Functional magnetic resonance imaging (fMRI) is an imaging technique used to measure brain activity, capturing information about neural activity by monitoring changes in blood oxygenation levels. fMRI can provide valuable information for diagnosing Alzheimer's disease. For example, fMRI can reveal changes in brain activity, particularly in areas related to memory, cognition, and executive functions. In Alzheimer's disease, these areas may show abnormal activity patterns or weakened functional connectivity. Furthermore, fMRI can help study the brain's connectivity networks, leading to a better understanding of changes in neural networks in Alzheimer's patients. Abnormal connections in these networks may be an early marker of Alzheimer's disease. Finally, fMRI may aid in the early diagnosis of Alzheimer's disease because it can detect neural changes that occur before symptoms appear.
[0004] The application of deep learning in fMRI data analysis offers several potential advantages for the diagnosis of Alzheimer's disease. Deep learning techniques can be used to identify complex patterns and features, thereby improving the accuracy of Alzheimer's diagnosis. These models can learn abstract features extracted from fMRI data, thus differentiating patients from healthy controls. Furthermore, deep learning models can build predictive models that predict an individual's risk of Alzheimer's disease based on brain activity patterns and connectivity. This facilitates early intervention and the development of personalized treatment plans.
[0005] Currently, many deep learning methods have emerged that utilize fMRI data to predict Alzheimer's disease types. In 2012, Dai et al. proposed a method that classifies AD and healthy subjects based on fMRI data using low-frequency amplitude wave (ALFF), region homogeneity (ReHo), and functional connectivity features. In 2014, Jie et al. proposed a classification technique for mild cognitive impairment (MCI), first extracting global topology and local connectivity as graph features, then using minimum absolute contraction for feature selection, and finally classifying using a multi-kernel SVM. Wang et al. used AdaBoost to classify AD and MCI, and compared the classification with normal subjects using an fMRI dataset. First, correlation coefficient feature vectors were extracted using the regions of interest (ROIs) of each subject; to minimize noise effects, linear discriminant analysis (LDA) was used to normalize these vectors. The final classification accuracy obtained using the proposed method was 75.8%. Fei et al. proposed an fMRI-based AD image classification model, which is based on an improved 3DPCANet model and classical correlation analysis (CCA). These preprocessed data were classified using support vector machines, achieving accuracies of 95.00%, 92.00%, and 91.30% for different stages of Alzheimer's disease (AD).
[0006] However, fMRI data contains time-series information that can capture changes in brain activity and help understand the evolution of brain function over time. The methods described above have limited consideration of the temporal information in fMRI data. Researchers typically construct dynamic functional connectivity (dFC) to capture the temporal changes in functional connectivity during fMRI acquisition. The input data for sliding window analysis is a set of time-series data representing the activity of brain regions. A time window is selected to divide the time series into different subsets, and then the correlation between each pair of time series is calculated as a correlation coefficient. The window is then moved by a certain step size, and the same calculation is repeated within time intervals. This process is repeated until the window crosses the end of the time series. For example, Kong et al. constructed a series of functional connectivity networks over time periods by applying sliding window processing to fMRI time-series data, and then used spatial graph attention convolution (SGAC) to learn the topology and temporal features. This method achieved a classification accuracy of over 84% on a major depressive disorder (MDD) classification task. However, the traditional dFC method based on sliding window correlation has several limitations, including increased data dimensionality, potential redundant information, and sensitivity to window parameters.
[0007] Furthermore, varying degrees of long-distance correlation exist between brain regions, influencing human cognition and behavior. Research indicates strong correlations between brain regions that are physically far apart. For example, gray matter volume in the hippocampus is significantly correlated with gray matter volume in other long-distance regions involved in the memory system, such as the amygdala and parahippocampal cortex, entorhinal cortex, and orbitofrontal cortex.
[0008] Existing methods rarely effectively utilize locational information in brain region time series, and even fewer fully consider the complex long-range interactions in fMRI data. Therefore, there is a need to design algorithms that can improve the prediction accuracy of disease types by mining temporal information in brain region time series and uncovering long-range associations across multiple datasets. Summary of the Invention
[0009] The technical problem to be solved by this invention is to provide a brain region time-series location encoding method for predicting Alzheimer's disease types. By mining the time location features in the time series and fusing long-range correlations in the data, the prediction accuracy of the disease type of the sample is further improved, thereby solving the above-mentioned technical problem.
[0010] The technical solution of this invention is: a method for encoding the temporal sequence location of brain regions for predicting Alzheimer's disease types, comprising:
[0011] Step 1: The functional magnetic resonance imaging (fMRI) was preprocessed using the DPARSF tool and the brain region time series was extracted. The brain regions were divided into multiple brain regions using the AAL-116 template.
[0012] Step 2: Integrate the relative and absolute temporal location information of brain region time series using an absolute location encoding function and a learnable location information matrix.
[0013] Step 3: Input the brain region time series containing location information obtained in Step 2 into the Transformer multi-head attention model to update the brain region time series features.
[0014] Step 4: Input the brain region time series features into a fully connected classifier to obtain the disease type prediction probability of the sample.
[0015] Step 1 specifically refers to:
[0016] Step 1.1: Convert the DICOM format fMRI data to the NIFTI format that DPARSF can recognize.
[0017] Step 1.2: Delete the first 10 time points of data for each sample.
[0018] Step 1.3: Align the brain scans to the same time point.
[0019] Step 1.4: Correct the image position offset caused by head movement during the scanning process to the correct position.
[0020] Step 1.5: Standardize brain imaging data for all subjects to match echo-planar imaging EPI templates.
[0021] Step 1.6: Perform Gaussian smoothing on the image with smoothing parameters FWHM of [6, 6, 6].
[0022] Step 1.7: Use a bandpass filter to preserve frequencies in the range of 0.01 to 0.08 Hz.
[0023] Step 1.8: Remove covariates from the image.
[0024] Step 1.9: Use the AAL-116 template to divide the brain into multiple brain regions, and calculate the average value of the blood oxygen level change signal of all voxels in each brain region to obtain the average time series of each brain region.
[0025] Step 2 specifically includes:
[0026] Step 2.1: Trim the brain region time series of all samples after preprocessing and extraction to the same length.
[0027] Step 2.2: Use Represents the time series of brain regions. This is a brain region time series with added time point location information, where N1 represents the number of brain regions and d represents the number of time points after pruning.
[0028] X b_pos Calculated using the following formula:
[0029] X b_pos =X b +PE(X b )
[0030] Where PE(·) represents the position encoding function, and its specific implementation equation is shown below:
[0031] PE(X b )=F(X b )+W b
[0032] Where F(·) is the absolute position encoding function, It has X b The specific implementation equation for the learnable temporal position embeddings of the same dimension, F(·), is shown below:
[0033] F(pos,2i)=sin(pos / 10000 2i / L )
[0034] F(pos,2i+1)=cos(pos / 10000 2i / L )
[0035] Where pos∈[0,d-1] represents X b The time point position is given by L = N1, which represents the number of brain regions, and i ∈ [0, L / 2) which represents the i-th component of the time series vector at position pos.
[0036] Step 3 specifically refers to:
[0037] X b_pos The input is fed into the multi-head attention model, and the specific implementation equation is shown below:
[0038] H=MA(Q,K,V)=norm(concat(SA1,SA2,…,SA h W)
[0039] in, The feature matrix is updated using multi-head attention, where c1 represents the feature dimension of the multi-head attention output, MA(·) represents the multi-head attention mechanism, h is the number of heads, W is the learnable parameter matrix, concat(·) represents the connection operation, norm(·) represents the normalization function, and SA represents single-head attention. The specific implementation equation is shown below:
[0040]
[0041] Q, K, and V represent query, key, and value, respectively. It is the scaling factor for dot product attention; the initial values of Q, K, and V are all brain region time series X. b_pos .
[0042] Step 4 specifically refers to:
[0043] The time series feature H is input into a fully connected classifier to obtain the disease type prediction probability of the sample. The specific implementation equation is shown below:
[0044] P = MLP(H)
[0045] Where MLP(·) represents a multilayer fully connected classifier, and P represents the predicted probability of a sample.
[0046] Finally, the binary cross-entropy loss is used as the loss function for model training, and the loss function is implemented by the following formula:
[0047]
[0048] Among them, y i p represents the true label of sample i. i is the predicted probability of sample i, and N is the number of samples.
[0049] The beneficial effects of this invention are as follows: This invention utilizes the fMRI data of each sample and then extracts time series data from multiple brain regions from the fMRI data; next, it uses positional encoding technology to mine the temporal location information in the brain region time series to obtain new time series; then, the new time series is input into a multi-head attention model composed of Transformers for long-range information mining; finally, the updated brain region features are input into a fully connected classifier to predict the disease type of the sample. Compared with traditional methods, the experimental results of this invention show that the method proposed in this invention can improve the performance of disease type prediction, demonstrating that extracting temporal location information from brain region time series and mining long-range correlations in the data can improve the accuracy of predicting the disease type of the sample. Attached Figure Description
[0050] Figure 1 This is a flowchart of the LEformer of the present invention;
[0051] Figure 2 This is a structural diagram of the position encoding algorithm used by the LEformer of this invention;
[0052] Figure 3 This is a diagram of the algorithm structure used by LEformer in this invention to fuse long-range correlations between multimodal data. Detailed Implementation
[0053] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0054] Example 1: As Figure 1-3 As shown, a method for encoding the temporal location of brain regions for predicting Alzheimer's disease types includes the following steps:
[0055] Step 1: Use the DPARSF tool to preprocess functional magnetic resonance imaging (fMRI) and extract brain region time series, where the brain regions are divided into multiple brain regions using the AAL-116 template.
[0056] DPARSF was used to process fMRI data and extract time series of regions of interest (ROIs) in the brain. First, the raw fMRI data was converted to a suitable format for further processing and analysis. Then, the first 10 time points in each sample's fMRI data were removed to help stabilize subsequent studies, as these data are often associated with transient effects. Subsequently, temporal effects in the data were corrected, and motion artifacts in the fMRI data caused by subject movement during the scan were corrected. Furthermore, the fMRI data was compared with a standard coordinate EPI template for comparison between subjects, and Gaussian smoothing with a full width at half maximum (FWHM) of [6, 6, 6] was applied to the fMRI data. Finally, a bandpass filter was used to preserve frequencies in the range of 0.01 to 0.08 Hz, and covariates such as global signal and noise in gray and white matter were removed. After preprocessing, the brain was divided into multiple brain regions using the AAL-116 template, and time series of each brain region were extracted.
[0057] Step 2: Trim the brain region time series of each sample to the same length, and then use position encoding technology to integrate the relative and absolute position information of the brain region time series;
[0058] To better capture changes in brain regions over time, an additional location encoding vector was added. This vector has the same dimension as the brain region's time series and reflects the location of each time point in the sequence. As a time sequence of primitive brain regions As a new brain region time series with added brain region time point location information, where N1 represents the number of brain regions and d represents the dimension of the feature. X b_pos It can be calculated using the following equation:
[0059] X b_pos =X b +PE(X b )
[0060] This embodiment proposes a novel location encoding method that combines absolute and relative temporal information at brain region nodes using a learnable location encoding function PE(·). The specific implementation equation is shown below:
[0061] PE(X b )=F(X b )+W b
[0062] Where F(·) is the absolute position encoding function. It has X b Learnable temporal position embeddings of the same dimension. The mathematical form of F(·) is shown below:
[0063] F(pos,2i)=sin(pos / 10000 2i / L )
[0064] F(pos,2i+1)=cos(pos / 10000 2i / L )
[0065] Where pos∈[0,d-1] represents X b The time point position, L = N1 represents the number of brain regions, and i ∈ [0, L / 2) represents the i-th component of the time series vector at position pos. It is worth noting that X b_pos With X b The fact that the PE(·) positional encoding function has the same dimension allows the model to easily learn relative positions. Because for any fixed offset k, PE... pos+k It can be represented as PE pos A linear function
[0066] Step 3: Input the brain region time series containing location information obtained in Step 2 into the Transformer multi-head attention model to update the brain region time series features;
[0067] X b_pos As input features to the Transformer multi-head attention model. The multi-head attention mechanism allows the model to interact and transfer information between different brain regions to achieve long-range information fusion, which captures richer feature representations of brain regions. MA(·) represents the multi-head attention mechanism, and the specific implementation equation is shown below:
[0068] H=MA(Q,K,V)=norm(concat(SA1,SA2,…,SA h W)
[0069] in, This is the feature matrix updated using multi-head attention, where c1 represents the feature dimension of the multi-head attention output. MA(·) represents the multi-head attention mechanism, h is the number of heads, W is the learnable parameter matrix, concat(·) represents the connection operation, and norm(·) represents the normalization function. SA represents single-head attention, and the specific implementation equation is shown below:
[0070]
[0071] Q, K, and V represent query, key, and value, respectively. It is the scaling factor for dot product attention. Q, K, and V are all assigned to the brain region's time-series feature X. b_pos .
[0072] Step 4: Input the brain region features obtained in Step 3 into a fully connected classifier to obtain the disease type prediction probability of the sample.
[0073] The feature H obtained in Step 3 is input into the fully connected classifier to obtain the disease type prediction probability of the sample. The specific implementation equation is shown below:
[0074] P = MLP(H)
[0075] Where MLP(·) represents a multilayer fully connected classifier, and P represents the predicted probability of a sample.
[0076] Finally, the binary cross-entropy loss is used as the loss function for model training, and the loss function is implemented by the following formula:
[0077]
[0078] Among them, y i p represents the true label of sample i. i is the predicted probability of sample i, and N is the number of samples.
[0079] Example 2:
[0080] To test the validity of this application, it was applied to a dataset from the ADNI study: the Alzheimer's Disease Neuroimaging Initiative (ADNI) is a longitudinal, multicenter study focused on the early detection and tracking of Alzheimer's disease. The study included participants with different cognitive states, including individuals with AD, cognitively normal controls (NC), and mild cognitive impairment (MCI). Mild cognitive impairment (MCI) is considered a potential precursor to AD and other types of dementia. Based on the new criteria defined by ADNI (ADNIgo, ADNI 2), patients with MCI were divided into two groups: early mild cognitive impairment (EMCI) and late mild cognitive impairment (LMCI). In this embodiment, 869 samples were obtained from ADNI, excluding samples affected by large head movements (exclusion criteria: maximum head movement of 2.0 mm and 2.0 degrees). The final sample of 817 included 222 AD patients, 213 cognitively normal controls (NC), 190 early mild cognitive impairment (EMCI) patients, and 192 late mild cognitive impairment (LMCI) patients.
[0081] Performance evaluation of a brain region time-series location encoding method for predicting Alzheimer's disease types.
[0082] To evaluate the performance of our model, LEformer was compared with four commonly used machine learning or deep learning methods: RF (Random Forest), SVM (Support Vector Machine), CNN (Convolutional Neural Network), and GCN (Graph Convolutional Network).
[0083] The LEformer model was evaluated on four independent classification tasks: AD diagnosis (NC vs. AD), early MCI diagnosis (NC vs. LMCI), differentiation of AD and its early subtypes (LMCI vs. AD), and differentiation of two MCI subtypes (EMCI vs. LMCI). For each independent classification task, all data were randomly split into training and test sets, with 80% of the data used for training and the remaining 20% for testing. Five-fold cross-validation was then used on the training set to optimize the hyperparameters of the LEformer model. All models were tuned according to their respective hyperparameters to ensure experimental fairness.
[0084] The hyperparameters involved in the LEformer model are set as follows: the dimension of the fMRI time series features d = 90, the number of heads in the multi-head attention model h = 2, the learning rate lr = 0.0001 during training, and the number of training iterations epoch = 100.
[0085] To evaluate the performance of the model and all comparison methods, the classification performance metrics used in this work include: accuracy (ACC), sensitivity (SEN), specificity (SPE), f1 score (F1), and area under the ROC curve (AUC). These metrics are calculated as follows:
[0086]
[0087]
[0088]
[0089]
[0090]
[0091] TP, FP, TN, and FN represent true positive, false positive, true negative, and false negative, respectively. Represents the false positive rate. This represents the true positive rate.
[0092] The results in Tables 1 and 2 demonstrate that LEformer achieved optimal classification performance in most of the five evaluation metrics across all four classification tasks. The bolded text in the tables represents the best results. These results indicate that integrating association information between brain regions and genes is beneficial for predictive performance. Furthermore, these results highlight the effectiveness of the location encoding and multi-head attention mechanisms proposed in LEformer, which enable the model to integrate relative and absolute location information from brain region time series and extract relevant information from global associations between brain regions.
[0093]
[0094] Table 1
[0095]
[0096] Table 2
[0097] Comparison of LEformer with various positional encoding methods:
[0098] LEformer enhances the feature representation of time series by combining relative and absolute location information of brain region time series through positional encoding. Therefore, to verify the effectiveness of the positional encoding method proposed by LEformer, the deep learning model structure was fixed, and only the positional encoding method was changed. The proposed positional encoding method was then compared with four popular positional encoding methods. The four positional encoding methods are:
[0099] Transformer: Encodes positional information in an input sequence using sine and cosine functions.
[0100] ViT: Adds a trainable position encoding matrix to represent position information.
[0101] Swin-T: Divides the input sequence into multiple regions and represents positional information by learning the relative positional deviations between the regions.
[0102] MAE: Encodes the two-dimensional positional information of the input sequence using sine and cosine functions.
[0103] The location encoding used in LEformer combines absolute and relative temporal information from brain region time series. Therefore, Transformer and ViT can be considered ablation methods of LEformer location encoding. The bolded text in the table represents the best results.
[0104]
[0105] Table 3
[0106]
[0107] Table 4
[0108] The results in Tables 3 and 4 show that the positional encoding used by the Transformer enables the model to perform better AD disease classification. ViT can also effectively extract the positional information of brain region functional signals over time. LEformer positional encoding combines the Transformer and ViT by combining absolute and relative positional information from time series, thereby effectively encoding brain region features associated with AD disease.
[0109] In summary, after comparison with other prediction methods, the effectiveness of the disease type prediction method based on location encoding and Transformer has been demonstrated.
[0110] The specific embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.
Claims
1. A method for encoding the temporal location of brain regions for predicting Alzheimer's disease types, characterized in that, include: Step 1: Use the DPARSF tool to preprocess the functional magnetic resonance imaging (fMRI) and extract the brain region time series. The brain regions are divided into multiple brain regions using the AAL-116 template. Step 2: Integrate the relative and absolute temporal location information of brain region time series using an absolute location encoding function and a learnable location information matrix; Step 3: Input the brain region time series containing location information obtained in Step 2 into the Transformer multi-head attention model to update the brain region time series features; Step 4: Input the brain region time series features into a fully connected classifier to obtain the disease type prediction probability of the sample; Step 2 specifically includes: Step 2.1: Trim the brain region time series data of all samples after preprocessing and extraction to the same length; Step 2.2: Use Represents the time series of brain regions. As a brain region time series with added time point location information, Indicates the number of brain regions. Indicates the number of time points after cropping; Calculated using the following formula: ; in, The positional encoding function is represented by the following equation: ; in, It is an absolute position encoding function. It has Learnable temporal embeddings of the same dimension The specific implementation equation is as follows: ; in, express The time point and location, Indicates the number of brain regions. Indicates position The first time series vector Each component.
2. The brain region time-series location encoding method for predicting Alzheimer's disease type according to claim 1, characterized in that, Step 1 specifically refers to: Step 1.1: Convert DICOM format fMRI data into NIFTI format that DPARSF can recognize; Step 1.2: Delete the first 10 time points of data for each sample; Step 1.3: Correct the brain scans to the same time point; Step 1.4: Correct the image position shift caused by head movement during the scanning process to the correct position; Step 1.5: Standardize brain imaging data for all subjects to match echo-planar imaging EPI templates; Step 1.6: Perform Gaussian smoothing on the image, with smoothing parameters FWHM of [6, 6, 6]; Step 1.7: Use a bandpass filter to preserve frequencies in the range of 0.01 to 0.08 Hz; Step 1.8: Remove covariates from the image; Step 1.9: Use the AAL-116 template to divide the brain into multiple brain regions, and calculate the average value of the blood oxygen level change signal of all voxels in each brain region to obtain the average time series of each brain region.
3. The brain region time-series location encoding method for predicting Alzheimer's disease type according to claim 1, characterized in that, Step 3 specifically refers to: Will The input is fed into the multi-head attention model, and the specific implementation equation is shown below: ; in, It uses a feature matrix updated with multi-head attention. The feature dimension representing the output of multi-head attention. This indicates a multi-head attention mechanism. It refers to the number of heads. It is a learnable parameter matrix. Indicates a connection operation. Represents the normalization function. This represents single-head attention, and the specific implementation equation is shown below: ; Q, and These represent the query, key, and value, respectively. It is the scaling factor for dot product attention. , and The initial values are all brain region time series. .
4. The brain region time-series location encoding method for predicting Alzheimer's disease type according to claim 1, characterized in that, Step 4 specifically refers to: The time series features The sample is input into a fully connected classifier to obtain the predicted probability of disease type. The specific implementation equation is shown below: ; in, This represents a multi-layer fully connected classifier. Indicates the predicted probability of a sample; Finally, the binary cross-entropy loss is used as the loss function for model training, and the loss function is implemented by the following formula: ; in, Indicates sample The true label, It is a sample The predicted probability, That is the number of samples.