Interpretable Alzheimer's disease deep learning prediction method and system

By employing a multimodal deep fusion prediction method that combines sMRI images, clinical data, and lifestyle data, and utilizing 3D-CNN, MLP embedding layers, XGBoost, and Transformer models, we have achieved multidimensional feature capture and risk prediction for Alzheimer's disease. This addresses the issues of insufficient data quality and model generalization in existing technologies, and provides more comprehensive support for disease progression prediction and treatment options.

CN121709247AInactive Publication Date: 2026-03-20HARBIN INST OF TECH WEIHAI RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511877643.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-12
Publication Date
2026-03-20
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing Alzheimer's disease detection technologies suffer from insufficient data quality and model generalization, lack of coverage of diverse populations, incomplete data integration, limited model functionality, and a lack of quantitative predictive ability for disease progression trends.

Method used

A multimodal deep fusion prediction method is adopted, which extracts features from sMRI images, clinical data and lifestyle data through 3D-CNN, MLP embedding layer and XGBoost, performs feature fusion using Transformer fusion model, and performs AD risk level classification and disease progression probability prediction through the result output layer, while generating a quantitative risk report.

Benefits of technology

It improves the ability to comprehensively utilize complex pathological information, provides AD risk level classification results and disease progression probability, improves the problem of single model function, enhances the quantitative prediction ability of disease progression trend, and provides a more comprehensive and quantitative basis for clinical diagnosis and treatment decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121709247A_ABST
    Figure CN121709247A_ABST
Patent Text Reader

Abstract

The invention relates to an interpretable Alzheimer's disease deep learning prediction method and system. The method comprises the following steps: acquiring multi-modal data of a candidate AD patient, wherein the multi-modal data comprises sMRI images and clinical and lifestyle data; processing the multi-modal data by using a pre-trained multi-modal deep fusion prediction model to obtain an AD risk level and a disease progress probability; wherein multi-modal features are extracted through the sub-modal feature extraction layer, deep fusion is carried out on the multi-modal features through the fusion modeling layer to obtain a global fusion feature vector, the global fusion feature vector is processed through the result output layer, and a risk level classification result and a disease progress probability in future preset time are obtained. By adopting the method, multi-modal heterogeneous data can be effectively fused, qualitative risk grading and quantitative disease course prediction are provided, and the diagnosis comprehensiveness is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of machine learning, and in particular relates to an interpretable deep learning prediction method and system for Alzheimer's disease. Background Technology

[0002] With the development of Alzheimer's disease (AD) prediction technology, it has moved from single-indicator analysis to a multi-dimensional integration stage, with multimodal data fusion and AI algorithm optimization being the research focus. This new technology can integrate up to 443 features, including demographics, medical history, neuropsychological tests, genetic markers (such as APOE-ε4), and structural MRI. Through training on large-scale multi-cohort data, AI models have achieved AUROC values ​​of 0.79 and 0.84 for predicting the core pathological markers Aβ and τ states of AD, respectively, with some deep learning models achieving an accuracy rate exceeding 85% for early AD prediction. Simultaneously, the increasing maturity of data acquisition technologies such as wearable devices, smartphone apps, and electronic medical record extraction has made this technology feasible.

[0003] Traditional methods for diagnosing and screening for Alzheimer's disease (AD) mainly include conventional cognitive assessments, cerebrospinal fluid testing, and recently commercialized high-precision testing technologies (such as Aβ-PET imaging agents). The latter allows for the detection of Aβ positivity 15-20 years before the onset of symptoms, ushering in an era of "pathologically accessible" early diagnosis of AD.

[0004] However, current detection methods or traditional approaches suffer from insufficient data quality and model generalization. Most model training data primarily consist of white participants, lacking diverse population coverage, and their applicability to different racial and regional populations remains to be verified; data integration is incomplete, with insufficient inclusion of plasma biomarkers, affecting the model's accuracy in capturing early pathological changes; and the models have limited functionality, focusing mainly on risk classification and lacking the ability to quantitatively predict disease progression trends. Summary of the Invention

[0005] Therefore, it is necessary to provide an interpretable deep learning prediction method and system for Alzheimer's disease to address the aforementioned technical problems.

[0006] Firstly, this application provides an interpretable deep learning prediction method for Alzheimer's disease, including:

[0007] Acquire multimodal data of the target group; the target group is candidate AD patients, and the multimodal data includes sMRI imaging data, clinical data, and lifestyle data;

[0008] A pre-trained multimodal deep fusion prediction model is used to process multimodal data to obtain AD risk level and disease progression probability;

[0009] Specifically, a pre-trained multimodal deep fusion prediction model is used to process multimodal data to obtain AD risk levels and disease progression probabilities, including:

[0010] Multimodal feature extraction is performed on multimodal data through a modal feature extraction layer to obtain multimodal features; the modal feature extraction layer includes 3D-CNN, MLP embedding layer, and XGBoost;

[0011] By using a fusion modeling layer, multimodal features are deeply fused to obtain a global fusion feature vector;

[0012] The output layer processes the global fusion feature vector to obtain the AD risk level classification result and the probability of disease progression within a preset time period.

[0013] In one embodiment, a multimodal feature extraction layer is used to extract features from the multimodal data to obtain multimodal features, including:

[0014] 3D-CNN is used to process sMRI image data to automatically learn the spatial structural information of sMRI image data and output brain region features; 3D-CNN includes convolutional layers and fully connected layers, where the calculation of the convolutional layer is expressed by the following formula:

[0015]

[0016] in, For the input 3D image data, This represents a 3D convolution operation. This is the 3D convolution kernel weight vector of the convolutional layer. This is the bias vector of the convolutional layer. It is the ReLU activation function;

[0017] The clinical molecular data in the clinical data is transformed using an MLP embedding layer to output a clinical feature vector;

[0018] XGBoost was used to perform feature filtering on lifestyle data to assess the importance of different lifestyle factors to the target population and output lifestyle features.

[0019] Brain region features, clinical feature vectors, and lifestyle features are integrated into multimodal features.

[0020] In one embodiment, a fusion modeling layer is used to perform deep fusion of multimodal features to obtain a global fused feature vector, including:

[0021] Multimodal features are input into an attention-based Transformer fusion model. The attention mechanism of the Transformer fusion model processes the multimodal features, and feature weights are dynamically assigned to the multimodal features using the following formula:

[0022]

[0023] in, For query matrix; The key matrix; It is a value matrix; Let be the dimension of the key matrix; Used to normalize the calculated scores into feature weights;

[0024] Based on feature weights, multimodal features are weighted and aggregated to generate a global fusion feature vector.

[0025] In one embodiment, the result output layer outputs the AD risk level classification result and the probability of disease progression within a preset future time period in parallel, including:

[0026] The global fused feature vector is input into the classification task branch, and the AD risk level classification result is output through the Softmax activation function.

[0027] The global fusion feature vector is input into the regression task branch, and the probability of disease progression within a preset time period is output through the linear output layer.

[0028] In one embodiment, the training steps of the multimodal deep fusion prediction model include:

[0029] The complete multimodal data used for training is divided into a training set, a validation set, and a test set;

[0030] The training set is preprocessed to obtain a balanced training set;

[0031] The multimodal deep fusion prediction model is iteratively trained using a balanced training and validation set, resulting in a new stage model after each iteration.

[0032] Use the validation set to evaluate the AUROC, accuracy, and high-risk recall of the staged model;

[0033] When AUROC, accuracy, and high-risk recall reach the preset optimal state, the current stage model is used as a pre-trained multimodal deep fusion prediction model.

[0034] In one embodiment, an interpretable deep learning prediction method for Alzheimer's disease further includes:

[0035] Calculate the weighted gradient of the AD risk level classification result relative to the activation map of the last convolutional layer in the 3D-CNN in the modality feature extraction layer;

[0036] Using weighted gradients, the activation maps of convolutional layers are summed in weights and the ReLU activation function is applied to generate a three-dimensional risk heatmap.

[0037] In the feature importance ranking based on the multimodal deep fusion prediction model, a list of key risk factors is extracted from the multimodal input data;

[0038] The three-dimensional risk heat map, the list of key risk factors, the AD risk level classification results, and the probability of disease progression within a preset time period are integrated into a report dataset;

[0039] Generate a quantitative risk report based on the report dataset.

[0040] Furthermore, an interpretable deep learning prediction method for Alzheimer's disease also includes:

[0041] Based on the quantitative risk report, the AD risk level classification results and key risk factors are extracted and combined to form a risk profile fact.

[0042] The risk profile facts are loaded into a rule engine that has a pre-configured intervention suggestion library. The rule engine compares the risk profile facts with the rule set in the intervention suggestion library to select intervention rules.

[0043] The specific intervention measures corresponding to the compiled intervention rules are used to generate structured, personalized intervention plans.

[0044] Secondly, this application also provides an interpretable deep learning prediction system for Alzheimer's disease, comprising:

[0045] The acquisition module is used to acquire multimodal data of the target object; the target object is the patient, and the multimodal data includes sMRI image data, clinical data, and lifestyle data.

[0046] The multimodal model module is used to process multimodal data using a pre-trained multimodal deep fusion prediction model. The processing procedure is as follows:

[0047] Multimodal feature extraction is performed on multimodal data through a modal feature extraction layer to obtain multimodal features; the modal feature extraction layer includes 3D-CNN, MLP embedding layer, and XGBoost;

[0048] By using a fusion modeling layer, multimodal features are deeply fused to obtain a global fusion feature vector;

[0049] The results output layer outputs AD risk level classification results and the probability of disease progression within a preset time period in parallel.

[0050] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps described in any of the above methods.

[0051] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps described in any of the above methods.

[0052] The aforementioned interpretable deep learning prediction method and system for Alzheimer's disease integrates sMRI images, clinical data, and lifestyle data, and processes them using a modal feature extraction layer and a deep fusion modeling layer. This allows for the capture of disease characteristics from multiple dimensions, improving the comprehensive utilization of complex pathological information. The output layer simultaneously provides AD risk level classification results and disease progression probabilities, addressing the problem of existing models having only a single function and focusing solely on risk classification. This dual-task output mechanism can provide more comprehensive and quantitative decision-making basis for clinical diagnosis and the development of prospective treatment plans. Attached Figure Description

[0053] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0054] Figure 1 This is a flowchart illustrating an interpretable deep learning prediction method for Alzheimer's disease according to one embodiment of the present invention.

[0055] Figure 2 This is a block diagram of an interpretable deep learning prediction system for Alzheimer's disease according to one embodiment of the present invention.

[0056] Figure 3 This is a block diagram of an interpretable deep learning prediction system for Alzheimer's disease according to another embodiment of the present invention. Detailed Implementation

[0057] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0058] In one embodiment, such as Figure 1 As shown, an interpretable deep learning prediction method for Alzheimer's disease is provided. This embodiment illustrates the application of this method to a processing device. It is understood that this method can also be applied to a server, and further to a system including both a processing device and a server, and implemented through the interaction between the processing device and the server. In this embodiment, the method may include:

[0059] Step 110: Obtain multimodal data of the target subject; the target subject is candidate AD patients, and the multimodal data includes sMRI imaging data, clinical data, and lifestyle data.

[0060] The data comprises sMRI imaging data (structural magnetic resonance imaging), which provides the processing device with information about the target subject's brain anatomy. Clinical data comes from the target subject's medical records and assessment results, such as demographic information, neuropsychological assessment scores, or biological indicators. Lifestyle data is information about the target subject's daily activities, dietary habits, or other behavioral patterns. The processing device integrates the sMRI imaging data, clinical data, and lifestyle data as input for predictive operations. The processing device verifies the integrity and format of the acquired data to ensure it can be used by subsequent predictive models.

[0061] Step 120: Use a pre-trained multimodal deep fusion prediction model to process the multimodal data to obtain the AD risk level and the probability of disease progression.

[0062] The training process of the multimodal deep fusion prediction model was completed before deployment, and the model parameters were fixed. The multimodal deep fusion prediction model is used to analyze the complex relationships between sMRI imaging data, clinical data, and lifestyle data. The AD risk level is a classification result, such as classifying the target population as high-risk, medium-risk, or low-risk. The disease progression probability is a regression result, representing the likelihood of cognitive decline in the target population within a predetermined timeframe. The processing device inputs the acquired multimodal data into this pre-trained multimodal deep fusion prediction model for processing. Specifically, the processing device outputs two results through the calculations of the multimodal deep fusion prediction model: the AD risk level and the disease progression probability. The processing device stores the AD risk level and the disease progression probability for subsequent report generation or intervention decisions.

[0063] Using a pre-trained multimodal deep fusion prediction model, multimodal data is processed to obtain AD risk levels and disease progression probabilities, which may include:

[0064] Step 121: Extract features from the multimodal data through a modal feature extraction layer to obtain multimodal features; the modal feature extraction layer includes 3D-CNN, MLP embedding layer, and XGBoost.

[0065] The system utilizes 3D-CNN (3D Convolutional Neural Network) to learn spatial structure information from 3D images. The MLP embedding layer (Multilayer Perceptron) is used to transform high-dimensional, sparse clinical molecular data into low-dimensional, dense clinical feature vectors. XGBoost is an open-source, high-performance machine learning library based on the eXtreme Gradient Boosting algorithm, used to filter lifestyle variables that contribute to the prediction results. The processing device uses 3D-CNN to process sMRI image data, MLP embedding layers to process clinical data, and XGBoost to process lifestyle data. The processing device combines the outputs of 3D-CNN, MLP embedding layers, and XGBoost to form a multimodal feature set containing image, clinical, and lifestyle information.

[0066] Step 122: The multimodal features are deeply fused through the fusion modeling layer to obtain a global fusion feature vector.

[0067] The fusion modeling layer effectively integrates feature information from different modalities. Internally, it uses a Transformer-based attention mechanism to calculate the correlations between different features. This attention mechanism dynamically assigns weights to features from different modalities, such as increasing attention to key pathological biomarkers. The global fusion feature vector is a unified, information-dense vector representation containing comprehensive information from all input modalities. The processing device inputs multimodal features into the fusion modeling layer, and through its calculations, outputs the final global fusion feature vector.

[0068] Step 123: The global fusion feature vector is processed through the result output layer to obtain the AD risk level classification result and the probability of disease progression within a preset time period.

[0069] The output layer comprises two parallel task branches. The first branch is a classification task branch, which uses the Softmax activation function to transform the global fused feature vector into a probability distribution, thereby outputting the AD risk level classification result. The second branch is a regression task branch, which uses a linear output layer to map the global fused feature vector into continuous values, thereby outputting the probability of disease progression within a preset future timeframe. The processing device simultaneously inputs the global fused feature vector into both the classification and regression task branches, outputting the AD risk level classification result and the disease progression probability as the final prediction result.

[0070] This embodiment provides an interpretable deep learning prediction method for Alzheimer's disease. By integrating sMRI images, clinical data, and lifestyle data, and processing them using a modal feature extraction layer and a deep fusion modeling layer, it can capture disease features from multiple dimensions, improving the comprehensive utilization of complex pathological information. Through its output layer, this method can simultaneously provide AD risk level classification results and disease progression probabilities, addressing the problem of existing technologies where models have a single function and only focus on risk classification. This dual-task output mechanism can provide more comprehensive and quantitative decision-making basis for clinical diagnosis and the development of prospective treatment plans.

[0071] In one embodiment, feature extraction is performed on the multimodal data through a modal feature extraction layer to obtain multimodal features, which may include:

[0072] Step 201: Process sMRI image data using 3D-CNN to automatically learn the spatial structural information of the sMRI image data and output brain region features; 3D-CNN includes convolutional layers and fully connected layers, wherein the calculation of the convolutional layer is expressed by the following formula:

[0073]

[0074] in, For the input 3D image data, This represents a 3D convolution operation. This is the 3D convolution kernel weight vector of the convolutional layer. This is the bias vector of the convolutional layer. This is the ReLU activation function.

[0075] The processing equipment will input 3D image data With the preset 3D convolution kernel weight vector Convolution is performed to capture spatial variations in the image data across its length, width, and height dimensions. The processing device then combines the convolution result with a bias vector. The data is added together, and the data distribution is adjusted. The processing device applies the ReLU activation function to the result of the addition, transforming the linear operation into a non-linear feature map, thereby generating a feature map containing local spatial structure information. The processing device inputs the feature map into a fully connected layer, and the weight matrix of the fully connected layer maps the high-dimensional spatial features into a one-dimensional feature vector. The processing device outputs a one-dimensional feature vector as brain region features, which contain 2048 numerical values ​​representing deep anatomical structure information learned from the original images.

[0076] Step 202: Use the MLP embedding layer to transform the clinical molecular data in the clinical data and output the clinical feature vector.

[0077] The MLP embedding layer is a neural network structure composed of a series of linear transformations and nonlinear activations. It transforms tabular clinical data into a mathematical form that computers can understand in conjunction with sMRI image features. The processing device uses the MLP embedding layer to process clinical molecular data. Specifically, it maps the clinical molecular data into a continuous vector space. Inside the MLP embedding layer, it performs matrix multiplication, multiplying the input data by the embedding matrix. Through this mapping operation, it transforms potentially sparse or discrete clinical molecular data into dense real-valued vectors. The mapping process preserves the potential correlations between data points. The processing device outputs a 192-dimensional clinical feature vector through the computation of this embedding layer.

[0078] Step 203: Use XGBoost to perform feature filtering on lifestyle data to assess the importance of different lifestyle factors to the target population and output lifestyle features.

[0079] XGBoost is used to select the K most important key features from the raw, complex lifestyle data based on their contribution, where K can be 64. XGBoost transforms these discrete, heterogeneous key features into a compact, high-density K-dimensional feature vector. The processing device calls the XGBoost model to analyze the lifestyle data. Specifically, the processing device constructs multiple decision trees within the XGBoost model. The processing device iteratively optimizes these decision trees using a gradient boosting algorithm. The processing device calculates the gain or coverage of each lifestyle factor during the decision tree construction process to assess the importance of each factor to the prediction target. Based on the evaluation results, the processing device retains the lifestyle factors with high importance and removes noisy data with low contribution to the prediction results. The processing device combines the selected lifestyle factors into a K-dimensional lifestyle feature vector and outputs the K-dimensional lifestyle feature vector.

[0080] Step 204: Integrate brain region features, clinical feature vectors, and lifestyle features into multimodal features.

[0081] The multimodal features simultaneously include spatial structural information from images, clinical molecular biological information, and lifestyle behavioral information. The processing device acquires brain region features output from a 3D-CNN, clinical feature vectors output from an MLP embedding layer, and lifestyle features output from XGBoost. The processing device then concatenates or stacks these three sets of feature vectors to construct a unified multimodal feature set.

[0082] This embodiment provides an interpretable deep learning prediction method for Alzheimer's disease. By employing 3D-CNN, MLP embedding layers, and XGBoost to extract features from sMRI images, clinical molecular data, and lifestyle data, differentiated processing can be achieved for different data modalities. This targeted extraction method can preserve the three-dimensional spatial structure information in sMRI images, transform sparse clinical molecular data into dense vectors for easier computation, and screen out lifestyle factors that actually contribute to disease prediction. By integrating these three heterogeneous features, this embodiment can enrich the information dimension of the input model, improve the model's ability to represent the complex pathological mechanisms of Alzheimer's disease, and thus improve the accuracy and robustness of the final prediction results.

[0083] In one embodiment, a fusion modeling layer is used to perform deep fusion of multimodal features to obtain a global fusion feature vector, which may include:

[0084] Step 301: Input the multimodal features into the Transformer fusion model based on the attention mechanism. Process the multimodal features through the attention mechanism of the Transformer fusion model, and calculate and dynamically assign feature weights to the multimodal features using the following formula:

[0085]

[0086] in, For query matrix; The key matrix; It is a value matrix; Let be the dimension of the key matrix; This is used to normalize the calculated scores into feature weights.

[0087] The processing device maps brain region features, clinical feature vectors, and lifestyle features into a unified hidden layer dimensional space to adapt to the input requirements of the Transformer fusion model. The processing device first performs matrix multiplication operations to calculate the query matrix. AND key matrix The dot product is the product of the transposes of the key matrices. The result of the dot product operation reflects the degree of matching or correlation strength between eigenvectors of different modalities. The processing device divides the dot product result by the square root of the dimension of the key matrix. The scaling operation is performed to adjust the numerical range, preventing the dot product result from becoming too large and causing the Softmax function to enter the saturation region of minimal gradients, thus maintaining gradient stability during model training. The processing device converts the numerical matrix into a normalized probability distribution form, i.e., the attention weight matrix, using the Softmax function. Each value in the attention weight matrix is ​​between zero and one, representing the model's attention to a specific combination of features. The processing device uses this calculation process to dynamically evaluate the contribution of each feature to the final prediction result; for example, it automatically assigns higher weight values ​​when identifying feature patterns highly correlated with AD pathology.

[0088] Step 302: Based on feature weights, perform weighted aggregation of multimodal features to generate a global fusion feature vector.

[0089] The global fusion feature vector is a high-dimensional numerical vector that comprehensively encodes key information and their interactions from imaging, clinical, and lifestyle data, serving as the input for the subsequent output layer. The processing device then calculates the attention weight matrix and value matrix... Matrix multiplication is performed. The value matrix containing the original information is then weighted and recombined through matrix multiplication. Specifically, the processing device enhances or suppresses feature values ​​according to their weights, generating a weighted feature representation. This weighted recombination process extracts key information and filters out noise. The processing device concatenates or sums the weighted results from different attention heads. The processing device then performs nonlinear transformations and further feature extraction on the fused information through a feedforward neural network layer. Before generation, the processing device applies layer normalization and residual connection operations to optimize the distribution of feature data and prevent information degradation during deep network propagation. Finally, the processing device generates a globally fused feature vector.

[0090] This embodiment provides an interpretable deep learning prediction method for Alzheimer's disease. By employing an attention-based Transformer model in the fusion modeling layer, intelligent deep fusion of multimodal features can be achieved. This method utilizes a specific attention calculation formula to accurately quantify the dependencies and importance between different modal features, thereby dynamically allocating feature weights according to the specific pathological patterns of the input data. This mechanism allows the model to automatically focus on core features with high indicative value for Alzheimer's disease diagnosis, such as hippocampal atrophy or abnormal specific protein indicators, while effectively suppressing noise interference that may exist in lifestyle data. This embodiment can solve the problems of intermodal information interference and weight equalization caused by traditional direct feature concatenation, enhance the model's ability to capture early, weak pathological signals, and thus improve the accuracy of disease risk prediction and progression assessment.

[0091] In one embodiment, the result output layer, which outputs AD risk level classification results and the probability of disease progression within a preset future time period in parallel, may include:

[0092] Step 401: Input the global fusion feature vector into the classification task branch, and output the AD risk level classification result through the Softmax activation function.

[0093] The processing device receives a global fusion feature vector generated by the fusion modeling layer and simultaneously inputs this vector into two parallel computation branches constructed within the result output layer. The processing device processes the global fusion feature vector in the classification task branch. Specifically, the processing device maps the high-dimensional features to dimensional spaces corresponding to different AD risk levels through the fully connected layer of the classification branch. The processing device applies the Softmax activation function to the mapped output, using the Softmax function to transform the original output values ​​of the neural network into a normalized probability distribution. Based on this probability distribution, the processing device determines the specific risk level classification of the target object: normal cognition, mild cognitive impairment, or Alzheimer's disease.

[0094] Step 402: Input the global fusion feature vector into the regression task branch, and output the probability of disease progression within a preset time period through the linear output layer.

[0095] The processing device processes the globally fused feature vector in the regression task branch. It extracts feature information related to the temporal evolution of the disease through the fully connected layer of the regression branch. A linear output layer is used at the end of the regression branch for computation. This linear output layer generates continuous real-valued results. The processing device represents these continuous real-valued results as the predicted change in the probability of disease progression or cognitive ability score of the target object within a preset future time period. The processing device outputs the probability of disease progression within the preset future time period, which can be within the next year.

[0096] This embodiment provides an interpretable deep learning prediction method for Alzheimer's disease. By constructing parallel classification and regression dual-task branches at the output layer, a comprehensive assessment of the patient's condition can be achieved. This embodiment utilizes the classification branch to clarify the patient's current risk level, and the regression branch to quantify the patient's disease progression trend within a specific future timeframe. This multi-task parallel processing mechanism fully leverages the expressive power of shared underlying features, overcoming the limitations of single-task models in information output. This method can provide clinicians with a comprehensive diagnostic basis of "current status + trend," improving the foresight in identifying high-risk patients and thus assisting in developing more precise early intervention plans.

[0097] In one embodiment, the training steps of the multimodal deep fusion prediction model may include:

[0098] Step 501: Divide the complete multimodal data used for training into a training set, a validation set, and a test set.

[0099] The processing device divides the dataset into three mutually exclusive subsets—training, validation, and test sets—in a pre-defined ratio of 7:1.5:1.5. The training set is designated as the foundational data for model parameter updates, the validation set as reference data for hyperparameter tuning and model state evaluation, and the test set as independent data for final performance assessment. This partitioning mechanism establishes the data foundation for model training and provides a standardized data input interface for subsequent iterative optimization.

[0100] Step 502: Preprocess the training set to obtain a balanced training set.

[0101] The processing device uses the SMOTE oversampling algorithm to interpolate minority class samples, generating synthetic minority class samples to augment the data volume. SMOTE oversampling is a data augmentation technique that generates artificially synthesized samples by linearly interpolating between minority class samples and their nearest neighbors. It addresses the severe imbalance between the number of diseased and healthy samples in medical data, preventing the model from ignoring the minority. Simultaneously, the processing device uses the Hard Negative Mining algorithm to filter majority class samples, retaining those samples that are difficult to classify correctly under the current model conditions. Hard Negative Mining is a strategy that specifically filters out negative samples that are misclassified by the model with high confidence during training and strengthens the training process, improving the model's ability to distinguish easily confused samples and significantly reducing the false positive rate. The processing device merges the synthetic minority class samples with the filtered majority class samples. The processing device outputs a balanced training set with a uniform class distribution.

[0102] Step 503: Use the balanced training and validation sets to iteratively train the multimodal deep fusion prediction model, thereby obtaining a new stage model after each iteration.

[0103] The processing device loads the initial architecture of the multimodal deep fusion prediction model and performs forward propagation computation on the model using a balanced training set. During computation, the device randomly discards some neuron connections, performing a dropout operation with a dropout rate of 0.3. It adds a sum-of-squares term to the loss function and performs L2 regularization with a regularization coefficient of 0.001. The device employs a five-fold cross-validation strategy, dividing the training process into five sub-stages for sequential validation. During iteration, the device dynamically adjusts the learning rate parameter, gradually decreasing it from 0.001 to 0.00001 in a stepwise manner. Finally, the device updates the weights of the multimodal deep fusion prediction model using the backpropagation algorithm and generates a stage model at the end of each iteration cycle.

[0104] Step 504: Use the validation set to evaluate the AUROC, accuracy, and high-risk recall of the staged model.

[0105] AUROC, also known as the Area Under the Receiver Operating Characteristic curve, is a core metric for measuring the generalization and discrimination capabilities of binary classification models. It assesses the overall ability of a staged model to correctly distinguish between positive and negative samples at various classification thresholds. The processing device inputs validation set data into the current staged model. The processing device obtains the staged model's prediction output for the validation set samples. The processing device calculates the difference between the staged model's prediction and the true labels. Based on the calculation results, the processing device statistically analyzes the AUROC value, overall accuracy, and recall rate for high-risk categories, serving as the core basis for measuring the current staged model's performance. The processing device records the values ​​for each evaluation, forming a staged model training log. Through this step, the processing device monitors the staged model's generalization ability on unseen data, preventing overfitting on the training set.

[0106] Step 505: When AUROC, accuracy, and high-risk recall reach the preset comprehensive optimal state, the current stage model is used as a pre-trained multimodal deep fusion prediction model.

[0107] The processing device reads a preset comprehensive optimal state threshold, such as an AUROC greater than or equal to 0.88 and an accuracy greater than or equal to 90%. The processing device compares the evaluation metrics of the current stage model with the comprehensive optimal state threshold. If the evaluation metrics meet or exceed the comprehensive optimal state threshold, the processing device issues a stop command to terminate the iterative training process. The processing device marks the current stage model as the final pre-trained multimodal deep fusion prediction model and saves its parameters. If the evaluation metrics do not reach the comprehensive optimal state threshold, the processing device continues to execute the next round of iterative training.

[0108] This embodiment presents an interpretable deep learning prediction method for Alzheimer's disease. By introducing strict data partitioning and balancing strategies during training, it addresses the common sample imbalance problem in medical data. This embodiment utilizes SMOTE and Hard Negative Mining techniques to improve the model's sensitivity in identifying minority classes (such as high-risk AD patients). By combining Dropout, L2 regularization, and dynamic learning rate adjustment, this method effectively suppresses overfitting and enhances the model's generalization ability on complex multimodal data. This embodiment optimizes AUROC and high-risk recall as optimization targets, guiding the model towards clinical applicability and thus improving the reliability and screening efficiency of the prediction model in real-world applications.

[0109] In one embodiment, an interpretable deep learning prediction method for Alzheimer's disease may further include:

[0110] Step 601: Calculate the weighted gradient of the AD risk level classification result relative to the activation map of the last convolutional layer in the 3D-CNN in the modality feature extraction layer.

[0111] The processing device acquires the AD risk level classification result generated by the output layer. It then determines the target category node corresponding to this classification result. Next, it locates the last convolutional layer of the 3D-CNN within the modality feature extraction layer and executes the backpropagation algorithm. Specifically, it calculates the gradient value of the predicted score of the target category node relative to each feature activation map in this last convolutional layer. The processing device performs global average pooling on the calculated gradient values. This global average pooling operation then calculates the weight coefficient corresponding to each feature activation map. The weight coefficient represents the importance of the corresponding feature activation map in the model's current AD risk level classification decision. Finally, the processing device uses the weight coefficient corresponding to each feature activation map as a weighted gradient.

[0112] Step 602: Using weighted gradients, the activation maps of the convolutional layers are summed in a weighted manner and the ReLU activation function is applied to generate a three-dimensional risk heatmap.

[0113] The 3D risk heatmap, in spatially distributed numerical form, marks anatomical regions in brain images that lead the model to classify them as having a specific risk level. The processing device linearly weights and sums the calculated weight coefficients with the corresponding feature activation maps. This generates a weighted composite map containing both positive and negative values, and the ReLU activation function is applied to this composite map. Specifically, the processing device sets all values ​​less than 0 in the composite map to 0, while retaining all values ​​greater than 0, thereby filtering out feature regions that negatively impact the target category prediction and retaining only regions that positively contribute to the prediction results. The processing device then upsamples and interpolates the ReLU-processed data to ensure that the spatial resolution of the processed data matches that of the original sMRI image data. The processing device outputs the upsampled and interpolated data as the 3D risk heatmap.

[0114] Step 603: In the feature importance ranking based on the multimodal deep fusion prediction model, extract a list of key risk factors from the multimodal input data.

[0115] The processing device accesses the internal parameter states of the multimodal deep fusion prediction model to obtain feature importance data generated during inference. It reads the feature gain values ​​from XGBoost and the attention weight distribution from the Transformer fusion model. The processing device ranks all input clinical and lifestyle data features by importance. Specifically, it sets an importance screening threshold, selecting features with importance scores higher than the threshold. The processing device identifies features with scores higher than the threshold as key risk factors leading to the current prediction result, generating a list of key risk factors containing feature names and their corresponding importance scores.

[0116] Step 604: Integrate the three-dimensional risk heat map, the list of key risk factors, the AD risk level classification results, and the probability of disease progression within a preset future time into a report dataset.

[0117] The processing device reads AD risk level classification results, disease progression probability values, generated 3D risk heatmap data, and an extracted list of key risk factors. It creates a unified data container object and writes the four types of heterogeneous data into this object according to a predefined data structure. The processing device then serializes this heterogeneous data, converting it into a standardized data exchange format, such as JSON or XML. Finally, it defines this structured data set as a report dataset.

[0118] Step 605: Generate a quantitative risk report based on the report dataset.

[0119] The processing device utilizes a document rendering engine to read a pre-set report template file. It then populates the text fields of the template with AD risk levels and disease progression probabilities. Specifically, it overlays a 3D risk heatmap onto a standard brain anatomy template to generate a visualized image of lesion distribution, which is then embedded into the template's image field. The device converts a list of key risk factors into charts or text lists and embeds them into the template's analysis field. Finally, it compiles the completed template to generate a portable document format (PDF) file. This file is then output as an exportable quantitative risk report.

[0120] This embodiment provides an interpretable deep learning prediction method for Alzheimer's disease. By introducing Gradient Weighted Class Activation Mapping (Grad-CAM) technology and feature importance analysis, it enables interpretability analysis of the deep learning model's prediction results. This embodiment generates a three-dimensional risk heatmap, mapping abstract algorithmic decisions back to specific brain anatomical regions, allowing clinicians to intuitively see the lesion locations the model focuses on (such as the hippocampus or temporal lobe). By extracting a list of key risk factors, this method can quantify the specific contribution of different clinical indicators and lifestyle factors to the risk of disease. This embodiment generates a comprehensive report containing visual charts and quantitative indicators, reducing the "black box" effect of deep learning models in medical applications, enhancing doctors' and patients' trust in AI-assisted diagnostic results, and thus providing pathologically grounded reference support for clinical decision-making.

[0121] Furthermore, an interpretable deep learning prediction method for Alzheimer's disease may also include:

[0122] Step 701: Based on the quantitative risk report, extract the AD risk level classification results and key risk factors, and combine them into a risk profile fact.

[0123] The processing device reads the quantitative risk report file generated in the preceding steps, parses its document structure, and locates the data fields storing AD risk level classification results and the data area storing the list of key risk factors. Specifically, the processing device extracts the text values ​​of the AD risk levels, such as "high risk" or "medium risk." It also extracts specific entries from the list of key risk factors, such as "APOE4 gene carrier" or "history of hypertension." The processing device instantiates risk profile fact objects in memory and assigns the extracted risk level values ​​and risk factor entries to the corresponding attributes of the fact objects. Finally, the processing device converts the fact objects into a standard input format recognizable by the rule engine, such as Java objects or JSON data packets, thus completing the transformation from unstructured or semi-structured report data to structured fact data that can be processed by the computer's logical reasoning unit.

[0124] Step 702: Load the risk profile facts into the rule engine that has a pre-configured intervention suggestion library. The rule engine compares the risk profile facts with the rule set in the intervention suggestion library to select intervention rules.

[0125] The intervention suggestion library contains a series of logical rules defined in a "condition-outcome" format. These rules encode intervention strategies from medical experts for different conditions. The processing device initializes a running instance of the inference rule engine, loading the pre-configured intervention suggestion library into the engine's production memory. The processing device then inserts the constructed risk profile fact objects into the engine's working memory, triggering the engine's pattern matching algorithm. Specifically, the processing device iterates through the rule set in production memory, comparing the condition portion of each rule with the risk profile fact attributes in working memory. The processing device identifies all rules whose conditions are met and marks them as triggered. Finally, the processing device places all triggered rules into the execution agenda, awaiting further processing.

[0126] Step 703: Compile the specific intervention measures corresponding to the intervention rules to generate a structured personalized intervention plan.

[0127] The processing device sequentially executes the consequences of the rules triggered in the agenda. It retrieves specific intervention data from each triggered rule. For example, the interventions obtained by the processing device include, but are not limited to, dietary adjustment suggestions, exercise plans of specific intensity and frequency, and targeted cognitive training tasks. The processing device aggregates all interventions obtained from multiple rules into a candidate list. It performs deduplication and conflict detection logic on the candidate list to eliminate duplicate suggestions or mutually exclusive instructions. The processing device then arranges the compiled interventions according to a predefined logical structure, such as dividing them into three sections: lifestyle habits, nutritional advice, and rehabilitation training, generating a structured, personalized intervention plan.

[0128] This embodiment provides an interpretable deep learning prediction method for Alzheimer's disease. By integrating a rule engine and an expert knowledge base, it enables automated closed-loop management from risk prediction to health intervention. This embodiment utilizes a rule matching mechanism to automatically transform complex medical diagnostic results (risk level and risk factors) into specific, actionable lifestyle recommendations, improving upon the inefficiencies of traditional diagnosis that often involves "diagnosis without treatment" or intervention plans relying on manual formulation. By customizing interventions based on individual risk profiles, this embodiment enhances the targeting and effectiveness of health management plans, strengthens patient adherence to early interventions, and ultimately slows the pathological progression of Alzheimer's disease and improves patients' quality of life.

[0129] The preferred embodiment of the present invention provides an interpretable deep learning prediction method for Alzheimer's disease, comprising the following steps:

[0130] Step 1: Divide the complete multimodal data used for training into training set, validation set, and test set.

[0131] Step 2: Perform SMOTE oversampling and Hard Negative Mining on the training set to obtain a balanced training set.

[0132] Step 3: Iteratively train the multimodal deep fusion prediction model using the balanced training and validation sets. Dropout and L2 regularization are applied during iterative training. Hyperparameters are optimized using 5-fold cross-validation, with the learning rate decreasing from 0.001 to 0.00001. A new stage model is obtained after each iteration.

[0133] Step 4: Use the validation set to evaluate the AUROC, accuracy, and high-risk recall of the staged model.

[0134] Step 5: Determine whether AUROC, accuracy, and high-risk recall have reached the preset optimal overall state. Stop iterative training when the preset optimal overall state is reached, and use the current stage model as the pre-trained multimodal deep fusion prediction model.

[0135] Step 6: Obtain multimodal data for the target group. The target group consists of candidate AD patients. Multimodal data includes sMRI imaging data, clinical data, and lifestyle data.

[0136] Step 7: Use 3D-CNN to process sMRI image data and output brain region features.

[0137] Step 8: Use the MLP embedding layer to transform the clinical molecular data in the clinical data and output the clinical feature vector.

[0138] Step 9: Use XGBoost to filter features from the lifestyle data and output the lifestyle features.

[0139] Step 10: Integrate brain region features, clinical feature vectors, and lifestyle features into multimodal features.

[0140] Step 11: Input the multimodal features into the Transformer fusion model based on the attention mechanism, process the multimodal features through the attention mechanism of the Transformer fusion model, and calculate and dynamically assign feature weights to the multimodal features.

[0141] Step 12: Based on the feature weights, perform weighted aggregation of multimodal features to generate a global fusion feature vector.

[0142] Step 13: Input the global fusion feature vector into the classification task branch and output the AD risk level classification result through the Softmax activation function; input the global fusion feature vector into the regression task branch and output the disease progression probability within a preset time period through the linear output layer.

[0143] Step 14: Calculate the weighted gradient of the AD risk level classification result relative to the activation map of the last convolutional layer in the 3D-CNN in the modality feature extraction layer.

[0144] Step 15: Use weighted gradients to perform weighted summation on the activation map of the convolutional layer and apply the ReLU activation function to generate a three-dimensional risk heatmap.

[0145] Step 16: Extract a list of key risk factors from the multimodal input data based on the feature importance ranking of the multimodal deep fusion prediction model.

[0146] Step 17: Integrate the three-dimensional risk heat map, the list of key risk factors, the AD risk level classification results, and the probability of disease progression within a preset future time into a report dataset.

[0147] Step 18: Generate a quantitative risk report based on the report dataset.

[0148] Step 19: Extract the AD risk level classification results and key risk factors from the quantitative risk report. The processing device combines the AD risk level classification results and key risk factors into a risk profile.

[0149] Step 20: Load the risk profile facts into the rule engine, which has a pre-configured intervention suggestion library. The rule engine compares the risk profile facts with the rule set in the intervention suggestion library. The rule engine then selects intervention rules.

[0150] Step 21: Compile the specific intervention measures corresponding to the intervention rules to generate a structured, personalized intervention plan.

[0151] The method provided in the preferred embodiment of this invention, by constructing an end-to-end deep learning architecture including modal feature extraction, Transformer attention fusion, and dual-task output, can achieve closed-loop processing of Alzheimer's disease from risk identification to intervention management. In the feature extraction stage, this scheme utilizes 3D-CNN, MLP embedding layers, and XGBoost to process sMRI images, clinical data, and lifestyle data respectively, fully mining the heterogeneous features of different modalities and avoiding the information loss problem caused by a single data source. In the feature fusion stage, the introduction of a Transformer-based attention mechanism, coupled with a specific weight calculation formula, allows for dynamic allocation of weights to features of different modalities. This solves the problem of feature mutual exclusion or redundancy caused by weight equalization in traditional fusion methods, enabling the model to automatically focus on core pathological features such as hippocampal atrophy or specific protein indicators, thereby improving prediction accuracy.

[0152] Furthermore, this solution outputs AD risk levels and disease progression probabilities in parallel through a dual-task output layer, simultaneously providing a qualitative assessment of the current state and a quantitative prediction of future trends, enriching the reference dimensions for clinical diagnosis. During model training, SMOTE and Hard Negative Mining address data imbalance, and the combination of Dropout and L2 regularization enhances the model's robustness and generalization ability, reducing the risk of overfitting. Further, this solution utilizes Grad-CAM technology to generate a 3D risk heatmap and extract key risk factors, addressing the "black box" problem of deep learning models and providing doctors with visualized pathological evidence, enhancing the interpretability of diagnostic results. Finally, a rule engine automatically transforms quantitative risk reports into structured, personalized intervention plans, achieving automated integration from risk prediction to health management. This provides patients with timely and customized interventions, helping to slow disease progression and improving the intelligence and practical application value of healthcare services.

[0153] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0154] Based on the same inventive concept, this application also provides a system for implementing the aforementioned interpretable deep learning prediction method for Alzheimer's disease. The solution provided by this system is similar to the implementation described in the above method. Therefore, the specific limitations of one or more embodiments of an interpretable deep learning prediction system for Alzheimer's disease provided below can be found in the above-described limitations of the interpretable deep learning prediction method for Alzheimer's disease, and will not be repeated here.

[0155] In one exemplary embodiment, such as Figure 2 As shown, an interpretable deep learning prediction system for Alzheimer's disease is provided, which may include:

[0156] The acquisition module 810 is used to acquire multimodal data of the target object; the target object is the patient, and the multimodal data includes sMRI image data, clinical data, and lifestyle data.

[0157] The multimodal model module 820 is used to process multimodal data using a pre-trained multimodal deep fusion prediction model.

[0158] Furthermore, such as Figure 3 As shown, the multimodal model module 820 may include:

[0159] The feature extraction module 821 is used to extract features from multimodal data through the modal feature extraction layer to obtain multimodal features; the modal feature extraction layer includes 3D-CNN, MLP embedding layer, and XGBoost.

[0160] The fusion modeling module 822 is used to perform deep fusion of multimodal features through the fusion modeling layer to obtain a global fusion feature vector.

[0161] The result output module 823 is used to output AD risk level classification results and the probability of disease progression within a preset time period in parallel through the result output layer.

[0162] In one embodiment, the feature extraction module 821 may include:

[0163] The 3D-CNN unit is used to process sMRI image data using 3D-CNN to automatically learn the spatial structural information of sMRI image data and output brain region features.

[0164] The MLP embedding layer unit is used to transform clinical molecular data in clinical data using the MLP embedding layer, and output clinical feature vectors.

[0165] The XGBoost unit is used to perform feature filtering on lifestyle data using XGBoost to assess the importance of different lifestyle factors to the target population and output lifestyle features.

[0166] The feature integration unit is used to integrate brain region features, clinical feature vectors, and lifestyle features into multimodal features.

[0167] In one embodiment, the fusion modeling module 822 may include:

[0168] The Transformer unit is used to input multimodal features into the attention-based Transformer fusion model. The attention mechanism of the Transformer fusion model processes the multimodal features and dynamically assigns feature weights to the multimodal features by formula calculation.

[0169] The fusion unit is used to perform weighted aggregation of multimodal features based on feature weights to generate a global fusion feature vector.

[0170] In one embodiment, the result output module 823 may include:

[0171] The AD risk level unit is used to input the global fused feature vector into the classification task branch, and output the AD risk level classification result through the Softmax activation function.

[0172] The disease progression probability unit is used to input the globally fused feature vector into the regression task branch and output the disease progression probability within a preset time period through a linear output layer.

[0173] In one embodiment, the multimodal model module 820 may further include:

[0174] The data partitioning unit is used to divide the complete multimodal data used for training into training, validation, and test sets.

[0175] The preprocessing unit is used to preprocess the training set to obtain a balanced training set.

[0176] The iterative training unit is used to iteratively train the multimodal deep fusion prediction model using a balanced training set and validation set, thereby obtaining a new stage model after each iteration.

[0177] The model evaluation unit is used to evaluate the AUROC, accuracy, and high-risk recall of the staged model using the validation set.

[0178] The detection unit is used to use the current stage model as a pre-trained multimodal deep fusion prediction model when AUROC, accuracy, and high-risk recall reach a preset comprehensive optimal state.

[0179] In one embodiment, an interpretable Alzheimer's disease deep learning prediction system may further include:

[0180] The weight calculation module is used to calculate the weighted gradient of the AD risk level classification result relative to the activation map of the last convolutional layer in the 3D-CNN in the modality feature extraction layer.

[0181] The heatmap module is used to generate a 3D risk heatmap by using weighted gradients to sum the activation maps of convolutional layers and applying the ReLU activation function.

[0182] The risk factors module is used to extract a list of key risk factors from multimodal input data in the feature importance ranking based on the multimodal deep fusion prediction model.

[0183] The report integration module is used to integrate the 3D risk heat map, the list of key risk factors, and the AD risk level classification results and the probability of disease progression within a preset time period into a report dataset.

[0184] In one embodiment, an interpretable Alzheimer's disease deep learning prediction system may further include:

[0185] The profile combination module is used to extract AD risk level classification results and key risk factors from the quantitative risk report and combine them into risk profile facts.

[0186] The rules engine module is used to load risk profile facts into a rules engine that has a pre-configured intervention suggestion library. The rules engine compares the risk profile facts with the set of rules in the intervention suggestion library to select intervention rules.

[0187] The plan generation module is used to compile the specific intervention measures corresponding to the intervention rules, thereby generating a structured and personalized intervention plan.

[0188] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of an interpretable deep learning prediction method for Alzheimer's disease as described above.

[0189] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.

[0190] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0191] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.

Claims

1. An interpretable deep learning prediction method for Alzheimer's disease, characterized in that, The method includes: Acquire multimodal data of the target object; the target object is a candidate AD patient, and the multimodal data includes sMRI imaging data, clinical data, and lifestyle data; The multimodal data is processed using a pre-trained multimodal deep fusion prediction model to obtain AD risk level and disease progression probability; The step of using a pre-trained multimodal deep fusion prediction model to process the multimodal data to obtain AD risk level and disease progression probability includes: Multimodal features are obtained by extracting features from the multimodal data through a modal feature extraction layer; the modal feature extraction layer includes 3D-CNN, MLP embedding layer, and XGBoost. By using a fusion modeling layer, the multimodal features are deeply fused to obtain a global fusion feature vector; The global fusion feature vector is processed through the output layer to obtain the AD risk level classification result and the probability of disease progression within a preset time period.

2. The method according to claim 1, characterized in that, The process of extracting features from the multimodal data through a modal feature extraction layer to obtain multimodal features includes: The sMRI image data is processed using a 3D-CNN to automatically learn the spatial structural information of the sMRI image data and output brain region features; the 3D-CNN includes convolutional layers and fully connected layers, wherein the calculation of the convolutional layer is expressed by the following formula: in, For the input 3D image data, This represents a 3D convolution operation. The three-dimensional convolution kernel weight vector of the convolutional layer. The bias vector of the convolutional layer. It is the ReLU activation function; The clinical molecular data in the clinical data is transformed using an MLP embedding layer to output a clinical feature vector; XGBoost was used to perform feature filtering on the lifestyle data to assess the importance of different lifestyle factors to the target population and output lifestyle features. The brain region features, the clinical feature vectors, and the lifestyle features are integrated into multimodal features.

3. The method according to claim 1, characterized in that, The process of performing deep fusion of the multimodal features through a fusion modeling layer to obtain a global fusion feature vector includes: The multimodal features are input into a Transformer fusion model based on an attention mechanism. The attention mechanism of the Transformer fusion model processes the multimodal features, and feature weights are dynamically assigned to the multimodal features using the following formula: in, For query matrix; The key matrix; It is a value matrix; Let be the dimension of the key matrix; Used to normalize the calculated score to the feature weights; Based on the feature weights, the multimodal features are weighted and aggregated to generate the global fusion feature vector.

4. The method according to claim 1, characterized in that, The process of outputting AD risk level classification results and the probability of disease progression within a preset time period through the result output layer includes: The global fusion feature vector is input into the classification task branch, and the AD risk level classification result is output through the Softmax activation function. The global fusion feature vector is input into the regression task branch, and the probability of disease progression within the future preset time period is output through the linear output layer.

5. The method according to claim 1, characterized in that, The training steps of the multimodal deep fusion prediction model include: The complete multimodal data used for training is divided into a training set, a validation set, and a test set; The training set is preprocessed to obtain a balanced training set; The multimodal deep fusion prediction model is iteratively trained using the balanced training set and the validation set, thereby obtaining a new stage model after each iteration. The validation set was used to evaluate the AUROC, accuracy, and high-risk recall of the staged model. When the AUROC, accuracy, and high-risk recall reach a preset optimal state, the current stage model is used as the pre-trained multimodal deep fusion prediction model.

6. The method according to claim 1, characterized in that, The method further includes: Calculate the weighted gradient of the AD risk level classification result relative to the activation map of the last convolutional layer in the 3D-CNN in the modality feature extraction layer; Using the weighted gradient, the activation map of the convolutional layer is weighted and summed, and the ReLU activation function is applied to generate a three-dimensional risk heat map. Based on the feature importance ranking of the multimodal deep fusion prediction model, a list of key risk factors is extracted from the multimodal input data; The three-dimensional risk heatmap, the list of key risk factors, the AD risk level classification results, and the probability of disease progression within a preset future time are integrated into a report dataset. A quantitative risk report is generated based on the aforementioned report dataset.

7. The method according to claim 6, characterized in that, The method further includes: Based on the quantitative risk report, the AD risk level classification results and the key risk factors are extracted and combined to form a risk profile fact. The risk profile facts are loaded into a rule engine that has a pre-configured intervention suggestion library. The rule engine compares the risk profile facts with the rule set in the intervention suggestion library to select intervention rules. The specific intervention measures corresponding to the intervention rules are compiled to generate a structured, personalized intervention plan.

8. An interpretable deep learning prediction system for Alzheimer's disease, characterized in that, The system includes: The acquisition module is used to acquire multimodal data of a target object; the target object is a patient, and the multimodal data includes sMRI image data, clinical data, and lifestyle data. The multimodal model module is used to process the multimodal data using a pre-trained multimodal deep fusion prediction model. The processing procedure is as follows: Multimodal features are obtained by extracting features from the multimodal data through a modal feature extraction layer; the modal feature extraction layer includes 3D-CNN, MLP embedding layer, and XGBoost. By using a fusion modeling layer, the multimodal features are deeply fused to obtain a global fusion feature vector; The results output layer outputs AD risk level classification results and the probability of disease progression within a preset time period in parallel.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.