Interpolation method for missing data in Alzheimer disease progress prediction

By integrating feature selection and conditional variational autoencoder (CVAE) with patient characteristics and time intervals, the impact of missing data in AD disease progression prediction is addressed, achieving high-precision disease progression prediction and the formulation of personalized treatment plans.

CN120656625APending Publication Date: 2025-09-16DONGHUA UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510719776.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing AD disease progression prediction methods are based on cross-sectional data and cannot effectively predict disease progression or provide early intervention guidance. Traditional interpolation methods fail to fully consider the conditional dependence of AD data, resulting in missing data affecting model accuracy.

Method used

An integrated feature selection module was used to screen out features closely related to disease progression, and a conditional variational autoencoder (CVAE) was used to combine patient-specific individual characteristics and time intervals to interpolate missing values ​​and construct a dynamic interpolation model.

Benefits of technology

It has significantly improved the accuracy and reliability of predicting the progression of Alzheimer's disease, can accurately predict the disease progression of patients, and provide a scientific basis for early diagnosis and personalized treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656625A_ABST
    Figure CN120656625A_ABST
Patent Text Reader

Abstract

The invention relates to an interpolation method for missing data in Alzheimer's disease progress prediction, which comprises an integrated feature selection module and a missing value interpolation module, and is characterized in that the integrated feature selection module adopts an isomorphic and heterogeneous integrated feature selection method for time heterogeneity and phenotypic heterogeneity respectively to screen out features closely related to disease progress; the missing value interpolation module obtains irregular missing time sequence information by calculating a real interval between follow-up visit data of each patient, takes specific individual features of the patient as intervention condition input of a variational auto-encoder, and learns conditional distribution of the features by using a conditional variational auto-encoder; according to the method, more refined conditional distribution modeling is carried out, and finally, missing data is generated by the model, so that the problems that the distribution characteristics of data are not fully considered and conditional dependence information behind a missing mechanism is difficult to capture in the current traditional interpolation method based on AD patient multi-source longitudinal data are solved, and the accuracy and reliability of Alzheimer's disease progress prediction are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a missing value interpolation method, and in particular to a missing value interpolation method for AD data of a distance-conditional variational autoencoder (CVAE). Background Art

[0002] Alzheimer's disease (AD) is a progressive neurodegenerative disorder that develops gradually. Clinically, AD has a long course and is generally divided into three stages: cognitively normal (CN), mild cognitive impairment (MCI), and AD. Currently, there are no medications or clinical treatments that can reverse or cure the disease. Clinical research widely recognizes that a deep understanding of disease progression is crucial for early diagnosis and intervention of AD, which helps slow further progression. Therefore, predicting a patient's future disease progression based on clinical diagnostic data is of great significance to AD research.

[0003] Current predictions of AD progression are mostly based on cross-sectional data, which cannot effectively predict disease progression or provide guidance for early intervention. To address these challenges, using computer technology to construct disease progression models based on longitudinal data has become a hot topic in current research.

[0004] Longitudinal data refers to the collection of disease-related information, such as neuropsychological assessment scores, brain imaging data, biomarker data, and clinical examination data, through repeated measurements of the same group of patients at multiple time points. Comprehensive analysis of this multi-source data provides a more comprehensive understanding of the disease's evolution, enabling more accurate predictions of a patient's future cognitive state and the stage of disease progression. Compared to traditional cross-sectional data, longitudinal data can reflect cognitive trends, reveal underlying patterns in disease progression, and provide more forward-looking clinical decision support.

[0005] Due to the irregular examination time of patients and the inconsistency of diagnostic items, AD patients have missing data at different time points. The reasons for the missing data may include the doctor's diagnosis and treatment decision and the patient's privacy considerations. The missing data of multi-source longitudinal diagnostic data of AD patients are not completely random, but conditional non-random missing data, that is, the probability of missing data is not only related to the data characteristics themselves, but also affected by external conditions such as the patient's medical decision-making and clinical symptoms. This irregularity undoubtedly brings greater complexity to the construction of accurate prediction models. At present, traditional static interpolation methods fail to fully consider the distribution characteristics of the data and it is difficult to capture the conditional dependency information behind the missing mechanism. Although the dynamic interpolation method based on variational auto-encoders (VAE) can dynamically interpolate from the perspective of capturing the potential distribution of data, it assumes that the development and changes of all data features are unconditionally random and identically distributed, and fails to fully consider the conditional dependency of the missing values ​​of different feature values ​​of AD data, thereby weakening the model's ability to capture key features. Summary of the Invention

[0006] To address the problem that traditional interpolation methods based on multi-source longitudinal data from Alzheimer's patients fail to fully consider the distributional characteristics of the data and struggle to capture the conditional dependencies underlying missing mechanisms, this paper first employs an integrated feature selection method to process high-dimensional data, ensuring the extraction of key features closely related to disease progression. Then, to address the issue of missing patient data, an interpolation method for missing data in Alzheimer's disease progression prediction is proposed.

[0007] Constructing a method to interpolate missing data for predicting Alzheimer's disease progression not only addresses the limitations of existing clinical assessments but also better reflects the dynamic evolution of the disease. By integrating multiple data sources, including neuropsychology, imaging, and biomarkers, and combining them with advanced computing techniques and predictive models, it is possible to accurately predict a patient's cognitive state, identify potential high-risk patients in advance, and thus provide a scientific basis for early intervention. The implementation of such a system will provide strong support for the early diagnosis of Alzheimer's disease, monitoring of disease progression, and the development of personalized treatment plans.

[0008] The technical solution of the present invention is:

[0009] A missing data interpolation method for predicting the progression of Alzheimer's disease includes two modules: an integrated feature selection module and a missing value interpolation module. First, the integrated feature selection module uses homogeneous and heterogeneous integrated feature selection methods for temporal heterogeneity and phenotypic heterogeneity, respectively, to screen out features closely related to disease progression and delete redundant information. Second, the missing value interpolation module obtains irregular missing time series information by calculating the true interval between each patient's follow-up data, and combines the patient's specific individual characteristics as the intervention condition input of the variational autoencoder, and uses the conditional variational autoencoder (CVAE) to learn the conditional distribution of features. By performing more refined conditional distribution modeling of the patient's time-dependent information and individual difference characteristics during the disease progression process, the final model generates three sets of outputs, namely, a complete eigenvector matrix, an uncertainty probability matrix, and a masking matrix. The complete eigenvector matrix is ​​the complete feature matrix after interpolation.

[0010] Furthermore, the integrated feature selection module optimizes the feature subset of the input data by reducing the dimension and filtering the features. The specific implementation is as follows:

[0011] A homogeneous ensemble feature selection method was used: first, the dataset was divided into three sub-datasets according to the disease development stage; then, the Shapley Additive Feature Interpretation method (SHAP value) was used to calculate the importance score of each feature in each sub-dataset to quantify the contribution of each feature to the model prediction; by calculating the contribution of each feature to the model prediction results, the relative importance of the feature was quantified;

[0012] Adopting heterogeneous ensemble feature selection method: treating all features as a whole dataset, using three feature selection methods: random forest, relief filtering method, and counterfactual feature selection method to evaluate feature importance;

[0013] Finally, the feature selection results of homogeneous and heterogeneous integration are normalized, and the recursive feature elimination method (RFE) is used to iteratively remove unimportant features to select the optimal feature subset.

[0014] Furthermore, the specific implementation of the intervention condition input in the missing value interpolation module is as follows:

[0015] Assume that the patient's visit sequence matrix is ​​expressed as:

[0016]

[0017] in, represents the j-th eigenvalue of patient n at time point t, t represents the patient's visit time, and j represents the value of the examination items performed; in addition, the patient's conditional feature matrix is ​​defined as:

[0018]

[0019] in, represents the kth conditional feature value of patient n at time point t;

[0020] We use feature importance to perform scoring and sorting, and select two key features with significantly higher scores as the basis for time interval calculation. For each patient's data, we check the missing status of the two key features and generate a masked Boolean matrix Bool. Each row indicates whether the two features are missing at that time point. If the feature is missing at a certain time point, the corresponding position is True, otherwise it is non-missing (False). The formula is as follows:

[0021]

[0022] Then, the mask matrix is ​​used to calculate the time interval between the current time point and the previous valid time point with non-missing features, and then a time interval label (D) is generated to capture the contextual relationship of the time series. The formula is as follows:

[0023]

[0024] Among them, t last Indicates the last non-missing time point; if at least one of the features of ADAS11 (Alzheimer's disease assessment scale-cognitive 11-item) and RAVLT_immediate (Rey's Auditory Verbal Learning Test immediate) at the current time point is non-missing, the data at that time point is valid and the time interval is 0; if both features are missing, the time interval from the current time point to the last non-missing time point is calculated;

[0025] Next, the true time interval is used as an additional condition and combined with the patient's personal characteristic information as a conditional variable, and input into the variational autoencoder for interpolation, thereby obtaining the missing completion inspection indicator matrix and its corresponding uncertainty.

[0026] Furthermore, in the missing value interpolation module, for missing data in the sub-rows of the original data, before the data is input into the encoder for learning, the missing values ​​are filled with -1, and a masking matrix M is generated accordingly. (n) To accurately record the location of missing values, the dimension of the masking matrix is ​​the same as the visit sequence matrix consistent.

[0027] Furthermore, in the missing value interpolation module, the model cooperates with the encoder and decoder to jointly complete the learning of latent variables and the generation of observation data;

[0028] The role of the encoder is to map the input data to the latent space, that is, the observed feature data and conditional features As input, it generates the distribution parameters of the latent variable: mean and log variance The formula is as follows:

[0029]

[0030] Among them, f encode It is the mapping function of the encoder network, which extracts and transforms the input data;

[0031] Then, we use the reparameterization technique to sample latent variables from the latent distribution The formula is as follows:

[0032]

[0033] Among them, the standard deviation ε is from the standard normal distribution Noise is used to introduce randomness;

[0034] The goal of the decoder is to utilize the latent variables and conditional features Reconstruct input features and output reconstructed feature values The formula is as follows:

[0035]

[0036] Among them, f decode is the mapping function of the decoder network, mapping the latent variables back to the original data space.

[0037] Furthermore, in the actual reconstruction process, if there are missing values ​​in the input features, The decoder will directly generate feature values And fill in the corresponding missing positions in the original matrix, the formula is as follows:

[0038]

[0039] Furthermore, in the missing value interpolation module, the optimization goal of VAE is to maximize the evidence lower bound ELBO; by maximizing the marginal likelihood of the observed data To affect the accuracy of latent variable learning and missing value generation; the ELBO definition formula is as follows:

[0040]

[0041] The first term is the regularization of the latent variable distribution, which measures the posterior distribution q φ With the prior distribution p θ The KL divergence between is used to limit the distribution of latent variables to match the standard normal distribution as much as possible to avoid overfitting; the second term is the reconstruction error, which represents the decoder generation value and the true value similarity; among them Indicates that the encoder is in the posterior distribution q φ Expectation under ; further bring the two into the following formula:

[0042]

[0043] Among them, l represents the dimension of the latent variable, L represents the maximum dimension; σ 2 is the reconstruction error variance of the decoder; finally, the model maximizes It is equivalent to minimizing the loss function of VAE to learn a better data representation, thereby generating high-quality reconstruction results. The formula of the loss function is as follows:

[0044]

[0045] Furthermore, in the missing value interpolation module, while generating data, the uncertainty measure p is additionally calculated. t , ensuring a better learning strategy for the model.

[0046] Furthermore, the complete eigenvector matrix is ​​the complete feature matrix after interpolation; the uncertainty probability matrix represents the probability of the accuracy of the interpolated value; and the mask matrix identifies which data is missing and which time points need to be estimated.

[0047] The beneficial effects of the present invention are:

[0048] The present invention can effectively process high-dimensional clinical data, neuroimaging data and cognitive assessment data of Alzheimer's patients, reduce the impact of missing data on the prediction model, and thus significantly improve the accuracy of disease progression prediction. Based on the integrated feature selection module of Alzheimer's Disease (AD) heterogeneity and the AD data missing value interpolation module based on the distance-conditional variational autoencoder (CVAE), the present invention can capture the complex disease progression pattern of Alzheimer's patients, and is particularly suitable for nonlinear and non-periodic changes with large individual differences, avoiding the limitations of traditional methods. At the same time, the present invention provides a data-driven early diagnosis longitudinal data monitoring and prediction system, which can accurately warn of disease development through comprehensive analysis of the patient's longitudinal data, and help clinicians formulate personalized treatment plans. Ultimately, the prediction results of the present invention help to improve the early diagnosis and intervention effect of Alzheimer's disease, reduce misdiagnosis and missed diagnosis, and provide a scientific basis for subsequent treatment and clinical decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 This is a general structural diagram of the missing data interpolation method for predicting the progression of Alzheimer's disease according to the present invention;

[0050] Figure 2 This is the basic framework diagram of the integrated feature selection module based on AD heterogeneity in the present invention;

[0051] Figure 3 This is the D-CVAE calculation flow chart of the present invention. DETAILED DESCRIPTION

[0052] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0053] A missing data interpolation method for predicting the progression of Alzheimer's disease is designed to solve the problems of time-dependent modeling, heterogeneous data processing, and missing data processing in multi-source longitudinal high-dimensional data. The solution includes two key modules: an integrated feature selection module based on AD heterogeneity (referred to as the integrated feature selection module) and an AD data missing value interpolation method based on a conditional variational autoencoder with time intervals (referred to as the missing value interpolation module). Figure 1First, the integrated feature selection module optimizes the feature subset of the input data by reducing the dimension and filtering features, thereby improving the model performance. Second, the missing value interpolation module obtains irregular missing time series information by calculating the true interval between each patient's follow-up data, and combines key demographic information (such as gender) as the intervention condition input of the variational autoencoder to generate missing data. The introduction of time interval labels enables the generated samples to consider the contextual time series information of the features and effectively capture irregular time series information. This method accurately predicts disease progression by modeling time dependency and nonlinear features. The overall solution significantly improves the accuracy and reliability of Alzheimer's disease progression prediction by effectively fusing multi-source data, processing missing data, and modeling time dependency and nonlinear features.

[0054] Integrated feature selection module: The process of this method is as follows Figure 2 As shown in the figure, homogeneous and heterogeneous ensemble feature selection methods are used to target temporal heterogeneity and phenotypic heterogeneity, respectively, to screen out features closely related to disease progression while removing redundant information. This approach not only overcomes the problem that single feature selection methods cannot fully capture the complex relationships between disease features and tend to retain excessive redundant features, but also fully accounts for the heterogeneous nature of AD disease data.

[0055] The temporal heterogeneity of AD manifests itself in patients progressing through different disease stages. Symptoms and manifestations vary within each stage, affecting examination focus and leading to variations in examination items and data collection. This can lead to irregular data sampling, which in turn increases the complexity of missing data. For example, in the early stages of the disease, certain features may be more predictive, while in later stages, other features may be more critical. Therefore, it is necessary to consider and identify the important features of different disease stages. To address this, the present invention employs a homogeneous ensemble feature selection method. First, the dataset is divided into three sub-datasets based on disease stage. Then, for each sub-dataset, the Shapley Additive Explanations (SHAP) method is used to calculate the importance score of each feature, thereby quantifying its contribution to model prediction. This method quantifies the relative importance of each feature by calculating its contribution to the model's predictions. This interpretable method helps identify which features have the greatest impact on predictions at different disease stages, which features' importance changes over the course of the disease, and which features remain relatively stable.

[0056] Phenotypic heterogeneity in AD manifests itself in the fact that even within the same disease stage, individual differences can lead to varying clinical manifestations and progression rates. This variability can also lead to irregularities in patient data sampling, resulting in missing data. Specifically, patients with mild MCI may undergo extensive testing to investigate the cause or treatment efficacy, while patients with mild cognitive impairment (MCI) may choose to forgo treatment, resulting in lost test data. Therefore, when mining features related to disease stage, it is important to consider the importance of features across the entire data set. To address this, we employ a heterogeneous ensemble feature selection approach. This approach treats all features as a single dataset and assesses feature importance using three feature selection methods: random forest, relevant features (relief), and counterfactual. Random forest is a decision tree-based approach suitable for capturing nonlinear relationships between AD features. The relief algorithm assesses the contribution of features to sample classification and is suitable for data with imbalanced classes. Counterfactual generation enhances causal analysis by generating counterfactual samples to assess the contribution of features to the predicted outcome, facilitating causal analysis in subsequent modeling. These methods can capture the complex nonlinear relationships between features from different perspectives of classification, regression, and causality. Their combined use can more comprehensively evaluate the importance of features and increase the flexibility of analysis.

[0057] Finally, the feature selection results of homogeneous and heterogeneous integration are normalized, and recursive feature elimination (RFE) is used to iteratively remove unimportant features to select the optimal feature subset.

[0058] Missing value interpolation module: In the actual clinical diagnosis process, due to various emergencies, patients' subjective abandonment and other factors, Alzheimer's disease data usually have a large proportion of missing data, and the time intervals are irregular. The severity of the disease varies from patient to patient, and the review time is also different accordingly, which causes the examination data to usually show non-completely random missing characteristics with certain conditions. There are three situations in which data is missing: one is that the data of the entire time point is missing, resulting in irregular sampling intervals; the second is that there is data at each time point, but due to asynchronous data collection, there are asynchronous sampling intervals; the third is a combination of the above two types of missing, resulting in irregular missing. Therefore, when interpolating missing values, not only conditional dependence should be considered, but also the interpolation data generated from the distribution should be guaranteed to have a certain degree of randomness. To solve this problem, the present invention proposes an AD data missing value interpolation method based on a conditional variational autoencoder based on time intervals. The method introduces patient-specific individual characteristics and true time intervals as conditional variables, and uses CVAE to learn the conditional distribution of features. By performing more refined conditional distribution modeling of the patient's time-dependent information and individual difference characteristics during the development of the disease, the accuracy of interpolation is significantly improved. The method flow is as follows: Figure 3 .

[0059] Assume that the patient's visit sequence matrix is ​​expressed as:

[0060]

[0061] in, represents the jth eigenvalue of patient n at time point t, where t represents the patient's visit time and j represents the value of the examination items performed. In addition, the patient's conditional feature matrix is ​​defined as:

[0062]

[0063] in, represents the kth conditional feature value of patient n at time point t.

[0064] Using the resulting feature importance ranking, we select two key features with significantly higher scores as the basis for time interval calculation. For each patient's data (i.e., six rows of data at six time points), we check for missingness of the two key features and generate a masked Boolean matrix (Bool). Each row indicates whether the two features are missing at that time point. If a feature is missing at a certain time point, the corresponding position is True; otherwise, it is non-missing (False). The formula is as follows:

[0065]

[0066] Then, the mask matrix is ​​used to calculate the time interval between the current time point and the previous valid time point with non-missing features, and then a time interval label (D) is generated to capture the contextual relationship of the time series. The formula is as follows:

[0067]

[0068] Among them, t last Represents the last non-missing time point. If at least one of the ADAS11 (Alzheimer's disease assessment scale-cognitive 11-item) and RAVLT_immediate (Rey's Auditory Verbal Learning Test immediate) features at the current time point is non-missing, the data at that time point is valid, and the time interval is 0. If both features are missing, the time interval from the current time point to the last non-missing time point is calculated.

[0069] Next, we added the true time interval as an additional condition, combined it with the patient's personal characteristics as a conditional variable, and fed it into the variational autoencoder for interpolation. This yielded a missing-completion inspection indicator matrix and its corresponding uncertainty. This process enabled the model to account for irregular time intervals during interpolation. The conditional variational autoencoder also generated interpolated data that relied on the input conditions while retaining randomness, effectively improving the diversity of the interpolated data and the model's generalization capabilities.

[0070] Since there are missing data in the sub-rows of the original data, before the data is input into the encoder for learning, the missing values ​​are filled with -1 and a masking matrix M is generated accordingly. (n) To accurately record the location of missing values, the dimension of the masking matrix is ​​the same as the visit sequence matrix consistent.

[0071] In order to interpolate unknown potential missing information, the model cooperates with the encoder and decoder to jointly complete the learning of latent variables and the generation of observation data.

[0072] The role of the encoder is to map the input data to the latent space, that is, the observed feature data and conditional features As input, it generates the distribution parameters of the latent variable: mean and log variance The formula is as follows:

[0073]

[0074] Among them, f encode It is the mapping function of the encoder network, which extracts features and transforms the input data.

[0075] Then, we use the reparameterization technique to sample latent variables from the latent distribution The formula is as follows:

[0076]

[0077] Among them, the standard deviation ε is from the standard normal distribution The noise is used to introduce randomness. It separates the randomness in the sampling process from the parameters of the neural network, so that during VAE training, the gradient can be smoothly transferred back to the parameters of the encoder, thereby achieving model optimization based on the backpropagation algorithm.

[0078] The goal of the decoder is to utilize the latent variables and conditional features Reconstruct input features and output reconstructed feature values The formula is as follows:

[0079]

[0080] Among them, f decode is the mapping function of the decoder network, mapping the latent variables back to the original data space.

[0081] In the actual reconstruction process, if there are missing values ​​in the input features (i.e. ), the decoder will directly generate the feature value And fill in the corresponding missing positions in the original matrix, the formula is as follows:

[0082]

[0083] In order to optimize the model, the optimization goal of VAE is to maximize the evidence lower bound (ELBO). To affect the accuracy of latent variable learning and missing value generation. The ELBO definition formula is as follows:

[0084]

[0085] The first term is the regularization of the latent variable distribution, which measures the posterior distribution q φ With the prior distribution p θ The KL divergence between is used to limit the distribution of latent variables to match the standard normal distribution as much as possible to avoid overfitting. The second term is the reconstruction error, which represents the decoder generation value and the true value The similarity of Indicates that the encoder is in the posterior distribution q φ Expectation under . Further substituting the two into the following formula:

[0086]

[0087] Among them, l represents the dimension of the latent variable, L represents the maximum dimension; σ 2 is the reconstruction error variance of the decoder; finally, the model maximizes It is equivalent to minimizing the loss function of VAE to learn a better data representation, thereby generating high-quality reconstruction results. The formula of the loss function is as follows:

[0088]

[0089] In addition, considering the randomness of interpolated data, there will be a certain degree of uncertainty. However, the prediction model usually assumes that all data are completely accurate and cannot directly obtain this information. Therefore, when generating data, we additionally calculate its uncertainty measure p t , ensuring a better learning strategy for the model.

[0090] The final model generates three sets of outputs: a complete eigenvector matrix, an uncertainty probability matrix, and a masking matrix. The complete eigenvector matrix is ​​the complete feature matrix after interpolation. The uncertainty probability matrix represents the probability of the accuracy of the interpolated value. The masking matrix identifies which data are missing and which time points need to be estimated. This method introduces patient-specific individual characteristics and true time intervals as conditional variables and uses CVAE to learn the conditional distribution of features. By modeling the conditional distribution of time-dependent information and individual differences in patients during disease progression, the accuracy of interpolation is significantly improved.

[0091] The above-described embodiment merely represents one embodiment of the present invention. While the description is relatively specific and detailed, it should not be construed as limiting the scope of the patent. It should be noted that a person skilled in the art would be able to make various modifications and improvements without departing from the spirit of the present invention, and these modifications and improvements fall within the scope of protection of the present invention. Therefore, the scope of protection of the patent for this invention shall be determined by the appended claims.

Claims

1. A method for interpolating missing data in predicting the progression of Alzheimer's disease, characterized in that: It includes two modules: integrated feature selection module and missing value interpolation module. First, the integrated feature selection module adopts homogeneous and heterogeneous integrated feature selection methods for temporal heterogeneity and phenotypic heterogeneity respectively to screen out features closely related to disease progression and delete redundant information at the same time; secondly, the missing value interpolation module obtains irregular missing time series information by calculating the true interval between the follow-up data of each patient, and combines the patient's specific individual characteristics as the intervention condition input of the variational autoencoder, and uses the conditional variational autoencoder CVAE to learn the conditional distribution of features; by performing more refined conditional distribution modeling of the patient's time-dependent information and individual difference characteristics during the disease development process, the final model generates three sets of outputs, namely, the complete feature vector matrix, the uncertainty probability matrix, and the masking matrix.

2. The missing data interpolation method for predicting the progression of Alzheimer's disease according to claim 1, characterized in that: The integrated feature selection module optimizes the feature subset of the input data by reducing the dimension and filtering the features. The specific implementation is as follows: A homogeneous ensemble feature selection method was used: first, the dataset was divided into three sub-datasets according to the disease development stage; then, the Shapley Additive Feature Interpretation method (SHAP) value was used to calculate the importance score of each feature in each sub-dataset to quantify the contribution of each feature to the model prediction; Quantify the relative importance of each feature by calculating its contribution to the model's prediction results; Adopting heterogeneous ensemble feature selection method: treating all features as a whole dataset, using three feature selection methods: random forest, relief filtering method, and counterfactual feature selection method to evaluate feature importance; Finally, the feature selection results of homogeneous and heterogeneous integration are normalized, and the recursive feature elimination method (RFE) is used to iteratively remove unimportant features to select the optimal feature subset.

3. The missing data interpolation method for predicting the progression of Alzheimer's disease according to claim 1, characterized in that: The specific implementation of intervention condition input in the missing value interpolation module is as follows: Assume that the patient's visit sequence matrix is ​​expressed as: in, represents the j-th eigenvalue of patient n at time point t, t represents the patient's visit time, and j represents the value of the examination items performed; in addition, the patient's conditional feature matrix is ​​defined as: in, represents the kth conditional feature value of patient n at time point t; We use feature importance to perform scoring and sorting, and select two key features with significantly higher scores as the basis for time interval calculation. For each patient's data, we check the missing status of the two key features and generate a masked Boolean matrix Bool. Each row indicates whether the two features are missing at that time point. If the feature is missing at a certain time point, the corresponding position is True, otherwise it is non-missing (False). The formula is as follows: Then, the mask matrix is ​​used to calculate the time interval between the current time point and the previous valid time point with non-missing features, and then a time interval label (D) is generated to capture the contextual relationship of the time series. The formula is as follows: Among them, t last Indicates the last non-missing time point; if at least one of the features in ADAS11 and RAVLT_immediate at the current time point is non-missing, the data at that time point is valid and the time interval is 0; if both features are missing, the time interval from the current time point to the last non-missing time point is calculated; Next, the true time interval is used as an additional condition and combined with the patient's personal characteristic information as a conditional variable, and input into the variational autoencoder for interpolation, thereby obtaining the missing completion inspection indicator matrix and its corresponding uncertainty.

4. The missing data interpolation method for predicting the progression of Alzheimer's disease according to claim 3, characterized in that: In the missing value interpolation module, if there is missing data in the sub-rows of the original data, before the data is input into the encoder for learning, the missing values ​​are filled with -1, and a masking matrix M is generated accordingly. (n) To accurately record the location of missing values, the dimension of the masking matrix is ​​the same as the visit sequence matrix H (n) consistent.

5. The missing data interpolation method for predicting the progression of Alzheimer's disease according to claim 1, characterized in that: In the missing value interpolation module, the model works together through the encoder and decoder to complete the learning of latent variables and the generation of observation data; The role of the encoder is to map the input data to the latent space, that is, the observed feature data and conditional features As input, it generates the distribution parameters of the latent variable: mean and log variance The formula is as follows: Among them, f encode It is the mapping function of the encoder network, which extracts and transforms the input data; Then, we use the reparameterization technique to sample latent variables from the latent distribution The formula is as follows: Among them, the standard deviation ε is from the standard normal distribution Noise is used to introduce randomness; The goal of the decoder is to utilize the latent variables and conditional features Reconstruct input features and output reconstructed feature values The formula is as follows: Among them, f decode is the mapping function of the decoder network, mapping the latent variables back to the original data space.

6. The method for interpolating missing data in predicting the progression of Alzheimer's disease according to claim 5, characterized in that: In the actual reconstruction process, if there are missing values ​​in the input features, The decoder will directly generate feature values And fill in the corresponding missing positions in the original matrix, the formula is as follows:

7. The missing data interpolation method for predicting the progression of Alzheimer's disease according to claim 1, characterized in that: In the missing value interpolation module, the optimization goal of VAE is to maximize the evidence lower bound ELBO; by maximizing the marginal likelihood of the observed data To affect the accuracy of latent variable learning and missing value generation; the ELBO definition formula is as follows: The first term is the regularization of the latent variable distribution, which measures the posterior distribution q φ With the prior distribution p θ The KL divergence between is used to limit the distribution of latent variables to match the standard normal distribution as much as possible to avoid overfitting; the second term is the reconstruction error, which represents the decoder generation value and the true value similarity; among them Indicates that the encoder is in the posterior distribution q φ The expectation under the above equations is: Among them, l represents the dimension of the latent variable, L represents the maximum value of the latent variable dimension; σ 2 is the reconstruction error variance of the decoder; finally, the model maximizes It is equivalent to minimizing the loss function of VAE to learn a better data representation, thereby generating high-quality reconstruction results. The formula of the loss function is as follows:

8. The missing data interpolation method for predicting the progression of Alzheimer's disease according to claim 1, characterized in that: In the missing value interpolation module, while generating data, its uncertainty measure p is additionally calculated. t , ensuring a better learning strategy for the model.

9. The missing data interpolation method for predicting the progression of Alzheimer's disease according to claim 1, characterized in that: The complete eigenvector matrix is ​​the complete feature matrix after interpolation; the uncertainty probability matrix represents the probability of the accuracy of the interpolated value; the mask matrix identifies which data is missing and which time points need to be estimated.