Brain amyloid protein deposition prediction tool suitable for early Alzheimer patient group
By developing a portable prediction tool that combines blood markers and clinically available variables, the problem of predicting amyloid deposition status in early Alzheimer's patients was solved, achieving high accuracy, low cost and low invasive prediction effects.
Patent Information
- Application Number
- CN202411968303.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art has difficulties in accurately predicting the state of brain amyloid deposition in early Alzheimer's patients, especially in the absence of effective solutions in non-invasive, low-cost and readily accessible methods.
Develop a portable prediction tool that combines blood markers and clinically available variables. Through multi-center data acquisition and processing, recursive feature elimination methods are used to select feature, build accurate models and simple models, use machine learning algorithms to train and evaluate models, and realize online prediction through web APP.
Accurate prediction of Aβ pathological status in patients with early Alzheimer's disease is achieved, reducing the cost and invasiveness of the detection, improving the accuracy and generalization of the prediction model, and is suitable for different medical conditions and application scenarios.
Smart Images

Figure CN120072283A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of neural computing modeling, and particularly relates to a prediction tool for cerebral amyloid deposition applicable to the early Alzheimer's disease population. Background Art
[0002] Alzheimer's Disease (AD) is a neurodegenerative disease characterized by progressive cognitive decline and neuronal loss. With the significant upward trend in the prevalence of AD, it has brought a heavy burden to society and families. The deposition of amyloid-β (Aβ) plaques is one of the main pathological features of AD and is considered to play a key role in the occurrence and progression of the disease. Therefore, accurately detecting the Aβ pathological state is of great significance for the early diagnosis, treatment decision-making, and prognosis evaluation of AD. However, the current gold standard methods for detecting the Aβ pathological state in clinical practice are mainly cerebrospinal fluid (CSF) analysis and positron emission tomography (PET). These methods have the following limitations: 1. High invasiveness: CSF detection requires lumbar puncture, which may cause adverse reactions such as headache and infection, and the acceptance rate of patients is low, especially in the elderly population; 2. High cost: PET scan equipment is costly, and radioactive tracers are required, so the examination cost is expensive and it is not suitable for large-scale screening and routine clinical applications; 3. Resource limitation: PET scan and CSF detection require professional equipment and technical personnel, and many primary medical institutions cannot provide these detection means, resulting in the risk that patients may face delayed diagnosis and treatment.
[0003] With the approval of anti-Aβ drugs such as lecanemab in China, there is an urgent need for an accurate, non-invasive, and easily accessible method to predict the Aβ pathological status of patients in order to timely screen out patients suitable for treatment. Therefore, developing a portable prediction tool that combines blood biomarkers and clinically available variables has important clinical significance and application value. In recent years, researchers have begun to explore the use of blood-based biomarkers (BBMs) such as plasma phosphorylated tau protein 181 (p-tau181) and the plasma Aβ42 / 40 ratio to predict the Aβ pathological status. These biomarkers have the advantages of being non-invasive, low-cost, and easily accessible, and have been proven to have high accuracy in AD diagnosis. However, using blood biomarkers or clinical evaluation indicators alone still faces challenges in clinical applications. For example, the test results may be affected by factors such as individual differences, comorbidities, and differences in testing platforms. In addition, current research mainly focuses on single centers, lacking national multi-center data support, and there is a lack of a dedicated prediction model for early AD (including mild cognitive impairment [MCI] and mild dementia) populations. And this part of the population is the main target for anti-Aβ treatment.
[0004] Currently, there is a lack of a precise prediction tool for brain Aβ deposition status in early AD populations based on national multi-center research using BBMs in China. The present invention hopes to precisely predict the brain amyloid protein deposition status in early AD populations based on national multi-center research using BBMs and clinically available variables. Summary of the Invention
[0005] In view of the above-mentioned defects and problems, the present invention provides a prediction tool for brain amyloid protein deposition applicable to early Alzheimer's disease populations, aiming to accurately predict the Aβ pathological status of early AD patients by using non-invasive blood biomarkers and clinically easily accessible variables.
[0006] The solution adopted by the present invention to solve its technical problems is as follows: A prediction tool for brain amyloid protein deposition applicable to early Alzheimer's disease populations, and the specific operation plan is as follows: (1) Data collection and processing: Collect data of MCI and mild dementia patients (Clinical Dementia Rating [CDR]=0.5) from multiple centers, and preprocess the data. The key variables collected include plasma p-tau181, plasma Aβ42 / 40 ratio, APOEε4 gene status, and MoCA score; (2) Feature Selection and Model Construction: Apply the recursive feature elimination method on the training set, and through ten-fold repeated cross-validation, identify the most important features for predicting Aβ positivity. Then, construct models based on the selected features and variables. The constructed models include an exact model and a parsimonious model; (3) Model Training and Evaluation. The specific operation steps are as follows: 1) Data Partitioning: Randomly divide the data into a training set and a validation set to ensure the independence of model training and testing; 2) Model Training: In the training set, construct a prediction model and use five-fold repeated cross-validation to determine the most suitable model; 3) Model Evaluation: On the validation set, calculate the sensitivity, specificity, accuracy, and the area under the receiver operating characteristic curve of the model. By comparing the predicted probability with the actual results, determine the optimal classification threshold; 4) External Validation: Conduct external validation on the model to evaluate the generalization ability and stability of the model; 5) Subgroup Analysis: Evaluate the model performance for different subgroups to ensure the applicability of the model in different populations; (4) Statistical Analysis: The model is evaluated using sensitivity, specificity, accuracy, and AUC metrics; (5) Develop a Web APP: Users input variable information through the APP. The system calculates the probability of a patient being Aβ positive based on the input data and prompts the user's risk level according to the set risk threshold, realizing the online deployment and real-time prediction functions of the model.
[0007] Further, in operation step (1), the data preprocessing methods include: Missing Value Handling: Use the multiple imputation method to handle missing values. For binary variables, use logistic regression imputation ('logreg'), and for multi-class variables, use polynomial regression imputation ('polyreg'); Outlier Handling: Remove outliers that are more than four standard deviations above the mean to avoid significant impacts on the accuracy and performance of the model.
[0008] Further, the basis for feature selection is the importance of variables and their availability in clinical practice. The recursive feature elimination strategy is used for feature selection. The specific configuration is as follows: I. Basic Algorithm: Use random forest as the basic algorithm (functions = rfFuncs); II. Validation Method: 10-fold cross-validation (number = 10), repeated 10 times (repeats = 10); III. Feature Evaluation Range: Evaluate the number of features decreasing from 11 to 1 in each iteration; Ⅳ. Evaluation Metrics: Accuracy, Kappa coefficient, and its AccuracySD are used as performance metrics for evaluation; Ⅴ. Screening Criteria: Screening criteria are formulated by combining prediction performance, stability, clinical availability, and practical value; Ⅵ. Feature Determination: Two optimal feature combinations are finally determined through RFE analysis and clinical availability evaluation: - Precise Model Feature Set: Plasma p-tau181, Plasma Aβ42 / 40, MoCA, APOE4; - Simplified Model Feature Set: Plasma p-tau181, MoCA, APOE4.
[0009] Furthermore, the model construction method is as follows: Precise Model: Incorporate APOEε4 gene status, plasma p-tau181, plasma Aβ42 / 40 ratio, and MoCA score, and construct it using a generalized linear model, which is applicable to medical institutions with complete detection conditions; Parsimonious Model: Based on APOE ε4 gene status, plasma p-tau181, and MoCA score, construct it using the gradient boosting machine algorithm, which is applicable to scenarios with limited detection conditions.
[0010] Furthermore, in step (iii), five machine learning algorithms are used for model construction and training: random forest, support vector machine, generalized linear model, LASSO regression, and gradient boosting machine to construct a prediction model.
[0011] Furthermore, the screening criteria for feature selection are as follows: ① Results of feature selection: Determine that the optimal number of features is 4 variables; ② Optimal model performance: Accuracy = 0.7495, Kappa = 0.2304118, AccuracySD = 0.06591; ③ Evaluation based on feature importance scores, the features with the highest importance are: Plasma p-tau181, importance score ≈ 13; Plasma Aβ42 / 40, importance score ≈ 4; MoCA, importance score ≈ 4; APOE4, importance score ≈ 2; Hippocampal atrophy, importance score ≈ 1.5.
[0012] Furthermore, during the statistical analysis process, all statistical analyses are performed in R version 4.3, and statistical significance is based on a two-sided p-value less than 0.05; multiple imputation is used to handle missing values to ensure data integrity; recursive feature elimination is used for feature selection, and ten-fold repeated cross-validation is used to evaluate the stability of the model.
[0013] Furthermore, the technical implementation of the web APP includes: Development environment: Based on the Shiny package of the R language, an interactive web application is built to achieve the online deployment and use of the model. Front-end design: Using the UI components of Shiny, an intuitive and user-friendly interface is created to facilitate users to input variables and view results. Backend logic: Integrate the prediction model on the server side, process user input in real time, calculate the probability of Aβ positivity, and provide the prediction results. The functions of the web APP include: Data input: Users can input information such as the APOE ε4 gene status, plasma p-tau181 concentration, and MoCA score. Result output: The system calculates the probability of Aβ positivity in patients in real time based on the input data, and prompts the user's risk level according to the set risk threshold. Data security: The HTTPS protocol is adopted to ensure the security of data transmission, and the user input data is not saved to protect user privacy.
[0014] Advantages of the present invention: The prediction tool model provided by the present invention is based on BBMs (such as p-tau181 and Aβ42 / 40 ratio), and combines common and highly accessible clinical indicators such as APOE ε4 gene status and MoCA score to predict the brain Aβ status, fully considering the factors that may affect the BBMs results in clinical practice, further improving the accuracy of the prediction model and its generalizability in the real world; the model uses blood tests and cognitive assessments, which have the advantages of lower cost and invasiveness compared to CSF and PET tests, thus having higher clinical practicability and generalizability; and provides an accurate model and a parsimonious model to adapt to different medical conditions and application scenarios, improving the practicability and flexibility of the tool. The web prediction tool provided by the present invention is the first portable prediction tool for brain Aβ deposition status in China. Compared with the prior art, the present invention has obvious advantages in terms of non-invasiveness, cost, and patient acceptance, and the prediction accuracy is higher than that of a single-index model, reducing the risk of misdiagnosis and missed diagnosis. Description of the Drawings
[0015] Figure 1 It is the English interface of the entry page of the web APP; Figure 2 It is the English interface of the model-related information of the web APP; Figure 3 It is the English interface of the prediction example of the APM1 model in the web APP; Figure 4 It is the English interface of the prediction example of the APM2 model in the web APP; Figure 5For the web APP to enter the Chinese interface of the page; Figure 6 For the Chinese interface of the information related to the web APP model; Figure 7 For the Chinese interface of the prediction example of the APM1 model in the web APP; Figure 8 For the Chinese interface of the prediction example of the APM2 model in the web APP. Specific implementation manner
[0016] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0017] Embodiment 1, the present invention provides a prediction tool for the Aβ pathological state of Alzheimer's disease, aiming to use non-invasive blood markers and clinically accessible variables to accurately predict the Aβ pathological state of early AD patients. The specific technical solutions are as follows: (1) Data collection and processing: 1) Multi-center data collection: Collect clinical data from four hospital research centers across the country, a total of 334 cases of MCI (mild cognitive impairment) and mild dementia patients (dementia rating scale [CDR]=0.5). The considered features include 11 key variables: - Demographic characteristics: age group, gender, BMI, education level; - Cognitive assessment indicators: MMSE score, MoCA score; - Biomarkers: gene marker APOEε4 status, plasma marker p-tau181 status, Aβ42 / 40 ratio status, NFL status; - Imaging features: hippocampal atrophy status.
[0018] The present invention uses blood markers to replace traditional CSF and PET detections, reducing the pain and risks of patients and improving the acceptance of patients. And by integrating blood markers, genetic information and cognitive assessment, the accuracy of prediction is improved and the limitations of single indicators are reduced.
[0019] 2) Data preprocessing: Ⅰ. Missing value processing: Use the multiple imputation method to process missing values. For binary variables, use logistic regression imputation ('logreg'), and for multi-class variables, use polynomial regression imputation ('polyreg'). This can effectively utilize the existing data and reduce the loss of sample size; Ⅱ. Outlier processing: Remove outliers that are more than four standard deviations above the mean to avoid significant impacts on the accuracy and performance of the model.
[0020] (2) Feature selection and model construction: 1) Feature Selection: Apply the Recursive Feature Elimination (RFE) method on the training set. Through ten-fold repeated cross-validation, identify the most important features for predicting Aβ positivity. The basis for feature selection is the importance of variables and their availability in clinical practice.
[0021] Adopt the Recursive Feature Elimination (RFE) strategy for feature selection, with the specific configuration as follows: Ⅰ. Base Algorithm: Use random forest as the base algorithm (functions = rfFuncs). The random forest algorithm is implemented through the following steps: ① Construct a random forest model: - Build multiple decision trees based on the training data, with the default being 500 trees. The construction process for each decision tree: - Use the Bootstrap method to sample from the training set to construct the training set for the tree. - For each node split, consider m randomly selected features, where m ≈ √p and p is the total number of features. ② Feature Importance Evaluation: - First, record the reduction in impurity for each node split in each decision tree. - Then calculate the Gini index: Gini(t) = 1 - Σ(pi)², where pi is the proportion of samples of the i-th class in node t. - Next, accumulate the importance scores of each feature in the tree. ③ Recursive Feature Elimination Process: Based on the feature evaluation results of the random forest: - Remove the feature with the lowest importance. - Retrain the model using the remaining feature subset (sizes = c(1:11)). - Through ten-fold cross-validation, repeat the performance evaluation 10 times. ④ Feature Subset Evaluation: For each subset with a specific number of features (1 - 11): - Record the average performance of cross-validation. - Evaluate the model stability through 10 repeated evaluations. - Select the optimal feature combination.
[0022] Ⅱ. Validation Method: Ten-fold cross-validation (number = 10), repeated 10 times (repeats = 10).
[0023] Ⅲ. Feature Evaluation Range: Evaluate the features from 11 to 1 in decreasing order. In each iteration: - Construct a random forest model to evaluate the current feature set. - Remove the least important feature based on the importance score. - Retrain the model using the remaining features.
[0024] Ⅳ. Evaluation Metrics: ① Use accuracy, Kappa coefficient, and its standard deviation (AccuracySD) as performance metrics to track the performance metrics of each feature subset: - Accuracy: 0.7311 - 0.7495, -Kappa: -0.0005587 - 0.2604998, -AccuracySD: 0.02886 - 0.06954, -KappaSD: 0.005587 - 0.196812; ② Evaluate the model stability through cross - validation; ③ Record the performance changes under different numbers of features.
[0025] Ⅴ. Screening criteria: Combining prediction performance, stability, clinical availability and practical value, the following screening criteria are formulated: ① The result of feature selection is to determine that the optimal number of features is 4 variables; ② Optimal model performance: Accuracy = 0.7495, Kappa = 0.2304118, AccuracySD = 0.06591; ③ Based on the evaluation of feature importance scores, the highest - importance features are: Plasma p - tau181 (importance score ≈ 13), Plasma Aβ42 / 40 (importance score ≈ 4), MoCA (importance score ≈ 4), APOE4 (importance score ≈ 2), Hippocampal atrophy (importance score ≈ 1.5).
[0026] Ⅵ. Feature determination: Through RFE analysis and clinical availability evaluation, two optimal feature combinations are finally determined: - Precise model feature set: Plasma p - tau181, Plasma Aβ42 / 40, MoCA, APOE4; - Simplified model feature set: Plasma p - tau181, MoCA, APOE4.
[0027] 2) Model construction, providing a dual - model strategy, developed two complementary prediction models: a precise model (APM1) and a parsimonious model (APM2), meeting the needs under different medical resource conditions and taking into account both high accuracy and ease of use.
[0028] Ⅰ. Precise model (APM1): Incorporate the APOE ε4 gene status, Plasma p - tau181, Plasma Aβ42 / 40 ratio and MoCA score, integrate all selected biomarkers, and construct using the Generalized Linear Model (GLM), applicable to medical institutions with complete detection conditions; Ⅱ. Parsimonious model (APM2): Based on the APOE ε4 gene status, Plasma p - tau181 and MoCA score, use a reduced feature set and construct using the Gradient Boosting Machine (GBM) algorithm, applicable to scenarios with limited detection conditions.
[0029] (III) Model Training and Evaluation: 1) Data Partitioning: Randomly divide the data into a training set and a validation set at a ratio of 8:2 to ensure the independence of model training and testing. Use stratified sampling to ensure the consistency of the sample distribution of each category.
[0030] 2) Model Training: In the training set, use five machine learning algorithms, namely random forest, support vector machine, generalized linear model, LASSO regression, and gradient boosting machine, to build a prediction model. Use five-fold repeated cross-validation to determine the most suitable model. The model construction and training process are as follows: Ⅰ. Cross-validation details of all models: ① Training control parameters (trainControl): -method = "repeatedcv": Repeated cross-validation, -number = 5: Five-fold cross-validation, -repeats = 5: Repeat 5 times, -search = "grid": Use grid search to optimize parameters, -sampling = "rose": Use ROSE to handle class imbalance, -classProbs = TRUE: Calculate class probabilities, -summaryFunction = twoClassSummary: Use binary classification evaluation metrics; ② Handling Class Imbalance: Use the ROSE (Random Over-Sampling Examples) sampling method. Sampling steps: - First, apply ROSE in each cross-validation fold. - Then, generate synthetic samples to balance positive and negative classes, maintaining the distribution characteristics of the original data. - Finally, dynamically adjust the class ratio during training.
[0031] Ⅱ. Comprehensively use five machine learning algorithms to train the model: ① Random Forest Model: - Model construction parameters: Learning task: Binary classification (amyloid_positive); Evaluation metric: Area under the ROC curve; Build the model using all predictor variables; - Cross-validation settings: Adopt five-fold cross-validation, repeat 5 times, use grid search to optimize parameters, and set a random seed to ensure reproducibility.
[0032] ② Gradient Boosting Machine: - Model construction parameters: Learning task: binary classification (amyloid_positive); Evaluation metric: area under the ROC curve; Use the default GBM parameter configuration; - Algorithm implementation details: Initialize the model with the mean of the training data for initial prediction; Iterative boosting process: Calculate the residuals of the current model, then train a new decision tree to fit the residuals, update the model through the step size (learning rate), and repeat the above steps until the specified number of iterations is reached; The final prediction combines the prediction results of all weak learners.
[0033] ③ Logistic regression model: - Model construction parameters: Learning task: binary classification (amyloid_positive); Evaluation metric: area under the ROC curve; Use the binomial distribution family (family = "binomial"); - Algorithm implementation details: Fit the logistic function through maximum likelihood estimation: P(Y = 1|X) = 1 / (1 + e^(-(β 0 + Σβᵢxᵢ))); Then perform risk prediction interpretation: Calculate the odds ratio for each feature (OR = exp(βᵢ)), evaluate the contribution of the feature to the prediction result, and intuitively explain the influence direction and degree of each feature through the magnitude of the βᵢ coefficient.
[0034] ④ LASSO regression: - Model construction parameters: Based on the generalized linear model framework; Use the binomial distribution family (family = "binomial"); alpha = 1 is specified for LASSO regression; - Regularization parameter setting: The lambda value grid search range is 0.001 to 0.1; Set 10 evenly distributed candidate values; Select the optimal lambda value through cross-validation; - Feature selection mechanism: Achieve feature sparsity through L1 regularization, automatically compress the coefficients of unimportant features to 0, and retain the features that have the most influence on the prediction.
[0035] ⑤ Support vector machine: - Model construction parameters: Kernel function: Use the radial basis kernel function (RBF, Radial Basis Function); Evaluation metric: area under the ROC curve; Parameter optimization length: 10 (tuneLength = 10); - Kernel function implementation: RBF kernel function: K(x,x') = exp(-γ||x - x'||²); Automatically optimize the kernel function parameter γ and the penalty parameter C; Determine the optimal parameter combination through a 10-point grid search.
[0036] 3) Model evaluation: On the validation set, calculate the sensitivity, specificity, accuracy, and the area under the receiver operating characteristic curve (AUC) of the model. By comparing the predicted probabilities with the actual results, determine the optimal classification threshold; - Main evaluation metrics: The AUC value and its 95% confidence interval, sensitivity, specificity, positive predictive value, negative predictive value, and overall accuracy under the optimal threshold; - Secondary evaluation metrics: Model complexity, computational efficiency, and clinical interpretability.
[0037] 4) External validation: Use independent samples from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database to externally validate the model and evaluate its generalization ability and stability. Validation strategies: - Internal validation: Five-fold cross-validation, - Evaluation of performance stability under multiple random seeds, - Threshold optimization analysis based on the ROC curve.
[0038] 5) Subgroup analysis: Conduct model performance evaluation for different subgroups, such as gender, comorbidities, etc., to ensure the applicability of the model in different populations. The following are the prediction performances of the model in different genders and different comorbidities: In the male subgroup: APM1: AUC = 0.86 [95%CI 0.71 - 1]; APM2: AUC = 0.69 [95%CI 0.48 - 0.91]; In the female subgroup: APM1: AUC = 0.95 [95%CI 0.87 - 1]; APM2: AUC = 0.9 [95%CI 0.8 - 1]; Carrying one or more comorbidities (hypertension, diabetes, stroke, cardiovascular disease, dyslipidemia): APM1: AUC = 0.83 [95%CI 0.68 - 0.99]; APM2: AUC = 0.81 [95%CI 0.6 - 0.1].
[0039] (IV) Statistical analysis methods: All statistical analyses were performed in R version 4.3. Statistical significance was defined as a two-sided p-value less than 0.05. Multiple imputation was used to handle missing values to ensure data integrity. Recursive feature elimination (RFE) was used for feature selection, and ten-fold repeated cross-validation was used to evaluate the stability of the model. The model was evaluated using metrics such as sensitivity, specificity, accuracy, and AUC.
[0040] Experimental verification and statistical analysis results: 2) Model performance metrics: - Accurate Prediction Model (APM1): On the validation set, the AUC reached 0.92, the accuracy was 88%, and both the sensitivity and specificity were 88%. - Parsimonious Prediction Model (APM2): The AUC was 0.85, the accuracy was 79%, the sensitivity was 75%, and the specificity was 80%. - External Validation: In the external validation using the ADNI database, the AUC of the APM1 model was 0.86, and the AUC of the APM2 model was 0.84, verifying the stability and generalization ability of the models. - Subgroup Analysis: For different subgroups such as gender, APOE ε4 status, and comorbidities, the models all showed high predictive performance, demonstrating the applicability of the models.
[0041] The model prediction tool provided by the present invention is trained and validated based on data from multiple national centers, has good generalization ability, and can be applied to different regions and populations. The above detailed parameter settings and feature selection processes illustrate the rigor and innovation of the present invention in technical implementation. At the same time, the selection of these parameter settings has been strictly experimentally verified to ensure the stability and reliability of the model in practical applications.
[0042] Example 2: Please refer to Figure 1-8 , on the basis of Example 1, this example provides a prediction tool for the Aβ pathological state of Alzheimer's disease based on a web application (Web APP), with the function of developing a web APP. It innovatively uses the Shiny package of the R language to develop the web application, without the need to install additional software, and users can directly use it online, which is convenient and fast.
[0043] Technical implementation part of the web APP development: 1) Development environment: Based on the Shiny package of the R language, an interactive web application is constructed to realize the online deployment and use of the model. 2) Front-end design: Using the UI components of Shiny, an intuitive and user-friendly interface is created to facilitate users to input variables and view results. 3) Back-end logic: Integrate the prediction model on the server side, process user input in real time, calculate the probability of Aβ positivity, and provide the prediction result. Function implementation part: 1) Data input: Users can input information such as APOE ε4 gene status, plasma p-tau181 concentration, optional plasma Aβ42 / 40 ratio, and MoCA score. 2) Result output: The system calculates the probability of Aβ positivity of the patient in real time according to the input data, and prompts the user's risk level (low risk or high risk) according to the set risk threshold. 3) Data security: The HTTPS protocol is adopted to ensure the security of data transmission. User input data is not saved to protect user privacy.
[0044] Both the above-mentioned model and the web APP are provided in two languages, Chinese and English, and users can switch the languages according to their needs.
[0045] Example 3: This example introduces the current application status and application prospects of the brain amyloid deposition prediction tool provided by the present invention for the early Alzheimer's population.
[0046] 1) Technical feasibility and practicality: Technical maturity: The adopted blood detection technology and cognitive assessment methods are both mature and easy to implement in clinical practice; User feedback: The preliminary application results show that clinicians have given positive evaluations on the accuracy and convenience of the tool, and the acceptance rate of patients is high.
[0047] 2) Technical expansion and future development: Model optimization: In the future, more biomarkers (such as GFAP, p-tau181), genetic and environmental factors can be added to further improve the prediction ability of the model; Mobile application: Develop a mobile application or mini-program, and other web APPs, mini-programs or other software with different UI interfaces can also be built to facilitate doctors and patients to use it anytime and anywhere, and improve the popularity of the tool; Artificial intelligence integration: The present invention can introduce advanced algorithms such as deep learning to continuously optimize the model performance and adapt to more complex data and application requirements.
[0048] 3) Industrialization prospects and social benefits: Clinical promotion: The prediction tool provided by the present invention can cooperate with hospitals, physical examination centers and primary medical institutions for promotion and application to improve the early diagnosis rate of AD; Scientific research cooperation: The present invention can provide an efficient subject screening tool for pharmaceutical companies and research institutions, and accelerate the R & D and clinical trial processes of anti-AD drugs; Social benefits: Through early screening and intervention, the burden on patients and families can be reduced, and the social medical cost can be lowered, with significant social and economic benefits.
[0049] In addition, based on the solution provided by the present invention, preventive and diagnostic measures for cognitive-related diseases other than Alzheimer's disease can also be developed, and the relationship between the prediction results of this model and the disease progression can also be developed.
[0050] The above are only the preferred embodiments of the present invention and do not limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A tool for predicting brain amyloid deposition in people with early Alzheimer's disease, characterized in that: The specific operation plan is as follows: (I) Data collection and processing: Data from MCI and mild dementia patients from multiple centers were collected and preprocessed. The key variables collected included plasma p-tau181, plasma Aβ42 / 40 ratio, APOEε4 gene status, and MoCA score; (ii) Feature selection and model construction: Recursive feature elimination was applied to the training set, and the most important features for predicting Aβ positivity were identified through ten-fold repeated cross-validation. Model construction was then carried out based on the selected features and variables. Taking economic benefits into consideration, two models were constructed, including an exact model and a parsimonious model. (III) Model training and evaluation. The specific steps are as follows: 1) Data partitioning: Randomly divide the data into training set and validation set to ensure that the training and testing of the model are independent of each other; 2) Model training: In the training set, a prediction model is constructed and the most suitable model is determined by using five-fold repeated cross validation; 3) Model evaluation: On the validation set, the sensitivity, specificity, accuracy, and area under the receiver operating characteristic curve of the model were calculated, and the optimal classification threshold was determined by comparing the predicted probability with the actual results; 4) External validation: external validation of the model to evaluate the generalization ability and stability of the model; 5) Subgroup analysis: Model performance evaluation is performed on different subgroups to ensure the applicability of the model in different populations; (IV) Statistical analysis: The model was evaluated using sensitivity, specificity, accuracy, and AUC indicators; (V) Development of a web APP: Users input variable information through the APP. The system calculates the probability of the patient being Aβ positive based on the input data and prompts the user's risk level based on the set risk threshold, thereby realizing the online deployment and real-time prediction function of the model.
2. A brain amyloid protein deposition prediction tool suitable for early Alzheimer's disease population according to claim 1, characterized in that: In addition to the above four variables, the key variables collected also included: age group, gender, BMI, education level, MMSE score, NFL status, and hippocampal atrophy status.
3. A brain amyloid protein deposition prediction tool suitable for early Alzheimer's disease population according to claim 1, characterized in that: In operation step (1), the data preprocessing method includes: Missing value processing: Multiple imputation methods were used to handle missing values, using logistic regression imputation ('logreg') for binary variables and polynomial regression imputation ('polyreg') for multi-class variables; Outlier processing: Outliers that exceed four standard deviations from the mean are removed to avoid significant impact on the accuracy and performance of the model.
4. A brain amyloid protein deposition prediction tool suitable for early Alzheimer's disease population according to claim 1 or 2, characterized in that: The basis for feature selection is the importance of variables and their availability in clinical practice. A recursive feature elimination strategy is used for feature selection. The specific configuration is as follows: Ⅰ. Basic algorithm: Use random forest as the basic algorithm (functions=rfFuncs); Ⅱ. Validation method: 10-fold cross validation (number=10), repeated 10 times (repeats=10); III. Feature evaluation range: the number of features is evaluated from 11 to 1 in each iteration; IV. Evaluation indicators: Accuracy, Kappa coefficient and its AccuracySD are used as performance indicators for evaluation; V. Screening criteria: The screening criteria were formulated based on the predictive performance, stability, clinical availability and practical value; VI. Feature determination: Two optimal feature combinations are finally determined through recursive feature elimination analysis and clinical usability evaluation: - Accurate model feature set: plasma p-tau181, plasma Aβ42 / 40, MoCA, APOE4; - Reduced model feature set: plasma p-tau181, MoCA, APOE4.
5. A brain amyloid protein deposition prediction tool suitable for early Alzheimer's disease population according to claim 1 or 3, characterized in that: The model is constructed as follows: Precise model: APOEε4 gene status, plasma p-tau181, plasma Aβ42 / 40 ratio and MoCA score were included and constructed using a generalized linear model. It is suitable for medical institutions with complete testing conditions. The minimalist model was constructed based on APOE ε4 gene status, plasma p-tau181, and MoCA score using a gradient boosting machine algorithm and was suitable for scenarios with limited testing conditions.
6. The brain amyloid protein deposition prediction tool suitable for early Alzheimer's disease population according to claim 1, characterized in that: The model construction and training in step (iii) adopt five machine learning algorithms: random forest, support vector machine, generalized linear model, LASSO regression, and gradient boosting machine to construct the prediction model.
7. The tool for predicting brain amyloid deposition in early-stage Alzheimer's disease patients according to claim 4, characterized in that: The screening criteria for feature selection are: ① The result of feature selection: the optimal number of features is determined to be 4 variables; ②Optimal model performance: Accuracy=0.7495, Kappa=0.2304118, AccuracySD=0.06591; ③Based on the evaluation of feature importance scores, the most important features are: Plasma p-tau181, importance score ≈ 13; Plasma Aβ42 / 40, importance score ≈ 4; MoCA, importance score ≈4; APOE4, importance score ≈ 2; Hippocampal atrophy, importance score ≈ 1.
5.
8. The tool for predicting brain amyloid deposition in early-stage Alzheimer's disease patients according to claim 1, characterized in that: During the statistical analysis, all statistical analyses were performed in R version 4.3, and statistical significance was based on a two-sided p value less than 0.
05. Multiple imputation was used to handle missing values to ensure data integrity; Recursive feature elimination was used for feature selection, and ten-fold repeated cross validation was used to assess the stability of the model.
9. The tool for predicting brain amyloid deposition in early-stage Alzheimer's disease patients according to claim 1, characterized in that: The technical implementation of the web APP includes: Development environment: Based on the Shiny package of R language, we build interactive web applications to realize online deployment and use of models; Front-end design: Use Shiny UI components to create an intuitive and friendly user interface to facilitate users to input variables and view results; Backend logic: Integrate the prediction model on the server side, process user input in real time, calculate the probability of Aβ positivity, and provide prediction results; The features of the web app include: Data input: Users can input information such as APOE ε4 gene status, plasma p-tau181 concentration, and MoCA score; Result output: The system calculates the probability of the patient being Aβ positive in real time based on the input data, and prompts the user's risk level based on the set risk threshold; Data security: HTTPS protocol is used to ensure the security of data transmission, user-entered data is not saved, and user privacy is protected.