Machine learning-based MCI and its evolution into a dementia prediction system

By combining machine learning methods with community population data, predictive factors were screened and a predictive model for the progression of MCI to dementia was constructed. This solved the problems of small sample size and data imbalance, and enabled accurate prediction of the progression of MCI to dementia in the community environment, providing an opportunity for early intervention.

CN119314687BActive Publication Date: 2025-10-31SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411326809.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-23
Publication Date
2025-10-31
Estimated Expiration
2044-09-23

AI Technical Summary

Technical Problem

Existing technologies for predicting the progression of mild cognitive impairment (MCI) to dementia suffer from problems such as small sample size, data imbalance, and lack of feature importance ranking. Furthermore, brain imaging and cerebrospinal fluid biomarker data are difficult to obtain in community settings.

Method used

Using machine learning methods, combined with non-brain imaging and non-cerebrospinal fluid biomarkers in the community population, and through epidemiological nested case-control design, time-series causal inference of directed acyclic graphs, and machine learning algorithms, predictive factors were screened to construct a predictive model for the progression of MCI to dementia. The model was trained using UK Biobank data from a large population cohort study.

Benefits of technology

It enables accurate prediction of the progression of MCI to dementia in a community setting, improves the interpretability and predictive accuracy of the model, provides opportunities for early intervention, and reduces the risk of developing dementia.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119314687B_ABST
    Figure CN119314687B_ABST
Patent Text Reader

Abstract

This invention discloses a machine learning-based system for predicting mild cognitive impairment (MCI) and its progression to dementia. The machine learning-based MCI prediction system includes: selecting predictive factors from a population known to have mild cognitive impairment by screening candidate factors; using this population as a dataset and dividing it into a training set and a test set; inputting the training set into all machine learning models to train them, resulting in a pre-trained machine learning model; inputting the test set into the pre-trained machine learning model to test it, and using the machine learning model whose test index value exceeds a set threshold as the final trained MCI machine learning model; and inputting several predictive factors of the subject to be tested into the final trained MCI machine learning model to obtain a prediction result indicating whether the subject has mild cognitive impairment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of disease prediction technology, and in particular to machine learning-based MCI and a predictive system for the progression of MCI to dementia. Background Technology

[0002] Mild cognitive impairment (MCI) is a disorder that affects cognitive abilities such as memory, language, and attention, and is often considered a precursor to dementia. With an aging population, the prevalence of MCI and dementia is projected to rise, making early detection and intervention crucial. Approximately 38% of MCI patients progress to Alzheimer's disease after 5 years or more of follow-up. Early detection of MCI and accurate prediction of its progression to dementia can provide opportunities for intervention, slowing the progression of cognitive decline and delaying the onset of dementia.

[0003] Previous models for predicting MCI suffer from problems such as small sample size, unresolved data imbalance, and lack of feature importance ranking. Furthermore, MCI outcome evaluation uses screening tools to identify high-risk groups, rather than relying on hospital diagnostic data. To address these issues, it is necessary to consider the proportion of physician-diagnosed MCI patients and non-patients in a large sample population, and to address data imbalance before constructing the predictive model. Additionally, ranking the predictor factors by feature importance enhances the model's interpretability.

[0004] Previous models predicting the progression of mild cognitive impairment (MCI) to dementia have primarily relied on magnetic resonance imaging (MRI) data or a combination of MRI data and cerebrospinal fluid (CSF) biomarker data. However, models using imaging or CSF biomarker data are not suitable for predicting the risk of MCI progressing to dementia in community populations because obtaining such data in a community setting is challenging. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a machine learning-based system for predicting MCI progression to dementia and considering alternative methods to predict which MCI patients in the community have a higher risk of dementia. For example, other types of data can be utilized, such as cognitive assessments, family history, medical history, genetic markers, sociodemographic characteristics, and lifestyle factors. These data are more readily available in a community setting and can help to more accurately predict the risk of MCI patients progressing to dementia, thereby providing more effective interventions.

[0006] On the one hand, a machine learning-based predictive system for mild cognitive impairment (MCI) is provided, including:

[0007] The predictor screening module is configured to: obtain predictors by screening candidate factors in a population with known mild cognitive impairment. The screening process includes: initially screening candidate factors related to mild cognitive impairment from several candidate factors in the population with mild cognitive impairment; then, using the variance inflation factor to perform a second screening from the initially screened candidate factors; and finally, using the Lasso regression algorithm to perform a third screening from the second-screened candidate factors to obtain several predictors.

[0008] The dataset construction module is configured to: use people with known mild cognitive impairment as the dataset, and divide the dataset into training and testing sets according to a set ratio;

[0009] The model training and testing module is configured to: input the training set into all machine learning models, train the models, and obtain the pre-trained machine learning models; input the test set into the pre-trained machine learning models, test the models, and use the machine learning models whose test index values ​​exceed the set threshold as the final trained mild cognitive impairment machine learning models.

[0010] The prediction module is configured to input several predictive factors of the subject to be tested into the finally trained mild cognitive impairment machine learning model to obtain a prediction result of whether the subject to be tested has mild cognitive impairment.

[0011] On the other hand, a machine learning-based predictive system for the progression of MCI to dementia is provided, including:

[0012] The predictor screening module is configured to: obtain predictor factors by screening candidate factors in a population known to have progressed from mild cognitive impairment to dementia. The screening process includes: initially screening candidate factors related to the progression from mild cognitive impairment to dementia from a number of candidate factors known to have progressed from mild cognitive impairment to dementia; then, using a variance inflation factor to perform a second screening from the initially screened factors; and finally, using the Boruta algorithm to perform a third screening from the second-screened factors to obtain the predictor factors.

[0013] The dataset construction module is configured to: take people with known mild cognitive impairment who have progressed to dementia as the dataset, and divide the dataset into training set and test set according to a set ratio;

[0014] The model training and testing module is configured to: input the training set into all machine learning models to train the models and obtain the pre-trained machine learning models; input the test set into the pre-trained machine learning models to test the models; and take the machine learning models whose test index values ​​exceed the set threshold as the final trained machine learning models for mild cognitive impairment progressing to dementia.

[0015] The prediction module is configured to input several predictive factors of the test subject who already has mild cognitive impairment into the finally trained machine learning model for progression from mild cognitive impairment to dementia, and obtain the prediction result of the test subject's progression from mild cognitive impairment to dementia.

[0016] The above technical solution has the following advantages or beneficial effects:

[0017] This invention uses deep learning on a training set to obtain a final trained machine learning model for mild cognitive impairment. Based on this model, it can provide an accurate prediction of whether a subject has mild cognitive impairment.

[0018] This invention utilizes data from the large-scale population cohort study UK Biobank (UKB) to construct a predictive model for the progression of microencephalopathy (MCI) to dementia based on non-brain imaging and non-cerebrospinal fluid biomarkers in community populations. By combining epidemiological nested case-control design, time-series causal inference using directed acyclic graphs, and machine learning algorithms, predictive factors for MCI progression to dementia are identified and a predictive model is developed. Attached Figure Description

[0019] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0020] Figure 1 Here is a flowchart of the predictor selection process in Example 1;

[0021] Figure 2 This is a schematic diagram of the data plan in Example 1;

[0022] Figure 3(a) shows the Receiver Operating Characteristic (ROC) curves used to predict MCI in the training set of Example 1;

[0023] Figure 3(b) shows the ROC curve of the cross-validation set prediction MCI in Example 1;

[0024] Figure 3(c) shows the ROC curve of the CatBoost model predicting MCI in the test set of Example 1;

[0025] Figure 3(d) shows the calibration curve of the CatBoost model predicting MCI in Example 1;

[0026] Figure 3(e) shows the evaluation metrics of the models for predicting MCI in Example 1;

[0027] Figure 3(f) is a summary diagram of the SHAP of the top 10 features of MCI predicted by the CatBoost model in Example 1;

[0028] Figure 4(a) shows the receiver operating characteristic (ROC) curves of subjects in the training set of Example 1 used to predict the progression of MCI to dementia;

[0029] Figure 4(b) shows the ROC curves for cross-validation set prediction of MCI progressing to dementia in Example 1;

[0030] Figure 4(c) shows the ROC curves of the RF model predicting the progression of MCI to dementia in the test set of Example 1;

[0031] Figure 4(d) is a calibration curve of the RF model in Example 1 predicting the progression of MCI to dementia;

[0032] Figure 4(e) shows the evaluation indicators of each model for predicting the progression of MCI to dementia in Example 1;

[0033] Figure 4(f) is a SHAP summary plot of the top 10 features used by the RF model in Example 1 to predict the progression of MCI to dementia; *Each data point in the SHAP plot represents a participant, and the color indicates the magnitude of the predictor. The horizontal range represents the overall predictive power of each predictor (the larger the range, the greater the contribution to the model). The magnitude and direction on the x-axis can be used to determine the probability of the outcome event occurring (the right side indicates a higher probability of experiencing the outcome event).

[0034] Figure 5(a) shows an example of an application platform for predicting MCI based on the CatBoost model in Example 1;

[0035] Figure 5(b) shows an example of an application platform for predicting MCI based on the CatBoost model in Example 1;

[0036] Figure 5(c) shows an example of an application platform for predicting MCI based on the CatBoost model in Example 1;

[0037] Figure 6(a) shows an example of the application platform for predicting the progression of MCI to dementia based on the Random Forest model in Example 1;

[0038] Figure 6(b) shows an example of the application platform for predicting the progression of MCI to dementia based on the Random Forest model in Example 1;

[0039] Figure 6(c) is an example of the application platform for predicting the progression of MCI to dementia based on the Random Forest model in Example 1. Detailed Implementation

[0040] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0041] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0042] All data acquisition in this embodiment is carried out in accordance with laws and regulations and with user consent, and the data is used legally.

[0043] Example 1

[0044] This embodiment provides an MCI prediction system based on machine learning;

[0045] like Figure 1 As shown, a machine learning-based mild cognitive impairment (MCI) prediction system includes:

[0046] The predictor screening module is configured to: obtain predictors by screening candidate factors in a population with known mild cognitive impairment. The screening process includes: initially screening candidate factors related to mild cognitive impairment from a number of candidate factors in the population with mild cognitive impairment; then, using the variance inflation factor to perform a second screening from the initially screened candidate factors; and finally, using the Lasso regression algorithm to perform a third screening from the second-screened candidate factors to obtain a number of predictors.

[0047] The dataset construction module is configured to: use people with known mild cognitive impairment as the dataset, and divide the dataset into training and testing sets according to a set ratio;

[0048] The model training and testing module is configured to: input the training set into all machine learning models, train the models, and obtain the pre-trained machine learning models; input the test set into the pre-trained machine learning models, test the models, and use the machine learning models whose test index values ​​exceed the set threshold as the final trained mild cognitive impairment machine learning models.

[0049] The prediction module is configured to input several predictive factors of the subject to be tested into the finally trained mild cognitive impairment machine learning model to obtain a prediction result of whether the subject to be tested has mild cognitive impairment.

[0050] Furthermore, several candidate factors for the aforementioned mild cognitive impairment population include:

[0051] Category 1: Genetic factors, including: the number of APOE e4 allele carriers and ABO blood type;

[0052] Category 2: Early life factors, including birth weight and breastfeeding;

[0053] The third category: demographic factors, including age, sex, race, education level, income level, and the Thomson Deprivation Index;

[0054] Category Four: Midlife Environmental Factors, which include: PM2.5 2.5 Pollution, PM 10 Pollution, NO2 pollution, noise pollution (L den );

[0055] Category 5: Lifestyle factors in middle age, including: smoking, drinking, physical activity, sedentary time, sleep duration, and diet;

[0056] Category 6: Midlife biological factors, which include: Body mass index (BMI), grip strength, diastolic blood pressure, systolic blood pressure, low-density lipoprotein, glycated hemoglobin level, insulin use, and vitamin D level.

[0057] Category 7: Social and psychological factors in middle age, including: depression, tension, anxiety, loneliness, and social isolation;

[0058] Category 8: Medical History; This medical history includes: disability, cardiovascular disease, metabolic disease, respiratory disease, digestive disease, musculoskeletal disease, and cancer; Disability includes hearing aid use, hearing loss, and physical weakness; cardiovascular disease includes coronary heart disease, stroke, heart failure, and hypertension; metabolic disease includes diabetes, obesity, and dyslipidemia; respiratory disease includes influenza and pneumonia, chronic obstructive pulmonary disease, and asthma; digestive disease includes inflammatory bowel disease, gastroesophageal reflux disease, peptic ulcer disease, and irritable bowel syndrome; musculoskeletal disease includes gout, osteoarthritis, and rheumatoid arthritis; and cancer includes lymphohematopoietic cancer, digestive system cancer, endocrine system cancer, reproductive system cancer, respiratory system cancer, and urinary system cancer.

[0059] Furthermore, the preliminary screening of candidate factors related to mild cognitive impairment from several candidate factors in individuals already suffering from mild cognitive impairment includes:

[0060] We used matched nested case-control studies to examine the relationship between candidate factors and mild cognitive impairment.

[0061] Furthermore, the use of matched nested case-control studies to examine the relationship between candidate factors and mild cognitive impairment includes:

[0062] Mild cognitive impairment cases and control samples were all from the same cohort, and the matching date was defined as the date on which the mild cognitive impairment occurred in the current case (indication date).

[0063] The control sample is defined as individuals who did not experience mild cognitive impairment on the indicated date.

[0064] Based on the principles that the difference between control and case ages should be within ±2 years and that they should be of the same sex, greedy matching was used, with a maximum of 10 controls matched for each case. Greedy matching is a commonly used matching method in case-control studies, designed to control confounding variables and improve the validity of research results. Greedy matching selects the most similar control sample to each case stepwise, calculates the distance, and ensures that matched controls are not selected repeatedly, thus matching 10 control samples for each case. The advantages of this method are that it can reduce confounding bias, improve statistical power, and enhance the interpretability of results.

[0065] For nested case-control studies, multivariate conditional logistic regression models are used to estimate the OR value, which represents the association strength between candidate traits and the risk of mild cognitive impairment.

[0066] The expression for the multivariate conditional logistic regression model is as follows:

[0067]

[0068] Y ij Y represents the outcome variable of the j-th observation in the i-th case. ij =1 indicates that the event has occurred (such as a case), Y ij =0 indicates that the event did not occur (as shown in the control). X kij Let represent the k-th independent variable (covariate) in the j-th observation of the i-th case, used to explain the change in the outcome variable. The independent variable can be continuous or categorical. β0 represents the intercept term, indicating the sum of all independent variables X. kij When β is zero, it represents the logarithmic probability of the event occurring. k Let X represent the regression coefficient of the k-th independent variable, and let X represent the independent variable X. kijThe change in the logarithmic probability of the event occurring for each additional unit.

[0069] Specifically, exp(β) k This represents the change in the relative risk of an event occurring when the independent variable increases by one unit.

[0070] The formula for calculating OR is:

[0071] OR = (a / b) / (c / d);

[0072] Where a represents the number of cases in the exposed group, b represents the number of controls in the exposed group, c represents the number of cases in the unexposed group, and d represents the number of controls in the unexposed group.

[0073] If the confidence interval of the OR value does not include 1, it indicates that the candidate factor is associated with mild cognitive impairment; if the confidence interval of the OR includes 1, it indicates that the candidate factor is not associated with mild cognitive impairment.

[0074] Furthermore, the second screening of candidate factors from the initial screening, using the variance inflation factor, includes:

[0075] For all candidate factors initially screened, calculate the variance inflation factor (VIF). If the VIF value of the current candidate factor is greater than the set threshold, the current candidate factor is deleted. If the VIF value of the current candidate factor is less than or equal to the set threshold, the current candidate factor is retained.

[0076] For example, to avoid multicollinearity, the variance inflation factor (VIF) is used to detect multicollinearity, and variables with a VIF value greater than 10 are removed.

[0077] Furthermore, from the candidate factors selected in the second screening, a third screening is performed using the Lasso regression algorithm to obtain several predictive factors, including:

[0078] Sparsity is achieved by adding a penalty term (L1 regularization) to the objective function, which makes the weights of many features in the coefficient vector become 0. By selecting the features corresponding to non-zero coefficients, the features with the greatest predictive power for the target variable can be screened out, thereby simplifying the model, improving the model's generalization ability, and obtaining several predictors.

[0079]

[0080] Where Minβ represents the objective function to be minimized, is a function of the coefficient β, N is the sample size, and p is the number of parameters. This represents L1 regularization (penalty term); under the condition that the sum of the absolute values ​​of the L1 penalty coefficients is less than a constant, some coefficients are forced to 0; y i x is the actual value of the i-th sample. ij It is the j-th feature value of the i-th sample, and β0 is the intercept term; β j λ is the coefficient of feature j, and λ is the regularization parameter that controls the weight of the penalty term.

[0081] Predictor selection flowchart as follows Figure 1 As shown.

[0082] Furthermore, the dataset is composed of individuals with known mild cognitive impairment (MCI). This dataset is divided into training and testing sets according to a predetermined ratio. When predicting MCI risk, 501,172 MCI and non-MCI patients are divided into training, cross-validation, and testing groups. First, a training set (representing 80% of the population) and a testing set (representing 20% ​​of the population) are created. Then, the training set is divided into 10 evenly sized partitions, forming 10 different training and cross-validation sets.

[0083] Furthermore, all the machine learning models mentioned include: Multi-Layer Perceptron (MLP), Categorical Boosting (CatBoost), Random Forest (RF), Light Gradient Boosting Machine (LightGBM), Gradient Boosting Decision Tree (GBDT), Decision Tree, Adaptive Boosting (AdaBoost), Logistic Regression (LG), Linear Support Vector Classification (LinearSVC), and eXtreme Gradient Boosting (XGBoost).

[0084] It should be understood that one deep learning model (MLP) and nine machine learning models (CatBoost, RF, LightGBM, GBDT, Decision Tree, AdaBoost, LG, LinearSVC, and XGBoost) were selected as candidate prediction models. Models requiring hyperparameter tuning were tuned using a random search with cross-validation.

[0085] In imbalanced data, many machine learning methods are "biased" towards the class with the larger proportion. In the prediction of MCI, the ratio of MCI to non-MCI patients is approximately 1:50, so the EasyEnsemble algorithm is used to address the extremely imbalanced dataset.

[0086] Further, the test set is input into the initially trained machine learning model to test the model. The machine learning model whose test index value exceeds a set threshold is used as the final trained machine learning model for mild cognitive impairment. The test index value includes:

[0087] The ROC (Receiver Operating Characteristic) curve includes the Area Under the Curve (AUC), accuracy, precision, recall, and F1 score.

[0088] For example, model evaluation and comparison: AUC is used to compare multiple models. Based on the AUC values ​​obtained from cross-validation, the model with the highest performance is selected for testing on the test set. AUC values ​​are categorized as follows: 0.5-0.7, 0.7-0.8, 0.8-0.9, and >0.9 represent acceptable, average, good, and excellent models, respectively. In addition, accuracy, precision, recall, and F1 score are calculated to evaluate the models.

[0089] Furthermore, the system also includes:

[0090] The model interpretation module is configured to input several predictors of individuals with mild cognitive impairment into the interpretable machine learning model SHAP (SHapley Additive Explanations) to obtain the importance value of each predictor to the model.

[0091] Furthermore, the step of inputting several predictors of individuals with mild cognitive impairment into the interpretable machine learning model SHAP to obtain the importance value of each predictor to the model is achieved by using the SHapley Additiveexplanation (SHAP) method to quantify the importance of features in the optimal model and visualize the results.

[0092] To enhance the interpretability of the model, this invention uses the SHAP method to quantify the importance of features, calculating the contribution of each input feature to each participant's MCI and MCI progression to dementia outcome, and assigning an importance score (i.e., SHAP value) to each feature. To visualize the results, honeycomb plots and bar charts of the predictors are drawn, and the features are ranked according to the absolute value of the SHAP values.

[0093] The model calibration performance was evaluated by comparing the true probabilities with the predicted probabilities, and a calibration curve was plotted; the closer the curve is to the diagonal, the better the model performance. Random forest was used to impute variables with a missing percentage of less than 10%.

[0094] like Figures 5(a) to 5(c) As shown, embodiments of this application also develop an online platform so that medical institutions and researchers can easily use these models for risk prediction of MCI.

[0095] By solving the above technical problems, this invention ultimately determined that CatBoost is the best model for predicting MCI, with an AUC of 0.839 and an accuracy of 77.7%.

[0096] This study included individuals with mild cognitive impairment (MCI) and those whose MCI progressed to dementia. Data source: This invention used data from a large community-based cohort study (UKB data), which recruited over 500,000 participants aged 40-70 years between 2006 and 2010. Written informed consent was obtained from participants when collecting questionnaires and biometric data. All participant data were linked to hospital data and national death registry records from England, Scotland, and Wales to determine the date of initial diagnosis of MCI and dementia following baseline assessment. The UKB has obtained approval from and supervised this study through the Northwest Multicenter Research Ethics Committee (https: / / www.ukbiobank.ac.uk / learn-more-about-uk-biobank / about-us / ethics). Initially, dementia patients at baseline were excluded. Further exclusion of baseline MCI patients resulted in 501,172 individuals when constructing the predictive model for MCI. 12,031 MCI patients at baseline and during follow-up were included in the construction of the predictive model for MCI progression to dementia.

[0097] Outcome definition: MCI patients were identified using International Classification of Diseases-10th (ICD-10) codes F067, G318, and R41. Dementia patients within the mild cognitive impairment (MCI) population were identified using algorithms listed in UKB Class 47. Outcomes for MCI and dementia were finalized by the UK Biobank Disease Outcomes Decision Panel.

[0098] Existing MCI prediction models are typically based on small samples of tens to hundreds of individuals. These predictors usually utilize routinely collected electronic health records, clinical imaging examinations, or community-obtained survey data, with AUC values ​​mostly between 0.62 and 0.75. While models using features from brain imaging tests (such as positron emission tomography) to predict MCI generally perform better, acquiring this data is both expensive and impractical in community settings. Furthermore, these models lack a ranking of feature importance, and their interpretability needs improvement.

[0099] This invention, based on a large sample of 501,172 individuals, constructed 10 models to predict the occurrence of MCI using 16 predictive factors from non-brain imaging data (APOE e4 allele carrier count, age, sleep duration, frailty, heart failure, stroke, hypertension, diabetes, dyslipidemia, influenza and pneumonia, chronic obstructive pulmonary disease, gastroesophageal reflux disease, osteoarthritis, lymphohematopoietic system cancers, digestive system cancers, and respiratory system cancers). The AUC value on the training set was 0.744-0.859, and the AUC value on the cross-validation set was 0.739-0.835. CatBoost was ultimately determined as the optimal model, with a test set AUC of 0.839, higher than the AUC metric of previous studies using non-image test data. The accuracy, precision, recall, and F1 score were 0.780, 0.972, 0.780, and 0.857, respectively. The calibration curve is close to the diagonal, indicating that the predicted probability is close to the true probability. Furthermore, the SHAP plot shows that age, hypertension, APOE e4 allele carrier count, diabetes, influenza and pneumonia, osteoarthritis, stroke, dyslipidemia, heart failure, and chronic obstructive pulmonary disease are the top 10 predictors of MCI. Specific model results are as follows... Figures 3(a) to 3(f) As shown.

[0100] Example 2

[0101] This embodiment provides a machine learning-based predictive system for the progression of MCI to dementia, including:

[0102] The predictor screening module is configured to: obtain predictor factors by screening candidate factors in a population known to have progressed from mild cognitive impairment to dementia. The screening process includes: initially screening candidate factors related to the progression from mild cognitive impairment to dementia from several candidate factors known to have progressed from mild cognitive impairment to dementia; then, using the variance inflation factor to perform a second screening from the initially screened candidate factors; and finally, using the Boruta algorithm to perform a third screening from the factors selected in the second screening to obtain the predictor factors.

[0103] The dataset construction module is configured to: take people with known mild cognitive impairment who have progressed to dementia as the dataset, and divide the dataset into training set and test set according to a set ratio;

[0104] The model training and testing module is configured to: input the training set into all machine learning models to train the models and obtain the pre-trained machine learning models; input the test set into the pre-trained machine learning models to test the models; and take the machine learning models whose test index values ​​exceed the set threshold as the final trained machine learning models for mild cognitive impairment progressing to dementia.

[0105] The prediction module is configured to input several predictive factors of the test subject who already has mild cognitive impairment into the finally trained machine learning model for progression from mild cognitive impairment to dementia, and obtain the prediction result of the test subject's progression from mild cognitive impairment to dementia.

[0106] Furthermore, the candidate factors for the progression of known mild cognitive impairment to dementia are consistent with the candidate factors for the mild cognitive impairment population described in Example 1.

[0107] Furthermore, from several candidate factors known to progress from mild cognitive impairment to dementia, factors related to the progression from mild cognitive impairment to dementia are initially screened out. The initial screening process is consistent with the screening process in Example 1.

[0108] Furthermore, the process of using the variance inflation factor to perform a second screening from the initially screened factors is consistent with the screening process in Example 1, which uses the variance inflation factor to perform a second screening from the initially screened candidate factors.

[0109] Furthermore, all the machine learning models described in Example 2 are consistent with all the machine learning models described in Example 1.

[0110] Furthermore, the system also includes a model interpretation module, which is configured to input several predictors of the progression of mild cognitive impairment to dementia into the interpretable machine learning model SHAP package to obtain the importance value of each predictor to the model.

[0111] Furthermore, from the factors selected in the second round of screening, a third screening using the Boruta algorithm is performed to obtain the predicted factors. The Boruta algorithm generates a random forest through multiple bootstrap resampling operations, calculates feature importance, and introduces shadow features (generated by shuffling the original features) for comparison. If the importance of the original feature is significantly higher than that of the shadow feature, it is considered important; otherwise, it is marked as irrelevant. This process iterates through multiple rounds to ultimately select the truly important features.

[0112] Furthermore, the dataset is divided into a training set and a test set according to a set ratio; when predicting the risk of MCI progressing to dementia, the dataset of 12,031 MCI patients is divided in the same way as in Example 1. The dataset division is as follows: Figure 2 As shown, in the prediction of MCI progression to dementia, the ratio of MCI patients who progress to dementia to those who do not is approximately 1:4. Random oversampling was used to address the data imbalance problem.

[0113] To enhance the interpretability of the model, this invention uses the SHAP method to quantify the importance of features, calculating the contribution of each input feature to the MCI progression to dementia outcome for each participant, and assigning an importance score (i.e., SHAP value) to each feature. To visualize the results, honeycomb plots and bar charts of the predictors are drawn, and the features are ranked according to the absolute value of the SHAP values.

[0114] In imbalanced data, many machine learning methods are "biased" towards the class with the larger proportion. In the prediction of MCI, the ratio of MCI patients to non-MCI patients is approximately 1:50, so the EasyEnsemble algorithm is used to address the extremely imbalanced dataset. In the prediction of MCI progressing to dementia, the ratio of MCI patients progressing to dementia to those not progressing to dementia is approximately 1:4, and random oversampling is used to address the data imbalance problem.

[0115] like Figures 6(a) to 6(c) As shown, embodiments of this application also develop an online platform so that medical institutions and researchers can easily use these models to predict the risk of MCI progressing to dementia.

[0116] By solving the above technical problems, this invention ultimately determined that Random Forest is the best model for predicting the progression of dementia in MCI, with an AUC of 0.887 and an accuracy of 80.1%.

[0117] Furthermore, an online platform was developed as a tool to facilitate the application of these models, providing important references for early intervention and treatment of MCI progression to dementia. Existing predictive models for MCI progression to dementia are mainly developed based on MRI brain imaging data in small sample populations of tens to thousands, making them unsuitable for widespread application in community populations. Predictive models based on non-brain imaging data from the community environment are fewer, and their AUC values ​​range from 0.36 to 0.72, indicating that model performance needs further improvement.

[0118] This invention, based on a large sample of 12,031 individuals, uses 17 predictive factors from non-brain imaging data (APOE e4 allele carrier count, glycated hemoglobin level, insulin use, social isolation, income level, smoking status, BMI, grip strength, coronary heart disease, influenza and pneumonia, osteoarthritis, education level, sedentary time, chronic obstructive pulmonary disease, heart failure, hypertension, and diabetes status) to construct 10 models to predict the progression of MCI to dementia. The AUC values ​​on the training set ranged from 0.636 to 0.980, and the AUC values ​​on the cross-validation set ranged from 0.627 to 0.876. The final RF model with good performance was determined, achieving a test set AUC of 0.875, higher than the AUC metrics of previous studies using non-imaging data. The accuracy, precision, recall, and F1 score were 0.751, 0.797, 0.793, and 0.792, respectively. The calibration curve is close to the diagonal, indicating that the predicted probability is close to the true probability. Furthermore, the SHAP plot shows that the number of APOE e4 allele carriers, income level, smoking status, BMI, coronary heart disease, education level, sedentary time, chronic obstructive pulmonary disease, heart failure, and diabetes status are the top 10 predictors of MCI progression to dementia. Specific model results are as follows... Figures 4(a) to 4(f) As shown.

[0119] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A machine learning-based system for predicting mild cognitive impairment, characterized by including: The predictor screening module is configured to: obtain predictors by screening candidate factors in a population with known mild cognitive impairment. The screening process includes: initially screening candidate factors related to mild cognitive impairment from several candidate factors in the population with mild cognitive impairment. Then, from the initial candidate factors, a second screening is performed using the variance inflation factor; finally, from the second screening candidate factors, a third screening is performed using the Lasso regression algorithm to obtain several predictive factors. The dataset construction module is configured to: use people with known mild cognitive impairment as the dataset, and divide the dataset into training and testing sets according to a set ratio; The model training and testing module is configured to: input the training set into all machine learning models, train the models, and obtain the pre-trained machine learning models; input the test set into the pre-trained machine learning models, test the models, and use the machine learning models whose test index values ​​exceed a set threshold as the final trained mild cognitive impairment machine learning models; the final trained mild cognitive impairment machine learning models are CatBoost models. The prediction module is configured to input several predictive factors of the subject to be tested, including: APOE e4 allele carrier count, age, sleep duration, frailty, heart failure, stroke, hypertension, diabetes, dyslipidemia, influenza and pneumonia, chronic obstructive pulmonary disease, gastroesophageal reflux disease, osteoarthritis, lymphohematopoietic system cancer, digestive system cancer, and respiratory system cancer; into the finally trained mild cognitive impairment machine learning model to obtain a prediction result of whether the subject to be tested has mild cognitive impairment.

2. The machine learning-based mild cognitive impairment prediction system as described in claim 1, characterized in that, The candidate factors for the population with mild cognitive impairment include: Category 1: genetic factors; Category 2: early life factors; Category 3: demographic factors; Category 4: mid-life environmental factors; Category 5: mid-life lifestyle factors; Category 6: mid-life biological factors; Category 7: mid-life psychosocial factors; and Category 8: medical history.

3. The machine learning-based mild cognitive impairment prediction system as described in claim 1, characterized in that, The preliminary screening of candidate factors related to mild cognitive impairment from several candidate factors in the population includes: using matched nested case-control studies to examine the relationship between candidate factors and mild cognitive impairment.

4. The machine learning-based mild cognitive impairment prediction system as described in claim 3, characterized in that, The use of matched nested case-control studies to examine the relationship between candidate factors and mild cognitive impairment includes: mild cognitive impairment cases and control samples are all from the same cohort, the matching date is defined as the date on which the current case developed mild cognitive impairment; the control sample is defined as individuals who did not experience mild cognitive impairment on the indicated date; for nested case-control studies, a multivariate conditional logistic regression model is used to estimate the association strength (OR) between candidate features and the risk of mild cognitive impairment.

5. The machine learning-based mild cognitive impairment prediction system as described in claim 1, characterized in that, The second screening, which involves using the variance inflation factor from the initially selected candidate factors, includes: For all candidate factors initially screened, calculate the variance inflation factor (VIF). If the VIF value of the current candidate factor is greater than the set threshold, the current candidate factor is deleted. If the VIF value of the current candidate factor is less than or equal to the set threshold, the current candidate factor is retained.

6. The machine learning-based mild cognitive impairment prediction system as described in claim 1, characterized in that, From the candidate factors selected in the second round, a third round of selection is performed using the Lasso regression algorithm to obtain several predictive factors, including: ; in, The objective to be minimized is about the coefficients. The function, where N is the sample size and p is the number of parameters. This represents L1 regularization; under the condition that the sum of the absolute values ​​of the L1 penalty coefficients is less than a constant, some coefficients are forced to be 0. It is the first The actual value of each sample It is the first The first sample 1 eigenvalue, It is a feature coefficient, It is a regularization parameter that controls the weight of the penalty term.

7. The machine learning-based mild cognitive impairment prediction system as described in claim 1, characterized in that, All machine learning models mentioned include: multilayer perceptron, class boosting model, random forest model, lightweight gradient boosting machine, gradient boosting decision tree model, decision tree model, adaptive boosting model, logistic regression model, linear support vector classification model, or extreme gradient boosting model.

8. The machine learning-based mild cognitive impairment prediction system as described in claim 1, characterized in that, The test set is input into the initially trained machine learning model to test the model. The machine learning model whose test index value exceeds the set threshold is used as the final trained machine learning model for mild cognitive impairment. The test index values ​​include: area under the ROC curve, accuracy, precision, recall, and F1 score.

9. The machine learning-based mild cognitive impairment prediction system as described in claim 1, characterized in that, The system also includes a model interpretation module, which is configured to input several predictors of people with mild cognitive impairment into an interpretable machine learning model SHAP to obtain the importance value of each predictor to the model.

10. Machine learning-based MCI advances into a dementia prediction system, characterized by including: The predictor screening module is configured to: obtain predictor factors by screening candidate factors in a population known to have progressed from mild cognitive impairment to dementia. The screening process includes: initially screening candidate factors related to the progression from mild cognitive impairment to dementia from a number of candidate factors known to have progressed from mild cognitive impairment to dementia; then, using a variance inflation factor to perform a second screening from the initially screened factors; and finally, using the Boruta algorithm to perform a third screening from the second-screened factors to obtain the predictor factors. The dataset construction module is configured to: take people with known mild cognitive impairment who have progressed to dementia as the dataset, and divide the dataset into training set and test set according to a set ratio; The model training and testing module is configured to: input the training set into all machine learning models to train them and obtain the pre-trained machine learning models; input the test set into the pre-trained machine learning models to test them; and use the machine learning models whose test index values ​​exceed a set threshold as the final trained machine learning models for mild cognitive impairment progressing to dementia. The final trained machine learning models for mild cognitive impairment progressing to dementia are called RF models. The prediction module is configured to input several predictive factors of a subject already suffering from mild cognitive impairment, including the number of APOE e4 allele carriers, glycated hemoglobin levels, insulin use, social isolation, income level, smoking status, BMI, grip strength, coronary heart disease, influenza and pneumonia, osteoarthritis, education level, sedentary time, chronic obstructive pulmonary disease, heart failure, hypertension, and diabetes status, into a machine learning model for the progression of mild cognitive impairment to dementia after final training, and obtain a prediction result of the subject's progression from mild cognitive impairment to dementia.

Citation Information

Patent Citations

  • System for predicting dementia or mild cognitive impairment

    CN114822850A

  • Visuospatial disorders detection in dementia using a computer-generated environment based on voting approach of machine learning algorithms

    US11250723B1