Methods for predicting disease onset risk and systems for same

AU2025224449A1Pending Publication Date: 2026-08-27RGT UNIV OF CALIFORNIA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
AU2025224449
Authority / Receiving Office
AU · AU
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-20
Filing Date
2025-02-19
Publication Date
2026-08-27

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Provided are methods of training a model to predict a disease onset risk. Aspects of the methods include obtaining a plurality of electronic health records, dividing the subjects into a first set for machine learning model training and a second set for machine learning evaluation, training a machine learning model using the electronic health records of the first set of subjects, evaluating the machine learning model using the electronic health records of the second set of subjects and identifying from the trained and evaluated machine learning model one or more features of the subjects that predict the disease onset risk. In some embodiments, the one or more features are one or more sex-specific features. Also provided are methods of predicting a disease onset risk in a subject, methods of clinically evaluating the presence of a disease in a subject and methods of treating a subject predicted to be at risk of developing a disease. Embodiments of the invention are computer-implemented methods. Systems and non-transitory computer-readable mediums for performing the subject methods are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

Statement Regarding Federally Sponsored Research This invention was made with government support under R01 AG060393, and P30 AG062422 awarded by the National Institutes of Health. The Government has certain rights in the invention. Cross-Reference Pursuant to 35 U.S.C. § 119 (e), this application claims priority to the filing date of U.S. Provisional Patent Application No. 63 / 555,699, filed February 20, 2024, which application is incorporated herein by reference in its entirety. Introduction Neurodegenerative disorders are devastating, heterogeneous and challenging to diagnose, and their burden on an aging population is expected to continue to grow (2022 Alzheimer’s disease facts and figures. Alzheimer’s Dement. 18, 700-789 (2022)). Among these, Alzheimer’s disease (AD) is the most common form of dementia after age 65, and its hallmark memory loss and other cognitive symptoms are costly and onerous to both patients and caregivers. Approaches to curb this impact are moving increasingly to targeting interventions in at-risk individuals prior to the onset of irreversible decline (Rasmussen, J. & Langerman, H. Alzheimer’s Disease - Why We Need Early Diagnosis. Degener. Neurol. Neuromuscul. Dis. Volume 9, 123-130 (2019); Kivipelto, M. Midlife vascular risk factors and Alzheimer’s disease in later life: longitudinal, population based study. BMJ322, 1447-1451 (2001); Niculescu, A. B. et al. Blood biomarkers for memory: toward early detection of risk for Alzheimer disease, pharmacogenomics, and repurposed drugs. Mol. Psychiatry 25, 1651-1672 (2020)). To this end, advancements in AD biomarkers, diagnostic tests, and neuroimaging have improved the detection and classification of AD, and disease-modifying treatments have been approved, but there is still no cure and much remains unknown about its pathogenesis (Alena V. Savonenko, Philip C. Wong, & Tong Li. Alzheimer diseases. (2023) doi:10.1016 / b978-0-323-85654-6.00022-8; Neugroschl, J. & Wang, S. Alzheimer’s Disease: Diagnosis and Treatment Across the Spectrum of Disease Severity. Mt. Sinai J. Med. N. Y. 78, 596-612 (2011)). Summary The inventors have realized that there is, therefore, a need for improved predictive models for disease onset (e.g., Alzheimer’s disease (AD) onset), and need for generating hypotheses of biological relationships between top predictors and disease (e.g., Alzheimer’s disease). The inventors of the present application have further realized that a promising approach of identifying disease onset risk involves utilizing electronic health record data as a source of rich longitudinal data that can be leveraged to understand and predict complex diseases, such as Alzheimer’s disease. Embodiments of the invention leverage this approach to solve this problem using machine learning techniques to predict features of individual health data, such as features of electronic health records, from training machine learning models based at least in part on electronic health record data. As described, early identification of Alzheimer’s disease (AD) risk can aid in interventions before disease progression. Embodiments of the present invention combine electronic health records (EHRs) with heterogeneous knowledge networks (e.g., SPOKE) to allow for (1) prediction of AD onset and (2) generation of biological hypotheses linking phenotypes with AD. Provided are methods of training a model to predict a disease onset risk. Aspects of the methods include obtaining a plurality of electronic health records, dividing the subjects into a first set for machine learning model training and a second set for machine learning evaluation, training a machine learning model using the electronic health records of the first set of subjects, evaluating the machine learning model using the electronic health records of the second set of subjects and identifying from the trained and evaluated machine learning model one or more features of the subjects that predict the disease onset risk. In some embodiments, the one or more features are one or more sex-specific features. Also provided are methods of predicting a disease onset risk in a subject, methods of clinically evaluating the presence of a disease in a subject and methods of treating a subject predicted to be at risk of developing a disease. Embodiments of the invention are computer-implemented methods. Systems and non-transitory computer-readable mediums for performing the subject methods are also provided. Brief Description of the Figures The invention may be best understood from the following detailed description when read in conjunction with the accompanying drawings. Included in the drawings are the following figures: FIG. 1 depicts a flow diagram for predicting a disease onset risk according to embodiments of the present invention. FIG. 2 depicts a functional block diagram for a computer system according to embodiments of the present invention. FIG. 3 depicts a general architecture of an example computing device according to embodiments of the present invention. FIGS. 4A-4C: Overview of participant selection and RF model performance. (FIG. 4A) From the UCSF EHRs and the UCSF Memory and Aging Center (MAC) database, participant and clinical information was extracted, filtered and prepared for time points before the index time. All clinical features extracted were one-hot encoded and trained on random forest (RF) models to predict future risk of AD diagnosis. Models were evaluated on a 30% held-out evaluation set to compute AUROC / AUPRC and interpreted based on feature importances and using a heterogeneous knowledge network (SPOKE). Top features were then further validated in external databases. (FIG. 4B) depicts a functional block diagram for a computer system according to certain embodiments. Filtering a consistent set of individuals with AD and controls from the UCSF EHR for model training and testing. Filtered participant cohorts are shown in Tables 1A and 1B and split with 30% held-out set for testing. (FIG. 4C) Bootstrapped performance of RF models on the held-out evaluation set (n = 300 bootstrapped iterations of 1,000 participants, prevalence of AD on held-out set = 0.003). Bootstrapped AUROC performance for models trained and tested on female strata and male strata are also shown. The box shows quartiles (25th, 50th and 75th percentiles), whiskers extend to 1.5 times the interquartile range, and the remaining points are outliers. FIGS. 5A-5E: Models trained on matched cohorts allow for identification of hypotheses for AD predictors. (FIG. 5A) Bootstrapped performance of models trained on cohorts matched by demographics and visit-related factors on the full held-out evaluation set (n = 300 bootstrapped iterations of 1,000 individuals, prevalence of AD on held-out set = 0.003). The box plot shows quartiles (25th, 50th and 75th percentiles), whiskers extend to 1.5 times the interquartile range, and the remaining points are outliers. (FIG. 5B) Top clinical phecode categories for matched models ranked by the average of the top five importance values for each phecode category. Sorting is based on this average across time models. (FIG. 5C) Top 50 phecodes (detailed features) across time models, with features clustered based on ward distance of rankings. (FIG. 5D) Bootstrapped performances of sex-stratified matched models on the held-out evaluation set (n = 300 bootstrapped iterations of 1,000 individuals for each sex; reference AUPRC = 0.0036 female, 0.0022 male). Each box shows quartiles (25th, 50th and 75th percentiles), and whiskers extend to 1.5 times the interquartile range, with remaining points as outliers. (FIG. 5E) Overlap of top matched model features for models trained on all individuals, female stratified individuals, and male stratified individuals, with model cutoff importance (RF average impurity decrease) greater than 1 x 10-6. Specific features are listed, with bold features indicating top features across all five time models and non-bolded features indicating top features across four time models. FIG. 6: SPOKE provides biological prioritization of hypotheses associated with shared clinical phenotypes. Combined SPOKE network of all shortest paths to AD node (Disease Ontology ID: 10652) for the top 25 input features (bolded) from matched AD model at every time point. Network is organized based on the number of time point model occurrences (y axis) and eccentricity of a node in the subnetwork (x axis). Specific time point model occurrences are colored by the pie chart within each node. FIGS. 7A-7D: The HLD and AD association is validated externally with APOE as a shared causal genetic link. (FIG. 7A) Kaplan-Meier curve on UC-wide EHR for HLD as the exposure (error bands show 95% Cl). Two-sided log-rank test is significant for all HLD versus controls (P = 2.4 x 10-85), female HLD versus female controls (P = 3.6 x 10-69), and male HLD versus male controls (P = 8.4 x 10-22). *P < 0.005. (FIG. 7B) First-degree and second-degree neighbors of HLD on the full network representing all shortest paths from the top 25 features per time model. (FIG. 7C) PheWAS for variant rs2075650 (ch19:44892362(hg38):A > G) on a shared locus associated with both HLD and AD, plotted based on multiple prior studies with variant phenotype associations with P value < 0.05 from the UK Biobank. The horizontal line indicates a Bonferroni-corrected significance level of 0.05 (191 phenotypes, Bonferroni P value = 0.00026), and the arrow direction represents the beta direction of effect of the alternative allele. (FIG. 7D) Plot of APOE protein expression colocalization with H4 (probability two associated traits share a causal variant) from Open Targets Genetics. Each dot represents a specific phenotype categorized based on trait (x axis). Each color represents an APOE molecular trait measured from blood plasma from refs. FIGS. 8A-8E: The association between osteoporosis and AD is validated externally with MS4A6A as a potential female-specific shared genetic link. (FIG. 8A) Kaplan-Meier curve on UC-wide EHR for osteoporosis as the exposure (error bands show 95% Cl). Two-sided log-rank test is significant for all osteoporosis-exposed individuals versus controls (P = 1.4 x 10-64) and osteoporosis-exposed female individuals versus controls (P = 7.2 x 10-72), but not male osteoporosis-exposed individuals versus controls (P = 0.46). *P < 0.005. (FIG. 8B) First-degree and second degree neighbors of osteoporosis node on the network representing all shortest paths from the top 25 features per time model. (FIG. 8C) P-P plots between summary statistics of AD GWAS (P value computed as described in ref. 35, n = 455,258) and sex-stratified HBMD GWAS (female n = 111,152, male HBMD n = 166,988, P value computed as described in Neale’s Lab GWAS version 3) of variants around the MS4A locus (left and middle plots) at region 60050000-60200000 of chr11 (locus plot on right). (FIG. 8D) MS4A6A gene expression (cis-eQTL, P values computed as described in ref. 104) association with AD GWAS (P value computed as described in ref. 35) and association with sex-stratified low HBMD (P value computed as described in Neale’s Lab GWAS version 3). (FIG. 8E) Open Targets Genetics associated phenotype graph for MS4A6A with association score computed based on a weighted harmonic sum across evidence. Circles are phenotypes colored by the association score, and boxes represent the most general categories. NS, not significant. FIG. 9 depicts a cross-validation approach. The full dataset was split into 70% for training and choosing the best model, and 30% was set aside as the held-out evaluation set. Model selection and optimization was performed with cross-validation on the 70% training set. All final models are then evaluated on the 30% held-out evaluation set. FIGS. 10A-10C: Top detailed features and phecodes from the random forest model. (FIG. 10A) Top detailed OMOP clinical features utilized in models for clinical feature only models (top), or clinical features + demographic + visit information models (bottom). Features within the drug / measurement categories are marked with a triangle, while demographic / visit features are marked with a circle. (FIG. 10B)Top phecode categories utilized in models, where importance is determined by the top 5 detailed features within each phecode mapping. The vertical order is based upon the average importance across time models. (FIG. 10C) Top 50 phecodes utilized in time models, clustered based on relative importance across time models. FIG. 11 depicts a comparison of age and visit-related factors between AD, controls, and matched controls. The plots demonstrate the distribution of continuous variables utilized in matching with error bands representing standard deviation. Orange represents AD patients at each time point. Dark blue represents all controls, while light blue represents controls that have been matched at each time point. FIGS. 12A-12C: Sex stratified model performance and top features. (FIG. 12A)The full performance of sex-stratified models is shown. The bootstrapped AUROC / AUPRC is determined by the male or female strata of the initial 30% heldout evaluation set (n = 300 bootstrapped iterations of 1000 patients for each sex, reference ALIPRC = 0.0036 female, 0.0022 male). The box shows quartiles (25%, 50%, 75%ile), and whiskers extend to 1.5*interquartile range, with remaining points as outliers. (FIG. 12B)Top phecode categories are listed by importance for all models, with inclusion of comparison with the general non-stratified model. Vertical ordering is determined by the average importance across time models. (FIG. 12C)Top 50 important phecodes clustered by relative importance across time models and across strata. FIGS. 13A-13E: Logistic regression models and top coefficients. (FIG. 13A) The full performance of logistic regression models. The bootstrapped AUROC / AUPRC is determined the 30% held-out evaluation set (n = 300 bootstrapped iterations of 1000 patients). The box shows quartiles (25%, 50%, 75%ile), and whiskers extend to 1.5*interquartile range, with remaining points as outliers. (FIG. 13B) Top detailed OMOP feature logistic regression coefficients are listed by importance for all model formulations. Top row shows coefficients from the model trained on all patients, while the bottom row shows coefficients from the model trained on matched cohorts. (FIG. 13C) The full performance of sex-stratified logistic regression models is shown. The bootstrapped AUROC / AUPRC is determined by the male or female strata of the initial 30% held-out evaluation set (n = 300 bootstrapped iterations of 1000 patients for each sex). The box shows quartiles (25%, 50%, 75%ile), and whiskers extend to 1.5*interquartile range, with remaining points as outliers. (FIG. 13D) Top phecode categories across time models and across strata, determined by the top 10 logistic regression coefficient magnitudes within each category. (FIG. 13E) Top 50 important phecodes clustered by average logistic regression coefficient across time models and across strata, where the average logistic regression coefficient is determined by the top 10 logistic regression coefficient magnitudes within each category. FIG. 14 depicts a comparison of the random forest model feature importance between the model trained on all patients (y-axis) and the model trained on demographics / care utilization matched cohorts (x-axis). The line represents no change in feature importance. Above the line represents a decrease in feature importance in the model trained on the full cohort compared to matched cohorts, and below the line represents features with increased importance for the model trained on matched cohorts. FIGS. 15A, 15B: Balanced accuracy and example permutation test. (FIG. 15A) Balanced accuracy on the 30% held-out evaluation set was computed for all random forest models. (FIG. 15B) A null distribution for AUROC (score) was computed based on retrained random forest models with permutations on the ground truth label (40 permutations). P-value is calculated by (C + 1) / (n_permutations + 1), where C represents the number of permutations that scored better than the non-permuted dataset. FIGS. 16A, 16B: External EHR validation support increased AD diagnosis with hyperlipidemia and osteoporosis exposure. (FIG. 16A) Sex-stratified combined Kaplan-Meier survival curves with hyperlipidemia (HLD) as the exposure (curve shows survival fraction, error bands show 95% confidence interval). Patient attrition is shown in the middle for each subgroup. Below, two-sided log rank test comparison results are shown. F = female, M = male. (FIG. 16B) Sex-stratified combined Kaplan-Meier survival curves with osteoporosis as the exposure (curve shows survival fraction, error bands show 95% confidence interval). Patient attrition is shown in the middle for each subgroup. Two-sided log rank test comparison results are shown. Detailed Description As reviewed above, provided are methods of training a model to predict a disease onset risk. Aspects of the methods include obtaining a plurality of electronic health records, dividing the subjects into a first set for machine learning model training and a second set for machine learning evaluation, training a machine learning model using the electronic health records of the first set of subjects, evaluating the machine learning model using the electronic health records of the second set of subjects and identifying from the trained and evaluated machine learning model one or more features of the subjects that predict the disease onset risk. In some embodiments, the one or more features are one or more sex-specific features. Also provided are methods of predicting a disease onset risk in a subject, methods of clinically evaluating the presence of a disease in a subject and methods of treating a subject predicted to be at risk of developing a disease. Embodiments of the invention are computer-implemented methods. Systems and non-transitory computer-readable mediums for performing the subject methods are also provided. Before the present invention is described in greater detail, it is to be understood that this invention is not limited to particular embodiments described, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present invention will be limited only by the appended claims. Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limit of that range and any other stated or intervening value in that stated range, is encompassed within the invention. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges and are also encompassed within the invention, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention. Certain ranges are presented herein with numerical values being preceded by the term "about." The term "about" is used herein to provide literal support for the exact number that it precedes, as well as a number that is near to or approximately the number that the term precedes. In determining whether a number is near to or approximately a specifically recited number, the near or approximating unrecited number may be a number which, in the context in which it is presented, provides the substantial equivalent of the specifically recited number. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present invention, representative illustrative methods and materials are now described. All publications and patents cited in this specification are herein incorporated by reference as if each individual publication or patent were specifically and individually indicated to be incorporated by reference and are incorporated herein by reference to disclose and describe the methods and / or materials in connection with which the publications are cited. The citation of any publication is for its disclosure prior to the filing date and should not be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. Further, the dates of publication provided may be different from the actual publication dates which may need to be independently confirmed. It is noted that, as used herein and in the appended claims, the singular forms “a”, “an”, and “the” include plural referents unless the context clearly dictates otherwise. It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as “solely,” “only” and the like in connection with the recitation of claim elements, or use of a “negative” limitation. As will be apparent to those of skill in the art upon reading this disclosure, each of the individual embodiments described and illustrated herein has discrete components and features which may be readily separated from or combined with the features of any of the other several embodiments without departing from the scope or spirit of the present invention. Any recited method can be carried out in the order of events recited or in any other order which is logically possible. While the apparatus and method has or will be described for the sake of grammatical fluidity with functional explanations, it is to be expressly understood that the claims, unless expressly formulated under 35 U.S.C. §112, are not to be construed as necessarily limited in any way by the construction of "means" or "steps" limitations, but are to be accorded the full scope of the meaning and equivalents of the definition provided by the claims under the judicial doctrine of equivalents, and in the case where the claims are expressly formulated under 35 U.S.C. §112 are to be accorded full statutory equivalents under 35 U.S.C. §112. Methods of Training a Model As reviewed above, aspects of the present disclosure include methods of training a model to predict a disease onset risk. Methods of training a model to predict a disease onset risk include obtaining a plurality of electronic health records comprising health information about a plurality of subjects, wherein the plurality of subjects comprise subjects diagnosed with the disease and subjects not diagnosed with the disease, dividing the subjects into a first set for machine learning model training and a second set for machine learning evaluation, training a machine learning model using the electronic health records of the first set of subjects, evaluating the machine learning model using the electronic health records of the second set of subjects and identifying from the trained and evaluated machine learning model one or more features of the subjects that predict the disease onset risk. Also provided are methods of training a model to predict a disease onset risk, where the method includes obtaining a plurality of electronic health records comprising health information about a plurality of subjects, wherein the plurality of subjects comprise subjects diagnosed with the disease and subjects not diagnosed with the disease, and wherein the electronic health records include sexes of the subjects, dividing the subjects into a first set for machine learning model training and a second set for machine learning evaluation, training a machine learning model using the electronic health records of the first set of subjects, evaluating the machine learning model using the electronic health records of the second set of subjects and identifying from the trained and evaluated machine learning model one or more sex-specific features of the subjects that predict the disease onset risk. FIG. 1 illustrates a flow diagram 100 for training a model to predict a disease onset risk. By “disease onset risk” it is meant the likelihood that a subject experiences an onset of a disease. In other words, predicting a disease onset risk may include predicting whether or not a subject is at risk for developing a disease. In some cases, predicting a disease onset risk includes identifying one or more features of an electronic health record of a subject that is indicative of an elevated risk of development of the disease by the subject. For example, predicting disease onset risk may include identifying features (e.g., sex-specific features) of a subject’s electronic health record that indicates the subject has an increased likelihood of developing the disease in the future, i.e., prior to such subject receiving a diagnosis for the disease. Models may be trained to predict a disease onset risk for any disease of interest. Diseases include, but are not limited to, neurological diseases, cardiovascular diseases, cancers, hereditary diseases and metabolic diseases. In some embodiments, the disease is a neurological disease (e.g., a neurodegenerative disease, a neuromuscular disease, a brain disease, a spine disease or a disease of the peripheral nerves). In some embodiments, the disease is a neurodegenerative disease. Neurodegenerative diseases are diseases characterized by a progressive loss of neurons. Neurodegenerative diseases include, but are not limited to, Alzheimer’s disease, Parkinson’s disease, amyotrophic lateral sclerosis, Huntington’s disease, multiple sclerosis and motor neuron disease. In some embodiments, the disease is Alzheimer’s disease. In FIG. 1, flow diagram 100 starts at step 101. At step 101, a plurality of electronic health records comprising health information about a plurality of subjects is obtained. In embodiments, the plurality of subjects include subjects diagnosed with the disease and subjects not diagnosed with the disease. In some embodiments, the electronic health records include sexes (i.e., gender) of the subject. Electronic health records may be obtained for any number of different subjects or individuals. In some instances, the subjects or individuals are “mammals” or "mammalian," where these terms are used broadly to describe organisms which are within the class mammalia, including the orders carnivore (e.g., dogs and cats), rodentia (e.g., mice, guinea pigs, and rats), and primates (e.g., humans, chimpanzees, and monkeys). In some instances, the subjects are humans, including humans of different ethnicities, ages, genders or other physiological or demographic characteristics. Electronic health records (EHRs) include health information about a plurality of subjects. Electronic health records of interest may be generated by a subject’s healthcare team, such as doctors, including general care practitioners or specialists or other medical professionals that encounter the subject. Electronic health records of interest may be obtained from medical centers, such as major medical centers including, for example, IIOSF or UCLA medical centers. Electronic health records of interest may be combined records from more than one institution or medical care provider. Electronic health records of interest may include clinical notes. Further details regarding health records of interest are provided in: Rudrapatna VA, Butte AJ. Opportunities and challenges in using real-world data for health care. J Clin Invest. 2020;130(2):565-574. doi:10.1172 / JCI129197; Pathak J, Kho AN, Denny JC. Electronic health records-driven phenotyping: challenges, recent advances, and perspectives. Journal of the American Medical Informatics Association. 2013;20(e2):e206-e211. doi:10.1136 / amiajnl-2013-002428; Hripcsak G, Albers DJ. Next-generation phenotyping of electronic health records. Journal of the American Medical Informatics Association. 2013;20(1 ):117-121. doi:10.1136 / amiajnl-2012-001145; Rajkomar A, Oren E, Chen K, et al. Scalable and accurate deep learning with electronic health records, npj Digital Med. 2018;1 (1 ):1-10. doi:10.1038 / s41746-018-0029-1, the disclosures of each of which are incorporated herein in their entirety. Electronic health records of interest may be available from, for example, Epic, including offerings such as MyChart. Electronic health records may include structured health record data, unstructured health record data or a combination thereof. In some embodiments, electronic health records include both structured and unstructured electronic health record data. By “structured health record,” it is meant that a subject’s health data is presented in a structured format, such as a database format, with, for example, labeled or tagged numerical results, or otherwise organized in a regular or structured format for ease of sorting, searching and analysis. By “unstructured health record,” it is meant a subject’s health data that is not presented in a structured format, such as notes, descriptions, drawings or other markings that are not, for example, numerical or responsive to a specific, structured database field. In embodiments, unstructured health record data comprises information in a free-text format. Unstructured health record data present in notes associated with a health record may be extracted using natural language processing (NLP) technology. In embodiments, unstructured health record data may be complementary to structured health record data in connection with training machine learning algorithms. In some embodiments, electronic health records include one or more of: demographics (birth year, gender, race and ethnicity), clinical concepts (conditions, drug exposures, abnormal measures), and visit-related features (age at prediction, first visit age, years in the electronic health records) capable of being extracted from the electronic health records database. In some embodiments, the electronic health records include sexes (i.e., gender) of the subjects. In some embodiments, electronic health record data may cover the entire patient health history through signs, symptoms, medical diagnosis and evaluation or aspects or subsets thereof. In some embodiments, electronic health record data may be incomplete with respect to certain aspects of a subject’s medical history. As reviewed above, a plurality of electronic health records include health information about a plurality of subjects. In some embodiments, the plurality of subjects includes subjects diagnosed with a disease (e.g., Alzheimer’s disease) and subjects not diagnosed with the same disease (e.g., Alzheimer’s disease). In other words, the health records include health records for subjects diagnosed with a disease (e.g., Alzheimer’s disease) and health records for subjects not diagnosed with the same disease (e.g., Alzheimer’s disease). Upon completing step 101 in FIG. 1 of flow diagram 100, the process next moves to step 102 of flow diagram 100. At step 102, the subjects are divided into a first set for machine learning model training and a second set for machine learning evaluation. Subjects may be divided into a first set for machine learning model training and a second set for machine learning evaluation according to any convenient ratio. The ratio of the number of subjects in the first set for machine learning model training to the number of subjects in the second set for machine learning evaluation may be 90:10, 80:20, 75:25, 70:30, 60:40, 50:50, 40:60, 30:70, 25:75, 20:80 or 10:90. In some embodiments, the ratio of the number of subjects in the first set for machine learning model training to the number of subjects in the second set for machine learning evaluation is 70:30. Upon completing step 102 in FIG. 1 of flow diagram 100, the process next moves to step 103 of flow diagram 100. At step 103, a machine learning model is trained using the electronic health records of the first set of subjects. Any convenient machine learning model may be trained using the electronic health records of the first set of subjects. Models may be trained using an unsupervised learning technique, a semi-supervised learning technique, a supervised learning technique, a parallel training technique, a round robin training technique, an attentionbased training technique and combinations thereof, as each such technique is known in the art. In some embodiments, the model is trained using one or more of: supervised training, unsupervised training or semi-supervised training. In some embodiments, the model is trained using unsupervised training. Unsupervised learning is a machine learning technique known in the art for training a model to, for example, identify or recognize patterns. Unsupervised learning comprises training a model where pre-assigned labels are not provided to the model with respect to data used to train the model. As a result, applying unsupervised learning to train a model entails the model itself discovering patterns among the training data. In other embodiments, the model is trained using semi-supervised learning. Semi-supervised learning is a machine learning technique known in the art for training a model to, for example, identify or recognize patterns. Semi-supervised learning comprises training a model using both labeled and unlabeled training data. In other embodiments, the model is trained using supervised learning. Supervised learning is a machine learning technique known in the art for training a model to, for example, identify or recognize patterns. Supervised learning involves training a model where labels are provided to the model. In some embodiments, the model is trained using round robin training. By “round robin training,” it is meant that input data to the model (i.e., data used to train the model) is divided into multiple partitions such that certain partitions are used to train the model and the remaining partitions are used to generate predictions using the trained model. Such processes may be iterated where partitions of data previously used to train the model are subsequently used to generate predictions using the trained model. Round robin training approaches may offer benefits including identifying which data sets used for training result in more accurate predictions. In embodiments, the type of machine learning model is selected based on an ease of interpretability of the model and / or an ability to capture nonlinear relationships. In some cases, interpretability is based on the ability to map diagnostic features to phecodes. Phecodes represent a high-throughput phenotyping tool based on international classification of diseases (ICD) codes. See, e.g., Bastarache, L. Using Phecodes for Research with the Electronic Health Record: From PheWAS to PheRS. Annu Rev Biomed Data Sci. 2021 Jul 20;4:1 -19., the disclosure of which is herein incorporated by reference. By non-linear relationships, it is meant relationships between features (e.g., features mapped to phecodes) and a disease (e.g., Alzheimer’s disease) that are non-linear. In some embodiments, the machine learning model comprises a binary classification time point model. By binary classification time point model, it is meant a model that categorizes data into two groups (e.g., at risk of disease onset or not at risk of disease onset) at various points in time before an index time. The index time may be the time point (e.g., age of the subject) at the onset of the disease (e.g., the onset of Alzheimer’s disease; i.e., clinical onset of the disease), or another time point of interest (e.g., one year prior to onset of the disease or time of list clinical visit, etc.). In other embodiments, the machine learning model comprises a random forest model. In still other embodiments, the machine learning model comprises a plurality of random forest models. Random forest models employ multiple decision trees to reach a single result. Each tree of the multiple decision trees may be averaged, or otherwise incorporated together with one or more separate decisions trees, to arrive at a result, e.g., a classification, such as a predicted disease onset risk. In embodiments, the machine learning model comprises a plurality of potential machine learning models, and training and evaluating the machine learning model comprises selecting or modifying one or more of the potential machine learning models. In some embodiments, training and evaluating the machine learning model comprises configuring the model to improve an accuracy of predicting disease onset risk. In other embodiments, training and evaluating the machine learning model comprises adjusting one or more parameters of the machine learning model. In some embodiments, training the machine learning model using the electronic health records of the first set of subjects comprises training the model using samples of subjects diagnosed with the disease and subjects not diagnosed with the disease. In certain embodiments, training the machine learning model using the electronic health records of the first set of subjects comprises training the model using: (a) clinical features only of the electronic health records, (b) clinical features and demographic information of the electronic health records, or (c) clinical features, demographic information and visit-related information of the electronic health records. In some cases, training the machine learning model using the electronic health records of the first set of subjects comprises training the model using samples of subjects matched by demographic information or hospital utilization or visit-related features. In other cases, matching subjects based on demographics comprises matching subjects based on one or more of: birth year, race and ethnicity and / or sex. In still other cases, matching subjects based on visit-related features comprises matching subjects based on one or more of: age, first visit age, years electronic health records, a value corresponding to a logarithm of the number of prior visits, a value corresponding to a logarithm of the number of prior concepts or a value corresponding to a logarithm of a number of days since a first clinical event. In embodiments, training the machine learning model using the electronic health records of the first set of subjects comprises training the model using samples of subjects matched using a propensity score match. In some embodiments, training the machine learning model using the electronic health records of the first set of subjects comprises training the model using samples of subjects of a single sex. In some embodiments, training the machine learning model comprises training the model using sex-stratified subgroups. Upon completing step 103 in FIG. 1 of flow diagram 100, the process next moves to step 104 of flow diagram 100. At step 104, the machine learning model is evaluated using the electronic health records of the second set of subjects. In other words, the electronic health records of the second set of subjects are evaluated using the second set of subjects (i.e., the set that was held out and not used for training the model). In some embodiments, evaluating the model includes assessing the accuracy of, or confidence in, predictions made by the model. That is, whether a prediction based on the model of the disease onset risk in a subject represents a true positive or a false positive and / or whether a prediction based on the model that the disease is not present in (or a referral is not warranted for) a subject represents a true negative or a false negative, as well as how frequently such errors occur. Upon completing step 104 in FIG. 1 of flow diagram 100, the process next moves to step 105 of flow diagram 100. At step 105, one or more features of the subjects that predict the disease onset risk are identified from the trained and evaluated machine learning model. In embodiments where the electronic health records used to train and evaluate include sexes of the subject, one or more sex-specific features of the subjects that predict the disease onset risk may be identified from the trained and evaluated machine learning model. By “features”, it is meant characteristics of the disease (e.g., Alzheimer’s disease) that are present in a subject’s health records. The terms “feature” or “indicator” in the context of disease feature or disease indicator may be used interchangeably throughout this disclosure. For example, in some embodiments, features may include clinically relevant signs or symptoms that may or may not be directly related to the disease. In some embodiments, features may include other diseases, disorders or conditions that may or may not be directly related to the disease. In some embodiments, the identified features of the subjects comprises one or more phenotypes presented in the electronic health records. In some embodiments, the identified features of the subjects comprises one or more diagnostic features that have been mapped to phecodes. By “sex-specific features” it is meant characteristics of the disease (e.g., Alzheimer’s disease) that are present in a subject’s health records, where the subject’s sex is known. In embodiments where one or more sex-specific features of the subjects predict the disease onset risk, the sex-specific features are specific to the prediction in a particular sex (e.g., male or female). For example, female-specific features and malespecific features may be the same or different. In embodiments where one or more sex-specific features of the subjects that predict the disease onset risk (e.g., Alzheimer’s disease onset risk) may be identified from the trained and evaluated machine learning model, the one or more sex-specific features may be selected from osteoporosis, major depressive disorder, allergic rhinitis, abnormal stool contents, chest pain, hypovolemia and prostate hyperplasia. In embodiments where the subjects are female, the one or more sex-specific features (e.g., female-specific features) may be selected from osteoporosis, major depressive disorder, allergic rhinitis and abnormal stool contents. In embodiments where the subjects are male, the one or more sex-specific features (e.g., male-specific features) are selected from chest pain, hypovolemia and prostate hyperplasia. In embodiments, identifying from the trained and evaluated machine learning model one or more features of the subjects that predict disease onset risk comprises identifying an average importance for a plurality of features across a plurality of time points. In some embodiments, identifying from the trained and evaluated machine learning model one or more features of the subjects that predict disease onset risk comprises ranking a plurality of features at a plurality of different time points. In other embodiments, identifying from the trained and evaluated machine learning model one or more features of the subjects that predict disease onset risk comprises identifying features of subjects at different time points that predict disease onset risk. Other Methods Also provided are methods of predicting a disease onset risk in a subject. In some embodiments, methods of predicting a disease onset risk in a subject include obtaining a health record comprising health information about a subject, where the health record includes a sex of the subject and applying the model trained according to the method of the disclosure to the health record of the subject to predict the disease onset risk in the subject. In some embodiments, methods of predicting a disease onset risk in a subject include obtaining a health record comprising health information about a subject, wherein the health record includes a sex of the subject and predicting the disease onset risk in a subject based on the presence of one or more sex-specific features in the health record that has been identified to predict the disease onset risk based on the methods of the present disclosure. In embodiments where the disease is Alzheimer’s disease, such sex-specific features include, but are not limited to those discussed above (e.g., osteoporosis, major depressive disorder, allergic rhinitis, abnormal stool contents, chest pain, hypovolemia and prostate hyperplasia). Also provided are methods of clinically evaluating the presence of a disease in a subject. In some embodiments, methods of clinically evaluating the presence of a disease in a subject include obtaining a health record comprising health information about a subject, wherein the health record includes a sex of the subject, applying the model trained according to the methods of the present disclosure to the health record of the subject to predict the risk of developing the disease and clinically evaluating the presence of the disease, in the event that the subject is predicted to be at risk of developing the disease. In some embodiments, methods of clinically evaluating the presence of a disease in a subject include obtaining a health record comprising health information about a subject, wherein the health record includes a sex of the subject, predicting the disease onset risk in a subject based on the presence of one or more sex-specific features in the health record that has been identified to predict disease onset risk based on the methods of the present disclosure and clinically evaluating the presence of the disease, in the event that subject is predicted to be at risk of developing the disease. In embodiments, clinically evaluating the presence of the disease comprises clinical evaluation of the subject by a specialist. In some cases, clinically evaluating the presence of the disease comprises conducting biochemical analysis (e.g., analysis of blood or analysis of cerebrospinal fluid). In some cases, clinically evaluating the presence of the disease includes conducting imaging (e.g., magnetic resonance imaging (MRI), computerized tomography (CT) or positron emission tomography (PET)). In some cases, clinically evaluating the presence of disease includes cognitive skill testing, memory testing and neuropsychological testing. In cases where the disease to be clinically evaluated is Alzheimer’s disease, clinical evaluation may include one or more of the following: biochemical analysis, imaging and testing (e.g., cognitive skill testing, memory testing and / or neuropsychological testing). Also provided are methods of treating a subject predicted to be at risk of developing a disease. In some embodiments, methods of treating a subject predicted to be at risk of developing a disease include obtaining a health record comprising health information about a subject, where the health record includes a sex of the subject, applying the model trained according to the methods of the present disclosure to the health record of the subject to predict the risk of developing the disease and administering a treatment to the subject in an amount effective to treat the disease, in the event that subject is predicted to be at risk of developing the disease. In some embodiments, methods of treating a subject predicted to be at risk of developing a disease include obtaining a health record comprising health information about a subject, where the health record includes a sex of the subject, predicting the disease onset risk in a subject based on the presence of one or more sex-specific features in the health record that has been identified to predict disease onset risk based on the methods of the present disclosure and administering a treatment to the subject in an amount effective to treat the disease, in the event that subject is predicted to be at risk of developing the disease. By “treat” or “treatment” is meant at least an amelioration of the symptoms associated with the disease (e.g., Alzheimer’s disease), where amelioration is used in a broad sense to refer to at least a reduction in the magnitude of a parameter, e.g., symptom, associated with the disease being treated. As such, treatment also includes situations where the disease, or at least symptoms associated therewith, are completely inhibited, e.g., prevented from happening, or stopped, e.g., terminated, such that the subject no longer suffers from the disease, or at least the symptoms that characterize the disease. The treatment to be administered to a subject may include known treatments for the disease. In cases where the disease is Alzheimer’s disease, treatments include, but are not limited to, cholinesterase inhibitors (e.g., donepezil, galantamine, rivastigmine, tacrine), memantine, lecanemab and donanemab. In some embodiments, the method is a method of identifying a candidate subject for an early intervention for the disease. In some embodiments, the method is a method of identifying a candidate subject for a disease intervention prior to disease diagnosis. Further details regarding aspects of embodiments of the present disclosure are discussed in: Tang, A.S., Rankin, K.P., Cerono, G. et al. Leveraging electronic health records and knowledge networks for Alzheimer’s disease prediction and sex-specific biological insights. Nat Aging 4, 379-395 (2024). https: / / doi.org / 10.1038 / s43587-024-00573-8, incorporated herein in its entirety. Computer-implemented Embodiments The various method and algorithm steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system applying a method according to the present disclosure. The described functionality can be implemented in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the disclosure. The various illustrative steps, components, and computing systems (such as devices, databases, interfaces, and engines) described in connection with the embodiments disclosed herein can be implemented or performed by a machine, such as a general purpose processor, a graphics processor unit, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor can be a microprocessor, but in the alternative, the processor can be a controller, microcontroller, or state machine, combinations of the same, or the like. A processor can also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Although described herein primarily with respect to digital technology, a processor can also include primarily analog components. A computing environment can include any type of computer system, including, but not limited to, a computer system based on a microprocessor, a graphics processor unit, a mainframe computer, a digital signal processor, a portable computing device, a personal organizer, a device controller, and a computational engine within an appliance, to name a few. The steps of a method, process, or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module, engine, and associated databases can reside in memory resources such as in RAM memory, FRAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, a removable disk, a CD-ROM, or any other form of non-transitory computer-readable storage medium, media, or physical computer storage known in the art. An external storage medium can be coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium can be integral to the processor. The processor and the storage medium can reside in an ASIC. The ASIC can reside in a user terminal. In the alternative, the processor and the storage medium can reside as discrete components in a user terminal. As described in detail above, embodiments of the present invention relate to computer-implemented methods for predicting a disease onset risk using electronic health records. Aspects of the present invention include methods of training a model to predict a disease onset risk comprising: obtaining a plurality of electronic health records comprising health information about a plurality of subjects, wherein the plurality of subjects comprise subjects diagnosed with the disease and subjects not diagnosed with the disease, dividing the subjects into a first set for machine learning model training and a second set for machine learning evaluation, training a machine learning model using the electronic health records of the first set of subjects, evaluating the machine learning model using the electronic health records of the second set of subjects, and identifying from the trained and evaluated machine learning model one or more features of the subjects that predict disease onset risk. Aspects of the present invention include methods of training a model to predict a disease onset risk comprising: obtaining a plurality of electronic health records comprising health information about a plurality of subjects, wherein the plurality of subjects comprise subjects diagnosed with the disease and subjects not diagnosed with the disease, and wherein the electronic health records include sexes of the subjects, dividing the subjects into a first set for machine learning model training and a second set for machine learning evaluation, training a machine learning model using the electronic health records of the first set of subjects, evaluating the machine learning model using the electronic health records of the second set of subjects and identifying from the trained and evaluated machine learning model one or more sex-specific features of the subjects that predict the disease onset risk. Aspects of the present invention further include methods of predicting a disease onset risk in a subject, methods of clinically evaluating the presence of a disease in a subject and methods of treating a subject predicted to be at risk of developing a disease. Systems for Training a Model to Predict a Disease Onset Risk Aspects of the present disclosure further include systems, such as computer-controlled systems, for practicing embodiments of the above methods. Embodiments of systems of the invention comprise a processor comprising memory operably coupled to the processor, wherein the memory comprises instructions stored thereon, which, when executed by the processor, cause the processor to: obtain a plurality of electronic health records comprising health information about a plurality of subjects, wherein the plurality of subjects comprise subjects diagnosed with the disease and subjects not diagnosed with the disease; divide the subjects into a first set for machine learning model training and a second set for machine learning evaluation; train a machine learning model using the electronic health records of the first set of subjects; evaluate the machine learning model using the electronic health records of the second set of subjects; and identify from the trained and evaluated machine learning model one or more features of the subjects that predict disease onset risk. Embodiments of systems of the invention comprise a processor comprising memory operably coupled to the processor, wherein the memory comprises instructions stored thereon, which, when executed by the processor, cause the processor to: obtain a plurality of electronic health records comprising health information about a plurality of subjects, wherein the plurality of subjects comprise subjects diagnosed with the disease and subjects not diagnosed with the disease, and wherein the electronic health records include sexes of the subjects; divide the subjects into a first set for machine learning model training and a second set for machine learning evaluation; train a machine learning model using the electronic health records of the first set of subjects; evaluate the machine learning model using the electronic health records of the second set of subjects; and identify from the trained and evaluated machine learning model one or more sex-specific features of the subjects that predict the disease onset risk. In some embodiments, the disease is a neurological disease (e.g., Alzheimer’s disease). In some embodiments, the one or more sex-specific features are selected from osteoporosis, major depressive disorder, allergic rhinitis, abnormal stool contents, chest pain, hypovolemia and prostate hyperplasia. In some embodiments, the subject is female and the one or more sex-specific features are selected from osteoporosis, major depressive disorder, allergic rhinitis and abnormal stool contents. In some embodiments, the subject is male and the one or more sex-specific features are selected from chest pain, hypovolemia and prostate hyperplasia. In some embodiments, the type of machine learning model is selected based on an ease of interpretability of the model and an ability to capture nonlinear relationships. In some embodiments, the machine learning model comprises a binary classification time point model. In some embodiments, the machine learning model comprises a random forest model. In some embodiments, the machine learning model comprises a plurality of random forest models. In some embodiments, training and evaluating the machine learning model comprises configuring the model to improve an accuracy of predicting disease onset risk. In some embodiments, training and evaluating the machine learning model comprises adjusting one or more parameters of the machine learning model. In some embodiments, training the machine learning model comprises training the model using sex-stratified subgroups. In some embodiments, training the machine learning model using the electronic health records of the first set of subjects comprises training the model using: (a) clinical features only of the electronic health records, (b) clinical features and demographic information of the electronic health records, or (c) clinical features, demographic information and visit-related information of the electronic health records. In some embodiments, training the machine learning model using the electronic health records of the first set of subjects comprises training the model using samples of subjects matched by demographic information or hospital utilization or visit-related features. In some embodiments, matching subjects based on demographics comprises matching subjects based on one or more of: birth year, race and ethnicity and / or sex. In some embodiments, matching subjects based on visit-related features comprises matching subjects based on one or more of: age, first visit age, years electronic health records, a value corresponding to a logarithm of the number of prior visits, a value corresponding to a logarithm of the number of prior concepts or a value corresponding to a logarithm of a number of days since a first clinical event. In some embodiments, training the machine learning model using the electronic health records of the first set of subjects comprises training the model using samples of subjects matched using a propensity score match. In some embodiments, training the machine learning model using the electronic health records of the first set of subjects comprises training the model using samples of subjects of a single sex. In some embodiments, identifying from the trained and evaluated machine learning model one or more features of the subjects that predict disease onset risk comprises identifying an average importance for a plurality of features across a plurality of time points. In some embodiments, identifying from the trained and evaluated machine learning model one or more features of the subjects that predict disease onset risk comprises ranking a plurality of features at a plurality of different time points. In some embodiments, identifying from the trained and evaluated machine learning model one or more features of the subjects that predict disease onset risk comprises identifying features of subjects at different time points that predict disease onset risk. In some embodiments, identifying from the trained and evaluated machine learning model one or more features of the subjects that predict disease onset risk comprises identifying features of subjects of points that predict disease onset risk corresponding to a sex of the subjects. In some embodiments, the identified features of the subjects comprises one or more phenotypes presented in the electronic health records. In some embodiments, the identified features of the subjects comprises one or more diagnostic features that have been mapped to phecodes. In some embodiments, the memory further comprises instructions stored thereon, which, when executed by the processor, cause the processor to: obtain a health record comprising health information about a subject, wherein the health record includes a sex of the subject; and apply the trained model to the health record of the subject to predict the disease onset risk in the subject. In some embodiments, the memory further comprises instructions stored thereon, which, when executed by the processor, cause the processor to: obtain a health record comprising health information about a subject, wherein the health record includes a sex of the subject; apply the trained model trained to the health record of the subject to predict the risk of developing the disease; and clinically evaluate the presence of the disease, in the event that subject is predicted to be at risk of developing the disease. In some embodiments, the memory further comprises instructions stored thereon, which, when executed by the processor, cause the processor to: obtain a health record comprising health information about a subject, wherein the health record includes a sex of the subject; apply the trained model to the health record of the subject to predict the risk of developing the disease; and administer a treatment to the subject in an amount effective to treat the disease, in the event that subject is predicted to be at risk of developing the disease. In some embodiments, the memory further comprises instructions stored thereon, which, when executed by the processor, cause the processor to: obtain a health record comprising health information about a subject, wherein the health record includes a sex of the subject; and predict the disease onset risk in a subject based on the presence of one or more sex-specific features in the health record that has been identified to predict the disease onset risk based on the trained model. In some embodiments, the memory further comprises instructions stored thereon, which, when executed by the processor, cause the processor to: obtain a health record comprising health information about a subject, wherein the health record includes a sex of the subject; predict the disease onset risk in a subject based on the presence of one or more sex-specific features in the health record that has been identified to predict disease onset risk based the trained model; and clinically evaluate the presence of the disease, in the event that subject is predicted to be at risk of developing the disease. In some embodiments, the memory further comprises instructions stored thereon, which, when executed by the processor, cause the processor to: obtain a health record comprising health information about a subject, wherein the health record includes a sex of the subject; predict the disease onset risk in a subject based on the presence of one or more sex-specific features in the health record that has been identified to predict disease onset risk based on the trained model; and administer a treatment to the subject in an amount effective to treat the disease, in the event that subject is predicted to be at risk of developing the disease. In some instances, the systems further include one or more computers for complete automation or partial automation of the methods described herein. In some embodiments, systems include a computer having a computer readable storage medium with a computer program stored thereon. In embodiments, the system includes an input module, a processing module and an output module. The subject systems may include both hardware and software components, where the hardware components may take the form of one or more platforms, e.g., in the form of servers, such that the functional elements, i.e., those elements of the system that carry out specific tasks (such as managing input and output of information, processing information, etc.) of the system may be carried out by the execution of software applications on and across the one or more computer platforms represented of the system. Systems may include a display and operator input device. Operator input devices may, for example, be a keyboard, mouse, or the like. The processing module includes a processor which has access to a memory having instructions stored thereon for performing the steps of the subject methods. The processing module may include an operating system, a graphical user interface (GUI) controller, a system memory, memory storage devices, and input-output controllers, cache memory, a data backup unit, and many other devices. The processor may be a commercially available processor or it may be one of other processors that are or will become available. The processor executes the operating system and the operating system interfaces with firmware and hardware in a well-known manner, and facilitates the processor in coordinating and executing the functions of various computer programs that may be written in a variety of programming languages, such as Java, Perl, C++, other high level or low level languages, as well as combinations thereof, as is known in the art. The operating system, typically in cooperation with the processor, coordinates and executes functions of the other components of the computer. The operating system also provides scheduling, input-output control, file and data management, memory management, and communication control and related services, all in accordance with known techniques. The processor may be any suitable analog or digital system. In some embodiments, processors include analog electronics which allows the user to manually align a light source with the flow stream based on the first and second light signals. In some embodiments, the processor includes analog electronics which provide feedback control, such as for example negative feedback control. The system memory may be any of a variety of known or future memory storage devices. Examples include any commonly available random access memory (RAM), magnetic medium such as a resident hard disk or tape, an optical medium such as a read and write compact disc, flash memory devices, or other memory storage device. The memory storage device may be any of a variety of known or future devices, including a compact disk drive, a tape drive, a removable hard disk drive, or a diskette drive. Such types of memory storage devices typically read from, and / or write to, a program storage medium (not shown) such as, respectively, a compact disk, magnetic tape, removable hard disk, or floppy diskette. Any of these program storage media, or others now in use or that may later be developed, may be considered a computer program product. As will be appreciated, these program storage media typically store a computer software program and / or data. Computer software programs, also called computer control logic, typically are stored in system memory and / or the program storage device used in conjunction with the memory storage device. In some embodiments, a computer program product is described comprising a computer usable medium having control logic (computer software program, including program code) stored therein. The control logic, when executed by the processor the computer, causes the processor to perform functions described herein. In other embodiments, some functions are implemented primarily in hardware using, for example, a hardware state machine. Implementation of the hardware state machine so as to perform the functions described herein will be apparent to those skilled in the relevant arts. Memory may be any suitable device in which the processor can store and retrieve data, such as magnetic, optical, or solid-state storage devices (including magnetic or optical disks or tape or RAM, or any other suitable device, either fixed or portable). The processor may include a general-purpose digital microprocessor suitably programmed from a computer readable medium carrying necessary program code. Programming can be provided remotely to processor through a communication channel, or previously saved in a computer program product such as memory or some other portable or fixed computer readable storage medium using any of those devices in connection with memory. For example, a magnetic or optical disk may carry the programming, and can be read by a disk writer / reader. Systems of the invention also include programming, e.g., in the form of computer program products, algorithms for use in practicing the methods as described above. Programming according to the present invention can be recorded on computer readable media, e.g., any medium that can be read and accessed directly by a computer. Such media include, but are not limited to: magnetic storage media, such as floppy discs, hard disc storage medium, and magnetic tape; optical storage media such as CD-ROM; electrical storage media such as RAM and ROM; portable flash drive; and hybrids of these categories such as magnetic / optical storage media. The processor may also have access to a communication channel to communicate with a user at a remote location. By remote location is meant the user is not directly in contact with the system and relays input information to an input manager from an external device, such as a computer connected to a Wide Area Network (“WAN”), telephone network, satellite network, or any other suitable communication channel, including a mobile telephone (i.e., smartphone). In some embodiments, systems according to the present disclosure may be configured to include a communication interface. In some embodiments, the communication interface includes a receiver and / or transmitter for communicating with a network and / or another device. The communication interface can be configured for wired or wireless communication, including, but not limited to, radio frequency (RF) communication (e.g., Radio-Frequency Identification (RFID), Zigbee communication protocols, WiFi, infrared, wireless Universal Serial Bus (USB), Ultra Wide Band (UWB), Bluetooth® communication protocols, and cellular communication, such as code division multiple access (CDMA) or Global System for Mobile communications (GSM). In one embodiment, the communication interface is configured to include one or more communication ports, e.g., physical ports or interfaces such as a USB port, an RS-232 port, or any other suitable electrical connection port to allow data communication between the subject systems and other external devices such as a computer terminal (for example, at a physician’s office or in hospital environment) that is configured for similar complementary data communication. In one embodiment, the communication interface is configured for infrared communication, Bluetooth® communication, or any other suitable wireless communication protocol to enable the subject systems to communicate with other devices such as computer terminals and / or networks, communication enabled mobile telephones, personal digital assistants, or any other communication devices which the user may use in conjunction. In one embodiment, the communication interface is configured to provide a connection for data transfer utilizing Internet Protocol (IP) through a cell phone network, Short Message Service (SMS), wireless connection to a personal computer (PC) on a Local Area Network (LAN) which is connected to the internet, or WiFi connection to the internet at a WiFi hotspot. In one embodiment, the subject systems are configured to wirelessly communicate with a server device via the communication interface, e.g., using a common standard such as 802.11 or Bluetooth® RF protocol, or an IrDA infrared protocol. The server device may be another portable device, such as a smart phone, Personal Digital Assistant (PDA) or notebook computer; or a larger device such as a desktop computer, appliance, etc. In some embodiments, the server device has a display, such as a liquid crystal display (LCD), as well as an input device, such as buttons, a keyboard, mouse or touch-screen. In some embodiments, the communication interface is configured to automatically or semi-automatically communicate data stored in the subject systems, e.g., in an optional data storage unit, with a network or server device using one or more of the communication protocols and / or mechanisms described above. Output controllers may include controllers for any of a variety of known display devices for presenting information to a user, whether a human or a machine, whether local or remote. If one of the display devices provides visual information, this information typically may be logically and / or physically organized as an array of picture elements. A graphical user interface (GUI) controller may include any of a variety of known or future software programs for providing graphical input and output interfaces between the system and a user, and for processing user inputs. The functional elements of the computer may communicate with each other via system bus. Some of these communications may be accomplished in alternative embodiments using network or other types of remote communications. The output manager may also provide information generated by the processing module to a user at a remote location, e.g., over the Internet, phone or satellite network, in accordance with known techniques. The presentation of data by the output manager may be implemented in accordance with a variety of known techniques. As some examples, data may include SQL, HTML or XML documents, email or other files, or data in other forms. The data may include Internet URL addresses so that a user may retrieve additional SQL, HTML, XML, or other documents or data from remote sources. The one or more platforms present in the subject systems may be any type of known computer platform or a type to be developed in the future, although they typically will be of a class of computer commonly referred to as servers. However, they may also be a main-frame computer, a workstation, or other computer type. They may be connected via any known or future type of cabling or other communication system including wireless systems, either networked or otherwise. They may be co-located or they may be physically separated. Various operating systems may be employed on any of the computer platforms, possibly depending on the type and / or make of computer platform chosen. Appropriate operating systems include Windows, iOS, Oracle Solaris, Linux, IBM i, Unix, and others. FIG. 2 shows a functional block diagram for one example of a computer system 200 for practicing methods of the present invention, i.e., a processor operably connected to memory, 202, for predicting a disease onset risk. A processor and memory 202 can be configured to implement a variety of processes for implementing and training a model, for example, in connection with predicting a disease onset risk. An apparatus, 212 can be configured to acquire training data, such as electronic health records of subjects. For example, apparatus 212 may be a remote database, and processor 202 may be operably connected to such remote database to acquire electronic health records for use in training and / or evaluating one or more models for predicting disease onset risk. A data communication channel can be included between the apparatus 212 and the processor 202. Electronic health record data can be provided to the processor 202 via the data communication channel. The processor 202 can be configured to provide a graphical display, such as one or more histograms or a first plot of data illustrating aspects of the health data or results of training model to predict disease onset risk to a display device 206. For example, processor 202 can be configured to cause a display device 206 to display one or more cluster diagrams or bar graphs or other probability distributions or the like related to results of training and evaluating a model. The processor 202 can be further configured to display data on the display device 206 related to predictive features of health records differently from other features that are not predictive or not as predictive (e.g., a coloring scale with gradations corresponding to a degree a health record feature is indicative of a disease onset risk). For example, the processor 202 can be configured to render the color of, for example, a feature of an electronic health record such as a first electronic health record diagnosis, e.g., a first phecode, to be distinct from the color of a second electronic health record diagnosis, e.g., a second phecode, where the second diagnosis is not as predictive of disease onset risk as the first diagnosis. The display device 206 can be implemented as a monitor, a tablet computer, a smartphone, or other electronic device configured to present graphical interfaces. The processor 202 can be configured to receive adjustments to configuration settings from a first input device. Such adjustments to configuration settings, when received from the first input device to the processor 202 can be used to select or adjust aspects of one or more models or one or more aspects of training a model, such as selecting using one or more training sets or one or more models. In embodiments, training data may be adjusted in order to avoid overfitting a data set. The first input device can be implemented as a mouse 210, or the first device can be implemented as the keyboard 208 or other means for providing an input signal to the processor 202 such as a touchscreen, a stylus, an optical detector, or a voice recognition system. Some input devices can include multiple inputting functions. In such implementations, the inputting functions can each be considered an input device. For example, as shown in FIG. 2, the mouse 210 can include a right mouse button and a left mouse button, each of which can generate a triggering event. Such triggering event can cause the processor 202 to alter the manner in which the data is displayed, which portions of the data is actually displayed on the display device 206, and / or provide input to further processing such as selection of additional training data or evaluation data. The processor 202 can be connected to a storage device 204. The storage device 204 can be configured to receive and store training data or evaluation data, such as electronic health records, or aspects of electronic health records, from the processor 202. The storage device 204 can be further configured to allow retrieval of training and / or evaluation data, such as electronic health record data, by the processor 202. The display device 206 can be further configured to alter the information presented according to input received from the processor 202 in conjunction with input from the apparatus 212, the storage device 204, the keyboard 208, and / or the mouse 210. In some implementations, the processor 202 can generate a user interface for use with training and evaluating a model for predicting disease onset risk. For example, the user interface can include a control for applying certain training data or evaluation data to a model. FIG. 3 depicts a general architecture of an example computing device 300 according to certain embodiments. The general architecture of the computing device 300 includes an arrangement of computer hardware and software components. The computing device 300 may include many more (or fewer) elements than those shown in FIG. 3. It is not necessary, however, that all of these generally conventional elements be shown in order to provide an enabling disclosure. As illustrated, the computing device 300 includes a processing unit 310, a network interface 320, a computer readable medium 330, an input / output device interface 340, a display 350, and an input device 360, all of which may communicate with one another by way of a communication bus. The network interface 320 may provide connectivity to one or more networks or computing systems. The processing unit 310 may thus receive information and instructions from other computing systems or services via a network. The processing unit 310 may also communicate to and from memory 370 and further provide output information for an optional display 350 via the input / output device interface 340. The input / output device interface 340 may also accept input from the optional input device 360, such as a keyboard, mouse, digital pen, microphone, touch screen, gesture recognition system, voice recognition system, gamepad, accelerometer, gyroscope, or other input device. The memory 370 may contain computer program instructions (grouped as modules or components in some embodiments) that the processing unit 310 executes in order to implement one or more embodiments. The memory 370 generally includes RAM, ROM and / or other persistent, auxiliary or non-transitory computer-readable media. The memory 370 may store an operating system 372 that provides computer program instructions for use by the processing unit 310 in the general administration and operation of the computing device 300. The memory 370 may further include computer program instructions and other information for implementing aspects of the present disclosure. For example, in one embodiment, the memory 370 includes an obtaining a plurality of electronic health records module 372 for obtaining a plurality of electronic health records comprising health information about a plurality of subjects, wherein the plurality of subjects comprise subjects diagnosed with the disease and subjects not diagnosed with the disease, e.g., by accessing a database comprising references to electronic health records in data store 390; a dividing subjects into a training set and an evaluation set module 374 for dividing the subjects into a first set for machine learning model training and a second set for machine learning evaluation; a training and evaluating a machine learning model 376 for training a machine learning model using the electronic health records of the first set of subjects and / or evaluating the machine learning model using the electronic health records of the second set of subjects; and an identifying features module 378 for identifying from the trained and evaluated machine learning model one or more features of the subjects that predict disease onset risk. Aspects of the present disclosure further include non-transitory computer readable storage media for training a model to predict a disease onset risk using electronic health records data. Non-transitory computer readable storage media according to certain embodiments comprise one or more algorithms corresponding to the subject methods described herein. Utility The subject methods and systems find use in a variety of applications where it is desirable to train a model to predict a disease onset risk. Embodiments of the invention provide techniques for identifying one or more features of subjects’ health characteristics, such as features of electronic health records, that are predictive of disease onset risk, such as onset risk of Alzheimer’s disease. Embodiments of the invention provide techniques for using models (e.g., machine learning models, such as random forest models), to predict a disease (e.g., Alzheimer’s disease) onset risk. Early prediction of a disease enables earlier clinical evaluation and / or earlier treatment of the subject, thus allowing for disease progression to be slowed down or halted. Embodiments of the invention may be applied to any disease capable of prediction from available data, e.g., electronic health record data. In at least some of the previously described embodiments, one or more elements used in an embodiment can interchangeably be used in another embodiment unless such a replacement is not technically feasible. It will be appreciated by those skilled in the art that various other omissions, additions and modifications may be made to the methods and structures described above without departing from the scope of the claimed subject matter. All such modifications and changes are intended to fall within the scope of the subject matter, as defined by the appended claims. It will be understood by those within the art that, in general, terms used herein, and especially in the appended claims {e.g., bodies of the appended claims) are generally intended as “open” terms {e.g., the term “including” should be interpreted as “including but not limited to,” the term “having” should be interpreted as “having at least,” the term “includes” should be interpreted as “includes but is not limited to,” etc.). It will be further understood by those within the art that if a specific number of an introduced claim recitation is intended, such an intent will be explicitly recited in the claim, and in the absence of such recitation no such intent is present. For example, as an aid to understanding, the following appended claims may contain usage of the introductory phrases “at least one” and “one or more” to introduce claim recitations. However, the use of such phrases should not be construed to imply that the introduction of a claim recitation by the indefinite articles “a” or “an” limits any particular claim containing such introduced claim recitation to embodiments containing only one such recitation, even when the same claim includes the introductory phrases “one or more” or “at least one” and indefinite articles such as “a” or “an” (e.g., “a” and / or “an” should be interpreted to mean “at least one” or “one or more”); the same holds true for the use of definite articles used to introduce claim recitations. In addition, even if a specific number of an introduced claim recitation is explicitly recited, those skilled in the art will recognize that such recitation should be interpreted to mean at least the recited number (e.g., the bare recitation of “two recitations,” without other modifiers, means at least two recitations, or two or more recitations). Furthermore, in those instances where a convention analogous to “at least one of A, B, and C, etc.” is used, in general such a construction is intended in the sense one having skill in the art would understand the convention (e.g., “ a system having at least one of A, B, and C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). In those instances where a convention analogous to “at least one of A, B, or C, etc.” is used, in general such a construction is intended in the sense one having skill in the art would understand the convention (e.g., “ a system having at least one of A, B, or C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). It will be further understood by those within the art that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the description, claims, or drawings, should be understood to contemplate the possibilities of including one of the terms, either of the terms, or both terms. For example, the phrase “A or B” will be understood to include the possibilities of “A” or “B” or “A and B.” In addition, where features or aspects of the disclosure are described in terms of Markush groups, those skilled in the art will recognize that the disclosure is also thereby described in terms of any individual member or subgroup of members of the Markush group. As will be understood by one skilled in the art, for any and all purposes, such as in terms of providing a written description, all ranges disclosed herein also encompass any and all possible sub-ranges and combinations of sub-ranges thereof. Any listed range can be easily recognized as sufficiently describing and enabling the same range being broken down into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each range discussed herein can be readily broken down into a lower third, middle third and upper third, etc. As will also be understood by one skilled in the art all language such as “up to,” “at least,” “greater than,” “less than,” and the like include the number recited and refer to ranges which can be subsequently broken down into sub-ranges as discussed above. Finally, as will be understood by one skilled in the art, a range includes each individual member. Thus, for example, a group having 1-3 articles refers to groups having 1,2, or 3 articles. Similarly, a group having 1 -5 articles refers to groups having 1,2, 3, 4, or 5 articles, and so forth. Although the foregoing invention has been described in some detail by way of illustration and example for purposes of clarity of understanding, it is readily apparent to those of ordinary skill in the art in light of the teachings of this invention that certain changes and modifications may be made thereto without departing from the spirit or scope of the appended claims. Accordingly, the preceding merely illustrates the principles of the invention. It will be appreciated that those skilled in the art will be able to devise various arrangements which, although not explicitly described or shown herein, embody the principles of the invention and are included within its spirit and scope. Furthermore, all examples and conditional language recited herein are principally intended to aid the reader in understanding the principles of the invention and the concepts contributed by the inventors to furthering the art, and are to be construed as being without limitation to such specifically recited examples and conditions. Moreover, all statements herein reciting principles, aspects, and embodiments of the invention as well as specific examples thereof, are intended to encompass both structural and functional equivalents thereof. Additionally, it is intended that such equivalents include both currently known equivalents and equivalents developed in the future, i.e., any elements developed that perform the same function, regardless of structure. Moreover, nothing disclosed herein is intended to be dedicated to the public regardless of whether such disclosure is explicitly recited in the claims. The scope of the present invention, therefore, is not intended to be limited to the exemplary embodiments shown and described herein. Rather, the scope and spirit of present invention is embodied by the appended claims. In the claims, 35 U.S.C. §112(f) or 35 U.S.C. §112(6) is expressly defined as being invoked for a limitation in the claim only when the exact phrase "means for" or the exact phrase "step for" is recited at the beginning of such limitation in the claim; if such exact phrase is not used in a limitation in the claim, then 35 U.S.C. § 112 (f) or 35 U.S.C. §112(6) is not invoked. The following examples are offered by way of illustration and not by way of limitation. Experimental Example 1 ML models based on clinical data can accurately predict Alzheimer’s Disease onset up to 7 years in advance From the UCSF EHR database of over 5 million patients from 1980-2021, 2,996 AD patients who had undergone dementia evaluation at the Memory and Aging Center and thus had expert-level clinical diagnoses were identified and mapped to the UCSF Observational Medical Outcomes Partnership (OMOP) EHR database. From the remaining patients, 823,671 control patients were extracted with over a year of visits and no dementia diagnosis. After identifying an index time representing AD onset (mean onset age (SD) 74 (5.6), see Methods) and filtering for availability of at least 7 years of longitudinal data, 749 AD patients and 250,545 control patients were identified (demographics shown in Tables 1A, Table 1B). From that, 30% was held-out for model evaluation and 70% utilized for model training (FIG. 4B, FIG. 9). For each time point and within sex strata, ML models were either trained for AD onset prediction or trained on the AD cohort and a subset of propensity-score matched controls for hypothesis generation, where balancing was performed on demographics (sex, race & ethnicity, birth year, age) and visit-related factors (years in EHR, first EHR visit age, number of visits, number of EHR concepts, and days since first EHR record, data not shown; matched example in Tables 1A, Table 1B). All Filtered Patients (pre-test / train split) Control AD 11 250545 749 Age of AD onset (SD) 74.0 (5.6) Birth year, mean (SD) 1945.5 (10.2) 1933.9 (5.3) First visit age, mean (SD) 51.2(11.4) 57.0(10.4) Sex, n (%) Female 139548 (55.7) 468 (62.5) Male 110829 (44.2) 281 (37.5) Nonbinary / Unknown 168 (0.1) R&E, n(%) Asian / NHPI 32427 (12.9) 151 (20.2) Black 17111 (6.8) 62 (8.3) Latinx 15036 (6.0) 53(7.1) Other / Unknown 28177 (11.2) 45 (6.0) White 157794 (63.0) 438 (58.5) Table 1 A. Demographics of patients used in models, and an example matched cohort for the -1 year model. The table shows characteristics of patients in the UCSF EHR with visits and concepts over 7 years prior to index time. Matched Train Patients for -1 year model Control AD SMD n 4184 523 Birth year, mean (SD) 1934.2 (5.6) 1934.0 (5.3) -0.042 First visit age, mean (SD) 57.2 (9.4) 56.9(10.5) -0.028 AD onset / index time age, mean (SD) 74.1 (5.8) 74.1 (5.8) -0.002 Years in EHR, mean (SD) 15.9 (7.8) 15.9 (7.9) -0.004 Log(n prev visits), mean (SD) 3.6 (1.5) 3.7 (1.6) 0.065 Log(n concepts), mean (SD) 3.1 (1.3) 3.3 (1.4) 0.108 Log(days since first event), mean (SD) 8.5 (0.4) 8.5 (0.4) 0.043 Sex, n (%)                        Female 2343 (56.0) 317 (60.6) 0.094 Male 1841 (44.0) 206 (39.4) R&E, n (% )                 Asian / NHPI 705 (16.8) 112(21.4) 0.219 Table 1B. Demographics of patients used in models, and an example matched cohort for the -1 year model. The table shows an example of training data where AD and controls are matched by the listed characteristics. Race & ethnicity (R&E) is a single variable derived from an algorithm developed by the UCSF Data Equity Taskforce84. Random forest (RF) models trained on only clinical features from time points between -7 years to -1 day to AD onset were evaluated on the held-out dataset with average bootstrapped Area Under the Receiver Operating Characteristic (AUROC) curve between 0.72 (median 0.75) for the -7 year time model to 0.81 (median 0.85) for the -1 day model. The RF models performed with Area Under the Precision Recall Curve (AUPRC) greater than the reference held-out evaluation set AD prevalence of 0.003 (average / median of 0.05 / 0.01 for -7 year model and 0.10 / 0.06 for -1 day model, FIG. 4C). With addition of demographics and visit-related features, RF model performance improved with average bootstrapped AU ROC between 0.86 (median 0.89) to 0.90 (median 0.94) and AUPRC between mean 0.06 (median 0.04) and 0.27 (median 0.14) for the -7 year to -1 day model, respectively (FIG. 4C). Top decision features across each time point model (see Methods) included features across clinical data domains, including vaccines, abnormal feces content, hypertension, hyperlipidemia (HLD), and cataracts (FIG. 10A). Demographic and visit-related features became predictive for AD diagnosis when added to the model, which is not unexpected since these features may contribute to confounding that influence the identified features and predicted risk of AD diagnosis (FIG. 10A). EHR diagnoses mapped to phecode categories29 (see Methods) identified sense organs, circulatory, and musculoskeletal phecode categories for early models, and mental disorder category for late models (FIG. 10B). Among the clusters of top 50 ranked phecodes, one cluster identified phecode features that maintain high relative importance throughout the time models (HLD, hypertension, dizziness, abnormal stool contents), and other clusters contain features with relative importance at specific time points (FIG. 10C). While some of these features support prior identified AD risk factors, the lack of adjustment may lead to feature identification as proxies for age in risk determination but not directly relevant to disease pathogenesis. Therefore, disease relevant features were identified by training models on patients matched on demographics and hospital utilization. Models trained on matched cohorts can identify hypotheses for biologically relevant AD predictors To train models that are robust for AD prediction for identifying predictors without demographic and visit-related confounding, time point models were trained on a matched set of participants at a 1:8 ratio between AD and controls. Sufficient balance was achieved on numerical covariates that were highly important in unmatched demographic models (FIG. 11 and data not shown). RF models trained on only clinical features from -7 years to -1 day performed with average bootstrapped held-out evaluation set AU ROC between .58 (median 0.57) for the -7 year time model to .77 (median 0.77) for the -1 day time model. The models performed with AU PRC greater than the held-out evaluation set AD prevalence of 0.003 with improvement closer to time 0 (mean / median of 0.02 / 0.008 for -7 year time model and 0.08 / 0.03 for -1 day model, FIG. 5A). When demographics and visit-related information were added as features, the models performed with minimal improvement, with average bootstrapped test set AUROC between 0.61 (median 0.61) to 0.71 (median 0.72) and similar AUPRC (mean / median of 0.02 / 0.009 for -7 year time model and 0.05 / 0.03 for -1 day model, FIG. 5A). For both the full and matched cohort models, the relative performances are consistent for balanced accuracy measures on the held-out evaluation, and an example permutation test demonstrates significance for the -1 day matched cohort model (FIG. 15). Among top features sorted by average importance across time models, top features include amnesia and cognitive concerns, HLD, dizziness, cataract, congestive heart failure, osteoarthritis, and others (FIG. 5B). These top features are consistently important even when demographics and visit information was added to the model, although demographic and visit features still had minimal influence on prediction (FIG. 5B). Compared to models trained on all patients, the models trained on matched cohorts have increased importance assigned to features like hyperlipidemia and amnesia, while decreasing importance of features like pain intensity rating scale and essential hypertension (FIG. 14). Since matching allows for the control of the influence of visit and demographic-related on AD prediction, the remaining diagnoses features can be identified for hypothesis generation with greater specificity for AD predictive risk. Top phecode categories include mental disorders, sense organs, and endocrine / metabolic categories (FIG. 5C). Among clusters of specific phecodes, one cluster included features with maintained predictive importance throughout time models (HLD and congestive heart failure), while other clusters include phecodes that are relatively predictive several years prior to AD onset (osteoarthritis, allergic rhinitis). A cluster of features emerges as important around -3 years (osteoporosis, dizziness, back pain, hemorrhoids, palpitations), and some features only emerge as important closer to the time of AD onset (memory loss, vitamin D deficiency, FIG. 5C). Together, this shows that the model can identify a combination of conditions that can lead to AD risk identification for a patient of a given age and hospital utilization burden. Stratification by sex allows identification of features that are predictive within a subgroup Since sex plays a role in AD risk, models were trained within male or female-identified sex groups to perform sex-specific prediction and identify sex-specific predictive features, without and with matching on demographics and hospital utilization (data not shown). Models trained on clinical features performed with average held-out evaluation set ALIROC between 0.75 (median 0.76) and 0.71 (median 0.71) for -7 year female and male models to 0.84 (median 0.86) and 0.82 (0.89) for -1 day female and male models. For AU PRC, the models performed greater than the held-out evaluation set prevalence (0.0036 for females, 0.0023 for males) with performance of 0.056-0.11 (median 0.0220.061) and 0.041-0.15(median 0.015-0.056) for females and male -7 years to -1 day time models respectively. With addition of demographics and visit-related features, AUROC / AUPRC improved considerably (FIG. 12A).Top features include sense organs and musculoskeletal phecode categories in female-only models, and circulatory system and digestive phecode categories as important among male-only models (FIG. 12B). To identify sex-specific biologically relevant clinical predictors for hypothesis generation, models were also trained by matching on demographic and visit-related factors within each subgroup (data not shown). Time point models trained only on clinical features performed with mean held-out evaluation set AU ROC between 0.60-0.68 (median 0.58-0.74) and 0.41-0.75 (median 0.43-0.84) for female and male models respectively (FIG. 5D). For AUPRC, models performed greater than held-out evaluation set prevalence with performance ranging from 0.031-0.095 (median 0.0076-0.046) and 0.0040-0.125 (0.0033-0.022) for female and male models respectively. Slight improvement in performance was observed with the addition of demographics and visit-related information (FIG. 5D). Top phecode categories in the female models include respiratory / circulatory system features earlier on, to musculoskeletal features in the -5 year model, to sense organs and mental disorders in the -1 year and -1 day model. Top categories in male models include endocrine / metabolic / circulatory disorders earlier, to digestive and genitourinary in -5 and -3 models, to mental disorders in -1 day model (FIG. 12B). When comparing specific phecodes, some are general across the subgroups such as HLD, congestive heart failure (early models), and memory / cognitive symptoms (later models) (FIG. 5E, FIG. 12C). Female-driven features across time models included osteoporosis, palpitations, allergic rhinitis, myocardial infarction, major depressive disorder, and abnormal stool contents. Male-driven features included chest pain, hypovolemia, sexual disorder, tobacco use disorder, and neoplasms (FIG. 5E). For all formulations of the prediction task, logistic regressions (LR) models performed comparably to random forests, and identify features with linear relationships with AD including some overlap with features identified from random forest models (FIG. 13). Nevertheless, for matched cohort models random forest performs better than logistic regression at the same time points (data not shown) and can identify decision features with nonlinear relationships with AD (e.g., RF identifies osteoporosis). Balanced accuracy measures for all of the random forest models support trends in performance between models, including lower overall performance for matched cohort models, and improvement in model performance closer to onset of AD (FIG. 15A and data not shown). As an example to evaluate the extent that clinical features meaningfully predict AD, random forest models were retrained on permutations of the groundtruth label for the -1 day matched cohort (40 permutations) and the trained model distribution was significant compared to the null distribution (p=0.024, FIG. 15B). Use of a knowledge graph allowed identification of potential biological explanations underlying predictive features Next, the SPOKE knowledge graph28 was utilized in order to utilize existing knowledge to explain biological relationships between groups of top clinical model features and AD. Biological features (genes, proteins, compounds, etc.) were mapped between top 25 clinical predictors (mapped to disease nodes) and AD node for each model (see Methods). Genes that appear in shortest path networks among matched models across multiple time include APOE, AKT1, INS, ALB, IL1B, INF, ALB, IL6, SOD1, etc. and compounds include atorvastatin, simvastatin, ergocalciferol, progesterone, estrogen, cyanocobalamin, and folic acid (FIG. 6). These genes and compounds also share relationships to multiple occurring model input nodes, particularly familial hyperlipidemia and osteoporosis among all time point models (FIG. 6). Notable nodes that appear over at least 2 models include C9orf72, TREM2, APP, MAPT with relationships to input nodes of musculoskeletal and joint disorders, deafness, and depression (FIG. 6). Hyperlipidemia validates as a top predictor of AD in external EHRs and a genetic link confirmed in APOE locus In order to further validate the utility of models to identify predictive disease associations, it was investigated whether using HLD as a top feature was a consistent predictor across all models. Utilizing a retrospective cohort study design in an EHR on five hospitals across the University of California system (University of California Data Discovery Platform (UCDDP)) with exclusion of UCSF, HLD-diagnosed patients (exposed group, n = 364,289) had a faster progression to AD event compared to matched unexposed patients (n = 364,289, data for matched demographics not shown) (FIG. 7A, FIG. 16A, log-rank test p-value<0.005). This was further confirmed with a Cox proportional hazards analysis (hazard ratio (HR) 1.52 (95% Confidence Interval (Cl) 1.461.57), visit / demographic adjusted HR (aHR) 1.26 (1.21 -1.31), p-value <0.005, Table 2). No Strata Strata: recruitment age Model Hazard Ratio 95% Cl p-value Hazard Ratio 95% Cl p-value Unadjusted 1.53 n.47< 1.5$ 2.18E-124 1.49     ^.44,1.54= S.79E-111 demographics adjusted visit adjusted 1 32 (“ .26, 1.371 9.22E-42 1.28      11,23.1,341 1.1 IE-32 visiVdcmographigs adjusted Table 2. UCDDP: Hyperlipidemia Exposure, AD Diagnosis Outcome. Hyperlipidemia exposure cox proportional hazard models for AD as the outcome, shown are the hazard ratios and 95% confidence intervals obtained from the exposure coefficient for unadjusted, demographic adjusted (gender, age, race, ethnicity), visit adjusted (first visit age, log(number of visits)), and demographic / visit adjusted. Right group shows computed hazard ratios with stratification by recruitment or starting age (age strata: <55, 55-60, 6065, 65-70, 70-75, 75-80, >80). In order to investigate potential relationships between HLD and AD, the HLD-specific knowledge network demonstrated shared gene associations with LSS, APOE, INS, SMAD3, ALB, and GFPT1 (FIG. 7B). Locus intersections between high LDL cholesterol and AD across two independent GWAS studies across 408,942 AD patients from Schwartzentruber et al.30 and 94,595 LDL Cholesterol patients from Wilier et al.31 respectively identified multiple shared variants, including ch19:44,892,362(hg38):A>G (rs2075650) and ch19:44,905,579(hg38):T>G (rs405509). PheWAS for rs2075650 on the UK Biobank verified significant associations with cholesterol levels, HLD, AD, and family history of AD (FIG. 7C). Colocalization H4 probability, a measure that determines the probability two traits are associated at a locus based on prior genetic studies, supports a causal link with locus variants for APOE protein QTL and both HLD traits and AD traits (FIG. 7D). Female-specific predictor of osteoporosis validates in an external EHR with potential explanations civen in SPOKE and genetic colocalization analysis From the Osteoporosis was identified as an important feature in the matched models as a female-specific clinical predictor of AD. In the UCDDP, osteoporosis-exposed patients (n=68,940) showed a quicker progression to AD compared to matched unexposed patients (n=68,940, data for matched demographics not shown) (FIG. 8A, FIG. 16B, log-rank test p-value<0.005). When stratified by sex, this progression is significant when comparing between female osteoporosis (n=57,486) vs female controls (n=58,636). Cox hazard analysis further supported osteoporosis as a general risk feature for AD (HR 1.81 (95% Cl 1.70-1.92), aHR 1.59 (1.45-1.70), p<.005, Table 3). No Strata Strata: recruitment age Model Hazard Ratio 95% Cl p-value Hazard Ratio 95% Cl p-value Unadjusted 1.81      [1.70-,1,92]     5.20E-82 1,71     [1,61,1.82]    7.10E-67 demographics adjusted visit adjusted 1.8«      (1.58, ^.83]     4.34B47 1.59     [1,48.1.72]    4,576-34 visit / demagraphics adjusted Table 3. LICDDP: Osteoporosis Exposure, AD Diagnosis Outcome. Osteoporosis exposure cox proportional hazard models for AD as the outcome, shown are the hazard ratios and 95% confidence intervals obtained from the exposure coefficient for unadjusted, demographic adjusted, visit adjusted, and demographic / visit adjusted. Right group shows computed hazard ratios with stratification by recruitment or starting age (age strata: <60, 60-65, 65-70, 70-75, 75-80, >80). Osteoporosis-specific SPOKE network demonstrated shared gene associations with IL6, SMAD3, TNF, HSPG2, GATA1, GFPT1, HFE, INS, and ALB (FIG. 8B). Based on previous GWAS studies across 472,868 AD patients from Schwartzentruber et al.30 and 426,824 heel bone mineral density (HBMD) patients from Morris et al.32, a shared risk locus was found in Chromosome 11 between HBMD and AD among the MS4A gene family, with the closest gene as MS4A6A. A comparison of prior GWAS of up to 71,880 AD patients from Jansen et al.33 and sex-stratified heel bone mineral density (HBMD) GWAS (111,152 Female, 166,988 Male) of UK Biobank patients from Neale Labs supports a female-specific association at the shared locus (FIG. 8C). Colocalization analysis supports a link between MS4A6A and AD (H4 = 0.987), female-specific HBMD with AD, and phenotypes with MS4A6A expression (FIG. 8D, AD vs Female HBMD H4 = 0.998, MS4A6A vs Female HBMD H4 = 0.997). This statistical significance is not replicated for male specific HBMD GWAS (FIG. 8D, AD vs Male HBMD H4 = 0.00263, MS4A6A vs Male HBMD H4 = 0.00266). MS4A6A weighted associations with other phenotypes from Open Targets Genetics found locus associations with many inflammatory phenotypes including c-reactive protein, lymphocyte percentage, and neutrophil count (FIG. 8E). Discussion While there is great potential in ML on clinical data, balancing clinical utility and biological interpretability can be challenging. To address this, thousands of EHR concepts were used to develop prediction models for expert-identified AD diagnosis, and selected an index time suggesting AD onset. Cohort selection and data preprocessing is a crucial first step to identify available clinical measures and optimal ground truth AD onset that is as close to biological AD and avoid overly optimistic model performance due to nonspecific groundtruth or improper data preprocessing34. This prediction model shows predictive power up to -7 years before the defined index time of AD onset with AU ROC of 0.72 (and up to AUROC 0.86 with additional demographic and care utilization features), which is comparable with other models in literature that utilize clinical data to predict less specific dementia or AD diagnosis1135. An application of the model trained on all patients includes determining early disease risk in primary care settings before time-consuming and costly detailed neuropsychological, biomarker, or neuroimaging assessments (after which imaging or biomarker classification models can be utilized13). The model may also identify at-risk patients for follow-up or inclusion in early intervention or clinical trials, with the -1 day model as suggesting possible AD onset to be considered at that visit to prevent underdiagnosis of AD. Furthermore, interpretable models, such as random forest models, can identify common decision point features and allow clinicians to understand what clinical features were used in determining prediction probability and assess the model output with greater trust compared to “black box” models. In order to identify early clinical predictors that may be biologically relevant for AD diagnosis, models were trained on patients matched by pre-identified confounding variables such as demographics and visit-related features so that these features have less influence in AD prediction. Machine learning models still retain the ability to predict AD diagnosis with mean AUROC over .70 after the -3 year time model for random forests. Inclusion of demographic and visit-related features minimally improved model performance, which is expected since matching increased the specificity of the task to predict AD onset controlled on demographics and visit-related features. In terms of clinical utility, the models trained on matched patients provide predictive power for a given clinical scenario between two patients with similar pre-test probability of AD risk (e.g. same age and disease burden), with application of this model as a tool for determining post-test probability of future AD risk. Furthermore, by balancing on pre-identified confounders such as demographics and visits, top features may be interpreted with more biological relevance for AD risk. For example, while essential hypertension was identified as an important feature in the models trained on the full cohort, this diagnosis became less important in the models trained on matched cohorts, suggesting hypertension may be nonspecific for AD and may instead be more directly related to aging or disease burden. The time models trained on matched cohorts identify or strengthen known or suggested hypotheses for early clinical predictors of AD, such as hyperlipidemia as a feature for all time point models. The relative importance of features is also identified years in advance, such as allergic rhinitis and atrial fibrillation as early predictors, osteoporosis and major depressive disorder as non-neurological predictors, and cognitive impairment and vitamin D deficiency as late predictors of AD. Some of these prior predictors, such as depression and vitamin D deficiency, have been previously implicated in AD risk36-38. These findings potentially support hypotheses suggesting AD can be associated with general aging or frailty, which might present in non-neurologic body systems either prior to or concurrent with AD39^3. Furthermore, interpretation of these models allows for the identification of high order groups of predictors that may contribute to disease heterogeneity or together, contribute to AD risk. Nevertheless, while these models can identify hypotheses of predictive features, EHR data can still capture clinical biases or misdiagnoses, and further studies can investigate the influence of behavioral bias vs biological relevance. Models were further trained on sex-stratified subgroups (female vs male), with and without matching on demographics and visit-related covariates, in order to identify sexspecific drivers of clinical predictors. Given evidence that sex may influence different pathways to AD diagnosis22’44 45, it is important to consider how patient heterogeneity may impact the training, utility, and interpretation of a prediction model. From the matched cohort models, clinical features were identified in each subgroup that were consistent with the general models, such as hyperlipidemia as important in every model and memory loss as important in late models. Furthermore, sex-specific features were identified, such as osteoporosis, major depressive disorder, allergic rhinitis, and abnormal stool contents as predictors enriched among women, and chest pain, hypovolemia, prostate hyperplasia, and sensorineural hearing loss as predictive among men. Further work can seek to disentangle the biological meaning of these sex-specific predictive features: whether they reflect sex-specific non-neurological manifestation of prodromal states, contributing risk factors, or even sex biases in clinician evaluation and treatment (e.g., bone density evaluation may arise more often after a fall). These models also demonstrate that for a heterogeneous disorder like AD, subgroup composition, like sex ratio of a cohort, can influence the performance and the features that are identified as important. Differences in subgroup size and prevalence of AD contribute to greater predictive performance among female strata models, and differences observed in AU PRC are impacted by AD prevalence which can influence interpretation of the positive predictive value of models within each sex strata. In terms of identified features, the higher preponderance of females lead to sex-specific predictive factor, osteoporosis, being identified as a general predictive variable in the general group. This further indicates that both generalizable models and subgroup-specific models can provide valuable insight, both general and personalized, for a complex disease. Furthermore, in the context of ML fairness, the performance and identified features of general models may be influenced by the demographic make-up of the training population, just like how greater number and AD prevalence among females influence greater female-strata performance and identification of osteoporosis in the general models. A heterogeneous knowledge network (SPOKE) was utilized to identify shared biological hypotheses underlying model-identified top clinical predictors and AD. By combining shortest paths in SPOKE between top predictors and AD, nodes (e.g., genes) that are consistently relevant for the high order combination of human data derived top clinical predictors and AD can be prioritized to give novel insight via prioritization and combination of relationships. First of all, known genetic associations with dementia were identified based upon top diagnoses, such as through identification of known autosomal dominant early AD genes such as APP and PSEN 1 / 246. Other genes identified with known associations with AD include APOE, HFE, and HSPG2 variants that impact AD risk47 51. An example of novel insight gained through SPOKE integration includes ACTB relating to AD5253, sensorineural hearing loss54, arthropathy, and arthritis55. The prediction model allows for the prioritization of ACTB for patients with the common comorbidities of sensorineural hearing loss and arthropathy / arthritis with risk of AD (where the connection through linking sensorineural hearing loss, arthropathy, arthritis, and AD all together through ACTB has not been previously implicated in literature). The SPOKE network can also be leveraged to propose biological explanations based on common nodes and shared associations between clinical predictors identified from human data and AD. For example, ALB is identified through SPOKE as a shared genetic association between congestive heart failure, malnutrition, hyperlipidemia, and AD. While prior relationships have been identified between ALB and many individual diseases, each of those diseases also have many implicated genetic relationships. Leveraging human data through the predictive models allows for the prioritization of abundant gene connection with multiple disease predictors. Given ALB roles in pathways such as heme biosynthesis (Reactome R-HSA-189445), HDL remodeling (Reactome R-HSA-8964058), and insulin-growth like factor regulation (Reactome R-HSA-8964058), prioritization of mechanistic hypotheses linking ALB related pathways with the pathophysiology of EHR-derived predictors (congestive heart failure, malnutrition, hyperlipidemia) can be explored in future studies. Another example insight includes INS as a shared association between osteoporosis56, hypertension57, hyperlipidemia58, and ^q59,6o prjor stuc|ies have identified potential mechanisms underlying the relationship between energy utilization, lipid levels, nutrition, and neurodegeneration (e.g., Reactome R-HSA-1266738, R-HSA-16368)61-63, and this analysis allows for prioritization of mechanistic hypotheses to be further explored. While these associations are included in the SPOKE network due to evidence in literature, the association of these genes with specific early clinical predictors is less established, and thus this analysis allowed for identification of a novel constellation of phenotypes and underlying genetic relationships observable in a clinical setting that, together, can lead a clinician to suspect future AD risk, prioritize molecular pathways for testing or personalized treatment, and guide biological hypotheses generation in AD pathogenesis for future studies. To validate a few top clinical predictors, a hypothesis-driven approach was utilized to support the relationship between two identified features (hyperlipidemia and osteoporosis) and progress to AD diagnosis in an external database across the University of California EHR system. For both phenotypes, the UC-wide EHR database supports a potential increased AD diagnosis risk due to evidence of decreased time to AD and increased hazard of AD diagnosis in patients exposed to the predictor of interest. The association between hyperlipidemia and AD has been identified in prior clinical studies and systematic reviews64-67. In particular, APOE is a well-established associated genetic locus68, and APOE polymorphism is known to modify AD risk, particularly in subjects carrying the £4 allele69. Many studies have also shown APOE association with elevated lipid levels and cardiovascular risk factors7071. The validation of these well-known associations not only show that the ML models on clinical data can pick up hyperlipidemia as a risk factor, but also by utilizing the SPOKE network known relationships in literature can be integrated to potentially explain the association between hyperlipidemia and AD and identify the APOE locus as a potential shared causal mechanism as demonstrated in the colocalization results. Beyond the ability to identify known relationships, the SPOKE network also proposes biological explanations of higher-order shared associations between clinical predictors, such as ALB as a shared genetic association between congestive heart failure, malnutrition, hyperlipidemia, and AD, or INS as a shared association between osteoporosis, hypertension, hyperlipidemia, and AD. Prior studies have identified potential mechanisms underlying the relationship between energy utilization, lipid levels, nutrition, and neurodegeneration59’60’72, although specific hypotheses of mechanistic relationships are an area for exploration in future studies. The association between osteoporosis and AD is also validated to a lesser extent in clinical studies and meta-analysis73 74, with unclear but possible sex-modification of this effect. This study identifies osteoporosis as a predictor for AD among females prior to AD, but shows less of a relative predictive effect for males compared to other clinical features. Nevertheless, it is still possible that shared relationships between osteoporosis and AD exist in males. A bone mineral density GWAS analysis of female patients shows p-value association with AD GWAS around the MS4A family locus, and this is further supported by MS4A6A eQTL colocalization with both Alzheimer and female HBMD. These findings of osteoporosis as a potential sex-specific predictor of AD, with shared relationships through MS4A6A, is a potential new and unexpected results identified from single hypothesis-driven follow-up from the prediction models. Prior studies have established the MS4A gene cluster as a risk for AD, with one study identifying the cluster based on mendelian randomization75, and another that identifies a stronger female-specific effect size for MS4A6A76. Some studies investigating the role of the MS4A family suggest mechanisms that involve immune function, particularly among microglia77. While this gene may not have been identified in SPOKE, SPOKE did capture direct pathways through known markers of inflammation such as IL6 and TNF, and MS4A6A is also seen as highly associated with measurements of immune cells in the blood. Further studies will be needed to validate the exact associative mechanism between osteoporosis and AD, although some prior hypotheses suggest the potential impact of genetic variants on osteoclast function, amyloid clearance, or oxidative stress response7879. While knowledge networks were utilized to leverage knowledge to explain relationships between groups of predictors, hypothesis-driven analysis was performed on independent EHRs and genetics to further explore and validate a few chosen predictors (hyperlipidemia, osteoporosis) with AD. Hypothesis-driven approaches can be applied to any other selected predictor or phenotype identified by the models to understand their relationships with AD onset that may not yet be represented by the knowledge graphs. This study has several limitations. First, EHR data complexity and quality can affect prediction models, and it is challenging to distinguish the influence of clinician / patient behavior, sociological factors, or underlying biology on identification of features. Matching can improve interpretability by removing influence of non-biological covariates, but followup validation of hypotheses across omics data types is needed. Due to changing patient demographics and societal factors, prediction models should be continuously trained, updated, and evaluated if implemented in the clinical setting to ensure effective utilization and account for biases that may have been learned from the data. Model utilization should investigate the impact of cohort selection biases and matching methods on model generalizability, and model retraining and calibration should be a continual aspect of model application to account for possible data drifts and changing clinical practice approaches that would arise in the future. Second, clinical EHR data is sometimes sparse and provides a superficial interval snapshot of a patient’s health, so the absence of a record may not necessarily reflect the absence of a condition and prior health information may not be available in the EHR. Therefore, the EHR provides a representation of an interval of a patient’s health history and is more likely to pick up diagnosis of chronic or common conditions, as well as common drugs or measurements. Future work can investigate the impact of variations in data representation that can account for data sparsity, continuous lab result outcomes, and best temporal assignment of diagnosis onset beyond binary representation or considering drug prescriptions for assignment of diagnoses. Third, survival models have extensive right censorship and do not take into account competing risks. Fourth, since AD is heterogeneous and differential diagnosis is nuanced and subjective even in expert hands, predictive performance can be limited by label quality and the signal from clinical features can be noisy, limiting performance and generalizability. Future work investigating heterogeneity may identify subgroup-specific features where subgroups can be divided based on biotype, dementia syndromes, racialization, and so on. Future applications with hierarchical models, transfer learning, or fine-tuning on a subpopulation can increase personalization of models. Fifth, the sex-stratified analysis was restricted to patients that identified as female or male. Future studies could explore AD patterns among intersex individuals. Lastly, predictive features identified are relevant prior to AD onset, and future work is needed to identify diagnosticrelevant AD comorbidities, or conditions that can occur after AD progression. Since predictive features are identified as hypotheses, the direct mechanism and causal pathway relating a phenotype to AD is not known. Future work can investigate causality with mendelian randomization or mechanistic studies. In this study, it was demonstrated how formulation of prediction models can influence utility for predictive application or biological interpretation. It is shown how models can be utilized to identify early predictors, and utilize SPOKE to explain relationships via shared biological associations. Lastly, it is shown that the models can pick up known associations with HLD through APOE, and identify a lesser known association with osteoporosis through MS4A6A that may be female-specific. This study contributes to the field of EHR integrative research that can inform AD care and research. Methods Patient Identification Alzheimer’s Disease (AD) patients were identified based on UCSF Memory and Aging Center database containing over 9000 patients mapped to the UCSF OMOP-format EHR. These patients have undergone dementia evaluation at the Memory and Aging Center and thus had expert-level clinical diagnoses. In clinical settings, since AD is often a syndromic diagnosis indicating general dementia for memory or cognitive concerns80-82, it was desired to identify a highly accurate cohort diagnosed by neurodegeneration specialists to obtain AD diagnosis that is closer to the biological ground truth83. The remaining control patients were obtained from the rest of the UCSF EHR, with over 1 year of records and no existing records of dementia diagnosis among the G

[123] * ICD-10 categories (data not shown). These controls include patients seen at the UCSF Memory and Aging Center with EHR data, but without a dementia diagnosis given. In order to best build models for prediction of AD onset, an index time was determined to identify input model features prior to first clinical indication of dementia. This was defined among the AD cohort as the first time of any AD diagnosis, dementia diagnosis, or prescription of cognitive drug (ATC codes N06D, data not shown) to be the first time point of possible biological AD manifestation. This approach was utilized since AD patients may be prescribed an anticholinesterase inhibitor or given an alternative dementia diagnosis before a formal confirmation of an AD diagnosis. For controls, the index time was defined as 1 year before the last recorded her visit date, with no dementia diagnosis given within that year. In order to maintain a consistent patient population for training and evaluation of machine learning models, the final AD and control cohort was identified by filtering to patients who are at least 55 years of age at the index time and have existing clinical visits and concepts 7 years prior to the index time. These patients were then split into 70% for model training and tuning, while the remaining 30% was held-out for model evaluation (FIG. 9). Data Extraction and Preparation Demographics (birth year, gender, race & ethnicity), clinical concepts (conditions, drug exposures, abnormal measures), and visit-related features (age at prediction, first visit age, years in UCSF EHR) were extracted before the index time for the AD and control cohort from the UCSF Observational Medical Outcomes Partnership (OMOP) EHR database. Race & ethnicity is a single variable derived from an algorithm developed by the UCSF Data Equity Taskforce to codify aggregated sociopolitical categorizations based on EHR self-reported identifiers84. To train models in advance of the index time, clinical information was extracted for each patient including all clinical data up to a time point X before the index time, where X includes -7 years, -5 years, -3 years, -1 years, and -1 day. These time points represent the knowledge of a patient’s clinical history leading up to time X before time. All existing clinical features (conditions, drug exposures, abnormal measurements) were one-hot encoded. Abnormal measures were extracted from the OMOP measurement table based on the numeric value falling either above range_high or below rangejow. and abnormal measures were binary encoded based on abnormal flagging, following the approach from Nelson et al.27. If a clinical feature did not exist or if the clinical measure was within normal range, the encoding is represented as a 0 and therefore assumed to be normal. Since the UCSF database only captures an interval of a patient’s interaction with the healthcare system, prior non-chronic conditions may not be captured within the EHR. Demographic and visit-related features (prediction age, first visit age, years in UCSF EHR, log(number prior visits), log(number prior concepts), log(days since first clinical event)) were scaled between 0-1 on the training data, where log indicates natural logarithm and feature scaling allows for multiple ML model approaches. Age at prediction is defined at the age of patient at which the model is applied (e.g., if a patient index time is at age 70, then the age of prediction for the -5 year model is 65). All features with no variance were removed for each model, with total number of features ranging from 5,211 features (-7 year model on matched cohorts) to 23,760 features (-1 day models on unmatched cohorts). Machine Learning Preparation and Training Binary classification time point models for AD were trained using the patient representation at each time point before the index time. The data was divided into two sets, 70% for model creation and 30% for evaluation. Training and optimal model selection (with hyperparameter tuning) was performed on the 70% split with crossvalidation, and 30% was held-out for evaluation and not seen during model training and selection in any way. Final selected model evaluation was performed on the 30% held-out evaluation set as the common dataset to obtain and compare the performance of all final models (diagram in FIG. 9). Models were trained with clinical features only (clinical model) and with clinical features + demographics and visit-related information (clinical + demo / visits model). Models were also trained on samples matched by demographics and hospital utilization to account for biases and confounding in prediction. In these models, control patients were matched to AD patients at a 1:8 ratio on demographics (birth year, race & ethnicity, sex) and visit-related features (age, first visit age, years in EHR, log(# prior visits), log(# prior concepts), log(days since first clinical event)) utilizing propensity score matching85 (propensity score estimated based upon a logistic regression model, nearest neighbor matching without replacement). While propensity score is often utilized to balance treatment probabilities in cohort studies, it has also been utilized for sample selection86 87, exposure likelihood88, or for outcome-based case-control studies7 89. Random forest models were primarily utilized for both predictive performance and interpretability that takes into account the high collinearity between clinical variables. Random forests were trained using scikit-learn package90, with balanced class weight parameter. Hyper-parameters were tuned (grid search) based on cross-validation performance (5 folds) of AU ROC on the 70% model training set to determine parameters of n_estimators (njeatures, n_features*2, n_features*3), max_depth (3, 5, 7, 9), and max_features (sqrt, Iog2). The number of estimators and max depth were tuned to balance between performance and overfitting, while a subset of features (max_features) was utilized per tree to help account for high correlation between features9192. Models were evaluated on bootstrapped subsamples (50-200 iterations, 1000 samples) of the 30% held-out evaluation set to determine AUROC (area under the receiver operating curve) and AUPRC (area under the precision-recall curve) for model comparability. Balanced accuracy scores were also computed on the 30% held-out evaluation set. An elastic net logistic regression model was also trained on both the full and matched cohorts for comparison. A permutation test was performed on the -1 day matched cohort model to determine the significance of AUROC compared to a null distribution of AUROC scores of models trained from permuted ground truth labels (40 permutations) to determine to the extent clinical features can be predictive of AD. Stratification-. Both models for full patient cohorts and matched cohorts were reperformed in sex strata in the same fashion based upon sex reported the UCSF EHR to augment the OMOP database. Models were trained on two sex subgroups: female and male, due to lack of other subgroups labelled in the EHR. For each strata, AD patients were re-matched to controls within each strata for the matched patient trained models. Models were evaluated similarly based on AUROC / AUPRC on the same bootstrapped held-out evaluation set, stratified by sex. Top Feature Interpretation Random forest models were investigated for feature interpretation due to the combined interpretable nature of the models (compared to neural networks) and the ability to capture nonlinear relationships (compared to logistic regression models)93. Average gini impurity decrease for each feature was utilized to evaluate the importance of each feature in the random forest models (feature importance). The average importance for each feature was taken across each time point models (-7yr, -5yr, -3yr, -1 yr, -1day) to obtain an across-model importance for each model type, and normalized by the maximum importance value across all time point models within each model type (e.g. random forest) and group (e.g. female strata). Feature importances are then ranked within each model to obtain relative importance within each of the time points. Since a patient’s exposure to a medication or a laboratory test is often a result of a diagnosis, it was desired to base interpretability on diagnostic features that have been mapped to phecodes, which is a semi-manual hierarchical aggregation of meaningful EHR phenotypes29. This allows for a lossy categorization of detailed OMOP features (OMOP IDs) to phecodes (OMOP ID —> SNOMED —> ICD10 —> phecode) and phecode category. SNOMED IDs were mapped to ICD10 based upon recommended rule-based mappings from the National Library of Medicine (NLM) September 2022 release (www.nlm.nih.gov / healthit / snomedct / us_edition.html). ICD10 codes were then mapped to phecodes based on the release from Wu et al.94 To obtain the importance within each phecode or phecode category, the average importance for the top 5 detailed OMOP features per phecode or phecode category was computed, and ranked between phecodes or categories. For phecodes across all models and sex-stratified models, the ranking of importance of phecodes across each time model was hierarchically clustered with Ward linkage. To compare top phecodes between sex-stratified models to identify sex-specific features, top random forest features over an average importance threshold of 1e-6 were identified per time model trained on matched participants. Upset plots were then generated for each time point based upon this overlap. Female-driven features are defined as features that exist in both the full model and female models, or only female models, and male-driven features defined analogously. UC-wide validation analysis with hypothesis-driven retrospective cohort analysis Two top clinical features were selected from the matched all patient model (hyperlipidemia) and matched sex-specific models (osteoporosis) and further followed up on an external EHR database to validate the feature as predictive and conferring risk for AD diagnosis. With these features defined as exposures, hypothesis-driven analysis was performed with a retrospective cohort study design on the University of California hospital EHR database (University of California Data Discovery Platform (UCDDP)) with exclusion of any patients seen at UCSF, so with included institutions consisting of UC Davis, UC Los Angeles, UC Riverside, UC San Diego, and UC Irvine. Exposed patients were identified with the exposure (hyperlipidemia or osteoporosis), which were identified by string-matching and mapping to all descendants or related concepts based on the OMOP relationship tables (final SNOMED codes not shown). Controls were identified among the remaining patients. Recruitment age was defined as the age of exposure diagnosis (for exposed cohort) or the first visit age in the visit_occurrence table (for unexposed or control cohort), which was then matched to represent the start of the cohort study timeline. All patients are then filtered to have at least 2 years of records in the EHR, and last visit age was utilized for right censorship. The outcome of interest was AD diagnosis, which was identified based on SNOMED codes 26929004, 416780008, 416975007 (data not shown). Exposed and control (unexposed) groups were then matched based on demographics (gender, race & ethnicity), birth year, and recruitment age (propensity score estimated based upon a logistic regression model, nearest neighbor matching without replacement). The genderjd column was utilized to identify sex, as the standard documentation intend for this column to represent biological sex. Note that only two options exist (female concept_id=8532 and male concept_id=8507), and that accurate sex and gender information may be limited depending on the institution or EHR collection of sex information. Analysis of time to AD diagnosis includes utilization Kaplan Meier survival curves fitted with 95% confidence interval and two-sided log-rank test to compare survival curves between groups. Sex-stratified curves were also fitted. Cox proportional hazard models were utilized to obtain unadjusted hazard ratios (HR) and adjusted hazard ratios by demographics and / or visit information (aHR), with and without stratification by recruitment age or birth year, and with 95% confidence intervals. Heterogeneous Network analysis Heterogeneous knowledge networks, such as SPOKE, integrate known relationships across biological and phenotypic data realms in databases and literature. Such a network could provide hypotheses to explain relationships between groups of phenotypes that may not be immediately known2126. Interpretation on the matched models was carried out, with the top 25 model features taken per time point and mapped to SPOKE nodes based on Nelson et al.27 Note that mappings may not be 1 to 1. All shortest paths were then computed from each input node to the Alzheimer’s Disease node (DOID: 10652), and shortest paths were filtered to exclude certain node types (Anatomy, SideEffect, AnatomyCellType,Nutrient) and edges (CONTRAINDICATES_CcD, CAUSES_CcSE, LOCALIZES_DIA, ISA_AiA, PARTOF_ApA, RESEMBLES_DrD). Edges were also filtered based on the following criteria: TREATS_CtD at least phase 3 clinical trial, UPREGULATES_KGuG / DOWNREGULATES_KGdG p-value at most 1E-4, PRESENTS DpS enrichment at least 5 and fisher p-value at most 1 E-4. If multiple detailed OMOP features map to the same node, the importance of the node was obtained by the average of OMOP feature importances. Networks for all time models were combined into a single network (union of nodes and edges), and total node importance was determined by the maximum across time. Network metrics were then computed with Cytoscape ‘Network Analyzer’ function95. The combined time model networks were then sorted by eccentricity metric on the x-axis (representing maximum distance to all other nodes, with lower number representing higher importance) and number of individual time model network occurrences in the y-axis (showing node importance persistence across time). With this layout, highly traversed nodes in the shortest paths between multiple EHR informed top model features and AD can be identified and prioritized for hypothesis generation and further investigation. Note that due to heterogeneous nature of edges and lack of edge weighting, distance in the figure is not meaningful. To focus on two selected features for the full matched model (hyperlipidemia (HLD)) and the female-specific matched model (osteoporosis), the combined network was filtered based on first and second degree neighbors of the starting feature of interest. This allows for visualization of associated genes and AD, as well as relationships with other top model features found from the clinical models. Validation with Genetic Datasets The association between clinical predictors and AD was explored by identifying shared genetic loci between top model phenotypes and AD, based on colocalization probability and weighted evidence association scores computed from Open Targets Genetics9697 (genetics.opentargets.org). Colocalization analysis is a method that determines if two independent signals at a locus share a causal variant, which helps increase the evidence that the two traits (e.g. hyperlipidemia and AD, or protein expression and AD) also share a causal mechanism. It is a Bayesian method which, for two traits, integrates evidence over all variants at a single locus to evaluate the following hypothesis that two associated traits share a causal variant. This is the H4 probability. The shared loci between the selected phenotypes (HLD or osteoporosis) and AD was investigated by identifying the genetic intersection between AD and related phenotypes in Open Targets Genetics. For HLD and AD, the Open Targets Genetics platform was utilized to identify overlapping variants and shared locus between LDL Cholesterol and Family History of AD or AD. PheWAS between a shared SNP and UK Biobank phenotypes were plotted and extracted from the Open Targets Genetics platform. Coloc analysis tables between the gene, molecular QTLs, and phenotypes were extracted, with protein QTLs for APOE specifically identified based on blood plasma data from Sun et al.98 and Suhre et al.99 Similarly for osteoporosis and AD, the Open Genetics platform was utilized to identify shared locus between heel bone mineral density (proxy for osteoporosis) and Family History of AD or AD. To further investigate the locus, GWAS summary statistics were extracted from Jansen et al.48 for AD and sex-stratified GWAS summary statistics for heel bone mineral density (HBMD) from Neale’s Lab GWAS round 2, Phenotype Code:3148, based on data from the UK Biobank100. Colocalization analysis was then conducted using the coloc method described in Giambartolomei et al.101, from R package coloc 5.1.0. Summary statistics for MS4A6A cis eQTLs in blood were extracted from eQTLGen102, and colocalization analysis was performed between AD, sex-stratified HBMD, and MS4A6A eQTLs on the Locus Region 60050000-60200000 of Chromosome 11. To investigate further associations with the locus, MS4A6A associations with all other phenotypes was extracted from Open Targets Genetics platform with inclusion of a weighted literature evidence association scores. References 1.    2022 Alzheimer’s disease facts and figures. Alzheimers Dement. 18, 700-789 (2022). 2. Rasmussen, J. & Langerman, H. Alzheimer’s Disease - Why We Need Early Diagnosis. Degener. Neurol. Neuromuscul. Dis. Volume 9, 123-130 (2019). 3. Kivipelto, M. Midlife vascular risk factors and Alzheimer’s disease in later life: longitudinal, population based study. BMJ322, 1447-1451 (2001). 4. Niculescu, A. B. et al. Blood biomarkers for memory: toward early detection of risk for Alzheimer disease, pharmacogenomics, and repurposed drugs. Mol. Psychiatry 25, 1651-1672 (2020). 5. Alena V. Savonenko, Philip C. Wong, & Tong Li. Alzheimer diseases. (2023) doi :10.1016 / b978-0-323-85654-6.00022-8. 6. Neugroschl, J. & Wang, S. Alzheimer’s Disease: Diagnosis and Treatment Across the Spectrum of Disease Severity. Mt. Sinai J. Med. N. Y. 78, 596-612 (2011). 7.    Tang, A. S. et al. Deep phenotyping of Alzheimer’s disease leveraging electronic medical records identifies sex-specific clinical associations. Nat. Commun. 13, 675 (2022). 8. Taubes, A. et al. Experimental and real-world evidence supporting the computational repurposing of bumetanide for APOE4-related Alzheimer’s disease. Nat. Aging 1,932-947 (2021). 9. Ben Miled, Z. et al. Predicting dementia with routine care EMR data. Artif. Intell. Med. 102, 101771 (2020). 10. Tang, A., Woldemariam, S., Roger, J. & Sirota, M. Translational Bioinformatics to Enable Precision Medicine for All: Elevating Equity across Molecular, Clinical, and Digital Realms. Yearb. Med. Inform. 31, 106-115 (2022). 11. Xu, J. et al. Data-driven discovery of probable Alzheimer’s disease and related dementia subphenotypes using electronic health records. Learn. Health Syst. 4, e10246 (2020). 12. Park, J. H. et al. Machine learning prediction of incidence of Alzheimer’s disease using large-scale administrative health data. Npj Digit. Med. 3, 46 (2020). 13. Qiu, S. et al. Multimodal deep learning for Alzheimer’s disease dementia assessment. Nat. Commun. 13, 3404 (2022). 14. Diogo, V. S., Ferreira, H. A., Prata, D., & for the Alzheimer’s Disease Neuroimaging Initiative. Early diagnosis of Alzheimer’s disease using machine learning: a multidiagnostic, generalizable approach. Alzheimers Res. Ther. 14, 107 (2022). 15. Ding, Y. et al. A Deep Learning Model to Predict a Diagnosis of Alzheimer Disease by Using 18 F-FDG PET of the Brain. Radiology 290, 456-464 (2019). 16. Popuri, K., Ma, D., Wang, L. & Beg, M. F. Using machine learning to quantify structural MRI neurodegeneration patterns of Alzheimer’s disease into dementia score: Independent validation on 8,834 images from ADNI, AIBL, OASIS, and MIRIAD databases. Hum. Brain Mapp. 41,4127-4147 (2020). 17. Chang, C.-H., Lin, C.-H. & Lane, H.-Y. Machine Learning and Novel Biomarkers for the Diagnosis of Alzheimer’s Disease. Int. J. Mol. Sci. 22, 2761 (2021). 18. Stamate, D. et al. A metabolite-based machine learning approach to diagnose Alzheimer-type dementia in blood: Results from the European Medical Information Framework for Alzheimer disease biomarker discovery cohort. Alzheimers Dement. Transl. Res. Clin. Interv. 5, 933-938 (2019). 19. Dubai, D. B. Chapter 16 - Sex difference in Alzheimer’s disease: An updated, balanced and emerging perspective on differing vulnerabilities, in Handbook of Clinical Neurology (eds. Lanzenberger, R., Kranz, G. S. & Savic, I.) vol. 175 261-273 (Elsevier, 2020). 20. Hampel, H. et al. Precision medicine and drug development in Alzheimer’s disease: the importance of sexual dimorphism and patient stratification. Front. Neuroendocrinol. 50, 31-51 (2018). 21. Nelson, C. A., Bove, R., Butte, A. J. & Baranzini, S. E. Embedding electronic health records onto a knowledge network recognizes prodromal features of multiple sclerosis and predicts diagnosis. J. Am. Med. Inform. Assoc. 29, 424-434 (2022). 22. Belonwu, S. A. et al. Sex-Stratified Single-Cell RNA-Seq Analysis Identifies SexSpecific and Cell Type-Specific Transcriptional Responses in Alzheimer’s Disease Across Two Brain Regions. Mol. Neurobiol. (2021) doi:10.1007 / s12035-021-02591-8. 23. Carlos A. Saura, Angel Deprada, Maria Dolores Capilla-Lopez, & Arnaldo Parra-Damas. Revealing cell vulnerability in Alzheimer’s disease by single-cell transcriptomics. Semin. Cell Dev. Biol. (2022) doi:10.1016 / j.semcdb.2022.05.007. 24. Leonenko, G. et al. Polygenic risk and hazard scores for Alzheimer’s disease prediction. Ann. Clin. Transl. Neurol. 6, 456-465 (2019). 25. Alzheimer’s Disease Neuroimaging Initiative et al. Multimodal Phenotyping of Alzheimer’s Disease with Longitudinal Magnetic Resonance Imaging and Cognitive Function Data. Sci. Rep. 10, 5527 (2020). 26. Himmelstein, D. S. etal. Systematic integration of biomedical knowledge prioritizes drugs for repurposing. eLife 6, e26726 (2017). 27. Nelson, C. A., Butte, A. J. & Baranzini, S. E. Integrating biomedical research and electronic health records to create knowledge-based biologically meaningful machine-readable embeddings. Nat. Commun. 10, 3045 (2019). 28. Morris, J. H. et al. The scalable precision medicine open knowledge engine (SPOKE): a massive knowledge graph of biomedical information. Bioinformatics 39, btad080 (2023). 29. Bastarache, L. Using Phecodes for Research with the Electronic Health Record: From PheWAS to PheRS. Annu. Rev. Biomed. Data Sei. 4, 1-19 (2021). 30. Schwartzentruber, J. et al. Genome-wide meta-analysis, fine-mapping and integrative prioritization implicate new Alzheimer’s disease risk genes. Nat. Genet. 53, 392-402 (2021). 31. Global Lipids Genetics Consortium. Discovery and refinement of loci associated with lipid levels. Nat. Genet. 45, 1274-1283 (2013). 32. Morris, J. A. et al. An atlas of genetic influences on osteoporosis in humans and mice. Nat. Genet. 51,258-266 (2019). 33. Jansen, W. J. et al. Association of Cerebral Amyloid-beta Aggregation With Cognitive Functioning in Persons Without Dementia. JAMA Psychiatry 75, 84 (2018). 34. Yagis, E. et al. Effect of data leakage in brain MRI classification using 2D convolutional neural networks. Sci. Rep. 11, 22544 (2021). 35. You, J. et al. Development of a novel dementia risk prediction model in the general population: A large, longitudinal, population-based machine-learning study. eClinicalMedicine 53, 101665 (2022). 36. Littlejohns, T. J. et al. Vitamin D and the risk of dementia and Alzheimer disease. Neurology 83, 920-928 (2014). 37. Elbejjani, M. et al. Depression, depressive symptoms, and rate of hippocampal atrophy in a longitudinal cohort of older men and women. Psychol. Med. 45, 1931-1944 (2015). 38. Goveas, J. S., Espeland, M. A., Woods, N. F., Wassertheil-Smoller, S. & Kotchen, J. M. Depressive Symptoms and Incidence of Mild Cognitive Impairment and Probable Dementia in Elderly Women: The Women’s Health Initiative Memory Study: DEPRESSION AND INCIDENT MCI AND DEMENTIA. J. Am. Geriatr. Soc. 59, 57-66 (2011). 39. Swerdlow, R. H. Is aging part of Alzheimer’s disease, or is Alzheimer’s disease part of aging? Neurobiol. Aging 28, 1465-1480 (2007). 40. Kosyreva, A. M., Sentyabreva, A. V., Tsvetkov, I. S. & Makarova, O. V. Alzheimer’s Disease and Inflammaging. Brain Sci. 12, 1237 (2022). 41. Wallace, L. M. K. et al. Investigation of frailty as a moderator of the relationship between neuropathology and dementia in Alzheimer’s disease: a cross-sectional analysis of data from the Rush Memory and Aging Project. Lancet Neurol. 18, 177-184 (2019). 42. Kojima, G., Taniguchi, Y., Iliffe, S. & Walters, K. Frailty as a Predictor of Alzheimer Disease, Vascular Dementia, and All Dementia Among Community-Dwelling Older People: A Systematic Review and Meta-Analysis. J. Am. Med. Dir. Assoc. 17, 881-888 (2016). 43. Wallace, L., Theou, 0., Rockwood, K. & Andrew, M. K. Relationship between frailty and Alzheimer’s disease biomarkers: A scoping review. Alzheimers Dement. Diagn. Assess. Dis. Monit. 10, 394-401 (2018). 44. Barnes, L. L. et al. Sex Differences in the Clinical Manifestations of Alzheimer Disease Pathology. Arch. Gen. Psychiatry 62, 685 (2005). 45. Davis, E. J. et al. Sex-Specific Association of the X Chromosome With Cognitive Change and Tau Pathology in Aging and Alzheimer Disease. JAMA Neurol. (2021) doi:10.1001 / jamaneurol.2021.2806. 46. Campion, D. et al. Early-onset autosomal dominant Alzheimer disease: prevalence, genetic heterogeneity, and mutation spectrum. Am. J. Hum. Genet. 65, 664670 (1999). 47. Liew, T. M. Subjective cognitive decline, APOE e4 allele, and the risk of neurocognitive disorders: Age- and sex-stratified cohort study. Aust. N. Z. J. Psychiatry (2022) doi: 10.1177 / 00048674221079217. 48. He, Z. et al. Genome-wide analysis of common and rare variants via multiple knockoffs at biobank scale, with an application to Alzheimer disease genetics. Am. J. Hum. Genet. (2021) doi:10.1016 / j.ajhg.2O21.10.009. 49. Nandar, W. & Connor, J. R. HFE Gene Variants Affect Iron in the Brain1-3. J. Nutr. 141, S729-S739 (2011). 50. Wang, Z. et al. Deep post-GWAS analysis identifies potential risk genes and risk variants for Alzheimer’s disease, providing new insights into its disease mechanisms. Sci. Rep. 11,20511 (2021). 51. livonen, S. et al. Heparan sulfate proteoglycan 2 polymorphism in Alzheimer’s disease and correlation with neuropathology. Neurosci. Lett. 352, 146-150 (2003). 52. Talwar, P. et al. Genomic convergence and network analysis approach to identify candidate genes in Alzheimer’s disease. BMC Genomics 15, 199 (2014). 53. Talwar, P. et al. Validating a Genomic Convergence and Network Analysis Approach Using Association Analysis of Identified Candidate Genes in Alzheimer’s Disease. Front. Genet. 12, 722221 (2021). 54. Zhu, M. et al. Mutations in the y-Actin Gene (ACTG1) Are Associated with Dominant Progressive Deafness (DFNA20 / 26). Am. J. Hum. Genet. 73, 1082-1091 (2003). 55. Vasilopoulos, Y., Gkretsi, V., Armaka, M., Aidinis, V. & Kollias, G. Actin cytoskeleton dynamics linked to synovial fibroblast activation as a novel pathogenic principle in TNF-driven arthritis. Ann. Rheum. Dis. 66, iii23-iii28 (2007). 56. Lee, W.-C., Guntur, A. FL, Long, F. & Rosen, C. J. Energy Metabolism of the Osteoblast: Implications for Osteoporosis. Endocr. Rev. 38, 255-266 (2017). 57. Wang, F., Han, L. & Hu, D. Fasting insulin, insulin resistance and risk of hypertension in the general population: A meta-analysis. Clin. Chim. Acta Int. J. Clin. Chern. 464, 57-63 (2017). 58. James, D. E., Stockli, J. & Birnbaum, M. J. The aetiology and molecular landscape of insulin resistance. Nat. Rev. Mol. Cell Biol. 22, 751-771 (2021). 59. Schrijvers, E. M. C. et al. Insulin metabolism and the risk of Alzheimer disease: The Rotterdam Study. Neurology 75, 1982-1987 (2010). 60. Ferreira, L. S. S., Fernandes, C. S., Vieira, M. N. N. & De Felice, F. G. Insulin Resistance in Alzheimer’s Disease. Front. Neurosci. 12, 830 (2018). 61. Rahman, S. O. et al. Association between insulin and Nrf2 signalling pathway in Alzheimer’s disease: A molecular landscape. Life Sci. 328, 121899 (2023). 62. Ataie-Ashtiani, S. & Forbes, B. A Review of the Biosynthesis and Structural Implications of Insulin Gene Mutations Linked to Human Disease. Cells 12, 1008 (2023). 63. Gillespie, M. et al. The reactome pathway knowledgebase 2022. Nucleic Acids Res. 50, D687-D692 (2022). 64. Bowman, G. L., Kaye, J. A. & Quinn, J. F. Dyslipidemia and Blood-Brain Barrier Integrity in Alzheimer’s Disease. Curr. Gerontol. Geriatr. Res. 2012, 1-5 (2012). 65. Reitz, C. Dyslipidemia and the risk of Alzheimer’s disease. Curr. Atheroscler. Rep. 15, 307 (2013). 66. Goldstein, F. C. et al. Effects of hypertension and hypercholesterolemia on cognitive functioning in patients with alzheimer disease. Alzheimer Dis. Assoc. Disord. 22, 336-342 (2008). 67. Saiz-Vazquez, 0., Puente-Martinez, A., Ubillos-Landa, S., Pacheco-Bonrostro, J. & Santabarbara, J. Cholesterol and Alzheimer’s Disease Risk: A Meta-Meta-Analysis. Brain Sci. 10, 386 (2020). 68. Bertram, L. & Tanzi, R. E. Genome-wide association studies in Alzheimer’s disease. Hum. Mol. Genet. 18, R137-R145 (2009). 69. Corder, E. H. et al. Gene Dose of Apolipoprotein E Type 4 Allele and the Risk of Alzheimer’s Disease in Late Onset Families. Science 261,921-923 (1993). 70. Garcia, A. R. etal. APOE4 is associated with elevated blood lipids and lower levels of innate immune biomarkers in a tropical Amerindian subsistence population. eLife 10, e68231 (2021). 71. Mahley, R. W. & Rall, S. C. Apolipoprotein E: Far More Than a Lipid Transport Protein. Annu. Rev. Genomics Hum. Genet. 1,507-537 (2000). 72. Kimura, R. etal. Albumin gene encoding free fatty acid and p-amyloid transporter is genetically associated with Alzheimer disease: Albumin gene and Alzheimer’s disease. Psychiatry Clin. Neurosci. 60, S34-S39 (2006). 73. Lv, X.-L. et al. Association between Osteoporosis, Bone Mineral Density Levels and Alzheimer’s Disease: A Systematic Review and Meta-analysis. Int. J. Gerontol. 12, 76-83 (2018). 74. Amouzougan, A. et al. High prevalence of dementia in women with osteoporosis. Joint Bone Spine 84, 611-614 (2017). 75. Liu, Y., Jin, G., Wang, X., Dong, Y. & Ding, F. Identification of New Genes and Loci Associated With Bone Mineral Density Based on Mendelian Randomization. Front. Genet. 12, 728563 (2021). 76. Fan, C. C. et al. Sex-dependent autosomal effects on clinical progression of Alzheimer’s disease. Brain 143, 2272-2280 (2020). 77. Deming, Y. et al. The MS4A gene cluster is a key modulator of soluble TREM2 and Alzheimer’s disease risk. Sci. Transl. Med. 11, eaau2291 (2019). 78. Chen, Y.-H. & Lo, R. Y. Alzheimer’s disease and osteoporosis. Ci Ji Yi Xue Za Zhi Tzu-Chi Med. J. 29, 138-142 (2017). 79. Li, S., Liu, B., Zhang, L. & Rong, L. Amyloid beta peptide is elevated in osteoporotic bone tissues and enhances osteoclast function. Bone 61, 164-175 (2014). 80. Gale, S. A. etal. Preclinical Alzheimer Disease and the Electronic Health Record: Balancing Confidentiality and Care. Neurology 99, 987-994 (2022). 81. Serrano-Pozo, A. et al. Mild to moderate Alzheimer dementia with insufficient neuropathological changes. Ann. Neurol. 75, 597-601 (2014). 82. Nelson, P. T. et al. Alzheimer’s disease is not ‘brain aging’: neuropathological, genetic, and epidemiological human studies. Acta NeuropathoL (Berl.) 121, 571-587 (2011). 83. Jack, C. R. et al. NIA-AA Research Framework: Toward a biological definition of Alzheimer’s disease. Alzheimers Dement. J. Alzheimers Assoc. 14, 535-562 (2018). 84. Data Equity Taskforce sponsored by the Health Equity Council at UCSF Health. UCSF Health’s equity-related variables user’s guide. (2021). 85. Austin, P. C. An Introduction to Propensity Score Methods for Reducing the Effects of Confounding in Observational Studies. Multivar. Behav. Res. 46, 399-424 (2011). 86. Karlin, L. et al. Use of the Propensity Score Matching Method to Reduce Recruitment Bias in Observational Studies: Application to the Estimation of Survival Benefit of Non-Myeloablative Allogeneic Transplantation In Patients with Multiple Myeloma Relapsing after a First Autologous Transplantation. Blood 112, 1133-1133 (2008). 87. Tipton, E. et al. Sample Selection in Randomized Experiments: A New Method Using Propensity Score Stratified Sampling. J. Res. Educ. Eff. 7, 114-135 (2014). 88. Bingenheimer, J. B., Brennan, R. T. & Earls, F. J. Firearm violence exposure and serious violent behavior. Science 308, 1323-1326 (2005). 89. Xia, Y. et al. Association between dietary patterns and metabolic syndrome in Chinese adults: a propensity score-matched case-control study. Sci. Rep. 6, 34748 (2016). 90. Pedregosa, F. etal. Scikit-learn: Machine Learning in Python. J. Mach. Learn. Res. (2012) doi:10.48550 / ARXIV.1201.0490. 91. scikit-learn developers. Scikit-Learn Documentation: Random Forest Parameters. https: / / scikit-learn.Org / stable / modules / ensemble.html#random-forest-parameters. 92. Breiman, L. Random Forests. Mach. Learn. 45, 5-32 (2001). 93. Azodi, C. B., Tang, J. & Shiu, S.-H. Opening the Black Box: Interpretable Machine Learning for Geneticists. Trends Genet. 36, 442-455 (2020). 94. Wu, P. et al. Mapping ICD-10 and ICD-10-CM Codes to Phecodes: Workflow Development and Initial Evaluation. JMIR Med. Inform. 7, e14325 (2019). 95. Assenov, Y., Ramirez, F., Schelhorn, S.-E., Lengauer, T. & Albrecht, M. Computing topological parameters of biological networks. Bioinformatics 24, 282-284 (2008). 96. Ghoussaini, M. et al. Open Targets Genetics: systematic identification of trait-associated genes using large-scale genetics and functional genomics. Nucleic Acids Res. 49, D1311-D1320 (2021). 97. Mountjoy, E. et al. An open approach to systematically prioritize causal variants and genes at all published human GWAS trait-associated loci. Nat. Genet. 53, 15271533 (2021). 98. Sun, B. B. et al. Genomic atlas of the human plasma proteome. Nature 558, 7379 (2018). 99. Suhre, K. et al. Connecting genetic risk to disease end points through the human blood plasma proteome. Nat. Commun. 8, 14357 (2017). 100. Neale Lab. UK Biobank GWAS Round 2. http: / / www.nealelab.is / uk-biobank / . 101. Giambartolomei, C. et al. Bayesian test for colocalisation between pairs of genetic association studies using summary statistics. PLoS Genet. 10, e1004383 (2014). 102. Vosa, U. et al. Large-scale cis- and trans-eQTL analyses identify thousands of genetic loci and polygenic scores that regulate blood gene expression. Nat. Genet. 53, 1300-1310 (2021). Although the foregoing invention has been described in some detail by way of illustration and example for purposes of clarity of understanding, it is readily apparent to those of ordinary skill in the art in light of the teachings of this invention that certain changes and modifications may be made thereto without departing from the spirit or scope of the appended claims. Accordingly, the preceding merely illustrates the principles of the invention. It will 5 be appreciated that those skilled in the art will be able to devise various arrangements which, although not explicitly described or shown herein, embody the principles of the invention and are included within its spirit and scope. Furthermore, all examples and conditional language recited herein are principally intended to aid the reader in understanding the principles of the invention and the concepts contributed by the 10 inventors to furthering the art and are to be construed as being without limitation to such specifically recited examples and conditions. Moreover, all statements herein reciting principles, aspects, and embodiments of the invention as well as specific examples thereof, are intended to encompass both structural and functional equivalents thereof. Additionally, it is intended that such equivalents include both currently known 15 equivalents and equivalents developed in the future, i.e., any elements developed that perform the same function, regardless of structure. The scope of the present invention, therefore, is not intended to be limited to the exemplary embodiments shown and described herein. Rather, the scope and spirit of the present invention is embodied by the appended claims.

Claims

1. A method of training a model to predict a disease onset risk, the method comprising:obtaining a plurality of electronic health records comprising health information about a plurality of subjects, wherein the plurality of subjects comprise subjects diagnosed with the disease and subjects not diagnosed with the disease;dividing the subjects into a first set for machine learning model training and a second set for machine learning evaluation;training a machine learning model using the electronic health records of the first set of subjects;evaluating the machine learning model using the electronic health records of the second set of subjects; andidentifying from the trained and evaluated machine learning model one or more features of the subjects that predict the disease onset risk.

2. A method of training a model to predict a disease onset risk, the method comprising:obtaining a plurality of electronic health records comprising health information about a plurality of subjects, wherein the plurality of subjects comprise subjects diagnosed with the disease and subjects not diagnosed with the disease, and wherein the electronic health records include sexes of the subjects;dividing the subjects into a first set for machine learning model training and a second set for machine learning evaluation;training a machine learning model using the electronic health records of the first set of subjects;evaluating the machine learning model using the electronic health records of the second set of subjects; andidentifying from the trained and evaluated machine learning model one or more sex-specific features of the subjects that predict the disease onset risk.

3. A method of predicting a disease onset risk in a subject, the method comprising:obtaining a health record comprising health information about a subject, wherein the health record includes a sex of the subject; andapplying the model trained according to the method of claim 2 to the health record of the subject to predict the disease onset risk in the subject.

4. A method of clinically evaluating the presence of a disease in a subject, the method comprising:obtaining a health record comprising health information about a subject, wherein the health record includes a sex of the subject;applying the model trained according to the method of claim 2 to the health record of the subject to predict the risk of developing the disease; andclinically evaluating the presence of the disease, in the event that subject is predicted to be at risk of developing the disease.

5. A method of treating a subject predicted to be at risk of developing a disease, the method comprising:obtaining a health record comprising health information about a subject, wherein the health record includes a sex of the subject;applying the model trained according to the method of claim 2 to the health record of the subject to predict the risk of developing the disease; andadministering a treatment to the subject in an amount effective to treat the disease, in the event that subject is predicted to be at risk of developing the disease.

6. A method of predicting a disease onset risk in a subject, the method comprising:obtaining a health record comprising health information about a subject, wherein the health record includes a sex of the subject; andpredicting the disease onset risk in a subject based on the presence of one or more sex-specific features in the health record that has been identified to predict the disease onset risk based on the method of claim 2.

7. A method of clinically evaluating the presence of a disease in a subject, the method comprising:obtaining a health record comprising health information about a subject, wherein the health record includes a sex of the subject;predicting the disease onset risk in a subject based on the presence of one or more sex-specific features in the health record that has been identified to predict disease onset risk based on the method of claim 2; andclinically evaluating the presence of the disease, in the event that subject is predicted to be at risk of developing the disease.

8. A method of treating a subject predicted to be at risk of developing a disease, the method comprising:obtaining a health record comprising health information about a subject, wherein the health record includes a sex of the subject;predicting the disease onset risk in a subject based on the presence of one or more sex-specific features in the health record that has been identified to predict disease onset risk based on the method of claim 2; andadministering a treatment to the subject in an amount effective to treat the disease, in the event that subject is predicted to be at risk of developing the disease.

9. The method of any of the previous claims, wherein the disease is a neurological disease.

10. The method of any of the previous claims, wherein the disease is Alzheimer’s disease.

11. The method of any one of claims 2 to 10, wherein the one or more sexspecific features are selected from osteoporosis, major depressive disorder, allergic rhinitis, abnormal stool contents, chest pain, hypovolemia and prostate hyperplasia.

12. The method of claim 11, wherein the subject is female and the one or more sex-specific features are selected from osteoporosis, major depressive disorder, allergic rhinitis and abnormal stool contents.

13. The method of claim 11, wherein the subject is male and the one or more sex-specific features are selected from chest pain, hypovolemia and prostate hyperplasia.

14. The method of any of the previous claims, wherein the type of machine learning model is selected based on an ease of interpretability of the model and an ability to capture nonlinear relationships.

15. The method of any of the previous claims, wherein the machine learning model comprises a binary classification time point model.

16. The method of any of the previous claims, wherein the machine learning model comprises a random forest model.

17. The method of any of the previous claims, wherein the machine learning model comprises a plurality of random forest models.

18. The method of any of the previous claims, wherein training and evaluating the machine learning model comprises configuring the model to improve an accuracy of predicting disease onset risk.

19. The method of any of the previous claims, wherein training and evaluating the machine learning model comprises adjusting one or more parameters of the machine learning model.

20. The method of any of the previous claims, wherein training the machine learning model comprises training the model using sex-stratified subgroups.

21. The method of any of the previous claims, wherein training the machine learning model using the electronic health records of the first set of subjects comprises training the model using:(a) clinical features only of the electronic health records,(b) clinical features and demographic information of the electronic health records, or(c) clinical features, demographic information and visit-related information of the electronic health records.

22. The method of any of the previous claims, wherein training the machine learning model using the electronic health records of the first set of subjects comprises training the model using samples of subjects matched by demographic information or hospital utilization or visit-related features.

23. The method of any of the previous claims, wherein matching subjects based on demographics comprises matching subjects based on one or more of: birth year, race and ethnicity and / or sex.

24. The method of any of the previous claims, wherein matching subjects based on visit-related features comprises matching subjects based on one or more of: age, first visit age, years electronic health records, a value corresponding to a logarithm of the number of prior visits, a value corresponding to a logarithm of the number of prior concepts or a value corresponding to a logarithm of a number of days since a first clinical event.

25. The method of any of the previous claims, wherein training the machine learning model using the electronic health records of the first set of subjects comprises training the model using samples of subjects matched using a propensity score match.

26. The method of any of the previous claims, wherein training the machine learning model using the electronic health records of the first set of subjects comprises training the model using samples of subjects of a single sex.

27. The method of any of the previous claims, wherein identifying from the trained and evaluated machine learning model one or more features of the subjects that predict disease onset risk comprises identifying an average importance for a plurality of features across a plurality of time points.

28. The method of any of the previous claims, wherein identifying from the trained and evaluated machine learning model one or more features of the subjects that predict disease onset risk comprises ranking a plurality of features at a plurality of different time points.

29. The method of any of the previous claims, wherein identifying from the trained and evaluated machine learning model one or more features of the subjects that predict disease onset risk comprises identifying features of subjects at different time points that predict disease onset risk.

30. The method of any of the previous claims, wherein the identified features of the subjects comprises one or more phenotypes presented in the electronic health records.

31. The method of any of the previous claims, wherein the identified features of the subjects comprises one or more diagnostic features that have been mapped to phecodes.

32. The method of any of the previous claims, wherein the method is a method of identifying a candidate subject for an early intervention for the disease.

33. The method of any of the previous claims, wherein the method is a method of identifying a candidate subject for a disease intervention prior to disease diagnosis.

34. The method of any of the previous claims, wherein the method is a computer-implemented method.

35. A system for training a model to predict a disease onset risk, the system comprising:a processor comprising memory operably coupled to the processor, wherein the memory comprises instructions stored thereon, which, when executed by the processor, cause the processor to:obtain a plurality of electronic health records comprising health information about a plurality of subjects, wherein the plurality of subjects comprise subjects diagnosed with the disease and subjects not diagnosed with the disease;divide the subjects into a first set for machine learning model training and a second set for machine learning evaluation;train a machine learning model using the electronic health records of the first set of subjects;evaluate the machine learning model using the electronic health records of the second set of subjects; andidentify from the trained and evaluated machine learning model one or more features of the subjects that predict disease onset risk.

36. A system for training a model to predict a disease onset risk, the system comprising:a processor comprising memory operably coupled to the processor, wherein the memory comprises instructions stored thereon, which, when executed by the processor, cause the processor to:obtain a plurality of electronic health records comprising health information about a plurality of subjects, wherein the plurality of subjects comprise subjects diagnosed with the disease and subjects not diagnosed with the disease, and wherein the electronic health records include sexes of the subjects;divide the subjects into a first set for machine learning model training and a second set for machine learning evaluation;train a machine learning model using the electronic health records of the first set of subjects;evaluate the machine learning model using the electronic health records of the second set of subjects; andidentify from the trained and evaluated machine learning model one or more sexspecific features of the subjects that predict the disease onset risk.

37. The system according to claim 36, wherein the memory further comprises instructions stored thereon, which, when executed by the processor, cause the processor to:obtain a health record comprising health information about a subject, wherein the health record includes a sex of the subject; andapply the trained model to the health record of the subject to predict the disease onset risk in the subject.

38. The system according to claim 36, wherein the memory further comprises instructions stored thereon, which, when executed by the processor, cause the processor to:obtain a health record comprising health information about a subject, wherein the health record includes a sex of the subject;apply the trained model trained to the health record of the subject to predict the risk of developing the disease; andclinically evaluate the presence of the disease, in the event that subject is predicted to be at risk of developing the disease.

39. The system according to claim 36, wherein the memory further comprises instructions stored thereon, which, when executed by the processor, cause the processor to:obtain a health record comprising health information about a subject, wherein the health record includes a sex of the subject;apply the trained model to the health record of the subject to predict the risk of developing the disease; andadminister a treatment to the subject in an amount effective to treat the disease, in the event that subject is predicted to be at risk of developing the disease.

40. The system according to claim 36, wherein the memory further comprises instructions stored thereon, which, when executed by the processor, cause the processor to:obtain a health record comprising health information about a subject, wherein the health record includes a sex of the subject; andpredict the disease onset risk in a subject based on the presence of one or more sex-specific features in the health record that has been identified to predict the disease onset risk based on the trained model.

41. The system according to claim 36, wherein the memory further comprises instructions stored thereon, which, when executed by the processor, cause the processor to:obtain a health record comprising health information about a subject, wherein the health record includes a sex of the subject;predict the disease onset risk in a subject based on the presence of one or more sex-specific features in the health record that has been identified to predict disease onset risk based the trained model; andclinically evaluate the presence of the disease, in the event that subject is predicted to be at risk of developing the disease.

42. The system according to claim 36, wherein the memory further comprises instructions stored thereon, which, when executed by the processor, cause the processor to:obtain a health record comprising health information about a subject, wherein the health record includes a sex of the subject;predict the disease onset risk in a subject based on the presence of one or more sex-specific features in the health record that has been identified to predict disease onset risk based on the trained model; andadminister a treatment to the subject in an amount effective to treat the disease, in the event that subject is predicted to be at risk of developing the disease.

43. The system of any one of claims 35 to 42, wherein the disease is a neurological disease.

44. The system of any one of claims 35 to 43, wherein the disease is Alzheimer’s disease.

45. The system of any one of claims 36 to 44, wherein the one or more sexspecific features are selected from osteoporosis, major depressive disorder, allergic rhinitis, abnormal stool contents, chest pain, hypovolemia and prostate hyperplasia.

46. The system of claim 45, wherein the subject is female and the one or more sex-specific features are selected from osteoporosis, major depressive disorder, allergic rhinitis and abnormal stool contents.

47. The system of claim 45, wherein the subject is male and the one or more sex-specific features are selected from chest pain, hypovolemia and prostate hyperplasia.

48. The system of any one of claims 35 to 47, wherein the type of machine learning model is selected based on an ease of interpretability of the model and an ability to capture nonlinear relationships.

49. The system of any one of claims 35 to 48, wherein the machine learning model comprises a binary classification time point model.

50. The system of any one of claims 35 to 49, wherein the machine learning model comprises a random forest model.

51. The system of any one of claims 35 to 50, wherein the machine learning model comprises a plurality of random forest models.

52. The system of any one of claims 35 to 51, wherein training and evaluating the machine learning model comprises configuring the model to improve an accuracy of predicting disease onset risk.

53. The system of any one of claims 35 to 52, wherein training and evaluating the machine learning model comprises adjusting one or more parameters of the machine learning model.

54. The system of any one of claims 35 to 53, wherein training the machine learning model comprises training the model using sex-stratified subgroups.

55. The system of any one of claims 35 to 54, wherein training the machine learning model using the electronic health records of the first set of subjects comprises training the model using:(a) clinical features only of the electronic health records,(b) clinical features and demographic information of the electronic health records, or(c) clinical features, demographic information and visit-related information of the electronic health records.

56. The system of any one of claims 35 to 55, wherein training the machine learning model using the electronic health records of the first set of subjects comprises training the model using samples of subjects matched by demographic information or hospital utilization or visit-related features.

57. The system of any one of claims 35 to 56, wherein matching subjects based on demographics comprises matching subjects based on one or more of: birth year, race and ethnicity and / or sex.

58. The system of any one of claims 35 to 57, wherein matching subjects based on visit-related features comprises matching subjects based on one or more of: age, first visit age, years electronic health records, a value corresponding to a logarithm of the number of prior visits, a value corresponding to a logarithm of the number of prior concepts or a value corresponding to a logarithm of a number of days since a first clinical event.

59. The system of any one of claims 35 to 58, wherein training the machine learning model using the electronic health records of the first set of subjects comprises training the model using samples of subjects matched using a propensity score match.

60. The system of any one of claims 35 to 59, wherein training the machine learning model using the electronic health records of the first set of subjects comprises training the model using samples of subjects of a single sex.

61. The system of any one of claims 35 to 60, wherein identifying from the trained and evaluated machine learning model one or more features of the subjects that predict disease onset risk comprises identifying an average importance for a plurality of features across a plurality of time points.

62. The system of any of claims 33 to 61, wherein identifying from the trained and evaluated machine learning model one or more features of the subjects that predictdisease onset risk comprises ranking a plurality of features at a plurality of different time points.

63. The system of any of claims 33 to 62, wherein identifying from the trained and evaluated machine learning model one or more features of the subjects that predict disease onset risk comprises identifying features of subjects at different time points that predict disease onset risk.

64. The system of any of claims 33 to 63, wherein identifying from the trained and evaluated machine learning model one or more features of the subjects that predict disease onset risk comprises identifying features of subjects of points that predict disease onset risk corresponding to a sex of the subjects.

65. The system of any of claims 33 to 64, wherein the identified features of the subjects comprises one or more phenotypes presented in the electronic health records.

66. The system of any of claims 33 to 65, wherein the identified features of the subjects comprises one or more diagnostic features that have been mapped to phecodes.

67. A non-transitory computer readable storage medium comprising instructions stored thereon, the instructions comprising:algorithm for obtaining a plurality of electronic health records comprising health information about a plurality of subjects, wherein the plurality of subjects comprise subjects diagnosed with the disease and subjects not diagnosed with the disease;algorithm for dividing the subjects into a first set for machine learning model training and a second set for machine learning evaluation;algorithm for training a machine learning model using the electronic health records of the first set of subjects;algorithm for evaluating the machine learning model using the electronic health records of the second set of subjects; andalgorithm for identifying from the trained and evaluated machine learning model one or more features of the subjects that predict the disease onset risk.

68. A non-transitory computer readable storage medium comprising instructions stored thereon, the instructions comprising:algorithm for obtaining a plurality of electronic health records comprising health information about a plurality of subjects, wherein the plurality of subjects comprise subjects diagnosed with the disease and subjects not diagnosed with the disease, and wherein the electronic health records include sexes of the subjects;algorithm for dividing the subjects into a first set for machine learning model training and a second set for machine learning evaluation;algorithm for training a machine learning model using the electronic health records of the first set of subjects;algorithm for evaluating the machine learning model using the electronic health records of the second set of subjects; andalgorithm for identifying from the trained and evaluated machine learning model one or more sex-specific features of the subjects that predict the disease onset risk.

69. The non-transitory computer readable storage medium of claim 68, the instructions stored thereon further comprising:algorithm for algorithm for obtaining a health record comprising health information about a subject, wherein the health record includes a sex of the subject; andalgorithm for algorithm for applying the trained model to the health record of the subject to predict the disease onset risk in the subject.

70. The non-transitory computer readable storage medium of claim 68, the instructions stored thereon further comprising:algorithm for obtaining a health record comprising health information about a subject, wherein the health record includes a sex of the subject;algorithm for applying the trained model to the health record of the subject to predict the risk of developing the disease; andalgorithm for clinically evaluating the presence of the disease, in the event that subject is predicted to be at risk of developing the disease.

71. The non-transitory computer readable storage medium of claim 68, the instructions stored thereon further comprising:algorithm for obtaining a health record comprising health information about a subject, wherein the health record includes a sex of the subject;algorithm for applying the trained model to the health record of the subject to predict the risk of developing the disease; andalgorithm for administering a treatment to the subject in an amount effective to treat the disease, in the event that subject is predicted to be at risk of developing the disease.

72. The non-transitory computer readable storage medium of claim 68, the instructions stored thereon further comprising:algorithm for obtaining a health record comprising health information about a subject, wherein the health record includes a sex of the subject; andalgorithm for predicting the disease onset risk in a subject based on the presence of one or more sex-specific features in the health record that has been identified to predict the disease onset risk based on the trained model.

73. The non-transitory computer readable storage medium of claim 68, the instructions stored thereon further comprising:algorithm for obtaining a health record comprising health information about a subject, wherein the health record includes a sex of the subject;algorithm for predicting the disease onset risk in a subject based on the presence of one or more sex-specific features in the health record that has been identified to predict disease onset risk based on the trained model; andalgorithm for clinically evaluating the presence of the disease, in the event that subject is predicted to be at risk of developing the disease.

74. The non-transitory computer readable storage medium of claim 68, the instructions stored thereon further comprising:algorithm for obtaining a health record comprising health information about a subject, wherein the health record includes a sex of the subject;algorithm for predicting the disease onset risk in a subject based on the presence of one or more sex-specific features in the health record that has been identified to predict disease onset risk based on the trained model; andalgorithm for administering a treatment to the subject in an amount effective to treat the disease, in the event that subject is predicted to be at risk of developing the disease.

75. The non-transitory computer readable storage medium of any one of claims 67 to 74, wherein the disease is a neurological disease.

76. The non-transitory computer readable storage medium of any one of claims 67 to 75, wherein the disease is Alzheimer’s disease.

77. The non-transitory computer readable storage medium of any one of claims 68 to 76, wherein the one or more sex-specific features are selected from osteoporosis, major depressive disorder, allergic rhinitis, abnormal stool contents, chest pain, hypovolemia and prostate hyperplasia.

78. The non-transitory computer readable storage medium of claim 77, wherein the subject is female and the one or more sex-specific features are selected from osteoporosis, major depressive disorder, allergic rhinitis and abnormal stool contents.

79. The non-transitory computer readable storage medium of claim 77, wherein the subject is male and the one or more sex-specific features are selected from chest pain, hypovolemia and prostate hyperplasia.

80. The non-transitory computer readable storage medium of any one of claims 67 to 79, wherein the type of machine learning model is selected based on an ease of interpretability of the model and an ability to capture nonlinear relationships.

81. The non-transitory computer readable storage medium of any one of claims 67 to 80, wherein the machine learning model comprises a binary classification time point model.

82. The non-transitory computer readable storage medium of any one of claims 67 to 81, wherein the machine learning model comprises a random forest model.

83. The non-transitory computer readable storage medium of any one of claims 67 to 82, wherein the machine learning model comprises a plurality of random forest models.

84. The non-transitory computer readable storage medium of any one of claims 67 to 83, wherein training and evaluating the machine learning model comprises configuring the model to improve an accuracy of predicting disease onset risk.

85. The non-transitory computer readable storage medium of any one of claims 67 to 84, wherein training and evaluating the machine learning model comprises adjusting one or more parameters of the machine learning model.

86. The non-transitory computer readable storage medium of any one of claims 67 to 85, wherein training the machine learning model comprises training the model using sex-stratified subgroups.

87. The non-transitory computer readable storage medium of any one of claims 67 to 86, wherein training the machine learning model using the electronic health records of the first set of subjects comprises training the model using:(a) clinical features only of the electronic health records,(b) clinical features and demographic information of the electronic health records, or(c) clinical features, demographic information and visit-related information of the electronic health records.

88. The non-transitory computer readable storage medium of any one of claims 67 to 87, wherein training the machine learning model using the electronic health records of the first set of subjects comprises training the model using samples of subjects matched by demographic information or hospital utilization or visit-related features.

89. The non-transitory computer readable storage medium of any one of claims 67 to 88, wherein matching subjects based on demographics comprises matching subjects based on one or more of: birth year, race and ethnicity and / or sex.

90. The non-transitory computer readable storage medium of any one of claims 67 to 89, wherein matching subjects based on visit-related features comprises matching subjects based on one or more of: age, first visit age, years electronic health records, a value corresponding to a logarithm of the number of prior visits, a value corresponding to a logarithm of the number of prior concepts or a value corresponding to a logarithm of a number of days since a first clinical event.

91. The non-transitory computer readable storage medium of any one of claims 67 to 90, wherein training the machine learning model using the electronic health records of the first set of subjects comprises training the model using samples of subjects matched using a propensity score match.

92. The non-transitory computer readable storage medium of any one of claims 67 to 91, wherein training the machine learning model using the electronic health records of the first set of subjects comprises training the model using samples of subjects of a single sex.

93. The non-transitory computer readable storage medium of any one of claims 67 to 92, wherein identifying from the trained and evaluated machine learning model one or more features of the subjects that predict disease onset risk comprises identifying an average importance for a plurality of features across a plurality of time points.

94. The non-transitory computer readable storage medium of any one of claims 67 to 93, wherein identifying from the trained and evaluated machine learning model one or more features of the subjects that predict disease onset risk comprises ranking a plurality of features at a plurality of different time points.

95. The non-transitory computer readable storage medium of any one of claims 67 to 94, wherein identifying from the trained and evaluated machine learning model one or more features of the subjects that predict disease onset risk comprises identifying features of subjects at different time points that predict disease onset risk.

96. The non-transitory computer readable storage medium of any one of claims 67 to 95, wherein the identified features of the subjects comprises one or more phenotypes presented in the electronic health records.

97. The non-transitory computer readable storage medium of any one of claims 67 to 96, wherein the identified features of the subjects comprises one or more diagnostic features that have been mapped to phecodes.