Using blood transcriptome analysis for alzheimer's disease diagnosis and patient stratification
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-01-30
- Publication Date
- 2026-08-13
Smart Images

Figure CN2026076141_13082026_PF_FP_ABST
Abstract
Description
USING BLOOD TRANSCRIPTOME ANALYSIS FOR ALZHEIMER′S DISEASE DIAGNOSIS AND PATIENT STRATIFICATIONCROSS-REFERENCES TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Application No. 63 / 754,355, filed February 5, 2025, the disclosure of which is incorporated by reference herein.BACKGROUND
[0002] Alzheimer’s disease (AD) , a prevalent and devastating neurodegenerative condition, presents a significant burden on global healthcare systems as a leading cause of mortality among aging individuals. Current diagnostic methodologies for AD rely heavily on subjective evaluations, often lacking the sensitivity needed for early disease detection and the specificity required for distinguishing AD from other forms of dementia. The discovery of blood-based biomarkers, including amyloid-β, phosphorylated tau, and neurofilament light polypeptide, has provided significant advancement for AD diagnosis. However, the complete representation of disease states during AD progression and the underlying molecular mechanisms are not yet understood.
[0003] Transcriptome analysis, which refers generally to analysis of RNA molecules in a cell population, is emerging as a robust tool for investigating the activity of genes across different biological conditions. Among other uses, transcriptome analysis has offered valuable insights into the molecular alterations within the brains of individuals with AD. Given the accessibility of blood samples and the emerging potential of blood-based diagnostics, there is a growing interest in utilizing blood transcriptome analysis for AD research.SUMMARY
[0004] Certain embodiments of the invention relate to assessment of Alzheimer's Disease (AD) risk based on transcriptome analysis applied to a selected subset of genes. The selected subset can be small. For example, just four genes can be selected from each of six gene modules, for a total of 24 genes. Expression levels of the 24 selected genes can be determined via transcriptome analysis from a blood sample obtained from a subject. The gene expression levels can be input (e.g., as a gene expression matrix) to a machine learning model that has been trained to generate an AD risk score based on an input gene expression matrix. The AD risk score generated for a particular subject by the machine learning model can be used to determine whether the subject has an increased risk of developing AD.
[0005] The following detailed description, together with the accompanying drawings, will provide a better understanding of the nature and advantages of the claimed invention.BRIEF DESCRIPTION OF THE DRAWINGS
[0006] FIG. 1 shows a listing of gene modules with their correlation to AD and the gene count for each gene module.
[0007] FIG. 2 shows a flow diagram of a process for training and using a machine learning model according to some embodiments.
[0008] FIG. 3 shows a flow diagram of a clinical process according to some embodiments.
[0009] FIGs. 4A and 4B show graphs of performance metrics for a trained model according to some embodiments.DETAILED DESCRIPTION
[0010] The following description of exemplary embodiments of the invention is presented for the purpose of illustration and description. It is not intended to be exhaustive or to limit the claimed invention to the precise form described, and persons skilled in the art will appreciate that many modifications and variations are possible. The embodiments have been chosen and described in order to best explain the principles of the invention and its practical applications to thereby enable others skilled in the art to best make and use the invention in various embodiments and with various modifications as are suited to the particular use contemplated.
[0011] A previous study by the present inventors (Zhong et al., Using blood transcriptome analysis for Alzheimer’s disease diagnosis and patient stratification, Alzheimer’s Dement. 2024; 20: 2469–84) explored the potential of utilizing blood transcriptome analysis diagnosis and characterization of Alzheimer’s Disease (AD) . Through bulk RNA-sequencing (RNA-seq) in blood, specific genes and gene modules (where a gene module comprises dozens to thousands of co-expressed genes) associated with AD were identified. The previous study showed that the six most significant AD-associated gene modules can be used to accurately classify AD and stratify individuals by disease status. However, these gene modules include a total of 9,484 genes, making the volume of data impractical for clinical application.
[0012] Certain embodiments of the invention relate to techniques for assessing AD risk based on a reduced set of genes. In some embodiments, the reduced set of genes can consist of just 24 genes. Expression levels of the 24 genes can be determined from a blood sample obtained from a subject. The gene expression levels can be formed into a matrix that can be input to a machine learning model that has been trained to generate an AD risk score based on an input gene expression matrix. The AD risk score generated for a particular subject by the machine learning model can be used to determine whether the subject has an increased risk of developing AD. Sample Preparation and Gene Expression Assessment
[0013] The first step of practicing the present invention is to obtain a blood sample from a subject being tested for assessing the risk of developing a neurodegenerative disorder (such as AD) or monitoring for the neurodegenerative disorder’s severity or progression. Samples of the same type should be taken from both a control group (cognitively normal individuals not suffering from AD and without increased risk for AD) and a test group (subjects being tested for possible AD or for increased risk for AD, for example) . Standard blood drawing procedures routinely employed in hospitals or clinics are typically followed for this purpose.
[0014] The blood sample from a subject for use in the present invention can be processed for measuring marker gene expression by well-known methods and as described in standard medical literature. In certain applications of this invention, an acelluar portion of the whole blood (e.g., serum or plasma) may be the preferred sample type. In other cases, a whole blood sample or a cellular portion (e.g., all blood cells or a sub-population of blood cells) may be used.
[0015] In some embodiments, the expression level of genes of interest (for example, one or more or all 24 genes in Table 1) is determined based on the level of their mRNA in a blood sample taken from an individual being tested. Often, a particular mRNA species is first extracted from a sample, its amount then quantified. The mRNA can be detected using other standard techniques, well known to those of skill in the art. Although the detection step is typically preceded by an amplification step, amplification is not required in all instances. For instance, the mRNA may be identified by size fractionation (e.g., gel electrophoresis) , whether or not proceeded by an amplification step. After running a sample in an agarose or polyacrylamide gel and labeling with ethidium bromide according to well-known techniques (see, e.g., Sambrook and Russell, Molecular Cloning, A Laboratory Manual (3rd ed. 2001) ) , the presence of a band of the same size as the standard comparison is an indication of the presence of a target mRNA, the amount of which may then be compared to the control based on the intensity of the band. One particular method for determining the mRNA level is an amplification-based method, e.g., by polymerase chain reaction (PCR) , especially reverse transcription-polymerase chain reaction (RT-PCR) and quantitative RT-PCR (qRT-PCR) . The general methods of PCR are well known in the art, see, e.g., Innis, et al., PCR Protocols: A Guide to Methods and Applications, Academic Press, Inc. N.Y., 1990. PCR reagents and protocols are also available from commercial vendors, such as Roche Molecular Systems. One preferred method for specific RNA detection and analysis is described in the 2024 Zhong et al. publication referenced above.
[0016] In other embodiments, the gene expression level is measured by assessing protein expression level in a patient blood sample. A variety of well-known analytical assays, especially immunological assays, may be used for this purpose. For example, a protein of interest (e.g., a protein encoded by any one gene selected from the 24 genes named in Table 1) can be detected by gel electrophoresis (such as 2-dimensional gel electrophoresis) and western blot analysis using a specific antibody. In some embodiments, an immunological assay known as sandwich assay can be performed by capturing a target polypeptide from a test sample with an antibody having specific binding affinity for the target polypeptide. Optionally, the antibody is conjugated to a detectable moiety, such that the presence of the target polypeptide hen can be readily detected based on the detectable signal. Such immunological assays can be carried out using microfluidic devices such as microarray chips. A protein of interest (e.g., a protein encoded by any one gene selected from the 24 genes named in Table 1) can also be detected by an immunoassay involving a secondary antibody that carries a detectable label and is thus readily detectable. For instance, the presence and quantity of a target protein X in a blood sample may be determined first by specific binding of X protein to a murine IgG anti-human X antibody, which is then detected by virtue of its specific binding to a rabbit anti-murine IgG antibody, a secondary antibody that is conjugated to a detectable moiety (for example, a label emitting a fluorescent or chemiluminescent signal) and therefore permits rapid, easy, and accurate quantitative identification of target protein X.
[0017] Other methods may also be employed for measuring the level of protein expression in practicing the present invention. For instance, a variety of methods have been developed based on the mass spectrometry technology to rapidly and accurately quantify target proteins even in a large number of samples. These methods involve highly sophisticated equipment such as the triple quadrupole (triple Q) instrument using the multiple reaction monitoring (MRM) technique, matrix assisted laser desorption / ionization time-of-flight tandem mass spectrometer (MALDI TOF / TOF) , an ion trap instrument using selective ion monitoring SIM) mode, and the electrospray ionization (ESI) based QTOP mass spectrometer. See, e.g., Pan et al., J Proteome Res. 2009 February; 8 (2) : 787–797. Reduced Set of Genes
[0018] In an example embodiment, the selection of a reduced set of genes is based on earlier work by the inventors (Zhong et al., referenced above) in which gene modules associated with AD were identified. As used herein, a “gene module” refers to a set of co-expressed genes, and a single gene module can include dozens to thousands of genes. Association of a gene or gene module with AD can be determined by identifying genes that are differentially expressed between subjects with AD and normal subjects, using known techniques such as DESeq2 (described in Love et al., Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2, Genome biology 2014; 15: 550) . Age, sex, and population structure can be included as covariates.
[0019] FIG. 1 shows a listing of gene modules (M01-M15) . For each gene module M01-M15, the gene count and a correlation coefficient to AD, as determined in the earlier work, is indicated. Of the sixteen gene modules M01-M15, the most significant (as determined by absolute value of the correlation coefficient) are M01, M02, M05, M09, M13, and M15, which contain a total of 9, 484 genes. This volume of data is impractical for clinical applications, and a reduction in the number of genes is therefore desirable.
[0020] According to some embodiments of the invention, a small number of genes from each of the most significant gene modules is selected. The number of selected genes per gene module can be a pre-selected number that strikes a desired balance between model simplicity and robustness. In some embodiments, four genes are selected from each of the six most significant gene modules. In one example of a selection process, an initial selection step identifies genes within each gene module that have the strongest association with AD. Strength of association can be determined by computing p-values using standard statistical techniques applied to a set of data samples representing individuals with and without AD. A subsequent selection step refines the initial selection by prioritizing the genes with the highest connectivity to other genes within their respective gene modules. ( “Connectivity” in this context refers to the degree of interaction between genes, which can be measured by the number of interactions one gene’s product has with other genes’ products. ) In this manner, four genes (or some other preselected number of genes) from each gene module can be selected.
[0021] By way of specific example, Table 1 lists a set of selected genes for an embodiment where four genes per gene module are selected from each of the six most significant gene modules for AD. For each selected gene, the p-values and connectivity score is also listed. TABLE 1 Machine Learning Model
[0022] According to some embodiments, a machine learning model can be trained to use the gene expression data to predict AD risk scores. For example, blood samples can be obtained from individual “training” subjects whose AD condition is known, and gene expression levels for each of the genes in the selected set of genes can be determined from the blood samples using transcriptome analysis techniques. Gene expression levels can be represented using structured data such as a matrix or vector (which is a one-dimensional matrix) . The particular data structure can be selected according to the type of machine learning model. The training data preferably includes data from individuals with AD and individuals without AD, as determined using cognitive assessments and optionally other information such as blood-based biomarkers. The known AD status of the training subjects can be used as ground truth labeling for supervised training.
[0023] Various machine learning models can be used, including logistic regression models, other generalized linear models, or other classifier models. The output can be, for example, a score reflecting a risk level or a binary score (e.g., whether or not the subject has AD) .
[0024] Training of the machine learning model involves automated processes to determine, or “learn, ” optimal values for internal parameters of the model, such as the weights for each node in a neural network model or coefficients of a parametric function such as a curve-fitting function or a transform function. A standard approach to training involves iteratively processing a set of training data samples through the model and adjusting the parameters of the model, with the goal of minimizing a loss function that characterizes a difference between the output of the model for a given input and an expected result (or ground truth) determined from a source other than the model, such as labels assigned to the training data samples. Loss functions can be selected based in part on the particular model and performance goals, and optimization of loss functions can proceed using various techniques. Training may occur across multiple “epochs, ” where each epoch corresponds to a pass through the training data set. Adjustment to parameters of the model (e.g., weights or coefficients) can occur multiple times during an epoch; for instance, the training data can be divided into “batches” or “mini-batches” and parameter adjustment can occur after each batch or mini-batch. In some embodiments, publicly available software such as the R Stats Package can be used to train a model such as a generalized linear model.
[0025] Once trained, the model can be used in inference mode to evaluate previously unseen subjects. For validation of the model, subjects whose AD condition is known can be used, provided that these subjects were not included in the training data set. In some embodiments, model performance can be assessed using metrics such as area under the receiver operating curve (auROC) . Other metrics can also be used. Example Process
[0026] FIG. 2 is a flow diagram of a process 200 for training and using a machine learning model according to some embodiments. Portions of process 200 can be implemented using one or more computer systems of generally conventional design executing program code to perform computations. Process 200 includes a gene selection phase 210, a training phase 230, and an inference phase 250.
[0027] Gene selection phase 210 can be used to identify a small number (e.g., fewer than 100 or fewer than 50 or about 30 or 24) of genes whose expression level is strongly associated with AD.
[0028] For example, at block 212, a set of gene modules that are most significantly associated with AD is identified. As described above, identification of the set of gene modules can be based on empirical data of gene expression levels determined from blood samples of individuals with and without AD; tools such as DESeq2 can be employed. The set of gene modules can include a preselected number (M) of gene modules, where M can be, e.g., fewer than 20 or fewer than 10 or about 5 or 6 or another number.
[0029] At block 214, an initial selection of genes within each gene module is performed. The initial selection can be based on determining a p-value for each gene in the gene module based on its correlation with AD, then ranking all of the genes within the gene module according to their p-values. A preliminary set of genes with the strongest associations can thereby be identified, for instance based on p-values below a certain threshold or a certain number of genes (e.g., top 100 or top 20 or top 10 or the like) .
[0030] At block 216, a refined selection for each gene module is made from the preliminary set of genes. For instance, the selection can be refined by selecting a number (N) of genes from the preliminary set that have the highest connectivity to other genes in the gene module. N can be a preselected number that is the same for each gene module; for instance, N can be 4 or 8 or 10 or the like.
[0031] In this manner, a set of M*N genes can be selected for analysis. The preselected numbers M and N can be chosen as desired, based on tradeoffs between robustness and clinical usability. Table 1 above shows an example of a specific set of 24 genes for an embodiment where M=6 and N=4.
[0032] Training phase 230 can begin at block 232 with obtaining gene expression data for a number of subjects to form a training data set. The subjects preferably include individuals whose AD status has been determined, including individuals with and without AD. Obtaining the gene expression data can include obtaining blood samples from the subjects and extracting RNA sequence data from the blood samples using conventional techniques. The RNA sequence data can be refined by filtering adapters and removing low-quality reads. The refined reads can be mapped to a reference genome (e.g., the GRCh38 reference genome) using conventional gene annotation software such as STAR (2.7.10b) (described in Dobin et al., STAR: ultrafast universal RNA-seq aligner, Bioinformatics 2013; 29: 15–21) . Transcriptomic abundance, which corresponds to gene expression levels, can be determined using conventional quantification software such as RSEM (1.3.3) (described in Li et al., RSEM: accurate transcript quantification from RNA-Seq data with or without a reference genome, BMC Bioinformatics 2011; 12: 323) .
[0033] At block 234, the gene expression data for each subject can be used to construct a gene expression matrix for the selected genes (i.e., the M*N genes selected in gene selection phase 210) . In some embodiments, the gene expression matrix can be a one-dimensional matrix (or vector) . Each gene expression matrix can be annotated with a ground-truth label indicating whether the subject has AD.
[0034] At block 236, a machine learning model is trained using the set of labeled gene expression matrices. The machine learning model can be, for instance, a generalized linear model, a logistic regression model, or other type of model. Neural networks can be used if desired, although the small size of the gene expression matrix may reduce any benefit of larger or more complex models. Training can proceed as described above. Through training, the machine learning model learns a set of parameters that can be used to compute an AD risk score from an input gene expression matrix. Depending on implementation, the AD risk score can be a binary score (AD or not-AD) or a variable score, e.g., a continuously-valued output indicating a likelihood that the subject has (or will develop) AD.
[0035] Inference phase 250 can begin at block 252 with obtaining gene expression data for a test subject, preferably a subject that was not included in the training data set. Gene expression data for the selected set of genes in the test subject can be acquired in the same manner as for the training subjects at block 232; however, in inference phase it is not necessary to establish ground truth of whether the subject has AD. At block 254, the gene expression data for the test subject is used to construct a gene expression matrix in the same manner as for the training subjects at block 234. At block 256, the gene expression matrix for the test subject is input to the trained machine learning model, and at block 258, the trained machine learning model outputs an AD risk score. The AD risk score can be used to determine whether the test subject has AD or risk of developing AD, alone or in combination with other diagnostic tools. Clinical Application
[0036] It is contemplated that the inference mode of process 200 can be applied in a clinical setting to determine an AD risk score for a patient whose condition is not initially known. For example, a clinician can be provided with a test kit that includes reagents to measure the gene expression level of each of the M*N selected genes (e.g., the genes listed in Table 1) . The test kit can facilitate obtaining a gene expression matrix for a particular patient, which can be input into a previously trained machine learning model to obtain an AD risk score for the patient.
[0037] FIG. 3 shows a flow diagram of a clinical process 300 according to some embodiments. At block 302, a clinician can obtain a blood sample from a patient. At block 304, the clinician can use the test kit to determine gene expression levels for each of the M*N selected genes. For instance, the kit can be designed as a lab-on-a-chip (LOC) with wells containing reagents for each gene and integrated electronics to provide quantitative data based on reactions with the reagents. At block 306, the gene expression levels can be input to a trained machine learning model. For example, the clinician can be provided with software that implements an inference mode of the machine learning model; inputting the gene expression levels. In some LOC implementations, the model computation can be incorporated into the chip, e.g., using appropriate digital logic circuitry. At block 308, the machine learning model can output an AD risk score.
[0038] A clinician can use the AD risk score, optionally in combination with other diagnostic tools, to diagnose AD. Treatment and Monitoring Methods for AD
[0039] In a related aspect, the present invention also supports and enables treatment for AD upon detection of the presence of the neurodegenerative disorder or a heightened risk of later developing the neurodegenerative disorder in a patient. In some embodiments, the method comprises, upon determining a subject as having an increased risk for AD, administering a treatment to said subject, for example, antibody drugs such as lecanemab, acetylcholinesterase inhibitors (such as donepezil, galantamine, rivastigmine) , memantine, glutamate receptor blockers, citalopram, fluoxetine, paroxeine, sertraline, trazodone, lorazepam, oxazepam, aripiprazole, clozapine, haloperidol, olanzapine, quetiapine, risperidone, ziprasidone, nortriptyline, tricyclic antidepressants, benzodiazepines, temazepam, zolpidem, zaleplon, chloral hydrate, coenzyme Q10, ubiquinone, coral calcium, Ginkgo biloba, huperzine A, omega-3 fatty acids, phosphatidylserine, or any combination thereof.
[0040] In some cases, when the diagnostic method steps described above and herein are completed, optionally with additional diagnostic examination performed to provide further confirmatory information (for example, by brain imaging via CT scan or other imaging techniques to show excessive loss of brain volume, or by testing cognitive capability to show an accelerated decline) , and a patient has been determined to either already have AD or is at a significantly increased risk of later developing AD, suitable therapeutic or prophylactic regimens may be ordered by physicians or other medical professionals to treat the patient, to manage and alleviate the ongoing symptoms, or to delay the future onset of the disease. The U.S. Food and Drug Administration (FDA) has approved a number of cholinesterase inhibitors, including donepezil (AriceptTM, the only cholinesterase inhibitor approved to treat all stages of AD, including moderate to severe) , rivastigmine (ExelonTM, approved to treat mild to moderate AD) , galantamine (RazadyneTM, mild to moderate patients) and memantine (NamendaTM) . Donepezil is the only cholinesterase inhibitor approved to treat all stages of AD, including moderate to severe. Any one or more of these drugs can be prescribed for treating patients who have been diagnosed with AD in accordance with the methods of this invention. Also, Biogen’s antibody drugs aducanumab and lecanemab recently received full approval by the FDA. Another possibility of treatment is administration of trazodone, which is currently approved for use as an antidepressant and has been reported as an effective agent for ameliorating AD symptoms.
[0041] For patients who are deemed at a high or an increased risk for developing AD in a future time but do not yet exhibit any clinical symptoms, continuous monitoring is also appropriate, especially at an increased frequency. For example, the patients may be subject to more frequently scheduled regular testing (e.g., once every six months, once a year, or once every two years) to detect any accelerated change in their cognitive capabilities. Methods suitable for such regular monitoring include General Practitioner Assessment of Cognition (GPCOG) , Mini-Cog, Eight-item Informant Interview to Differentiate Aging and Dementia (AD8) , and Short Informant Questionnaire on Cognitive Decline in the Elderly (IQCODE) . Furthermore, prophylactic treatment with trazodone may also be recommended. Working Example
[0042] To illustrate effectiveness of the methods described above, blood samples were obtained from a group of patients with AD and a group of normal controls (NCs) , all with age over 60 years. The samples were divided into two sets: a training set (n = 422) including 214 AD patients and 208 NCs, and a validation set (n = 64) including 26 AD patients and 38 NCs. Each of the blood samples was analyzed using techniques described above to determine gene expression levels based on RNA sequences. In particular, gene expression levels were determined for each of the 24 genes listed in Table 1. A gene expression matrix was constructed from the gene expression levels for each of the 24 genes listed in Table 1. Gene expression matrices from the training set were used to train a generalized linear model. The trained model was then applied, in inference mode, to gene expression matrices for individuals in the validation set.
[0043] FIGs. 4A and 4B show graphs of auROC results for the trained model for the training set (FIG. 4A) and the validation set (FIG. 4B) in this example. As can be seen in the figures, the trained model exhibited relatively high accuracy in distinguishing AD patients from NCs in both the training (auROC = 0.920) and validation sets (auROC = 0.810) while using only a total of 24 genes as described above. Additional Embodiments
[0044] While the invention has been described with reference to specific embodiments, those skilled in the art will appreciate that variations and modifications are possible. For instance, the number of gene modules (M) and number of genes within each gene module (N) can be varied as desired. Further, while the genes listed in Table 1 have been observed to provide high accuracy, other combinations of genes may also be used. Various types of machine learning models can be used, including but not limited to generalized linear models, logistic regression models, or the like.
[0045] All processes described herein are also illustrative and can be modified. Operations can be performed in a different order from that described, to the extent that logic permits; operations described above may be omitted or combined; and operations not expressly described above may be added.
[0046] Data analysis and computational operations of the kind described herein can be implemented in computer systems that may be of generally conventional design, such as a desktop computer, laptop computer, tablet computer, mobile device (e.g., smart phone) , or the like. Such systems may include one or more processors to execute program code (e.g., general-purpose microprocessors usable as a central processing unit (CPU) and / or special-purpose processors such as graphics processors (GPUs) that may provide enhanced parallel-processing capability) ; memory and other storage devices to store program code and data; user input devices (e.g., keyboards, pointing devices such as a mouse or touchpad, microphones) ; user output devices (e.g., display devices, speakers, printers) ; combined input / output devices (e.g., touchscreen displays) ; signal input / output ports; network communication interfaces (e.g., wired network interfaces such as Ethernet interfaces and / or wireless network communication interfaces such as Wi-Fi) ; and so on.
[0047] Computer programs incorporating features of the present invention that can be implemented using program code may be encoded and stored on various computer readable storage media; suitable media include magnetic disk or tape, optical storage media such as compact disk (CD) or DVD (digital versatile disk) , flash memory, and other non-transitory media. (It is understood that “storage” of data is distinct from propagation of data using transitory media such as carrier waves. ) Computer readable media encoded with the program code may include an internal storage medium of a compatible electronic device and / or external storage media readable by the electronic device that can execute the code. In some instances, program code can be supplied to the electronic device via Internet download or other transmission paths.
[0048] Accordingly, although the invention has been described with respect to specific embodiments, it will be appreciated that the invention is intended to cover all modifications and equivalents within the scope of the following claims.
[0049] All patents, patent applications, and other publications, including GenBank Accession Numbers and equivalents, cited in this application are incorporated by reference in the entirety for all purposes.
Claims
1.A method for assessing risk for Alzheimer's Disease, the method comprising:(a) determining, from a blood sample from a subject, an expression level of each of the genes in Table 1;(b) computing a gene expression matrix representing the expression level of each of the genes in Table 1;(c) inputting the gene expression matrix to a machine learning model that has been trained to generate an Alzheimer's Disease risk score based on a gene expression matrix;(d) obtaining a predicted Alzheimer's Disease risk score from the machine learning model; and(e) determining that the subject has an increased risk for Alzheimer's disease based on the risk score.2.The method of claim 1 wherein the machine learning model is based on a generalized linear model.3.The method of claim 1 wherein the machine learning model is based on a logistic regression model.4.The method of claim 1 wherein the blood sample is a sample of all blood cells.5.The method of claim 1 wherein the expression level is an mRNA level.6.A kit for assessing Alzheimer's Disease risk, the kit comprising reagents for determining an expression level of each of the genes in Table 1.7.A method of manufacturing a kit for assessing Alzheimer’s Disease risk, the method comprising:including in the kit a plurality of reagents for detecting an expression level of each of the genes in Table 1.