Protein markers for neurodegenerative diseases
The use of protein markers like p-Tau217, CD33, FAM3B, KYNU, and NELL1 with logistic regression addresses the challenge of late-stage AD diagnoses by enabling early detection and intervention.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- THE HONG KONG UNIV OF SCI & TECH
- Filing Date
- 2025-09-01
- Publication Date
- 2026-05-15
AI Technical Summary
Current diagnostic methods for neurodegenerative diseases like Alzheimer's disease (AD) are ineffective for early detection, as they only alleviate symptoms and do not address the underlying pathology, and there is a lack of reliable tools for early diagnosis, leading to late-stage diagnoses.
A method using protein markers such as p-Tau217, CD33, FAM3B, KYNU, and NELL1, combined with logistic regression, to calculate individual risk scores for neurodegenerative diseases, allowing for early detection and intervention.
Enables accurate early diagnosis of neurodegenerative diseases by assessing individual risk scores, facilitating timely intervention and monitoring disease progression.
Smart Images

Figure CN2025118154_15052026_PF_FP_ABST
Abstract
Description
PROTEIN MARKERS FOR NEURODEGENERATIVE DISEASESRELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 692,297, filed September 9, 2024, the contents of which are hereby incorporated by reference in the entirety for all purposes.BACKGROUND OF THE INVENTION
[0002] Alzheimer’ disease (AD) is one of the most common forms of dementia in the world, accounting for 60-70%of all dementia cases. It is an irreversible degenerative brain disease and a leading cause of mortality among the elderly. The hallmarks of this disease are deposition of extracellular β-amyloid (Aβ) plaques and intracellular neurofibrillary tangles, which result in declining memory, reasoning, judgment, and locomotion abilities, with symptoms worsening over time.
[0003] Currently, an estimated 35 million people worldwide are afflicted with AD. This figure is expected to rise significantly to 100 million by 2050 due to longer life expectancies. There is no cure for AD; and the pathophysiology of the disease is still relatively unknown. There are only five drugs approved by the US Food and Drug Administration (FDA) to treat AD, but these only alleviate symptoms rather than alter disease pathology, as they cannot reverse the condition or prevent further deterioration, and are ineffective in severe conditions. Thus, early diagnosis and early therapeutic intervention is critical in the management of AD. Research has confirmed that AD affects the brain long before actual symptoms of memory loss or cognitive decline actually manifest. To this date, however, there are no effective and reliable diagnostic tools for early detection of AD; by the time a patient is diagnosed with AD using standard methods currently in use, which involves subjective clinical assessment, the pathological symptoms are already at an advanced stage. The present disclosure provides high performance diagnostic methods utilizing one or more protein markers for assessing the risk of neurodegenerative diseases including AD to aid early diagnosis. BRIEF SUMMARY OF THE INVENTION
[0004] Neurodegenerative diseases are caused by the progressive and often irreversible loss of neuron structures and functions, a process known as neurodegeneration and commonly seen within an older population. Due to the lack of cure for neurodegenerative diseases as well as the devastating social and economical impact of such diseases, there exists an urgent need for new, efficient, and effective methods for early diagnosis, which can then potentially support early intervention for these conditions. The present invention fulfills this and other related needs.
[0005] In a first aspect, the present invention provides a method for assessing risk for a neurodegenerative disease with brain amyloid pathology in an individual. The method includes these steps: first, calculating the individual’s risk score for the disease based on the concentrations of proteins p-Tau217 in combination with any one or more of CD33, FAM3B, KYNU, and NELL1 as measured in a biological sample taken from the individual, wherein each protein is assigned a weighted coefficient (βi) , and the risk score is computed using these coefficients along with an intercept (ε) derived from training data involving diagnosed patients who have been diagnosed of the disease and healthy controls who do not have the neurodegenerative disease with brain amyloid pathology. Second, the risk score from the first step is then compared against predefined thresholds to classify the individual’s risk as low, intermediate, or high. In some embodiments, the method includes these steps: measuring the concentrations of proteins in a biological sample from the individual, wherein the proteins are p-Tau217 in combination with any one or more of CD33, FAM3B, KYNU, and NELL1; assigning a weighted coefficient (βi) to each measured protein concentration; computing an individual risk score using the weighted coefficients (βi) and an intercept (ε) derived from training data involving diagnosed patients and healthy controls; and comparing the computed risk score against predefined thresholds to classify the individual's risk as low, intermediate, or high. Optionally, there is a step of obtaining the biological sample from the individual prior to the measuring step.
[0006] An exemplary method of the present invention for assessing the risk of a neurodegenerative disease with brain amyloid pathology in an individual includes these steps: (1) calculating an individual risk score by inputting a set of values into the formula and (2) determining the individual who has a risk score below 28 as having a low risk for the neurodegenerative disease with brain amyloid pathology, determining the individual who has a risk score above 45 as having a high risk for the neurodegenerative disease with brain amyloid pathology, and determining the individual who has a risk score between 28 and 45 as having an intermediate risk for the neurodegenerative disease with brain amyloid pathology, wherein the set of values comprises the plasma, serum, or whole blood level of p-Tau217 in combination with any one or more of CD33, FAM3B, KYNU, and NELL1 (e.g., all of CD33, FAM3B, KYNU, NELL1, and p-Tau217) , and the weighted coefficient (βi) and intercept (ε) for each protein are calculated from the plasma, serum, or whole blood level of p-Tau217 in combination with any one or more of CD33, FAM3B, KYNU, and NELL1 (e.g., all of CD33, FAM3B, KYNU, NELL1, and p-Tau217) in a group of subjects with a diagnosis of a neurodegenerative disease (diagnosis score = 1) or in a group of healthy control subjects without a neurodegenerative disease (diagnosis score = 0) using the logistic regression model:
[0007] In some embodiments, the neurodegenerative disease is Alzheimer’s Disease (AD) or mild cognitive impairment (MCI) . In some embodiments, the weighted coefficient (βi) and intercept (ε) for each protein are provided in Table 2.
[0008] In a second aspect, the present invention provides a method for assessing risk of brain amyloid pathology in an individual who has been diagnosed with a neurodegenerative disease. The method includes these steps: first, calculating the individual’s risk score for brain amyloid pathology based on the concentrations of proteins p-Tau217 in combination with any one or more of CD33, FAM3B, KYNU, and NELL1 as measured in a biological sample taken from the individual, wherein each protein is assigned a weighted coefficient (βi) , and the risk score is computed using these coefficients along with an intercept (ε) derived from training data involving diagnosed patients who have been diagnosed of the neurodegenerative disease with brain amyloid pathology and control subjects who have been diagnosed of the neurodegenerative disease but without brain amyloid pathology. Second, the risk score from the first step is then compared against predefined thresholds to classify the individual’s risk as low, intermediate, or high. In some embodiments, the method includes these steps: measuring the concentrations of proteins in a biological sample from the individual, wherein the proteins are p-Tau217 in combination with any one or more of CD33, FAM3B, KYNU, and NELL1; assigning a weighted coefficient (βi) to each measured protein concentration; computing an individual risk score using the weighted coefficients (βi) and an intercept (ε) derived from training data involving diagnosed patients who have been diagnosed of the neurodegenerative disease with brain amyloid pathology and control subjects who have been diagnosed of the neurodegenerative disease but without brain amyloid pathology; and comparing the computed risk score against predefined thresholds to classify the individual's risk as low, intermediate, or high. Optionally, there is a step of obtaining the biological sample from the individual prior to the measuring step.
[0009] An exemplary method of this invention for assessing the risk of brain amyloid pathology in an individual who has been diagnosed with a neurodegenerative disease including these steps: (1) calculating an individual risk score by inputting a set of values into the formula and (2) determining the individual who has a risk score below 28 as having a low risk for brain amyloid pathology, determining the individual who has a risk score above 45 as having a high risk for brain amyloid pathology, and determining the individual who has a risk score between 28 and 45 as having an intermediate risk for brain amyloid pathology, wherein the set of values comprises the plasma, serum, or whole blood level of p-Tau217 in combination with any one or more of CD33, FAM3B, KYNU, and NELL1 (e.g., all of CD33, FAM3B, KYNU, NELL1, and p-Tau217) , and the weighted coefficient (βi) and intercept (ε) for each protein are calculated from the plasma, serum, or whole blood level of p-Tau217 in combination with any one or more of CD33, FAM3B, KYNU, and NELL1 (e.g., all of CD33, FAM3B, KYNU, NELL1, and p-Tau217) in a group of subjects with a diagnosis of a neurodegenerative disease (diagnosis score = 1) or in a group of healthy control subjects without a neurodegenerative disease (diagnosis score = 0) using the logistic regression model:
[0010] In some embodiments, the neurodegenerative disease is Alzheimer’s Disease (AD) or mild cognitive impairment (MCI) . In some embodiments, the weighted coefficient (βi) and intercept (ε) for each protein are provided in Table 2.
[0011] In a third aspect, the present invention provides a method for assessing the relative risk for a neurodegenerative disease with brain amyloid pathology in two individuals. The method includes these steps: (1) calculating a risk score for each individual by inputting a set of values into the formula and (2) determining the first individual, whose a risk score is lower than the risk score of the second individual, as having a lower risk for the neurodegenerative disease with amyloid pathology in comparison to the second individual. In this method, the set of values comprises the plasma, serum, or whole blood level of p-Tau217 in combination with any one or more of CD33, FAM3B, KYNU, and NELL1 (e.g., all of CD33, FAM3B, KYNU, NELL1, and p-Tau217) , and the weighted coefficient (βi) and intercept (ε) for each protein are calculated from the plasma, serum, or whole blood level of p-Tau217 in combination with any one or more of CD33, FAM3B, KYNU, and NELL1 (e.g., all of CD33, FAM3B, KYNU, NELL1, and p-Tau217) in a group of subjects with a diagnosis of a neurodegenerative disease with amyloid pathology (diagnosis score = 1) or in a group of healthy control subjects who does not have any neurodegenerative disease with amyloid pathology (diagnosis score = 0) using the logistic regression model:
[0012] In some embodiments, the neurodegenerative disease is Alzheimer’s Disease (AD) or mild cognitive impairment (MCI) . In some embodiments, the weighted coefficient (βi) and intercept (ε) for each protein are provided in Table 2.
[0013] In a fourth aspect, the present invention provides a method for assessing or monitoring the status or progression of a neurodegenerative disease with brain amyloid pathology in an individual. The method includes these steps: (1) calculating a first risk score at an earlier time by inputting a set of values into the formula: (2) repeating step (1) at a later time to calculate a second risk score using the same formula; and (3) determining the individual whose second risk score (i.e., at a later time) is higher than the first risk score (i.e., at an earlier time) as having a worsening / progressive / deteriorating neurodegenerative disease with amyloid pathology, and determining the individual whose second risk score is no higher than the first risk score as having a stable or non-progressing neurodegenerative disease with amyloid pathology, wherein the set of values comprises the plasma, serum, or whole blood level of p-Tau217 in combination with any one or more of CD33, FAM3B, KYNU, and NELL1 (e.g., all of CD33, FAM3B, KYNU, NELL1, and p-Tau217) , and the weighted coefficient (βi) and intercept (ε) for each protein are determined from the plasma, serum, or whole blood level of p-Tau217 in combination with any one or more of CD33, FAM3B, KYNU, and NELL1 (e.g., all of CD33, FAM3B, KYNU, NELL1, and p-Tau217) in a group of subjects with a diagnosis of a neurodegenerative disease with brain amyloid pathology (diagnosis score = 1) or in a group of healthy control subjects without any neurodegenerative disease (diagnosis score = 0) using the following logistic regression model:
[0014] In some embodiments, the neurodegenerative disease is Alzheimer’s Disease (AD) or mild cognitive impairment (MCI) . In some embodiments, the weighted coefficient (βi) and intercept (ε) for each protein are provided in Table 2.
[0015] In some embodiments of any of the above described methods, when an individual is determined to have a high or intermediate risk for a neurodegenerative disease, is determined to have a relatively higher risk for a neurodegenerative disease in comparison to a second individual, or is determined to have a worsening / progressive / deteriorating neurodegenerative disease over a period of time (e.g., between two time points when two disease risk scores were calculated) , the individual may be treated for the neurodegenerative disorder, for example, by administration of anti-AD drugs including antibody drugs such as lecanemab, acetylcholinesterase inhibitors (such as donepezil, galantamine, rivastigmine) , memantine, glutamate receptor blockers, citalopram, fluoxetine, paroxeine, sertraline, trazodone, lorazepam, oxazepam, aripiprazole, clozapine, haloperidol, olanzapine, quetiapine, risperidone, ziprasidone, nortriptyline, tricyclic antidepressants, benzodiazepines, temazepam, zolpidem, zaleplon, chloral hydrate, coenzyme Q10, ubiquinone, coral calcium, Ginkgo biloba, huperzine A, omega-3 fatty acids, phosphatidylserine, or any combination thereof. For those at risk of MCI or developing MCI, certain remedial regimens can be administered, including regular physical exercise, adopting a diet low in fat and rich in fruits and vegetables, supplementation of omega-3 fatty acids, engagement in a mentally and socially active life style, up to professionally administered memory training and other cognitive training programs. In some embodiments of these methods, the neurodegenerative disease is Alzheimer’s Disease (AD) with brain amyloid pathology. In some embodiments, the neurodegenerative disease is mild cognitive impairment (MCI) with brain amyloid pathology.
[0016] In a fifth aspect, the present invention provides a kit for assessing presence or risk of neurodegenerative diseases such as mild cognitive impairment (MCI) and Alzheimer’s disease (AD) in patients, as well as for distinguishing among these diseases. Typically, such a kit comprises a series of containers, each containing a reagent for detecting the presence and quantity of a protein marker selected from a panel of protein markers newly identified by the present inventors and disclosed herein. Agents for the detection of at least two, optionally more, of such protein markers are included in the claimed kit. In some kits, these two or more protein markers are selected from CD33, FAM3B, KYNU, NELL1, and p-Tau217, for example, p-Tau217 in combination with any one or more of CD33, FAM3B, KYNU, and NELL1, including the combination of all of these 5 proteins. In some embodiments, the claimed kit furthur includes a users’ instruction manual to provide detailed description on how the protein markers should be analyzed in biological samples taken from a patient and the level of protein markers should be interpreted to indicate the patient’s status with regard to a potential diagnosis of a neurodegenerative disorder.
[0017] In the sixth aspect, the present invention provides a detection chip for assessing the presence or risk of a neurodegenerative disease (especially with brain amyloid pathology) , such as mild cognitive impairment (MCI) and Alzheimer’s disease (AD) , in patients, as well as for distinguishing among these diseases. Typically, the chip comprises a solid substrate and a reagent capable of determining the individual’s plasma, serum, or whole blood level of a protein marker, which may be among two or more proteins independently selected from CD33, FAM3B, KYNU, NELL1, and p-Tau217 (for example, p-Tau217 in combination with any one or more of CD33, FAM3B, KYNU, and NELL1, including the combination of all of these 5 proteins) , with each of the reagents immobilized at an addressable location on the substrate such that the presence of a particular marker protein in a plasma, serum, or whole blood sample take from an individual (such as one being tested for a potential neurodegenerative disease or being monitored for his disease status) can be immediately determined based on the location of the reagent on the solid substrate.
[0018] In the seventh aspect, the present invention provides machine learning systems and methods for assessing risk for a neurodegenerative disease with brain amyloid pathology in an individual. To train a machine learning model, a dataset is gathered that includes levels of certain proteins from a group of subjects. This group includes both healthy subjects and those already diagnosed with the disease. The machine learning model is used to predict the risk levels (e.g., risk scores or classified risks) for these people based on their protein levels. The predicted risk levels are compared to the actual known risk levels to see how accurate the predictions are. The model's internal parameters (weights) are adjusted to improve its accuracy based on the differences (losses) between predicted and actual risks. Once trained, an individual is assessed by obtaining the protein levels for the individual, providing the protein levels to the machine learning model as input, and generating an output risk level (e.g., risk score or classified risk) to predict this person's risk of having the neurodegenerative disease based on their protein levels.BRIEF DESCRIPTION OF THE DRAWINGS
[0019] FIG. 1 illustrates an example machine learning model functioning as a risk score generator used to generate a risk score based on a set of protein marker levels.
[0020] FIG. 2 illustrates an example machine learning model functioning as a risk classifier used to generate a classified risk based on a set of protein marker levels.
[0021] FIG. 3 illustrates an example system for training a machine learning model such as a risk score generator, which may be trained to generate a risk score based on a set of protein marker levels.
[0022] FIG. 4 illustrates an example system for training a machine learning model such as a risk classifier, which may be trained to generate a classified risk based on a set of protein marker levels.
[0023] FIG. 5 illustrates an example method for assessing risk for a neurodegenerative disease with brain amyloid pathology in an individual.
[0024] FIGS. 6A-6D illustrate a model integrating four blood proteins and blood p-Tau217 distinguished MCI and AD with brain amyloid pathology from healthy controls. FIG. 6A illustrates a boxplot showing the individual’s risk scores assigned by the model integrating four blood proteins (i.e., CD33, FAM3B, KYNU, and NELL1) and blood p-Tau217 in the HK Chinese cohort, stratified by diagnoses and brain amyloid pathology (n = 8 Aβ-HC, 9 Aβ+ MCI and 18 Aβ+ AD, respectively) . The dashed lines in orange (risk score = 28) and red (risk score = 45) represent the cut-off values of having low, intermediate or high risks of developing Aβ+ MCI or Aβ+ AD. *p <0.05, **p <0.01, ***p < 0.001. FIGS. 6B-6D illustrate receiver operating characteristic (ROC) curves showing the performance of the model integrating four blood proteins and blood p-Tau217 (red line) or the model solely using blood p-Tau217 (orange dash line) for distinguishing Aβ+ MCI (FIG. 6B) , or Aβ+ AD (FIG. 6C) or overall Aβ+ individuals (i.e., Aβ+ MCI and Aβ+ AD; d) from Aβ-HC individuals. Numbers in brackets indicate the area under the ROC (AUC) , which indicate the model's performance in the corresponding classification.
[0025] FIGS. 7A-7C illustrate a model integrating four blood proteins and blood p-Tau217 distinguished MCI with brain amyloid pathology from MCI without brain amyloid pathology. FIG. 7A illustrates a boxplot showing the MCI individual’s risk scores assigned by the models integrating four blood proteins (i.e., CD33, FAM3B, KYNU, and NELL1) and blood p-Tau217 in the HK Chinese cohort, stratified by diagnoses and brain amyloid pathology (n = 21 Aβ-MCI and 9 Aβ+ MCI, respectively) . The dashed lines in orange (risk score = 28) and red (risk score = 45) represent the cut-off values of having low, intermediate, or high risks of developing brain amyloid pathology for MCI individuals. *p <0.05, **p <0.01, ***p < 0.001. FIGS. 7B-7C illustrate receiver operating characteristic (ROC) curves showing the performance of the model integrating four blood proteins and blood p-Tau217 (solid line) or the model solely using blood p-Tau217 (dash line) for distinguishing Aβ+ MCI individuals from Aβ-MCI (FIG. 7B) , or overall Aβ-individuals (i.e., Aβ-HC and Aβ-MCI; FIG. 7C) . Numbers in brackets indicate the area under the ROC (AUC) , which indicate the model's performance in the corresponding classification.
[0026] FIG. 8 illustrates correlation between the model and the progression of brain amyloid pathology. Scatterplot showing the correlations between the individual’s risk scores assigned by the model integrating four blood proteins (i.e., CD33, FAM3B, KYNU, and NELL1) and blood p-Tau217 and the progression of brain amyloid pathology indicated by amyloid-PET in global cortical regions in the HK Chinese cohort (n = 56 individuals) . Data in regression lines are presented as the slop (red) and 95%confidence intervals (gray) . r2, Pearson's correlation coefficient; SUVR, standardized uptake value ratio. DEFINITIONS
[0027] “Polypeptide, ” “peptide, ” and “protein” are used interchangeably herein to refer to a polymer of amino acid residues. All three terms apply to amino acid polymers in which one or more amino acid residue is an artificial chemical mimetic of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers and non-naturally occurring amino acid polymers. As used herein, the terms encompass amino acid chains of any length, including full-length proteins, wherein the amino acid residues are linked by covalent peptide bonds.
[0028] In this disclosure the term "biological sample" or “sample” includes sections of tissues such as biopsy and autopsy samples, and frozen sections taken for histologic purposes, or processed forms of any of such samples. Biological samples include blood and blood fractions or products (e.g., whole blood, acellular fraction of blood (serum, plasma) , and blood cells) , sputum or saliva, lymph and tongue tissue, cultured cells, e.g., primary cultures, explants, and transformed cells, stool, urine, stomach biopsy tissue etc. A biological sample is typically obtained from a eukaryotic organism, which may be a mammal, may be a primate and may be a human subject.
[0029] The term “immunoglobulin” or “antibody” (used interchangeably herein) refers to an antigen-binding protein having a basic four-polypeptide chain structure consisting of two heavy and two light chains, said chains being stabilized, for example, by interchain disulfide bonds, which has the ability to specifically bind antigen. Both heavy and light chains are folded into domains.
[0030] The term “antibody” also refers to antigen-and epitope-binding fragments of antibodies, e.g., Fab fragments, that can be used in immunological affinity assays. There are a number of well characterized antibody fragments. Thus, for example, pepsin digests an antibody C-terminal to the disulfide linkages in the hinge region to produce F (ab) '2, a dimer of Fab which itself is a light chain joined to VH-CH1 by a disulfide bond. The F (ab) '2 can be reduced under mild conditions to break the disulfide linkage in the hinge region thereby converting the (Fab') 2 dimer into an Fab' monomer. The Fab' monomer is essentially a Fab with part of the hinge region (see, e.g., Fundamental Immunology, Paul, ed., Raven Press, N.Y. (1993) , for a more detailed description of other antibody fragments) . While various antibody fragments are defined in terms of the digestion of an intact antibody, one of skill will appreciate that fragments can be synthesized de novo either chemically or by utilizing recombinant DNA methodology. Thus, the term antibody also includes antibody fragments either produced by the modification of whole antibodies or synthesized using recombinant DNA methodologies.
[0031] The phrase "specifically binds, " when used in the context of describing a binding relationship of a particular molecule to a protein or peptide, refers to a binding reaction that is determinative of the presence of the protein in a heterogeneous population of proteins and other biologics. Thus, under designated binding assay conditions, the specified binding agent (e.g., an antibody) binds to a particular protein at least two times the background and does not substantially bind in a significant amount to other proteins present in the sample. Specific binding of an antibody under such conditions may require an antibody that is selected for its specificity for a particular protein or a protein but not its similar "sister" proteins. A variety of immunoassay formats may be used to select antibodies specifically immunoreactive with a particular protein or in a particular form. For example, solid-phase ELISA immunoassays are routinely used to select antibodies specifically immunoreactive with a protein (see, e.g., Harlow &Lane, Antibodies, A Laboratory Manual (1988) for a description of immunoassay formats and conditions that can be used to determine specific immunoreactivity) . Typically a specific or selective binding reaction will be at least twice background signal or noise and more typically more than 10 to 100 times background. On the other hand, the term “specifically bind” when used in the context of referring to a polynucleotide sequence forming a double-stranded complex with another polynucleotide sequence describes “polynucleotide hybridization” based on the Watson-Crick base-pairing, as provided in the definition for the term “polynucleotide hybridization method. ”
[0032] As used in this application, an "increase" or a "decrease" refers to a detectable positive or negative change in quantity from a comparison control, e.g., an established standard control (such as an average level / amount of a particular protein found in samples from healthy subjects who has not been diagnosed with, and has no increased risk for, a neurodegenerative disease such as AD or MCI, especially with brain amyloid pathology) . An increase is a positive change that is typically at least 10%, or at least 20%, or 50%, or 100%, and can be as high as at least 2-fold or at least 5-fold or even 10-fold of the control value. Similarly, a decrease is a negative change that is typically at least 10%, or at least 20%, 30%, or 50%, or even as high as at least 80%or 90%of the control value. Other terms indicating quantitative changes or differences from a comparative basis, such as "more, " "less, " "higher, " and "lower, " are used in this application in the same fashion as described above. In contrast, the term "substantially the same" or "substantially lack of change" indicates little to no change in quantity from the standard control value, typically within ± 10%of the standard control, or within ± 5%, 2%, or even less variation from the standard control.
[0033] A "label, " "detectable label, " or "detectable moiety" is a composition detectable by spectroscopic, photochemical, biochemical, immunochemical, chemical, or other physical means. For example, useful labels include 32P, fluorescent dyes, electron-dense reagents, enzymes (e.g., as commonly used in an ELISA) , biotin, digoxigenin, or haptens and proteins that can be made detectable, e.g., by incorporating a radioactive component into the protein or used to detect antibodies specifically reactive with the protein. Typically a detectable label is attached to a probe or a molecule with defined binding characteristics (e.g., an antibody with a known binding specificity to a polypeptide antigen) , so as to allow the presence of the probe (and therefore its binding target) to be readily detectable.
[0034] The term "amount" as used in this application refers to the quantity of a substance of interest, such as a polypeptide of interest, present in a sample. Such quantity may be expressed in the absolute terms, i.e., the total quantity of the substance in the sample, or in the relative terms, i.e., the concentration of the substance in the sample.
[0035] The term "subject" or "subject in need of treatment, " as used herein, includes individuals who seek medical attention due to risk of (e.g., with family history) , or having been diagnosed of, a neurodegenerative disease, such as AD or MCI, especially with brain amyloid pathology. Subjects also include individuals currently undergoing therapy that seek manipulation of the therapeutic regimen. Subjects or individuals in need of treatment include those that demonstrate symptoms of the neurodegenerative disease or are at risk of suffering from the neurodegenerative disease or its symptoms. For example, a subject in need of treatment includes individuals with a genetic predisposition or family history for AD, those that have suffered relevant symptoms in the past, those that have been exposed to a triggering substance or event, as well as those suffering from chronic or acute symptoms of the condition. A “subject in need of treatment” may be at any age of life.
[0036] “Inhibitors, ” “activators, ” and “modulators” of a target protein are used to refer to inhibitory, activating, or modulating molecules, respectively, identified using in vitro and in vivo assays for the protein binding or signaling, e.g., ligands, agonists, antagonists, and their homologs and mimetics. The term “modulator” includes inhibitors and activators. Inhibitors are agents that, e.g., partially or totally block, decrease, prevent, delay activation, inactivate, desensitize, or down regulate the activity of the target protein. In some cases, the inhibitor directly or indirectly binds to the protein, such as a neutralizing antibody. Inhibitors, as used herein, are synonymous with inactivators and antagonists. Activators are agents that, e.g., stimulate, increase, facilitate, enhance activation, sensitize or up regulate the activity of the target protein. Modulators include the target protein’s ligands or binding partners, including modifications of naturally-occurring ligands and synthetically-designed ligands, antibodies and antibody fragments, antagonists, agonists, small molecules including carbohydrate-containing molecules, siRNAs, RNA aptamers, and the like.
[0037] The term "treat" or "treating, " as used in this application, describes an act that leads to the elimination, reduction, alleviation, reversal, prevention and / or delay of onset or recurrence of any symptom of a predetermined medical condition. In other words, "treating" a condition encompasses both therapeutic and prophylactic intervention against the condition.
[0038] The term “effective amount, ” as used herein, refers to an amount that produces therapeutic effects for which a substance is administered. The effects include the prevention, correction, or inhibition of progression of the symptoms of a disease / condition and related complications to any detectable extent. The exact amount will depend on the purpose of the treatment, and will be ascertainable by one skilled in the art using known techniques (see, e.g., Lieberman, Pharmaceutical Dosage Forms (vols. 1-3, 1992) ; Lloyd, The Art, Science and Technology of Pharmaceutical Compounding (1999) ; and Pickar, Dosage Calculations (1999) ) .
[0039] The term "standard control, " as used herein, refers to a sample comprising an analyte of a predetermined amount to indicate the quantity or concentration of this analyte present in this type of sample (e.g., a predetermined DNA / mRNA or protein) taken from an average healthy subject not suffering from or at risk of developing a predetermined disease or condition (e.g., a neurodegenerative disease such AD or MCI, especially with brain amyloid pathology) . When used in the context of describing a value, this term may also be used to simply refer to the quantity or concentration of this analyte present in a “standard control” sample.
[0040] The term "average, " as used in the context of describing a healthy subject who does not suffer from and is not at risk of developing a relevant disease or disorders (e.g., AD) refers to certain characteristics, such as the level of a pertinent protein in the person's sample (e.g., serum or plasma or whole blood) , that are representative of a randomly selected group of healthy humans who are not suffering from and is not at risk of developing the disease or disorder. This selected group should comprise a sufficient number of human subjects such that the average amount or concentration of the analyte of interest among these individuals reflects, with reasonable accuracy, the corresponding profile in the general population of healthy people. Optionally, the selected group of subjects may be chosen to have a similar background to that of a person whose is tested for indication or risk of the relevant disease or disorder, for example, matching or comparable age, gender, ethnicity, and medical history, etc.
[0041] The term "inhibiting" or "inhibition, " as used herein, refers to any detectable negative effect on a target biological process or on the level of a biomarker (e.g., a protein) . Typically, an inhibition is reflected in a decrease of at least 10%, 20%, 30%, 40%, or 50%in one or more parameters indicative of the biological process or its downstream effect or the level of biomarker when compared to a control where no such inhibition is present. The term “enhancing” or “enhancement” is defined in a similar manner, except for indicating a positive effect, i.e., the positive change is at least 10%, 20%, 30%, 40%, 50%, 80%, 100%, 200%, 300%or even more in comparison with a control. The terms “inhibitor” and “enhancer” are used to describe an agent that exhibits inhibiting or enhancing effects as described above, respectively. Also used in a similar fashion in this disclosure are the terms “increase, ” “decrease, ” “more, ” and “less, ” which are meant to indicate positive changes in one or more predetermined parameters by at least 10%, 20%, 30%, 40%, 50%, 80%, 100%, 200%, 300%or even more, or negative changes of at least 10%, 20%, 30%, 40%, 50%, 80%or even more in one or more predetermined parameters.
[0042] As used herein, the term “brain amyloid pathology” refers to a neurological pathology characterized by aggregation of β-amyloid plaques in the brain. The β-amyloid (Aβ) peptide is a 4 kDa fragment of the amyloid precursor protein (APP) , a larger precursor molecule widely produced by brain neurons, vascular and blood cells (including platelets) . Two subsequent proteolytic cleavages of APP by β-secretase (β-APP-cleaving enzyme-1 (BACE1) ) at the ectodomain and γ-secretase at intra-membranous sites generate Aβ. Brain amyloid pathology can be measured by positron emission tomography (PET) tracers or indirectly by a reduction of the β-amyloid peptide in cerebrospinal fluid (CSF) .DETAILED DESCRIPTION OF THE INVENTIONI. INTRODUCTION
[0043] Neurodegenerative diseases are highly prevalent around the world, especially among the aging population in the industrialized nations. For instance, Alzheimer’s disease (AD) is the primary cause of dementia, affecting around 45 million individuals worldwide and is ranked as the fifth leading cause of death globally. In the United States alone, an estimated 6 million individuals live with AD dementia today, with this number expected to grow to 13.8 million by 2050. Similarly, in Western Europe, dementia affects about 2.5%of people aged 65–69 years, escalating to about 40%of those aged 90–94 years. By 2050, there will likely be up to 18.9 million patients with dementia in Europe and 36.5 million in East Asian countries.
[0044] To date, drugs approved for the treatment of AD are labeled for the disease’s clinical dementia stage and target the neurochemical systems underlying cognitive dysfunction and behavioral symptoms, with only short-term symptomatic effects. In the last 30 years, translational studies-including experimental animal and human neuropathological, genetic, and in vivo biomarker-based evidence-support a descriptive hypothetical model of AD pathophysiology characterized by the upstream brain accumulation of Aβ species and plaques, which precedes spreading of tau, neuronal loss and ultimately clinical manifestations by up to 20–30 years. The present invention relates to the use of a novel marker protein panel for the purpose of assessing individual patient’s risk of developing neurodegenerative diseases such AD or MCI, especially one with brain amyloid pathology, or assessing disease risk or disease status among multiple patients. High performance diagnostic methods are devised from this disclosure for effective early detection and early intervention for neurodegenerative disorders. II. QUANTITATION OF MARKER PROTEINS A. Obtaining Samples
[0045] The first step of practicing the present invention is to obtain a blood sample from a subject being tested for assessing the risk of developing a neurodegenerative disease such as AD or MCI or monitoring for disease severity or progression. Samples of the same type should be taken from both a control group (normal individuals not suffering from a neurodegenerative disease and without increased risk for such disease) and a test group (subjects being tested for possible neurodegenerative disease or for increased risk for the disease, for example) . Standard procedures routinely employed in hospitals or clinics are typically followed for this purpose.
[0046] For the purpose of detecting the presence / quantity of marker proteins or assessing the risk of developing a neurodegenerative disease (such as AD or MCI) , especially that with brain amyloid pathology, in test subjects, individual patients’ blood samples are taken, and the serum / plasma or whole blood level of pertinent marker proteins (e.g., one or more of CD33, FAM3B, KYNU, NELL1, and p-Tau217, such as p-Tau217 in combination with any one or more of CD33, FAM3B, KYNU, and NELL1, including the combination of all of these 5 proteins) may be measured and then each individual patient’s risk score for neurodegenerative disease (such as AD or MCI) with brain amyloid pathology is calculated by the formula described above and herein. Subsequently, these risk scores are compared to thresholds indicative of high, low, and intermediate risks such that the test subjects’ disease risk is determined individually.
[0047] In some cases, a same individual’s risk score may be calculated based on the marker protein levels measured at two or more different times. An increase in the risk score at a later time when compared with the earlier risk score is considered an indication of the increased risk for the neurodegenerative disease or the worsening or progression in the neurological disease, whereas a decrease in the risk score at a later time in comparison to an earlier score is considered an indication of reduced risk for the neurodegenerative disease or lessening or improvement with regard to the disease. For the purpose of monitoring disease progression or assessing therapeutic effectiveness in patients already diagnosed with a neurodegenerative disease such as AD or MCI, especially with brain amyloid pathology, individual patient’s blood samples are typically taken at different time points in order to calculate risk scores over the passage of time. A lack of substantial change in a patient’s disease risk scores would indicate a lack of change in the status of neurodegenerative disease and ineffectiveness of the therapy given to the patient.
[0048] For the practical application of this invention, the present inventors have devised novel calculation methods to produce a composite risk score based on multiple marker protein levels (e.g., two, three, four, or all five of CD33, FAM3B, KYNU, NELL1, and p-Tau217, such as p-Tau217 in combination with any one or more of CD33, FAM3B, KYNU, and NELL1, including the combination of all of these 5 proteins) to assess the risk of an individual to develop a neurodegenerative disease such as AD or MCI, especially with brain amyloid pathology, or to assess the relative disease risks or relative disease status among multiple (two or more) individuals. Possible marker combinations suitable for use in various applications of the present invention include but are not limited to: p-Tau217 with CD33, FAM3B, KYNU, NELL1 p-Tau217 with CD33, FAM3B, KYNU p-Tau217 with CD33, FAM3B, NELL1 p-Tau217 with CD33, FAM3B p-Tau217 with CD33, KYNU, NELL1 p-Tau217 with CD33, KYNU p-Tau217 with CD33, NELL1 p-Tau217 with CD33 p-Tau217 with FAM3B, KYNU, NELL1 p-Tau217 with FAM3B, KYNU p-Tau217 with FAM3B, NELL1 p-Tau217 with FAM3B p-Tau217 with KYNU, NELL1 p-Tau217 with KYNU p-Tau217 with NELL1 B. Preparing Samples for Protein Detection
[0049] The blood sample from a subject is suitable for the present invention and can be obtained by well-known methods and as described in standard medical literature. In certain applications of this invention, serum or plasma or whole blood may be the preferred sample type. In other cases, whole blood samples may be used.
[0050] A blood sample is obtained from a person to be tested or monitored for a neurodegenerative disease such as AD or MCI, especially one with brain amyloid pathology, using a method of the present invention. Collection of blood sample from an individual is performed in accordance with the standard protocol hospitals or clinics generally follow. An appropriate amount of blood is collected and may be stored according to standard procedures prior to further preparation.
[0051] The analysis of marker protein (s) found in a patient's sample according to the present invention may be performed using, e.g., serum or plasma or whole blood. The methods for preparing patient samples for protein extraction / quantitative detection are well known among those of skill in the art. C. Determining the Level of Marker Proteins
[0052] A protein of any particular identity, such as any one of more of CD33, FAM3B, KYNU, NELL1, and p-Tau217, can be detected using a variety of immunological assays. In some embodiments, a sandwich assay can be performed by capturing the protein from a test sample with an antibody having specific binding affinity for the protein. The protein then can be detected with a labeled antibody having specific binding affinity for it. Such immunological assays can be carried out using microfluidic devices such as microarray protein chips. A protein of interest (e.g., any one of CD33, FAM3B, KYNU, NELL1, and p-Tau217) can also be detected by gel electrophoresis (such as 2-dimensional gel electrophoresis) and western blot analysis using specific antibodies. Alternatively, standard immunohistochemical techniques can be used to detect a pre-selected protein (e.g., any one of CD33, FAM3B, KYNU, NELL1, and p-Tau217) , using the appropriate antibodies. Both monoclonal and polyclonal antibodies (including antibody fragment with desired binding specificity) can be used for specific detection of the polypeptide. Such antibodies and their binding fragments with specific binding affinity to a particular protein (e.g., any one or more of CD33, FAM3B, KYNU, NELL1, and p-Tau217) can be generated by known techniques.
[0053] Other methods may also be employed for measuring the level of marker protein (s) in practicing the present invention. For instance, a variety of methods have been developed based on the mass spectrometry technology to rapidly and accurately quantify target proteins even in a large number of samples. These methods involve highly sophisticated equipment such as the triple quadrupole (triple Q) instrument using the multiple reaction monitoring (MRM) technique, matrix assisted laser desorption / ionization time-of-flight tandem mass spectrometer (MALDI TOF / TOF) , an ion trap instrument using selective ion monitoring SIM) mode, and the electrospray ionization (ESI) based QTOP mass spectrometer. See, e.g., Pan et al., J Proteome Res. 2009 February; 8 (2) : 787–797. III. ESTABLISHING A STANDARD CONTROL
[0054] In order to establish a standard control for practicing the method of this invention, a group of healthy persons free of any neurodegenerative disease (such as AD or MCI, especially with brain amyloid pathology) or increased risk for developing such a disease as conventionally defined is first selected. These individuals are within the appropriate parameters, if applicable, for the purpose of screening for and / or monitoring neurodegenerative diseases using the methods of the present invention. Optionally, the individuals are of same gender, similar age, or similar ethnic background to the test subjects.
[0055] The healthy status of the selected individuals is confirmed by well-established, routinely employed methods including but not limited to general physical examination of the individuals and general review of their medical history.
[0056] Furthermore, the selected group of healthy individuals must be of a reasonable size, such that the average amount / concentration of marker protein (s) in the serum or plasma or whole blood sample obtained from the group can be reasonably regarded as representative of the normal or average level among the general population of healthy people without any neurodegenerative disease or increased risk for such disease. Preferably, the selected group comprises at least 10, 20, 30, or 50 human subjects.
[0057] Once an average value for the marker protein (s) is established based on the individual values found in each subject of the selected healthy control group, this average or median or representative value or profile is considered a standard control. A standard deviation is also determined during the same process. In some cases, separate standard controls may be established for separately defined groups having distinct characteristics such as age, gender, or ethnic background. IV. MONITORING AND TREATMENT
[0058] In a related aspect, the present invention also provides treatment methods for neurodegenerative diseases (such as AD or MCI, especially with brain amyloid pathology) patients upon detection of the disease or a heightened risk of later developing such a disease in a patient. In some embodiments, the method comprises, upon determining a subject as having an increased risk for a neurodegenerative disease (such as AD or MCI, especially with brain amyloid pathology) , administering a treatment to said subject, for example, an acetylcholinesterase inhibitor (such as donepezil, galantamine, rivastigmine) , memantine, a glutamate receptor blocker, citalopram, fluoxetine, paroxeine, sertraline, trazodone, lorazepam, oxazepam, aripiprazole, clozapine, haloperidol, olanzapine, quetiapine, risperidone, ziprasidone, nortriptyline, tricyclic antidepressants, benzodiazepines, temazepam, zolpidem, zaleplon, chloral hydrate, coenzyme Q10, ubiquinone, coral calcium, Ginkgo biloba, huperzine A, omega-3 fatty acids, phosphatidylserine, or any combination thereof.
[0059] In some cases, when the diagnostic method steps described above and herein are completed, optionally with additional diagnostic examination performed to provide further confirmatory information (for example, by brain imaging via CT scan or other imaging techniques to show excessive loss of brain volume, or by testing cognitive capability to show an accelerated decline) , and a patient has been determined to either already suffer from a neurodegenerative disease (such as AD or MCI, especially with brain amyloid pathology) or is at a significantly increased risk of later developing the disease, suitable therapeutic or prophylactic regimens may be ordered by physicians or other medical professionals to treat the patient, to manage / alleviate the ongoing symptoms, or to delay the future onset of the disease. The U.S. Food and Drug Administration (FDA) has approved a number of cholinesterase inhibitors, including donepezil (AriceptTM, the only cholinesterase inhibitor approved to treat all stages of AD, including moderate to severe) , rivastigmine (ExelonTM, approved to treat mild to moderate AD) , galantamine (RazadyneTM, mild to moderate patients) and memantine (NamendaTM) . Donepezil is the only cholinesterase inhibitor approved to treat all stages of AD, including moderate to severe. Any one or more of these drugs can be prescribed for treating patients who have been diagnosed with AD in accordance with the methods of this invention. Another possibility of treatment is administration of trazodone, which is currently approved for use as an antidepressant and has been reported as an effective agent for ameliorating AD symptoms.
[0060] For patients who are deemed at high or increased risk for developing a neurodegenerative disease (such as AD or MCI, especially with brain amyloid pathology) in a future time but do not yet exhibit any clinical symptoms, continuous monitoring is also appropriate, especially at an increased frequency. For example, the patients may be subject to more frequently scheduled regular testing (e.g., once every six months, once a year, or once every two years) to detect any accelerated change in their cognitive capabilities. Methods suitable for such regular monitoring include General Practitioner Assessment of Cognition (GPCOG) , Mini-Cog, Eight-item Informant Interview to Differentiate Aging and Dementia (AD8) , and Short Informant Questionnaire on Cognitive Decline in the Elderly (IQCODE) . Furthermore, prophylactic treatment with trazodone may also be recommended. V. KITS AND DEVICES
[0061] The invention provides compositions and kits for practicing the methods described herein to assess the pertinent marker protein level in a subject’s serum / plasma or whole blood, which can be used for various purposes such as detecting or diagnosing the presence of a neurodegenerative disease (such as AD or MCI, especially with brain amyloid pathology) , determining the risk of developing the disease, and monitoring progression of the disease in a patient, including assessing the therapeutic efficacy of a therapy administered for the disease among patients who have received a diagnosis of the disease and have undergone treatment.
[0062] Kits for carrying out assays for determining marker protein levels typically include at least one antibody useful for specific binding to the marker protein amino acid sequence. Optionally, this antibody is labeled with a detectable moiety. The antibody can be either a monoclonal antibody or a polyclonal antibody. In some cases, the kits may include at least two different antibodies, one for specific binding to a marker protein (i.e., the primary antibody) and the other for detection of the primary antibody (i.e., the secondary antibody) , which is often attached to a detectable moiety.
[0063] Typically, the kits also include an appropriate standard control. The standard controls indicate the average value of marker protein (s) in the serum or plasma or whole blood of healthy subjects not suffering from or at increased risk of developing any neurodegenerative disorder. In some cases, such standard control may be provided in the form of a set value. In addition, the kits of this invention may provide instruction manuals to guide users in analyzing test samples and assessing the presence or risk of neurodegenerative diseases (such as AD or MCI, especially with brain amyloid pathology) , or disease status / progression in a test subject.
[0064] In a further aspect, the present invention can also be embodied in a device or a system comprising one or more such devices, which is capable of carrying out all or some of the method steps described herein. For instance, in some cases, the device or system performs the following steps upon receiving a serum or plasma or whole blood sample taken from a subject being tested for detecting a neurodegenerative disease (such as AD or MCI, especially with brain amyloid pathology) , assessing the risk of developing such a disease, or assessing the disease status / progression: (a) determining in sample the amount or concentration of one or more marker proteins such as CD33, FAM3B, KYNU, NELL1, and p-Tau217; (b) calculating the disease risk score based on the formula of this invention; and (c) providing an output indicating whether the disease risk in the subject is high, low, or intermediate, or whether one patient has a higher risk of later developing the disease relative to another patient being tested; or whether a patient’s disease risk is increasing or decreasing or the patient’s disease status is improving or worsening over a time course. In other cases, the device or system of the invention performs the task of steps (b) and (c) , after step (a) has been performed and the amount or concentration from (a) has been entered into the device. Preferably, the device or system is partially or fully automated. VI. MACHINE LEARNING METHODS
[0065] The invention provides machine learning techniques for practicing the methods described herein to detect or diagnose the presence of neurogenerative diseases with brain amyloid pathology such as AD or MCI, determine the risk of developing the condition, and monitor progression of the condition in a patient, including assessing the therapeutic efficacy of a therapy administered for the condition among patients who have received a diagnosis of the disease and have undergone treatment.
[0066] FIG. 1 illustrates an example machine learning model functioning as a risk score generator 110 used to generate a risk score 112 based on a set (e.g., 3, 5, 10, 100, etc. ) of protein marker levels 102. Risk score generator 110 may be a neural network (e.g., feedforward neural network (FNN) , convolutional neural network (CNN) , recurrent neural network (RNN) , long short-term memory (LSTM) , etc. ) , a decision tree, a random forest, a support vector machine (SVM) , or an ensemble model. Risk score generator 110 may include a set of weights 108 and other parameters that may be adjusted during training to minimize the error in the model’s predictions. In some examples, weights 108 represent the strength of the connections between nodes 104 in the model and determine how input data is transformed as it passes through different layers 106 in the model. Layers 106 consist of groups of nodes 104 that operate at the same level in risk score generator 110.
[0067] During training or inference, protein marker levels 102 (e.g., comprising a set of values) for an individual are provided as input to risk score generator 110. In some examples, protein marker levels 102 are processed through a neural network layer by layer. Each node of each layer performs a weighted sum of its inputs based on weights 108, adds a bias, and passes the result through an activation function. This process continues until the final output layer produces a single output value as risk score 112. It is to be understood that while the illustrated example shows only two layers 106, risk score generator 110 may include any number of layers 106 (e.g., hundreds) , nodes 104 (e.g., thousands) , and weights 108 (e.g., millions or billions) . In general, each layer is associated with a set of weighted sum operations performed by a computer system, a set of bias additions performed by the computer system, and / or a set of activation function computations by the computer system.
[0068] The computer system used to execute risk score generator 110 (during training or inference) may include hardware elements and software elements to handle the computational demands. The computer system may include a central processing unit (CPU) having one or more cores (e.g., 8 or 16 cores) for handling parallel processing tasks. The computer system may include one or more graphics processing units (GPUs) or domain-specific processors such as neural network processors (NNPs) , neural processing units (NPUs) , artificial intelligence (AI) accelerators, or tensor processing units (TPUs) . The computer system may include memory and storage elements including volatile memory such as random-access memory (RAM) and non-volatile memory such as solid-state drive (SSD) memory to store the weights, the input data, the output data, as well as any intermediate data.
[0069] In some examples, risk score generator 110 may be used to assess the risk of an individual having a neurodegenerative disease with brain amyloid pathology by comparing risk score 112 to one or more thresholds, such as a lower threshold 114 and / or an upper threshold 116. If risk score 112 is less than lower threshold 114, it may be determined that the individual has a low risk for the neurodegenerative disease. If risk score 112 is greater than upper threshold 116, it may be determined that the individual has a high risk for the neurodegenerative disease. If risk score 112 is between lower threshold 114 and upper threshold 116, it may be determined that the individual has an intermediate risk for the neurodegenerative disease.
[0070] FIG. 2 illustrates an example machine learning model functioning as a risk classifier 240 used to generate a classified risk 218 based on a set (e.g., 3, 5, 10, 100, etc. ) of protein marker levels 202. Risk classifier 240 may be a neural network (e.g., FNN, CNN, RNN, LSTM, etc. ) , a decision tree, a random forest, an SVM, or an ensemble model. Risk classifier 240 may include a set of weights 208 and other parameters that may be adjusted during training to minimize the error in the model’s predictions. In some examples, weights 208 represent the strength of the connections between nodes 204 in the model and determine how input data is transformed as it passes through different layers 206 in the model. Layers 206 consist of groups of nodes 204 that operate at the same level in risk classifier 240.
[0071] During training or inference, protein marker levels 202 (e.g., comprising a set of values) for an individual are provided as input to risk classifier 240. In some examples, protein marker levels 202 are processed through a neural network layer by layer. Each node of each layer performs a weighted sum of its inputs based on weights 208, adds a bias, and passes the result through an activation function. This process continues until the final output layer produces a single output value or multiple (e.g., 3) output values representing possible classifications. For example, each of the output values may correspond to a particular classified risk (e.g., low, intermediate, or high risk) and the maximum output value may indicate the predicted classified risk (e.g., intermediate risk) . It is to be understood that while the illustrated example shows only two layers 206, risk classifier 240 may include any number of layers 206 (e.g., hundreds) , nodes 204 (e.g., thousands) , and weights 208 (e.g., millions or billions) . In general, each layer is associated with a set of weighted sum operations performed by a computer system, a set of bias additions performed by the computer system, and / or a set of activation function computations by the computer system.
[0072] The computer system used to execute risk classifier 240 (during training or inference) may include hardware elements and software elements to handle the computational demands. The computer system may include a CPU having one or more cores (e.g., 8 or 16 cores) for handling parallel processing tasks. The computer system may include one or more GPUs or domain-specific processors such as NNPs, AI accelerators, or TPUs. The computer system may include memory and storage elements including volatile memory such as RAM and non-volatile memory such as SSD memory to store the weights, the input data, the output data, as well as any intermediate data.
[0073] FIG. 3 illustrates an example system for training a machine learning model such as a risk score generator 310, which may be trained to generate a risk score 312 based on a set (e.g., 3, 5, 10, 100, etc. ) of protein marker levels 302. In some examples, a training dataset 320 may be obtained having protein marker levels 302 for a group of subjects with known risk scores (ground truth risk score 322) for a neurodegenerative disease with brain amyloid pathology. The group of subjects may include multiple healthy subjects (having low risk scores) and multiple subjects with a diagnosis of the neurodegenerative disease (having high risk scores) . As described in reference to FIG. 1, risk score generator 310 may be a neural network, a decision tree, a random forest, an SVM, or an ensemble model. Risk score generator 310 may include a set of weights 308 and other parameters, and a set of nodes 304 arranged in multiple layers 306.
[0074] During training, for each subject in training dataset 320, protein marker levels 302 and a ground truth risk score 322 for that subject may be retrieved from training dataset 320. Forward propagation is performed by providing protein marker levels 302 as input to risk score generator 310, which processes this input data using weights 308. For example, protein marker levels 302 may be processed layer by layer through a neural network, where each node of each layer performs a weighted sum of its inputs based on weights 308, adds a bias, and passes the result through an activation function. Such computations may be performed by a computer system such as that described in reference to FIG. 1. This process continues until the final output layer produces a single output value as risk score 312.
[0075] In some examples, training risk score generator 310 involves adjusting weights 308 of the model to minimize the error between the predicted output (risk score 312) and the actual target (ground truth risk score 322) . This process relies on a loss calculator 324 and a weight adjuster 328. In some examples, loss calculator 324 measures the discrepancy between risk score 312 and ground truth risk score 322 to calculate a loss 326. In some examples, loss calculator 324 may implement a loss function, which may include calculating a difference, a mean squared error (MSE) , a cross-entropy loss, among other possibilities.
[0076] Weight adjuster 328 uses loss 326 to make a set of weight adjustments 330 to weights 308, thereby updating weights 308 to minimize loss 326. In some examples, this process involves backpropagation and gradient descent. Backpropagation computes the gradients of the loss function with respect to each weight in the model. This may be achieved by applying the chain rule of calculus to propagate the error backward through the model. Once the gradients are computed, gradient descent is used to update the weights. The weights are adjusted in the direction that reduces the loss. The step size of weight adjustments 330 may be controlled by adjusting a learning rate. The process of forward propagation, loss calculation, backpropagation, and weight update may be repeated for multiple iterations (epochs) . During each epoch, the entire training dataset 320 is passed through risk score generator 310, and weights 308 are adjusted accordingly. This iterative process continues until loss 326 converges to a minimum value, indicating that risk score generator 310 has learned the underlying patterns in the data.
[0077] FIG. 4 illustrates an example system for training a machine learning model such as a risk classifier 440, which may be trained to generate a classified risk 418 based on a set (e.g., 3, 5, 10, 100, etc. ) of protein marker levels 402. In some examples, a training dataset 420 may be obtained having protein marker levels 402 for a group of subjects with known classified risks (ground truth classified risk 432) for a neurodegenerative disease with brain amyloid pathology. The group of subjects may include multiple healthy subjects (having low risk classifications) and multiple subjects with a diagnosis of the neurodegenerative disease (having high risk classifications) . As described in reference to FIG. 2, risk classifier 440 may be a neural network, a decision tree, a random forest, an SVM, or an ensemble model. Risk classifier 440 may include a set of weights 408 and other parameters, and a set of nodes 404 arranged in multiple layers 406.
[0078] During training, for each subject in training dataset 420, protein marker levels 402 and a ground truth classified risk 432 for that subject may be retrieved from training dataset 420. Forward propagation is performed by providing protein marker levels 402 as input to risk classifier 440, which processes this input data using weights 408. For example, protein marker levels 402 may be processed layer by layer through a neural network, where each node of each layer performs a weighted sum of its inputs based on weights 408, adds a bias, and passes the result through an activation function. Such computations may be performed by a computer system such as that described in reference to FIG. 2. This process continues until the final output layer produces a single output value as classified risk 418.
[0079] In some examples, training risk classifier 440 involves adjusting weights 408 of the model to minimize the error between the predicted output (classified risk 418) and the actual target (ground truth classified risk 432) . This process relies on a loss calculator 424 and a weight adjuster 428. In some examples, loss calculator 424 measures the discrepancy between classified risk 418 and ground truth classified risk 432 to calculate a loss 426. In some examples, loss calculator 424 may implement a loss function, which may include calculating a difference, an MSE, a cross-entropy loss, among other possibilities.
[0080] Weight adjuster 428 uses loss 426 to make a set of weight adjustments 430 to weights 408, thereby updating weights 408 to minimize loss 426. In some examples, this process involves backpropagation and gradient descent. Backpropagation computes the gradients of the loss function with respect to each weight in the model. This may be achieved by applying the chain rule of calculus to propagate the error backward through the network. Once the gradients are computed, gradient descent is used to update the weights. The weights are adjusted in the direction that reduces the loss. The step size of weight adjustments 430 may be controlled by adjusting a learning rate. The process of forward propagation, loss calculation, backpropagation, and weight update may be repeated for multiple iterations (epochs) . During each epoch, the entire training dataset 420 is passed through risk classifier 440, and weights 408 are adjusted accordingly. This iterative process continues until loss 426 converges to a minimum value, indicating that risk classifier 440 has learned the underlying patterns in the data.
[0081] FIG. 5 illustrates an example method 500 for assessing risk for a neurodegenerative disease with brain amyloid pathology in an individual. One or more steps of method 500 may be omitted during performance of method 500, and steps of method 500 may be performed in any order and / or in parallel. One or more steps of method 500 may be performed by one or more processors, such as those included in a computer system. Method 500 may be implemented as a computer-readable medium or computer program product comprising instructions which, when the program is executed by one or more computers, cause the one or more computers to carry out the steps of method 500. Such computer program products can be transmitted, over a wired or wireless network, in a data carrier signal carrying the computer program product.
[0082] At step 502, a machine learning model (e.g., risk score generators 110, 310 or risk classifiers 240, 440) having a set of weights (e.g., weights 108, 208, 308, 408) is trained. Step 502 may include obtaining a training dataset (e.g., training datasets 320, 420) comprising protein marker levels (e.g., protein marker levels 302, 402) for a group of subjects with known risk scores (e.g., ground truth risk score 322) or known classified risks (e.g., ground truth classified risk 432) for the neurodegenerative disease with brain amyloid pathology. The group of subjects may include healthy subjects and subjects with a diagnosis of the neurodegenerative disease with brain amyloid pathology. Step 502 may further include generating, by the machine learning model, a set of risk scores (e.g., risk score 312) or a set of classified risks (e.g., classified risk 418) for the group of subjects for the neurodegenerative disease with brain amyloid pathology by providing the protein marker levels to the machine learning model as inputs. Step 502 may further include comparing the set of risk scores or the set of classified risks for the group of subjects to the known risk scores or the known classified risks to calculate losses (e.g., losses 326, 426) . Step 502 may further include adjusting the set of weights of the machine learning model by a set of weight adjustments (e.g., weight adjustments 330, 430) based on the losses.
[0083] At step 504, a set of protein marker levels (e.g., protein marker levels 102, 202) for the individual are obtained.
[0084] At step 506, the machine learning model generates a risk score (e.g., risk score 112) or a classified risk (e.g., classified risk 218) for the individual for the neurodegenerative disease with brain amyloid pathology by providing the set of protein marker levels to the machine learning model as inputs.
[0085] At step 508, it is determined that the individual who has a risk score below a lower threshold (e.g., lower threshold 114) has a low risk for the neurodegenerative disease with brain amyloid pathology, the individual who has a risk score above an upper threshold (upper threshold 116) has a high risk for the neurodegenerative disease with brain amyloid pathology, and the individual who has a risk score between the lower threshold and the upper threshold has an intermediate risk for the neurodegenerative disease with brain amyloid pathology. EXAMPLES
[0086] The following examples are provided by way of illustration only and not by way of limitation. Those of skill in the art will readily recognize a variety of non-critical parameters that could be changed or modified to yield essentially the same or similar results. INTRODUCTION
[0087] Neurodegenerative diseases are devastating conditions of the brain that affect a large subset of the population. Many are highly debilitating, currently incurable, and often result in progressive deterioration of brain structure and cognitive function. For instance, Alzheimer’s disease (AD) constitutes 60 –80 %of dementia cases and represents a leading cause of death in the elderly. Characterized by progressive cognitive decline, AD is an age-related, progressive neurodegenerative disorder that is currently affecting 46.8 million people worldwide, ~10%of people aged 65 or above, with nearly 10 million new cases ever y year. The pathological hallmarks of this chronic disease include the accumulation of amyloid-beta plaques and neurofibrillary tangles in the brain, together with synaptic dysfunction and neuronal loss, which trigger inflammatory responses in the brain. The most common AD symptoms include memory problems, difficulty communicating, impaired reasoning and judgement, and reduced locomotor abilities. Like many neurodegenerative diseases and neuroinflammatory disorders, current challenges in AD diagnosis attribute to limited understanding of the disease pathophysiology. Meanwhile, currently available treatments are ineffective and can only provide transient effects on symptom alleviation, while patients still suffer severely from the diseases.
[0088] Characterized as a transitional state between normal cognition and dementia such as AD, the mild cognitive impairment (MCI) , particularly those with brain amyloid pathology (Aβ+ MCI) accounts for ~ 10%-20%in people over 65 years of age. Individuals with Aβ+MCI are cognitively impaired but not demented, while at a greater risk of developing AD or other dementia compared to those with normal cognition, with a cumulative probability of conversion to AD at 33 -50%. Whilst Aβ+ MCI is considered a symptomatic pre-dementia stage, its reversion to normal cognition is deemed possible. Hence, to treat Aβ+ MCI and prevent or delay the conversion of Aβ+ MCI to AD, the recognition of Aβ+ MCI is crucial. Specifically, the early diagnosis and timely intervention of AD and pre-AD MCI, for example by the recent approved anti-amyloid drugs, will advance the prevention and treatment of AD, which is expected to effectively delay the progression of AD and reduce the medical, socioeconomic and psychological burden on families and the society as a whole.
[0089] With the increase in global life expectancy, the incidence of Aβ+ MCI and AD is predicted to skyrocket in the coming decades. In particular, the global prevalence of AD is estimated at 75 million in 2030 and 131 million by 2050, while soaring in China where the largest elderly population resides. In fact, the number of AD cases in China doubled from 3.7 million to 9.2 million from 1990-2010, with a projected 22.5 million cases by 2050. Similarly, the Hong Kong population is also aging rapidly, with an estimated 100,000 people currently living with dementia, and ~ 333 000 or 11%of the population predicted to suffer from dementia in 2039. While increasing with age, the prevalence of MCI ranges from 9.74 –27.8%and 2.48 –35.5%for the Chinese and European populations, respectively. Despite the devastating impacts, MCI and AD remain largely underdiagnosed in primary care, with only 19%of patients with a confirmed dementia diagnosis as a result of routine medical care, while the proportion of diagnosis or medical consultation sought is estimated to be even lower in Hong Kong. As increasing studies indicate that the disease process of AD begins up to 20 years before the appearance of recognizable symptoms, such underdiagnoses thus attribute primarily to the challenges and limitations associated with the current diagnosis for dementia, which either relies on subjective assessment of apparent expressed symptoms, costly brain imaging or invasive sampling from the cerebral spinal fluid. This together highlights the importance of the time lag between disease onset and diagnosis for any potential intervention, thus the urgent need for early diagnosis, and the significance of novel biomarkers for the detection of Aβ+ MCI, the progression of Aβ+ MCI to AD, and AD. Moreover, it is imporatnt to differentiate Aβ+ MCI from those Aβ-MCI, which is esseential for selecting the suitable subjects for the treatment by the recent anti-amyloid drugs. Close monitoring the brain amyloid pathology and disease progression will also help better evaluate the drug response.
[0090] To address the current lacks in objective diagnostic tools for early detection and monitoring, we leveraged the recent technological advances in plasma proteomics approach, and identified novel protein biomarkers for assessing brain amyloid pathology in MCI and AD. Through large-scale analysis of the blood samples of Hong Kong Chinese participants that comprising individuals with mild cognitive impairment (MCI) individuals with AD and age-and sex-matched healthy and cognitively normal people, we identified 4 novel blood proteins (i.e., CD33, FAM3B, KYNU, and NELL1) , together with the tau pathology marker plasma p-Tau217 to form a protein panel, of which the level changes are associated with progression of brain amyloid pathology in individuals with MCI or AD. This simple, ultrasensitive, non-invasive and affordable blood-based technology for assessing brain amyloid pathology in MCI and AD will potentiate the early and accurate detection and monitoring of the disease.
[0091] The present invention addresses this and other related needs by disclosing novel methods and kits related to the use of plasma or serum or whole blood protein markers or their combinations, to assess brain amyloid pathology in MCI and AD. The invention relates to the discovery of novel blood protein markers associated with Aβ pathology in MCI and AD. The invention thus provides methods and compositions useful for assessing brain amyloid pathology in a subject with MCI or AD. It also provided as methods for differentiating MCI subjectives with or without brain amyloid pathology, and for indicating the progression of brain amyloid pathology in a subject with MCI or AD. As such, in a first aspect, the present invention provides a method for assessing a subject’s risk of developing Aβ+ MCI or Aβ+ AD at the current stage. The method includes the following steps: (1) measuring the plasma or serum or whole blood level of a panel of proteins (i.e., two or more proteins from protein panel that includes CD33, FAM3B, KYNU, and NELL1) . (2) measuring the plasma or serum or whole blood level of p-Tau217. (3) calculating the risk scores for individuals based on the measured protein levels and the corresponding prediction models; and (4) determining the subject as having increased risk for developing Aβ+ MCI or Aβ+ AD. In some embodiments, when the subject is developing MCI, the subject is then determined if having brain amyloid pathology (Aβ+) or not (Aβ-) . In some embodiments, when the subject is dermined in step (4) as having increased risk for developing Aβ+ MCI or Aβ+ AD, the subject is then provided follow-up monitoring of the progression of brain amyloid pathology. 5 Proteins -Summary:
[0092] 1. Protein CD33, also known as Siglec-3, is a transmembrane protein that is primarily expressed on the surface of myeloid cells, including monocytes, macrophages, and dendritic cells. It belongs to the family of sialic acid-binding immunoglobulin-like lectins (Siglecs) and plays a role in regulating immune responses. CD33 binds to sialic acid residues on glycoproteins and glycolipids, which modulates cell signaling and adhesion. CD33 is involved in various physiological processes, including phagocytosis, antigen presentation, and cytokine production. Dysregulation of CD33 expression has been associated with various diseases, including Alzheimer's disease, acute myeloid leukemia, and autoimmune disorders. CD33 has been proposed as a potential therapeutic target for these diseases due to its involvement in regulating immune function and cell signaling.
[0093] 2. Protein FAM3B, also known as FAM3 Metabolism Regulating Signaling Molecule B, is a member of the FAM3 family and is also known by its other name, PANDER (PANcreatic DERived factor) . FAM3B is primarily expressed in the pancreas islets, intestines and salivary gland, and is involved in the regulation of glucose metabolism. It acts as a signaling molecule influencing insulin secretion from beta cells and has been shown to have effects on liver glucose output. The role of FAM3B in metabolic regulation suggests that it may have potential implications in the development of type 2 diabetes mellitus.
[0094] 3. Protein KYNU, also known as kynureninase, is an enzyme that is involved in the metabolism of tryptophan, an essential amino acid. KYNU catalyzes the conversion of kynurenine to anthranilic acid in the kynurenine pathway, which is the major pathway for tryptophan metabolism in humans. This pathway plays a crucial role in regulating immune function, inflammation, and neurotransmitter synthesis in the brain. Dysregulation of KYNU activity has been implicated in the pathogenesis of various diseases, including autoimmune disorders, neurodegenerative diseases, and cancer. KYNU has also been proposed as a potential therapeutic target for these diseases due to its involvement in modulating immune responses and inflammation.
[0095] 4. Protein NELL1, also known as Neural Epidermal Growth Factor-Like 1, is a secreted osteoinductive protein that is involved in cell signaling and tissue development. It is especially important in osteogenesis, where it promotes bone growth and mineralization. NELL1 is expressed in various tissues but is most notably involved in the development and regeneration of bone and cartilage. Its role in the nervous system is also being explored, with studies suggesting involvement in neural development.
[0096] 5. Phosphorylated Tau at threonine 217 (pTau-217) is a post-translational modification of the Tau protein that has gained prominence as a biomarker for AD and MCI due to its high diagnostic and prognostic capabilities. The hyperphosphorylation of Tau protein, including at threonine 217, leads to the formation of neurofibrillary tangles, which are a pathological hallmark of AD. The accumulation of pTau-217 is strongly correlated with the progression of AD pathology and cognitive decline, making it a potential biomarker for early diagnosis and monitoring of disease progression.
[0097] The effectiveness of pTau-217 as a biomarker is highlighted by its ability to identify Aβ deposition, another key pathology of AD. Elevated levels of pTau-217 in blood plasma are strongly associated with both amyloid plaques and neurofibrillary tangles. This association facilitates the early detection, diagnosis, and monitoring of AD, allowing for the differentiation of AD from other neurodegenerative disorders with accuracy comparable to that of cerebrospinal fluid (CSF) biomarkers and positron emission tomography (PET) imaging.
[0098] While pTau-217 demonstrates high specificity and sensitivity in general AD and MCI populations, its effectiveness in differentiating Aβ status specifically within the MCI subgroup has shown some inconsistencies. In certain studies, pTau-217 alone may not reliably distinguish between Aβ-positive and Aβ-negative individuals among MCI patients. However, the overall predictive value of pTau-217 for progression from MCI to AD dementia remains robust. This underscores the utility of pTau-217 in longitudinal monitoring and potentially enhances patient outcomes through early intervention and tailored treatment approaches, especially in the context of emerging AD therapies targeting amyloid pathology. MATERIALS AND METHODS
[0099] Participant recruitment for the Hong Kong Chinese cohort: A total of 56 Hong Kong Chinese individuals aged 60 years or over, comprising of 18 individuals with Alzheimer’s disease (AD) and brain amyloid pathology (Aβ+) , 9 Aβ+ individuals with mild cognitive impairment (MCI) , and 8 cognitively normal healthy controls (HC) without brain amyloid pathology (Aβ-) , 21 Aβ-individuals with MCI, who visited the Neurology Department of the Prince of Wales Hospital of the Chinese University of Hong Kong were recruited. Participants were clinically diagnosed with AD or MCI based on the American Psychiatric Association’s Diagnostic and Statistical Manual of Mental Disorders, Fifth Edition1. All participants underwent medical history assessment, clinical assessment, cognitive and functional assessment via the Montreal Cognitive Assessment (MoCA) 2, and neuroimaging assessment by amyloid positron emission tomography (PET) imaging using 11C-Pittsburgh Compound B (PiB) . Participants with global cortical-to-cerebellum standardized uptake value ratio ≥1.4 were defined as amyloid PET positive. Individuals with neurological diseases other than AD or psychiatric diseases were excluded from enrollment. The age, sex, body mass index (BMI) , years of education, and medical history of each participant were recorded. This study was approved by the Prince of Wales Hospital of the Chinese University of Hong Kong and the Hong Kong University of Science and Technology. All participants provided written informed consent for study participation and sample collection.
[0100] Plasma preparation from blood samples: Whole blood (3 mL) was collected in K3EDTA tubes (VACUETTE) and centrifuged at 2,000g for 15 minutes to separate the cell pellet and plasma. The plasma was collected, aliquoted, and stored at -80℃ until use.
[0101] Measurement of blood protein levels: Blood abundance of p-Tau217 was quantified in prepared plasma samples, using the Quanterix ALZpath pTau-217 CARe Advantage Kit on a Quanterix HD-X Automated Immunoassay analyzer. Blood abundance of 4 proteins (i.e., CD33, FAM3B, KYNU, and NELL1) were quantified in prepared plasma samples, via Proximity Extension Assay technology of Olink Proteomics biomarker panels, including Immune Response, Neuro Exploratory, Neurology, and Oncology III.
[0102] Calculation of risk scores: For the model, the weighted coefficient (βi) of candidate proteins and intercept (ε) were calculated by fitting the blood levels of candidate proteins (in log form) and AD diagnosis into the following logistic regression model:
[0103] Individual risk scores were calculated using the blood levels of candidate proteins (in log form) and corresponding weighted coefficient (βi) and intercept (ε) using the following linear model:
[0104] Evaluation of prediction accuracy: The auc () function from the R pROC package was used to evaluate the accuracy of each prediction model by calculating the areas under the curve (AUCs) of receiver operating characteristic (ROC) curves.
[0105] Data visualization: The investigators performing blood protein measurements were blinded to the diagnosis and phenotypes of participants. All statistical plots were generated using Prism v8.0 (GraphPad) . Example I: A model integrating four blood proteins and blood p-Tau217 distinguishes MCI and AD with brain amyloid pathology from healthy controls
[0106] We developed an algorithm to integrate the level changes of candidate blood proteins with blood p-Tau217, which can improve its performance in classification of AD and brain amyloid pathology from healthy controls. We showed that adding either one or multiple blood proteins from the four candidate blood proteins can achieve better performance than solely using blood p-Tau217 (Table 1) . In particular, a model integrating the levels of a blood protein panel comprising all four proteins (i.e., CD33, FAM3B, KYNU, and NELL1) and blood p-Tau217 has high performance in assessing brain amyloid pathology in MCI and AD (Table 2) . The model generated risk scores that can significantly distinguish MCI with brain amyloid pathology (Aβ+ MCI) and AD with brain amyloid pathology (Aβ+ AD) from healthy controls (Aβ-HC) (FIG. 6A) . Notably, the accuracies of this model in distinguishing Aβ+ MCI or Aβ+ AD or overall Aβ+ individuals (Aβ+ MCI and Aβ+ AD) from Aβ-HC are 95.22%, 100%, and 98.07%, respectively, which are significantly higher than that of solely using blood p-Tau217 (all p < 0.05; FIGS. 6B-6D) . Individuals with risk scores lower than 28 will have low risk of developing Aβ+ MCI or Aβ+ AD; by comparison, individuals with risk scores in the range of 28 to 45 will have intermediate risks of developing Aβ+ MCI or Aβ+ AD, and individuals with risk scores larger than 45 will have high risks of developing Aβ+ MCI or Aβ+ AD. Example II: A model integrating four blood proteins and blood p-Tau217 distinguishes MCI with brain amyloid pathology from MCI without brain amyloid pathology
[0107] It is important to identify Aβ+ MCI from MCI without brain amyloid pathology (Aβ-MCI) , who are at the early stage of AD and can be effectively treated by AD drugs such as anti-amyloid therapy. Notably, our model can accurately distinguish Aβ+ MCI from Aβ-MCI in the MCI population, with an accuracy of 93.12% (FIG. 7A) , which is significantly higher than that of solely using blood p-Tau217 (p < 0.05; FIG. 7B) . Moreover, the accuracies of the model in distinguishing Aβ+ MCI from overall Aβ-individuals (Aβ-HC and Aβ-MCI) is 94.25%, which is significantly higher than that of solely using blood p-Tau217 (p < 0.05; FIG. 7C) . MCI individuals with risk scores lower than 28 will have low risk of developing brain amyloid pathology; by comparison, MCI individuals with risk scores in the range of 28 to 45 will have intermediate risks of developing brain amyloid pathology, and MCI individuals with risk scores larger than 45 will have high risks of developing brain amyloid pathology. Example III: A model integrating four blood proteins and blood p-Tau217 indicates the progression of brain amyloid pathology
[0108] Furthermore, we showed that the risk scores generated by the model are significantly correlated with the progression of brain amyloid pathology (p < 0.0001, r2 = 0.6284) , as assessed by amyloid-PET (FIG. 8) . Again, this correlation is stronger than that of solely using blood p-Tau217 in indicating the progression of brain amyloid pathology (p < 0.0001, r2 = 0.5554) . These results together demonstrated that this model integrating four blood proteins and blood p-Tau217 has superior performance than solely using blood p-Tau217 in assessing brain amyloid pathology, which can be utilized to monitor the AD progression and provide a real-time assessment of disease status. Table 1. Performance of the models integrating different combinations of blood proteins in classifying AD and brain amyloid positivity. Table 2. Weighted coefficients (βi) and intercept (ε) for the model integrating four blood proteins and blood p-Tau217 REFERENCES 1. American Psychiatric Association. "American Psychiatric Association: Diagnostic and Statistical Manual of Mental Disorders, Arlington. " (2013) . 2. Nasreddine, Ziad S., et al. "The Montreal Cognitive Assessment, MoCA: a brief screening tool for mild cognitive impairment. " Journal of the American Geriatrics Society 53.4 (2005) : 695-699.
[0109] All patents, patent applications, and other publications, including GenBank Accession Numbers and the like, cited in this application are incorporated by reference in the entirety for all purposes.
Claims
1.A method for assessing risk for a neurodegenerative disease with amyloid pathology in an individual, comprising:calculating an individual risk score based on the concentrations of p-Tau217 plus any one or more of CD33, FAM3B, KYNU, and NELL1 in a biological sample from the individual, wherein each protein is assigned a weighted coefficient (βi) , and the risk score is computed using these coefficients along with an intercept (ε) derived from training data involving patients who have been diagnosed of the neurodegenerative disease and healthy controls who do not have the disease; andcomparing the risk score against predefined thresholds to classify the individual’s risk as low, intermediate, or high.2.The method of claim 1, further comprising, prior to the calculating step, obtaining the biological sample from the individual and measuring concentration of each of the proteins in the sample.3.The method of claim 1, comprising:(1) calculating an individual risk score by inputting a set of values into the formula(2) determining the individual who has a risk score below 28 as having a low risk for the neurodegenerative disease with brain amyloid pathology, determining the individual who has a risk score above 45 value as having a high risk for the neurodegenerative disease with brain amyloid pathology, and determining the individual who has a risk score between 28 and 45 as having an intermediate risk for the neurodegenerative disease with brain amyloid pathology,wherein the set of values comprises the plasma, serum, or whole blood level of p-Tau217 plus any one or more of CD33, FAM3B, KYNU, and NELL1,and the weighted coefficient (βi) and intercept (ε) for each protein are obtained from the plasma, serum, or whole blood level of each protein in a group of subjects with a diagnosis of a neurodegenerative disease or in a group of healthy control subjects using the following logistic regression model (neurodegenerative disease = 1; healthy control = 0) :4.The method of claim 1, wherein the neurodegenerative disease is Alzheimer’s Disease (AD) or mild cognitive impairment (MCI) .5.The method of claim 1, wherein the weighted coefficient (βi) and intercept (ε) for each protein are set forth in Table 2.6.A method for assessing risk of brain amyloid pathology in an individual who has been diagnosed with a neurodegenerative disease, comprising:calculating an individual risk score for brain amyloid pathology based on the concentrations of p-Tau217 plus any one or more of CD33, FAM3B, KYNU, and NELL1 in a biological sample from the individual, wherein each protein is assigned a weighted coefficient (βi) , and the risk score is computed using these coefficients along with an intercept (ε) derived from training data involving diagnosed patients who have been diagnosed of the neurodegenerative disease with brain amyloid pathology and control subjects who have been diagnosed of the neurodegenerative disease but without brain amyloid pathology; andcomparing the risk score against predefined thresholds to classify the individual’s risk for brain amyloid pathology as low, intermediate, or high.7.The method of claim 6, further comprising, prior to the calculating step, obtaining the biological sample from the individual and measuring concentration of each of the proteins in the sample.8.The method of claim 6, comprising:(1) calculating an individual risk score by inputting a set of values into the formula(2) determining the individual who has a risk score below 28 as having a low risk for brain amyloid pathology, determining the individual who has a risk score above 45 as having a high risk for brain amyloid pathology, and determining the individual who has a risk score between 28 and 45 as having an intermediate risk for brain amyloid pathology,wherein the set of values comprises the plasma, serum, or whole blood level of p-Tau217 plus any one or more of CD33, FAM3B, KYNU, and NELL1,and the weighted coefficient (βi) and intercept (ε) for each protein are obtained from the plasma, serum, or whole blood level of each protein in a group of subjects with a diagnosis of a neurodegenerative disease or in a group of healthy control subjects using the following logistic regression model (neurodegenerative disease = 1; healthy control = 0) :9.The method of claim 6, wherein the neurodegenerative disease is Alzheimer’s Disease (AD) or mild cognitive impairment (MCI) .10.The method of claim 6, wherein the weighted coefficient (βi) and intercept (ε) for each protein are set forth in Table 2.11.A method for assessing risk for a neurodegenerative disease with brain amyloid pathology in two individuals, comprising:(1) calculating a risk score for each individual by inputting a set of values into the formula(2) determining the first individual whose a risk score is lower than that of the second individual as having a lower risk for the neurodegenerative disease with amyloid pathology,wherein the set of values comprises the plasma, serum, or whole blood level of p-Tau217 plus any one or more of CD33, FAM3B, KYNU, and NELL1,and the weighted coefficient (βi) and intercept (ε) for each protein are obtained from the plasma, serum, or whole blood level of each protein in a group of subjects with a diagnosis of a neurodegenerative disease or in a group of healthy control subjects using the following logistic regression model (neurodegenerative disease = 1; healthy control = 0) :12.The method of claim 11, wherein the neurodegenerative disease is Alzheimer’s Disease (AD) or mild cognitive impairment (MCI) .13.The method of claim 11, wherein the weighted coefficient (βi) and intercept (ε) for each protein are set forth in Table 2.14.A method for assessing progression of a neurodegenerative disease with brain amyloid pathology in an individual, comprising:(1) calculating a first risk score at an earlier time by inputting a set of values into the formula(2) repeating step (1) at a later time to calculate a second risk score; and(3) determining the individual whose second risk score is higher than the first risk score as having a worsening neurodegenerative disease with amyloid pathology, and determining the individual whose second risk score is no higher than the first risk score as having a stable or non-progressing neurodegenerative disease with amyloid pathology,wherein the set of values comprises the plasma, serum, or whole blood level of p-Tau217 plus any one or more of CD33, FAM3B, KYNU, and NELL1,and the weighted coefficient (βi) and intercept (ε) for each protein are determined from the plasma, serum, or whole blood level of each protein in a group of subjects with a diagnosis of a neurodegenerative disease or in a group of healthy control subjects using the following logistic regression model (neurodegenerative disease = 1; healthy control = 0) :15.The method of claim 14, wherein the neurodegenerative disease is Alzheimer’s Disease (AD) or mild cognitive impairment (MCI) .16.The method of claim 14, wherein the weighted coefficient (βi) and intercept (ε) for each protein are set forth in Table 2.17.A kit for assessing risk for a neurodegenerative disease with brain amyloid pathology in an individual, comprising multiple containers each containing a reagent capable of determining the individual’s plasma, serum, or whole blood level of p-Tau217 plus any one or more of CD33, FAM3B, KYNU, and NELL1.18.The kit of claim 17, wherein the neurodegenerative disease is Alzheimer’s Disease (AD) or mild cognitive impairment (MCI) .19.A detection chip for risk for a neurodegenerative disease with brain amyloid pathology in an individual, comprising a solid substrate and a reagent capable of determining the individual’s plasma, serum, or whole blood level of p-Tau217 plus any one or more of CD33, FAM3B, KYNU, and NELL1, wherein each reagent is immobilized at an addressable location on the substrate.20.The chip of claim 19, wherein the neurodegenerative disease is Alzheimer’s Disease (AD) or mild cognitive impairment (MCI) .21.A computer-implemented method for assessing risk for a neurodegenerative disease with brain amyloid pathology in an individual, comprising:training a machine learning model having a set of weights by:obtaining a training dataset comprising protein marker levels for a group of subjects with known risk scores for the neurodegenerative disease with brain amyloid pathology, wherein the group of subjects includes healthy subjects and subjects with a diagnosis of the neurodegenerative disease with brain amyloid pathology;generating, by the machine learning model, a set of risk scores for the group of subjects for the neurodegenerative disease with brain amyloid pathology by providing the protein marker levels to the machine learning model as inputs;comparing the set of risk scores for the group of subjects to the known risk scores to calculate losses for the set of risk scores; andadjusting the set of weights of the machine learning model based on the losses;obtaining a set of protein marker levels for the individual; andgenerating, by the machine learning model being executed on a computer system, a risk score for the individual for the neurodegenerative disease with brain amyloid pathology by providing the set of protein marker levels to the machine learning model as inputs.22.The computer-implemented method of claim 21, further comprising:determining the individual who has a risk score below a lower threshold as having a low risk for the neurodegenerative disease with brain amyloid pathology, determining the individual who has a risk score above an upper threshold as having a high risk for the neurodegenerative disease with brain amyloid pathology, and determining the individual who has a risk score between the lower threshold and the upper threshold as having an intermediate risk for the neurodegenerative disease with brain amyloid pathology.23.The computer-implemented method of claim 21, wherein the protein marker levels for the group of subjects and the set of protein marker levels for the individual comprise the plasma, serum, or whole blood level of p-Tau217 plus any one or more of CD33, FAM3B, KYNU, and NELL1.24.The computer-implemented method of claim 21, wherein the machine learning model comprises a neural network, a decision tree, a random forest, a support vector machine (SVM) , or an ensemble model.25.The computer-implemented method of claim 21, wherein the neurodegenerative disease is Alzheimer’s Disease (AD) or mild cognitive impairment (MCI) .26.A computer-implemented method for assessing risk for a neurodegenerative disease with brain amyloid pathology in an individual, comprising:training a machine learning model having a set of weights by:obtaining a training dataset comprising protein marker levels for a group of subjects with known classified risks for the neurodegenerative disease with brain amyloid pathology, wherein the group of subjects includes healthy subjects and subjects with a diagnosis of the neurodegenerative disease with brain amyloid pathology, wherein the protein marker levels for the group of subjects comprise the plasma, serum, or whole blood level of p-Tau217 plus any one or more of CD33, FAM3B, KYNU, and NELL1;generating, by the machine learning model, a set of classified risks for the group of subjects for the neurodegenerative disease with brain amyloid pathology by providing the protein marker levels to the machine learning model as inputs;comparing the set of classified risks for the group of subjects to the known classified risks to calculate losses for the set of classified risks; andadjusting the set of weights of the machine learning model based on the losses;obtaining a set of protein marker levels for the individual, wherein the protein marker levels for the individual comprise the plasma, serum, or whole blood level of p-Tau217 plus any one or more of CD33, FAM3B, KYNU, and NELL1; andgenerating, by the machine learning model being executed on a computer system, a classified risk for the individual for the neurodegenerative disease with brain amyloid pathology by providing the set of protein marker levels to the machine learning model as inputs.27.The computer-implemented method of claim 26, further comprising:determining the individual who has a risk score below a lower threshold as having a low risk for the neurodegenerative disease with brain amyloid pathology, determining the individual who has a risk score above an upper threshold as having a high risk for the neurodegenerative disease with brain amyloid pathology, and determining the individual who has a risk score between the lower threshold and the upper threshold as having an intermediate risk for the neurodegenerative disease with brain amyloid pathology.28.The computer-implemented method of claim 26, wherein the machine learning model comprises a neural network, a decision tree, a random forest, a support vector machine (SVM) , a transformer, or an ensemble model.29.The computer-implemented method of claim 26, wherein the neurodegenerative disease is Alzheimer’sDisease (AD) or mild cognitive impairment (MCI) .