How to determine your risk of developing Alzheimer's disease dementia
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-03-10
- Publication Date
- 2026-03-16
AI Technical Summary
The prior art is difficult to accurately and economically identify and predict the risks of Alzheimer's disease (AD), especially in the early stages of dementia or mild cognitive impairment, and the commonly used diagnostic methods are highly invasive, costly and inadequately accurate.
By running a classification model, the model is able to process biomarker screening data (such as mitochondrial methylation data) and other relevant clinical data (such as MMSE and SOB) in blood samples to calculate and determine risk scores for Alzheimer's disease dementia (ADD) and classify subjects at different risk levels. This method utilizes newly discovered mitochondrial methylation sites, especially the CHH sites of the ND1 gene, in combination with efficient methylation detection techniques.
A rapid, non-invasive, effective method to determine the risk of ADD is achieved, improving the accuracy and reliability of the diagnosis, able to effectively identify the risk of AD at an early stage, and guiding subject screening and treatment selection in clinical trials.
Smart Images

Figure 00000073_0000 
Figure 00000073_0001 
Figure 00000073_0002
Abstract
Description
[Technical field]
[0001] The present invention relates to the field of medicine and to the field of diagnosing or determining the risk of developing a neurodegenerative disease, in particular to a method for diagnosing or determining the risk of developing Alzheimer's dementia. [Background technology]
[0002] Alzheimer's disease (AD) is a neurodegenerative disorder estimated to affect approximately 40 million people worldwide, resulting in substantial social and economic burdens. Currently, there is no cure or treatment that can slow the progression of AD, and only four drugs are available for symptomatic treatment. Furthermore, clinical trials evaluating treatments for AD have faced unprecedented failures of over 99.6% since 2002, despite numerous trials aiming to find novel treatments for AD. This lack of success is, in part, a result of recruiting late-stage AD patients who suffer primarily from irreversible neurodegeneration. Thus, there is a need to shift the design of clinical trials to include predementia patients with early predementia, i.e., mild cognitive impairment (MCI).
[0003] However, there are new challenges in identifying patients at risk of developing AD. Patients diagnosed with MCI are likely to actually develop AD, but only about 1 / 3 will progress to AD. Thus, current diagnostic methods, such as PET-amyloid imaging, PET-FDG, lumbar puncture, magnetic resonance imaging, or behavioral testing (e.g., Clinical Dementia Rating Scale or Mini-Mental State Examination), while improving recruitment efficiency, still fail to optimally determine a patient's risk of developing AD. For example, about 1 / 3 of patients diagnosed with MCI recruited by a positive PET test will revert to a cognitively normal state or develop other types of dementia. Moreover, most of the methods are highly invasive, incurring significant financial costs to patients and healthcare systems. Thus, it seems clear that available diagnostic methods do not meet the accuracy required for clinical trials to successfully evaluate potential AD treatments, nor do they meet other desirable characteristics, such as non-invasiveness.
[0004] As a result, optimal stratification of patients with high risk of developing AD during recruitment for clinical trials remains a major challenge.Therefore, there is a strong unmet need to develop accurate, reliable, yet feasible and cost-effective screening tools that can identify subjects with high risk of developing AD, allowing accurate recruitment and ultimately the development of therapeutic drugs for AD.In addition, the above-mentioned desirable screening approach also allows differentiation between AD patients and patients suffering from other dementias, and achieves accurate diagnosis, when available, which facilitates the selection of the most appropriate treatment option.
[0005] The pathophysiology of Alzheimer's disease is associated with alterations in mitochondrial functionality and mitochondrial DNA (mtDNA), such as inherited somatic mutations that are usually located in mtDNA regulatory elements. Consistently, there is substantial evidence of defects in oxidative phosphorylation (OXPHOS) in AD, in which polypeptides encoded by mtDNA play a key role. Thus, AD diagnostic methods have been described based on the identification of mutations in mitochondrial DNA via restriction fragment length polymorphism (RFLP) technology or other related techniques.
[0006] WO 2015 / 144964 discloses the use of mitochondrial methylation patterns to diagnose or determine the risk of developing neurodegenerative diseases, such as AD and Parkinson's disease, based on the analysis of postmortem subject brain samples. The methylation patterns disclosed in WO 2015 / 144964 come from a very early research stage. Thus, there is still a need for feasible, effective and non-invasive methods that allow the determination of said risk at the early stage of dementia in patients or even at the healthy stage of subjects. The data disclosed in WO 2015 / 144964 A2 are also discussed in Blanch et al. (2016), who reach the same conclusion and emphasize that methylated mtDNA represents only a small portion of total mtDNA.
[0007] The mtDNA methylation patterns in AD patients have been further studied by Stoccoro et al. (2017), which disclosed reduced levels of D-loop methylation in the peripheral blood of late-onset AD patients compared to controls. These are surprisingly different results compared to earlier studies. Finally, an abstract by Stoccoro et al. (2022) disclosed that patients diagnosed with MCI exhibited higher methylation levels in the D-loop region than controls and AD patients.
[0008] Although there are evidences that suggest that mtDNA methylation can provide information about the Alzheimer's dementia stage of subjects, its role and pattern in AD, as well as its significance, are far from being clearly and clearly defined.Therefore, its role and usefulness for determining the risk of developing Alzheimer's disease or diagnosing the disease at an early stage remains unclear.Therefore, there is a great need to accurately classify subjects so as to be able to determine the risk of developing AD and ultimately ensure improved recruitment efficiency in clinical trials and appropriate treatment. Summary of the Invention
[0009] One problem solved by the present invention is to provide a method for diagnosing or determining the risk that a subject will develop Alzheimer's Disease Dementia (herein referred to as ADD).
[0010] The present invention discloses a method that can calculate or determine a score to quantify the risk that a subject develops ADD, and then classify the subject according to said risk.The method includes implementing a classification model that can process more than one data set, including biomarker screening data (i.e., mitochondrial methylation data) and other related clinical data (e.g., MMSE and SOB).The biomarker screening data is obtained from blood samples, thus enabling a rapid, non-invasive and effective method for determining the risk.
[0011] The use of mitochondrial markers to diagnose AD has already been described in International Patent Publication No. 2015 / 144964. However, the examples in International Patent Publication No. 2015 / 144964 disclose mitochondrial methylation patterns obtained through the analysis of a small number of brain samples (obtained post-mortem) of subjects known to suffer from AD (N=16) as well as controls (N=8). These examples analyze the methylation of a total of 89 sites that have been identified as differentially methylated by statistical methods. These sites include CpG, CHG, and CHH in the D-loop region, as well as CpG and CHG sites in the ND1 gene.
[0012] Surprisingly, the inventors found that the methylation sites that contribute most to determining the risk of developing ADD have not been disclosed in the prior art. As shown in the present invention, some novel methylation sites show very important methylation patterns for determining such risk, which mostly correspond to CHH sites of ND1 gene. Furthermore, the inventors of the present invention have developed a set of primers that are much more efficient for detecting methylation in mtDNA extracted from blood samples than commonly designed primers that mainly focus on CpG sites (see Example 1). Furthermore, the present embodiment collects information from blood samples instead of brain samples, and includes a larger number of samples for developing the method disclosed herein.
[0013] Stoccoro et al. (2017 & 2022) et al. disclose that patients diagnosed with MCI show high levels of methylation in the mitochondrial D-loop region, while in AD patients, said methylation levels are reduced. The abstract of Stoccoro et al. (2022) does not distinguish between subjects with early dementia (MCI) and does not disclose any information on how said pattern may be useful in classifying subjects according to their risk of developing ADD. Moreover, again, Stoccoro et al. (2017 & 2022) does not refer to any site in the ND1 gene, but only to a portion of the D-loop region disclosed herein.
[0014] The examples herein provide detailed experimental data that demonstrate the efficient processing of blood samples for the detection and calculation of mtDNA methylation.Furthermore, the information on mitochondrial methylation is shown to have very high performance when combined with other relevant clinical data and processed together by a classification model.As a result, the method provided herein determines a score that corresponds to the risk of developing ADD, thereby clearly classifying subjects.
[0015] Example 1 shows a method for detecting mtDNA methylation, which includes collecting blood samples, extracting and treating (bisulfite treatment) DNA, and preparing an amplicon library to detect, quantify and normalize the methylation of mtDNA sites of interest. The use of degenerate primers resulted in exceptionally high sensitivity in detecting mtDNA methylation for both regions (i.e., D-loop region and ND1 gene) and in all three contexts (i.e., CpG, CHG, CHH) to the point that these results exceed any possible expectation. Furthermore, comparison of methylation levels between groups for different contexts and regions resulted in a large number of significant, differentially methylated comparisons.
[0016] Example 2 shows the development of a classification model that considers not only the data on the methylation site of interest, but also other related clinical data (e.g. MMSE, SOB). The model assigns a specific weight to each variable by statistical methods according to the training data input. Thus, the model can calculate the score corresponding to the risk of developing ADD, with a remarkably high performance as shown by the overall accuracy score of 0.76 and the kappa value of 0.63. Thus, the classification model can calculate the risk of developing ADDD of any subject by rapidly processing their individual information (i.e. mitochondrial methylation and clinical data) with a remarkably good performance.
[0017] It is noteworthy that the method disclosed in Example 2 does not include any clinical variables that can be obtained through invasive or costly techniques such as PET.In the current clinical field, PET used to detect amyloid plaque is actually considered to be a highly informative diagnostic technique for AD.However, the classification model developed herein does not use such information, but still can clearly predict the risk of developing ADD with a highly efficient indicator.
[0018] Furthermore, Example 3 shows the development of a classification model that takes into account the data on the methylation sites of interest shown in Example 2 and other relevant clinical data. In this case, the clinical data included the above-mentioned PET variables (positive or negative) used to detect amyloid plaques. This classification model was performed on subjects who had already undergone PET in order to take advantage of this additional information (positive or negative) on amyloid PET testing. This model is able to calculate a score that corresponds to the risk of developing ADD, thus again classifying patients with a significantly high performance, as shown by the overall accuracy score of 0.89 and the kappa value of 0.84.
[0019] Example 4 shows the development of a classification model that considers data on methylation sites of interest and related clinical data on the above-mentioned PET variables (positive or negative) used to detect amyloid plaque.This model can calculate a score that corresponds to the risk of developing ADD, and thus classifies patients with a significantly high performance again, as shown by the total accuracy score of 0.756 and the kappa value of 0.63.Thus, this classification model illustrates the possibility of developing a high-performance classification model that considers a small number of clinical variables, but significantly contributes to the determination of a reliable score.
[0020] Thus, a first aspect of the present invention relates to a method of determining the risk of developing Alzheimer's disease dementia in a subject, comprising the step of applying a classification model to a methylation pattern of the D-loop region of mitochondrial DNA and / or the ND1 gene of a sample from the subject, wherein the classification model assigns a Dementia Stage Class (DSC) to the subject selected from progression to ADD or non-progression to ADD.
[0021] A second aspect of the invention relates to a method of identifying a subject suitable for treatment with a particular AD therapy, comprising the step of applying a classification model to a methylation pattern of the D-loop region of mitochondrial DNA and / or the ND1 gene of a sample from the subject, wherein the classification model assigns to the subject a dementia stage class selected from progression to ADD or not progression to ADD, wherein a dementia stage class consisting of progression to ADD indicates that a particular AD therapy may be administered to the subject.
[0022] A third aspect of the invention relates to a method of treating a subject, particularly a subject diagnosed with MCI or CDR0.5, comprising the step of administering to the subject a specific AD therapy or other treatment for dementia, wherein prior to administration the subject is assigned to a DSC, determined by applying a classification model to the methylation pattern of the D-loop region of mitochondrial DNA and / or the ND1 gene of a sample from the subject, the DSC being selected from progression to ADD or non-progression to ADD.
[0023] A fourth aspect of the invention relates to a method of providing an individualized therapy to a subject at high risk of developing ADD, comprising the steps described herein.
[0024] A fifth aspect of the present invention relates to a classification model for determining a subject's risk of developing ADD, which uses data relating to methylation patterns from a sample of the subject and data relating to the subject's clinical variables to identify the subject as belonging to a class selected from the group consisting of progressing to ADD and not progressing to ADD, wherein identification as progressing to ADD indicates that the subject is at risk of developing ADD.
[0025] Another embodiment is a computer-implemented method applicable to the methods described herein, for example for obtaining a risk score for developing ADD, comprising the following steps: (a) Input data: (1) the methylation pattern of the D-loop region of the mitochondrial DNA and / or the ND1 gene of the subject, and optionally (2) at least one clinical variable of the subject described herein Providing or receiving (b) combining and weighting the methylation pattern(s) and the clinical variable(s) using a classification model to obtain a risk score; The present invention relates to a computer-implemented method,
[0026] In some examples, input data may be received that allows a computer or other data processing system to derive the methylation pattern of the D-loop region of the mitochondrial DNA and / or the methylation pattern of the ND1 gene of the subject.Thus, the computer-implemented method may further include receiving at least one clinical variable of the subject as described herein, and using a classification model to combine and weight the methylation pattern(s) and the clinical variable(s) to obtain a risk score.
[0027] One aspect of the present invention relates to the use of an oligonucleotide having a length of 15 to 100 nucleotides, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, and SEQ ID NO:4.
[0028] One aspect of the present invention relates to the use of an oligonucleotide of 15 to 100 nucleotides in length comprising a nucleic acid sequence selected from the group consisting of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, and SEQ ID NO:4 for determining the methylation pattern of mitochondrial DNA.
[0029] In one aspect, the present invention relates to a kit comprising at least one oligonucleotide capable of specifically hybridizing to mitochondrial DNA comprising the D-loop region or the ND1 gene.
[0030] In one aspect, the present invention relates to the use of the kit as defined above for determining the methylation pattern of mitochondrial DNA. In another aspect, the present invention relates to the use of the kit as defined above for determining the methylation pattern of mitochondrial DNA to determine the risk of developing Alzheimer's disease dementia in a subject. In another aspect, the present invention relates to the use of the kit as defined above according to the method described herein.
[0031] Throughout the specification and claims, the term "comprise" and variations thereof are not intended to exclude other technical features, additives, ingredients, or steps. Additional objects, advantages, and properties of the present invention will become apparent to those skilled in the art upon examination of the description or may be learned by practice of the invention. Moreover, the present invention encompasses all possible combinations of the specific and preferred embodiments described herein. The following examples and figures are provided herein for illustrative purposes and are not intended to limit the present invention. [Brief description of the drawings]
[0032] [Figure 1] Figure 1 shows the detection of methylation of mtDNA in the D-loop region in a CpG context using degenerate (Deg) and non-degenerate (NoDeg) primers. Figure 1 shows five subjects (e.g., FIS004) as examples with triplicate samples shown on the horizontal axis. [Diagram 2] Figure 2 shows the detection of methylation of mtDNA in the D-loop region in the context of CHG using degenerate primers (Deg) and non-degenerate primers (NoDeg). Figure 2 shows five subjects (e.g., FIS004) as examples with triplicate samples shown on the horizontal axis. [Diagram 3] Figure 3 shows the detection of methylation of mtDNA in the D-loop region in the context of CHH using degenerate primers (Deg) and non-degenerate primers (NoDeg). Figure 3 shows five subjects (e.g., FIS004) as examples with triplicate samples indicated on the horizontal axis. [Figure 4] Figure 4 shows the detection of methylation of mtDNA in the ND1 gene in a CpG context using degenerate (Deg) and non-degenerate (NoDeg) primers. Figure 4 shows five subjects (e.g., FIS004) as examples with triplicate samples shown on the horizontal axis. [Diagram 5] Figure 5 shows the detection of methylation of mtDNA in the ND1 gene in the context of CHG using degenerate (Deg) and non-degenerate (NoDeg) primers. Figure 5 shows five subjects (e.g., FIS004) as examples with triplicate samples shown on the horizontal axis. [Figure 6] Figure 6 shows the detection of methylation of mtDNA in the ND1 gene in the context of CHH using degenerate (Deg) and non-degenerate primers (NoDeg). Figure 6 shows five subjects (e.g., FIS004) as examples with triplicate samples shown on the horizontal axis. [Figure 7] Figure 7 is a bar graph of the variables dementia stage classification, gender, and beta amyloid PET. "DS" refers to dementia stage, "N" refers to "number," "C" refers to "control," "NP" refers to "non-progressive," "P" refers to "progressive," "F" refers to "female," "M" refers to "male," and "S" refers to "gender." [Figure 8]Figure 8 shows a violin diagram for the variables age, MMSE, and SOB. "DS" refers to dementia stage, "A" refers to "age", "C" refers to "control", "NP" refers to "non-progressive", and "P" refers to "progressive". [Figure 9] FIG. 9 shows the accuracy and kappa metrics of the supervised learning model of Example 2. [Figure 10] 1 shows the ROC curve for comparing progressor vs. non-progressor groups using a model constructed using the random forest method of Example 2. [Figure 11] Figure 11 shows a bar graph of PET for beta amyloid. "N" refers to "number", "Neg" refers to negative, and "Pos" refers to "positive". [Figure 12] FIG. 12 shows the accuracy and kappa metrics of the supervised learning model used in Example 3. [Figure 13] FIG. 13 shows the ROC curve for comparing progressor vs. non-progressor groups using the model constructed using the random forest method of Example 3. [Figure 14] FIG. 14 shows the accuracy and kappa metrics of the supervised learning model used in Example 4. [Figure 15] FIG. 15 shows the ROC curve for comparing progressor vs. non-progressor groups using the model constructed using the random forest method of Example 4. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0033] For the avoidance of doubt, the methods provided herein do not include diagnostics performed on the human or animal body. The methods of the invention are particularly performed on samples previously extracted from a subject. The kits provided herein may include means for extracting a sample from a subject.
[0034] definition Diagnosis: The term "diagnosis" refers to both the process of attempting to determine and / or identify a possible disease in a subject, i.e., the diagnostic procedure, as well as the opinion reached through this process, i.e., the diagnostic opinion. It may therefore also be seen as an attempt to classify an individual's condition into separate and distinct categories that allow medical decisions about treatment and prognosis to be made. As the skilled artisan will appreciate, such a diagnosis may not be accurate for 100% of subjects diagnosed, although it is preferred that such a diagnosis is accurate for 100% of subjects diagnosed. However, the term requires that a statistically significant portion of a plurality of subjects may be identified as having Alzheimer's disease or a predisposition thereto in the context of the present invention. The skilled artisan may use different well-known statistical evaluation tools to determine whether a portion is statistically significant, for example, by determining a confidence interval, a value of p-value, Student's t-test, Mann-Whitney test, etc. A particular confidence interval may be at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95%. In particular, the P value is 0.05, 0.025, 0.001, or less.
[0035] Risk of developing Alzheimer's disease dementia: The term "risk of developing Alzheimer's disease dementia (ADD)" is used herein interchangeably with "risk of progressing to Alzheimer's disease dementia" and refers to the predisposition, susceptibility, or tendency of a subject to develop ADD. The risk of developing ADD implies that there is a high or low risk or a higher or lower risk in general. Thus, a subject with a high risk of developing ADD has at least 50%, or at least 60%, or at least 70%, or at least 80%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99%, or at least 100% likelihood of developing this dementia. Similarly, a subject with a low risk of developing ADD is a subject with at least one chance of developing dementia of up to 1%, or up to 2%, or up to 3%, or up to 5%, or up to 10%, or up to 20%, or up to 30%, or up to 40%, or up to 49%.
[0036] Thus, "determining the risk of developing Alzheimer's disease dementia" refers to the probability of progressing to Alzheimer's disease dementia, including all possible grades of dementia within the disease.
[0037] The terms "progression to ADD" or "progressed ADD" or "ADD progression" or any other similar expressions are used interchangeably with "developing ADD" or "developed ADD" or "ADD development" or "development of ADD".
[0038] Alzheimer's Disease: The term "Alzheimer's disease" or "senile dementia" or AD refers to mental impairment associated with a specific degenerative brain disease characterized by the appearance of senile plaques, neuritic tangles, and progressive neuronal loss that is clinically manifested in progressive memory loss, confusion, behavioral problems, inability to care for oneself, gradual physical deterioration, and ultimately death. Alzheimer's disease can progress through the following stages according to the Braak staging system: Stage I-II: The brain area is affected by the presence of neurofibrillary tangles that correspond to the entorhinal region of the brain. Stage III-IV: The affected brain regions extend to areas of the limbic system, such as the hippocampus. Stages V-VI: Affected brain areas include the neocortex can be classified into:
[0039] This neuropathological stage classification correlates with the clinically progressive changes of the existing disease, with parallels between memory decline and the formation of neurofibrillary tangles and neuritic plaques in the entorhinal cortex and hippocampus (stages I-IV). Furthermore, their presence in the neocortex (stages V and VI) correlates with clinically severe changes. The entorhinal status (I-II) corresponds to a clinically silent period of the disease. The limbic status (III-IV) corresponds to clinically incident AD. The neocortical status corresponds to fully developed AD.
[0040] Alzheimer's Disease Dementia: The term "Alzheimer's Disease Dementia" or "ADD" refers to a set of symptoms, including memory loss and difficulties with thinking, problem-solving, or language, that develop as a result of the degenerative brain damage and progressive neuronal loss characteristic of Alzheimer's disease.
[0041] In clinical terms, subjects at risk of developing / progressing to Alzheimer's disease in the future are said to be at risk of developing / progressing to Alzheimer's disease dementia (ADD) because this is a set of symptoms that may occur or that a subject at risk may have. Thus, the present invention is directed to methods for determining the risk of developing / progressing to ADD, methods for identifying subjects at risk of developing / progressing to ADD, and other methods related to the risk of developing / progressing to ADD.
[0042] Subject: The terms "subject," "patient," "individual," and variants thereof, are used interchangeably herein and refer to any mammalian subject, particularly a human subject. The terms do not denote a particular age or sex.
[0043] Sample containing mitochondrial DNA: The expression "sample containing mitochondrial DNA", as used herein, refers to any sample that can be obtained from a subject in which mitochondrial genetic material suitable for detection of methylation patterns is present.
[0044] Mitochondrial DNA: The term "mitochondrial DNA" or "mtDNA" as used herein refers to the genetic material located in the mitochondria of an organism. It is a closed circular double-stranded molecule. In humans, it consists of 16,569 base pairs and contains a small number of genes distributed between the H and L chains. Mitochondrial DNA codes for 37 genes: two ribosomal DNAs, 22 transfer RNAs, and 13 proteins involved in oxidative phosphorylation.
[0045] Methylation pattern and methylation state: The term "methylation pattern" as used herein refers to, but is not limited to, the presence or absence of methylation of one or more nucleotides, particularly the methylation of cytosine. Thus, one or more nucleotides are contained in a single nucleic acid molecule. One or more nucleotides can be methylated or not methylated. The term "methylation state" can also be used when considering only a single nucleotide. Methylation pattern can be quantified when considering more than one nucleic acid molecule.
[0046] D-loop region: The term "D-loop region" as used herein refers to a region of non-coding mtDNA that acts as a promoter for both the heavy and light chains of mDNA and contains essential transcription and replication elements. The D-loop region contains approximately 1120 base pairs, can be visualized under an electron microscope, and is generated during H-chain replication for the synthesis of the short segment 7S DNA of the heavy chain. The sequence of the human D-loop region has been deposited in the GenBank database under the accession number NC_012920.1.
[0047] ND1 gene: The term "ND1 gene" or "NADH dehydrogenase 1", or "ND1mt", as used herein, refers to a gene located in the mitochondrial genome that encodes the protein NADH dehydrogenase 1 or ND1. The human ND1 gene sequence has been deposited in the GenBank database under the accession number NC_012920.1. The ND1 protein is part of an enzyme complex called complex I, which is active in mitochondria and involved in the process of oxidative phosphorylation. In some embodiments, the term "ND1 gene" may refer to the above gene further comprising about 50 additional base pairs at one or both ends of the sequence.
[0048] CpG site: The term "CpG site" is used herein to distinguish this single-stranded linear sequence from the CG base pairing of cytosine and guanine in double-stranded sequences. "CpG" is an abbreviation for "C-phosphate-G", i.e., cytosine and guanine separated by only one phosphate, which is bound together to any two nucleosides in DNA. The term "CpG" is used to distinguish this linear sequence of CG base pairs of guanine and cytosine. Cytosine in a CpG dinucleotide can be methylated to form 5-methylcytosine.
[0049] CHG site: The term "CHG site" as used herein refers to a region of DNA, particularly a mitochondrial DNA region, where a cytosine nucleotide and a guanine nucleotide are separated by a variable nucleotide, which may be adenine, cytosine, or thymine. The cytosine at the CHG site may be methylated to form 5-methylcytosine.
[0050] CHH site: The term "CHH site" as used herein refers to a region of DNA, particularly mitochondrial DNA, in which a cytosine nucleotide is followed by a first and a second variable nucleotide (H), which may be adenine, cytosine, or thymine. The cytosine at the CHG site may be methylated to form 5-methylcytosine.
[0051] Determining the methylation pattern at a CpG site: The term "determining the methylation pattern at a CpG site" as used herein refers to determining the methylation status of a particular CpG site. Determining the methylation pattern at a CpG site can be performed by several processes known to those skilled in the art.
[0052] Determining the methylation pattern at a CHG site: The term "determining the methylation pattern at a CHG site" as used herein refers to determining the methylation status of a particular CHG site. Determining the methylation pattern at a CHG site can be performed by a number of processes known to those skilled in the art.
[0053] Determining the methylation pattern at a CHH site: The term "determining the methylation pattern at a CHH site" as used herein refers to determining the methylation status of a particular CHH site. Determining the methylation pattern at a CHG site can be performed by several processes known to those skilled in the art.
[0054] To determine the methylation pattern of mitochondrial DNA, a sample can be chemically treated so that all cytosine unmethylated bases are modified with uracil bases or with another base whose base pairing behavior is different from cytosine, while 5-methylcytosine bases remain unchanged. The term "modify" as used herein means the conversion of unmethylated cytosine to another nucleotide that distinguishes between methylated and unmethylated cytosine. The conversion of unmethylated cytosine bases in a sample containing mitochondrial DNA, without converting methylated cytosine, is carried out using a conversion agent. The term "conversion agent" or "conversion reagent" as used herein refers to a reagent capable of converting unmethylated cytosine to uracil or another base that is differentially detectable with respect to cytosine in terms of hybridization properties. The conversion agent is in particular a hydrogen sulfate, such as bisulfite or bisulfite. However, other agents that do not modify methylated cytosine but similarly modify unmethylated cytosine, such as bisulfite, can also be used in this method of the invention. This reaction is carried out by standard procedures (Frommer et al., 1992, Proc. Natl. Acad. Sci. USA 89: 1827-1831; Olek, 1996, Nucleic Acids Res. 24: 5064-6; EP 1394172). It is also possible to carry out this conversion enzymatically, for example using cytidine diaminase specific methylation.
[0055] Reference sample: The term "reference sample" refers to a sample containing mitochondrial DNA obtained from a subject not suffering from AD. In particular, the term refers to a low number of 5-methylcytosines in one or more CpG sites in the D-loop region shown in Table 1, one or more CpG sites in the ND1 gene shown in Table 2, one or more CHG sites in the D-loop region shown in Table 3, one or more CHG sites in the ND1 gene shown in Table 4, one or more CHH sites in the D-loop region shown in Table 5, and / or one or more CHH sites in the ND1 gene shown in Table 6, compared to the relative amount of 5-methylcytosines present in the one or more CpG sites, one or more CHG sites, and / or one or more CHH sites in the reference sample.
[0056] Treatment of Alzheimer's disease: The term "treatment of AD" as used herein refers to the treatment of the disease or any associated symptoms. Such treatments may include drug therapy, epigenetic treatments, or any cognitive stimulation treatments. Some cognitive stimulation treatments are digital cognitive treatments (i.e., using digital devices). This term may include any treatments known in the art for AD or future developments. Treatments for AD may include, but are not limited to, antioxidants, anti-inflammatory drugs, ginkgo, vitamins or dietary supplements, and antibodies against beta-amyloid, such as aducanumab, or other similar.
[0057] Furthermore, in the present invention, a subject may be classified as being at risk of developing ADD in the future. Some of the treatments for AD mentioned above may be useful for treating these patients to delay or prevent the onset of AD symptoms. Other treatments may be useful specifically to delay or prevent the onset of AD symptoms. Treatments for AD administered to subjects who have not yet developed ADD but are at risk of developing ADD (i.e., as determined by the methods disclosed herein) may be referred to as "preventive treatments for AD" as used herein. Thus, as used herein, the term "treatments for AD" may also refer to "preventive treatments for AD".
[0058] Preventive treatment of AD: The term "preventive treatment" as used herein refers to a combination of preventive or prophylactic measures to prevent or delay the onset of symptoms and to reduce or alleviate their clinical symptoms. In particular, this term refers to a set of measures to prevent or prevent the onset of clinical symptoms associated with Alzheimer's disease, to delay or alleviate them. Desirable clinical outcomes associated with administering treatment to a subject may include, but are not limited to, stabilization of the pathological stage of the disease, delaying the progression of the disease, and improving the physiological state of the subject. Suitable preventive treatments aimed at preventing or delaying the onset of symptoms of Alzheimer's disease include, but are not limited to, cholinesterase inhibitors, such as donepezil hydrochloride (Arecept), rivastigmine (Exelon), and galantamine (Reminyl), or antagonists N-methyl-D-aspartic acid (NMDA).
[0059] Dementia Staging Classification (DSC): The term "dementia staging classification" as used herein refers to classification of subjects based on prediction of dementia occurrence. This classification is derived from the risk score of progression to ADD as determined by the classification model disclosed herein. The classification model according to an embodiment of the present invention classifies subjects who are predicted to develop Alzheimer's disease dementia or diagnosed to develop Alzheimer's disease dementia (progression to ADD) and subjects who are predicted not to develop Alzheimer's disease dementia or diagnosed not to develop Alzheimer's disease dementia (non-progression to ADD). Subjects who do not develop Alzheimer's disease dementia may remit dementia, establish mild stage dementia, or develop other types of dementia. Thus, "dementia staging classification" includes two classes: progression to ADD and non-progression to ADD.
[0060] In addition, dementia stage classification (DSC) is also referred to in the development of classification model, especially during the training process of development.In this case, dementia stage classification (DSC) can include a third class: control.This is because the training data set used to develop classification model uses subjects who actually develop ADD, subjects who do not progress to ADD, and controls who do not suffer from any mild dementia and do not progress to ADD.Therefore, the classes included in dementia stage classification (DSC) can also be called ADD progress (i.e. progress to ADD) and ADD non-progression (i.e. not progress to ADD).
[0061] Specific hybridization: The phrase "specific hybridization" or "capable of hybridizing in a specific manner" as used herein refers to the ability of an oligonucleotide or polynucleotide to specifically recognize a specific sequence of interest, such as the D-loop region or the ND1 gene. The sequence of interest may refer to a reference sequence or a sequence resulting from a specific modification treatment, such as bisulfite treatment, in which unmethylated cytosine is modified to uracil. As used herein, the term "hybridization" refers to the process of combining two nucleic acid molecules or single-stranded molecules with a high degree of similarity, resulting in a simple double-stranded molecule by specific pairing between complementary bases. Usually, hybridization occurs under highly stringent or moderately stringent conditions.
[0062] Oligonucleotide: The term "oligonucleotide" as used herein is used interchangeably with "primer" and "nucleic acid sequence" and refers to a DNA molecule or short RNA having a length of up to 100 bases. The oligonucleotide of the present invention is in particular a DNA molecule at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, or 100 bases in length.
[0063] Score: The term "score" refers to one or more values, particularly a single value, that can be used as a component of a classification model for determining a subject's risk of developing a disease. Such a single value can be calculated or determined (i.e., estimated) by combining values of descriptive features processed by an interpretation function or algorithm. In an embodiment of the present invention, a score associated with the risk of developing ADD (also referred to as a "risk score") refers to a probability for the risk of developing ADD. The risk score can be a score between 0.0 and 1, where 0 refers to the lowest risk of developing the disease and 1 refers to the highest risk of developing the disease. The risk scores can be classified into groups, i.e., classes such as none, low, intermediate, and high.
[0064] Classifying a subject according to risk of developing Alzheimer's disease dementia: The term "classifying a subject according to risk of developing Alzheimer's disease dementia" means assigning a diagnostic subcategory, referred to herein as a dementia staging classification, which may include at least two categories: progression to ADD (referring to subjects at risk of progressing to ADD) and non-progression to ADD (referring to subjects not at risk of progressing to ADD).
[0065] Computer-Implemented Method: The term "computer-implemented method" refers to a method in which all or some of the steps of the method are performed by a computer, another programmable device, or a network of computers.
[0066] Supervised Machine Learning Methods: The term "supervised learning methods" refers to methods that use a training set to create a desired output to create a model. The training data set includes inputs and accurate outputs that train the model over time. The accuracy is measured via a loss function, and the model can be adjusted until the generalization error is sufficiently minimized.
[0067] Non-supervised machine learning methods: The term "non-supervised machine learning methods," also known as "unsupervised learning methods," refers to methods that use machine learning algorithms to analyze unlabeled data sets. These methods recognize similarities and differences in information to detect clusters or patterns in the data. These methods are commonly used for clustering, association, and dimensionality reduction of data sets.
[0068] Deep Learning: The term "deep learning" refers to a subfield of machine learning that uses neural network models with multiple layers to automatically extract and learn features and patterns from raw data. A neural network's layers are a set of interconnected artificial neurons (also known as nodes). These neurons receive input data and apply a certain type of computational task to it. Typically, these tasks include weighted sums and mathematical activation functions. The output of a layer is transmitted to the next layer, and so on. This process allows the network to gradually learn more complex features, attributes, and patterns in the data. Some examples are fully connected layers, convolutional layers, pooling layers, and recurrent layers.
[0069] Artificial Intelligence Methods: The term "artificial intelligence methods" refers to methods that use artificial intelligence, defined as the ability of a computer or a computer-controlled robot to mimic the ability of a human to respond to certain stimuli.
[0070] Classification model: The term "classification model" is interchangeably referred to herein as "classifying model". A classification model is, for example, a model developed using a type of supervised learning method that is capable of accurately assigning test data / new observations to a specific category based on training data. The model is trained to learn from a given training data set, and as a result, can classify new data sets into a specific score / number or class / group. Examples of classification algorithms are Linear Discriminant Analysis (LDA), Classification and Regression Trees (CART), k-Nearest Neighbors (kNN), Naive Bayes (NB), Support Vector Machines with Linear Kernel (SVM), Random Forests (RF), and Neural Networks (NNET).
[0071] Transform: The term "transform" refers to subjecting one or more descriptor features to an interpretation function or algorithm for a predictive model of a disease, particularly Alzheimer's disease. In some embodiments, the interpretation function may also result from multiple predictive models. In one embodiment, the predictive models include regression models and Bayesian classifiers or scores. In one embodiment, the interpretation function includes one or more terms related to one or more biomarkers or sets of biomarkers. In one embodiment, the interpretation function includes one or more terms related to the presence or absence or spatial distribution of specific cell types disclosed herein. In one embodiment, the interpretation function includes one or more terms related to the presence, absence, amount, intensity, or spatial distribution of morphological features of cells in the cell sample. In one embodiment, the interpretation function includes one or more terms related to the presence, absence, amount, intensity, or spatial distribution of descriptor features of cells in the cell sample.
[0072] Methods and classification model One aspect of the invention relates to a method of determining a subject's risk of developing Alzheimer's disease dementia, comprising applying a classification model to a methylation pattern of the D-loop region of mitochondrial DNA and / or the ND1 gene of a sample from the subject, wherein the classification model assigns the subject to a Dementia Stage Class (DSC) selected from progression to ADD or not progression to ADD. A computer-implemented method for determining a subject's risk of developing Alzheimer's disease dementia may comprise receiving a methylation pattern of the D-loop region of mitochondrial DNA and / or the ND1 gene of a sample from the subject, and the classification model assigning a DSC to the subject.
[0073] In some embodiments, the method comprises applying a classification model to at least one clinical variable.This method can be devised as a method for predicting the onset of ADD in a subject or for determining / identifying / assigning the dementia stage class to a subject; or for distinguishing between high and low risk of developing ADD in a subject.
[0074] Alternatively, the present invention relates in particular to a method, in particular a computer-implemented method, for determining a subject's risk of developing ADD, comprising the steps of: (a) determining a methylation pattern of (a) the D-loop region and / or (b) the ND1 gene of mitochondrial DNA of a sample obtained from the subject; and (b) determining a risk score indicative of the risk of developing ADD, wherein the risk score is calculated using a classification model configured to combine the methylation pattern of one or more sites determined in step (a) with at least one clinical variable of the subject as described herein. This method may alternatively be devised as a method for identifying a human subject at risk of developing ADD.
[0075] The methods described herein (based on determining the risk of developing Alzheimer's disease dementia in a subject) are useful for diagnosing ADD in a subject. The predictive methods described herein can be used in the clinic to make treatment decisions by selecting the most appropriate treatment modality for any particular patient. The methods are also useful for classifying subjects according to their risk of developing ADD, for example, to participate in clinical trials or receive appropriate treatment at an early stage, i.e., before the onset of severe dementia corresponding to the onset of ADD. Thus, application of the methods disclosed herein can improve clinical outcomes by matching patients to treatment, and can improve the accuracy in selecting patients necessary for successful clinical trials in the evaluation of potential AD treatments. The methods described herein are also useful for monitoring the progression to ADD in a subject over time, or the risk of developing ADD in a subject over time. Such methods can also be referred to as methods of determining prognosis.
[0076] In this sense, another embodiment relates to a method, particularly a computer-implemented method, for identifying a subject suitable for treatment of AD, comprising applying a classification model to the methylation pattern of the D-loop region of mitochondrial DNA and / or ND1 gene of a sample from the subject, and the classification model assigns the subject to a dementia stage class selected from progression to ADD or non-progression to ADD, and the dementia stage class consisting of progression to ADD indicates that the subject can be treated for AD. If the subject is assigned to non-progression to ADD, the subject may not be subjected to any treatment, but may undergo clinical follow-up. Alternatively, the subject may be treated for other dementias, if necessary. This embodiment may alternatively be described as related to a method for classifying a subject according to the risk of developing ADD, comprising following the steps described herein; a method for selecting an appropriate treatment for treating a subject at high risk of developing ADD, comprising the steps described herein.
[0077] Alternatively, the present invention relates to a method, in particular a computer-implemented method, for selecting a human subject for treatment (e.g. preventive treatment) for Alzheimer's disease, comprising the steps of (a) determining a methylation pattern of (a) the D-loop region and / or (b) the ND1 gene of mitochondrial DNA of a sample obtained from the subject, and (b) determining a risk score indicative of the risk of developing Alzheimer's disease, wherein the risk score is calculated using a classification model configured to combine the methylation pattern of each site determined in step (a) with at least one clinical variable of the subject as described herein.
[0078] In addition, another aspect relates to a method, particularly a computer-implemented method, for treating a subject, particularly a subject diagnosed with MCT or CDR0.5, administering a treatment for AD, and prior to administration, assigning the subject a DSC, determined by applying a classification model to the methylation pattern of the D-loop region of mitochondrial DNA and / or the ND1 gene of a sample from the subject, and the DSC is selected from progression to ADD or non-progression to ADD. In another particular embodiment, the subject has AD or is at risk of developing ADD. In one embodiment, if the subject is assigned a DSC progression to ADD, the subject is suitable to be administered a treatment for AD. In another embodiment, if the subject is assigned a non-progression to ADD, the subject may not be subjected to any treatment, but may undergo clinical follow-up. Alternatively, the subject may be treated for other dementias if necessary. This may alternatively be a method, particularly a computer-implemented method, for treating a subject, comprising: (i) prior to administration, applying a classification model to a methylation pattern of the D-loop region of mitochondrial DNA and / or the ND1 gene of a sample from the subject to assign the subject to a DCS, where the DCS is selected from progression to ADD or non-progression to ADD; (ii) administering to the subject a treatment for AD if DSC is progression to AFF, or no treatment or clinical follow-up for ADD, or treatment for another dementia, if DSC is non-progression to ADD. It may be devised as a method, particularly a computer-implemented method, including:
[0079] Alternatively, the present invention provides a method, particularly a computer-implemented method, for treating a subject having Alzheimer's disease or at risk of developing ADD, comprising the steps of: (a) determining a methylation pattern of (a) the D-loop region and / or (b) the ND1 gene of mitochondrial DNA of a sample obtained from the subject; (b) determining a risk score indicative of the risk of developing ADD, wherein the risk score is calculated using a classification model configured to combine the methylation pattern of one or more sites determined in step (a) with at least one clinical variable of the subject as described herein; and (c) administering treatment to the subject if the risk score indicates that the subject is at risk of developing ADD. The present invention relates to a method, particularly a computer-implemented method, including:
[0080] Another aspect relates to a method, particularly a computer-implemented method, of providing individualized therapy to a subject at high risk of developing ADD, the method comprising the steps described herein.
[0081] Another aspect relates to a method, particularly a computer-implemented method, for monitoring the progression of Alzheimer's disease in a subject or for monitoring the risk of developing ADD in a subject, comprising the steps of: (a) determining the methylation pattern of (a) D-loop region and / or (b) ND1 gene of mitochondrial DNA of a sample obtained from the subject; (b) determining a risk score indicative of the risk of developing ADD, wherein the risk score is calculated using a classification model configured to combine the methylation pattern of one or more sites determined in step (a) with at least one clinical variable of the subject described herein; and (c) comparing the risk score determined in step (b) with the risk score obtained at an earlier stage of the disease. A higher risk score than the previous risk score indicates the progression of ADD and thus a worse prognosis. The risk score can be monitored, for example, once a year.
[0082] In some embodiments, the method described above further comprises: a) determining the methylation pattern of the D-loop region of mitochondrial DNA and / or the ND1 gene in a sample of a subject comprising mitochondrial DNA; b) combining the methylation pattern data with at least one clinical variable of the subject described herein, said combining being performed using a classification model to determine a risk score that correlates with the subject's risk of developing ADD; Includes.
[0083] These steps may alternatively be: a) determining the methylation pattern of the D-loop region of mitochondrial DNA and / or the ND1 gene in a sample of a subject comprising mitochondrial DNA; and b) determining a risk score indicative of the risk of developing ADD, the risk score being calculated using a classification model configured to combine the methylation pattern of one or more sites determined in step (a) with at least one clinical variable of the subject as described herein. It can be conceived as:
[0084] In some embodiments, the methylation pattern of the D-loop region of mitochondrial DNA and / or the ND1 gene is (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 The present invention is characterized in that the endothelial cell is at least one site selected from the group consisting of:
[0085] In some embodiments, the methylation pattern is determined using at least one oligonucleotide capable of specifically hybridizing to a mitochondrial DNA sequence comprising a methylation site selected from the group consisting of (i) to (vi).
[0086] [Table 1]
[0087] [Table 2]
[0088] [Table 3]
[0089] [Table 4]
[0090] [Table 5]
[0091] [Table 6]
[0092] In some embodiments, the methylation pattern is determined using at least one oligonucleotide capable of specifically hybridizing to a mitochondrial DNA sequence comprising a methylation site selected from the group consisting of (i) to (vi). In particular, the at least one oligonucleotide has a length of 15 to 100 nucleotides and comprises a sequence selected from the group consisting of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, and SEQ ID NO:4.
[0093] In one embodiment, the methods described herein include: a) determining a methylation pattern of the D-loop region of mitochondrial DNA and / or the ND1 gene in a sample of a subject comprising mitochondrial DNA, the methylation pattern being: (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 and determining at least one site selected from the group consisting of: b) using the classification model to determine a risk score correlating the methylation pattern to the subject's risk of developing ADD; Includes.
[0094] In one embodiment, the methods described herein include: a) determining a methylation pattern of the D-loop region of mitochondrial DNA and / or the ND1 gene in a sample of a subject comprising mitochondrial DNA, the methylation pattern being: (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 is determined at at least one site selected from the group consisting of The methylation pattern correlates with the subject's risk of developing ADD. Includes.
[0095] In some embodiments, the method of the invention further comprises: a) determining a methylation pattern of the D-loop region of mitochondrial DNA and / or the ND1 gene in a sample of a subject comprising mitochondrial DNA, the methylation pattern being: (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 and determining at least one site selected from the group consisting of: b) combining the methylation pattern of one or more sites determined in step (a) with at least one clinical variable of the subject as described herein, said combining being performed using a classification model to determine a risk score that correlates with the subject's risk of developing ADD; Includes.
[0096] In particular, the methylation pattern is determined using at least one oligonucleotide / primer capable of specifically hybridizing to a mitochondrial DNA sequence containing a methylation site selected from the group consisting of (i) to (vi).
[0097] Classification Models In another embodiment, the present invention provides a classification model capable of classifying subjects into classes of progression to ADD and non-progression to ADD, which relate to subjects at risk of developing ADD and subjects not at risk of developing ADD, respectively.
[0098] Thus, the present invention provides a classification model for determining a subject's risk of developing ADD, wherein the classification model uses data relating to the methylation pattern obtained from the subject's sample and data relating to the subject's clinical variables to identify the subject as belonging to a class from the group consisting of progressing to ADD and not progressing to ADD, wherein identification of progressing to ADD indicates that the subject is at risk of developing into ADD.
[0099] In some embodiments, the methods disclosed herein include determining a risk score indicative of the risk of developing ADD, wherein the risk score is calculated or determined using a classification model configured to combine the methylation pattern with at least one clinical variable of the subject described herein.
[0100] In some embodiments, the methylation pattern comprises a methylation pattern of at least one of the D-loop region of mitochondrial DNA and / or the ND1 gene, wherein the methylation pattern is (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 The at least one site selected from the group consisting of:
[0101] In some embodiments, the classification model is obtained by artificial intelligence methods. In particular embodiments, the classification model is obtained by machine learning methods. In more particular embodiments, the classification model is obtained by supervised machine learning methods. In more particular embodiments, the supervised machine learning method is selected from the group consisting of Linear Discriminant Analysis (LDA), Penalized Multinomial Regression (PMR), Classification and Regression Trees (CART), k-Nearest Neighbors (kNN), Naive Bayes (NB), Support Vector Machines with a Linear Kernel (SVM), Support Vector Machines with a Radial Basis Function Kernel (SVM. Radial), Random Forests (RF), and Neural Networks (NNET), Logistic Regression, Artificial Neural Networks (ANN), GBoost (XGB; an implementation of gradient boosting decision trees designed for speed and performance), Glmnet (a package that fits generalized linear models via restricted maximum likelihood), cforest (an implementation of random forest and bagging ensemble algorithms that utilize conditional inference trees as base learners), Treebag (bagging, i.e. bootstrap aggregation, an algorithm for improving model accuracy in regression and classification problems that builds multiple models from individual subsets of training data and builds a final aggregate model), or combinations thereof. More specifically, the supervised machine learning method is selected from the group consisting of Linear Discriminant Analysis (LDA), Penalized Multinomial Regression (PMR), Classification and Regression Trees (CART), k-Nearest Neighbors (kNN), Naive Bayes (NB), Support Vector Machine with Linear Kernel (SVM), Support Vector Machine with Radial Basis Function Kernel (SVM. Radial), Random Forest (RF), and Neural Network (NNET). More specifically, the supervised machine learning method is Random Forest. In another embodiment, the classification model is obtained by a non-supervised machine learning method.In certain embodiments, the unsupervised machine learning method is selected from the group consisting of k-means, K-Medoids, Fuzzy C-Means, Agglomerative Hierarchical Clustering, Gaussian Mixture Models (GMM), Neural Networks, Hidden Markov Models (HMM), Mean-Shift, DBSCAN Clustering, Apriori Algorithm, Principal Component Analysis (PCA), Independent Component Analysis (ICA), Linear Discriminant Analysis (LDA), Singular Value Decomposition (SVD), Linear Semantic Analysis (LSA), t-SNE, Nonlinear Multidimensional Scalling, Principal Curves, k-Nearest Neighbors (kNN), Locally Linear Embedding (LKE), and Autoencoder.
[0102] Alternatively, the classification model is obtained by a deep learning method. In a particular embodiment, the deep learning method can be supervised or unsupervised. In another embodiment, the deep learning method is selected from the group consisting of a convolutional neural network (CNN), a long short-term memory network (LSTM), a recurrent neural network (RNN), a generative adversarial network (GAN), a RBF network (RBFN), a multi-layer perceptron (MLP), a self-organizing map (SOM), a deep belief network (DBN), a restricted Boltzmann machine (RBM), and an autoencoder.
[0103] In some embodiments, the classification model is trained or has been trained with a training set comprising mitochondrial methylation patterns at methylation sites defined herein in a plurality of samples associated with a plurality of subjects, and clinical variables associated with a plurality of subjects, and each subject is assigned a dementia stage classification. In certain embodiments, the dementia stage classification is selected from control, ADD progressing, and ADD non-progressing.
[0104] In some embodiments, a subject classified as a control has a CDR score of 0 and is characterized by a clinical follow-up of more than 10 years (i.e., clinical follow-up of more than 10 years). In some embodiments, a subject classified as an ADD non-progressor has a CDR score of 0.5 and is characterized by a clinical follow-up of more than 36 months without progression of symptoms. In some embodiments, a subject classified as an ADD progressor has a CDR score of 0.5 to 1 after progression.
[0105] In some embodiments, the training dataset comprises an accuracy output corresponding to the dementia stage class assigned to each subject, where the dementia stage classes are control, progression to ADD, and non-progression to ADD.
[0106] In some embodiments, data is pre-processed or has been pre-processed before the classification model is trained. This step is performed to ensure and enhance the performance of the model training process. Data pre-processing includes: (1) creating dummy variables; (2) removing zero-variance and near-zero-variance variables; (4) splitting data into training data set and test data set; (5) centering and scaling; and (6) testing and visualizing the training data set.
[0107] In particular, (1) a process of creating dummy variables is performed to handle categorical data. Essentially, each categorical variable is converted to a numerical variable by creating a dummy variable using a procedure called the "One-Hot Encoding" approach (i.e., each new variable is forced to have a value of 0 or 1, representing the presence or absence of the attribute). This process is performed to ensure that the variables are coded consistently, i.e., coded to ensure that there is no linear dependency between the new attributes, thus avoiding the dummy variable trap. Additionally, the data is reviewed to ensure that all categorical one-coded variables do not exhibit unusual linear combinations, and if there are any abnormalities, redundant variables are removed until the linear combinations are eliminated.
[0108] (2) The process of removing zero-variance and near-zero-variance variables is performed to remove variables that exhibit a single unique value and variables that have a few unique values that are significantly unbalanced. Otherwise, these predictors may cause instability problems during the fitting process or model crashes.
[0109] (3) The process of dimensionality reduction is carried out to reduce the number of variables (i.e. features or attributes) in a dataset while preserving as much relevant information as possible. In other words, the objective is to remove redundant or irrelevant features with the aim of improving the potency and effectiveness of learning models applied to build classification models. In this context, there are two main approaches: feature selection and feature extraction.
[0110] (i) Feature Selection Techniques: The basic idea of these methods is to select a subset of variables based on some criteria, for example by identifying and removing correlated variables. This process is performed with the aim of reducing highly correlated variables. To perform this step, a correlation matrix is calculated. Usually, the correlation measure applied is the Pearson correlation coefficient. Thus, to detect highly correlated variables based on the absolute value of the pairwise correlation, if two variables are highly correlated, the average absolute correlation of each variable is examined and the variable with the maximum average absolute correlation is removed. In this context, usually, a cutoff of the pairwise absolute correlation can be set, for example, by testing a linear regression between each pair of variables. However, this approach could not reveal the correlation of additional features. For this reason, on the other hand, other monotonic functional methods for identifying correlations between variables can also be contemplated. For example, the non-parametric methods Spearman's rank correlation and Kendall's tau correlation measure the association between each pair of variables based on the rank of the observations and not their actual values. They are usually applied when a non-linear relationship between attributes (variables), skewness measures or ordinal scales is reasonably suspected. Another type of method could be distance correlation (aka dCor). This is done to measure the dependency between each pair of variables in a way that is sensitive to non-linear relationships. Most of these methods are considered to test the relationship between continuous and / or ordinal variables, but there are also methods that test the strength of the relationship between dichotomous and continuous variables (e.g. point biserial correlation) or when both variables are dichotomous (Phy coefficient). On the other hand, it is recommended to use an objective approach such as cross-validation to determine the appropriate threshold for detecting significantly correlated variables. That is, to determine the required cutoff, the data must first be split into a training data set and a test data set using a random number seed to ensure reproducibility. Second, the training data set is used to train the model and the test set is used to evaluate its performance. Third, the threshold for considering variables as highly correlated is changed and then the second step is repeated for each threshold.Fourth, select the threshold that gives the best performance on the test set, such as the one with the highest accuracy (or lowest error rate). Fifth, after the cutoff is selected, it must be applied to the entire data set and the model re-trained using all the data.
[0111] Other feature selection methods that may be applied are genetic algorithms (GA), L1-regularized logistic regression, Lasso regression, or hybrid methods.
[0112] (ii) Feature Extraction Techniques: These approaches create new variables by combining or applying transformations on the original features into a space with fewer dimensions. Some examples are Principal Components (PCA), Multifactor Analysis (MFA), t-SNE, or alternatively UMAP and / or Multidimensional Scaling (MDS) approaches, Partial Least Squares-Discriminant Analysis (PLS-DA), or autoencoders.
[0113] (4) A data splitting process is performed to randomly split the data into two main subsets: one for training the model (80% of the samples) and one for testing the classification model (20% of the samples). A random sampling process is driven within each class to preserve the overall class distribution of the data. A random number seed is designed to ensure reproducibility.
[0114] (5) A centering and scaling process is applied to the continuous features (variables) of the training dataset for the purpose of estimating centering and scaling factors that must be applied to both datasets to create a normalized dataset for the training and testing processes of the classification model.
[0115] (6) Examining and visualizing the training data set is a process that takes place after the previous preprocessing tasks. This is a second exploratory data analysis (EDA) that is performed on the training data set and is guided to review and confirm the absence of bias from the original data. Traditional statistical descriptive methods are applied (i.e., univariate, bivariate, and multivariate descriptive methods).
[0116] The classification models disclosed herein can be trained with data corresponding to a set of samples from which methylation data corresponding to a set of methylation sites were obtained. For example, the training set includes data of methylation patterns from the methylation sites presented in Tables 1-6 or any combination thereof.In some embodiments, the methylation pattern data is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73 , 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 84, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 8, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196 , 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, or 250 methylation sites. In some embodiments, the methylation pattern data comprises data for more than 50 methylation sites.In some embodiments, the methylation pattern data comprises data for more than 100 methylation sites. In some embodiments, the methylation pattern data comprises data for more than 200 methylation sites. In some embodiments, the methylation pattern data comprises data for between about 10 and about 20, between about 20 and about 30, between about 30 and about 40, between about 40 and about 50, between about 50 and about 60, between about 60 and about 70, between about 70 and about 80, between about 80 and about 90, between about 90 and about 100, between about 110 and about 120, between about 120 and about 130, between about The methylation sites are between 130 and about 140, between about 140 and about 150, between about 150 and about 160, between about 160 and about 170, between about 170 and about 180, between about 180 and about 190, between about 190 and about 200, between about 200 and about 210, between about 210 and about 220, between about 220 and about 230, between about 230 and about 240, and between about 240 and about 250.
[0117] In some embodiments, the training data set includes additional clinical variables for each subject, such as the classification of the subject according to a classification model disclosed herein, hi other embodiments, the training data includes data about the subjects, such as weight, ethnicity, the presence or absence of biomarkers, medications, etc.
[0118] In some embodiments, the training set comprises a reference population of at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 110, at least about 120, at least about 130, at least about 140, at least about 150, at least about 160, at least about 170, at least about 180, at least about 190, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, at least about 500, at least about 600, at least about 700, at least about 800, at least about 900, or at least about 1000 subjects. In other embodiments, the training set comprises more than 1000 subjects.
[0119] In some embodiments, the classification model comprises determining a (relative) weight for each methylation pattern and each clinical variable considered.
[0120] In some embodiments, the classification model uses data indicating the (relative) weight of each methylation pattern and each clinical variable for determining the risk of developing ADD.
[0121] In some embodiments, determining the risk score comprises correlating each of the at least one methylation pattern and each of the at least one clinical variable with the determined weights.
[0122] The classification model described herein may include different sets and combinations of methylation patterns and / or clinical variables. The classification model selects methylation patterns and / or clinical variables according to the contribution or importance (i.e., determined weight) associated with each methylation pattern and / or clinical variable. That is, the classification model is configured to combine methylation patterns of sites that correlate with the determined weights (e.g., 0.25, 0.5, 1, 2, 2.5, 5, etc.) and / or clinical variables that correlate with at least one determined weight (e.g., 0.25, 0.5, 1, 2, 2.5, 5, etc.).
[0123] In some embodiments, the determined weights correlating to the methylation patterns or clinical variables are in the range of 0 to 100. Alternatively, the determined weights correlating to the methylation patterns or clinical variables are selected from 0 to 100. The value "0" corresponds to the least variable importance or the least determined weight, and the value "100" corresponds to the most variable importance or the most determined weight. In certain embodiments, the determined weights are 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51 , 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, and 100.
[0124] In some embodiments, the determined weight is at least 0.25. In particular, the determined weight is selected from the group consisting of 0.25, 0.3, 0.35, 0.40, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, 0.95, and 1. In another embodiment, the determined weight is at least 1. In particular, the determined weight is selected from the group consisting of 1, 1.25, 1.5, 1.75, 2, 2.25, 2.5, 2.75, 3, 3.25, 3.5, 3.75, 4, 4.25, 4.5, 4.75, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, and 100. In another embodiment, the determined weight is at least 2. In particular, the determined weights are selected from the group consisting of 2, 2.25, 2.5, 2.75, 3, 3.25, 3.5, 3.75, 4, 4.25, 4.5, 4.75, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, and 100. In another embodiment, the determined weights are selected from the group consisting of at least 1, at least 2, at least 2.5, at least 3, at least 3.5, at least 4, at least 4.5, at least 5, at least 7.5, at least 10, at least 15, at least 20, at least 25, at least 50, and at least 75.
[0125] In some embodiments, determining the subject's risk of developing ADD comprises considering the methylation patterns of sites that correlate with a determined weight of at least 0.5. In another embodiment, determining the subject's risk of developing ADD comprises combining the methylation patterns of sites that correlate with a determined weight of at least 0.5.
[0126] In another embodiment, determining the subject's risk of developing ADD comprises considering clinical variables that correlate with a determined weight of at least 0.5. In another embodiment, determining the subject's risk of developing ADD comprises combining clinical variables that correlate with a determined weight of at least 0.5.
[0127] In another embodiment, determining the risk of developing ADD in a subject comprises combining a methylation pattern that correlates with a determined weight of at least 0.5 with a clinical variable of a site that correlates with a determined weight of at least 0.5. Alternatively, determining the risk of developing ADD in a subject comprises combining or considering a methylation pattern and / or a clinical variable that correlates with a determined weight of at least 0.5.
[0128] In some embodiments, determining the subject's risk of developing ADD comprises considering the methylation pattern of sites that correlate with at least one determined weight. In another embodiment, determining the subject's risk of developing ADD comprises combining the methylation patterns of sites that correlate with at least one determined weight.
[0129] In another embodiment, determining the subject's risk of developing ADD comprises considering clinical variables that correlate with at least one determined weight. In another embodiment, determining the subject's risk of developing ADD comprises combining clinical variables that correlate with at least one determined weight.
[0130] In another embodiment, determining the subject's risk of developing ADD comprises combining a methylation pattern of a site that correlates with at least one determined weight with a clinical variable that correlates with at least one determined weight. Alternatively, determining the subject's risk of developing ADD comprises combining or considering a methylation pattern and / or a clinical variable that correlates with at least one determined weight.
[0131] In some embodiments, determining the subject's risk of developing ADD comprises considering the methylation pattern of sites that correlate with at least two of the determined weights. In another embodiment, determining the subject's risk of developing ADD comprises combining the methylation pattern of sites that correlate with at least two of the determined weights.
[0132] In another embodiment, determining the subject's risk of developing ADD comprises considering clinical variables that correlate with at least two of the determined weights. In another embodiment, determining the subject's risk of developing ADD comprises combining clinical variables that correlate with at least two of the determined weights.
[0133] In another embodiment, determining the subject's risk of developing ADD comprises combining a methylation pattern of sites that correlate with at least two determined weights with a clinical variable that correlates with at least two determined weights. Alternatively, determining the subject's risk of developing ADD comprises combining or considering a methylation pattern and / or a clinical variable that correlates with at least two determined weights.
[0134] In some embodiments, determining the subject's risk of developing ADD comprises considering the methylation patterns of sites that correlate with a determined weight of at least 2.5. In another embodiment, determining the subject's risk of developing ADD comprises combining the methylation patterns of sites that correlate with a determined weight of at least 2.5.
[0135] In another embodiment, determining the subject's risk of developing ADD comprises considering clinical variables that correlate with a determined weight of at least 2.5. In another embodiment, determining the subject's risk of developing ADD comprises combining clinical variables that correlate with a determined weight of at least 2.5.
[0136] In another embodiment, determining the risk of developing ADD in a subject comprises combining a methylation pattern of a site that correlates with a determined weight of at least 2.5 with a clinical variable that correlates with a determined weight of at least 2.5. Alternatively, determining the risk of developing ADD in a subject comprises combining or considering a methylation pattern and / or a clinical variable that correlates with a determined weight of at least 2.5.
[0137] In some embodiments, determining the subject's risk of developing ADD comprises considering methylation patterns of sites that correlate with a determined weight of at least 5. In another embodiment, determining the subject's risk of developing ADD comprises combining methylation patterns of sites that correlate with a determined weight of at least 5.
[0138] In another embodiment, determining the subject's risk of developing ADD comprises considering clinical variables that correlate with at least 5 of the determined weights. In another embodiment, determining the subject's risk of developing ADD comprises combining clinical variables that correlate with at least 5 of the determined weights.
[0139] In another embodiment, determining the subject's risk of developing ADD comprises combining a methylation pattern of sites that correlate with a determined weight of at least 5 with a clinical variable that correlates with a determined weight of at least 5. Alternatively, determining the subject's risk of developing ADD comprises combining or considering a methylation pattern and / or a clinical variable that correlates with a determined weight of at least 5.
[0140] Alternatively, the methylation patterns and / or clinical variables combined or considered by the classification model for determining the risk of developing ADD are selected according to their determined weights. In particular, the methylation patterns and / or clinical variables combined by the classification model are selected according to their determined weights, the determined weights being at least 1. In another particular embodiment, the methylation patterns and / or clinical variables combined by the classification model are selected according to their determined weights, the determined weights being at least 2. In some embodiments, the methylation patterns and / or clinical variables combined by the classification model for determining the risk of developing ADD are correlated with a determined weight selected from the group consisting of at least 0.25, at least 0.5, at least 1, at least 2, at least 2.5, at least 3, at least 3.5, at least 4, at least 4.5, at least 5, at least 7.5, at least 10, at least 15, at least 20, at least 25, at least 50, and at least 75.
[0141] In some embodiments, the classification model can classify subjects into two categories: progression to ADD and non-progression to ADD.In some embodiments, the classification model calculates a risk score for each category, which corresponds to the probability of the subject being assigned to each category.In some embodiments, the probabilities for each category sum to 1.
[0142] Herein above, the weights correlating to the methylation patterns or clinical variables are in the range or selected from 0 to 100, i.e. a scale of 0 to 100 is used. In this scale, minimum values of certain weights are defined herein. However, it is clear that any other suitable scale may be used, for example a scale of 0 to 1 or a scale of 0 to 1000 or a scale of 0 to 20. With such a different scale, the minimum values of weights mentioned here above may be changed proportionately.
[0143] The classification model created by the machine learning method (e.g., random forest) disclosed herein can then be evaluated by determining the characteristics of the classifier that correctly classifies each test subject. In some embodiments, the subjects of the training population used to derive the model are different from the subjects of the test population used to test the model. As will be understood by those skilled in the art, this allows prediction of the ability of the data set used to train the classifier with respect to its ability to properly classify subjects whose output classification (e.g., dementia stage classification, i.e., progression to ADD or not progression to ADD) is unknown.
[0144] In some embodiments, the classification model is evaluated for the features that properly classify each subject of the training population using methods known to those skilled in the art.For example, the classification model can be evaluated using cross-validation, leave-one-out cross-validation (LOOCV), n-fold cross-validation, or jackknife analysis using standard statistical methods.In other embodiments, each classifier is evaluated for its ability to properly characterize the subjects of the training population that were not used to create the classifier.
[0145] In some embodiments, the method used to evaluate the classification model with respect to its ability to properly classify each subject in the training population is a method to evaluate the sensitivity (TPF, true positive rate and 1-specificity (FPF, false positive rate) of the classification model. In one embodiment, the method used to test the classifier is the Receiver Operating Characteristic ("ROC"), which provides several parameters to evaluate both the sensitivity and specificity of the results of the created classification model, e.g., a model derived from the application of a random forest.
[0146] In some embodiments, the metrics used to evaluate the classification model for its ability to properly classify each subject in the training population include classification accuracy (ACC), area under the receiver operating characteristic curve (AUC ROC), sensitivity (true positive rate, TPF), specificity (true negative rate, TNF), positive predictive value (PPV), negative predictive value (NPV), or a combination thereof. In other embodiments, the metrics used to evaluate the classification model for its ability to properly classify each subject in the training population are classification accuracy (ACC), area under the receiver operating characteristic curve (AUC ROC), sensitivity (true positive rate, TPF), specificity (true negative rate, TNF), positive predictive value (PPV), and negative predictive value (NPV).
[0147] Another aspect of the invention is a computer implemented method applicable to the methods described herein, for example for obtaining a risk score for developing Alzheimer's disease dementia, comprising the following steps: (a) As input data (1) The methylation pattern of the D-loop region of the mitochondrial DNA and / or the ND1 gene of the subject (the methylation pattern is (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 at least one site selected from the group consisting of: (2) at least one clinical variable of the subject described herein Providing or receiving (b) combining and weighting the methylation patterns and clinical variables using a classification model to obtain a risk score; The present invention relates to a method comprising the steps of:
[0148] In one embodiment, a computer-implemented method for determining a subject's risk of developing Alzheimer's disease dementia comprises: a) receiving data relating to a methylation pattern of (a) the D-loop region and / or (b) the ND1 gene of the mitochondrial DNA of the subject, (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 and b) determining a risk score indicative of risk of developing Alzheimer's disease dementia, the risk score being calculated using a classification model configured to combine the methylation pattern of the one or more sites determined in step (a) with at least one clinical variable of the subject selected from the group consisting of sex, total box score, Mini-Mental State Examination, Positron Emission Tomography, presence or absence of beta amyloid protein, age, apolipoprotein E genotype level, and reclassified apolipoprotein E genotype level; Includes.
[0149] In particular, the at least one clinical variable of the subject is selected from the group consisting of sex, age, APOE, alE4, PET, presence or absence of beta amyloid protein, SOB, and MMSE.
[0150] Score In some embodiments, the risk score is a score between 0 and 1, with 0 indicating the lowest risk of developing ADD and 1 indicating the highest risk of developing ADD in a subject. In certain embodiments, a risk score of 0.5 or greater than 0.5 indicates that the subject is likely to develop ADD. Alternatively, a risk score of 0.5 or greater than 0.5 indicates that the subject is at high risk of developing ADD. In certain embodiments, a risk score of 0.75 or greater than 0.75 indicates that the subject is at very high risk of developing / will develop ADD. In some embodiments, a risk score of less than 0.5 indicates that the subject is unlikely to develop ADD. Alternatively, a risk score of less than 0.5 or greater indicates that the subject is at low risk of developing ADD. In certain embodiments, a risk score of 0.25 or greater than 0.25 indicates that the subject is at very low risk of developing / will not develop ADD.
[0151] In some embodiments, the risk score is a score of 0-100, with 0 indicating the subject is at least at risk for developing ADD and 100 indicating the subject is at highest risk for developing ADD. In some embodiments, the risk score can be, but is not limited to, 1-2, 1-5, 1-10, 1-100, 0-10, and 0-100. It is understood that the risk score in the present invention can be any range of values that serves to correlate with the subject's risk for developing ADD.
[0152] subject In some embodiments, the subject is suspected of having dementia, is at risk of developing dementia, or has been diagnosed with dementia.In some embodiments, the subject has a clinical dementia rating scale score of 0.5.In particular, the subject is diagnosed with mild cognitive impairment.In some embodiments, the subject has at least one symptom of mild dementia.In particular, the subject has at least one symptom selected from the group consisting of mild memory loss, mild attention loss, difficulty in reasoning, planning, or problem solving, language difficulty, and reduced depth perception.
[0153] The "Clinical Dementia Rating Scale" is a global summary obtained through semi-structured interviews of the patient and informant, in which the cognitive status of the subject is graded in six domains of functioning, including memory, orientation, judgment, and problem solving, community situations, home and hobbies, and daily care. Each domain is assigned a score of 0 to 3, and the results are computerized through an algorithm to obtain a CDR global score. The CDR global score ranges from 0 to 3 and allows for grouping of subjects according to the severity of their dementia, where CDR=0 corresponds to no cognitive impairment, CDR=0.5 is doubtful or very mild dementia, CDR=1 is mild dementia, CDR=2 is moderate dementia, and CDR=3 is severe dementia.
[0154] Subjects with a CDR of 0.5 are diagnosed with mild cognitive impairment (MCI), which is an early memory loss or loss of other cognitive abilities, such as language or visual / spatial perception, in subjects who maintain the ability to independently perform most activities of daily living.Subjects with a CDR of 1 or more are considered to have already progressed to ADD.
[0155] Methylation patterns and their measurement The methylation pattern of mitochondrial DNA (mtDNA) sites described herein can be determined using any method in the art.For example, the methylation of mtDNA can be determined by treating a sample with bisulfite and sequencing the treated sample.
[0156] In some embodiments, determining the methylation pattern comprises: (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 The method comprises determining the methylation pattern of at least one site selected from the group consisting of:
[0157] In some embodiments, determining the methylation pattern comprises determining the methylation pattern of (vi) the CHH sites of the ND1 region as shown in Table 6, and (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4, and (v) a CHH site in the D-loop region shown in Table 5 at least one site selected from the group consisting of:
[0158] In some embodiments, determining the methylation pattern comprises determining the methylation pattern of at least one of the CHH sites of the ND1 gene shown in Table 6. In some embodiments, determining the methylation pattern comprises determining the methylation pattern of at least 10 of the CHH sites of the ND1 gene shown in Table 6. In some embodiments, determining the methylation pattern comprises determining the methylation pattern of at least 25 of the CHH sites of the ND1 gene shown in Table 6. In some embodiments, determining the methylation pattern comprises determining the methylation pattern of at least 50 of the CHH sites of the ND1 gene shown in Table 6. In some embodiments, determining the methylation pattern comprises determining the methylation pattern of at least 100 of the CHH sites of the ND1 gene shown in Table 6. In some embodiments, determining the methylation pattern comprises determining the methylation pattern of all of the CHH sites of the ND1 gene shown in Table 6.
[0159] In some embodiments, determining the methylation pattern comprises determining the methylation pattern of at least one CHH site in the D-loop region as set forth in Table 5. In some embodiments, determining the methylation pattern comprises determining the methylation pattern of all CHH sites in the D-loop region as set forth in Table 5.
[0160] In some embodiments, determining the methylation pattern comprises determining the methylation pattern of at least one CHG site of the ND1 gene shown in Table 4. In some embodiments, determining the methylation pattern comprises determining the methylation pattern of all CHG sites of the ND1 gene shown in Table 4.
[0161] In some embodiments, determining the methylation pattern comprises determining the methylation pattern of at least one CHG site in the D-loop region as set forth in Table 3. In some embodiments, determining the methylation pattern comprises determining the methylation pattern of all CHG sites in the D-loop region as set forth in Table 3.
[0162] In some embodiments, determining the methylation pattern comprises determining the methylation pattern of at least one CpG site of the ND1 gene shown in Table 2. In some embodiments, determining the methylation pattern comprises determining the methylation pattern of all CpG sites of the ND1 gene shown in Table 2.
[0163] In some embodiments, determining the methylation pattern comprises determining the methylation pattern of at least one CpG site of the D-loop region shown in Table 1. In some embodiments, determining the methylation pattern comprises determining the methylation pattern of all CpG sites of the D-loop region shown in Table 1.
[0164] In some embodiments, determining the methylation pattern comprises determining the methylation of all CpG sites, CHG sites, and CHH sites in the D-loop region. In another embodiment, determining the methylation pattern comprises determining the methylation of all CpG sites, CHG sites, and CHH sites in the ND1 gene. In some embodiments, determining the methylation pattern comprises determining the methylation of all CpG sites in the D-loop region and the ND1 gene. In some embodiments, determining the methylation pattern comprises determining the methylation of all CHG sites in the D-loop region and the ND1 gene. In some embodiments, determining the methylation pattern comprises determining the methylation of all CHH sites in the D-loop region and the ND1 gene. In some embodiments, determining the methylation pattern comprises determining the methylation of all CpG sites, CHG sites, and CHH sites in the D-loop region and the ND1 gene.
[0165] The methylation pattern may be determined by any method known in the art. In some embodiments, the methylation pattern is determined by a technique selected from the group consisting of a bisulfite treatment-based technique, a biological identification-based technique, and a bisulfite-free and enzyme-free technique-based technique.
[0166] In some embodiments, the bisulfite treatment-based techniques include, but are not limited to, sequence-based analysis, melting temperature and interaction analysis-based analysis. In some embodiments, the sequence-based analysis includes, but is not limited to, bisulfite sequencing, methylation-specific PCR (MS-PCR), methylation-sensitive single nucleotide primer extension (Ms-SnuPE), and reduced representation bisulfite sequencing (RRBS). In some embodiments, the melting temperature-based analysis includes, but is not limited to, methylation-specific denaturing gradient gel electrophoresis (MS-DGGE), methylation-specific melting curve analysis (MS-MCA), and methylation-specific high-resolution melting curve analysis (MS-HRM). In some embodiments, the interaction-based analysis includes, but is not limited to, combined bisulfite-restriction analysis (COBRA) and Methylight assay.
[0167] In some embodiments, the biological identification-based techniques include, but are not limited to, enzymatic digestion and organism-dependent reaction-based methods. In some embodiments, the enzymatic digestion-based methods include, but are not limited to, restriction landmark genome scanning (RLGS), online monitoring and methylation-sensitive restriction enzyme-PCR (MS-RE-PCR / Southern). In certain embodiments, the organism-dependent reaction is methyl capture using methyl-CpG binding domain (MBD) proteins.
[0168] In some embodiments, the bisulfite-free and enzyme-free techniques include, but are not limited to, direct oxidation-based assays and oxidative chemical decomposition-based assays. In certain embodiments, the direct oxidation-based assay is multi-walled carbon nanotubes supported with a choline chloride monolayer (MWCNTs / Ch / GCE). In certain embodiments, the oxidative chemical decomposition-based assay is Na1O4 / LiBr.
[0169] In certain embodiments, the methylation pattern is determined by a technique based on bisulfite treatment. In particular, the methylation pattern is determined by sequence-based analysis. More particularly, the methylation pattern is determined by bisulfite sequencing.
[0170] In some embodiments, the methylation patterns are analyzed by methylation specific PCR, methylation specific PCR (MS-PCR), quantitative methylation specific polymerase chain reaction (qMSP), bisulfite sequencing, pyrosequencing, nanopore sequencing, MassArray, methylation sensitive single nucleotide primer extension (Ms-SnuPE), reduced representation bisulfite sequencing (RRBS), methylation specific denaturing gradient gel electrophoresis (MS-DGGE), methylation specific melting analysis (MS-MCA), methylation specific high resolution melting analysis (MS-HRM), combined bisulfite-restriction NMR (COBRA), or a combination of ... quantitative methylation specific polymerase chain reaction (qMSP), quantitative methylation specific polymerase chain reaction (qMSP), quantitative methylation specific polymerase chain reaction (qMSP), quantitative methylation specific polymerase chain reaction (qMSP analysis), and a sequencing technique selected from the group consisting of Methylight assay, methylation-specific restriction endonuclease analysis (MSRE), methylation-sensitive restriction enzyme sequencing (MRE-seq), restriction landmark genomic scanning (RLGS), methylated DNA immunoprecipitation (MeDIP), or MeDIP-seq, methyl capture using methyl-CpG binding domain (MBD) proteins, ChIP assay, methylation array, multi-walled carbon nanotubes bearing a choline chloride monolayer (MWCNTs / ch / CGE), and oxidative chemical degradation-based analysis (NaIO4 / LiBr).
[0171] In certain embodiments, the methylation pattern is determined by bisulfite sequencing. In some embodiments, bisulfite sequencing comprises treating the sample with bisulfite and sequencing the bisulfite-treated sample by PCR. In certain embodiments, the bisulfite-treated sample is sequenced using a kit that may be, but is not limited to, a kit produced by Illumina. More specifically, the bisulfite-treated sample is sequenced using a kit selected from the group consisting of MiSeq reagent Kit v3-600-cycles (#MS-102-3003, Illumina), MiSeq reagent Kit v2-500-cycles (# MS-102-2003, Illumina) and MiSeq reagent Nano Kit v2-500 cycles (# MS-103-1003, Illumina).
[0172] Sequencing of the sample can be performed using any method known in the art. Sequencing platforms include, but are not limited to, Roche, Illumina, Life Technologies, Polonator, Helicos Bioscience, Pacific Biosciences, HTG Molecular Diagnostic, Singular Genomics, Element Biosciences, Oxford Nanopore, and Nanostring Technology.
[0173] In some embodiments, determining the methylation pattern comprises a step of library quantification. In some embodiments, the library quantification is performed using a fluorometric quantification method. Alternatively, the quantification of the methylation pattern is determined using fluorescence. In a particular embodiment, the fluorometric quantification method is characterized by using a kit comprising a dsDNA binding dye. In a particular embodiment, the library quantification is performed using a Qubit® 3.0 fluorometer produced by Thermo Fisher Scientific and a kit produced by Thermo Fisher Scientific. More specifically, the library quantification is performed using a kit selected from the group consisting of Qubit™ dsDNA HS Assay Kit (# Q32854, Thermo Fisher Scientific and Qubit™ dsDNA BR Assay Kit and # Q32850, Thermo Fisher Scientific). More specifically, quantification of the libraries is performed using a kit selected from the group consisting of Qubit™ dsDNA Assay Kits (# Q32854, Thermo Fisher Scientific and Qubit™ dsDNA BR Assay Kit and # Q32850, Thermo Fisher Scientific).
[0174] Quantification of libraries using the fluorometric quantification method of the present invention can be performed using other brand fluorometers and kits containing dsDNA binding dyes other than Thermo Fisher Scientific. Examples of other suitable fluorometers include, but are not limited to, the QFX fluorometer produced by DeNovix and the Quantus™ fluorometer produced by Promega. Examples of other suitable dsDNA fluorescence kits include, but are not limited to, the QuantiFluor® Dye Systems and QuantiFluor® dsDNA produced by Promega. In addition, the QFX fluorometer from DeNovix works with DeNovix's own DeNovix dsDNA Fluorescence Quantification Kit and any of the other common commercially available assays.
[0175] In the present example, the analysis to compare the methylation levels between each methylation site is performed using the Dispersion Shrinkage for Sequencing data (DSS) Bioconductor package. In other embodiments, any other suitable method known in the art may be used. In some embodiments, the analysis is performed using a β-binomial based model. In some embodiments, the analysis is performed using a non-β-binomial based model.
[0176] In the present example, the threshold p-value for establishing differential methylation is 0.05. In other embodiments, the threshold p-value is established at 0.25. In other embodiments, the threshold p-value is established at 0.2. In other embodiments, the threshold p-value is established at 0.15. In other embodiments, the threshold p-value is established at 0.1. In other embodiments, the threshold p-value is established at 0.09. In other embodiments, the threshold p-value is established at 0.08. In other embodiments, the threshold p-value is established at 0.07. In other embodiments, the threshold p-value is established at 0.06. In other embodiments, the threshold p-value is established at 0.05. In other embodiments, the threshold p-value is established at 0.04. In other embodiments, the threshold p-value is established at 0.03. In other embodiments, the threshold p-value is established at 0.02. In other embodiments, the threshold p-value is established at 0.01.
[0177] In another aspect, the present invention provides a classification model for use in identifying a human subject at risk for developing ADD, using a classification model configured to combine the methylation pattern of (a) the D-loop region and / or (b) the ND1 gene of mitochondrial DNA of a sample obtained from the subject with at least one clinical variable of the subject as described herein: (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 The present invention relates to a methylation site panel comprising at least one site selected from the group consisting of:
[0178] Clinical variables The classification model and method described herein, for example, the method for determining the risk of developing ADD in a subject, uses input data including clinical variables of the subject.Furthermore, the classification model is trained with a training data set that includes clinical variables related to a plurality of subjects.
[0179] The clinical variable may include any clinical variable known in the art. In some embodiments, the method comprises combining at least one clinical variable selected from the group consisting of sum of box scores (SOB), mini-mental state examination (MMSE), positron emission tomography (PET), presence or absence of beta amyloid protein, sex, age, apolipoprotein E genotype level (APOE), apolipoprotein reclassified genotype level (alE4), beta-amyloid-42 protein, beta-amyloid-40 protein, tau-T protein, tau-P protein, glial fibrillary acidic protein (GFAP), chitinase-3-like protein 1 (YKL-40), p53, and neurofilament light chain (NfL) with the methylation pattern described herein.
[0180] In some embodiments, the method comprises combining a methylation pattern described herein with at least one clinical variable of the subject selected from the group consisting of SOB, MMSE, presence or absence of beta amyloid protein, sex, age, APOE, alE4, Aβ-40, Aβ-42, tau-T, and tau-P. In some embodiments, the method comprises combining a methylation pattern described herein with at least one clinical variable of the subject selected from the group consisting of SOB, MMSE, PET, sex, age, APOE, alE4, Aβ-42, tau-T, and tau-P.
[0181] In another embodiment, the at least one clinical variable of the subject is selected from the group consisting of gender, SOB, MMSE, PET, the presence or absence of beta amyloid protein, age, APOE, and alE4.
[0182] In a particular embodiment, at least one clinical variable of the subject is selected from the group consisting of SOB, MMSE, PET, sex, age, APOE, and alE4. In another embodiment, at least one clinical variable of the subject is selected from the group consisting of SOB, MMSE, sex, age, APOE, and alE4. In another embodiment, at least one clinical variable of the subject is selected from the group consisting of SOB, MMSE, sex, and age. In another embodiment, at least one clinical variable is the presence or absence of beta amyloid protein. In a particular embodiment, at least one clinical variable is SOB. In another embodiment, at least one clinical variable is MMSE. In another embodiment, the clinical variable is at least SOB and MMSE. In another embodiment, at least one variable is PET, in particular beta amyloid-PET.
[0183] "Sex" is a categorical variable with two possible categories: female or male.
[0184] "Age" is a numerical variable corresponding to the subject's age.
[0185] "APOE" is a categorical variable that includes categories corresponding to different genotype levels of apolipoprotein E: E2.E2, E2.E3, E2.E4, E3.E3, E3.E4, and E4.E4.
[0186] "alE4" is a categorical variable corresponding to the reclassification of APOE genotype into the following categories: 0 (including E2.E2, E2.E3, and E3.E3), 1 (including E2.E4 and E3.E4), and (including E4.E4).
[0187] "SOB" is a numerical variable that refers to the "sum of box scores," a score ranging from 0 to 18 obtained by summing each of the above-mentioned domain box scores for the calculation of the CDR global score.
[0188] "MMSE" is a numerical variable that refers to the "Mini-Mental State Examination," an assessment of five main items: orientation, gaze, concentration and calculation, memory and language, and interpretation, with an output score of 1 to 30.
[0189] "PET" refers to positron emission tomography, a type of nuclear medicine procedure that measures the metabolic activity of cells in body tissues. PET can be performed with different types of tracers, each used for a specific purpose, test, or detection. For example, PET may measure glucose levels, beta amyloid plaques, or tau protein. FDG-PET refers to a PET designed to detect glucose and consequently analyze the metabolic activity of a tissue or body part. In another example, PET is used to determine the presence of beta amyloid plaques in the brain. In one embodiment, PET is a categorical variable that refers to the presence or absence of beta amyloid determined via positron emission tomography.
[0190] "Aβ-42" is a categorical variable that refers to the presence or absence of the protein beta amyloid-42 (Aβ-42) in a sample of cerebrospinal fluid or any other suitable biological sample or fluid, such as blood, or via radioimaging techniques (e.g., PET).
[0191] "Aβ-40" is a categorical variable that refers to the presence or absence of the protein beta amyloid-40 (Aβ-40) in a sample of cerebrospinal fluid or any other suitable biological sample or fluid, such as blood, or via radioimaging techniques (e.g., PET).
[0192] "Tau-T" is a categorical variable that refers to the presence or absence of the protein tau in a sample of cerebrospinal fluid or any other suitable biological sample or fluid, such as blood, or via radioimaging techniques.
[0193] "Tau-P" is a categorical variable that refers to the presence or absence of phosphorylated protein Tau in a sample of cerebrospinal fluid or any other suitable biological sample or fluid, such as blood, or via radioimaging techniques. The protein Tau-P can be phosphorylated at one or more positions, such as 181, 217, 231, among other options (i.e., p-tau-181, p-tau-217, p-tau-231).
[0194] "GFAP" is a categorical variable that refers to the presence or absence of glial fibrillary acidic protein in a sample of cerebrospinal fluid or any other suitable biological sample or fluid, such as blood.
[0195] "YKL-40" is a categorical variable that refers to the presence or absence of chitinase-3-like protein in a sample of cerebrospinal fluid or any other suitable biological sample or fluid, such as blood.
[0196] "p53" is the gene encoding the tumor protein p53, which can adopt multiple structural and functional states, including altered conformational states that may contribute to the development of neurodegenerative diseases such as AD. Thus, "p53" is a categorical variable that refers to the detection of conformational variants of p53 associated with the development of Alzheimer's disease.
[0197] "NfL" is a categorical variable that refers to the presence or absence of neurofilament light chain (NfL) in plasma or any other appropriate biological sample or fluid.
[0198] In some embodiments, the clinical variable comprises the determination of the presence or absence of β-amyloid protein. The presence or absence of β-amyloid protein can be determined using different techniques, including PET scan, analysis of cerebrospinal fluid (CSF), retinal screening and blood test. Thus, in some embodiments, the clinical variable comprises the presence or absence of β-amyloid protein. In particular, the clinical variable comprises the presence or absence of β-amyloid protein determined by PET scan, analysis of CSF, blood test, and / or retinal screening. In certain embodiments, the clinical variable comprises the presence or absence of β-amyloid protein determined by PET scan and / or analysis of CSF.
[0199] As described above, clinical variables may include biomarkers known in the art, as well as other neuropsychological tests and radioimaging techniques known in the art. In some embodiments, clinical variables include the presence, absence, or levels of biomarkers known in the art (e.g., Aβ-42). In some embodiments, clinical variables include the presence or absence of biomarkers known in the art.
[0200] In some embodiments, the biomarkers may be measured in any sample or fluid of the subject. In certain embodiments, the biomarkers are measured in cerebrospinal fluid (CSF) samples and / or blood samples. In particular, the biomarkers are measured in cerebrospinal fluid (CSF) samples and / or blood samples and are selected from the group consisting of Aβ-40, Aβ-42, neurofilament light chain (NfL), tau-T, tau-P (e.g., p-tau-181, p-tau-217, p-tau-231, etc.), GFAP, p-53, and / or YKL-40.
[0201] In other embodiments, the biomarkers may be detected using radioimaging techniques, particularly positron emission tomography (PET). In another embodiment, the biomarkers are detected using radioimaging techniques, and the biomarkers are selected from the group consisting of beta amyloid protein and tau protein, particularly Aβ-40, Aβ-42, tau-T, and tau-P (e.g., p-tau-181, p-tau-217, p-tau-231, etc.). In some embodiments, the clinical variables include data obtained via invasive methods, such as PET or Aβ-40, Aβ-42, tau-T, tau-P, and GFAP, which may require analysis of cerebrospinal fluid samples. In other embodiments, the clinical variables include only data obtained via non-invasive methods.
[0202] In some embodiments, the clinical variables may include other neuroimaging techniques. In other embodiments, the clinical variables may be derived from medical imaging, such as PET or magnetic resonance imaging (MRI). In some embodiments, PET is used to measure glucose, beta amyloid protein, and / or tau protein. In certain embodiments, PET is used to measure beta amyloid protein.
[0203] In other embodiments, the clinical variables include retinal screening. In particular, the retinal screening determines the presence or absence of a biomarker. More particularly, the retinal screening determines the presence or absence of beta amyloid protein.
[0204] In some embodiments, the clinical variables may include variables relating to the subject's current treatment or medication, particularly for dementia or dementia-related symptoms.
[0205] In one embodiment, a method for determining a subject's risk of developing Alzheimer's disease dementia comprises: a) determining a methylation pattern of (a) the D-loop region and / or (b) the ND1 gene of mitochondrial DNA of a sample obtained from the subject, the methylation pattern being: (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 and determining at least one site selected from the group consisting of: b) determining a risk score indicative of risk of developing Alzheimer's disease dementia, the risk score being calculated using a classification model configured to combine the methylation pattern of the one or more sites determined in step (a) with at least one clinical variable of the subject selected from the group consisting of sex, total box score, Mini-Mental State Examination, Positron Emission Tomography, presence or absence of beta amyloid protein, genotype, age, apolipoprotein E genotype level, and reclassified apolipoprotein E genotype level; Includes.
[0206] Samples and sample processing The term "biological sample" or "sample" as used herein refers to biological material isolated from a subject. A biological sample can include any biological material suitable for determining methylation patterns, for example, by processing and sequencing nucleic acids.
[0207] In some embodiments, the sample is selected from a biological fluid or a biopsy of a solid tissue. In particular, the sample is selected from the group consisting of blood, plasma, saliva, cerebrospinal fluid, a brain sample, a skin sample, and urine. In particular, the sample is blood, in particular peripheral blood.
[0208] The source of the sample can be solid tissue, e.g., from intact, frozen and / or preserved organs, tissue samples, biopsies, or aspirates. In some embodiments, the sample is acellular, e.g., comprising cell-free nucleic acid (e.g., DNA or RNA). The sample, in some embodiments, comprises compounds that are not naturally mixed with tissue in nature, such as preservatives, anticoagulants, buffers, fixatives, nutrients, antimicrobial agents, etc.
[0209] In some embodiments, the method includes obtaining a sample. In a particular embodiment, the sample is blood or plasma, and the sample is extracted using a needle. In another particular embodiment, the sample is saliva, and the sample is obtained using a method selected from the group consisting of drainage, spitting, aspiration, and swabbing. In some embodiments, the sample may be obtained, for example, from surgical material or a biopsy. In some embodiments, the biopsy may be preserved tissue from a previous selection therapy. In some embodiments, the biopsy may be from tissue that has not been treated.
[0210] In some embodiments, the sample is frozen or preserved. In some embodiments, the sample is preserved as a frozen sample or as a formalin, formaldehyde or paraformaldehyde fixed paraffin embedded (FFPE) tissue preparation. For example, the sample can be embedded in a matrix, such as an FFPE block or a frozen sample. In some embodiments, the sample can include bone marrow aspirate; scraping; bone marrow specimen; tissue biopsy specimen; surgical specimen, etc. In some embodiments, the sample is or includes cells obtained from an individual, such as from the individual from whom the sample is obtained.
[0211] In some embodiments, the sample is a fresh sample (or not a preserved sample) or a preserved sample. As used herein, the terms "fresh sample", "non-preserved sample" and grammatical variations thereof refer to a sample that has been treated prior to a predetermined time, e.g., one week after extraction from a subject. In some embodiments, a fresh sample is not frozen. In some embodiments, a fresh sample is not fixed. In some embodiments, a fresh sample is stored for less than about 2 weeks, less than about one week, or less than 6, 5, 4, 3, or 2 days prior to processing. As used herein, the term "preserved sample" and grammatical variations thereof refer to a sample that has been treated after a predetermined time, one week after extraction from a subject. In some embodiments, a preserved sample is frozen. In some embodiments, a preserved sample is fixed. In some embodiments, a preserved sample has a known diagnosis and / or treatment history. In some embodiments, a preserved sample is stored for at least one week, at least one month, at least six months, or at least one year prior to processing.
[0212] In another aspect, the present invention provides an enriched sample obtained from a subject at risk for developing ADD, comprising: (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 The present invention relates to an enriched sample comprising mitochondrial DNA suitable for use in determining the methylation pattern of at least one site selected from the group consisting of:
[0213] Oligonucleotides and Kits Oligonucleotides As discussed, in some embodiments, the methylation pattern is determined using at least one oligonucleotide that can specifically hybridize with the mitochondrial DNA sequence that comprises the D-loop region or the ND1 gene. In particular, the oligonucleotide can specifically hybridize under high stringency conditions.
[0214] The sequence of interest may refer to a reference sequence or a sequence resulting from a specific modification treatment, such as bisulfite treatment, in which unmethylated cytosine is modified to uracil. In some embodiments, the oligonucleotide hybridizes to a reference mitochondrial DNA sequence that includes the D-loop region or the ND1 gene. In other embodiments, the oligonucleotide hybridizes to a modified mitochondrial DNA sequence that includes the D-loop region or the ND1 gene. In certain embodiments, the modified mitochondrial DNA sequence is modified by bisulfite treatment. In more particular embodiments, the modified mitochondrial DNA sequence is modified by bisulfite treatment, in which unmethylated cytosine is modified to uracil.
[0215] In some embodiments, the methylation pattern is (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 The methylation site is determined using at least one oligonucleotide / primer capable of specifically hybridizing to mitochondrial DNA containing a methylation site selected from the group consisting of:
[0216] In some embodiments, the oligonucleotide is capable of specifically hybridizing with a mitochondrial DNA sequence comprising a D-loop region. In certain embodiments, the oligonucleotide is capable of specifically hybridizing with a mitochondrial DNA sequence comprising at least one site selected from the group consisting of the CpG site of the D-loop region shown in Table 1, the CHG site of the D-loop region shown in Table 3, and the CHH site of the D-loop region shown in Table 5. More specifically, the oligonucleotide is capable of specifically hybridizing with a mitochondrial DNA sequence comprising all of the CpG sites of the D-loop region shown in Table 1, the CHG sites of the D-loop region shown in Table 3, and the CHH sites of the D-loop region shown in Table 5.
[0217] In some embodiments, the oligonucleotide can specifically hybridize with the mitochondrial DNA sequence comprising ND1 gene.In certain embodiments, the oligonucleotide can specifically hybridize with the mitochondrial DNA sequence comprising at least one site selected from the group consisting of the CpG site of ND1 gene shown in Table 2, the CHG site of ND1 gene shown in Table 4, and the CHH site of ND1 region shown in Table 4.More specifically, the oligonucleotide can specifically hybridize with the mitochondrial DNA sequence comprising all sites of the CpG site of ND1 gene shown in Table 2, the CHG site of ND1 gene shown in Table 4, and the CHH site of ND1 region shown in Table 4.
[0218] In one embodiment, the oligonucleotide is a DNA sequence.
[0219] The inventors herein have designed primers that contain a minimum number of cytosines. In some embodiments, the primers are modified to include all possible methylation and non-methylation scenarios due to unknown C / U conversion of several cytosine residues contained in the sequence. These primers are a mixture of oligonucleotide sequences that contain several possible nucleotide bases at specific positions. As a result, the probability of detecting mitochondrial methylation is higher. The degenerated forward primer contains Y, which refers to either C or T (Y=C / T), at any position where the reference sequence is C. Furthermore, as is known in the art, the reverse primer does not correspond to the reference sequence, but corresponds to the inverted complementary sequence of the reference sequence. Thus, the reverse primer does not contain the C site of the reference sequence, but contains their complementary G site. Thus, the degenerated reverse primer contains R, which refers to either C or A (R=A / G), at any position where the reference sequence is C or the complementary sequence contains G.
[0220] In certain embodiments, the methylation pattern is determined using at least one oligonucleotide selected from the group consisting of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, and SEQ ID NO:4. The oligonucleotides may contain additional nucleotides at their ends to be suitable for use, for example, for sequencing (i.e., sequencing adapters). In certain embodiments, the methylation pattern is determined using at least one oligonucleotide having a length of 15-100 nucleotides and comprising a sequence selected from the group consisting of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, and SEQ ID NO:4.
[0221] In a more specific embodiment, the methylation pattern is determined using oligonucleotides having a length of 15 to 100 nucleotides and including the sequences SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, and SEQ ID NO:4. In a particular embodiment, the methylation pattern is determined using oligonucleotides SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, and SEQ ID NO:4.
[0222] One aspect of the invention relates to an oligonucleotide having a length of 15-100 nucleotides and comprising a sequence selected from the group consisting of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, and SEQ ID NO:4. In certain embodiments, the oligonucleotide is selected from the group consisting of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, and SEQ ID NO:4. As described, in one embodiment, the nucleic acid sequence comprises a sequencing adaptor at the end of SEQ ID NO:1, 2, 3, and / or 4. In certain embodiments, the oligonucleotide is SEQ ID NO:1. In another particular embodiment, the oligonucleotide is SEQ ID NO:2. In another particular embodiment, the oligonucleotide is SEQ ID NO:3. In another particular embodiment, the oligonucleotide is SEQ ID NO:4.
[0223] One aspect of the present invention relates to the use of an oligonucleotide having a length of 15 to 100 nucleotides and comprising a sequence selected from the group consisting of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, and SEQ ID NO: 4, for determining the methylation pattern of mitochondrial DNA.
[0224] In a particular embodiment, the present invention relates to the use of an oligonucleotide having a length of 15-100 nucleotides and comprising (particularly consisting of) the sequence SEQ ID NO: 1 for the determination of the methylation pattern of a mitochondrial DNA sequence comprising the D-loop region. In another embodiment, the present invention relates to the use of an oligonucleotide having a length of 15-100 nucleotides and comprising (particularly consisting of) the sequence SEQ ID NO: 2 for the determination of the methylation pattern of a mitochondrial DNA sequence comprising the D-loop region. In another embodiment, the present invention relates to the use of an oligonucleotide having a length of 15-100 nucleotides and comprising (particularly consisting of) the sequence SEQ ID NO: 3 for the determination of the methylation pattern of a mitochondrial DNA sequence comprising the ND1 gene. In another embodiment, the present invention relates to the use of an oligonucleotide having a length of 15-100 nucleotides and comprising (particularly consisting of) the sequence SEQ ID NO: 4 for the determination of the methylation pattern of a mitochondrial DNA sequence comprising the ND1 gene.
[0225] In some embodiments, the present invention relates to the use of an oligonucleotide having a length of 15 to 100 nucleotides and comprising a sequence selected from the group consisting of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, and SEQ ID NO:4, for determining the methylation pattern of mitochondrial DNA to determine the risk of developing ADD in a subject.
[0226] kit In one aspect, the present invention relates to a kit comprising at least one oligonucleotide capable of specifically hybridizing with a mitochondrial DNA sequence comprising the D-loop region or the ND1 gene. In some embodiments, the kit comprises the oligonucleotide defined above.
[0227] In particular, the kit comprises at least one oligonucleotide having a length of 15-100 nucleotides and comprising a sequence selected from the group consisting of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, and SEQ ID NO:4. In a particular embodiment, the present invention relates to a kit comprising an oligonucleotide having a length of 15-100 nucleotides and comprising the nucleic acid sequences of SEQ ID NO:1 and SEQ ID NO:2. In another embodiment, the present invention relates to a kit comprising an oligonucleotide having a length of 15-100 nucleotides and comprising the nucleic acid sequences of SEQ ID NO:3 and SEQ ID NO:4. In a particular embodiment, the present invention relates to a kit comprising an oligonucleotide having a length of 15-100 nucleotides and comprising the nucleic acid sequences of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, and SEQ ID NO:4.
[0228] In a particular embodiment, the kit comprises an oligonucleotide having a length of 15-100 nucleotides and comprising SEQ ID NO: 1 and SEQ ID NO: 2, for determining the methylation pattern of a mitochondrial DNA sequence comprising the D-loop region. In another embodiment, the present invention comprises an oligonucleotide having a length of 15-100 nucleotides and comprising SEQ ID NO: 3 and SEQ ID NO: 4, for determining the methylation pattern of a mitochondrial DNA sequence comprising the ND1 gene.
[0229] In one aspect, the present invention relates to the use of the kit as defined above for determining the methylation pattern of mitochondrial DNA.In another aspect, the present invention relates to the use of the kit as defined above for determining the methylation pattern of mitochondrial DNA to determine the risk of developing ADD in a subject.In another aspect, the present invention relates to the use of the kit as defined above according to the method described herein.
[0230] Such kits can include multiple containers, each containing one or more of the various reagents (e.g., in concentrated form) utilized in the methods, including, for example, one or more oligonucleotides (e.g., oligonucleotides having SEQ ID NOs: 1-4 provided herein), and the kits can provide reagents, buffers, and / or equipment to aid in the practice of the methods provided herein.
[0231] The kit provided by the present invention may also include a pamphlet or instructions that explain the methods disclosed herein, or their practical application for determining the risk of developing ADD in a subject. The instructions included in the kit may be attached to packaging materials or may be included as a package insert. The instructions are typically written or printed materials, but are not limited to such. Any medium capable of storing such instructions and transmitting them to an end user is contemplated. Such medium includes, but is not limited to, electronic storage media (e.g., magnetic disks, tapes, cartridges, chips), optical media (e.g., CD-ROMs), and the like. As used herein, the term "instructions" may include the address of an internet site that provides instructions.
[0232] In some embodiments, the kit is an Illumina sequencing kit.More specifically, the kit is selected from the group consisting of MiSeq reagent Kit v3-600-cycles (# MS-102-3003, Illumina), MiSeq reagent Kit v2-500-cycles (# MS-102-2003, Illumina) and MiSeq reagent Nano Kit v2-500 cycles (# MS-103-1003, Illumina).
[0233] In another aspect, the present invention provides a method for producing a composition comprising: a) a reagent for determining a methylation pattern of (a) the D-loop region and / or (b) the ND1 gene of mitochondrial DNA of a sample obtained from a subject, the methylation pattern being (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 A reagent, wherein the at least one site is determined from the group consisting of: b) Optionally, instructions for using the reagent; The present invention relates to a kit comprising:
[0234] Comparative diagnosis system The methods disclosed herein may be provided as comparative diagnostics, available, for example, via a web server, to inform a clinician or patient of potential treatment options or for patient selection for clinical trials. The methods disclosed herein may include collecting or otherwise obtaining a biological sample and performing an analytical method disclosed herein to determine a subject's risk of developing ADD.
[0235] In one aspect of the invention, there is provided a computing system comprising suitable means for performing any of the computer-implemented methods described herein.
[0236] At least some embodiments of the methods described herein may be implemented using a computer due to the complexity of the computations involved. In some embodiments, a computer system includes hardware components electrically connected via a bus, including a processor, input devices, output devices, storage, computer-readable storage medium readers, communication systems, processing acceleration (e.g., DSPs or dedicated processors), and / or memory. The computer-readable storage medium readers may further be connected to computer-readable media, this combination being inclusive of remote storage, local storage, fixed storage, and / or removable storage+storage media, memory, etc. for temporarily and / or more permanently containing computer-readable information, which may include storage, memory, and / or any other such accessible system resources.
[0237] A single architecture may be utilized to implement one or more servers that may be further configured according to currently desired protocols, protocol variations, extensions, and the like. However, it will be apparent to one of ordinary skill in the art that embodiments may be better utilized according to more specific application requirements. Customized hardware may also be utilized, and / or particular components may be implemented in hardware, software, firmware, or a combination thereof. Additionally, connections to other computing devices, such as network input / output devices (not shown), may be used, although it should be understood that wired, wireless, modem, and / or other connections to other computing devices may also be utilized.
[0238] In one embodiment, the system further includes one or more devices for providing input data to the one or more processors. The system further includes a memory for storing the data set of ranked data elements. In another embodiment, the device for providing input data includes a detector for detecting features of the data elements, such as a fluorescent plate reader, a mass spectrometer, or a gene chip reader.
[0239] The system may further include a database management system. A user's request or query may be formatted in an appropriate language understood by the database management system, which processes the query to extract relevant information from the training set database. The system may be connectable to a network to which a network server and one or more clients are connected. The network may be a local area network (LAN) or a wide area network (WAN), as known in the art. In particular, the server includes the necessary hardware to execute a computer program product (e.g., software) that accesses data from the database to process the user's request. The system may communicate with an input device to provide data (e.g., methylation patterns) regarding the data elements to the system.
[0240] In a further aspect, the present invention is directed to a computer program product comprising instructions which, when the program is executed by a computer, cause the computer to perform any of the computer-implemented methods described herein.
[0241] Some embodiments described herein may be implemented to include a computer program product. The computer program product may include a computer readable medium having computer readable program code embedded therein for executing an application program on a computer with a database. As used herein, a "computer program product" refers to an organized set of instructions in the form of statements of a natural language or programming language that are contained on a physical medium of any nature (e.g., written, electronic, magnetic, optical, or other nature) and that can be used by a computer or other automated data processing system. Such programming language statements, when executed by a computer or data processing system, cause the computer or data processing system to operate according to the specific content of the statements.
[0242] In some embodiments, the invention is a computer program product comprising a computer readable medium embodied with program code executable by a processor of a computing device or system, the program code comprising code for executing a classification model (or other method described herein) for identifying a human subject at risk for developing ADD, configured, for example, to combine a methylation pattern with at least one clinical variable as described herein, the methylation pattern comprising a methylation pattern of (a) the D-loop region and / or (b) the ND1 gene of mitochondrial DNA of a sample obtained from the subject, the methylation pattern comprising: (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 determined at at least one site selected from the group consisting of: Applies to computer program products.
[0243] In one embodiment, the invention is a computer program product comprising a computer readable medium embodied with program code executable by a processor of a computing device or system, the program code comprising code for executing, for example, a classification model (or other method described herein) for identifying a human subject at risk of developing ADD, the model configured to identify a human subject at risk of developing ADD, the classification model configured to combine a methylation pattern with at least one clinical variable of the subject as described herein, the methylation pattern comprising a methylation pattern of (a) the D-loop region and / or (b) the ND1 gene of mitochondrial DNA of a sample obtained from the subject, the methylation pattern being: (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 determined at at least one site selected from the group consisting of: Related to computer program products.
[0244] Computer program products include, but are not limited to, programs in source and object code embedded in a computer readable medium and / or test or data libraries. Furthermore, computer program products that enable a computer system or data processing system to operate in a preselected manner may be provided in many forms, including, but not limited to, original source code, assembly code, object code, machine language, encrypted or compressed versions of the above, and any equivalents. In one aspect, computer program products are provided to implement the treatment, diagnosis, methods disclosed herein, for example, to determine whether to perform a particular treatment based on the score obtained.
[0245] A computer program product is a program product comprising: (a) code for obtaining data attributable to a biological sample of a subject, the data comprising a methylation pattern corresponding to the methylation sites of Tables 1-6 in the biological sample, or a methylation pattern can be derived from the data; and (b) code for implementing a classification method, for example, indicating whether a therapeutic agent should be administered to a patient in need thereof based on the score obtained; The present invention includes a computer-readable medium having program code embodied therein that is executable by a processor of a computing device or system, including:
[0246] While various embodiments are described as methods or apparatus, it should be understood that the embodiments may be implemented via code in connection with a computer, such as code residing on or accessible by a computer. For example, software and databases may be utilized to perform many of the methods discussed above. Thus, it should be noted that in addition to embodiments that are accomplished through hardware, these embodiments may be accomplished through the use of an article of manufacture consisting of a computer usable medium having computer readable program code embodied therein that enables the functionality disclosed in this description.
[0247] Further, some embodiments may be code stored in virtually any type of computer readable memory, including but not limited to RAM, ROM, magnetic media, optical media, or magneto-optical media. Even more generally, some embodiments may be implemented in software, including but not limited to software running on a general purpose processor, microcode, PLA, or ASIC, or in hardware, or any combination thereof.
[0248] It is also contemplated that some embodiments may be achieved as a computer signal embodied in a carrier wave and a signal (e.g., electrical and optical signals) propagated through a transmission medium. Thus, the various types of information discussed above may be formatted in structures, such as data structures, and transmitted as an electrical signal through a transmission medium or stored on a computer-readable medium.
[0249] Performance of the method In some embodiments, the sample may be requested, for example, by a health care provider (e.g., a doctor) or a health care benefit provider, or may be obtained and / or processed by the same or a different health care provider (e.g., a nurse, a hospital), or a clinical laboratory, and after processing, the results may be sent to the original health care provider or to an additional health care provider, a health care benefit provider, or the patient. Similarly, the determination of the methylation pattern disclosed herein; the application of the classification model; the determination of the score; the treatment decision; the clinical trial participation decision; or combinations thereof may be performed by one or more health care providers, health care benefit providers, and / or clinical laboratories.
[0250] As used herein, the term "healthcare provider" refers to an individual or institution that directly interacts with or administers to a living subject, such as a human patient. Non-limiting examples of healthcare providers include doctors, nurses, technicians, therapists, pharmacists, counselors, other medical professionals, medical facilities, doctor's offices, hospitals, emergency rooms, clinics, urgent care centers, other medical clinics / facilities, and any other entity that provides general and / or specialized treatment, diagnosis, evaluation, maintenance, treatment, medication, and / or advice related to all or part of a patient's health condition, including, but not limited to, medical, specialty care, surgery, and / or any other type of treatment, diagnosis, evaluation, maintenance, treatment, medication, and / or advice. Healthcare provider, as used herein, also refers to pharmaceutical companies or their providers / intermediaries (e.g., CROs) involved in the development of clinical trials.
[0251] As used herein, the term "clinical laboratory" refers to a facility for the investigation or processing of materials derived from subjects. These investigations may also include procedures for collecting or otherwise obtaining samples, preparing, determining, measuring, or describing the presence or absence of various substances (e.g., mtDNA methylation patterns or biomarkers used herein as clinical variables) in the subject's body or in samples obtained from the subject's body. Investigations may also include procedures such as medical imaging procedures (e.g., PET, MRI) to obtain clinical variable data.
[0252] As used herein, the term "healthcare benefits provider" includes a separate entity, organization, or group that provides, offers, furnishes, pays in whole or in part, or is involved in the provision of patient access to one or more health care benefits, benefit plans, health insurance, and / or health care accounting programs.
[0253] The healthcare provider performs the following actions: obtaining samples / clinical variables, processing samples / clinical variables, submitting samples / clinical variables, receiving samples / clinical variables, transferring samples / clinical variables, analyzing or measuring samples / clinical variables (e.g., thereby obtaining a methylation pattern), quantifying samples / clinical variables, providing results obtained after analysis / measurement / quantification of samples / clinical variables, receiving results obtained after analysis / measurement / quantification of samples / clinical variables, applying a classification model, scoring results obtained after analysis / measurement / quantification of one or more sample / clinical variables, and providing a score from one or more samples. The patient may perform or instruct another health care provider or the patient to obtain a score from one or more samples, administer a treatment, initiate administration of a treatment, stop administration of a treatment, continue administration of a treatment, temporarily discontinue administration of a treatment, increase the amount of a therapeutic agent administered, decrease the amount of a therapeutic agent administered, continue administration of an amount of a therapeutic agent, decrease the frequency of administration of a therapeutic agent, maintain the same administration frequency for a therapeutic agent, replace a treatment or therapeutic agent with at least another treatment or therapeutic agent, or combine a treatment or therapeutic agent with at least another treatment or additional therapeutic agents.
[0254] In some embodiments, the healthcare benefits provider may perform a number of steps, such as collecting the sample / clinical variables, processing the sample / clinical variables, submitting the sample / clinical variables, receiving the sample / clinical variables, shipping the sample / clinical variables, analyzing or measuring the sample / clinical variables (e.g., to obtain a methylation pattern), quantifying the sample / clinical variables, applying a classification model, providing results obtained after analyzing / measuring / quantifying the sample / clinical variables, communicating results obtained after analyzing / measuring / quantifying one or more sample / clinical variables, scoring results obtained after analyzing / measuring / quantifying one or more sample / clinical variables, and / or performing a classification process for the sample / clinical variables. The health care provider may authorize or deny communication of the scores of the above samples / clinical variables, administration of a treatment or therapeutic agent, initiation of administration of a treatment or therapeutic agent, stopping administration of a treatment or therapeutic agent, continuing administration of a treatment or therapeutic agent, temporarily interrupting administration of a treatment or therapeutic agent, increasing the amount of a therapeutic agent administered, decreasing the amount of a therapeutic agent administered, continuing administration of an amount of a therapeutic agent, increasing the frequency of administration of a therapeutic agent, decreasing the frequency of administration of a therapeutic agent, maintaining the same administration frequency for a therapeutic agent, substituting the treatment or therapeutic agent with at least another treatment or therapeutic agent, or combining the treatment or therapeutic agent with at least another treatment or additional therapeutic agent. Additionally, the health care provider may, for example, authorize or deny prescription of a treatment, authorize or deny coverage of a treatment, authorize or deny reimbursement for the cost of a treatment, determine or deny eligibility for a treatment, etc.
[0255] In some embodiments, a clinical laboratory may, for example, collect or obtain samples / clinical variables, process samples / clinical variables, submit samples / clinical variables, receive samples / clinical variables, ship samples / clinical variables, analyze or measure samples / clinical variables (e.g., to obtain methylation patterns), quantify samples / clinical variables, apply classification models, provide results obtained after analyzing / measuring / quantifying samples / clinical variables, receive results obtained after analyzing / measuring / quantifying samples / clinical variables, score results obtained after analyzing / measuring / quantifying one or more sample / clinical variables, provide scores from one or more sample / clinical variables, obtain scores from one or more sample / clinical variables, or perform other related activities.
[0256] In some embodiments, the sample / clinical variables can be obtained by the healthcare professional treating or diagnosing the patient, by the healthcare provider, or by a clinical laboratory. The measurement of the sample (e.g., by using certain assays described herein) and obtaining the clinical variables (e.g., by medical imaging techniques) can be performed by the same or different healthcare provider or clinical laboratory from the healthcare provider or clinical laboratory that obtained the sample / clinical variables. The classification model can be applied by the healthcare provider or a different healthcare provider or clinical laboratory. The obtained scores and results are finally sent to the first healthcare professional or healthcare provider that treated or diagnosed the patient. Thus, in some embodiments, the healthcare provider or clinical laboratory can advise the healthcare professional / provider regarding the diagnosis or whether the patient will benefit from the treatment. In some embodiments, the healthcare provider is a pharmaceutical company or one of its providers / intermediaries (e.g., CROs) involved in the development of clinical trials. All steps described herein can be performed by the pharmaceutical company and / or one of its providers / intermediaries (e.g., CROs) or can be partially performed, for example, by a clinical laboratory or a different healthcare provider.
[0257] Specific Embodiments As will be apparent to one of ordinary skill in the art upon reading this description, each of the individual embodiments described and illustrated herein has separate components and features that may be combined with the features of any of the other several embodiments without departing from the scope or spirit of the invention. Specific combinations of the above embodiments, detailed in different sections, are described herein.
[0258] In one embodiment, the present invention provides a method for determining a subject's risk of developing ADD, comprising: a) determining a methylation pattern of the D-loop region of mitochondrial DNA and / or the ND1 gene in a sample of a subject comprising mitochondrial DNA, the methylation pattern being: (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 determined at at least one site selected from the group consisting of: Steps and b) combining the methylation pattern of one or more sites determined in step (a) with at least one clinical variable of the subject as described herein, said combining being performed using a classification model to determine a risk score that correlates with the subject's risk of developing ADD; The present invention relates to a method comprising the steps of:
[0259] In one embodiment, the present invention provides a method for determining a subject's risk of developing ADD, comprising: a) determining a methylation pattern of (a) the D-loop region and / or (b) the ND1 gene of mitochondrial DNA of a sample obtained from the subject, the methylation pattern being: (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 determined at at least one site selected from the group consisting of: Steps and b) determining a risk score indicative of the risk of developing ADD, the risk score being calculated using a classification model configured to combine the methylation pattern of the one or more sites determined in step (a) with at least one clinical variable of the subject as described herein; The present invention relates to a method comprising the steps of:
[0260] In one embodiment, the present invention provides a method for determining a subject's risk of developing ADD, comprising: Determining a methylation pattern of (a) the D-loop region and / or (b) the ND1 gene of mitochondrial DNA of a sample obtained from the subject, the methylation pattern being: (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 determined at at least one site selected from the group consisting of: Including steps, The methylation pattern is combined with at least one clinical variable of the subject described herein to determine a risk score indicative of the risk of developing ADD using a classification model. It concerns the method.
[0261] In one embodiment, the present invention provides a method for determining a subject's risk of developing ADD, comprising: determining a risk score indicative of risk for developing ADD using a classification model configured to combine at least one clinical variable and a methylation pattern as described herein; The methylation pattern comprises a methylation pattern of (a) a D-loop region and / or (b) an ND1 gene of mitochondrial DNA of a sample obtained from the subject, the methylation pattern being: (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 determined at at least one site selected from the group consisting of: It concerns the method.
[0262] In one embodiment, the invention provides a method of treating a subject having or at risk of developing AD, comprising: a) determining a methylation pattern of (a) the D-loop region and / or (b) the ND1 gene of mitochondrial DNA of a sample obtained from the subject, the methylation pattern being: (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 determined at at least one site selected from the group consisting of: Steps and b) determining a risk score indicative of the risk of developing ADD, the risk score being calculated using a classification model configured to combine the methylation pattern of the one or more sites determined in step (a) with at least one clinical variable of the subject as described herein; c) administering a treatment to the subject if the risk score indicates that the subject is at risk for developing ADD. The present invention relates to a method comprising the steps of:
[0263] In one embodiment, the invention provides a method of treating a subject having or at risk of developing AD, comprising: a) determining a methylation pattern of (a) the D-loop region and / or (b) the ND1 gene of mitochondrial DNA of a sample obtained from the subject, the methylation pattern being: (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 and wherein the at least one site selected from the group consisting of The methylation pattern is combined with at least one clinical variable of the subject described herein to determine a risk score indicative of the risk of developing ADD using a classification model. Steps and b) administering a treatment to the subject if the risk score indicates that the subject is at risk for developing ADD. The present invention relates to a method comprising the steps of:
[0264] In one embodiment, the invention provides a method of treating a subject having or at risk of developing AD, comprising: a) determining a risk score indicative of the risk of developing ADD using a classification model configured to combine at least one clinical variable and a methylation pattern of a subject as described herein, The methylation pattern comprises a methylation pattern of (a) a D-loop region and / or (b) an ND1 gene of mitochondrial DNA of a sample obtained from the subject, the methylation pattern being: (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 and determining at least one site selected from the group consisting of: b) administering a treatment to the subject if the risk score indicates that the subject is at risk for developing ADD. The present invention relates to a method comprising the steps of:
[0265] In one embodiment, the invention provides a method of treating a subject having or at risk of developing AD, comprising: administering a treatment to the subject if the risk score indicates that the subject is at risk for developing ADD; A risk score indicative of the risk of developing ADD is calculated using a classification model configured to combine the methylation pattern of (a) the D-loop region and / or (b) the ND1 gene of the mitochondrial DNA of the sample obtained from the subject with at least one clinical variable of the subject as described herein; The methylation pattern is (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 determined at at least one site selected from the group consisting of: , It concerns the method.
[0266] In one embodiment, the invention provides a method for identifying a human subject at risk for developing ADD, comprising: a) determining a methylation pattern of (a) the D-loop region and / or (b) the ND1 gene of mitochondrial DNA of a sample obtained from the subject, the methylation pattern being: (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 determined at at least one site selected from the group consisting of: Steps and b) determining a risk score indicative of the risk of developing ADD, the risk score being calculated using a classification model configured to combine the methylation pattern of the one or more sites determined in step (a) with at least one clinical variable of the subject as described herein; The present invention relates to a method comprising the steps of:
[0267] In one embodiment, the invention provides a method for identifying a human subject at risk for developing ADD, comprising: Determining a methylation pattern of (a) the D-loop region and / or (b) the ND1 gene of mitochondrial DNA of a sample obtained from the subject, the methylation pattern being: (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 and wherein the at least one site selected from the group consisting of The methylation pattern is combined with at least one clinical variable of the subject described herein to determine a risk score indicative of the risk of developing ADD using a classification model. The present invention relates to a method comprising the steps of:
[0268] In one embodiment, the invention provides a method for identifying a human subject at risk for developing ADD, comprising: determining a risk score indicative of risk of developing ADD using a classification model configured to combine the methylation pattern with at least one clinical variable of the subject described herein, The methylation pattern comprises a methylation pattern of (a) a D-loop region and / or (b) an ND1 gene of mitochondrial DNA of a sample obtained from the subject, the methylation pattern being: (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 determined at at least one site selected from the group consisting of: The present invention relates to a method comprising the steps of:
[0269] In one embodiment, the invention provides a method for selecting a human subject for treatment (e.g., prophylactic treatment) for AD, comprising: a) determining a methylation pattern of (a) the D-loop region and / or (b) the ND1 gene of mitochondrial DNA of a sample obtained from the subject, the methylation pattern being: (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 determined at at least one site selected from the group consisting of: Steps and b) determining a risk score indicative of the risk of developing ADD, the risk score being calculated using a classification model configured to combine the methylation pattern of the one or more sites determined in step (a) with at least one clinical variable of the subject as described herein; The present invention relates to a method comprising the steps of:
[0270] In one embodiment, the invention provides a method for selecting a human subject for treatment (e.g., prophylactic treatment) for AD, comprising: Determining a methylation pattern of (a) the D-loop region and / or (b) the ND1 gene of mitochondrial DNA of a sample obtained from the subject, the methylation pattern being: (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 and wherein the at least one site selected from the group consisting of The methylation pattern is combined with at least one clinical variable of the subject described herein to determine a risk score indicative of the risk of developing ADD using a classification model. The present invention relates to a method comprising the steps of:
[0271] In one embodiment, the invention provides a method for selecting a human subject for treatment (e.g., prophylactic treatment) for AD, comprising: determining a risk score indicative of risk of developing ADD using a classification model configured to combine the methylation pattern with at least one clinical variable of the subject described herein, The methylation pattern comprises a methylation pattern of (a) a D-loop region and / or (b) an ND1 gene of mitochondrial DNA of a sample obtained from the subject, the methylation pattern being: (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 determined at at least one site selected from the group consisting of: The present invention relates to a method comprising the steps of:
[0272] Alternatively, another aspect of the present invention is a combined biomarker for identifying a human subject at risk of developing ADD, the combined biomarker comprising a classification model configured to combine a methylation pattern with at least one clinical variable of the subject as described herein, The methylation pattern comprises a methylation pattern of (a) a D-loop region and / or (b) an ND1 gene of mitochondrial DNA of a sample obtained from the subject, the methylation pattern being: (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 determined at at least one site selected from the group consisting of: Concerning biomarkers.
[0273] In one embodiment, the present invention provides a classification model for identifying human subjects at risk of developing ADD, configured to combine a methylation pattern with at least one clinical variable of a subject as described herein, comprising: The methylation pattern comprises a methylation pattern of (a) a D-loop region and / or (b) an ND1 gene of mitochondrial DNA of a sample obtained from the subject, the methylation pattern being: (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 determined at at least one site selected from the group consisting of: Regarding classification models.
[0274] In one embodiment, the present invention provides a classification model for identifying a human subject at risk of developing ADD, the model being configured to identify a human subject at risk of developing ADD, the classification model being configured to combine a methylation pattern with at least one clinical variable of the subject as described herein, The methylation pattern comprises a methylation pattern of (a) a D-loop region and / or (b) an ND1 gene of mitochondrial DNA of a sample obtained from the subject, the methylation pattern being: (i) a CpG site in the D-loop region shown in Table 1; (ii) a CpG site in the ND1 gene shown in Table 2; (iii) a CHG site in the D-loop region shown in Table 3; (iv) a CHG site in the ND1 gene as shown in Table 4; (v) a CHH site in the D-loop region as shown in Table 5, and (vi) CHH site in the ND1 region shown in Table 6 determined at at least one site selected from the group consisting of: Regarding classification models.
[0275] For example, specific embodiments described in different sections of this document as classification models, clinical variables, samples, or means for determining methylation patterns apply equally to the above-mentioned embodiments.
[0276] Working Example Example 1: Detection of mtDNA methylation in blood samples 1.1 Materials and methods 1) Collection of blood samples Human blood samples were collected in EDTA tubes to prevent blood clotting. After obtaining the samples in the laboratory, the blood was either directly processed for DNA extraction or aliquoted and stored at -80°C until processing.
[0277] A total of 304 subjects were sampled, recruited from two different cohorts: (Cohort A corresponds to the Australian Imaging, Biomarker & Lifestyle Flagship Study of Ageing (AIBL) cohort, and Cohort B corresponds to MCI patients recruited between 2015 and 2019 at the Bellvitge Hospital in Barcelona). These subjects were classified into three groups: controls (35.5%), MCI subjects who did not progress to ADD (30.6%), and MCI subjects who progressed to ADD (33.9%).
[0278] Control subjects were available only in cohort A, and MCI subjects were available in both cohorts.
[0279] 2) Total DNA extraction Whole blood samples were processed to obtain the extraction of total DNA and the co-purification of both genomic and mitochondrial DNA. DNA was isolated from human whole blood samples using Wizard® Genomic DNA Purification Kit (# A1620, Promega) according to the manufacturer's instructions. Alternatively, samples were processed with Maxwell® RSC Instrument, which provides a simple method for efficient automated purification of DNA from samples. Capture, washing, and purification of DNA samples were performed using paramagnetic beads. Maxwell® RSC Blood DNA Kit (# AS1400, Promega) was used according to the manufacturer's specifications. The quality and quantity of purified DNA was determined using a Thermo Fisher Scientific NanoDropTM One spectrophotometer.
[0280] 3) Bisulfite treatment Bisulfite conversion consists of the deamination of unmodified cytosines to uracil, resulting in the intact modified base 5-mc, i.e., methylated cytosine. A sample of total DNA (300 ng) was treated with bisulfite reagent using the EZ DNA Methylation Kit (# D5001, Zymo Research) according to the manufacturer's protocol. To obtain good bisulfite conversion, the incubation conditions in step 2 of the protocol, consisting of a 15 min incubation at 37°C, were replaced by a 30 min incubation at 42°C, as indicated in Appendix 1.A of the manufacturer's protocol. These last conditions are recommended to minimize incomplete C to T conversion. The treated DNA was finally resuspended in 30 μL of nuclease-free water.
[0281] 4) Preparation of amplicon libraries The workflow for amplicon library construction was based on Illumina's "16S Metagenomic Sequencing Library Preparation Protocol," which can be used to sequence 16S rRNA genes and other targeted regions of amplicon sequences of interest. Preparation of the amplicon library allowed for the acquisition of amplicons of interest in mtDNA and their preparation for processing on the Illumina MiSeq System.
[0282] 4.1) First PCR: Amplicon PCR The mtDNA regions of interest were amplified by PCR with specific degenerate primers (see Results section 1.2.1) corresponding to the sequences of SEQ ID NO: 1-4, which further contained an overhanging Illumina adapter. When designing primers for the regions of interest, the overhanging adapter sequence had to be added to the locus-specific primers of the targeted region, as described in the Illumina protocol.
[0283] The Illumina® overhang adapter sequences added to the locus-specific sequences are (SEQ ID NOs: 5, 6): Forward overhang: 5' TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG-[locus-specific sequence] Reverse overhang: 5' GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG-[locus-specific sequence] It is.
[0284] Amplification of bisulfite converted DNA was performed using the FastStart™ High Fidelity PCR System (# 3553400001, Roche). The final PCR mixture (25 μL) contained 5 μL of bisulfite treated DNA, 1× FastStart Buffer # 2; 0.05 U FastStart HiFi Polymerase; 0.8 mM total dNTPs (0.2 mM each dNTP); and 0.4 μM each of forward and reverse primers. The ND1 amplicon reaction also contained 5% DMSO. The final volume was adjusted with nuclease-free water.
[0285] Amplification was carried out in an Applied Biosystems SimpliAmp™ Thermal Cycler.
[0286] To evaluate the resulting products of the first PCR, 3 μL of each PCR product was analyzed by electrophoresis on a 1.5% agarose gel stained with SybrSafe™ DNA Gel Stain (# S33102, Thermo Fisher Scientific) to confirm the presence of bands from the amplicons.
[0287] 4.2) First PCR Cleanup AMPure XP beads (#A63881, Beckman Coulter) were used to purify the amplicons and separate them from free primer and primer-dimer species. Steps were performed according to Illumina's "16S Metagenomic Sequencing Library Preparation Protocol". PCR amplicon products were purified using a 0.8x ratio of AMpure Beads. Elution of the beads was performed with 14 μL of Buffer EB (# 19086, Qiagen) and 12 μL was recovered from the beads.
[0288] 4.3) Second PCR: Index PCR An index PCR was performed to attach Illumina unique dual indexes (UDIs) and sequencing adapters. The PCR index reaction was performed in a 50 μL reaction containing 5 μL of DNA purified from the first PCR cleanup, 25 μL of KAPA HiFi HotStart Ready Mix (2x) (# 7958935001, Roche), 10 μL of Illumina UDIs (Unic dual indexes), and 10 μL of nuclease-free water. This was performed using an Applied Biosystems™ SimpliAmp™ Thermal Cycler.
[0289] 4.4) Second PCR Cleanup A second PCR cleanup was performed using AMPure XP beads to clean up the final live library before quantification. 50 μL of the second PCR reaction was purified following the steps described in Illumina's 16S Metagenomic Sequencing Library Preparation Protocol. A 1.12x ratio of AMPure XP beads was used and the final elution was performed with 27.5 μL of Buffer EB, and 25 μL was recovered from the beads.
[0290] 4.5) Quantifying, normalizing, and pooling libraries Libraries were quantified using a fluorometric quantification method using a dsDNA binding dye with a Thermo Fisher Scientific Qubit® 3.0 fluorometer. Quantification was performed using the Qubit™ dsDNA HS Assay Kit according to the manufacturer's instructions. After obtaining Qubit quantification values in ng / μL, DNA concentrations were calculated using the following formula: (concentration (ng / μL)) / (660 g / mol × average library size) × 10 6 = Concentrate (nM) The results were calculated in nM based on the size of the DNA amplicon as determined by Agilent Technologies 2100 Bioanalyzer traces using the NMR spectroscopy.
[0291] For normalization, the final library was diluted to 10 nM using Buffer EB and a final 4 nM amplicon library pool was prepared in Buffer EB in a final volume of 20 μl.
[0292] In parallel with the 10 nM PhiX library (# FC-110-3001, Illumina), dilutions of the 4 nM PhiX library were prepared in buffer EB in a final volume of 5 μL (each dilution run had to contain a minimum of 5% Phix to serve as an internal standard for these low density libraries).
[0293] To prepare for cluster generation and sequencing, the pooled amplicon library was denatured with NaOH and diluted with HT1 buffer as follows: 5 μL of 4 nM amplicon library and 5 μL of 0.2 N NaOH (freshly prepared) were introduced into a microcentrifuge tube, mixed briefly using a vortex mixer, and centrifuged at 280×g for 1 min at 20° C. A 5 min incubation was performed at room temperature to denature the DNA, and 990 μL of pre-chilled HT1 buffer was added to 10 μL of denatured DNA. Finally, the HT1 result was added to 20 pM denatured amplicon library in 1 mM NaOH. The denatured DNA was placed on ice until proceeding to the final dilution.
[0294] The same steps were repeated with 5 μL of 4 nM PhiX library, denaturing and diluting PhiX, resulting in a 20 pM PhiX denatured library.
[0295] A 7 pM denatured amplicon library was prepared by mixing 210 μL of 20 pM denatured amplicon library with 390 μL of pre-chilled HT1 buffer in a final volume of 600 μL. A 10 pM denatured PhiX library was then prepared by mixing 300 μL of 20 pM denatured PhiX library with 300 μL of pre-chilled HT1 buffer in a final volume of 600 μL.
[0296] Finally, the 25% denatured PhiX library and the 75% denatured amplicon library were combined in a final volume of 600 μL (150 μL of the 7 pM denatured amplicon library was discarded and replaced with 150 μL of the 10 pM denatured Phix library).
[0297] The combined amplicon library and Phix controls were placed on ice until ready for heat denaturation, which was performed immediately prior to loading the library onto the MiSeq reagent cartridge to ensure efficient template loading onto the MiSeq flow cell.
[0298] The combined library and Phix control tubes were incubated at 96°C for 2 minutes using a heat block. After incubation, the tubes were mixed by inversion 1-2 times and immediately placed on ice. The tubes were kept on ice for 5 minutes.
[0299] 5) Loading the template and setting up a run on the MiSeq instrument Sequencing on a MiSeq instrument using 300 bp read pairs was prepared using the MiSeq reagent Kit v3 (#MS-102-3003, Illumina).
[0300] When the Illumina v3 reagent cartridge was fully thawed and ready to use, the prepared libraries were loaded onto the cartridge and run on the MiSeq instrument according to the manufacturer's instructions.
[0301] 6) Calculation of mitochondrial methylation percentage The percentage (%) of methylation for each cytosine site is calculated by the beta value (β), which is the ratio of methylated reads per site and the combined sum of methylated and unmethylated reads per site, i.e.: β i =M / (M+U) where M is the number of methylated reads at site (i) and U is the number of unmethylated reads at the same site (i). It is.
[0302] The beta value ranges from 0 to 1, where 0 is fully unmethylated and 1 is fully methylated.
[0303] The percent (%) of methylation at each site is calculated by multiplying the β value by 100: i =βi * It is obtained by 100.
[0304] 7) Differential methylation analysis Analysis to compare the level of methylation between each methylation site was performed using the DSS (Dispersion Shrinkage for Sequencing data) Bioconductor package, which is intended to identify differentially methylated loci / sites (DML / DMS) on bisulfite sequencing (BS-seq). A Bayesian hierarchical model was performed to estimate and shrink the site-specific variance of each context, and a Wald test for beta-binomial distribution was applied to each of these context-specific variances to test the null hypothesis that there is no differential methylation between groups of samples. The model is set to have a single main effect. Furthermore, if there is a technical bias due to batch effects or some kind of technical variation, these undesirable effects are contemplated in the model.
[0305] Raw P values were adjusted for multiple testing using both false discovery rate (FDR) and family-wise error rate (FWER) approaches. Any site with a corrected p value less than 0.05 was considered differentially methylated.
[0306] 1.2 Results 1.2.1 Optimization of mt methylation detection The primers used herein (mentioned in section 4-Preparation of library of amplicons) were designed for optimized detection of mtDNA methylation. To avoid biases derived from the general assumption in the field that non-CpG cytosines are mainly unmethylated, we designed primers herein that contain a minimum number of cytosines. Furthermore, the primers were degenerated to cover all possible methylation and unmethylation scenarios with unknown C / U conversion of several cytosine residues contained in the sequence. These primers were a mixture of oligonucleotide sequences containing several possible nucleotide bases at specific positions. As a result, the probability of detecting mitochondrial methylation was higher.
[0307] As a result, the forward degenerate primer contains Y, which points to either C or T (Y=C / T) at any position where the reference sequence is C. Furthermore, as is known in the art, the reverse primer does not correspond to the reference sequence, but to the reversed complementary sequence of the reference sequence. Thus, the reverse primer does not contain the C sites of the reference sequence, but their complementary G sites. Thus, the reverse degenerate primer contains R, which points to either G or A (R=A / G) at any position where the reference sequence is C or the complementary sequence contains G. The four degenerate primers are: D-loop region: Forward primer: YAYTTGGGGGTAGYTAAAGTGAAYTG (SEQ ID NO: 1) Reverse primer: TCCTACAARCATTAATTAATTAACACAC (SEQ ID NO: 2) ND1 gene: Forward primer: ATAAAAYTTAAAAYTTTAYAGTYAGAG (SEQ ID NO: 3) Reverse primer: TTRARTTTRATRCTCACCCTRATCA (SEQ ID NO: 4) It is.
[0308] As shown in Figures 1-6, the use of degenerate primers resulted in extremely high sensitivity for the detection of mtDNA methylation in both regions (i.e., the D-loop region and the ND1 gene) and in all three contexts (i.e., CpG, CHG, CHH). These results were particularly unexpected and exceeded any expectations for high methylation detection.
[0309] 1.2.2 Mitochondrial DNA methylation patterns Comparison of methylation levels between groups for different contexts and regions resulted in a larger number of significant differentially methylated comparisons.
[0310] Example 2: Development of a classification model A model was developed to classify subjects diagnosed with mild cognitive impairment (i.e. CDR=0.5) that is likely to progress to Alzheimer's disease (ADD). Individual data of about 200 controls, including both clinical data and mitochondrial methylation information, was introduced into this model. The data of most patients was used to train the model, and the rest was used to evaluate the performance of the trained model developed. The trained model was shown to be able to calculate the risk of a subject developing ADD, and as a result, classify said patients into the corresponding category between non-progression to ADD and progression to ADD according to the evaluation of performance.
[0311] 2.1 Materials and Methods The classification model of the present invention was developed through processing of individual data (clinical data and mitochondrial methylation measurements generated in-house as described in Example 1).
[0312] 1) Target Data After excluding subjects with missing data, a total of 199 subjects were recruited from the two cohorts described in Section 1.1 of Example 1. (Cohort A: 133 subjects, Cohort B: 66 subjects). These subjects were divided into three groups: Control: Clinical Dementia Rating (CDR) score 0 Clinical follow-up >10 years 79 subjects (39.7%) were characterized as having ADD non-progress: A Clinical Dementia Rating (CDR) score of 0.5 (associated with having early stage MCI) Clinical follow-up >36 months (without progression of symptoms) 47 subjects (23.6%) were characterized as having ADD progression: Clinical Dementia Rating (CDR) score progressed from 0.5 (MCI) to 1 (ADD) 73 subjects (36.7%) were characterized as having were classified into:
[0313] Two main sources of information were included in the analysis: clinical variables and in-house generated methylation measurements (described in Example 1).
[0314] 1.1) Clinical variables The prototype described herein includes testing of seven clinical variables: Dementia stage*: corresponds to the CDR global score (control is CDR=0, ADD non-progression is CDR=0.5, and ADD progression is CDR=1 or higher) (*Dementia stage is used as a known precise output variable to train the model) Gender: Female and Male. Age: The target age. APOE: Apolipoprotein E genotype levels, i.e., E2.E2, E2.E3, E2.E4, E3.E3, E3.E4, and E4.E4. alE4: Reclassification of APOE genotype levels into 0 (E2.E2, E2.E3, and E3.E3), 1 (E2.E4, and E3.E4), and 2 (E4.E4). SOB: Total Box Score is a score ranging from 0 to 18 obtained by adding up each of the domain box scores described above to calculate the CDR global score. MMSE: Mini-Mental State Examination. Five main items: orientation, gaze, concentration and calculation, memory and language, and interpretation assessment with output scores from 1 to 30.
[0315] 1.2) Methylation measurements More than 200 variables were considered, each collecting the percentage of methylation of a single specific cytosine site as described in Example 1. Cytosine sites in three different contexts: CpG, CHG, and CHH, within one of two loci: the D-loop region and the ND1 gene, were considered.
[0316] 2) Exploratory data analysis Exploratory data analysis (EDA) was intended to account for any variables tested in the analysis. A balanced number of subjects were represented by gender, 103 (51.8%) females and 96 (48.2%) males. Table 7 shows a frequency distribution table listing the categorical clinical variables: dementia stage, cohort, APOE, and alE4. Figure 7 shows a bar graph of dementia stage, cohort, APOE, and alE4.
[0317] Furthermore, the recruited individuals were 72.06 ± 6.7 years old, the individuals showed a SOB of 1.27 ± 1.46 and an MMSE of 27.37 ± 2.50. Figure 8 shows a violin plot of the variables age, MMSE, and SOB. Note that the distribution of MMSE and SOB in ADD-progressing patients is more dispersed than in ADD-nonprogressing patients.
[0318] [Table 7]
[0319] 3) Data preprocessing This step was performed to ensure and enhance the performance of the model training process. It consisted of creating dummy variables, removing zero-variance and near-zero-variance variables, identifying and removing correlated variables, splitting the data into training and test datasets, centering and scaling both datasets, and testing and visualizing the training dataset. For the identification and removal of correlated variables, a pairwise correlation analysis based on Pearson's correlation coefficient was performed. For those pairs showing a high level of absolute correlation value (>0.65), the variable with the highest average absolute correlation was removed from the dataset. In this regard, more than 2900 pairs of variables were identified with absolute correlations higher than 0.65. Furthermore, a multifactor analysis (MFA) was performed to test the relationship between individuals described by a set of quantitative and qualitative variables structured in the group of clinical and molecular weight data. However, the features derived from the MFA were not involved in building the classifier. The original calibration data was randomly split into two main subsets: one for model training (80% of samples) and one for testing the classification model (20% of samples). A random sampling process was driven within each class to preserve the overall class distribution of the data.
[0320] The training dataset included inputs and accurate outputs to train the model over time. The accurate outputs referred to the dementia stage of the subject (i.e., control, ADD progression, and ADD non-progression), and the input data included data on methylation patterns and all remaining clinical variables.
[0321] 4) Training the classification model Several supervised learning methods were considered to build classification models according to their ability to handle data with specific features. Thus, all the selected methods were supervised classification methods capable of handling continuous and categorical data. These methods included Linear Discriminant Analysis (LDA), Classification and Regression Trees (CART), k-Nearest Neighbors (kNN), Naive Bayes (NB), Support Vector Machines (SVM) with Linear Kernel, Random Forests (RF), and Neural Networks (NNET).
[0322] All methods were performed using 15-fold cross-validation, and the total number of parameter combinations evaluated for SVM, RF, and NNET was 5. NNET was estimated using a backpropagation approach with 1000 iterations.
[0323] 5) Evaluation of the performance of the classification model To measure the performance of the trained model, we estimated the following metrics: Accuracy: Overall agreement averaged over cross-validation iterations. Kappa; Cohen's unweighted kappa statistic averaged over the resampled results Other common metrics: sensitivity, specificity, positive predictive value, negative predictive value, accuracy, prevalence, F1 score, detection rate, and detection prevalence.
[0324] For each of these statistics, the mean, median, minimum, maximum, and first and third quartiles were estimated. Furthermore, to evaluate the performance of the classification prediction of the model, a confusion matrix was constructed to show the cross-tabulation of observed and predicted classes. This table was accompanied by the accuracy and kappa coefficient, as well as the remaining common metrics used to evaluate the classification model. Furthermore, a ROC curve was constructed to measure the performance of the model in classifying subjects with ADD progression versus those without ADD progression.
[0325] 2.2 Results 2.2.1 Performance of the trained model Overall, all classification models showed adequate performance in terms of accuracy, with values above 0.5, except for LDA, which had an accuracy of 0.48 (see Figure 9). In particular, models trained using RF, RVM, and NNET, and CART showed significantly higher accuracy values above 0.65. Furthermore, these four models also had significantly higher kappa values ranging from 0.46 to 0.56. The best performing model was trained using Random Forest, which showed excellent performance values of 0.72 and 0.56 for accuracy and kappa, respectively. It is noteworthy that the Random Forest method is not only the best performing model, but also able to classify borderline cases well.
[0326] As a result, a model trained based on the Random Forest method was chosen to incorporate the remaining training data and thus validate its performance, which had been initially determined based only on the training data.
[0327] As a result, the prediction of the random forest model on the test data showed an overall accuracy score of 0.76 with a 95% confidence interval of 0.60 to 0.89, and a kappa value of 0.63. Table 8 presents a confusion matrix comparing the prediction of the model with the actual events of each group. The sensitivity and specificity of the model for classifying MCI subjects as ADD progression were 0.86 and 0.7, respectively. The accuracy of the model was 0.63, and the F1 score was 0.72. Figure 10 shows the ROC curves showing the performance of the classification model for ADD progression patients at all classification thresholds. In summary, the classification model developed herein showed very high performance in identifying progressive MCI.
[0328] [Table 8]
[0329] Thus, Table 9 shows some examples of predictions of the classification model, where the output scores correspond to risk scores for progression to ADD.
[0330] [Table 9]
[0331] 2.2.2 Importance of variables In terms of the importance of each variable introduced into the classification model, this describes the contribution of each clinical variable and each site's methylation pattern to determining the risk of developing ADD.
[0332] First, it is noteworthy that the model disclosed in Example 2 did not include any clinical variables that can be obtained via invasive or expensive techniques such as PET. In the current clinical setting, PET to detect β-amyloid plaques is indeed considered to be a highly informative diagnostic technique for AD diagnosis. However, the classification model developed herein, despite not using such information, performed with very high performance indices. It was possible to clearly predict the risk of developing ADD.
[0333] Regarding the importance of the variables, the SOB score and the MMSE were the two variables that contributed most to the determination of the risk score for developing ADD. As mentioned above, they are related to the semi-structured interview and the Mini-Mental State Examination, respectively. Furthermore, age was the fourth variable with a high importance. Thus, these results support the importance of combining clinical variables with variables related to mitochondrial methylation to adequately determine the risk of developing ADD. Furthermore, 15 of the 20 variables with the highest importance ratio corresponded to mitochondrial methylation data of the CHH site of the ND1 gene, whereas only two of these variables were related to the CHH site of the D-loop region. This confirms the significant contribution of the CHH site of the ND1 gene in determining such risk.
[0334] Finally, variables associated with the apolipoprotein E polymorphism showed minimal contribution to determining the risk score for progression to ADD. These results appear surprising since the apolipoprotein E polymorphism is a known major genetic risk determinant of late-onset ADD, and the results herein suggest that said polymorphism may not be significant for detecting subjects at risk of developing ADD at early stages of dementia (i.e. MCI).
[0335] Example 3: Development of a classification model including PET as a clinical variable A model was developed to classify subjects diagnosed with mild cognitive impairment (i.e., CDR=0.5) prone to progression to Alzheimer's Disease Dementia (ADD) as described in Example 1, further including PET to detect beta-amyloid plaques as a clinical variable.
[0336] In current clinical practice, patients suspected of suffering from MCI or ADD undergo amyloid PET testing, which can detect plaques of beta amyloid protein in the brain.However, amyloid PET cannot distinguish between subjects diagnosed with MCI that progress to ADD or those that do not progress to ADD.Therefore, in subjects who have already undergone PET testing, the inventors aim to develop a classification model using the data provided by amyloid PET testing (positive or negative).
[0337] As described in Example 2, about 200 subjects' individual data, including both clinical data and mitochondrial methylation information, are introduced into the model.The data of most patients are used to train the model, and the remaining is used to evaluate the performance of the trained model that is developed.The trained model is shown to be able to calculate the risk of developing ADD, and as a result, it is able to classify said patients into the corresponding category between non-progression to ADD and progression to ADD according to performance evaluation.
[0338] All subjects included in Example 2 were included in this example because there was data available for all of them on amyloid PET studies.
[0339] 3.1 Materials and Methods The classification model of the present invention was developed through processing of individual data (clinical data and in-house generated mitochondrial methylation measurements as described in Example 1).
[0340] 1) Target Data The subjects included in this experiment correspond to those included in Example 2.
[0341] Again, two main sources of information were included in the analysis: clinical variables and in-house generated methylation measurements (described in Example 1).
[0342] With regard to the clinical variables, the prototype described herein includes 10 variables, 9 of which correspond to the clinical variables described in Example 2. The tenth variable corresponds to PET Amyloid: a positron emission tomography measurement to detect the level of amyloid protein aggregates in the brain: POS (positive) or NEG (negative).
[0343] 1.2) Methylation measurements More than 200 variables, each collecting the percentage of methylation of a single specific cytosine site. Three different contexts of cytosine sites in either of the two gene D-loops and the ND1 gene were considered: CpG, CHG, and CHH.
[0344] 2) Exploratory data analysis Exploratory data analysis (EDA) was performed as described in the corresponding section of Example 2. As mentioned above, the subjects of Example 3 are the same subjects as those of Example 2. Thus, the results of the EDA can be found in Figures 7-8 and Table 7. Additionally, Tables 10 and 11 show the frequency of positive and negative β-amyloid PET results.
[0345] [Table 10]
[0346] 3) Data preprocessing As described in Example 2, this step was performed to ensure and enhance the performance of the model training process, which includes splitting the data into a training dataset and a test dataset, and testing and visualizing the training dataset.
[0347] 4) Training the classification model As described in Example 2, several supervised learning methods were examined to construct classification models according to their ability to process data with specific characteristics. Thus, all the selected methods were supervised classification methods capable of processing continuous and categorical data. These methods included Linear Discriminant Analysis (LDA), Classification and Regression Trees (CART), k-Nearest Neighbors (kNN), Naive Bayes (NB), Support Vector Machines (SVM) with Linear Kernel, Random Forests (RF), and Neural Networks (NNET).
[0348] All methods were performed using 15-fold cross-validation, and the total number of parameter combinations evaluated for SVM, RF, and NNET was 5. NNET was estimated using a backpropagation approach with 1000 iterations.
[0349] 5) Evaluation of the performance of the classification model Evaluation of the performance of the classification models was performed as described in the corresponding section of Example 2.
[0350] 3.2 Results 3.2.1. Performance of the trained model Overall, all the trained models showed adequate performance, with most of them having significantly high performance values, as shown in Figure 12. With regard to accuracy, all the classification models showed values above 0.5, except for the LDA model, which had an accuracy of 0.47. Kappa values were low, with only CART, SVM, and RF having values above 0.5.
[0351] The results of training the classification model show that the best performance appears to be the CART model, with an average accuracy of 0.87 and a kappa value of 0.80. These results clearly indicate that the training model developed using the CART method is a good model for classifying MCI patients. However, these types of models are sensitive to outlier observations and may result in misclassifying subjects.
[0352] Furthermore, the trained classification model produced using the Random Forest method also showed very good performance values, with an average accuracy value of 0.83 and a Kappa value of 0.73. It is noteworthy that the Random Forest method is able to classify borderline cases well. As a result, the model trained based on the Random Forest method was selected to introduce the remaining test data and thus verify its performance, which was initially determined based only on the training data.
[0353] As a result, the prediction of the random forest model on the test data showed an overall accuracy score of 0.89, with a 95% confidence interval of 0.75-0.97 and a kappa value of 0.84. Table 11 presents a confusion matrix comparing the prediction of the model with the actual events of each group. The sensitivity and specificity for classifying MCI subjects as ADD progression are 1 and 0.83, respectively. The accuracy of the model is 0.78, and the F1 score is 0.86. Figure 13 shows the ROC curves showing the performance of the classification model for ADD progression patients at all classification thresholds. In summary, the classification model developed herein shows very high performance in identifying ADD progression.
[0354] [Table 11]
[0355] Thus, Table 11 shows some examples of predictions of the classification model, with output values corresponding to the risk scores for progressing to ADD.
[0356] [Table 12]
[0357] 3.2.2 Importance of variables As described in Example 2, the importance of each variable introduced into the classification model describes its contribution to determining the risk of developing ADD.
[0358] In the classification model of this example, again, SOB score and MMSE were the two variables that were the first and third most contributing factors to the risk score of developing ADD, respectively. However, in this classification model, a positive PET result contributed greatly to such a decision and was the second most determinant variable.
[0359] As in Example 2, most of the variables with the highest importance ratios correspond to mitochondrial methylation data of the CHH site of the ND1 gene, again confirming the significant contribution of the CHH site of the ND1 gene in determining such risk.
[0360] Finally, again as in Example 2, the variables associated with the apolipoprotein E polymorphism showed the least contribution to determining the risk score for progressing to ADD.
[0361] Example 4: Development of a classification model including only PET as a clinical variable A model was developed to classify subjects diagnosed with mild cognitive impairment (i.e. CDR=0.5) that are likely to progress to ADD. Individual data of 211 subjects, including both clinical data (including only data on PET study) and mitochondrial methylation information, was introduced into this model. Data on the majority of patients was used to train the model, and the remaining was used to evaluate the performance of the developed and trained model. It was shown that the trained model can calculate the risk of developing ADD for a subject, and as a result, it can classify said patients into the corresponding category between non-progression to ADD and progression to ADD according to the evaluation of performance.
[0362] In this embodiment, the number of subjects used to train and test the model is larger than in previous embodiments (199 to 211 subjects), resulting in more samples and information being available.In particular, such a large number of subjects allows the development of a more reliable classification model.Therefore, it is expected that the inclusion of a large number of subjects in the classification model described herein will achieve even higher performance.
[0363] 4.1 Materials and Methods The classification model of the present invention was developed through processing of individual data: clinical data (PET studies) and mitochondrial methylation measurements.
[0364] 1) Target data After excluding subjects with missing data, a total of 211 subjects were recruited from three cohorts (A, B, and C). Cohort A corresponds to the Australian Imaging, Biomarker & Lifestyle Flagship Study of Ageing (AIBL) cohort (129 subjects), Cohort B corresponds to MCI patients recruited between 2015 and 2019 at the Bellvitge Hospital in Barcelona (74 subjects), and Cohort C corresponds to MCI patients recruited between 2014 and 2019 at the Hospital Clinic de Barcelona (8 subjects). These subjects were divided into three groups: Control: Clinical Dementia Rating (CDR) score 0 Clinical follow-up >10 years 68 subjects (32.23%) were characterized as having ADD non-progress: A Clinical Dementia Rating (CDR) score of 0.5 (associated with having early stage MCI) Clinical follow-up >36 months (without progression of symptoms) 58 subjects (27.49%) were characterized as having ADD progression: Clinical Dementia Rating (CDR) score has progressed from 0.5 (MCI) to 1 (ADD) 85 subjects (40.28%) were characterized as having were classified into:
[0365] Control subjects were available only in cohort A, and MCI subjects were available in all three cohorts.
[0366] Two main types of sources of information were included in the analysis: clinical variables and in-house generated methylation measurements (described in Example 1).
[0367] 1.1) Clinical variables The prototype described herein involves the examination of two clinical variables: Dementia stage*: corresponds to the CDR global score (control is CDR=0, ADD non-progression is CDR=0.5, and ADD progression is CDR=1 or higher) (*Dementia stage is used as a known precise output variable to train the model) PET Study (POS or NEG): corresponds to an amyloid PET study performed on a subject to detect the presence of plaques of beta amyloid protein in the brain.
[0368] As discussed in Example 2, in current clinical practice, patients suspected of suffering from MCI or ADD undergo amyloid PET testing, which can detect plaques of beta amyloid protein in the brain. However, amyloid PET cannot distinguish between subjects diagnosed with MCI who progress to ADD or those who do not progress to ADD, therefore, in subjects who have already undergone PET testing, the inventors aimed to develop a classification model that can use such data provided by amyloid PET testing (positive or negative).
[0369] 1.2) Methylation measurements More than 200 variables were considered, each collecting the percentage of methylation of a single specific cytosine site as described in Example 1. Three different context cytosine sites were considered: CpG, CHG, and CHH, contained in two loci: the D-loop region and one of the ND1 genes.
[0370] 2) Exploratory data analysis Exploratory data analysis (EDA) was intended to account for any variables tested in the analysis. A balanced number of subjects were represented by gender, 110 (52.13%) female and 101 (47.86%) male. Table 13 shows a frequency distribution table describing the categorical clinical variables: dementia stage and PET.
[0371] [Table 13]
[0372] Furthermore, the recruited individuals were aged 72.2±7 years, specifically 69.3±5.28 years in the control group, 71.22±7.67 years in the non-ADD group, and 75.18±6.62 years in the ADD-progressing group.
[0373] 3) Data preprocessing This step was carried out as in Example 2 to ensure and enhance the performance of the model training process, which includes splitting the data into a training dataset and a test dataset, testing and visualizing the training dataset.
[0374] The training data set included inputs and accurate outputs, which allowed the model to learn over time. The accurate outputs referred to the dementia stage of the subject (i.e., control, ADD progression, and ADD non-progression), and the input data included data on methylation patterns and PET studies as clinical variables.
[0375] 4) Training the classification model Several supervised learning methods were considered to build classification models according to their ability to handle data with specific features. Thus, all the selected methods were supervised classification methods that can handle continuous categorical data. These methods included Linear Discriminant Analysis (LDA), Penalized Multinomial Regression (PMR), Classification and Regression Trees (CART), k-Nearest Neighbors (kNN), Naive Bayes (NB), Support Vector Machine with Radial Basis Function Kernel (SVM. Radial), Random Forest (RF), and Neural Networks (NNET).
[0376] All methods were performed using 10-fold cross-validation with 3 repetitions, and the total number of parameter combinations evaluated for SVM, RF, and NNET was 5. NNET was estimated using a backpropagation approach with 1000 repetitions.
[0377] 5) Evaluation of the performance of the classification model To measure the performance of the trained model, we estimated the following metrics: Accuracy: Overall agreement averaged over cross-validation iterations. Kappa; Cohen's unweighted kappa coefficient averaged over the resampled outcomes.
[0378] For each of these statistics, the mean, median, minimum, maximum, and the first and third quartiles were estimated. To evaluate the performance of the model's classification prediction, a confusion matrix was constructed to show the cross-tabulation of observed and predicted classes. This table was accompanied by accuracy and kappa coefficients, as well as common metrics used to evaluate classification models, such as sensitivity, specificity, positive predictive value, negative predictive value, accuracy, prevalence, F1 score, detection rate, and detection prevalence. In addition, a ROC curve was constructed to measure the performance of the model in classifying subjects with ADD progression versus those without ADD progression.
[0379] 4.2 Results 4.2.1 Performance of the trained model Overall, most of the classification models showed adequate performance in terms of accuracy, with values above 0.5 in the case of RF, CART, SVM.Radial, PMR, and NNET.
[0380] In particular, the models trained using RF and CART showed significantly higher accuracy values, above 0.73. Moreover, these two models also have high kappa values, above 0.59. The best performing model was trained using Random Forest, which showed superior performance values for accuracy and kappa, 0.76 and 0.636, respectively. It is noteworthy that the Random Forest method is not only the best performing model, but also can classify the borderline cases well.
[0381] As a result, a model trained based on the Random Forest method was chosen to be run on the remaining test data and thus validate its performance, which was initially determined based only on the training data.
[0382] As a result, the random forest model predictions on the test data showed an overall accuracy score of 0.756, with a 95% confidence interval of 0.597-0.876 and a kappa value of 0.63. Table 14 presents the confusion matrix comparing the model predictions with the actual events in each group.
[0383] [Table 14]
[0384] The sensitivity and specificity of the model for classifying MCI subjects as ADD progression are 0.76 and 0.92, respectively. The accuracy of the model is 0.87, and the F1 score is 0.81. The positive predictive value is 0.87, and the negative predictive value is 0.85. Figure 15 shows the ROC curves that show the performance of the classification model for ADD progression patients at all classification thresholds, with AUC=0.791. In summary, the classification model developed herein shows very high performance in identifying advanced MCI.
[0385] Thus, Table 15 shows some examples of predictions of the classification model, with output values corresponding to risk scores for progression to ADD.
[0386] [Table 15]
[0387] 4.2.2 Importance of variables As described in Examples 2 and 3, the importance of each variable introduced into the classification model describes its contribution to determining the risk of developing ADD.
[0388] In the classification model of this example, only one clinical variable is considered for determining the risk of developing ADD: PET scan. Consistent with Example 3 of the present example, a positive PET result contributed significantly to such a decision and was the largest determining variable.
[0389] Mitochondrial methylation of two CHH sites in the D-loop region were the second and fourth most contributing variables, corresponding to a contribution of 33% and 8.3%, respectively. Furthermore, two CpG sites in the ND1 gene were the third and fifth most contributing variables, with a contribution of 12.2% and 7.6%, respectively. Moreover, the majority of variables with the highest importance ratios corresponded to mitochondrial methylation data of the CHH sites in the ND1 gene, which accounted for half of the 20 highest contributing variables.
[0390] Strikingly, 86 of the 142 variables contributed more than 1% to determining the risk of developing ADD, 35 of 142 contributed more than 2%, 7 of 142 contributed more than 5%, and 3 of 142 contributed more than 10%.
[0391] This example clearly shows that the developed classification model that contains only one clinical variable can achieve significantly high performance, and demonstrates its ability to identify the subjects suffering from MCI who develop late stage ADD.In this case, PET scan is selected as the clinical variable that is included in classification.However, PET scan should not be interpreted as the essential clinical variable that is required for the development of classification model to achieve good results.As shown in Example 2, the construction of the model that does not contain PET scan as clinical variable also leads to a very high performance ratio.
[0392] 4.2.3 Variable importance. Differences between classification models Regarding the variables related to mitochondrial methylation patterns at different sites, the variables that contribute most to determining the risk of developing ADD are different between Examples 2 to 4. This is because the contribution of each variable depends on several aspects, such as the number of variables (e.g., clinical variables included) considered in the model, the number of subjects used in training the classification model (to avoid overfitting / underfitting problems), the parameter customization process during the model training process, etc.
[0393] Thus, the variables included in each classification model depend on the information available from each subject, and as a result, the importance / contribution of each selected variable varies between classification models.Therefore, it is difficult to define the exact set of variables or the specific number of variables that are considered when constructing a classification model.On the contrary, it is desirable to develop a classification model that has the ability to adapt to available information, which can be constructed using variables selected according to their importance / contribution in each specific situation, as illustrated in this embodiment.In other words, the variables included in each classification model are defined by their contribution during the training process, rather than arbitrarily configuring a predefined set of variables or a minimum number of variables that may not contribute much in other scenarios.In any case, the main objective is to achieve the highest possible accuracy.
[0394] Finally, as mentioned above, the more subjects are included in the training process, the more reliable the model can be. However, obtaining samples and clinical information of subjects is a complicated process for many reasons: suitable subjects are limited, they must be monitored for a long period of time; accessibility to samples is limited and expensive; quality of samples may be compromised; information about subjects includes missing data, etc. Overall, the data available for developing classification models should not be taken lightly, as it is the result of a cumbersome and difficult process.
[0395] References International Patent Publication No. 2015 / 144964 European Patent Publication No. 394172 M Blanch, JL Mosquera, B Ansoleaga, I Ferrer, M Barrachina, “Altered Mitochondrial DNA methylation pattern in Alzheimer Disease-related pathology and in Parkinson disease, The American Journal of Pathology, 2016 Feb;186(2):385-97. A Stoccoro, G Siciliano, L Migliore, F Coppede, “Decreased methylation of the mitocondrial D-Loop region in late-onset Alzheimer‘s disease”, Journal of Alzheimer’s Disease 2017;59(2):559-564. A Stoccoro, F Baldacci, F Coppede, L Migliore, “Mitochondrial DNA methylation levels are altered in individuals with mild cognitive impairment”, Abstract OO016 / #1584 On-demand symposium: AD diagnosis & clinical trials & advances in drug development 1. M Frommer, LE McDonald, DS Millar, CM Collis, F. Watt, GW Grigg, PL Molloy, CL Paul, “A genomic sequencing protocol that yields a positive display of 5-methylcytosine residues in individual DNA strands”, Proc Natl Acad Sci U S A. 1992 Mar 1; 89(5): 1827-1831. A Olek, J Oswald, and J Walter, “A modified and improved method for bisulphite based cytosine methylation analysis”, Nucleic Acids Res. 1996 Dec 15; 24(24): 5064-5066.
Claims
1. A method for determining the risk of developing Alzheimer's disease dementia, a) A step of determining the methylation pattern of (a) the D-loop region and / or (b) the ND1 gene of the mitochondrial DNA of a sample obtained from the subject, wherein the methylation pattern is (i) CpG sites in the D-loop region shown in Table 1, (ii) The CpG sites of the ND1 gene shown in Table 2, (iii) CHG sites in the D-loop region shown in Table 3, (iv) CHG sites of the ND1 gene shown in Table 4, (v) CHH region of the D-loop region shown in Table 5, and (vi) CHH sites in the ND1 region shown in Table 6 A step determined by at least one site selected from the group consisting of, b) A step of determining a risk score indicating the risk of developing Alzheimer's disease dementia, wherein the risk score is calculated using a classification model configured to combine the methylation patterns of one or more sites determined in step (a) with at least one clinical variable of the subject selected from the group consisting of sex, sum of box scores, Mini-Mental State Examination, positron emission tomography, presence or absence of β-amyloid protein, age, apolipoprotein E genotype level, and reclassified apolipoprotein E genotype level, Includes, The aforementioned subjects are human subjects diagnosed with mild cognitive impairment. method.
2. The method according to claim 1, wherein the methylation pattern is determined using at least one oligonucleotide having a length of 15 to 100 nucleotides that is capable of specifically hybridizing with a mitochondrial DNA sequence containing a methylation site selected from the group consisting of sequences (i)-(vi), and in particular, the oligonucleotide comprises at least one sequence selected from the group consisting of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, and SEQ ID NO:
4.
3. The method according to claim 1, wherein the step of determining the methylation pattern is determined by bisulfite sequencing.
4. The method according to claim 1, wherein the risk score is related to progression to or non-progression to Alzheimer's disease dementia.
5. The method according to claim 1, wherein the methylation pattern is determined at least all CHH sites in the ND1 region shown in Table 6.
6. The method according to claim 1, wherein the methylation pattern is determined at all CpG, CHG, and CHH sites in the D-loop region and the ND1 region.
7. The method according to claim 1, wherein the classification model is developed using supervised machine learning.
8. The method according to claim 7, wherein the supervised machine learning method is selected from the group consisting of linear discriminant analysis (LDA), penalized multinomial regression (PMR), classification regression tree (CART), k-nearest neighbors (kNN), naive Bayes (NB), support vector machine (SVM) with a linear kernel, support vector machine (SVM.Radial) with a radial basis function kernel, random forest (RF), and neural network (NNET), and is in particular random forest (RF).
9. The method according to claim 7 or 8, wherein the classification model includes mitochondrial methylation patterns of each methylation site in multiple samples related to multiple subjects, is trained with a training set including clinical variables related to multiple subjects, and each subject is assigned to a dementia stage classification selected from a group consisting of control, advanced Alzheimer's disease dementia, and non-advanced Alzheimer's disease dementia.
10. The method according to claim 1, wherein the step of determining the risk score includes correlating each of the at least one methylation pattern and each of the at least one clinical variable determined in step (a) with their weights determined during training of the classification model.
11. The method according to claim 1, wherein the subject is a human subject diagnosed with having a Clinical Dementia Rating Scale score of 0.5, corresponding to mild cognitive impairment.
12. The method according to claim 1, wherein the sample is a biological fluid selected from the group consisting of blood, plasma, saliva, cerebrospinal fluid, brain sample, skin sample, and urine.
13. A kit comprising an oligonucleotide having a length of 15 to 100 nucleotides, comprising the nucleic acid sequences of SEQ ID NO: 1 and SEQ ID NO: 2 or the nucleic acid sequences of SEQ ID NO: 3 and SEQ ID NO:
4.
14. The kit according to claim 13, comprising an oligonucleotide having a length of 15 to 100 nucleotides, comprising the nucleic acid sequences of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, and SEQ ID NO:
4.
15. A computer implementation method for determining the risk of developing Alzheimer's disease dementia, a) A step of receiving data relating to the methylation pattern of the (a) D-loop region and / or (b) ND1 gene of the target mitochondrial DNA, wherein the methylation pattern is (i) CpG sites in the D-loop region shown in Table 1, (ii) The CpG sites of the ND1 gene shown in Table 2, (iii) CHG sites in the D-loop region shown in Table 3, (iv) CHG sites of the ND1 gene shown in Table 4, (v) CHH region of the D-loop region shown in Table 5, and (vi) CHH sites in the ND1 region shown in Table 6 A step determined by at least one site selected from the group consisting of, b) A step of determining a risk score indicating the risk of developing Alzheimer's disease dementia, wherein the risk score is calculated using a classification model configured to combine the methylation patterns of one or more sites determined in step (a) with at least one clinical variable of the subject selected from the group consisting of sex, sum of box scores, Mini-Mental State Examination, positron emission tomography, presence or absence of β-amyloid protein, age, apolipoprotein E genotype level, and reclassified apolipoprotein E genotype level, Includes, The aforementioned subjects are human subjects diagnosed with mild cognitive impairment. method.