Method for Identifying Target Levy-Type Dementia
By analyzing mitochondrial DNA methylation patterns in the D-loop region and ND1 gene, the method provides a non-invasive and accurate diagnosis of Lewy body dementia, addressing the limitations of current diagnostic methods and improving treatment selection.
Patent Information
- Application Number
- JP2024573762
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-04
- Filing Date
- 2023-07-04
- Publication Date
- 2025-07-28
AI Technical Summary
Current diagnostic methods for Lewy body dementia (DLB) are insufficiently reliable and can lead to inaccurate diagnoses, posing risks for patients due to misdiagnosis with neuroleptic drugs, and there is a lack of understanding regarding the role of mitochondrial methylation patterns in DLB diagnosis.
A method utilizing mitochondrial DNA methylation patterns, specifically in the D-loop region and ND1 gene, combined with clinical data, to develop a classification model for diagnosing DLB through blood samples, providing a non-invasive and accurate identification of the disease.
The method achieves high accuracy in distinguishing DLB patients from controls, with an overall accuracy score of 0.83 and kappa value of 0.65, enabling effective diagnosis and treatment selection.
Smart Images

Figure 2025524277000016 
Figure 2025524277000017 
Figure 2025524277000018
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medicaments and diagnosis or identification for a target neurodegenerative disease, in particular to a method for the diagnosis and identification of a target Lewy body dementia.
Background Art
[0002] Lewy body dementia (DLB), i.e., Lewy body dementia, is the second most common cause of degenerative dementia after Alzheimer's disease (AD), and is characterized by non-motor cognitive function changes preceding parkinsonism. DLB belongs to a heterogeneous group of disorders known as Lewy body disease (LBD), a synucleinopathy with heterogeneous clinical symptoms. LBD is characterized by abnormal accumulation and aggregation of misfolded and aggregated α-synuclein (α-syn) that gives rise to Lewy bodies and Lewy neurites.
[0003] DLB shares some features with both Alzheimer's disease and Parkinson's disease (PD). Compared with AD, people with DLB have a higher risk of falls, lower quality of life, greater caregiver burden, and higher mortality, in addition to other differences in clinical features. On the other hand, DLB and Parkinson's disease with dementia (PDD) present overlapping clinical symptoms, including symptoms of attention deficit, visual hallucinations, parkinsonism, fluctuating cognitive impairment, and rapid eye movement (REM) sleep behavior disorder. However, DLB patients are often affected by dementia before the first signs of parkinsonism symptoms, while PD patients are usually the opposite.
[0004] In this regard, regarding the accurate identification and diagnosis of neurodegenerative dementia, new challenges arise due to its shared characteristics and common clinical symptoms. Currently, DLB is diagnosed using a combination of clinical and neuropsychological tests (e.g., Mini-Mental State Examination, F18-fluorodeoxyglucose PET, etc.). However, these diagnostic methods are sometimes insufficient to reach a reliable and clear diagnosis and may lead to inaccurate diagnoses, which can be very dangerous for DLB patients. Patients suffering from DLB are very sensitive to typical neuroleptic drugs used to treat delusions and hallucinations. Therefore, misdiagnosing DLB patients as AD and consequently treating them with typical neuroleptic drugs can cause patients to have significant side effects including high levels of confusion, exacerbation of parkinsonism, extreme drowsiness, and in the most severe cases, malignant syndrome. Thus, it is extremely important to develop an efficient yet achievable and non-invasive diagnostic tool to improve the DLB diagnosis and ultimately reach a clear and definitive diagnosis of DLB.
[0005] Epigenetic mechanisms, such as DNA methylation, regulate the brain transcriptome and have an important role in neurodegeneration. In this regard, genetic and epigenetic diversity of the genes of apolipoprotein E and α-synuclein have been observed in DLB, and changes in the methylation of the promoter and intron 1 of the SNCA gene encoding α-synuclein have also been observed in brain and blood samples from cases of PD and DLB (De Boni et al., 2011; Desplats et al., 2011; Funahashi et al., 2017).
[0006] Furthermore, Urbizu et al. (2020) conducted a detailed re-examination of the epigenetic modifications identified in DLB in postmortem human cells and peripheral tissues, thereby highlighting the fact that there are actually no studies focusing on DLB. Additionally, Chouliaras et al. (2020) summarized different methylation tests in DLB that identify differentially methylated regions containing both the APOE gene and the SNCA gene, along with other novel potential targets.
[0007] In a wide range of epigenomic tests analyzing several tissues and conditions, Fernandez et al. (2012) identified a group of genes that could help distinguish DLB samples from controls in cortical samples. Additionally, Sanchez-Mut et al. (2016) developed a cohort of postmortem brain samples using whole-genome bisulfite sequencing and identified a set of genes that were differentially methylated in DLB and other neurodegenerative conditions. They were able to validate these findings using bisulfite pyrosequencing of a larger cohort, resulting in 1428 differentially methylated regions shared by both PD and DLB subjects when compared to controls.
[0008] Conversely, recent studies have compared the blood methylomes of DLB patients and PDD patients and found significant differences in blood methylation, suggesting that the clinical variability in DLB can be reflected in the blood epigenome. The LASSO method of regularized regression was used at 26 significant differentially methylated positions, and linear discriminant analysis was used to identify the best predictors for distinguishing between DLB cases and PDD cases (Nasamran et al., 2021).
[0009] In conclusion, little is known about the role and characteristics of epigenetics in DLB. As described above, all studies published to date have focused on the epigenetics of the nuclear genome, and notably, due to the still unclear significance and potential similarities or differences in methylation patterns in DLB and other neurodegenerative diseases, no clear outcomes can be derived from such studies.
[0010] On the other hand, International Patent Publication No. WO 2015 / 144964 discloses potential mitochondrial methylation patterns for both AD and PD resulting from the analysis of brain samples from postmortem subjects, and the above results correspond to very early research. Furthermore, the data disclosed in International Patent Publication No. WO 2015 / 144964 are also discussed in Blanch et al. (2016). Similarly, Stoccoro et al. (2017 & 2022) presented the mitochondrial methylation patterns of patients with late-onset AD compared to controls, and later compared the above patterns among subjects diagnosed with mild cognitive impairment, AD patients, and controls.
[0011] Currently, there are no cures or treatments available that can slow down the progression of DLB, and only a few drugs (such as cholinesterase inhibitors) are available for symptomatic treatment. Therefore, it is highly necessary to conduct clinical trials that require accurate recruitment to succeed and ultimately lead to the discovery of new treatment options for DLB patients. In summary, an accurate, reliable, non-invasive, and feasible method for diagnosing and identifying Lewy body dementia with Lewy bodies (referred to as DLB in this specification) is highly desirable to improve the selection of the most appropriate treatment options for patients, improve the recruitment efficiency of clinical trials, and accelerate the clinical development of DLB. SUMMARY OF THE INVENTION
[0012] One problem solved by the present invention is to provide a method for diagnosing or identifying Lewy body dementia with Lewy bodies (referred to as DLB in this specification) in a subject.
[0013] The present invention discloses a method capable of identifying a sample from a subject suffering from DLB and, as a result, diagnosing DLB of a subject in need thereof. The discrimination enabling the diagnosis of DLB is obtained through the processing of a feasible and accessible sample such as a blood sample. Further, the present invention also discloses a method capable of identifying / diagnosing the presence of DLB in a subject and, as a result, calculating or determining a score for classifying the subject according to the diagnosis. The method disclosed herein is further capable of identifying a subject who may not yet be progressing to DLB but is at high risk of developing DLB. This method involves the execution of a classification model capable of processing more than one data set including biomarker screening data (i.e., mitochondrial methylation data) and other relevant clinical data (such as MMSE). The biomarker screening data is obtained from a blood sample, thus enabling a rapid, non-invasive, and effective methodology for the diagnosis / identification of the disease.
[0014] The use of mitochondrial markers for diagnosing other neurodegenerative diseases such as AD and PD is disclosed in International Patent Publication No. WO 2015 / 144964. However, neither the information nor the significance regarding the potential role of mitochondrial markers in DLB has been previously described, nor has its potential usefulness in the diagnosis or identification of the disease been described.
[0015] From the perspective of other epigenetic markers related to DLB, there are very few published studies in this regard, and all of them focus on the analysis of the epigenetics of the nuclear genome. However, no clear conclusions can be drawn from such studies. Some studies have identified that the methylation of specific positions / regions in samples from DLB subjects is different compared to samples from subjects with other neurodegenerative diseases and control subjects, while other studies have suggested a shared methylation pattern of the nuclear genome for DLB and other neurodegenerative diseases (such as PD). In summary, no clues have been disclosed so far regarding the characteristics or significance of the potential methylation pattern of mitochondrial DNA (referred to as mtDNA herein) in DLB.
[0016] Surprisingly, the inventors have found that samples from DLB patients actually exhibit a characteristic pattern of mtDNA methylation in specific regions of mtDNA, such as the D-loop region, while other regions tested in this specification do not contribute to the recognition of DLB patients when compared to controls. As shown in the present invention, the differences in mitochondrial methylation of sites contained in the D-loop region and the ND1 gene, including all three possible contexts (i.e., CpG, CHG, and CHH sites), are significantly statistically significant for DLB subjects.
[0017] As shown in this specification, some methylation sites showing significantly significant methylation patterns correspond to CHG sites in the D-loop region. Considering that epigenetics research commonly focuses on the methylation of CpG sites, these results are worthy of attention. Methylation of non-CpG sites is usually considered to be less or not related to the methylation of CpG sites, potentially leading to biased results. To avoid the above bias, the inventors of the present invention used a set of primers that equally consider the potential methylation of all sites. Furthermore, the primers used in this specification are extremely efficient in detecting the methylation of mtDNA extracted from blood samples (see Example 1). Finally, the examples of the present invention collect information from blood samples, which are examples of non-invasive and feasible samples that do not require expensive processing.
[0018] The working examples in this specification provide detailed experimental data demonstrating the efficient processing of blood samples for the detection and calculation of mtDNA methylation. As a result, an efficient diagnosis of DLB for subjects in need thereof is achievable. Furthermore, the information regarding the methylation of the mitochondria is combined with other relevant clinical data and processed together by a classification model. As a result, the method provided in this specification determines a score corresponding to the presence (i.e., diagnosis) of DLB in a subject.
[0019] Example 1 shows a method for detecting the methylation of mtDNA, including the steps of collecting a blood sample, extracting and treating DNA (bisulfite treatment), and preparing an amplicon library to detect, quantify, and normalize the methylation of the target mtDNA site. The use of degenerate primers has brought extremely high sensitivity in detecting the methylation of mtDNA in all three contexts (i.e., CpG, CHG, CHH). Furthermore, the comparison of methylation levels between DLB subjects and controls in all different contexts of the D-loop region has led to a number of significant differential methylation sites.
[0020] Example 2 shows the development of a classification model that takes into account not only data on the targeted methylation sites but also other relevant clinical data (e.g., MMSE). Several supervised learning methods are applied to select the most appropriate model (i.e., random forest). Basically, after preprocessing the data to handle appropriate levels of categorical variables, missing values, outliers, and normalize continuous variables, the dataset is split into a training set and a validation set. Each method constructs the classification model differently. In the case of random forest, an ensemble of decision trees is constructed using a random subset of the training data, where each decision tree uses a random subset of features (i.e., clinical variables and / or methylation variables) at each split. The decision trees are grown to a maximum depth or until a stopping criterion is met. When classifying new findings, these are passed through each decision tree of the random forest, and the final predicted label is determined based on a majority vote from all the trees. The performance of this model is evaluated using metrics such as accuracy and kappa, which measure the proportion of correctly classified individuals and the agreement between the predicted and true labels, respectively. After constructing the random forest model, the test data is classified by assigning a label (i.e., DLB or control) based on the highest probability predicted by the model. Performance statistics, such as accuracy, kappa, sensitivity, specificity, precision, recall, and F1-score, are computed by a computer to evaluate the classification results. These measurements provide an assessment of accuracy and agreement of the model, as well as its ability to accurately identify positive and negative examples.
[0021] Example 3 shows the methylation patterns of a large number of samples (84). Comparing the methylation levels between DLB subjects and controls in all different contexts of the D-loop region and the ND1 gene resulted in a large number of significant differential methylation sites. Example 1 also shows the development of a classification model considering the data on the methylation sites provided herein. In this example, during the preprocessing step, two methods were used to reduce the dimensionality of the variables (i.e., the measurements of methylation), namely, first, the selection of aggregation features based on Spearman's rank correlation, and then the application of principal component analysis. Based on the top 10 components resulting from PCA, a model was constructed according to the same method as described in Example 2. That is, several supervised methods were implemented to select the most appropriate model (i.e., random forest). The resulting trained model showed excellent high performance as suggested by an overall accuracy score of 0.83 and a kappa value of 0.65. Thus, this classification model can identify the target Lewy body dementia by quickly processing its individual information with significantly good performance.
[0022] Furthermore, it should be noted that this first prototype uses only methylation data and that the methylation patterns provided herein are strong enough to distinguish between DLB patients and healthy controls. It is predicted that if a large number of samples can be obtained and different neuropsychological variables can be introduced into this algorithm, an even more efficient classification model will become available.
[0023] Regarding this, the acquisition of the target samples and clinical information is a complex process for many reasons: the appropriate targets are limited, and the targets have to be monitored for a long period; the accessibility to the samples is limited; the quality of the samples can be impaired; the information regarding the targets contains missing data or is heterogeneous due to, for example, variations among neuropsychological tests used in different countries. In summary, the data available for the development of the classification model should not be regarded as a matter of course because it is the result of a complicated and difficult process.
[0024] Therefore, a first aspect of the present invention is a method for identifying Lewy body dementia in a subject, a) determining a methylation pattern of a D-loop region of mitochondrial DNA and / or an ND1 gene in a sample of the subject containing mitochondrial DNA, wherein the methylation pattern is (i) a CpG site in the D-loop region shown in Table 1, (ii) a CHG site in the D-loop region shown in Table 3, (iii) a CHH site in the D-loop region shown in Table 5, (iv) a CpG site in the ND1 gene shown in Table 2, (v) a CHG site in the ND1 gene shown in Table 4, and (vi) a CHH site in the ND1 gene shown in Table 6 determined at at least one site selected from the group consisting of, relating to the method.
[0025] A second aspect is a method for identifying Lewy body dementia in a subject, a) determining a methylation pattern of a D-loop region of mitochondrial DNA and / or an ND1 gene in a sample of the subject containing mitochondrial DNA, wherein the methylation pattern is (i) a CpG site in the D-loop region shown in Table 1, (ii) The CHG site of the D-loop region shown in Table 3, (iii) The CHH site of the D-loop region shown in Table 5, (iv) The CpG site of the ND1 gene shown in Table 2, (v) The CHG site of the ND1 gene shown in Table 4, and (vi) The CHH site of the ND1 gene shown in Table 6 determined by at least one site selected from the group consisting of: (b) Optionally, combining the methylation pattern of one or more sites determined in step (a) with at least one clinical variable of the subject, wherein said combination is performed using a classification model for determining a score that correlates with the identification of Lewy body dementia in the subject. relates to a method comprising.
[0026] Another aspect of the present invention relates to a method for diagnosing DLB in a subject; a method for identifying a subject as having DLB; a method for identifying a subject suitable for treatment of DLB; a method for selecting a subject to receive treatment for DLB; a method for selecting a treatment suitable for treatment of a subject having DLB; a method for classifying a subject based on the presence or absence of DLB; a method for treating DLB in a subject comprising administering a treatment for DLB when the subject is identified as having DLB; a method for treating DLB in a subject comprising determining the presence or absence of said dementia; and a method for treating DLB in a subject comprising assigning a dementia classification status to the subject using the methods described above prior to administration. Another aspect of the present invention relates to a method for monitoring the progression of DLB in a subject using the methods described above, as well as a method for monitoring and treating a subject afflicted with DLB.
[0027] Another aspect of the present invention is an oligonucleotide and a kit comprising the oligonucleotide for use in determining the methylation pattern of mitochondrial DNA to identify DLB in a subject, wherein said methylation pattern is (i) The CpG site of the D-loop region shown in Table 1, (ii) The CHG sites in the D-loop region shown in Table 3, (iii) The CHH sites in the D-loop region shown in Table 5, (iv) The CpG sites of the ND1 gene shown in Table 2, (v) The CHG sites of the ND1 gene shown in Table 4, and (vi) The CHH sites of the ND1 gene shown in Table 6 determined by at least one methylation site selected from the group consisting of oligonucleotides and kits.
[0028] Throughout the description and claims, the word "comprise" and its variations are not intended to exclude other technical features, additives, components, or steps. Further objects, advantages, and features of the present invention will become apparent to those skilled in the art upon examination of the description, and can be learned by practice of the present invention. Furthermore, the present invention encompasses all possible combinations of the specific preferred embodiments described herein. The following examples and drawings are provided herein for illustrative purposes without intending to limit the present invention.
Brief Description of the Drawings
[0029]
Figure 1
Figure 2
Figure 3
Figure 4
Figures 5A - 5D
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figures 11A - B
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Modes for Carrying Out the Invention
[0030] Detailed Description of the Invention To avoid doubts, the methods provided herein do not include diagnoses performed on the human or animal body. The methods of the invention are, in particular, performed on samples previously taken from a subject. The kits provided herein may include means for extracting a sample from a subject.
[0031] Definitions Diagnosis: The term "diagnosis" refers to both the process of attempting to determine and / or identify a potential disease in a subject, i.e., a diagnostic procedure, and the findings achieved from this process, i.e., diagnostic findings. Thus, this can also be seen as an attempt to classify the state of an individual into distinct and separate categories that enable medical decisions regarding the treatment to be administered and the prognosis. As will be understood by those skilled in the art, such a diagnosis need not be accurate for 100% of the subjects being diagnosed, although 100% accuracy is preferred. However, this term requires that a statistically significant portion of the subjects can be identified as having DLB or a predisposition thereto in the context of the present invention. Those skilled in the art can use various well-known statistical evaluation tools to determine whether a portion is statistically significant, for example, by determining confidence intervals, p-values, adjusted p-values, Student's t-test, Mann-Whitney test. Specific confidence intervals are at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95%. The p-value or adjusted p-value is, in particular, 0.1, 0.05, 0.025, 0.001, or less.
[0032] Lewy body dementia: The term "Lewy body dementia" or "Lewy body dementia" or DLB refers to a degenerative dementia characterized by a slowly progressive decline in cognitive function. Other characteristic features of DLB are spontaneous parkinsonism, recurrent visual hallucinations, cognitive fluctuations, rapid eye movement (REM) sleep behavior disorder (RBD), and severe sensitivity to antipsychotic medication in the development of extrapyramidal symptoms. DLB is a highly variable and fluctuating disease, and its progression can vary considerably between patients, so it does not include a standardized disease stage like other neurodegenerative diseases. The symptoms not only vary greatly between patients, but the symptoms of a patient can also fluctuate considerably over a short period of time.
[0033] The main symptoms of DLB are usually accompanied by dysfunction of visual and spatial cognition (i.e., difficulty in accurately perceiving distance and depth, as well as misidentification of objects), and further by dementia with difficulties in cognitive functions such as planning, multitasking, problem-solving, and reasoning. Dementia can also include changes in mood and behavior, faulty judgment, loss of initiative, mental confusion from the perspective of time and place, and difficulties with language and numbers. Furthermore, the symptoms of DLB also include cognitive fluctuations and hallucinations, and memory loss can occur in the later stages of this disease. Cognitive fluctuations are common and include unpredictable changes in concentration, attention, focus, and arousal, which can vary significantly from day to day or even within the same day. Hallucinations can be visual and / or auditory, but visual hallucinations are more common, affecting up to 80% of DLB patients and generally being very realistic and detailed. Hallucinations are perceptions in the absence of external stimuli that have the nature of real perceptions.
[0034] Regarding motor symptoms, parkinsonism is a common symptom in patients with DLB and can affect each patient very differently. Parkinsonism in DLB causes slowness of movement, difficulty walking, rigidity, and postural instability. Patients with DLB may also exhibit REM sleep behavior disorder, which is a sleep accompaniment with abnormal behavior consistent with dreams during REM sleep. This can include vivid dreams, talking in sleep, and aggressive movements. Additionally, some DLB patients have autonomic neuropathy or autonomic dysfunction due to dysfunction of the autonomic nervous system (ANS).
[0035] Subject: The terms "subject", "patient", "individual", and variations thereof are used interchangeably herein and refer to any mammalian subject, particularly a human subject. This term does not indicate a specific age or gender.
[0036] Sample containing mitochondrial DNA: As used herein, the expression "sample containing mitochondrial DNA" refers to any sample that can be obtained from a subject in which there is genetic material from mitochondria suitable for detection of methylation patterns.
[0037] Mitochondrial DNA: As used herein, the term "mitochondrial DNA" or "mtDNA" refers to the genetic material located in the mitochondria of a living organism. This is a closed-circular double-stranded molecule. In humans, it consists of 16,569 base pairs and contains a small number of genes distributed between the H-strand and the L-strand. Mitochondrial DNA encodes 37 genes: two ribosomal RNAs, 22 transfer RNAs, and 13 proteins involved in oxidative phosphorylation.
[0038] Methylation pattern or methylation state: As used herein, the term "methylation pattern", without limitation, refers to the presence or absence of methylation of one or more nucleotides, particularly methylation of cytosine. Thus, the one or more nucleotides are contained in a single nucleic acid molecule. The one or more nucleotides can or cannot be methylated. The term "methylation state" can also be used when considering only a single nucleotide. The methylation pattern can be quantified when considering more than one nucleic acid molecule.
[0039] D-loop region: As used herein, this term refers to a region of non-coding mtDNA that acts as a promoter for both the heavy and light strands of mtDNA and contains essential transcriptional and replication elements. The D-loop region contains approximately 1120 base pairs, is visible by electron microscopy, and is created during replication of the H-strand for synthesis of a short segment of the heavy strand, 7S DNA. The human D-loop region sequence is deposited in the GenBank database under accession number NC_012920.1.
[0040] ND1 gene: As used herein, the term "ND1 gene" or "NADH dehydrogenase 1" or "ND1mt" refers to a gene located in the mitochondrial genome that encodes the protein NADH dehydrogenase 1 or ND1. The human ND1 gene sequence is deposited in the GenBank database under accession number NC_012920.1. The ND1 protein is active in mitochondria and is part of an enzyme complex called complex I that is involved in the process of oxidative phosphorylation. In some embodiments, the term "ND1 gene" can refer to the above gene further comprising approximately 50 additional base pairs at one or both ends of the sequence.
[0041] CpG site: This term is used herein to distinguish this single-stranded linear sequence from the CG base pairs of cytosine and guanine in a double-stranded sequence. "CpG" is an abbreviation for "C-phosphate-G", i.e., cytosine and guanine separated by only one phosphate; the phosphate covalently links any two nucleosides in DNA. The term "CpG" is used to distinguish the formation of CG base pairs of guanine and cytosine from this linear sequence. The cytosine of a CpG dinucleotide can be methylated to form 5-methylcytosine.
[0042] CHG site: This term, as used herein, refers to a DNA region, particularly a mitochondrial DNA region, in which a cytosine nucleotide and a guanine nucleotide are separated by a variable nucleotide (H) which can be adenine, cytosine, or thymine. The cytosine of a CHG site can be methylated to form 5-methylcytosine.
[0043] CHH site: This term, as used herein, refers to a DNA region, particularly a mitochondrial DNA region, in which a cytosine nucleotide is followed by first and second variable nucleotides (H) which can be adenine, cytosine, or thymine. The cytosine of CHH can be methylated to form 5-methylcytosine.
[0044] Determination of the methylation pattern of a CpG site: The term "determination of the methylation pattern of a CpG site", as used herein, refers to the determination of the methylation state of a specific CpG site. The determination of the methylation pattern of a CpG site can be performed by a plurality of processes known to those skilled in the art.
[0045] Determination of the methylation pattern of a CHG site: This term, as used herein, refers to the determination of the methylation state of a specific CHG site. The determination of the methylation state of a CHG site can be performed by a plurality of processes known to those skilled in the art.
[0046] Determination of the methylation pattern at CHH sites: This term, as used herein, refers to the determination of the methylation status of a specific CHH site. The determination of the methylation status of a CHH site can be performed by a plurality of processes known to those skilled in the art.
[0047] To determine the methylation pattern of mitochondrial DNA, the sample can be chemically treated such that all unmethylated cytosine bases are modified to a uracil base or another base that is different from cytosine in terms of base pairing behavior, while the base of 5-methylcytosine remains unchanged. The term "modify ~" as used herein means the conversion of an unmethylated cytosine to another nucleotide that distinguishes the unmethylated cytosine from the methylated cytosine. The conversion of unmethylated cytosine bases that do not convert methylated cytosine bases in a sample containing mitochondrial DNA is carried out by a conversion agent. The term "conversion agent" or "conversion reagent" as used herein refers to a reagent capable of converting an unmethylated cytosine to uracil or another base that is differentially detectable from cytosine in terms of hybridization characteristics. The conversion agent is bisulfite, such as bisulfite or hydrogen sulfite. However, other agents that similarly modify unmethylated cytosine but do not modify methylated cytosine, such as hydrogen sulfite, can also be used in this method of the present invention. This reaction is carried out according to standard procedures. Also, for example, using methylation specific to cytidine deaminase, it is also possible to perform the conversion enzymatically.
[0048] Reference sample: As used herein, this term refers to a sample containing mitochondrial DNA obtained from a subject not suffering from DLB. Thus, the reference DNA sequence is NCBI reference sequence: NC_012920.1. In particular, the term refers to a few 5-methylated cytosines in one or more CpG sites, one or more CHG sites, and / or one or more CHH sites in the mitochondrial DNA sequence, compared to the relative amount of 5-methylcytosine present in one or more CpG sites, one or more CHG sites, and / or one or more CHH sites in the subject sample, in one or more CpG sites in the D-loop region shown in Table 1, one or more CpG sites in the ND1 gene shown in Table 2, one or more CHG sites in the D-loop region shown in Table 3, one or more CHG sites in the ND1 gene shown in Table 4, one or more CHH sites in the D-loop region shown in Table 5, and / or one or more CHH sites in the ND1 gene shown in Table 6.
[0049] Treatment of Lewy body dementia: As used herein, this term refers to the treatment of this disease or any related symptoms. Such treatment may include drug therapy, epigenetic treatment, or any treatment that stimulates cognition. Some treatments that stimulate cognition are digital cognitive treatments (i.e., using digital devices). This term may include any treatment known or to be developed in the future for DLB. Treatments for DLB may include, but are not limited to, drugs targeted at treating symptoms of dementia (which may include cholinesterase inhibitors such as rivastigmine), drugs targeted at treating parkinsonism (such as levodopa), or drugs targeted at treating symptoms of psychosis such as hallucinations (such as quetiapine). Treatments for DLB may include preventive treatment for subjects identified as being at high risk of developing DLB but not yet in the stage of dementia. Preventive treatment includes any method known in the field of cognitive stimulation. Preventive treatment may include appropriate future developments.
[0050] Dementia Staging Classification (DSC): As used herein, the term "Dementia Staging Classification" refers to the classification of a subject based on a diagnosis or identification of dementia. This classification results from a score determined by a classification model disclosed herein that determines the presence or absence of dementia. The classification models according to embodiments of the present invention classify subjects with DLB and subjects without DLB. Subjects without DLB are not necessarily healthy subjects and may be affected by other diseases such as Alzheimer's disease. Thus, "Dementia Staging Classification" includes two classes: subjects with Lewy body dementia and subjects without DLB.
[0051] Specifically hybridize: The expressions "specifically hybridize" or "capable of hybridizing in a specific form" refer to the property of an oligonucleotide or polynucleotide that specifically recognizes a specific sequence of interest, such as a D-loop region or an ND1 gene. The sequence of interest may refer to a reference sequence or a sequence resulting from a specific modification treatment, such as bisulfite treatment in which unmethylated cytosine is modified to uracil. As used herein, the term "hybridization" refers to a process in which two nucleic acid molecules or single-stranded molecules with a high degree of similarity are combined to result in a simple double-stranded molecule by specific pairing between complementary bases. Usually, hybridization occurs under very stringent conditions or moderately stringent conditions.
[0052] Oligonucleotide: As used herein, the term "oligonucleotide" is used without distinction from "primer" and "nucleic acid sequence" and refers to a DNA molecule or short RNA up to the maximum base length. The oligonucleotides of the present invention are, in particular, DNA molecules having a base length of at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 30, at least 40, at least 45, or 50.
[0053] Score: As used herein, this term refers to one or more values, particularly a single value, that can be used as a component of a classification model for the diagnosis or identification of a target disease. Such a single value can be calculated or determined (i.e., estimated) by combining values of descriptive characteristics that are processed by an interpretation function or algorithm. The score can be a score between 0.0 and 1, where a score above 0.5 indicates the presence of the disease in the subject, and 0 indicates the absence of the disease. The score can be classified into multiple groups, i.e., multiple classes, such as none, low, intermediate, and high.
[0054] Computer-implemented method: The term "computer-implemented method" refers to a method in which all or some steps of the method are performed by a computer, another programmable device, or a network of computers.
[0055] Supervised machine learning method: This term refers to a method that uses a training set to create a desired output for developing a model. This training dataset includes inputs and correct outputs, thereby enabling the model to learn over time. The accuracy is measured via a loss function, and the model can be adjusted until the generalization error is sufficiently minimized.
[0056] Unsupervised machine learning method: The term "unsupervised learning methods (also referred to as non-supervised learning methods)" refers to a method that uses machine learning algorithms to analyze an unlabeled dataset. These methods recognize similarities and differences in information to detect groups or patterns in the data. These methods are generally used for clustering, association, and dimensionality reduction of datasets.
[0057] Deep learning method: This term refers to a subfield of machine learning that uses neural networks with multiple layers to automatically extract and learn features and patterns from raw data. The layers of a neural network are sets of interconnected artificial neurons (also known as nodes). These neurons receive input data and apply a task on a specific type of computer to this data. Usually, these tasks include a weighted sum and a mathematical activation function. The output of a layer is transmitted to the next layer, and so on. Through this process, the network gradually learns more complex features, attributes, and patterns in the data. There are different types of layers in a neural network. Some examples are fully connected layers, convolutional layers, pooling layers, and regression layers.
[0058] Artificial intelligence method: This term refers to a method that uses artificial intelligence, which is defined as the ability of a computer or computer-controlled robot to mimic the ability of a human to respond to specific stimuli.
[0059] Classification model: The term "Classification model" is referred to in this specification without distinction as "classifying model". A classification model can be developed, for example, using a type of supervised learning method that can accurately assign test data / new findings to specific categories based on training data. This model is trained to learn from a given training dataset and, as a result, can classify a new dataset into specific scores / numbers or classes / groups. Examples of classification algorithms are logistic regression, linear discriminant analysis (LDA), CART (Classification and Regression Trees), k-nearest neighbor method (kNN), naive Bayes (NB), support vector machine (SVM) using a linear kernel, random forest (RF), neural network (NNET), Generalized Boosted Regression Model (GBM), and binary logistic regression (GLM).
[0060] Conversion: As used herein, this term means subjecting one or more described characteristics to an interpretation function or algorithm related to a predictive model for a disease, particularly DLB. In some embodiments, the interpretation function can also be created by multiple predictive models. In one embodiment, the predictive model includes a regression model and a Bayesian classifier or score. In one embodiment, the interpretation function includes one or more terms related to one or more biomarkers or a set of biomarkers. In one embodiment, the interpretation function includes one or more terms related to the presence, absence, or spatial distribution of a particular cell type disclosed herein. In one embodiment, the interpretation function includes one or more terms related to the presence, absence, amount, intensity, or spatial distribution of morphological characteristics of cells in a cell sample. In one embodiment, the interpretation function includes one or more terms related to the presence, absence, amount, intensity, or spatial distribution of descriptive characteristics of cells in a cell sample.
[0061] Method for identifying DLB One aspect of the present invention is a method for identifying DLB in a subject, comprising determining a methylation pattern of the D-loop region and / or the ND1 gene in a sample from the subject containing mitochondrial DNA, wherein the methylation pattern is (i) CpG sites in the D-loop region shown in Table 1, (ii) CHG sites in the D-loop region shown in Table 3, (iii) CHH sites in the D-loop region shown in Table 5, (iv) CpG sites in the ND1 gene shown in Table 2, (v) CHG sites in the ND1 gene shown in Table 4, and (vi) CHH sites in the ND1 gene shown in Table 6 Relates to a method comprising steps determined by at least one methylation site selected from the group consisting of. A computer-implemented method for identifying DLB in a subject may include the step of receiving a methylation pattern of the D-loop region of mitochondrial DNA and / or the ND1 gene from a sample from the subject, and a classification model for assigning a score to the subject.
[0062] In one embodiment, the method is (a) In a sample of a subject containing mitochondrial DNA, determining a methylation pattern of the D-loop region of mitochondrial DNA and / or the ND1 gene, wherein the methylation pattern is (i) CpG sites in the D-loop region shown in Table 1, (ii) CHG sites in the D-loop region shown in Table 3, (iii) CHH sites in the D-loop region shown in Table 5, (iv) CpG sites in the ND1 gene shown in Table 2, and (v) CHG sites in the ND1 gene shown in Table 4 determined by at least one site selected from the group consisting of the methylation pattern is determined by at least one of the CHH sites in the ND1 gene shown in Table 6, step including.
[0063] Alternatively, the method of the present invention can be devised as a method for diagnosing DLB in a subject, including the steps described above. Alternatively, the method of the present invention can be devised as a method for identifying a subject having DLB, including the steps described above.
[0064] In some embodiments, the method, particularly the computer-implemented method, includes a step of determining a methylation pattern of a D-loop region, wherein hypomethylation of at least one site of CpG sites in the D-loop region, hypomethylation of at least one site of CHG sites in the D-loop region, and hypomethylation of at least one site of CHH sites in the D-loop region indicate that the subject has DLB. Alternatively, these methylation patterns indicate that the subject has a high risk of developing DLB.
[0065] In one embodiment, the method includes a step of determining a methylation pattern of the ND1 gene, wherein the methylation pattern is (iv) CpG sites of the ND1 gene shown in Table 2, (v) CHG sites of the ND1 gene shown in Table 4, and (vi) CHH sites of the ND1 gene shown in Table 6 and is determined by at least one methylation site selected from the group consisting of and further includes the step.
[0066] In some embodiments, the method includes a step of determining a methylation pattern of the ND1 gene, and no significant difference in the methylation level of at least one site of the CpG sites of the ND1 gene, no significant difference in the methylation level of at least one site of the CHG sites of the ND1 gene, and / or no significant difference in the methylation level of at least one site of the CHH sites of the ND1 gene indicate that the subject has DLB. Alternatively, these methylation patterns indicate that the subject has a high risk of developing DLB.
[0067] The statistical significance of the methylation pattern at each methylation site can vary depending on the set significance threshold, i.e., the corrected p-value. In some embodiments, the p-value is adjusted using the false discovery rate (FDR) method or the family-wise error rate (FWER) method. In certain embodiments, the p-value is adjusted using the FDR method. The difference in the methylation pattern can be considered statistically significant or not depending on the established significance threshold, which can be, for example, a corrected p-value ≥ 0.05 or a corrected p-value ≥ 0.1. In some embodiments, a difference in the methylation pattern is not considered statistically significant if the corrected p-value ≥ 0.05. In other embodiments, a difference in the methylation pattern is not considered statistically significant if the corrected p-value ≥ 0.1. In another embodiment, the method comprises determining the methylation pattern of the ND1 gene, wherein there is no significant difference in the methylation level at at least one site of the CpG site of the ND1 gene, no significant difference in the methylation level at at least one site of the CHG site of the ND1 gene, and / or hypomethylation at at least one site of the CHH site of the ND1 gene indicates that the subject has DLB. Alternatively, these methylation patterns indicate that the subject has a high risk of developing DLB.
[0068] In some embodiments, the method comprises determining the methylation patterns of the D-loop region and the ND1 gene, wherein hypomethylation at at least one site of the CpG site of the D-loop region, hypomethylation at at least one site of the CHG site of the D-loop region, and / or hypomethylation at at least one site of the CHH site of the D-loop region, and there is no significant difference in the methylation level at at least one site of the CpG site of the ND1 gene, there is no significant difference in the methylation level at at least one site of the CHG site of the ND1 gene, and / or There is no significant difference in the methylation level or hypomethylation at at least one site of the CHH site of the ND1 gene indicating that the subject has DLB including steps
[0069] The methods described herein are useful for diagnosing DLB in a subject and can be used clinically to make treatment decisions by selecting the most appropriate treatment modality for any particular patient. The method can also be useful, for example, for classifying subjects for participation in clinical trials or for receiving appropriate treatment early. Thus, the applications disclosed herein can improve clinical outcomes by matching patients to treatment and can also improve the accuracy rate in the selection of patients required for clinical trials to succeed in evaluating potential DLB treatments.
[0070] In this regard, another aspect is a method, particularly a computer-implemented method, for identifying a subject suitable for the treatment of DLB, the method comprising determining a methylation pattern in a sample from the subject comprising mitochondrial DNA, the methylation pattern being (i) CpG sites in the D-loop region shown in Table 1 (ii) CHG sites in the D-loop region shown in Table 3 (iii) CHH sites in the D-loop region shown in Table 5 (iv) CpG sites of the ND1 gene shown in Table 2 (v) CHG sites of the ND1 gene shown in Table 4, and (vi) CHH sites of the ND1 gene shown in Table 6 determined by at least one methylation site selected from the group consisting of relates to a method including steps
[0071] Alternatively, this aspect can be devised as a method for selecting a subject to receive treatment for DLB, including the steps described above.
[0072] In some embodiments, the method (i.e., a method for identifying a subject suitable for treatment of DLB or a method for selecting a subject to receive treatment of DLB) further comprises administering treatment of DLB to the subject.
[0073] In some embodiments, the method comprises determining a methylation pattern of a D-loop region, wherein hypomethylation at at least one site of CpG sites of the D-loop region, hypomethylation at at least one site of CHG sites of the D-loop region, and / or hypomethylation at at least one site of CHH sites of the D-loop region indicates that the subject has DLB or has a high risk of developing DLB.
[0074] In some embodiments, the method further comprises determining a methylation pattern of the ND1 gene as described above.
[0075] Diagnosis or identification of DLB in a subject indicates that treatment of DLB can be administered to the subject. If the subject is not diagnosed or identified as having DLB, the subject cannot be subjected to any treatment related to DLB and is suspected of having another disease, thus allowing clinical follow-up and / or other diagnostic methods. Alternatively, the subject can be treated for other dementias if necessary.
[0076] Thus, another aspect of the present invention is a method, particularly a computer-implemented method, for selecting an appropriate treatment for treating a subject having DLB, comprising determining a methylation pattern of a D-loop region and / or an ND1 gene in a sample from the subject containing mitochondrial DNA, wherein the methylation pattern is (i) CpG sites of the D-loop region shown in Table 1, (ii) CHG sites of the D-loop region shown in Table 3, (iii) CHH sites of the D-loop region shown in Table 5, (iv) the CpG sites of the ND1 gene shown in Table 2, (v) the CHG sites of the ND1 gene shown in Table 4, and (vi) the CHH sites of the ND1 gene shown in Table 6 determined by at least one methylation site selected from the group consisting of a method comprising the steps is targeted.
[0077] Alternatively, an aspect of the present invention is a method for classifying a subject according to the presence or absence of DLB, which can be assembled as a method including the above steps.
[0078] In some embodiments, the method (i.e., the method for selecting a treatment suitable for treating a subject suffering from DLB and the method for classifying a subject according to the presence or absence of DLB) further comprises administering a treatment for DLB to the subject when the subject is classified as having DLB.
[0079] In some embodiments, the method further comprises determining the methylation pattern of the ND1 gene as described above.
[0080] Another aspect of the present invention is a method for treating DLB in a subject, particularly a computer-implemented method, comprising administering a treatment for DLB when the subject is identified as having DLB by determining the methylation pattern of the D-loop region and / or the ND1 gene in a sample from the subject containing mitochondrial DNA, wherein the methylation pattern is (i) the CpG sites of the D-loop region shown in Table 1, (ii) the CHG sites of the D-loop region shown in Table 3, (iii) the CHH sites of the D-loop region shown in Table 5, (iv) the CpG sites of the ND1 gene shown in Table 2, (v) the CHG sites of the ND1 gene shown in Table 4, and (vi) the CHH sites of the ND1 gene shown in Table 6 Determined by at least one site selected from the group consisting of, A method, particularly a computer-implemented method, comprising steps.
[0081] Alternatively, this aspect of the invention is a method for treating DLB, comprising determining the methylation pattern of the D-loop region and / or the ND1 gene in a sample from the subject containing mitochondrial DNA to determine the presence or absence of the dementia, wherein the methylation pattern is (i) CpG sites in the D-loop region shown in Table 1, (ii) CHG sites in the D-loop region shown in Table 3, (iii) CHH sites in the D-loop region shown in Table 5, (iv) CpG sites in the ND1 gene shown in Table 2, (v) CHG sites in the ND1 gene shown in Table 4, and (vi) CHH sites in the ND1 gene shown in Table 6 Determined by at least one site selected from the group consisting of, and comprising the step of The treatment for DLB is administered when the subject is identified as having DLB. Can be devised as a method.
[0082] In some embodiments, the method comprises determining the methylation pattern of the D-loop region, wherein hypomethylation at at least one site of the CpG sites in the D-loop region, hypomethylation at at least one site of the CHG sites in the D-loop region, and / or hypomethylation at at least one site of the CHH sites in the D-loop region indicates that the subject is suffering from DLB.
[0083] In some embodiments, the method further comprises determining the methylation pattern of the ND gene, comprising the steps described above.
[0084] In addition, other aspects are methods for treating a subject, particularly a subject diagnosed with DLB, by administering a treatment for DLB, wherein, prior to administration, a dementia state classification (DSC) determined by applying a classification model to a methylation pattern of the D-loop region of mitochondrial DNA and / or the ND1 gene from a sample from the subject is assigned to the subject, and the DSC is selected from subjects with DLB and subjects without DLB. In one embodiment, if the subject is assigned to the presence of DLB in the DSC, the subject is suitable for receiving a treatment for DLB. In another embodiment, if the subject is assigned as not suffering from DLB, the subject cannot be subjected to any treatment for DLB, but can be under clinical observation or tested for another disease such as AD. Alternatively, the subject can be treated for other dementias if necessary.
[0085] Alternatively, this aspect of the invention is a method for treating a subject, particularly a computer-implemented method, comprising (i) prior to administration, assigning a DCS to the subject by applying to a classification model a methylation pattern of the D-loop region of mitochondrial DNA and / or the ND1 gene from a sample from the subject, wherein the DSC is selected from subjects with DLB and subjects without DLB; and (ii) if the DSC is a subject with DLB, administering a treatment for DLB to the subject, or if the DSC is a subject without DLB, not administering a treatment for DLB or administering clinical observation or treatment for another dementia. which can be devised as a computer-implemented method.
[0086] In some embodiments, the methylation pattern of the D-loop region of mitochondrial DNA is (i) the CpG sites of the D-loop region shown in Table 1, (ii) the CHG sites of the D-loop region shown in Table 3, and (iii) the CHH sites of the D-loop region shown in Table 5 determined by at least one site selected from the group consisting of
[0087] In other embodiments, the method (iv) the CpG sites of the ND1 gene shown in Table 2, (v) the CHG sites of the ND1 gene shown in Table 4, and (vi) the CHH sites of the ND1 gene shown in Table 6 further comprises determining the methylation pattern of the ND1 gene of mitochondrial DNA at at least one site selected from the group consisting of
[0088] In some embodiments, the methods described herein (a) determining the methylation pattern of the D-loop region of mitochondrial DNA and / or the ND1 gene in a sample of a subject comprising mitochondrial DNA, wherein the methylation pattern is at least one of the CHH sites of the ND1 gene shown in Table 6, and (i) the CpG sites of the D-loop region shown in Table 1, (ii) the CHG sites of the D-loop region shown in Table 3, (iii) the CHH sites of the D-loop region shown in Table 5, (iv) the CpG sites of the ND1 gene shown in Table 2, and (v) the CHG sites of the ND1 gene shown in Table 4 determined by at least one site selected from the group consisting of comprising
[0089]
Table 1
[0090]
Table 2
[0091]
Table 3
[0092]
Table 4
[0093]
Table 5
[0094]
Table 6
[0095] Another aspect of the present invention is a method, particularly a computer-implemented method, for monitoring the progression of DLB in a subject, comprising: (a) determining the methylation pattern described herein in a sample from the subject comprising mitochondrial DNA; (b) comparing the methylation pattern determined in step (a) with the methylation pattern obtained at the early stage of the disease; A higher risk score than the previous risk score indicates progression of DLB and thus a poor prognosis. The risk score can be monitored, for example, once a year.
[0096] In some embodiments, an increase in hypomethylation at at least one site of the CpG site in the D-loop region, an increase in hypomethylation at at least one site of the CHG site in the D-loop region, and / or an increase in hypomethylation at at least one site of the CHH site in the D-loop region over time indicates progression of DLB.
[0097] Alternatively, this aspect is a method, particularly a computer-implemented method, for monitoring and treating a subject suffering from DLB, comprising: (a) administering a treatment for DLB to the subject; (b) If the treatment is effective, continue administering the treatment to the subject, or (c) If the treatment is ineffective, discontinue administering the treatment to the subject comprising (i) If, with respect to the methylation level before administration of the treatment, hypomethylation at at least one site of the CpG site in the D-loop region does not increase, hypomethylation at at least one site of the CHG site in the D-loop region does not increase, and / or hypomethylation at at least one site of the CHH site in the D-loop region does not increase, it indicates that the treatment is effective, (ii) If, with respect to the methylation level before administration of the treatment, hypomethylation at at least one site of the CpG site in the D-loop region increases, hypomethylation at at least one site of the CHG site in the D-loop region increases, and / or hypomethylation at at least one site of the CHH site in the D-loop region increases, it indicates that the treatment is ineffective, The method can be devised as a computer-implemented method.
[0098] Another aspect of the present invention is a method for monitoring the progression of DLB in a subject, comprising (a) determining a methylation pattern in a sample from the subject comprising mitochondrial DNA, wherein the methylation pattern is (i) the CpG site of the D-loop region shown in Table 1, (ii) the CHG site of the D-loop region shown in Table 3, and (iii) the CHH site of the D-loop region shown in Table 5 determined at at least one site selected from the group consisting of: (b) combining the methylation pattern of one or more sites determined in step (a) with at least one clinical variable of the subject, wherein the combination is performed using a classification model for determining a score that correlates with the presence of DLB in the subject, (c) comparing the score pattern determined in step (b) with the score obtained at the initial stage of the disease A method is targeted.
[0099] In one embodiment, the change in score over time is related to the progression of the disease. In some embodiments, a high score over time indicates the progression of the disease. In another embodiment, a low score over time or no difference in the score indicates the non - progression of the disease.
[0100] Alternatively, this aspect of the invention is a method for monitoring and treating a subject suffering from DLB, (a) administering a treatment for DLB to the subject; (b) if the treatment is effective, continuing to administer the treatment to the subject, or (c) if the treatment is ineffective, discontinuing the administration of the treatment to the subject comprising (i) a low or equivalent score indicates that the treatment is effective; (ii) a high score indicates that the treatment is ineffective, A method is targeted.
[0101] Subject In some embodiments, the subject is an animal, particularly a mammal. In certain embodiments, the subject is selected from the group consisting of rats, mice, cats, dogs, chimpanzees, and humans. More specifically, the subject is a human.
[0102] In some embodiments, the subject is suspected of having dementia. In particular, the subject is suspected of having DLB. In another embodiment, the subject is at risk of developing dementia or has been diagnosed with dementia. In some embodiments, the subject has at least one symptom of mild cognitive impairment. In particular, the subject has at least one symptom selected from the group consisting of mild memory loss, mild loss of attention, difficulty with reasoning, planning, or problem-solving, difficulty with language, and reduced depth perception. In other embodiments, the subject has at least one symptom selected from the group consisting of parkinsonism, visual and / or auditory hallucinations, and REM sleep behavior disorder.
[0103] In some embodiments, the subject has been diagnosed with a dementia other than DLB. In particular, the subject has been misdiagnosed with a dementia other than DLB. In some embodiments, the subject is suspected of being misdiagnosed. In some embodiments, the subject has a Clinical Dementia Rating (CDR) score of 0.5. In particular, the subject is diagnosed with mild cognitive impairment. In another embodiment, the subject has a CDR score greater than 0.5. In particular, the subject has CDR1. In another embodiment, the subject has a score greater than CDR1. In particular, the subject has CDR2 or 3.
[0104] The "Clinical Dementia Rating" is an overall summary obtained through semi-structured interviews of the patient and informant, and grades the subject's cognitive status into six functional domains including memory, orientation, judgment and problem-solving, social adaptation, family situation and interests, and caregiving situation. Each domain is assigned a score from 0 to 3, and the results are computer-processed via an algorithm to obtain the CDR global score. The CDR global score ranges from 0 to 3 and enables grouping of subjects according to the severity of their dementia, where CDR = 0 corresponds to the absence of cognitive impairment, CDR = 0.5 corresponds to suspected or very mild dementia, CDR = 1 corresponds to mild dementia, CDR = 2 corresponds to moderate dementia, and CDR = 3 corresponds to severe dementia.
[0105] Subjects with a CDR of 0.5 are diagnosed with mild cognitive impairment (MCI), which is early memory loss or loss of other cognitive abilities such as language or visual / spatial perception in subjects who maintain the ability to perform most activities of daily life independently. Subjects with a CDR of 1 or more are considered to already be in the progression of dementia.
[0106] Methylation patterns and their measurement Methylation patterns The methylation patterns of the mitochondrial DNA sites described herein can be determined using any method known in the art or future methods developed in the art. For example, methylation of mtDNA can be determined by treating the sample with bisulfite and sequencing the treated sample.
[0107] In some embodiments, determining the methylation pattern comprises (i) the CpG sites of the D-loop region shown in Table 1, (ii) the CHG sites of the D-loop region shown in Table 3, (iii) the CHH sites of the D-loop region shown in Table 5 and comprises the step of determining the methylation pattern of the D-loop region of at least one site selected from the group consisting of.
[0108] In some embodiments, determining the methylation pattern comprises (iv) the CpG sites of the ND1 gene shown in Table 2, (v) the CHG sites of the ND1 gene shown in Table 4, and (vi) the CHH sites of the ND1 gene shown in Table 6 and comprises the step of determining the methylation pattern of the ND1 region of at least one site selected from the group consisting of.
[0109] In some embodiments, determining the methylation pattern includes determining the methylation pattern of at least one CpG site in the D-loop region of Table 1. In some embodiments, determining the methylation pattern includes determining the methylation pattern of at least five CpG sites in the D-loop region of Table 1. In some embodiments, determining the methylation pattern includes determining the methylation pattern of at least ten CpG sites in the D-loop region of Table 1. In some embodiments, determining the methylation pattern includes determining the methylation pattern of all CpG sites in the D-loop region of Table 1.
[0110] In some embodiments, determining the methylation pattern includes determining the methylation pattern of at least one CHG site in the D-loop region of Table 3. In some embodiments, determining the methylation pattern includes determining the methylation pattern of at least five CHG sites in the D-loop region of Table 3. In some embodiments, determining the methylation pattern includes determining the methylation pattern of at least ten CHG sites in the D-loop region of Table 3. In some embodiments, determining the methylation pattern includes determining the methylation pattern of all CHG sites in the D-loop region of Table 3.
[0111] In some embodiments, determining the methylation pattern includes determining the methylation pattern of at least one CHH site in the D-loop region of Table 5. In some embodiments, determining the methylation pattern includes determining the methylation pattern of at least five CHH sites in the D-loop region of Table 5. In some embodiments, determining the methylation pattern includes determining the methylation pattern of at least ten CHH sites in the D-loop region of Table 5. In some embodiments, determining the methylation pattern includes determining the methylation pattern of at least twenty-five CHH sites in the D-loop region of Table 5. In some embodiments, determining the methylation pattern includes determining the methylation pattern of at least fifty CHH sites in the D-loop region of Table 5. In some embodiments, determining the methylation pattern includes determining the methylation pattern of all CHH sites in the D-loop region shown in Table 5.
[0112] In some embodiments, determining the methylation pattern includes determining the methylation pattern of at least one CHG site in the D-loop region of Table 3. In particular, determining the methylation pattern includes determining the methylation pattern of all CHG sites in the D-loop region of Table 3. In certain embodiments, hypomethylation at at least one site of the CHG sites in the D-loop region indicates that the subject has DLB. In some embodiments, determining the methylation pattern further includes determining the methylation pattern of at least one CHH site in the D-loop region of Table 5 and / or at least one CpG site in the D-loop region of Table 1. In particular, determining the methylation pattern further includes determining the methylation pattern of all CHH sites in the D-loop region of Table 5 and / or all CpG sites in the D-loop region of Table 1. In certain embodiments, hypomethylation at at least one site of the CHH sites in the D-loop region and hypomethylation at at least one site of the CpG sites in the D-loop region indicate that the subject has DLB.
[0113] In some embodiments, determining the methylation pattern includes determining the methylation pattern of at least one CpG site of the ND1 gene in Table 2. In some embodiments, determining the methylation pattern includes determining the methylation pattern of at least five CpG sites of the ND1 gene in Table 2. In some embodiments, determining the methylation pattern includes determining the methylation pattern of at least ten CpG sites of the ND1 gene in Table 2. In some embodiments, determining the methylation pattern includes determining the methylation pattern of all CpG sites of the ND1 gene in Table 2.
[0114] In some embodiments, determining the methylation pattern includes determining the methylation pattern of at least one CHG site of the ND1 gene in Table 4. In some embodiments, determining the methylation pattern includes determining the methylation pattern of at least five CHG sites of the ND1 gene in Table 4. In some embodiments, determining the methylation pattern includes determining the methylation pattern of all CHG sites of the ND1 gene in Table 4.
[0115] In some embodiments, determining the methylation pattern includes determining the methylation pattern of at least one CHH site of the ND1 gene in Table 6. In some embodiments, determining the methylation pattern includes determining the methylation pattern of at least five CHH sites of the ND1 gene in Table 6. In some embodiments, determining the methylation pattern includes determining the methylation pattern of at least ten CHH sites of the ND1 gene in Table 6. In some embodiments, determining the methylation pattern includes determining the methylation pattern of at least twenty-five CHH sites of the ND1 gene in Table 6. In some embodiments, determining the methylation pattern includes determining the methylation pattern of at least fifty CHH sites of the ND1 gene in Table 6. In some embodiments, determining the methylation pattern includes determining the methylation pattern of all CHH sites of the ND1 gene shown in Table 6.
[0116] In some embodiments, determining the methylation pattern includes determining the methylation pattern of at least one CHH site of the ND1 gene in Table 6. In particular, determining the methylation pattern includes determining the methylation pattern of all CHH sites of the ND1 gene in Table 6. In certain embodiments, no significant difference in the methylation level of at least one site of the CHH site of the ND1 gene indicates that the subject has DLB. Alternatively, hypomethylation of at least one site of the CHH site of the ND1 gene indicates that the subject has DLB. In some embodiments, determining the methylation pattern further includes determining the methylation pattern of at least one CpG site of the ND1 gene in Table 2 and / or at least one CHG site of the ND1 gene in Table 4. In particular, determining the methylation pattern includes determining the methylation pattern of all CpG sites of the ND1 gene in Table 2 and / or all CHG sites of the ND1 gene in Table 4. In certain embodiments, no significant difference in the methylation level at at least one site of the CpG site of the ND1 gene and no significant difference in the methylation level at at least one site of the CHG site of the ND1 gene indicate that the subject has DLB.
[0117] In some embodiments, determining the methylation pattern includes determining the methylation of all CpG sites, CHG sites, and CHH sites in the D-loop regions shown in Tables 1, 3, and 5.
[0118] In some embodiments, determining the methylation pattern includes determining the methylation of all CpG sites, CHG sites, and CHH sites of the ND1 gene shown in Tables 2, 4, and 6.
[0119] In some embodiments, determining the methylation pattern includes determining the methylation pattern of all CpG sites, CHG sites, and CHH sites in the D-loop region and the ND1 gene.
[0120] The methylation pattern can be determined by any method known in the art. In some embodiments, the methylation pattern is determined by a technique selected from the group consisting of bisulfite treatment-based techniques, biology identification-based techniques, and bisulfite-free and enzyme-free techniques.
[0121] In some embodiments, examples of bisulfite treatment-based techniques include, but are not limited to, sequence-based analysis, melting temperature-based analysis, and interaction-based analysis. In some embodiments, examples of sequence-based analysis include, but are not limited to, bisulfite sequencing, methylation-specific PCR (MS-PCR), methylation-sensitive single-stranded nucleotide primer extension method (Ms-SnuPE), and RRBS (reduced representation bisulfite sequencing). In some embodiments, examples of melting temperature-based analysis include, but are not limited to, methylation-specific denaturant gradient gel electrophoresis (MS-DGGE), methylation-specific melting curve analysis (MS-MCA), and methylation-specific HRM (High-Resolution Melting) (MS-HRM). In some embodiments, examples of interaction-based analysis include, but are not limited to, COBRA (combined bisulfite-restriction analysis), and Methylight assay.
[0122] In some embodiments, examples of biology identification-based techniques include, but are not limited to, methods based on enzymatic digestion and methods based on biotic-dependent reactions. In some embodiments, examples of methods based on enzymatic digestion include, but are not limited to, RLGS (Restriction-landmark genomic scanning), online monitoring, and methylation-sensitive restriction enzyme-PCR (MS-RE-PCR / Southern). In certain embodiments, the biotic-dependent reaction is methyl capture using a methyl-CpG binding domain (MBD) protein.
[0123] In some embodiments, bisulfite-free and enzyme-free techniques include, but are not limited to, analysis based on direct oxidation and analysis based on chemical decomposition of oxidation. In certain embodiments, the analysis based on direct oxidation is MWCNTs / Ch / GCE (chloride monolayer supported multiwalled carbon nanotubes). In certain embodiments, the analysis based on chemical decomposition of oxidation is Na1O4 / LiBr.
[0124] In certain embodiments, the methylation pattern is determined by techniques based on bisulfite treatment. In particular, the methylation pattern is determined by sequence-based analysis. More specifically, the methylation pattern is determined by bisulfite sequencing.
[0125] In some embodiments, the methylation pattern is determined by a sequencing technique selected from the group consisting of methylation-specific PCR (MS-PCR), quantitative methylation-specific polymerase chain reaction (qMSP), bisulfite sequencing, pyrosequencing, nanopore sequencing, MassArray, methylation-sensitive single nucleotide primer extension (Ms-SnuPE), RRBS (reduced representation bisulfite sequencing), methylation-specific denaturing gradient gel electrophoresis (MS-DGGE), methylation-specific melting curve analysis (MS-MCA), methylation-specific HRM (MS-HRM), COBRA (combined bisulfite-restriction analysis), and Methylight assay, methylation-specific restriction endonuclease analysis (MSRE), methylation-sensitive restriction enzyme sequencing (MRE-seq), RLGS (Restriction-landmark genomic scanning), methylation-DNA immunoprecipitation MeDIP or MeDIP-seq, methyl capture using methyl-CpG binding domain (MBD) proteins, ChIP assay, methylation array, MWCNTs / ch / CGE (choline chloride monolayer supported multiwalled carbon nanotubes), and chemical oxidative degradation (NaIO4 / LiBr).
[0126] In certain embodiments, the methylation pattern is determined by bisulfite sequencing. In some embodiments, bisulfite sequencing includes treating the sample with bisulfite and separately sequencing the bisulfite-treated sample by PCR. In certain embodiments, the bisulfite-treated sample is sequenced using a kit, which can be, but is not limited to, a kit manufactured by Illumina. More specifically, the bisulfite-treated sample is sequenced using a kit selected from the group consisting of MiSeq reagent Kit v3-600-cycles (#MS-102-3003, Illumina), MiSeq reagent Kit v2-500-cycles (# MS-102-2003, Illumina), and MiSeq reagent Nano Kit v2-500 cycles (# MS-103-1003, Illumina).
[0127] Sequencing of the sample can be performed using any method known in the art. Examples of sequencing platforms include, but are not limited to, Roche, Illumina, Life Technologies, Polonator, Helicos Bioscience, Pacific Biosciences, HTG Molecular Diagnostic, Singular Genomics, Element Biosciences, Oxford Nanopore, and Nanostring Technology.
[0128] In some embodiments, determining the methylation pattern includes a step of quantifying the library. In some embodiments, the quantification of the library is performed using fluorescence quantification methods. Alternatively, the quantification of the methylation pattern is performed using fluorescence. In certain embodiments, the fluorescence quantification method is characterized by using a kit containing a dsDNA-binding dye. In certain embodiments, the quantification of the library is performed using a Qubit® 3.0 Fluorometer manufactured by Thermofisher Scientific or other kits manufactured by Thermofisher Scientific. More specifically, the quantification of the library is performed using the Qubit® dsDNA Assay Kit. In particular, the Qubit® dsDNA Assay Kit is selected from the group consisting of Qubit® dsDNA HS Assay Kit #Q32854 and Qubit® dsDNA BR Assay Kit #Q32850.
[0129] It is understood that the quantification of the library using the fluorescence quantitative quantification method in the present invention can be performed using a fluorometer from other brands and a kit containing a dDNA-binding dye other than Thermofisher Scientific. An example of another suitable fluorometer is, but not limited to, the QFX fluorometer manufactured by DeNovix and the Quantus® fluorometer manufactured by Promega. Examples of other suitable dsDNA fluorescence kits are, but not limited to, QuantiFluor® Dye Systems and QuantiFluor® dsDNA manufactured by Promega. Furthermore, the QFX fluorometer manufactured by DeNovix operates with either its own DeNovix dsDNA Fluorescence Quantification Kit or other common commercially available assays.
[0130] In an embodiment of the present invention, the analysis for comparing the methylation levels of each methylation site is performed using the DSS (Dispersion Shrinkage for Sequencing data) Bioconductor package. In other embodiments, any other suitable methodology used in the art can be used. In some embodiments, the analysis is performed using a beta-binomial based model. In certain embodiments, this model is a beta-binomial generalized linear model using a Logit link function. In another specific embodiment, the model performs local smoothing within each sample and then applies a beta regression using a probit link function. In some embodiments, the analysis is performed using an empirical Bayes method, a standard maximum likelihood estimation, or another approach applying other methods known in the art. In some embodiments, the analysis includes making inferences using standard techniques known in the art or other suitable techniques to be developed in the future. In certain embodiments, the inferences are made using standard techniques selected from the group consisting of Wald test, likelihood ratio test, permutation test, and ANOVA.
[0131] In an embodiment of the present invention, the threshold corrected p-value for establishing differential methylation is 0.05. In other embodiments, the threshold corrected p-value is established at 0.25. In other embodiments, the threshold corrected p-value is established at 0.2. In other embodiments, the threshold corrected p-value is established at 0.15. In other embodiments, the threshold corrected p-value is established at 0.1. In other embodiments, the threshold corrected p-value is established at 0.09. In other embodiments, the threshold corrected p-value is established at 0.08. In other embodiments, the threshold corrected p-value is established at 0.07. In other embodiments, the threshold corrected p-value is established at 0.06. In other embodiments, the threshold corrected p-value is established at 0.05. In other embodiments, the threshold corrected p-value is established at 0.04. In other embodiments, the threshold corrected p-value is established at 0.03. In other embodiments, the threshold corrected p-value is established at 0.02. In other embodiments, the threshold corrected p-value is established at 0.01.
[0132] Samples and Sample Processing As used herein, the term "biological sample" or "sample" refers to a biological material isolated from a subject. A biological sample can include any biological material suitable for determining methylation patterns, for example, by treating and sequencing nucleic acids.
[0133] In some embodiments, the sample is selected from a biological fluid or a biopsy of solid tissue. In particular, the sample is selected from the group consisting of blood, plasma, saliva, cerebrospinal fluid, brain sample, skin sample, and urine. In particular, the sample is blood, especially peripheral blood.
[0134] The source of the sample can be, for example, solid tissue, tissue sample, biopsy, or aspirate from fresh, frozen and / or preserved organs. In some embodiments, the sample is a cell-free sample, for example, containing cell-free nucleic acids (such as DNA or RNA). The sample can, in some embodiments, contain compounds that do not naturally mix with tissue, such as preservatives, anticoagulants, buffers, fixatives, nutrients, antibacterial agents, etc.
[0135] In some embodiments, the method includes the step of obtaining a sample. In certain embodiments, the sample is blood or plasma, and the sample is extracted by using a needle. In another embodiment, the sample is saliva, and the sample is obtained by using a method selected from the group consisting of a draining method, a spitting method, a suction method, and a swab method. In some embodiments, the sample can be obtained, for example, from surgical materials or a biopsy. In some embodiments, the biopsy can be tissue stored from a previous treatment line. In some embodiments, the biopsy can be from untreated tissue.
[0136] In some embodiments, the sample is frozen or stored. In some embodiments, the sample is stored as a frozen sample or as a paraffin-embedded (FFPE) tissue preparation fixed with formalin, formaldehyde, or paraformaldehyde. For example, the sample can be embedded in a matrix, such as a DDPE block or a frozen sample. In some embodiments, the sample can include aspirates, scrapings, bone marrow specimens; tissue biopsy specimens; surgical specimens, etc. In some embodiments, the sample is or includes cells obtained from an individual, such as the individual from whom the sample was obtained.
[0137] In some embodiments, the sample is a fresh sample (or an unpreserved sample) or a stored sample. As used herein, the terms “fresh sample,” “unpreserved sample,” and their grammatical variations refer to a sample that has been processed more than a predetermined period, such as more than one week, after being extracted from a subject. In some embodiments, a fresh sample is not frozen. In some embodiments, a fresh sample is not fixed. In some embodiments, a fresh sample is stored for less than about two weeks, less than about one week, or less than six days, five days, four days, three days, or two days before processing. As used herein, the terms “stored sample” and “their grammatical variations” refer to a sample that has been processed after a predetermined period, such as after one week, after being extracted from a subject. In some embodiments, a stored sample is frozen. In some embodiments, a stored sample is fixed. In some embodiments, a stored sample has a known diagnostic history and / or treatment history. In some embodiments, a stored sample is stored for at least one week, at least one month, at least six months, or at least one year before processing.
[0138] Oligonucleotides and Kits In some embodiments, in any of the methods described herein, the methylation pattern is determined using at least one oligonucleotide capable of specifically hybridizing to mitochondrial DNA comprising a D-loop region and / or the ND1 gene. In particular, the oligonucleotide is capable of specifically hybridizing under high stringency conditions.
[0139] The sequence of interest refers to a reference sequence or a sequence resulting from a specific modification treatment, such as a bisulfite treatment in which unmethylated cytosine is modified to uracil. In some embodiments, the oligonucleotide hybridizes to a reference mitochondrial DNA sequence comprising a D-loop region and / or the ND1 gene. In certain embodiments, the oligonucleotide hybridizes to a modified mitochondrial DNA sequence comprising a D-loop region. Alternatively, the oligonucleotide hybridizes to a modified mitochondrial DNA comprising the ND1 gene. In certain embodiments, the modified mitochondrial DNA sequence has been modified by bisulfite treatment. In more specific embodiments, the modified mitochondrial DNA sequence has been modified by bisulfite treatment, where unmethylated cytosine has been modified to uracil.
[0140] In some embodiments, the methylation pattern is (i) CpG sites in the D-loop region shown in Table 1, (ii) CHG sites in the D-loop region shown in Table 3, (iii) CHH sites in the D-loop region shown in Table 5, (iv) CpG sites in the ND1 gene shown in Table 2, (v) CHG sites in the ND1 gene shown in Table 4, and (vi) CHH sites in the ND1 gene shown in Table 6 determined using at least one oligonucleotide capable of specifically hybridizing to mitochondrial DNA comprising at least one methylation site selected from the group consisting of.
[0141] Alternatively, the oligonucleotide of the present invention is referred to as a nucleic acid sequence. In some embodiments, the oligonucleotide is a DNA sequence. In some embodiments, the oligonucleotide has a length of 15 to 100 nucleotides.
[0142] In one embodiment, the methylation pattern is determined using an oligonucleotide that can specifically hybridize to and amplify a mitochondrial DNA sequence containing nucleotides 16,465 to 230 (corresponding to the D-loop region) of the NCBI reference sequence: NC_012920.1, and / or nucleotides 3,257 to 3,682 (corresponding to the ND1 gene) of the NCBI reference sequence: NC_012920.1.
[0143] The inventors have designed primers herein (also referred to as oligonucleotides or nucleic acid sequences herein) that contain a minimum number of cytosines. In some embodiments, the primers are degenerate to encompass all possible methylated and non-methylated scenarios due to the ambiguous C / U conversion of several cytosine residues contained in the sequence. These primers are a mixture of oligonucleotide sequences containing several possible nucleotide bases at specific positions. As a result, the probability of detecting mitochondrial methylation is higher. The degenerate forward primer contains a Y that refers to either C or T (Y = C / T) at any position where the reference sequence is C. Further, as is known in the art, the reverse primer corresponds not to the reference sequence but to the complementary reverse sequence of the reference sequence. Thus, the reverse primer does not contain the C site of the reference sequence but contains their complementary G site. Thus, the degenerate reverse primer contains an R that refers to either G or A (R = A / G) at any position where the reference sequence is C or the complementary sequence contains G.
[0144] Thus, in some embodiments, the oligonucleotide is a degenerate oligonucleotide. A degenerate oligonucleotide / primer is a mixture of oligonucleotide sequences containing multiple possible bases that provide a population of oligonucleotides with similar sequences where some positions contain all possible combinations of nucleotides.
[0145] In certain embodiments, the methylation pattern is determined using at least one oligonucleotide having a length of 15 to 100 nucleotides, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, and SEQ ID NO:4. The oligonucleotides can include additional nucleotides (i.e., sequencing adapters) at their ends to be suitable for use, for example, in sequencing.
[0146] In certain embodiments, the methylation pattern is determined using at least one oligonucleotide selected from the group consisting of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, and SEQ ID NO:4. In more specific embodiments, the methylation pattern is determined using oligonucleotides of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, and SEQ ID NO:4. In another embodiment, the methylation pattern is determined using at least one oligonucleotide selected from SEQ ID NO:1 and SEQ ID NO:2. In some embodiments, the methylation pattern is determined using oligonucleotides of SEQ ID NO:1 and SEQ ID NO:2.
[0147] Another aspect of the present invention relates to an oligonucleotide having a length of 15 to 100 nucleotides and comprising a nucleic acid sequence selected from the group consisting of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, and SEQ ID NO: 4. As mentioned, in one embodiment, the oligonucleotide comprises a sequencing adapter at the ends of SEQ ID NO: 1, 2, 3, and / or 4. In particular, the oligonucleotide is selected from the group consisting of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, and SEQ ID NO: 4. In certain embodiments, the nucleic acid sequence is SEQ ID NO: 1. In another specific embodiment, the oligonucleotide is SEQ ID NO: 2. In another specific embodiment, the oligonucleotide is SEQ ID NO: 3. In another specific embodiment, the oligonucleotide is SEQ ID NO: 4.
[0148] Another aspect of the present invention relates to the use of an oligonucleotide having a length of 15 to 100 nucleotides and comprising a nucleic acid sequence selected from the group consisting of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, and SEQ ID NO: 4 for determining the methylation pattern of mitochondrial DNA. In particular, the oligonucleotide is selected from the group consisting of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, and SEQ ID NO: 4. In certain embodiments, the present invention relates to the use of an oligonucleotide having SEQ ID NO: 1 for determining the methylation pattern of mitochondrial DNA comprising a D-loop region. In another embodiment, the present invention relates to the use of an oligonucleotide having SEQ ID NO: 2 for determining the methylation pattern of mitochondrial DNA comprising a D-loop region. In another embodiment, the present invention relates to the use of an oligonucleotide having SEQ ID NO: 3 for determining the methylation pattern of mitochondrial DNA comprising the ND1 gene. In another embodiment, the present invention relates to the use of an oligonucleotide having SEQ ID NO: 4 for determining the methylation pattern of mitochondrial DNA comprising the ND1 gene.
[0149] In some embodiments, the present invention relates to the use of oligonucleotides selected from SEQ ID NO: 1 and SEQ ID NO: 2 for determining the methylation pattern of mitochondrial DNA sequences containing the D-loop region. In another embodiment, the present invention relates to the use of oligonucleotides selected from SEQ ID NO: 1 and SEQ ID NO: 2 for determining the methylation pattern for diagnosing or identifying DLB in a subject.
[0150] In other embodiments, the present invention relates to the use of oligonucleotides having a length of 15 to 100 nucleotides, comprising nucleic acid sequences selected from the group consisting of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, and SEQ ID NO: 4 (particularly with SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, and SEQ ID NO: 4) for determining the methylation pattern of mitochondrial DNA sequences containing the D-loop region and / or the ND1 gene. In another embodiment, the present invention relates to the use of said oligonucleotides for determining the methylation pattern for diagnosing or identifying DLB in a subject.
[0151] In some embodiments, the oligonucleotides defined herein are included in a kit. In this regard, another aspect of the present invention relates to the use of a kit for determining the methylation pattern of mitochondrial DNA. In another aspect, the present invention relates to the use of the kit as defined above for determining the methylation pattern of mitochondrial DNA for diagnosing or determining DLB in a subject. In another aspect, the present invention relates to the use of the kit as defined above according to the methods described herein.
[0152] Such a kit may include a plurality of containers each containing one or more of the various reagents (e.g., in concentrated form) utilized in the method, such as one or more oligonucleotides (e.g., oligonucleotides having SEQ ID NOs: 1-4 provided herein, particularly oligonucleotides having SEQ ID NOs: 1-2 provided herein). The kit may also provide reagents, buffers, and / or instruments for assisting in the practice of the methods provided herein.
[0153] The kit provided according to the present invention may also include a brochure or instruction manual describing the methods disclosed herein for diagnosing or identifying DLB in a subject or their practical applications. The instructions included in the kit may be attached to the packaging material or may be included as an accompanying document. The instructions are usually written or printed materials, but these are not limited to such materials. All media capable of storing such instructions and transmitting them to an end user are contemplated. Such media include, but are not limited to, electronic storage media (e.g., magnetic disks, tapes, cartridges, chips), optical media (e.g., CDROM), etc. As used herein, the term "instructions" may include the address of an Internet site that provides the instructions.
[0154] In some embodiments, the kit is an Illumina sequencing kit. More specifically, the kit is selected from the group consisting of MiSeq reagent Kit v3-600-cycles (#MS-102-3003, Illumina), MiSeq reagent Kit v2-500-cycles (# MS-102-2003, Illumina), and MiSeq reagent Nano Kit v2-500 cycles (# MS-103-1003, Illumina).
[0155] Clinical variables The methods described herein are capable of diagnosing or identifying DLB in a subject. Information derived from the clinical variables of the subject can be used in the methods described herein to improve or complement the diagnosis. Further, the clinical variables can be used as input data for the classification models described herein. Further, the classification models can be trained with a training data set that includes clinical variables related to a plurality of subjects.
[0156] Thus, in some embodiments, any method described herein further comprises the step of considering at least one clinical variable of the subject. In other embodiments, any method described herein further comprises the step of combining a methylation pattern described herein with at least one clinical variable of the subject. In particular, any method of the invention further comprises the step of considering at least one clinical variable of the subject to identify or diagnose DLB of the subject or to identify a subject having DLB.
[0157] Clinical variables can include any clinical variables known in the art. In some embodiments, the method comprises the step of considering at least one clinical variable of the subject or the step of combining a methylation pattern described herein with at least one clinical variable of the subject. In some embodiments, the clinical variable is selected from the group consisting of demographic variables, neuropsychological variables, clinical finding variables, and clinical test variables.
[0158] Demographic variables "Gender" is a categorical variable that includes two possible variables: female or male.
[0159] "Age" is a numerical variable corresponding to the age of the subject.
[0160] "Race" is a categorical variable that includes two or more possible categories: white or Caucasian, black, American Indian, Latin or Hispanic, Asian, etc. "Race" is a category of people who share certain specific physical characteristics.
[0161] "Education" is a numerical variable corresponding to the number of years of education since age 6.
[0162] "Family history" is a categorical variable that includes two possible categories: a history of dementia within the family or no history of dementia within the family.
[0163] In one embodiment, the demographic variables are selected from gender, age, race, education, and family history.
[0164] Neuropsychological variables Neuropsychological tests may be useful for stratifying the level of cognitive impairment and may be useful when a clinician suspects a diagnosis of DLB or attempts to distinguish DLB from other types of dementia. Given the cognitive profile of DLB, neuropsychological assessment should focus on attention, executive function, visuospatial ability, language function, and memory (immediate recall and delayed recall). Neuropsychological tests may vary across facilities / countries but assess the same cognitive functions.
[0165] Some of the neuropsychological tests are the Clinical Dementia Rating (CDR), CDR-SOB (Clinical Dementia Rating Scale Sum of Boxes), Mini-Mental State Examination (MMSE), MoCA (Montreal Cognitive Assessment), GDS (Global Deterioration Scale), FAQ (Functional Activities Questionnaire), GDS-Yesavage (Geriatric Depression Scale Yesavage), NPI (Neuropsychiatric Inventory score), Delayed Verbal Memory score, Verbal Learning Curve score, Verbal Recognition score, Semantic Verbal Fluency score, Constructional Praxis score, Clock drawing test-Luria.
[0166] "MMSE" (Mini-Mental State Examination) is a numerical variable that refers to the assessment of five main items: orientation, registration, attention and calculation, memory and language, and recall, with an output score of 1 - 30.
[0167] When the subject is diagnosed with MCI but the diagnosis is unclear, the following neuropsychological variables can be considered: CDR is a global summary obtained through semi-structured interviews of the patient and informant, and is a numerical variable referring to the "Clinical Dementia Rating Scale" that classifies the subject's cognitive state into six domains of functions including memory, orientation, judgment, and problem-solving, social adaptation, family situation and interests, and caregiving situation. A score of 0-3 is assigned to each domain, and the results are calculated by a computer via an algorithm to obtain the CDR global score. The CDR global score ranges from 0 to 3 and enables grouping of subjects according to the severity of dementia (where CDR = 0 corresponds to the absence of cognitive impairment, CDR = 0.5 is suspected or very mild dementia, CDR = 1 is mild dementia, CDR = 2 is moderate dementia, and CDR = 3 is severe dementia).
[0168] Subjects with a CDR of 0.5 are diagnosed with MCI, which is early memory loss or loss of other cognitive abilities such as language or visual / spatial perception in subjects who maintain the ability to perform most activities of daily life independently.
[0169] CDR-SOB (Clinical Dementia Rating Scale Sum of Boxes) is a numerical variable referring to the "Sum of Boxes Score", which is a score in the range of 0 to 18 obtained by summing each of the above domain box scores for the calculation of the CDR global score.
[0170] In one embodiment, the neuropsychological variable is selected from the group consisting of MMSE, CDR, and CDR-SOB. More specifically, the neuropsychological variable is MMSE.
[0171] In one embodiment, the neuropsychological variable is a praxis test, a Luria test, CDR, GDS, MMSE, or CDR-SOB.
[0172] Clinical finding variables "Neuroleptic intolerance" is a categorical variable that includes two possible variables: neuroleptic intolerance or non-neuroleptic intolerance. This variable corresponds to the specific sensitivity of DLB patients to neuroleptics regarding the onset of extrapyramidal symptoms.
[0173] "REM sleep behavior disorder" is a categorical variable that includes two possible variables: the presence or absence of REM sleep behavior disorder. The behavior during REM sleep is a parasomnia accompanied by abnormal behavior that coincides with dreams during REM sleep.
[0174] "Autonomic neuropathy" is a categorical variable that includes two possible categories: the presence or absence of autonomic neuropathy. Autonomic neuropathy or autonomic dysfunction is a state in which the function of the autonomic nervous system (ANS) is impaired.
[0175] "Parkinsonism" is a categorical variable that includes two possible categories: the presence or absence of parkinsonism. Parkinsonism is a clinical symptom characterized by tremors, slowed movements, rigidity, and postural instability.
[0176] "Visual hallucinations" is a categorical variable that includes two possible categories: the presence or absence of visual hallucinations. Visual hallucinations are visual perceptions in the absence of external visual stimuli that have the quality of real visual perceptions.
[0177] "Fluctuations in cognitive function" is a categorical variable that includes two possible categories: the presence or absence of fluctuations in cognitive function. Fluctuations in cognitive function are spontaneous changes in concentration, attention, focus, and arousal that can vary significantly from day to day or even within the same day.
[0178] A categorical variable with only two possible categories and a clinical variable as defined herein may be considered a categorical variable with more than two options in cases where not only the presence or absence of such a variable is determined, but also the degree, grade, or rank of the variable. Similarly, a categorical variable with two or more possible categories and a clinical variable as defined herein may also be considered a numerical variable.
[0179] In one embodiment, the clinical finding variables are neuroleptic intolerance, REM sleep behavior disorder, autonomic neuropathy, parkinsonism, visual hallucinations, and cognitive fluctuations.
[0180] Clinical test variables In some embodiments, the clinical variable may include a clinical test variable such as a test obtained via a neuroimaging technique or a medical imaging technique. In some embodiments, the clinical variable is selected from the group consisting of MRI, DaTSCAN, and 18F-FDG-PET.
[0181] "MRI" (magnetic resonance imaging) is a categorical variable that includes two possible categories: positive MRI and negative MRI. MRI is a type of scan that uses a strong magnetic field and radio waves to produce detailed images of the inside of the body.
[0182] "DaTSCAN" (dopamine transporter scan) is a categorical variable that includes two possible categories: positive DaTSCAN and negative DaTSCAN. DaTSCAN is a diagnostic method used to determine the loss of dopaminergic neurons in the striatum.
[0183] Positron emission tomography (PET) is a type of nuclear medicine procedure that measures the metabolic activity of cells in body tissues. PET can be performed using different types of tracers, each of which is used for a specific purpose, test, or detection. For example, PET can measure glucose levels, beta-amyloid plaques, or tau proteins. FDG-PET refers to PET designed to detect glucose in tissues or body parts and, consequently, analyze metabolic activity. In another example, PET is used to determine the presence of beta-amyloid plaques in the brain. PET can also be used to measure glucose, beta-amyloid, and / or tau.
[0184] "18F-FDG-PET" (positron emission tomography with F18-fluorodeoxyglucose) is a categorical variable that includes two possible categories: positive 18F-FDG-PET and negative 18F-FDG-PET. "18F-FDG-PET" is an imaging metabolic technique that uses a radioactive tracer and refers to PET designed to detect glucose and, consequently, analyze the metabolic activity of tissues or body parts.
[0185] Clinical variables defined in this specification as categorical variables with only two possible categories can be considered categorical variables with more than two options in cases where not only the presence or absence of such a variable is determined, but also the degree, grade, or rank of the variable. Similarly, clinical variables defined in this specification as categorical variables with more than two possible categories can also be considered numerical variables.
[0186] The following clinical test variables can be contemplated when a subject has been diagnosed with MCI, but the diagnosis is still unclear.
[0187] "Amyloid-PET" is a categorical variable that includes two possible categories: positive amyloid PET and negative amyloid PET, each indicating the presence or absence of beta-amyloid. Amyloid PET is used to determine the presence of beta-amyloid plaques in the brain.
[0188] Clinical variables can include biomarkers known in the art, as well as other neuropsychological tests and other radiimaging techniques known in the art. In some embodiments, the clinical variables include the presence, absence, or level of a biomarker known in the art (e.g., APOE). In particular, the clinical variables include the presence or absence of said biomarker. In some embodiments, the biomarker can be measured in any sample or body fluid from the subject, particularly in a cerebrospinal fluid sample. In other embodiments, the biomarker is measured in a blood sample. In other embodiments, the biomarker can be detected using other radiimaging techniques, such as positron emission tomography (PET). In some embodiments, the clinical variables include APOE, alE4, β-42, tau-T, and / or tau-P.
[0189] "APOE" is a categorical variable that includes categories corresponding to different genotypes of apolipoprotein E: E2.E2, E2.E3, E2.E4, E3.E3, E3.E4, and E4.E4.
[0190] "alE4" is a categorical variable corresponding to the recategorization of the APOE genotype into the following categories: 0 (including E2.E2, E2.E3, and E3.E3), 1 (including E2.E4 and E3.E4), and 2 (including E4.E4).
[0191] "β-42" is a categorical variable indicating the presence or absence of the β-42 peptide in a cerebrospinal fluid sample.
[0192] "tau-T" is a categorical variable indicating the presence or absence of the protein tau in a cerebrospinal fluid sample.
[0193] "Tau-P" is a categorical variable indicating the presence or absence of phosphorylated protein tau in a cerebrospinal fluid sample.
[0194] In some embodiments, the clinical test variables are magnetic resonance imaging (MRI), dopamine transporter scan (DaTSCAN), positron emission tomography with 18F-fluorodeoxyglucose (18F-FDG-PET), amyloid PET, apolipoprotein E (APOE) genotype, alE4, β-42, tau-T, and tau-P.
[0195] In some embodiments, the clinical variables can include variables related to the current treatment or drug therapy of the subject for dementia or the current treatment or drug therapy of the subject for symptoms related to dementia (i.e., combination drug therapy).
[0196] In some embodiments, the clinical variables are selected from the group consisting of gender, age, race, education, family history, MMSE, neuroleptic intolerance, REM sleep behavior disorder, autonomic neuropathy, parkinsonism, visual hallucinations, cognitive fluctuations, MRI, DaTSCAN, and 18F-FDG-PET.
[0197] In another embodiment, the clinical variables are selected from the group consisting of CDR, SOB, amyloid PET, APOE, alE4, β-42, tau-T, and tau-P.
[0198] In some embodiments, the clinical variables are selected from the group consisting of gender, age, race, education, family history, praxis test, Luria test, CDR, GDS, MMSE, CDR-SOB, neuroleptic intolerance, REM sleep behavior disorder, autonomic neuropathy, parkinsonism, visual hallucinations, cognitive fluctuations, magnetic resonance imaging (MRI), dopamine transporter scan (DaTSCAN), positron emission tomography with 18F-fluorodeoxyglucose (18F-FDG-PET), amyloid PET, apolipoprotein E (APOE) genotype, alE4, β-42, tau-T, and tau-P.
[0199] In certain embodiments, the clinical variables are selected from the group consisting of age, CDR, GDS, REM sleep behavior disorder, autonomic neuropathy, parkinsonism, visual hallucinations, cognitive fluctuations, PET-FDG, and DatTSCAN.
[0200] Classification model In another aspect, the present invention provides a classification model capable of classifying a subject into two categories: subjects having DLB and subjects not having DLB. These classes are associated with subjects suffering from DLB and subjects different from DLB. Subjects not classified as having DLB can be healthy subjects, subjects with earlier dementia, or subjects suffering from other types of dementia or other diseases.
[0201] Thus, the present invention provides a classification model for identifying a target with DLB, which uses data on methylation patterns obtained from a sample of the target and / or data on clinical variables of the target to identify the target as belonging to a class from the group consisting of targets with DLB and targets without DLB, and being identified as a target with DLB indicates that the target is suffering from DLB. In some embodiments, the classification model uses the methylation pattern obtained from a sample of the target as data to identify the target as belonging to a class from the group consisting of targets with DLB and targets without DLB, and being identified as a target with DLB indicates that the target is suffering from DLB. In other embodiments, the classification model uses the clinical variables of the target as data to identify the target as belonging to a class from the group consisting of targets with DLB and targets without DLB, and being identified as a target with DLB indicates that the target is suffering from DLB.
[0202] Thus, the present invention provides a method for identifying Lewy body dementia in a subject, comprising: (a) determining, in a sample of the subject containing mitochondrial DNA, a methylation pattern of the D-loop region and / or the ND1 gene of the mitochondrial DNA, wherein the methylation pattern is (i) a CpG site in the D-loop region shown in Table 1, (ii) a CHG site in the D-loop region shown in Table 3, (iii) a CHH site in the D-loop region shown in Table 5, (iv) a CpG site in the ND1 gene shown in Table 2, (v) a CHG site in the ND1 gene shown in Table 4, and (vi) a CHH site in the ND1 gene shown in Table 6 determined at at least one site selected from the group consisting of: (b) Optionally, combining the methylation patterns of the one or more sites determined in step (a) with at least one clinical variable of the subject, wherein the combination is performed using a classification model for determining a score that correlates with the identification of Lewy body dementia in the subject, step and Provided is a method comprising.
[0203] In some embodiments, the classification model is obtained by an artificial intelligence method. In certain embodiments, the classification model is obtained by a machine learning method. In more specific embodiments, the classification model is obtained by a supervised machine learning method or related software. In more specific embodiments, the supervised machine learning method or related software includes linear discriminant analysis (LDA), CART (Classification and Regression Trees), k-nearest neighbor method (kNN), naive Bayes (NB), support vector machine using a linear kernel (SVM), random forest (RF), neural network (NNET), generalized boosted regression model (GBM), binary logistic regression (GLM), artificial neural network (ANN), GBoost (XGB; an implementation of gradient boosting designed for speed and performance), GLMNET (a generalized linear model via penalized maximum likelihood, a package suitable for implementing a generalized linear model via penalized maximum likelihood, for example logistic regression), cforest (an implementation of a random forest and bagging ensemble algorithm using conditional inference trees as the base learner), Treebag (bagging, i.e., bootstrap aggregation, an algorithm for building multiple models from individual subsets of training data and constructing a final aggregated model to improve the accuracy of the model in regression and classification problems), AdaBoost (AdaBoost Classification Trees), PMR (Penalized Multinomial Regression), or a combination thereof. In another embodiment, the supervised machine learning method is any suitable method known in the art or any other suitable method to be developed in the future.
[0204] In some embodiments, the classification model is obtained by a supervised machine learning method. In particular, the supervised machine learning method is selected from the group consisting of LDA, CART, kNN, NB, SVM using a linear kernel, RF, NNET, GBM, GLM, GLMNET, AdaBoost, and PMR. In certain embodiments, the supervised machine learning method is selected from the group consisting of LDA, CART, kNN, NB, SVM, RF, NNET, GLMNET, AdaBoost, and PMR. In particular, this method is RF, GLMNET, or AdaBoost, and more specifically RF.
[0205] In another embodiment, the classification model is obtained by an unsupervised machine learning method. In certain embodiments, the unsupervised machine learning method is selected from the group consisting of k-means, K-Medoids method, Fuzzy C-means, agglomerative hierarchical clustering, Gaussian mixture model (GMM), neural network, hidden Markov model (HMM), mean shift, DBSCAN clustering, Apriori algorithm, principal component analysis (PCA), independent component analysis (ICA), linear discriminant analysis (LDA), singular value decomposition (SVD), linear semantic analysis (LSA), t-distributed stochastic neighbor embedding (t-SNE), non-linear multidimensional scaling, Principal Curves, k-nearest neighbor method (kNN), locally linear embedding, and autoencoder.
[0206] Alternatively, the classification model is obtained by using a deep learning algorithm as a set of machine learning methods. In certain embodiments, the deep learning method is selected from the group consisting of a convolutional neural network (CNN), a long short-term memory network (LSTM), a recurrent neural network (RNN), a generative adversarial network (GAN), a radial basis function network (RBFN), a multi-layer perceptron (MLP), a self-organizing map (SOM), deep belief networks (DBN), a restricted Boltzmann machine (RBM), and an autoencoder.
[0207] In some embodiments, the classification model is trained or has been trained with a training set that includes the mitochondrial methylation pattern for methylation sites as defined herein in a plurality of samples related to a plurality of subjects. In certain embodiments, the training set further includes clinical variables related to the plurality of subjects. More specifically, each subject in the training set is assigned a dementia stage classification. In certain embodiments, the dementia stage classification is selected from the group consisting of controls and subjects with DLB.
[0208] In another embodiment, the classification model is trained or has been trained with a training set that includes clinical variables related to a plurality of subjects. In certain embodiments, the training set further includes the mitochondrial methylation pattern for each methylation site in a plurality of samples related to the plurality of subjects. In particular, each subject in the training set is assigned a dementia stage classification. More specifically, the dementia stage classification is selected from the group consisting of controls and subjects with DLB.
[0209] In some embodiments, subjects classified as controls are characterized by having a CDR score of 0 and a clinical follow-up observation longer than 10 years (i.e., a clinical follow-up observation of more than 10 years). In some embodiments, subjects classified as having DLB are clinically diagnosed with DLB.
[0210] In some embodiments, the training dataset includes the correct output corresponding to the dementia stage class assigned to each subject, where the dementia stage class is a subject with a control and DLB.
[0211] In some embodiments, the data is preprocessed or has been preprocessed before training the classification model. This step is performed to ensure and update the performance of the model's training process. The preprocessing of the data includes (1) creating dummy variables, (2) removing zero and near-zero variance variables, (3) splitting the data into a training dataset and a test dataset, (4) centering and scaling, (5) dimensionality reduction (i.e., identifying and removing correlated variables), and (6) testing and visualizing the training dataset and includes.
[0212] (1) The process of creating dummy variables is performed to handle categorical data. Basically, each categorical variable is converted into a numerical variable by creating dummy variables using the "One-hot encoding" technique (i.e., each new variable is forced to have a value of either 0 or 1, representing the presence or absence of that attribute). This process is performed to ensure that the variables are encoded in a consistent manner. That is, this is encoded to ensure that there is no linear dependence between the new attributes, thus avoiding the dummy variable trap. Further, the data is reviewed to ensure that all one-coded categorical variables do not exhibit any abnormal linear combinations, and if they do, redundant variables are removed until the linear combination is excluded.
[0213] (2) The process of removing zero and near-zero variance variables is done to remove variables that exhibit a single unique value and variables that have a few unique numerical values that are highly imbalanced. Otherwise, these predictors can cause instability issues during the fitting process and model crashes.
[0214] (3) The process of data splitting is done so that it is randomly split into two main subsets: a subset for training the model (80% of the samples) and a subset for testing the classification model (20% of the samples). The random sampling process is driven within each class to preserve the overall class distribution of the data. A random number seed is contemplated to ensure reproducibility.
[0215] (4) The process of centering and scaling is applied to the continuous features (variables) of the training dataset for the purpose of estimating the centering factor and scaling factor that must be applied to both data to create a normalized dataset for performing the training process and testing process of the classification model.
[0216] (5) The dimensionality reduction process is done to reduce the number of variables (i.e., features or attributes) of the dataset while retaining as much relevant information as possible. In other words, this objective is to remove redundant or irrelevant features to improve the efficiency and effectiveness of the learning method applied to build the classification model. In this context, there are two main approaches: feature selection and feature extraction. (i) Feature selection techniques: The basic idea of these methods is to select a subset of variables based on certain criteria, for example, by identifying and removing correlated variables. This process is carried out to reduce highly correlated variables. To perform this step, a correlation matrix is calculated. Usually, the correlation measure applied is the Pearson correlation coefficient. Next, based on the absolute value of the pairwise correlation, when two variables have a high correlation, to detect variables that are highly correlated individually, the average absolute correlation of each variable is examined, and the variable with the maximum average absolute correlation is removed. In this context, usually, the cutoff of the pairwise absolute correlation can be set, for example, by testing the linear regression between each pair of variables. However, this approach was unable to reveal further correlations of features. For this reason, on the one hand, other monotonic methods can be attempted to identify correlations between variables. For example, Spearman's rank correlation and Kendall's rank correlation of non-parametric methods measure the relationship between each pair of variables based on the rank of the observations rather than their actual values. These are usually applied when a non-linear relationship is reasonably suspected between attributes (variables), skewness, or ordinal scales. Another type of method can be distance correlation (aka dCor). This is done to measure the dependence between each pair of variables by a method that is sensitive to non-linear relationships. Most of these methods are considered for testing the relationship between continuous variables and / or ranked variables. Similarly, there are also methods for testing the strength of the relationship between dichotomous variables and continuous variables (e.g., point-biserial correlation) or for testing the strength of the relationship when both variables are dichotomous (phi coefficient). On the other hand, it is recommended to use objective techniques such as cross-validation to determine a threshold suitable for detecting highly correlated variables. That is, to determine the necessary cutoff, first, the data must be split into a training dataset and a test dataset using a random number seed to ensure reproducibility. Second, the model is trained using the training dataset, and its performance is evaluated using the test set. Third, the threshold is varied to consider variables as highly correlated, and then the second step is repeated for each value of the threshold.Fourthly, select a threshold value that yields the best performance in the test set, for example, a threshold value with the highest correct answer rate (or the lowest error rate). Fifthly, after selecting the cut-off, this must be applied to the entire dataset and the model must be retrained using all the data.
[0217] Other applicable feature selection methods include genetic algorithms (GA), L1-regularized logistic regression, lasso regression, or hybrid methods.
[0218] (ii) Feature extraction techniques: These techniques create new variables by combining or applying transformations on the original features into a space with reduced dimensions. Some examples are principal component analysis (PCA), multi-factor analysis (MFA), t-SNE or UMAP, and / or multi-dimensional scaling (MDS), partial least squares discriminant analysis (PLS-DA), or autoencoders.
[0219] (6) Testing and visualizing the training dataset is a process that is carried out after performing the previous preprocessing tasks. This is done on the training dataset and is the second exploratory data analysis (EDA) conducted to re-investigate and confirm whether there is no bias in the original data. Classical statistical description methods are applied (i.e., univariate description methods, bivariate description methods, and multivariate description methods).
[0220] In one embodiment, data preprocessing includes (1) Data splitting, (2) Centering and scaling (3) Identification and removal of correlated variables, and (4) Testing and visualizing the training dataset and includes.
[0221] (1) In data splitting, the original corrected data is randomly split into a main subset: a subset for training the model (e.g., 80% of the samples) and a subset for testing the classification model (e.g., 20% of the samples). The random sampling process is driven within each class to preserve the overall class distribution of the data.
[0222] (2) In centering and scaling, centering factors and scaling factors are estimated that are applied to both data sets to create a normalized data set for performing the training process and the testing process of the classification model using continuous variables from the training data set.
[0223] (4) Identification and removal of correlated variables are performed to reduce the level of correlation between variables. For example, pairwise correlation analysis is performed based on the Spearman rank correlation coefficient. For pairs showing a high level of absolute correlation value (e.g., > 0.65), the variable with the maximum mean absolute correlation is removed from the data set. Furthermore, principal component analysis (PCA) can be applied to reduce the dimensionality of the features to, for example, up to 10 variables.
[0224] (4) Testing and visualizing the training data set includes applying exploratory data analysis to the training data set to re-investigate and confirm that there is no bias from the original data.
[0225] The classification model disclosed in this specification can be trained with data for a set of samples for which methylation data corresponding to a set of methylation sites has been obtained. For example, the training set includes data on methylation patterns consisting of the methylation sites presented in Tables 1 to 6, or any combination thereof. In particular, the training data set includes data on methylation patterns from the methylation sites presented in Tables 1 to 6.In some embodiments, the data of the methylation pattern includes data of methylation sites of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 84, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, or 250. In some embodiments, the data of the methylation pattern includes data of more than 50 methylation sites.In some embodiments, the data on methylation patterns includes data on more than 100 methylation sites. In some embodiments, the data on methylation patterns includes data on more than 200 methylation sites. In some embodiments, the data on methylation patterns includes about 10 to about 20, about 20 to about 30, about 30 to about 40, about 40 to about 50, about 50 to about 60, about 60 to about 70, about 70 to about 80, about 80 to about 90, about 90 to about 100, about 100 to about 110, about 110 to about 120, about 120 to about 130, about 130 to about 140, about 140 to about 150, about 150 to about 160, about 160 to about 170, about 170 to about 180, about 180 to about 190, about 190 to about 200, about 200 to about 210, about 210 to about 220, about 220 to about 230, about 230 to about 240, about 240 to about 250 methylation sites selected from Tables 1 - 6.
[0226] In some embodiments, the training dataset includes clinical variables for each subject, such as subject classification according to the classification models disclosed herein. In other embodiments, the training data includes data about the subject, such as weight, presence or absence of biomarkers, drug therapy, etc.
[0227] In some embodiments, the training set includes a reference population of at least about 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900 subjects, or at least about 1000 subjects. In other embodiments, the training set includes more than 1000 subjects.
[0228] In some embodiments, the classification model includes determining (relative) weights for each methylation pattern and each clinical variable considered. In other embodiments, the classification model includes determining (relative) weights for each methylation pattern and / or each clinical variable. In other embodiments, the classification model includes determining weights for each methylation pattern, each clinical variable, and / or any additional variable.
[0229] In some embodiments, the classification model uses data indicating the (relative) weights of each methylation pattern and each clinical variable for the identification of the subject's DLB. In other embodiments, the classification model uses data indicating the (relative) weights of each methylation pattern and / or each clinical variable for the identification of the subject's DLB.
[0230] In some embodiments, determining the score includes correlating each of at least one determined methylation pattern and each of at least one clinical variable with its determined weight. In some embodiments, determining the score includes correlating each of at least one determined methylation pattern and / or each of at least one clinical variable with its determined weight.
[0231] From the perspective of variables regarding methylation patterns of mitochondria in different sites, the variables that most contribute to the identification of the subject's DLB may be different. This is because the contribution of each variable depends on several aspects, such as the number of variables contemplated in the model (including clinical variables), the number of subjects used in the training of the classification model (to avoid overfitting / underfitting problems), the parameter customization process during the model's training process, and so on.
[0232] In this regard, the variables included in each classification model depend on the information available from each subject, and as a result, the importance / contribution of each selected variable varies among classification models. Therefore, it is difficult to define an exact set or a specific number of variables contemplated when constructing a classification model. In contrast, it is desirable to develop a classification model that has the ability to adapt to the available information and is constructed using variables selected according to their importance / contribution in each specific situation. In other words, the variables included in each classification model are defined by their contributions during the training process rather than any predefined set or minimum number of variables that may not contribute much in other scenarios. In any case, the main objective is to achieve the highest possible accuracy rate.
[0233] The classification models described in this specification may include different sets and combinations of methylation patterns and / or clinical variables. The classification model selects methylation patterns and / or clinical variables according to the contributions or importance (i.e., determined weights) associated with each methylation pattern and / or clinical variable. That is, the classification model combines methylation patterns of sites correlated with determined weights (such as 0.25, 0.5, 1, 2, 2.5, 5, etc.), and / or clinical variables correlated with at least one determined weight (such as 0.25, 0.5, 1, 2, 2.5, 5, etc.).
[0234] In some embodiments, the determined weights correlated with a methylation pattern or clinical variable, or a combination of measured variables (such as principal components) are between 0 and 100. Alternatively, the determined weights correlated with a methylation pattern or clinical variable are selected from 0 to 100. The value "0" corresponds to the minimum variable importance or minimum determined weight, and the value "100" corresponds to the maximum variable importance or maximum determined weight. In certain embodiments, the determined weights are selected from the group consisting of 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, and 100.
[0235] In some embodiments, the determined weight is at least 0.25. In particular, the determined weight is selected from the group consisting of 0.25, 0.3, 0.35, 0.40, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, 0.95, and 1. In another embodiment, the determined weight is at least 1. In particular, the determined weight is selected from the group consisting of 1, 1.25, 1.5, 1.75, 2, 2.25, 2.5, 2.75, 3, 3.25, 3.5, 3.75, 4, 4.25, 4.5, 4.75, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, and 100. In another embodiment, the determined weight is at least 2. In particular, the determined weight is selected from the group consisting of 2, 2.25, 2.5, 2.75, 3, 3.25, 3.5, 3.75, 4, 4.25, 4.5, 4.75, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, and 100. In another embodiment, the determined weight is selected from the group consisting of at least 1, at least 2, at least 2.5, at least 3, at least 3.5, at least 4, at least 4.5, at least 5, at least 7.5, at least 10, at least 15, at least 20, at least 25, at least 50, and at least 75.
[0236] In some embodiments, determining the subject's DLB includes considering the methylation pattern of a site correlated with a determined weight of at least 0.5. In another embodiment, determining the subject's DLB includes combining the methylation patterns of sites correlated with a determined weight of at least 0.5.
[0237] In another embodiment, determining the subject's DLB includes considering clinical variables correlated with a determined weight of at least 0.5. In another embodiment, determining the subject's DLB includes combining clinical variables correlated with a determined weight of at least 0.5.
[0238] In another embodiment, determining the subject's DLB comprises combining the methylation pattern of a site correlated with a determined weight of at least 0.5 with a clinical variable correlated with a determined weight of at least 0.5. Alternatively, determining the subject's DLB comprises combining or considering a determined weight of at least 0.5 with a correlated methylation pattern and / or clinical variable.
[0239] In some embodiments, determining the subject's DLB comprises considering the methylation pattern of a site correlated with a determined weight of at least 1. In another embodiment, determining the subject's DLB comprises combining the methylation pattern of a site correlated with a determined weight of at least 1.
[0240] In another embodiment, determining the subject's DLB comprises considering a clinical variable correlated with a determined weight of at least 1. In another embodiment, determining the subject's DLB comprises combining a clinical variable correlated with a determined weight of at least 1.
[0241] In another embodiment, determining the subject's DLB comprises combining the methylation pattern of a site correlated with a determined weight of at least 1 with a clinical variable correlated with a determined weight of at least 1. Alternatively, determining the subject's DLB comprises combining or considering a determined weight of at least 1 with a correlated methylation pattern and / or clinical variable.
[0242] In some embodiments, determining the subject's DLB comprises considering the methylation pattern of a site correlated with a determined weight of at least 2. In another embodiment, determining the subject's DLB comprises combining the methylation pattern of a site correlated with a determined weight of at least 2.
[0243] In another embodiment, determining the subject's DLB includes considering clinical variables that correlate with at least two determined weights. In another embodiment, determining the subject's DLB includes combining clinical variables that correlate with at least two determined weights.
[0244] In another embodiment, determining the subject's DLB includes combining the methylation pattern of a site that correlates with at least two determined weights with clinical variables that correlate with at least two determined weights. Alternatively, determining the subject's DLB includes combining or considering methylation patterns and / or clinical variables that correlate with at least two determined weights.
[0245] In some embodiments, determining the subject's DLB includes considering the methylation pattern of a site that correlates with at least 2.5 determined weights. In another embodiment, determining the subject's DLB includes combining the methylation pattern of a site that correlates with at least 2.5 determined weights.
[0246] In another embodiment, determining the subject's DLB includes considering clinical variables that correlate with at least 2.5 determined weights. In another embodiment, determining the subject's DLB includes combining clinical variables that correlate with at least 2.5 determined weights.
[0247] In another embodiment, determining the subject's DLB includes combining the methylation pattern of a site that correlates with at least 2.5 determined weights with clinical variables that correlate with at least 2.5 determined weights. Alternatively, determining the subject's DLB includes combining or considering methylation patterns and / or clinical variables that correlate with at least 2.5 determined weights.
[0248] In some embodiments, determining the subject's DLB involves considering methylation patterns of sites correlated with at least five determined weights. In another embodiment, determining the subject's DLB involves combining methylation patterns of sites correlated with at least five determined weights.
[0249] In another embodiment, determining the subject's DLB involves considering clinical variables correlated with at least five determined weights. In another embodiment, determining the subject's DLB involves combining clinical variables correlated with at least five determined weights.
[0250] In another embodiment, determining the subject's DLB involves combining methylation patterns of sites correlated with at least five determined weights with clinical variables correlated with at least five determined weights. Alternatively, determining the subject's DLB involves combining or considering methylation patterns and / or clinical variables correlated with at least five determined weights.
[0251] Alternatively, the methylation patterns and / or clinical variables that are combined or considered by a classification model to determine a subject's DLB are selected according to their determined weights. In particular, the methylation patterns and / or clinical variables combined by a classification model are selected according to their determined weights, and the determined weights are at least 1. In another specific embodiment, the methylation patterns and / or clinical variables combined by a classification model are selected according to their determined weights, and the determined weights are at least 2. In some embodiments, the methylation patterns and / or clinical variables combined by a classification model to determine a subject's DLB are correlated with determined weights selected from the group consisting of at least 0.25, at least 0.5, at least 1, at least 2, at least 2.5, at least 3, at least 3.5, at least 4, at least 4.5, at least 5, at least 7.5, at least 10, at least 15, at least 20, at least 25, at least 50, and at least 75.
[0252] In some embodiments, the classification model is capable of classifying a subject into two categories: CTL or DLB. In some embodiments, the classification model calculates a risk score for each category corresponding to the probability of a subject being assigned to each category. In some embodiments, the probabilities for each category sum to 1.
[0253] In some embodiments, the classification model is capable of classifying a subject into two categories: subjects having DLB and control subjects. In some embodiments, the classification model calculates a score for each category corresponding to the probability of a subject being assigned to each category. In some embodiments, the probabilities for each category sum to 1.
[0254] In the above, the weights correlated with the methylation pattern or clinical variables are in the range of 0 to 100, or are selected from 0 to 100, that is, a scale of 0 to 100 is used. For this scale, the minimum value of a specific weight is defined herein. However, it is obvious that any other appropriate scale can be used, such as a scale of 0 to 1, or a scale of 0 to 1000, or a scale of 0 to 20. With such different scales, the minimum value of the weights described above in this specification can be proportionally changed.
[0255] The classification models created by the machine learning methods disclosed herein or any other supervised methods known in the art (such as LDA, CART, kNN, NB, SVM using a linear kernel, RF, NNET, GBM, and GLM) can then be evaluated by determining the ability of the model to accurately classify each test subject (i.e., the classifier). In some embodiments, the subjects of the training population used to derive the model are different from the subjects of the test population used to test the model. As will be understood by those skilled in the art, this thereby predicts the ability of the dataset used to train the classifier with respect to their ability to appropriately classify subjects for which the output classification (such as dementia stage classification, i.e., subjects with DLB and control subjects) is unknown.
[0256] In some embodiments, the classification model is evaluated with respect to its property of appropriately classifying each subject in the training population using methods known to those skilled in the art. For example, cross-validation (CV), leave-one-out cross-validation (LOOCV), k-fold cross-validation, or Jackknife analysis using standard statistical methods are used to evaluate the classification model. In other embodiments, each classifier is evaluated with respect to its ability to appropriately characterize subjects in the training population that were not used to create the classifier.
[0257] In some embodiments, the metrics used to evaluate a classification model with respect to its ability to appropriately classify each subject in a training population are selected from the group consisting of classification accuracy (ACC), kappa coefficient, error rate (i.e., misclassification rate), sensitivity (true positive rate (Fraction or rate), TPF or TPR), specificity (true negative rate (Fraction or rate), TNF or TNR), positive predictive value (PPV), negative predictive value (NPV), or any combination thereof. In some embodiments, the metrics used to evaluate a classification model with respect to its ability to appropriately classify each subject in a training population further include recall, precision, F-value (F-measure or F1 score), confusion matrix (or contingency table), and / or area under the ROC curve (AUC ROC).
[0258] In certain embodiments, the metrics used to evaluate a classification model with respect to its ability to appropriately classify each subject in a training population include the sensitivity of the model (TPF, true positive rate) and 1 - specificity (FPF, false positive rate). In one embodiment, the method used to test a classifier is the receiver operating characteristic (“ROC”), which provides several parameters for evaluating both the sensitivity and specificity of the results of the classification model created.
[0259] Another aspect of the present invention is a computer - implemented method for the identification of DLB in a subject, comprising: (a) providing or receiving data related to the methylation pattern of the D - loop region of the subject's mitochondrial DNA and / or the ND1 gene, and optionally at least one clinical variable of the subject, wherein the methylation pattern is (i) the CpG sites of the D - loop region shown in Table 1, (ii) the CHG sites of the D - loop region shown in Table 3, (iii) the CHH sites of the D - loop region shown in Table 5, (iv) the CpG sites of the ND1 gene shown in Table 2, (v) the CHG sites of the ND1 gene shown in Table 4, and (vi) the CHH sites of the ND1 gene shown in Table 6 at least one site selected from the group consisting of, determined steps, (b) determining a risk score correlated with the identification of DLB in the subject, wherein the risk score is calculated or determined using a classification model configured to combine the methylation patterns of one or more sites of step (a) and optionally at least one clinical variable of the subject, steps A computer-implemented method comprising.
[0260] Another aspect of the present invention is a computer-implemented method for obtaining a score correlated with the identification of DLB in a subject, the following steps: (a) As input data: (1) The methylation pattern of the D-loop region of the mitochondrial DNA of the subject and / or the ND1 gene, (i) CpG sites in the D-loop region shown in Table 1, (ii) CHG sites in the D-loop region shown in Table 3, (iii) CHH sites in the D-loop region shown in Table 5, (iv) CpG sites in the ND1 gene shown in Table 2, (v) CHG sites in the ND1 gene shown in Table 4, and (vi) CHH sites in the ND1 gene shown in Table 6, determined at at least one site selected from the group consisting of, providing or receiving a methylation pattern, and optionally, (b) weighting the methylation pattern using a classification model to obtain a score A method comprising.
[0261] Another aspect of the present invention is a computer-implemented method for obtaining a score correlated with the identification of DLB in a subject, the following steps: (a) As input data: (1) at least one clinical variable of the subject providing or receiving steps, (b) obtaining a score by weighting the clinical variables using a classification model relates to a method comprising the same.
[0262] In some examples, input data can be received from which a computer or other data processing system can derive the methylation pattern of the D-loop region of the mitochondrial DNA of the subject and / or the methylation pattern of the ND1 gene. The computer-implemented method may then further comprise receiving at least one clinical variable of the subject as described herein and weighting the methylation pattern and the clinical variable in combination using a classification model to obtain a risk score.
[0263] Score In some embodiments, the score is a score from 0 to 1, where 0 represents the lowest probability of the subject having DLB and 1 represents the highest probability of the subject having DLB. In certain embodiments, a score of 0.5 or greater indicates that the subject may have DLB. Alternatively, a score of 0.5 or greater indicates that the subject has a high probability of having DLB. In certain embodiments, a score of 0.75 or greater indicates that the subject has a very high probability of having DLB. In some embodiments, a score of less than 0.5 indicates that the subject has a low probability of having DLB. In certain embodiments, a risk score of 0.25 or less indicates that the subject has a very low probability of having DLB.
[0264] In some embodiments, the score is a score from 0 to 100, where 0 is the lowest probability of a subject having DLB and 100 represents the highest probability of a subject having DLB. In some embodiments, the score can be, but is not limited to, 1 to 2, 1 to 5, 1 to 10, 1 to 100, 0 to 10, and 0 to 100. It should be understood that the scores of the present invention can represent any range of values that function to correlate with the probability of a subject having DLB.
[0265] Companion diagnostic system The methods disclosed herein can be provided as companion diagnostics, for example, made available via a web server, for informing clinicians or patients about potential treatment options or for patient selection for clinical trials. The methods disclosed herein can include the step of collecting or obtaining a biological sample and the step of performing the analytical methods disclosed herein to diagnose DLB in a subject.
[0266] In one aspect of the invention, a computing system is provided that includes suitable means for performing any of the computer-implemented methods described herein.
[0267] At least some embodiments of the methods described herein can be implemented by the use of a computer due to the complexity of the calculations involved. In some embodiments, the computer system includes hardware components that are electrically connected via a bus, including a processor, an input device, an output device, a storage device, a computer-readable storage medium reader, a communication system, processing acceleration (e.g., a DSP or a dedicated processor), and memory. The computer-readable storage medium reader can be further connected to a computer-readable medium, and this combination comprehensively represents remote storage devices, local storage devices, fixed storage devices, and / or removable storage devices, plus storage media, memory, etc., for temporarily and / or more persistently containing computer-readable information, which can include storage devices, memory, and any other such accessible system resources.
[0268] One or more servers may be implemented using a single architecture and may be further configured by currently desirable protocols, protocol variations, extensions, etc. However, it will be apparent to those skilled in the art that embodiments can be advantageously utilized according to more specific application requirements. Customized hardware may be used and / or certain elements may be implemented in hardware, software, firmware, or combinations thereof. Further, connections to other computing devices (not shown), such as network input / output devices, can be used, but it should be understood that wireless, wired, modem, and / or other connections to other computing devices may also be utilized.
[0269] In one embodiment, the system further includes one or more devices for providing input data to one or more processors. The system further includes a memory for storing a dataset of ranked data elements. In another embodiment, the device for providing input data includes a detector for detecting characteristics of the data elements, such as a fluorescence plate reader, a mass spectrometer, or a gene chip reader.
[0270] The system may further include a database management system. A user request or query can be formatted in an appropriate language understood by a database management system that processes the query to extract relevant information from a database of a training set. The system may be connectable to a network to which a network server and one or more clients are connected. This network can be a local area network (LAN) or a wide area network (WAN) known in the art. In particular, the server includes the hardware necessary for the execution of a computer program product (e.g., software) to access data in a database for processing user requests. This system can communicate with an input device for providing data (e.g., methylation patterns) regarding data elements to the system.
[0271] In a further aspect, the present invention is directed to a computer program product that includes instructions for causing a computer to perform any of the computer-implemented methods described herein when the program is executed by the computer.
[0272] Some embodiments described herein may be implemented to include a computer program product. The computer program product may include a computer-readable medium having computer-readable program code embodied therein for causing an application program to be executed on a computer having a database. As used herein, "computer program product" refers to an ordered set of instructions embodied in the form of statements of a natural or programming language that are included in a physical medium of any nature (e.g., a writing medium, an electronic medium, a magnetic medium, an optical medium, or other medium) and can be used in a computer or other automated data processing system. When such statements of a programming language are executed by a computer or data processing system, they cause the computer or data processing system to operate in accordance with the specific content of the statements.
[0273] The computer program product may include, but is not limited to, source code and object code embedded in a computer-readable medium, and / or programs of a test library or a data library. Further, a computer program product that enables a computer system or a data processing device to operate in a preselected manner may be provided in many forms including, but not limited to, original source code, assembly code, object code, machine language, the aforementioned encrypted or compressed versions, and all equivalents. In one aspect, the computer program product is provided to implement the procedures, diagnostics, and methods disclosed herein, for example, to determine whether to perform a particular treatment based on a score obtained.
[0274] A computer program product includes a computer-readable medium embodying program code executable by a processor of a computer device or system, the program code including: (a) obtaining data attributable to a biological sample from a subject, the data including a methylation pattern corresponding to at least one methylation site of the methylation sites of Tables 1-6 in the biological sample, or the methylation pattern being derivable from the data; combining these values with values corresponding to clinical variables; and (b) code for performing a classification method that indicates whether to administer a therapeutic agent to a patient in need thereof, for example, based on a score obtained.
[0275] Although various embodiments have been described as methods or apparatuses, it should be understood that the embodiments can be implemented via code coupled to a computer, such as code present on or accessible by the computer. For example, software and databases can be utilized to implement many of the methods described above. Thus, in addition to embodiments achieved by hardware, it should also be noted that these embodiments can be achieved through the use of a product consisting of a computer-usable medium having computer-readable program code embodied therein that provides the usability of the functions disclosed herein.
[0276] Furthermore, some embodiments can be, but are not limited to, code stored in substantially any type of computer-readable memory including RAM, ROM, magnetic media, optical media, or magneto-optical media. Even more generally, some embodiments can be implemented in software, hardware, or any combination thereof, including, but not limited to, software executed on a general-purpose processor, microcode, PLA, or ASIC.
[0277] One embodiment can also be envisioned as being achieved as a computer signal embedded in a carrier wave and signals propagated through a transmission medium (e.g., electrical and optical). Thus, the various types of information described above can be formatted into a structure such as a data structure and transmitted as an electrical signal through a transmission medium or stored on a computer-readable medium.
[0278] Execution of the method In some embodiments, a sample can be requested, for example, by a healthcare provider (e.g., a physician) or a healthcare payer, and obtained and / or processed by the same or a different healthcare provider (e.g., a nurse, a hospital) or a clinical laboratory, and after processing, these results can be sent to the original healthcare provider or yet another different healthcare provider, a healthcare payer, or a patient. Similarly, determining the methylation patterns disclosed herein; obtaining clinical variables from a subject; combining methylation patterns with clinical information; identifying / diagnosing DLB in a subject; applying a classification model; determining a score; making a diagnosis / prognosis determination; making a treatment determination; determining incorporation into a clinical trial; or combinations thereof can be performed by one or more healthcare providers, healthcare payers, and / or clinical laboratories.
[0279] As used herein, the term "healthcare provider" refers to an individual or institution that interacts directly with a living subject, such as a human patient, and administers to the subject. Non-limiting examples of healthcare providers include physicians, nurses, technicians, therapists, pharmacists, counselors, alternative medical practitioners, healthcare facilities, examination rooms, hospitals, emergency rooms, clinics, urgent care centers, other medical clinics / facilities, and any other organization that provides general and / or specialized procedures, diagnoses, evaluations, maintenance, treatments, drug therapies, and / or advice related to all or any part of a patient's health condition, including, but not limited to, general medical specialties, surgical, and / or any other type of procedure, diagnosis, evaluation, maintenance, treatment, drug therapy, and / or advice. The term "healthcare provider" also refers herein to a pharmaceutical company or its supplier / intermediary (e.g., CRO) involved in the development of clinical trials.
[0280] As used herein, the term "clinical laboratory" refers to a facility for testing or processing materials derived from a subject. These tests can also include procedures for collecting or otherwise obtaining and preparing samples and determining, measuring, or characterizing the presence or absence of various substances (e.g., methylation patterns of mtDNA or biomarkers used as clinical variables herein) in a subject's body or in samples obtained from the subject's body. The tests can also include procedures such as medical imaging procedures (e.g., PET, MRI) for obtaining clinical variable data.
[0281] As used herein, the term "healthcare payor" encompasses individual groups, organizations, or groups related to providing, presenting, proposing, paying, or otherwise giving patients access to all or part of one or more healthcare benefits, benefit plans, medical insurance, and / or healthcare cost accounting programs.
[0282] A healthcare provider may, to another healthcare provider or a patient, for example, the following actions: obtaining samples / clinical variables, processing samples / clinical variables, submitting samples / clinical variables, receiving samples / clinical variables, sending samples / clinical variables, analyzing or measuring samples / clinical variables (e.g., to obtain methylation patterns), quantifying samples / clinical variables, providing results obtained after analysis / measurement / quantification of samples / clinical variables, receiving results obtained after analysis / measurement / quantification of samples / clinical variables, applying a classification model, obtaining results after analysis / measurement / quantification of one or more samples / clinical variables, providing results, diagnosing / identifying DLB in a subject, identifying a subject suitable for treatment of DLB, selecting a subject to receive treatment for DLB, selecting a treatment suitable for treatment of a subject with DLB, classifying a subject according to the presence or absence of DLB; treating DLB in a subject, including administering treatment for DLB if the subject is identified as having DLB, administering treatment, initiating administration of treatment, interrupting administration of treatment, continuing administration of treatment, temporarily interrupting administration of treatment, increasing the amount of therapeutic agent administered, decreasing the amount of therapeutic agent administered, continuing administration of a certain amount of therapeutic agent, increasing the frequency of administration of the therapeutic agent, decreasing the frequency of administration of the therapeutic agent, maintaining the same frequency of administration of the therapeutic agent, replacing a treatment or therapeutic agent with at least another treatment or therapeutic agent, combining a treatment or therapeutic agent with at least another treatment or additional therapeutic agent, and / or instructing to perform or conduct monitoring of the progression of DLB.
[0283] In some embodiments, the healthcare provider can approve or deny any of the above actions, such as collection of sample / clinical variables, processing of sample / clinical variables, submission of sample / clinical variables, receipt of sample / clinical variables, sending of sample / clinical variables, analysis or measurement of sample / clinical variables (e.g., to obtain methylation patterns), quantification of sample / clinical variables, application of classification models, provision of results obtained after analysis / measurement / quantification of sample / clinical variables, sending of results obtained after analysis / measurement / quantification of sample / clinical variables, scoring of results obtained after analysis / measurement / quantification of one or more sample / clinical variables, sending of results from one or more sample / clinical variables, administration of treatment or therapeutic agent, initiation of administration of treatment or therapeutic agent, interruption of administration of treatment or therapeutic agent, continuation of administration of treatment or therapeutic agent, temporary interruption of administration of treatment or therapeutic agent, increase in the amount of therapeutic agent administered, decrease in the amount of therapeutic agent administered, continuation of administration of a certain amount of therapeutic agent, increase in the frequency of administration of therapeutic agent, decrease in the frequency of administration of therapeutic agent, maintenance of the same frequency of administration for the therapeutic agent, replacement of the treatment or therapeutic agent with at least another treatment or therapeutic agent, or combination of the treatment or therapeutic agent with at least another treatment or additional therapeutic agent. Further, the healthcare provider can, for example, approve or deny the prescribing of treatment, approve or deny the scope of treatment, approve or deny reimbursement for the cost of treatment, and determine or deny the eligibility for treatment.
[0284] In some embodiments, a clinical laboratory can, for example, collect or obtain sample / clinical variables, process sample / clinical variables, submit sample / clinical variables, receive sample / clinical variables, send sample / clinical variables, analyze or measure sample / clinical variables (e.g., to obtain methylation patterns), quantify sample / clinical variables, apply a classification model, provide results obtained after analysis / measurement / quantification of sample / clinical variables, receive results obtained after analysis / measurement / quantification of sample / clinical variables, obtain results after analysis / measurement / quantification of one or more sample / clinical variables, provide these results, diagnose / identify the subject's DLB, identify subjects suitable for treatment of DLB, select subjects to receive treatment for DLB, classify subjects based on the presence or absence of DLB, and monitor the progression or other related activities of DLB.
[0285] In some embodiments, sample / clinical variables can be obtained by a medical professional treating or diagnosing a patient, by a healthcare provider, or by a clinical laboratory. The measurement of samples (e.g., using specific assays described herein) and the obtaining of clinical variables (e.g., by medical imaging techniques) can be performed by the same or a different healthcare provider or clinical laboratory than the one that obtained the sample / clinical variables. The classification model can be applied by a healthcare provider or a different healthcare provider, or by a clinical laboratory. The results obtained are ultimately sent to the original medical professional or healthcare provider who treated or diagnosed the patient. Thus, in some embodiments, a healthcare provider or clinical laboratory can advise the medical professional / provider regarding diagnosis / prognosis or whether the patient can benefit from treatment. In some embodiments, the healthcare provider is one of a pharmaceutical company involved in the development of a clinical trial or its supplier / intermediary (e.g., a CRO). All steps described herein can be performed by one of the pharmaceutical company and / or its supplier / intermediary, or can be performed in part by, for example, a clinical laboratory or a different healthcare provider.
[0286] Example Example 1: Detection of mtDNA Methylation in Blood Samples 1.1 Materials and Methods 1) Collection of Blood Samples Human blood samples were collected in EDTA tubes to prevent blood clotting. After obtaining the samples in the laboratory, the blood was either processed directly for DNA extraction or aliquoted and stored at -80 °C until processing.
[0287] Samples were extracted from a total of 36 subjects recruited from three different cohorts. Controls were recruited from two different cohorts: cohort A corresponded to the Australian Imaging, Biomarker & Lifestyle Flagship Study of Ageing (AIBL), and cohort B corresponded to the Fundacion CITA (Centro de Investigacion y Terapias Avanzadas) in San Sebastian, Spain. Additionally, cohort C corresponded to DLB patients recruited at the Bellvitge Hospital in Barcelona from 2015 to 2019. Overall, the subjects were classified into two groups: controls (50%, N = 18) and DLB subjects (50%, N = 18).
[0288] 2) Total DNA Extraction The whole blood sample was processed for total DNA extraction and co-purification of both genomic DNA and mitochondrial DNA. DNA was isolated from human whole blood samples using the Wizard® Genomic DNA Purification Kit (# A1620, Promega) according to the manufacturer's instructions. Alternatively, the samples were processed using the Maxwell® RSC Instrument, which provides a simple method for the efficient and automated purification of DNA from the samples. Capture, washing, and purification of the DNA samples were performed using paramagnetic beads. The Maxwell® RSC Blood DNA Kit (# AS1400, Promega) was used according to the following manufacturer's specifications. The quality and quantity of the purified DNA were determined using a NanoDrop™ One Spectrophotometer manufactured by Thermofisher Scientific.
[0289] 3) Bisulfite treatment Bisulfite conversion consists of the deamination of non-modified cytosine to uracil, leaving the modified base 5-mC, i.e., methylated cytosine intact. Samples of total DNA (300 ng) were treated with bisulfite reagent using the EZ DNA Methylation Kit (# D5001, Zymo Research) according to the manufacturer's protocol. To obtain better bisulfite conversion, the incubation condition of step 2 of the protocol, which consists of a 15-minute incubation at 37 °C, was replaced with a 30-minute incubation at 42 °C as shown in Addendum 1.A of the manufacturer's protocol. These final conditions are recommended to minimize incomplete C to T conversion. The treated DNA was finally resuspended in 30 μL of nuclease-free water.
[0290] 4) Preparation of amplicon libraries The workflow for constructing the amplicon library was based on Illumina's "16S Metagenomic Sequencing Library Preparation Protocol" and could be used to sequence regions of the 16S rRNA gene and targeted amplicon sequences for other purposes. Preparation of the amplicon library enabled the acquisition of the mtDNA amplicons of interest and their preparation for processing on Illumina's MiSeq system.
[0291] 4.1) First PCR: Amplicon PCR The mtDNA region of interest, corresponding to the sequences of SEQ ID NOs: 1-4 and further containing the overhanging Illumina adapters, was amplified by PCR using specific degenerate primers (see Section 1.2.1 of the Results). When designing primers for the region of interest, the overhanging adapter sequences had to be added to the locus-specific primers of the region to be targeted, as described in the Illumina protocol.
[0292] The Illumina® overhanging adapter sequences added to the locus-specific sequences are (SEQ ID NOs: 5, 6): Forward overhang: 5’TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG-[locus-specific sequence] Reverse overhang: 5’GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG-[locus-specific sequence] That is.
[0293] The amplification of bisulfite-converted DNA was performed using the FastStart™ High Fidelity PCR System (# 3553400001, Roche). The final PCR mixture (25 μL) contained 5 μL of bisulfite-treated DNA, 1× FastStart Buffer # 2; 0.05 U of FastStart HiFi Polymerase; 0.8 mM total dNTP (0.2 mM dNTP each); and 0.4 μM of forward primer and reverse primer each. The reaction for the ND1 amplicon also contained 5% DMSO. The final volume was adjusted with nuclease-free water.
[0294] Amplification was performed in a SimpliAmp™ Thermal Cycler manufactured by Applied Biosystems.
[0295] To evaluate the samples obtained from the first PCR, 3 μL of each PCR product was analyzed by electrophoresis on a 1.5% agarose gel stained with SybrSafe™ DNA Gel Stain (# S33102, Thermofisher Scientific) to verify the presence of bands from the amplicons.
[0296] 4.2) Cleanup of the first PCR The amplicons were purified using AMPure XP beads (#A63881, Beckman Coulter) and separated from free primer species and primer dimer species. The steps were performed according to Illumina's "16S Metagenomic Sequencing Library Preparation Protocol". A ratio of 0.8× AMpure Beads was used to purify the PCR amplicon products. Elution of the beads was performed with 14 μL of Buffer EB (# 19086, Qiagen), and 12 μL was recovered from the beads.
[0297] 4.3) Second PCR: Index PCR Index PCR was performed to ligate Illumina's UDI (unique dual indices) and sequencing adapters. The PCR index reaction was carried out in a 50 μL reaction containing 5 μL of the DNA purified by the clean-up of the first PCR, 25 μL of KAPA HiFi HotStart Ready Mix (2x) (# 7958935001, Roche), 10 μL of Illumina's Unic dual indexes, and 10 μL of nuclease-free water. This was performed using an Applied Biosystems(™) SimpliAmp(™) Thermal Cycler.
[0298] 4.4) Clean-up of the second PCR The clean-up of the second PCR was performed using AMPure XP beads to clean up the final library and then quantification was carried out. The 50 μL of the second PCR reaction was purified according to the steps described in Illumina's "16S Metagenomic Sequencing Library Preparation Protocol". A ratio of 1.12× of AMPure XP beads was used and the final elution was carried out with 27.5 μL of Buffer EB, and 25 μL was recovered from the beads.
[0299] 4.5) Quantification, normalization, and pooling of the library Quantification of the library was performed using a fluorescence quantification method using a dsDNS binding dye with a Thermo Fisher Scientific Qubit(®) 3.0 fluorometer. Quantification was carried out according to the manufacturer's instructions with the Qubit(™) dsDNA HS Assay Kit. After obtaining the Qubit quantification value in ng / μL, the DNA concentration was calculated using the following formula: (Concentration in ng / μL) / (660 g / mol × average library size) × 10 6 = Concentration in nM Based on the size of DNA amplicons determined by the Agilent Technologies 2100 Bioanalyzer trace using [specific method], it was calculated in nM.
[0300] For normalization, the final library was diluted to 10 nM using Buffer EB, and a final 4 nM pool of the amplicon library was prepared in a final volume of 20 μL with Buffer EB.
[0301] In parallel, a 4 nM dilution of the PhiX library (# FC-110-3001, Illumina) from a 10 nM PhiX library was prepared in a final volume of 5 μL with Buffer EB (each run had to contain at least 5% PhiX as an internal standard for these low-density libraries).
[0302] For cluster generation and sequencing preparation, the pooled amplicon library was denatured with NaOH and diluted with HT1 buffer as follows: 5 μL of 4 nM amplicon library and 5 μL of 0.2 N NaOH (freshly prepared) were introduced into a microcentrifuge tube, briefly mixed using a vortex mixer, and centrifuged at 280×g at 20 °C for 1 minute. Incubation was carried out at room temperature for 5 minutes to denature the DNA, and 990 μL of pre-cooled HT1 buffer was added to 10 μL of denatured DNA. Finally, the HT1 result was added to a 20 pM denatured amplicon library in 1 mM NaOH. The denatured DNA was placed on ice until proceeding to the final dilution.
[0303] The same steps were repeated for 5 μL of 4 nM PhiX library to denature and dilute PhiX, resulting in a 20 pM PhiX denatured library.
[0304] A 7 pM denatured amplicon library was prepared by mixing 210 μL of a 20 pM denatured amplicon library and 390 μL of pre-cooled HT1 buffer to a final volume of 600 μL. Subsequently, a 10 pM denatured PhiX library was prepared by mixing 300 μL of a 20 pM denatured PhiX library with 300 μl of pre-cooled HT1 buffer to a final volume of 600 μL.
[0305] Finally, a 25% denatured PhiX library and 75% denatured amplicon library were combined to a final volume of 600 μL (150 μL of the 7 pM denatured amplicon library was discarded and replaced with 150 μL of the 10 pM denatured PhiX library).
[0306] The combined amplicon library and PhiX control were placed on ice until ready for heat denaturation. Immediately after the heat denaturation step, the library was loaded into the MiSeq reagent cartridge to ensure efficient template loading onto the MiSeq flow cell.
[0307] Using a heat block, the combined library and PhiX control tubes were incubated at 96 °C for 2 minutes. After incubation, the tubes were inverted 1 - 2 times to mix and immediately placed on ice. The tubes were maintained on ice for 5 minutes.
[0308] 5) Template loading and run setup on the MiSeq instrument Sequencing on the MiSeq instrument using 300 bp paired reads was prepared with the MiSeq reagent Kit v3 (#MS - 102 - 3003, Illumina).
[0309] When the Illumina v3 reagent cartridge was completely thawed and ready for use, the prepared library was loaded into the cartridge and the run was set up on the MiSeq instrument according to the manufacturer's instructions.
[0310] 6) Calculation of the Methylation Percentage of Mitochondria The percentage (%) of methylation at each cytosine site is calculated by the β value (β). The β value is the ratio of the methylated reads per site to the total of methylated reads and unmethylated reads per site, that is β i = M / (M + U) (where M is the number of methylated reads at site (i) and U is the number of unmethylated reads at the same site) is.
[0311] The β value ranges from 0 to 1, where 0 is completely unmethylated and 1 is completely methylated.
[0312] The percentage (%) of methylation at each site is obtained by multiplying 100 by the β value, that is % i = β i * 100 is obtained by.
[0313] 7) Differential Methylation Analysis To identify differentially methylated sites (DMSs, also called DMLs in the case of differentially methylated loci), a standard pipeline consisting of the following five main steps was performed: 1. Quality control of raw data: Evaluation of the quality of reads generated by Illumina instruments during the sequencing process. 2. Preprocessing of raw data: Preparation of data for the alignment process by filtering, clipping, and trimming raw data according to different criteria. 3. Alignment process: Mapping of preprocessed reads to the reference genome. 4. Methylation calling (quantification): Identification of the positions of cytosines in each context (i.e., CpG, CHG, and CHH) and the number of methylated and unmethylated reads per site, region, and sample. 5. Methylation profile pile-up: Compilation and construction of a matrix of methylation, unmethylation, and percentage to be analyzed (rows are samples and columns are identified positions), and execution of quality control of the reported measurements.
[0314] Based on a matrix of mitochondrial methylation percentages (also called methylation measurements), two data analyses were performed: (1) exploratory data analysis and (2) identification of DMS.
[0315] 1). Exploratory data analysis (EDA) is intended to describe the distribution of measurements in both samples and positions with respect to groups, contexts, and regions of each genotype. For this purpose, two different plots are reported. The box plot of the data (also called box-and-whisker plot or box-and-whisker diagram) graphically shows the distribution of the percentage of methylation through the quartiles of the different samples at each position in the comparison of DLB vs CTL samples in each context. CRL samples are represented in black, and DLB samples are represented in gray. In the box plot, the central box is the interquartile range and represents the central 50% of the data. In each box, the median (the middle quartile) that shows the midpoint 50% of the data is represented by a line that divides the box into two parts. The lower whisker represents quartile 25, the upper whisker represents quartile 75, and the extreme whiskers represent the maximum and minimum values. Also, the overall median is represented by a horizontal gray line, which is the overall median of all the existing values in each group. Furthermore, the overall mean is also represented by a horizontal gray dashed line that shows the overall mean of the percentage of methylation in the CTL group and the DLB group. In the second type of figure, the plot of the mean and confidence interval (95%) represents the mean and its confidence interval (95%) of the percentage of methylation at all positions for each subject of CTL and DLB. The mean is represented by a black circle in CLL and a gray circle in DLB, and the confidence interval is represented by vertical bars. The overall median is also represented by a horizontal gray line that shows the overall median of all the existing values in each group of samples in the CTL group and the DLB group. The overall mean is also represented by a horizontal gray dashed line that shows the overall mean of the percentage of methylation in the CTL group and the DLB group.
[0316] 2) Identification of differential methylation sites (DMS): Analyses to compare methylation levels between groups (e.g., subjects with DLB and controls) at each methylation site were performed using the DSS (Dispersion Shrinkage for Sequencing data) Bioconductor package. This package is intended to identify differential methylation loci / sites (DML / DMS) on bisulfite sequencing (BS-seq) data. The core of DSS is a hierarchical Bayesian model-based procedure for estimating and shrinking the dispersion of cytosine site-specific contexts, and then a Wald test is performed with a beta-binomial distribution to detect differential methylation. Furthermore, for general experimental designs, DSS is based on a beta-binomial regression model that takes into account the arcsine link function, and model fitting is performed on the transformed data using generalized least squares.
[0317] Problems of multiple tests were addressed by adjusting the Benjamini-Hochberg false discovery rate (FDR), i.e., the p-values were adjusted using the FDR method. Other methods, such as using the family-wise error rate (FWER), can be used to adjust the p-values. Cytosine site-specific contexts with corrected p-values (i.e., FDR) less than 0.05 were identified as differentially methylated. This model is set with a single main effect. Furthermore, if there are technical biases due to batch effects or certain technical variabilities, these unwanted effects are considered in the model.
[0318] 1.2 Results 1.2.1 Optimization of mtDNA methylation detection The primers used in this specification (referred to in Section 4 - Preparation of amplicon libraries) were designed to optimize the detection of mtDNA methylation. To avoid bias resulting from the general assumption in the art that non - CpG cytosines are mainly not methylated, the inventors used primers containing a minimum number of cytosines. Further, the primers were degenerate such that they cover all possible methylation scenarios and not the unmethylated scenarios, thereby addressing the problem resulting from the uncertainty of C / U conversion that can affect several cytosine residues contained in the sequence. Said degeneracy includes the inclusion of a mixture of oligonucleotide sequences containing several possible nucleotide bases at each specific position, thus significantly increasing the probability of detecting mitochondrial methylation.
[0319] In summary, the degenerate forward primer contains Y (Y = C / T) which refers to either C or T at any position where the reference sequence is C. Further, as is known in the art, the reverse primer does not correspond to the reference sequence but to the reverse complement of the reference sequence. Thus, the reverse primer does not contain the C sites of the reference sequence but contains their complementary G sites. Consistent with this, the degenerate reverse primer contains R (R = A / G) which refers to either G or A at any position where the reference sequence is C or the complementary sequence contains G. As a result, the four degenerate primers are as follows: D - loop region: Forward primer: YAYTTGGGGGTAGYTAAAGTGAAYTG (SEQ ID NO: 1) Reverse primer: TCCTACAARCATTAATTAATTAACACAC (SEQ ID NO: 2) ND1 gene: Forward primer: ATAAAAYTTAAAAYTTTAYAGTYAGAG (SEQ ID NO: 3) Reverse primer: TTRARTTTRATRCTCACCCTRATCA (SEQ ID NO: 4)
[0320] 1.2.2 Methylation pattern of mitochondrial DNA The methylation levels were compared between DLB subjects (N = 18) and control subjects (N = 18) with respect to two different regions, namely the ND1 gene and the D-loop region, in three different contexts: CpG, CHG, and CHH.
[0321] For the ND1 gene, the results are presented in box plots for all three contexts CpG, CHG, and CHH corresponding to Figures 1, 3, and 5A - D herein. As shown in Figure 7, no significant differences in the methylation profile distribution were found when comparing control samples with samples from DLB subjects in any of the three contexts for the ND1 gene (adjusted p-value < 0.05). Thus, figures of the mean values and confidence intervals (95%) corresponding to the display of the average percentage of methylation for each patient, as presented in Figures 2, 4, and 6 respectively, show no significant differences between DLB subjects and CTL subjects in any of the three contexts CpG, CHG, and CHH (adjusted p-value < 0.05). However, it is noteworthy that when the threshold for statistical significance was set at an adjusted p-value < 0.1, significant differences were found at all CHH positions of the ND1 gene. Thus, the methylation pattern at the CHH sites in the ND1 region is actually different between subjects with DLB and controls.
[0322] On the other hand, when comparing methylation levels in three contexts in the D-loop region, a significant difference was observed. Surprisingly, the DLB subjects showed highly significant hypomethylation compared to the control subjects in all three contexts, namely CpG, CHG, and CHH, as shown in Figures 7, 9, and 11A - B, and Table 8 respectively (adjusted p-value < 0.05). Furthermore, the figures of the mean values and confidence intervals (95%) also showed a strong and statistically significant difference in the average percentage of methylation between the DLB subjects and the control subjects in all three contexts, namely CpG, CHG, CHH (Figures 8, 10, and 12 respectively). In particular, as shown in Table 8, the adjusted p-values of these differences were well below the threshold of statistical significance set at adjusted p-value < 0.05, and the highest adjusted p-value was only 1.04E - 11. In summary, these results clearly supported the existence of a specific mito - epigenetic signature in the D-loop of the mitochondrial control region that characterizes DLB subjects, and thus could be considered a promising biomarker for DLB.
[0323]
Table 7
[0324]
Table 8
[0325] Example 2: Development of a Classification Model The classification model of the present invention is constructed based on two main types of information sources: (1) clinical variables and (2) methylation measurements prepared in-house as described in Example 1.
[0326] This construction procedure is carried out using a supervised learning method. This process includes four main steps: (1) exploratory data analysis (EDA), (2) data preprocessing, (3) model training process, and (4) model performance evaluation, which are described below.
[0327] 1) Exploratory Data Analysis (EDA) EDA is intended to explain all variables being tested in the analysis. For categorical variables, a frequency distribution table and a bar chart are created, and for continuous variables, central measures, measures of dispersion, and measures of symmetry are estimated. Further, a histogram and a violin plot are provided for each continuous variable.
[0328] 2) Data Preprocessing This step is performed to ensure and enhance the performance of the model's training process. Data preprocessing consists of the following six steps: 1. The step of creating dummy variables using the one-hot encoding technique to ensure that there is no linear dependence between new attributes, thus avoiding the dummy variable trap. 2. The step of removing zero and near-zero variance variables to avoid instability problems during the fitting process and model crashes. 3. The step of splitting the data into two separate datasets: a training dataset and a test dataset. The original corrected data is randomly split into two main subsets: a subset for training the model (80% of the samples) and another subset for validating the classification model (20% of the samples). The random sampling process is driven within each class to preserve the overall class distribution of the data. The percentage depends on the results of the EDA. 4. The step of centering and scaling the two datasets. Using the continuous variables from the training dataset, centering and scaling coefficients are estimated, and then this is applied to both datasets to create a normalized dataset for performing the training process and test process of the classification model. 5. Identify and remove correlated variables together with confounding factors. Furthermore, pairwise correlation analysis based on Pearson's correlation coefficient was performed to benefit from reducing the level of correlation between variables. For pairs showing a high level of absolute correlation value, the variable with the maximum mean absolute correlation is removed from the dataset. On the other hand, a generalized linear model (GLM) is applied to each variable to identify and control potential confounding factors. Furthermore, an unsupervised multivariate approach is performed to examine the relationships between individuals, explain this by a set of quantitative and qualitative variables, and structure it in groups of clinical data and molecular data. As a result, identify the variables included in the model. 6. Re-examine and confirm that there are no potential biases from the original data by testing and visualizing the training dataset.
[0329] 3) Training of the model Build a classification model using nine supervised learning methods. These methods are linear discriminant analysis (LDA), CART (Classification and Regression Trees), k-nearest neighbor method (kNN), naive Bayes (NB), support vector machine with linear kernel (SVM), random forest (RF), neural network (NNET), generalized boosted regression model (GBM), and binary logistic regression (GLM). These methods are selected with respect to supervised classification methods and their ability to handle both continuous data and categorical data. Furthermore, examining the criteria for measuring the importance of these methods is useful for identifying key variables that can be modified and used in further analysis. Alternative algorithmic methods can be considered for the development of the model training process. All selected methods are executed using k-fold cross-validation, and the total number of combinations of parameters to be evaluated, the number of iterations of backpropagation, and other types of hyperparameters are adapted and modified during the adjustment process.
[0330] 4) Performance evaluation of the classification model To measure the performance of the training model, the following measurement criteria are estimated: Accuracy: The overall agreement rate averaged over the iterations of cross-validation Kappa: Cohen's unweighted kappa coefficient averaged over the results of resampling.
[0331] For each of these statistics, the mean, median, minimum, maximum, and first and third quartiles are estimated.
[0332] To measure the performance of the classification prediction of the trained model, a confusion matrix is constructed to show the cross-tabulation of the observed and predicted classes. This table comes with the accuracy and kappa coefficient, as well as common measurement criteria used to evaluate the model: sensitivity, specificity, positive predictive value, negative predictive value, precision, prevalence, F1-score, detection rate, and detection prevalence. Additionally, an ROC curve is created to measure the performance of the model that classifies DLB subjects vs CTL subjects.
[0333] Example 3: mtDNA Methylation Patterns in Additional Samples and Development of Classification Model 2 3.1 Methylation Patterns of Mitochondrial DNA A total of 84 samples were analyzed by the method described in Example 1.
[0334] The methylation levels in two different regions, namely the ND1 gene and the D-loop region, in three different contexts: CpG, CHG, and CHH, were compared between DLB subjects (N = 42) and control subjects (N = 42).
[0335] Figures 13 to 18 and Tables 9 to 10 show these results. For the ND1 gene, the results are presented herein as box plots for all three contexts CpG, CHG, and CHH. For the ND1 gene, significant differences in the methylation profile distribution were found when comparing control samples with samples from DLB subjects in all three contexts (adjusted p-value < 0.05). Thus, figures of the mean values and confidence intervals (95%) corresponding to the display of the average percentage of methylation for each patient show significant differences between DLB subjects and CTL subjects in all three contexts CpG, CHG, and CHH (adjusted p-value < 0.05).
[0336] On the other hand, when comparing methylation levels for the three contexts in the D-loop region, highly significant differences were observed. DLB subjects show highly significant differences in methylation compared to control subjects in all three contexts, namely CpG, CHG, and CHH (adjusted p-value < 0.05). Furthermore, figures of the mean values and confidence intervals (95%) also show a strong and statistically significant difference in the average percentage of methylation between DLB subjects and control subjects in all three contexts, namely CpG, CHG, and CHH. In particular, the adjusted p-values for such differences are well below the statistically significant threshold set at adjusted p-value < 0.05. In summary, these results clearly support the existence of specific mitoepigenetic signatures in the D-loop and ND1 gene of the mitochondrial control region that characterize DLB subjects, and thus can be considered promising biomarkers for DLB.
[0337]
Table 9
[0338]
Table 10
[0339] 3.2 Development of Classification Model A prototype model was developed to classify subjects diagnosed with DLB.
[0340] Materials and Methods 1) Target Data Raw data was verified to ensure the performance of the modeling process. For this reason, subjects with missing values were excluded from the analysis. In this way, a total of 84 subjects were ADmit Cohort (15 subjects): Hospital Universitari de Bellvitge, Barcelona, Spain; AIBL Cohort (39 subjects): Australian Imaging, Biomarker & Lifestyle Flagship Study of Ageing, Australia; CITA Cohort (3 subjects): Center for Research and Advanced Therapies, CITA - Alzheimer Foundation, Donostia - San Sebastian, Spain; Hospital Clinic Cohort (3 subjects): Hospital Clinic de Barcelona, Spain; and MAP - AD Cohort (24 subjects): Hospital Universitari de Bellvitge, Hospital Clinic de Barcelona, Hospital General de l’Hospitalet, and Hospital de Sant Joan Despi Moises Broggi in Barcelona, Spain recruited from.
[0341] In this regard, two groups of individuals were considered: - Control (CTL): 42 (50%) subjects with a Clinical Dementia Rating (CDR) scale score of 0 and a clinical follow - up of over 10 years - DLB: 42 subjects (50%) diagnosed with DLB.
[0342] Control subjects were only available from the AIBL cohort and CITA cohort, as well as patients from the ADmit cohort, Hospital Clinic cohort and MAP-AD cohort. Mainly two types of information sources: (1) clinical variables and (2) in-house generated methylation measurements were considered for the analysis. Clinical variables were used to describe the data considered for the analysis, i.e., only methylation measurements were involved in the construction of the classification model.
[0343] 2) Clinical variables In this first prototype, three clinical variables were tested: Stage (response variable): consisting of three levels of control and DLB as described above Gender: indicating two gender levels, female and male. Age: the age of the patient at the first diagnosis.
[0344] 3) Methylation measurements 224 variables, each collecting the percentage of methylation at a single specific cytosine site in one of three contexts (i.e., CpG, CHG, and CHH) for each gene (i.e., D-loop and ND1) as described in Example 1.
[0345] 4) Methods Exploratory data analysis Data preprocessing Model training process, and Performance evaluation A workflow was implemented that was divided into four main steps:
[0346] 5) Exploratory data analysis Exploratory data analysis (EDA) was intended to describe all single variables tested in the analysis. For categorical variables, a frequency distribution table and a bar chart were created. For continuous variables, measures of central tendency (mean, median, and mode if applicable), measures of dispersion (standard deviation, standard error, mean absolute deviation, and range), and measures of symmetry (skewness and kurtosis) were estimated. Additionally, for each continuous variable, a histogram and a violin plot were provided.
[0347] 6) Data preprocessing This step was performed to ensure and enhance the performance of the model's training process. This consisted of splitting the data into a training dataset and a test dataset, centering and scaling both datasets, identifying appropriate feature variables, and testing and visualizing the training dataset.
[0348] 1. Data splitting: The original corrected data was randomly split into two main subsets: a subset for training the model (80% of the samples) and another subset for testing the classification model (20% of the samples). The random sampling process was driven within each class to preserve the overall class distribution of the data.
[0349] 2. Centering and scaling: Using the continuous variables from the training dataset, centering coefficients and scaling coefficients were estimated and applied to both datasets to create a normalized dataset for performing the training process and the testing process of the classification model.
[0350] 3. Identification and removal of correlated variables: To benefit from reducing the level of correlation between variables, a pairwise correlation analysis based on Spearman's rank correlation coefficient was performed. For these pairs showing a high level of absolute correlation value (>0.65), the variable with the maximum mean absolute correlation was removed from the dataset. Additionally, principal component analysis (PCA) was applied to reduce the dimensionality of the features. In this case, the top 10 variables were considered.
[0351] 4. Testing and Visualization of the Training Dataset. After applying the previous preprocessing tasks, a second EDA was applied to the training dataset to re-investigate and confirm that there is no bias in the original data.
[0352] 7) Model Training Ten supervised learning methods were selected to build the prototype. These methods are GLMNET: Generalized Linear Model (Logistic Regression) via Penalized Maximum Likelihood Estimation LDA: Linear Discriminant Analysis PMR: Penalized Multinomial Regression AdaBoost: AdaBoost Classification Trees CART: Classification and Regression Trees kNN: k-Nearest Neighbors NB: Naive Bayes SVM: Support Vector Machine with a Linear Kernel RF: Random Forest NNET: Neural Network respectively.
[0353] These were selected for three main reasons: 1. These techniques are supervised classification methods. 2. They can handle continuous categorical data. 3. The measurement criteria of importance can assist in identifying important variables that can be modified and used in further analysis.
[0354] All methods were run using 10-fold cross-validation repeated 3 times. The number of parameter combinations evaluated for SVM, RF, and NNET was 5. NNET was evaluated using the backpropagation method with 1000 repetitions.
[0355] 8) Evaluation of the Performance of the Classification Model To measure the performance of the training model, the following measurement criteria were estimated: Accuracy: The overall agreement rate averaged over the iterations of cross-validation Kappa: Cohen's unweighted kappa coefficient averaged over the results of resampling
[0356] For each of these statistics, the mean, median, minimum, maximum, as well as the first and third quartiles were estimated. To measure the performance of the classification prediction, a confusion matrix was constructed to show the cross-tabulation of the observed and predicted classes. This table was accompanied by the accuracy and kappa coefficient, as well as the common measurement criteria used to evaluate the model: sensitivity, specificity, positive predictive value, negative predictive value, precision, prevalence, F1-score, detection rate, and detected prevalence rate. Also, an ROC curve was constructed to measure the performance of the model for classifying DLB patients versus control patients.
[0357]
Table 11
[0358] Results Exploratory data analysis Briefly stated, according to the results of the EDA, the gender showed 35 (41.67%) women and 49 (58.33%) men in an equal number of subjects. The percentage of sex was still balanced in each stage group, 17 (40.48%) women and 25 (59.52%) men in the control, and 18 (42.86%) women and 24 (57.14%) men in DLB. Furthermore, the recruited individuals were 72.01 ± 5.56 years old. The age of the subjects in each stage was 70.57 ± 5.4 years in the control and 73.45 ± 5.4 years in DLB.
[0359] Data preprocessing By splitting the data, a first subset of 69 subjects for the training model process and a second subset of 16 individuals for testing the classification model were created. More than 2900 pairs of variables with an absolute correlation exceeding 0.65 were identified. This fact led to removing more than 40% of the quantitative variables from the data, leaving 132 potential predictors. After applying the correlation method, PCA was applied to reduce the dimensionality of the variables.
[0360] Model training The results of the training process indicate that the best performance seems to be achieved by the RF model with an average accuracy of 0.83 and a kappa value of 0.65. These values suggest that this is a good model for classifying DLB patients. Tables 12 and 13 show the accuracy and kappa metrics of the implemented training models, and Figure 19 shows an overview of these tables.
[0361] [Table 12]
[0362] [Table 13]
[0363] Evaluation of model performance The classification accuracy for the test data of the RF model was 0.81, with a 95% confidence interval of 0.54 - 0.95, and a kappa value of 0.63. The sensitivity and specificity of the model for classifying the control as DLB were 0.75 and 0.875, respectively. The precision of the model was 0.86, and the F1 score was 0.8. Figure 20 shows the ROC curve indicating the performance of the DLB patient classification model at all classification thresholds.
[0364] In summary, the prototype of the classification model of the present invention seems to do a good job of identifying DLB patients. Nevertheless, this requires reexamination with a larger sample size to improve the adjustment process during model training. This fact may slightly modify the estimates provided in this report and its related metrics for evaluating the performance of classification.
[0365] References Non-Patent Literature De Boni, L., Tierling, S., Roeber, S., Walter, J., Giese, A., & Kretzschmar, H. A. (2011). Next-generation sequencing reveals regional differences of the α-synuclein methylation state independent of Lewy body disease. NeuroMolecular Medicine, 13(4), 310-320. Desplats, P., Spencer, B., Coffee, E., Patel, P., Michael, S., Patrick, C., Adame, A., Rockenstein, E., & Masliah, E. (2011). α-synuclein sequesters Dnmt1 from the nucleus: A novel mechanism for epigenetic alterations in Lewy body diseases. Journal of Biological Chemistry, 286(11), 9031-9037. Funahashi, Y., Yoshino, Y., Yamazaki, K., Mori, Y., Mori, T., Ozaki, Y., Sao, T., Ochi, S., Iga, J. I., & Ueno, S. I. (2017). DNA methylation changes at SNCA intron 1 in patients with dementia with Lewy bodies. Psychiatry and Clinical Neurosciences, 71(1), 28-35. Urbizu, A., & Beyer, K. (2020). Epigenetics in lewy body diseases: Impact on gene expression, utility as a biomarker, and possibilities for therapy. International Journal of Molecular Sciences, 21(13), 1-31. Chouliaras, L., Kumar, G. S., Thomas, A. J., Lunnon, K., Chinnery, P. F., & O’Brien, J. T. (2020). Epigenetic regulation in the pathophysiology of Lewy body dementia. Progress in Neurobiology, 192. Fernandez, A. F., Assenov, Y., Martin-Subero, J. I., Balint, B., Siebert, R., Taniguchi, H., Yamamoto, H., Hidalgo, M., Tan, A. C., Galm, O., Ferrer, I., Sanchez-Cespedes, M., Villanueva, A., Carmona, J., Sanchez-Mut, J. V., Berdasco, M., Moreno, V., Capella, G., Monk, D., Esteller, M. (2012). A DNA methylation fingerprint of 1628 human samples. Genome Research, 22(2), 407-419. Sanchez-Mut, J. V., Heyn, H., Vidal, E., Moran, S., Sayols, S., Delgado-Morales, R., Schultz, M. D., Ansoleaga, B., Garcia-Esparcia, P., Pons-Espinal, M., De Lagran, M. M., Dopazo, J., Rabano, A., Avila, J., Dierssen, M., Lott, I., Ferrer, I., Ecker, J. R., & Esteller, M. (2016). Human DNA methylomes of neurodegenerative diseases show common epigenomic patterns. Translational Psychiatry, 6(1), e718-8. Nasamran, C. A., Sachan, A. N. S., Mott, J., Kuras, Y. I., Scherzer, C. R., Ricciardelli, E., Jepsen, K., Edland, S. D., Fisch, K. M., & Desplats, P. (2020). Differential blood DNA methylation across lewy body dementias. Alzheimer’s and Dementia: Diagnosis, Assessment and Disease Monitoring, 13(1), 1-12. Blanch, M., Mosquera, JL., Ansoleaga, B., Ferrer, I., Barrachina, M. (2016). Altered Mitochondrial DNA methylation pattern in Alzheimer Disease-related pathology and in Parkinson disease, The American Journal of Pathology, 186(2):385-97. Stoccoro, A., Siciliano, G., Migliore, L., Coppede, F. (2017). Decreased methylation of the mitocondrial D-Loop region in late-onset Alzheimer’s disease, Journal of Alzheimer’s Disease, 59(2):559-564. Stoccoro, A., Baldacci, F., Coppede, F., Migliore, L. (2022) Mitochondrial DNA methylation levels are altered in individuals with mild cognitive impairment, Abstract OO016 / #1584 On-demand symposium: AD diagnosis & clinical trials & advances in drug development 1. Patent Document International Patent Publication No. 2015 / 144964
Claims
**Claim 1** A method for identifying a target Lewy body dementia, comprising: a) determining a methylation pattern of a D-loop region and / or an ND1 gene of mitochondrial DNA in a sample of the subject containing mitochondrial DNA, wherein the methylation pattern is determined at at least one site selected from the group consisting of: (i) CpG sites in the D-loop region shown in Table 1; (ii) CHG sites in the D-loop region shown in Table 3; (iii) CHH sites in the D-loop region shown in Table 5; (iv) CpG sites in the ND1 gene shown in Table 2; and (v) CHG sites in the ND1 gene shown in Table 4; and the methylation pattern is determined at at least one CHH site in the ND1 gene shown in Table 6. A method **Claim 2** The method according to claim 1, wherein the methylation pattern is determined at all CHH sites of the ND1 gene shown in Table 6. **Claim 3** The method according to any one of claims 1 to 2, wherein the methylation pattern is determined at all CHG sites of the D-loop region gene shown in Table 3. **Claim 4** The method according to any one of claims 1 to 3, wherein hypomethylation at at least one site of the CpG sites in the D-loop region, hypomethylation at at least one site of the CHG sites in the D-loop region, and / or hypomethylation at at least one site of the CHH sites in the D-loop region indicates that the subject has Lewy body dementia. **Claim 5** The method according to any one of claims 1 to 4, wherein the methylation pattern is determined at all CpG sites, CHG sites, and CHH sites of the D-loop region shown in Tables 1, 3, and 5. **Claim 6** The methylation pattern is determined using at least one oligonucleotide capable of specifically hybridizing to a mitochondrial DNA sequence comprising at least one methylation site selected from the group consisting of (i) to (v), the oligonucleotide having a length of 15 to 100 nucleotides, a mitochondrial DNA sequence comprising nucleotides 16,435 to 230 of the NCBI reference sequence: NC_012920.1 corresponding to the D-loop region, and / or a mitochondrial DNA sequence comprising nucleotides 3,257 to 3,682 of the NCBI reference sequence: NC_012920.1 corresponding to the ND1 gene, and being capable of specifically hybridizing thereto. The method according to any one of claims 1 to 5.
7. The method according to claim 6, wherein the oligonucleotide is a degenerate oligonucleotide.
8. The method according to claim 7, wherein the at least one oligonucleotide is selected from the group consisting of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, and SEQ ID NO:
4.
9. The method according to any one of claims 1 to 8, wherein the step of determining the methylation pattern is determined by bisulfite sequencing.
10. The method comprises (b) optionally combining the methylation pattern of one or more sites determined in step (a) with at least one clinical variable of the subject selected from the group consisting of demographic variables, neuropsychological variables, clinical finding variables, and clinical laboratory variables, wherein the combination is performed using a classification model for determining a score correlated with the identification of Lewy body dementia in the subject. The method according to any one of claims 1 to 9, further comprising the step.
11. The method according to claim 10, wherein at least one clinical variable of the subject is selected from the group consisting of gender, age, race, education, family history, praxis test, Luria test, clinical dementia rating scale, GDS (Global Deterioration Scale), Mini-Mental State Examination, CDR-SOB (Clinical Dementia Rating Scale Sum of Boxes), neuroleptic intolerance, REM sleep behavior disorder, autonomic neuropathy, parkinsonism, visual hallucination, cognitive fluctuation, magnetic resonance imaging diagnosis, dopamine transporter scan, positron emission tomography with 18F-fluorodeoxyglucose, amyloid positron emission tomography, apolipoprotein E genotype, alE4, β-42, tau-T, and tau-P.
12. The method according to claim 10, wherein the change over time of the score is related to the progression of the disease.
13. The method according to any one of claims 10 to 12, wherein the classification model is developed using a supervised machine learning method.
14. The method according to any one of claims 1 to 13, wherein the sample is a biological fluid selected from the group consisting of blood, plasma, saliva, cerebrospinal fluid, brain sample, skin sample, and urine.
15. A computer-implemented method for identifying DLB in a subject, comprising: (a) receiving data related to the methylation pattern of the D-loop region of the subject's mitochondrial DNA and / or the ND1 gene and optionally at least one clinical variable of the subject, wherein the methylation pattern is determined at least at one site selected from the group consisting of: (i) CpG sites in the D-loop region shown in Table 1, (ii) CHG sites in the D-loop region shown in Table 3, (iii) CHH sites in the D-loop region shown in Table 5, (iv) CpG sites in the ND1 gene shown in Table 2, and (v) CHG sites in the ND1 gene shown in Table 4, and the methylation pattern is determined at least at one CHH site in the ND1 gene shown in Table 6. (b)Determining a risk score correlated with the identification of the DLB of interest, wherein the risk score is calculated using a classification model configured to combine the methylation patterns of one or more sites of step (a) and optionally at least one clinical variable of the subject, step and A method comprising.