Predictive biomarkers and their use for treating Parkinson's disease
BMP and hemi-BMP biomarkers are used to identify PD patients with wild-type LRRK2 who can benefit from LRRK2 inhibitors, addressing the inadequacy of current treatments by enabling targeted therapy for Parkinson's disease.
Patent Information
- Application Number
- JP2025505811
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-02
- Filing Date
- 2023-08-02
- Publication Date
- 2025-09-02
AI Technical Summary
Current treatments for Parkinson's disease associated with wild-type LRRK2 are inadequate due to the inability to identify patients who would benefit from LRRK2 inhibitors, as these inhibitors cannot be indiscriminately administered without risking harm to patients without pathological LRRK2 activity.
Utilizing biomarkers such as bis(monoacylglycerol)phosphate (BMP) and/or hemi-bis(monoacylglycerol)phosphate (hemi-BMP) levels to determine whether PD patients with wild-type LRRK2 are likely to respond to LRRK2 inhibitor treatment, allowing for targeted therapy.
Enables the identification and treatment of a subset of PD patients with wild-type LRRK2 who can benefit from LRRK2 inhibitors, improving therapeutic outcomes for these patients.
Smart Images

Figure 2025528770000001_ABST
Abstract
Description
[Technical Field]
[0001] FIELD OF THE INVENTION The present invention relates to methods of treating and diagnosing patients with Parkinson's disease associated with wild-type leucine-rich repeat kinase 2 (LRRK2). [Background technology]
[0002] background Parkinson's disease (PD) is a progressive neurodegenerative disorder that affects over 6 million people worldwide. PD is usually first recognized by movement disorders, with cardinal symptoms being tremor, stiffness, slowness of movement, and difficulty walking. In later stages, PD also results in neuropsychiatric disorders, including dementia, depression, and anxiety. PD affects over 1% of people over the age of 60 and causes over 100,000 deaths annually.
[0003] PD is thought to result from a confluence of genetic and environmental factors. Although numerous mutations associated with familial PD have been identified, 85–90% of PD cases are idiopathic. Among PD cases potentially linked to known genetic factors, mutations in the LRRK2 gene are the most common cause of both familial and idiopathic PD. LRRK2 encodes a protein kinase expressed in multiple tissues, including brain regions associated with PD, such as the basal ganglia, and disease-causing mutations result in enhanced kinase activity. However, recent evidence indicates that some cases of PD are associated with increased activity of wild-type, or non-mutant, LRRK2.
[0004] Because there is no cure for PD, current treatments focus on alleviating symptoms, particularly movement disorders. For decades, the predominant approach has been to enhance dopaminergic function using the dopamine precursor levodopa, dopamine agonists, or monoamine oxidase inhibitors. However, such medications lose their effectiveness as the disease progresses, and eventually their side effects outweigh their benefits. Summary of the Invention
[0005] overview More recently, the use of LRRK2 inhibitors has been investigated for the treatment of PD cases associated with mutant forms of the LRRK2 kinase. However, in the majority of PD cases, mutations in LRRK2 cannot be identified. Unfortunately, for PD patients with wild-type LRRK2, there is no way to identify the subset of patients whose disease is associated with elevated LRRK2 activity, and LRRK2 inhibitors cannot be indiscriminately administered to PD patients due to the risk of harming patients without pathological LRRK2 activity. As a result, existing treatments for most PD patients are inadequate, and millions of people continue to suffer from the progressive and debilitating effects of this disease.
[0006] The present invention solves this problem by using biomarkers to determine whether PD patients with wild-type LRRK2 are likely to benefit from LRRK2 inhibitor treatment. The present invention provides a method for using predictive biomarkers to determine whether PD patients with wild-type LRRK2 are more likely to respond to LRRK2 inhibitors. The present invention recognizes that biomarkers including bis(monoacylglycerol)phosphate (BMP) and / or hemi-bis(monoacylglycerol)phosphate (hemi-BMP) can be relied upon as biomarkers for determining whether LRRK2 inhibitor treatment is effective in treating PD. The present invention recognizes that BMP and / or hemi-BMP levels can be indicators of LRRK2 kinase levels and activity. Thus, the present invention recognizes that BMP and / or hemi-BMP levels serve as indicators for determining whether LRRK2 inhibitor treatment is appropriate for a given individual. The methods of the present invention are useful for both identifying PD patients as candidates for LRRK2 inhibitor treatment and treating such patients.
[0007] In one aspect, the present invention provides a method of treating a patient with Parkinson's disease associated with LRRK2, comprising providing one or more LRRK2 inhibitors to a patient exhibiting PD with wild-type LRRK2 and elevated levels of BMP and / or hemi-BMP compared to BMP and / or hemi-BMP levels in subjects not affected by the neurological disease, thereby treating the PD associated with wild-type LRRK2. In certain embodiments, the BMP may be selected from the group consisting of (i) di-22:6 BMP, (ii) di-18:1 BMP, (iii) 16:0 / 18:1 BMP, (iv) di-20:4 BMP, (v) 18:0 / 20:4 BMP, (vi) 2,2' di-22:6 BMP, and (vii) any combination thereof. In certain embodiments, the hemi-BMP may be selected from the group consisting of (i) hemi-BMP (18:1 / 18:1)_16:0, (ii) hemi-BMP (14:0 / 14:0)_14:0, (iii) hemi-BMP (18:1 / 18:1)_18:0, (iv) hemi-BMP (18:1 / 18:1)_18:1, and (iv) any combination thereof. BMP levels are measured in any biological fluid from a patient. In certain embodiments, the biological fluid for measuring BMP and / or hemi-BMP levels is urine, blood, cerebrospinal fluid (CSF), bile, or saliva. In certain preferred embodiments, the biological fluid used to measure BMP and / or hemi-BMP is urine or CSF. In certain embodiments, the elevated levels of BMP and / or hemi-BMP in the patient are at concentrations that indicate the patient is responsive to one or more LRRK2 inhibitors. In certain embodiments, the one or more LRRK2 inhibitors for treatment are selected from the group consisting of CZC-25146, CZC-54252, DNL151, DNL201, GNE-7915, GSK2578215A, HG-10-102-01, JH-II-127, K252A, K252B, LRRK2-IN-1, MLi-2, PF-06447475, and staurosporine. In certain embodiments, the LRRK2 inhibitor is represented by Formulas (I), (II), (III), and (IV): [ka] (In the formula, A is NH, O, S, C=O, NR 3 or CR 4 R 5 and X is an optionally substituted arylene, heteroarylene, cycloalkylene, heterocycloalkylene, alkylcycloalkylene, heteroalkylcycloalkylene, aralkylene, or heteroaralkylene group; R 1 is an optionally substituted alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 2 is a hydrogen atom, a halogen atom, NO, N, OH, SH, NH, or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 3 is an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkyl-cycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 4 is a hydrogen atom, NO, N, OH, SH, NH, or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 5 is a hydrogen atom, NO, N, OH, SH, NH, or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; B is NH, O, S, C=O, NR 14 or CR 15 R 16 and R 11 is an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 12 is an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 12 is attached to the pyrimidine ring of formula (II) via a carbon-carbon bond, R 13 is a hydrogen atom, a halogen atom, NO, N, OH, SH, NH, or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 14 is an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkyl-cycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 15 is a hydrogen atom, NO, N, OH, SH, NH, or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 16is a hydrogen atom, NO, N, OH, SH, NH, or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 21 each of which is optionally substituted aryl or heteroaryl; R 22 H, halo, OH, CN, CF3, C 1~6 Alkyl, C 1~6 Alkoxy, C 1~6 Haloalkyl, C 1~6 Thioalkyl, C 3~8 Cycloalkyl, C 2~8 heterocycloalkyl, aryl, or heteroaryl; Y is aryl or 5- or 6-membered heteroaryl; C 1~6 Alkyl, C 1~6 Alkoxy, C 1~6 Haloalkyl, C 1~6 Thioalkyl, C 3~8 Cycloalkyl, C 2~8 Each of heterocycloalkyl, aryl, and heteroaryl is selected from halo, OH, CN, CF, NH, NO, C 1~6 Alkyl, C 1~6 Haloalkyl, C 1~6 Thioalkyl, C 3~8 Cycloalkyl, C 2~8 Heterocycloalkyl, C 2~8 Heterocycloalkenyl, C 2~6 Alkenyl, C 2~6 Alkynyl, C 1~6 Alkoxy, C 1~6 Haloalkoxy, C 1~6 Alkylamino, C 2~6 Dialkylamino, C 7~12 Aralkyl, C 1~12optionally substituted with one or more moieties selected from the group consisting of heteroaralkyl, aryl, heteroaryl, —C(O)R, —C(O)OR, —C(O)NRR′, —C(O)NRS(O)R′, —C(O)NRS(O)NR′R″, —OR, —OC(O)NRR′, —NRR′, —NRC(O)R′, —NRC(O)NR′R″, —NRS(O)R′, —NRS(O)NR′R″, —S(O)R, and —S(O)NRR; Each of R, R' and R" is independently H, halo, OH, C 1~6 Alkyl, C 1~6 Haloalkyl, C 1~6 Alkoxy, C 3~8 Cycloalkyl, C 2~8 heterocycloalkyl, aryl, or heteroaryl, or R and R′ or R′ and R″ together with the nitrogen to which they are attached are C 2~8 forming a heterocycloalkyl, R 31 is C(O)CH2R 33 , optionally substituted cycloalkyl, optionally substituted cycloheteroalkyl, optionally substituted cycloalkenyl, optionally substituted cycloheteroalkenyl, optionally substituted aryl, or optionally substituted heteroaryl; R 32 each instance of is independently halo, haloalkyl, optionally substituted alkoxyl, optionally substituted alkyl, optionally substituted heteroalkyl, optionally substituted alkenyl, optionally substituted heteroalkenyl; R 33 is optionally substituted cycloalkyl, optionally substituted cycloheteroalkyl, optionally substituted cycloalkenyl, optionally substituted cycloheteroalkenyl, optionally substituted aryl, or optionally substituted heteroaryl; Z is cycloalkyl, cycloheteroalkyl, cycloalkenyl, cycloheteroalkenyl, aryl or heteroaryl, Z can be aryl substituted with 2 or 3 instances of R, Z can be phenyl substituted with 2 or 3 instances of R, Z can be heteroaryl substituted with 2 or 3 instances of R, Z can be a 6-membered heteroaryl substituted with 2 or 3 instances of R; n is 0 to 5) or a pharmaceutically acceptable salt of any of the above compounds.
[0008] In another aspect, the present invention provides a method for determining whether a patient with PD associated with wild-type LRRK2 will respond to an LRRK2 inhibitor, the method comprising: performing an assay to measure the patient's BMP and / or hemi-BMP levels; generating a report identifying the patient's BMP and / or hemi-BMP levels compared to the BMP and / or hemi-BMP levels of a subject not afflicted with PD and having wild-type LRRK2; and, if the report indicates elevated levels of BMP and / or hemi-BMP in the patient compared to the subject, providing the report to a physician so that the physician may prescribe or provide one or more LRRK2 inhibitors to the patient. In certain embodiments, the BMP may be selected from the group consisting of (i) di-22:6 BMP, (ii) di-18:1 BMP, (iii) 16:0 / 18:1 BMP, (iv) di-20:4 BMP, (v) 18:0 / 20:4 BMP, (vi) 2,2' di-22:6 BMP, and (vii) any combination thereof. In certain embodiments, the hemi-BMP may be selected from the group consisting of (i) hemi-BMP (18:1 / 18:1)_16:0, (ii) hemi-BMP (14:0 / 14:0)_14:0, (iii) hemi-BMP (18:1 / 18:1)_18:0, (iv) hemi-BMP (18:1 / 18:1)_18:1, and (iv) any combination thereof. BMP and / or hemi-BMP levels are measured in any biological fluid from a patient. In certain embodiments, the biological fluid for measuring BMP and / or hemi-BMP levels is urine, blood, cerebrospinal fluid (CSF), bile, or saliva. In certain preferred embodiments, the biological fluid used to measure BMP and / or hemi-BMP is urine or CSF. In certain embodiments, the elevated levels of BMP and / or hemi-BMP in the patient are at concentrations that indicate the patient is responsive to one or more LRRK2 inhibitors. The LRRK2 inhibitor may be any of those described above.
[0009] In another aspect, the present invention provides a method of treating a patient with PD associated with wild-type LRRK2, comprising receiving data identifying the patient's BMP and / or hemi-BMP levels, comparing the data with BMP and / or hemi-BMP levels from a subject without PD who has wild-type LRRK2, and prescribing or providing one or more LRRK2 inhibitors to the patient if the patient has elevated levels of BMP and / or hemi-BMP compared to the subject. In certain embodiments, the BMP may be selected from the group consisting of (i) di-22:6 BMP, (ii) di-18:1 BMP, (iii) 16:0 / 18:1 BMP, (iv) di-20:4 BMP, (v) 18:0 / 20:4 BMP, (vi) 2,2' di-22:6 BMP, and (vii) any combination thereof. In certain embodiments, the hemi-BMP may be selected from the group consisting of (i) hemi-BMP (18:1 / 18:1)_16:0, (ii) hemi-BMP (14:0 / 14:0)_14:0, (iii) hemi-BMP (18:1 / 18:1)_18:0, (iv) hemi-BMP (18:1 / 18:1)_18:1, and (iv) any combination thereof. BMP and / or hemi-BMP levels are measured in any biological fluid from a patient. In certain embodiments, the biological fluid for measuring BMP and / or hemi-BMP levels is urine, blood, cerebrospinal fluid (CSF), bile, or saliva. In certain preferred embodiments, the biological fluid used to measure BMP and / or hemi-BMP is urine or CSF. In certain embodiments, the elevated levels of BMP and / or hemi-BMP in the patient are at concentrations that indicate the patient is responsive to one or more LRRK2 inhibitors. The LRRK2 inhibitor may be any of those described above. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 provides total urinary di-18:1-BMP levels in different patient population sets. [Figure 2]FIG. 2 provides total urinary di-22:6-BMP levels in different patient population sets. [Figure 3] FIG. 3 provides total urinary 2,2'di-22:6-BMP levels in different patient population sets. [Figure 4] Figure 4 provides a classification of the population with Parkinson's disease and LRRK2 status. [Figure 5] FIG. 5 provides urinary BMP (uBMP) levels in patients with LRRK2-active and LRRK2-normal PD. [Figure 6A] Figures 6A and 6B provide that the LRRK2 activity fraction in idiopathic Parkinson's disease is stable among different cohorts of AMP-PD. [Figure 6B] Figure 6B and Figure 6B provide that the LRRK2 activity fraction in idiopathic Parkinson's disease is stable across different cohorts of AMP-PD. DETAILED DESCRIPTION OF THE INVENTION
[0011] Detailed Description Parkinson's disease (PD) is a progressive neurodegenerative disorder caused by both genetic and environmental factors. One gene that plays a role in the development of some cases of PD is LRRK2, which encodes a kinase expressed in multiple tissues, including brain regions associated with PD, such as the basal ganglia. While mutations in LRRK2 are the most common known genetic cause of PD, patients with LRRK2 mutations comprise a small fraction of the total number of PD cases. Nevertheless, the pathology of some patients with wild-type, or non-mutant, LRRK2 appears to resemble that of patients with mutant LRRK2. Notably, disease-causing mutations in LRRK2 result in increased activity of the LRRK2 kinase, and it has recently been shown that LRRK2 activity is elevated in some PD patients with wild-type LRRK2.
[0012] Various LRRK2 inhibitors are currently being investigated as potential treatments for PD. These drugs are promising for PD patients with LRRK2 mutations. However, the use of LRRK2 inhibitors to treat PD patients with wild-type LRRK2 is problematic due to the variable etiology of the disease. While patients with enhanced wild-type LRRK2 activity would benefit from LRRK2 inhibitors, LRRK2 inhibition may not be effective in PD patients with normal levels of LRRK2 activity whose disease pathology is due to alterations in other molecular pathways. Because LRRK2-expressing neurons are located in the midbrain and are extremely difficult to access, kinase activity cannot be assessed in living patients. As a result, until now, there has been no means to identify a subset of PD patients who have wild-type LRRK2 but who could still benefit from LRRK2 inhibition.
[0013] The present invention solves this problem by using biomarkers to determine whether PD patients with wild-type LRRK2 are likely to benefit from LRRK2 inhibitors. As a result, the methods of the present invention allow for the identification of candidates for LRRK2 drug therapy based on genetic data that can be readily obtained from the patient. Thus, for a subset of PD patients, the present invention reveals the therapeutic potential of classes of drugs that have not previously been recommended for PD patients.
[0014] Parkinson's disease and its treatment Parkinson's disease (PD) is a progressive neurodegenerative disorder of the central nervous system. In the early stages, the disease affects the motor system, and cardinal symptoms are tremor, rigidity, slowness of movement, and difficulty walking. Cognitive and behavioral symptoms, such as dementia, depression, and anxiety, often appear in the later stages of PD. PD typically occurs in people over 60 years of age, affecting approximately 1% of those affected, although so-called early-onset PD can occur before age 50.
[0015] PD is characterized by the death of cells in the basal ganglia, including dopamine-secreting neurons, astrocytes, and microglia in the substantia nigra. Five mechanisms for neuronal death in PD have been proposed. First, oligomerization of proteins such as alpha-synuclein into aggregates called Lewy bodies can directly lead to cell death. A second proposed cause is dysregulation of autophagy, particularly mitochondrial degradation. Another proposed mechanism is that mitochondrial dysfunction leads to decreased energy production and increased reactive oxygen species. A fourth proposed mechanism is neuroinflammation resulting from the secretion of pro-inflammatory factors by microglia. Finally, it has been proposed that disruption of the blood-brain barrier allows plasma proteins to leak into the substantia nigra, promoting apoptosis.
[0016] PD is thought to result from a combination of genetic and environmental factors. In some cases, genetic mutations that increase the risk of PD are inherited, and approximately 10–15% of individuals with PD have a first-degree relative with the disease. However, most cases of PD are idiopathic or "sporadic." Genes with mutations implicated in PD include CHCHD2, DJ1 / PARK7, DNAJC13, EIF4G1, GBA, LRRK2 / PARK8, PINK1, PRKN, SNCA, UCHL1, and VPS35. The most common known cause of both familial and sporadic PD is mutations in LRRK2. Disease-causing mutations in LRRK2 result in a form of the kinase with increased activity. Enhanced wild-type LRRK2 activity has also recently been implicated in idiopathic PD. The role of LRRK2 in PD has been discussed, for example, in Chen et al., Leucine-Rich Repeat Kinase 2 in Parkinson's Disease: Updated from Pathogenesis to Potential Therapeutic Target, Eur Neurol. 2018;79(5-6):256-265, doi:10.1159 / 000488938. Epub 2018 Apr 27; Di Maio et al., LRRK2 activation in idiopathic Parkinson's disease, Sci Transl Med. 2018 Jul 25;10(451):eaar5429, doi:10.1126 / scitranslmed.aar5429; Taymans and Greggio, LRRK2 Kinase Inhibition as a Therapeutic Strategy for Parkinson's Disease, Where Do We Stand? Curr Neuropharmacol. 2016;14(3):214-25, doi:10.2174 / 1570159x13666151030102847, the contents of each of which are incorporated herein by reference.
[0017] Several behavioral and environmental conditions are known to increase the risk of developing PD. Risk factors associated with PD include exposure to pesticides and a history of head trauma. Caffeine consumption and tobacco use are associated with a decreased risk of PD. Low blood uric acid levels are associated with an increased risk of PD.
[0018] Management of PD usually involves pharmacological stimulation of the dopaminergic system. The most widely used drug for treating PD is levodopa, which is enzymatically converted to dopamine in dopaminergic neurons. Dopamine agonists such as bromocriptine, pergolide, pramipexole, ropinirole, piribedil, cabergoline, apomorphine, and lisuride can also be used to treat PD. A third class of drugs for the treatment of PD includes inhibitors of monoamine oxidase, such as selegiline and rasagiline.
[0019] BMP and hemi-BMP BMPs and their isoforms (including hemi-BMPs) are localized within the inner membranes of late endosomes (multivesicular bodies) and lysosomes, where they contribute to the multivesicular / lamellar morphology of the endosomal network. BMPs are structural isomers of phosphatidylglycerol (PG) and are synthesized through a series of acylation and deacylation steps involving transacylases that reorient the glycerol backbone. They possess a unique sn-1-glycerophospho-sn-1'-glycerol conformation, making them highly resistant to degradation by phospholipases in acidic organelles. Because they are negatively charged at lysosomal pH, they can act as docking platforms to recruit positively charged lipid hydrolases to ILVs, thereby facilitating the degradation of lipid cargo. Furthermore, BMPs are important cofactors in lysosomal cholesterol and sphingolipid metabolism through interactions with cholesterol transport and sphingolipid activating proteins. Based on these versatile functions, BMPs are believed to be key activators of lipid sorting and digestion.
[0020] Bis(acylglycero)phosphates can contain multiple acyl chains, i.e., two (BMPs), three (known as hemi-BMPs), or four (bis(diacylglycero)phosphate, BDP). A wide variety of BMP and hemi-BMP species exist. Examples of BMP and hemi-BMP species are provided in Showalter et al. (Int. J. Mol. Sci. 2020, 21, 8067), which is incorporated by reference in its entirety.
[0021] Exemplary BMP and hemi-BMP species of the present invention are di-22:6 BMP, di-18:1 BMP, 16:0 / 18:1 BMP, di-20:4 BMP, 18:0 / 20:4 BMP, 2,2'di-22:6 BMP, hemi-BMP(18:1 / 18:1)_16:0, hemi-BMP(14:0 / 14:0)_14:0, hemi-BMP(18:1 / 18:1)_18:0, and hemi-BMP(18:1 / 18:1)_18:1.
[0022] The present invention also provides that lipids closely related to BMPs can also be used as biomarkers for predicting responsiveness to LRRK2 inhibitor treatment.
[0023] Measurement of BMP levels The present invention provides that BMP levels can be measured in any biological fluid from a patient. BMP levels can be measured in any biological fluid from a patient. Biological fluids from a patient include urine, blood, cerebrospinal fluid (CSF), bile, or saliva. In particular, the present invention provides that BMP levels are measured in CSF or urine. In a more preferred embodiment, BMP levels are measured in urine. A wide variety of techniques can be used to measure BMP levels. For example, techniques used to measure BMP and / or hemi-BMP include high-pressure liquid chromatography (HPLC) and mass spectrometry (MS). Overviews of techniques used to measure BMPs and / or hemi-BMPs are provided in Luquain et al., High-Performance Liquid Chromatography Determination of Bis(monoacylglycerol)Phosphate and Other Lysophospholipids, 2001, 296(1), 41-48, and Pan et al., Quantitative Analysis of Polyphosphoinositide, Bis(monoacylglycero)phosphate and Phosphatidylglycerol Species by Shotgun Lipidomics after Methylation, 2021, 2306:77-91, which are incorporated by reference in their entireties.
[0024] BMPs as biomarkers for LRRK2 inhibitor therapy The present invention provides a method for using predictive biomarkers to determine whether PD patients with wild-type LRRK2 are more likely to respond to LRRK2 inhibitors. The present invention recognizes that biomarkers including bis(monoacylglycerol)phosphate (BMP) and / or hemi-bis(monoacylglycerol)phosphate (hemi-BMP) can be relied upon as biomarkers for determining whether LRRK2 inhibitor therapy is effective in treating PD. The present invention recognizes that the level of BMP and / or hemi-BMP can be an indicator of the level and activity of LRRK2 kinase. Thus, BMP and / or hemi-BMP levels serve as indicators for determining whether LRRK2 inhibitor therapy is appropriate for a given individual. The methods of the present invention are useful for both identifying PD patients as candidates for LRRK2 inhibitor therapy and treating such patients.
[0025] In one aspect, the present invention provides a method of treating a patient with Parkinson's disease associated with LRRK2, comprising providing one or more LRRK2 inhibitors to a patient exhibiting PD with wild-type LRRK2 and elevated levels of BMP and / or hemi-BMP compared to BMP and / or hemi-BMP levels in subjects not affected by the neurological disease, thereby treating the PD associated with wild-type LRRK2. In certain embodiments, the BMP may be selected from the group consisting of (i) di-22:6 BMP, (ii) di-18:1 BMP, (iii) 16:0 / 18:1 BMP, (iv) di-20:4 BMP, (v) 18:0 / 20:4 BMP, (vi) 2,2' di-22:6 BMP, and (vii) any combination thereof. In certain embodiments, the hemi-BMP may be selected from the group consisting of (i) hemi-BMP (18:1 / 18:1)_16:0, (ii) hemi-BMP (14:0 / 14:0)_14:0, (iii) hemi-BMP (18:1 / 18:1)_18:0, (iv) hemi-BMP (18:1 / 18:1)_18:1, and (iv) any combination thereof. BMP levels are measured in any biological fluid from a patient. In certain embodiments, the biological fluid for measuring BMP and / or hemi-BMP levels is urine, blood, cerebrospinal fluid (CSF), bile, or saliva. In certain preferred embodiments, the biological fluid used to measure BMP and / or hemi-BMP is urine or CSF. In certain embodiments, the elevated levels of BMP and / or hemi-BMP in the patient are at concentrations that indicate the patient is responsive to one or more LRRK2 inhibitors. The LRRK2 inhibitor of the present invention can be any inhibitor of LRRK2. Exemplary LRRK2 inhibitors useful in the present invention can be selected from the group consisting of CZC-25146, CZC-54252, DNL151, DNL201, GNE-7915, GSK2578215A, HG-10-102-01, JH-II-127, K252A, K252B, LRRK2-IN-1, MLi-2, PF-06447475, and staurosporine. In certain embodiments, the LRRK2 inhibitor is represented by Formulas (I), (II), (III), and (IV): [ka] (In the formula, A is NH, O, S, C=O, NR 3 or CR 4 R 5 and X is an optionally substituted arylene, heteroarylene, cycloalkylene, heterocycloalkylene, alkylcycloalkylene, heteroalkylcycloalkylene, aralkylene, or heteroaralkylene group; R 1 is an optionally substituted alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 2 is a hydrogen atom, a halogen atom, NO, N, OH, SH, NH, or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 3 is an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkyl-cycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 4 is a hydrogen atom, NO, N, OH, SH, NH, or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 5is a hydrogen atom, NO, N, OH, SH, NH, or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; B is NH, O, S, C=O, NR 14 or CR 15 R 16 and R 11 is an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 12 is an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 12 is attached to the pyrimidine ring of formula (II) via a carbon-carbon bond, R 13 is a hydrogen atom, a halogen atom, NO, N, OH, SH, NH, or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 14 is an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkyl-cycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 15 is a hydrogen atom, NO, N, OH, SH, NH, or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 16 is a hydrogen atom, NO, N, OH, SH, NH, or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 21 each of which is optionally substituted aryl or heteroaryl; R 22 H, halo, OH, CN, CF3, C 1~6 Alkyl, C 1~6 Alkoxy, C 1~6 Haloalkyl, C 1~6 Thioalkyl, C 3~8 Cycloalkyl, C 2~8 heterocycloalkyl, aryl, or heteroaryl; Y is aryl or 5- or 6-membered heteroaryl; C 1~6 Alkyl, C 1~6 Alkoxy, C 1~6 Haloalkyl, C 1~6 Thioalkyl, C 3~8 Cycloalkyl, C 2~8 Each of heterocycloalkyl, aryl, and heteroaryl is selected from halo, OH, CN, CF, NH, NO, C 1~6 Alkyl, C 1~6 Haloalkyl, C 1~6 Thioalkyl, C 3~8 Cycloalkyl, C 2~8 Heterocycloalkyl, C 2~8 Heterocycloalkenyl, C 2~6 Alkenyl, C 2~6 Alkynyl, C 1~6 Alkoxy, C 1~6 Haloalkoxy, C 1~6 Alkylamino, C 2~6 Dialkylamino, C 7~12 Aralkyl, C 1~12optionally substituted with one or more moieties selected from the group consisting of heteroaralkyl, aryl, heteroaryl, —C(O)R, —C(O)OR, —C(O)NRR′, —C(O)NRS(O)R′, —C(O)NRS(O)NR′R″, —OR, —OC(O)NRR′, —NRR′, —NRC(O)R′, —NRC(O)NR′R″, —NRS(O)R′, —NRS(O)NR′R″, —S(O)R, and —S(O)NRR; Each of R, R' and R" is independently H, halo, OH, C 1~6 Alkyl, C 1~6 Haloalkyl, C 1~6 Alkoxy, C 3~8 Cycloalkyl, C 2~8 heterocycloalkyl, aryl, or heteroaryl, or R and R′ or R′ and R″ together with the nitrogen to which they are attached are C 2~8 forming a heterocycloalkyl, R 31 is C(O)CH2R 33 , optionally substituted cycloalkyl, optionally substituted cycloheteroalkyl, optionally substituted cycloalkenyl, optionally substituted cycloheteroalkenyl, optionally substituted aryl, or optionally substituted heteroaryl; R 32 each instance of is independently halo, haloalkyl, optionally substituted alkoxyl, optionally substituted alkyl, optionally substituted heteroalkyl, optionally substituted alkenyl, optionally substituted heteroalkenyl; R 33 is optionally substituted cycloalkyl, optionally substituted cycloheteroalkyl, optionally substituted cycloalkenyl, optionally substituted cycloheteroalkenyl, optionally substituted aryl, or optionally substituted heteroaryl; Z is cycloalkyl, cycloheteroalkyl, cycloalkenyl, cycloheteroalkenyl, aryl or heteroaryl, Z can be aryl substituted with 2 or 3 instances of R, Z can be phenyl substituted with 2 or 3 instances of R, Z can be heteroaryl substituted with 2 or 3 instances of R, Z can be a 6-membered heteroaryl substituted with 2 or 3 instances of R; n is 0 to 5) or a pharmaceutically acceptable salt of any of the above compounds.
[0026] In another aspect, the present invention provides a method for determining whether a patient having PD associated with wild-type LRRK2 will respond to an LRRK2 inhibitor, the method comprising: performing an assay to measure BMP and / or hemi-BMP levels in the patient; generating a report identifying the patient's BMP and / or hemi-BMP levels compared to BMP and / or hemi-BMP levels in a subject not afflicted with PD and having wild-type LRRK2; and if the report indicates elevated levels of BMP and / or hemi-BMP in the patient compared to the subject, providing the report to a physician so that the physician may prescribe or provide one or more LRRK2 inhibitors to the patient. In certain embodiments, the BMP may be selected from the group consisting of (i) di-22:6 BMP, (ii) di-18:1 BMP, (iii) 16:0 / 18:1 BMP, (iv) di-20:4 BMP, (v) 18:0 / 20:4 BMP, (vi) 2,2' di-22:6 BMP, and (vii) any combination thereof. In certain embodiments, the hemi-BMP may be selected from the group consisting of (i) hemi-BMP (18:1 / 18:1) 16:0, (ii) hemi-BMP (14:0 / 14:0) 14:0, (iii) hemi-BMP (18:1 / 18:1) 18:0, (iv) hemi-BMP (18:1 / 18:1) 18:1, and (iv) any combination thereof. BMP and / or hemi-BMP levels are measured in any biological fluid from a patient. In certain embodiments, the biological fluid for measuring BMP and / or hemi-BMP levels is urine, blood, cerebrospinal fluid (CSF), bile, or saliva. In certain preferred embodiments, the biological fluid used to measure BMP and / or hemi-BMP is urine or CSF. In certain embodiments, the elevated level of BMP and / or hemi-BMP in a patient is at a concentration that indicates that the patient is responsive to one or more LRRK2 inhibitors. The LRRK2 inhibitor may be any of the above.
[0027] In another aspect, the present invention provides a method of treating a patient with PD associated with wild-type LRRK2, comprising receiving data identifying the patient's BMP and / or hemi-BMP levels, comparing the data with BMP and / or hemi-BMP levels from a subject without PD who has wild-type LRRK2, and prescribing or providing one or more LRRK2 inhibitors to the patient if the patient has elevated levels of BMP and / or hemi-BMP compared to the subject. In certain embodiments, the BMP may be selected from the group consisting of (i) di-22:6 BMP, (ii) di-18:1 BMP, (iii) 16:0 / 18:1 BMP, (iv) di-20:4 BMP, (v) 18:0 / 20:4 BMP, (vi) 2,2' di-22:6 BMP, and (vii) any combination thereof. In certain embodiments, the hemi-BMP may be selected from the group consisting of (i) hemi-BMP (18:1 / 18:1)_16:0, (ii) hemi-BMP (14:0 / 14:0)_14:0, (iii) hemi-BMP (18:1 / 18:1)_18:0, (iv) hemi-BMP (18:1 / 18:1)_18:1, and (iv) any combination thereof. BMP and / or hemi-BMP levels are measured in any biological fluid from a patient. In certain embodiments, the biological fluid for measuring BMP and / or hemi-BMP levels is urine, blood, cerebrospinal fluid (CSF), bile, or saliva. In certain preferred embodiments, the biological fluid used to measure BMP and / or hemi-BMP is urine or CSF. In certain embodiments, the elevated levels of BMP and / or hemi-BMP in the patient are at concentrations that indicate the patient is responsive to one or more LRRK2 inhibitors. The LRRK2 inhibitor may be any of those described above.
[0028] The present invention provides methods and systems for predicting a subject's responsiveness to an LRRK2 inhibitor based on the subject's BMP and / or hemi-BMP levels. In some embodiments, the methods and systems of the present invention use a diagnostic signature to predict responsiveness. The diagnostic predictor can be based on any suitable pattern recognition method that receives input data representing (1) BMP levels in PD patients who are carriers of LRRK2 deleterious variants, (2) PD of unknown mechanism, and (3) appropriate controls, and provides an output indicating the probability that the subject will respond to an LRRK2 inhibitor. The diagnostic predictor can be trained with data from multiple individuals whose BMP and / or hemi-BMP levels, medical interventions, and LRRK2 inhibitor response outcomes are known. The multiple individuals used to train the diagnostic predictor are also known as a training population. For each individual in the training population, the training data includes (a) data representing the patient's BMP and / or hemi-BMP levels, (b) medical interventions, and (c) LRRK2 inhibitor response information. The LRRK2 inhibitor response outcome may not be required to generate a diagnostic signature. LRRK2 inhibitor response can be evaluated in a prospectively selected patient population.Various diagnostic predictors that can be used in conjunction with the present invention are described below.In some embodiments, additional individuals with known trait profiles and LRRK2 response outcomes can be used to test the accuracy of diagnostic predictors obtained using the training population.Such additional patients are known as the test population.
[0029] In certain embodiments, the methods of the present invention use a diagnostic predictor, also called a classifier, to determine the probability of responding to LRRK2 inhibition. As described above, the diagnostic predictor can be based on any suitable pattern recognition method that receives a profile, such as a profile based on multiple phenotypic traits, and provides an output containing data indicating whether a patient is likely or unlikely to respond to an LRRK2 inhibitor, and may include the potential risks and benefits of treatment with such an inhibitor. The profile can be obtained by completing a questionnaire containing questions about specific phenotypic traits, or by collecting a biological sample to obtain genotypic data or a combination thereof. The diagnostic predictor is trained with training data from a training population of individuals with known phenotypic traits, medical interventions, and LRRK2 inhibitor response outcomes.
[0030] Diagnostic predictors based on any of these methods can be constructed using the profiles and diagnostic data of training patients.Then, these diagnostic predictors can be used to predict the response of subjects to LRRK2 inhibitors based on the profiles of phenotype traits, genotype traits, or both.This method can also be used to identify the traits that distinguish between responding and not responding to LRRK2 inhibition using the trait profiles and diagnostic data of training population.
[0031] In one embodiment, a diagnostic predictor can be prepared by (a) generating a reference set of individuals with known phenotypic traits, medical interventions, and LRRK2 response outcomes; (b) for each trait, determining a metric of correlation between the trait and LRRK2 response outcomes in a plurality of individuals with known LRRK2 response outcomes at a given time point; (c) selecting one or more traits based on the level of association; and (d) training a diagnostic predictor, wherein the diagnostic predictor receives data representing the traits selected in the previous step and provides an output indicative of the probability of responding to LRRK2 inhibition, and the training data from the reference set of subjects includes assessments of the traits obtained from individuals.
[0032] In some embodiments, the diagnostic predictor is based on a regression model, preferably a logistic regression model. Such a regression model includes coefficients for each of the markers in the selected set of markers of the present invention. In such embodiments, the coefficients of the regression model are calculated, for example, using a maximum likelihood approach.
[0033] Cox proportional hazards regression also includes the coefficient for each of the markers in the selected set of markers of the present invention.Cox proportional hazards regression incorporates censored data (individuals in the reference set who did not return to treatment).In such an embodiment, the coefficients of the regression model are calculated, for example, using maximum partial likelihood approach.
[0034] Some embodiments of the present invention provide a generalization of logistic regression models to handle multi-category (multinomial) responses. Such embodiments can be used to differentiate organisms into one, three, or more diagnostic groups. Such regression models use a multicategory logit model, which simultaneously looks at all pairs of categories and describes the odds of a response in one category rather than another. Once the model specifies the logits for a particular (J-1) pair of categories, the rest become redundant. See, for example, Agresti, An Introduction to Categorical Data Analysis, John Wiley & Sons, Inc., 1996, New York, Chapter 8, incorporated herein by reference. Linear discriminant analysis (LDA) attempts to classify objects into one of two categories based on specific object characteristics. In other words, LDA tests whether experimentally measured object attributes predict the object's classification. LDA typically requires continuous independent variables and a dichotomous categorical dependent variable. In the present invention, the selected phenotypic traits serve as the necessary continuous independent variables, and the diagnostic group classification of each member of the training population serves as the dichotomous categorical dependent variable.
[0035] LDA uses grouping information to find a linear combination of variables that maximizes the ratio of inter-group variance to intra-group variance. Implicitly, the linear weights used by LDA depend on how the selected phenotypic trait appears in two groups (e.g., a group that responds to LRRK2 inhibition and a group that does not respond to LRRK2 inhibition) and how the selected trait correlates with the expression of other traits. For example, LDA can be applied to a data matrix of N members in a training sample by K genes in the gene combination described in the present invention. The linear discriminant of each member of the training population is then plotted. Ideally, members of the training population representing a first subgroup (e.g., subjects that do not respond to LRRK2 inhibition) are clustered into one range of linear discriminant values (e.g., negative), and members of the training population representing a second subgroup (e.g., subjects that respond to LRRK2 inhibition) are clustered into a second range of linear discriminant values (e.g., positive). LDA is considered more successful when the separation between clusters of discriminant values is greater. For more information on linear discriminant analysis, see Duda, Pattern Classification, 2nd ed., 2001, John Wiley & Sons, Inc.; and Hastie, 2001, The Elements of Statistical Learning, Springer, New York; Venables & Ripley, 1997, Modern Applied Statistics with s-plus, Springer, New York.
[0036] Quadratic discriminant analysis (QDA) takes the same input parameters and returns the same results as LDA. QDA uses a quadratic equation rather than a linear equation to generate its results. LDA and QDA are interchangeable, and which one to use is a matter of preference and / or availability of software to support the analysis. Logistic regression takes the same input parameters and returns the same results as LDA and QDA.
[0037] In some embodiments of the present invention, a decision tree is used to classify patients using expression data of a selected set of molecular markers of the present invention. Decision tree algorithms belong to the class of supervised learning algorithms. The purpose of a decision tree is to derive a classifier (tree) from real-world example data. This tree can be used to classify unknown examples that were not used to derive the decision tree.
[0038] The decision tree is derived from training data. An example includes values of different attributes and which class the example belongs to. In one embodiment, the training data is data representing multiple phenotypic traits, medical interventions, and LRRK2 inhibition response outcomes.
[0039] The following algorithm describes the decision tree derivation:
number
number
number
number
[0040] The information gain of a particular attribute A is calculated as the difference between the information content of the class and the residual of attribute A.
number
[0041] In general, there are many different decision tree algorithms, many of which are described in Duda, Pattern Classification, Second Edition, 2001, John Wiley & Sons, Inc. Decision tree algorithms often require consideration of feature processing, impurity measures, stopping criteria, and pruning. Specific decision tree algorithms include, but are not limited to, classification and regression trees (CART), multivariate decision trees, ID3, and C4.5.
[0042] In one approach, when using exemplary embodiments of a decision tree, data representing multiple phenotypic traits across a training population are standardized to have a mean of zero and unit variance. Members of the training population are randomly divided into a training set and a test set. For example, in one embodiment, 2 / 3 of the members of the training population are placed in the training set, and 1 / 3 of the members of the training population are placed in the test set. The formula values of the selected trait combinations are used to construct a decision tree. The ability of the decision tree to correctly classify members in the test set is then determined. In some embodiments, this calculation is performed several times for a given combination of molecular markers. In each iteration of the calculation, members of the training population are randomly assigned to the training set and the test set. The quality of the trait combination is then taken as the average of each such iteration of the decision tree calculation.
[0043] In some embodiments, phenotypic traits and / or genotypic data are used to cluster a training set. For example, consider the case of using 10 genes as described in the present invention. Each member m of the training set has an expression value for each of the 10 genes. Such values from member m in the training population define the following vector: X 1m X 2m X 3m X 4m X 5m X 6m X 7m X 8m X 9m X 10m (In the formula, X im(where is the expression level of the i-th gene in organism m). If there are m organisms in the training set, then the selection of i genes defines m vectors. Note that in the methods of the present invention, it is not necessary for each expression value of every single trait used in the vectors to be represented in every single vector m. In other words, data from a subject in which one of the i-th traits is missing can still be used for clustering. In such cases, the missing expression value is assigned either "0" or some other normalized value. In some embodiments, prior to clustering, the trait expression values are normalized to have a mean of 0 and unit variance.
[0044] The members of the training group that show similar expression patterns across the training group tend to cluster together.A particular combination of traits of the present invention is considered to be a good classifier in this aspect of the present invention if its vector clusters into the trait group found in the training group.For example, if the training group includes patients with good prognosis or poor prognosis, the clustering classifier will cluster the group into two groups, and each group will uniquely represent either good prognosis or poor prognosis.
[0045] Clustering is described in Duda and Hart, Pattern Classification and Scene Analysis, 1973, John Wiley & Sons, Inc., New York, pp. 211-256. As described in section 6.7 of Duda, the clustering problem is described as one of finding natural groupings within a data set. To identify natural groupings, two problems are addressed. First, a method for measuring the similarity (or dissimilarity) between two samples is determined. This metric (similarity measure) is used to ensure that samples in one cluster are more similar to each other than to samples in other clusters. Second, a mechanism for dividing the data into clusters using the similarity measure is determined.
[0046] Similarity measures are described in Section 6.7 of Duda, which states that one way to begin a clustering study is to define a distance function and calculate a matrix of distances between all pairs of samples in a dataset. If distance is a good measure of similarity, the distance between samples in the same cluster will be significantly smaller than the distance between samples in different clusters. However, as described on page 215 of Duda, clustering does not require the use of a distance metric. For example, a non-metric similarity function s(x, x') can be used to compare two vectors x and x'. Traditionally, s(x, x') is a symmetric function that is large if x and x' are somehow "similar." An example of a non-metric similarity function s(x, x') is provided on page 216 of Duda.
[0047] Once a method for measuring "similarity" or "dissimilarity" between points in a data set has been chosen, clustering requires a criterion function that measures the clustering quality of any partition of the data. The partition of the data set that extremizes the criterion function is used to cluster the data. See Duda, page 217. Criterion functions are described in section 6.8 of Duda.
[0048] More recently, Duda et al., Pattern Classification, 2nd Edition, John Wiley & Sons, Inc. New York has been published. Pages 537-563 provide a detailed description of clustering. Further information on clustering techniques can be found in Kaufman and Rousseeuw, 1990, Finding Groups in Data: An Introduction to Cluster Analysis, Wiley, New York, NY; Everitt, 1993, Cluster analysis (3rd Edition), Wiley, New York, NY; and Backer, 1995, Computer-Assisted Reasoning in Cluster Analysis, Prentice Hall, Upper Saddle River, NJ. Specific exemplary clustering techniques that can be used in the present invention include, but are not limited to, hierarchical clustering (agglomerative clustering using nearest neighbor algorithm, nearest neighbor algorithm, average linkage algorithm, centroid algorithm, or sum-of-squares algorithm), k-means clustering, fuzzy k-means clustering algorithm, and Jarvis-Patrick clustering.
[0049] Nearest neighbor classifiers are memory-based and do not require a model to be fitted. Given a query point x0, they find the k training points x0 that are closest to x0 in distance. (r)、 r,...,k are identified and then the point x0 is classified using its k nearest neighbors. Ties can be broken randomly. In some embodiments, Euclidean distance in feature space is used to determine the distance as follows:
number
[0050] Typically, when a nearest neighbor algorithm is used, the expression data used to calculate linear discrimination is standardized to have a mean of 0 and a variance of 1. In the present invention, members of a training population are randomly divided into a training set and a test set. For example, in one embodiment, 2 / 3 of the members of the training population are placed in the training set, and 1 / 3 of the members of the training population are placed in the test set. A profile represents the feature space in which the members of the test set are plotted. Then, the ability of the training set to correctly characterize the members of the test set is calculated. In some embodiments, nearest neighbor calculations are performed several times for a given combination of phenotypic traits. In each iteration of calculation, members of the training population are randomly assigned to a training set and a test set. The quality of the combination of traits is then considered as the average of each such iteration of nearest neighbor calculations.
[0051] The nearest neighbor rule can be refined to address issues of unequal class priors, differential misclassification costs, and feature selection. Many of these refinements involve some form of weighted voting over neighbors. For more information on nearest neighbor analysis, see Duda, Pattern Classification, 2nd ed., 2001, John Wiley & Sons, Inc.; and Hastie, 2001, The Elements of Statistical Learning, Springer, New York.
[0052] The pattern classification and statistical techniques described above are merely examples of the types of models that can be used to build models for classification. It should be understood that any statistical method can be used in accordance with the present invention. Combinations of these methods can also be used. Further details regarding other statistical methods and their implementation can be found in U.S. Patent No. 10,181,009, which is incorporated herein by reference in its entirety.
[0053] It is understood that during the course of treatment, individuals comprising the reference set may drop out before their LRRK2 inhibitor response is determined. It is unknown whether these individuals will ultimately respond to LRRK2 inhibition. Simply excluding these individuals from the reference set would bias the reference data set by excluding the characteristics of individuals with poor prognosis for response. Such bias would result in reporting an overly optimistic probability of response to treatment with an LRRK2 inhibitor.
[0054] In the system and method of the present invention, rather than omitting these subjects on a large scale, the present invention utilizes certain methods of statistical analysis to account for dropouts. For example, the Kaplan-Meier method can be used to censor or exclude data from individuals in the reference set who do not return to treatment. Other forms of statistical analysis can be used in accordance with the present invention to compile the data in the reference set. For example, logistic regression, ordinal logistic regression, Cox proportional hazards regression, and other methods can all be used to compile the data in the reference set. In addition, it is contemplated that the reference set can omit or account for dropouts based on the traits of the individual, rather than making blanket assumptions about the responsiveness of dropouts. For example, rather than simply assuming that dropouts have the same chance of responding as individuals who continue treatment, or that dropouts have no chance of responding, the present invention can evaluate the traits of dropouts and omit dropouts based on such information. In this way, overly optimistic estimates (resulting from the assumption that all dropouts are equally likely to respond) or overly conservative estimates (resulting from the assumption that dropouts have no chance of responding) are avoided.
[0055] In certain aspects, the present invention incorporates the use of artificial censoring to account for dropout. In artificial censoring, participants are censored if they meet a predetermined test criterion, such as exposure to an intervention, non-compliance with a treatment regimen, or the occurrence of a competing outcome. Further analysis methods, such as inverse-probability-of-censoring weights (IPCW), can then be used to determine whether the survival experience of the artificially censored participant would have been the same if they had not been exposed to the intervention, had complied, or had not experienced the competing outcome. In some embodiments, methods that include the use of artificial censoring, as well as the use of IPCW, are encompassed by the present invention to account for dropout in the reference set. Further details regarding the use of artificial censoring and the use of IPCW are described in Howe et al., "Limitation of inverse probability-of-censoring weights in estimating survival in the presence of strong selection bias," Am J Epidemiology, 2011, which is incorporated herein by reference in its entirety.
[0056] Aspects of the invention described herein may be implemented using any type of computing device, such as a computer including a processor, e.g., a central processing unit, or any combination of computing devices, each device performing at least a portion of a process or method. In some embodiments, the systems and methods described herein may be implemented using a handheld device, e.g., a smart tablet, or a smartphone, or a specialized device manufactured for the system.
[0057] The methods of the present invention can be implemented using software, hardware, firmware, hardwiring, or any combination thereof. The features that implement the functionality can also be physically located in various locations, including being distributed such that some of the functionality is implemented in different physical locations (e.g., an imaging device in one room and a host workstation in another room, or imaging devices in separate buildings with, for example, wireless or wired connections).
[0058] Processors suitable for executing a computer program include, by way of example, both general-purpose and special-purpose microprocessors, as well as any one or more processors of any kind of digital computer. Generally, a processor receives instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes one or more mass storage devices, such as magnetic, magneto-optical, or optical disks, for storing data, or is operatively coupled to receive data from or transfer data to them, or both. Information carriers suitable for embodying computer program instructions and data include, by way of example, all forms of non-volatile memory, including semiconductor memory devices (e.g., EPROM, EEPROM, solid-state drives (SSDs), and flash memory devices); magnetic disks (e.g., internal or removable hard disks); magneto-optical disks; and optical disks (e.g., CD and DVD disks). The processor and memory may be supplemented by, or incorporated in, special-purpose logic circuitry.
[0059] To provide for user interaction, the subject matter described herein can be implemented on a computer having I / O devices such as a CRT, LCD, LED, or projection device for displaying information to the user, and input or output devices such as a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide for user interaction. For example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0060] The subject matter described herein can be implemented in a computing system including a back-end component (e.g., a data server), a middleware component (e.g., an application server), or a front-end component (e.g., a client computer having a graphical user interface or web browser through which a user can interact with an implementation of the subject matter described herein), or any combination of such back-end, middleware, and front-end components. The components of the system can be interconnected via a network by any form or medium of digital data communication, e.g., a communication network. For example, a reference dataset may be stored in a remote location, and the computer communicates via the network to access the reference set and compare data derived from the subject with the reference set. However, in other embodiments, the reference set is stored locally within the computer, and the computer accesses the reference set in the CPU to compare the subject data with the reference set. Examples of communication networks include a cellular network (e.g., 3G or 4G), a local area network (LAN), and a wide area network (WAN), e.g., the Internet.
[0061] The subject matter described herein may be implemented as one or more computer program products, such as one or more computer programs tangibly embodied in an information carrier (e.g., in a non-transitory computer-readable medium) for execution by or control the operation of a data processing device (e.g., a programmable processor, computer, or multiple computers). The computer programs (also known as programs, software, software applications, applications, macros, or code) can be written in any form of programming language, including compiled or interpreted languages (e.g., C, C++, Perl), and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. The systems and methods of the present invention may include instructions written in any suitable programming language known in the art, including, but not limited to, C, C++, Perl, Java, ActiveX, HTML5, Visual Basic, or JavaScript.
[0062] A computer program does not necessarily correspond to a file. A program can be stored in a file or part of a file that holds other programs or data, in a single file dedicated to the program in question, or in multiple cooperating files (e.g., files storing one or more modules, subprograms, or portions of code). A computer program can be deployed to be executed on one computer, on multiple computers at one site, or distributed across multiple sites and interconnected by a communications network.
[0063] A file may be, for example, a digital file stored on a hard drive, SSD, CD, or other tangible, non-transitory medium. A file may be transmitted from one device to another over a network (e.g., from a server to a client, as packets transmitted via, for example, a network interface card, modem, wireless card, or the like).
[0064] Writing a file in accordance with the present invention involves transforming a tangible, non-transitory computer-readable medium, for example, by adding, removing, or rearranging particles (e.g., using net charge or dipole moment to a magnetization pattern by a read / write head), where the pattern represents a new collocation of information about a target physical phenomenon desired and useful to the user. In some embodiments, writing involves a physical transformation of material in the tangible, non-transitory computer-readable medium (e.g., having specific optical properties so that an optical read / write device can subsequently read the new, useful collocation of information (e.g., burned to a CD-ROM)). In some embodiments, writing a file involves transforming a physical flash memory device, such as a NAND flash memory device, and storing information by transforming physical elements within an array of memory cells comprised of floating-gate transistors. Methods of writing files are well known in the art and can be invoked manually or automatically, for example, by a program, by a save command from software, or by a write command from a programming language.
[0065] Suitable computing devices typically include mass memory, at least one graphical user interface, at least one display device, and typically include communication between devices. Mass memory refers to a type of computer-readable medium, i.e., computer storage media. Computer storage media may include volatile, nonvolatile, removable, and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include RAM, ROM, EEPROM, flash memory, or other memory technology, CD-ROM, digital versatile disks (DVDs), or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage, or other magnetic storage devices, radio frequency identification tags or chips, or any other medium that can be used to store desired information and that can be accessed by a computing device.
[0066] As those skilled in the art will recognize as necessary or best suited for practicing the methods of the present invention, a computer system or machine of the present invention includes one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), or both), a main memory, and a static memory that communicate with each other via a bus.
[0067] The methods of the present invention may utilize machine learning systems, which may learn, for example, in a supervised, unsupervised, semi-supervised, or reinforcement learning manner.
[0068] In unsupervised or autonomous models, a machine learning system is provided with only input training data without paired output data to autonomously identify patterns. Unsupervised models identify underlying patterns or structures in the training data to make predictions for the test data. Unsupervised models are advantageous for clustering data, detecting anomalies, and independently discovering rules in the data. The accuracy of unsupervised models is more difficult to evaluate because there is no predetermined output variable that the system is optimizing for. Autonomous models may employ both supervised and unsupervised learning techniques to optimize predictions. Unsupervised models are advantageous for training machine learning systems to cluster data into clusters when labeled training data is not available. Unsupervised models may use principal component analysis (PCA) and uniform manifold approximation and projection (UMAP). Discriminant analysis may also be used when the groups in the training and test data are already known. Discriminant analysis may include linear discriminant analysis (LDA) and quadratic discriminant analysis (QDA).
[0069] In a semi-supervised model, a machine learning system is provided with training data including input variables, and output variable pairs are available for only a limited pool of input variables. The model uses the input variables with output variable pairs and the remaining input training data to learn patterns and make inferences to generate predictions about previously unseen test data. Semi-supervised models may advantageously query a user for additional paired output data based on unpaired data. Semi-supervised models are advantageous for training machine learning systems when only an incomplete training dataset is available.
[0070] In reinforcement learning models, the machine learning system is not given input or output variables. Rather, the model provides a "reward" term and then attempts to maximize the cumulative reward term through trial and error. Reinforcement learning models are Markov decision processes. Supervised, unsupervised, semi-supervised, and reinforcement models are described in Jordan and Mirchell, 2015, Machine learning, Trends, perspectives, and prospects, Science 349(6245):255-260, which is incorporated by reference.
[0071] An example of a supervised learning model is a "decision tree." A decision tree is a nonparametric supervised learning model that uses simple decision rules to infer a classification of test data from features in the test data. In classification trees, the test data takes on a finite set of discrete values or classes, while in regression trees, the test data can take on continuous values such as real numbers. Decision trees have several advantages in that they are easy to understand and can be visualized as a tree that begins with a root (usually a single node) and repeatedly branches into leaves (multiple nodes) associated with classifications. See Criminisi, 2012, "Decision Forests: A unified framework for classification, regression, density estimation, manifold learning and semi-supervised learning," Foundations and Trends in Computer Graphics and Vision 7(2-3):81-227, incorporated by reference.
[0072] Another supervised learning model is the "support vector machine" (SVM), "support vector network" (SVN), or support vector classifier (SVC), which is a supervised learning model for classification and regression problems. When used to classify new data into one of two categories, an SVM creates a hyperplane in a multidimensional space that separates the data points into one category or the other. While the original problem can be expressed in terms requiring only a finite-dimensional space, linear separation of data between categories may not be possible in a finite-dimensional space. As a result, the multidimensional space is selected to allow for the construction of a hyperplane that results in a clear separation of the data points. See Press, W.H. et al., Section 16.5. Support Vector Machines. Numerical Recipes: The Art of Scientific Computing (3rd ed.). New York: Cambridge University (2007), incorporated by reference. When output variable pairs are not available for the input variables in the training data, an SVM can be designed as an unsupervised or semi-supervised learning model using support vector clustering. See Ben-Hur, 2001, Support Vector Clustering, J Mach Learning Res 2:125-137, which is incorporated by reference. SVM models can be advantageous for machine learning systems where the test data falls into a limited number of possible categories. Furthermore, SVM models can be advantageous when there is only a limited set of training data available to the machine learning system.
[0073] Logistic regression analysis is another statistical process that can be used by machine learning systems to find patterns in training and test data to make predictions. It involves techniques for modeling and analyzing the relationships between multiple variables. Specifically, regression analysis focuses on the change in a dependent variable in response to changes in a single independent variable. Regression analysis can be used to estimate the conditional expectation of the dependent variable given the independent variable. The variation of the dependent variable can be characterized around a regression function and described by a probability distribution. The parameters of the regression model can be estimated using, for example, least squares, Bayesian methods, percent regression, least absolute deviation, nonparametric regression, or distance metric learning. Regression models also offer the advantage of being effectively implemented by a variety of tools, and the model can be easily updated to identify new particles.
[0074] SVM and logistic regression systems may use stochastic gradient descent (SGD) techniques to fit data, which is advantageous for optimizing machine learning systems.
[0075] Bayesian algorithms can also be used to discover patterns in training and test data to make predictions. A Bayesian network is a probabilistic graphical model that represents a set of random variables and their conditional dependencies via a directed acyclic graph (DAG). The DAG has nodes that represent random variables, which can be observables, latent variables, unknown parameters, or hypotheses. Edges represent conditional dependencies; unconnected nodes represent variables that are conditionally independent of each other. Each is associated with a probability function that takes as input a specific set of values for the node's parent variables and gives (as output) the probability (or probability distribution, if applicable) of the variable represented by the node. Bayesian models generally offer the advantage of requiring less training data than other models.
[0076] Some models may rely on clustering training and test data to find patterns and make predictions. The "k-nearest neighbor" (k-NN) model is a supervised, nonparametric learning model for classification and regression problems. The k-nearest neighbor model assumes that similar data exists in close proximity and assigns a category or value to each data point based on its k nearest neighbors. The k-NN model can be advantageous when the data has few outliers and can be defined by homogeneous features. Furthermore, the k-NN model offers the advantage of continuously learning from the test data and does not require a training period before identifying material from the training data.
[0077] An example of an unsupervised learning model using clustering is the "k-means" clustering model. The k-means model aims to discover clusters of data in input data and test data. The k-means model is advantageous when a defined number of clusters are known to exist in the data, and also when the test data has few outliers and uniform features can be defined. Additional models for clustering training data include, for example, farthest-neighbor, centroid, sum of squares, fuzzy k-means, and Jarvis-Patrick clustering. K-means and other unsupervised clustering models are advantageous when training data is unavailable or limited.
[0078] A trained machine learning model can be a "stable learner." A stable learner is a model that is less sensitive to perturbations in predictions based on new training data. A stable learner can be advantageous when test data is stable, but may be less advantageous when the system must continually improve its performance to accurately predict new test data that may be less stable. Thus, a stable learning model can be advantageous for use by a machine learning system when the types of data that may be introduced are known and unlikely to change.
[0079] Several machine learning system types can be combined into a final predictive model known as an ensemble. Ensembles can be divided into two types: homogenous ensembles and heterogeneous ensembles. Homogeneous ensembles combine multiple machine learning models of the same type. Heterogeneous ensembles combine multiple machine learning models of different types. Ensembles can provide advantages because they can be more accurate than any of the individual base member models (“members”) within the ensemble. The number of members combined in an ensemble can affect the accuracy of the final prediction. Therefore, it is advantageous to determine the optimal number of members when designing an ensemble system to be used by a machine learning system.
[0080] Ensembles used by machine learning systems may combine or aggregate outputs from individual members by using "voting"-type methods for classification systems and "averaging"-type methods for regression systems. In "majority voting" methods, each member makes a prediction on test data, and the prediction that receives a majority of votes is the final output of the ensemble. If none of the predictions receives more than half of the votes, it may be determined that the ensemble is unable to make stable predictions. In "plurality voting" methods, the most voted prediction may be considered the final output of the ensemble, even if it receives less than half of the votes. In "weighted voting" methods, the votes of more accurate members are multiplied by a weight assigned to each member based on its accuracy. In "simple average" methods, each member makes a prediction on test data, and the average of the outputs is calculated. This method may be advantageous for reducing overfitting and creating a smoother regression model. In "weighted average" methods, the predicted output of each member is multiplied by a weight assigned to each member based on its accuracy. Voting, averaging, and weighting methods can be combined to improve the accuracy of the ensemble used by the machine learning system.
[0081] Members in an ensemble used by a machine learning system can be trained independently, or new members can be trained using information from previously trained members. In a "parallel ensemble," the ensemble attempts to provide greater accuracy than the individual members by exploiting the independence between members, for example, by training multiple members simultaneously to identify and aggregate the outputs from the members. In a "sequential ensemble system," the ensemble attempts to provide greater accuracy than the individual members by exploiting the dependency between members, for example, by using information from a first member regarding the identification of data to improve the training of a second member to identify data and weight the outputs from the members.
[0082] The overall accuracy of an ensemble used by a machine learning system can be optimized by using ensemble meta-algorithms, such as "bagging" algorithms to reduce variance, "boosting" algorithms to reduce bias, or "stacking" algorithms to improve predictions.
[0083] Boosting algorithms can be used to reduce bias and improve less accurate or "weakly learned" models. A member can be considered a "weakly learned" model if it has a substantial error rate but its performance is non-random. Boosting algorithms gradually build an ensemble by sequentially training each member on the same training dataset, examining the prediction error of the test data, and assigning weights to the training data based on the member's difficulty in making accurate predictions. With each successive member trained, the algorithm emphasizes training data that previous members found difficult. Members are then weighted based on the accuracy of their prediction output, taking into account the weights applied to the training data. Predictions from each member may be combined using weighted voting or weighted averaging methods. Boosting algorithms are advantageous when combining multiple weakly learned models. However, boosting algorithms can result in overfitting the test data to the training data. Examples of boosting algorithms include AdaBoost, gradient boosting, and eXtreme Gradient Boost (XGBoost). See Freund, 1997, A decision-theoretic generalization of online learning and an application to boosting, J Comp Sys Sci 55:119; and Chen, 2016, XGBoost: A Scalable Tree Boosting System, arXiv:1603.02754, both of which are incorporated by reference.
[0084] Bagging or "bootstrap aggregation" algorithms reduce variance by averaging together multiple estimates from members. Bagging algorithms provide each member with a random subsample of the full training data set, with each random subsample known as a "bootstrap" sample. In a bootstrap sample, some data from the training data set may appear multiple times, and some data from the training data set may be absent. Because the subsamples can be generated independently of each other, training can occur in parallel. Predictions of the test data from each member are then aggregated, such as by voting or averaging methods.
[0085] One example of a bagging algorithm that can be used by machine learning systems is "random forest." In random forests, an ensemble combines multiple randomized decision tree models. Each decision tree model is trained from bootstrap samples from a training set for test data. The training set itself may be a random subset of features from a larger training set. By providing a random subset of the larger training set at each split in the learning process, spurious correlations that may arise from the presence of individual features that are strong predictors of the output variable are reduced. Averaging predictions of the test data reduces the variance of the ensemble and improves predictions of the test data. Random forests can be autonomous models and can include both supervised and unsupervised learning periods. Because stable learning systems tend to provide generalized outputs with less variability across bootstrap samples, bagging may be less advantageous in optimizing ensembles that combine stable learning systems. Random forests are advantageous for use by machine learning systems to identify data by providing a high degree of versatility in identifying test data and reducing false identifications by the machine learning system. See Breiman, 2001, Random Forests, Machine Learning 45:5-32, which is incorporated by reference.
[0086] Stacking algorithms, or "stacked generalization" algorithms, improve predictions by combining and building ensembles using meta-machine learning models. In stacking algorithms, base member models are trained on a training dataset and generate a new dataset as output. This new dataset is then used as a training dataset for the meta-machine learning model to build the ensemble. Stacking algorithms are generally advantageous for use by machine learning systems to identify test data when building heterogeneous ensembles. Ensembles are described in Villaverde et al., 2019, "On the adaptability of ensemble methods for distribution classification systems: A comparative analysis," International Journal of Distributed Sensor Networks 15(7); and Heitor et al., 2017, "A Survey of Ensemble Learning for Data Stream Classification," 50(2):Art. 23, each of which is incorporated by reference.
[0087] Neural networks modeled on the human brain enable information processing and machine learning. Neural networks include nodes that mimic the functions of individual neurons, and the nodes are organized into layers. A neural network includes an input layer, an output layer, and one or more hidden layers that define connections from the input layer to the output layer. The systems and methods of the present invention may include any neural network that facilitates machine learning. The system may include known neural network architectures such as GoogLeNet (Szegedy et al., Going deeper with convolutions, in CVPR2015, 2015); AlexNet (Krizhevsky et al., Imagenet classification with deep convolutional neural networks, Pereira et al., Eds., Advances in Neural Information Processing Systems 25, pp. 1097-3105, Curran Associates, Inc., 2012); VGG16 (Simonyan and Zisserman, Very deep convolutional networks for large-scale image recognition, CoRR, abs / 3409.1556, 2014); or FaceNet (Wang et al., Face Search at Scale: 80 Million Gallery, 2015) (each of the above references is incorporated by reference). An advantage of using machine learning systems based on neural network architectures is that neural networks can learn patterns and correlations by themselves and generate outputs that are not limited by the training data provided to them.
[0088] Deep learning neural networks (also known as deep structured learning, hierarchical learning, or deep machine learning) comprise a class of machine learning operations that can be used by classifiers, which use a cascade of many layers of nonlinear processing units for feature extraction and transformation. Each successive layer uses the output from the previous layer as input. Algorithms can be supervised or unsupervised, and applications include pattern analysis (unsupervised) and classification (supervised). Certain embodiments are based on unsupervised learning of multiple levels of features or representations of data. Higher-level features are derived from lower-level features to form hierarchical representations. Deep learning with neural networks involves learning multiple levels of representations corresponding to different levels of abstraction, where the levels form a hierarchy of concepts. In some embodiments, neural networks include at least five, and preferably more than ten, hidden layers. The many layers between the input and output allow the system to operate through multiple processing layers.
[0089] In neural networks used by machine learning systems, nodes are connected in layers, and signals travel from the input layer to the output layer. Each node in the input layer may correspond to a feature from the training data. The hidden layer nodes are calculated as a function of a bias term and a weighted sum of the input layer nodes, and each connection between an input layer node and a hidden layer node is assigned a respective weight. The bias terms and weights between the input layer and the hidden layer are advantageously learned autonomously during neural network training. Networks can contain thousands or millions of nodes and connections. Typically, the signals and states of artificial neurons are real numbers, typically between 0 and 1. Optionally, each connection and unit itself may have a threshold or limit function so that a signal must exceed a limit before propagating. Backpropagation is the use of forward stimuli to modify connection weights, sometimes performed to train a network using known correct outputs. See WO 2016 / 182551, U.S. Patent Application Publication No. 2016 / 0174902, U.S. Patent No. 8,639,043 and U.S. Patent Application Publication No. 2017 / 0053398, each of which is incorporated herein by reference.
[0090] Features from test or training data can be represented by a deep learning network in many ways, such as a vector of intensity values for each pixel in an image, or in a more abstract way, such as a set of edges, regions of a particular shape, etc. These features are represented by nodes in the network. Preferably, each feature is structured as a numerical feature or vector representing the image feature. This provides, for example, a numerical representation of an object from an image, as such a representation facilitates processing and statistical analysis. Numerical features are often combined with weights using dot products to construct a linear prediction function that is used to determine a score to make a prediction.
[0091] The vector space associated with these feature vectors may be referred to as the feature space. Dimensionality reduction may be employed by the network used by the classifier to reduce the dimensionality of the feature space. Higher-level features may be obtained from already available features and added to the feature vector in a process called feature construction. Feature construction is the application of a set of constructive operators to a set of existing features that results in the construction of new features. For example, a machine learning system based on a neural network architecture may be provided with image data from an image sensor. Early layers of the neural network may identify horizontal and vertical lines in the image data. Later layers in the network may then use the identified lines to obtain edges, which are higher-level features for particles in the image.
[0092] The deep learning neural network may be a multi-layer perceptron (MLP), a convolutional neural network (CNN), or a recurrent neural network (RNN).
[0093] FIG. 1 provides total urinary di-18:1-BMP levels in different patient population sets.
[0094] FIG. 2 provides total urinary di-22:6-BMP levels in different patient population sets.
[0095] Figure 3 provides total urinary 2,2'di-22:6-BMP levels in different patient population sets. As is evident from the data, total BMP levels are a good indicator of the efficacy of LRRK2 inhibitor treatment.
[0096] Figure 4 provides the distribution of PD patients and their LRRK2 activity. The data provided in Figure 4 shows that the LRRK2 of PD patients can be classified as LRRK2 active, LRRK2 rare-variant, and LRRK2 normal.
[0097] Figure 5 provides data on total urinary di-22:6-BMP concentrations in patients with LRRK2-normal, LRRK2-predicted, and LRRK2-rare variants. The data demonstrate that urinary di-22:6-BMP levels are a good surrogate for predicting LRRK2 pathway activity. LRRK2-predicted patients identified by the methods provided herein exhibited LRRK2 レア-バリアント Patients and LRRK2 正常 Patients with consistent urinary BMP levels may be more responsive to treatment with an LRRK2 inhibitor.
[0098] Figures 6A and 6B show LRRK2 in idiopathic Parkinson's disease. 活性 We demonstrate that the fraction is stable across different cohorts of Accelerating Medicines Partnership-PD (AMP-PD) patients. Therefore, the methods provided herein for isolating patient populations are important for identifying patients who are more likely to respond to treatment with LRRK2 inhibitors.
[0099] Identifying gene modifiers from genetic data The present invention provides a method for determining whether PD patients with wild-type LRRK2 are more likely to respond to LRRK2 inhibitors by using genetic modifiers of LRRK2 in the patient's genome as an indicator. The present invention recognizes that genetic modifiers of LRRK2 can cause changes in the level or activity of LRRK2 kinase, such as an increase or decrease, or can alter the LRRK2 signaling pathway through upstream or downstream regulators, thereby contributing to PD pathogenesis. As a result, PD patients with one or more such modifiers may benefit from drug therapy using an LRRK2 inhibitor, despite having an LRRK2 allele that produces a normal form of the kinase. Therefore, genetic modifiers of LRRK2 activity serve as an indicator for determining whether LRRK2 inhibitor treatment is appropriate for a given individual. The method of the present invention is useful for both identifying PD patients as candidates for LRRK2 inhibitor treatment and treating such patients.
[0100] In one aspect, the present invention provides a method of treating a subject having Parkinson's disease associated with wild-type LRRK2 by providing an LRRK2 inhibitor to a subject exhibiting Parkinson's disease and having wild-type LRRK2 and genetic modifiers of wild-type LRRK2, such that the subject responds to the LRRK2 inhibitor, thereby treating Parkinson's disease associated with wild-type LRRK2 in the subject.
[0101] Genetic data can include any type of data regarding the composition and / or expression of one or more genes in a subject. Genetic data can include one or more of exome sequence, genome sequence, genotype sequence, proteome sequence, and transcriptome data.
[0102] A genetic modifier can be any genetic element that alters LRRK2 expression or activity, correlates with changes in activity, or causes changes in protein levels associated with disease burden (whether increased or decreased). Genetic modifiers can increase or decrease LRRK2 expression and / or activity; genetic modifiers can also increase or decrease LRRK2 degradation. Genetic modifiers can be amplifications, deletions, duplications, fusions, insertions, inversions, rearrangements, single nucleotide polymorphisms (SNPs), substitutions, or translocations. Genetic modifiers can be within coding or non-coding regions within the subject's genome. Genetic modifiers can be associated with family history and genetically confirmed Ashkenazi status.
[0103] In certain embodiments, the genetic modifier can be any of the genetic modifiers provided in PCT / US2021 / 056443, which is incorporated by reference in its entirety. In certain embodiments, the SNP can be any of the SNPs described in PCT / US2021 / 056443, which is incorporated by reference in its entirety.
[0104] <h2 style=";text-align:left;direction:ltr">SNPは、rs10784722、rs10877877、rs10879122、rs11181542、rs113111234、r s113736300、rs12230765、rs12816484、rs12829831、rs13377670、rs14155 1396、rs144377852、rs149173058、rs17580794、rs17621741、rs1838354、r s184120094、rs188535877、rs188583486、rs188604552、rs189517205、rs20 0611801、rs200907772、rs201889643、rs201944175、rs2406426、rs240686 0、rs285561、rs34566033、rs368141132、rs369084695、rs371700002、rs371 905892, rs373439540, rs376468815, rs377104202, rs377627337, rs384234, rs61920964, rs6581941, rs6650226, rs71078241, rs7304080, rs7308892 6、rs74434364、rs74842215、rs75043969、rs78468120、rs7960429、rs7979 420、rs76904798、rs57025360、rs112515153、rs10877877、rs10784722、rs4 272849、rs2404832、rs117534366、rs1838343、rs10880342、rs11177660、r s183028452、rs116912628、rs147755361、rs11584630、rs3793397、rs11179 4893、rs4931640、rs526507、rs79307177、rs187116363、rs71609573、rs74 390551、rs144665441、rs1718880、rs1991401、rs11052225、rs145801597、r s72907976、rs147286120、rs378690、rs73188365、rs610037、rs75479531、 rs1112191556、rs308303、rs10790282、rs3729912、rs4326638、rs4414548、The LRRK2 inhibitor may be rs13009437, rs56045011, rs6858566, rs4425, rs11052253, or any other SNP in linkage disequilibrium (LD) with these SNPs that would be suitable as a surrogate for these SNPs. The LRRK2 inhibitor may be any of the inhibitors listed in this application.
[0105] In certain embodiments, the SNPs are rs33939927, rs35801418, rs34805604, rs34637584, rs35870237, rs34995376, rs34778348, rs11611119, rs6581439, rs12296462, rs549790, rs12423473, rs555740, rs2242367, rs7295598, rs2253736, rs11564274, rs2708419, rs17519419, rs76904798, rs17519573, rs11175847, rs10878452, rs17444612, rs117929583, rs11564235, rs17128233, rs 7960976, rs1918942, rs611829, rs7531501, rs9793102, rs3755541, rs7578955, rs17738103, rs4676776, rs1259475, rs10 477505, rs7380062, rs4636028, rs12704998, rs2768282, rs10283642, rs3750779, rs1362993, rs7089200, rs9783486, rs 4763946, rs16920645, rs12312400, rs2708494, rs12813279, rs11613339, rs7300813, rs166806, rs3794253, rs7960429, r rs7955116, rs10491998, rs12814145, rs2708078, rs1922761, rs1373422, rs17691793, rs7301498, rs7967809, rs1866074, rs10507535, rs9300705, rs1642819, rs8048361, rs1152838, rs2331796, rs10515969, rs6566942, or rs2422956.
[0106] In another aspect, the present invention provides a method for determining whether a subject with Parkinson's disease associated with wild-type LRRK2 will respond to an LRRK2 inhibitor. The method includes: performing an assay on a sample from the subject with Parkinson's disease associated with wild-type LRRK2 to obtain genetic data from the subject; generating a report identifying one or more genetic modifiers of LRRK2 in the genetic data, wherein the one or more genetic modifiers in the LRRK2 network indicate that the subject with Parkinson's disease associated with wild-type LRRK2 will be responsive to an LRRK2 inhibitor; and providing the report to a physician so that the physician can prescribe or provide the LRRK2 inhibitor to the subject. The genetic data can be any of the types of genetic data described above. The genetic modifier can be any of the types of genetic modifiers of LRRK2 described above. The genetic modifier can be any of the SNPs described above. The LRRK2 inhibitor can be any of the inhibitors listed in this application.
[0107] In another aspect, the present invention provides a method for treating a subject with PD associated with wild-type LRRK2. The method includes receiving genetic data identifying one or more genetic modifiers of LRRK2, where the one or more genetic modifiers indicate that the subject with Parkinson's disease associated with wild-type LRRK2 will be responsive to an LRRK2 inhibitor, and prescribing or providing an LRRK2 inhibitor to the subject. The genetic data can be any type of genetic data described above. The genetic modifier can be any type of genetic modifier of LRRK2 described above. The genetic modifier can be any of the SNPs described above.
[0108] The present invention also recognizes that genetic modifiers of LRRK2 serve as indicators that PD patients with wild-type LRRK2 may benefit from drug therapy using one or more LRRK2 inhibitors. Genetic modifiers of LRRK2 can be one or more genetic elements (e.g., a single genetic element alone or any combination(s) of genetic elements) that operatively modify LRRK2 (e.g., wild-type LRRK2), such as altering the expression, degradation, localization (e.g., within a cell or across cell types), binding, or activity of LRRK2 in a subject, including the LRRK2 gene, transcripts of the LRRK2 gene, and polypeptide products of the LRRK2 gene. For example, but not limited to, genetic modifiers can alter, e.g., increase or decrease, the expression, activity, stability, binding, localization, degradation, transcription, or translation of LRRK2, including the LRRK2 gene, transcripts of the LRRK2 gene, and polypeptide products of the LRRK2 gene. In certain embodiments, genetic modifiers of LRRK2 can be structural variations in the subject's genome. For example, but not limited to, a genetic modifier can be an amplification, deletion, duplication, fusion, insertion, inversion, rearrangement, single nucleotide polymorphism (SNP), substitution, or translocation. SNPs that can be genetic modifiers of LRRK2 are listed in Example 1. In addition, any other SNPs in linkage disequilibrium (LD) with the SNPs listed in Example 1 can be used as genetic modifiers. A genetic modifier can be a cis-regulatory element such as a promoter, enhancer, silencer, or operator. A cis-regulatory element can regulate the binding of one or more proteins to DNA adjacent to LRRK2. A cis-regulatory element can affect the binding of histones, transcription factors, initiation factors, helicases, polymerases, or components of any of the aforementioned proteins. A genetic modifier can be a trans-acting factor. A trans-acting factor can affect the transcription or translation of LRRK2. A genetic modifier can be present in any region of a subject's genome. A genetic modifier can be present in a coding or non-coding region of a subject's genome. The coding region may be within LRRK2 or another gene.A genetic modifier can be present within the LRRK2 coding region but does not alter the sequence of the LRRK2 polypeptide, the size of the LRRK2 polypeptide, or both.
[0109] The methods of the present invention may include identifying or analyzing one or more genetic modifiers of LRRK2 in genetic data obtained from a subject. The genetic data may include any type of data regarding the composition and / or expression of one or more genes in a subject. The genetic data may include one or more exome sequences, genome sequences, genotype sequences, proteome sequences, and transcriptome data. The genetic data may include data regarding one or more genes known to be associated with PD, such as any of the genes listed above.
[0110] Genetic modifiers can be identified from genetic data using any suitable method. In some embodiments, genetic data collected from a subject is compared to a reference set of data to provide a probability of responsiveness to an LRRK2 inhibitor. The reference set can include data collected from individuals without PD. Phenotypic data from the subject and reference individuals can also be used. Phenotypic data can include traits associated with PD, including PD symptoms or PD risk factors (such as those described above). Data can include outcomes such as whether the individual responded to LRRK2 inhibitor treatment.
[0111] The present invention provides methods and systems for predicting a subject's responsiveness to an LRRK2 inhibitor based on the subject's phenotypic traits and / or genotypic data. In some embodiments, the methods and systems of the present invention use a diagnostic signature to predict responsiveness. The diagnostic predictor can be based on any suitable pattern recognition method that receives input data representing multiple responsiveness-associated phenotypic traits, such as (1) LRRK2-like expression of PD observed in carriers of LRRK2 deleterious variants, (2) PD of apparently unknown mechanism, and (3) molecular signatures of appropriate controls, and provides an output indicating the probability that the subject will respond to an LRRK2 inhibitor. The diagnostic predictor can be trained with data from multiple individuals whose phenotypic traits, medical interventions, and LRRK2 inhibitor response outcomes are known. The multiple individuals used to train the diagnostic predictor are also known as a training population. For each individual in the training population, the training data includes (a) data representing multiple phenotypic traits, (b) medical interventions, and (c) LRRK2 inhibitor response information. LRRK2 inhibitor response outcomes may not be required to generate a diagnostic signature. LRRK2 inhibitor response can be evaluated in a prospectively selected patient population.Various diagnostic predictors that can be used in conjunction with the present invention are described below.In some embodiments, additional individuals with known trait profiles and LRRK2 response outcomes can be used to test the accuracy of diagnostic predictors obtained using the training population.Such additional patients are known as the test population.
[0112] In certain embodiments, the methods of the present invention use a diagnostic predictor, also called a classifier, to determine the probability of responding to LRRK2 inhibition. As described above, the diagnostic predictor can be based on any suitable pattern recognition method that receives a profile, such as a profile based on multiple phenotypic traits, and provides an output containing data indicating whether a patient is likely or unlikely to respond to an LRRK2 inhibitor, and may include the potential risks and benefits of treatment with such an inhibitor. The profile can be obtained by completing a questionnaire containing questions about specific phenotypic traits, or by collecting a biological sample to obtain genotypic data or a combination thereof. The diagnostic predictor is trained with training data from a training population of individuals with known phenotypic traits, medical interventions, and LRRK2 inhibitor response outcomes.
[0113] Diagnostic predictors based on any of these methods can be constructed using the profiles and diagnostic data of training patients.Then, these diagnostic predictors can be used to predict the response of subjects to LRRK2 inhibitors based on the profiles of phenotype traits, genotype traits, or both.This method can also be used to identify the traits that distinguish between responding and not responding to LRRK2 inhibition using the trait profiles and diagnostic data of training population.
[0114] In one embodiment, a diagnostic predictor can be prepared by (a) generating a reference set of individuals with known phenotypic traits, medical interventions, and LRRK2 response outcomes; (b) for each trait, determining a metric of correlation between the trait and LRRK2 response outcomes in a plurality of individuals with known LRRK2 response outcomes at a given time point; (c) selecting one or more traits based on the level of association; and (d) training a diagnostic predictor, wherein the diagnostic predictor receives data representing the traits selected in the previous step and provides an output indicative of the probability of responding to LRRK2 inhibition, and the training data from the reference set of subjects includes assessments of the traits obtained from individuals.
[0115] A variety of known statistical pattern recognition methods can be used in conjunction with the present invention, and are described in detail above.
[0116] Assays for obtaining genetic data Identifying or analyzing one or more genetic modifiers of LRRK2 may involve performing an assay on a sample obtained from a subject. The sample may be any type of sample containing genetic material, such as DNA or RNA. For example, but not limited to, the sample may be from amniotic fluid, biopsy, blood, body fluid, cells, cerebrospinal fluid, lymph, mouthwash, needle aspiration biopsy, hair, phlegm, plasma, pus, saliva, semen, serum, sputum, stool, swab, sweat, synovial fluid, tears, tissue, urine, or any combination of the above samples. For example, but not limited to, the tissue sample may be derived from bone marrow tissue, CNS tissue, eye tissue, gastrointestinal tissue, urogenital tissue, hair, kidney tissue, liver tissue, mammary gland tissue, musculoskeletal tissue, nail, nasal passage tissue, nervous tissue, placenta tissue, or skin tissue.
[0117] The subject can be any type of subject. The subject can be human. The subject can exhibit one or more symptoms of Parkinson's disease, or the subject can be asymptomatic. The patient can be associated with a PD patient. The subject can be a pediatric patient, a newborn, a neonate, an infant, a child, an adolescent, a preteen, a teenager, an adult, or an elderly subject. The subject can exhibit one or more symptoms of Parkinson's disease, or the subject can be asymptomatic. The patient can be associated with a PD patient.
[0118] Genetic analysis methods are known in the art.In certain embodiments, the known single nucleotide polymorphism at specific position can be detected by the single base extension of the primer that binds to the sample DNA adjacent to this position, as described in, for example, United States Patent No. 6,566,101, the entire contents of which are incorporated herein by reference.In other embodiments, the hybridization probe that overlaps with the SNP of interest and selectively hybridizes to the sample nucleic acid that comprises specific nucleotide at this position can be used, as described in, for example, United States Patent No. 6,214,558 and United States Patent No. 6,300,077, the entire contents of which are incorporated herein by reference.
[0119] In certain embodiments, nucleic acid is sequenced to detect variants (i.e., mutations) in nucleic acid compared with wild type and / or non-mutated form of sequence.Nucleic acid can comprise multiple nucleic acids derived from multiple genetic elements.Methods for detecting sequence variants are known in the art, and sequence variants can be detected by any sequencing method known in the art, such as ensemble sequencing or single molecule sequencing.
[0120] Sequencing can be by any method known in the art.DNA sequencing techniques include classical dideoxy sequencing reaction (Sanger method) using labeled terminators or primers and gel separation in slabs or capillaries, sequencing by synthesis using reversibly terminated labeled nucleotides, pyrosequencing, 454 sequencing, allele-specific hybridization of labeled oligonucleotide probes to a library, sequencing by synthesis using allele-specific hybridization of labeled clones to a library followed by ligation, real-time monitoring of the incorporation of labeled nucleotides during the polymerization process, polony sequencing, and SOLiD sequencing.Recently, sequencing of separated molecules has been demonstrated by sequential or single extension reaction using polymerase or ligase, and single or sequential differential hybridization with a library of probes.
[0121] One traditional method of performing sequencing is by chain termination and gel separation, as described, for example, in Sanger et al., Proc.Natl.Acad.Sci.USA, 74(12):5463 67(1977).Another traditional sequencing method involves chemical decomposition of nucleic acid fragments, as described, for example, in Maxam et al., Proc.Natl.Acad.Sci., 74:560 564(1977).Finally, a method based on sequencing by hybridization has been developed, as described, for example, in US Patent Application Publication No. 2009 / 0156412.The contents of each reference are incorporated herein by reference in their entirety.
[0122] Sequencing techniques that can be used in the methods of the provided invention include, for example, Harris TD et al., Single-Molecule DNA Sequencing of a Viral Genome, (2008) Science 320:106-109. True single-molecule sequencing (tSMS) techniques involve cleaving a DNA sample into strands of approximately 100-200 nucleotides, and adding a polyA sequence to the 3' end of each DNA strand. Each strand is labeled by the addition of a fluorescently labeled adenosine nucleotide. The DNA strands are then hybridized to a flow cell containing millions of oligo-T capture sites immobilized on the flow cell surface. Templates are sequenced at approximately 100 million templates / cm. 2 The flow cell is then loaded into an instrument, such as a HeliScope™ sequencer, where a laser illuminates the surface of the flow cell, revealing the location of each template. A CCD camera can map the location of the templates on the flow cell surface. The template fluorescent label is then cleaved and washed away. The sequencing reaction is initiated by introducing DNA polymerase and fluorescently labeled nucleotides. Oligo-T nucleic acids serve as primers. The polymerase incorporates the labeled nucleotides into the primers in a template-directed manner. The polymerase and unincorporated nucleotides are removed. Templates with directional incorporation of the fluorescently labeled nucleotides are detected by imaging the flow cell surface. After imaging, a cleavage step removes the fluorescent label, and the process is repeated with other fluorescently labeled nucleotides until the desired read length is achieved. Sequence information is collected after each nucleotide addition step. Further description of tSMS is provided, for example, in U.S. Pat. Nos. 7,169,560, 6,818,395, and 7,282,337; U.S. Patent Publication Nos. 2009 / 0191565 and 2002 / 0164629; and Braslavsky et al., PNAS (USA), 100:3960-3964 (2003), the contents of each of which are incorporated herein by reference in their entirety.
[0123] Another example of a DNA sequencing technology that can be used in the methods of the provided invention is 454 sequencing (Roche), as described, for example, in Margulies, M. et al., 2005, Nature, 437, 376-380. 454 sequencing involves two steps. In the first step, DNA is sheared into fragments of approximately 300-800 base pairs, and the fragments are blunt-ended. Oligonucleotide adapters are then ligated to the ends of the fragments. The adapters serve as primers for fragment amplification and sequencing. The fragments can be bound to DNA capture beads, e.g., streptavidin-coated beads, using adapter B, e.g., containing a 5'-biotin tag. The bead-bound fragments are PCR-amplified within droplets of an oil-water emulsion. The result is multiple copies of clonally amplified DNA fragments on each bead. In the second step, the beads are captured in wells (picoliter size). Pyrosequencing is performed in parallel for each DNA fragment. The addition of one or more nucleotides generates a light signal that is recorded by a CCD camera in the sequencing instrument. The signal intensity is proportional to the number of nucleotides incorporated. Pyrosequencing utilizes pyrophosphate (PPi), which is released upon nucleotide addition. PPi is converted to ATP by ATP sulfurylase in the presence of adenosine 5' phosphosulfate. Luciferase uses ATP to convert luciferin to oxyluciferin, and this reaction generates light that is detected and analyzed.
[0124] Another example of a DNA sequencing technology that can be used in the methods of the provided invention is the SOLiD technology (Applied Biosystems). In SOLiD sequencing, genomic DNA is sheared into fragments and adapters are attached to the 5' and 3' ends of the fragments to generate a fragment library. Alternatively, internal adapters can be introduced by ligating adapters to the 5' and 3' ends of the fragments, circularizing the fragments, digesting the circularized fragments to generate internal adapters, and attaching adapters to the 5' and 3' ends of the resulting fragments to generate a mate-paired library. Next, a clonal bead population is prepared in a microreactor containing beads, primers, templates, and PCR components. After PCR, the templates are denatured and the beads are enriched to separate beads with extended templates. The templates on selected beads are subjected to a 3' modification that allows them to be attached to a glass slide. Sequences can be determined by sequential hybridization and ligation of partially random oligonucleotides to a central determinant base (or base pair) identified by a specific fluorophore. After recording the color, the ligated oligonucleotides are cleaved and removed, and the process is then repeated.
[0125] Another example of a DNA sequencing technology that can be used in the methods of the present invention provided is Ion Torrent sequencing, as described in U.S. Patent Application Publication Nos. 2009 / 0026082, 2009 / 0127589, 2010 / 0035252, 2010 / 0137143, 2010 / 0188073, 2010 / 0197507, 2010 / 0282617, 2010 / 0300559, 2010 / 0300895, 2010 / 0301398, and 2010 / 0304982, the contents of each of which are incorporated herein by reference in their entirety. In Ion Torrent sequencing, DNA is sheared into fragments of approximately 300-800 base pairs and the fragments are blunt-ended. Oligonucleotide adapters are then ligated to the ends of the fragments. The adapters serve as primers for amplification and sequencing of the fragments. The fragments can be attached to a surface and bound at a resolution such that the fragments can be individually separated. The addition of one or more nucleotides releases protons (H + ), which transmits a signal that is detected and recorded in the sequencing instrument. The signal strength is proportional to the number of nucleotides incorporated.
[0126] Another example of a sequencing technology that can be used in the provided methods of the present invention is Illumina sequencing. Illumina sequencing is based on amplification of DNA on a solid surface using fold-over PCR and immobilized primers. Genomic DNA is fragmented, and adapters are added to the 5' and 3' ends of the fragments. DNA fragments bound to the surface of the flow cell channel are extended and cross-linked amplified. The fragments become double-stranded, and the double-stranded molecules are denatured. Multiple cycles of solid-phase amplification followed by denaturation can create millions of clusters of approximately 1,000 copies of single-stranded DNA molecules of the same template in each channel of the flow cell. Sequential sequencing is performed using primers, DNA polymerase, and four fluorophore-labeled reversibly terminating nucleotides. After nucleotide incorporation, a laser is used to excite the fluorophore, an image is captured, and the identity of the first base is recorded. The 3' terminator and fluorophore from each incorporated base are removed, and the incorporation, detection, and identification steps are repeated.
[0127] Another example of a sequencing technology that can be used in the methods of the provided invention is Pacific Biosciences' single molecule real-time (SMRT) technology. In SMRT, each of the four DNA bases is attached to one of four different fluorescent dyes. These dyes are linked by phosphates. A single DNA polymerase is immobilized at the bottom of a zero-mode waveguide (ZMW) with a single molecule of template single-stranded DNA. The ZMW is a confining structure that allows observation of the incorporation of a single base by the DNA polymerase against a background of fluorescent nucleotides that rapidly diffuse out of the ZMW (microseconds). Incorporation of a nucleotide into the growing strand takes several milliseconds. During this time, the fluorescent label is excited, generating a fluorescent signal, and the fluorescent tag is cleaved. Detection of the corresponding fluorescence of the dye indicates which base has been incorporated. The process is then repeated.
[0128] Another example of sequencing technology that can be used in the provided method of the present invention is nanopore sequencing, as described in, for example, Soni GV and Meller A. (2007) Clin Chem 53:1996-2001. A nanopore is a small hole with a diameter of about 1 nanometer. When a nanopore is immersed in a conductive fluid and a potential is applied across it, a small current is generated due to the conduction of ions through the nanopore. The amount of current that flows is sensitive to the size of the nanopore. When a DNA molecule passes through the nanopore, each nucleotide on the DNA molecule obstructs the nanopore to a different extent. Therefore, when a DNA molecule passes through the nanopore, the change in the current passing through the nanopore represents the reading of the DNA sequence.
[0129] Another example of a sequencing technology that can be used in the methods of the provided invention involves sequencing DNA using a chemically sensitive field effect transistor (chemFET) array, as described, for example, in U.S. Patent Application Publication No. 20090026082. In one example of this technology, DNA molecules can be placed in a reaction chamber, and template molecules can be hybridized to a sequencing primer bound to a polymerase. The incorporation of one or more triphosphates into a new nucleic acid strand at the 3' end of the sequencing primer can be detected by a change in current flow via the chemFET. The array can have multiple chemFET sensors. In another example, a single nucleic acid can be bound to a bead, the nucleic acid can be amplified on the bead, and individual beads can be transferred to individual reaction chambers on a chemFET array, each chamber containing a chemFET sensor, allowing the nucleic acid to be sequenced.
[0130] Another example of sequencing technology that can be used in the method of the provided invention includes using electron microscope, as described in, for example, Moudrianakis EN and Beer M. Proc Natl Acad Sci USA.1965 March;53:564-71.In one example of this technology, individual DNA molecules are labeled with metal labels that can be distinguished using electron microscope.Then, these molecules are stretched on a flat surface, and are imaged using electron microscope to measure sequence.
[0131] If the nucleic acid from the sample is degraded or only minimal amounts of nucleic acid are available from the sample, PCR can be performed on the nucleic acid to obtain sufficient amounts of nucleic acid for sequencing, as described, for example, in U.S. Pat. No. 4,683,195, the contents of which are incorporated herein by reference in their entirety.
[0132] Methods for detecting levels of gene products (eg, RNA or protein) are known in the art.
[0133] The commonly used methods known in the art for quantifying mRNA expression in samples include, for example, Northern blotting and in situ hybridization, as described in Parker & Barnes, Methods in Molecular Biology 106:247-283(1999) (the contents of which are incorporated herein by reference in their entirety); RNAse protection assay, Hod, Biotechniques 13:852 854(1992) (the contents of which are incorporated herein by reference in their entirety); and PCR-based methods, such as reverse transcription polymerase chain reaction (RT-PCR), Weis et al., Trends in Genetics 8:263 264(1992) (the contents of which are incorporated herein by reference in their entirety). Alternatively, antibodies can be used to recognize specific duplexes, including RNA duplexes, DNA-RNA hybrid duplexes or DNA-protein duplexes. Other methods known in the art for measuring gene expression (e.g., RNA or protein abundance) are provided, for example, in U.S. Patent Application Publication No. 2006 / 0195269, the contents of which are incorporated herein by reference in their entirety.
[0134] A differentially or abnormally expressed gene refers to a gene whose expression is activated to a higher or lower level in subjects suffering from disorders such as PD, compared with its expression in normal or control subjects.This term also includes genes whose expression is activated to a higher or lower level in different stages of the same disorder.It is also understood that a differentially expressed gene can be activated or inhibited at the nucleic acid level or protein level, or can be subjected to alternative splicing, resulting in different polypeptide products.Such differences can be demonstrated, for example, by changes in the mRNA level, surface expression, secretion or other distribution of polypeptides.
[0135] Differential gene expression can involve comparing the expression between two or more genes or their gene products, or comparing the expression ratio between two or more genes or their gene products, or even comparing two differently processed products of the same gene, which differ between normal subjects and subjects suffering from a disorder such as PD, or between different stages of the same disorder. Differential expression includes both quantitative and qualitative differences in the temporal or cellular expression patterns of a gene or its expression products. Differential gene expression (increases and decreases in expression) is based on the percentage or fold change relative to expression in normal cells. An increase can be 1, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, or 200% increase compared to expression levels in normal cells. Alternatively, the fold increase can be a 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5 or 10-fold increase over the expression level in a normal cell. The decrease can be a 1, 5, 10, 20, 30, 40, 50, 55, 60, 65, 70, 75, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 99 or 100% decrease compared to the expression level in a normal cell.
[0136] In certain embodiments, gene expression is measured using reverse transcriptase PCR (RT-PCR), a quantitative method that can be used to characterize patterns of gene expression, distinguish closely related mRNAs, and analyze RNA structure, as well as compare mRNA levels in different sample populations.
[0137] The first step is the isolation of mRNA from a target sample. The starting material is typically total RNA isolated from human tissues or body fluids.
[0138] General methods for mRNA extraction are well known in the art and are disclosed in standard molecular biology textbooks, including Ausubel et al., Current Protocols of Molecular Biology, John Wiley and Sons (1997). Methods for RNA extraction from paraffin-embedded tissues are disclosed, for example, in Rupp and Locker, Lab Invest. 56:A67 (1987) and De Andres et al., BioTechniques 18:42044 (1995). The contents of each of these references are incorporated herein by reference in their entirety. In particular, RNA isolation can be performed using purification kits, buffer sets, and proteases from commercial manufacturers such as Qiagen, according to the manufacturer's instructions. For example, total RNA from cultured cells can be isolated using Qiagen RNeasy mini-columns. Other commercially available RNA isolation kits include the MASTERPURE Complete DNA and RNA Purification Kit (EPICENTRE, Madison, Wisconsin) and the Paraffin Block RNA Isolation Kit (Ambion, Inc.). Total RNA from tissue samples can be isolated using RNA Stat-60 (Tel-Test). RNA prepared from tumors can be isolated, for example, by cesium chloride density gradient centrifugation.
[0139] The first step in gene expression profiling by RT-PCR is reverse transcription of the RNA template into cDNA, followed by exponential amplification in a PCR reaction. The two most commonly used reverse transcriptases are aviromyeloblastosis virus reverse transcriptase (AMV-RT) and Moloney murine leukemia virus reverse transcriptase (MMLV-RT). The reverse transcription step is typically primed using specific primers, random hexamers, or oligo-dT primers, depending on the situation and the goal of expression profiling. For example, extracted RNA can be reverse transcribed using a GeneAmp RNA PCR kit (Perkin Elmer, California, USA) according to the manufacturer's instructions. The derived cDNA can then be used as a template in the subsequent PCR reaction.
[0140] The PCR process can use a variety of thermostable, DNA-dependent DNA polymerases, but typically employs Taq DNA polymerase, which possesses 5'-3' nuclease activity but lacks 3'-5' proofreading endonuclease activity. Thus, TaqMan® PCR typically utilizes the 5' nuclease activity of Taq polymerase to hydrolyze hybridization probes bound to its target amplicon, although any enzyme with comparable 5' nuclease activity can be used. Two oligonucleotide primers are used to generate the amplicon typical of a PCR reaction. A third oligonucleotide, or probe, is designed to detect the nucleotide sequence located between the two PCR primers. The probe is non-extendible by the Taq DNA polymerase enzyme and is labeled with a reporter and a quencher fluorescent dye. Laser-induced emission from the reporter dye is quenched by the quencher dye when the two dyes are positioned close to each other on the probe. During the amplification reaction, the Taq DNA polymerase enzyme cleaves the probe in a template-dependent manner. The resulting probe fragments dissociate in solution, and the signal from the released reporter dye is free from the quenching effect of the second fluorophore. One molecule of reporter dye is liberated for each new molecule synthesized, and detection of the unquenched reporter dye provides the basis for quantitative interpretation of the data.
[0141] TaqMan® RT-PCR can be performed using commercially available equipment, such as the ABI PRISM 7700™ Sequence Detection System (Perkin-Elmer-Applied Biosystems, Foster City, CA, USA) or the Lightcycler (Roche Molecular Biochemicals, Mannheim, Germany). In a specific embodiment, the 5' nuclease procedure is performed on a real-time quantitative PCR device such as the ABI PRISM 7700™ Sequence Detection System™. This system consists of a thermocycler, laser, charge-coupled device (CCD), camera, and computer. The system amplifies samples in a 96-well format in the thermocycler. During amplification, laser-induced fluorescence signals are collected in real time for all 96 wells via fiber optic cables and detected by the CCD. The system includes software for operating the instrument and analyzing the data.
[0142] 5'-nuclease assay data are initially expressed as Ct, or threshold cycle. As described above, a fluorescence value is recorded during each cycle and represents the amount of product amplified up to that point in the amplification reaction. The point at which the fluorescence signal is first recorded as statistically significant is the threshold cycle (Ct).
[0143] To minimize errors and the effects of inter-sample variation, RT-PCR is usually performed using an internal standard. An ideal internal standard is expressed at a consistent level across different tissues and is unaffected by experimental treatments. The RNAs most frequently used to normalize patterns of gene expression are the mRNAs of the housekeeping genes glyceraldehyde-3-phosphate dehydrogenase (GAPDH) and actin, beta (ACTB). For analyses performed on preimplantation embryos and oocytes, conserved helix-loop-helix ubiquitous kinase (CHUK) is the gene used for normalization.
[0144] A more recent variation of the RT-PCR technique is real-time quantitative PCR, which measures PCR product accumulation via a dual-labeled fluorogenic probe (i.e., TaqMan® probe). Real-time PCR is compatible with both quantitative competitive PCR, which uses an internal competitor for each target sequence for normalization, and quantitative comparative PCR, which uses a normalization gene contained in the sample or a housekeeping gene for RT-PCR. For further details, see, for example, Held et al., Genome Research 6:986 994 (1996), the contents of which are incorporated herein by reference in their entirety.
[0145] In another embodiment, gene expression is measured using a MassARRAY-based gene expression profiling method. In this MassARRAY-based gene expression profiling method, developed by Sequenom, Inc. (San Diego, CA), after RNA isolation and reverse transcription, the resulting cDNA is spiked with a synthetic DNA molecule (competitor), which matches the target cDNA region at all positions except a single base and serves as an internal standard. The cDNA / competitor mixture is PCR-amplified and subjected to post-PCR shrimp alkaline phosphatase (SAP) enzyme treatment, which results in dephosphorylation of the remaining nucleotides. After inactivation of the alkaline phosphatase, the competitor-derived PCR products and the cDNA-derived PCR products are subjected to primer extension, which results in different mass signals for the competitor-derived and cDNA-derived PCR products. After purification, these products are dispensed onto a chip array pre-packaged with the components required for analysis by matrix-assisted laser desorption / ionization time-of-flight mass spectrometry (MALDI-TOF MS). The amount of cDNA present in the reaction is then quantified by analyzing the ratio of peak areas in the resulting mass spectra. For further details, see, for example, Ding and Cantor, Proc. Natl. Acad. Sci. USA 100:3059 3064 (2003).
[0146] Additional PCR-based techniques include, for example, differential display (Liang and Pardee, Science 257:967-971 (1992)); amplified fragment length polymorphism (iAFLP) (Kawamoto et al., Genome Res. 12:1305-1312 (1999)); BeadArray™ technology (Illumina, San Diego, CA; Oliphant et al., Discovery of Markers for Disease (Supplement to Biotechniques), June 2002; Ferguson et al., Analytical Chemistry 72:5618 (2000)); BeadsArray for detection of gene expression (BADGE) (Yang et al., Genome Res. 11:1888), which uses the commercially available Luminex 100 LabMAP system and multiple color-coded microspheres (Luminex Corp., Austin, TX) in a rapid assay of gene expression. 1898 (2001); and high-coverage expression profiling (HiCEP) analysis (Fukumura et al., Nucl. Acids. Res. 31(16)e94 (2003)), the contents of each of which are incorporated herein by reference in their entirety.
[0147] In certain embodiments, differential gene expression can be identified or confirmed by using microarray technology.In this method, the polynucleotide sequence of interest (including cDNA and oligonucleotide) is plated or arrayed on a microchip substrate.The arrayed sequence is then hybridized with specific DNA probes derived from target cells or tissues.The method for making microarrays and determining gene product expression (for example, RNA or protein) is described in US Patent Application Publication No. 2006 / 0195269, the contents of which are incorporated herein by reference in their entirety.
[0148] In a specific embodiment of microarray technology, PCR-amplified inserts of cDNA clones are applied to a substrate in a dense array, for example, at least 10,000 nucleotide sequences. The microarrayed genes, each with 10,000 elements, are immobilized on a microchip and suitable for hybridization under stringent conditions. Fluorescently labeled cDNA probes can be generated by incorporating fluorescent nucleotides through reverse transcription of RNA extracted from the tissue of interest. The labeled cDNA probes applied to the chip specifically hybridize to each DNA spot on the array. After rigorous washing to remove nonspecifically bound probes, the chip is scanned by confocal laser microscopy or another detection method, such as a CCD camera. Quantitation of hybridization of each arrayed element allows assessment of the corresponding mRNA abundance. Using dual-color fluorescence, separately labeled cDNA probes generated from two RNA sources are hybridized pairwise to the array. Thus, the relative abundance of transcripts from the two sources corresponding to each specific gene is simultaneously determined. By reducing the scale of hybridization, the expression pattern of a large number of genes can be easily and quickly evaluated.Such method has been shown to have the necessary sensitivity to detect the rare transcripts that are expressed in a few copies per cell, and to reproducibly detect at least approximately two-fold difference in expression level, as described, for example, in Schena et al., Proc.Natl.Acad.Sci.USA 93(2):106 149(1996) (the entire contents of which are incorporated herein by reference).Microarray analysis can be carried out by commercially available equipment according to manufacturer's protocol, such as by using Affymetrix GenChip technology or Incyte's microarray technology.
[0149] Alternatively, protein levels can be determined by constructing an antibody microarray, whose binding sites contain fixed, preferably monoclonal, antibodies specific to multiple protein species encoded by the cellular genome. Preferably, the antibodies are present against a substantial portion of the proteins of interest. Methods for producing monoclonal antibodies are well known (see, for example, Harlow and Lane, 1988, ANTIBODIES: A LABORATORY MANUAL, Cold Spring Harbor, NY, incorporated in its entirety for all purposes). In one embodiment, monoclonal antibodies are produced against synthetic peptide fragments designed based on the genomic sequence of the cell. Using such an antibody array, proteins from the cell are contacted with the array, and their binding is assayed using assays known in the art. Generally, the expression and expression levels of proteins of diagnostic or prognostic interest can be detected by immunohistochemical staining of tissue sections or sections.
[0150] Finally, the transcript levels of marker genes in several tissue specimens can be characterized using "tissue arrays," as described, for example, in Kononen et al., Nat.Med 4(7):844-7(1998). In tissue arrays, multiple tissue samples are evaluated on the same microarray. Arrays allow in situ detection of RNA and protein levels, and serial sections allow for the simultaneous analysis of multiple samples.
[0151] In another embodiment, gene expression is measured using sequence analysis of gene expression (SAGE). Sequence analysis of gene expression (SAGE) is a method that allows for the simultaneous quantitative analysis of multiple gene transcripts without the need to provide individual hybridization probes for each transcript. First, short sequence tags (approximately 10-14 bp) are generated that contain enough information to uniquely identify each transcript, as long as the tag is obtained from a unique location within each transcript. Many transcripts are then linked together to form a long serial molecule that can be sequenced, revealing the identity of multiple tags simultaneously. By determining the abundance of individual tags and identifying the genes corresponding to each tag, the expression pattern of any population of transcripts can be quantitatively assessed. For further details, see, for example, Velculescu et al., Science 270:484-487 (1995); and Velculescu et al., Cell 88:243-51 (1997), the contents of each of which are incorporated herein by reference in their entirety.
[0152] In another embodiment, gene expression is measured using massively parallel signature sequencing (MPSS). This method, described in Brenner et al., Nature Biotechnology 18:630 634 (2000), is a sequencing approach that combines non-gel-based signature sequencing with in vitro cloning of millions of templates on individual 5 μm diameter microbeads. First, a microbead library of DNA templates is constructed by in vitro cloning. This is followed by the formation of a high-density (typically 3 × 10) planar array of template-containing microbeads in a flow cell. 6 Microbeads / cm 2 The free ends of the cloned templates on each microbead are simultaneously analyzed using a fluorescence-based signature sequencing method that does not require DNA fragment separation. This method has been shown to simultaneously and accurately generate hundreds of thousands of gene signature sequences from a yeast cDNA library in a single run.
[0153] Immunohistochemical methods are also suitable for detecting the expression level of the gene products of the present invention. Therefore, expression is detected using antibodies (monoclonal or polyclonal) or antisera, such as polyclonal antisera, specific to each marker. The antibody can be detected by direct labeling of the antibody itself, for example, with a radioactive label, a fluorescent label, a hapten label such as biotin, or an enzyme such as horseradish peroxidase or alkaline phosphatase. Alternatively, an unlabeled primary antibody is used in combination with a labeled secondary antibody, including an antiserum, polyclonal antiserum, or monoclonal antibody specific to the primary antibody. Immunohistochemical protocols and kits are well known in the art and are commercially available.
[0154] In certain embodiments, a proteomics approach is used to measure gene expression. Proteome refers to the entire protein present in a sample (e.g., tissue, organism, or cell culture) at a given time point. Proteomics includes, among other things, the study of the overall changes in protein expression in a sample (also called expression proteomics). Proteomics typically involves the following steps: (1) separating individual proteins in a sample by 2-D gel electrophoresis (2-D PAGE); (2) identifying the individual proteins recovered from the gel, for example, by mass spectrometry or N-terminal sequencing; and (3) analyzing the data using bioinformatics. Proteomics methods are a valuable complement to other methods of gene expression profiling, and can be used alone or in combination with other methods to detect the products of the diagnostic markers of the present invention.
[0155] In some embodiments, mass spectrometry (MS) analysis can be used alone or in combination with other methods (e.g., immunoassays or RNA measurement assays) to determine the presence and / or quantity of one or more biomarkers disclosed herein in a biological sample. In some embodiments, the MS analysis comprises matrix-assisted laser desorption / ionization (MALDI) time-of-flight (TOF) MS analysis, such as direct spot MALDI-TOF or liquid chromatography MALDI-TOF mass spectrometry. In some embodiments, the MS analysis comprises electrospray ionization (ESI) MS (e.g., liquid chromatography (LC) ESI-MS, etc.). Mass spectrometry can be accomplished using commercially available spectrometers. Methods utilizing MS analysis, including MALDI-TOF MS and ESI-MS, to detect the presence and quantity of biomarker peptides in biological samples are known in the art. See, e.g., U.S. Patent Nos. 6,925,389, 6,989,100, and 6,890,763, each of which is incorporated herein by reference in its entirety.
[0156] Report on genetic modifiers of LRRK2 The method of the present invention may include providing a report about a subject. The report may identify one or more genetic modifiers of LRRK2 in genetic data from the subject. The report may include additional information about the subject, such as age, sex, weight, height, genetic data, genomic data, or other health or medical information. The report may also include other information about PD. For example, but not limited to, the report may include information about PD symptoms or genes associated with PD, such as the symptoms and genes listed above.
[0157] The report may be provided in any suitable manner, for example, but not limited to, the report may be provided on paper or on a display device such as a computer monitor, a telephone, a portable electronic device, or the like.
[0158] The report may be provided to a healthcare provider, such as a doctor or nurse. The report may provide guidance to the healthcare provider regarding whether treating the subject with an LRRK2 inhibitor is appropriate. The report may provide the healthcare provider with instructions or recommendations for treating the subject with an LRRK2 inhibitor. The report may recommend that the healthcare provider prescribe or provide the subject with an LRRK2 inhibitor, or otherwise instruct the subject to obtain and take an LRRK2 inhibitor.
[0159] The report may include guidance on whether to use a second agent in addition to the LRRK2 inhibitor to treat the subject. The second agent may be a known therapeutic agent for the treatment of PD, such as any of those listed above.
[0160] LRRK2 inhibitors The methods of the invention may include providing one or more LRRK2 inhibitors to a subject or recommending that a subject take one or more LRRK2 inhibitors. LRRK2 inhibitors are known in the art and are described in, for example, International Patent Publication Nos. WO2012 / 028629, WO2012 / 058193, WO2012 / 118679, WO2012 / 143143, WO2012 / 143144, WO2014 / 001973, WO2014 / 060112, WO2014 / 060113, WO2014 / 145909, WO2014 / 160430, WO2014 / 170248, WO2015 / 092592, WO2015 / 113451, WO2015 / 113452, WO2016 / 130920, WO2017 / 012576, WO2017 / 046675, WO2017 / 087905, WO20 17 / 106771, WO2017 / 156493, WO2017 / 218843, WO2018 / 137573, WO2018 / 137593, WO2018 / 137618, WO2018 / 137619, WO2 No. 018 / 163030, WO2018 / 163066, WO2018 / 217946, WO2019 / 012093, WO2019 / 104086, WO2019 / 112269, WO2019 / 126383, WO2020 / 149723, WO2020 / 170205, and WO2020 / 210684; U.S. Patent No. 9,499,535; and co-pending U.S. patent applications Ser. Nos. 63 / 050,385 and 63 / 133,523. Nos. 63 / 113,533, 63 / 137,814, 63 / 137816, and 63 / 142009; and co-pending international applications PCT / IB2020 / 000727, PCT / IB2020 / 000730, PCT / US2021 / 041270, and PCT / US2021 / 041271, the contents of each of which are incorporated herein by reference in their entireties. Any LRRK2 disclosed in any of the foregoing references may be used in the methods of the present invention.
[0161] For example, and without limitation, the LRRK2 inhibitor can be CZC-25146, CZC-54252, DNL151, DNL201, GNE-7915, GSK2578215A, HG-10-102-01, JH-II-127, K252A, K252B, LRRK2-IN-1, MLi-2, PF-06447475, or staurosporine.
[0162] In some methods of the present invention, the LRRK2 inhibitor is represented by Formula (I), (II), (III), and (IV): [ka] (In the formula, A is NH, O, S, C=O, NR 3 or CR 4 R 5 and X is an optionally substituted arylene, heteroarylene, cycloalkylene, heterocycloalkylene, alkylcycloalkylene, heteroalkylcycloalkylene, aralkylene, or heteroaralkylene group; R 1 is an optionally substituted alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 2 is a hydrogen atom, a halogen atom, NO, N, OH, SH, NH, or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 3 is an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkyl-cycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 4is a hydrogen atom, NO, N, OH, SH, NH, or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 5 is a hydrogen atom, NO, N, OH, SH, NH, or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; B is NH, O, S, C=O, NR 14 or CR 15 R 16 and R 11 is an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 12 is an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 12 is attached to the pyrimidine ring of formula (II) via a carbon-carbon bond, R 13 is a hydrogen atom, a halogen atom, NO, N, OH, SH, NH, or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 14 is an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkyl-cycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 15 is a hydrogen atom, NO, N, OH, SH, NH, or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 16 is a hydrogen atom, NO, N, OH, SH, NH, or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 21 each of which is optionally substituted aryl or heteroaryl; R 22 H, halo, OH, CN, CF3, C 1~6 Alkyl, C 1~6 Alkoxy, C 1~6 Haloalkyl, C 1~6 Thioalkyl, C 3~8 Cycloalkyl, C 2~8 heterocycloalkyl, aryl, or heteroaryl; Y is aryl or 5- or 6-membered heteroaryl; C 1~6 Alkyl, C 1~6 Alkoxy, C 1~6 Haloalkyl, C 1~6 Thioalkyl, C 3~8 Cycloalkyl, C 2~8 Each of heterocycloalkyl, aryl, and heteroaryl is selected from halo, OH, CN, CF, NH, NO, C 1~6 Alkyl, C 1~6 Haloalkyl, C 1~6 Thioalkyl, C 3~8 Cycloalkyl, C 2~8 Heterocycloalkyl, C 2~8 Heterocycloalkenyl, C 2~6 Alkenyl, C 2~6 Alkynyl, C 1~6 Alkoxy, C 1~6 Haloalkoxy, C 1~6Alkylamino, C 2~6 Dialkylamino, C 7~12 Aralkyl, C 1~12 optionally substituted with one or more moieties selected from the group consisting of heteroaralkyl, aryl, heteroaryl, —C(O)R, —C(O)OR, —C(O)NRR′, —C(O)NRS(O)R′, —C(O)NRS(O)NR′R″, —OR, —OC(O)NRR′, —NRR′, —NRC(O)R′, —NRC(O)NR′R″, —NRS(O)R′, —NRS(O)NR′R″, —S(O)R, and —S(O)NRR; Each of R, R' and R" is independently H, halo, OH, C 1~6 Alkyl, C 1~6 Haloalkyl, C 1~6 Alkoxy, C 3~8 Cycloalkyl, C 2~8 heterocycloalkyl, aryl, or heteroaryl, or R and R′ or R′ and R″ together with the nitrogen to which they are attached are C 2~8 forming a heterocycloalkyl, R 31 is C(O)CH2R 33 , optionally substituted cycloalkyl, optionally substituted cycloheteroalkyl, optionally substituted cycloalkenyl, optionally substituted cycloheteroalkenyl, optionally substituted aryl, or optionally substituted heteroaryl; R 32 each instance of is independently halo, haloalkyl, optionally substituted alkoxyl, optionally substituted alkyl, optionally substituted heteroalkyl, optionally substituted alkenyl, optionally substituted heteroalkenyl; R 33 is optionally substituted cycloalkyl, optionally substituted cycloheteroalkyl, optionally substituted cycloalkenyl, optionally substituted cycloheteroalkenyl, optionally substituted aryl, or optionally substituted heteroaryl; Z is cycloalkyl, cycloheteroalkyl, cycloalkenyl, cycloheteroalkenyl, aryl or heteroaryl, Z can be aryl substituted with 2 or 3 instances of R, Z can be phenyl substituted with 2 or 3 instances of R, Z can be heteroaryl substituted with 2 or 3 instances of R, Z can be a 6-membered heteroaryl substituted with 2 or 3 instances of R; n is 0 to 5) or a pharmaceutically acceptable salt of any of the above compounds.
[0163] The LRRK2 inhibitor may be provided to a subject as a pharmaceutical composition. The pharmaceutical composition may contain a therapeutically effective amount of the LRRK2 inhibitor. A therapeutically effective amount refers to an amount effective to prevent, alleviate, or ameliorate symptoms of a disease such as PD, or to prolong the survival of the subject being treated. Determining a therapeutically effective amount is within the skill of the art. The therapeutically effective amount or dosage of an LRRK2 inhibitor may vary within a wide range and may be determined by methods known in the art. Such dosage may be adjusted to the individual requirements of each particular case, including the specific compound administered, the route of administration, the condition being treated, and the patient being treated.
[0164] For oral administration, such therapeutically useful agents can be administered by one of the following routes: orally, for example, in the form of tablets, sugar-coated tablets, coated tablets, pills, semisolids, soft or hard capsules, such as soft and hard gelatin capsules, aqueous or oily solutions, emulsions, suspensions, or syrups; parenterally, including intravenous, intramuscular, and subcutaneous injections, for example, as injectable solutions or suspensions, as suppositories, by inhalation or insufflation, for example, as powder formulations, microcrystalline preparations, or as sprays (e.g., liquid aerosols); transdermally, for example, via transdermal delivery systems (TDS), such as plasters containing the active ingredient, or intranasally. For the preparation of such tablets, pills, semisolids, coated tablets, sugar-coated tablets, and hard, e.g., gelatin, capsules, the therapeutically useful products can be mixed with pharmaceutically inert inorganic or organic excipients, such as lactose, sucrose, glucose, gelatin, malt, silica gel, starch or derivatives thereof, talc, stearic acid or its salts, dried skim milk, etc. For the preparation of soft capsules, excipients such as vegetable, petroleum, animal, or synthetic oils, waxes, fats, and polyols may be used. For the preparation of liquid solutions, emulsions, suspensions, or syrups, excipients such as water, alcohol, aqueous saline, aqueous dextrose, polyols, glycerin, lipids, phospholipids, cyclodextrins, vegetable, petroleum, animal, or synthetic oils may be used. Lipids such as phospholipids (e.g., of natural origin and / or with a particle size between 300 and 350 nm) in phosphate-buffered saline (pH 7-8, e.g., 7.4) are particularly useful. For suppositories, excipients such as vegetable, petroleum, animal, or synthetic oils, waxes, fats, and polyols may be used as they are. For aerosol formulations, compressed gases suitable for this purpose, such as oxygen, nitrogen, and carbon dioxide, may be used as they are. Pharmaceutically useful agents may also contain additives for preservation and stabilization, such as UV stabilizers, emulsifiers, sweeteners, flavoring agents, salts for varying osmotic pressure, buffers, coating additives, and antioxidants.
[0165] Provision of LRRK2 inhibitors to subjects The method of the present invention can include providing a subject with an LRRK2 inhibitor. The LRRK2 inhibitor can be provided by any suitable route or mode of administration. For example, but not limited to, the compound can be provided bucally, dermally, enterally, intraarterially, intramuscularly, intraocularly, intravenously, nasally, orally, parenterally, pulmonary, rectally, subcutaneously, topically, transdermally, by injection, or with or on an implantable medical device (e.g., a stent or a drug-eluting stent or a balloon equivalent).
[0166] The LRRK2 inhibitor can be provided according to a dosing regimen, which can include the amount of administration, the frequency of administration, or both.
[0167] Dosage can be provided at any suitable interval.For example, but not limited to, dosage can be provided once a day, twice a day, three times a day, four times a day, five times a day, six times a day, eight times a day, once every 48 hours, once every 36 hours, once every 24 hours, once every 12 hours, once every 8 hours, once every 6 hours, once every 4 hours, once every 3 hours, once every 2 days, once every 3 days, once every 4 days, once every 5 days, once every week, twice a week, three times a week, four times a week or five times a week.
[0168] The dose may be provided in a single dosage, i.e., the dose may be provided as a single tablet, capsule, pill, etc. Alternatively, the dose may be provided in split doses, i.e., the dose may be provided as multiple tablets, capsules, pills, etc.
[0169] Administration can continue for a specified period of time. For example, but not limited to, doses can be provided for at least 1 week, at least 2 weeks, at least 3 weeks, at least 4 weeks, at least 6 weeks, at least 8 weeks, at least 10 weeks, at least 12 weeks or more.
[0170] The subject can be any type of subject, for example, any of the subjects described above in connection with assays for obtaining genetic data.
[0171] The present invention includes a combination therapy in which an LRRK2 inhibitor is provided to a subject in combination with a second drug, such as any of the drugs described above in the PD section.The LRRK2 inhibitor and the second drug may be provided in a single composition, or in separate compositions.The LRRK2 inhibitor and the second drug may be provided according to the same administration regimen, or according to different administration regimens. [Example]
[0172] Example 1 The likelihood of response to LRRK2 inhibitors was analyzed in a human population. The dataset included the complete dataset from the Accelerating Medicines Partnership-Parkinson's Disease (AMP-PD). The input data were quality-controlled Parkinson's disease case data, focusing on baseline clinical, demographic, RNA, and DNA sequencing data for available samples as of June 1, 2020. Whole-genome sequencing and RNA sequencing were processed using the standard pipeline described on the AMP-PD website. The analysis was restricted to samples with data missiness rates below 15% after consensus quality control. The analysis was also rerun with adjustment for a European population substructure, yielding nearly identical results for the same set of over 1,000 cases.
[0173] To identify potential modifiers, we used the open-source automated machine learning package GenoML. This package performed feature selection / weighting and normalization, and then competed algorithms on a randomly validated 70% training set and 30% test set. The best-performing algorithm, in terms of balanced accuracy, was then selected for further tuning and cross-validation. The best-performing algorithm then underwent hyperparameter tuning using a randomized grid search method and 10-fold cross-validation, with the focus of this tuning process being to maximize balanced accuracy. Results were coded as 0 / 1, with 1 indicating harboring a known LRRK2 causative variant. A matrix of probabilities for WT LRRK2 cases was exported, indicating how "LRRK2+-like" they were at the molecular, clinical, and demographic levels. The most significant features across all iterations of the model were used as potential modifiers.
[0174] The results are shown in Table 1.
[0175] [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4] [Table 1-5]
[0176] Example 2 The likelihood of response to an LRRK2 inhibitor was analyzed in a population of human subjects. Figures 1, 2, and 3 provide total urinary di-18:1 BMP, total urinary di-22:6 BMP, and 2,2' 22:6 BMP, respectively, for different patient populations.
[0177] Incorporation by Reference References and citations to other documents, such as patents, patent applications, patent publications, journals, books, articles, web content, etc. are made throughout this disclosure, and all such documents are incorporated herein by reference in their entirety for all purposes.
[0178] equivalent Various modifications of the invention and many further embodiments thereof, in addition to those shown and described herein, will become apparent to those skilled in the art from the entire contents of this document, including references to the scientific and patent literature cited herein. The subject matter of this specification contains important information, exemplification, and guidance that can be adapted to the practice of this invention in its various embodiments and equivalents thereof.
Claims
1. A method of treating a patient with Parkinson's disease associated with wild-type leucine-rich repeat kinase (LRRK2), comprising: A method comprising providing one or more LRRK2 inhibitors to a patient exhibiting Parkinson's disease who has wild-type LRRK2 and elevated levels of bis(monoacylglycerol)phosphate (BMP) and / or hemi-bis(monoacylglycerol)phosphate (hemi-BMP) compared to the BMP and / or hemi-BMP levels of a subject not suffering from Parkinson's disease, and who has wild-type LRRK2, thereby treating Parkinson's disease associated with wild-type LRRK2.
2. BMP, (i) di-22:6 BMP, (ii) di-18:1 BMP, (iii) 16:0 / 18:1 BMP, (iv) di-20:4 BMP, (v) 18:0 / 20:4 BMP, (vi) 2,2' di-22:6 BMP, and (vii) any combination thereof; 2. The method of claim 1, wherein the hemi-BMP is selected from the group consisting of: (i) hemi-BMP (18:1 / 18:1) 16:0, (ii) hemi-BMP (14:0 / 14:0) 14:0, (iii) hemi-BMP (18:1 / 18:1) 18:0, (iv) hemi-BMP (18:1 / 18:1) 18:1, and (iv) any combination thereof.
3. The method of claim 1 , wherein the BMP level is measured in a biological fluid of the patient.
4. 4. The method of claim 3, wherein the biological fluid is urine, blood, cerebrospinal fluid (CSF), bile, or saliva.
5. 10. The method of claim 1, wherein said elevated levels of BMP and / or hemi-BMP are at concentrations that indicate said patient is responsive to said one or more LRRK2 inhibitors.
6. 2. The method of claim 1, wherein the one or more LRRK2 inhibitors are selected from the group consisting of CZC-25146, CZC-54252, DNL151, DNL201, GNE-7915, GSK2578215A, HG-10-102-01, JH-II-127, K252A, K252B, LRRK2-IN-1, MLi-2, PF-06447475, and staurosporine.
7. The one or more LRRK2 inhibitors are represented by formula (I), (II), (III) and (IV): 【Chemistry 4】 (In the formula, A is NH, O, S, C=O, NR 3 or CR 4 R 5 and X is an optionally substituted arylene, heteroarylene, cycloalkylene, heterocycloalkylene, alkylcycloalkylene, heteroalkylcycloalkylene, aralkylene, or heteroaralkylene group; R 1 is an optionally substituted alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 2 is a hydrogen atom, a halogen atom, NO 2 , N 3 , OH, SH, NH 2 or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl or heteroaralkyl group; R 3 is an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkyl-cycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 4 is a hydrogen atom, NO 2 , N 3 , OH, SH, NH 2 or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl or heteroaralkyl group; R 5 is a hydrogen atom, NO 2 , N 3 , OH, SH, NH 2 or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl or heteroaralkyl group; B is NH, O, S, C=O, NR 14 or CR 15 R 16 and R 11 is an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 12 is an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 12 is attached to the pyrimidine ring of formula (II) via a carbon-carbon bond, R 13 is a hydrogen atom, a halogen atom, NO 2 , N 3 , OH, SH, NH 2 or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl or heteroaralkyl group; R 14 is an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkyl-cycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 15 is a hydrogen atom, NO 2 , N 3 , OH, SH, NH 2 or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl or heteroaralkyl group; R 16 is a hydrogen atom, NO 2 , N 3 , OH, SH, NH 2 or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl or heteroaralkyl group; R 21 each of which is optionally substituted aryl or heteroaryl; R 22 H, halo, OH, CN, CF 3 , C 1~6 Alkyl, C 1~6 Alkoxy, C 1~6 Haloalkyl, C 1~6 Thioalkyl, C 3~8 Cycloalkyl, C 2~8 heterocycloalkyl, aryl, or heteroaryl; Y is aryl or 5- or 6-membered heteroaryl; C 1~6 Alkyl, C 1~6 Alkoxy, C 1~6 Haloalkyl, C 1~6 Thioalkyl, C 3~8 Cycloalkyl, C 2~8 Each of heterocycloalkyl, aryl and heteroaryl is selected from halo, OH, CN, CF 3 , N.H. 2 , NO 2 , C 1~6 Alkyl, C 1~6 Haloalkyl, C 1~6 Thioalkyl, C 3~8 Cycloalkyl, C 2~8 Heterocycloalkyl, C 2~8 Heterocycloalkenyl, C 2~6 Alkenyl, C 2~6 Alkynyl, C 1~6 Alkoxy, C 1~6 Haloalkoxy, C 1~6 Alkylamino, C 2~6 Dialkylamino, C 7~12 Aralkyl, C 1~12 Heteroaralkyl, aryl, heteroaryl, —C(O)R, —C(O)OR, —C(O)NRR′, —C(O)NRS(O) 2 R', -C(O)NRS(O) 2 NR'R", -OR, -OC(O)NRR', -NRR', -NRC(O)R', -NRC(O)NR'R", -NRS(O) 2 R', -NRS(O) 2 NR'R", -S(O) 2 R, and -S(O) 2 and optionally substituted with one or more moieties selected from the group consisting of NRR; Each of R, R′, and R″ is independently H, halo, OH, C 1~6 Alkyl, C 1~6 Haloalkyl, C 1~6 Alkoxy, C 3~8 Cycloalkyl, C 2~8 heterocycloalkyl, aryl, or heteroaryl, or R and R′ or R′ and R″ together with the nitrogen to which they are attached are C 2~8 forming a heterocycloalkyl, R 31 is C(O)CH 2 R 33 , optionally substituted cycloalkyl, optionally substituted cycloheteroalkyl, optionally substituted cycloalkenyl, optionally substituted cycloheteroalkenyl, optionally substituted aryl, or optionally substituted heteroaryl; R 32 each instance of is independently halo, haloalkyl, optionally substituted alkoxyl, optionally substituted alkyl, optionally substituted heteroalkyl, optionally substituted alkenyl, optionally substituted heteroalkenyl; R 33 is optionally substituted cycloalkyl, optionally substituted cycloheteroalkyl, optionally substituted cycloalkenyl, optionally substituted cycloheteroalkenyl, optionally substituted aryl, or optionally substituted heteroaryl; Z is cycloalkyl, cycloheteroalkyl, cycloalkenyl, cycloheteroalkenyl, aryl, or heteroaryl, and Z is selected from two or three instances of R 2 Z may be aryl substituted with two or three instances of R 2 Z may be phenyl substituted with two or three instances of R 2 Z may be heteroaryl substituted with two or three instances of R 2 may be a 6-membered heteroaryl substituted with n is 0 to 5. or a pharmaceutically acceptable salt of any of the above compounds.
8. 1. A method for determining whether a patient having Parkinson's disease associated with wild-type LRRK2 will respond to an LRRK2 inhibitor, comprising: conducting an assay to measure bis(monoacylglycerol)phosphate (BMP) and / or hemi-bis(monoacylglycerol)phosphate (hemi-BMP) levels in said patient; generating a report identifying the patient's levels of BMP and / or hemi-BMP compared to BMP and / or hemi-BMP levels in subjects not afflicted with Parkinson's disease and having wild-type LRRK2; and providing said report to a physician so that if said report indicates elevated levels of BMP and / or hemi-BMP in said patient compared to said subject, said physician may prescribe or provide one or more LRRK2 inhibitors to said patient.
9. BMP, (i) di-22:6 BMP, (ii) di-18:1 BMP, (iii) 16:0 / 18:1 BMP, (iv) di-20:4 BMP, (v) 18:0 / 20:4 BMP, (vi) 2,2' di-22:6 BMP, and (vii) any combination thereof; 9. The method of claim 8, wherein the hemi-BMP is selected from the group consisting of: (i) hemi-BMP (18:1 / 18:1) 16:0, (ii) hemi-BMP (14:0 / 14:0) 14:0, (iii) hemi-BMP (18:1 / 18:1) 18:0, (iv) hemi-BMP (18:1 / 18:1) 18:1, and (iv) any combination thereof.
10. The method of claim 8 , wherein the BMP level is measured in a biological fluid of the patient.
11. 11. The method of claim 10, wherein the biological fluid is urine, blood, cerebrospinal fluid (CSF), bile, or saliva.
12. 9. The method of claim 8, wherein elevated levels of BMP and / or hemi-BMP in the patient indicate that the patient is responsive to one or more LRRK2 inhibitors.
13. 9. The method of claim 8, wherein the one or more LRRK2 inhibitors are selected from the group consisting of CZC-25146, CZC-54252, DNL151, DNL201, GNE-7915, GSK2578215A, HG-10-102-01, JH-II-127, K252A, K252B, LRRK2-IN-1, MLi-2, PF-06447475, and staurosporine.
14. The one or more LRRK2 inhibitors are represented by formula (I), (II), (III) and (IV): 【Chemistry 5】 (In the formula, A is NH, O, S, C=O, NR 3 or CR 4 R 5 and X is an optionally substituted arylene, heteroarylene, cycloalkylene, heterocycloalkylene, alkylcycloalkylene, heteroalkylcycloalkylene, aralkylene, or heteroaralkylene group; R 1 is an optionally substituted alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 2 is a hydrogen atom, a halogen atom, NO 2 , N 3 , OH, SH, NH 2 or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl or heteroaralkyl group; R 3 is an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkyl-cycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 4 is a hydrogen atom, NO 2 , N 3 , OH, SH, NH 2 or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl or heteroaralkyl group; R 5 is a hydrogen atom, NO 2 , N 3 , OH, SH, NH 2 or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl or heteroaralkyl group; B is NH, O, S, C=O, NR 14 or CR 15 R 16 and R 11 is an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 12 is an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 12 is attached to the pyrimidine ring of formula (II) via a carbon-carbon bond, R 13 is a hydrogen atom, a halogen atom, NO 2 , N 3 , OH, SH, NH 2 or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl or heteroaralkyl group; R 14 is an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkyl-cycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 15 is a hydrogen atom, NO 2 , N 3 , OH, SH, NH 2 or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl or heteroaralkyl group; R 16 is a hydrogen atom, NO 2 , N 3 , OH, SH, NH 2 or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl or heteroaralkyl group; R 21 each of which is optionally substituted aryl or heteroaryl; R 22 H, halo, OH, CN, CF 3 , C 1~6 Alkyl, C 1~6 Alkoxy, C 1~6 Haloalkyl, C 1~6 Thioalkyl, C 3~8 Cycloalkyl, C 2~8 heterocycloalkyl, aryl, or heteroaryl; Y is aryl or 5- or 6-membered heteroaryl; C 1~6 Alkyl, C 1~6 Alkoxy, C 1~6 Haloalkyl, C 1~6 Thioalkyl, C 3~8 Cycloalkyl, C 2~8 Each of heterocycloalkyl, aryl and heteroaryl is selected from halo, OH, CN, CF 3 , N.H. 2 , NO 2 , C 1~6 Alkyl, C 1~6 Haloalkyl, C 1~6 Thioalkyl, C 3~8 Cycloalkyl, C 2~8 Heterocycloalkyl, C 2~8 Heterocycloalkenyl, C 2~6 Alkenyl, C 2~6 Alkynyl, C 1~6 Alkoxy, C 1~6 Haloalkoxy, C 1~6 Alkylamino, C 2~6 Dialkylamino, C 7~12 Aralkyl, C 1~12 Heteroaralkyl, aryl, heteroaryl, —C(O)R, —C(O)OR, —C(O)NRR′, —C(O)NRS(O) 2 R', -C(O)NRS(O) 2 NR'R", -OR, -OC(O)NRR', -NRR', -NRC(O)R', -NRC(O)NR'R", -NRS(O) 2 R', -NRS(O) 2 NR'R", -S(O) 2 R, and -S(O) 2 and optionally substituted with one or more moieties selected from the group consisting of NRR; Each of R, R′, and R″ is independently H, halo, OH, C 1~6 Alkyl, C 1~6 Haloalkyl, C 1~6 Alkoxy, C 3~8 Cycloalkyl, C 2~8 heterocycloalkyl, aryl, or heteroaryl, or R and R′ or R′ and R″ together with the nitrogen to which they are attached are C 2~8 forming a heterocycloalkyl, R 31 is C(O)CH 2 R 33 , optionally substituted cycloalkyl, optionally substituted cycloheteroalkyl, optionally substituted cycloalkenyl, optionally substituted cycloheteroalkenyl, optionally substituted aryl, or optionally substituted heteroaryl; R 32 each instance of is independently halo, haloalkyl, optionally substituted alkoxyl, optionally substituted alkyl, optionally substituted heteroalkyl, optionally substituted alkenyl, optionally substituted heteroalkenyl; R 33 is optionally substituted cycloalkyl, optionally substituted cycloheteroalkyl, optionally substituted cycloalkenyl, optionally substituted cycloheteroalkenyl, optionally substituted aryl, or optionally substituted heteroaryl; Z is cycloalkyl, cycloheteroalkyl, cycloalkenyl, cycloheteroalkenyl, aryl, or heteroaryl, and Z is selected from two or three instances of R 2 Z may be aryl substituted with two or three instances of R 2 Z may be phenyl substituted with two or three instances of R 2 Z may be heteroaryl substituted with two or three instances of R 2 may be a 6-membered heteroaryl substituted with n is 0 to 5. or a pharmaceutically acceptable salt of any of the above compounds.
15. 1. A method of treating a patient having Parkinson's disease associated with wild-type LRRK2, comprising: receiving data identifying bis(monoacylglycerol)phosphate (BMP) and / or hemi-bis(monoacylglycerol)phosphate (hemi-BMP) levels in said patient; comparing said data with BMP and / or hemi-BMP levels from subjects without a neurological condition and having wild-type LRRK2; and if said patient has elevated levels of BMP and / or hemi-BMP compared to said subject, prescribing or providing one or more LRRK2 inhibitors to said patient.
16. The method of claim 15, wherein the BMP levels are measured in biological fluids of the patient and the subject.
17. 17. The method of claim 16, wherein the biological fluid is urine, blood, cerebrospinal fluid (CSF), bile, or saliva.
18. BMP, (i) di-22:6 BMP, (ii) di-18:1 BMP, (iii) 16:0 / 18:1 BMP, (iv) di-20:4 BMP, (v) 18:0 / 20:4 BMP, (vi) 2,2' di-22:6 BMP, and (vii) any combination thereof; 16. The method of claim 15, wherein the hemi-BMP is selected from the group consisting of: (i) hemi-BMP (18:1 / 18:1)_16:0, (ii) hemi-BMP (14:0 / 14:0)_14:0, (iii) hemi-BMP (18:1 / 18:1)_18:0, (iv) hemi-BMP (18:1 / 18:1)_18:1, and (iv) any combination thereof.
19. 16. The method of claim 15, wherein elevated levels of BMP and / or hemi-BMP in the patient indicate that the patient is responsive to one or more LRRK2 inhibitors.
20. 16. The method of claim 15, wherein the one or more LRRK2 inhibitors are selected from the group consisting of CZC-25146, CZC-54252, DNL151, DNL201, GNE-7915, GSK2578215A, HG-10-102-01, JH-II-127, K252A, K252B, LRRK2-IN-1, MLi-2, PF-06447475, and staurosporine.
21. The one or more LRRK2 inhibitors are represented by formula (I), (II), (III) and (IV): 【Chemistry 6】 【Chemistry 7】 (In the formula, A is NH, O, S, C=O, NR 3 or CR 4 R 5 and X is an optionally substituted arylene, heteroarylene, cycloalkylene, heterocycloalkylene, alkylcycloalkylene, heteroalkylcycloalkylene, aralkylene, or heteroaralkylene group; R 1 is an optionally substituted alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 2 is a hydrogen atom, a halogen atom, NO 2 , N 3 , OH, SH, NH 2 or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl or heteroaralkyl group; R 3 is an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkyl-cycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 4 is a hydrogen atom, NO 2 , N 3 , OH, SH, NH 2 or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl or heteroaralkyl group; R 5 is a hydrogen atom, NO 2 , N 3 , OH, SH, NH 2 or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl or heteroaralkyl group; B is NH, O, S, C=O, NR 14 or CR 15 R 16 and R 11 is an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 12 is an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 12 is attached to the pyrimidine ring of formula (II) via a carbon-carbon bond, R 13 is a hydrogen atom, a halogen atom, NO 2 , N 3 , OH, SH, NH 2 or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl or heteroaralkyl group; R 14 is an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkyl-cycloalkyl, heterocycloalkyl, aralkyl, or heteroaralkyl group; R 15 is a hydrogen atom, NO 2 , N 3 , OH, SH, NH 2 or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl or heteroaralkyl group; R 16 is a hydrogen atom, NO 2 , N 3 , OH, SH, NH 2 or an alkyl, alkenyl, alkynyl, heteroalkyl, aryl, heteroaryl, cycloalkyl, alkylcycloalkyl, heteroalkylcycloalkyl, heterocycloalkyl, aralkyl or heteroaralkyl group; R 21 each of which is optionally substituted aryl or heteroaryl; R 22 H, halo, OH, CN, CF 3 , C 1~6 Alkyl, C 1~6 Alkoxy, C 1~6 Haloalkyl, C 1~6 Thioalkyl, C 3~8 Cycloalkyl, C 2~8 heterocycloalkyl, aryl, or heteroaryl; Y is aryl or 5- or 6-membered heteroaryl; C 1~6 Alkyl, C 1~6 Alkoxy, C 1~6 Haloalkyl, C 1~6 Thioalkyl, C 3~8 Cycloalkyl, C 2~8 Each of heterocycloalkyl, aryl and heteroaryl is selected from halo, OH, CN, CF 3 , N.H. 2 , NO 2 , C 1~6 Alkyl, C 1~6 Haloalkyl, C 1~6 Thioalkyl, C 3~8 Cycloalkyl, C 2~8 Heterocycloalkyl, C 2~8 Heterocycloalkenyl, C 2~6 Alkenyl, C 2~6 Alkynyl, C 1~6 Alkoxy, C 1~6 Haloalkoxy, C 1~6 Alkylamino, C 2~6 Dialkylamino, C 7~12 Aralkyl, C 1~12 Heteroaralkyl, aryl, heteroaryl, —C(O)R, —C(O)OR, —C(O)NRR′, —C(O)NRS(O) 2 R', -C(O)NRS(O) 2 NR'R", -OR, -OC(O)NRR', -NRR', -NRC(O)R', -NRC(O)NR'R", -NRS(O) 2 R', -NRS(O) 2 NR'R", -S(O) 2 R, and -S(O) 2 and optionally substituted with one or more moieties selected from the group consisting of NRR; Each of R, R′, and R″ is independently H, halo, OH, C 1~6 Alkyl, C 1~6 Haloalkyl, C 1~6 Alkoxy, C 3~8 Cycloalkyl, C 2~8 heterocycloalkyl, aryl, or heteroaryl, or R and R′ or R′ and R″ together with the nitrogen to which they are attached are C 2~8 forming a heterocycloalkyl, R 31 is C(O)CH 2 R 33 , optionally substituted cycloalkyl, optionally substituted cycloheteroalkyl, optionally substituted cycloalkenyl, optionally substituted cycloheteroalkenyl, optionally substituted aryl, or optionally substituted heteroaryl; R 32 each instance of is independently halo, haloalkyl, optionally substituted alkoxyl, optionally substituted alkyl, optionally substituted heteroalkyl, optionally substituted alkenyl, optionally substituted heteroalkenyl; R 33 is optionally substituted cycloalkyl, optionally substituted cycloheteroalkyl, optionally substituted cycloalkenyl, optionally substituted cycloheteroalkenyl, optionally substituted aryl, or optionally substituted heteroaryl; Z is cycloalkyl, cycloheteroalkyl, cycloalkenyl, cycloheteroalkenyl, aryl, or heteroaryl, and Z is selected from two or three instances of R 2 Z may be aryl substituted with two or three instances of R 2 Z may be phenyl substituted with two or three instances of R 2 Z may be heteroaryl substituted with two or three instances of R 2 may be a 6-membered heteroaryl substituted with n is 0 to 5. or a pharmaceutically acceptable salt of any of the above compounds.