Method and system for determining autism spectrum disorder risk
By analyzing blood metabolite distribution tails and using threshold values, the method addresses the delay in ASD diagnosis, enabling early risk prediction and timely interventions.
Patent Information
- Application Number
- JP2025140823
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2015-02-27
- Filing Date
- 2025-08-26
- Publication Date
- 2025-11-26
AI Technical Summary
Current diagnostic methods for autism spectrum disorder (ASD) are delayed, often occurring after 54 months, missing the critical early intervention window, and require intensive specialist evaluations.
Analyzing population distribution curves of blood metabolite levels to identify tail effects, using threshold values for metabolites that are enriched in the distribution tails, allowing for early risk prediction of ASD through a statistical analysis that incorporates metabolites with low mutual information.
Enhances the predictability of ASD risk assessment, enabling earlier diagnosis with a median age of around 36 months, improving the chances of timely therapeutic interventions.
Smart Images

Figure 2025172836000029 
Figure 2025172836000030 
Figure 2025172836000031
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Patent Application No. 14 / 633,558, a continuation-in-part of U.S. Patent Application No. 14 / 493,141, filed September 22, 2014, which claims the benefit of U.S. Provisional Patent Application No. 61 / 978,773, filed April 11, 2014, and U.S. Provisional Patent Application No. 62 / 002,169, filed May 22, 2014. The contents of each of these applications are incorporated herein by reference in their entirety.
[0002] Technical Field The present invention relates generally to predicting risk for autism spectrum disorder (ASD) and other disorders. [Background technology]
[0003] Autism spectrum disorder (ASD) is a pervasive developmental disorder characterized by deficits in interactive social interactions, language difficulties, as well as repetitive behaviors and restricted interests, often manifesting during the first three years of life. The etiology of ASD is poorly understood but is thought to be multifactorial, with both genetic and environmental factors contributing to the development of the disorder.
[0004] Data show that the average age at which parents begin to suspect ASD in their child is 20 months, yet the median age of diagnosis is after 54 months. From a clinical perspective, a key challenge is to determine as early as possible whether a child has an ASD and whether they require specialist referral for an autism treatment plan. Summary of the Invention [Means for solving the problem]
[0005] A diagnosis of ASD is typically made by developmental pediatricians and other specialists only after a careful assessment of the child using criteria detailed in the Diagnostic and Statistical Manual of Mental Disorders. A reliable diagnosis often requires intensive subject evaluation by multiple specialists, including developmental pediatricians, neurologists, psychiatrists, psychologists, speech and hearing specialists, and occupational therapists. Furthermore, despite the fact that the average age at which parents suspect ASD is as early as 20 months, the median age at ASD diagnosis is 54 months. According to the Centers for Disease Control (CDC), only 18% of children diagnosed with ASD are identified by the age of 36 months. Unfortunately, young children with undiagnosed ASD miss the opportunity to benefit from early therapeutic intervention during a critical period of childhood development. Medical diagnostic testing is needed to reliably determine ASD risk, particularly to identify younger children earlier, when therapeutic interventions are more likely to be effective.
[0006] The present invention is based on the discovery that analyzing distribution curves of analytes, such as metabolites, measured within and across populations provides information that can be used to build or improve classifiers for predicting the risk of conditions or disorders, such as ASD. In particular, analyzing population distribution curves of blood metabolite levels can facilitate predicting the risk of autism spectrum disorder (ASD) in subjects. For example, analyzing population distribution curves of blood metabolite levels can be used to distinguish between autism spectrum disorder (ASD) and non-ASD developmental disorders, such as developmental delay (DD), that are not caused by autism spectrum disorder in subjects.
[0007] Statistical analysis of biomarkers that distinguish two groups usually assumes that the two populations differ in their mean biomarker levels, and that the variation around this mean is due to experimental and / or population variations, which are best characterized by a Gaussian distribution. Contrary to this baseline model, it has been observed herein that for some analytes, but not others, the distribution in ASD, or sometimes DD, is best characterized by being composed of multiple sub-distributions: one sub-distribution that is essentially undifferentiated from other health conditions (for example, when the ASD distribution and the DD distribution are not differentiated), and another sub-distribution that is far from the mean in a minority of subjects, such as the "tail" of the combined distribution for that population. This insight leads to a significantly different analytical framework from the baseline, and it has been found that for certain analytes, better results can be achieved by defining thresholds based on the upper or lower part of the population distribution, for example, by establishing a ranking that does not require a standard Gaussian distribution model.
[0008] Thus, herein, a metabolite is described as exhibiting "tail enrichment" or a "tail" effect when samples from a particular population (e.g., either ASD or DD) are enriched in the distal portion of the distribution curve of metabolite levels for that metabolite. Information from assessing the presence, absence, and / or direction (higher or lower) of the tail effect in a metabolite distribution curve can be used to predict the risk of ASD. For certain metabolites, metabolite levels corresponding to the upper or lower portions (e.g., deciles) of the distribution curve, i.e., within the "tail" of the distribution curve (whether in the "right tail" or "left tail"), have been found to be highly informative for the presence or absence of ASD.
[0009] Furthermore, risk prediction has been found to improve as multiple metabolites with low degrees of overlapping mutual information are incorporated. For example, for the assessment of ASD, there are certain groups of metabolites that provide complementary diagnostic / risk assessment information. That is, ASD-positive individuals identifiable by analysis of the levels of a first metabolite (e.g., individuals within the identified tail of the first metabolite) are not the same as ASD-positive individuals identifiable by analysis of a second metabolite (or there may be a low but non-zero degree of overlap). While not wishing to be bound by any particular theory, this finding may reflect the multifaceted nature of ASD itself.
[0010] Thus, in certain embodiments, a risk assessment method involves identifying whether a subject falls into any of a number of identified metabolite tails that include multiple metabolites, e.g., when predictor variables in different metabolite tails are at least partially exclusive, e.g., when multiple metabolites have low mutual information, such that incorporating them improves risk prediction. A classifier has a predetermined level of predictability, e.g., in the form of AUC, i.e., the area under the ROC curve, which plots the false positive rate (1-specificity) against the true positive rate (sensitivity) for that classifier, and the AUC increases when metabolites with low mutual information are added to the classifier that represent tail effects.
[0011] In some embodiments, the present invention arises from the discovery that certain threshold values of metabolite levels in blood can be used to facilitate risk prediction of autism spectrum disorder (ASD) in a subject. In certain embodiments, these threshold values of metabolites, inferred from assessing the presence, absence, and / or direction (higher or lower) of tail effects in the distribution curve of metabolites, are used to predict the risk of ASD. In certain embodiments, these threshold values can be at either the upper or lower end of the distribution of metabolite levels in a population. For certain metabolites, it has been discovered that metabolite levels above the upper threshold value and / or below the lower threshold value are highly informative for the presence or absence of ASD.
[0012] In some embodiments, the levels of these metabolites are useful in distinguishing ASD from other forms of developmental delay (eg, developmental delay (DD) not attributable to autism spectrum disorder).
[0013] In one aspect, the invention provides a method of distinguishing between autism spectrum disorder (ASD) and non-ASD developmental delay (DD) in a subject, comprising: (i) measuring a level of a first metabolite of a plurality of metabolites from a sample obtained from the subject, wherein a population distribution of the first metabolite has been previously characterized in a first population of subjects with ASD and a second population of subjects with non-ASD developmental delay (DD), wherein the first metabolite has been predetermined to represent an ASD tail effect and / or a DD tail effect, each tail effect comprising an associated right tail or left tail enriched in members of the corresponding (ASD or DD) population, wherein the first metabolite represents an ASD tail effect with a right tail; and wherein if the level of the first metabolite in the sample is greater than a predetermined upper (minimum) threshold defining a right tail enriched in members of the first (ASD) population, then measuring a level of the first metabolite in the sample. the level of the first metabolite in the sample is within the ASD tail, where the first metabolite represents an ASD tail effect with a left tail, if the level of the first metabolite in the sample is less than a predetermined lower (maximum) threshold defining the left tail enriched in members of the first (ASD) population, and the level of the first metabolite in the sample is within the ASD tail, where the first metabolite represents a DD tail effect with a right tail, if the level of the first metabolite in the sample is greater than a predetermined upper (minimum) threshold defining the right tail enriched in members of the second (DD) population, and the level of the first metabolite in the sample is within the DD tail, where the first metabolite represents a DD tail effect with a left tail, if the level of the first metabolite in the sample is less than a predetermined lower (maximum) threshold defining the left tail enriched in members of the second (DD) population;(ii) measuring the level of at least one additional metabolite of the plurality of metabolites from the sample, wherein the population distribution of each of the at least one additional metabolite has been previously characterized in the first population and the second population and predetermined to exhibit at least one of an ASD tail effect and a DD tail effect, and for each of the at least one additional metabolite, identifying whether the level of said metabolite in the sample is within the corresponding ASD tail and / or DD tail according to step (i); and (iii) determining, using a predetermined level of predictability based on the identified ASD tail and / or identified DD tail within which the sample falls for the metabolites analyzed in steps (i) and (ii), (a) that the subject has ASD and does not have DD, or (b) that the subject has DD and does not have ASD;
[0014] In certain embodiments, the first metabolite is predetermined to represent an ASD tail effect with an associated upper (minimum) or lower (maximum) threshold, the threshold being predetermined so that the odds of an unknown classification sample (previously uncharacterized sample) meeting this criterion being ASD relative to DD are 1.6:1 or greater with p≦0.3. In certain embodiments, the odds are 2:1 or greater, or 2.5:1 or greater, or 2.75:1 or greater, or 3:1 or greater, or 3.25:1 or greater, or 3.5:1 or greater, or 3.75:1 or greater, or 4:1 or greater. In any of the foregoing, the p-value (value of statistical significance) meets p≦0.3, or p≦0.25, or p≦0.2, or p≦0.15, or p≦0.1, or p≦0.05.
[0015] In certain embodiments, the first metabolite is predetermined to exhibit a DD tail effect with an associated upper (minimum) or lower (maximum) threshold, the threshold being predetermined such that the odds of an unknown classification sample (previously uncharacterized sample) meeting this criterion being DD relative to ASD are 1.6:1 or greater with p≦0.3. In certain embodiments, the odds are 2:1 or greater, or 2.5:1 or greater, or 2.75:1 or greater, or 3:1 or greater, or 3.25:1 or greater, or 3.5:1 or greater, or 3.75:1 or greater, or 4:1 or greater. In any of the foregoing, the p-value (value of statistical significance) meets p≦0.3, or p≦0.25, or p≦0.2, or p≦0.15, or p≦0.1, or p≦0.05.
[0016] In certain embodiments, the predetermined level of predictability corresponds to a receiver operating characteristic (ROC) curve plotting the false positive rate (1-specificity) against the true positive rate (sensitivity) with an AUC (area under the curve) of at least 0.70.
[0017] In certain embodiments, the predetermined upper (minimum) threshold for one or more of the metabolites is a percentile between the 85th and 95th percentile (e.g., about the 90th percentile, or about the 85th, 86th, 87th, 88th, 89th, 91st, 92nd, 93rd, 94th, or 95th percentile, rounded down to the nearest integer), and the predetermined lower (maximum) threshold for one or more of the metabolites is a percentile between the 10th and 20th percentile (e.g., about the 15th percentile, or about the 10th, 11th, 12th, 13th, 14th, 16th, 17th, 18th, 19th, or 20th percentile, rounded down to the nearest integer).
[0018] In certain embodiments, the multiple metabolites are 5-hydroxyindole acetate (5-HIAA), 1,5-anhydroglucitol (1,5-AG), 3-(3-hydroxyphenyl)propionate, 3-carboxy-4-methyl-5-propyl-2-furanpropanoate (CMPF), 3-indoxyl sulfate, 4-ethylphenyl sulfate, 8-hydroxyoctanoate, gamma-CEHC, hydroxyisovaleroylcarnitine (C5), indole acetate , isovalerylglycine, lactate, N1-methyl-2-pyridone-5-carboxamide, p-cresol sulfate, pantothenate (vitamin B5), phenylacetylglutamine, pipecolate, xanthine, hydroxychlorothalonil, octenoylcarnitine, and 3-hydroxyhipprete. The compound contains at least two metabolites that are metabolized.
[0019] In certain embodiments, the plurality of metabolites comprises at least two metabolites selected from the group consisting of phenylacetylglutamine, xanthine, octenoylcarnitine, p-cresol sulfate, isovalerylglycine, gamma-CEHC, indole acetate, pipecolate, 1,5-anhydroglucitol (1,5-AG), lactate, 3-(3-hydroxyphenyl)propionate, 3-indoxyl sulfate, pantothenate (vitamin B5), and hydroxy-chlorothalonil.
[0020] In certain embodiments, the plurality of metabolites comprises at least three metabolites selected from the group consisting of phenylacetylglutamine, xanthine, octenoylcarnitine, p-cresol sulfate, isovalerylglycine, gamma-CEHC, indole acetate, pipecolate, 1,5-anhydroglucitol (1,5-AG), lactate, 3-(3-hydroxyphenyl)propionate, 3-indoxyl sulfate, pantothenate (vitamin B5), and hydroxy-chlorothalonil.
[0021] In certain embodiments, the plurality of metabolites comprises at least one metabolite pair selected from the pairs listed in Table 6.
[0022] In certain embodiments, the plurality of metabolites comprises at least one metabolite triad selected from the triads listed in Table 7.
[0023] In certain embodiments, the plurality of metabolites includes at least one pair of metabolites that, when combined together as a set of two metabolites, provides an AUC of at least 0.62 (e.g., at least about 0.63, 0.64, or 0.65), where AUC is the area under the ROC curve plotting the false positive rate (1 - specificity) against the true positive rate (sensitivity) for a classifier based only on the set of two metabolites.
[0024] In certain embodiments, the plurality of metabolites includes at least one metabolite triplet that, when combined together as a set of three metabolites, provides an AUC of at least 0.66 (e.g., at least about 0.67 or 0.68), where AUC is the area under the ROC curve plotting the false positive rate (1 - specificity) against the true positive rate (sensitivity) for a classifier based solely on that set of three metabolites.
[0025] In another aspect, the invention provides a method of determining risk of autism spectrum disorder (ASD) in a subject, comprising: (i) analyzing a level of a first metabolite of a plurality of metabolites from a sample obtained from the subject, wherein the population distribution of the first metabolite has been previously characterized in a reference population of subjects of known classification, wherein the first metabolite has been predetermined to represent an ASD tail effect with an associated right tail or left tail enriched in ASD members, wherein the first metabolite represents an ASD tail effect with a right tail, and wherein the level of the first metabolite in the sample is within the ASD tail, wherein the first metabolite represents an ASD tail effect with a left tail, if the level of the first metabolite in the sample is greater than a predetermined upper (minimum) threshold defining a right tail enriched in members of the ASD population. (ii) measuring the level of at least one additional metabolite of the plurality of metabolites from the sample, wherein the population distribution of each of the at least one additional metabolite has been previously characterized in a reference population and predetermined to represent an ASD tail effect, and for each of the at least one additional metabolite, identifying whether the level of said metabolite in the sample is within the corresponding ASD tail according to step (i); and (iii) determining the subject's risk of having an ASD using a predetermined level of predictability based on the identified ASD tail within which the sample falls for the metabolites analyzed in steps (i) and (ii).
[0026] In certain embodiments, the first metabolite is predetermined to represent an ASD tail effect with an associated upper (minimum) or lower (maximum) threshold, the threshold being predetermined so that the odds of an unknown classification sample (previously uncharacterized sample) meeting this criterion being ASD relative to DD are 1.6:1 or greater with p≦0.3. In certain embodiments, the odds are 2:1 or greater, or 2.5:1 or greater, or 2.75:1 or greater, or 3:1 or greater, or 3.25:1 or greater, or 3.5:1 or greater, or 3.75:1 or greater, or 4:1 or greater. In any of the foregoing, the p-value (value of statistical significance) meets p≦0.3, or p≦0.25, or p≦0.2, or p≦0.15, or p≦0.1, or p≦0.05.
[0027] In certain embodiments, the predetermined level of predictability corresponds to a receiver operating characteristic (ROC) curve plotting the false positive rate (1-specificity) against the true positive rate (sensitivity) with an AUC (area under the curve) of at least 0.70.
[0028] In certain embodiments, the plurality of metabolites comprises at least two metabolites selected from the group consisting of 5-hydroxyindole acetate (5-HIAA), 1,5-anhydroglucitol (1,5-AG), 3-(3-hydroxyphenyl)propionate, 3-carboxy-4-methyl-5-propyl-2-furanpropanoate (CMPF), 3-indoxyl sulfate, 4-ethylphenyl sulfate, 8-hydroxyoctanoate, gamma-CEHC, hydroxyisovaleroylcarnitine (C5), indole acetate, isovalerylglycine, lactate, N1-methyl-2-pyridone-5-carboxamide, p-cresol sulfate, pantothenate (vitamin B5), phenylacetylglutamine, pipecolate, xanthine, hydroxy-chlorothalonil, octenoylcarnitine, and 3-hydroxyhipprete.
[0029] In another aspect, the invention provides a method for determining risk of autism spectrum disorder (ASD) in a subject, comprising: (i) analyzing levels of a plurality of metabolites in a sample obtained from the subject, wherein the plurality of metabolites are 5-hydroxyindole acetate (5-HIAA), 1,5-anhydroglucitol (1,5-AG), 3-(3-hydroxyphenyl)propionate, 3-carboxy-4-methyl-5-propyl-2-furanpropanoate (CMPF), 3-indoxyl sulfate, 4-ethylphenyl sulfate, 8-hydroxyoctanoate, gamma-CEHC, and (ii) determining the subject's risk of having ASD based on the quantified levels of the plurality of metabolites.
[0030] In certain embodiments, the subject is about 54 months of age or younger, hi certain embodiments, the subject is about 36 months of age or younger.
[0031] In certain embodiments, the plurality of metabolites comprises at least two metabolites selected from the group consisting of phenylacetylglutamine, xanthine, octenoylcarnitine, p-cresol sulfate, isovalerylglycine, gamma-CEHC, indole acetate, pipecolate, 1,5-anhydroglucitol (1,5-AG), lactate, 3-(3-hydroxyphenyl)propionate, 3-indoxyl sulfate, pantothenate (vitamin B5), and hydroxy-chlorothalonil.
[0032] In certain embodiments, the plurality of metabolites comprises at least three metabolites selected from the group consisting of phenylacetylglutamine, xanthine, octenoylcarnitine, p-cresol sulfate, isovalerylglycine, gamma-CEHC, indole acetate, pipecolate, 1,5-anhydroglucitol (1,5-AG), lactate, 3-(3-hydroxyphenyl)propionate, 3-indoxyl sulfate, pantothenate (vitamin B5), and hydroxy-chlorothalonil.
[0033] In certain embodiments, the plurality of metabolites comprises at least one metabolite pair selected from the pairs listed in Table 6.
[0034] In certain embodiments, the plurality of metabolites comprises at least one metabolite triad selected from the triads listed in Table 7.
[0035] In certain embodiments, the plurality of metabolites includes at least one pair of metabolites that, when combined together as a set of two metabolites, provides an AUC of at least 0.62 (e.g., at least about 0.63, 0.64, or 0.65), where AUC is the area under the ROC curve plotting the false positive rate (1 - specificity) against the true positive rate (sensitivity) for a classifier based only on the set of two metabolites.
[0036] In certain embodiments, the plurality of metabolites includes at least one metabolite triplet that, when combined together as a set of three metabolites, provides an AUC of at least 0.66 (e.g., at least about 0.67 or 0.68), where AUC is the area under the ROC curve plotting the false positive rate (1 - specificity) against the true positive rate (sensitivity) for a classifier based solely on that set of three metabolites.
[0037] In certain embodiments, the sample is a plasma sample.
[0038] In certain embodiments, measuring the level of a metabolite comprises performing mass spectrometry. In certain embodiments, performing mass spectrometry comprises performing one or more members selected from the group consisting of pyrolysis mass spectrometry, Fourier transform infrared spectrometry, Raman spectrometry, gas chromatography mass spectrometry, high pressure liquid chromatography / mass spectrometry (HPLC / MS), liquid chromatography (LC)-electrospray mass spectrometry, cap-LC-tandem electrospray mass spectrometry, and ultra-high performance liquid chromatography / electrospray ionization tandem mass spectrometry.
[0039] In another aspect, the present invention provides a method for distinguishing between autism spectrum disorder (ASD) and non-ASD developmental delay (DD) in a subject, comprising: (i) analyzing levels of a plurality of metabolites in a sample obtained from the subject, wherein the plurality of metabolites are 5-hydroxyindole acetate (5-HIAA), 1,5-anhydroglucitol (1,5-AG), 3-(3-hydroxyphenyl)propionate, 3-carboxy-4-methyl-5-propyl-2-furanpropanoate (CMPF), 3-indoxyl sulfate, 4-ethylphenyl sulfate, 8-hydroxyoctanoate, gamma-CEHC, hydroxyisovaleroylcarnitine (C5), indole acetate, isovalerylglycine, lactate, N1-methyl-2-pyridone-5-carboxamide ... and (ii) determining (a) that the subject has ASD and does not have DD, or (b) that the subject has DD and does not have ASD by comparing levels of the plurality of metabolites from a sample from a subject with a predetermined threshold (e.g., a threshold determined from a reference population of samples with known classification) with the predictability of the predetermined levels.
[0040] In certain embodiments, the present invention provides methods for analyzing metabolites by assigning weights to different metabolites to reflect their respective functions in risk prediction. In some embodiments, the weight assignment can be inferred from the biological function of the metabolites (e.g., the pathway to which they belong), their clinical utility, or their importance based on statistical or epidemiological analysis.
[0041] In certain embodiments, the present invention provides methods for measuring metabolites using a variety of techniques, including, but not limited to, chromatographic assays, mass spectrometry assays, fluorimetric assays, electrophoretic assays, immunoaffinity assays, and immunochemical assays.
[0042] In certain embodiments, the present invention provides a method for determining risk of autism spectrum disorder (ASD) in a subject, the method comprising: analyzing levels of a plurality of metabolites derived from a sample from the subject; and determining, based on the quantified levels of the plurality of metabolites using a predetermined level predictability, whether the subject has an ASD rather than a non-ASD developmental disorder.
[0043] In certain embodiments, the plurality of metabolites includes at least one metabolite selected from the group consisting of 5-hydroxyindole acetate (5-HIAA), 1,5-anhydroglucitol (1,5-AG), 3-(3-hydroxyphenyl)propionate, 3-carboxy-4-methyl-5-propyl-2-furanpropanoate (CMPF), 3-indoxyl sulfate, 4-ethylphenyl sulfate, 8-hydroxyoctanoate, gamma-CEHC, hydroxyisovaleroylcarnitine (C5), indole acetate, isovalerylglycine, lactate, N1-methyl-2-pyridone-5-carboxamide, p-cresol sulfate, pantothenate (vitamin B5), phenylacetylglutamine, pipecolate, xanthine, hydroxy-chlorothalonil, octenoylcarnitine, 3-hydroxyhiprate, and combinations thereof.
[0044] In certain embodiments, the plurality of metabolites includes at least two metabolites selected from the group consisting of 5-hydroxyindole acetate (5-HIAA), 1,5-anhydroglucitol (1,5-AG), 3-(3-hydroxyphenyl)propionate, 3-carboxy-4-methyl-5-propyl-2-furanpropanoate (CMPF), 3-indoxyl sulfate, 4-ethylphenyl sulfate, 8-hydroxyoctanoate, gamma-CEHC, hydroxyisovaleroylcarnitine (C5), indole acetate, isovalerylglycine, lactate, N1-methyl-2-pyridone-5-carboxamide, p-cresol sulfate, pantothenate (vitamin B5), phenylacetylglutamine, pipecolate, xanthine, hydroxy-chlorothalonil, octenoylcarnitine, 3-hydroxyhiprate, and combinations thereof.
[0045] In certain embodiments, the plurality of metabolites are 5-hydroxyindole acetate (5-HIAA), 1,5-anhydroglucitol (1,5-AG), 3-(3-hydroxyphenyl)propionate, 3-carboxy-4-methyl-5-propyl-2-furanpropanoate (CMPF), 3-indoxyl sulfate, 4-ethylphenyl sulfate, 8-hydroxyoctanoate, gamma-CEHC, hydroxyisovaleroylcarnitine (C5), indole acetate, isovalerate, or hydroxyisovaleroylcarnitine (C5). The metabolites include at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten metabolites selected from the group consisting of hydroxylglycine, lactate, N1-methyl-2-pyridone-5-carboxamide, p-cresol sulfate, pantothenate (vitamin B5), phenylacetylglutamine, pipecolate, xanthine, hydroxychlorothalonil, octenoylcarnitine, 3-hydroxyhipprete, and combinations thereof.
[0046] In certain embodiments, the plurality of metabolites includes additional metabolites, hi some embodiments, the plurality of metabolites includes more than 21 metabolites.
[0047] In certain embodiments, the present invention provides a method for distinguishing between an autism spectrum disorder (ASD) and a non-ASD developmental disorder in a subject, the method comprising analyzing levels of a plurality of metabolites from a sample from the subject, comparing the levels of the metabolites to their respective population distributions in a reference population, and determining whether the subject has an ASD as opposed to a non-ASD developmental disorder by comparing the levels of the plurality of metabolites from the sample from the subject to pre-characterized levels and / or population distributions of the plurality of metabolites in the reference population using a predetermined level of predictability.
[0048] For example, in certain embodiments, the present invention provides a diagnostic standard comprising at least one metabolite capable of predicting the risk of ASD in a subject, wherein the ROC curve has an AUC of at least 0.60, at least 0.65, at least 0.70, at least 0.75, at least 0.80, at least 0.85, or at least 0.90. The AUC is the area under the ROC curve, which plots the false positive rate (1-specificity) against the true positive rate (sensitivity) for a classifier.
[0049] In certain embodiments, the at least one metabolite for analysis is selected from the group consisting of 5-hydroxyindole acetate (5-HIAA), 1,5-anhydroglucitol (1,5-AG), 3-(3-hydroxyphenyl)propionate, 3-carboxy-4-methyl-5-propyl-2-furanpropanoate (CMPF), 3-indoxyl sulfate, 4-ethylphenyl sulfate, 8-hydroxyoctanoate, gamma-CEHC, hydroxyisovaleroylcarnitine (C5), indole acetate, isovalerylglycine, lactate, N1-methyl-2-pyridone-5-carboxamide, p-cresol sulfate, pantothenate (vitamin B5), phenylacetylglutamine, pipecolate, xanthine, hydroxy-chlorothalonil, octenoylcarnitine, 3-hydroxyhipprete, and combinations thereof.
[0050] In certain embodiments, the at least one metabolite for analysis is 5-hydroxyindole acetate (5-HIAA), 1,5-anhydroglucitol (1,5-AG), 3-(3-hydroxyphenyl)propionate, 3-carboxy-4-methyl-5-propyl-2-furanpropanoate (CMPF), 3-indoxyl sulfate, 4-ethylphenyl sulfate, 8-hydroxyoctanoate, gamma-CEHC, hydroxyisovaleroylcarnitine (C5), indole acetate, isovalerylglycine, lactate, N1-methyl-2-pyridone-5-carbohydrate, N1-methyl-2-pyridone-5-carboxylate ... The metabolites include at least two or more (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21) members selected from the group consisting of voxamide, p-cresol sulfate, pantothenate (vitamin B5), phenylacetylglutamine, pipecolate, xanthine, hydroxychlorothalonil, octenoylcarnitine, and 3-hydroxyhipprete, and non-ASD and ASD population distribution curves are established for each of the metabolites (e.g., each of the metabolites demonstrating a tail effect).
[0051] In certain embodiments, the metabolites for analysis are selected from the group consisting of gamma-CEHC, xanthine, p-cresol sulfate, octenoylcarnitine, phenylacetylglutamine, and combinations thereof.
[0052] In certain embodiments, the metabolite for analysis is gamma-CEHC.
[0053] In certain embodiments, the metabolite for analysis is a xanthine.
[0054] In certain embodiments, the metabolite for analysis is p-cresol sulfate.
[0055] In certain embodiments, the metabolite for analysis is octenoylcarnitine.
[0056] In certain embodiments, the metabolite for analysis is phenylacetylglutamine.
[0057] In certain embodiments, the metabolite for analysis is isovalerylglycine.
[0058] In certain embodiments, the metabolite for analysis is pipecolate.
[0059] In certain embodiments, the metabolite for analysis is indole acetate.
[0060] In certain embodiments, the metabolite for analysis is octenoylcarnitine.
[0061] In certain embodiments, the metabolite for analysis is hydroxy-chlorothalonil.
[0062] In certain embodiments, the plurality of metabolites includes at least a first metabolite and a second metabolite that are complementary (e.g., the ASD tail samples for the first and second metabolites are substantially non-overlapping, such that the predictors provided by the metabolites are partially exclusive and have low mutual information. In certain embodiments, risk prediction is improved by incorporating multiple metabolites with low mutual information.
[0063] In certain embodiments, the plurality of metabolites includes two metabolites, wherein the two metabolites, combined together as a set of two metabolites, provide an AUC of at least 0.62, 0.63, 0.64, or 0.65.
[0064] In certain embodiments, the plurality of metabolites includes three metabolites, wherein the three metabolites, combined together as a set of three metabolites, provide an AUC of at least 0.66, 0.67, or 0.68.
[0065] In certain embodiments, the present invention provides a method for distinguishing between autism spectrum disorder (ASD) and non-ASD developmental disorders in a subject by analyzing the levels of two groups of predefined metabolites. In certain embodiments, the first group of metabolites represents metabolites closely associated with ASD, while the second group of metabolites represents metabolites associated with a control condition (e.g., DD). By analyzing both groups of metabolites from a sample from a subject, the subject's risk of having ASD rather than a control condition can be determined by various methods described in this disclosure. For example, this can be achieved by comparing the aggregated ASD tail effect for the first group of metabolites with the aggregated non-ASD tail effect for the second group of metabolites.
[0066] In certain embodiments, the present invention provides a method for distinguishing between autism spectrum disorder (ASD) and non-ASD developmental delay (DD) in a subject, comprising the steps of: (i) measuring levels of a plurality of metabolites in a sample obtained from the subject, wherein the plurality of metabolites include xanthine, gamma-CEHC, hydroxy-chlorothalonil, 5-hydroxyindole acetate (5-HIAA), indole acetate, p-cresol sulfate, 1,5-anhydroglucitol (1,5-AG), 3-(3-hydroxyphenyl)propionate, 3-carboxy-4-methyl-5-propyl-2-furanpropanoate (CMPF), 3-indoxyl sulfate, 4-ethylphenyl sulfate, hydroxyisovaleroylcarnitine (C5), isovalerylglycine, lactate, N1-methyl-2-pyridone-5-carboxamide, pantothenate (vitamin B5), phenylalanine, and phenylalanine. and (ii) calculating the number of metabolites in the sample having levels at or below a predetermined threshold concentration that are (a) indicative of an ASD as defined in Table 9A (ASD left-tail effect) or (b) indicative of a DD as defined in Table 9B (DD left-tail effect); and / or (iii) calculating the number of metabolites in the sample having levels at or above a predetermined threshold concentration that are (a) indicative of an ASD as defined in Table 9A (ASD right-tail effect) or (b) indicative of a DD as defined in Table 9B (DD right-tail effect); and (iv) determining that the subject has an ASD or DD based on the numbers obtained in steps (ii) and / or (iii).
[0067] In certain embodiments, the present invention provides a method for determining whether a subject has or is at risk for ASD, comprising: (i) measuring levels of a plurality of metabolites in a sample obtained from the subject, the plurality of metabolites being xanthine, gamma-CEHC, hydroxy-chlorothalonil, 5-hydroxyindole acetate (5-HIAA), indole acetate, p-cresol sulfate, 1,5-anhydroglucitol (1,5-AG), 3-carboxy-4-methyl-5-propyl-2-furanpropanoic acid (ASD), methylparaben, ... and (ii) at least two metabolites selected from the group consisting of (a) xanthine at a level of 182.7 ng / ml or greater; (b) 20. (c) hydroxyl-chlorothalonil at a level of 3 ng / ml or greater; (c) 5-hydroxyindoleacetate at a level of 28.5 ng / ml or greater; (d) lactate at a level of 686,600 ng / ml or greater; (e) pantothenate at a level of 63.3 ng / ml or greater; (f) pipecolate at a level of 303.6 ng / ml or greater; (g) gamma-CEHC at a level of 32.0 ng / ml or less; (h) indoleacetate at a level of 141.4 ng / ml or less. (i) p-cresol sulfate at a level of 182.7 ng / ml or less; (j) 1,5-anhydroglucitol (1,5-AG) at a level of 10.3 ng / ml or less; (k) 3-carboxy-4-methyl-5-propyl-2-furanpropanoate (CMPF) at a level of 7.98 ng / ml or less; (l) 3-indoxyl sulfate at a level of 256.7 ng / ml or less; (m) 4-ethylphenyl sulfate at a level of 3.0 ng / ml or less; (n) 12.Two or more of the following: hydroxyisovaleroylcarnitine (C5) at a level of 9 ng / ml or less; (o) N1-methyl-2-pyridone-5-carboxamide at a level of 124.82 ng / ml or less; and (p) phenylacetylglutamine at a level of 166.4 ng / ml or less. and (iii) determining that the subject has or is at risk of having an ASD based on the metabolite levels detected in step (ii).
[0068] In some embodiments, the metabolite level detected in step (ii) is an approximate level.
[0069] In some embodiments, the present invention provides a method of diagnosing a subject as having or at risk of having an autism spectrum disorder (ASD), comprising the steps of: (i) determining levels of one or more metabolites in a sample obtained from the subject; (ii) comparing the determined levels of one or more metabolites to predetermined levels of said one or more metabolites, wherein said predetermined levels are indicative of the subject having or at risk of having an ASD; and (iii) diagnosing the subject as having or at risk of having an ASD based on a difference between the determined and predetermined levels of said one or more metabolites, wherein said one or more metabolites are 5-hydroxyindoline, 5-hydroxypropyltrimonial ... In one embodiment, the compound is selected from the group consisting of indole acetate (5-HIAA), 1,5-anhydroglucitol (1,5-AG), 3-(3-hydroxyphenyl)propionate, 3-carboxy-4-methyl-5-propyl-2-furanpropanoate (CMPF), 3-indoxyl sulfate, 4-ethylphenyl sulfate, 8-hydroxyoctanoate, gamma-CEHC, hydroxyisovaleroylcarnitine (C5), indole acetate, isovalerylglycine, lactate, N1-methyl-2-pyridone-5-carboxamide, p-cresol sulfate, pantothenate (vitamin B5), phenylacetylglutamine, pipecolate, xanthine, hydroxy-chlorothalonil, octenoylcarnitine, and 3-hydroxyhippreate.
[0070] In some embodiments, the present invention provides a method of diagnosing a subject as having or at risk of having an autism spectrum disorder (ASD), comprising the steps of: (i) determining a level of one or more metabolites in a sample obtained from the subject; (ii) comparing the determined level of one or more metabolites to a predetermined level of said one or more metabolites, wherein said predetermined level is indicative of the subject having or at risk of having an ASD;and (iii) diagnosing the subject as having or at risk of having an ASD based on a difference between the determined level and a predetermined level of the one or more metabolites, wherein the one or more metabolites are 5-hydroxyindole acetate (5-HIAA), 1,5-anhydroglucitol (1,5-AG), 3-(3-hydroxyphenyl)propionate, 3-carboxy-4-methyl-5-propyl-2-furanpropanoate (CMPF), 3-indoxyl sulfate, 4-ethylphenyl sulfate, 8-hydroxyoctanoate, gamma-CEHC, hydroxyisovaleroylcarnitine (C5), indole acetate, isovalerylglycine, lactate, N1-methyl-2-pyridone-5-carboxamide, p-cresol sulfate, pantothenate (vitamin B5), phenylacetylglutamine, pipecolate, xanthine, hydroxyisovaleroylcarnitine (C5), indole acetate, isovalerylglycine, hydroxyisovalero ... and 3-hydroxyhiprate, wherein the metabolite is selected from the group consisting of 5-hydroxyindole acetate (5-HIAA), 1,5-anhydroglucitol (1,5-AG), 3-(3-hydroxyphenyl)propionate, 3-carboxy-4-methyl-5-propyl-2-furanpropanoate (CMPF), 3-indoxyl sulfate, 4-ethylphenyl sulfate, 8- The method is not hydroxyoctanoate, gamma-CEHC, hydroxyisovaleroylcarnitine (C5), indole acetate, isovalerylglycine, lactate, N1-methyl-2-pyridone-5-carboxamide, p-cresol sulfate, pantothenate (vitamin B5), phenylacetylglutamine, pipecolate, xanthine, hydroxychlorothalonil, octenoylcarnitine, or 3-hydroxyhipprete;
[0071] In certain embodiments, the present invention provides a method for determining ASD risk in a subject by measuring both the levels of certain metabolites and genetic information from the subject. In some embodiments, the genetic information includes copy number variation (CNV) and / or fragile X (FXS) testing.
[0072] In additional embodiments, limitations described with respect to one particular aspect of the invention may apply to other aspects of the invention, for example, limitations of claims dependent on one independent claim may, in some embodiments, apply to another independent claim. For example, in an embodiment of the present invention, the following items are provided: (Item 1) 1. A method for distinguishing between autism spectrum disorder (ASD) and non-ASD developmental delay (DD) in a subject, comprising: (i) measuring the level of a plurality of metabolites in a sample obtained from said subject, the plurality of metabolites comprises at least two metabolites selected from the group consisting of xanthine, gamma-CEHC, hydroxy-chlorothalonil, 5-hydroxyindole acetate (5-HIAA), indole acetate, p-cresol sulfate, 1,5-anhydroglucitol (1,5-AG), 3-(3-hydroxyphenyl)propionate, 3-carboxy-4-methyl-5-propyl-2-furanpropanoate (CMPF), 3-indoxyl sulfate, 4-ethylphenyl sulfate, hydroxyisovaleroylcarnitine (C5), isovalerylglycine, lactate, N1-methyl-2-pyridone-5-carboxamide, pantothenate (vitamin B5), phenylacetylglutamine, pipecolate, 3-hydroxyhipprete, and combinations thereof; and (ii) (a) is indicative of ASD (ASD left-tail effect) as defined in Table 9A; or (b) is an indicator of DD (DD left-tail effect) as defined in Table 9B; calculating the number of metabolites in said sample having levels at or below a predetermined threshold concentration; and / or (iii) (a) is indicative of ASD (ASD right-tail effect) as defined in Table 9A; or (b) is an indicator of DD (DD right-tail effect) as defined in Table 9B; Calculating the number of metabolites in the sample having levels at or above a predetermined threshold concentration; and (iv) determining that the subject has ASD or DD based on the numbers obtained in steps (ii) and / or (iii). A method comprising: (Item 2) 2. The method of claim 1, wherein the plurality of metabolites comprises xanthine and gamma-CEHC. (Item 3) 2. The method of claim 1, wherein the sample is a plasma sample. (Item 4) 2. The method of claim 1, wherein the level of the metabolite is measured by mass spectrometry. (Item 5) Item 10. The method of item 1, wherein the subject is about 54 months of age or younger. (Item 6) Item 10. The method of item 1, wherein the subject is about 36 months of age or younger. (Item 7) 1. A method for determining whether a subject has or is at risk for ASD, comprising: (i) measuring the level of a plurality of metabolites in a sample obtained from said subject, the plurality of metabolites comprises at least two metabolites selected from the group consisting of xanthine, gamma-CEHC, hydroxy-chlorothalonil, 5-hydroxyindole acetate (5-HIAA), indole acetate, p-cresol sulfate, 1,5-anhydroglucitol (1,5-AG), 3-carboxy-4-methyl-5-propyl-2-furanpropanoate (CMPF), 3-indoxyl sulfate, 4-ethylphenyl sulfate, hydroxyisovaleroylcarnitine (C5), isovalerylglycine, lactate, N1-methyl-2-pyridone-5-carboxamide, pantothenate (vitamin B5), phenylacetylglutamine, pipecolate, 3-hydroxyhipprete, and combinations thereof; and (ii) (a) xanthine at a level of 182.7 ng / ml or greater; (b) hydroxyl-chlorothalonil at a level of 20.3 ng / ml or greater; (c) 5-hydroxyindole acetate at a level of 28.5 ng / ml or greater; (d) lactate at a level of 686,600.0 ng / ml or greater; (e) pantothenate at a level of 63.3 ng / ml or greater; (f) pipecolate at a level of 303.6 ng / ml or greater; (g) gamma-CEHC at levels of 32.0 ng / ml or less; (h) indole acetate at a level of 141.4 ng / ml or less; (i) p-cresol sulfate at a level of 182.7 ng / ml or less; (j) 1,5-anhydroglucitol (1,5-AG) at a level of 11910.3 ng / ml or less; (k) 3-carboxy-4-methyl-5-propyl-2-furanpropanoate (CMPF) at a level of 7.98 ng / ml or less; (l) 3-indoxyl sulfate at a level of 256.7 ng / ml or less; (m) 4-ethylphenyl sulfate at a level of 3.0 ng / ml or less; (n) hydroxyisovaleroylcarnitine (C5) at a level of 12.9 ng / ml or less; (o) N1-methyl-2-pyridone-5-carboxamide at a level of 124.82 ng / ml or less; and (p) Phenyacetylglutamine at a level of 166.4 ng / ml or less detecting two or more of: (iii) determining that the subject has or is at risk of having an ASD based on the metabolite levels detected in step (ii). A method comprising: (Item 8) 8. The method of claim 7, wherein the plurality of metabolites comprises xanthine and gamma-CEHC. (Item 9) 8. The method of claim 7, wherein the sample is a plasma sample. (Item 10) 8. The method of claim 7, wherein the level of the metabolite is measured by mass spectrometry. (Item 11) 8. The method of item 7, wherein the subject is about 54 months of age or younger. (Item 12) 8. The method of item 7, wherein the subject is about 36 months of age or younger. [Brief explanation of the drawings]
[0073] [Figure 1] FIG. 1 illustrates the distribution of an exemplary metabolite in two populations (eg, ASD and DD) and the mean shift of this metabolite between these two populations.
[0074] [Figure 2]Figure 2 illustrates the distribution of an exemplary metabolite in two populations (e.g., ASD and DD) and the tail effect of this metabolite between these two populations (e.g., the ASD distribution has a more densely populated tail).
[0075] [Figure 3] Figure 3 illustrates the distribution of the metabolite 5-hydroxyindole acetate in two populations (e.g., ASD and DD) that exhibits a statistically significant mean shift (t-test; p<0.01) and a statistically significant right-tail effect between the two populations ("extreme values" indicate a tail effect, p=0.001 in Fisher's test).
[0076] [Figure 4] Figure 4 illustrates the distribution of the metabolite gamma-CEHC in two populations (e.g., ASD and DD) that exhibits a statistically significant left-tail effect between the two populations ("extreme values" indicate a tail effect, p=0.008 in Fisher's test).
[0077] [Figure 5] Figure 5 illustrates the distribution of the metabolite phenylacetylglutamine in two populations (e.g., ASD and DD), which exhibits a statistically significant mean shift (t-test; p=0.001) and statistically significant left and right tail effects between the two populations ("extreme values" indicate tail effects, p=0.0001 in Fisher's test). The distributions are depicted as shifted Gaussian curves in the two populations.
[0078] [Figure 6] FIG. 6 illustrates the correlation of two exemplary metabolites, demonstrating that these two metabolites possess distinct and complementary profiles of tail effects.
[0079] [Figure 7]FIG. 7 illustrates the tail effect of 12 exemplary metabolites in 180 subjects and their predictive power for ASD and DD.
[0080] [Figure 8A] FIG. 8A illustrates a plot of the ASD and non-ASD tail effect for 180 samples using an exemplary panel of 12 metabolites, demonstrating that samples from ASD patients show an aggregation of the ASD tail effect.
[0081] [Figure 8B] FIG. 8B illustrates a plot of ASD and non-ASD tail effects for 180 samples using an exemplary panel of 12 metabolites, along with an exemplary method for binning the data.
[0082] [Figure 8C] FIG. 8C illustrates a plot of the ASD and non-ASD tail effect for 180 samples using an exemplary panel of 21 metabolites, demonstrating that samples from ASD patients show an aggregation of the ASD tail effect.
[0083] [Figure 9] FIG. 9 illustrates the increasing predictability of ASD for an exemplary panel of 12 metabolites as a function of the number of metabolites assessed as increased.
[0084] [Figure 10A] FIG. 10A illustrates the effect of trichotomization on the predictability of ASD using an exemplary panel of 12 metabolites.
[0085] [Figure 10B] FIG. 10B illustrates the effect of triangulation on the predictability of ASD in the analysis of an exemplary panel of 21 metabolites.
[0086] [Figure 11A]FIG. 11A illustrates the improvement in predictability of ASD using a voting method compared to a non-voting method for analysis of an exemplary panel of 12 metabolites.
[0087] [Figure 11B] FIG. 11B illustrates the improvement in predictability of ASD using a voting method compared to a non-voting method using an exemplary panel of 21 metabolites.
[0088] [Figure 12] FIG. 12 illustrates the validation process for using an exemplary 12-metabolite panel to achieve high predictability of ASD.
[0089] [Figure 13A] Figures 13A-13U illustrate the population distribution of 21 exemplary metabolites in ASD and non-ASD populations. [Figure 13B] Figures 13A-13U illustrate the population distribution of 21 exemplary metabolites in ASD and non-ASD populations. [Figure 13C] Figures 13A-13U illustrate the population distribution of 21 exemplary metabolites in ASD and non-ASD populations. [Figure 13D] Figures 13A-13U illustrate the population distribution of 21 exemplary metabolites in ASD and non-ASD populations. [Figure 13E] Figures 13A-13U illustrate the population distribution of 21 exemplary metabolites in ASD and non-ASD populations. [Figure 13F] Figures 13A-13U illustrate the population distribution of 21 exemplary metabolites in ASD and non-ASD populations. [Figure 13G] Figures 13A-13U illustrate the population distribution of 21 exemplary metabolites in ASD and non-ASD populations. [Figure 13H] Figures 13A-13U illustrate the population distribution of 21 exemplary metabolites in ASD and non-ASD populations. [Figure 13I] Figures 13A-13U illustrate the population distribution of 21 exemplary metabolites in ASD and non-ASD populations. [Figure 13J] Figures 13A-13U illustrate the population distribution of 21 exemplary metabolites in ASD and non-ASD populations. [Figure 13K] Figures 13A-13U illustrate the population distribution of 21 exemplary metabolites in ASD and non-ASD populations. [Figure 13L] Figures 13A-13U illustrate the population distribution of 21 exemplary metabolites in ASD and non-ASD populations. [Figure 13M] Figures 13A-13U illustrate the population distribution of 21 exemplary metabolites in ASD and non-ASD populations. [Figure 13N] Figures 13A-13U illustrate the population distribution of 21 exemplary metabolites in ASD and non-ASD populations. [Figure 13O] Figures 13A-13U illustrate the population distribution of 21 exemplary metabolites in ASD and non-ASD populations. [Figure 13P] Figures 13A-13U illustrate the population distribution of 21 exemplary metabolites in ASD and non-ASD populations. [Figure 13Q] Figures 13A-13U illustrate the population distribution of 21 exemplary metabolites in ASD and non-ASD populations. [Figure 13R] Figures 13A-13U illustrate the population distribution of 21 exemplary metabolites in ASD and non-ASD populations. [Figure 13S] Figures 13A-13U illustrate the population distribution of 21 exemplary metabolites in ASD and non-ASD populations. [Figure 13T] Figures 13A-13U illustrate the population distribution of 21 exemplary metabolites in ASD and non-ASD populations. [Figure 13U] Figures 13A-13U illustrate the population distribution of 21 exemplary metabolites in ASD and non-ASD populations.
[0090] [Figure 14A] 14A-B illustrate the effect on ASD predictability of inclusion and exclusion of an exemplary 12-metabolite panel, an exemplary 21-metabolite panel, and a set of 84 candidate metabolites from a total of 600 metabolites, as assessed by tail effect and mean shift analysis. (Blacklist = excluded; Whitelist = included; mx_12 = exemplary 12-metabolite panel; mx_targeted21 = exemplary 21-metabolite panel; mx_total_candidates = 84 candidate metabolites; Total_features = total set of 600 metabolites.) [Figure 14B] 14A-B illustrate the effect on ASD predictability of inclusion and exclusion of an exemplary 12-metabolite panel, an exemplary 21-metabolite panel, and a set of 84 candidate metabolites from a total of 600 metabolites, as assessed by tail effect and mean shift analysis. (Blacklist = excluded; Whitelist = included; mx_12 = exemplary 12-metabolite panel; mx_targeted21 = exemplary 21-metabolite panel; mx_total_candidates = 84 candidate metabolites; Total_features = total set of 600 metabolites.)
[0091] [Figure 14C] Figures 14C-D illustrate the effect on the predictability of ASD by including (whitelist) and excluding (blacklist) an exemplary 12-metabolite panel and an exemplary 21-metabolite panel from a total of 600 metabolites, as assessed by tail effect and mean shift analysis, and by comparing logistic regression with Bayesian analysis in two cohorts of samples (i.e., "Christmas" and "Easter"). (Blacklist = excluded, whitelist = included, mx_12 = exemplary 12-metabolite panel, mx_targeted21 = exemplary 21-metabolite panel, mx_total_candidates = 84 candidate metabolites, total_features = total set of 600 metabolites). [Figure 14D] Figures 14C-D illustrate the effect on the predictability of ASD by including (whitelist) and excluding (blacklist) an exemplary 12-metabolite panel and an exemplary 21-metabolite panel from a total of 600 metabolites, as assessed by tail effect and mean shift analysis, and by comparing logistic regression with Bayesian analysis in two cohorts of samples (i.e., "Christmas" and "Easter"). (Blacklist = excluded, whitelist = included, mx_12 = exemplary 12-metabolite panel, mx_targeted21 = exemplary 21-metabolite panel, mx_total_candidates = 84 candidate metabolites, total_features = total set of 600 metabolites).
[0092] [Figure 15] FIG. 15 illustrates the effect on the predictability of ASD by using an increasing number of metabolites selected from a subset of an exemplary panel of 21 metabolites.
[0093] [Figure 16A] FIG. 16A illustrates the effect of adding genetic information to tail effect analysis using an exemplary panel of 12 metabolites, demonstrating improved power to separate ASD from non-ASD.
[0094] [Figure 16B] FIG. 16B illustrates the effect of adding genetic information to tail effect analysis using an exemplary panel of 21 metabolites, demonstrating improved power to separate ASD from non-ASD.
[0095] [Figure 17A]17A-B illustrate the effect on the predictability of ASD of including and excluding an exemplary panel of 21 metabolites from the total number of metabolites, of comparing tail effect analysis with mean-shift analysis, and of comparing logistic regression with Bayesian analysis. (Blacklist = excluded, Whitelist = included, mx_12 = exemplary panel of 12 metabolites, mx_targeted21 = exemplary panel of 21 metabolites, mx_total candidates = 84 candidate metabolites, Total features = total set of 600 metabolites). [Figure 17B] 17A-B illustrate the effect on the predictability of ASD of including and excluding an exemplary panel of 21 metabolites from the total number of metabolites, of comparing tail effect analysis with mean-shift analysis, and of comparing logistic regression with Bayesian analysis. (Blacklist = excluded, Whitelist = included, mx_12 = exemplary panel of 12 metabolites, mx_targeted21 = exemplary panel of 21 metabolites, mx_total candidates = 84 candidate metabolites, Total features = total set of 600 metabolites).
[0096] [Figure 18A] 18A-B illustrate the effect on the predictability of ASD by including and excluding an exemplary 21-metabolite panel from the total number of metabolites, by comparing tail effect analysis with mean-shift analysis, and by using logistic regression in two cohorts (i.e., "Christmas" and "Easter"). (Blacklist = excluded, Whitelist = included, mx_12 = exemplary 12-metabolite panel, mx_targeted21 = exemplary 21-metabolite panel, mx_total_candidates = 84 candidate metabolites, Total_features = total set of 600 metabolites.) [Figure 18B]18A-B illustrate the effect on the predictability of ASD by including and excluding an exemplary 21-metabolite panel from the total number of metabolites, by comparing tail effect analysis with mean-shift analysis, and by using logistic regression in two cohorts (i.e., "Christmas" and "Easter"). (Blacklist = excluded, Whitelist = included, mx_12 = exemplary 12-metabolite panel, mx_targeted21 = exemplary 21-metabolite panel, mx_total_candidates = 84 candidate metabolites, Total_features = total set of 600 metabolites.)
[0097] [Figure 19A] 19A-D illustrate the effect on the predictability of ASD of including and excluding an exemplary 21-metabolite panel from the total number of metabolites, comparing tail effect analysis with mean-shift analysis, and comparing logistic regression with Bayesian analysis using either the "Christmas" or "Easter" cohorts, or a combination of both. (Blacklist = excluded; Whitelist = included; mx_12 = exemplary 12-metabolite panel; mx_targeted21 = exemplary 21-metabolite panel; mx_totalcandidates = 84 candidate metabolites; Total features = total set of 600 metabolites.) [Figure 19B] 19A-D illustrate the effect on the predictability of ASD of including and excluding an exemplary 21-metabolite panel from the total number of metabolites, comparing tail effect analysis with mean-shift analysis, and comparing logistic regression with Bayesian analysis using either the "Christmas" or "Easter" cohorts, or a combination of both. (Blacklist = excluded; Whitelist = included; mx_12 = exemplary 12-metabolite panel; mx_targeted21 = exemplary 21-metabolite panel; mx_totalcandidates = 84 candidate metabolites; Total features = total set of 600 metabolites.) [Figure 19C]19A-D illustrate the effect on the predictability of ASD of including and excluding an exemplary 21-metabolite panel from the total number of metabolites, comparing tail effect analysis with mean-shift analysis, and comparing logistic regression with Bayesian analysis using either the "Christmas" or "Easter" cohorts, or a combination of both. (Blacklist = excluded; Whitelist = included; mx_12 = exemplary 12-metabolite panel; mx_targeted21 = exemplary 21-metabolite panel; mx_totalcandidates = 84 candidate metabolites; Total features = total set of 600 metabolites.) [Figure 19D] 19A-D illustrate the effect on the predictability of ASD of including and excluding an exemplary 21-metabolite panel from the total number of metabolites, comparing tail effect analysis with mean-shift analysis, and comparing logistic regression with Bayesian analysis using either the "Christmas" or "Easter" cohorts, or a combination of both. (Blacklist = excluded; Whitelist = included; mx_12 = exemplary 12-metabolite panel; mx_targeted21 = exemplary 21-metabolite panel; mx_totalcandidates = 84 candidate metabolites; Total features = total set of 600 metabolites.)
[0098] [Figure 20] FIG. 20 illustrates representative plots of the specificity and sensitivity of tail effect analysis for an exemplary panel of 21 metabolites for ASD prediction.
[0099] [Figure 21] FIG. 21 illustrates a scoring system in which a risk score for ASD or DD is calculated based on the sum of the log2 values of the odds ratios for the predictive metabolites of ASD and DD. DETAILED DESCRIPTION OF THE INVENTION
[0100] definition In order that the present invention may be more readily understood, certain terms are first defined below. Additional definitions of these terms and other terms appear throughout the specification.
[0101] In this application, unless otherwise clear from the context, (i) the term "a" may be understood to mean "at least one," (ii) the term "or" may be understood to mean "and / or," (iii) the terms "comprise" and "including" may be understood to include the recited components or steps, whether presented alone or with one or more additional components or steps, (iv) the terms "about" and "approximately" may be understood to allow for standard deviation that would be understood by one of ordinary skill in the art, and (v) when ranges are provided, the endpoints are included.
[0102] Agent: The term "agent," as used herein, may refer to any chemical class of compound or entity, such as, for example, a polypeptide, a nucleic acid, a saccharide, a lipid, a small molecule, a metal, or a combination thereof.
[0103] Approximately: As used herein, the terms "approximately," "approximately," or "about" are intended to encompass normal statistical variations that would be understood by one of ordinary skill in the art as appropriate to the relevant context. In certain embodiments, the terms "approximately," "approximately," or "about" refer to a range of values that falls within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1% or less in either direction (greater or less) of a stated reference value, or that are otherwise apparent from the context, unless otherwise specified (unless such number is expected to exceed 100% of the possible values).
[0104] Area Under the Curve (AUC): A classifier has an associated ROC curve (Receiver Operating Characteristic curve) that plots the false positive rate (1 - specificity) against the true positive rate (sensitivity). The area under the ROC curve (AUC) is a measure of how well a classifier can distinguish between two diagnostic groups. A perfect classifier will have an AUC of 1.0 compared to a random classifier that has an AUC of 0.5.
[0105] Associated with: As used herein, this term refers to two events or entities being "associated" with one another when the presence, level, and / or form of one correlates with that of the other. For example, if the presence, level, and / or form of a particular entity correlates with the incidence and / or susceptibility of a particular disease, disorder, or condition (e.g., across a relevant population), then that entity is considered to be associated with that disease, disorder, or condition.
[0106] Autism spectrum disorder: As used herein, the term "autism spectrum disorder" is recognized by those skilled in the art as a developmental disorder on the autism "spectrum" characterized by one or more of the following: deficits in interactive social interaction, language difficulties, repetitive behaviors, and restricted interests. Autism spectrum disorder is characterized in the DSM-V (May 2013) as a disorder that includes a range of symptoms, such as communication deficits, e.g., inappropriate conversational responses, misleading nonverbal interactions, difficulty establishing age-appropriate friendships, excessive dependency in routine tasks, extreme sensitivity to changes in their environment, and / or intense preoccupation with inappropriate topics. Autism spectrum disorder is additionally characterized, for example, by the DSM-IV-TR, to include autistic disorder, Asperger's disorder, Rett's disorder, childhood disintegrative disorder, and pervasive developmental disorder not otherwise specified (e.g., atypical autism). In some embodiments, autism spectrum disorder (ASD) is characterized using standardized testing instruments, such as questionnaires and observation schedules.For example, in some embodiments, ASD is characterized by (i) a score that meets the autism cutoff for the communication and social interaction total on the Autism Diagnostic Observation Schedule (ADOS) and a score that meets the cutoffs for social interaction, communication, behavioral patterns, and developmental abnormalities at <36 months on the Autism Diagnostic Interview-Revised (ADI-R); and / or (ii) a score that meets the ASD cutoff for communication and social interaction total on the ADOS and a score that meets the cutoffs for social interaction, communication, behavioral patterns, and developmental abnormalities at <36 months on the ADI-R, and (ii)(a) a score that meets the cutoff for social interaction and communication on the ADI-R, or (ii)(b) a score that meets the cutoff for social interaction or communication and is within two points of the cutoff for social interaction or communication on the ADI-R (neither of which meets the cutoff), or (ii)(c) a score that is within one point of the cutoff for social interaction and communication on the ADI-R.
[0107] Classification: As used herein, "classification" is the process of learning to separate data points into different classes by discovering common features among collected data points within known classes and then using mathematical or other methods to assign the data points to one of the different classes. In statistics, classification is the problem of identifying the subpopulation to which a new observation, whose subpopulation identity is unknown, belongs, based on a training set of data containing observations whose subpopulations are known. The requirement is therefore that new individual items be classified into groups based on quantitative information about one or more measurements, traits, or characteristics, and the like, and based on a training set in which predetermined groupings have already been established. Classification has many applications. In some cases, classification is employed as a data mining procedure, while in other cases more detailed statistical modeling is used.
[0108] Classifier: As used herein, a "classifier" is a method, algorithm, computer program, or system for performing data classification. Examples of commonly used classifiers include, but are not limited to, neural networks (multilayer perceptrons), logistic regression, support vector machines, k-nearest neighbors, Gaussian mixture models, Gaussian naive Bayes, decision trees, partial least squares discriminant analysis (PSL-DA), Fisher's linear discriminant, logistic regression, naive Bayes classifiers, perceptrons, support vector machines, quadratic classifiers, kernel estimation, boosting, neural networks, Bayesian networks, hidden Markov models, and learning vector quantization.
[0109] Determining: Many of the techniques described herein involve a "determining" step. Those skilled in the art will understand, upon reading this specification, that such "determining" may utilize or be accomplished through the use of any of a variety of techniques available to those skilled in the art, such as, for example, the specific techniques explicitly mentioned herein. In some embodiments, determining involves the manipulation of the physical sample. In some embodiments, determining involves the consideration and / or manipulation of data or information, for example, utilizing a computer or other processing unit adapted to perform relevant analyses. In some embodiments, determining involves receiving relevant information and / or material from a source. In some embodiments, determining involves comparing one or more features of the sample or entity to a comparable reference.
[0110] Determining risk: As used herein, determining risk encompasses calculating or quantifying the probability that a given subject has or does not have a particular condition or disorder. In some embodiments, a positive or negative diagnosis for a disorder or condition, such as autism spectrum disorder (ASD) or developmental delay (DD), can be made based in whole or in part on the determined risk or risk score (e.g., odds ratio or range).
[0111] Developmental delay: As used herein, the phrase developmental delay (DD) refers to ongoing, severe or mild delay in one or more processes of childhood development, such as physical development, cognitive development, communication development, social or emotional development, or adaptive development, that is not attributable to autism spectrum disorder. While individuals with ASD may be considered developmentally delayed, the classification of ASD as used herein will be considered to supersede the classification of DD, such that the classifications of ASD and DD are mutually exclusive. In other words, unless otherwise specified, a classification of DD is assumed to refer to non-ASD developmental delay. In some embodiments, DD is non-autistic (AU) and non-ASD, further characterized by (i) a score of 69 or lower on the Mullen Scale, a score of 69 or lower on the Vineland Scale, and a score of 14 or lower on the SCQ, or (ii) a score of 69 or lower on either the Mullen or Vineland Scale and within half a standard deviation of the cutoff value (score of 77 or lower) on the other assessment.
[0112] Diagnostic information: As used herein, diagnostic information or information for use in diagnosis is any information useful in determining whether a patient has a disease or condition, and / or useful in classifying a disease or condition into a phenotypic category or any category that has prognostic significance for the disease or condition, or that may respond to treatment (either general or any specific treatment) for the disease or condition. Similarly, diagnosis refers to providing any type of diagnostic information, such as, but not limited to, whether a subject is likely to have a disease or condition (e.g., autism spectrum disorder), the status, stage, or characteristics of the disease or condition manifested in the subject, information on the nature or classification of the disorder, information on prognosis, and / or information useful in selecting an appropriate treatment. Treatment selection can include selecting a specific therapeutic agent or other treatment modality, such as behavioral therapy, dietary modification, whether to withhold or implement therapy, and selection of a dosage regimen (e.g., the frequency or level of one or more doses of a specific therapeutic agent or combination of therapeutic agents).
[0113] Marker: As used herein, marker refers to an agent whose presence or level is associated with or correlates with a specific disease or condition. Alternatively, or in addition, in some embodiments, the presence or level of a specific marker correlates with the activity (or activity level) of a specific signal transduction pathway, which may be characteristic of a specific disorder, for example. A marker may or may not play an etiological role in a disease or condition. The statistical significance of the presence or absence of a marker may vary depending on the specific marker. In some embodiments, the detection of a marker is highly specific in that it reflects a high probability that the disorder belongs to a specific subclass. According to the present invention, a useful marker does not need to distinguish between specific subclasses of disorders with 100% accuracy.
[0114] Metabolite: As used herein, the term metabolite refers to a substance produced during a chemical or physical process in the body. The term "metabolite" encompasses any chemical or biochemical product of a metabolic process, such as any compound produced by the processing, cleavage, or consumption of a biomolecule. Examples of such molecules include, but are not limited to, acids and related compounds; mono-, di-, and tricarboxylic acids (saturated, unsaturated aliphatic and cyclic, aryl, alkaryl); aldo-acids, keto-acids; lactones; di- and di-carboxylic acids; and carboxylic acids. Gibbereillin; abscisic acid; alcohols, polyols, derivatives, and related compounds; ethyl alcohol, benzyl alcohol, menthanol; propylene glycol, glycerol, phytol; inositol, furfuryl alcohol, menthol; aldehydes, ketones, quinones, derivatives, and related compounds; acetaldehyde, butyraldehyde, benzaldehyde, acrolein, furfural, glyoxal; acetone, butanone; anthraquinones; carbohydrates; mono Sugars, disaccharides, and trisaccharides; alkaloids, amines, and other bases; pyridines (nicotinic acid, nicotinamide, etc.); pyrimidines (cytidine, thymine, etc.); purines (guanine, adenine, xanthine / hypoxanthine, kinetin, etc.); pyrroles; quinolines (isoquinolines, etc.); morphinans, tropanes, and cinchonans; nucleotides, oligonucleotides, derivatives, and related compounds; guanosine, cytosine, adenosine, thymidine, and inosine; amino acids, oligopeptides, and derivatives Examples of metabolites include: esters, phenols, and related compounds; esters; phenols and related compounds; heterocyclic compounds and derivatives; pyrroles, tetrapyrroles (corrinoids and porphines / porphyrins, with or without metal ions); flavonoids; indoles; lipids (such as fatty acids and triglycerides), derivatives, and related compounds; carotenoids, phytoene; and isoprenoids such as sterols and terpenes; and modified forms of the above molecules. In some embodiments, the metabolites are metabolic products of endogenous substances. In some embodiments, the metabolites are metabolic products of exogenous substances. In some embodiments, the metabolites are metabolic products of endogenous and exogenous substances. As used herein, the term "metabolome" refers to the chemical profile or fingerprint of the metabolites in a body fluid, cell, tissue, organ, or organism.
[0115] Metabolite distribution curve: As used herein, a metabolite distribution curve is a probability distribution curve defined by a function derived from metabolite levels plotted against population density (e.g., ASD or DD). In some embodiments, the distribution curve is a standard curve fit of the data. In some embodiments, the distribution curve is a least-squares polynomial curve fit. In some embodiments, the distribution curve is asymmetric or non-Gaussian. In some embodiments, the distribution curve is simply a plot (e.g., a "rug plot") of cases with associated diagnostic category versus metabolite value, in which case there is no curve fitting.
[0116] Mutual Information: As used herein, mutual information refers to a measure of the interdependence of two variables (i.e., the degree to which knowing one variable reduces the uncertainty about the other variable). High mutual information indicates a large reduction in uncertainty, low mutual information indicates a small reduction, and zero mutual information between two random variables means the variables are independent.
[0117] Non-autism spectrum disorder (non-ASD): As used herein, non-autism spectrum disorder (non-ASD) refers to a classification that does not belong to children or adults with autism spectrum disorder. In some embodiments, "non-ASD" refers to a subject that develops normally. In some embodiments, the non-ASD population consists of or includes subjects with developmental delay (DD). In some embodiments, "non-ASD" consists of or includes both DD and normally developing subjects.
[0118] Patient: As used herein, the term "patient" or "subject" refers to any living organism to which a test or composition is or can be administered, e.g., for experimental, diagnostic, prophylactic, and / or therapeutic purposes. In some embodiments, a patient is suffering from or susceptible to one or more disorders or conditions. In some embodiments, a patient exhibits one or more symptoms of a disorder or condition. In some embodiments, a patient is suspected of having one or more disorders or conditions.
[0119] Predictability: As used herein, predictability refers to the degree to which a correct prediction or forecast of a subject's disease state can be made, either qualitatively or quantitatively. Perfect predictability implies strict determinism, but lack of predictability does not necessarily imply lack of determinism. Limitations to predictability may be caused by factors such as lack of information or excessive complexity.
[0120] Prognostic and predictive information: As used herein, the terms prognostic and predictive information are used interchangeably and refer to any information that can be used to indicate any aspect of the course of a disease or condition, either in the absence or presence of treatment. Such information can include, but is not limited to, the likelihood that a patient's disease will be cured, or the likelihood that a patient's disease will respond to a particular therapy (where response can be defined in any of a variety of ways). Prognostic and predictive information is encompassed within the broad category of diagnostic information.
[0121] Reference: The term "reference" is often used herein to describe a standard or control agent, individual, population, sample, sequence, or value to which an agent, individual, population, sample, sequence, or value of interest is compared. In some embodiments, the reference agent, individual, population, sample, sequence, or value is tested and / or determined substantially contemporaneously with the testing or determination of the agent, individual, population, sample, sequence, or value of interest. In some embodiments, the reference agent, individual, population, sample, sequence, or value is a historical reference, optionally embodied in tangible media. Typically, as will be understood by one of skill in the art, the reference agent, individual, population, sample, sequence, or value is determined or characterized under conditions comparable to those utilized to determine or characterize the agent, individual, population, sample, sequence, or value of interest.
[0122] Regression analysis: As used herein, "regression analysis" encompasses any technique for modeling and analyzing several variables when the focus is on the relationship between the dependent variable and one or more independent variables. More specifically, regression analysis helps understand how the typical value of a dependent variable changes when any one of the independent variables varies while the other independent variables remain fixed. Most commonly, regression analysis estimates the conditional expectation of the dependent variable given the independent variables, i.e., the mean value of the dependent variable when the independent variables remain fixed. Less commonly, the focus is on the quantiles or other location parameters of the conditional distribution of the dependent variable given the independent variables. In all cases, the target of estimation is a function of the independent variables, called the regression function. In regression analysis, it is also interesting to characterize the variation of the dependent variable around the regression function, which can be described by a probability distribution. Regression analysis is widely used to predict and forecast, where its use has substantial overlap with the field of machine learning. Regression analysis is also used to understand which independent variables are related to a dependent variable and to investigate the form of these relationships. In limited circumstances, regression analysis can be used to infer causal relationships between independent and dependent variables. Numerous techniques have been developed for performing regression analysis. Well-known methods such as linear regression and ordinary least squares regression are parametric in that the regression function is defined in terms of a finite number of unknown parameters that are estimated from the data. Nonparametric regression refers to techniques that make the regression function exist in a specific set of functions, which can be infinite-dimensional.
[0123] Risk: As will be understood from the context, the "risk" of a disease, disorder, or condition is the likelihood that a particular individual will be diagnosed with or develop a disease, disorder, or condition. In some embodiments, the risk is expressed as a percentage. In some embodiments, the risk is from 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 to 100%. In some embodiments, the risk is expressed as the risk compared to the risk associated with a reference sample or group of reference samples. In some embodiments, the reference sample or group of reference samples has a known risk of the disease, disorder, or condition. In some embodiments, the reference sample or group of reference samples is derived from an individual comparable to the particular individual. In some embodiments, the relative risk is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more. In some embodiments, the relative risk can be expressed as a relative risk (RR) or odds ratio (OR).
[0124] Sample: As used herein, the term "sample" typically refers to a biological sample obtained or derived from a source of interest, as described herein. In some embodiments, the source of interest includes a living organism, such as an animal or a human. In some embodiments, the biological sample is or comprises a biological tissue or bodily fluid. In some embodiments, the biological sample may be or comprise bone marrow; blood; plasma; serum; blood cells; ascites; tissue or fine needle biopsy sample; cell-containing bodily fluid; free nucleic acid; sputum; saliva; urine; cerebrospinal fluid, peritoneal fluid; pleural fluid; stool; lymphatic fluid; gynecological fluid; skin swab; vaginal swab; oral swab; nasal swab; lavage or lavage fluid, such as ductal washing or bronchoalveolar lavage; aspirate; scraping; bone marrow specimen; tissue biopsy specimen; surgical specimen; stool, other bodily fluid, secretion, and / or excretion; and / or cells therefrom, etc. In some embodiments, the biological sample is or comprises cells obtained from an individual. In some embodiments, the obtained cells are or include cells from the individual from whom the sample was obtained. In some embodiments, the sample is a "primary sample" obtained directly from the source of interest by any suitable means. For example, in some embodiments, a primary biological sample is obtained by a method selected from the group consisting of biopsy (e.g., fine needle aspiration or tissue biopsy), surgery, collection of bodily fluids (e.g., blood, lymph, stool, etc.), and the like. In some embodiments, as will be clear from the context, the term "sample" refers to a preparation obtained by processing the primary sample (e.g., by removing one or more components of the primary sample and / or by adding one or more agents to the primary sample), such as by filtration using a semipermeable membrane. Such a "processed sample" may include, for example, nucleic acids or proteins extracted from the sample, or nucleic acids or proteins obtained by subjecting the primary sample to techniques such as mRNA amplification or reverse transcription, isolating and / or purifying certain components, etc.
[0125] Subject: "Subject" means a mammal (e.g., a human, including in some embodiments prenatal human forms). In some embodiments, the subject is suffering from the relevant disease, disorder, or condition. In some embodiments, the subject is susceptible to the disease, disorder, or condition. In some embodiments, the subject exhibits one or more symptoms or characteristics of the disease, disorder, or condition. In some embodiments, the subject does not exhibit any symptoms or characteristics of the disease, disorder, or condition. In some embodiments, the subject is one who possesses one or more features characteristic of susceptibility to or risk for a disease, disorder, or condition. A subject may be a patient, which refers to a human who sees a health care provider for diagnosis or treatment of a disease. In some embodiments, the subject is an individual to whom a therapy is administered.
[0126] Substantially: As used herein, the term "substantially" refers to the qualitative condition of complete or near-completeness of the degree or extent of a desired feature or characteristic. Those skilled in the biological arts will understand that biological and chemical phenomena rarely, if ever, proceed to completion and / or perfection or achieve or avoid absolute results. Thus, the term "substantially" is used herein to capture the possible lack of perfection inherent in many biological and chemical phenomena.
[0127] Suffering from: An individual "suffering from" a disease, disorder, or condition has been diagnosed with the disease, disorder, or condition and / or exhibits or is exhibiting one or more symptoms or characteristics of the disease, disorder, or condition.
[0128] Susceptible to: An individual who is "susceptible to" a disease, disorder, or condition is at risk of developing the disease, disorder, or condition. In some embodiments, such individuals are known to have one or more susceptibility factors that statistically correlate with an increased risk of developing the associated disease, disorder, and / or condition. In some embodiments, an individual who is susceptible to a disease, disorder, or condition does not exhibit any symptoms of the disease, disorder, or condition. In some embodiments, an individual who is susceptible to a disease, disorder, or condition has not been diagnosed with or has not yet been diagnosed with the disease, disorder, and / or condition. In some embodiments, an individual who is susceptible to a disease, disorder, or condition is an individual who has been exposed to a condition associated with the development of the disease, disorder, or condition. In some embodiments, the risk of developing a disease, disorder, and / or condition is a population-based risk (e.g., a family member of an individual with an allergy).
[0129] Tail enrichment and tail effect: As used herein, the term "tail enrichment" or "tail effect" refers to the characteristic of enhancing classification represented by metabolites (or other analytes) with relatively high concentrations in samples from a particular population in the distal portion of the distribution curve of metabolite levels. "Upper tail" or "right tail" refers to the distal portion of the distribution curve that is above the mean. "Lower tail" or "left tail" refers to the distal portion of the distribution curve that is below the mean. In some embodiments, the tail is determined by a predetermined threshold value based on ranking. For example, if the measured value of a sample for a particular metabolite is higher than the value corresponding to the 85th to 95th (e.g., 90th) percentile in the population for that metabolite, or lower than the value corresponding to the 10th to 20th (e.g., 15th) percentile in the population for that metabolite, the sample is said to be in the tail.
[0130] Therapeutic Agent: As used herein, the phrase "therapeutic agent" refers to any agent that has a therapeutic effect and / or elicits a desired biological and / or pharmacological effect when administered to a subject. In some embodiments, an agent is considered a therapeutic agent if administration of the agent to a relevant population statistically correlates with a desired or beneficial therapeutic outcome in the population, regardless of whether the particular subject to whom the agent is administered experiences the desired or beneficial therapeutic outcome.
[0131] Training Set: As used herein, a "training set" is a set of data used in various fields of information science to discover possible predictive relationships. Training sets are used in artificial intelligence, machine learning, genetic programming, intelligent systems, and statistics. In all of these fields, training sets have much the same role and are often used in conjunction with a test set.
[0132] Test Set: As used herein, a "test set" is a set of data used in various fields of information science to assess the strength and usefulness of predictive relationships. Test sets are used in artificial intelligence, machine learning, genetic programming, intelligent systems, and statistics. In all of these fields, test sets have much the same role.
[0133] Treatment: As used herein, the term "treatment" (also "treat" or "treating") refers to any administration of a substance or therapy (e.g., behavioral therapy) that partially or completely alleviates, improves, alleviates, suppresses, delays the onset of, reduces the severity of, and / or reduces the frequency, incidence, or severity of one or more symptoms, features, and / or causes of a particular disease, disorder, and / or condition. Such treatment may be treatment of a subject who does not exhibit symptoms of the associated disease, disorder, and / or condition and / or who exhibits only early signs of the disease, disorder, and / or condition. Alternatively, or in addition, such treatment may be treatment of a subject who exhibits one or more established signs of the associated disease, disorder, and / or condition. In some embodiments, treatment may be treatment of a subject who has been diagnosed with the associated disease, disorder, and / or condition. In some embodiments, the treatment may be treatment of a subject known to have one or more susceptibility factors that statistically correlate with an increased risk of developing the relevant disease, disorder, and / or condition.
[0134] Detailed Description The present invention provides methods and systems for determining the risk of autism spectrum disorder (ASD) in a subject based on specific analysis of metabolite levels in a sample, such as a blood sample or plasma sample. The following sections describe various aspects of the present invention in detail. The use of chapters and headings does not imply limitations on the present invention. Each chapter may apply to any aspect of the present invention. In this application, the use of "or" means "and / or" unless it is clear otherwise.
[0135] Autism spectrum disorder The criteria for the clinical diagnosis of autism spectrum disorder (ASD) are set out in the Diagnostic and Statistical Manual of Mental Disorders, Version 5 (DSM-V, published May 2013). ) is described in
[0136] ASD has been further characterized by, for example, DSM-IV-TR to include autistic disorder, Asperger's disorder, Rett's disorder, childhood disintegrative disorder, and pervasive developmental disorder not otherwise specified (such as atypical autism).
[0137] In some embodiments, ASD is characterized by (i) a score meeting the autism cutoff for the communication and social interaction total on the ADOS and a score meeting the cutoffs for social interaction, communication, behavioral patterns, and developmental abnormalities at <36 months on the ADI-R; and / or (ii) a score meeting the ASD cutoff for communication and social interaction total on the ADOS and a score meeting the cutoffs for social interaction, communication, behavioral patterns, and developmental abnormalities at <36 months on the ADI-R, and (ii)(a) a score meeting the cutoff for social interaction and communication on the ADI-R, or (ii)(b) a score meeting the cutoff for social interaction or communication and within two points (neither of which meets the cutoff) of the social interaction or communication cutoff on the ADI-R, or (ii)(c) a score within one point of the cutoff for social interaction and communication on the ADI-R.
[0138] developmental delay Developmental delay is a severe or mild delay in one or more processes of childhood development, such as physical development, cognitive development, communication development, social or emotional development, or adaptive development, not attributable to ASD. In some embodiments, DD is non-autistic (AU) and non-ASD, and is characterized by (i) a score of 69 or lower on the Mullen Scale, a score of 69 or lower on the Vineland Scale, and a score of 14 or lower on the SCQ, or (ii) a score of 69 or lower on either the Mullen or Vineland Scale, and within half a standard deviation of the cutoff value (score of 77 or lower) on other assessments. Although individuals with ASD may be considered developmentally delayed, the classification of ASD used herein will be considered to supersede the classification of DD, so that the classifications of ASD and DD are mutually exclusive.
[0139] ASD risk assessment Children who exhibit symptoms of impaired language, behavior, or social development often see clinicians, most commonly in primary care settings, but the clinician is unable to determine whether the child has ASD or some other condition, disorder, or classification (e.g., DD). Diagnosing children, especially those before extensive language development, is difficult, and many primary care physicians lack the ability or competence to differentially diagnose their patients. For example, ASD may not be easily distinguished from other developmental disorders, conditions, or classifications, such as DD.
[0140] It is useful to assess a subject's risk for ASD (e.g., probability of non-ASD and DD) to distinguish ASD from DD. ASD risk assessment provides opportunities for early intervention and treatment. For example, non-specialist physicians can use ASD risk assessment to suggest referral to specialists. Specialists can use ASD risk assessment to prioritize further evaluation of patients. ASD risk assessment may also be used to establish a provisional diagnosis before a definitive diagnosis, during which time facilitative services can be provided to high-risk children and their families.
[0141] Described herein is a method for determining the risk of ASD in a subject.In some embodiments, determining ASD risk comprises determining that the subject has more than about 50% chance of having ASD.In some embodiments, determining ASD risk comprises determining that the subject has more than about 60%, 65%, 70%, 74%, 80%, 85%, 90%, 95%, or 98% chance of having ASD.In some embodiments, determining ASD risk comprises determining that the subject has ASD.In some embodiments, determining ASD risk comprises determining that the subject does not have ASD (i.e., is non-ASD).
[0142] In some embodiments, the present invention provides a method for distinguishing ASD from a non-ASD classification (e.g., DD) in a subject. In some embodiments, distinguishing ASD from a non-ASD classification / condition comprises determining that the subject has more than about 60%, 65%, 70%, 74%, 80%, 85%, 90%, 95%, or 98% probability of having ASD (i.e., probability of having ASD and not having a non-ASD classification) rather than a non-ASD classification. In some embodiments, the non-ASD classification is DD. In some embodiments, the non-ASD classification is "normal".
[0143] In some embodiments, the present invention provides methods for determining that a subject does not have ASD or DD.
[0144] Analysis method The method for assessing ASD risk or distinguishing ASD from other non-ASD developmental disorders is described herein.In some embodiments, risk assessment is based (at least in part) on measuring and characterizing metabolites in sample (for example, blood sample) from subject.In some embodiments, plasma sample is obtained from blood sample, and plasma sample is analyzed.
[0145] Metabolites can be detected by a variety of methods, including chromatography and / or mass spectrometry, fluorometry, electrophoresis, immunoaffinity, hybridization, immunochemistry, ultraviolet spectroscopy (UV), fluorimetry, radiochemistry, near-infrared spectroscopy (near-IR), nuclear magnetic resonance spectroscopy (NMR), light scattering spectroscopy (LS), and nephelometry-based assays.
[0146] In some embodiments, metabolites are analyzed by liquid or gas chromatography or ion mobility (electrophoresis), alone or in combination with mass spectrometry, or by mass spectrometry alone. Such methods have been used to identify and quantify biomolecules, such as cellular metabolites. (See, e.g., Li et al., 2000; Rowley et al., 2000; and Kuster and Mann, 1998). Mass spectrometry methods include, for example, single, dual, or triple mass-to-charge scanning and / or filtering (MS, MS / MS, or MS). 3The mass spectrometry procedure may be based on quadrupole, ion trap, or time-of-flight mass spectrometry using a quadrupole, ion trap, or time-of-flight mass spectrometry instrument, which may be preceded by an appropriate ionization method, such as electrospray ionization, atmospheric pressure chemical ionization, atmospheric pressure photoionization, matrix-assisted laser desorption ionization (MALDI), or surface-enhanced laser desorption ionization (SELDI). (See, for example, International Patent Application Publications WO2004056456 and WO2004088309.) In some embodiments, the initial sorting of metabolites from a biological sample can be achieved by using gas or liquid chromatography or ion mobility / electrophoresis. In some embodiments, the ionization for the mass spectrometry procedure can be achieved by electrospray ionization, atmospheric pressure chemical ionization, or atmospheric pressure photoionization. In some embodiments, the mass spectrometry instrument includes a quadrupole, ion trap, time-of-flight, or Fourier transform instrument.
[0147] In some embodiments, metabolites are analyzed on a mass scale via untargeted ultrafast liquid or gas chromatography / electrospray or atmospheric pressure chemical ionization tandem mass spectrometry platforms optimized for the identification and relative quantification of small molecule complements in biological systems (see, e.g., Evans et al., Anal. Chem., 2009, 81, 6656-6659). (See page 6667).
[0148] In some embodiments, the first selection of metabolites from biological samples can be achieved by using gas or liquid chromatography or ion mobility / electrophoresis.In some embodiments, the ionization for mass spectrometry can be achieved by electrospray ionization, atmospheric pressure chemical ionization, or atmospheric pressure photoionization.In some embodiments, mass spectrometry instruments include quadrupole, ion trap, or time-of-flight, or Fourier transform instruments.
[0149] In some embodiments, a blood sample containing the metabolite of interest is centrifuged to separate the plasma from other blood components. In certain embodiments, an internal standard is unnecessary. In some embodiments, a predetermined amount of internal standard is added to a portion of the plasma, and then methanol is added to precipitate plasma components such as proteins. The precipitate is separated from the supernatant by centrifugation, and the supernatant is collected. If the concentration of the metabolite of interest needs to be increased for more accurate detection, the supernatant is evaporated and the residue is dissolved in an appropriate amount of solvent. If the concentration of the metabolite of interest is unnecessarily high, the supernatant is diluted with an appropriate solvent. An appropriate amount of the metabolite-containing sample is loaded onto a liquid chromatography column equilibrated with an appropriate mixture of mobile phase A and mobile phase B. In reversed-phase liquid chromatography, mobile phase A is typically water with or without a small amount of additive such as formic acid, and mobile phase B is typically methanol or acetonitrile. An appropriate gradient of mobile phase A and mobile phase B is pumped onto the column to achieve separation of the metabolite of interest based on retention time or elution time from the column. When metabolites elute from the column, they are ionized and transported into the gas phase, where the ions are detected and quantified by mass spectrometry. Detection specificity is achieved by double filtering of specific precursor ions and specific product ions generated from the precursor ions. Absolute quantification can be achieved by normalizing the ion count derived from the metabolite of interest to the ion count derived from a known amount of an internal standard for the given metabolite, and comparing the normalized ion count to a calibration curve established using known amounts of pure metabolite and the internal standard. The internal standard is typically a stable isotope-labeled form of the pure metabolite, or a pure form of a structural analog of the metabolite. Alternatively, the relative quantification of a given metabolite in arbitrary units may be calculated by normalizing to a selected internal reference value (e.g., the median metabolite level in all samples generated from a given group).
[0150] In some embodiments, one or more metabolites are measured by immunoassay. Numerous specific immunoassay formats and variations thereof are available for measuring metabolites. (See, e.g., E. Maggio, Enzyme-Immunoassay, (1980) (CRCPress, Inc., Boca Raton, Fla.); U.S. Pat. No. 4,727,022.) Issue 2 “Methods for Modulating Ligand-Receptor Interactions and their Application” U.S. Patent No. 4,659,678 "Immunoassay of Antigens"; U.S. Patent No. 4,376,110 "Immunometric Assays Using Monoclonal Antibodies"; U.S. Patent No. 4,27 No. 5,149 “MacromolecularEnvironmentControlin Specific Receptor Assays”; US US Patent No. 4,233,402 "Reagents and Method Employing Channeling"; and (See also U.S. Patent No. 4,230,767, "Heterogenous Specific Binding Assay Employing a Coenzyme as Label.") Antibodies may be bound to the antibody according to known techniques, such as immobilized binding. , or conjugated to a solid support suitable for diagnostic assays (e.g., beads such as protein A or protein G agarose, microspheres, plates, slides, or wells formed of materials such as latex or polystyrene). The antibodies described herein can also be radiolabeled (e.g., 35 S, 125 I, 131 I), enzyme labels (e.g., horseradish peroxidase, alkaline phosphatase), and fluorescent labels (e.g., fluorescein, Alexa, green fluorescent protein).
[0151] Determining ASD risk In some embodiments, the methods of the present invention enable one of skill in the art to identify, diagnose, or otherwise assess a subject based at least in part on measuring metabolite levels in a sample obtained from a subject who may not currently exhibit signs or symptoms of ASD and / or other developmental disorders, but who may nevertheless have or be at risk for developing ASD and / or other developmental disorders.
[0152] In certain embodiments, metabolite or other analyte levels (e.g., proteomic or genomic information) can be measured in a test sample and compared to normal control levels or levels in subjects with a non-ASD developmental disorder, condition, or classification (e.g., non-ASD developmental delay, DD). In some embodiments, the term "normal control level" refers to a level or index of one or more metabolites or other analytes typically found in subjects who do not suffer from ASD or who are unlikely to have ASD or other developmental disorders. In some embodiments, the normal control level is a range or index. In some embodiments, the normal control level is determined from a database of previously tested subjects. A difference in the level of one or more metabolites or other analytes compared to the normal control level may indicate that the subject has or is at risk for developing ASD. Conversely, no difference in the level of one or more metabolites compared to the normal control level of one or more metabolites or other analytes may indicate that the subject does not have or is at low risk for developing ASD.
[0153] In some embodiments, the reference value is a value obtained from a control subject or population with a known diagnosis (i.e., diagnosed or identified as having an ASD, or not diagnosed or identified as having an ASD). In some embodiments, the reference value is an index value or baseline value, such as a "normal control level" as described herein. In some embodiments, the reference sample or index value or baseline value may be taken or derived from one or more subjects undergoing treatment for an ASD, or from one or more subjects at low risk of developing an ASD, or from subjects who have shown improvement in ASD risk factors as a result of receiving treatment. In some embodiments, the reference sample or index value or baseline value is taken or derived from one or more subjects not undergoing treatment for an ASD. In some embodiments, samples are collected from subjects undergoing initial treatment for an ASD and / or subsequent treatment for an ASD to monitor treatment progress. In some embodiments, the reference value is derived from a risk prediction algorithm or computer-generated index from a population study of ASD. In some embodiments, the reference value is from a subject or population with a disease or disorder other than ASD, such as another developmental disorder, eg, a non-ASD developmental delay (DD).
[0154] In some embodiments, the difference in metabolite level measured by the method of the present invention comprises an increase or decrease in the level of the metabolite compared to a normal control level, reference value, index value, or baseline value. In some embodiments, an increase or decrease in the level of the metabolite relative to a reference value from a normal control population, the general population, or a population with another disease is indicative of the presence of ASD, the progression of ASD, the worsening of ASD, or the improvement of ASD or ASD symptoms. In some embodiments, an increase or decrease in the level of the metabolite relative to a reference value from a normal control population, the general population, or a population with another disease is indicative of an increase or decrease in the risk of developing ASD or its associated complications. The increase or decrease may be indicative of the success of one or more treatment regimens for ASD, or may indicate an improvement or regression of ASD risk factors. The increase or decrease may be, for example, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, or at least 50% of the reference value.
[0155] In some embodiments, the difference in metabolite levels described herein is a statistically significant difference. "Statistically significant" refers to a difference that is greater than the difference expected to occur by chance alone. Statistical significance can be determined by any method known in the art. For example, statistical significance can be determined by p-value. The p-value is a measure of the probability that a difference between groups in an experiment occurs by chance. For example, a p-value of 0.01 means that the result occurs by chance 1 in 100 times. The smaller the p-value, the more likely it is that the difference measured between groups is not due to chance. A difference is considered statistically significant if the p-value is 0.05 or less. In some embodiments, a statistically significant p-value is 0.04, 0.03, 0.02, 0.01, 0.005, or 0.001 or less. In some embodiments, a statistically significant p-value is 0.30, 0.25, 0.20, 0.15, or 0.10 or less (e.g., in the case of identifying whether a single particular metabolite has additional predictive value when used in a classifier that includes other metabolites). In some embodiments, the p-value is determined by a t-test. In some embodiments, the p-value is obtained by Fisher's test. In some embodiments, statistical significance is achieved by analyzing a combination of several metabolites in a panel, and in combination with a mathematical algorithm, a statistically significant risk prediction is achieved.
[0156] A classification test, assay, or method has an associated ROC curve (Receiver Operating Characteristic curve), which plots the false positive rate (1 - specificity) against the true positive rate (sensitivity). The area under the ROC curve (AUC) is a measure of how well a classifier can distinguish between two diagnostic groups. The maximum AUC is 1.0 (a perfect test), and the minimum is 0.5 (e.g., no difference between normal and disease). It is observed that as the AUC approaches 1, the accuracy of the test increases.
[0157] In some embodiments, a high degree of risk prediction accuracy is a test or assay having an AUC of at least 0.60, hi some embodiments, a high degree of risk prediction accuracy is a test or assay having an AUC of at least 0.65, at least 0.70, at least 0.75, at least 0.80, at least 0.85, at least 0.90, or at least 0.95.
[0158] Predicting ASD risk by assessing tail effects In some embodiments, the mean difference in metabolite levels is assessed within or between populations, for example, between an ASD population and a DD population, or compared with a normal control population. In some embodiments, metabolites from samples from a given population (i.e., ASD) are assessed for enrichment in the tail of the distribution curve. That is, this is to determine whether a greater proportion of samples from a specified population (e.g., ASD) is present in the tail of the distribution curve compared with a second population (e.g., DD) (i.e., "tail effect"). In some embodiments, both the mean difference and the tail effect are identified and utilized. In some embodiments, the tail is determined by a predetermined threshold value. For example, if the measured value of a sample for a particular metabolite is higher than the value corresponding to the 90th percentile in the population for that metabolite (right tail or upper tail), or lower than the value corresponding to the 15th percentile (left tail or lower tail), the sample is said to be in the tail. In some embodiments, the right (upper) tail threshold for a given metabolite is a value corresponding to the 80th, 81st, 82nd, 83rd, 84th, 85th, 86th, 87th, 88th, 89th, 90th, 91st, 92nd, 93rd, 94th, 95th, 96th, 97th, 98th, or 99th percentile (e.g., in this case, if a sample's measured value for a given metabolite is higher than the value associated with this percentile, the sample is said to be in the right tail). In some embodiments, the left (lower) tail threshold for a given metabolite is a value corresponding to the 25th, 24th, 23rd, 22nd, 21st, 20th, 19th, 18th, 17th, 16th, 15th, 14th, 13th, 12th, 11th, 10th, 9th, 8th, 7th, 6th, 5th, 4th, 3rd, 2nd, or 1st percentile (e.g., in this case, if a sample's measured value for a given metabolite is lower than the value associated with this percentile, the sample is said to be in the left tail). The percentile values shown include decimal values.
[0159] In some embodiments, distribution curves are generated from plots of metabolite levels in one or more populations. In some embodiments, distribution curves are generated from a single reference population, such as the general population. In some embodiments, distribution curves are generated from two populations, such as an ASD population and a non-ASD population, such as DD. In some embodiments, distribution curves are generated from three or more populations, such as an ASD population, a non-ASD population with another developmental disorder / condition / classification, such as DD, and a healthy (e.g., developmentally undisturbed) control population. The distribution curves of metabolites from each population can be used to perform more than one risk assessment (e.g., diagnosing ASD, diagnosing DD, or distinguishing between ASD and DD). The methods for assessment utilizing the tail effect described herein can be applied to more than two populations.
[0160] In some embodiments, multiple metabolites and their distributions are used in risk assessment. In some embodiments, levels of two or more metabolites are utilized to predict ASD risk. In some embodiments, at least two of the metabolites are selected from the metabolites listed in Table 1. In some embodiments, at least three of the metabolites are selected from the metabolites listed in Table 1. In some embodiments, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or 21 metabolites selected from the metabolites listed in Table 1 are used to predict ASD risk.
[0161] Further discussion of Table 1 (Tables 1A to 1C) is provided in the Examples section below. [Table 1A] [Table 1B] [Table 1C]
[0162] In some embodiments, the at least two metabolites for analysis are selected from the group consisting of phenylacetylglutamine, xanthine, octenoylcarnitine, p-cresol sulfate, isovalerylglycine, gamma-CEHC, indole acetate, pipecolate, 1,5-anhydroglucitol (1,5-AG), lactate, 3-(3-hydroxyphenyl)propionate, 3-indoxyl sulfate, pantothenate (vitamin B5), hydroxy-chlorothalonil, and combinations thereof.
[0163] In some embodiments, the at least three metabolites for analysis are selected from the group consisting of phenylacetylglutamine, xanthine, octenoylcarnitine, p-cresol sulfate, isovalerylglycine, gamma-CEHC, indole acetate, pipecolate, 1,5-anhydroglucitol (1,5-AG), lactate, 3-(3-hydroxyphenyl)propionate, 3-indoxyl sulfate, pantothenate (vitamin B5), hydroxy-chlorothalonil, and combinations thereof.
[0164] In some embodiments, information about the absence of a tail effect for a particular set of metabolites is used in risk assessment. In some embodiments, the absence of a tail effect is determined to provide a null result (i.e., absence of information, as opposed to negative information). In some embodiments, the absence of a tail effect is determined to be indicative of one classification over another (e.g., more indicative of DD than ASD).
[0165] In some embodiments, the distribution curve is asymmetric or non-Gaussian, hi some embodiments, the distribution curve does not follow a parametric distribution pattern.
[0166] In some embodiments, information from the mean difference (e.g., mean shift) is combined with information on tail effects for risk assessment. In some embodiments, information from the mean difference is used for risk assessment without using information on tail effects.
[0167] In some embodiments, analysis of metabolites is combined with other types of information, such as genetic information, demographic information, and / or behavioral assessments, to determine a subject's risk for ASD or other disorders.
[0168] In some embodiments, ASD risk assessment is based at least in part on the measured amount of a certain metabolite in a biological sample (e.g., blood, plasma, urine, saliva, stool) obtained from a subject, where the certain metabolite is found to exhibit a "tail effect." The inventors have found that there is not necessarily a statistically significant mean shift between two populations associated with the tail effect. Thus, the tail effect is a special phenomenon that is different from the mean shift.
[0169] In certain embodiments, a particular metabolite exhibits a right tail effect that is indicative of ASD relative to a non-ASD population (e.g., a DD population) if the metabolite is characterized as follows: · Non-ASD population distribution curves are established for metabolites in a non-ASD population (e.g., a DD population), with the x-axis showing the level of the first metabolite and the y-axis showing the corresponding population; An ASD population distribution curve is established for a metabolite in an ASD population, with the x-axis representing the level of a first metabolite and the y-axis representing the corresponding population; and The non-ASD population distribution curve and the ASD population distribution curve are characterized by the maintenance of one or both of (A) and (B): (A) the ratio of (i) the area under the ASD population distribution curve where x is greater than level n of the metabolite to (ii) the area under the non-ASD population distribution curve where x is greater than level n of the metabolite is greater than 150% (e.g., >200%, >300%, >500%, >1000%, etc.), thereby providing predictive utility for distinguishing between ASD and non-ASD classifications for samples greater than level n of the metabolite; and (B) If n' is the minimum threshold metabolite level corresponding to the top decile (or any cutoff from about 5% to about 20%) of the combined non-ASD and ASD populations used to create the distribution curve, then an unknown sample (e.g., a random sample selected from a population having equal numbers of ASD and non-ASD members) has a metabolite level of at least n', and the odds of the sample being ASD versus non-ASD are 1.6:1 or greater (e.g., 2:1 or greater, 3:1 or greater, 4:1 or greater, 5:1 or greater, 6:1 or greater, 7:1 or greater, 8:1 or greater, 9:1 or greater, or 10:1 or greater) (e.g., p<0.3, p<0.2, p<0.1, p<0.05, p<0.03, or p<0.01 is, for example, a statistically significant classification), thereby providing predictive utility for distinguishing between ASD and non-ASD classifications for samples with metabolite levels greater than n'.
[0170] In certain embodiments, a particular metabolite exhibits a left-tail effect that is indicative of ASD relative to a non-ASD population (e.g., a DD population) if the metabolite is characterized as follows: · Non-ASD population distribution curves are established for metabolites in a non-ASD population (e.g., a DD population), with the x-axis showing the level of the first metabolite and the y-axis showing the corresponding population; An ASD population distribution curve is established for a metabolite in an ASD population, with the x-axis representing the level of a first metabolite and the y-axis representing the corresponding population; and The non-ASD population distribution curve and the ASD population distribution curve are characterized by the maintenance of one or both of (A) and (B): (A) the ratio of (i) the area under the ASD population distribution curve when x is less than the metabolite level m to (ii) the area under the non-ASD population distribution curve when x is less than the metabolite level m is greater than 150% (e.g., >200%, >300%, >500%, >1000%, etc.), thereby providing predictive utility for distinguishing between ASD and non-ASD classifications for samples with metabolite levels less than m; and (B) If m' is the maximum threshold metabolite level corresponding to the bottom decile (or any cutoff between about 5% and about 20%) of the combined non-ASD and ASD populations used to create the distribution curve, then an unknown sample (e.g., a random sample selected from a population having equal numbers of ASD and non-ASD members) has a metabolite level below m', and the odds of the sample being ASD versus non-ASD are 1.6:1 or greater (e.g., 2:1 or greater, 3:1 or greater, 4:1 or greater, 5:1 or greater, 6:1 or greater, 7:1 or greater, 8:1 or greater, 9:1 or greater, or 10:1 or greater) (e.g., p<0.3, p<0.2, p<0.1, p<0.05, p<0.03, or p<0.01 is, for example, a statistically significant classification), thereby providing predictive utility for distinguishing between ASD and non-ASD classifications for samples with metabolite levels below m'.
[0171] In certain embodiments, a particular metabolite exhibits a right tail effect that is indicative of a non-ASD (e.g., DD) versus ASD population if the metabolite is characterized as follows: · Non-ASD population distribution curves are established for metabolites in a non-ASD population (e.g., a DD population), with the x-axis showing the level of the first metabolite and the y-axis showing the corresponding population; An ASD population distribution curve is established for a metabolite in an ASD population, with the x-axis representing the level of a first metabolite and the y-axis representing the corresponding population; and The non-ASD population distribution curve and the ASD population distribution curve are characterized by the maintenance of one or both of (A) and (B): (A) the ratio of (i) the area under the non-ASD population distribution curve where x is greater than level n of the metabolite to (ii) the area under the ASD population distribution curve where x is greater than level n of the metabolite is greater than 150% (e.g., >200%, >300%, >500%, >1000%, etc.), thereby providing predictive utility for distinguishing between non-ASD and ASD classifications for samples with levels greater than n of the metabolite; and (B) If n' is the minimum threshold metabolite level corresponding to the top decile (or any cutoff between about 5% and about 20%) of the combined non-ASD and ASD populations used to create the distribution curve, then an unknown sample (e.g., a random sample selected from a population having equal numbers of ASD and non-ASD members) has a metabolite level greater than n', and the odds of the sample being non-ASD versus ASD are 1.6:1 or greater (e.g., 2:1 or greater, 3:1 or greater, 4:1 or greater, 5:1 or greater, 6:1 or greater, 7:1 or greater, 8:1 or greater, 9:1 or greater, or 10:1 or greater) (e.g., p<0.3, p<0.2, p<0.1, p<0.05, p<0.03, or p<0.01 is, for example, a statistically significant classification), thereby providing predictive utility for distinguishing between non-ASD and ASD classifications for samples with metabolite levels greater than n'.
[0172] In certain embodiments, a particular metabolite exhibits a left-tail effect that is indicative of a non-ASD (e.g., DD) versus ASD population if the metabolite is characterized as follows: · Non-ASD population distribution curves are established for metabolites in a non-ASD population (e.g., a DD population), with the x-axis showing the level of the first metabolite and the y-axis showing the corresponding population; An ASD population distribution curve is established for a metabolite in an ASD population, with the x-axis representing the level of a first metabolite and the y-axis representing the corresponding population; and The non-ASD population distribution curve and the ASD population distribution curve are characterized by the maintenance of one or both of (A) and (B): (A) the ratio of (i) the area under the non-ASD population distribution curve where x is less than the level m of the metabolite to (ii) the area under the ASD population distribution curve where x is less than the level m of the metabolite is greater than 150% (e.g., >200%, >300%, >500%, >1000%, etc.), thereby providing predictive utility for distinguishing between non-ASD and ASD classifications for samples with levels less than m of the metabolite; and (B) If m' is the maximum threshold metabolite level corresponding to the bottom decile (or any cutoff between about 5% and about 20%) of the combined non-ASD and ASD populations used to create the distribution curve, then an unknown sample (e.g., a random sample selected from a population having equal numbers of ASD and non-ASD members) has a metabolite level below m', and the odds of the sample being non-ASD versus ASD are 1.6:1 or greater (e.g., 2:1 or greater, 3:1 or greater, 4:1 or greater, 5:1 or greater, 6:1 or greater, 7:1 or greater, 8:1 or greater, 9:1 or greater, or 10:1 or greater) (e.g., p<0.3, p<0.2, p<0.1, p<0.05, p<0.03, or p<0.01 is, for example, a statistically significant classification), thereby providing predictive utility for distinguishing between non-ASD and ASD classifications for samples with metabolite levels below m'.
[0173] In certain embodiments, risk assessment is performed using multiple metabolites that exhibit tail effect.It has been observed that there are certain groups of metabolites (for example, two or more metabolites) that provide complementary diagnostic / risk assessment information for the assessment of ASD.For example, the ASD-positive individuals (for example, individuals within the identified tail of the first metabolite) that can be identified by analyzing the level of the first metabolite are not the same ASD-positive individuals that can be identified by analyzing the level of the second metabolite (or there may be a low but non-zero degree of overlap).The tail of the first metabolite predicts certain ASD individuals, while the tail of the second metabolite predicts other ASD individuals.Without wishing to be bound by a particular theory, this finding may reflect the multifaceted nature of ASD itself.
[0174] Thus, in certain embodiments, the risk assessment method involves identifying whether a subject falls into any of a number of identified metabolite tails that include multiple metabolites, e.g., when the predictor variables in different metabolite tails are at least partially exclusive, e.g., when multiple metabolites have low mutual information, such that incorporating multiple metabolites with low mutual information improves risk prediction.
[0175] In some embodiments, the invention provides one or more binding members for use in a method of determining / diagnosing a subject having or at risk of having an ASD, wherein each binding member is selected from the group consisting of 5-hydroxyindole acetate (5-HIAA), 1,5-anhydroglucitol (1,5-AG), 3-(3-hydroxyphenyl)propionate, 3-carboxy-4-methyl-5-propyl-2-furanpropanoate (CMPF), 3-indoxyl sulfate, 4-ethylphenyl sulfate, 4-hydroxybenzoate ... The compound is capable of specifically binding to a metabolite selected from the group consisting of acetone, 8-hydroxyoctanoate, gamma-CEHC, hydroxyisovaleroylcarnitine (C5), indole acetate, isovalerylglycine, lactate, N1-methyl-2-pyridone-5-carboxamide, p-cresol sulfate, pantothenate (vitamin B5), phenylacetylglutamine, pipecolate, xanthine, hydroxychlorothalonil, octenoylcarnitine, and 3-hydroxyhipprete.
[0176] In some embodiments, the present invention relates to the use of one or more binding members for diagnosing / determining whether a subject has or is at risk of having ASD, wherein said binding members are selected from the group consisting of 5-hydroxyindole acetate (5-HIAA), 1,5-anhydroglucitol (1,5-AG), 3-(3-hydroxyphenyl)propionate, 3-carboxy-4-methyl-5-propyl-2-furanpropanoate (CMPF), 3-indoxyl sulfate, 4-ethylphenyl sulfate, 8-hydroxyindole acetate, 4-hydroxybenzoate ... The present invention provides a method for the preparation of a compound capable of specifically binding to a metabolite selected from the group consisting of: hydroxyoctanoate, gamma-CEHC, hydroxyisovaleroylcarnitine (C5), indole acetate, isovalerylglycine, lactate, N1-methyl-2-pyridone-5-carboxamide, p-cresol sulfate, pantothenate (vitamin B5), phenylacetylglutamine, pipecolate, xanthine, hydroxychlorothalonil, octenoylcarnitine, and 3-hydroxyhipprete.
[0177] The binding member may be selected from the group consisting of a nucleic acid molecule, a protein, a peptide, an antibody or fragments thereof, all of which are capable of binding to a specific metabolite, hi certain embodiments, the binding member may be labeled to aid in the detection and determination of the level of bound metabolite.
[0178] The present invention provides a plurality of binding members immobilized on a solid support, including 5-hydroxyindole acetate (5-HIAA), 1,5-anhydroglucitol (1,5-AG), 3-(3-hydroxyphenyl)propionate, 3-carboxy-4-methyl-5-propyl-2-furanpropanoate (CMPF), 3-indoxyl sulfate, 4-ethylphenyl sulfate, 8-hydroxyoctanoate, gamma-CEHC, hydroxyisovaleroylcarnitine (C5), indole acetate, , isovalerylglycine, lactate, N1-methyl-2-pyridone-5-carboxamide, p-cresol sulfate, pantothenate (vitamin B5), phenylacetylglutamine, pipecolate, xanthine, hydroxy-chlorothalonil, octenoylcarnitine, and 3-hydroxyhipprete, and comprising at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the population of binding members on said solid support.
[0179] In some embodiments, the present invention provides kits for carrying out the methods described herein, particularly for determining / diagnosing a subject as having or at risk of having an ASD. The kits allow a user to determine the levels of 5-hydroxyindole acetate (5-HIAA), 1,5-anhydroglucitol (1,5-AG), 3-(3-hydroxyphenyl)propionate, 3-carboxy-4-methyl-5-propyl-2-furanpropanoate (CMPF), 3-indoxyl sulfate, 4-ethylphenyl sulfate, 8-hydroxyoctanoate, gamma-CEHC, hydroxyisovaleroylcarnitine (C5), indole acetate, isovalerylglycine, lactate, N1-methyl-2-pyridone-5-carboxamide, p-cresol sulfonate, and the like in a sample under test. The kit enables the determination of the presence, level (elevated or decreased) of one or more metabolites selected from the group consisting of thiamin mononitrate, pantothenate (vitamin B5), phenylacetylglutamine, pipecolate, xanthine, hydroxychlorothalonil, octenoylcarnitine, and 3-hydroxyhipprete, and the kit includes: (a) a solid support having multiple binding members immobilized thereon, each binding member being independently specific for one of the one or more metabolites; (b) a developing agent comprising a label; and, optionally, (c) one or more components selected from the group consisting of wash solutions, diluents, and buffers. [Example]
[0180] subject Blood samples were collected from subjects between 18 and 60 months of age who were referred to 19 developmental assessment centers for evaluation of possible developmental disorders other than isolated motor disorders. Informed consent was obtained from all subjects. Subjects with a prior diagnosis of ASD from a clinic specializing in child developmental assessment or who were unable or unwilling to complete the study procedures were excluded from the study.
[0181] Subjects were enrolled in the SynapDx Autism Spectrum Disorder Gene Expression Analysis (STORY) study. The STORY study was conducted in accordance with current ICH guidelines for Good Clinical Practice (GCP) and applicable regulatory requirements. GCP is an international ethical and scientific quality standard for designing, conducting, recording, and reporting research involving the participation of human subjects. Adherence to this standard, consistent with guidelines derived from the Declaration of Helsinki, provides public assurance that the rights, safety, and satisfactory conditions of research subjects are protected and that clinical research data are trustworthy.
[0182] The results shown in Figures 1 through 12 are based on 180 blood samples from men in the STORY study. The sample set included 122 ASD samples and 58 DD (non-ASD) samples. ASD diagnoses followed DSM-V diagnostic criteria. Additional results are based on a broader set of 299 blood samples from male subjects in the STORY study. The broader sample set included 198 ASD samples and 101 DD samples.
[0183] For all studies, approximately 3 mL of blood samples were collected in EDTA tubes, and plasma was prepared by centrifuging the tubes. The plasma was then frozen and sent to the laboratory for analysis. In the laboratory, methanol extraction of the samples was performed, and the extracts were analyzed by optimized ultra-high performance liquid or gas chromatography / tandem mass spectrometry (UHPLC / MS / MS or GC / MS / MS) methods (e.g., Anal. Chem., 2009, Vol. 81, 6). See pages 656-6667).
[0184] Data analysis The metabolites in blood samples of both male and female subjects were quantified.Samples were assayed for metabolite levels, and quantified as arbitrary units of concentration, normalized to the median concentration of all samples measured on a given day.For example, a unit greater than 1 refers to the amount of metabolite that is greater than the median of the samples on that day, and a unit less than 1 refers to the amount that is less than the median.Then, cross-validation was performed, in which samples were randomly divided into non-overlapping training / test sets, and the unbiased performance of machine learning classifiers was evaluated on these sets.21 metabolites were identified that individually and collectively have high informative value for ASD prediction, especially in male subjects.
[0185] Example 1 Identifying metabolite-level information This example demonstrates that useful information for ASD risk assessment can be discerned from the identification and analysis of tail effects in the distribution of samples that would otherwise be missed in traditional analyses (e.g., mean-shift-based analyses).
[0186] Once metabolite levels are determined, there are several ways to implement the information for risk assessment, such as mean shift and tail effects. Mean shift alone was found and, while not optimal, provided some predictive information. An exemplary mean shift is shown in Figure 1. In this figure, the ASD distribution is shifted to the right of the non-ASD distribution (DD).
[0187] In addition to traditional mean shift analysis, we identified additional information from the samples. By plotting the distribution curves of metabolites for ASD and non-ASD (here, DD) samples, we found that for a subset of measured metabolites, samples from either the ASD or DD population were enriched in the right (upper) or left (lower) tail (i.e., tail effect). Figure 2 shows a representative tail effect. Notably, the two distributions shared nearly identical mean values (i.e., minimal or no mean shift). Therefore, the predictive value of metabolites would not be discernible from traditional mean shift analysis.
[0188] Metabolites may exhibit a right (upper) tail effect, a left (lower) tail effect, or both. Figure 3 shows the distribution curves for ASD and non-ASD (here, DD) for a representative metabolite, 5-HIAA. A clear right tail effect is observed; for example, the ASD distribution has a larger AUC in the right tail. Therefore, samples with high levels of this metabolite demonstrate that they are highly enriched in members of the ASD population. With this metabolite, both the mean shift (indicated by the t-test value) and the right tail (indicated by the "extreme" Fisher's test value) are statistically significant.
[0189] Figure 4 shows the distribution curves of ASD and non-ASD (here, DD) for another exemplary metabolite, gamma-CEHC. A clear left tail effect is observed, for example, the ASD distribution has a larger AUC at the left tail. Therefore, it is demonstrated that samples with low levels of this metabolite are highly enriched in members of the ASD population. With this metabolite, the mean shift (shown by t-test value) is not statistically significant, but the left tail is statistically significant.
[0190] These data indicate that the identification and analysis of tail effects provides additional information for risk assessment that cannot be obtained through traditional mean-shift analysis.
[0191] Example 2 Robust prediction of ASD from selected metabolites demonstrating tail effects This example illustrates the assessment of tail effects for the prediction of ASD. We identified statistically significant tail effects for several metabolites in samples obtained from male subjects. The tail effects individually and cumulatively provided information about which group a subject belonged to, i.e., ASD or DD. Table 1 shows an exemplary panel of 21 metabolites that exhibited tail effects of ASD versus DD with high predictive power.
[0192] Table 1B shows metabolites from a panel of 21 metabolites with tail effects predictive of ASD. The statistical significance (p-value) of each tail effect is shown, as well as its position on the distribution curve (i.e., left-tail or right-tail effect). An odds ratio greater than 1 indicates predictive power for ASD. For example, 5HIAA has a right tail with an odds ratio of 4.91, indicating that in the STORY study dataset (ASD to DD samples in a 2:1 ratio), approximately 10 ASD samples were in the right tail for every DD sample. Confidence intervals were estimated by the bootstrap method. 1000 individual bootstraps were generated from the STORY data by sampling with replacement. For each bootstrap, the tail position and corresponding odds ratio were determined. 90% confidence intervals were calculated from the distribution of observed odds ratios.
[0193] Based on these criteria, 19 metabolites from a panel of 21 metabolites were found to be predictive of ASD.
[0194] Table 1C shows metabolites with tail effects predictive of DD. The statistical significance (p-value) of each tail effect is shown, as well as its position on the distribution curve (i.e., left-tail effect or right-tail effect). An odds ratio of less than 1 indicates predictive power for DD. Based on these criteria, 8 metabolites in a panel of 21 metabolites were found to predict DD. As with ASD, odds ratios and 90% confidence intervals were determined considering a 1:2 ratio of DD to ASD samples in the STORY study.
[0195] Notably, certain metabolites demonstrated a single tail effect (either left or right) that had predictive power for either ASD or DD, whereas other metabolites demonstrated both left and right tail effects, providing combined predictive power for both ASD and DD. For example, phenylacetylglutamine and p-cresol sulfate demonstrated both right and left tail effects.
[0196] The tail effects of the 21 metabolites listed in Table 1 are shown individually in the graphs in Figures 13A to 13U. Each graph shows the distribution of one metabolite in both the ASD and DD populations. The legend above each panel indicates the statistical significance of the left and right tails for the metabolite (p-values generated by Fisher's test).
[0197] Some metabolites, such as phenylacetylglutamine, exhibit mean shift and tail effect. As shown in Figure 5, phenylacetylglutamine exhibits a statistically significant mean shift (t-test; p=0.001) and statistically significant left and right tail effect between the two populations ("extreme values" indicate tail effect, p=0.0001 in Fisher's test). The distribution is represented as a shifted Gaussian curve between the ASD population and the DD population.
[0198] Table 2 shows the threshold values used to determine the tail effect of a panel of 21 metabolites based on the baseline population distribution of each metabolite in ASD and non-ASD populations. Illustratively, the upper threshold value corresponds to the 90th percentile distribution, while the lower threshold value corresponds to the 15th percentile distribution. Absolute measurements of threshold values (e.g., ng / mL, nM, etc.) can be calculated by using the values in Table 2 along with the mean concentration of the metabolite in the population. [Table 2]
[0199] Example 3 Predicting ASD using multiple metabolites The information provided by multiple metabolites (e.g., those listed in Table 1) can be used individually or as a group to aid in the prediction of disease risk. In particular, an informative set of metabolites encompasses members that are not highly correlated with each other and have low collinearity (i.e., low correlation). For example, Figure 6 shows 5HIAA levels compared to gamma-CEHC levels, demonstrating the lack of correlation between the informative levels of the two metabolites. For example, ASD individuals identified in the 5HIAA tail (Figure 3) are generally not the same ASD individuals identified in the gamma-CEHC tail. Therefore, the metabolites 5HIAA and gamma-CEHC are considered to provide complementary information. Low-correlation metabolites enriched in the tail provide complementary classification information.
[0200] 7 is a chart showing, for each of the 180 samples, whether the sample is within or not within each of the metabolite tails of a panel of 12 metabolites. In this exemplary panel, the tails of two metabolites, xanthine and p-cresol sulfate, predict non-ASD (e.g., DD), while the tails of the other 10 metabolites predict ASD.
[0201] When multiple metabolites are assessed, the number of combinations of aggregated tail effect counts increases, as do the possible aggregated tail effect counts. The distribution of aggregated tail effect counts from ASD and non-ASD populations can be plotted, and the resulting distribution can be used to determine the appropriate sorting between ASD and non-ASD when unknown samples are measured. As shown in Figure 8A, ASD and non-ASD (here, DD) samples can be further analyzed by employing a voting (e.g., binning) scheme to further exploit the complementary information provided by metabolites for which tail effects were observed. Data for a total of 12 metabolites are shown. In one particular scheme, for a given sample, the number of metabolites for which the sample fell within the ASD predicted tail was summed, as well as the number of metabolites for which the sample fell within the non-ASD (here, DD) predicted tail. These two values are plotted and shown as x and y coordinates (Figure 8A). Notably, as the number of ASD-enriched metabolites increases (higher on the y-axis) and the number of non-ASD-enriched metabolites decreases (lower on the x-axis), fewer non-ASD dots appear to be intermingled with ASD dots, suggesting, for example, a lower likelihood of a false-positive diagnosis of ASD. Conversely, as the number of ASD-enriched metabolites decreases (lower on the y-axis) and the number of non-ASD-enriched metabolites increases (higher on the x-axis), fewer ASD dots appear to be intermingled with non-ASD dots, suggesting, for example, a lower likelihood of a false-positive diagnosis of DD.
[0202] The samples were divided into four distinct intervals, which are shown in Figure 8B. The upper and lower right intervals in particular showed clear selection, facilitating risk assessment for ASD or DD.
[0203] Of the four intervals shown in Figure 8, the interval most strongly predictive of ASD encompassed samples with two or more ASD-enriched features and either zero or one non-ASD-enriched feature. Intervals with one ASD-enriched feature and either zero or one non-ASD-enriched feature also predicted ASD, but less strongly. Intervals with one or more non-ASD-enriched features and zero ASD-enriched features strongly predicted non-ASD. Intervals with samples with no ASD-enriched features and no non-ASD-enriched features may also provide predictive information in some circumstances.
[0204] In one exemplary voting scheme, votes are tallied for a given sample, e.g., metabolites enriched in ASD score 1 point, and metabolites enriched in non-ASD subtract 1 point. Samples with a positive result (e.g., equal to or greater than 1) may be considered ASD (or have a significant risk of ASD), and samples with a negative result (equal to or less than -1) may be considered non-ASD (or have a significant likelihood of non-ASD). Samples with a zero result may be considered non-ASD or likely ASD, depending on the distribution of ASD vs. non-ASD in the sample, or may be returned as unresolved or "unclassified result" samples. Similarly, Figure 8C shows the voting results for a panel of 21 metabolites listed in Table 1.
[0205] In another exemplary scoring system (shown in Figure 21), the log2 values of the odds ratios (log2 OR) for ASD and DD features were summed for each metabolite to calculate a risk score for ASD or DD.
[0206] Tail effect information may be used to distinguish between subjects with ASD or non-ASD conditions. Similarly, tail effect information may be used to predict a subject's risk for another disease or condition, such as DD.
[0207] For example, the tail effect distribution of a non-ASD population, e.g., DD, can be used to establish a reference value for the average tail effect sum for a given number of metabolites in that population, as shown in Figures 8A and 8C. This average value can be used as a reference for comparison with the average tail effect sum from samples from unknown subjects and can be used to assess a subject's risk for ASD without having to obtain population distribution curves for the metabolites in both ASD and non-ASD populations.
[0208] Tail effect information may also be combined with traditional mean shift information and / or other classification information for improved classification results, for example, as described in the exemplary voting scheme above or similar schemes.
[0209] It is demonstrated herein that analysis of certain metabolite combinations can increase the predictability of ASD risk. For example, Figures 9-11 and 13-14A-D illustrate how the use of a voting scheme can increase the AUC of a classifier and improve its predictive ability. Using subsets of a 12-metabolite panel, the ASD predictive power (y-axis) increased as the number of metabolites in the subset increased (from 1 to 12) (Figure 9). The use of different classifiers (i.e., logistic regression, naive Bayes, or support vector machine (SVM)) and the selection of different features also impact the AUC (Figure 9). Figure 10A shows triangulated predictions of ASD risk using different features and classifiers for the same population using a 12-metabolite panel, while Figure 10B shows results using a 21-metabolite panel. Figures 11A and 11B show the improvement in ASD risk prediction using a voting scheme for a panel of 12 metabolites (Figure 11A) and a panel of 21 metabolites (Figure 11B). Taken together, these analyses demonstrate that by selecting targeted metabolites and using appropriate statistical tools, a high degree of certainty in ASD risk assessment can be achieved. For example, as shown in Figure 12, an AUC of at least 0.74 was obtained following the method described above using 12 metabolites.
[0210] Example 4 Selection of high-impact metabolites from metabolomics data Samples from ASD and DD subjects were screened for detection of approximately 600 known metabolites (shown in Table 3). From the initial set of 600, 84 candidate metabolites were identified as representing a tail effect. Subsets of the 84 metabolites detected in the samples were identified and are identified by name in Table 4. Panels of metabolites (e.g., 12- and 21-member panels) were selected from the set of 84 candidate metabolites based on high individual metabolite AUC. Certain candidate metabolites were eliminated from the panels based on factors such as association with medication or age. [Table 3-1] Table 3-2 Table 3-3 Table 3-4 Table 3-5 Table 3-6 Table 3-7 Table 3-8 Table 3-9 Table 3-10 Table 3-11 Table 3-12 Table 4-1 Table 4-2
[0211] Two panels of metabolites (a 12-metabolite panel composed of metabolites from Figure 7 and a 21-metabolite panel composed of metabolites from Table 1) were tested for ASD risk prediction. Results indicate that the 12- and 21-metabolite panels strongly contributed to ASD prediction. Figures 14A-D outline the effect of including and excluding metabolites from the 12- or 21-metabolite panels on ASD prediction. The whitelist represents the AUC values of classifiers using data from only the 12- or 21-metabolite panels, while the blacklist represents the AUC values of classifiers that excluded the 12- or 21-metabolite panels but used other metabolites from either the set of 84 candidate metabolites or the full set of 600 metabolites (total candidates = 84 candidate metabolites; total features = 600 metabolites). Mean-shift (top panel) and tail analysis (bottom panel) were performed. These data indicate that the information for predicting ASD is attributable to metabolites within the 12 or 21 metabolite panel, regardless of whether assessment was made by mean-shift or tail analysis. Thus, metabolites observed to exhibit strong tail effects (those in the 12 and 21 metabolite groups) have significantly greater predictive power for ASD versus DD than other metabolites from the 600 metabolite panel that do not exhibit strong tail effects.
[0212] Figures 14C-D expand on the results of Figures 14A-B, including additional analyses using naive Bayes analysis in addition to logistic regression. Additionally, Figure 14B shows the results of splitting samples into different cohorts (i.e., "Christmas" and "Easter"). The leftmost panel shows the AUC results for training a classifier on 192 samples and cross-validating with the Christmas cohort only. The center left panel shows the AUC results for training a classifier on 299 samples and cross-validating with the Christmas and Easter cohorts. The center right panel shows the AUC results for training a classifier on samples from the Easter cohort only and cross-validating with the Easter cohort only. The rightmost panel shows the AUC results for training a classifier on samples from the Christmas and Easter cohorts and cross-validating with the Easter cohort. The highest AUC was achieved using metabolites within the 12- or 21-metabolite panel (e.g., metabolites representing tail effects).
[0213] Figures 17A-B, 18A-B, and 19A-D expand on the results of Figure 14A-D by showing AUC predictions by including (whitelist) and excluding (blacklist) a panel of 12 or 21 metabolites according to the number of features added to the statistical analysis. The top panels show results from the mean shift analysis, and the bottom panels show the tail effect analysis. Within each individual panel, the bars represent different metabolite panels, as indicated by the symbols and legend below.
[0214] Figure 15 shows an exemplary plot illustrating the cumulative AUC of ASD risk prediction when a total of 21 metabolite subsets are assessed. In this figure, the x-axis indicates the number of metabolites from the subset selected from the group of 21 metabolites. The y-axis indicates the predictive power of ASD. For each value on the x-axis, several random metabolite combinations were analyzed, and their AUC values were plotted (dots). The curve shows the increase in AUC due to the increase in the number of metabolites (selected from the group of 21) used. Meanwhile, this figure demonstrates that even subsets with a small number of metabolites (e.g., 3 or 5) exhibit high AUC. Therefore, certain metabolites appear to have particularly important predictive tails.
[0215] Table 5 shows an exemplary table illustrating representative subsets of 21 metabolites from Table 1 containing 3, 4, 5, 6, and 7 metabolites that yielded high AUC values. For each subset size (3, 4, 5, 6, or 7), 50 random selections of metabolite sets were analyzed. For example, for a 3-subset panel of 21 metabolites, 50 random combinations of the 3-subset metabolites (out of a total of 1,330 possible permutations) were assessed. The combination from the 50 random sets with the highest AUC is shown. Therefore, certain metabolite combinations containing fewer than 21 metabolites yielded high AUC values. Metabolites such as gamma-CEHC, p-cresol sulfate, xanthine, phenylacetylglutamine, isovalerylglycine, octenoylcarnitine, and hydroxychlorothalonil appeared in multiple subsets that yielded high AUC values, indicating that these metabolites may be closely related to the patient's ASD status. Thus, these metabolites, alone or in combination with each other or additional metabolites, are believed to be particularly useful in predicting a patient's risk of ASD. [Table 5-1] [Table 5-2]
[0216] A subset of two metabolites of the 21 metabolites from Table 1 was assessed in pairwise combinations for their ability to predict ASD. Table 6 shows representative pairwise combinations with robust AUCs. Similarly, a subset of three metabolites of the 21 metabolites from Table 1 was assessed in triplet combinations for their ability to predict ASD. Table 7 shows representative triplet combinations with robust AUCs. [Table 6-1] [Table 6-2] [Table 7-1] [Table 7-2]
[0217] Example 5 Validation of classifiers on a panel of 12 metabolites Using data from 180 tested samples, approximately two-thirds of which were ASD, a classifier was generated based on 12 highly informative metabolites shown in Figure 7. The classifier was tested for its ability to distinguish ASD from non-ASD (here, DD) in a second cohort of 130 samples. This method yielded an unbiased estimate of true predictive performance, corresponding to an AUC of 0.74. Figure 12 shows a schematic representation of the process.
[0218] Example 6 Adding genetic information to metabolites may improve ASD risk prediction It has been found that adding genetic information to metabolite information improves ASD risk prediction for certain groups. For example, combining copy number variation (CNV) data with metabolite information significantly reduces the confidence interval of ASD risk prediction, as shown in Figures 16A and 16B. As Figure 16A demonstrates, adding genetic information further strengthens the separation between ASD and non-ASD groups. In addition to CNV, other genetic information, such as, but not limited to, fragile X (FXS) status, may further contribute to diagnostic tests that enable ASD risk prediction with improved accuracy and reduced type I and / or type II errors. As shown in Figure 16B, including such additional information (e.g., "PathoCV") enhanced the separation between ASD and DD groups, thus helping to distinguish these two conditions.
[0219] Example 7 Important biological pathways emerged from metabolite analysis Further analysis of the metabolite information revealed clusters of metabolites that play important roles in distinct biological pathways, as presented in Table 1. For example, seven of the 21 metabolites (33%) were associated with gut microbial activity, as shown in Table 8. All seven are amino acid metabolites. Six of the seven are metabolites of aromatic amino acids and contain a benzene ring. [Table 8-1] [Table 8-2]
[0220] Analysis of metabolites strongly associated with ASD reveals relationships with certain biological pathways, as shown in Table 1. For example, certain metabolites providing information on ASD prediction suggested defects in phase II biotransformation, impaired ability to metabolize benzene rings, dysregulated reabsorption in the kidney, dysregulated carnitine metabolism, and imbalances in the transport of large neutral amino acids to the brain. Biological pathway information can be further utilized to improve ASD risk assessment and / or investigate the etiology and pathophysiology of ASD. Such information can also be used to develop pharmaceutical therapeutics for the treatment of ASD.
[0221] Example 8 Elucidation of metabolite concentrations in blood Absolute metabolite concentrations in plasma samples were determined by mass spectrometry for 19 of the 21 metabolites described in Example 2. Absolute metabolite concentrations (ng / ml) in plasma were calculated using calibration curves generated from standard samples containing known amounts of metabolites.
[0222] Table 9A shows 17 metabolites that predict ASD. The direction of the tail effect (left or right), the threshold value for determining the presence of a tail effect (i.e., the 15th percentile for a left tail effect and the 90th percentile for a right tail effect), and the odds ratio (log2) are provided. A positive odds ratio indicates that the tail effect of the metabolite predicts ASD.
[0223] Table 9B shows seven metabolites that predict DD. The direction of the tail effect (left or right), the threshold value for determining the presence of a tail effect (i.e., 15th percentile for a left tail effect and 90th percentile for a right tail effect), and the odds ratio (log2) are provided. A negative odds ratio indicates that the tail effect of the metabolite predicts DD. [Table 9A] [Table 9B]
Claims
1. 1. A method for using levels of a plurality of metabolites as indicators for distinguishing between autism spectrum disorder (ASD) and non-ASD developmental delay (DD) in a subject, comprising: (i) measuring the level of a plurality of metabolites in a sample obtained from said subject, said plurality of metabolites comprising: 3-(3-hydroxyphenyl)propionate, and at least one metabolite selected from the group consisting of hydroxychlorothalonil, 5-hydroxyindole acetate (5-HIAA), indole acetate, p-cresol sulfate, 1,5-anhydroglucitol (1,5-AG), 3-carboxy-4-methyl-5-propyl-2-furanpropanoate (CMPF), 3-indoxyl sulfate, 4-ethylphenyl sulfate, hydroxyisovaleroylcarnitine (C5), isovalerylglycine, lactate, N1-methyl-2-pyridone-5-carboxamide, pantothenate (vitamin B5), phenylacetylglutamine, pipecolate, 3-hydroxyhipprete, and combinations thereof and (ii) (a) is indicative of ASD (ASD left-tail effect) as defined in Table 9A; or (b) is an index of DD (DD left-tail effect) as defined in Table 9B; Calculating the number of metabolites in the sample having levels at or below a predetermined threshold concentration; and / or (iii) (a) is indicative of ASD (ASD right-tail effect) as defined in Table 9A; or (b) is an index of DD (DD right-tail effect) as defined in Table 9B; calculating the number of metabolites in said sample having levels at or above a predetermined threshold concentration. wherein the number obtained in step (ii) and / or (iii) is indicative that the subject has ASD or DD.
2. 2. The method of claim 1, wherein the plurality of metabolites further comprises one or both of 3-indoxyl sulfate and 4-ethylphenyl sulfate.
3. The method according to any one of claims 1 to 2, wherein the sample is a plasma sample.
4. The method of any one of claims 1 to 3, wherein the level of the metabolite is measured by mass spectrometry.
5. The method of any one of claims 1 to 4, wherein the subject is about 54 months of age or younger.
6. The method of any one of claims 1 to 4, wherein the subject is about 36 months of age or younger.
7. 1. A method for using the level of one or more metabolites as an indicator for distinguishing between autism spectrum disorder (ASD) and non-ASD developmental delay (DD) in a subject, comprising: (i) measuring the level of one or more metabolites in a sample obtained from the subject, wherein the one or more metabolites comprise 3-(3-hydroxyphenyl)propionate; and (ii) The levels of the metabolites listed below (a) xanthine at a level of 182.7 ng / ml or greater; (b) hydroxyl-chlorothalonil at a level of 20.3 ng / ml or greater; (c) 5-hydroxyindole acetate at a level of 28.5 ng / ml or greater; (d) lactate at a level of 686,600.0 ng / ml or greater; (e) pantothenate at a level of 63.3 ng / ml or greater; (f) pipecolate at a level of 303.6 ng / ml or greater; (g) gamma-CEHC at a level of 32.0 ng / ml or less; (h) indole acetate at a level of 141.4 ng / ml or less; (i) p-cresol sulfate at a level of 182.7 ng / ml or less; (j) 1,5-anhydroglucitol (1,5-AG) at a level of 11910.3 ng / ml or less; (k) 3-carboxy-4-methyl-5-propyl-2-furanpropanoate (CMPF) at a level of 7.98 ng / ml or less; (l) 3-indoxyl sulfate at a level of 256.7 ng / ml or less; (m) 4-ethylphenyl sulfate at a level of 3.0 ng / ml or less; (n) hydroxyisovaleroylcarnitine (C5) at a level of 12.9 ng / ml or less; (o) N1-methyl-2-pyridone-5-carboxamide at a level of 124.82 ng / ml or less; and (p) phenylacetylglutamine at a level of 166.4 ng / ml or less to determine whether the subject has or is at risk for ASD; wherein the detected level of 3-(3-hydroxyphenyl)propionate of less than 5.0 ng / ml is an indication that the subject has or is at risk for DD.
8. 8. The method of claim 7, wherein the one or more metabolites comprise one or both of 3-indoxyl sulfate and 4-ethylphenyl sulfate.
9. The method according to any one of claims 7 to 8, wherein the sample is a plasma sample.
10. The method of any one of claims 7 to 9, wherein the level of the metabolite is measured by mass spectrometry.
11. 10. The method of any one of claims 7 to 9, wherein the subject is about 54 months of age or younger.
12. The method of any one of claims 7 to 9, wherein the subject is about 36 months of age or younger.
13. 1. A method for using levels of a plurality of metabolites as indicators for distinguishing between autism spectrum disorder (ASD) and non-ASD developmental delay (DD) in a subject, comprising: (i) measuring the level of a plurality of metabolites in a sample obtained from said subject, said plurality of metabolites comprising: 3-hydroxyhiprate, and at least one metabolite selected from the group consisting of hydroxychlorothalonil, 5-hydroxyindole acetate (5-HIAA), indole acetate, p-cresol sulfate, 1,5-anhydroglucitol (1,5-AG), 3-(3-hydroxyphenyl)propionate, 3-carboxy-4-methyl-5-propyl-2-furanpropanoate (CMPF), 3-indoxyl sulfate, 4-ethylphenyl sulfate, hydroxyisovaleroylcarnitine (C5), isovalerylglycine, lactate, N1-methyl-2-pyridone-5-carboxamide, pantothenate (vitamin B5), phenylacetylglutamine, pipecolate, and combinations thereof; and (ii) (a) is indicative of ASD (ASD left-tail effect) as defined in Table 9A; or (b) is an index of DD (DD left-tail effect) as defined in Table 9B; Calculating the number of metabolites in the sample having levels at or below a predetermined threshold concentration; and / or (iii) (a) is indicative of ASD (ASD right-tail effect) as defined in Table 9A; or (b) is an index of DD (DD right-tail effect) as defined in Table 9B; calculating the number of metabolites in said sample having levels at or above a predetermined threshold concentration. wherein the number obtained in step (ii) and / or (iii) is indicative that the subject has ASD or DD.
14. 14. The method of claim 13, wherein the plurality of metabolites further comprises one or both of 3-indoxyl sulfate and 4-ethylphenyl sulfate.
15. The method according to any one of claims 13 to 14, wherein the sample is a plasma sample.
16. The method of any one of claims 13 to 15, wherein the level of the metabolite is measured by mass spectrometry.
17. 17. The method of any one of claims 13 to 16, wherein the subject is about 54 months of age or younger.
18. The method of any one of claims 13 to 16, wherein the subject is about 36 months of age or younger.
19. 1. A method for using the level of one or more metabolites as an indicator for distinguishing between autism spectrum disorder (ASD) and non-ASD developmental delay (DD) in a subject, comprising: (i) measuring the level of one or more metabolites in a sample obtained from the subject, wherein the one or more metabolites include 3-hydroxyhippreate; and (ii) The levels of the metabolites listed below (a) xanthine at a level of 182.7 ng / ml or greater; (b) hydroxyl-chlorothalonil at a level of 20.3 ng / ml or greater; (c) 5-hydroxyindole acetate at a level of 28.5 ng / ml or greater; (d) lactate at a level of 686,600.0 ng / ml or greater; (e) pantothenate at a level of 63.3 ng / ml or greater; (f) pipecolate at a level of 303.6 ng / ml or greater; (g) gamma-CEHC at a level of 32.0 ng / ml or less; (h) indole acetate at a level of 141.4 ng / ml or less; (i) p-cresol sulfate at a level of 182.7 ng / ml or less; (j) 1,5-anhydroglucitol (1,5-AG) at a level of 11910.3 ng / ml or less; (k) 3-carboxy-4-methyl-5-propyl-2-furanpropanoate (CMPF) at a level of 7.98 ng / ml or less; (l) 3-indoxyl sulfate at a level of 256.7 ng / ml or less; (m) 4-ethylphenyl sulfate at a level of 3.0 ng / ml or less; (n) hydroxyisovaleroylcarnitine (C5) at a level of 12.9 ng / ml or less; (o) N1-methyl-2-pyridone-5-carboxamide at a level of 124.82 ng / ml or less; and (p) phenylacetylglutamine at a level of 166.4 ng / ml or less to determine whether the subject has or is at risk for ASD; wherein the detected level of 3-hydroxyhippreate of less than 0.86 ng / ml is an indication that the subject has or is at risk for DD.
20. 20. The method of claim 19, wherein the one or more metabolites comprise one or both of 3-indoxyl sulfate and 4-ethylphenyl sulfate.