Biomarker panels and methods for predicting preeclampsia
Patent Information
- Application Number
- JP2024548619
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-02-18
- Filing Date
- 2022-08-15
- Publication Date
- 2025-08-22
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Application No. 63 / 311,845, filed February 18, 2022, the entire contents of which are incorporated herein by reference.
[0002] Sequence Listing This application contains a Sequence Listing that has been submitted via the Patent Center. The Sequence Listing, entitled 203123-028002_PCT_SL.xml, created on August 15, 2022, is 45,061 bytes in size and is hereby incorporated by reference in its entirety.
[0003] Field The present disclosure relates generally to the field of personalized medicine, and more specifically to compositions and methods for determining the probability of preeclampsia in a pregnant woman. [Background technology]
[0004] background Preeclampsia (PE), a pregnancy-specific multisystem disorder characterized by hypertension and excessive protein excretion in the urine, is a leading cause of maternal and fetal morbidity and mortality worldwide. Preeclampsia affects at least 5-8% of all pregnancies and accounts for nearly 18% of maternal deaths in the United States. Although the disorder is likely multifactorial, most cases of preeclampsia are characterized by aberrant maternal uterine vascular remodeling by fetally derived placental trophoblast cells.
[0005] Complications of preeclampsia may include impaired placental blood flow, placental abruption, eclampsia, HELLP syndrome (hemolysis, elevated liver enzymes and low platelet count), acute renal failure, cerebral hemorrhage, liver failure or rupture, pulmonary edema, disseminated intravascular coagulation and future cardiovascular disease. Symptoms may include elevated blood pressure, swelling, sudden weight gain, headaches and vision changes, although some women remain asymptomatic.
[0006] Management of preeclampsia generally involves two options: delivery or observation. Management decisions may depend on the gestational age at which preeclampsia is diagnosed and the relative health of the fetus. The only currently effective treatment for preeclampsia is delivery of the fetus and placenta. However, the decision to deliver involves balancing the potential benefits to the fetus of further development in utero against the fetal and maternal risks of progressive disease, including the development of eclampsia, a form of preeclampsia complicated by maternal insults.
[0007] There is a great need to identify women at risk for preeclampsia because most of the currently available tests cannot predict the majority of women who will ultimately develop preeclampsia. Women identified as high risk can be scheduled for more intensive prenatal surveillance and preventive interventions. Reliable early detection of preeclampsia should allow for planning of appropriate monitoring and clinical management and potentially provide early identification of disease complications. Such monitoring and management may include more frequent assessment of blood pressure and urinary protein concentrations, uterine artery Doppler measurements, ultrasound assessment of fetal growth, and prophylactic treatment with aspirin. Finally, reliable prenatal identification of preeclampsia is also important for cost-effective allocation of monitoring resources.
[0008] The present disclosure addresses this need by providing compositions and methods for determining whether a pregnant woman is at risk for developing preeclampsia, e.g., premature onset preeclampsia or preeclampsia of any gestational age. Related advantages are also provided as well. Summary of the Invention
[0009] overview The present disclosure provides compositions and methods for predicting the probability of early-onset preeclampsia or preeclampsia at any gestational age in a pregnant woman.
[0010] In one aspect, the invention provides a panel of isolated biomarkers comprising N of the biomarkers listed in Table 1, Table 6, Table 7, Table 8 or Table 9. In some embodiments, N is a number selected from the group consisting of 2 to 12. In additional embodiments, the biomarker panel comprises at least two, at least three, or at least four isolated biomarkers selected from the group consisting of the exemplary peptides listed in Table 1.
[0011] In some embodiments, the invention provides a biomarker panel comprising at least two, at least three, or at least four isolated biomarkers selected from the group consisting of AFAM, CD14, LYAM1, IBP4, INHBC, PRG2, ENPP2, PEDF, PAPP1, SHBG, CBPN, CSH, and a ratio of IBP4 levels to SHBG levels. In some embodiments, the invention provides a biomarker panel comprising at least two, at least three, or at least four isolated biomarkers selected from the group consisting of INHBC, CD14, PEDF, AFAM, IBP4 / SHBG (a ratio of the levels of two protein biomarkers), CBPN, CSH, PRG2, SHBG, and PAPP1. In some embodiments, the invention provides a biomarker panel comprising at least two, at least three, or at least four isolated biomarkers selected from the group consisting of INHBC, CD14, PEDF, AFAM, CBPN, CSH, SHBG, and IBP4 / SHBG (a ratio of the levels of two protein biomarkers).
[0012] Also provided by the present invention is a method of determining the probability of early onset preeclampsia or preeclampsia at any gestational age in a pregnant woman, comprising: detecting a measurable feature of each of N biomarkers selected from the biomarkers listed in Table 1, Table 6, Table 7, Table 8 or Table 9 in a biological sample obtained from the pregnant woman; and analyzing the measurable feature to determine the probability of early onset preeclampsia or preeclampsia at any gestational age in the pregnant woman. In some embodiments, the measurable feature comprises a fragment or derivative of each of the N biomarkers selected from the biomarkers listed in Table 1, Table 6, Table 7, Table 8 or Table 9. In some embodiments of the disclosed methods, detecting the measurable feature comprises quantifying the amount of each of the N biomarkers selected from the biomarkers listed in Table 1, Table 6, Table 7, Table 8 or Table 9, combinations or portions and / or derivatives thereof, in a biological sample obtained from the pregnant woman. In additional embodiments, the disclosed methods of determining the probability of early-onset preeclampsia or preeclampsia at any gestational age in a pregnant woman further comprise detecting a measurable characteristic of one or more risk indicators associated with preeclampsia.
[0013] In some embodiments, the disclosed methods of determining the probability of early onset preeclampsia or preeclampsia at any gestational age in a pregnant woman include detecting a measurable feature of each of N biomarkers, where N is selected from the group consisting of 2 to 12. In further embodiments, the disclosed methods of determining the probability of early onset preeclampsia or preeclampsia at any gestational age in a pregnant woman include detecting a measurable feature of each of at least two, at least three, or at least four isolated biomarkers selected from the group consisting of the exemplary peptides listed in Table 1, Table 6, Table 7, Table 8, or Table 9.
[0014] In other embodiments, the disclosed method of determining the probability of early onset preeclampsia or preeclampsia at any gestational age in a pregnant woman comprises detecting a measurable feature of each of at least two, at least three, or at least four isolated biomarkers selected from the group consisting of AFAM, CD14, insulin-like growth factor binding protein 4 (IBP4) / sex hormone binding globulin (SHBG), INHBC, PAPP1, PEDF, CBPN, CSH, SHBG, and PRG2.
[0015] In some embodiments of the methods of determining the probability of early-onset preeclampsia or preeclampsia at any gestational age of a pregnant woman, the probability of early-onset preeclampsia or preeclampsia at any gestational age of a pregnant woman is calculated based on the quantified amounts of each of N biomarkers selected from the biomarkers listed in Table 1, Table 6, Table 7, Table 8 or Table 9. In some embodiments, the disclosed methods for determining the probability of early-onset preeclampsia or preeclampsia at any gestational age include detecting and / or quantifying one or more biomarkers using mass spectrometry, a capture agent or a combination thereof.
[0016] In some embodiments, the disclosed methods of determining the probability of early-onset preeclampsia or preeclampsia at any gestational age in a pregnant woman include an initial step of providing a biomarker panel comprising N of the biomarkers listed in Table 1, Table 6, Table 7, Table 8 or Table 9. In additional embodiments, the disclosed methods of determining the probability of early-onset preeclampsia or preeclampsia at any gestational age in a pregnant woman include an initial step of providing a biological sample from the pregnant woman.
[0017] In some embodiments, the disclosed methods of determining the probability of early onset preeclampsia or preeclampsia at any gestational age of a pregnant woman include communicating the probability to a health care provider. In additional embodiments, the communication informs a subsequent treatment decision for the pregnant woman, or in some embodiments, the method can further include a subsequent treatment decision for the pregnant woman. In further embodiments, the treatment decision includes one or more selected from the group consisting of elements of larger outpatient care, variously referred to as case management, care management, and outpatient management. Elements include oral and written patient education related to the course of hypertensive disease during pregnancy, as well as self-care measures, collection of biometric data (i.e., automated blood pressure measurements, qualitative urinary protein, and telephone transmission to a nursing call center or clinic). Treatment may also include more frequent direct contact as described above, education, and further clinical evaluation of risk factors and signs and symptoms of hypertensive disorders, such as evaluation of blood pressure and urinary protein concentration, complete blood count including platelet count and liver and renal function, presence of thrombocytopenia, oliguria, cerebral or visual symptoms, pulmonary edema or cyanosis, epigastric or right upper quadrant pain, uterine artery Doppler measurement, ultrasound evaluation of fetal growth, and amniotic fluid volume estimation. Treatment options may also include prophylactic treatment with aspirin, and oral antihypertensive therapy, which generally includes calcium channel blockers, including oral labetalol and oral long-acting nifedipine.
[0018] In some embodiments, the methods provided herein relate to determining the probability of early onset preeclampsia, hi some embodiments, the methods provided herein relate to determining the probability of preeclampsia at any gestational age.
[0019] In further embodiments, the disclosed methods of determining the probability of early onset preeclampsia or preeclampsia at any gestational age of a pregnant woman include analyzing a measurable characteristic of one or more isolated biomarkers using a predictive model. In some embodiments of the disclosed methods, the measurable characteristic of one or more isolated biomarkers is compared to a reference characteristic.
[0020] In additional embodiments, the disclosed method of determining the probability of early onset preeclampsia or preeclampsia at any gestational age in a pregnant woman comprises using one or more analyses selected from a linear discriminant analysis model, a support vector machine classification algorithm, a recursive feature elimination model, a predictive analysis of a microarray model, a logistic regression model, a CART algorithm, a flextree algorithm, a conditional interference tree model, a LART algorithm, a random forest algorithm, a MART algorithm, a machine learning algorithm, a penalized regression method, and combinations thereof. In one embodiment, the disclosed method of determining the probability of early onset preeclampsia or preeclampsia at any gestational age in a pregnant woman comprises logistic regression.
[0021] In some embodiments, a method is provided for determining the probability of early onset preeclampsia or preeclampsia at any gestational age in a pregnant woman, comprising quantifying the amount of each of N biomarkers selected from the biomarkers listed in Table 1, Table 6, Table 7, Table 8 or Table 9 in a biological sample obtained from the pregnant woman, multiplying the amounts by a predetermined coefficient, and adding the individual products to obtain a total risk score corresponding to the probability.
[0022] Kits comprising one or more agents for detecting one or more biomarkers or fragments or derivatives thereof are also provided by the present invention. In some embodiments, the one or more biomarkers are selected from the group consisting of AFAM, CD14, LYAM1, IBP4, INHBC, PRG2, ENPP2, PEDF, PAPP1, CBPN, CSH, SHBG, and the ratio of insulin-like growth factor binding protein 4 (IBP4) / sex hormone binding globulin (SHBG). In some embodiments, the one or more biomarkers are selected from the biomarkers listed in Table 1, Table 6, Table 7, Table 8, or Table 9. In some embodiments, the one or more biomarkers are selected from the group consisting of exemplary peptides listed in Table 1.
[0023] In another aspect, the present invention provides a biochip for the detection of one or more biomarkers or fragments or derivatives thereof. In some embodiments, the one or more biomarkers are selected from the group consisting of AFAM, CD14, LYAM1, IBP4, INHBC, PRG2, ENPP2, PEDF, PAPP1, CBPN, CSH, SHBG, and the ratio of insulin-like growth factor binding protein 4 (IBP4) / sex hormone binding globulin (SHBG). In some embodiments, the one or more biomarkers are selected from the biomarkers listed in Table 1, Table 6, Table 7, Table 8, or Table 9. In some embodiments, the one or more biomarkers are selected from the group consisting of the exemplary peptides listed in Table 1.
[0024] In yet another aspect, the present invention provides biomarkers for use in determining the probability of early onset preeclampsia or preeclampsia at any gestational age of a pregnant woman. In some embodiments, the biomarkers are selected from the group consisting of AFAM, CD14, LYAM1, IBP4, INHBC, PRG2, ENPP2, PEDF, PAPP1, SHBG, CBPN, CSH, and the ratio of IBP4 levels to SHBG levels. In some embodiments, the biomarkers are selected from the biomarkers listed in Table 1, Table 6, Table 7, Table 8, or Table 9. In some embodiments, the biomarkers are selected from the group consisting of exemplary peptides listed in Table 1.
[0025] Further provided by the present invention is the use of AFAM, CD14, LYAM1, IBP4, INHBC, PRG2, ENPP2, PEDF, PAPP1, SHBG, CBPN, CSH, and / or the ratio of IBP4 levels to SHBG levels as one or more biomarkers for determining the probability of early onset preeclampsia or preeclampsia at any gestational age in a pregnant woman.
[0026] In another aspect, the invention provides for the use of N biomarkers selected from the biomarkers listed in Table 1, Table 6, Table 7, Table 8 or Table 9 to determine the probability of early onset preeclampsia or preeclampsia at any gestational age in a pregnant woman. In some embodiments, N is a number selected from the group consisting of 2 to 12. In some embodiments, the N biomarkers include at least two, at least three, or at least four isolated biomarkers selected from the group consisting of exemplary peptides listed in Table 1. In other embodiments, the N biomarkers include at least two, at least three, or at least four isolated biomarkers selected from the group consisting of AFAM, CD14, LYAM1, IBP4, INHBC, PRG2, ENPP2, PEDF, PAPP1, SHBG, CBPN, CSH, and the ratio of IBP4 levels to SHBG levels. In other embodiments, the N biomarkers comprise at least three isolated biomarkers selected from the group consisting of AFAM, CD14, LYAM1, IBP4, INHBC, PRG2, ENPP2, PEDF, PAPP1, SHBG, CBPN, CSH, and a ratio of IBP4 levels to SHBG levels.
[0027] Other features and advantages of the invention will become apparent from the detailed description and claims. [Brief description of the drawings]
[0028] BRIEF DESCRIPTION OF THE DRAWINGS [Figure 1] Figure 1 shows the contribution of clinical factors and protein biomarkers to the prediction of preeclampsia (PE), individually and in pairs. The average log-likelihood ratio (logLR) across PAPR (NCT01371019) and TREETOP (NCT02787213) for the contribution of individual factors is shown on the diagonal. The logLR of pairs of factors (one from the x-axis and one from the y-axis) is shown in the triangles below the diagonal for PAPR and above the diagonal for P for TRETOP. Greyscale: scale of logLR.
[0029] [Diagram 2] Figure 2 shows additive predictive performance for prior PE by protein biomarkers (+). Seven biomarkers are shown in the left panel and five biomarkers in the right panel. Horizontal dashed lines indicate the performance of prior PE alone in PAPR (dark grey) or TREETOP (light grey). Abbreviations: HTN, hypertension; DM, diabetes mellitus; full-term, full-term delivery. Proteins are indicated by their official gene symbols.
[0030] [Diagram 3] Figure 3 shows the performance of seven biomarkers for predicting preeclampsia. Seven of the 31 protein biomarkers investigated are shown (left panel) that showed logLR and significance for predicting PREE similar to those of earlier PREE (Figure 1; logLR 8.1-31.5, p<0.005). Five were similarly predictive of early-onset PREE (right panel).
[0031] [Figure 4] FIG. 4 illustrates an exemplary conditional interference tree model with four terminal nodes for determining risk of preeclampsia.
[0032] [Diagram 5] FIG. 5 illustrates an exemplary conditional interference tree model with 10 terminal nodes for determining risk of preeclampsia. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0033] Detailed Description The present disclosure is based in part on the discovery that certain proteins and peptides in biological samples obtained from pregnant women are differentially expressed in pregnant women who are at high risk of developing preeclampsia in the future or who currently suffer from preeclampsia compared to matched controls.The present disclosure is further based in part on the unexpected discovery that a panel combining one or more of these proteins and peptides can be utilized in a method for determining the probability of early-onset preeclampsia or preeclampsia at any gestational age of a pregnant woman with relatively high sensitivity and specificity.These proteins and peptides disclosed herein serve as biomarkers for classifying test samples, predicting the probability of early-onset preeclampsia or preeclampsia at any gestational age, and monitoring the progression of early-onset preeclampsia or preeclampsia at any gestational age in pregnant women, either individually or as a panel of biomarkers. It should be noted that, in this specification, references to predicting the probability of preeclampsia may also encompass predicting the probability of early-onset preeclampsia and / or preeclampsia at any gestational age.
[0034] As used in this application, including the appended claims, the singular forms "a," "an," and "the" are used interchangeably with "at least one" and "one or more," including plural references, unless the content clearly dictates otherwise.
[0035] The term "about" is meant to encompass a deviation of plus or minus 5 percent, particularly with respect to a given quantity.
[0036] As used herein, the terms "comprises," "comprising," "includes," "including," "contains," "containing," and any variations thereof, are intended to cover a non-exclusive inclusion whereby a process, method, product-by-process, or composition of matter that comprises, includes, or contains an element or list of elements does not include only those elements, but may include other elements not expressly enumerated or inherent to such process, method, product-by-process, or composition of matter.
[0037] The term "amount" or "level" as used herein refers to the quantity of a biomarker that is detectable or measurable in a biological sample and / or control. The amount of a biomarker can be, for example, the amount of a polypeptide, the amount of a nucleic acid, or the amount of a fragment or surrogate. Alternatively, the term can include combinations thereof. The term "amount" or "level" of a biomarker is a measurable characteristic of that biomarker.
[0038] The term "biomarker" refers to a biological molecule or a fragment of a biological molecule whose measurable characteristics (such as detected presence or level, structure or sequence) can be correlated with a particular physical state or condition. The terms "marker" and "biomarker" are used interchangeably throughout this disclosure. For example, the biomarkers of the present invention are correlated with an increased likelihood of preeclampsia. Such biomarkers include, but are not limited to, biological molecules that contain nucleotides, nucleic acids, nucleosides, amino acids, sugars, fatty acids, steroids, metabolites, peptides, polypeptides, proteins, carbohydrates, lipids, hormones, antibodies, regions of interest that function as surrogates for biological macromolecules, and combinations thereof (e.g., glycoproteins, ribonucleoproteins, lipoproteins). The term also encompasses portions or fragments of biological molecules, such as peptide fragments of proteins or polypeptides comprising at least 5 contiguous amino acid residues, at least 6 contiguous amino acid residues, at least 7 contiguous amino acid residues, at least 8 contiguous amino acid residues, at least 9 contiguous amino acid residues, at least 10 contiguous amino acid residues, at least 11 contiguous amino acid residues, at least 12 contiguous amino acid residues, at least 13 contiguous amino acid residues, at least 14 contiguous amino acid residues, at least 15 contiguous amino acid residues, at least 5 contiguous amino acid residues, at least 16 contiguous amino acid residues, at least 17 contiguous amino acid residues, at least 18 contiguous amino acid residues, at least 19 contiguous amino acid residues, at least 20 contiguous amino acid residues, at least 21 contiguous amino acid residues, at least 22 contiguous amino acid residues, at least 23 contiguous amino acid residues, at least 24 contiguous amino acid residues, at least 25 contiguous amino acid residues, or more contiguous amino acid residues. The disclosed methods, kits, compositions, and panels may also refer to or include a profile or index of the expression pattern of said biomarker(s) described herein.
[0039] As used herein, unless otherwise indicated, the terms "isolated" and "purified" generally refer to a composition of matter that is removed from its original environment (e.g., the natural environment if it occurs in nature) and thus is altered from its natural state by the hand of man. An isolated protein or nucleic acid is different from the manner in which it occurs in nature.
[0040] "Measurable characteristic" is any characteristic, feature or aspect that can be determined and correlated with the probability of preeclampsia in a subject.In the case of biomarkers, such measurable characteristics include, for example, the presence, absence, or concentration or level of the biomarker, or a fragment thereof, in a biological sample; the structure of the biomarker (including, for example, altered structure, such as the presence or amount of post-translational modification, such as oxidation or glycosylation, at one or more positions on the amino acid sequence of the biomarker), or the presence of altered conformation, for example, compared with the conformation of the biomarker in a normal control subject, and / or the presence, amount, or altered structure of the biomarker as part of a profile of one or more biomarkers.In addition to biomarkers, measurable characteristics can also include risk indicators, including, for example, maternal age, race, ethnicity, medical history, past pregnancy history, obstetric history. For risk indicators, measurable characteristics can include, for example, age, pre-pregnancy weight, ethnicity, race; presence, absence or severity of diabetes, hypertension, heart disease, renal disease; incidence and / or frequency of previous preeclampsia, previous preeclampsia; presence, absence, frequency or severity of current or past smoking, illicit drug use, alcohol use; presence, absence or severity of bleeding after the 12th gestational week; cervical cerclage and transvaginal cervical length.
[0041] As used herein, the term "panel" refers to a composition, such as an array or collection, that includes one or more biomarkers. The number of biomarkers useful in a biomarker panel is based on the sensitivity and specificity values for a particular combination of biomarker values.
[0042] As used herein, the term "risk score" refers to a score that can be assigned based on comparing the measurable feature(s) of one or more biomarkers in a biological sample obtained from a pregnant woman to a standard or reference score that represents the average or normal measurable feature(s) of one or more biomarkers measured in biological samples obtained from a reference (e.g., random) pool of pregnant women. Since the measurable features of the biomarkers may not be static throughout pregnancy, in some embodiments, a standard or reference score has been obtained for a gestational time point that corresponds to the gestational time point of the pregnant woman at the time the sample was taken. The standard or reference score can be predetermined and incorporated into the predictor model such that the comparison is indirect rather than being made in practice each time a probability is determined for a subject. The risk score can be a standard (e.g., a number) or a graphic threshold (e.g., a line on a graph). The value of the risk score correlates to an upward or downward deviation from a reference measurable feature of one or more biomarkers measured in biological samples obtained from a reference pool of pregnant women. In certain embodiments, if the risk score is greater than a standard or baseline risk score (e.g., a threshold number associated with increased likelihood of preeclampsia), the pregnant woman has a high likelihood of preeclampsia (e.g., may be identified or diagnosed as having preeclampsia). In some embodiments, the magnitude of a pregnant woman's risk score, or the amount by which it exceeds a baseline risk score or threshold, may indicate or correlate with the risk level of that pregnant woman.
[0043] In the context of the present invention, the term "biological sample" encompasses any sample taken from a pregnant woman and may contain one or more biomarkers listed in Table 1, Table 6, Table 7, Table 8, or Table 9. Suitable samples in the context of the present invention include, for example, blood, plasma, serum, amniotic fluid, vaginal secretions, saliva, and urine. In some embodiments, the biological sample is selected from the group consisting of whole blood, plasma, and serum. As will be appreciated by those skilled in the art, the biological sample may include any fraction or component of blood, including, but not limited to, T cells, monocytes, neutrophils, red blood cells, platelets, and microvesicles, such as exosomes and exosome-like vesicles. In certain embodiments, the biological sample is serum.
[0044] The present disclosure provides a biomarker panel, method and kit for determining the probability of preeclampsia in pregnant women. One major advantage of the present disclosure is that the risk of developing preeclampsia can be assessed early in pregnancy so that management of the condition can be initiated in a timely manner. Sibai, Hypertension. In: Gabbe et al., eds. Obstetrics: Normal and Problem Pregnancies. 6th ed. Philadelphia, Pa: Saunders Elsevier; 2012: chap 35. The present invention is particularly beneficial for asymptomatic women who would otherwise not be identified and treated.
[0045] By way of example, the present disclosure includes a method for generating results useful in determining the probability of preeclampsia in a pregnant woman by obtaining a dataset associated with a sample, where the dataset includes at least quantitative data regarding biomarkers and panels of biomarkers identified as predictive of preeclampsia, and inputting the dataset into an analytical process that uses the dataset to generate results useful for determining the probability of preeclampsia in a pregnant woman. As described further below, this quantitative data may include amino acids, peptides, polypeptides, proteins, nucleotides, nucleic acids, nucleosides, sugars, fatty acids, steroids, metabolites, carbohydrates, lipids, hormones, antibodies, regions of interest that serve as surrogates for biological macromolecules, and combinations thereof.
[0046] In addition to the specific biomarkers identified in this disclosure, for example, by accession number, sequence or reference, the present invention also contemplates the use of biomarker variants that are at least 90% or at least 95% or at least 97% identical to the exemplified sequences, either now known or later discovered, and have utility in the methods of the present invention. These variants may represent polymorphisms, splice variants, mutations, and the like. In this regard, the present specification discloses a number of art-known proteins in the context of the present invention and provides exemplary accession numbers associated with one or more public databases, as well as exemplary references to published journal articles relating to these art-known proteins. However, one of skill in the art will readily be able to identify additional accession numbers and journal articles that can provide further characteristics of the disclosed biomarkers, and will understand that the exemplified references are in no way limiting with respect to the disclosed biomarkers. As described herein, a variety of techniques and reagents find use in the methods of the present invention. Suitable samples in the context of the present invention include, for example, blood, plasma, serum, amniotic fluid, vaginal secretions, saliva, and urine. In some embodiments, the biological sample is selected from the group consisting of whole blood, plasma, and serum. In certain embodiments, the biological sample is serum.As described herein, biomarkers can be detected by various assays and techniques known in the art.As described herein further, such assays include, but are not limited to, mass spectrometry (MS)-based assays, antibody-based assays, and the combination of these two aspects.
[0047] Protein biomarkers associated with the probability of preeclampsia in pregnant women in this disclosure include, but are not limited to, one or more isolated biomarkers listed in Table 1, Table 6, Table 7, Table 8 or Table 9. In addition to the specific biomarkers, the present disclosure further includes biomarker variants that are about 90%, about 95%, or about 97% identical to the exemplified sequences. Variants as used herein include polymorphisms, splice variants, mutations, and the like. Table 1. Biomarkers associated with preeclampsia. [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4] [Table 1-5] [Table 1-6] [Table 1-7] 1 SOM2_CSH, listed elsewhere herein, indicates that the peptide fragment corresponds to the biomarkers SOMA, SOM2, CSH1 and CSH2. 2 CSH listed elsewhere herein indicates that the peptide fragments correspond to the biomarkers CSH1 and CSH2.
[0048] Additional markers can be selected from one or more risk indicators, including but not limited to maternal age, race, ethnicity, medical history, past pregnancy history, and obstetric history.Such additional markers can include, for example, age, pre-pregnancy weight, ethnicity, race; the presence, absence or severity of diabetes, hypertension, heart disease, kidney disease; the incidence and / or frequency of previous preeclampsia, previous preeclampsia; the presence, absence, frequency or severity of current or past smoking, illicit drug use, alcohol use; the presence, absence or severity of bleeding after the 12th gestational week; cervical cerclage and transvaginal cervical length.Additional risk indicators useful as markers can be identified using learning algorithms known to those skilled in the art, such as linear discriminant analysis, support vector machine classification, recursive feature elimination, predictive analysis of microarrays, logistic regression, CART, FlexTree, LART, random forest, MART, and / or survival analysis regression, which are known to those skilled in the art and are further described herein.
[0049] Additionally, Table 1 shows the results of MRM ("multiple reaction monitoring") or SRM ("selected reaction monitoring") assays of certain biomarkers. The inclusion of stable isotope standards (SIS) can be incorporated into the methods described herein with a level of precision and can be used to quantify the corresponding unknown analytes. An additional level of specificity is provided by the co-elution of the unknown analyte and its corresponding SIS as well as the characteristics of their transitions (e.g., the similarity of the ratio of the levels of the two transitions of the unknown and the ratio of the two transitions of its corresponding SIS). In other words, the information provided herein can be used for mass spectrometry measurements of peptides ("transitions") or label surrogates ("SIS transitions"). Further information regarding MRM and SRM assays is provided herein.
[0050] Provided herein are panels of isolated biomarkers comprising N biomarkers selected from the group listed in Table 1, Table 6, Table 7, Table 8 or Table 9. In the disclosed panels of biomarkers, N can be a number selected from the group consisting of 2 to 12. In the disclosed methods, the number of biomarkers detected and the levels of which are determined can be 2, or more than 2, such as 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 12, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 or more, at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12 of which are from the group listed in Table 1, Table 6, Table 7, Table 8 or Table 9. In some embodiments, the biomarkers comprise protein biomarkers. In some embodiments, the protein biomarkers are used in combination with one or more clinical factors to determine or predict preeclampsia, e.g., previous preeclampsia. The methods of the present disclosure are useful for determining the probability (or likelihood or risk) of preeclampsia in a pregnant woman.
[0051] While certain of the biomarkers listed in Table 1, Table 6, Table 7, Table 8 or Table 9 are useful alone for determining the probability of preeclampsia in a pregnant woman, methods are also described herein for grouping multiple subsets of biomarkers, each useful as a panel of three or more biomarkers. In some embodiments, the invention provides a panel comprising N biomarkers, where N is at least three biomarkers. In other embodiments, N is selected to be any number between three and eight biomarkers.
[0052] In yet another embodiment, N is selected to be any number from 2 to 3, 2 to 4, 2 to 5, 2 to 6, 2 to 7, 2 to 8, 2 to 9, 2 to 10, 2 to 11, or 2 to 12. In another embodiment, N is selected to be any number from 3 to 4, 3 to 5, 3 to 6, 3 to 7, 3 to 8, 3 to 9, 3 to 10, 3 to 11, or 3 to 12. In another embodiment, N is selected to be any number from 4 to 5, 4 to 6, 4 to 7, 4 to 8, 4 to 9, 4 to 10, 4 to 11, or 4 to 12. In another embodiment, N is selected to be any number from 5 to 6, 5 to 7, 5 to 8, 5 to 9, 5 to 10, 5 to 11, or 5 to 12. In another embodiment, N is selected to be any number from 6 to 7, 6 to 8, 6 to 9, 6 to 10, 6 to 11, or 6 to 12. In other embodiments, N is selected to be any number from 7-8, 7-9, 7-10, 7-11, or 7-12. In other embodiments, N is selected to be any number from 8-9, 8-10, 8-11, or 8-12. In other embodiments, N is selected to be any number from 9-10, 9-11, or 9-12. In other embodiments, N is selected to be any number from 10-11, or 10-12. In other embodiments, N is selected to be any number from 11-12. It will be appreciated that N can be selected to encompass similar but higher ranges.
[0053] In certain embodiments, the panel of isolated biomarkers comprises one or more, two or more, three or more, four or more, or five isolated biomarkers comprising amino acid sequences selected from the exemplary peptides listed in Table 1.
[0054] In further embodiments, the biomarker panel comprises at least two, at least three, or at least four isolated biomarkers selected from the group consisting of AFAM, CD14, LYAM1, IBP4, INHBC, PRG2, ENPP2, PEDF, PAPP1, SHBG, CBPN, CSH, and a ratio of IBP4 levels to SHBG levels. In some embodiments, the invention provides a biomarker panel comprising at least two, at least three, or at least four isolated biomarkers selected from the group consisting of INHBC, CD14, PEDF, AFAM, IBP4 / SHBG (a ratio of the levels of two protein biomarkers), PRG2, CBPN, CSH, SHBG, and PAPP1. In some embodiments, the invention provides a biomarker panel comprising at least two, at least three, or at least four isolated biomarkers selected from the group consisting of INHBC, CD14, PEDF, AFAM, CBPN, CSH, and IBP4 / SHBG (a ratio of the levels of two protein biomarkers).
[0055] In further embodiments, the biomarker panel comprises at least two, at least three, or at least four isolated biomarkers selected from the group consisting of afamin (AFAM), monocyte differentiation antigen CD14 (CD14), ectonucleotide pyrophosphatase / phosphodiesterase family member 2 (ENPP2), L-selectin (LYAM1), insulin-like growth factor binding protein 4 (IBP4), inhibin beta C chain (INHBC), papalysin-1 (PAPP1), pigment epithelium-derived factor (PEDF), myeloid proteoglycan (PRG2), sex hormone binding globulin (SHBG). In another embodiment, the present invention provides a biomarker panel comprising at least three isolated biomarkers selected from the group consisting of sex hormone binding globulin (SHBG), afamin (AFAM) monocyte differentiation antigen CD14 (CD14), ectonucleotide pyrophosphatase / phosphodiesterase family member 2 (ENPP2), L-selectin (LYAM1), insulin-like growth factor binding protein 4 (IBP4), inhibin beta C chain (INHBC), papalysin-1 (PAPP1), pigment epithelium-derived factor (PEDF), bone marrow proteoglycan (PRG2), carboxypeptidase N catalytic chain (CBPN) and chorionic somatomammotropic hormone 1 and 2 (CSH).
[0056] In some embodiments, the panel of isolated biomarkers comprises one or more peptides comprising fragments from (a) monocyte differentiation antigen CD14 (CD14), (b) L-selectin (LYAM1), (c) insulin-like growth factor binding protein 4 (IBP4), (d) inhibin beta C chain (INHBC), (e) myeloid proteoglycan (PRG2), (f) ectonucleotide pyrophosphatase / phosphodiesterase family member 2 (ENPP2), (g) pigment epithelium-derived factor (PEDF), (h) papalysin-1 (PAPP1), (i) sex hormone-binding globulin (SHBG), (j) afamin (AFAM), (k) carboxypeptidase N catalytic chain (CBPN), (l) chorionic somatomammotropic hormone 1 and 2 (CSH).
[0057] The present invention also provides a method of determining a probability of preeclampsia in a pregnant woman, the method comprising: detecting a measurable feature of each of N biomarkers selected from the biomarkers listed in Table 1, Table 6, Table 7, Table 8 or Table 9 in a biological sample obtained from the pregnant woman; and analyzing the measurable feature to determine the probability of preeclampsia in the pregnant woman. As disclosed herein, the measurable feature comprises a fragment or derivative of each of the N biomarkers selected from the biomarkers listed in Table 1, Table 6, Table 7, Table 8 or Table 9. In some embodiments of the disclosed method, detecting the measurable feature comprises quantifying the amount of each of the N biomarkers selected from the biomarkers listed in Table 1, Table 6, Table 7, Table 8 or Table 9, combinations or portions and / or derivatives thereof, in a biological sample obtained from the pregnant woman.
[0058] In some embodiments, the present invention describes a method for predicting time to onset of preeclampsia in a pregnant woman, the method comprising: (a) providing (or obtaining) a biological sample from the pregnant woman; (b) quantifying the amount of each of N biomarkers selected from the biomarkers listed in Table 1, Table 6, Table 7, Table 8 or Table 9 in the biological sample; (c) multiplying or thresholding the amount by a predetermined factor to provide a plurality of individual products; and (d) determining the predicted onset of preeclampsia in the pregnant woman, comprising combining (e.g., adding) the individual products to obtain a composite risk score corresponding to the predicted onset of preeclampsia in the pregnant woman. Although described and exemplified with reference to a method for determining the probability of preeclampsia in a pregnant woman, the present disclosure is equally applicable to a method for predicting time to onset of preeclampsia in a pregnant woman. It will be apparent to one skilled in the art that each of the above methods has certain substantial utility and advantages with respect to maternal-fetal health considerations.
[0059] In some embodiments, the method of determining the probability of early onset preeclampsia or preeclampsia at any gestational age in a pregnant woman comprises detecting a measurable feature of each of N biomarkers, where N is selected from the group consisting of 2 to 12. In further embodiments, the disclosed method of determining the probability of early onset preeclampsia or preeclampsia at any gestational age in a pregnant woman comprises detecting a measurable feature of each of at least two, at least three, or at least four isolated biomarkers selected from the group consisting of the exemplary peptides listed in Table 1.
[0060] In additional embodiments, a method of determining the probability of early onset preeclampsia or preeclampsia at any gestational age in a pregnant woman comprises detecting a measurable feature of each of at least two, at least three, or at least four isolated biomarkers selected from the group consisting of AFAM, CD14, LYAM1, IBP4, INHBC, PRG2, ENPP2, PEDF, PAPP1, SHBG, CBPN, CSH, and the ratio of IBP4 levels to SHBG levels. In some embodiments, the invention provides a biomarker panel comprising at least two, at least three, or at least four isolated biomarkers selected from the group consisting of INHBC, CD14, PEDF, AFAM, IBP4 / SHBG (the ratio of the levels of two protein biomarkers), PRG2, CBPN, CSH, SHBG, and PAPP1. In some embodiments, the invention provides biomarker panels comprising at least two, at least three, or at least four isolated biomarkers selected from the group consisting of INHBC, CD14, PEDF, AFAM, CBPN, CSH, SHBG, and IBP4 / SHBG (the ratio of the levels of two protein biomarkers).
[0061] In an additional embodiment, a method for determining the probability of early onset preeclampsia or preeclampsia at any gestational age in a pregnant woman comprises detecting a measurable characteristic of each of at least two, at least three, or at least four isolated biomarkers selected from the group consisting of afamin (AFAM), monocyte differentiation antigen CD14 (CD14), ectonucleotide pyrophosphatase / phosphodiesterase family member 2 (ENPP2), L-selectin (LYAM1), insulin-like growth factor binding protein 4 (IBP4), inhibin beta C chain (INHBC), papalysin-1 (PAPP1), pigment epithelium-derived factor (PEDF), bone marrow proteoglycan (PRG2), sex hormone binding globulin (SHBG), (CBPN), and (CSH). In another embodiment, the invention provides a biomarker panel comprising at least three isolated biomarkers selected from the group consisting of monocyte differentiation antigen CD14 (CD14), ectonucleotide pyrophosphatase / phosphodiesterase family member 2 (ENPP2), L-selectin (LYAM1), insulin-like growth factor binding protein 4 (IBP4), inhibin beta C chain (INHBC), papalysin-1 (PAPP1), pigment epithelium-derived factor (PEDF), bone marrow proteoglycan (BMPG) (PRG2), carboxypeptidase N catalytic chain (CBPN), sex hormone-binding globulin (SHBG), and chorionic somatomammotropic hormone 1 and 2 (CSH).
[0062] In an additional embodiment, a method of determining the probability of early onset preeclampsia or preeclampsia at any gestational age in a pregnant woman comprises detecting a measurable feature of each of at least two, at least three, or at least four isolated biomarkers selected from the group consisting of afamin (AFAM), monocyte differentiation antigen CD14 (CD14), ectonucleotide pyrophosphatase / phosphodiesterase family member 2 (ENPP2), L-selectin (LYAM1), insulin-like growth factor binding protein 4 (IBP4), inhibin beta C chain (INHBC), papalysin-1 (PAPP1), pigment epithelium-derived factor (PEDF), bone marrow proteoglycan (PRG2), sex hormone binding globulin (SHBG), carboxypeptidase N catalytic chain (CBPN), and chorionic somatomammotropic hormone 1 and 2 (CSH). In another embodiment, the invention provides a biomarker panel comprising at least three isolated biomarkers selected from the group consisting of sex hormone binding globulin (SHBG), monocyte differentiation antigen CD14 (CD14), ectonucleotide pyrophosphatase / phosphodiesterase family member 2 (ENPP2), L-selectin (LYAM1), insulin-like growth factor binding protein 4 (IBP4), inhibin beta C chain (INHBC), papalysin-1 (PAPP1), pigment epithelium-derived factor (PEDF), myeloid proteoglycan (PRG2), carboxypeptidase N catalytic chain (CBPN), and chorionic somatomammotropic hormone 1 and 2 (CSH).
[0063] In some embodiments, the biomarkers detected or measured include one or more peptides comprising fragments from (a) monocyte differentiation antigen CD14 (CD14), (b) L-selectin (LYAM1), (c) insulin-like growth factor binding protein 4 (IBP4), (d) inhibin beta C chain (INHBC), (e) myeloid proteoglycan (PRG2), (f) ectonucleotide pyrophosphatase / phosphodiesterase family member 2 (ENPP2), (g) pigment epithelium-derived factor (PEDF), (h) papalysin-1 (PAPP1), (i) sex hormone-binding globulin (SHBG), (j) afamin (AFAM), (k) carboxypeptidase N catalytic chain (CBPN), (l) chorionic somatomammotropic hormone 1 and 2 (CSH).
[0064] In additional embodiments, the method of determining the probability of early onset preeclampsia or preeclampsia at any gestational age in a pregnant woman further comprises detecting a measurable characteristic of one or more risk indicia associated with early onset preeclampsia or preeclampsia at any gestational age. In additional embodiments, the risk indicia is selected from the group consisting of a history of preeclampsia, early onset preeclampsia, first pregnancy, age, obesity, diabetes, gestational diabetes, hypertension, renal disease, multiple pregnancy, pregnancy spacing, migraine headaches, rheumatoid arthritis, and lupus.
[0065] In some embodiments of the disclosed methods of determining the probability of early-onset preeclampsia or preeclampsia at any gestational age of a pregnant woman, the probability of early-onset preeclampsia or preeclampsia at any gestational age of a pregnant woman is calculated based on the quantified amounts of each of N biomarkers selected from the biomarkers listed in Table 1, Table 6, Table 7, Table 8 or Table 9. In some embodiments, the disclosed methods for determining the probability of early-onset preeclampsia or preeclampsia at any gestational age include detecting and / or quantifying one or more biomarkers using mass spectrometry, a capture agent, or a combination thereof.
[0066] In some embodiments, the disclosed methods of determining the probability of early onset preeclampsia or preeclampsia at any gestational age of a pregnant woman include an initial step of providing a biomarker panel including N of the biomarkers listed in Table 1, Table 6, Table 7, Table 8, or Table 9. In some embodiments, N includes between 2-12 biomarkers linked to Table 1, Table 6, Table 7, Table 8, Table 9, or any combination thereof. In some embodiments, N includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 of the biomarkers listed in one or more of Table 1, Table 6, Table 7, Table 8, and Table 9. In additional embodiments, the disclosed methods of determining the probability of early onset preeclampsia or preeclampsia at any gestational age of a pregnant woman include an initial step of providing a biological sample from the pregnant woman.
[0067] In some embodiments, the disclosed methods of determining the probability of early onset preeclampsia or preeclampsia at any gestational age of a pregnant woman include communicating the probability to a health care provider, hi additional embodiments, the communicating informs subsequent treatment decisions for the pregnant woman.
[0068] In some embodiments, the method of determining the probability of early onset preeclampsia or preeclampsia at any gestational age in a pregnant woman includes an additional feature that expresses the probability as a risk score.
[0069] Preeclampsia refers to a condition characterized by high blood pressure and excess protein in the urine (proteinuria), usually after 20 weeks of pregnancy, in women who previously had normal blood pressure. Preeclampsia encompasses eclampsia, a more severe form of preeclampsia that is further characterized by seizures. Preeclampsia can be further classified as mild or severe depending on the severity of the clinical symptoms. Preeclampsia usually develops during the second half of pregnancy (after 20 weeks), but can also develop shortly after birth or before 20 weeks of pregnancy.
[0070] Preeclampsia is recognized as a complex disorder that likely includes distinct pathophysiological subtypes. Such subtypes have been defined by gestational age (e.g., early onset vs. late onset) or severity (based on blood pressure, degree of proteinuria, or clinical findings). Because severity is often based on subjective clinical assessment and the onset of the disease is often unknown, gestational age at the time of delivery of the newborn is accepted as a surrogate. In particular, "indicated" delivery of an infant that effectively cures preeclampsia occurs when, in the physician's judgment, the health of the infant or mother would be at risk by continuing the pregnancy. Pregnancies requiring delivery before 37 weeks of gestation, defined herein as "early onset preeclampsia," are associated with adverse perinatal outcomes due to preterm birth and are in greatest need of advanced clinical management. For this reason, we evaluated biomarker predictive performance for preeclampsia at delivery before 37 weeks in addition to preeclampsia at various gestational ages. Because prediction at earlier gestational age cutoffs at birth (e.g., <34 weeks) may be affected by small sample sizes, we assessed the association of biomarkers and predictors with gestational age at birth as a continuous variable.
[0071] As described herein, "preeclampsia at any gestational age" refers to, for example, preeclampsia occurring at any time during pregnancy. This includes, but is not limited to, preeclampsia occurring during pregnancies requiring delivery at or after 37 weeks of gestation. In some embodiments, preeclampsia at any gestational age includes, but is not limited to, between 17 and 28 weeks of gestation at the time the biological sample is collected. In other embodiments, preeclampsia at any gestational age includes, but is not limited to, between 16 and 29 weeks, between 17 and 28 weeks, between 18 and 27 weeks, between 19 and 26 weeks, between 20 and 25 weeks, between 21 and 24 weeks, or between 22 and 23 weeks of gestation at the time the biological sample is collected. In further embodiments, preeclampsia at any gestational age includes, but is not limited to, between about 17-22 weeks, between about 16-22 weeks, between about 22-25 weeks, between about 13-25 weeks, between about 26-28 weeks, or between about 26-29 weeks of gestation at the time the biological sample is collected. Thus, preeclampsia at any gestational age can refer to, for example, 15 weeks, 16 weeks, 17 weeks, 18 weeks, 19 weeks, 20 weeks, 21 weeks, 22 weeks, 23 weeks, 24 weeks, 25 weeks, 26 weeks, 27 weeks, 28 weeks, 29 weeks, or 30 weeks of gestation. Preeclampsia at any gestational age can also include any of the pathophysiological or severity subtypes disclosed herein.
[0072] Preeclampsia has been characterized by some researchers as two distinct disease entities: early-onset preeclampsia and late-onset preeclampsia, both of which are intended to be encompassed by reference to preeclampsia herein. Early-onset preeclampsia is usually defined as preeclampsia that develops before 34 weeks of pregnancy, whereas late-onset preeclampsia develops at or after 34 weeks of pregnancy. Preeclampsia also includes postpartum preeclampsia, a less common condition that occurs when a woman has high blood pressure and excess protein in the urine immediately after delivery. Most cases of postpartum preeclampsia develop within 48 hours of delivery. However, postpartum preeclampsia can develop up to 4-6 weeks after delivery. This is known as late postpartum preeclampsia.
[0073] Clinical criteria for the diagnosis of preeclampsia are well established, for example, a blood pressure of at least 140 / 90 mmHg on two occasions, each 4-6 hours apart, and a urinary excretion of at least 0.3 grams of protein in a 24-hour urine protein excretion (or at least +1 or greater on a dipstick test). Severe preeclampsia generally refers to a blood pressure of at least 160 / 110 mmHg on at least two occasions, 6 hours apart, and a urinary protein excretion of more than 5 grams of protein in a 24-hour urine protein excretion or persistent +3 proteinuria on a dipstick test. Preeclampsia may include HELLP syndrome (hemolysis, elevated liver enzymes, low platelet count). Other elements of preeclampsia may include intrauterine growth restriction (IUGR) below the 10th percentile for US census data, persistent neurologic symptoms (headache, visual disturbances), epigastric pain, oliguria (<500 mL / 24 hours), serum creatinine greater than 1.0 mg / dL, elevated liver enzymes (>2x normal), and thrombocytopenia (<100,000 cells / μL).
[0074] In some embodiments, the pregnant woman was between 17 and 28 weeks of gestation at the time the biological sample was collected. In other embodiments, the pregnant woman was between 16 and 29 weeks, between 17 and 28 weeks, between 18 and 27 weeks, between 19 and 26 weeks, between 20 and 25 weeks, between 21 and 24 weeks, or between 22 and 23 weeks of gestation at the time the biological sample was collected. In further embodiments, the pregnant woman was between about 17 and 22 weeks, between about 16 and 22 weeks, between about 22 and 25 weeks, between about 13 and 25 weeks, between about 26 and 28 weeks, or between about 26 and 29 weeks of gestation at the time the biological sample was collected. Thus, the gestational age of the pregnant woman at the time the biological sample is taken may be 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 weeks.
[0075] In some embodiments of the claimed methods, the measurable feature comprises a fragment or derivative of each of the N biomarkers selected from the biomarkers listed in Table 1, Table 6, Table 7, Table 8 or Table 9. In additional embodiments of the claimed methods, detecting the measurable feature comprises quantifying the amount of each of the N biomarkers selected from the biomarkers listed in Table 1, Table 6, Table 7, Table 8 or Table 9, combinations or portions and / or derivatives thereof, in a biological sample obtained from the pregnant woman.
[0076] In some embodiments, calculating the probability of early onset preeclampsia or preeclampsia at any gestational age of the pregnant woman is based on the quantified amount of each of N biomarkers selected from the biomarkers listed in Table 1, Table 6, Table 7, Table 8 or Table 9. Any separation, detection and quantification method known to one of skill in the art can be used herein to measure the presence or absence (e.g., readout is present vs. absent, or detectable vs. undetectable amount) and / or amount (e.g., readout is absolute amount or relative amount, such as absolute concentration or relative concentration) of biomarkers, peptides, polypeptides, proteins and / or fragments thereof, and optionally one or more other biomarkers or fragments thereof in a sample. In some embodiments, the detection and / or quantification of one or more biomarkers comprises an assay that utilizes a capture agent. In further embodiments, the capture agent is an antibody, an antibody fragment, a nucleic acid-based protein binding reagent, a small molecule or a variant thereof. In additional embodiments, the assay is an enzyme immunoassay (EIA), an enzyme-linked immunosorbent assay (ELISA), and a radioimmunoassay (RIA). In some embodiments, the detection and / or quantification of the one or more biomarkers further comprises mass spectrometry (MS). In yet further embodiments, the mass spectrometry is co-immunoprecipitation-mass spectrometry (co-IP MS), where co-immunoprecipitation (a technique suitable for isolating total protein complexes) is followed by analysis by mass spectrometry.
[0077] As used herein, the term "mass spectrometer" refers to an instrument that can volatilize / ionize analytes to form gas phase ions and determine their absolute or relative molecular mass. Suitable methods of volatilization / ionization are matrix-assisted laser desorption / ionization (MALDI), electrospray, laser / light, thermal, electrical, atomization / nebulization, etc., or combinations thereof. Suitable forms of mass spectrometry include, but are not limited to, ion trap instruments, quadrupole instruments, electrostatic and magnetic sector instruments, time-of-flight instruments, time-of-flight tandem mass spectrometers (TOF MS / MS), Fourier transform mass spectrometers, Orbitraps, and hybrid instruments composed of various combinations of these types of mass analyzers. These instruments can then be interfaced with a variety of other instruments that fractionate the sample (e.g., liquid chromatography or solid-phase adsorption techniques based on chemical or biological properties) and ionize the sample for introduction into the mass spectrometer, including matrix-assisted laser desorption (MALDI), electrospray, or nanospray ionization (ESI), or combinations thereof.
[0078] In general, any mass spectrometry (MS) technique capable of providing accurate information on the mass of peptides, preferably fragmentation and / or (partial) amino acid sequence of selected peptides (e.g., tandem mass spectrometry, MS / MS; or post-source decay, TOF MS) can be used in the methods disclosed herein. Suitable peptide MS and MS / MS techniques and systems are well known per se (see, for example, Methods in Molecular Biology, vol. 146: "Mass Spectrometry of Proteins and Peptides", by Chapman, ed., Humana Press 2000; Biemann 1990. Methods Enzymol 193: 455-79; or Methods in Enzymology, vol. 402: "Biological Mass Spectrometry", by Burlingame, ed., Academic Press 2005) and can be used in carrying out the methods disclosed herein. Thus, in some embodiments, the disclosed methods comprise performing quantitative MS to measure one or more biomarkers. Such quantification methods can be performed in an automated (Villanueva et al., Nature Protocols (2006) 1(2):880-891) or semi-automated format. In certain embodiments, the MS can be operably linked to a liquid chromatography device (LC-MS / MS or LC-MS) or a gas chromatography device (GC-MS or GC-MS / MS). Other methods useful in this regard include isotope-coded affinity tagging (ICAT), followed by chromatography and MS / MS.
[0079] As used herein, the terms "multiple reaction monitoring (MRM)" or "selected reaction monitoring (SRM)" refer to MS-based quantification methods that are particularly useful for quantifying low abundance analytes. In an SRM experiment, a predefined precursor ion and one or more of its fragments are selected by two mass filters of a triple quadrupole instrument and monitored over time for accurate quantification. By performing an MRM experiment by rapidly switching between different precursor / fragment pairs, multiple SRM precursor and fragment ion pairs can be measured within the same experiment on the chromatographic time scale. A series of transitions (precursor / fragment ion pairs) in combination with the retention time of the target analyte (e.g., peptides or small molecules, e.g., chemical entities, steroids, hormones) can constitute a definitive assay. A large number of analytes can be quantified during a single LC-MS experiment. The terms "scheduled" or "dynamic" with respect to MRM or SRM refer to a variant of the assay in which transitions for a particular analyte are acquired only in a time window around the expected retention time, since retention time is a property that depends on the physical properties of the analyte, significantly increasing the number of analytes that can be detected and quantified in a single LC-MS experiment and contributing to the selectivity of the test. A single analyte can also be monitored in more than one transition. Finally, the assay can include a standard that corresponds to the analyte of interest (e.g., the same amino acid sequence) but differs by including a stable isotope. Stable isotope standards (SIS) can be incorporated into the assay at precise levels and used to quantify the corresponding unknown analyte. An additional level of specificity is provided by the coelution of the unknown analyte and its corresponding SIS as well as the properties of their transitions (e.g., the similarity of the ratio of the levels of two transitions of the unknown and the ratio of the two transitions of its corresponding SIS).
[0080] Mass spectrometry assays, instruments and systems suitable for biomarker peptide analysis include: Matrix-assisted laser desorption / ionization time-of-flight (MALDI-TOF) MS; MALDI-TOF post-source decay (PSD); MALDI-TOF / TOF; Surface-enhanced laser desorption / ionization time-of-flight mass spectrometry (SELDI-TOF) MS; Electrospray ionization mass spectrometry (ESI-MS); ESI-MS / MS; ESI-MS / (MS) n (n is an integer greater than 0); ESI 3D or linear (2D) ion trap MS; ESI triple quadrupole MS; ESI quadrupole orthogonal TOF (Q-TOF); ESI Fourier transform MS system; Desorption / ionization on silicon (DIOS); Secondary ion mass spectrometry (SIMS); Atmospheric pressure chemical ionization mass spectrometry (APCI-MS); APCI-MS / MS; APCI-(MS) n Atmospheric Pressure Photoionization Mass Spectrometry (APPI-MS); APPI-MS / MS; and APPI-(MS) n Examples of methods that can be used include, but are not limited to, tandem MS (MS / MS) configurations. Peptide ion fragmentation in tandem MS (MS / MS) configurations can be achieved using methods established in the art, such as, for example, collision-induced dissociation (CID). As described herein, detection and quantification of biomarkers by mass spectrometry can include multiple reaction monitoring (MRM), as described, inter alia, by Kuhn et al., Proteomics 4:1175-86 (2004). Scheduled multiple reaction monitoring (scheduled MRM) mode acquisition during LC-MS / MS analysis enhances the sensitivity and accuracy of peptide quantification. Anderson and Hunter, Molecular and Cellular Proteomics 5(4):573 (2006). As described herein, mass spectrometry-based assays can be advantageously combined with upstream peptide or protein separation or fractionation methods, such as with chromatography and other methods described herein below.
[0081] Those skilled in the art will appreciate that the amount of a biomarker can be determined using several methods, including mass spectrometry approaches such as MS / MS, LC-MS / MS, multiple reaction monitoring (MRM) or SRM and product ion monitoring (PIM), as well as antibody-based methods such as Western blot, enzyme-linked immunosorbent assay (ELISA), immunoprecipitation, immunohistochemistry, immunofluorescence, radioimmunoassay, dot blotting and immunoassays such as fluorescence-activated cell sorting (FACS). Thus, in some embodiments, determining the level of at least one biomarker comprises using immunoassays and / or mass spectrometry. In additional embodiments, the mass spectrometry method is selected from MS, MS / MS, LC-MS / MS, SRM, PIM, and other such methods known in the art. In other embodiments, the LC-MS / MS further comprises 1D LC-MS / MS, 2D LC-MS / MS or 3D LC-MS / MS. Immunoassay techniques and protocols are generally known to those of skill in the art (Price and Newman, Principles and Practice of Immunoassay, 2nd Edition, Grove's Dictionaries, 1997; and Gosling, Immunoassays: A Practical Approach, Oxford University Press, 2000). A variety of immunoassay techniques can be used, including competitive and noncompetitive immunoassays (Self et al., Curr. Opin. Biotechnol., 7:60-65 (1996)).
[0082] In further embodiments, the immunoassay is selected from Western blot, ELISA, immunoprecipitation, immunohistochemistry, immunofluorescence, radioimmunoassay (RIA), dot blotting and FACS. In certain embodiments, the immunoassay is ELISA. In still further embodiments, the ELISA is direct ELISA (enzyme-linked immunosorbent assay), indirect ELISA, sandwich ELISA, competitive ELISA, multiplex ELISA, ELISPOT technology, and other similar techniques known in the art. The principles of these immunoassay methods are known in the art, for example, John R. Crowther, The ELISA Guidebook, 1st ed. Humana Press 2000, ISBN 0896037282. Typically, ELISA is performed using antibodies, but can be performed using any capture agent that can specifically bind and detect one or more biomarkers of the present invention. Multiplex ELISAs allow the simultaneous detection of two or more analytes within a single compartment (e.g., a microplate well), usually at multiple array addresses (Nielsen and Geierstanger 2004. J Immunol Methods 290:107-20 (2004) and Ling et al. 2007. Expert Rev Mol Diagn 7:87-98 (2007)).
[0083] In some embodiments, the methods of the invention can detect one or more biomarkers using a radioimmunoassay (RIA). RIA is a competition-based assay well known in the art that involves the use of a known amount of a radioactive label (e.g., 125 I or 131It involves mixing a target analyte (I-labeled) with an antibody specific for the analyte, then adding unlabeled analyte from the sample and measuring the amount of labeled analyte displaced (see, for example, An Introduction to Radioimmunoassay and Related Techniques, by Chard T, ed. Elsevier Science 1995, ISBN 0444821198 for guidance).
[0084] Detectable labels can be used in the assays described herein for direct or indirect detection of biomarkers in the methods of the invention. A wide variety of detectable labels can be used, with the label being selected depending on the required sensitivity, ease of conjugation with the antibody, stability requirements, and available instrumentation and disposal provisions. Those skilled in the art will be familiar with the selection of appropriate detectable labels based on the assay detection of biomarkers in the methods of the invention. Suitable detectable labels include, but are not limited to, fluorescent dyes (e.g., fluorescein, fluorescein isothiocyanate (FITC), Oregon Green™, rhodamine, Texas Red, tetrarhodimine isothiocynate (TRITC), Cy3, Cy5, etc.), fluorescent markers (e.g., green fluorescent protein (GFP), phycoerythrin, etc.), enzymes (e.g., luciferase, horseradish peroxidase, alkaline phosphatase, etc.), nanoparticles, biotin, digoxigenin, metals, etc.
[0085] For mass spectrometry-based analysis, differential tagging with isotopic reagents, such as isotope-coded affinity tags (ICAT), or more recent variations using the isobaric tagging reagents iTRAQ (Applied Biosystems, Foster City, Calif.) or tandem mass tags TMT (Thermo Scientific, Rockford, Ill.), followed by multidimensional liquid chromatography (LC) and tandem mass spectrometry (MS / MS) analysis, can provide further methodologies in carrying out the methods of the invention.
[0086] Chemiluminescence assays using chemiluminescent antibodies can be used for sensitive non-radioactive detection of protein levels. Antibodies labeled with fluorescent dyes may also be suitable. Examples of fluorescent dyes include, but are not limited to, DAPI, fluorescein, Hoechst 33258, R-phycocyanin, B-phycoerythrin, R-phycoerythrin, rhodamine, Texas Red, and Lissamine. Indirect labels include various enzymes known in the art, such as, for example, horseradish peroxidase (HRP), alkaline phosphatase (AP), beta-galactosidase, urease, etc. Detection systems using suitable substrates for horseradish peroxidase, alkaline phosphatase, and beta-galactosidase are well known in the art.
[0087] Signals from direct or indirect labels can be detected using, for example, a spectrophotometer to detect color from a chromogenic substrate; 125 The assay can be performed using a radiation counter to detect radiation, such as a gamma counter to detect I; or a fluorometer to detect fluorescence in the presence of light of a particular wavelength. For detection of enzyme-linked antibodies, a spectrophotometer such as the EMAX Microplate Reader (Molecular Devices; Menlo Park, CA) can be used to perform quantitative analysis according to the manufacturer's instructions. If desired, the assays used to practice the invention can be automated or performed robotically, and signals from multiple samples can be detected simultaneously.
[0088] In some embodiments, the methods described herein include quantification of biomarkers using mass spectrometry (MS). In further embodiments, the mass spectrometry can be liquid chromatography mass spectrometry (LC-MS), multiple reaction monitoring (MRM) or selected reaction monitoring (SRM). In additional embodiments, MRM or SRM can further include scheduled MRM or scheduled SRM, respectively.
[0089] As mentioned above, chromatography can also be used to carry out the method of the present invention. Chromatography encompasses methods for separating chemical substances and generally involves a process in which a mixture of analytes is carried by a moving flow of liquid or gas ("mobile phase") and separated into components as a result of differential distribution of the analytes between the mobile phase and a stationary phase as they flow around or over the stationary liquid or solid phase (said "stationary phase"). The stationary phase can usually be a finely divided solid, a sheet of filter material, or a thin film of liquid on the surface of a solid, etc. Chromatography is well understood by those skilled in the art as a technique applicable to the separation of chemical compounds of biological origin, such as, for example, amino acids, proteins, fragments of proteins or peptides, etc.
[0090] Chromatography can be column (i.e., the stationary phase is deposited or packed within a column), preferably liquid chromatography, even more preferably high performance liquid chromatography (HPLC) or ultra-high performance / pressure liquid chromatography (UHPLC). Details of chromatography are well known in the art (Bidlingmeyer, Practical HPLC Methodology and Applications, John Wiley & Sons Inc., 1993). Exemplary types of chromatography include, but are not limited to, high performance liquid chromatography (HPLC), UHPLC, normal phase HPLC (NP-HPLC), reverse phase HPLC (RP-HPLC), ion exchange chromatography (IEC), such as cation or anion exchange chromatography, hydrophilic interaction chromatography (HILIC), hydrophobic interaction chromatography (HIC), size exclusion chromatography (SEC), including gel filtration or gel permeation chromatography, chromatofocusing, affinity chromatography, such as immunoaffinity, immobilized metal affinity chromatography, and the like. Chromatography, including one, two or more dimensions, can be used as a peptide fractionation method in combination with further peptide analysis methods, such as, for example, downstream mass spectrometry analysis as described elsewhere herein.
[0091] To measure the biomarkers in the present disclosure, additional peptide or polypeptide separation, identification or quantification methods can be used in conjunction with any of the above analytical methods as needed. Such methods include, but are not limited to, chemical extraction partitioning, isoelectric focusing (IEF), including capillary isoelectric focusing (CIEF), capillary isotachophoresis (CITP), capillary electrochromatography (CEC), etc., one-dimensional polyacrylamide gel electrophoresis (PAGE), two-dimensional polyacrylamide gel electrophoresis (2D-PAGE), capillary gel electrophoresis (CGE), capillary zone electrophoresis (CZE), micellar electrokinetic chromatography (MEKC), free-flow electrophoresis (FFE), etc.
[0092] In the context of the present invention, the term "capture agent" refers to a compound capable of specifically binding to a target, in particular a biomarker. This term includes antibodies, antibody fragments, nucleic acid-based protein binding reagents (e.g., aptamers, slow off-rate modified aptamers (SOMAmers™)), protein capture agents, natural ligands (i.e., hormones for their receptors or vice versa), small molecules or variants thereof.
[0093] The capture agent can be configured to specifically bind to a target, particularly a biomarker. The capture agent can include, but is not limited to, organic molecules, such as polypeptides, polynucleotides, and other non-polymeric molecules that can be identified by those skilled in the art. In the embodiments disclosed herein, the capture agent includes any agent that can be used to detect, purify, isolate, or enrich a target, particularly a biomarker. Any art-known affinity capture technique can be used to selectively isolate and enrich / concentrate a biomarker that is a component of a complex mixture of biological media for use in the disclosed methods.
[0094] Antibody capture agents that specifically bind to biomarkers can be prepared using any suitable method known in the art. See, for example, Coligan, Current Protocols in Immunology (1991); Harlow & Lane, Antibodies: A Laboratory Manual (1988); Goding, Monoclonal Antibodies: Principles and Practice (2d ed. 1986).Antibody capture agents can be any immunoglobulin or derivative thereof, whether natural or wholly or partially synthetically produced. All derivatives thereof that maintain specific binding ability are also included in the term.Antibody capture agents have binding domains that are homologous or largely homologous to immunoglobulin binding domains and can be derived from natural sources or can be partially or fully synthetically produced.Antibody capture agents can be monoclonal or polyclonal antibodies. In some embodiments, the antibody is a single chain antibody. Those skilled in the art will appreciate that antibodies can be provided in any of a variety of forms, including, for example, humanized, partially humanized, chimeric, chimeric humanized, and the like. The antibody capture agent may be an antibody fragment, including but not limited to Fab, Fab', F(ab')2, scFv, Fv, dsFv bispecific antibody, and Fd fragments. The antibody capture agent may be produced by any means. For example, the antibody capture agent may be enzymatically or chemically produced by fragmentation of an intact antibody and / or recombinantly produced from a gene encoding a partial antibody sequence. The antibody capture agent may include a single chain antibody fragment. Alternatively or additionally, the antibody capture agent may include multiple chains linked together, for example, by disulfide bonds, and any functional fragment obtained from such molecules, such fragments retaining the specific binding properties of the parent antibody molecule. Due to their smaller size as a functional component of the whole molecule, antibody fragments may offer advantages over intact antibodies for use in certain immunochemical techniques and experimental applications.
[0095] Suitable capture agents useful in carrying out the present invention also include aptamers. Aptamers are oligonucleotide sequences that can specifically bind to their targets through a unique three-dimensional (3-D) structure. Aptamers can include any suitable number of nucleotides, and different aptamers can have either the same or different number of nucleotides. Aptamers can be DNA or RNA or chemically modified nucleic acids, and can be single-stranded, double-stranded, or can include double-stranded regions and higher order structures. Aptamers can also be photoaptamers, where photoreactive or chemically reactive functional groups are included in the aptamer to allow it to covalently bind to the corresponding target. The use of aptamer capture agents can include the use of two or more aptamers that specifically bind to the same biomarker. Aptamers can include tags. Aptamers can be identified using any known method, including the SELEX (Systematic Evolution of Ligands by Exponential Enrichment) process. Once identified, aptamers can be prepared or synthesized according to any known method, including chemical synthesis and enzymatic synthesis, and can be used in various applications for biomarker detection. Liu et al., Curr Med Chem.18(27):4117-25(2011). Capture agents useful for carrying out the method of the present invention also include SOMAmers (slow off-rate modified aptamers), which are known in the art to have improved off-rate properties. Brody et al., J Mol Biol.422(5):595-606(2012). SOMAmers can be generated using any known method, including SELEX.
[0096] Various sample collection and preparation techniques can be used in embodiments of the present disclosure. As a non-limiting example, maternal whole blood can be processed into serum within 2 hours after collection. Serum aliquots can be barcoded and frozen at -80°C or kept on dry ice for 2.5 hours. Samples can be shipped overnight on dry ice in temperature monitored shippers. In some embodiments, thawed or hemolyzed (hemoglobin ≥ 100 mg / dL per standardized color scale) samples may not be acceptable.
[0097] The abundant proteins can then be depleted from the samples, for example using the Human 14 Multiple Affinity Removal System (MARS 14), which removes 14 of the most abundant proteins that are essentially uninformative with respect to identifying disease-associated changes in the serum proteome. For this purpose, an equal volume of each clinical or control sample can be diluted with column buffer and filtered to remove precipitates. The filtered samples can be depleted using a MARS-14 column (4.6 x 100 mm, Cat. No. 5188-6558, Agilent Technologies). The samples can be cooled in an autosampler, for example to 4°C, the depletion column can be run, for example at room temperature, and the collected fractions can be kept, for example, at 4°C until further analysis. The unbound fraction can be collected for further analysis.
[0098] A second aliquot of each clinical serum sample and each control can be diluted in ammonium bicarbonate buffer and depleted of 14 highly and approximately 60 additional moderately abundant proteins using IgY14-SuperMix (Sigma) hand-packed columns composed of 10 mL of bulk material (50% slurry, Sigma). Shi et al., Methods, 56(2):246-53 (2012). Samples can be cooled, e.g., to 4°C in an autosampler, the depletion column can be run, e.g., at room temperature, and collected fractions can be kept, e.g., at 4°C until further analysis. Unbound fractions can be collected for further analysis.
[0099] The depleted serum sample can be denatured, e.g., with trifluoroethanol, reduced, e.g., with dithiothreitol, alkylated, e.g., with iodoacetamide, and then digested, e.g., with trypsin at a trypsin:protein ratio of 1:10. After trypsin digestion, the sample can be desalted on a C18 column and the eluate lyophilized to dryness. The desalted sample can be resolubilized in a reconstitution solution containing, e.g., five internal standard peptides.
[0100] The depleted and trypsin digested samples can be analyzed using multiple reaction monitoring (sMRM) as discussed herein. Peptides can be separated, for example, on a 150mm x 0.32mm Bio-Basic C18 column (ThermoFisher) using a Waters Nano Acquity UPLC at a flow rate of 5μl / min, and eluted, for example, using an acetonitrile gradient on an AB SCIEX QTRAP 5500 with a Turbo V source (AB SCIEX, Framingham, MA). The sMRM assay can be configured to measure over 1,000 transitions corresponding to hundreds of peptides and hundreds of corresponding proteins. Chromatographic peaks can be integrated, for example, using Rosetta Elucidator software (Ceiba Solutions).
[0101] Transitions can be excluded from the analysis if their intensity area counts are below some absolute cutoff or multiple of the signal-to-noise and if they are missing in more than a few samples per batch. Intensity area counts can be log-transformed and mass spectrometry run order trends and depletion batch effects can be minimized using regression analysis.
[0102] It will be understood by those skilled in the art that biomarkers may be modified prior to analysis to improve their detection or measurement or to determine their identity. For example, biomarkers may be subjected to proteolytic digestion prior to analysis. Any protease may be used. Proteases that may cleave the biomarker into a distinct number of fragments, such as trypsin, are particularly useful. The fragments resulting from digestion serve as fingerprints of the biomarkers, thereby allowing their detection indirectly. This is particularly useful when there are biomarkers with similar molecular weights that may be confused with the biomarker in question. Proteolytic fragmentation is also useful for high molecular weight biomarkers, as smaller biomarkers are more easily resolved by mass spectrometry. In another example, biomarkers may be modified to improve detection resolution. For example, neuraminidase may be used to remove terminal sialic acid residues from glycoproteins to improve binding to anionic adsorbents and improve detection resolution. In another example, biomarkers may be modified by attachment of tags of specific molecular weights that specifically bind to molecular biomarkers, allowing them to be further differentiated. If desired, after detecting such modified biomarkers, the identity of the biomarkers can be further determined by matching the physical and chemical properties of the modified biomarkers in protein databases (e.g., UniProt, SwissProt, NCBI).
[0103] It is further recognized in the art that biomarkers in a sample can be captured on a substrate for detection. Conventional substrates include 96-well plates or nitrocellulose membranes coated with antibodies, which are then probed for the presence of proteins. Alternatively, protein-binding molecules bound to microspheres, microparticles, microbeads, beads or other particles can be used for the capture and detection of biomarkers. The protein-binding molecules can be antibodies, peptides, peptoids, aptamers, small molecule ligands or other protein-binding capture agents bound to the surface of the particles. Each protein-binding molecule can contain a unique detectable label that is coded so that it can be distinguished from other detectable labels bound to other protein-binding molecules, to enable detection of biomarkers in multiplex assays. Examples include, but are not limited to, color-coded microspheres with known fluorescence intensities (see, e.g., microspheres using xMAP technology manufactured by Luminex, Austin, TX), microspheres containing quantum dot nanocrystals (e.g., different ratios or combinations of quantum dot colors) (see, e.g., Qdot nanocrystals manufactured by Life Technologies, Carlsbad, CA), glass-coated metal nanoparticles (see, e.g., SERS nanotags manufactured by Nanoplex Technologies, Inc., Mountain View, CA), barcode materials (see, e.g., submicron-sized striped metal bars such as Nanobarcodes manufactured by Nanoplex Technologies, Inc.), microparticles encoded with colored barcodes (see, e.g., CellCards manufactured by Vitra Bioscience, vitrabio.com), glass microparticles with digital holographic code images (see, e.g., CyVera microbeads manufactured by Illumina, San Diego, CA), and the like.
[0104] In another embodiment, biochips can be used for the capture and detection of the biomarkers of the present invention. Many protein biochips are known in the art. These include, for example, protein biochips manufactured by Packard BioScience Company (Meriden Corning, Conn.), Zyomyx (Hayward, Calif.) and Phylos (Lexington, Mass.). In general, protein biochips include a substrate having a surface. A capture reagent or adsorbent is attached to the substrate's surface. Often, the surface includes multiple addressable locations, each of which has a capture agent bound thereto. The capture agent can be a biological molecule, such as a polypeptide or a nucleic acid, that captures other biomarkers in a specific manner. Alternatively, the capture agent can be a chromatographic material, such as an anion exchange material or a hydrophilic material. Examples of protein biochips are well known in the art.
[0105] Measurement of mRNA in a biological sample can be used as a surrogate for detecting the level of the corresponding protein biomarker in the biological sample. Thus, any of the biomarkers or biomarker panels described herein can also be detected by detecting the appropriate RNA. mRNA levels can be measured by reverse transcription quantitative polymerase chain reaction (RT-PCR followed by qPCR). RT-PCR is used to generate cDNA from the mRNA. The cDNA can be used in a qPCR assay to generate fluorescence as the DNA amplification process proceeds. By comparison with a standard curve, qPCR can yield absolute measurements such as the number of copies of mRNA per cell. Northern blots, microarrays, Invader assays, and RT-PCR combined with capillary electrophoresis have all been used to measure the expression level of mRNA in a sample. See Gene Expression Profiling: Methods and Protocols, Richard A. Shimkets, editor, Humana Press, 2004.
[0106] Some embodiments disclosed herein relate to diagnostic and prognostic methods for determining the probability of preeclampsia in pregnant women.The measurable feature(s) of one or more biomarkers (e.g., detecting expression levels) and / or the comparison of the measurable features of two or more biomarkers (e.g., determining the ratio of the levels of two biomarkers) can be used to determine the probability of preeclampsia in pregnant women.For example, such detection methods can be used for early diagnosis of conditions to determine whether a subject is predisposed to preeclampsia, to monitor the progression of preeclampsia or the progress of treatment protocols, to evaluate the severity of preeclampsia, to predict the outcome of preeclampsia and / or the prospects of recovery or birth at full term, or to help determine the appropriate treatment of preeclampsia.
[0107] The quantitative determination of biomarkers in biological samples can be determined by, but not limited to, the above-mentioned methods as well as any other methods known in the art.The quantitative data thus obtained can then be subjected to an analytical classification process.In such a process, raw data is manipulated according to an algorithm, where the algorithm is predefined by a training set of data, for example as described in the examples provided herein.The algorithm can utilize the training data set provided herein, or can utilize the guidelines provided herein to generate an algorithm with a different data set.
[0108] In some embodiments, analyzing the measurable features to determine the probability of preeclampsia in the pregnant woman comprises using a predictive model. In further embodiments, analyzing the measurable features to determine the probability of preeclampsia in the pregnant woman comprises comparing the measurable features to a reference feature. As will be appreciated by those skilled in the art, such comparisons can be direct comparisons with the reference feature or indirect comparisons where the reference feature is incorporated into a predictive model. In further embodiments, analyzing the measurable features to determine the probability of preeclampsia in the pregnant woman comprises one or more of a linear discriminant analysis model, a support vector machine classification algorithm, a recursive feature elimination model, a predictive analysis of a microarray model, a logistic regression model, a CART algorithm, a flextree algorithm, a LART algorithm, a random forest algorithm, a MART algorithm, a machine learning algorithm, a penalized regression method, or a combination thereof. In certain embodiments, the analysis comprises logistic regression.
[0109] The analytical classification process can use any one of a variety of statistical analysis methods to manipulate the quantitative data to provide a classification of the samples. Examples of useful methods include linear discriminant analysis, recursive feature elimination, predictive analysis of microarrays, logistic regression, CART algorithms, FlexTree algorithms, LART algorithms, random forest algorithms, MART algorithms, machine learning algorithms, etc.
[0110] Classification can be performed according to a predictive modeling method that sets a threshold value to determine the probability that a sample belongs to a given class. The probability is preferably at least 50%, or at least 60%, or at least 70%, or at least 80% or higher. Classification can also be performed by determining whether the comparison between the obtained data set and the reference data set produces a statistically significant difference. If so, the sample from which the data set is obtained is classified as not belonging to the reference data set class. Conversely, if such comparison is not statistically significantly different from the reference data set, the sample from which the data set is obtained is classified as belonging to the reference data set class.
[0111] The predictive ability of a model may be evaluated according to its ability to provide a quality metric, such as the AUC (area under the curve) or accuracy for a particular value or range of values. The area under the curve measure is useful for comparing the accuracy of classifiers across the complete data range. A classifier with a higher AUC has a higher ability to correctly classify unknowns between two groups of interest. In some embodiments, the desired quality threshold is a predictive model that classifies samples with an accuracy of at least about 0.7, at least about 0.75, at least about 0.8, at least about 0.85, at least about 0.9, at least about 0.95, or higher. As an alternative measure, the desired quality threshold can refer to a predictive model that classifies samples with an AUC of at least about 0.7, at least about 0.75, at least about 0.8, at least about 0.85, at least about 0.9, or higher.
[0112] As known in the art, the relative sensitivity and specificity of a predictive model can be adjusted to favor either the selectivity metric or the sensitivity metric, and the two metrics have an inverse relationship.The limits of such models can be adjusted to provide a selected sensitivity or specificity level depending on the specific requirements of the test being performed.One or both of the sensitivity and specificity can be at least about 0.7, at least about 0.75, at least about 0.8, at least about 0.85, at least about 0.9, or higher.
[0113] The raw data may be first analyzed by measuring the value of each biomarker, in some embodiments in triplicate or multiple triplicates. The data may be manipulated, e.g., the raw data may be transformed using a standard curve, and the average of the triplicate measurements may be used to calculate the mean and standard deviation for each patient. These values may be transformed before being used in the model, e.g., log-transformed and Box-Cox transformed (Box and Cox, Royal Stat. Soc., Series B, 26:211-246 (1964)). The data is then input into a predictive model to classify the sample according to state. The information obtained may be communicated to the patient or health care provider.
[0114] To generate a prediction model of preeclampsia, a robust data set including known control samples and samples corresponding to the target preeclampsia classification can be used in the training set.The sample size can be selected using generally accepted criteria.As mentioned above, different statistical methods can be used to obtain a highly accurate prediction model.Examples of such analysis are provided in Examples 1 and 2.
[0115] In one embodiment, hierarchical clustering is performed in the derivation of the predictive model, and Pearson correlation is employed as the clustering metric. One approach is to consider the preeclampsia dataset as a "training sample" in a "supervised learning" problem. CART is a standard in medical applications (Singer, Recursive Partitioning in the Health Sciences, Springer (1999)) and is used to convert any qualitative feature into a quantitative feature; Hotelling's T 2 classifying them by the achieved significance level, assessed by the sample reuse method for statistics; and by the appropriate application of the Lasso method. The prediction problem is transformed into a regression problem by the appropriate use of the Gini criterion for classification in assessing the quality of the regression in practice, without losing sight of the prediction.
[0116] This approach resulted in what is called FlexTree (Huang, Proc. Nat. Acad. Sci. USA 101:10529-10534 (2004)). FlexTree performs very well in simulations and is useful for implementing the claimed methods when applied to multiple forms of data. Software has been developed to automate FlexTree. Alternatively, LARTree or LART can be used (Classification Trees with Subset Analysis Selection by Turnbull (2005) Stanford University, Lasso). The names reflect a binary tree, as in CART and FlexTree; Lasso, as previously mentioned; an implementation of Lasso via what is called LARS by Efron et al. (2004) Annals of Statistics 32:407-451 (2004). See also Huang et al., Proc. Natl. Acad. Sci. USA. 101(29):10529-34 (2004). Other analytical methods that may be used include logistic regression. One method of logistic regression, Ruczinski, Journal of Computational and Graphical Statistics 12:475-512 (2003). Logistic regression is similar to CART in that its classifiers can be viewed as binary trees. It differs in that each node has a Boolean statement about the features that is more general than the simple "and" statements produced by CART.
[0117] Another approach is that of nearest shrink centroid (Tibshirani, Proc. Natl. Acad. Sci. USA 99:6567-72 (2002)). This technique is like k-means, but has the advantage of automatically selecting features as in the lasso by shrinking cluster centers to focus on a small number of informative features. This approach is available as PAM software and is widely used. Two further sets of algorithms that can be used are Random Forest (Breiman, Machine Learning 45:5-32 (2001)) and MART (Hastie, The Elements of Statistical Learning, Springer (2001)). These two methods are known in the art as "committee methods" that involve predictors that "vote" on the outcome.
[0118] To provide the significance order, a false discovery rate (FDR) can be determined. First, a null distribution of a set of dissimilarity values is generated. In one embodiment, the observed profile values are permuted to create a distribution of a set of correlation coefficients obtained by chance, thereby creating an appropriate set of null distributions of correlation coefficients (Tusher et al., Proc. Natl. Acad. Sci. USA 98, 5116-21 (2001)). The set of null distributions is obtained by permuting the values of each profile for all available profiles; calculating pairwise correlation coefficients for all profiles; calculating the probability density function of the correlation coefficients for this permutation, and repeating the procedure N times (N is a large number, typically 300). The N distributions are used to calculate an appropriate measure (mean, median, etc.) of the count of correlation coefficient values whose values exceed the (similarity) value obtained from the distribution of experimentally observed similarity values at a given significance level.
[0119] The FDR is the ratio of the number of expected false significant correlations (estimated from correlations greater than this selected Pearson correlation in a set of randomized data) to the number of correlations greater than this selected Pearson correlation (significant correlations) in the empirical data. This cutoff correlation value can be applied to the correlations between experimental profiles. Using the aforementioned distribution, a level of confidence is selected for significance. This is used to determine the lowest value of the correlation coefficient that exceeds the results obtained by chance. Using this method, a threshold is obtained for positive correlation, negative correlation, or both. Using this threshold, the user can filter the pairwise correlation coefficient observations and eliminate those that do not exceed the threshold(s). Furthermore, an estimate of the false positive rate can be obtained for a given threshold. For each of the individual "random correlation" distributions, it can be found how many observations are outside the threshold range. This procedure provides a series of counts. The mean and standard deviation of the array provide the average number of potential false positives and their standard deviation.
[0120] In an alternative analytical approach, the variables selected in the cross-sectional analysis are used separately as predictors in a time-to-event analysis (survival analysis) in which the event is the occurrence of preeclampsia and subjects without events are considered censored at birth. Given the specific pregnancy outcomes (preeclampsia events or no events), the random length of time each patient is observed, and the choice of proteomic and other features, a parametric approach to analyze survival may be better than the widely applied semi-parametric Cox model. The Weibull parametric fit of survival allows for hazard rates to be monotonically increasing, decreasing, or constant, and also has proportional hazards and accelerated failure time representations (as with the Cox model). All standard tools available in obtaining approximate maximum likelihood estimators of regression coefficients and corresponding functions are available with this model.
[0121] Furthermore, Cox model can be used, and especially, reducing the number of covariates to a manageable size with Lasso significantly simplifies the analysis, allowing the possibility of non-parametric or semi-parametric approach to predicting the time to preeclampsia.These statistical tools are known in the art and can be applied to all manner of proteomics data.A set of biomarkers, clinical data and genetic data that can be easily determined and are highly informative regarding the probability of preeclampsia in said pregnant woman and the predicted time to preeclampsia events is provided.Algorithm also provides information regarding the probability of preeclampsia in pregnant woman.
[0122] In developing a predictive model, it may be desirable to select a subset of markers, i.e., at least 3, at least 4, at least 5, at least 6, up to a complete set of markers. Usually, a subset of markers is selected that provides the needs of quantitative sample analysis, e.g., availability of reagents, convenience of quantification, etc., while maintaining a highly accurate predictive model. The selection of some informative markers for building a classification model requires the definition of a performance metric and a user-defined threshold for generating a model with useful predictive ability based on this metric. For example, the performance metric may be the AUROC, the sensitivity and / or specificity of the prediction, and the overall accuracy of the predictive model.
[0123] As will be appreciated by those skilled in the art, the analytical classification process can use any one of a variety of statistical analysis methods to manipulate the quantitative data and provide a classification of the sample. Examples of useful methods include, but are not limited to, linear discriminant analysis, recursive feature elimination, predictive analysis of microarrays, logistic regression, CART algorithm, FlexTree algorithm, LART algorithm, random forest algorithm, MART algorithm, and machine learning algorithms.
[0124] As described in Examples 1 and 2, various methods can be used in training models. The selection of a subset of markers can be for forward selection or reverse selection of a subset of markers. The number of markers that optimizes the performance of the model can be selected without using all markers. One way to define the optimal number of terms is to select the number of terms that produces a model with the desired predictive ability (e.g., AUC>0.75, or an equivalent measure of sensitivity / specificity) that is within one standard error from the maximum value obtained for this metric, using any combination and number of terms used in a given algorithm.
[0125] In some embodiments, one or more isolated biomarkers described herein are combined with clinical and / or demographic variables into a numerical combination score (sometimes referred to herein as a "classifier") that can be trained on a dataset to predict a patient's probability of a particular clinical outcome (e.g., early-onset preeclampsia or preeclampsia at any gestational age). In some embodiments, multiple clinical factors are combined into one variable. For example, as illustrated in Example 3, three clinical factors (history of preeclampsia, existing hypertension, and / or existing diabetes) can be combined into one variable "Clin3," which can be positive if any of the clinical risk factors apply to the subject. In some embodiments, one or more variables or components of the combination score are multiplied by a factor or otherwise mathematically transformed (e.g., logarithmic). In some embodiments, the total combination score can be multiplied by a factor or otherwise mathematically transformed, for example, by scaling the score to run from 0 to 100. Such a combination score can be used as a score or value (or corresponding reference score or reference value) for any of the panels, methods, and kits described herein, such as in calculating the score of a classifier. In some embodiments, the combination score is calculated according to any of the models in Table 8 of Example 3 or Tables 9-10 of Example 4.
[0126] In some embodiments, a classifier may not have all of the specified coefficients, or may have a value of 0 for one or more coefficients (thus not incorporating the corresponding variable(s)), or 1 for the coefficients. In some embodiments, a classifier includes multiple individual biomarkers, each of which has its own coefficient, e.g., A * log(IBP4 / SHBG) or B * log(CD14 / SHBG), all of which can be combined additively or multiplexed. In some embodiments, the term "log([BIOMARKER])" is the natural logarithm of the measured level or concentration of the biomarker in a sample. In some embodiments, including the classifiers listed in Table 10, the acceptable range of coefficient fold change is within a particular value (e.g., 1 / 3, 1 / 2, 1 / sqrt(2), 1, sqrt(2), 2, 3 fold change).
[0127] In some embodiments, the biomarker panels, methods, compositions, and kits described herein comprise a combination of isolated biomarkers as set forth in Table 8. In some embodiments, the biomarker panels, methods, compositions, and kits described herein comprise a combination of isolated biomarkers as set forth in Table 9. In some embodiments, the biomarker panels, methods, compositions, or kits described herein comprise a combination of CD14 and SHBG. In some embodiments, the biomarker panels, methods, compositions, or kits described herein comprise a combination of IBP4, SHBG, CD14, and PRG2. In some embodiments, the biomarker panels, methods, compositions, or kits described herein comprise a combination of AFAM and SHBG. In some embodiments, the biomarker panels, methods, compositions, or kits described herein comprise a combination of IBP4, SHBG, AFAM, and PRG2. In some embodiments, the biomarker panels, methods, compositions, or kits described herein comprise a combination of INHBC and SHBG. In some embodiments, the biomarker panels, methods, compositions, or kits described herein comprise a combination of IBP4, SHBG, AFAM, and CSH. In some embodiments, the biomarker panels, methods, compositions, or kits described herein comprise a combination of IBP4, SHBG, PEDF, and CSH. In some embodiments, the biomarker panels, methods, compositions, or kits described herein comprise a combination of IBP4, SHBG, PEDF, and CBPN. In some embodiments, the biomarker panels, methods, compositions, or kits described herein comprise a combination of IBP4, SHBG, CD14, and CBPN.
[0128] In some embodiments, the biomarker panels, methods, compositions, and kits described herein comprise a ratio of the levels of two protein biomarkers or a combination of two ratios of isolated biomarkers as set forth in Table 8. In some embodiments, the biomarker panels, methods, compositions, and kits described herein comprise a ratio of the levels of two protein biomarkers or a combination of two ratios of isolated biomarkers as set forth in Table 9. In some embodiments, the biomarker panels, methods, compositions, or kits described herein comprise a ratio of the levels of CD14 to SHBG (e.g., CD14 / SHBG). In some embodiments, the biomarker panels, methods, compositions, or kits described herein comprise a combination of a ratio of IBP4 to SHBG (e.g., IBP4 / SHBG) and a ratio of CD14 to PRG2 (e.g., CD14 / PRG2). In some embodiments, the biomarker panels, methods, compositions, or kits described herein comprise a ratio of the levels of AFAM and SHBG (e.g., AFAM / SHBG). In some embodiments, the biomarker panels, methods, compositions, or kits described herein include a combination of a ratio of IBP4 to SHBG (e.g., IBP4 / SHBG) and a ratio of AFAM to PRG2 (e.g., AFAM / PRG2). In some embodiments, the biomarker panels, methods, compositions, or kits described herein include a ratio of INHBC levels to SHBG levels (e.g., INHBC / SHBG). In some embodiments, the biomarker panels, methods, compositions, or kits described herein include a combination of a ratio of IBP4 to SHBG (e.g., IBP4 / SHBG) and a ratio of AFAM to CSH (e.g., AFAM / CSH). In some embodiments, the biomarker panels, methods, compositions, or kits described herein include a combination of a ratio of IBP4 to SHBG (e.g., IBP4 / SHBG) and a ratio of PEDF to CSH (e.g., PEDF / CSH).In some embodiments, a biomarker panel, method, composition, or kit described herein comprises a combination of a ratio of IBP4 to SHBG (e.g., IBP4 / SHBG) and a ratio of PEDF to CBPN (e.g., PEDF / CBPN). In some embodiments, a biomarker panel, method, composition, or kit described herein comprises a combination of a ratio of IBP4 to SHBG (e.g., IBP4 / SHBG) and a ratio of CD14 to CBPN (e.g., CD14 / CBPN).
[0129] In yet another aspect, the present invention provides a kit for determining the probability of preeclampsia, which can be used to detect N isolated biomarkers listed in Table 1, Table 6, Table 7, Table 8, or Table 9. For example, the kit can be used to detect one or more, two or more, three or more, four or more, or five isolated biomarkers selected from the group consisting of the exemplary peptides listed in Table 1. In another aspect, the kit can detect one or more isolated biomarkers selected from the group consisting of afamin (AFAM), monocyte differentiation antigen CD14 (CD14), ectonucleotide pyrophosphatase / phosphodiesterase family member 2 (ENPP2), L-selectin (LYAM1), insulin-like growth factor binding protein 4 (IBP4), inhibin beta C chain (INHBC), papalysin-1 (PAPP1), pigment epithelium-derived factor (PEDF), bone marrow proteoglycan (BMPG) (proteoglycan 2) (PRG2), sex hormone binding globulin (SHBG), catecholamine (CAA ... The present invention may be used to detect one or more, two or more, three or more, four or more, five or more, six or more, seven or more, or eight isolated biomarkers selected from the group consisting of carboxypeptidase N catalytic chain (CBPN), insulin-like growth factor binding protein 4 (IBP4) / sex hormone binding globulin (SHBG), and the ratio of chorionic somatomammotropin hormone 1 and 2 (CSH).
[0130] The kit may include one or more agents for detecting the biomarkers, a container for holding an isolated biological sample from a pregnant woman, and printed instructions for reacting the agents with the biological sample or a portion of the biological sample to detect the presence or amount of the isolated biomarker in the biological sample. The agents may be packaged in separate containers. The kit may further include one or more control reference samples and reagents for performing an immunoassay.
[0131] In one embodiment, the kit comprises agents for measuring the levels of at least N isolated biomarkers listed in Table 1 or Table 2. The kit can include antibodies that specifically bind to these biomarkers, for example, the kit can include at least one of an antibody that specifically binds to afamin (AFAM), an antibody that specifically binds to monocyte differentiation antigen CD14 (CD14), an antibody that specifically binds to ectonucleotide pyrophosphatase / phosphodiesterase family member 2 (ENPP2), an antibody that specifically binds to L-selectin (LYAM1), an antibody that specifically binds to insulin-like growth factor binding protein 4 (IBP4), an antibody that specifically binds to inhibin beta C chain (INHBC), an antibody that specifically binds to papalysin-1 (PAPP1), an antibody that specifically binds to pigment epithelium-derived factor (PEDF), an antibody that specifically binds to bone marrow proteoglycan (BMPG) (proteoglycan 2) (PRG2), an antibody that specifically binds to carboxypeptidase N catalytic chain (CBPN), an antibody that specifically binds to chorionic somatomammotropic hormone 1 and 2 (CSH), and an antibody that specifically binds to sex hormone-binding globulin (SHBG).
[0132] The kit can include one or more containers for the composition included in the kit. The composition can be in liquid form or can be lyophilized. Suitable containers for the composition include, for example, bottles, vials, syringes and test tubes. The containers can be made of a variety of materials, including glass or plastic. The kit can also include a package insert that includes written instructions for the method of determining the probability of preeclampsia.
[0133] In some embodiments of the present disclosure, the determination of the risk (or probability or likelihood) of preeclampsia is based on a tree model. To allow for non-linear relationships between preeclampsia and the biomarkers discussed herein, a panel can be constructed from the biomarkers in the form of a binary decision tree model. Non-limiting examples of binary decision tree models useful as described herein for determining the risk (or probability or likelihood) of preeclampsia are shown in Figures 4 and 5. The conditional inference trees in these examples were used via the party package in R (versions >= 1.3-5). In some embodiments, conditional inference trees are preferred over traditional CART models, since conditional inference uses statistical significance to create new branches, while CART uses information entropy partitioning to maximize homogeneity, and in some cases exhibits overfitting and selection bias. For further discussion, see Hothorn T, Hornik K, Zeileis A, 2006, Unbiased Recursive Partitioning: A Conditional Inference Framework, incorporated by reference.
[0134] The nodes of the exemplary tree models of Figures 4 and 5 have the following meanings: PriorPE.Y / N: "Was the subject diagnosed with preeclampsia in a previous pregnancy?"; answers No=0 and Yes=1. DM.Y / N: "Was the subject diagnosed with diabetes prior to blood draw?"; answers No=0 and Yes=1. ObRisk.Y / N: "If nulliparous, has the subject had a previous pregnancy loss? If parous, has the subject had an equal or greater number of preterm births than term births?"; answers No=0 and Yes=1. HTN.Y / N: "Was the subject diagnosed with hypertension prior to blood draw?"; answers No=0 and Yes=1. Parity: "Was the subject previously carrying a pregnancy to a viable gestational age?"; answers No=0 and Yes=1. An asterisk ( * Each protein biomarker with a marker number (SEQ ID NO: 7) indicates the concentration of the protein in the sample. In some embodiments, the concentration is measured as the ratio of the area count of the endogenous fragment to the isotope standard, and the amino acid fragments detected or measured for each protein can be as follows: INHBC (LDFHFSSDR (SEQ ID NO: 7)), PRG2 (WNFAYWAAHQPWSR (SEQ ID NO: 8)), CD14 (LTVGAAQVPAQLLVGALR (SEQ ID NO: 3)), ENPP2 (TEFLSNYLTNVDDITLVPGTLGR (SEQ ID NO: 9)), PEDF (TVQAVLTVPK (SEQ ID NO: 11)).
[0135] The performance of the tree model of FIG. 4 and Example 2 was as follows in Table 2: [Table 2]
[0136] The performance of the tree model of FIG. 5 and Example 2 was as follows in Table 3: [Table 3]
[0137] In some embodiments of the present disclosure, the determination of the risk (or probability or likelihood) of preeclampsia is based on the above regression model.Non-limiting examples of regression models useful as described herein for determining the risk (or probability or likelihood) of preeclampsia are as follows: Equation 1: PE Risk=(1.79×PriorPE.Y / N)+(0.82×ObRisk.Y / N)+(0.33×log([IBP4 * ] / [SHBG * ]))+(1.42×log[INHBC * ])-(0.77 * log([PRG2 * ]) Equation 2 (Nulliparous patients): PE Risk = (1.32 × DM.Y / N) + (0.084 × ObRisk.Y / N) + (0.90 × log([IBP4 * ] / [SHBG * ]))+(1.72×log[CD14 * ])-(1.24×log[LYAM1 * ]) Equation 3 (parous patients): PE Risk = (2.20 × PriorPE.Y / N) + (1.66 × ObRisk.Y / N) + (0.14 × log([IBP4 * ] / [SHBG * ]))+(1.50×log[CD14 * ])-(1.09×log[LYAM1 * ]) The variables in the above formula have the following meaning: "PriorPE.Y / N" means "Was the subject diagnosed with preeclampsia in a previous pregnancy?" Responses No=0 and Yes=1. "DM.Y / N" means "Was the subject diagnosed with diabetes prior to blood draw?" Responses No=0 and Yes=1. "ObRisk.Y / N" means "If nulliparous, has the subject had a previous pregnancy loss? If parous, has the subject had an equal or greater number of preterm births than term births?" Responses No=0 and Yes=1. An asterisk ( *Each protein biomarker with a marker number (A) indicates the concentration of the protein in the sample. In some embodiments, the concentration is measured as the ratio of the area count of the endogenous fragment to the isotope standard, and the amino acid fragments detected or measured for each protein can be as follows: INHBC (LDFHFSSDR (SEQ ID NO: 7)), PRG2 (WNFAYWAAHQPWSR (SEQ ID NO: 8)), CD14 (LTVGAAQVPAQLLVGALR (SEQ ID NO: 3)), LYAM1 (SYYWIGIR (SEQ ID NO: 5)), IBP4 (QCHPALDGQR (SEQ ID NO: 6)), SHBG (IALGGLLFPASNLR (SEQ ID NO: 13)). The performance of the regression model according to Equation 1 and Example 2 was as follows in Table 4. [Table 4]
[0138] The performance of the regression model according to Equations 2 and 3 and Example 2 was as follows in Table 5. [Table 5]
[0139] In Tables 2, 3, 4, and 5, "PE" refers to "preeclampsia with delivery at any gestational age" and AUC is the area under the curve for predicting preeclampsia. "PTBPE" refers to "preeclampsia with preterm birth" and AUC is the area under the curve for predicting preeclampsia with preterm birth.
[0140] In some embodiments of the above disclosure, the plurality of biomarkers is selected from the group listed in Table 6. Table 6 shows the univariate predictive performance of protein biomarkers and bivariate predictive performance of protein biomarkers in combination with a history of prior preeclampsia found in Example 2 to predict preeclampsia. Table 6. Predictive performance of protein biomarkers. [Table 6]
[0141] More specific subsets of markers in Table 6 may be used as described in this disclosure and are identified by an asterisk ( * ), i.e. Prior PREE+INHBC+CD14+PEDF+AFAM+IBP4 / SHBG (the ratio of the levels of the two protein biomarkers). The performance of this subset of Example 2 in predicting preeclampsia and early-onset preeclampsia is shown in Table 2, in comparison with the performance of prior history of preeclampsia as a predictor. Table 7. Predictive performance of previous PREE and previous PREE combined with protein biomarkers. [Table 7]
[0142] The following examples are offered by way of illustration and not by way of limitation. EXAMPLES
[0143] Example 1. Development of a sample set for biomarker discovery and validation for preeclampsia PAPR
[0144] We developed a standard protocol to govern the conduct of the Proteomic Assessment of Early Onset Risk (PAPR) clinical study. This protocol also offered the option that samples and clinical information could be used to study other pregnancy complications. Specimens were obtained from women at 11 Internal Review Board (IRB) approved facilities across the United States. After providing informed consent, serum and plasma samples were obtained, as well as pertinent information regarding patient demographic characteristics, past medical and pregnancy history, current pregnancy history and concomitant medications. 0 / 7 -28 6 / 7A total of 5,501 women were enrolled who had their blood drawn during the week. After delivery, data on maternal and infant status and complications were collected. Serum samples were processed according to a standardized protocol requiring refrigerated centrifugation, and samples were aliquoted into 0.5 ml 2D-barcoded cryovials that were then frozen at -80°C.
[0145] TREETOP
[0146] The Multicenter Evaluation of Risk Predictors of Spontaneous Preterm Birth (TREETOP) was a prospective observational study at 18 centers across the United States (ClinicalTrials.gov identifier: NCT02787213). The study was approved by each center's Institutional Review Board. The study enrolled women aged 18 years or older with a singleton pregnancy who were at low risk for PTB and had not experienced symptoms of preterm labor or membrane rupture. 0 / 7 Planned birth before the week, major anomalies or chromosomal disorders, planned cerclage or pregnancy 13 6 / 7 Women who had used progesterone after 17 weeks of pregnancy were excluded. 0 / 7 ~21 6 / 7 Pregnancy was assessed by first trimester ultrasound and determined by American College of Obstetrics and Gynecology guidelines: Committee on Obstetric Practice, American Society of Ultrasound in Medicine, Society of Maternal-Fetal Medicine; Committee Opinion No. 700: Methods for Estimating the Due Date of Delivery; Obstet. Gynecol. (2017) 129: e150-4.
[0147] Participant Selection
[0148] The subset of PAPR studies used for preeclampsia product development included selected long-term PAPR repositories. In total, the repository contained 17 0 / 7 ~28 6 / 7The study included a sample of approximately 3500 patients who underwent blood draws over a 19-week period. For TREETOP, participants were randomly assigned by a third-party statistician to Phase 1 (approximately 30% of the study population) and Phase 2 (approximately 70% of the study population). Each phase reflected the TREETOP study population as a whole in both clinical and demographic factors. The prespecified range of gestational age at blood draw for this substudy was based on a previously validated blood draw range ( 1 / 7 ~20 6 / 7 The total available subjects in this gestational age range were 909 and 1251 in the PAPR and TREETOP, respectively.
[0149] Clinical Data Collection
[0150] Clinical data were recorded using an electronic case report form. Data collected were monitored centrally and on-site and were subject to source document verification. Body mass index (BMI) was calculated using self-reported pre-pregnancy weight. Outcomes plus any complications were recorded. Births were determined to be full-term (≥ 37 years) with a specific gestational age at the time of capture delivery. 0 / 7 weeks) or preterm birth (<37 0 / 7 Neonatal outcomes were collected over the first 28 days of life. Outcome classification was by physician adjudication based on commonly used definitions.
[0151] Laboratory Methods
[0152] Samples were analyzed in a Clinical Laboratory Improvement Amendments (CLIA) and College of American Pathologists (CAP) certified laboratory using analytically validated methods. Briefly, serum was depleted of abundant proteins using a MARS-14 column (Agilent catalog number 51886558), trypsin digested, enriched with stable isotope standard (SIS) peptides, desalted, and analyzed using liquid chromatography-multiple reaction monitoring mass spectrometry. Response ratios (RR) were calculated by dividing the peak area of the endogenous peptide by the peak area of the SIS peptide. Aliquots of pregnant and non-pregnant pooled serum were included for quality control. Bradford et al., Analytical validation of protein biomarkers for risk of spontaneous preterm birth; Clin. Mass. Spectrom. (2017) 3:25-38. Routine clinical trial quality metrics monitoring analytical performance were applied to all samples.
[0153] Example 2. Analysis of transitions to identify PE biomarkers Methods. This was a secondary analysis of two large pregnancy cohorts, PAPR (NCT01371019; 11 US sites, 2011-2013) and TREETOP (NCT02787213; 18 US sites, 2016-2018), described in Example 1. Outcomes were adjudicated based on common definitions. Analyses used serum at 18-20 weeks and clinical data at enrollment. The performance of each factor was assessed using the positive likelihood ratio (LR+) of the top 10% of predictions. Goodness-of-fit logLR tests quantified the significance of each factor's contribution to prediction (corrected p<0.05). The logLR reflects the change in odds of predicting a condition, e.g., a logLR of 3 reflects a 20-fold improvement.
[0154] Results: PAPR and TREETOP contributed 909 (69 PE) and 1251 (87 PE) subjects, respectively, with 34 cases of early-onset PE each. Of the six clinical factors (Figure 1), only previous PE was a significant predictor of PE and early-onset PE in both studies (p<0.001) (PE: logLR 19.8, 37.7; LR+4.2, 5.8; early-onset PE: logLR 18.2, 20.6; LR+5.4, 6.1 in PAPR, TREETOP, respectively). Seven of the 31 biomarkers had significant logLR for PE in both studies (Figures 1 and 2) (logLR 8.1-31.5, p<0.005), and five had significant logLR for early-onset PE (INHBC, PEDF, CD14, AFM, and IBP4 / SHBG). LR+s increased with severity (PE 2–4, early-onset PE 2–6). Biomarkers paired with prior PE showed additive logLR (29.1–60.6) with higher LR+s for PE (4–7) and early-onset PE (5–10).
[0155] Figure 1 shows the contribution of clinical factors and protein biomarkers, individually and pairwise, to the prediction of preeclampsia (PE). The mean log-likelihood ratios (logLR) of PAPR (NCT01371019) and TREETOP (NCT02787213) for the contribution of individual factors are shown on the diagonal. The logLR of pairs of factors (one from the x-axis and one from the y-axis) are shown in the triangles below the diagonal for PAPR and above the diagonal for TREETOP. Color bar: scale of logLR.
[0156] Figure 2 shows additive predictive performance for prior PE by protein biomarkers (+) (seven biomarkers on the left, five biomarkers on the right). Horizontal dashed lines indicate the performance of prior PE alone in PAPR (dark grey) or TREETOP (light grey). Abbreviations: HTN, hypertension; DM, diabetes; full-term, full-term delivery. Proteins are indicated by their official gene symbols.
[0157] Figure 3 shows the performance of seven of the 31 protein biomarkers investigated (left panel), which showed logLR and significance for prediction of PREE similar to that of previous PREE (Figure 1; logLR 8.1-31.5, p<0.005). Five were similarly predictive of early-onset PREE (right panel). Pairing of protein biomarkers with previous PREE showed approximately additive logLR, demonstrating independent contribution (Figure 2). Thus, prediction improved by adding protein biomarkers to previous PREE (Table 7: PREE: PAPR ΔlogLR+26.3, ΔAUC+0.16; TREETOP ΔlogLR+22.7, ΔAUC+0.13. Early-onset PREE: PAPR ΔlogLR+15.9, ΔAUC+0.18; TREETOP ΔlogLR+12.2, ΔAUC+0.16).
[0158] Conclusions: Consistent with established evidence, previous PE predicted PE and early-onset PE in two independent studies. Other clinical factors were not consistently predictive. Although clinical factors may not always be well validated (e.g., definition of previous PE) or applicable (e.g., no previous pregnancy), it is demonstrated here that protein biomarkers are objective and biomarker panels of five and seven markers, alone and in combination with each other, are reproducible and valid as clinical risk factors. Importantly, this is true even when clinical risk factors are unavailable.
[0159] Example 3. Discovery and validation of regression-based preeclampsia predictors This study reports the discovery / validation of early-onset preeclampsia risk predictors. Discovery and validation used serum samples from blood collected from pregnant women from the PAPR (NCT01371019) and TREETOP (NCT02787213) studies between the 18th and end of the 20th week of pregnancy. PAPR and TREETOP contributed 1352 (102 PE) and 1251 (87 PE) subjects, respectively; including 45 and 34 early-onset PE cases, respectively.
[0160] The classifiers were regression models with up to two additional protein analytes in the form of ratio X / Y, possibly with the addition of the ratio IBP4 / SHBG. The log ratio of IBP4 / SHBG was included in the classifier model unless SHBG was present in an additional ratio. The three clinical factors were combined into one variable (Clin3), which is positive if any one of the clinical risk factors (history of preeclampsia, existing hypertension and / or existing diabetes) was present in the subject. This strategy resulted in the following classifier model form: 1.log(IBP4 / SHBG)+Clin3+log(X / Y) 2. Clin3+log(X / SHBG) where X and Y are protein analytes.
[0161] As a result, models have up to three coefficients. These models were trained on PAPR preeclampsia subject data (discovery) and then trained with model parameters adjusted on TREETOP (validation) for prediction of early-onset preeclampsia. These studies confirmed that 396 classifiers were highly predictive of early-onset preeclampsia. See Table 8 for classifier definitions and predictive performance. Base protein biomarkers are identified as Uniprot entry names appended to the amino acid sequences of peptides quantified by mass spectrometry. Table 8. Classifier definitions and regression-based predictive performance of early-onset preeclampsia risk predictors [Table 8-1] [Table 8-2] [Table 8-3] [Table 8-4] [Table 8-5] [Table 8-6] [Table 8-7] [Table 8-8] [Table 8-9] [Table 8-10] [Table 8-11] [Table 8-12] [Table 8-13] [Table 8-14] [Table 8-15] Key: TT = TREETOP study sample results; se75sp = specificity at 75% sensitivity; PAPR = PAPR study sample results; corPEGAB = correlation of model scores to gestational age at birth for preeclampsia sample data; IBP4 / SHBG = IBP4_QCHPALDGQR / SHBG_IALGGLLFPASNLR, SEQ ID NOs: 6 and 13
[0162] Example 4. Validation of regression-based preeclampsia predictors This study validated a generalizable classifier of early-onset preeclampsia risk with clinically valuable performance. We evaluated the reproducibility and independent predictive power of clinical risk factors and protein biomarkers in TREETOP (NCT02787213) samples, independent of those used for discovery and validation (Example 3). TREETOP validation serum samples collected from the 18th to the end of the 20th week of pregnancy included 2289 samples, including 176 preeclampsia samples, of which 62 were early-onset preeclampsia.
[0163] The classifier was evaluated for clinical validity by statistically significant enrichment of early-onset preeclampsia subjects at the classifier score threshold. Secondly, prediction of preeclampsia severity was evaluated using correlation of classifier score with gestational age at birth (GAB) among preeclampsia subjects.
[0164] Protocol Overview The parameters used to select classifiers for validation from the 396 discovery / validation candidates in Example 3 included AUC for early-onset preeclampsia, correlation of classifier probability scores with GAB among preeclampsia cases, specificity at 75% sensitivity in all subjects, and significance of novel analytes in regressions with and without clinical variables. Type I errors were controlled by fixed sequence hypothesis testing. All data used for validation remained blinded until validation. The nine selected and validated classifiers and their performance characteristics are shown in Table 9. Additional protein biomarkers are identified as Uniprot entry names appended to the amino acid sequences of peptides quantified by mass spectrometry. In modeling the presence or absence of an outcome such as early-onset preeclampsia, the regression trains a predictor of the log odds of the event occurring, as shown in the first column of Table 10, labeled "Model to Calculate Log Odds Score." The regression-fitted coefficients are applied to the relative abundance of the analytes and "Clin3" to create a score that corresponds linearly to the predicted log odds of the event for each individual. Table 9. Selected classifiers for validating regression-based preeclampsia predictors and performance characteristics [Table 9] Key: corPEGAB = correlation of model score to gestational age at birth for preeclampsia sample data; cor p-value = p-value for correlation of model score to gestational age at birth for preeclampsia sample data; IBP4 / SHBG = IBP4_QCHPALDGQR / SHBG_IALGGLLFPASNLR, SEQ ID NOs: 6 and 13 Table 10. Validated models with illustrative coefficients [Table 10-1] [Table 10-2] Key: IBP4 / SHBG=IBP4_QCHPALDGQR / SHBG_IALGGLLFPASNLR, SEQ ID NO: 6 and 13
[0165] The predicted log odds of an event are converted to probabilities using known mathematical transformations, as shown in the third column of Table 10.
[0166] Coefficients up to 3x smaller to 3x larger than the fitted regression coefficients maintain within-assay performance, as shown in column 4 of Table 10. The coefficients and their acceptable ranges are provided as non-limiting examples, as these are all assay dependent.
[0167] The coefficients from any one of the regressions described in Table 10, the analyte relative abundance, and Clin3 generate a score that corresponds linearly to the log odds of the predicted event for each individual.
[0168] All patents and publications mentioned in this specification are herein incorporated by reference to the same extent as if each individual patent and publication was specifically and individually indicated to be incorporated by reference.
[0169] From the above description, it is apparent that variations and modifications can be made to the invention described herein to adapt it to various applications and conditions, such embodiments still falling within the scope of the following claims.
[0170] The recitation of a list of elements in any definition of a variable herein includes definitions of that variable as any single element or combination (or subcombination) of the listed elements. The recitation of an embodiment herein includes that embodiment as any single embodiment or in combination with any other embodiment or portion thereof.
Claims
1. A panel of isolated biomarkers comprising at least two isolated biomarkers selected from the group consisting of INHBC, AFAM, CD14, LYAM1, IBP4, PRG2, ENPP2, PEDF, PAPP1, SHBG, CBPN, CSH and the ratio of IBP4 levels to SHBG levels.
2. 2. The panel of claim 1, wherein the panel comprises at least three isolated biomarkers selected from the group consisting of INHBC, AFAM, CD14, LYAM1, IBP4, PRG2, ENPP2, PEDF, PAPP1, SHBG, CBPN, CSH, and a ratio of IBP4 levels to SHBG levels.
3. 1. A method of analyzing a measurable feature of each of at least two isolated biomarkers selected from the group consisting of INHBC, AFAM, CD14, LYAM1, IBP4, PRG2, ENPP2, PEDF, PAPP1, SHBG, CBPN, CSH, and a ratio of IBP4 level to SHBG level in a biological sample obtained from said pregnant woman as an indicator of early-onset preeclampsia or the probability of preeclampsia at any gestational age in said pregnant woman, said method comprising: and analyzing the measurable features, wherein the analyzed measurable features are indicative of the probability of early-onset preeclampsia or preeclampsia at a gestational age in the pregnant woman. (a) the measurable characteristic comprises a fragment or derivative of each of the at least two isolated biomarkers selected from the group consisting of INHBC, AFAM, CD14, LYAM1, IBP4, PRG2, ENPP2, PEDF, PAPP1, SHBG, CBPN, CSH and the ratio of IBP4 levels to SHBG levels; (b) detecting the measurable characteristic comprises quantifying the amount of each of at least two isolated biomarkers selected from the group consisting of INHBC, AFAM, CD14, LYAM1, IBP4, PRG2, ENPP2, PEDF, PAPP1, SHBG, CBPN, CSH and a ratio of IBP4 levels to SHBG levels, combinations or portions and / or derivatives thereof, in the biological sample obtained from the pregnant woman; (c) the method further comprises calculating the probability of early-onset preeclampsia or preeclampsia at any gestational age of the pregnant woman based on the quantified amount of each of at least two isolated biomarkers selected from the group consisting of INHBC, AFAM, CD14, LYAM1, IBP4, PRG2, ENPP2, PEDF, PAPP1, SHBG, CBPN, CSH, and a ratio of IBP4 level to SHBG level. (d) the method further comprises an initial step of providing a biomarker panel comprising at least two isolated biomarkers selected from the group consisting of INHBC, AFAM, CD14, LYAM1, IBP4, PRG2, ENPP2, PEDF, PAPP1, SHBG, CBPN, CSH, and a ratio of IBP4 levels to SHBG levels; (e) further comprising the initial step of providing a biological sample from said pregnant woman; and / or (f) communicating the probability to a health care provider. The method of claim 3.
5. The method of claim 3 , wherein the analysis comprises the use of a predictive model.
6. The method of claim 5 , wherein the analyzing comprises comparing the measurable characteristic to a reference characteristic.
7. 7. The method of claim 6, wherein the analysis comprises using one or more selected from the group consisting of a linear discriminant analysis model, a support vector machine classification algorithm, a recursive feature elimination model, predictive analysis of microarray models, a logistic regression model, a CART algorithm, a FlexTree algorithm, a LART algorithm, a random forest algorithm, a MART algorithm, a machine learning algorithm, a penalized regression method, and combinations thereof.
8. The method of claim 7 , wherein the analysis comprises logistic regression.
9. The method of claim 3 , wherein the probability is expressed as a risk score.
10. 4. The method of claim 3, wherein the biological sample is selected from the group consisting of whole blood, plasma, and serum.
11. The method of claim 10, wherein the biological sample is serum.
12. The method of claim 3 , wherein the quantifying comprises mass spectrometry (MS).
13. 13. The method of claim 12, wherein the MS comprises (i) liquid chromatography-mass spectrometry (LC-MS), or (ii) multiple reaction monitoring (MRM) or selected reaction monitoring (SRM).
14. 14. The method of claim 13, wherein the MRM comprises a scheduled MRM or the SRM comprises a scheduled SRM.
15. The method of claim 3 , wherein the quantifying comprises an assay that utilizes a capture agent.
16. (a) the capture agent is selected from the group consisting of an antibody, an antibody fragment, a nucleic acid-based protein binding reagent, a small molecule or a variant thereof; or (b) the assay is selected from the group consisting of an enzyme immunoassay (EIA), an enzyme-linked immunosorbent assay (ELISA), and a radioimmunoassay (RIA); 16. The method of claim 15.
17. 17. The method of claim 16, wherein said quantifying further comprises mass spectrometry (MS).
18. 18. The method of claim 17, wherein said quantifying comprises co-immunoprecipitation mass spectrometry (co-IP MS).
19. 4. The method of claim 3, further comprising detecting one or more measurable characteristics of a risk indicator.
20. 20. The method of claim 19, wherein the one or more risk indicators are selected from the group consisting of a history of early-onset preeclampsia, severe preeclampsia, preeclampsia at any gestational age, first pregnancy, age, obesity, diabetes, gestational diabetes, hypertension, renal disease, multiple pregnancy, pregnancy spacing, new father, migraine, rheumatoid arthritis, and lupus.
21. 1. A method for obtaining a comprehensive risk score as an indication of the probability of early-onset preeclampsia or preeclampsia at any gestational age in a pregnant woman, the method comprising: (a) quantifying in a biological sample obtained from the pregnant woman the amount of each of at least two isolated biomarkers selected from the group consisting of INHBC, AFAM, CD14, LYAM1, IBP4, PRG2, ENPP2, PEDF, PAPP1, SHBG, CBPN, CSH and a ratio of IBP4 level to SHBG level; (b) multiplying the amounts by a predetermined coefficient; and (c) adding the individual products to obtain the comprehensive risk score corresponding to the probability, wherein the comprehensive risk score is indicative of the probability of early-onset preeclampsia or preeclampsia at any gestational age in the pregnant woman. (a) the probability of early-onset preeclampsia in a pregnant woman is indicated; or (b) the probability of preeclampsia in a pregnant woman at any gestational age is indicated; 22. The method of claim 3 or 21.
23. A kit comprising one or more agents for detecting at least two biomarkers or fragments or derivatives thereof, wherein the at least two biomarkers are selected from the group consisting of INHBC, AFAM, CD14, LYAM1, PRG2, ENPP2, PEDF, PAPP1, CBPN, CSH and the ratio of IBP4 levels to SHBG levels. (a) the kit is for use in determining the probability of early-onset preeclampsia in a pregnant woman; or (b) the kit is for use in determining the probability of preeclampsia at any gestational age in a pregnant woman; 24. The kit of claim 23.
25. A biochip for the detection of at least two biomarkers or fragments or derivatives thereof, wherein the at least two biomarkers are selected from the group consisting of INHBC, AFAM, CD14, LYAM1, IBP4, PRG2, ENPP2, PEDF, PAPP1, CBPN, CSH and the ratio of IBP4 levels to SHBG levels.
26. (a) The biochip is for use in determining the probability of early-onset preeclampsia in a pregnant woman; or (b) the biochip is for use in determining the probability of preeclampsia at any gestational age in a pregnant woman; 26. The biochip of claim 25.
27. A composition for use in determining the probability of early-onset preeclampsia or preeclampsia at any gestational age in a pregnant woman, comprising a biomarker, wherein the biomarker is selected from the group consisting of INHBC, AFAM, CD14, LYAM1, IBP4, PRG2, ENPP2, PEDF, PAPP1, SHBG, CBPN, CSH and the ratio of IBP4 levels to SHBG levels.
28. (a) The composition is for use in determining the probability of early-onset preeclampsia in a pregnant woman; or (b) the composition is for use in determining the probability of preeclampsia at any gestational age in a pregnant woman; 28. The composition of claim 27.
29. A composition for determining the probability of early-onset preeclampsia or preeclampsia at any gestational age in a pregnant woman, comprising at least two of the following biomarkers: INHBC, AFAM, CD14, LYAM1, IBP4, PRG2, ENPP2, PEDF, PAPP1, SHBG, CBPN, CSH, and the ratio of IBP4 levels to SHBG levels. (a) the composition is for determining the probability of early-onset preeclampsia in a pregnant woman; or (b) the composition is for determining the probability of preeclampsia at any gestational age in a pregnant woman; 30. The composition of claim 29. (a) the panel is for use in determining the probability of early-onset preeclampsia in a pregnant woman; or (b) the panel is for use in determining the probability of preeclampsia at any gestational age in a pregnant woman; The panel of claim 1.