A preeclampsia-specific circulating RNA signature
By identifying specific C-RNA molecules and their protein-coding sequences, the method enhances the early detection and risk assessment of preeclampsia, facilitating timely interventions and improving maternal and fetal outcomes.
Patent Information
- Application Number
- JP2021558675
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-11-22
- Filing Date
- 2020-11-20
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2040-11-20
AI Technical Summary
Current biomarkers for preeclampsia lack discriminatory and predictive power for early detection, posing a risk for untreated severe complications in both mother and infant.
The method involves identifying specific circulating RNA (C-RNA) molecules in a biosample from pregnant women, including a list of multiple C-RNA molecules associated with preeclampsia, and analyzing their protein-coding sequences to detect or determine the risk of preeclampsia through hybridization, PCR, microarray chip analysis, or sequencing.
This approach allows for early detection and risk assessment of preeclampsia, enabling timely intervention and reducing maternal and fetal complications.
Smart Images

Figure 0007797199000050 
Figure 0007797199000051 
Figure 0007797199000052
Abstract
Description
[Technical Field]
[0001] The present invention relates generally to methods and materials for use in the detection and early risk assessment of the pregnancy complication pre-eclampsia.
[0002] Continuing Application Data This application claims the benefit of U.S. Provisional Patent Application No. 62 / 939,324, filed November 22, 2019, which is incorporated herein by reference. [Background technology]
[0003] Preeclampsia is a condition that occurs exclusively during pregnancy and affects 5% to 8% of all pregnancies. It is the direct cause of 10% to 15% of maternal deaths and 40% of fetal deaths. The three main symptoms of preeclampsia include high blood pressure, swelling of the hands and feet, and excess protein in the urine (proteinuria), which occur after 20 weeks of pregnancy. Other signs and symptoms of preeclampsia include severe headache, vision changes (including temporary vision loss, blurred vision, or light sensitivity), nausea or vomiting, decreased urine output, decreased platelet levels (thrombocytopenia), liver dysfunction, and shortness of breath caused by fluid in the lungs.
[0004] The more severe the preeclampsia and the earlier it occurs in pregnancy, the greater the risks to the mother and infant. Preeclampsia may require induction of labor and delivery or delivery by Caesarean section. If left untreated, preeclampsia can lead to serious, even fatal, complications for both the mother and infant. Complications of preeclampsia include fetal growth restriction, low birth weight, preterm delivery, placental abruption, HELLP syndrome (hemolysis, high liver enzymes, and low platelet count syndrome), eclampsia (a severe form of preeclampsia that leads to seizures), organ damage including kidney, liver, lung, heart, or eye damage, stroke, or other brain trauma. See, e.g., "Preeclampsia - Symptoms and Causes - Mayo Clinic" (April 3, 2018), available on the World Wide Web at mayoclinic.org / diseases-conditions / preeclampsia / symptoms-causes / syc-20355745. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] U.S. Provisional Patent Application No. 62 / 939,324 [Non-patent literature]
[0006] [Non-Patent Document 1] "Preeclampsia - Symptoms and Causes - Mayo Clinic" (April 3, 2018) [Non-patent document 2] Karumanchi and Granger,2016,Hypertension;67(2):238-242 Summary of the Invention [Problem to be solved by the invention]
[0007] Early detection and treatment allow most women to give birth to healthy infants if preeclampsia is detected early and treated with regular prenatal care. Although various protein biomarkers show altered levels in maternal serum during the presymptomatic stage, these biomarkers lack discriminatory and predictive power in individual patients (Karumanchi and Granger, 2016, Hypertension; 67(2):238-242). Therefore, identifying biomarkers for early detection of preeclampsia is important for early diagnosis and treatment of preeclampsia. [Means for solving the problem]
[0008] The present invention includes a method for detecting pre-eclampsia and / or determining an increased risk of pre-eclampsia in a pregnant woman, the method comprising: identifying a plurality of circulating RNA (C-RNA) molecules in a biosample obtained from the pregnant woman; Multiple C-RNA molecules (a) Any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 ≥ 23, ≥ 24, ≥ 25, ≥ 75, ARRDC2, JUN, SKIL, ATP13A3, PDE8B, GSTA3, PAPPA2, TIPARP, LEP, RGP1, USP54, CLEC4C, MRPS35, ARHGEF25, CUX2, HEATR9, FSTL3, DDI2, ZMYM6, ST6GALNAC3, GBP2, NES, ETV3, ADAM1 7, ATOH8, SLC4A3, TRAF3IP1, TTC21A, HEG1, ASTE1, TMEM108, ENC1, SCAMP1, ARRDC3, SLC26A2, SLIT3, CLIC5, TNFR SF21, PPP1R17, TPST1, GATSL2, SPDYE5, HIPK2, MTRNR2L6, CLCN1, GINS4, CRH, C10orf2, TRUB1, PRG2, ACY3, FAR2, a plurality of C-RNA molecules encoding at least a portion of a protein selected from CD63, CKAP4, TPCN1, RNF6, THTPA, FOS, PARN, ORAI3, ELMO3, SMPD3, SERPINF1, TMEM11, PSMD11, EBI3, CLEC4M, CCDC151, CPAMD8, CNFN, LILRA4, ADA, C22orf39, PI4KAP1, and ARFGAP3; or (b) Any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12 or more, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 23 or more, any 24 or more, any 25 or more, any 2 a plurality of C-RNA molecules encoding at least a portion of a protein selected from six or more, or all twenty-seven, of the following proteins: TIMP4, FLG, HTRA4, AMPH, LCN6, CRH, TEAD4, ARMS2, PAPPA2, SEMA3G, ADAMTS1, ALOX15B, SLC9A3R2, TIMP3, IGFBP5, HSPA12B, CLEC4C, KRT5, PRG2, PRX, ARHGEF25, ADAMTS2, DAAM2, FAM107A, LEP, NES, and VSIG4; or (c) Any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 23 or more, any 24 or more, any 25 or more, up to a maximum of 122 CYP26B1, IRF6, MYH14, and PODX L, PPP1R3C, SH3RF2, TMC7, ZNF366, ADCY1, C6, FAM219A, HAO2, IGIP, IL1R2, NTRK2, SH3PXD2A, SSUH2, SULT2A1, FMO3, FSTL3, GATA5, HTRA1, C8B, H19, MN1 , NFE2L1, PRDM16, AP3B2, EMP1, FLNC, STAG3, CPB2, TENC1, RP1L1, A1CF, NPR1, TEK, ERRFI1, ARHGEF15, CD34, RSPO3, ALPK3, SAMD4A, ZCCHC24, LEAP2, MYL 2, NRG3, ZBTB16, SERPINA3, AQP7, SRPX, UACA, ANO1, FKBP5, SCN5A, PTPN21, CACNA1C, ERG, SOX17, WWTR1, AIF1L, CA3, HRG, TAT, AQP7P1, ADRA2C, SYNPO, F N1, GPR116, KRT17, AZGP1, BCL6B, KIF1C, CLIC5, GPR4, GJA5, OLAH, C14orf37, ZEB1, JAG2, KIF26A, APOLD1, PNMT, MYOM3, PITPNM3, TIMP4, HTRA4, AMPH, L a plurality of C-RNA molecules encoding at least a portion of a protein selected from CN6, CRH, TEAD4, ARMS2, PAPPA2, SEMA3G, ADAMTS1, ALOX15B, SLC9A3R2, TIMP3, IGFBP5, HSPA12B, PRG2, PRX, ARHGEF25, ADAMTS2, DAAM2, FAM107A, LEP, NES, VSIG4, HBG2, CADM2, LAMP5, PTGDR2, NOMO1, NXF3, PLD4, BPIFB3, PACSIN1, CUX2, FLG, CLEC4C, and KRT5; or (d) any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12 or more, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 23 or more, any 24 or more, any 25 or more, any 26 or more, any 27 or more, any 28 or more, a plurality of C-RNA molecules encoding at least a portion of a protein selected from any 29 or more, or all 30, of VSIG4, ADAMTS2, NES, FAM107A, LEP, DAAM2, ARHGEF25, TIMP3, PRX, ALOX15B, HSPA12B, IGFBP5, CLEC4C, SLC9A3R2, ADAMTS1, SEMA3G, KRT5, AMPH, PRG2, PAPPA2, TEAD4, CRH, PITPNM3, TIMP4, PNMT, ZEB1, APOLD1, PLD4, CUX2, and HTRA4; or (e) any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12 or more, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 23 or more, any 24 or more, any 25 or more, or a plurality of C-RNA molecules encoding at least a portion of a protein selected from all 26 of ADAMTS1, ADAMTS2, ALOX15B, AMPH, ARHGEF25, CELF4, DAAM2, FAM107A, HSPA12B, HTRA4, IGFBP5, KCNA5, KRT5, LCN6, LEP, LRRC26, NES, OLAH, PACSIN1, PAPPA2, PRX, PTGDR2, SEMA3G, SLC9A3R2, TIMP3, and VSIG4; or (f) a plurality of C-RNA molecules encoding at least a portion of proteins selected from any one or more, any two or more, any three or more, any four or more, any five or more, any six or more, any seven or more, any eight or more, any nine or more, any ten or more, any eleven or more, any twelve or more, any thirteen or more, any fourteen or more, any fifteen or more, any sixteen or more, any seventeen or more, any eighteen or more, any nineteen or more, any twenty or more, any twenty one or more, or all twenty-two of ADAMTS1, ADAMTS2, ALOX15B, ARHGEF25, CELF4, DAAM2, FAM107A, HTRA4, IGFBP5, KCNA5, KRT5, LCN6, LEP, LRRC26, NES, OLAH, PRX, PTGDR2, SEMA3G, SLC9A3R2, TIMP3, and VSIG4; or (g) any one or more, any two or more, any three or more, any four or more, any five or more, any six or more, any seven or more, any eight or more, any nine or more, any ten or more, or all eleven of CLEC4C, ARHGEF25, ADAMTS2, LEP, ARRDC2, SKIL, PAPPA2, VSIG4, ARRDC4, CRH, and NES (in some embodiments, seven of ADAMTS2, ARHGEF25, ARRDC2, CLEC4C, LEP, PAPPA2, and VSIG4, and eight of ADAMTS2, ARHGEF25, ARRDC2, CLEC4C, LEP, PAPPA2, SKIL, and NES) and VSIG4; eight ADAMTS2, ARHGEF25, ARRDC4, CLEC4C, LEP, NES, SKIL, and VSIG4; ten ADAMTS2, ARHGEF25, ARRDC2, ARRDC4, CLEC4C, CRH, LEP, PAPPA2, SKIL, and VSIG4; six ADAMTS2, ARHGEF25, ARRDC2, CLEC4C, LEP, and SKIL; or eight ADAMTS2, ARHGEF25, ARRDC2, ARRDC4, CLEC4C, LEP, PAPPA2, and SKIL; or (h) Any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12 or more, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 23 a plurality of C-RNA molecules encoding at least a portion of a protein selected from one or all twenty-four of LEP, PAPPA2, KCNA5, ADAMTS2, MYOM3, ATP13A3, ARHGEF25, ADA, HTRA4, NES, CRH, ACY3, PLD4, SCT, NOX4, PACSIN1, SERPINF1, SKIL, SEMA3G, TIPARP, LRRC26, PHEX, LILRA4, and PER1; or (i) any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 23 or more, any 24 or more, any 25 or more, any 26 or more, any 27 or more, any 28 or more, a plurality of C-RNA molecules encoding at least a portion of a protein selected from any 29 or more, any 30 or more, any 31 or more, any 32 or more, any 33 or more, any 34 or more, any 35 or more, any 36 or more, any 37 or more, any 38 or more, any 39 or more, any 40 or more, any 41 or more, any 42 or more, any 43 or more, any 44 or more, any 45 or more, any 46 or more, any 47 or more, any 48 or more, or all 49 of those listed in Table S9 of Example 7; or (j) a plurality of C-RNA molecules encoding at least a portion of a protein selected from any one or more, any two or more, any three or more, any four or more, any five or more, any six or more, any seven or more, any eight or more, any nine or more, any ten or more, any eleven or more, any twelve or more, or all thirteen of AKAP2, ARRB1, CPSF7, INO80C, JAG1, MSMP, NR4A2, PLEK, RAP1GAP2, SPEG, TRPS1, UBE2Q1, and ZNF768; It indicates pre-eclampsia and / or an increased risk of pre-eclampsia in pregnant women.
[0009] The present invention includes a method for detecting pre-eclampsia and / or determining an increased risk of pre-eclampsia in a pregnant woman, the method comprising: Obtaining a biosample from a pregnant woman; purifying a population of circulating RNA (C-RNA) molecules from the biosample; identifying protein-coding sequences encoded by C-RNA molecules within the population of purified C-RNA molecules; The protein coding sequence encoded by the C-RNA molecule encoding at least a portion of a protein is (a) any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any Any 22 or more, any 23 or more, any 24 or more, any 25 or more, any 50 or more, any 70 or more, or up to 75 of the following: ARRDC2, JUN, SKIL, ATP13A3, PDE8B, GSTA3, PAPPA2, TIPARP, LEP, RGP1, USP54, CLEC4C, MRPS35, ARHGEF25, CUX2, HEATR9, FSTL3, DDI2, ZMYM6, S T6GALNAC3, GBP2, NES, ETV3, ADAM17, ATOH8, SLC4A3, TRAF3IP1, TTC21A, HEG1, ASTE1, TMEM108, ENC1, SCAMP1, ARRDC3, SLC26A2, SLIT3, CLIC5, TNFRSF21, PPP1R17, TPST1, GATSL2, SPDYE5, HIPK2, MTRNR2L6, CLCN1, GINS4, CRH, C10orf2, TRUB1, PRG2, ACY3, FAR2, CD63, CKAP4, TPCN1, RNF6, THTPA, FOS, PARN, ORAI3, ELMO3, SMPD3, SER PINF1, TMEM11, PSMD11, EBI3, CLEC4M, CCDC151, CPAMD8, CNFN, LILRA4, ADA, C22orf39, PI4KAP1, and ARFGAP3, or (b) any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12 or more, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 23 or more, any 2 four or more, any 25 or more, any 26 or more, or all 27 of the following: TIMP4, FLG, HTRA4, AMPH, LCN6, CRH, TEAD4, ARMS2, PAPPA2, SEMA3G, ADAMTS1, ALOX15B, SLC9A3R2, TIMP3, IGFBP5, HSPA12B, CLEC4C, KRT5, PRG2, PRX, ARHGEF25, ADAMTS2, DAAM2, FAM107A, LEP, NES, and VSIG4; or (c) any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 23 or more, any 24 or more, any 25 or more, any 50 or more, any 75 or more, any 100 or more, or All 122 CYP26B1, IRF6, MYH14, PODXL, PPP1R3C, SH3RF2, TMC7, ZNF366, ADCY1, C6, FAM219A, HAO2, IGIP, IL1R2, NTRK2, SH3PXD2A, SSUH2, SULT2A1, FM O3, FSTL3, GATA5, HTRA1, C8B, H19, MN1, NFE2L1, PRDM16, AP3B2, EMP1, FLNC, STAG3, CPB2, TENC1, RP1L1, A1CF, NPR1, TEK, ERRFI1, ARHGEF15, CD34, RSP O3, ALPK3, SAMD4A, ZCCHC24, LEAP2, MYL2, NRG3, ZBTB16, SERPINA3, AQP7, SRPX, UACA, ANO1, FKBP5, SCN5A, PTPN21, CACNA1C, ERG, SOX17, WWTR1, AIF1L , CA3, HRG, TAT, AQP7P1, ADRA2C, SYNPO, FN1, GPR116, KRT17, AZGP1, BCL6B, KIF1C, CLIC5, GPR4, GJA5, OLAH, C14orf37, ZEB1, JAG2, KIF26A, APOLD1, PN MT, MYOM3, PITPNM3, TIMP4, HTRA4, AMPH, LCN6, CRH, TEAD4, ARMS2, PAPPA2, SEMA3G, ADAMTS1, ALOX15B, SLC9A3R2, TIMP3, IGFBP5, HSPA12B, PRG2, PRX, ARHGEF25, ADAMTS2, DAAM2, FAM107A, LEP, NES, VSIG4, HBG2, CADM2, LAMP5, PTGDR2, NOMO1, NXF3, PLD4, BPIFB3, PACSIN1, CUX2, FLG, CLEC4C, and KRT5, or (d) any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12 or more, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 23 or more, any 24 or more, any 25 or more, any 26 or more, any 27 or more, any 28 or more, any 29 or more, or all 30 of VSIG4, ADAMTS2, NES, FAM107A, LEP, DAAM2, ARHGEF25, TIMP3, PRX, ALOX15B, HSPA12B, IGFBP5, CLEC4C, SLC9A3R2, ADAMTS1, SEMA3G, KRT5, AMPH, PRG2, PAPPA2, TEAD4, CRH, PITPNM3, TIMP4, PNMT, ZEB1, APOLD1, PLD4, CUX2, and HTRA4; or (e) any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12 or more, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 23 or more, any 24 or more, any 25 or more, or all 26 of ADAMTS1, ADAMTS2, ALOX15B, AMPH, ARHGEF25, CELF4, DAAM2, FAM107A, HSPA12B, HTRA4, IGFBP5, KCNA5, KRT5, LCN6, LEP, LRRC26, NES, OLAH, PACSIN1, PAPPA2, PRX, PTGDR2, SEMA3G, SLC9A3R2, TIMP3, and VSIG4; or (f) any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12 or more, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, or all 22 of ADAMTS1, ADAMTS2, ALOX15B, ARHGEF25, CELF4, DAAM2, FAM107A, HTRA4, IGFBP5, KCNA5, KRT5, LCN6, LEP, LRRC26, NES, OLAH, PRX, PTGDR2, SEMA3G, SLC9A3R2, TIMP3, and VSIG4; or (g) any one or more, any two or more, any three or more, any four or more, any five or more, any six or more, any seven or more, any eight or more, any nine or more, any ten or more, or all eleven of CLEC4C, ARHGEF25, ADAMTS2, LEP, ARRDC2, SKIL, PAPPA2, VSIG4, ARRDC4, CRH, and NES (in some embodiments, seven of ADAMTS2, ARHGEF25, ARRDC2, CLEC4C, LEP, PAPPA2, and VSIG4, and eight of ADAMTS2, ARHGEF25, ARRDC2, CLEC4C, , LEP, PAPPA2, SKIL, and VSIG4, 8 ADAMTS2, ARHGEF25, ARRDC4, CLEC4C, LEP, NES, SKIL, and VSIG4, 10 ADAMTS2, ARHGEF25, ARRDC2, ARRDC4, CLEC4C, CRH, LEP, PAPPA2, SKIL, and VSIG4, 6 ADAMTS2, ARHGEF25, ARRDC2, CLEC4C, LEP, and SKIL, or 8 ADAMTS2, ARHGEF25, ARRDC2, ARRDC4, CLEC4C, LEP, PAPPA2, and SKIL). (h) any one or more, any two or more, any three or more, any four or more, any five or more, any six or more, any seven or more, any eight or more, any nine or more, any ten or more, any eleven or more, any twelve or more, any thirteen or more, any fourteen or more, any fifteen or more, any sixteen or more, any seventeen or more, any eighteen or more, any nineteen or more, any twenty or more, any twenty-one or more, any twenty-two or more, any twenty-three or more, or all twenty-four of LEP, PAPPA2, KCNA5, ADAMTS2, MYOM3, ATP13A3, ARHGEF25, ADA, HTRA4, NES, CRH, ACY3, PLD4, SCT, NOX4, PACSIN1, SERPINF1, SKIL, SEMA3G, TIPARP, LRRC26, PHEX, LILRA4, and PER1; or (i) any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 23 or more, any 24 or more, any 25 or more, any 26 or more, any 27 or more, any 28 or more, any 29 or more, any 30 or more, any 31 or more, any 32 or more, any 33 or more, any 34 or more, any 35 or more, any 36 or more, any 37 or more, any 38 or more, any 39 or more, any 40 or more, any 41 or more, any 42 or more, any 43 or more, any 44 or more, any 45 or more, any 46 or more, any 47 or more, any 48 or more, or all 49 of those listed in Table S9 of Example 7; or (j) selected from any one or more, any two or more, any three or more, any four or more, any five or more, any six or more, any seven or more, any eight or more, any nine or more, any ten or more, any eleven or more, any twelve or more, or all thirteen of AKAP2, ARRB1, CPSF7, INO80C, JAG1, MSMP, NR4A2, PLEK, RAP1GAP2, SPEG, TRPS1, UBE2Q1, and ZNF768; It indicates pre-eclampsia and / or an increased risk of pre-eclampsia in pregnant women.
[0010] In some embodiments, identifying protein-coding sequences encoded by C-RNA molecules in the biosample comprises hybridization, reverse transcriptase PCR, microarray chip analysis, or sequencing.
[0011] In some embodiments, identifying protein-coding sequences encoded by C-RNA molecules within a biosample comprises sequencing, including, for example, massively parallel sequencing of clonally amplified molecules and / or RNA sequencing.
[0012] In some embodiments, the methods further include removing intact cells from the biosample, treating the biosample with deoxynuclease (DNase) to remove cell-free DNA (cfDNA), synthesizing complementary DNA (cDNA) from C-RNA molecules in the biosample, and / or enriching cDNA sequences for protein-encoding DNA sequences by exome enrichment before identifying protein-coding sequences encoded by the circular RNA (C-RNA) molecules.
[0013] The present invention includes a method for detecting pre-eclampsia and / or determining an increased risk of pre-eclampsia in a pregnant woman, the method comprising: Obtaining biological samples from pregnant women; removing intact cells from the biosample; Treating the biosample with deoxynuclease (DNase) to remove cell-free DNA (cfDNA); synthesizing complementary DNA (cDNA) from RNA molecules in the biosample; Enrichment of cDNA sequences for protein-coding DNA sequences (exome enrichment); Sequencing the resulting enriched cDNA sequences; and identifying protein-coding sequences encoded by the enriched C-RNA molecules; (a) Any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 23 or more, any 24 or more, any 25 or more, or up to all 75 of the following: ARRDC2, JUN, SKIL, ATP13A3, PDE8B, GSTA3, PAPPA2, TIPARP, LEP, RGP1, USP54, CLEC4C, MRPS35, ARHGEF25, CUX2, HEATR9, FSTL3, DDI2, ZMYM6, ST6GALNAC3 , GBP2, NES, ETV3, ADAM17, ATOH8, SLC4A3, TRAF3IP1, TTC21A, HEG1, ASTE1, TMEM108, ENC1, SCAMP1, ARRDC3, SLC26A2, SLIT3, CLIC5, TNFRSF21, PPP1R17, TPST1, GATSL2, SPDYE5, HIPK2, MTRNR2L6, CLCN1, GINS4, CRH, C 10orf2, TRUB1, PRG2, ACY3, FAR2, CD63, CKAP4, TPCN1, RNF6, THTPA, FOS, PARN, ORAI3, ELMO3, SMPD3, SERPIN F1, TMEM11, PSMD11, EBI3, CLEC4M, CCDC151, CPAMD8, CNFN, LILRA4, ADA, C22orf39, PI4KAP1, and ARFGAP3, or (b) any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 23 or more, any 24 any 25 or more, any 26 or more, or all 27 of TIMP4, FLG, HTRA4, AMPH, LCN6, CRH, TEAD4, ARMS2, PAPPA2, SEMA3G, ADAMTS1, ALOX15B, SLC9A3R2, TIMP3, IGFBP5, HSPA12B, CLEC4C, KRT5, PRG2, PRX, ARHGEF25, ADAMTS2, DAAM2, FAM107A, LEP, NES, and VSIG4; or (c) any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12 or more, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 23 or more, any 24 or more, any 25 or more, or all 122 CYP26B1 or IRF 6, MYH14, PODXL, PPP1R3C, SH3RF2, TMC7, ZNF366, ADCY1, C6, FAM219A, HAO2, IGIP, IL1R2, NTRK2, SH3PXD2A, SSUH2, SULT2A1, FMO3, FSTL3, GATA5, H TRA1, C8B, H19, MN1, NFE2L1, PRDM16, AP3B2, EMP1, FLNC, STAG3, CPB2, TENC1, RP1L1, A1CF, NPR1, TEK, ERRFI1, ARHGEF15, CD34, RSPO3, ALPK3, SAMD 4A, ZCCHC24, LEAP2, MYL2, NRG3, ZBTB16, SERPINA3, AQP7, SRPX, UACA, ANO1, FKBP5, SCN5A, PTPN21, CACNA1C, ERG, SOX17, WWTR1, AIF1L, CA3, HRG, T AT, AQP7P1, ADRA2C, SYNPO, FN1, GPR116, KRT17, AZGP1, BCL6B, KIF1C, CLIC5, GPR4, GJA5, OLAH, C14orf37, ZEB1, JAG2, KIF26A, APOLD1, PNMT, MYOM 3, PITPNM3, TIMP4, HTRA4, AMPH, LCN6, CRH, TEAD4, ARMS2, PAPPA2, SEMA3G, ADAMTS1, ALOX15B, SLC9A3R2, TIMP3, IGFBP5, HSPA12B, PRG2, PRX, ARHG EF25, ADAMTS2, DAAM2, FAM107A, LEP, NES, VSIG4, HBG2, CADM2, LAMP5, PTGDR2, NOMO1, NXF3, PLD4, BPIFB3, PACSIN1, CUX2, FLG, CLEC4C, and KRT5, or (d) any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12 or more, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 23 or more, any 24 or more, any 25 or more, any 26 or more, any 27 or more, any 28 or more, any 29 or more, or all 30 of VSIG4, ADAMTS2, NES, FAM107A, LEP, DAAM2, ARHGEF25, TIMP3, PRX, ALOX15B, HSPA12B, IGFBP5, CLEC4C, SLC9A3R2, ADAMTS1, SEMA3G, KRT5, AMPH, PRG2, PAPPA2, TEAD4, CRH, PITPNM3, TIMP4, PNMT, ZEB1, APOLD1, PLD4, CUX2, and HTRA4; or (e) any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12 or more, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 23 or more, any 24 or more, any 25 or more, or all 26 of ADAMTS1, ADAMTS2, ALOX15B, AMPH, ARHGEF25, CELF4, DAAM2, FAM107A, HSPA12B, HTRA4, IGFBP5, KCNA5, KRT5, LCN6, LEP, LRRC26, NES, OLAH, PACSIN1, PAPPA2, PRX, PTGDR2, SEMA3G, SLC9A3R2, TIMP3, and VSIG4; or (f) any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12 or more, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, or all 22 of ADAMTS1, ADAMTS2, ALOX15B, ARHGEF25, CELF4, DAAM2, FAM107A, HTRA4, IGFBP5, KCNA5, KRT5, LCN6, LEP, LRRC26, NES, OLAH, PRX, PTGDR2, SEMA3G, SLC9A3R2, TIMP3, and VSIG4; or (g) any one or more, any two or more, any three or more, any four or more, any five or more, any six or more, any seven or more, any eight or more, any nine or more, any ten or more, or all eleven of CLEC4C, ARHGEF25, ADAMTS2, LEP, ARRDC2, SKIL, PAPPA2, VSIG4, ARRDC4, CRH, and NES (in some embodiments, seven of ADAMTS2, ARHGEF25, ARRDC2, CLEC4C, LEP, PAPPA2, and VSIG4, and eight of ADAMTS2, ARHGEF25, ARRDC2, CLEC4C, , LEP, PAPPA2, SKIL, and VSIG4, 8 ADAMTS2, ARHGEF25, ARRDC4, CLEC4C, LEP, NES, SKIL, and VSIG4, 10 ADAMTS2, ARHGEF25, ARRDC2, ARRDC4, CLEC4C, CRH, LEP, PAPPA2, SKIL, and VSIG4, 6 ADAMTS2, ARHGEF25, ARRDC2, CLEC4C, LEP, and SKIL, or 8 ADAMTS2, ARHGEF25, ARRDC2, ARRDC4, CLEC4C, LEP, PAPPA2, and SKIL). (h) any one or more, any two or more, any three or more, any four or more, any five or more, any six or more, any seven or more, any eight or more, any nine or more, any ten or more, any eleven or more, any twelve or more, any thirteen or more, any fourteen or more, any fifteen or more, any sixteen or more, any seventeen or more, any eighteen or more, any nineteen or more, any twenty or more, any twenty-one or more, any twenty-two or more, any twenty-three or more, or all twenty-four of LEP, PAPPA2, KCNA5, ADAMTS2, MYOM3, ATP13A3, ARHGEF25, ADA, HTRA4, NES, CRH, ACY3, PLD4, SCT, NOX4, PACSIN1, SERPINF1, SKIL, SEMA3G, TIPARP, LRRC26, PHEX, LILRA4, and PER1; or (i) any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 23 or more, any 24 or more, any 25 or more, any 26 or more, any 27 or more, any 28 or more, any 29 or more, any 30 or more, any 31 or more, any 32 or more, any 33 or more, any 34 or more, any 35 or more, any 36 or more, any 37 or more, any 38 or more, any 39 or more, any 40 or more, any 41 or more, any 42 or more, any 43 or more, any 44 or more, any 45 or more, any 46 or more, any 47 or more, any 48 or more, or all 49 of those listed in Table S9 of Example 7; or (j) a protein coding sequence encoded by the C-RNA molecule encoding at least a portion of a protein selected from any one or more, any two or more, any three or more, any four or more, any five or more, any six or more, any seven or more, any eight or more, any nine or more, any ten or more, any eleven or more, any twelve or more, or all thirteen of AKAP2, ARRB1, CPSF7, INO80C, JAG1, MSMP, NR4A2, PLEK, RAP1GAP2, SPEG, TRPS1, UBE2Q1, and ZNF768 is It indicates pre-eclampsia and / or an increased risk of pre-eclampsia in pregnant women.
[0014] The present invention includes a method for identifying circulating RNA signatures associated with an increased risk of pre-eclampsia, the method comprising obtaining a biological sample from a pregnant woman, removing intact cells from the biological sample, treating the biological sample with deoxynuclease (DNase) to remove cell-free DNA (cfDNA), synthesizing complementary DNA (cDNA) from RNA molecules in the biological sample, enriching the cDNA sequences for protein-coding DNA sequences (exome enrichment), sequencing the resulting enriched cDNA sequences, and identifying protein-coding sequences encoded by the enriched C-RNA molecules.
[0015] The present invention provides Obtaining biological samples from pregnant women; removing intact cells from the biosample; Treating the biosample with deoxynuclease (DNase) to remove cell-free DNA (cfDNA); synthesizing complementary DNA (cDNA) from RNA molecules in the biosample; Enrichment of cDNA sequences for protein-coding DNA sequences (exome enrichment); Sequencing the resulting enriched cDNA sequences; and identifying a protein-coding sequence encoded by the enriched C-RNA molecules, The protein coding sequence is (a) Any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 23 or more, any 24 or more, any 25, up to a maximum of 75 of ARRDC2, JUN, SKIL, ATP13A3, PDE8B, GSTA3, PAPPA2, TIPARP, LEP, RGP1, USP54, CLEC4C, MRPS35, ARHGEF25, CUX2, HEATR9, FSTL3, DDI2, ZMYM6, ST6GALNAC3, GB P2, NES, ETV3, ADAM17, ATOH8, SLC4A3, TRAF3IP1, TTC21A, HEG1, ASTE1, TMEM108, ENC1, SCAMP1, ARRDC3, SL C26A2, SLIT3, CLIC5, TNFRSF21, PPP1R17, TPST1, GATSL2, SPDYE5, HIPK2, MTRNR2L6, CLCN1, GINS4, CRH, C1 0orf2, TRUB1, PRG2, ACY3, FAR2, CD63, CKAP4, TPCN1, RNF6, THTPA, FOS, PARN, ORAI3, ELMO3, SMPD3, SERPINF1, TMEM11, PSMD11, EBI3, CLEC4M, CCDC151, CPAMD8, CNFN, LILRA4, ADA, C22orf39, PI4KAP1, and ARFGAP3, or (b) any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12 or more, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 23 or more, any 2 four or more, any 25 or more, any 26 or more, or all 27 of the following: TIMP4, FLG, HTRA4, AMPH, LCN6, CRH, TEAD4, ARMS2, PAPPA2, SEMA3G, ADAMTS1, ALOX15B, SLC9A3R2, TIMP3, IGFBP5, HSPA12B, CLEC4C, KRT5, PRG2, PRX, ARHGEF25, ADAMTS2, DAAM2, FAM107A, LEP, NES, and VSIG4; or (c) any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 23 or more, any 24 or more, any 25 or more, or all 122 CYP26B1, IRF6, MYH14, PODXL, PPP1R3C, SH3RF2, TMC7, ZNF366, ADCY1, C6, FAM219A, HAO2, IGIP, IL1R2, NTRK2, SH3PXD2A, SSUH2, SULT2A1, FMO3, FSTL3, GATA5, HT RA1, C8B, H19, MN1, NFE2L1, PRDM16, AP3B2, EMP1, FLNC, STAG3, CPB2, TENC1, RP1L1, A1CF, NPR1, TEK, ERRFI1, ARHGEF15, CD34, RSPO3, ALPK3, SAMD4 A, ZCCHC24, LEAP2, MYL2, NRG3, ZBTB16, SERPINA3, AQP7, SRPX, UACA, ANO1, FKBP5, SCN5A, PTPN21, CACNA1C, ERG, SOX17, WWTR1, AIF1L, CA3, HRG, T AT, AQP7P1, ADRA2C, SYNPO, FN1, GPR116, KRT17, AZGP1, BCL6B, KIF1C, CLIC5, GPR4, GJA5, OLAH, C14orf37, ZEB1, JAG2, KIF26A, APOLD1, PNMT, MYOM 3, PITPNM3, TIMP4, HTRA4, AMPH, LCN6, CRH, TEAD4, ARMS2, PAPPA2, SEMA3G, ADAMTS1, ALOX15B, SLC9A3R2, TIMP3, IGFBP5, HSPA12B, PRG2, PRX, ARHG EF25, ADAMTS2, DAAM2, FAM107A, LEP, NES, VSIG4, HBG2, CADM2, LAMP5, PTGDR2, NOMO1, NXF3, PLD4, BPIFB3, PACSIN1, CUX2, FLG, CLEC4C, and KRT5, or (d) any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12 or more, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 23 or more, any 24 or more, any 25 or more, any 26 or more, any 27 or more, any 28 or more, any 29 or more, or all 30 of VSIG4, ADAMTS2, NES, FAM107A, LEP, DAAM2, ARHGEF25, TIMP3, PRX, ALOX15B, HSPA12B, IGFBP5, CLEC4C, SLC9A3R2, ADAMTS1, SEMA3G, KRT5, AMPH, PRG2, PAPPA2, TEAD4, CRH, PITPNM3, TIMP4, PNMT, ZEB1, APOLD1, PLD4, CUX2, and HTRA4; or (e) any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12 or more, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 23 or more, any 24 or more, any 25 or more, or all 26 of ADAMTS1, ADAMTS2, ALOX15B, AMPH, ARHGEF25, CELF4, DAAM2, FAM107A, HSPA12B, HTRA4, IGFBP5, KCNA5, KRT5, LCN6, LEP, LRRC26, NES, OLAH, PACSIN1, PAPPA2, PRX, PTGDR2, SEMA3G, SLC9A3R2, TIMP3, and VSIG4; or (f) any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12 or more, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, or all 22 of ADAMTS1, ADAMTS2, ALOX15B, ARHGEF25, CELF4, DAAM2, FAM107A, HTRA4, IGFBP5, KCNA5, KRT5, LCN6, LEP, LRRC26, NES, OLAH, PRX, PTGDR2, SEMA3G, SLC9A3R2, TIMP3, and VSIG4; or (g) any one or more, any two or more, any three or more, any four or more, any five or more, any six or more, any seven or more, any eight or more, any nine or more, any ten or more, or all eleven of CLEC4C, ARHGEF25, ADAMTS2, LEP, ARRDC2, SKIL, PAPPA2, VSIG4, ARRDC4, CRH, and NES (in some embodiments, seven of ADAMTS2, ARHGEF25, ARRDC2, CLEC4C, LEP, PAPPA2, and VSIG4, and eight of ADAMTS2, ARHGEF25, ARRDC2, CLEC4C, , LEP, PAPPA2, SKIL, and VSIG4, 8 ADAMTS2, ARHGEF25, ARRDC4, CLEC4C, LEP, NES, SKIL, and VSIG4, 10 ADAMTS2, ARHGEF25, ARRDC2, ARRDC4, CLEC4C, CRH, LEP, PAPPA2, SKIL, and VSIG4, 6 ADAMTS2, ARHGEF25, ARRDC2, CLEC4C, LEP, and SKIL, or 8 ADAMTS2, ARHGEF25, ARRDC2, ARRDC4, CLEC4C, LEP, PAPPA2, and SKIL). (h) any one or more, any two or more, any three or more, any four or more, any five or more, any six or more, any seven or more, any eight or more, any nine or more, any ten or more, any eleven or more, any twelve or more, any thirteen or more, any fourteen or more, any fifteen or more, any sixteen or more, any seventeen or more, any eighteen or more, any nineteen or more, any twenty or more, any twenty-one or more, any twenty-two or more, any twenty-three or more, or all twenty-four of LEP, PAPPA2, KCNA5, ADAMTS2, MYOM3, ATP13A3, ARHGEF25, ADA, HTRA4, NES, CRH, ACY3, PLD4, SCT, NOX4, PACSIN1, SERPINF1, SKIL, SEMA3G, TIPARP, LRRC26, PHEX, LILRA4, and PER1; or (i) any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 23 or more, any 24 or more, any 25 or more, any 26 or more, any 27 or more, any 28 or more, any 29 or more, any 30 or more, any 31 or more, any 32 or more, any 33 or more, any 34 or more, any 35 or more, any 36 or more, any 37 or more, any 38 or more, any 39 or more, any 40 or more, any 41 or more, any 42 or more, any 43 or more, any 44 or more, any 45 or more, any 46 or more, any 47 or more, any 48 or more, or all 49 of those listed in Table S9 of Example 7; or (j) Comprises at least a portion of any one or more, any two or more, any three or more, any four or more, any five or more, any six or more, any seven or more, any eight or more, any nine or more, any ten or more, any eleven or more, any twelve or more, or all thirteen of the following proteins: AKAP2, ARRB1, CPSF7, INO80C, JAG1, MSMP, NR4A2, PLEK, RAP1GAP2, SPEG, TRPS1, UBE2Q1, and ZNF768.
[0016] In some embodiments, the biosample comprises plasma.
[0017] In some embodiments, the biosample is obtained from a pregnant woman at less than 16 weeks gestational age or less than 20 weeks gestational age.
[0018] In some embodiments, the biosample is obtained from a pregnant woman at a gestational age of less than 20 weeks.
[0019] The present invention provides circulating RNA (C-RNA) signatures for high risk of pre-eclampsia, any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more Above, any 20 or more, any 21 or more, any 22 or more, any 23 or more, any 24 or more, any 25 or more, any 50 or more, any 70 or more, up to 75 of all ARRDC2, JUN, SKIL, ATP13A3, PDE8B, GSTA3, PAPPA2, TIPARP, LEP, RGP1, USP54, CLEC4C, MRPS35, ARHGEF25, CUX2, HEATR9, FSTL3, DDI2, ZMYM 6, ST6GALNAC3, GBP2, NES, ETV3, ADAM17, ATOH8, SLC4A3, TRAF3IP1, TTC21A, HEG1, ASTE1, TMEM108, ENC1, SCAMP1, ARRD C3, SLC26A2, SLIT3, CLIC5, TNFRSF21, PPP1R17, TPST1, GATSL2, SPDYE5, HIPK2, MTRNR2L6, CLCN1, GINS4, CRH, C10orf2, The C-RNA signatures include those encoding at least a portion of TRUB1, PRG2, ACY3, FAR2, CD63, CKAP4, TPCN1, RNF6, THTPA, FOS, PARN, ORAI3, ELMO3, SMPD3, SERPINF1, TMEM11, PSMD11, EBI3, CLEC4M, CCDC151, CPAMD8, CNFN, LILRA4, ADA, C22orf39, PI4KAP1, and ARFGAP3.
[0020] The present invention provides circulating RNA (C-RNA) signatures for high risk of pre-eclampsia, any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12 or more, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 23 or more , any 24 or more, any 25 or more, any 26 or more, or all 27 of TIMP4, FLG, HTRA4, AMPH, LCN6, CRH, TEAD4, ARMS2, PAPPA2, SEMA3G, ADAMTS1, ALOX15B, SLC9A3R2, TIMP3, IGFBP5, HSPA12B, CLEC4C, KRT5, PRG2, PRX, ARHGEF25, ADAMTS2, DAAM2, FAM107A, LEP, NES, and C-RNA signatures encoding at least a portion of VSIG4.
[0021] The present invention provides a circulating RNA (C-RNA) signature for high risk of preeclampsia, including multiple CYP26B1, IRF6, MYH14, PODXL, PPP1R3C, SH3RF2, TMC7, ZNF366, ADCY1, C6, FAM219A, HAO2, IGIP, IL1R2, NTRK2, SH3PXD2A, SSUH2, SULT2A1, FMO3, FSTL3, GATA5, HTRA1, C8B, H19, MN1, NFE2L1, PRDM1. 16, AP3B2, EMP1, FLNC, STAG3, CPB2, TENC1, RP1L1, A1CF, NPR1, TEK, ERRFI1, ARHGEF15, CD34, RSPO3, ALPK3, SAMD4A, ZCCH C24, LEAP2, MYL2, NRG3, ZBTB16, SERPINA3, AQP7, SRPX, UACA, ANO1, FKBP5, SCN5A, PTPN21, CACNA1C, ERG, SOX17, WWTR1, AI F1L, CA3, HRG, TAT, AQP7P1, ADRA2C, SYNPO, FN1, GPR116, KRT17, AZGP1, BCL6B, KIF1C, CLIC5, GPR4, GJA5, OLAH, C14orf37 , ZEB1, JAG2, KIF26A, APOLD1, PNMT, MYOM3, PITPNM3, TIMP4, HTRA4, AMPH, LCN6, CRH, TEAD4, ARMS2, PAPPA2, SEMA3G, ADAMT The C-RNA signatures include those encoding at least a portion of S1, ALOX15B, SLC9A3R2, TIMP3, IGFBP5, HSPA12B, PRG2, PRX, ARHGEF25, ADAMTS2, DAAM2, FAM107A, LEP, NES, VSIG4, HBG2, CADM2, LAMP5, PTGDR2, NOMO1, NXF3, PLD4, BPIFB3, PACSIN1, CUX2, FLG, CLEC4C, and KRT5.
[0022] The present invention provides circulating RNA (C-RNA) signatures for high risk of pre-eclampsia, any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12 or more, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 23 or more, any 24 or more, any 25 or more, any 26 or more any 27 or more, any 28 or more, any 29 or more, or all 30 C-RNA signatures encoding at least a portion of VSIG4, ADAMTS2, NES, FAM107A, LEP, DAAM2, ARHGEF25, TIMP3, PRX, ALOX15B, HSPA12B, IGFBP5, CLEC4C, SLC9A3R2, ADAMTS1, SEMA3G, KRT5, AMPH, PRG2, PAPPA2, TEAD4, CRH, PITPNM3, TIMP4, PNMT, ZEB1, APOLD1, PLD4, CUX2, and HTRA4.
[0023] The present invention provides circulating RNA (C-RNA) signatures for high risk of pre-eclampsia, any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12 or more, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 2 The C-RNA signatures include C-RNA signatures encoding at least a portion of three or more, any 24 or more, any 25 or more, or all 26 of ADAMTS1, ADAMTS2, ALOX15B, AMPH, ARHGEF25, CELF4, DAAM2, FAM107A, HSPA12B, HTRA4, IGFBP5, KCNA5, KRT5, LCN6, LEP, LRRC26, NES, OLAH, PACSIN1, PAPPA2, PRX, PTGDR2, SEMA3G, SLC9A3R2, TIMP3, and VSIG4.
[0024] The present invention provides circulating RNA (C-RNA) signatures for high risk of pre-eclampsia, any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12 or more, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more , any 20 or more, any 21 or more, or all 22 of ADAMTS1, ADAMTS2, ALOX15B, ARHGEF25, CELF4, DAAM2, FAM107A, HTRA4, IGFBP5, KCNA5, KRT5, LCN6, LEP, LRRC26, NES, OLAH, PRX, PTGDR2, SEMA3G, SLC9A3R2, TIMP3, and VSIG4.
[0025] The present invention provides circulating RNA (C-RNA) signatures for high risk of pre-eclampsia, including any one or more, any two or more, any three or more, any four or more, any five or more, any six or more, any seven or more, any eight or more, any nine or more, any ten or more, or all eleven of the following: CLEC4C, ARHGEF25, ADAMTS2, LEP, ARRDC2, SKIL, PAPPA2, VSIG4, ARRDC4, CRH, and in some embodiments, seven of ADAMTS2, ARHGEF25, ARRDC2, CLEC4C, LEP, PAPPA2, and VSIG4; and eight of ADAMTS2, ARHGEF25, ARRDC2, CLEC4C, LEP, PAPPA2, and VSIG4. and C-RNA signatures encoding at least a portion of NES, including 4C, LEP, PAPPA2, SKIL, and VSIG4, 8 ADAMTS2, ARHGEF25, ARRDC4, CLEC4C, LEP, NES, SKIL, and VSIG4, 10 ADAMTS2, ARHGEF25, ARRDC2, ARRDC4, CLEC4C, CRH, LEP, PAPPA2, SKIL, and VSIG4, 6 ADAMTS2, ARHGEF25, ARRDC2, CLEC4C, LEP, and SKIL, or 8 ADAMTS2, ARHGEF25, ARRDC2, ARRDC4, CLEC4C, LEP, PAPPA2, and SKIL.
[0026] The present invention provides circulating RNA (C-RNA) signatures for high risk of pre-eclampsia, any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12 or more, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any any 21 or more, any 22 or more, any 23 or more, or all 24 of the following C-RNA signatures: LEP, PAPPA2, KCNA5, ADAMTS2, MYOM3, ATP13A3, ARHGEF25, ADA, HTRA4, NES, CRH, ACY3, PLD4, SCT, NOX4, PACSIN1, SERPINF1, SKIL, SEMA3G, TIPARP, LRRC26, PHEX, LILRA4, and PER1.
[0027] The present invention provides circulating RNA (C-RNA) signatures for high risk of pre-eclampsia, any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 23 or more, any 24 or more, any 25 or more, any 26 or more or more, any 27 or more, any 28 or more, any 29 or more, any 30 or more, any 31 or more, any 32 or more, any 33 or more, any 34 or more, any 35 or more, any 36 or more, any 37 or more, any 38 or more, any 39 or more, any 40 or more, any 41 or more, any 42 or more, any 43 or more, any 44 or more, any 45 or more, any 46 or more, any 47 or more, any 48 or more, or all 49 of the C-RNA signatures encoding at least a portion of those listed in Table S9 of Example 7.
[0028] The present invention includes circulating RNA (C-RNA) signatures for high risk of pre-eclampsia, C-RNA signatures encoding at least a portion of any one or more, any two or more, any three or more, any four or more, any five or more, any six or more, any seven or more, any eight or more, any nine or more, any ten or more, any eleven or more, any twelve or more, or all thirteen of AKAP2, ARRB1, CPSF7, INO80C, JAG1, MSMP, NR4A2, PLEK, RAP1GAP2, SPEG, TRPS1, UBE2Q1, and ZNF768.
[0029] The present invention includes solid support arrays comprising a plurality of agents capable of binding and / or identifying the C-RNA signatures described herein.
[0030] The present invention includes kits comprising a plurality of probes capable of binding and / or identifying the C-RNA signatures described herein.
[0031] The present invention includes kits that include a plurality of primers for selectively amplifying the C-RNA signatures described herein.
[0032] As used herein, the term "nucleic acid" is intended to be consistent with its use in the art and includes naturally occurring nucleic acids or functional analogs thereof. Particularly useful functional analogs are capable of hybridizing to nucleic acids in a sequence-specific manner or can be used as templates for replicating specific nucleotide sequences. Naturally occurring nucleic acids generally have backbones containing phosphodiester bonds. Analog structures can have alternative backbone linkages, including any of a variety known in the art. Naturally occurring nucleic acids generally have deoxyribose sugars (e.g., found in deoxyribonucleic acid (DNA)) or ribose sugars (e.g., found in ribonucleic acid (RNA)). Nucleic acids can contain any of a variety of analogs of these sugar moieties known in the art. Nucleic acids can include natural or unnatural bases. In this regard, naturally occurring deoxyribonucleic acids can have one or more bases selected from the group consisting of adenine, thymine, cytosine, or guanine, and ribonucleic acids can have one or more bases selected from the group consisting of uracil, adenine, cytosine, or guanine. Useful unnatural bases that can be included in nucleic acids are known in the art. The terms "template" and "target," when used in reference to a nucleic acid, are intended as semantic identifiers of the nucleic acid in the context of the methods or compositions described herein and do not necessarily limit the structure or function of the nucleic acid beyond what is otherwise expressly indicated.
[0033] As used herein, "amplification," "amplifying," or "amplification reaction," and derivatives thereof, generally refer to any act or process in which at least a portion of a nucleic acid molecule is duplicated or copied onto at least one additional nucleic acid molecule. The additional nucleic acid molecule optionally comprises a sequence that is substantially identical to or substantially complementary to at least a portion of a target nucleic acid molecule. The target nucleic acid molecule may be single-stranded or double-stranded, and the additional nucleic acid molecules may independently be single-stranded or double-stranded. Amplification optionally involves linear or exponential replication of nucleic acid molecules. In some embodiments, such amplification can be performed using isothermal conditions; in other embodiments, such amplification can involve thermal cycling. In some embodiments, amplification is multiplex amplification, which involves simultaneous amplification of multiple target sequences in a single amplification reaction. In some embodiments, "amplification" includes amplifying at least a portion of DNA- and RNA-based nucleic acids, alone or in combination. The amplification reaction can include any amplification process known to those of skill in the art. In some embodiments, the amplification reaction involves polymerase chain reaction (PCR).
[0034] As used herein, "amplification conditions" and its derivatives generally refer to conditions suitable for amplifying one or more nucleic acid sequences. Such amplification can be linear or exponential. In some embodiments, amplification conditions can include isothermal conditions, or can include thermal cycling conditions, or a combination of isothermal and thermal cycling conditions. In some embodiments, conditions suitable for amplifying one or more nucleic acid sequences include polymerase chain reaction (PCR) conditions. Typically, amplification conditions refer to a reaction mixture sufficient to amplify a nucleic acid, such as one or more target sequences, or an amplification target sequence appended with one or more adapters, e.g., an adapted amplification target sequence. Generally, amplification conditions include a catalyst for amplification, or nucleic acid synthesis, e.g., a polymerase, primers having a degree of complementarity to the nucleic acid to be amplified, and nucleotides, such as deoxyribonucleotide triphosphates (dNTPs), to facilitate primer extension when hybridized to the nucleic acid. Amplification conditions may require hybridization or annealing of a primer to a nucleic acid, extension of the primer, and a denaturation step in which the extended primer is separated from the nucleic acid sequence undergoing amplification. Typically, but not necessarily, amplification conditions may include thermal cycling, although in some embodiments, amplification conditions include multiple cycles in which the annealing, extension, and separation steps are repeated. Typically, amplification conditions include Mg ++ or Mn ++ and may also include various modifiers of ionic strength.
[0035] As used herein, the term "polymerase chain reaction" (PCR) refers to the method of K.B. Mullis in U.S. Pat. Nos. 4,683,195 and 4,683,202, which describes a method for increasing the concentration of a polynucleotide segment of interest in a mixture of genomic DNA without cloning or purification. This process for amplifying a polynucleotide of interest involves introducing a large excess of two oligonucleotide primers into a DNA mixture containing the desired polynucleotide of interest, followed by a series of thermal cycling steps in the presence of a DNA polymerase. The two primers are complementary to each strand of the double-stranded polynucleotide of interest. The mixture is first denatured at a higher temperature, and then the primers are annealed to complementary sequences within the polynucleotide of the molecule of interest. After annealing, the primers are extended with a polymerase to form a new pair of complementary strands. The steps of denaturation, primer annealing, and polymerase extension can be repeated multiple times (called thermal cycling) to obtain a highly concentrated amplified segment of the desired polynucleotide of interest. The length of the amplified segment (amplicon) of the desired target polynucleotide is determined by the relative positions of the primers with respect to each other, and therefore this length is a controllable parameter. By repeating this process, the method is called a "polymerase chain reaction" (hereinafter "PCR"). Because the desired amplified segment of the target polynucleotide becomes the predominant nucleic acid sequence (in terms of concentration) in the mixture, it is said to be "PCR amplified." In a modification of the above method, the target nucleic acid molecule can be PCR amplified using multiple different primer pairs, and in some cases, more than one primer pair per target nucleic acid molecule of interest, thereby forming a multiplex PCR reaction.
[0036] As used herein, the term "primer" and its derivatives generally refer to any polynucleotide that can hybridize to a target sequence of interest. Typically, a primer serves as a substrate onto which nucleotides can be polymerized by a polymerase. In some embodiments, a primer can be incorporated into a synthesized nucleic acid strand, providing a site to which another primer can hybridize and prime synthesis of a new strand complementary to the synthesized nucleic acid molecule. A primer can contain any combination of nucleotides or their analogs. In some embodiments, a primer is a single-stranded oligonucleotide or polynucleotide. The terms "polynucleotide" and "oligonucleotide" are used interchangeably herein to refer to polymeric forms of nucleotides of any length and can include ribonucleotides, deoxyribonucleotides, their analogs, or mixtures thereof. It should be understood that these terms include, as equivalents, analogs of either DNA or RNA made from nucleotide analogs and are applicable to single-stranded (such as sense or antisense) and double-stranded polynucleotides. As used herein, the term also encompasses cDNA, which is complementary or copy DNA produced from an RNA template, for example, by the action of reverse transcriptase. The term refers only to the primary structure of the molecule. Thus, the term includes triple-, double-, and single-stranded deoxyribonucleic acid ("DNA"), as well as triple-, double-, and single-stranded ribonucleic acid ("RNA").
[0037] As used herein, the terms "library" and "sequencing library" refer to a collection or plurality of template molecules that share a common sequence at their 5' end and a common sequence at their 3' end. A collection of template molecules containing known common sequences at their 3' and 5' ends may also be referred to as a 3' and 5' modified library.
[0038] As used herein, the term "flow cell" refers to a chamber that includes a solid surface through which one or more fluidic reagents can flow. Examples of flow cells and related fluidic systems and detection platforms that can be easily used in the methods of the present disclosure are described, for example, in Bentley et al., Nature 456:53-59 (2008), WO 04 / 018497, U.S. Patent No. 7,057,026, WO 91 / 06678, U.S. Patent No. 07 / 123744, U.S. Patent No. 7,329,492, U.S. Patent No. 7,211,414, U.S. Patent No. 7,315,019, U.S. Patent No. 7,405,281, and U.S. Patent Application Publication No. 2008 / 0108082.
[0039] As used herein, the term "amplicon," when used with reference to a nucleic acid, refers to a product of copying a nucleic acid, which product has a nucleotide sequence that is the same as or complementary to at least a portion of the nucleotide sequence of the nucleic acid. An amplicon can be produced by any of a variety of amplification methods using a nucleic acid or its amplicon as a template, including, for example, PCR, rolling circle amplification (RCA), ligation extension, or ligation chain reaction. An amplicon can be a nucleic acid molecule having a single copy of a particular nucleotide sequence (e.g., a PCR product) or multiple copies of a nucleotide sequence (e.g., a concatemeric product of RCA). A first amplicon of a target nucleic acid is typically a complementary copy. Subsequent amplicons are copies made from the target nucleic acid or the first amplicon after the generation of the first amplicon. Subsequent amplicons can have a sequence that is substantially complementary to or substantially identical to the target nucleic acid.
[0040] As used herein, the term "array" refers to a collection of sites that can be distinguished from one another according to their relative positions. Different molecules at different sites of an array can be distinguished from one another according to the site's position within the array. Each site of an array can contain one or more molecules of a particular type. For example, a site can contain a single target nucleic acid molecule having a particular sequence, or a site can contain several nucleic acid molecules having the same sequence (and / or its complementary sequence). The sites of an array can be different features located on the same substrate. Exemplary features include, but are not limited to, wells in a substrate, beads (or other particles) in or on a substrate, protrusions from a substrate, ridges on a substrate, or channels within a substrate. The sites of an array can be separate substrates, each with a different molecule. The different molecules attached to the separate substrates can be identified according to the position of the substrate on a surface to which the substrates are associated, or according to the position of the substrate within a liquid or gel. An exemplary array in which separate substrates are located on a surface includes, but is not limited to, beads in wells.
[0041] As used herein, the term "Next Generation Sequencing (NGS)" refers to sequencing methods that enable massively parallel sequencing of clonally amplified molecules and single nucleic acid molecules. Non-limiting examples of NGS include sequencing-by-synthesis using reversible dye terminators and sequencing-by-ligation.
[0042] As used herein, the term "sensitivity" is equal to the number of true positives divided by the sum of the true positives and false negatives.
[0043] As used herein, the term "specificity" is equal to the number of true negatives divided by the sum of the true negatives and false positives.
[0044] As used herein, the term "enriched" refers to a process of amplifying nucleic acids contained in a portion of a sample. Enrichment includes specific enrichment, which targets specific sequences, e.g., polymorphic sequences, and non-specific enrichment, which amplifies entire genomes of DNA fragments in a sample.
[0045] As used herein, the term "each," when used in reference to a collection of items, is intended to identify each individual item in the set, but does not necessarily refer to every item in the set, unless the context clearly dictates otherwise.
[0046] As used herein, "providing" in the context of a composition, article, nucleic acid, or nucleus means making the composition, article, nucleic acid, or nucleus, purchasing the composition, article, nucleic acid, or nucleus, or otherwise obtaining the compound, composition, article, or nucleus.
[0047] The term "and / or" means one or all of the listed elements or a combination of any two or more of the listed elements.
[0048] The words "preferred" and "preferably" refer to embodiments of the present disclosure that may offer certain benefits, under particular circumstances. However, other embodiments may also be preferred, under the same or other circumstances. Furthermore, the recitation of one or more preferred embodiments does not imply that other embodiments are not useful, and is not intended to exclude other embodiments from the scope of the present disclosure.
[0049] The terms "comprises" and variations thereof do not have a limiting meaning where these terms appear in the description and claims.
[0050] As used herein, wherever words such as "include," "includes," or "including" are used herein, it is understood that analogous embodiments described with the terms "consisting of" and / or "consisting essentially of" are also provided.
[0051] Unless otherwise noted, "a," "an," "the," and "at least one" are used interchangeably and mean one or more.
[0052] As used herein, the recitations of numerical ranges by endpoints include all numbers subsumed within that range (eg, 1 to 5 includes 1, 1.5, 2, 2.75, 3, 3.80, 4, 5, etc.).
[0053] References to "one embodiment," "an embodiment," "particular embodiments," or "some embodiments" mean that the particular feature, configuration, composition, or characteristic described in connection with this embodiment is included in at least one embodiment of the present disclosure. Thus, the appearances of such phrases in various places throughout this specification do not necessarily refer to the same embodiment of the present disclosure. Furthermore, the particular features, configurations, compositions, or characteristics may be combined in any suitable manner in one or more embodiments.
[0054] In any method disclosed herein that includes distinct steps, the steps may be performed in any practicable order, and, suitably, any combination of two or more steps may be performed simultaneously.
[0055] The above summary of the present disclosure is not intended to describe each disclosed embodiment or every implementation of the present disclosure. The following description more particularly exemplifies exemplary embodiments. In several places throughout the application, guidance is provided through lists of examples, which examples can be used in various combinations. In each instance, the recited list serves only as a representative group and should not be interpreted as an exclusive list. [Brief explanation of the drawings]
[0056] [Figure 1] Schematic representation of the relationship between placental health, maternal response, and fetal response. [Figure 2] Origin of circular RNA (C-RNA). [Figure 3] Library preparation workflow for c-RNA. [Figure 4] Validation of the C-RNA method comparing late-pregnancy and non-pregnant samples. [Figure 5] Validation of the C-RNA method using longitudinal pregnancy samples. [Figure 6] Clinical trial description. [Figure 7] Sequencing data characteristics. [Figure 8] Classification of PEs without gene selection and relying on the entire dataset. [Figure 9] Explanation of the bootstrap method. [Figure 10] Classification of preeclampsia samples using bootstrap techniques. [Figure 11] Excessive pre-eclampsia gene testing. [Figure 12] Standard Adaboost model. [Figure 13] An independent cohort will allow further validation of the preeclampsia signature. [Figure 14] Performance of the standard Adaboost model in classifying preeclampsia. [Figure 15] Classification of preeclampsia using the standard DEX TREAT assay. [Figure 16]Gene selection and classification of preeclampsia using the jackknife approach. [Figure 17] Validation of TREAT, bootstrap, and jackknife methods in an independent PEARL Biobank cohort. [Figure 18] Diagram of the bioinformatics approach to building the AdaBoost Refined model. [Figure 19] Relative abundance of genes utilized by the AdaBoost Refined model and their predictive ability on an independent dataset. [Figure 20] Identification of pre-eclampsia specific C-RNA signatures in Nextera Flex generated libraries using standard TREAT analysis and jackknife methods. [Figure 21] Relative abundance of genes utilized by AdaBoost Refined models on Nextera Flex-generated libraries and their predictive power in the RGH14 dataset. [Figure 22A] Validation of a clinically relevant whole-exome C-RNA analysis method. Figure 22A is a schematic diagram of the sequencing library preparation method. All steps after blood collection can be performed in a centralized processing laboratory. Temporal changes in altered transcripts throughout pregnancy (Figure 22B). Overlap of genes identified in the C-RNA pregnancy progression study (Figure 22C). Tissues expressing 91 genes unique to the pregnancy progression study (Figure 22D). [Figure 22B] Validation of a clinically relevant whole-exome C-RNA analysis method. Figure 22A is a schematic diagram of the sequencing library preparation method. All steps after blood collection can be performed in a centralized processing laboratory. Temporal changes in altered transcripts throughout pregnancy (Figure 22B). Overlap of genes identified in the C-RNA pregnancy progression study (Figure 22C). Tissues expressing 91 genes unique to the pregnancy progression study (Figure 22D). [Figure 22C]Validation of a clinically relevant whole-exome C-RNA analysis method. Figure 22A is a schematic diagram of the sequencing library preparation method. All steps after blood collection can be performed in a centralized processing laboratory. Temporal changes in altered transcripts throughout pregnancy (Figure 22B). Overlap of genes identified in the C-RNA pregnancy progression study (Figure 22C). Tissues expressing 91 genes unique to the pregnancy progression study (Figure 22D). [Figure 22D] Validation of a clinically relevant whole-exome C-RNA analysis method. Figure 22A is a schematic diagram of the sequencing library preparation method. All steps after blood collection can be performed in a centralized processing laboratory. Temporal changes in altered transcripts throughout pregnancy (Figure 22B). Overlap of genes identified in the C-RNA pregnancy progression study (Figure 22C). Tissues expressing 91 genes unique to the pregnancy progression study (Figure 22D). [Figure 23A] Sample collection for PE clinical trials. The panels show the time of blood collection (triangles) and gestational age at birth (squares) for each individual in the iPC study (Figure 23A) and the PEARL study (Figure 23B). The red line indicates the threshold for full-term delivery. The rate of early term delivery is significantly higher in the preterm PE cohort (Figure 23C). ***p<0.001 by Fisher's exact test. [Figure 23B] Sample collection for PE clinical trials. The panels show the time of blood collection (triangles) and gestational age at birth (squares) for each individual in the iPC study (Figure 23A) and the PEARL study (Figure 23B). The red line indicates the threshold for full-term delivery. The rate of early term delivery is significantly higher in the preterm PE cohort (Figure 23C). ***p<0.001 by Fisher's exact test. [Figure 23C] Sample collection for PE clinical trials. The panels show the time of blood collection (triangles) and gestational age at birth (squares) for each individual in the iPC study (Figure 23A) and the PEARL study (Figure 23B). The red line indicates the threshold for full-term delivery. The rate of early term delivery is significantly higher in the preterm PE cohort (Figure 23C). ***p<0.001 by Fisher's exact test. [Figure 24A]Differential analysis of C-RNA identifies preeclampsia biomarkers. Fold change and abundance of altered transcripts in PE (Figure 24A). One-sided confidence p-intervals were calculated after jackknife analysis for each gene detected by standard analysis (Figure 24B). Fold change in transcript abundance determined by whole exome sequencing and qPCR of (21) genes (Figure 24C). *p<0.05 by Student's T-test. Tissue distribution of disease genes (Figure 24D). Hierarchical clustering (average linkage, squared Euclidean distance) of iPC samples (Figure 24E). Clustering of early PE (Figure 24F) and late PE (Figure 24G) samples from the PEARL study. [Figure 24B] Differential analysis of C-RNA identifies preeclampsia biomarkers. Fold change and abundance of altered transcripts in PE (Figure 24A). One-sided confidence p-intervals were calculated after jackknife analysis for each gene detected by standard analysis (Figure 24B). Fold change in transcript abundance determined by whole exome sequencing and qPCR of (21) genes (Figure 24C). *p<0.05 by Student's T-test. Tissue distribution of disease genes (Figure 24D). Hierarchical clustering (average linkage, squared Euclidean distance) of iPC samples (Figure 24E). Clustering of early PE (Figure 24F) and late PE (Figure 24G) samples from the PEARL study. [Figure 24C] Differential analysis of C-RNA identifies preeclampsia biomarkers. Fold change and abundance of altered transcripts in PE (Figure 24A). One-sided confidence p-intervals were calculated after jackknife analysis for each gene detected by standard analysis (Figure 24B). Fold change in transcript abundance determined by whole exome sequencing and qPCR of (21) genes (Figure 24C). *p<0.05 by Student's T-test. Tissue distribution of disease genes (Figure 24D). Hierarchical clustering (average linkage, squared Euclidean distance) of iPC samples (Figure 24E). Clustering of early PE (Figure 24F) and late PE (Figure 24G) samples from the PEARL study. [Figure 24D]Differential analysis of C-RNA identifies preeclampsia biomarkers. Fold change and abundance of altered transcripts in PE (Figure 24A). One-sided confidence p-intervals were calculated after jackknife analysis for each gene detected by standard analysis (Figure 24B). Fold change in transcript abundance determined by whole exome sequencing and qPCR of (21) genes (Figure 24C). *p<0.05 by Student's T-test. Tissue distribution of disease genes (Figure 24D). Hierarchical clustering (average linkage, squared Euclidean distance) of iPC samples (Figure 24E). Clustering of early PE (Figure 24F) and late PE (Figure 24G) samples from the PEARL study. [Figure 24E] Differential analysis of C-RNA identifies preeclampsia biomarkers. Fold change and abundance of altered transcripts in PE (Figure 24A). One-sided confidence p-intervals were calculated after jackknife analysis for each gene detected by standard analysis (Figure 24B). Fold change in transcript abundance determined by whole exome sequencing and qPCR of (21) genes (Figure 24C). *p<0.05 by Student's T-test. Tissue distribution of disease genes (Figure 24D). Hierarchical clustering (average linkage, squared Euclidean distance) of iPC samples (Figure 24E). Clustering of early PE (Figure 24F) and late PE (Figure 24G) samples from the PEARL study. [Figure 24F] Differential analysis of C-RNA identifies preeclampsia biomarkers. Fold change and abundance of altered transcripts in PE (Figure 24A). One-sided confidence p-intervals were calculated after jackknife analysis for each gene detected by standard analysis (Figure 24B). Fold change in transcript abundance determined by whole exome sequencing and qPCR of (21) genes (Figure 24C). *p<0.05 by Student's T-test. Tissue distribution of disease genes (Figure 24D). Hierarchical clustering (average linkage, squared Euclidean distance) of iPC samples (Figure 24E). Clustering of early PE (Figure 24F) and late PE (Figure 24G) samples from the PEARL study. [Figure 24G]Differential analysis of C-RNA identifies preeclampsia biomarkers. Fold change and abundance of altered transcripts in PE (Figure 24A). One-sided confidence p-intervals were calculated after jackknife analysis for each gene detected by standard analysis (Figure 24B). Fold change in transcript abundance determined by whole exome sequencing and qPCR of (21) genes (Figure 24C). *p<0.05 by Student's T-test. Tissue distribution of disease genes (Figure 24D). Hierarchical clustering (average linkage, squared Euclidean distance) of iPC samples (Figure 24E). Clustering of early PE (Figure 24F) and late PE (Figure 24G) samples from the PEARL study. [Figure 25A] AdaBoost classifies preeclampsia samples across cohorts. Heatmap showing the relative abundance of transcripts used by machine learning in each cohort (Figure 25A). The height of each block reflects the importance of each gene. ROC curves for each dataset (Figure 25B). Distribution of AdaBoost scores (KDE). The orange line indicates the optimal boundary for distinguishing PE and control samples (Figure 25C). Concordance between genes identified by differential analysis and genes used in AdaBoost (Figure 25D). Tissue distribution of AdaBoost genes (Figure 25E). [Figure 25B] AdaBoost classifies preeclampsia samples across cohorts. Heatmap showing the relative abundance of transcripts used by machine learning in each cohort (Figure 25A). The height of each block reflects the importance of each gene. ROC curves for each dataset (Figure 25B). Distribution of AdaBoost scores (KDE). The orange line indicates the optimal boundary for distinguishing PE and control samples (Figure 25C). Concordance between genes identified by differential analysis and genes used in AdaBoost (Figure 25D). Tissue distribution of AdaBoost genes (Figure 25E). [Figure 25C]AdaBoost classifies preeclampsia samples across cohorts. Heatmap showing the relative abundance of transcripts used by machine learning in each cohort (Figure 25A). The height of each block reflects the importance of each gene. ROC curves for each dataset (Figure 25B). Distribution of AdaBoost scores (KDE). The orange line indicates the optimal boundary for distinguishing PE and control samples (Figure 25C). Concordance between genes identified by differential analysis and genes used in AdaBoost (Figure 25D). Tissue distribution of AdaBoost genes (Figure 25E). [Figure 25D] AdaBoost classifies preeclampsia samples across cohorts. Heatmap showing the relative abundance of transcripts used by machine learning in each cohort (Figure 25A). The height of each block reflects the importance of each gene. ROC curves for each dataset (Figure 25B). Distribution of AdaBoost scores (KDE). The orange line indicates the optimal boundary for distinguishing PE and control samples (Figure 25C). Concordance between genes identified by differential analysis and genes used in AdaBoost (Figure 25D). Tissue distribution of AdaBoost genes (Figure 25E). [Figure 25E] AdaBoost classifies preeclampsia samples across cohorts. Heatmap showing the relative abundance of transcripts used by machine learning in each cohort (Figure 25A). The height of each block reflects the importance of each gene. ROC curves for each dataset (Figure 25B). Distribution of AdaBoost scores (KDE). The orange line indicates the optimal boundary for distinguishing PE and control samples (Figure 25C). Concordance between genes identified by differential analysis and genes used in AdaBoost (Figure 25D). Tissue distribution of AdaBoost genes (Figure 25E). [Figure 26A]C-RNA data integrity when blood is stored in different collection tubes. Abundance of previously detected C-RNA pregnancy markers from blood stored overnight in different tube types is compared to immediate processing after collection in EDTA tubes (Figure 26A). Scatter plot comparing transcript FPKM values of C-RNA prepared from the same individual after different blood storage periods (Figure 26B). Pearson's correlation coefficient R is more variable when using EDTA tubes (see cell-free) (Figure 26C). [Figure 26B] C-RNA data integrity when blood is stored in different collection tubes. Abundance of previously detected C-RNA pregnancy markers from blood stored overnight in different tube types is compared to immediate processing after collection in EDTA tubes (Figure 26A). Scatter plot comparing transcript FPKM values of C-RNA prepared from the same individual after different blood storage periods (Figure 26B). Pearson's correlation coefficient R is more variable when using EDTA tubes (see cell-free) (Figure 26C). [Figure 26C] C-RNA data integrity when blood is stored in different collection tubes. Abundance of previously detected C-RNA pregnancy markers from blood stored overnight in different tube types is compared to immediate processing after collection in EDTA tubes (Figure 26A). Scatter plot comparing transcript FPKM values of C-RNA prepared from the same individual after different blood storage periods (Figure 26B). Pearson's correlation coefficient R is more variable when using EDTA tubes (see cell-free) (Figure 26C). [Figure 27A] Effect of plasma volume on C-RNA data quality. A meta-analysis was performed using data from nine independent studies to determine the appropriate plasma input for the protocol. Noise (biological coefficient of variation, EdgeR) was calculated from biological replicates within each study (Figure 27A). Library complexity (edge population, Preseq) was calculated for each sample (Figure 27B). **p<0.01, ***p<0.001 by ANOVA with Tukey's HSD correction using study as the blocking variable. [Figure 27B]Effect of plasma volume on C-RNA data quality. A meta-analysis was performed using data from nine independent studies to determine the appropriate plasma input for the protocol. Noise (biological coefficient of variation, EdgeR) was calculated from biological replicates within each study (Figure 27A). Library complexity (edge population, Preseq) was calculated for each sample (Figure 27B). **p<0.01, ***p<0.001 by ANOVA with Tukey's HSD correction using study as the blocking variable. [Figure 28A] Pregnancy marker tissue specificity. Pie charts showing the tissue specificity of genes detected during pregnancy by three independent tests using either the full set of altered genes (Figure 28A), transcripts specific to each test (Figure 28B), or the intersection gene set (Figure 28C). [Figure 28B] Pregnancy marker tissue specificity. Pie charts showing the tissue specificity of genes detected during pregnancy by three independent tests using either the full set of altered genes (Figure 28A), transcripts specific to each test (Figure 28B), or the intersection gene set (Figure 28C). [Figure 28C] Pregnancy marker tissue specificity. Pie charts showing the tissue specificity of genes detected during pregnancy by three independent tests using either the full set of altered genes (Figure 28A), transcripts specific to each test (Figure 28B), or the intersection gene set (Figure 28C). [Figure 29A]The jackknife method excludes genes that are not universally altered in preeclampsia. Schematic of the jackknife method used to determine how consistently transcripts were altered across PE samples (Figure 29A). Mean abundance and noise for each differentially abundant gene (Figure 29B). The ROC area under the curve value for each disease transcript provides an indication of how separated the C-RNA transcript abundance distributions are for control and PE samples (Figure 29C). *P<0.05 by Mann-Whitney U test. Hierarchical clustering of iPC samples using genes excluded after jackknife method (Figure 29D). Tissue distribution of excluded transcripts (Figure 29E). The reduced contribution of fetus and placenta may suggest that the maternal component of PE is most variable across individuals. [Figure 29B] The jackknife method excludes genes that are not universally altered in preeclampsia. Schematic of the jackknife method used to determine how consistently transcripts were altered across PE samples (Figure 29A). Mean abundance and noise for each differentially abundant gene (Figure 29B). The ROC area under the curve value for each disease transcript provides an indication of how separated the C-RNA transcript abundance distributions are for control and PE samples (Figure 29C). *P<0.05 by Mann-Whitney U test. Hierarchical clustering of iPC samples using genes excluded after jackknife method (Figure 29D). Tissue distribution of excluded transcripts (Figure 29E). The reduced contribution of fetus and placenta may suggest that the maternal component of PE is most variable across individuals. [Figure 29C]The jackknife method excludes genes that are not universally altered in preeclampsia. Schematic of the jackknife method used to determine how consistently transcripts were altered across PE samples (Figure 29A). Mean abundance and noise for each differentially abundant gene (Figure 29B). The ROC area under the curve value for each disease transcript provides an indication of how separated the C-RNA transcript abundance distributions are for control and PE samples (Figure 29C). *P<0.05 by Mann-Whitney U test. Hierarchical clustering of iPC samples using genes excluded after jackknife method (Figure 29D). Tissue distribution of excluded transcripts (Figure 29E). The reduced contribution of fetus and placenta may suggest that the maternal component of PE is most variable across individuals. [Figure 29D] The jackknife method excludes genes that are not universally altered in preeclampsia. Schematic of the jackknife method used to determine how consistently transcripts were altered across PE samples (Figure 29A). Mean abundance and noise for each differentially abundant gene (Figure 29B). The ROC area under the curve value for each disease transcript provides an indication of how separated the C-RNA transcript abundance distributions are for control and PE samples (Figure 29C). *P<0.05 by Mann-Whitney U test. Hierarchical clustering of iPC samples using genes excluded after jackknife method (Figure 29D). Tissue distribution of excluded transcripts (Figure 29E). The reduced contribution of fetus and placenta may suggest that the maternal component of PE is most variable across individuals. [Figure 29E]The jackknife method excludes genes that are not universally altered in preeclampsia. Schematic of the jackknife method used to determine how consistently transcripts were altered across PE samples (Figure 29A). Mean abundance and noise for each differentially abundant gene (Figure 29B). The ROC area under the curve value for each disease transcript provides an indication of how separated the C-RNA transcript abundance distributions are for control and PE samples (Figure 29C). *P<0.05 by Mann-Whitney U test. Hierarchical clustering of iPC samples using genes excluded after jackknife method (Figure 29D). Tissue distribution of excluded transcripts (Figure 29E). The reduced contribution of fetus and placenta may suggest that the maternal component of PE is most variable across individuals. [Figure 30A]AdaBoost model development strategy. The RGH014 dataset was divided into six pieces (Figure 30A). The "holdout subset," which included 10% of the samples (randomly selected) and three samples that were incorrectly clustered using differentially abundant genes (as in Figure 24C), was completely excluded from model building. The remaining samples were randomly divided into five uniformly sized "test subsets." For each test subset, the training data consisted of all non-holdout samples and non-test samples. Gene counts for the training and test data were TMM-normalized in edgeR and then standardized to a mean of 0 and a standard deviation of 1 for each gene. For each training / test sample set, AdaBoost models (90 estimators, 1.6 learning rate) were built from the training data 10 times (Figure 30B). Feature pruning was performed to remove genes below gradually increasing importance thresholds, and performance was evaluated by the Matthews correlation coefficient when predicting the test data. The model with the best performance (and in case of ties, the one with the fewest genes) was retained. Estimators from all 50 independent models were combined into a single AdaBoost model (Figure 30C). Feature pruning was then performed on the resulting ensemble, this time using the percentage of models that used genes to set a threshold and performance, as measured by the average log-loss value across the test subset. ROC curve after applying the final AdaBoost model to the holdout data (Figure 30D). All samples were accurately separated, except for two of the three samples, which were similarly misclustered by HCA. [Figure 30B]AdaBoost model development strategy. The RGH014 dataset was divided into six pieces (Figure 30A). The "holdout subset," which included 10% of the samples (randomly selected) and three samples that were incorrectly clustered using differentially abundant genes (as in Figure 24C), was completely excluded from model building. The remaining samples were randomly divided into five uniformly sized "test subsets." For each test subset, the training data consisted of all non-holdout samples and non-test samples. Gene counts for the training and test data were TMM-normalized in edgeR and then standardized to a mean of 0 and a standard deviation of 1 for each gene. For each training / test sample set, AdaBoost models (90 estimators, 1.6 learning rate) were built from the training data 10 times (Figure 30B). Feature pruning was performed to remove genes below gradually increasing importance thresholds, and performance was evaluated by the Matthews correlation coefficient when predicting the test data. The model with the best performance (and in case of ties, the one with the fewest genes) was retained. Estimators from all 50 independent models were combined into a single AdaBoost model (Figure 30C). Feature pruning was then performed on the resulting ensemble, this time using the percentage of models that used genes to set a threshold and performance, as measured by the average log-loss value across the test subset. ROC curve after applying the final AdaBoost model to the holdout data (Figure 30D). All samples were accurately separated, except for two of the three samples, which were similarly misclustered by HCA. [Figure 30C]AdaBoost model development strategy. The RGH014 dataset was divided into six pieces (Figure 30A). The "holdout subset," which included 10% of the samples (randomly selected) and three samples that were incorrectly clustered using differentially abundant genes (as in Figure 24C), was completely excluded from model building. The remaining samples were randomly divided into five uniformly sized "test subsets." For each test subset, the training data consisted of all non-holdout samples and non-test samples. Gene counts for the training and test data were TMM-normalized in edgeR and then standardized to a mean of 0 and a standard deviation of 1 for each gene. For each training / test sample set, AdaBoost models (90 estimators, 1.6 learning rate) were built from the training data 10 times (Figure 30B). Feature pruning was performed to remove genes below gradually increasing importance thresholds, and performance was evaluated by the Matthews correlation coefficient when predicting the test data. The model with the best performance (and in case of ties, the one with the fewest genes) was retained. Estimators from all 50 independent models were combined into a single AdaBoost model (Figure 30C). Feature pruning was then performed on the resulting ensemble, this time using the percentage of models that used genes to set a threshold and performance, as measured by the average log-loss value across the test subset. ROC curve after applying the final AdaBoost model to the holdout data (Figure 30D). All samples were accurately separated, except for two of the three samples, which were similarly misclustered by HCA. [Figure 30D]AdaBoost model development strategy. The RGH014 dataset was divided into six pieces (Figure 30A). The "holdout subset," which included 10% of the samples (randomly selected) and three samples that were incorrectly clustered using differentially abundant genes (as in Figure 24C), was completely excluded from model building. The remaining samples were randomly divided into five uniformly sized "test subsets." For each test subset, the training data consisted of all non-holdout samples and non-test samples. Gene counts for the training and test data were TMM-normalized in edgeR and then standardized to a mean of 0 and a standard deviation of 1 for each gene. For each training / test sample set, AdaBoost models (90 estimators, 1.6 learning rate) were built from the training data 10 times (Figure 30B). Feature pruning was performed to remove genes below gradually increasing importance thresholds, and performance was evaluated by the Matthews correlation coefficient when predicting the test data. The model with the best performance (and in case of ties, the one with the fewest genes) was retained. Estimators from all 50 independent models were combined into a single AdaBoost model (Figure 30C). Feature pruning was then performed on the resulting ensemble, this time using the percentage of models that used genes to set a threshold and performance, as measured by the average log-loss value across the test subset. ROC curve after applying the final AdaBoost model to the holdout data (Figure 30D). All samples were accurately separated, except for two of the three samples, which were similarly misclustered by HCA. [Figure 31A]Impact of hyperparameter selection and feature pruning on machine learning performance. Heatmap of grid search to identify optimal hyperparameters for AdaBoost (Fig. 31A). Matthews correlation coefficient was used as a performance metric. Flattened plot for each hyperparameter (Fig. 31B). Arrows indicate values selected for model construction. Fig. 31C shows the impact of pruning individual AdaBoost models on performance (as in Fig. 30B). The solid line is the average of all 10 models, and the shaded area indicates the standard deviation. Number of AdaBoost models using each gene observed in the pre-pruned ensemble (Fig. 31D). Model performance when pruning the combined AdaBoost ensemble (Fig. 31E). The orange lines in Figs. 31D and 31E indicate the threshold applied to generate the final AdaBoost model. [Figure 31B] Impact of hyperparameter selection and feature pruning on machine learning performance. Heatmap of grid search to identify optimal hyperparameters for AdaBoost (Fig. 31A). Matthews correlation coefficient was used as a performance metric. Flattened plot for each hyperparameter (Fig. 31B). Arrows indicate values selected for model construction. Fig. 31C shows the impact of pruning individual AdaBoost models on performance (as in Fig. 30B). The solid line is the average of all 10 models, and the shaded area indicates the standard deviation. Number of AdaBoost models using each gene observed in the pre-pruned ensemble (Fig. 31D). Model performance when pruning the combined AdaBoost ensemble (Fig. 31E). The orange lines in Figs. 31D and 31E indicate the threshold applied to generate the final AdaBoost model. [Figure 31C]Impact of hyperparameter selection and feature pruning on machine learning performance. Heatmap of grid search to identify optimal hyperparameters for AdaBoost (Fig. 31A). Matthews correlation coefficient was used as a performance metric. Flattened plot for each hyperparameter (Fig. 31B). Arrows indicate values selected for model construction. Fig. 31C shows the impact of pruning individual AdaBoost models on performance (as in Fig. 30B). The solid line is the average of all 10 models, and the shaded area indicates the standard deviation. Number of AdaBoost models using each gene observed in the pre-pruned ensemble (Fig. 31D). Model performance when pruning the combined AdaBoost ensemble (Fig. 31E). The orange lines in Figs. 31D and 31E indicate the threshold applied to generate the final AdaBoost model. [Figure 31D] Impact of hyperparameter selection and feature pruning on machine learning performance. Heatmap of grid search to identify optimal hyperparameters for AdaBoost (Fig. 31A). Matthews correlation coefficient was used as a performance metric. Flattened plot for each hyperparameter (Fig. 31B). Arrows indicate values selected for model construction. Fig. 31C shows the impact of pruning individual AdaBoost models on performance (as in Fig. 30B). The solid line is the average of all 10 models, and the shaded area indicates the standard deviation. Number of AdaBoost models using each gene observed in the pre-pruned ensemble (Fig. 31D). Model performance when pruning the combined AdaBoost ensemble (Fig. 31E). The orange lines in Figs. 31D and 31E indicate the threshold applied to generate the final AdaBoost model. [Figure 31E]Impact of hyperparameter selection and feature pruning on machine learning performance. Heatmap of grid search to identify optimal hyperparameters for AdaBoost (Fig. 31A). Matthews correlation coefficient was used as a performance metric. Flattened plot for each hyperparameter (Fig. 31B). Arrows indicate values selected for model construction. Fig. 31C shows the impact of pruning individual AdaBoost models on performance (as in Fig. 30B). The solid line is the average of all 10 models, and the shaded area indicates the standard deviation. Number of AdaBoost models using each gene observed in the pre-pruned ensemble (Fig. 31D). Model performance when pruning the combined AdaBoost ensemble (Fig. 31E). The orange lines in Figs. 31D and 31E indicate the threshold applied to generate the final AdaBoost model. [Figure 32A] Changes in C-RNA transcriptome tracks with pregnancy progression. Figure 32A shows the temporal changes of significantly altered transcripts throughout pregnancy. Each column corresponds to a transcript, whose abundance was normalized across all samples (N=152) before clustering. Orange indicates high abundance, and purple indicates decreased abundance. Figure 32B shows the overlap of transcripts identified in three independent C-RNA pregnancy progression analyses. N=number of plasma samples collected from pregnant women in the cohort. Figure 32C shows tissues expressing the 91 genes detected exclusively in the PEARL HCC cohort. [Figure 32B] Changes in C-RNA transcriptome tracks with pregnancy progression. Figure 32A shows the temporal changes of significantly altered transcripts throughout pregnancy. Each column corresponds to a transcript, whose abundance was normalized across all samples (N=152) before clustering. Orange indicates high abundance, and purple indicates decreased abundance. Figure 32B shows the overlap of transcripts identified in three independent C-RNA pregnancy progression analyses. N=number of plasma samples collected from pregnant women in the cohort. Figure 32C shows tissues expressing the 91 genes detected exclusively in the PEARL HCC cohort. [Figure 32C] Changes in C-RNA transcriptome tracks with pregnancy progression. Figure 32A shows the temporal changes of significantly altered transcripts throughout pregnancy. Each column corresponds to a transcript, whose abundance was normalized across all samples (N=152) before clustering. Orange indicates high abundance, and purple indicates decreased abundance. Figure 32B shows the overlap of transcripts identified in three independent C-RNA pregnancy progression analyses. N=number of plasma samples collected from pregnant women in the cohort. Figure 32C shows tissues expressing the 91 genes detected exclusively in the PEARL HCC cohort. [Figure 33A] For control and PE samples, sample collection was consistent, but clinical outcomes were not. The panels show the time of blood collection (triangles) and gestational age at birth (squares) for each individual in the iPEC study (Figure 33A) and PEARL PEC study (Figure 33B). The red line indicates the 37-week live-term threshold. As shown in Figure 33C, the rate of preterm birth was significantly higher in the preterm PE cohort. ***p<0.001 by Fisher's exact test. iPEC, control N=73, PE N=40; PEARL PEC, N=12 for each group. [Figure 33B] For control and PE samples, sample collection was consistent, but clinical outcomes were not. The panels show the time of blood collection (triangles) and gestational age at birth (squares) for each individual in the iPEC study (Figure 33A) and PEARL PEC study (Figure 33B). The red line indicates the 37-week live-term threshold. As shown in Figure 33C, the rate of preterm birth was significantly higher in the preterm PE cohort. ***p<0.001 by Fisher's exact test. iPEC, control N=73, PE N=40; PEARL PEC, N=12 for each group. [Figure 33C]For control and PE samples, sample collection was consistent, but clinical outcomes were not. The panels show the time of blood collection (triangles) and gestational age at birth (squares) for each individual in the iPEC study (Figure 33A) and PEARL PEC study (Figure 33B). The red line indicates the 37-week live-term threshold. As shown in Figure 33C, the rate of preterm birth was significantly higher in the preterm PE cohort. ***p<0.001 by Fisher's exact test. iPEC, control N=73, PE N=40; PEARL PEC, N=12 for each group. [Figure 34A] Applying the jackknife method to differential expression analysis eliminates genes with low sensitivity for identifying PE samples. Figure 34A shows the fold change and abundance of altered transcripts in PE. Figure 34B shows the one-sided, normal-based 95% confidence interval of the p-value for each transcript detected as altered by standard analysis. Figure 34C is a schematic diagram of the jackknife method, in which 90% of samples were randomly selected for analysis across many replicates to quantify p-value stability. Figure 34D shows the mean abundance and noise of each differentially abundant gene (p>0.05 for both variables by Mann-Whitney U test; N=30, 12). Figure 34E shows the ROC area under the curve for each disease transcript, reflecting how the C-RNA transcript abundance distribution separates for control versus PE samples (*p<0.05 by Mann-Whitney U test; include, N=30; exclude, N=12). Figure 34F shows hierarchical clustering of iPEC samples using genes excluded after jackknife analysis (sensitivity = 73%, specificity = 99%, N = 113). In Figures 34A, 34B, 34D, and 34E, orange and blue data points reflect transcripts that would be statistically altered in PE by standard differential expression analysis, but only orange data points were identified by jackknife analysis. Sample status is indicated by blue (PE) and gray (control) rectangles along the right side of the heart map. [Figure 34B]Applying the jackknife method to differential expression analysis eliminates genes with low sensitivity for identifying PE samples. Figure 34A shows the fold change and abundance of altered transcripts in PE. Figure 34B shows the one-sided, normal-based 95% confidence interval of the p-value for each transcript detected as altered by standard analysis. Figure 34C is a schematic diagram of the jackknife method, in which 90% of samples were randomly selected for analysis across many replicates to quantify p-value stability. Figure 34D shows the mean abundance and noise of each differentially abundant gene (p>0.05 for both variables by Mann-Whitney U test; N=30, 12). Figure 34E shows the ROC area under the curve for each disease transcript, reflecting how the C-RNA transcript abundance distribution separates for control versus PE samples (*p<0.05 by Mann-Whitney U test; include, N=30; exclude, N=12). Figure 34F shows hierarchical clustering of iPEC samples using genes excluded after jackknife analysis (sensitivity = 73%, specificity = 99%, N = 113). In Figures 34A, 34B, 34D, and 34E, orange and blue data points reflect transcripts that would be statistically altered in PE by standard differential expression analysis, but only orange data points were identified by jackknife analysis. Sample status is indicated by blue (PE) and gray (control) rectangles along the right side of the heart map. [Figure 34C]Applying the jackknife method to differential expression analysis eliminates genes with low sensitivity for identifying PE samples. Figure 34A shows the fold change and abundance of altered transcripts in PE. Figure 34B shows the one-sided, normal-based 95% confidence interval of the p-value for each transcript detected as altered by standard analysis. Figure 34C is a schematic diagram of the jackknife method, in which 90% of samples were randomly selected for analysis across many replicates to quantify p-value stability. Figure 34D shows the mean abundance and noise of each differentially abundant gene (p>0.05 for both variables by Mann-Whitney U test; N=30, 12). Figure 34E shows the ROC area under the curve for each disease transcript, reflecting how the C-RNA transcript abundance distribution separates for control versus PE samples (*p<0.05 by Mann-Whitney U test; include, N=30; exclude, N=12). Figure 34F shows hierarchical clustering of iPEC samples using genes excluded after jackknife analysis (sensitivity = 73%, specificity = 99%, N = 113). In Figures 34A, 34B, 34D, and 34E, orange and blue data points reflect transcripts that would be statistically altered in PE by standard differential expression analysis, but only orange data points were identified by jackknife analysis. Sample status is indicated by blue (PE) and gray (control) rectangles along the right side of the heart map. [Figure 34D]Applying the jackknife method to differential expression analysis eliminates genes with low sensitivity for identifying PE samples. Figure 34A shows the fold change and abundance of altered transcripts in PE. Figure 34B shows the one-sided, normal-based 95% confidence interval of the p-value for each transcript detected as altered by standard analysis. Figure 34C is a schematic diagram of the jackknife method, in which 90% of samples were randomly selected for analysis across many replicates to quantify p-value stability. Figure 34D shows the mean abundance and noise of each differentially abundant gene (p>0.05 for both variables by Mann-Whitney U test; N=30, 12). Figure 34E shows the ROC area under the curve for each disease transcript, reflecting how the C-RNA transcript abundance distribution separates for control versus PE samples (*p<0.05 by Mann-Whitney U test; include, N=30; exclude, N=12). Figure 34F shows hierarchical clustering of iPEC samples using genes excluded after jackknife analysis (sensitivity = 73%, specificity = 99%, N = 113). In Figures 34A, 34B, 34D, and 34E, orange and blue data points reflect transcripts that would be statistically altered in PE by standard differential expression analysis, but only orange data points were identified by jackknife analysis. Sample status is indicated by blue (PE) and gray (control) rectangles along the right side of the heart map. [Figure 34E]Applying the jackknife method to differential expression analysis eliminates genes with low sensitivity for identifying PE samples. Figure 34A shows the fold change and abundance of altered transcripts in PE. Figure 34B shows the one-sided, normal-based 95% confidence interval of the p-value for each transcript detected as altered by standard analysis. Figure 34C is a schematic diagram of the jackknife method, in which 90% of samples were randomly selected for analysis across many replicates to quantify p-value stability. Figure 34D shows the mean abundance and noise of each differentially abundant gene (p>0.05 for both variables by Mann-Whitney U test; N=30, 12). Figure 34E shows the ROC area under the curve for each disease transcript, reflecting how the C-RNA transcript abundance distribution separates for control versus PE samples (*p<0.05 by Mann-Whitney U test; include, N=30; exclude, N=12). Figure 34F shows hierarchical clustering of iPEC samples using genes excluded after jackknife analysis (sensitivity = 73%, specificity = 99%, N = 113). In Figures 34A, 34B, 34D, and 34E, orange and blue data points reflect transcripts that would be statistically altered in PE by standard differential expression analysis, but only orange data points were identified by jackknife analysis. Sample status is indicated by blue (PE) and gray (control) rectangles along the right side of the heart map. [Figure 34F]Applying the jackknife method to differential expression analysis eliminates genes with low sensitivity for identifying PE samples. Figure 34A shows the fold change and abundance of altered transcripts in PE. Figure 34B shows the one-sided, normal-based 95% confidence interval of the p-value for each transcript detected as altered by standard analysis. Figure 34C is a schematic diagram of the jackknife method, in which 90% of samples were randomly selected for analysis across many replicates to quantify p-value stability. Figure 34D shows the mean abundance and noise of each differentially abundant gene (p>0.05 for both variables by Mann-Whitney U test; N=30, 12). Figure 34E shows the ROC area under the curve for each disease transcript, reflecting how the C-RNA transcript abundance distribution separates for control versus PE samples (*p<0.05 by Mann-Whitney U test; include, N=30; exclude, N=12). Figure 34F shows hierarchical clustering of iPEC samples using genes excluded after jackknife analysis (sensitivity = 73%, specificity = 99%, N = 113). In Figures 34A, 34B, 34D, and 34E, orange and blue data points reflect transcripts that would be statistically altered in PE by standard differential expression analysis, but only orange data points were identified by jackknife analysis. Sample status is indicated by blue (PE) and gray (control) rectangles along the right side of the heart map. [Figure 35A] Invariably modified C-RNA transcripts separate early PE samples from controls. Figure 35A shows the fold change between PE and control pregnant women assessed by both sequencing (orange) and qPCR (purple) of 20 transcripts (*p<0.05 by Student's T-test, N=19 for controls and PE). Figure 35B shows tissue expression of disease genes. Figure 35C shows hierarchical clustering of iPEC samples (average linkage, squared Euclidean distance, PE, N=40; control, N=73). Clustering of early (Figure 35D) and late (Figure 35E) PE from PEARL PEC and control pregnancy samples (N=12 for each group). [Figure 35B]Invariably modified C-RNA transcripts separate early PE samples from controls. Figure 35A shows the fold change between PE and control pregnant women assessed by both sequencing (orange) and qPCR (purple) of 20 transcripts (*p<0.05 by Student's T-test, N=19 for controls and PE). Figure 35B shows tissue expression of disease genes. Figure 35C shows hierarchical clustering of iPEC samples (average linkage, squared Euclidean distance, PE, N=40; control, N=73). Clustering of early (Figure 35D) and late (Figure 35E) PE from PEARL PEC and control pregnancy samples (N=12 for each group). [Figure 35C] Invariably modified C-RNA transcripts separate early PE samples from controls. Figure 35A shows the fold change between PE and control pregnant women assessed by both sequencing (orange) and qPCR (purple) of 20 transcripts (*p<0.05 by Student's T-test, N=19 for controls and PE). Figure 35B shows tissue expression of disease genes. Figure 35C shows hierarchical clustering of iPEC samples (average linkage, squared Euclidean distance, PE, N=40; control, N=73). Clustering of early (Figure 35D) and late (Figure 35E) PE from PEARL PEC and control pregnancy samples (N=12 for each group). [Figure 35D] Invariably modified C-RNA transcripts separate early PE samples from controls. Figure 35A shows the fold change between PE and control pregnant women assessed by both sequencing (orange) and qPCR (purple) of 20 transcripts (*p<0.05 by Student's T-test, N=19 for controls and PE). Figure 35B shows tissue expression of disease genes. Figure 35C shows hierarchical clustering of iPEC samples (average linkage, squared Euclidean distance, PE, N=40; control, N=73). Clustering of early (Figure 35D) and late (Figure 35E) PE from PEARL PEC and control pregnancy samples (N=12 for each group). [Figure 35E]Invariably modified C-RNA transcripts separate early PE samples from controls. Figure 35A shows the fold change between PE and control pregnant women assessed by both sequencing (orange) and qPCR (purple) of 20 transcripts (*p<0.05 by Student's T-test, N=19 for controls and PE). Figure 35B shows tissue expression of disease genes. Figure 35C shows hierarchical clustering of iPEC samples (average linkage, squared Euclidean distance, PE, N=40; control, N=73). Clustering of early (Figure 35D) and late (Figure 35E) PE from PEARL PEC and control pregnancy samples (N=12 for each group). [Figure 36A] Machine learning accurately classifies PE across independent cohorts. Figure 36A shows the mean ROC curve for iPEC validation samples (dashed line = SD; N = 10). Figure 36B shows accuracy, sensitivity, and specificity measurements using the iPEC holdout sample and independent PEARL PEC samples (N = 10). Figure 36C shows a heart map of relative transcript abundance in the iPEC cohort for genes used by the AdaBoost model. The graph on the right shows how many cross-validated models a given transcript appeared in. Figure 36D shows the concordance of transcripts identified by differential analysis and AdaBoost. Figure 36E shows tissues expressing high levels of transcripts selected by the AdaBoost model. [Figure 36B]Machine learning accurately classifies PE across independent cohorts. Figure 36A shows the mean ROC curve for iPEC validation samples (dashed line = SD; N = 10). Figure 36B shows accuracy, sensitivity, and specificity measurements using the iPEC holdout sample and independent PEARL PEC samples (N = 10). Figure 36C shows a heart map of relative transcript abundance in the iPEC cohort for genes used by the AdaBoost model. The graph on the right shows how many cross-validated models a given transcript appeared in. Figure 36D shows the concordance of transcripts identified by differential analysis and AdaBoost. Figure 36E shows tissues expressing high levels of transcripts selected by the AdaBoost model. [Figure 36C] Machine learning accurately classifies PE across independent cohorts. Figure 36A shows the mean ROC curve for iPEC validation samples (dashed line = SD; N = 10). Figure 36B shows accuracy, sensitivity, and specificity measurements using the iPEC holdout sample and independent PEARL PEC samples (N = 10). Figure 36C shows a heart map of relative transcript abundance in the iPEC cohort for genes used by the AdaBoost model. The graph on the right shows how many cross-validated models a given transcript appeared in. Figure 36D shows the concordance of transcripts identified by differential analysis and AdaBoost. Figure 36E shows tissues expressing high levels of transcripts selected by the AdaBoost model. [Figure 36D]Machine learning accurately classifies PE across independent cohorts. Figure 36A shows the mean ROC curve for iPEC validation samples (dashed line = SD; N = 10). Figure 36B shows accuracy, sensitivity, and specificity measurements using the iPEC holdout sample and independent PEARL PEC samples (N = 10). Figure 36C shows a heart map of relative transcript abundance in the iPEC cohort for genes used by the AdaBoost model. The graph on the right shows how many cross-validated models a given transcript appeared in. Figure 36D shows the concordance of transcripts identified by differential analysis and AdaBoost. Figure 36E shows tissues expressing high levels of transcripts selected by the AdaBoost model. [Figure 36E] Machine learning accurately classifies PE across independent cohorts. Figure 36A shows the mean ROC curve for iPEC validation samples (dashed line = SD; N = 10). Figure 36B shows accuracy, sensitivity, and specificity measurements using the iPEC holdout sample and independent PEARL PEC samples (N = 10). Figure 36C shows a heart map of relative transcript abundance in the iPEC cohort for genes used by the AdaBoost model. The graph on the right shows how many cross-validated models a given transcript appeared in. Figure 36D shows the concordance of transcripts identified by differential analysis and AdaBoost. Figure 36E shows tissues expressing high levels of transcripts selected by the AdaBoost model. [Figure 37A]C-RNA sample preparation workflow. Figure 37A shows the approach used for sequencing library preparation. Blood is shipped overnight before plasma processing and nucleic acid extraction. cfDNA is digested with DNase, and then cDNA is synthesized from total RNA. Whole-transcriptome enrichment is performed before sequencing. In Figure 37B, three methods were evaluated for a C-RNA transcriptome analysis session. rRNA depletion did not consistently enrich the exonic C-RNA fraction. Many libraries contained numerous unaligned reads. Similarly, rRNA overwhelmed the sequencing dataset if not removed. Enrichment produced libraries with the highest proportion of reads from exonic C-RNA. In all bar graphs, orange indicates reads aligned to the human genome, gray indicates reads aligned to rRNA sequences, and pink indicates reads not aligned to the human genome (including both non-human RNA and low-quality sequences). [Figure 37B] C-RNA sample preparation workflow. Figure 37A shows the approach used for sequencing library preparation. Blood is shipped overnight before plasma processing and nucleic acid extraction. cfDNA is digested with DNase, and then cDNA is synthesized from total RNA. Whole-transcriptome enrichment is performed before sequencing. In Figure 37B, three methods were evaluated for a C-RNA transcriptome analysis session. rRNA depletion did not consistently enrich the exonic C-RNA fraction. Many libraries contained numerous unaligned reads. Similarly, rRNA overwhelmed the sequencing dataset if not removed. Enrichment produced libraries with the highest proportion of reads from exonic C-RNA. In all bar graphs, orange indicates reads aligned to the human genome, gray indicates reads aligned to rRNA sequences, and pink indicates reads not aligned to the human genome (including both non-human RNA and low-quality sequences). [Figure 38A]The effect of plasma volume on C-RNA data quality. Figure 38A shows the C-RNA yield in plasma from 122 samples quantified using the Quant-iT RiboGreen assay (Thermo Fisher). Measurements from 23 samples were below the detection threshold and are excluded from the graph. Bars represent the mean ± SD from two technical replicates. Data from nine independent experiments were used in a meta-analysis to evaluate the effect of plasma input on data quality. Figure 38B shows the noise (biological coefficient of variation, edgeR) calculated from biological replicates within each study. Figure 38C shows the library complexity (combined population, preseq) for each sample. **p<0.01, ***p<0.001 by ANOVA with Tukey's HSD correction using study as the blocking variable. For Figures 38B and 38C, respectively, 0.5 mL, N = 8, 95; 1 mL, N = 7, 83; 2 mL, N = 7, 33; and 4 mL, N = 17, 267. [Figure 38B] The effect of plasma volume on C-RNA data quality. Figure 38A shows the C-RNA yield in plasma from 122 samples quantified using the Quant-iT RiboGreen assay (Thermo Fisher). Measurements from 23 samples were below the detection threshold and are excluded from the graph. Bars represent the mean ± SD from two technical replicates. Data from nine independent experiments were used in a meta-analysis to evaluate the effect of plasma input on data quality. Figure 38B shows the noise (biological coefficient of variation, edgeR) calculated from biological replicates within each study. Figure 38C shows the library complexity (combined population, preseq) for each sample. **p<0.01, ***p<0.001 by ANOVA with Tukey's HSD correction using study as the blocking variable. For Figures 38B and 38C, respectively, 0.5 mL, N = 8, 95; 1 mL, N = 7, 83; 2 mL, N = 7, 33; and 4 mL, N = 17, 267. [Figure 38C]The effect of plasma volume on C-RNA data quality. Figure 38A shows the C-RNA yield in plasma from 122 samples quantified using the Quant-iT RiboGreen assay (Thermo Fisher). Measurements from 23 samples were below the detection threshold and are excluded from the graph. Bars represent the mean ± SD from two technical replicates. Data from nine independent experiments were used in a meta-analysis to evaluate the effect of plasma input on data quality. Figure 38B shows the noise (biological coefficient of variation, edgeR) calculated from biological replicates within each study. Figure 38C shows the library complexity (combined population, preseq) for each sample. **p<0.01, ***p<0.001 by ANOVA with Tukey's HSD correction using study as the blocking variable. For Figures 38B and 38C, respectively, 0.5 mL, N = 8, 95; 1 mL, N = 7, 83; 2 mL, N = 7, 33; and 4 mL, N = 17, 267. [Figure 39A]Integrity of the C-RNA pregnancy signal after storage in different BCTs. Figure 39A shows a heat map of the abundance of known C-RNA pregnancy markers after overnight storage in 4BCT compared to immediate processing from EDTA BCT. Figure 39B shows that the integrated pregnancy signal index obtained by summing the transcript abundances in Figure 39A discriminates between pregnant and non-pregnant samples. **p<0.01, ***p<0.001 by ANOVA with Tukey's HSD correction. For non-pregnant and pregnant groups, respectively, N=4, 8 for EDTA immediate; N=7, 7 for EDTA overnight; N=16, 16 for ACD overnight; N=10, 9 for cell-free RNA overnight; and N=8, 8 for cell-free DNA overnight. Figure 39C shows the correlation of transcriptomic profiles from blood samples collected from the same individual and stored in Cell-Free DNA BCT (Streck, Inc.) for 0, 1, or 5 days before processing. Bars indicate mean ± range; N=2. AdaBoost scores assigned to control samples (Figure 39D) or PE samples from the iPEC cohort (Figure 39E) versus the number of days blood was stored at room temperature before plasma processing. AdaBoost scores are normalized to a range of -1 to +1, with control samples expected to have a score <0 and PE samples expected to have a score >0. No significant differences in AdaBoost scores were observed (ANOVA for control, t-test for PE). Control 1 day, N=60; 2 days, N=4; 3 days, N=1; 5 days, N=2. PE 1 day, N=37; 2 days, N=3. [Figure 39B]Integrity of the C-RNA pregnancy signal after storage in different BCTs. Figure 39A shows a heat map of the abundance of known C-RNA pregnancy markers after overnight storage in 4BCT compared to immediate processing from EDTA BCT. Figure 39B shows that the integrated pregnancy signal index obtained by summing the transcript abundances in Figure 39A discriminates between pregnant and non-pregnant samples. **p<0.01, ***p<0.001 by ANOVA with Tukey's HSD correction. For non-pregnant and pregnant groups, respectively, N=4, 8 for EDTA immediate; N=7, 7 for EDTA overnight; N=16, 16 for ACD overnight; N=10, 9 for cell-free RNA overnight; and N=8, 8 for cell-free DNA overnight. Figure 39C shows the correlation of transcriptomic profiles from blood samples collected from the same individual and stored in Cell-Free DNA BCT (Streck, Inc.) for 0, 1, or 5 days before processing. Bars indicate mean ± range; N=2. AdaBoost scores assigned to control samples (Figure 39D) or PE samples from the iPEC cohort (Figure 39E) versus the number of days blood was stored at room temperature before plasma processing. AdaBoost scores are normalized to a range of -1 to +1, with control samples expected to have a score <0 and PE samples expected to have a score >0. No significant differences in AdaBoost scores were observed (ANOVA for control, t-test for PE). Control 1 day, N=60; 2 days, N=4; 3 days, N=1; 5 days, N=2. PE 1 day, N=37; 2 days, N=3. [Figure 39C]Integrity of the C-RNA pregnancy signal after storage in different BCTs. Figure 39A shows a heat map of the abundance of known C-RNA pregnancy markers after overnight storage in 4BCT compared to immediate processing from EDTA BCT. Figure 39B shows that the integrated pregnancy signal index obtained by summing the transcript abundances in Figure 39A discriminates between pregnant and non-pregnant samples. **p<0.01, ***p<0.001 by ANOVA with Tukey's HSD correction. For non-pregnant and pregnant groups, respectively, N=4, 8 for EDTA immediate; N=7, 7 for EDTA overnight; N=16, 16 for ACD overnight; N=10, 9 for cell-free RNA overnight; and N=8, 8 for cell-free DNA overnight. Figure 39C shows the correlation of transcriptomic profiles from blood samples collected from the same individual and stored in Cell-Free DNA BCT (Streck, Inc.) for 0, 1, or 5 days before processing. Bars indicate mean ± range; N=2. AdaBoost scores assigned to control samples (Figure 39D) or PE samples from the iPEC cohort (Figure 39E) versus the number of days blood was stored at room temperature before plasma processing. AdaBoost scores are normalized to a range of -1 to +1, with control samples expected to have a score <0 and PE samples expected to have a score >0. No significant differences in AdaBoost scores were observed (ANOVA for control, t-test for PE). Control 1 day, N=60; 2 days, N=4; 3 days, N=1; 5 days, N=2. PE 1 day, N=37; 2 days, N=3. [Figure 39D]Integrity of the C-RNA pregnancy signal after storage in different BCTs. Figure 39A shows a heat map of the abundance of known C-RNA pregnancy markers after overnight storage in 4BCT compared to immediate processing from EDTA BCT. Figure 39B shows that the integrated pregnancy signal index obtained by summing the transcript abundances in Figure 39A discriminates between pregnant and non-pregnant samples. **p<0.01, ***p<0.001 by ANOVA with Tukey's HSD correction. For non-pregnant and pregnant groups, respectively, N=4, 8 for EDTA immediate; N=7, 7 for EDTA overnight; N=16, 16 for ACD overnight; N=10, 9 for cell-free RNA overnight; and N=8, 8 for cell-free DNA overnight. Figure 39C shows the correlation of transcriptomic profiles from blood samples collected from the same individual and stored in Cell-Free DNA BCT (Streck, Inc.) for 0, 1, or 5 days before processing. Bars indicate mean ± range; N=2. AdaBoost scores assigned to control samples (Figure 39D) or PE samples from the iPEC cohort (Figure 39E) versus the number of days blood was stored at room temperature before plasma processing. AdaBoost scores are normalized to a range of -1 to +1, with control samples expected to have a score <0 and PE samples expected to have a score >0. No significant differences in AdaBoost scores were observed (ANOVA for control, t-test for PE). Control 1 day, N=60; 2 days, N=4; 3 days, N=1; 5 days, N=2. PE 1 day, N=37; 2 days, N=3. [Figure 39E]Integrity of the C-RNA pregnancy signal after storage in different BCTs. Figure 39A shows a heat map of the abundance of known C-RNA pregnancy markers after overnight storage in 4BCT compared to immediate processing from EDTA BCT. Figure 39B shows that the integrated pregnancy signal index obtained by summing the transcript abundances in Figure 39A discriminates between pregnant and non-pregnant samples. **p<0.01, ***p<0.001 by ANOVA with Tukey's HSD correction. For non-pregnant and pregnant groups, respectively, N=4, 8 for EDTA immediate; N=7, 7 for EDTA overnight; N=16, 16 for ACD overnight; N=10, 9 for cell-free RNA overnight; and N=8, 8 for cell-free DNA overnight. Figure 39C shows the correlation of transcriptomic profiles from blood samples collected from the same individual and stored in Cell-Free DNA BCT (Streck, Inc.) for 0, 1, or 5 days before processing. Bars indicate mean ± range; N=2. AdaBoost scores assigned to control samples (Figure 39D) or PE samples from the iPEC cohort (Figure 39E) versus the number of days blood was stored at room temperature before plasma processing. AdaBoost scores are normalized to a range of -1 to +1, with control samples expected to have a score <0 and PE samples expected to have a score >0. No significant differences in AdaBoost scores were observed (ANOVA for control, t-test for PE). Control 1 day, N=60; 2 days, N=4; 3 days, N=1; 5 days, N=2. PE 1 day, N=37; 2 days, N=3. [Figure 40A]C-RNA transcripts can change during specific stages of pregnancy. Dynamic changes in transcripts that change primarily during early pregnancy (Figure 40A, approximately 14 weeks), throughout pregnancy (Figure 40B), or primarily during late pregnancy (Figure 40C, approximately 33 weeks). Note how transcripts that change during early-mid pregnancy do not return to baseline levels but remain at altered abundance for the remainder of pregnancy. Figure 40D shows ontology and pathway enrichment analysis of transcripts that change during healthy pregnancy. Each filled box indicates significant enrichment for the corresponding term or pathway. "All genes" indicates analysis of all 156 differentially abundant transcripts. Too few genes were altered during pregnancy to perform ontology analysis. [Figure 40B] C-RNA transcripts can change during specific stages of pregnancy. Dynamic changes in transcripts that change primarily during early pregnancy (Figure 40A, approximately 14 weeks), throughout pregnancy (Figure 40B), or primarily during late pregnancy (Figure 40C, approximately 33 weeks). Note how transcripts that change during early-mid pregnancy do not return to baseline levels but remain at altered abundance for the remainder of pregnancy. Figure 40D shows ontology and pathway enrichment analysis of transcripts that change during healthy pregnancy. Each filled box indicates significant enrichment for the corresponding term or pathway. "All genes" indicates analysis of all 156 differentially abundant transcripts. Too few genes were altered during pregnancy to perform ontology analysis. [Figure 40C]C-RNA transcripts can change during specific stages of pregnancy. Dynamic changes in transcripts that change primarily during early pregnancy (Figure 40A, approximately 14 weeks), throughout pregnancy (Figure 40B), or primarily during late pregnancy (Figure 40C, approximately 33 weeks). Note how transcripts that change during early-mid pregnancy do not return to baseline levels but remain at altered abundance for the remainder of pregnancy. Figure 40D shows ontology and pathway enrichment analysis of transcripts that change during healthy pregnancy. Each filled box indicates significant enrichment for the corresponding term or pathway. "All genes" indicates analysis of all 156 differentially abundant transcripts. Too few genes were altered during pregnancy to perform ontology analysis. [Figure 40D] C-RNA transcripts can change during specific stages of pregnancy. Dynamic changes in transcripts that change primarily during early pregnancy (Figure 40A, approximately 14 weeks), throughout pregnancy (Figure 40B), or primarily during late pregnancy (Figure 40C, approximately 33 weeks). Note how transcripts that change during early-mid pregnancy do not return to baseline levels but remain at altered abundance for the remainder of pregnancy. Figure 40D shows ontology and pathway enrichment analysis of transcripts that change during healthy pregnancy. Each filled box indicates significant enrichment for the corresponding term or pathway. "All genes" indicates analysis of all 156 differentially abundant transcripts. Too few genes were altered during pregnancy to perform ontology analysis. [Figure 41A] Comparison of pregnancy-related transcriptional tissue specificity of three independent C-RNA tests. Figure 41A shows the tissue specificity for the complete set of genes detected in each test. Figure 41B shows the transcripts unique to each test. Figure 41C shows the intersecting gene sets. [Figure 41B] Comparison of pregnancy-related transcriptional tissue specificity of three independent C-RNA tests. Figure 41A shows the tissue specificity for the complete set of genes detected in each test. Figure 41B shows the transcripts unique to each test. Figure 41C shows the intersecting gene sets. [Figure 41C]Comparison of pregnancy-related transcriptional tissue specificity of three independent C-RNA tests. Figure 41A shows the tissue specificity for the complete set of genes detected in each test. Figure 41B shows the transcripts unique to each test. Figure 41C shows the intersecting gene sets. [Figure 42A] AdaBoost hyperparameter optimization. Performance measured by Matthews correlation coefficient, pairwise estimates (Fig. 42A) or learning rate (Fig. 42B). Each dot represents the mean value obtained from 3-fold cross-validation during random search. [Figure 42B] AdaBoost hyperparameter optimization. Performance measured by Matthews correlation coefficient, pairwise estimates (Fig. 42A) or learning rate (Fig. 42B). Each dot represents the mean value obtained from 3-fold cross-validation during random search. [Figure 43A] AdaBoost Training Strategy. In Figure 43A, the training samples are divided into five equally sized subsets, corresponding to five iterations of model building. In the first iteration, samples from subset 1 are used for pruning, while samples in subsets 2, 3, 4, and 5 are combined for AdaBoost fitting. In the second iteration, samples from subset 2 are used for pruning, and samples from subsets 1, 3, 4, and 5 are used for AdaBoost fitting, and so on for the remaining third iteration. For each iteration, an AdaBoost model is fitted to the "fitting samples," as shown in Figure 43B. The impact of removing genes below increasing importance thresholds on classification performance is then evaluated in the "pruning samples." The model with the best performance and fewest genes is retained. In Figure 43C, the process of Figure 43B is repeated 10 times for each set of fitting and pruning samples, generating a total of 50 models. In Figure 43D, estimates from all models are aggregated into a single AdaBoost ensemble. Feature pruning of the agglomerative model is performed to identify the minimal gene set required for optimal classification. [Figure 43B]AdaBoost Training Strategy. In Figure 43A, the training samples are divided into five equally sized subsets, corresponding to five iterations of model building. In the first iteration, samples from subset 1 are used for pruning, while samples in subsets 2, 3, 4, and 5 are combined for AdaBoost fitting. In the second iteration, samples from subset 2 are used for pruning, and samples from subsets 1, 3, 4, and 5 are used for AdaBoost fitting, and so on for the remaining third iteration. For each iteration, an AdaBoost model is fitted to the "fitting samples," as shown in Figure 43B. The impact of removing genes below increasing importance thresholds on classification performance is then evaluated in the "pruning samples." The model with the best performance and fewest genes is retained. In Figure 43C, the process of Figure 43B is repeated 10 times for each set of fitting and pruning samples, generating a total of 50 models. In Figure 43D, estimates from all models are aggregated into a single AdaBoost ensemble. Feature pruning of the agglomerative model is performed to identify the minimal gene set required for optimal classification. [Figure 43C]AdaBoost Training Strategy. In Figure 43A, the training samples are divided into five equally sized subsets, corresponding to five iterations of model building. In the first iteration, samples from subset 1 are used for pruning, while samples in subsets 2, 3, 4, and 5 are combined for AdaBoost fitting. In the second iteration, samples from subset 2 are used for pruning, and samples from subsets 1, 3, 4, and 5 are used for AdaBoost fitting, and so on for the remaining third iteration. For each iteration, an AdaBoost model is fitted to the "fitting samples," as shown in Figure 43B. The impact of removing genes below increasing importance thresholds on classification performance is then evaluated in the "pruning samples." The model with the best performance and fewest genes is retained. In Figure 43C, the process of Figure 43B is repeated 10 times for each set of fitting and pruning samples, generating a total of 50 models. In Figure 43D, estimates from all models are aggregated into a single AdaBoost ensemble. Feature pruning of the agglomerative model is performed to identify the minimal gene set required for optimal classification. [Figure 43D]AdaBoost Training Strategy. In Figure 43A, the training samples are divided into five equally sized subsets, corresponding to five iterations of model building. In the first iteration, samples from subset 1 are used for pruning, while samples in subsets 2, 3, 4, and 5 are combined for AdaBoost fitting. In the second iteration, samples from subset 2 are used for pruning, and samples from subsets 1, 3, 4, and 5 are used for AdaBoost fitting, and so on for the remaining third iteration. For each iteration, an AdaBoost model is fitted to the "fitting samples," as shown in Figure 43B. The impact of removing genes below increasing importance thresholds on classification performance is then evaluated in the "pruning samples." The model with the best performance and fewest genes is retained. In Figure 43C, the process of Figure 43B is repeated 10 times for each set of fitting and pruning samples, generating a total of 50 models. In Figure 43D, estimates from all models are aggregated into a single AdaBoost ensemble. Feature pruning of the agglomerative model is performed to identify the minimal gene set required for optimal classification. [Figure 44A]AdaBoost output is significantly affected by sample selection. Figure 44A shows the classification performance (log loss) of individual models generated from each AdaBoost subset (*p<0.05, ***p<0.001 by ANOVA with Tukey's HSD correction, N=10 each). Figure 44B shows the frequency of each transcript included in one of 50 distinct AdaBoost models. Figure 44C shows a Venn diagram of the transcripts incorporated into the models for each training subset. Five transcripts are utilized in the models from all subsets, while 40 are unique to a single subset. Figure 44D shows the impact of the pruning estimator on the classification performance (log loss) of the final, fully aggregated AdaBoost model, showing different behavior for each sample set. These trends are particularly striking considering that 75% of the data used to fit the AdaBoost models was shared between one of the two sample sets used to fit AdaBoost. Note that these data were generated separately from the final machine learning analysis shown in Figure 41. [Figure 44B]AdaBoost output is significantly affected by sample selection. Figure 44A shows the classification performance (log loss) of individual models generated from each AdaBoost subset (*p<0.05, ***p<0.001 by ANOVA with Tukey's HSD correction, N=10 each). Figure 44B shows the frequency of each transcript included in one of 50 distinct AdaBoost models. Figure 44C shows a Venn diagram of the transcripts incorporated into the models for each training subset. Five transcripts are utilized in the models from all subsets, while 40 are unique to a single subset. Figure 44D shows the impact of the pruning estimator on the classification performance (log loss) of the final, fully aggregated AdaBoost model, showing different behavior for each sample set. These trends are particularly striking considering that 75% of the data used to fit the AdaBoost models was shared between one of the two sample sets used to fit AdaBoost. Note that these data were generated separately from the final machine learning analysis shown in Figure 41. [Figure 44C]AdaBoost output is significantly affected by sample selection. Figure 44A shows the classification performance (log loss) of individual models generated from each AdaBoost subset (*p<0.05, ***p<0.001 by ANOVA with Tukey's HSD correction, N=10 each). Figure 44B shows the frequency of each transcript included in one of 50 distinct AdaBoost models. Figure 44C shows a Venn diagram of the transcripts incorporated into the models for each training subset. Five transcripts are utilized in the models from all subsets, while 40 are unique to a single subset. Figure 44D shows the impact of the pruning estimator on the classification performance (log loss) of the final, fully aggregated AdaBoost model, showing different behavior for each sample set. These trends are particularly striking considering that 75% of the data used to fit the AdaBoost models was shared between one of the two sample sets used to fit AdaBoost. Note that these data were generated separately from the final machine learning analysis shown in Figure 41. [Figure 44D]AdaBoost output is significantly affected by sample selection. Figure 44A shows the classification performance (log loss) of individual models generated from each AdaBoost subset (*p<0.05, ***p<0.001 by ANOVA with Tukey's HSD correction, N=10 each). Figure 44B shows the frequency of each transcript included in one of 50 distinct AdaBoost models. Figure 44C shows a Venn diagram of the transcripts incorporated into the models for each training subset. Five transcripts are utilized in the models from all subsets, while 40 are unique to a single subset. Figure 44D shows the impact of the pruning estimator on the classification performance (log loss) of the final, fully aggregated AdaBoost model, showing different behavior for each sample set. These trends are particularly striking considering that 75% of the data used to fit the AdaBoost models was shared between one of the two sample sets used to fit AdaBoost. Note that these data were generated separately from the final machine learning analysis shown in Figure 41.
[0057] Schematic drawings are not necessarily to scale. Like numbers used in the drawings may refer to like components. However, it will be understood that the use of a number to refer to a component in a given figure is not intended to limit the component in another figure labeled with the same number. Furthermore, the use of different numbers to refer to a component is not intended to indicate that the differently numbered component may not be the same as or similar to the other numbered component. DETAILED DESCRIPTION OF THE INVENTION
[0058] Provided herein are signatures of circulating RNAs found in the maternal circulation that are specific for pre-eclampsia, and the use of such signatures in non-invasive methods for diagnosing pre-eclampsia and identifying pregnant women at risk for developing pre-eclampsia.
[0059] While the majority of DNA and RNA in the body is located intracellularly, extracellular nucleic acids are also found circulating freely in the blood. Circulating RNA, also referred to herein as "C-RNA," refers to the extracellular segments of RNA found in the bloodstream. C-RNA molecules primarily originate from two sources: those released into the circulation from dying cells undergoing apoptosis, and those contained within exosomes released by viable cells into the circulation. Exosomes are small membranous vesicles approximately 30–150 nm in diameter that are released into the extracellular space by many cell types and found in various body fluids (including serum, urine, and breast milk) and carry proteins, mRNAs, and microRNAs. The lipid bilayer structure of exosomes protects the RNA contained within them from degradation by RNases and provides stability in the blood. See, e.g., Huang et al., 2013, BMC Genomics;14:319, and Li et al., 2017, Mol Cancer;16:145). There is increasing evidence that exosomes have specialized functions and play a role in processes such as coagulation, cell-cell signaling, and waste management (van der Pol et al., 2012, Pharmacol Rev;64(3):676-705). See also Samos et al., 2006, Ann NY Acad Sci;1075:165-173; Zernecke et al., 2009, Sci Signal;2:ra81; Ma et al., 2012, J Exp Clin Cancer Res;31:38; and Sato-Kuwabara et al., 2015, Int J Oncol;46:17-27.
[0060] In the method described herein, the C-RNA molecules found in the maternal circulatory system serve as biomarkers for fetal, placental and maternal health, and provide the opportunity to know the progress of pregnancy.The C-RNA signatures in the maternal circulatory system that indicate pregnancy, the C-RNA signatures in the maternal circulatory system that are temporally related to gestational age, and the C-RNA signatures in the maternal circulatory system that indicate the pregnancy complication preeclampsia are described herein.
[0061] The maternal circulatory C-RNA signature indicative of preeclampsia includes ARRDC2, JUN, SKIL, ATP13A3, PDE8B, GSTA3, PAPPA2, TIPARP, LEP, RGP1, USP54, CLEC4C, MRPS35, ARHGEF25, CUX2, HEATR9, FSTL3, DDI2, ZMYM6, ST6GALNAC3, GBP2, NES, ETV3, ADAM17, ATOH8, SLC4A3, TRAF3IP1, TTC21A, HEG1, ASTE1, TMEM108, ENC1, SCAMP1, ARRDC3, SLC26A2, SLIT3, CLIC5, TNFRSF21, and PP The C-RNA signature comprises a plurality of C-RNA molecules encoding at least a portion of a plurality of proteins selected from P1R17, TPST1, GATSL2, SPDYE5, HIPK2, MTRNR2L6, CLCN1, GINS4, CRH, C10orf2, TRUB1, PRG2, ACY3, FAR2, CD63, CKAP4, TPCN1, RNF6, THTPA, FOS, PARN, ORAI3, ELMO3, SMPD3, SERPINF1, TMEM11, PSMD11, EBI3, CLEC4M, CCDC151, CPAMD8, CNFN, LILRA4, ADA, C22orf39, PI4KAP1, and ARFGAP3. This C-RNA signature is the Adaboost General signature obtained using the TruSeq library preparation method set forth in Table 1 below, also referred to herein as "List (a)" or "(a)."
[0062] The maternal circulatory C-RNA signature indicative of preeclampsia includes a plurality of C-RNA molecules encoding at least a portion of a plurality of proteins selected from TIMP4, FLG, HTRA4, AMPH, LCN6, CRH, TEAD4, ARMS2, PAPPA2, SEMA3G, ADAMTS1, ALOX15B, SLC9A3R2, TIMP3, IGFBP5, HSPA12B, CLEC4C, KRT5, PRG2, PRX, ARHGEF25, ADAMTS2, DAAM2, FAM107A, LEP, NES, and VSIG4. This C-RNA signature is a bootstrap signature obtained using the TruSeq library preparation method shown in Table 1 below, also referred to herein as "List (b)" or "(b)."
[0063] C-RNA signatures in the maternal circulation indicative of preeclampsia include CYP26B1, IRF6, MYH14, PODXL, PPP1R3C, SH3RF2, TMC7, ZNF366, ADCY1, C6, FAM219A, HAO2, IGIP, IL1R2, NTRK2, SH3PXD2A, SSUH2, SULT2A1, FMO3, FSTL3, GATA5, HTRA1, C8B, H19, MN1, NFE2L1, PRDM16, AP3B2, and EMP. 1, FLNC, STAG3, CPB2, TENC1, RP1L1, A1CF, NPR1, TEK, ERRFI1, ARHGEF15, CD34, RSPO3, ALPK3, SAMD4A, ZCCHC24, LEAP2, MY L2, NRG3, ZBTB16, SERPINA3, AQP7, SRPX, UACA, ANO1, FKBP5, SCN5A, PTPN21, CACNA1C, ERG, SOX17, WWTR1, AIF1L, CA3, HRG, TAT, AQP7P1, ADRA2C, SYNPO, FN1, GPR116, KRT17, AZGP1, BCL6B, KIF1C, CLIC5, GPR4, GJA5, OLAH, C14orf37, ZEB1, JAG2, K IF26A, APOLD1, PNMT, MYOM3, PITPNM3, TIMP4, HTRA4, AMPH, LCN6, CRH, TEAD4, ARMS2, PAPPA2, SEMA3G, ADAMTS1, ALOX15B, The C-RNA signature comprises a plurality of C-RNA molecules encoding at least a portion of a protein selected from SLC9A3R2, TIMP3, IGFBP5, HSPA12B, PRG2, PRX, ARHGEF25, ADAMTS2, DAAM2, FAM107A, LEP, NES, VSIG4, HBG2, CADM2, LAMP5, PTGDR2, NOMO1, NXF3, PLD4, BPIFB3, PACSIN1, CUX2, FLG, CLEC4C, and KRT5. This C-RNA signature is a standard DEX Treat signature obtained using the TruSeq library preparation method shown in Table 1 below, also referred to herein as "List (c)" or "(c)".
[0064] A maternal circulatory C-RNA signature indicative of preeclampsia includes multiple C-RNA molecules encoding at least a portion of a protein selected from VSIG4, ADAMTS2, NES, FAM107A, LEP, DAAM2, ARHGEF25, TIMP3, PRX, ALOX15B, HSPA12B, IGFBP5, CLEC4C, SLC9A3R2, ADAMTS1, SEMA3G, KRT5, AMPH, PRG2, PAPPA2, TEAD4, CRH, PITPNM3, TIMP4, PNMT, ZEB1, APOLD1, PLD4, CUX2, and HTRA4. This C-RNA signature is a jackknife signature obtained using the TruSeq library preparation method shown in Table 1 below, also referred to herein as "List (d)" or "(d)."
[0065] The maternal circulatory C-RNA signature indicative of preeclampsia includes a plurality of C-RNA molecules encoding at least a portion of a protein selected from ADAMTS1, ADAMTS2, ALOX15B, AMPH, ARHGEF25, CELF4, DAAM2, FAM107A, HSPA12B, HTRA4, IGFBP5, KCNA5, KRT5, LCN6, LEP, LRRC26, NES, OLAH, PACSIN1, PAPPA2, PRX, PTGDR2, SEMA3G, SLC9A3R2, TIMP3, and VSIG4. This C-RNA signature is a standard DEX Treat signature obtained using the Nextera Flex for Enrichment library preparation method shown in Table 1 below, also referred to herein as "List (e)" or "(e)."
[0066] A C-RNA signature in the maternal circulation indicative of preeclampsia includes a plurality of C-RNA molecules encoding at least a portion of a protein selected from ADAMTS1, ADAMTS2, ALOX15B, ARHGEF25, CELF4, DAAM2, FAM107A, HTRA4, IGFBP5, KCNA5, KRT5, LCN6, LEP, LRRC26, NES, OLAH, PRX, PTGDR2, SEMA3G, SLC9A3R2, TIMP3, and VSIG4. This C-RNA signature is a jackknife signature obtained using the Nextera Flex Enrichment library preparation method shown in Table 1 below, also referred to herein as "List (f)" or "(f)."
[0067] A C-RNA signature in the maternal circulation indicative of preeclampsia includes multiple C-RNA molecules encoding at least a portion of a protein selected from CLEC4C, ARHGEF25, ADAMTS2, LEP, ARRDC2, SKIL, PAPPA2, VSIG4, ARRDC4, CRH, and NES. This C-RNA signature is an Adaboost Refined TruSeq signature obtained using the TruSeq library preparation method shown in Table 1 below, also referred to herein as "AdaBoost Refined 1," "List (g)," or "(g)."
[0068] In some embodiments, C-RNA signatures in the maternal circulation that are indicative of pre-eclampsia include ADAMTS2, ARHGEF25, ARRDC2, CLEC4C, LEP, PAPPA2, and VSIG4 (also referred to herein as "AdaBoost Refined 2"), ADAMTS2, ARHGEF25, ARRDC2, CLEC4C, LEP, PAPPA2, SKIL, and VSIG4 (also referred to herein as "AdaBoost Refined 3"), ADAMTS2, ARHGEF25, ARRDC4, CLEC4C, LEP, NES, SKIL, and VSIG4 (also referred to herein as "AdaBoost Refined 4"), ADAMTS2, ARHGEF25, ARRDC2, ARRDC4, CLEC4C, CRH, LEP, PAPPA2, SKIL, and VSIG4 (also referred to herein as "AdaBoost Refined 5"). The C-RNA molecules include C-RNA molecules encoding at least a portion of a protein selected from ADAMTS2, ARHGEF25, ARRDC2, CLEC4C, LEP, and SKIL (also referred to herein as "AdaBoost Refined 5"), ADAMTS2, ARHGEF25, ARRDC2, CLEC4C, LEP, and SKIL (also referred to herein as "AdaBoost Refined 6"), or ADAMTS2, ARHGEF25, ARRDC2, ARRDC4, CLEC4C, LEP, PAPPA2, and SKIL (also referred to herein as "AdaBoost Refined 7").
[0069] A C-RNA signature in the maternal circulation indicative of preeclampsia comprises a plurality of C-RNA molecules encoding at least a portion of a protein selected from LEP, PAPPA2, KCNA5, ADAMTS2, MYOM3, ATP13A3, ARHGEF25, ADA, HTRA4, NES, CRH, ACY3, PLD4, SCT, NOX4, PACSIN1, SERPINF1, SKIL, SEMA3G, TIPARP, LRRC26, PHEX, LILRA4, and PER1. This C-RNA signature is an Adaboost Refined Nextera Flex signature obtained using the Nextera Flex for Enrichment library preparation method set forth in Table 1 below, also referred to herein as "List(h)" or "(h)."
[0070] A C-RNA signature in the maternal circulatory system that is indicative of pre-eclampsia includes a plurality of C-RNA molecules that encode at least a portion of a protein selected from any of those shown in Table S9 of Example 7, also referred to herein as "List (i)" or "(i)."
[0071] A C-RNA signature in the maternal circulatory system indicative of preeclampsia, also referred to herein as "List (j)" or "(j)", includes a plurality of C-RNA molecules encoding at least a portion of a protein selected from AKAP2, ARRB1, CPSF7, INO80C, JAG1, MSMP, NR4A2, PLEK, RAP1GAP2, SPEG, TRPS1, UBE2Q1, and ZNF768.
[0072] In some embodiments, a C-RNA signature in the maternal circulatory system that is indicative of pre-eclampsia comprises a plurality of C-RNA molecules encoding at least a portion of a protein selected from any one or more of any of (a), (b), (c), (d), (e), (f), (g), (h), (i), and / or (j), in combination with any one or more of any of (a), (b), (c), (d), (e), (f), (g), (h), (i), and / or (j).
[0073] The examples provided herein illustrate the eight gene lists summarized above that distinguish preeclampsia from control pregnancies. Each was identified by using different analytical methods and / or separate data sets. However, there is a high degree of concordance between many of these gene sets. Identifying a transcript as altered in preeclampsia C-RNA using multiple methods indicates that the transcript has a higher predictive value for classifying this disease. Therefore, the importance of the transcripts identified by all differential expression analyses and all AdaBoost models was combined and ranked. Genes assigned lower ranks are not important or informative, and they may be less robust for classifying preeclampsia across cohorts and sample preparations.
[0074] First, transcripts identified using all differential expression analyses (standard DEX Treat, bootstrap, and jackknife) for both library preparation methods (TruSeq and Nextera Flex for Enrichment) were combined. Table 2 below shows the relative importance of all 125 transcripts identified by the different analysis methods. Transcripts identified across all analysis methods and both library preparations are the most powerful classifiers and are assigned an importance ranking of 1. Transcripts identified by three or more analysis methods and detected in both library preparations were assigned an importance ranking of 2. Transcripts identified by the most stringent analysis method, jackknife, but only by one of the library preparations, were assigned an importance ranking of 3. Transcripts identified by two of the five analysis methods were assigned an importance ranking of 4. Transcripts identified only by the standard DEX Treat method, the broadest and most comprehensive analysis, were assigned the lowest importance ranking of 5.
[0075] The 91 transcripts identified across all AdaBoost models (AdaBoost General and AdaBoost Refined) and both library preparations (Table 3 below) were then combined. When generating refined AdaBoost models for each library preparation, slight variations were observed in the resulting gene sets each time a model was built from the same data. This is a natural consequence of the randomness used by AdaBoost to search through the large whole-exome C-RNA data. To obtain a representative list of genes, model building for refined AdaBoost was performed a minimum of nine separate times, and all genes used by one or more models were reported. The percentage of models that included each transcript is reported in Table 3 (frequency of use by AdaBoost). AdaBoost assigns its own "importance" value to each transcript, which reflects how much the transcript's abundance influences whether a sample is derived from a preeclamptic patient. These AdaBoost importance values were averaged across each refined AdaBoost model that used a given transcript (Table 3, average AdaBoost model importance).
[0076] Transcripts identified across all AdaBoost analyses and library preparations were assigned the highest importance ranking of 1. Transcripts identified in the refined AdaBoost model for a single library preparation method with a frequency greater than 90% used by AdaBoost were assigned an importance ranking of 2. In general, these transcripts have higher AdaBoost model importance, consistent with improved predictive ability. Transcripts identified in the refined AdaBoost model for a single library preparation method, but used by less than 90% of the AdaBoost models, were assigned an importance ranking of 3. Transcripts identified only in the general AdaBoost model for a single library preparation were given the lowest importance ranking of 4.
[0077] Table 2 lists all genes identified by DEX analysis across all analytical methods and library preparations. Rank 1 = transcripts identified across all analytical and library preparation methods. Rank 2 = transcripts identified in both library preparations and 3 / 5 analytical methods. Rank 3 = identified in one library preparation method, the Jackknife method, which is the most stringent analysis. Rank 4 = identified in 2 / 5 analyses. And Rank 5 = identified only in the standard DEX Treat method, which is the least stringent analysis method.
[0078] Table 3 lists all genes identified by Adaboost analysis across both library preparations: Rank 1 = identified in both library preparation methods and the refined adaboost model; Rank 2 = identified in one library preparation method that was present in the refined adaboost model with high model importance and frequency; Rank 3 = identified in one library preparation method that was present in the refined adaboost model with moderate model importance and frequency; and Rank 4 = identified in one library preparation that was absent from the refined adaboost model.
[0079] Table 4 below is a glossary of all of the terms for the various genes listed herein. Information was obtained from the Hugo Gene Nomenclature Committee of the European Bioinformatics Institute.
[0080] [Table 1-1]
[0081] [Table 1-2]
[0082] [Table 1-3]
[0083] Table 1-4
[0084] Table 1-5
[0085] Table 1-6
[0086] Table 2-1
[0087] Table 2-2
[0088] Table 2-3
[0089] Table 2-4
[0090] Table 2-5
[0091] Table 3-1
[0092] Table 3-2
[0093] Table 3-3
[0094] Table 3-4
[0095] Table 4-1
[0096] Table 4-2
[0097] Table 4-3
[0098] Table 4-4
[0099] Table 4-5
[0100] Table 4-6
[0101] Table 4-7
[0102] Table 4-8
[0103] Table 4-9
[0104] The term "multiple" means two or more elements. For example, the term is used herein in reference to several C-RNA molecules that function as a signature indicative of pre-eclampsia.
[0105] A plurality may be any 2, any 3, any 4, any 5, any 6, any 7, any 8, any 9, any 10, any 11, any 12, any 13, any 14, any 15, any 16, any 17, any 18, any 19, any 20, any 21, any 22, any 23, any 24, any 25, any 26, any 27, any 28, any 29, any of the molecules listed herein. 30, any 31, any 32, any 33, any 34, any 35, any 36, any 37, any 38, any 39, any 40, any 41, any 42, any 43, any 44, any 45, any 46, any 47, any 48, any 49, any 50, any 51, any 52, any 53, any 54, any 55, any 56, any 57, any 58, any 59, any 60, any 61, any 62, any 63, any 64, any 65, any 66, any 67, any 68, any 69, any 70, any 71, any 72, any 73, any 74, any 75, any 76, any 77, any 78, any 79, any 80, any 81, any 82, any 83, any 84, any 85, any 86, any 87, any 88, any 89, any 90, any 91, any 92, any 93, any 94, any 95, any 96, any 97, any 98, any 99, any 100, any 101, any 102, any 103, any 104, any 105, any 106, any 107, any 108, any 109, any 110, any 111, any 112, any 113, any 114, any 115, any 116, any 117, any 118, any 119, any 120,The plurality may include any 121 or any 122 of the biomarkers. The plurality may include at least any of the numbers listed above. The plurality may include more than any of the numbers listed above. The plurality may include any of the ranges listed above. In some embodiments, a C-RNA signature indicative of pre-eclampsia includes only one of the biomarkers listed above.
[0106] Identification and / or quantification of one of these C-RNA signatures in a sample obtained from a subject can be used to determine whether the subject is suffering from or at risk of developing pre-eclampsia.
[0107] The sample may be a biological or biosample, including, but not limited to, blood, serum, plasma, sweat, tears, urine, sputum, lymph, saliva, amniotic fluid, tissue biopsy, swab, or smear (e.g., including, but not limited to, a placental tissue sample). In some preferred embodiments, the biological sample is a cell-free plasma sample. The biological sample may be a maternal sample obtained from a pregnant female subject.
[0108] As used herein, the term "subject" refers to human subjects as well as non-human mammalian subjects. Although the examples herein relate to humans and the language primarily relates to humans, the concepts of the present disclosure are applicable to any mammal and are useful in veterinary medicine, animal science fields, research laboratories, and the like.
[0109] The subject may be a pregnant woman, including a pregnant woman at any stage of pregnancy. The stage of pregnancy may be, for example, early pregnancy, mid-pregnancy (including the second half of the second trimester), or late pregnancy (including the first half of the third trimester). The stage of pregnancy may be, for example, before 16 weeks, before 20 weeks, or after 20 weeks. The stage of pregnancy may be, for example, 8-18 weeks, 10-14 weeks, 11-14 weeks, 11-13 weeks, or 12-13 weeks.
[0110] The discovery of cell-free fetal nucleic acids in maternal plasma has opened up new possibilities for non-invasive prenatal diagnosis. Over the last few years, several approaches have been demonstrated to enable such circulating fetal nucleic acids to be used for the prenatal detection of chromosomal aneuploidies. For example, Poon et al.,2000,Clin Chem;1832-4, Poon et al.,2001,Ann NY Acad Sci;945:207-10, Ng et al.,2003,Clin Chem;49(5):727-31, Ng et al.,2003,Proc Natl Acad Sci US A.;100(8):4748-53, Tsui et al.,2004,J Med Genet;41(6):461-7, Go et al.,2004,Clin Chem;50(8):1413-4, Smets et al.,2006,Clin Chim Acta;364(1-2):22-32, Tsui et al. al.,2006,Methods Mol Biol;336:123-34,Purwosunu et al.,2007,Clin Chem;53(3):399-404, Chim et al.,2008,Clin Chem;54(3):482-90, Tsui and Lo,2008,Methods Mol Biol;444:275-89;Lo,2008,Ann NY Acad Sci;1137:140-143, Miura et al. al.,2010,Prenat Diagn;30(9):849-61, Li et al.,2012,Clin Chim Acta;413(5-6):568-76, Williams et al.,2013,Proc Natl Acad Sci USA;110(11):4255-60, Tsui et al.,2014,Clin Chem;60(7):954-62, Tsang Any of the methods described in U.S. Patent Application Publication No. 2014 / 0243212, and U.S. Patent Application Publication No. 2014 / 0243212, may be used in the methods described herein.
[0111] Detection and identification of biomarkers of C-RNA signatures in the maternal circulation that are indicative of pre-eclampsia or the risk of developing pre-eclampsia can involve any of a variety of techniques. For example, biomarkers may be detected in serum by radioimmunoassay, or polymerase chain reaction (PCR) techniques can be used.
[0112] In various embodiments, identifying biomarkers of C-RNA signatures in the maternal circulatory system that are indicative of pre-eclampsia or the risk of developing pre-eclampsia can include sequencing the C-RNA molecules. Any of a number of sequencing techniques can be used, including but not limited to any of a variety of high-throughput sequencing techniques.
[0113] In some embodiments, the C-RNA population in the maternal biological sample can be subjected to enrichment for RNA sequences containing protein-coding sequences prior to sequencing. Any of a variety of platforms available for whole exome enrichment and sequencing may be used, including, but not limited to, the Agilent SureSelect Human All Exon platform (Chen et al., 2015a, Cold Spring Harb Protoc; 2015(7):626-33. doi:10.1101 / pdb.prot083659), the Roche NimbleGen SeqCap EZ Exome Library SR platform (Chen et al., 2015b, Cold Spring Harb Protoc; 2015(7):634-41. doi:10.1101 / pdb.prot084855), or the Illumina TruSeq Exome Enrichment platform (Chen et al., 2015c, Cold Spring Harb Protoc; 2015(7):642-8. doi:10.1101 / pdb.prot084863). See also "TruSeq™ Exome Enrichment Guide," Catalog #FC-930-1012 Part #15013230 Rev. B November 2010 and Illumina's "TruSeq™ RNA Sample Preparation Guide," Catalog #RS-122-9001DOC Part #15026495 Rev. F March 2014.
[0114] In certain embodiments, biomarkers of C-RNA signatures in the maternal circulatory system that indicate pre-eclampsia or the risk of developing pre-eclampsia can be detected and identified using microarray technology.In this method, the polynucleotide sequences of interest are arranged or arrayed on a microchip substrate.The arrayed sequences are then hybridized with maternal biosamples, or purified and / or enriched portions thereof.Microarrays can include various solid supports, including but not limited to beads, microscope slides, glass wafers, gold, silicon, microchips, and other plastic, metal, ceramic, or biological surfaces.Microarray analysis can be performed by commercially available equipment, for example, by using Illumina technology, according to the manufacturer's protocol.
[0115] With respect to the collection, transport, storage, and / or processing of blood samples for preparation of circulating RNA, steps may be taken to stabilize the sample and / or prevent rupture of cell membranes, which would result in the release of cellular RNA into the sample. For example, in some embodiments, blood samples may be collected, transported, and / or stored in tubes with cell and DNA stabilizing properties, such as Streck Cell-Free DNA BCT® blood collection tubes, before being processed into plasma. In some embodiments, blood samples are not exposed to EDTA. See, e.g., Qin et al., 2013, BMC Research Notes; 6:380 and Medina Diaz et al., 2016, PLoS ONE; 11(11):e0166354.
[0116] In some embodiments, the blood sample is processed into plasma within about 24 to about 72 hours of collection, and in some embodiments, within about 24 hours of collection.
[0117] In some embodiments, the blood sample is maintained, stored, and / or transported at room temperature before being processed into plasma.
[0118] In some embodiments, blood samples are maintained, stored, and / or transported without exposure to low temperatures (e.g., on ice) or without freezing before being processed into plasma.
[0119] The present disclosure includes a kit for use in diagnosing preeclampsia and identifying pregnant women at risk of developing preeclampsia.The kit is any product (e.g., package or container) that includes at least one reagent, such as a probe, for specifically detecting the C-RNA signature in the maternal circulatory system described herein, which indicates preeclampsia or the risk of developing preeclampsia.The kit can be promoted, distributed, or sold as a unit for carrying out the method of the present disclosure.
[0120] The use of a preeclampsia-specific circulating RNA signature found in the maternal circulation in a non-invasive method for diagnosing preeclampsia and identifying pregnant women at risk for developing preeclampsia can be combined with appropriate monitoring and medical management. For example, further testing can be indicated. Such testing can include, for example, blood tests to measure liver function, kidney function, and / or platelets and various coagulation proteins; urine analysis to measure protein or creatinine levels; fetal ultrasound to monitor fetal growth, weight, and amniotic fluid; non-stress testing to measure fetal heart rate with fetal movement; and / or biophysical profiling using ultrasound to measure fetal breathing, muscle tone, and movement and amniotic fluid volume. Therapeutic interventions can include, for example, increased frequency of prenatal visits, antihypertensive drugs to reduce blood pressure, corticosteroids, anticonvulsants, bed rest, hospitalization, and / or premature delivery. See, e.g., Townsend et al., 2016 "Current best practice in the management of hypertensive disorders in pregnancy," Integr Blood Press Control; 9:79-94.
[0121] Therapeutic intervention may include administering low-dose aspirin to pregnant women identified as being at risk for developing pre-eclampsia. A recent multicenter, double-blind, controlled trial demonstrated that treatment of women at high risk for early-term pre-eclampsia with low-dose aspirin reduced the incidence of this diagnosis compared with placebo (Rolnik et al., 2017, "Aspirin versus Placebo in Pregnancies at High Risk for Preterm Preeclampsia," N Engl J Med;377(7):613-622). Doses of low-dose aspirin include, but are not limited to, about 50 to about 150 mg per day, about 60 to about 80 mg per day, about 100 mg or more per day, or about 150 mg per day. Administration may begin, for example, at or before 16 weeks of gestation or between 11 and 14 weeks of gestation. Administration may continue until 36 weeks of gestation.
[0122] The present invention is defined in the claims. However, below is provided a non-exhaustive list of non-limiting embodiments. Any one or more of the features of these embodiments may be combined with any one or more features of any other example, embodiment, or aspect described herein.
[0123] Embodiment 1 includes a method of detecting pre-eclampsia and / or determining an increased risk of pre-eclampsia in a pregnant woman, the method comprising: identifying a plurality of circulating RNA (C-RNA) molecules in a biosample obtained from the pregnant woman; Any 1 or more, Any 2 or more, Any 3 or more, Any 4 or more, Any 5 or more, Any 6 or more, Any 7 or more, Any 8 or more, Any 9 or more, Any 10 or more, Any 11 or more, Any 12, Any 13 or more, Any 14 or more, Any 15 or more, Any 16 or more, Any 17 or more, Any 18 or more, Any 19 or more, Any 20 or more, Any 21 or more, Any 22 or more, Any 23 or more, Any 24 or more, Any 25 or more, Any 26 or more, Any 27 or more, Any 28 or more, Any or a plurality of C-RNA molecules encoding at least a portion of a protein selected from 29 or more, any 30 or more, any 31 or more, any 32 or more, any 33 or more, any 34 or more, any 35 or more, any 36 or more, any 37 or more, any 38 or more, any 39 or more, any 40 or more, any 41 or more, any 42 or more, any 43 or more, any 44 or more, any 45 or more, any 46 or more, any 47 or more, any 48 or more, or all 49 of those listed in Table S9 of Example 7; A plurality of C-RNA molecules selected from a plurality of C-RNA molecules encoding at least a portion of a protein selected from any one or more, any two or more, any three or more, any four or more, any five or more, any six or more, any seven or more, any eight or more, any nine or more, any ten or more, any eleven or more, any twelve or more, or all thirteen of AKAP2, ARRB1, CPSF7, INO80C, JAG1, MSMP, NR4A2, PLEK, RAP1GAP2, SPEG, TRPS1, UBE2Q1, and ZNF768, indicate pre-eclampsia and / or an increased risk of pre-eclampsia in a pregnant woman.
[0124] Embodiment 2 includes a method of detecting pre-eclampsia and / or determining an increased risk of pre-eclampsia in a pregnant woman, the method comprising: Obtaining a biosample from a pregnant woman; purifying a population of circulating RNA (C-RNA) molecules from the biosample; identifying protein-coding sequences encoded by C-RNA molecules within the population of purified C-RNA molecules; The protein coding sequence encoded by the C-RNA molecule encoding at least a portion of a protein is Any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 23 or more, any 24 or more, any 25 or more, any 26 or more, any or any 27 or more, any 28 or more, any 29 or more, any 30 or more, any 31 or more, any 32 or more, any 33 or more, any 34 or more, any 35 or more, any 36 or more, any 37 or more, any 38 or more, any 39 or more, any 40 or more, any 41 or more, any 42 or more, any 43 or more, any 44 or more, any 45 or more, any 46 or more, any 47 or more, any 48 or more, or all 49 of those listed in Table S9 of Example 7; Any one or more, any two or more, any three or more, any four or more, any five or more, any six or more, any seven or more, any eight or more, any nine or more, any ten or more, any eleven or more, any twelve or more, or all thirteen of AKAP2, ARRB1, CPSF7, INO80C, JAG1, MSMP, NR4A2, PLEK, RAP1GAP2, SPEG, TRPS1, UBE2Q1, and ZNF768.
[0125] Embodiment 3 includes the method of embodiment 1 or 2, wherein identifying protein-coding sequences encoded by C-RNA molecules in the biosample includes hybridization, reverse transcriptase PCR, microarray chip analysis, or sequencing.
[0126] Embodiment 4 includes the method of embodiment 1 or 2, wherein identifying protein-coding sequences encoded by C-RNA molecules in the biosample includes sequencing.
[0127] Embodiment 4 includes the method of embodiment 4, wherein the sequencing comprises massively parallel sequencing of clonally amplified molecules.
[0128] Embodiment 6 includes the method of embodiment 4 or 5, wherein the sequencing comprises RNA sequencing.
[0129] Embodiment 7 includes the method of any one of embodiments 1 to 6, wherein prior to identifying the protein-coding sequence encoded by the circular RNA (C-RNA molecule): removing intact cells from the biosample; Treating the biosample with deoxynuclease (DNase) to remove cell-free DNA (cfDNA); synthesizing complementary DNA (cDNA) from the c-RNA molecules in the biosample; and / or It further includes enriching cDNA sequences for DNA sequences encoding proteins by exome enrichment.
[0130] Embodiment 8 includes a method of detecting pre-eclampsia and / or determining an increased risk of pre-eclampsia in a pregnant woman, the method comprising: Obtaining biological samples from pregnant women; removing intact cells from the biosample; Treating the biosample with deoxynuclease (DNase) to remove cell-free DNA (cfDNA); synthesizing complementary DNA (cDNA) from RNA molecules in the biosample; Enrichment of cDNA sequences for protein-coding DNA sequences (exome enrichment); Sequencing the resulting enriched cDNA sequences; and identifying protein-coding sequences encoded by the enriched C-RNA molecules; Any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 23 or more, any 24 or more, any 25 or more, any 26 or more, any or any 27 or more, any 28 or more, any 29 or more, any 30 or more, any 31 or more, any 32 or more, any 33 or more, any 34 or more, any 35 or more, any 36 or more, any 37 or more, any 38 or more, any 39 or more, any 40 or more, any 41 or more, any 42 or more, any 43 or more, any 44 or more, any 45 or more, any 46 or more, any 47 or more, any 48 or more, or all 49 of those listed in Table S9 of Example 7; A protein coding sequence encoded by a C-RNA molecule encoding at least a portion of a protein selected from any one or more, any two or more, any three or more, any four or more, any five or more, any six or more, any seven or more, any eight or more, any nine or more, any ten or more, any eleven or more, any twelve or more, or all thirteen of AKAP2, ARRB1, CPSF7, INO80C, JAG1, MSMP, NR4A2, PLEK, RAP1GAP2, SPEG, TRPS1, UBE2Q1, and ZNF768 indicates pre-eclampsia and / or an increased risk of pre-eclampsia in a pregnant woman.
[0131] Embodiment 9 is Obtaining biological samples from pregnant women; removing intact cells from the biosample; Treating the biosample with deoxynuclease (DNase) to remove cell-free DNA (cfDNA); synthesizing complementary DNA (cDNA) from RNA molecules in the biosample; Enrichment of cDNA sequences for protein-coding DNA sequences (exome enrichment); Sequencing the resulting enriched cDNA sequences; and identifying a protein-coding sequence encoded by the enriched C-RNA molecules, The protein coding sequence encoded by the C-RNA molecule is Any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 23 or more, any 24 or more, any 25 or more, any 26 or more, any or any 27 or more, any 28 or more, any 29 or more, any 30 or more, any 31 or more, any 32 or more, any 33 or more, any 34 or more, any 35 or more, any 36 or more, any 37 or more, any 38 or more, any 39 or more, any 40 or more, any 41 or more, any 42 or more, any 43 or more, any 44 or more, any 45 or more, any 46 or more, any 47 or more, any 48 or more, or all 49 of those listed in Table S9 of Example 7; The protein comprises at least a portion of any one or more, any two or more, any three or more, any four or more, any five or more, any six or more, any seven or more, any eight or more, any nine or more, any ten or more, any eleven or more, any twelve or more, or all thirteen of the proteins selected from AKAP2, ARRB1, CPSF7, INO80C, JAG1, MSMP, NR4A2, PLEK, RAP1GAP2, SPEG, TRPS1, UBE2Q1, and ZNF768.
[0132] Embodiment 10 includes the method of any one of embodiments 1 to 9, wherein the biosample comprises plasma.
[0133] Embodiment 11 includes the method of any one of Embodiments 1-10, wherein the biosample is obtained from a pregnant woman at less than 16 weeks' gestational age or less than 20 weeks' gestational age.
[0134] Embodiment 12 includes the method of any one of Embodiments 1-10, wherein the biosample is obtained from a pregnant woman at greater than 20 weeks' gestation.
[0135] Embodiment 13 is a circulating RNA (C-RNA) signature for high risk of pre-eclampsia, any 1 or more, any 2 or more, any 3 or more, any 4 or more, any 5 or more, any 6 or more, any 7 or more, any 8 or more, any 9 or more, any 10 or more, any 11 or more, any 12, any 13 or more, any 14 or more, any 15 or more, any 16 or more, any 17 or more, any 18 or more, any 19 or more, any 20 or more, any 21 or more, any 22 or more, any 23 or more, any 24 or more, any 25 or more or any 26 or more, any 27 or more, any 28 or more, any 29 or more, any 30 or more, any 31 or more, any 32 or more, any 33 or more, any 34 or more, any 35 or more, any 36 or more, any 37 or more, any 38 or more, any 39 or more, any 40 or more, any 41 or more, any 42 or more, any 43 or more, any 44 or more, any 45 or more, any 46 or more, any 47 or more, any 48 or more, or all 49 C-RNA signatures including those listed in Table S9 of Example 7.
[0136] Embodiment 14 includes a circulating RNA (C-RNA) signature for high risk of pre-eclampsia, a C-RNA signature including any one or more, any two or more, any three or more, any four or more, any five or more, any six or more, any seven or more, any eight or more, any nine or more, any ten or more, any eleven or more, any twelve or more, or all thirteen of AKAP2, ARRB1, CPSF7, INO80C, JAG1, MSMP, NR4A2, PLEK, RAP1GAP2, SPEG, TRPS1, UBE2Q1, and ZNF768.
[0137] Embodiment 15 comprises a solid support array comprising a plurality of agents capable of binding and / or identifying the C-RNA signature of embodiment 13 or 14.
[0138] Embodiment 16 comprises a kit comprising a plurality of probes capable of binding and / or identifying the C-RNA signature of embodiment 13 or 14.
[0139] Embodiment 17 comprises a kit comprising a plurality of primers for selectively amplifying the C-RNA signature of embodiment 13 or 14.
[0140] Embodiment 18 includes the method of any one of embodiments 1 to 12, wherein the sample is a blood sample, and the blood sample is collected, transported, and / or stored in a tube with cell and DNA stabilizing properties before the blood sample is processed into plasma.
[0141] Embodiment 19 includes the method of embodiment 18, wherein the tube comprises a Streck Cell-Free DNA BCT® blood collection tube.
[0142] The present invention is illustrated by the following examples, it being understood that the particular examples, materials, amounts, and procedures are to be interpreted broadly in accordance with the scope and spirit of the invention described herein.
[0143] Example Example 1 Pregnancy-specific C-RNA signature The presence of circulating nucleic acids in maternal plasma provides insight into fetal and placental progression and health (Figure 1). Circulating RNA (C-RNA) is detected in the maternal circulation and originates from two main sources. A significant fraction of C-RNA is derived from apoptotic cells, which release C-RNA-containing vesicles into the bloodstream. C-RNA also enters the maternal circulation through the release of active signaling vesicles, such as exosomes and microvesicles, from various cell types. Thus, as shown in Figure 2, C-RNA is composed of both by-products of cell death and active signaling products. Characteristics of C-RNA include its production by a common process, its release from cells throughout the body, and its stability and vesicle-contained nature. This represents the circulating transcriptome, reflecting tissue-specific changes in gene expression, signaling, and cell death.
[0144] C-RNA has the potential to be a good biomarker for at least the following reasons: 1) All C-RNA is fairly stable in the blood because it is contained within membrane-bound vesicles that protect the C-RNA from degradation. 2) C-RNA originates from all cell types. For example, C-RNA has been shown to contain transcripts from both the placenta and the developing fetus. The diverse origins of C-RNA make it a potential treasure trove for obtaining information about both fetal and overall maternal health.
[0145] C-RNA libraries were prepared from plasma samples using standard Illumina library preparation and whole-exome enrichment techniques, as shown in Figure 3. Specifically, Illumina TruSeq™ library preparation and RNA Access Enrichment were used. Using this approach, libraries with 90% of reads aligning to human coding regions were prepared (Figures 3 and 7). Samples were downsampled to 50M reads, and ≥40M mapped reads were used for downstream analysis. Samples were processed using the C-RNA workflow shown in Figure 3. Dual-indexed libraries were sequenced at 50x50 on a Hiseq2000.
[0146] As shown in Figure 4, comparing the results of plasma samples from third-trimester pregnant women with plasma samples from non-pregnant women reveals a clear pregnancy-specific signature. The top 20 differentially abundant genes in this signature are CSHL1, CSH2, KISS1, CGA, PLAC4, PSG1, GH2, PSG3, PSG4, PSG7, PSG11, CSH1, PSG2, HSD3B1, GRHL2, LGALS14, FCGR1C, PSG5, LGALS13, and GCM1. The majority of the genes identified in the pregnancy signature are expressed in the placenta, correlating with published data. These results also confirm that placental RNA can be obtained in the maternal circulation.
[0147] Example 2 C-RNA signature across gestational age This example characterizes the C-RNA signature across different gestational ages during pregnancy.It is expected that the longitudinal change of C-RNA signature at different time points throughout pregnancy will be less than the difference between the C-RNA signatures of pregnant and non-pregnant samples shown in Example 1.As shown in Figure 5, a clear time course of the C-RNA profile of signature genes is observed as pregnancy progresses, with a clear group of genes that are up-regulated in early pregnancy and a clear group of genes that increase in late pregnancy.
[0148] These genes included CGB8, CGB5, ZSCAN23, HSPA1A, PMAIP1, C8orf4, ITM2B, IFIT2, CD74, HSPA6, TFAP2A, TRPV6, EXPH5, CAPN6, ALDH3B2, RAB3B, MUC15, GSTA3, GRHL2, and CSHL1, as listed in Figure 5.
[0149] These genes may also include CSHL1, CSH2, KISS1, CGA, PLAC4, PSG1, GH2, PSG3, PSG4, PSG7, PSG11, CSH1, PSG2, HSD3B1, GRHL2, LGALS14, FCGR1C, PSG5, LGALS13, and GCM1.
[0150] These changes throughout pregnancy correlate with published data from both Steve Quake and Dennis Lo. See, e.g., Maron et al., 2007, "Gene expression analysis in pregnant women and their infants identifies unique fetal biomarkers that circulate in maternal blood," J Clin Invest; 117(10):3007-3019; Koh et al., 2014, "Noninvasive in vivo monitoring of tissue-specific global gene expression in humans," Proc Natl Acad Sci USA; 111(20):7361-6; and Ngo et al., 2018, "Noninvasive blood tests for fetal development predict gestational age and preterm delivery," Science; 360(6393):1133-1136. C-RNA signatures correlated with patterns of placental gene expression. This technique can therefore detect subtle changes during pregnancy and provides a non-invasive means to monitor placental health.
[0151] Example 3 C-RNA signature of preeclampsia In this example, a C-RNA signature specific to preeclampsia was identified. The C-RNA signature was determined and assayed in samples collected from pregnant women diagnosed with preeclampsia from two studies: the RGH14 study (registered at clinicaltrials.gov under NCT0208494) and the Pearl study (also referred to herein as Pearl Biobank under NCT02379832) (FIG. 6). Two tubes of blood were collected at the time of diagnosis of preeclampsia. To minimize transcriptional variability unrelated to the pathology of preeclampsia and control for gestational age differences in the C-RNA signature, 80 gestational age-matched control samples were collected. Using samples from the RGH14 study, a set of biologically relevant genes was identified, and the predictive value of these biomarkers was validated in an independent cohort of samples from Pearl Biobank.
[0152] In analyzing the RGH14 data, C-RNA signatures specific to preeclampsia (PE) were identified using four different methods: TREAT, bootstrap, jackknife, and Adaboost. Example 3 focuses on the first three analytical methods, and Example 4 focuses on Adaboost.
[0153] The t-test against threshold (TREAT) statistical method, utilizing the EDGR program, allows researchers to formally test (with associated p-values) whether expression differences in microarray experiments are greater than a given (biologically significant) threshold. For a more detailed description of the TREAT statistical method, see McCarthy and Smyth, 2009, "Testing significance relative to a fold-change threshold is a TREAT," Bioinformatics; 25(6): 765-71; and for a more detailed description of the EDGR program, see Robinson et al., 2010, "edgeR: a Bioconductor package for differential expression analysis of digital gene expression data," Bioinformatics; 26: 139-140. For a more detailed description of the Adaboost method, see Freund and Schapire, 1997, "A Decision-Theoretic Generalization of On-Line Learning and an Application to Boosting", Journal of Computer and Systems Sciences; 55(1): 119-139 and Pedregosa et al., 2011, "Scikit-learn: Machine Learning in Python", JMLR; 12: 2825-2830. The Adaboost method is described in Example 4.
[0154] In the first method, standard statistical tests (TREAT method) were used to identify genes that were statistically different in the RGH14 preeclampsia cohort of 40 patients compared to the subset of matched controls (40 patients). 122 genes were identified as statistically different in the preeclampsia cohort (40 patients) compared to the subset of matched controls (40 patients) (Figure 8, right panel). These genes include CYP26B1, IRF6, MYH14, PODXL, PPP1R3C, SH3RF2, TMC7, ZNF366, ADCY1, C6, FAM219A, HAO2, IGIP, IL1R2, NTRK2, SH3PXD2A, SSUH2, SULT2A1, FMO3, FSTL3, GATA5, HTRA1, C8B, H19, MN1, NFE2L1, PRDM16, AP3B2, EMP1, and FLN C, STAG3, CPB2, TENC1, RP1L1, A1CF, NPR1, TEK, ERRFI1, ARHGEF15, CD34, RSPO3, ALPK3, SAMD4A, ZCCHC24, LEAP2, MYL2, NRG3, ZBTB16, SERPINA3, AQP7, SRPX, UACA, ANO1, FKBP5, SCN5A, PTPN21, CACNA1C, ERG, SOX17, WWTR1, AIF1L , CA3, HRG, TAT, AQP7P1, ADRA2C, SYNPO, FN1, GPR116, KRT17, AZGP1, BCL6B, KIF1C, CLIC5, GPR4, GJA5, OLAH, C14o rf37, ZEB1, JAG2, KIF26A, APOLD1, PNMT, MYOM3, PITPNM3, TIMP4, HTRA4, AMPH, LCN6, CRH, TEAD4, ARMS2, PAPPA2, S EMA3G, ADAMTS1, ALOX15B, SLC9A3R2, TIMP3, IGFBP5, HSPA12B, PRG2, PRX, ARHGEF25, ADAMTS2, DAAM2, FAM107A, L Including EP, NES, VSIG4, HBG2, CADM2, LAMP5, PTGDR2, NOMO1, NXF3, PLD4, BPIFB3, PACSIN1, CUX2, FLG, CLEC4C, and KRT5.
[0155] The TREAT method did not identify a set of genes that accurately classified preeclampsia patients into distinct groups with 100% accuracy (Figure 15). However, focusing on these identified genes improved classification compared to using the entire dataset of all measured genes (Figure 8, left panel). This highlights the value of focusing on a subset of genes for prediction. However, with the TREAT method, a significant amount of variation was observed in the identified genes depending on the choice of control. To address this biological variability and further improve the predictive value of the gene list, a second bootstrap method was developed.
[0156] The RGH14 study had more control samples (80) available than preeclampsia patient samples (40). Therefore, we compared the RGH14 cohort of 40 preeclampsia patient samples with 40 randomly selected control samples (matched for gestational age) to identify a list of genes that were statistically different in the preeclampsia cohort. As shown in Figure 9, this was then repeated 1,000 times to identify the frequency with which a set of genes was identified. A significant subset of genes appeared in fewer than 10 of the 1,000 repeats (less than 1% of the 1,000 repeats). These low-frequency genes are most likely due to biological noise and may not reflect genes broadly specific to preeclampsia. Therefore, we further narrowed the gene list selection by requiring that a gene be considered statistically different in the preeclampsia cohort only if it was identified in 50% of the 1,000 repeats performed (Figure 9, right panel). As shown in Figure 10, differential transcript abundance with additional bootstrap selection distinguishes pre-eclampsia samples from healthy controls. Using this additional requirement helped address biological variability and further improved our ability to accurately classify pre-eclampsia samples.
[0157] Using this bootstrap method, 27 genes were identified as statistically associated with preeclampsia. These genes include TIMP4, FLG, HTRA4, AMPH, LCN6, CRH, TEAD4, ARMS2, PAPPA2, SEMA3G, ADAMTS1, ALOX15B, SLC9A3R2, TIMP3, IGFBP5, HSPA12B, CLEC4C, KRT5, PRG2, PRX, ARHGEF25, ADAMTS2, DAAM2, FAM107A, LEP, NES, and VSIG4. The genes identified by this bootstrap method had excellent agreement with published data. Approximately 75% of these genes are expressed by the placenta. As shown in Figure 11, there is overlap with known markers of preeclampsia, including PAPPA and CRH. Furthermore, a significant number of these genes are involved in embryonic development, extracellular matrix remodeling, immune regulation, and cardiovascular function, all pathways known to be dysregulated in pre-eclampsia.
[0158] A third jackknife method was also developed to capture a subset of genes with the highest predictive value. This approach is similar to the bootstrap method. Patients from both the preeclampsia and control groups were randomly subsampled to identify differentially abundant genes 1,000 times. Instead of using the frequency at which genes were identified as statistically different, the jackknife method calculated a confidence interval (95%, one-sided) for the p-value of each transcript. Genes with this confidence interval exceeding 0.05 were excluded (Figure 16, left panel).
[0159] Using the jackknife method, 30 genes were identified as predictive of preeclampsia: VSIG4, ADAMTS2, NES, FAM107A, LEP, DAAM2, ARHGEF25, TIMP3, PRX, ALOX15B, HSPA12B, IGFBP5, CLEC4C, SLC9A3R2, ADAMTS1, SEMA3G, KRT5, AMPH, PRG2, PAPPA2, TEAD4, CRH, PITPNM3, TIMP4, PNMT, ZEB1, APOLD1, PLD4, CUX2, and HTRA4.
[0160] As shown in the right panel of Figure 16, this approach gave good classification of pre-eclampsia patients in the RGH14 dataset (compare Figure 15 (TREAT), Figure 10 (bootstrap method), and Figure 16 (jackknife method)). Each identified gene list was also used to classify pre-eclampsia samples in an independent Pearl Biobank dataset. As shown in Figure 17, each gene list was able to classify pre-eclampsia samples.
[0161] All genes identified by the bootstrap and jackknife methods are represented in the 122 TREAT genes (Table 2, DEX analysis, TruSeq library preparation method). The bootstrap and jackknife gene lists are highly concordant, with over 70% of genes in common. Approximately 90% of transcripts identified by either method showed increased transcript abundance in preeclamptic patients, consistent with increased signaling and / or cell death in this disease.
[0162] Example 4 Identification of C-RNA signatures using Adaboost In this example, we used an alternative approach, a published machine learning algorithm called adaboost, to identify specific C-RNA signatures associated with preeclampsia. As shown in Figure 12, this approach identifies a set of genes with the highest predictive ability for classifying samples as preeclamptic (PE) or normal. Using this gene list, we observed the most significant separation of the preeclamptic cohort from healthy controls. However, this approach can also be highly prone to overfitting the samples used to build the model. Therefore, we validated the prediction model using a completely independent dataset from the PEARL study (Figure 13). Using this Adaboost gene list, we correctly classified 85% of the preeclamptic samples with 85% specificity (Figure 14). Overall, the Adaboost machine learning approach constructed the most accurate predictive model for preeclampsia.
[0163] Using the AdaBoost method, 75 genes were identified as statistically associated with preeclampsia (Table 3, AdaBoost analysis, TruSeq library preparation methods). These genes include ARRDC2, JUN, SKIL, ATP13A3, PDE8B, GSTA3, PAPPA2, TIPARP, LEP, RGP1, USP54, CLEC4C, MRPS35, ARHGEF25, CUX2, HEATR9, FSTL3, DDI2, ZMYM6, ST6GALNAC3, GBP2, NES, ETV3, ADAM17, ATOH8, SLC4A3, TRAF3IP1, TTC21A, HEG1, ASTE1, TMEM108, ENC1, SCAMP1, ARRDC3, SLC26A2, SLIT3, CLIC5, and T These include NFRSF21, PPP1R17, TPST1, GATSL2, SPDYE5, HIPK2, MTRNR2L6, CLCN1, GINS4, CRH, C10orf2, TRUB1, PRG2, ACY3, FAR2, CD63, CKAP4, TPCN1, RNF6, THTPA, FOS, PARN, ORAI3, ELMO3, SMPD3, SERPINF1, TMEM11, PSMD11, EBI3, CLEC4M, CCDC151, CPAMD8, CNFN, LILRA4, ADA, C22orf39, PI4KAP1, and ARFGAP3.
[0164] A refined AdaBoost model was also developed for robust classification of PE samples. To generate a generalized machine learning model that can accurately predict new samples, we used a rigorous approach that avoided overfitting with a single dataset and validated the final classifier using samples not used in model construction. As shown in Figure 18, the RGH14 dataset was randomly divided into six pieces: a holdout subset containing 12% of the samples excluded from model construction, and five uniformly sized test subsets. For each iteration, a subset was designated as training data or test samples. This process, starting with building an AdaBoost model, was repeated a minimum of 10 times for the data subsets. After building 50 high-performing models for the five test-training subsets, the estimators from all models were combined into a single AdaBoost model.
[0165] Using the refined AdaBoost model, 11 genes were identified as being statically associated with preeclampsia. These genes include CLEC4C, ARHGEF25, ADAMTS2, LEP, ARRDC2, SKIL, PAPPA2, VSIG4, ARRDC4, CRH, and NES. The performance of this predictive model was validated using holdout datasets from RGH14 and a completely independent Pearl Biobank cohort (Figure 19).
[0166] Description of AdaBoost Model Generation. The AdaBoost classification method was refined to obtain more specific gene sets (AdaBoost Refined 1-7) by the following procedure, also shown in Figure 18. The RGH14 dataset was divided into six pieces by random selection: a holdout subset containing 12% of the samples excluded from model building, and five uniformly sized test subsets.
[0167] For each test subset, training data was assigned as all samples that were neither holdout nor test samples. Gene counts for the test and training samples were TMM normalized in edgeR, and then the training data was standardized to have a mean of 0 and a standard deviation of 1 for each gene. An AdaBoost model with 90 estimators and a learning rate of 1.6 was then fitted to the training data. Feature pruning was then performed by determining the feature importance of each gene in the model and testing the effect of removing estimators with genes with importance below the threshold. The threshold that yielded the best performance (as measured by the Matthews correlation coefficient for test data classification) with the fewest genes was selected, and that model was retained. This process, starting with building an AdaBoost model, was repeated a minimum of 10 times for this data subset.
[0168] After building all 50+ models for the five test-training subsets, the estimators from all models were combined into a single AdaBoost model. Feature pruning was performed again, this time using the percentage of models incorporating genes for the threshold, and performance was evaluated using the average negative log-loss value for classification of each test subset. The model with the fewest genes and the highest negative log-loss value was selected as the final AdaBoost model.
[0169] AdaBoost Gene List. Through this iteration of the process, slight variations were observed in the genes selected for the final model due to the inherent randomization in the AdaBoost algorithm implementation, but performance for predicting the test data, holdout data, and an independent (Pearl) data set remained high.
[0170] A total of 11 genes were observed in at least one of the 14 AdaBoost Refined models generated: ADAMTS2, ARHGEF25, ARRDC2, ARRDC4, CLEC4C, CRH, LEP, NES, PAPPA2, SKIL, VSIG4 (AdaBoost Refined 1), but no model was generated that simultaneously included all of them.
[0171] Two observed gene sets gave the best performance for classification of the independent data: AdaBoost Refined 2: ADAMTS2, ARHGEF25, ARRDC2, CLEC4C, LEP, PAPPA2, VSIG4 and AdaBoost Refined 3: ADAMTS2, ARHGEF25, ARRDC2, CLEC4C, LEP, PAPPA2, SKIL, VSIG4.
[0172] Four additional gene sets performed nearly as well as AdaBoost Refined 2-3: AdaBoost Refined 4: ADAMTS2, ARHGEF25, ARRDC4, CLEC4C, LEP, NES, SKIL, VSIG4; AdaBoost Refined 5: ADAMTS2, ARHGEF25, ARRDC2, ARRDC4, CLEC4C, CRH, LEP, PAPPA2, SKIL, VSIG4; AdaBoost Refined 6: ADAMTS2, ARHGEF25, ARRDC2, CLEC4C, LEP, SKIL; and AdaBoost Refined 7: ADAMTS2, ARHGEF25, ARRDC2, ARRDC4, CLEC4C, LEP, PAPPA2, SKIL.
[0173] Example 5 Identification of C-RNA signatures using transposome-based library preparation RGH14 sample was also processed by Illumina Nextera Flex for Enrichment protocol, enriched for whole exome, and sequenced to over 40 million reads.This method is more sensitive and robust at low input, and therefore has a high possibility of identifying additional genes that predict preeclampsia.This data set was run by three analysis methods: standard differential expression analysis (TREAT), jackknife method, and refined Adaboost model.For detailed description of these analysis methods, please refer to Example 3 and Example 4.
[0174] By changing the library preparation method, the genes detected in all three analysis methods changed. For the TREAT method, 26 genes were identified as differentially enriched in preeclampsia, most of which also showed increased abundance in preeclampsia (see Table 2, DEX analysis, Nextera Flex for Enrichment library preparation method). These genes include ADAMTS1, ADAMTS2, ALOX15B, AMPH, ARHGEF25, CELF4, DAAM2, FAM107A, HSPA12B, HTRA4, IGFBP5, KCNA5, KRT5, LCN6, LEP, LRRC26, NES, OLAH, PACSIN1, PAPPA2, PRX, PTGDR2, SEMA3G, SLC9A3R2, TIMP3, and VSIG4. Figure 20 shows the classification of RGH14 samples within this gene list.
[0175] By applying the jackknife analysis method, the TREAT list was narrowed to 22 genes identified as differentially enriched in preeclampsia. These genes included ADAMTS1, ADAMTS2, ALOX15B, ARHGEF25, CELF4, DAAM2, FAM107A, HTRA4, IGFBP5, KCNA5, KRT5, LCN6, LEP, LRRC26, NES, OLAH, PRX, PTGDR2, SEMA3G, SLC9A3R2, TIMP3, and VSIG4. The improved performance of this list is shown in Figure 20.
[0176] A refined AdaBoost model approach was applied to this data, as described in Example 4. Using this method, 24 genes are identified as statistically associated with preeclampsia (Table 3, AdaBoost analysis, Nextera Flex for Enrichment library preparation method). These genes include LEP, PAPPA2, KCNA5, ADAMTS2, MYOM3, ATP13A3, ARHGEF25, ADA, HTRA4, NES, CRH, ACY3, PLD4, SCT, NOX4, PACSIN1, SERPINF1, SKIL, SEMAG3, TIPARP, LRRC26, PHEX, LILRA4, and PER1. The predictive model is shown in Figure 21.
[0177] Example 6 Circulating transcriptome measurements from maternal blood detect early-stage preeclampsia signatures Molecular tools for noninvasively monitoring maternal health from conception to delivery would enable accurate detection of pregnant women at risk for adverse outcomes. Circulating RNA (c-RNA) is released into the bloodstream by all tissues, providing an accessible, comprehensive measure of placental, fetal, and maternal health (Koh et al., 2014, Proceedings of the National Academy of Sciences; 111:7361-7366; and Tsui et al., 2014, Clinical Chemistry; 60:954-962). Preeclampsia (PE), a common and potentially life-threatening pregnancy complication, originates in the placenta but spreads to a significant maternal component as the disease progresses (Staff et al., 2013, Hypertension; 61:932-942; and Chaiworapongsa et al., 2014, Nature Reviews Nephrology; 10:466-480). Furthermore, biomarkers have shown limited clinical utility (Poon and Nicolaides, 2014, Obstetrics and Gynecology International; 2014:1-11; Zeisler et al., 2016, N Engl J Med; 374:13-22; and Duhig et al., 2018, F1000 Research; 7:242). Assuming that characterization of the circulating transcriptome could identify better biomarkers, we analyzed C-RNAs from 113 pregnant women (40 at the time of early PE diagnosis). Using a novel workflow, we identified differential abundances of 30 transcripts consistent with PE biology and representing placental, fetal, and maternal contributions. Furthermore, we developed a machine learning model and demonstrated that only seven C-RNA transcripts were required to classify PE in two independent cohorts (92-98% accuracy). The global measurement of C-RNA disclosed in this example highlights its usefulness in monitoring both maternal and fetal health and holds great promise for the diagnosis and prediction of at-risk pregnancies.
[0178] Several studies have been initiated to investigate and identify potential biomarkers in C-RNA for various pregnancy complications (Pan et al., 2017, Clinical Chemistry; 63:1695-1704; Whitehead et al., 2016, Prenatal Diagnosis; 36:997-1008; Tsang et al., 2017, Proc Natl Acad Sci USA; 114:E7786-E7795; and Ngo et al., 2018, Science; 360:1133-1136). However, these studies have involved only a small number of patients and are limited to monitoring a small number of genes—mostly placental and fetal transcripts. Measuring the entire circulating transcriptome is challenging because it requires specific upfront sample collection and processing to minimize variability and contamination from cell lysis (Chiu et al., 2001, Clinical Chemistry; 47:1607-1613, and Page et al., 2013, PLoS ONE; 8:e77963). This complex workflow makes large-scale clinical sample collection difficult because the work required for immediate processing of blood samples is not feasible in many clinics (Marton and Weiner, 2013, BioMed Research International; 2013:891391). Therefore, in this example, a method was established that allows overnight transport of blood to a processing laboratory, where all steps of sample preparation are performed in a controlled environment, providing a scalable platform for clinical trial-level evaluation (Figure 22A).
[0179] A key feature of this method is the ability to transport blood overnight to a processing laboratory. C-RNA pregnancy signals were assessed after overnight room-temperature transport in several tube types (Figures 26A-26C). Blood stored in EDTA tubes, the gold standard used in previous C-RNA studies, showed reduced abundance of pregnancy-associated transcripts and overall instability of the transcriptomic profile (Qin et al., 2013, BMC Research Notes; 6:380). In contrast, Cell-Free DNA BCT (Streck), the primary tube type used for noninvasive prenatal testing (NIPT), retained signals from placental transcripts and had improved technical reproducibility (Figure 26B) (Medina Diaz et al., 2016, PLoS ONE; 11:e0166354).
[0180] Blood transport allowed for easy acquisition of an average of 5 mL of plasma per patient from a single tube of blood. The difference in C-RNA data quality was evaluated when using variable plasma volumes, and it was determined that using less than 2 mL of plasma significantly increased noise and reduced library complexity (Figures 27A and 27B). Therefore, 4 mL of plasma was used in this example study to maximize confidence in data quality.
[0181] We demonstrated this novel workflow by repeating a previous study monitoring the C-RNA dynamics of over 10,000 transcripts per healthy pregnant woman from early to late pregnancy. Using 152 consecutively collected samples from 45 healthy pregnant women (Pre-Eclampsia and Growth Restriction Longitudinal Study Control Cohort-PEARL; NCT02379832; Table 5), we identified 156 significantly altered transcripts, the majority of which increased in abundance as pregnancy progressed (Figure 22B). 42% of the altered genes were identified in previous C-RNA studies (Figure 22C) (Koh et al., 2014, Proceedings of the National Academy of Sciences; 111:7361-7366, and Tsui et al., 2014, Clinical Chemistry; 60:954-962). Of the 91 transcripts identified in this study alone, 64% are expressed by placental and / or fetal tissues (Figures 22D and 28A-28C). The remaining genes likely reflect maternal responses to pregnancy.
[0182] Test Design For the next phase of the study, the workflow was applied to clinical samples to measure changes in C-RNA in PE (iPC, Illumina Preeclampsia Cohort). PE is a heterogeneous disorder associated with different severity and patient outcomes based on whether it presents before (early) or after (late) 34 weeks of gestation (Staff et al., 2013, Hypertension; 61:932-942; Chaiworapongsa et al., 2014, Nature Reviews Nephrology; 10:466-4803; and Dadelszen et al., 2003, Hypertension in Pregnancy; 22:143-148). This study focused on the more severe, early-stage form of PE and defined strict diagnostic criteria with clear inclusion and exclusion criteria (most importantly, excluding individuals with a history of chronic hypertension) to obtain a complete cohort (Table 6) (Nakanishi et al., 2017, Pregnancy Hypertension;7:39-43, and Hiltunen et al., 2017, PLoS ONE;12:e0187729). Maternal characteristics, pregnancy outcomes, and current medications were recorded throughout the study period (Table 7). 113 samples were collected at eight centers (Table 8), including 40 at the time of PE diagnosis and 73 within one week of diagnosis as gestational age-matched controls (Figure 23A). All but one individual with PE delivered preterm, in contrast to 9.5% of controls, confirming that these diagnostic criteria identify individuals severely affected by this disease (Figure 23C).
[0183] All samples were randomly distributed across multiple processing batches and then sequenced to >40M reads. Standard differential expression analysis using the full cohort identified 42 altered transcripts, 37 of which were increased in PE (Figure 24A, blue and orange). However, of concern was the high variability observed in the genes detected as altered when different subsets of controls were selected for analysis.
[0184] To address this discrepancy, a jackknife approach was implemented, allowing for the identification of the most consistently altered genes (Figures 24A and 24B, orange). By repeating the differential analysis 1,000 times using randomly selected sample subsets, we were able to construct confidence intervals for the p-values associated with each putatively altered transcript (Figure 29A). We excluded 12 genes with confidence intervals exceeding 0.05 (Figure 24B). These genes were not excluded by simply setting thresholds for baseline abundance or biological variance (Figure 29B), yet these transcripts were observed to have lower predictive value (Figure 29C). Hierarchical clustering indicates that these genes are not universally altered in the PE cohort and therefore lack the sensitivity (73%) to accurately classify this condition (Figure 29D).
[0185] Analysis then focused on a refined set of 30 genes, 60% of which had previously been associated with PE (Namli et al., 2018, Hypertension in Pregnancy; 37:9-17; Than et al., 2018, Frontiers in Immunology; 9:1661; Kramer et al., 2016, Placenta; 37:19-25; Winn et al., 2008, Endocrinology; 150:452-462; and Liu et al., 2018, Molecular Medicine Reports; 18:2937-2944). qPCR analysis confirmed that 19 of the 20 genes were significantly altered in PE (Figure 24C, Table 9). Importantly, 40% of these genes encode extracellular or secreted protein products. In addition, almost all genes are involved in PE-related processes, including extracellular matrix (ECM) remodeling, gestational age, placental / fetal development, angiogenesis, and hypoxic response (Table 10). 67% of these transcripts were expressed by the placenta and / or fetus (Figure 24D). Vascular and immune functions were well represented among the remaining maternal-expressed transcripts (Table 10). Hierarchical clustering of these genes effectively separated PE and control samples with 98% sensitivity and 97% specificity (Figure 24E). Interestingly, clinical data from two misidentified controls indicated potential confounding health issues, as suggested by the use of hypertension medications (Table 7).
[0186] We evaluated the ability of iPC-identified genes to cluster a cohort of samples from an independent biobank—the Pre-Eclampsia and Growth Restriction Longitudinal Study (PEARL; NCT02379832; Figures 23B and 23C, Table 11). This cohort consisted of early-onset (diagnosed at <34 weeks) and late-onset PE cases and gestational age-matched controls. Early-onset PE samples clustered separately from matched controls with 83% sensitivity and 92% specificity, further validating the association of these transcripts (Figure 24F). In contrast, no clustering was observed for late-onset PE and matched control samples (Figure 24G).
[0187] We then used the iPC data to build an AdaBoost model for robust classification of PE samples. To generate a generalized machine learning model that could accurately predict new samples, we used a rigorous approach that avoided overfitting with a single dataset, and validated the final classifier with samples not used in model construction (Figures 30A-30D and 31A-31E). Surprisingly, the final model utilized only seven genes, three of which had not been previously reported (Figure 25A). For the entire iPC cohort, the model classified samples with very high accuracy (AUC = 0.99, sensitivity = 98%, specificity = 99%; Figures 25B and 25C, blue). It also accurately classified early-stage PE PEARL samples (AUC = 0.88, sensitivity = 100%, specificity = 83%; Figures 25B and 25C, pink). Unexpectedly, late-onset PE PEARL samples were also classified with reasonable accuracy (AUC=0.74, sensitivity=75%, specificity=67%, Figures 25B and 25C, green).
[0188] This gene set was highly concordant with the transcripts identified by differential abundance analysis (Figure 25D, Table 10). The classifier relied on both placental and maternal expressed transcripts (Figure 25E). All genes used by the model form protein products that are either extracellular or membrane-bound. Despite the small number of genes selected by AdaBoost, a diversity of PE-related functions was observed, specifically cardiovascular function and angiogenesis, immune regulation, fetal development, and ECM remodeling.
[0189] method Prospective Clinical Sample Collection. Pregnant women were recruited through an Illumina-Sponsored clinical trial protocol that adhered to the International Conference on Harmonization for Good Clinical Practice. Following informed consent, 20 mL whole blood samples were collected from 40 pregnant women diagnosed with preeclampsia before 34 weeks of gestation with severe features as defined by ACOG guidelines (Table 6). Samples from 76 healthy pregnant women were also collected, and gestational age was matched to the preeclampsia group. Three control samples developed preeclampsia after blood collection and were excluded from data analysis. For detailed inclusion / exclusion criteria, see Table 6. Patient clinical history, treatment, and birth outcome information were also recorded (Table 7).
[0190] Patients were recruited at eight different clinical sites, including the University of Texas Medical Branch (Galveston, Texas), Tufts Medical Center (Boston, Massachusetts), Columbia University Irving Medical Center (New York, New York), Winthrop University Hospital (Mineola, New York), St. Peter's University Hospital (New Brunswick, New Jersey), Christiana Care (Newark, Delaware), Rutgers University Robert Wood Johnson Medical School (New Brunswick, New Jersey), and New York Presbyterian / Queens (New York, New York). Clinical protocols and informed consent were approved by the institutional review boards at each clinical site. See Table 8 for patient distribution across clinical sites.
[0191] PEARL Validation Cohort Study Design. Illumina obtained plasma samples from the Preeclampsia and Growth Restriction Longitudinal study (PEARL; NCT02379832) for use as an independent validation cohort. Plasma samples were obtained after the study was completed. PEARL samples were collected at the Centre hospitalier universitaire de Quebec (CHU de Quebec) by principal investigator Emmanual Bujold, MD, MSc. The study recruited 45 control and 45 case pregnancies, and written informed consent was obtained from all patients. Only participants over the age of 18 were eligible, and all pregnancies were singletons.
[0192] Preeclampsia group. Preeclampsia was defined based on the June 2014 criteria of the Society of Obstetricians and Gynecologists of Canada (SOGC), with a gestational age requirement of 20 to 41 weeks. A single blood sample was collected at the time of diagnosis.
[0193] Control group. Forty-five pregnant women expected to have a normal pregnancy were recruited between 11 and 13 weeks of gestation. Each enrolled patient underwent blood sampling at four time points throughout pregnancy and delivery, and were followed longitudinally. The control women were divided into three subgroups, and subsequent follow-up blood sampling was staggered to cover the entire range of gestational ages during pregnancy (Table 5).
[0194] The PEARL control sample served two purposes: 153 longitudinal samples from 45 individual women were used to monitor placental dynamics during pregnancy, and control samples were selected for comparison with the preeclampsia cohort and matched for gestational age to validate the model.
[0195] Study sample processing. All samples from the Illumina prospective collection and PEARL samples were processed identically by investigators blinded to disease status. Two blood tubes per patient were collected into Cell-Free DNA BCT tubes (Streck) according to the manufacturer's instructions. Blood samples were stored overnight at room temperature for transport and processed within 72 hours. Blood was centrifuged at 1,600 × g for 20 minutes at room temperature, and the plasma was transferred to a new tube and centrifuged at 16,000 × g for an additional 10 minutes to remove residual cells. Plasma was stored at -80°C until use. Circulating RNA was extracted from 4.5 mL of plasma using a Circulating Nucleic Acid Kit (Qiagen) followed by DNAse I digestion (Thermofisher) according to the manufacturer's instructions.
[0196] cDNA synthesis and library preparation. Circulating RNA was disrupted at 94°C for 8 minutes, followed by random hexamer-primed cDNA synthesis using the Illumina TruSight Tumor 170 Library Prep Kit (Illumina). Illumina sequencing library preparation was performed according to the TST170 Tumor Library Prep Kit for RNA, with the following modifications to accommodate low RNA input: all reactions were reduced to 25% of their original volume, and ligation adapters were used at a 1:10 dilution. Library quality was assessed using the High Sensitivity DNA Analysis kit on an Agilent Bioanalyzer 2100 (Agilent).
[0197] Whole-exome enrichment. Sequencing libraries were quantified using the Quant-iT PicoGreen dsDNA kit (ThermoFisher Scientific), normalized to 200 ng input, and pooled into four samples per enrichment reaction. Whole-exome enrichment was performed according to the TruSeq RNA Access Library Prep guide (Illumina). Additionally, 5' biotin-free blocking oligos designed for the hemoglobin genes HBA1, HBA2, and HBB were added to the enrichment reaction to reduce the enrichment of these genes in the sequencing library. The final enriched libraries were quantified, normalized, and pooled using the Quant-IT Picogreen dsDNA kit (ThermoFisher Scientific), and subjected to paired-end 50x50 sequencing on an Illumina HiSeq2000 platform with a minimum of 40 million reads per sample.
[0198] Data Analysis. All statistical tests were two-sided unless otherwise noted. Nonparametric tests were used when data were not normally distributed. Sequencing reads were mapped to the human reference genome (hg19) using tophat (v2.0.13), and transcript abundance was quantified using featureCounts (subread-1.4.6) against RefGene coordinates (retrieved October 27, 2014). Tissue expression data were obtained from Body Atlas (CorrelationEngine, BaseSpace, Illumina, Inc.) (Kupershmidt et al., 2010, PLoS ONE 5;10.1371 / journal.pone.0013066). vGenes that showed expression ≥2-fold higher than the median in all placental tissues or in any of the fetal tissues (brain, liver, lung, thyroid) were assigned to that group. Subcellular localization was obtained from UniProt.
[0199] Differential expression analysis was performed using R (v3.4.2) with edgeR (v3.20.9) after filtering out genes with a CPM ≤ 0.5 in < 25% of samples. Datasets were normalized using the TMM method, and differentially enriched genes were identified using the glmTreat test for log fold change ≥ 1 and Bonferroni-Holm p-value correction. The same process was performed for each jackknife iteration using 90% of the samples from each group selected by random sampling without replacement. After 1,000 jackknife iterations, one-sided 95% confidence intervals for gene-wise p-values were calculated using StatsModels (v0.8.0). Hierarchical clustering analysis was performed using squared Euclidean distance and average linkage.
[0200] AdaBoost was performed using scikit-learn (v0.19.1, sklearn.ensemble.AdaBoostClassifier) in Python. Optimal hyperparameter values (90 estimates, learning rate 1.6) were determined by grid search, and performance was quantified using the Matthews correlation coefficient. The overall AdaBoost model development strategy is shown in Figures 31A-31E. Datasets (TMM-normalized log CPM values for genes with CPM ≤ 0.5 in < 25% of samples) were standardized (sklearn.preprocessing.StandardScaler) before fitting the classifier. The same scaler fitted to the training data was applied to the corresponding test data set, and all five scalers across the five training data sets were averaged and used for the final model. This decision_function score was used to generate a ROC curve and determine the sample classification.
[0201] RT-qPCR Validation Assay and Analysis. C-RNA was isolated from 2 ml of plasma from 19 randomly selected preeclamptic (PE) cases and 19 matched controls and converted to cDNA. The cDNA was pre-amplified for 16 cycles using TaqMan Preamp Master Mix (Cat. No. 4488593) and diluted 10-fold to a final volume of 500 μL. For qPCR, a reaction mixture containing 5 μL of diluted pre-amplified cDNA, 10 μL of TaqMan Gene Expression Master Mix (Cat. No. 4369542), 1 μL of TaqMan probe, and 4 μL of water was used, according to the manufacturer's instructions. For each TaqMan probe (Table 9), triplicate qPCR reactions were performed for each diluted cDNA sample, and Cq values were determined using Bio-Rad CFX manager software. To determine the gene abundance of each target gene, the average Cq value (ref C) across five reference gene probes was used. qavg ) and ΔΔCq=2^-(target Cq-ref C qavg ) was calculated. To determine the fold change (PE / CTRL) for each probe, the ΔΔCq value for each sample was divided by the mean ΔΔCq value of the matched controls.
[0202] Tube type testing. To evaluate the effect of tube type and overnight shipping on circulating RNA quality, blood was collected from pregnant and non-pregnant women in the following tube types: K2 EDTA (Beckton Dickinson), ACD (Beckton Dickinson), Cell-Free RNA BCT tubes (Streck), and Cell-Free DNA BCT tubes (Streck). Eight milliliters of blood was collected in each tube and transported overnight either on ice (EDTA and ACD) or at room temperature (Cell-Free RNA and DNA BCT tubes). All transported blood tubes were processed into plasma within 24 hours of collection. As a reference, 8 mL of blood was also collected in K2 EDTA tubes, processed into plasma in situ within 4 hours, and shipped as plasma on dry ice. All plasma processing and circulating RNA extraction were performed as described in the Methods section. Three milliliters of plasma per condition was used to generate sequencing libraries for enrichment using the Illumina protocol described.
[0203] Reproducibility study: Plasma was obtained from 10 individuals and divided into volumes of 4 mL, 1 mL, and 0.5 mL, with each volume replicated three times. Circulating RNA extraction (Qiagen Circulating Nucleic Acid Kit) and random-primed cDNA synthesis were performed for all samples as described above. For libraries using 4.5 mL of plasma input, sequencing libraries were generated using the TST170 tumor library preparation kit described above. For libraries using 1 mL and 0.5 mL inputs, libraries were prepared using the Accel-NGS 1 S Plus DNA Library Kit (Swift Biosciences). Whole-exome enrichment and sequencing were performed on all samples using the same procedures.
[0204] Consideration This study focused on identifying universal differences in early-stage PE, supporting the ultimate goal of clinically actionable biomarker discovery. To achieve this, analytical methods needed to be tailored to account for data variability. This variability stems from both substantial biological noise in C-RNA measurements and the phenotypic diversity of PE. C-RNA is inherently more variable than single-tissue transcriptomics because it represents a combination of cell death, signaling, and gene expression in all organs. Furthermore, PE has diverse maternal and fetal outcomes, which may be associated with different molecular causes. While the excluded genes may be biologically relevant in PE, they were not universal in our cohort. Interestingly, the excluded transcripts were elevated in certain women, which may represent molecular subsets of PE. Using a larger cohort will clarify whether C-RNA can identify PE subtypes, which is important for understanding the diverse pathophysiology of this disease.
[0205] The most common set of transcripts was identified using AdaBoost. The success of this method was highlighted by the highly accurate classification of an independent early-onset PE cohort (PEARL). These samples were from a different population with significantly relaxed inclusion and exclusion criteria, including the inclusion of women with chronic hypertension, gestational diabetes, and Alport syndrome in the control group; however, none of these individuals were misidentified as having PE. In contrast to hierarchical clustering, 17 of 24 individuals from the late-onset PE cohort were accurately classified by this machine learning model, strikingly suggesting that early- and late-onset PE are distinct conditions. The findings of this example suggest that there may be some pathways that are universally altered in all PE.
[0206] In all assessments, C-RNA revealed changes in placental, fetal, and maternal-expressed transcripts. One of the most striking trends observed in PE samples was the increased abundance of numerous ECM remodeling and cell migration / invasion proteins (FAM107A, SLC9A3R2, TIMP4, ADAMTS1, PRG2, TIMP3, LEP, ADAMTS2, ZEB1, HSPA12B), tracked by the infiltration of dysfunctional extravillous trophoblast cells, and remodeling of maternal vascular characteristics in this disease. The maternal side of preterm PE manifests as cardiovascular dysfunction, inflammation, and preterm labor (PNMT, ZEB1, CRH), all of which demonstrate molecular signatures of abnormal behavior in this example data.
[0207] [Table 5]
[0208] [Table 6]
[0209] [Table 7]
[0210] [Table 8]
[0211] [Table 9]
[0212] [Table 10-1]
[0213] [Table 10-2]
[0214] important increase * indicates that the change is not statistically different. The CorrelationEngine Body Atlas was used to find the top three tissues expressing genes in the "other" category. UniProt was used to determine subcellular location as a memo. All "membrane" classifications were combined into one category (plasma membrane, ER membrane, etc. were not distinguished).
[0215] [Table 11]
[0216] Example 7 Circulating transcriptome measurements from maternal blood detect molecular signatures of early-stage preeclampsia Circulating RNA (C-RNA) is continuously released into the bloodstream from tissues throughout the body, allowing noninvasive monitoring of pregnancy health from conception to delivery. In this example, we determine that C-RNA analysis can detect abnormalities in patients diagnosed with preeclampsia (PE), a prevalent and potentially fatal pregnancy complication. As an initial test, we sequenced the circulating transcriptomes from 40 pregnancies at the time of diagnosis of severe, early-stage PE, along with 73 gestational age-matched controls. Thirty transcripts consistent with PE biology were identified, likely representing contributions to placental, fetal, and maternal disease. Furthermore, machine learning identified C-RNA transcript combinations that reliably classified PE patients in two independent cohorts (accuracy 85%-89%). The ability of C-RNA to reflect maternal, placental, and fetal health holds great promise for improving the diagnosis and identification of at-risk pregnancies. In summary, the circulating transcriptome reflects biologically relevant changes in patients with early-stage severe preeclampsia and can be used to accurately classify patient status.
[0217] Preeclampsia (PE) is one of the most common and serious complications of pregnancy, affecting an estimated 4-5% of pregnancies worldwide (Abalos et al., 2013, Eur J Obstet Gynecol Reprod Biol; 170:1-7, and Ananth et al., 2013, BMJ; 347:f6564) and is associated with substantial maternal and perinatal morbidity and mortality (Kuklina et al., 2009, Obstet Gynecol; 113:1299, and Basso et al., 2006, JAMA; 296:1357-1362). In the United States, the incidence of PE is increasing due to an aging maternal population and the increasing prevalence of comorbid conditions such as obesity (Spradley et al., 2015, Biomolecules; 5: 3142-3176), costing the US healthcare system an estimated $2 billion (Stevens et al., 2017, Am J Obstet Gynecol; 217: 237-248.e16).
[0218] PE is diagnosed as new-onset hypertension associated with maternal peripheral organ damage occurring after 20 weeks of pregnancy (Hypertension in Pregnancy: Executive Summary, 2013, Obstet Gynecol; 122:1122, and Tranquilli et al., 2014, Pregnancy Hypertens; 4(2):97-104). However, there is significant heterogeneity in the presentation and course of PE, including the time of disease onset, symptom severity, clinical manifestations, and maternal and neonatal outcomes (Lisonkova and Joseph, 2013, Am J Obstet Gynecol; 209:544.e1-544.e12). PE is primarily described based on whether it occurs before 34 weeks (early) or after 34 weeks (late), or whether it presents with severe symptoms such as persistently elevated blood pressure (≥160 / 110 mmHg), neurological symptoms, and / or severe liver or kidney damage (American College of Obstetricians and Gynecologists, Task Force on Hypertension in Pregnancy, 2013, Obstet Gynecol;122:1122-1131).
[0219] The pathophysiology of early PE is not fully understood, but it is thought to occur in two phases (Phipps et al., 2019, Nature Reviews Nephrology;15:275). Early PE results from abnormalities in implantation and placentation during early pregnancy and is associated with maternal immunodeficiency (Hiby et al., 2010, J Clin Invest; 120:4102-4110; Ratsep et al., 2015, Reproduction; 149:R91-R102; and Girardi, 2018, Semin Immunopathol; 40:103-111), incomplete cytotrophoblast differentiation (Zhou et al., 1997, J Clin Invest; 99:2152-2164), and / or oxidative stress at the maternal-placental interface (Burton and Jauniaux, 2011, Best Pract Res Clin Obstet Gynaecol; 25:287-299), resulting in incomplete remodeling of the maternal spiral arteries and failure to establish the definitive uterofetal circulation (Lyall et al., 2014). (Hecht et al., 2017, Hypertens Pregnancy; 36:259-268; Young et al., 2010, Annu Rev Pathol; 5:173-192; and Backes et al., 2011, J Pregnancy; 2011: doi:10.1155 / 2011 / 214365). This leads to inadequate placental perfusion after 20 weeks of gestation. Consequently, placental dysregulation triggers a second stage, manifesting primarily as maternal systemic vascular insufficiency with negative consequences for the fetus, including fetal growth restriction and iatrogenic preterm birth (Hecht et al., 2017, Hypertens Pregnancy; 36:259-268; Young et al., 2010, Annu Rev Pathol; 5:173-192; and Backes et al., 2011, J Pregnancy; 2011: doi:10.1155 / 2011 / 214365).In contrast, placental insufficiency in late-onset PE is thought to be due not to abnormalities in placentation but to impaired placental perfusion secondary to maternal vascular disease, such as that seen in patients with chronic hypertension, pregestational diabetes (Vambergue and Fajardy, 2011, World J Diabetes; 2:196-203), and collagen vasculopathy ("Placental pathology in maternal autoimmune diseases—new insights and clinical implications," 2017, International Journal of Reproduction, Contraception, Obstetrics and Gynecology; 6:4090-4097).
[0220] The heterogeneity and complexity of this disease make diagnosis, risk prediction, and treatment development challenging. Furthermore, the placenta offers limited molecular characterization of disease progression because the primary diseased organ cannot be easily examined. Circulating RNA (c-RNA) has shown great promise for noninvasively monitoring maternal, placental, and fetal dynamics during pregnancy (Tsui et al., 2014, Clin Chem; 60:954-962, and Koh et al., 2014, Proc Natl Acad Sci USA; 111:7361-7366). C-RNA is released into the bloodstream from many tissues via multiple cellular processes, including apoptosis, microvesicle shedding, and exosome signaling (van Niel et al., 2018, Nat Rev Mol Cell Biol; 19:213-228). Due to these diverse sources, measurement of C-RNA reflects tissue-specific changes in gene expression, cell-cell signaling, and the extent of cell death occurring in different tissues throughout the body. Therefore, C-RNA may elucidate the molecular basis of PE and ultimately identify predictive, prognostic, and diagnostic biomarkers for the disease (Hahn et al., 2011, Placenta;32:S17-20).
[0221] Several studies have begun to investigate and identify potential C-RNA-based biomarkers for a range of pregnancy complications, including preterm birth, PE, and infectious diseases (Pan et al., 2017, Clin Chem; 63:1695-1704; Ngo et al., 2019, Science; 360:1133-1136; and Whitehead et al., 2016, Prenat Diagn; 36:997-1008). However, significant interindividual variability in this sample type can obscure subtle changes in disease-specific biomarkers (Meder et al. 2014, Clin Chem; 60:1200-1208). Therefore, many PE-focused C-RNA studies have measured the RNA of previously identified serum protein biomarkers, including soluble FLT1, soluble endoglin, and oxidative stress and angiogenesis markers (Nakamura et al., 2009, Prenat Diagn;29:691-696, Purwosunu et al., 2009, Reprod Sci;16:857-864, and Paiva et al., 2011, J Clin Endocrinol Metab;96:E1807-1815). Although protein measurements of these serum markers are known to be altered in PE (Maynard et al., 2003, J Clin Invest; 111:649-658, Venkatesha et al., 2006, Nat Med; 12:642-649, and Rana et al., 2018, Pregnancy Hypertens; 13:100-106), it is unclear whether they are the most effective predictors of C-RNA, and a more extensive exploration approach is warranted.
[0222] In this example, global measurements of the circulating transcriptome detect unique molecular signatures specific to early-stage severe PE. To facilitate discovery, whole-transcriptome enrichment performance was optimized for high-throughput sequencing, enabling the measurement of >14,000 C-RNA transcripts per sample with high confidence. C-RNA profiles were then globally characterized from a preliminary cohort of 113 pregnancies, 40 of which were diagnosed with early-stage severe PE. All analytical methods were tailored to address the high biological variance inherent in C-RNA, identifying altered transcripts consistent with PE biology that could be classified across the cohort with high accuracy, highlighting that this sample type provides a means for developing a robust test for assessing preeclampsia.
[0223] result Establishing a reproducible whole transcriptome workflow for c-RNA.
[0224] C-RNAs are present in plasma at relatively low abundance, dominated by ribosomal (rRNA) and globin RNAs, and are a mixture of fragmented and full-length transcripts (Crescitelli et al., 2013, J Extracell Vesicles;2:doi:10.3402 / jev.v2i0.20677), all of which can affect the efficiency of library preparation methods for next-generation sequencing. Therefore, we optimized the workflow to minimize variability and maximize exonic C-RNA signals (Figure 37A).
[0225] High-abundance globin and rRNA do not inform biomarker discovery and must be removed. However, standard depletion methods, such as Ribo-Zero (Illumina, Inc.) or NEBNext rRNA Depletion (New England Biolabs), are not suitable for low starting amounts of RNA (Adiconis et al., 2013, Nat Methods 10:623-629). Pre-depletion by these methods was able to remove unwanted ribosomal sequences from C-RNA (Figure 37B, gray), but did not consistently increase the signal of exonic C-RNA in sequencing libraries (Figure 37B, orange). Instead, samples differed in the proportion of reads mapping to complex and variable populations of non-human RNA sequences, such as GB Virus C (Figure 37B, pink) (Manso et al., 2017, Sci Rep;7:doi:10.1038 / s41598-017-02239-5, and Whittle et al., 2019, Front Microbiol;9:doi:10.3389 / fmicb.2018.03266). In addition, removal of high-abundance rRNA and globulin RNA results in extremely low RNA input, which increases the destructive rate of bias-based library preparation methods. To circumvent these issues, we opted for a probe-assisted enrichment method that generates libraries from all C-RNAs followed by targeting the entire human exome (Figure 37A). This method consistently generated high-quality sequencing libraries composed of >90% exonic C-RNAs (Figure 37B, orange).
[0226] There is significant interindividual variability in plasma concentrations of C-RNA (Figure 38A), which can vary by an order of magnitude (mean 1.1 ng / mL plasma, SD 0.7, range <0.1–5 ng / mL plasma). To ensure reproducible results, we evaluated the impact of plasma volume input on C-RNA data quality. Using less than 2 mL of plasma significantly increases the biological coefficient of variation and reduces library complexity (Figures 38B and 38C), resulting in reduced sensitivity. Therefore, we selected a plasma input of 4 mL to minimize noise, maximize confidence in data quality, and ensure normal data generation from all individuals, regardless of C-RNA plasma concentration.
[0227] To achieve consistent handling of all samples across diverse collection sites, all processing was centralized, and collected blood samples had to be transported to a single laboratory. Therefore, the impact of overnight transport on C-RNA data quality needs to be evaluated. Blood from non-pregnant and pregnant women (gestational age, GA, >28 weeks) was collected into four blood collection tubes (BCT) and stored overnight at the manufacturer's recommended temperature before processing: BD Vacutainer K2EDTA (Beckton Dickinson, 4°C), BD Vacutainer ACD-A (Beckton Dickinson, 4°C), Cell-Free DNA BCT (Streck, Inc., room temperature), and Cell-Free RNA BCT (Streck, Inc., room temperature). The set of samples collected in the EDTA BCT was processed in the clinic within 2 hours to obtain a baseline. The C-RNA pregnancy signal for each sample was determined by summing the normalized abundance levels of 155 transcripts identified as pregnancy markers in two previous C-RNA publications (Tsui et al., 2014, Clin Chem; 60:954-962, and Koh et al., 2014, Proc Natl Acad Sci USA; 111:7361-7366). After overnight storage, the EDTA RNA pregnancy signal was clearly detectable in the pregnant samples, despite a decrease in overall signal intensity compared to the immediate processing of the EDTA samples (Figures 39A and 39B). No significant differences were observed between the different BCTs after overnight storage, indicating that all were suitable for C-RNA analysis. Cell-Free DNA BCTs (Streck, Inc.) were selected for subsequent sample collection, allowing for room-temperature transport. Furthermore, correlation of transcriptomic profiles in a time course experiment confirmed that these BCTs showed no increased technical variability after room temperature storage for up to 5 days (Figure 39C).
[0228] The complete workflow was demonstrated by replicating a previous study monitoring C-RNA dynamics in healthy pregnant women from early to late pregnancy. Using 152 consecutively collected samples from 41 healthy pregnant women (Pre-Eclampsia and Growth Restriction Longitudinal Study Control Cohort - PEARL CC; NCT02379832; Table 13), we identified 156 significantly altered transcripts, the majority of which increased in abundance as pregnancy progressed (Figure 32A, Table 14). Of the identified transcripts, 51% were primarily altered during early pregnancy, 6% during late pregnancy, and 43% were differentially regulated throughout pregnancy (Figure 40, Table 14). Early pregnancy genes were enriched for regulation of placental steroidogenesis and trophoblast differentiation, while late pregnancy genes were involved in the onset of labor. Transcripts that increase during pregnancy are associated with the development and morphogenesis of tissues and organs (Chatuphonprasert et al., 2018, Front Pharmacol;9:doi:10.3389 / fphar.2018.01027; Debieve et al., 2011, Mol Hum Reprod;17:702-9; Grammatopoulos and Hillhouse, 1999, Lancet;354:1546-1549; and Marshall et al., 2017, Reprod Sci;24:342-354). The PEARL CC results were highly consistent with the literature, as 42% of the altered genes had been identified in previous C-RNA studies (Figure 32B) (Tsui et al., 2014, Clin Chem; 60:954-962, and Koh et al., 2014, Proc Natl Acad Sci USA; 111:7361-7366). Of the 91 transcripts identified exclusively in this study, 64% are expressed by placental and / or fetal tissues (tissue specificity defined as >2-fold higher than the median for all tissues in the Body Atlas) (Figure 32C and Figure 17) (Kupershmidt et al., 2010, PLOS ONE; 5:e13066). The remaining genes are hypothesized to reflect maternal tissue responses to pregnancy (Table 14).
[0229] Clinical trial design for early severe PE After confirming that this workflow robustly detects pregnancy-associated C-RNA dynamics, we identified C-RNA changes associated with pregnancy complications. We applied the workflow to samples collected from two independent PE cohorts: the Illumina Preeclampsia Cohort (iPEC; NCT02808494) and the PEARL Preeclampsia Cohort (PEARL PEC; NCT02379832). While iPEC was used for biomarker identification, PEARL PEC was used for independent confirmation of our findings. Importantly, all samples in both cohorts were collected in Streck Cell-Free DNA BCTs, and libraries were prepared using the same method as discussed in the previous section.
[0230] The IPEC study focused on early-stage PE with severe features and excluded women diagnosed with additional health complications, such as chronic hypertension or diabetes, to prevent further heterogeneity by obscuring consistent PE-associated C-RNA signals (Table 15). One hundred and thirteen samples were collected across 40 centers at the time of initial PE diagnosis (Table 17), and 73 gestational age-matched controls were collected within one week (Figure 33A). Maternal characteristics, pregnancy outcomes, and medications were recorded throughout the study (Tables 12 and 16). There were no significant differences between the PE and control groups in fetal sex, maternal age, and nulliparity. In contrast, BMI was significantly higher in the PE cohort (p=0.0007) (O'Brien et al., 2003, Epidemiology;14:368-374). All but one patient with PE was born preterm, in contrast to 9.5% of the control group, confirming that our diagnostic criteria identify individuals severely affected by the disease ( Figure 33C ).
[0231] The PEARL PEC sample was collected by an independent institution (CHU de Quebec-Universite Laval) and consisted of 12 early-onset and 12 late-onset PE pregnancies and an equal number of gestational age-matched controls (Figure 33B). Maternal characteristics, pregnancy outcomes, and medications in use were recorded throughout the study (Table 18). Similar to iPEC, 100% of early-onset patients delivered preterm, while 75% of the late-onset cohort delivered at term, confirming the difference in severity associated with early- and late-onset PE. Chronic hypertension, diabetes, and other maternal health conditions were not warranted for exclusion, making this cohort more representative of the inherent heterogeneity in the pregnant population.
[0232] Identification of transcripts consistently altered in early-stage PE Standard differential expression analysis (Robinson et al., 2010, Bioinformatics; 26:139-140) using the entire cohort of iPECs identified 42 transcripts with altered abundance in plasma, 37 of which were increased in PE (Figure 34A, blue and orange). However, when different sample subsets were selected for analysis, variations in differential transcript abundance were observed. Therefore, we incorporated a jackknife algorithm (Library, 1958, Ann Math Statist; 29:614-623) to identify the most consistently altered transcript abundances when comparing PE and control samples (Figures 34A and 34B, orange). We performed 10,000 repetitions of the differential analysis with randomly selected subsets of PE and control samples, resulting in the construction of confidence intervals for the p-values associated with each transiently altered transcript (Figure 34C). We subsequently excluded 12 transcripts with p-value confidence intervals greater than 0.05 (Figure 34B). These transcripts were not excluded by simply setting thresholds for baseline abundance or biological variance (Figure 34D), but they were observed to have lower predictive value (Figure 34E). Hierarchical clustering indicates that these transcripts are not universally altered in the PE cohort and therefore lack the sensitivity (73%) for accurate classification of this condition (Figure 34F).
[0233] A representative set of 20 transcripts altered in PE was independently quantified by qPCR in a subset of affected and control iPEC patient samples. The fold changes measured by qPCR were highly concordant with the sequencing data, validating our findings (Figure 35A and Table 19). 58% of the transcripts in the refined list have been previously associated with PE (Table 20). In addition, nearly all genes could be linked to PE-related processes, including extracellular matrix (ECM) remodeling, gestational age, placental / fetal development, angiogenesis, and hypoxic response (Table 20). 67% of these genes were expressed by the placenta and / or fetus (Figure 35B). Among the remaining maternal-expressed genes, cardiovascular and immune functions were well represented, both of which are altered in PE (Table 20) (Phipps et al., 2019, Nature Reviews Nephrology; 15:275).
[0234] Hierarchical clustering of these transcripts effectively separated PE and control samples from iPEC with 98% sensitivity and 97% specificity (Figure 35C). The refined list of 30 transcripts was validated in an independent PEARL PEC. Early PE samples clustered separately from matched controls with 83% sensitivity and 92% specificity, further validating the association of these transcripts (Figure 35D). In contrast, no clustering was observed for late PE and matched control samples, suggesting that late PE may have a weaker or distinct C-RNA signature (Figure 35E) (Redman, 2017, An International Journal of Women's Cardiovascular Health; 7:58, and Hahn et al., 2015, Expert Rev Mol Diagn; 15:617-629).
[0235] Upon review of the clinical data for the three incorrectly clustered iPEC samples, we discovered that two controls had potentially confounding health problems, including hypertension and, in one case, preterm delivery. These controls, according to our strict exclusion criteria, should not have been enrolled in iPEC and were therefore excluded from further analysis. The incorrectly clustered PE sample showed no clinical abnormalities and was retained in our iPEC dataset.
[0236] Developing a robust machine learning classifier for early-stage PE Differential expression analysis confirmed that C-RNAs detect biologically relevant changes in PE patients. To evaluate whether C-RNA signatures can reliably classify PE, we constructed AdaBoost models using data from the iPEC cohort (Freund and Schapire, 1997, J Comp Sys Sci; 55:119-139; McPherson et al., 2011, PLoS Comput Biol; 7:doi:10.1371 / journal.pcbi.1001138; and Lu et al., 2015, PLOS ONE; 10:e0130622). Ten percent of randomly selected samples were excluded from the entire machine learning process to evaluate final model performance. A nested cross-validation approach was then used to optimize the hyperparameters (Figure 42) and build the AdaBoost model (Figure 43) (Cawley and Talbot, 2010, J Machine Learning Res; 11:2079-2107).
[0237] While developing the machine learning method, we observed a high degree of variability in AdaBoost performance and selected genes based on the samples included in training (Figure 44). These observations indicate that different subsets of samples significantly affect model building (Assessing and Improving the Stability of Chemometric Models When Sample Numbers Are Low | SpringerLink (available on the World Wide Web at link.springer.com / article / 10.1007%2Fs00216-007-1818-6)), which may be due in part to heterogeneity in PEs. To account for this heterogeneity, AdaBoost models were fitted to multiple combinations of training samples. The estimates from these orthogonally generated models were then combined into a single ensemble to obtain a minimal gene set (Figure 43) (Martinez-Muñoz and Suarez, 2007, Pattern Recognition Letters; 28:156-165, and AveBoost2: Boosting for Noisy Data | SpringerLink (available on the World Wide Web at link.springer.com / chapter / 10.1007 / 978-3-540-25966-4_3)). This allowed for the capture of the broad diversity of PE expression in a refined machine learning model with the potential to accurately classify independent samples from highly pregnant populations.
[0238] Within each fold of cross-validation, a threshold AdaBoost score was identified to distinguish PE and control samples that maximized both sensitivity and specificity. Across all 10 folds, we obtained an average ROC of 0.964 (+ / - 0.068 SD) (Figure 36A). We first evaluated performance on holdout iPEC samples, yielding an accuracy of 89% (+ / - 5% SD), sensitivity of 88% (+ / - 13% SD), and specificity of 92% (+ / - 6% SD) (Figure 36B, blue). AdaBoost classification performance was unaffected by the amount of time prior to plasma treatment, further supporting the robustness of our sample preparation protocol and analysis (Figures 40D and 40E). We then investigated the model's ability to classify an independent PEARL PEC cohort. Early-onset PEARL PEC samples achieved 85% (+ / - 4% SD) accuracy, 77% (+ / - 9% SD) sensitivity, and 92% (+ / - 7% SD) specificity (Figure 36B, pink). Unexpectedly, late-onset PEARL PEC samples were also classified with reasonable accuracy of 72% (+ / - 6% SD), 59% (+ / - 10% SD) sensitivity, and 80% (+ / - 10% SD) specificity (Figure 36B, green).
[0239] A total of 49 transcripts were used by AdaBoost, with 63% selected in at least two rounds of model building (Figure 36C). In conventional analysis, 40% of the genes identified in the jackknife analysis were also used in machine learning (Figure 36D, Table 21). 38% of the transcripts used by the classifier were upregulated in the placenta and / or fetus (Figure 36E). Transcripts reflecting the diversity of PE-related pathways were observed, particularly genes related to immune regulation and fetal development (Table 21).
[0240] Consideration Whole-transcriptome C-RNA analysis captures a molecular snapshot of the diverse and complex interactions of pregnancy at a single time point, casting a broad net necessary for effective biomarker discovery. We tailored our workflow and analysis to minimize technical noise, obtain high-quality C-RNA measurements, and ultimately detect biologically relevant changes. Importantly, molecular alterations specific to the complex pathophysiology of early-stage severe PE were detected at diagnosis, supporting robust classification across cohorts. The identified altered C-RNA transcripts represent contributions from maternal, placental, and fetal tissues, many of which are not captured in studies focusing on placental tissue collected after delivery. These findings highlight the power of C-RNA to comprehensively monitor signals contributed by diverse tissues of origin during ongoing pregnancy.
[0241] To identify the best method capable of detecting global and potentially subtle changes in pregnancy, we determined the effects of plasma input, library preparation method, and BCT on C-RNA data quality. Because the majority of transcripts are likely present in plasma at low abundance, a 4 mL plasma input was selected in the protocol to minimize noise due to sampling error and insufficient library conversion, which are problematic in low-input sequencing applications. Pre-depletion of abundant RNAs did not remove all of the diverse contaminating RNA species present in the sample population. This targeted depletion was not feasible, and therefore a whole-transcript enrichment method was selected to consistently isolate the targeted exonic C-RNA signal. Overnight transport was a logistical requirement of the protocol, but it risked both introducing signal loss due to C-RNA degradation and further RNA contamination from cell lysis. Streck Cell-Free DNA BCT (Zhao et al., 2019, J Clin Lab Anal;33:e22670) was selected because it inhibits cell lysis at room temperature. Although this BCT was not specifically designed for RNA stabilization, little evidence of C-RNA degradation was observed after several days of storage. Previous studies have shown that C-RNA receives sufficient endogenous protection from extracellular nucleases (Tsui et al., 2002, Clin Chem; 48:1647-1653), therefore, no additional precautions to protect the RNA are necessary. These optimizations identified a workflow that maximized C-RNA transcriptome signal and minimized technical variability, as indicated by the numerous biologically relevant changes observed across healthy and PE pregnancies.
[0242] Next, our analysis focused on identifying differences in the circulating transcriptome ubiquitously present in the most extreme phenotype of the disease, i.e., early-stage PE with severe features. This approach needs to be adjusted to account for the variability observed in our data, resulting from both substantial biological noise in C-RNA measurements and the phenotypic diversity of PE. C-RNA is inherently more variable than single-tissue transcriptomics because it interrogates RNA from diverse tissues and biological processes and detects not only changes in gene expression but also differences in the rates of cell death and intercellular signaling. Furthermore, PE exhibits a wide range of maternal and fetal symptoms and outcomes that may be related to various underlying molecular causes and responses. Genes excluded after jackknife analysis may be biologically relevant in PE but were not universally altered in the affected cohort. These transcripts may represent molecular subsets of the disease, and larger cohorts will help clarify whether C-RNA can further delineate PE subtypes, which is important for understanding the diverse pathophysiology of this syndrome.
[0243] The transcripts identified by jackknife represent a diversity of functions across the maternal-fetal interface. The majority of the identified changes are related to placental dysfunction and altered fetal development. One of the most striking trends was the increased abundance of ECM remodeling and cell migration proteins (N=10), which tracked the dysfunction of trophoblast invasion outside the follicle, characteristic of early-onset PE (Yang et al., 2019, Gene; 683:225-232; Zhu et al., 2012, Rev Obstet Gynecol; 5:e137-e143; and Wang et al., 2019, Scientific Reports; 9:2728). Twenty percent of the repressed genes encoded angiogenic proteins, consistent with numerous observations that the balance of angiogenic factors plays an important role in regulating placental vascularization (Cerdeira et al., 2012, Cold Spring Harbor Perspectives in Medicine;2:a006585-a006585) and can identify early PE with severe features (Zeisler et al., 2016, N Engl J Med;374:13-22). As shown in the data presented herein, early-stage severe PE also impairs fetal growth and development, as evidenced by increased abundance of four transcripts encoding regulators of IGF signaling (Argente et al., 2017, EMBO Mol Med;9:1338-1345, and Weyer and Glerup, 2011, Biol. Reprod;84:1077-1086), a pathway important for fetal development (Forbes and Westwood, 2008, Horm Res;69(3):129-137). The remaining transcripts captured maternal components of PE, namely immune and cardiovascular dysregulation.Evidence of the maternal immune imbalance that characterizes PE manifests as changes in immune tolerance and levels of pro- and anti-inflammatory factors (Chistiakov et al., 2014, Front Physiol; 5 (2014), doi:10.3389 / fphys.2014.00279; Kumar et al., 2012, Cancers (Basel); 4: 1252-1299; Qi et al., 2003, Nature Medicine; 9: 407; and Yang et al., 2014, Biochim Biophys Acta; 1840: 3483-3493). Altered transcripts important for blood pressure regulation and several genes associated with atherosclerosis were also identified in PE c-RNA profiles, consistent with maternal vascular disease as an underlying mechanism predisposing some patients to PE (Calo et al., 2014, J Hypertens;32:331-338, and Magnusson et al., 2012, PLOS ONE;7:e43142). The identified transcripts capture the diversity of PE-associated functions and highlight the ability of c-RNA to simultaneously monitor multiple molecular processes involved in complex diseases.
[0244] Next, we determined whether C-RNA could not only detect biologically relevant changes but also accurately classify pregnant women affected by early-stage severe PE. The careful approach to AdaBoost model building described herein identified transcript combinations that could classify across different patient subsets while excluding features that could overfit our model and introduce bias. Although 76% of the transcripts used by AdaBoost were not identified as differentially abundant, these still reflect the same PE-related pathways captured in our jackknife analysis. The success of this strategy was demonstrated by the highly accurate classification of independent early-stage PEARL PECs. These samples were collected at diagnosis from a different population than used for training, and the inclusion / exclusion criteria were less stringent. For example, this cohort included five women with chronic hypertension or gestational diabetes, only one of whom was misclassified by the minor AdaBoost model, indicating that the C-RNA changes utilized in the machine learning were highly specific for PE.
[0245] This example provides an important step toward improving the understanding and diagnosis of PE. Due to the limited size of our clinical cohort, it does not capture the phenotypic diversity of the entire pregnant population and the global pregnant population, which is reflected in the lower values obtained for sensitivity than specificity in the classification analysis. A larger cohort would better encompass the heterogeneity of PE and identify signals for diverse manifestations of the disease. A second limitation is the targeted nature of whole-transcriptome enrichment, which therefore does not capture the full range of non-coding or non-human transcripts present in plasma. Certain infections and miRNAs have been associated with PE (Nourollahpour Shiadeh et al., 2017, Infection; 45:589-600, and Skalis et al., 2019, Microrna; 8:28-35). Therefore, future studies aim to incorporate these measurements using RNA transcriptome data. While the findings in this example are highly consistent with those reported in the literature, no changes in transcripts of widely reported serum protein biomarkers, such as soluble FLT1 (sFLT), vascular endothelial growth factor (VEGF), or placental growth factor (PIGF), were observed (Phipps et al., 2019, Nature Reviews Nephrology;15:275). This result is not surprising given that gene expression does not necessarily correlate with protein abundance or release into circulation.
[0246] iPEC sampling was particularly focused on early-onset PE, which has severe features at diagnosis. While this represents the most extreme phenotype of the disease, it represents a small percentage of PE cases and merits further exploration of the clinical PE spectrum. As a preliminary test, we applied the AdaBoost model to the late-onset PEARL PEC cohort and achieved reasonably good accuracy (72%). This is surprising given substantial evidence that early-onset PE and late-onset PE are distinct conditions, despite sharing the ultimate phenotype of placental insufficiency (Burton et al., 2019, BMJ;366:12381). Therefore, some of the transcripts identified by AdaBoost may reflect a response to uteroplacental insufficiency rather than their source and therefore may not have early predictive value for disease progression. In contrast, changes in transcripts involved in angiogenesis and trophoblast invasion were observed (Xie et al., 2018, Res Commun;506:692-697; Hunkapiller et al., 2011, Development;138:2987-2998; and Chrzanowska-Wodnicka, 2017, Curr Opin Hematol;24;248:255), known molecular drivers of PE development and potentially predictive of early stages of disease progression. Regardless of whether the detected changes represent cause or effect, or even a combination thereof, the methods and findings described herein demonstrate that c-RNA provides a unique opportunity to build robust diagnostic algorithms and investigate previously unconsidered disease mechanisms.
[0247] The successful classification of PE patients at the time of diagnosis indicates that C-RNA profiles can be used to reliably monitor maternal, fetal, and placental function in real time. Future studies should focus on early pregnancy to evaluate the potential of this approach to improve prognosis and outcome prediction for women with PE. Indeed, such studies hold great promise for early stratification of all at-risk pregnant women and identifying predictive biomarkers for preventive intervention or more careful monitoring of the pregnancy. The application of C-RNA will ultimately provide comprehensive molecular monitoring of maternal and fetal health throughout pregnancy.
[0248] Materials and Methods Test Design The objective of this example was to determine whether C-RNA can detect molecular markers associated with severe early-stage PE. This goal was achieved by (i) optimizing a protocol for obtaining robust whole-transcriptome C-RNA measurements, (ii) analyzing plasma C-RNA profiles from patient- and gestational age-matched control pregnancies at the time of PE diagnosis, and (iii) validating findings using C-RNA data generated from an independent cohort. The iPEC study clinical protocol and informed consent forms were approved by the Institutional Review Boards of each clinical site, and the inclusion and exclusion criteria are listed in Table 15. Investigators were blinded to sample status throughout the bioinformatics processing of the sequencing data.
[0249] Clinical sample collection iPEC. Pregnant patients were recruited under an Illumina-sponsored clinical trial protocol (NCT02808494) in compliance with the International Conference on Harmonization of Clinical Conduct. Participants were recruited at eight different clinical sites: University of Texas Medical Branch (Galveston, Texas), Tufts Medical Center (Boston, Massachusetts), Columbia University Irving Medical Center (New York, New York), Winthrop University Hospital (Mineola, New York), St. Peter's University Hospital (New Brunswick, New Jersey), Christiana Care (Newark, Delaware), Rutgers University Robert Wood Johnson Medical School (New Brunswick, New Jersey), and New York Presbyterian / Queens (New York, New York).
[0250] After informed consent, 20 mL whole blood samples were collected from 40 singleton pregnant women diagnosed with PE before 34 weeks of gestation with severe features as defined by ACOG guidelines (Table 15) (Hypertension in Pregnancy: Executive Summary, 2013, Obstet Gynecol;122:1122). Samples from 76 healthy pregnant women were also collected, and gestational age was matched to the PE group. Three control samples developed PE after blood collection and were excluded from the study. Maternal characteristics, birth outcomes (Table 32), and medications in use (Table 16) were all recorded during the study.
[0251] PEARL. Plasma samples from the PEARL trial (NCT02379832) were used as an independent validation cohort. PEARL samples were collected at the Centre Hospitalier Universitaire de Quebec (CHU DE Quebec). Only participants over the age of 18 were eligible, and all pregnancies were singletons. A group of 45 control pregnancies (PEARL Healthy Control Cohort, PEARL HCC) and 45 case pregnancies (PEARL Preeclampsia Cohort, PEARL PEC) were recruited in this study, and written informed consent was obtained from all patients. Selected plasma samples were obtained after study completion.
[0252] PE criteria were defined based on the Society of Obstetrics and Gynaecologists of Canada (SOGC) June 2014 criteria for PE, with a gestational age requirement of 20 to 41 weeks, encompassing both early (diagnosed at <34 weeks, N = 12) and late (diagnosed at >34 weeks, N = 12) PE. From the PEARL PEC sample, a blood sample was collected once at the time of diagnosis. The PEARL HCC included 45 pregnant women recruited between 11 and 13 weeks of gestation, with prospective normal pregnancies. Each enrolled patient underwent blood sampling at four time points throughout pregnancy and delivery, and was followed longitudinally. Control women were divided into three subgroups, and subsequent follow-up blood sampling was staggered to cover the full range of gestational ages during pregnancy (Table 13). In addition to assessing C-RNA in healthy pregnancies using PEARL HCC samples, samples from 24 unique individuals in PEARL HCC were selected to serve as gestational age-matched controls for both early- and late-onset PE cohorts.
[0253] Sample preparation Plasma Processing. All samples from the iPEC and PEARL cohorts were randomly processed by investigators blinded to disease status. Two blood samples per patient were collected into Cell-Free DNA BCT tubes (Streck, Inc.). iPEC blood samples were stored overnight at room temperature for transport and processed within 120 hours. PEARL blood samples were collected, processed into plasma within 24 hours, and stored at -80°C until transported to Illumina on dry ice. All blood samples were centrifuged at 1,600 × g for 20 minutes at room temperature, and the plasma was transferred to a new tube and centrifuged at 16,000 × g for an additional 10 minutes. The plasma supernatant was stored at -80°C until use.
[0254] Preparation of sequencing libraries. C-RNA was extracted from 4.5 mL of plasma using a Circulating Nucleic Acid Kit (Qiagen) followed by DNAse I digestion (Thermo Fisher Scientific) according to the manufacturer's instructions. C-RNA was disrupted at 94°C for 8 minutes, followed by random hexamer-primed cDNA synthesis using the Illumina TruSight Tumor 170 Library Prep Kit (TST170, Illumina, Inc.). Illumina sequencing library preparation was performed according to the TST170 kit for RNA with two modifications: reducing all reactions to 25% of the original volume and using a 1:10 dilution of ligation adapters. Library quality was assessed using a high-sensitivity DNA analysis chip on an Agilent Bioanalyzer 2100 (Agilent Technologies).
[0255] Whole-transcriptome enrichment. Sequencing libraries were quantified using the Quant-iT PicoGreen dsDNA kit (Thermo Fisher Scientific), normalized to 200 ng input, and pooled into four samples per enrichment reaction. Global exome enrichment was performed using the Illumina TruSeq RNA Enrichment Kit (Illumina, Inc.). Briefly, biotinylated exome-targeting oligonucleotides were hybridized to the sequencing library and pulled down by magnetic streptavidin beads to enrich the library for exonic RNA. This process was performed twice to maximize exon enrichment. The final enriched library was then reamplified by PCR to obtain a yield sufficient for sequencing. The enrichment reaction included 5' biotin-free block oligos designed against the hemoglobin genes HBA1, HBA2, and HBB to minimize contributions from abundant hemoglobin. The final enriched libraries were quantified using the Quant-IT Picogreen dsDNA kit (Fisher Scientific), normalized, pooled, and subjected to paired-end 50 × 50 sequencing on an Illumina HiSeq2000 platform with a minimum of 40 million reads per sample.
[0256] Sequencing data analysis Bioinformatics processing of sequencing data. Fastq files containing >50 million reads were downsampled to 50 million reads using seqtk (v1.2-r102-dirty). Sequencing reads were mapped to the human reference genome (hg19) using TopHat2 (v2.0.13) (Yang et al., 2014, Biochim Biophys Acta; 1840:3483-3493), and transcript abundance was quantified using featureCounts (subread-1.4.6) (Calo et al., 2014, J Hypertens; 32:331-338) against RefGene coordinates (obtained October 27, 2014). Tissue expression data were obtained from Body Atlas (Correlation Engine, Base Space, Illumina, Inc.) (Whittle et al., 2019, Front Microbiol;9:doi:10.3389 / fmicb.2018.03266). Genes with expression ≥2-fold higher than the median in all placental tissues or in any fetal tissue (brain, liver, lung, thyroid) were assigned to that group. Subcellular localization was obtained from UniProt (Magnusson et al., 2012, PLOS ONE 7:e43142). Functional enrichment analysis was performed in gProfiler (v e97_eg44_p13_d22abce) (Raudvere et al., 2019, Nucleic Acids Res;47:W191-W198).
[0257] Differential expression analysis was performed using R (v3.4.2) with edgeR (v3.20.9) after filtering out genes with a CPM (counts per million reads sequenced) of ≤0.5 in >25% of samples. Datasets were normalized using the TMM method (Debieve et al., 2011, Mol Hum Reprod;17:702-9), and differentially abundant genes were identified using a glmTreat test (Shiadeh et al., 2017, Infection;45:589-600) with a log fold change of ≥1, followed by Bonferroni-Holm p-value correction. For the iPEC data, this same process was performed for each jackknife iteration using 90% of the samples from each group selected by random sampling without replacement. After 1,000 jackknife iterations, one-sided, normal-based 95% confidence intervals for gene-specific p-values were calculated using StatsModels (v0.8.0) (Skalis et al., 2019, Microrna;8:28-35). Because only transcripts with p-values <0.05 were included, one-sided calculations were used. Hierarchical clustering analysis was performed using squared Euclidean distance and average linkage.
[0258] AdaBoost. AdaBoost was performed using scikit-learn (v0.19.1, sklearn.ensemble.AdaBoostClassifier) in Python (Burton et al., 2019, BMJ;366:12381). A holdout subset of 10% of iPEC samples was excluded from all machine learning activities. The remaining samples available for training the machine learning were filtered to remove genes with CPM ≤ 0.5 and TMM normalized in > 25% of samples. The logarithm of transcript (CPM) values were then normalized (sklearn.preprocessing.StandardScaler) to a mean of 0 and standard deviation of 1 before fitting the classifier.
[0259] Optimal hyperparameter values were determined by random search over 1,000 iterations using three-fold stratified cross-validation, and performance was quantified using the Matthews correlation coefficient (Figure 42) (Bergstra et al., 2012, J Machine Learning Res;13:281-305). The number of estimates was sampled from a geometric distribution (scipy.stats.geom, p=0.004, loc=7), and the learning rate was sampled from an exponential distribution (scipy.stats.expon, loc=0.08, scale=2). Three iterations of the search showed the best performance, and the median value for each hyperparameter was selected for further use (500 estimates and a 1.6 learning rate).
[0260] AdaBoost models were trained using 10-fold stratified cross-validation to obtain robust estimates of PE classification ability. The complete AdaBoost model training strategy is shown in Figure 43. This strategy was performed in two stages. In the first stage, five subsets of samples were created from the training data. Four subsets (80% of the training data) were combined to fit AdaBoost, and one subset (20% of the training data) was used to evaluate performance during feature pruning. During feature pruning, the performance measure was the matrix correlation coefficient, and the model with the highest value was retained. In the case of ties, the model with the fewest transcripts was retained. Due to the stochastic nature of AdaBoost, model composition changed each time a model was fitted. Therefore, to increase the likelihood of including the most robust estimates, fitting and feature pruning were repeated 10 times, generating 10 models for each subset. This process was repeated five times, each time removing one subset for pruning and combining the other four subsets for fitting. This process ultimately produces 50 total models.
[0261] In the second step, estimates from all 50 models were combined to generate a single aggregated AdaBoost model. This ensemble was then subjected to feature pruning, with the importance measure for each transcript being the number of models in which it appeared, and the performance metric being logarithmic loss. The model with the best logarithmic-scaled average performance among all pruned subsets was selected as the final ensemble.
[0262] The validation samples within each fold were completely excluded from model training but were used to construct ROC curves and determine the score threshold that maximized both sensitivity and specificity. The status of the iPEC holdout samples and the PEARL PEC independent cohort was then classified into each of 10 AdaBoost models from cross-validation.
[0263] RT-qPCR Twenty-five TaqMan probes (Table 19, Thermo Fisher Scientific) were selected to validate sequencing for a subset of patients from the iPEC cohort (N = 19 PE, N = 19 control). Five reference probes were used to normalize for differences in fold change. These targeted transcript sets were consistent between control and PE samples and covered an abundance range of 0.2–20 CPM. 20 indicates probes selected to span exon junctions.
[0264] C-RNA was isolated from 2 mL of plasma and converted to cDNA. The cDNA was preamplified for 16 cycles using TaqMan Preamp master mix (Thermo Fisher Scientific) and then diluted 10-fold. Triplicate TaqMan qPCR reactions for all probes were performed according to the manufacturer's protocol (Thermo Fisher Scientific). Cq values were determined using Bio-Rad CFX Manager software. To determine transcript abundance, the mean Cq values of the reference probes were used to calculate ΔΔCq. To determine the fold change of the PE samples for each probe, the mean ΔΔCq value of the PE samples was divided by the mean ΔΔCq value of the matched control samples.
[0265] Sample preparation protocol optimization rRNA and globin depletion. C-RNA was extracted from 2 mL of plasma and treated with DNAse before depletion. The TruSeq Total RNA Library Kit using RiboZero (Illumina, Inc.) was used according to the manufacturer's protocol. RNAse H depletion was performed according to a previously published protocol (Crescitelli et al., 2013, J Extracell Vesicles;2:doi:10.3402 / jev.v2i0.20677), except that hybridization was performed in a total volume of 6 μL with a final concentration of 125 pM depleted oligos per oligo.
[0266] C-RNA quantification. C-RNA was extracted from 4.5 mL of plasma and treated with DNAse. One-tenth of the extracted C-RNA was quantified using the Quant-iT RiboGreen RNA Kit (Thermo Fisher Scientific). C-RNA was diluted 100-fold and quantified against the manufacturer's recommended low-range standard curve.
[0267] Plasma Input Comparison. A single experiment simultaneously evaluated the use of 0.5, 1, 2, and 4 mL plasma inputs, but multiple datasets utilizing different plasma inputs were generated during protocol optimization. A meta-analysis of data from eight separate experiments was performed to assess the impact of this variable. The biological coefficient of variation (BCV, edgeR) was used to quantify noise (Debieve et al., 2011, Mol Hum Reprod;17:702-9). In all experiments, BCV measurements were obtained for each set of samples constituting biologically distinct groups. The combined population function from Preseq (v2.0.0) provided library complexity estimates for each individual sample (Xie et al., 2018, Res Commun;506:692-697). All sample preparation was performed as previously described with one exception: Libraries were generated using Accel-NGS 1 S Plus DNA Library (Swift Biosciences) for 1 mL and 0.5 mL inputs according to the manufacturer's instructions.
[0268] BCT comparison. 8 mL of blood was collected from pregnant and non-pregnant women in the following tube types: K2 EDTA (Beckton Dickinson), ACD (Beckton Dickinson), Cell-Free RNA BCT tubes (Streck), and Cell-Free DNA BCT tubes (Streck, Inc.). Blood was transported overnight either on ice (EDTA and ACD) or at room temperature (Cell-Free RNA and DNA BCT tubes). As a reference, 8 mL of blood was collected in a K2 EDTA tube, processed into plasma on-site within 4 hours, and transported as plasma on dry ice. All other blood samples were processed after transport after 1 or 5 days of storage at the desired temperature. 3 mL of plasma per condition was used to generate sequencing libraries for enrichment using the previously described Illumina protocol. To compare pregnant and non-pregnant samples, pregnancy signals were quantified using the following formula using 155 transcripts reported before previous C-RNA pregnancy studies (Tsui et al., 2014, Clin Chem; 60:954-962, and Koh et al., 2014, Proc Natl Acad Sci USA; 111:7361-7366):
[0269]
number
[0270] In the formula, i represents a single transcript, x represents the log (CPM) value of the subject sample, and m and s represent the mean and standard deviation of the non-pregnant samples taken at the same BCT, respectively.
[0271] Pregnancy time course analysis. Differential expression analysis was performed as previously described without jackknife analysis. Transcripts that changed during specific stages of pregnancy were identified as follows: CPM values for each transcript were normalized within each patient using the first trimester sample (11-14 weeks' gestational age) as the baseline. Consensus values across all patients were obtained by Lowess smoothing (statsmodels.nonparametric.smoothers_lowess.lowess). Transcripts were classified as previously altered before pregnancy if the slope of the lower curve was ≥2-fold higher in absolute magnitude at 14 weeks than at 34 weeks. Transcripts were classified as altered after pregnancy if the slope of the lower curve was ≥2-fold higher in absolute magnitude at 34 weeks than at 14 weeks. The remaining transcripts were classified as altered during pregnancy.
[0272] statistical analysis All statistical tests were two-sided unless otherwise noted. Nonparametric tests were used when data were not normally distributed. P values were adjusted for multiple comparisons via Bonferroni-Holm or Tukey HSD calculations.
[0273] [Table 12]
[0274] [Table 13]
[0275] [Table 14-1]
[0276] [Table 14-2]
[0277] [Table 14-3]
[0278] Table 14-4
[0279] Table 15
[0280] Table 16
[0281] Table 17
[0282] Table 18
[0283] Table 19
[0284] Table 20-1
[0285] Table 20-2
[0286] Table 21-1
[0287] Table 21-2
[0288] [Table 21-3]
[0289] The complete disclosures of all patents, patent applications, and publications, as well as electronically available materials cited herein (e.g., nucleotide sequence submissions in GenBank and RefSeq, amino acid sequence submissions in SwissProt, PIR, PRF, PDB, and translations from annotated coding regions in GenBank and RefSeq), are incorporated by reference in their entirety. Supplementary materials referenced in publications (such as supplementary tables, figures, materials and methods, and / or experimental data) are likewise incorporated by reference in their entirety. In the event of a conflict between the disclosure of this application and the disclosure of a document incorporated by reference herein, the disclosure of this application shall control. The foregoing detailed description and examples are provided for clarity of understanding only. No unnecessary limitations should be understood therefrom. The disclosure is not limited to the exact details shown and described, since variations obvious to those skilled in the art are included in the disclosure defined by the claims.
[0290] Unless otherwise indicated, all numbers expressing quantities of ingredients, molecular weights, and the like used in the specification and claims should be understood as being modified in all instances by the term "about." Accordingly, unless otherwise indicated, the numerical parameters set forth in the specification and claims are approximations that may vary depending upon the desired properties sought to be obtained by the present disclosure. At the very least, and not as an attempt to limit the doctrine of equivalents to the scope of the claims, each numerical parameter should, at the very least, be construed in light of the number of reported significant digits and by applying ordinary rounding techniques.
[0291] Notwithstanding that the numerical ranges and parameters setting forth the broad scope of the present disclosure are approximations, the numerical values set forth in the specific examples are reported as precisely as possible, however, all numerical values inherently contain ranges necessarily resulting from the standard deviations found in their respective testing measurements.
[0292] All headings are for the convenience of the reader and should not be used to limit the meaning of the text that follows the heading, unless specifically stated.
Claims
1. 1. A method for detecting severe early pre-eclampsia and / or determining an increased risk of severe early pre-eclampsia in a pregnant woman, said method comprising: determining in a biosample obtained from a pregnant woman the level of circulating RNA (C-RNA) molecules encoding at least a portion of the protein AKAP2; the biosample comprises blood, plasma, or serum; A method wherein identification of a level of a C-RNA molecule encoding at least a portion of the protein AKAP2 that is increased by at least 1.95-fold relative to the level in gestational age-matched control pregnant human women without pre-eclampsia is indicative of severe early-stage pre-eclampsia and / or an increased risk of severe early-stage pre-eclampsia in the pregnant woman.
2. 1. A method for detecting severe early pre-eclampsia and / or determining an increased risk of severe early pre-eclampsia in a pregnant woman, said method comprising: purifying a population of circulating RNA (C-RNA) molecules from a biosample obtained from said pregnant woman, said biosample comprising blood, plasma or serum; identifying protein-coding sequences encoded by said C-RNA molecules within said purified population of C-RNA molecules; Including, A method wherein identification of a C-RNA molecule encoding at least a portion of the protein AKAP2 that is increased by at least 1.95-fold relative to levels in gestational age-matched control pregnant human women without pre-eclampsia is indicative of severe early-stage pre-eclampsia and / or an increased risk of severe early-stage pre-eclampsia in the pregnant woman.
3. 3. The method of claim 1, wherein identifying the protein-coding sequence encoded by the C-RNA molecule in the biosample comprises hybridization, reverse transcriptase PCR, microarray chip analysis, or sequencing.
4. The method of claim 1 or 2, wherein identifying the protein-coding sequence encoded by the C-RNA molecule in the biosample comprises sequencing.
5. 5. The method of claim 4, wherein the sequencing comprises massively parallel sequencing of clonally amplified molecules.
6. The method of claim 4 or 5, wherein the sequencing comprises RNA sequencing.
7. Prior to identifying the protein-coding sequence encoded by said circular RNA (C-RNA) molecule, removing intact cells from the biosample; treating the biosample with deoxyribonuclease (DNAse) to remove cell-free DNA (cfDNA); synthesizing complementary DNA (cDNA) from the C-RNA molecules in the biosample; and / or Enriching cDNA sequences of protein-coding DNA sequences by exome enrichment; The method of any one of claims 1 to 6, further comprising:
8. 1. A method for detecting severe early pre-eclampsia and / or determining an increased risk of severe early pre-eclampsia in a pregnant woman, said method comprising: removing intact cells from a biosample obtained from said pregnant woman, said biosample comprising blood, plasma or serum; treating the biosample with deoxyribonuclease (DNAse) to remove cell-free DNA (cfDNA); synthesizing complementary DNA (cDNA) from RNA molecules in the biosample; enriching said cDNA sequences for DNA sequences encoding proteins (exome enrichment); sequencing the obtained enriched cDNA sequences; and Identifying protein-coding sequences encoded by enriched C-RNA molecules; Including, A method wherein identification of a C-RNA molecule encoding at least a portion of the protein AKAP2 that is increased by at least 1.95-fold relative to levels in gestational age-matched control pregnant human women without pre-eclampsia is indicative of severe early-stage pre-eclampsia and / or an increased risk of severe early-stage pre-eclampsia in the pregnant woman.
9. The method of any one of claims 1 to 8, wherein the biosample comprises plasma.
10. 10. The method of any one of claims 1 to 9, wherein the biosample is obtained from a pregnant woman at a gestational age of at least 20 weeks.
11. 11. The method of any one of claims 1 to 10, wherein the sample is a blood sample and the blood sample is collected, transported and / or stored in a tube with cell and DNA stabilizing properties before the blood sample is processed into plasma.
12. 12. The method of claim 11, wherein the tube comprises a Streck Cell-Free DNA BCT® blood collection tube.
13. the biosample is a blood sample that is processed into plasma; The blood sample Not exposed to EDTA before being processed into plasma; processed into plasma within about 24 to about 72 hours of blood collection; maintained, stored, and / or transported at room temperature before being processed into plasma; and / or maintained, stored, and / or transported without exposure to low temperatures or freezing prior to being processed into plasma; The method according to any one of claims 1 to 12.
14. 14. The method of any one of claims 1 to 13, further comprising detecting a C-RNA molecule encoding at least a part of the proteins ARRB1, CPSF7, INO80C, JAG1, MSMP, NR4A2, PLEK, RAP1GAP2, SPEG, TRPS1, UBE2Q1, and / or ZNF768 in a biosample obtained from the pregnant woman.
Citation Information
Patent Citations
Serum mRNA as a diagnostic marker for pregnancy disorders
JP2006515517A
Markers for prenatal diagnosis and monitoring
JP2008524993A
Transcriptome analysis of maternal plasma by ultra-parallel RNA sequencing
JP2016511641A
Gene expression related to preeclampsia
US20110171650A1
Additional circulating RNA signatures specific to preeclampsia
US62939324P0