Compositions and methods for detecting gynecological cancer
Novel DMRs in DNA samples improve the detection of gynecological cancers by distinguishing between different types and subtypes, enabling early diagnosis and personalized treatment.
Patent Information
- Application Number
- JP2025513305
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-02
- Filing Date
- 2023-09-01
- Publication Date
- 2025-09-17
AI Technical Summary
Current screening methods for gynecological cancers are limited, particularly for types other than cervical cancer, leading to late diagnoses and inadequate treatment strategies due to the lack of reliable diagnostic tools for multiple gynecological cancers in a single biological sample.
The use of novel differentially methylated regions (DMRs) in DNA samples to distinguish between various types and subtypes of gynecological cancers, including cervical, ovarian, and endometrial cancers, through methylation-specific analysis of CpG sites in specific genes and regions.
Enhances the sensitivity and accuracy of cancer detection by identifying multiple gynecological cancers early, allowing for more precise patient stratification and tailored treatment strategies.
Smart Images

Figure 2025530795000026 
Figure 2025530795000027 
Figure 2025530795000028
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 374,415, filed September 2, 2022, which is incorporated herein by reference in its entirety for all purposes.
[0002] Incorporation by Reference of Electronically Submitted Materials The computer-readable nucleotide / amino acid sequence listing, submitted concurrently herewith and identified as follows: one 57,580 byte ASCII (text) file (filename "40960-601_SEQUENCE_LISTING", created September 1, 2023), is hereby incorporated by reference in its entirety.
[0003] The present disclosure relates to the detection of one or more types of gynecological cancer in a biological sample from a subject. In particular, the present disclosure provides compositions and methods for detecting the presence or absence of one or more types of gynecological cancer (e.g., cervical cancer, ovarian cancer, endometrial cancer) in a biological sample from a subject having or suspected of having gynecological cancer. [Background technology]
[0004] Compared with other types of cancer (e.g., breast or colon cancer), gynecological cancers are less common, affecting approximately 100,000 women each year in the United States. However, all women are at risk of developing gynecological cancer, and this risk increases with age. The five major types of gynecological cancer include cervical, ovarian, uterine, vaginal, and vulvar cancers. The sixth type of gynecological cancer is fallopian tube cancer, which is extremely rare. Despite evidence that early detection of gynecological cancers is particularly important for improving survival, currently, only cervical cancer has a clinically relevant screening test available. The lack of simple and reliable methods for screening for multiple gynecological cancers makes recognizing precursors and reducing risk particularly important. Furthermore, while gynecological cancer screening programs (e.g., HPV testing, Pap testing) aim to improve survival rates through early detection, these tests are typically available only to a subset of the population (e.g., those at highest risk) and are limited to a small number of cancers (e.g., cervical cancer). Therefore, medical professionals are often only able to make an accurate cancer diagnosis after symptoms occur, which may result in effective treatment being administered too late.Therefore, there is an urgent need for improved diagnostic tools to detect multiple types or subtypes of gynecological cancer in a single biological sample, not only for earlier detection but also for more accurate patient stratification and to provide greater insight into treatment strategies. Summary of the Invention
[0005] Embodiments of the present disclosure provide methods, compositions, and systems for screening for multiple types of gynecological cancer from a biological sample. According to these embodiments, the present disclosure includes, but is not limited to, methods and compositions for detecting the presence of multiple types or subtypes of gynecological cancer from a biological sample. In some embodiments, the biological sample is a tissue sample, a blood sample, a plasma sample, a serum sample, a whole blood sample, a secretion sample, an organ secretion sample, a cerebrospinal fluid (CSF) sample, a saliva sample, a urine sample, and / or a stool sample. In some embodiments, the tissue sample is a gynecological tissue sample including one or more of vaginal tissue, vaginal cells, cervical tissue, cervical cells, endometrial tissue, endometrial cells, ovarian tissue, and ovarian cells. In some embodiments, the tissue sample is an ovarian tissue sample, an endometrial tissue sample, or a cervical tissue sample. In some embodiments, the secretion sample is a gynecological secretion sample. In some embodiments, the subject is a human.
[0006] As further described herein, embodiments of the present disclosure include novel differentially methylated regions (DMRs), each of which can independently distinguish between specific types of gynecological cancer (i.e., endometrial cancer (EC), or ovarian cancer (OC), or cervical cancer (CC)) and benign gynecological tissue samples. According to these embodiments, these novel DMRs include ADAM8, ADHFE1, AES, AGBL2, AIM1, AK5, ALKBH3, ARAP1, ARHGAP20, ASCL2, BCAT1, BEGAIN, BEND4_3696, BMP6, C12orf68, C13orf18, C14orf169_7694, C14orf169_8382, C18orf18, C1orf61, C20orf195, C4orf31, C5orf52, C6orf147, C7orf169_8382, C8orf169_9382, C9orf169_10382, C10orf169_11382, C11orf169_12382, C12orf68, C13orf18, C14orf169_13382, C14orf169_14382, C15orf169_15382, C16orf169_16382, C17orf169_17382, C18orf18, C1orf61, C20orf195, C4orf31, C5orf52, C6orf147, C7orf169_11382, C8orf169_12382, C8orf169_13382, C8orf169_14382, C8orf169_ rf51, CD14, CELF2, CHCHD5, CHMP2A, CHST10, CLIC6, CLIP4, COL13A1, COL19A1, COL6A2, COPZ2, CREB3L1, CXCL2, CXXC5, CYTH2, DAB2 IP, DGKZ, DLGAP3, DNASE2, DSCAML1, EBF1, EDARADD, EGR2, EIF5A2, ELMO1, ELMOD1, ELOVL4, EME2, EML6, EPSTI1, FADS2, FAM109B, FA M126A, FAM174B, FGF18, FKBP11, FLI1, FLOT1, FOXD3, FYN, GAL3ST2, GALR3, GAS7, GATA2_5878, GLT25D2, GNB2, HDAC7, HIC1, HLA-F, HNRNPF, HPDL, HS3ST4, HSPA1A, IDUA, IGSF9B, IL12RB2, IRAK3, IRF7, IRF8, ITPKA, KCNA2, KCNC3_6487, KCNC3_7105, KCNC4, KCNH8, KDM2B, LBX2, LCMT2, LOC100129726, LOC100287216, LOC255130, LOC339290, LOC729678, LPPR3, LRRC41, LRRC8D_8856, LTBP2, LYPL AL1, MAST4, MAX.chr1.2152, HIVEP3, GRAMD1B, MAX.chr11.0394, MAX.chr11.3750, FAT3, SLC16A7, MTUS2, LINC02323, MAX.chr14.7696、MCTP2、LOC107984974、TRIM80P、MAX.chr19.5552、ZNF433-AS1、ZNF254、MAX.chr19.0548、B3GALT1、MAX.chr2.8918、MAX.chr2.4778、MAX.chr20.3853、MAX.chr20.2903、MAX.chr21.5011、DSCR9、MAX.chr22.5665、MAX.chr3.6408、LINC02028、LINC02084、MAX.chr5.3588、CTD-2532K18.1、HS3ST5、ARHGAP18、GRM4、LINC01004、MAX.chr8.5938、MAX.chr9.4007、MAX.chr9.2025、TRPM3、MED12L、MIAT、MLH1_4513、MLH1_5193、MMP16、MRPS21、MSI1、MT1E、MX1、MYC、MYH10、MYO15B、N4BP 2L1、NBR1、NDRG2、NEGR1、NEU1、NOL3、NR3C1_2223、NR3C1_4614、NRP2、NTN1、NTNG1、PAPL、PAQR9、PDE10A、PDE3B , PDE4A, PDXK, PER1, PISD, PLEC, PLIN2, PLXND1, PPM1E, PPP1R9A, PPP2R5C, PRDM5, PTP4A3, PYCARD, RAB3C, RAI1, RARG, RASA3, RPRM, RREB1, S100A6, SAMD5, SBNO2, SDC2, SDK2, SELM, SERP2, SFMBT2_2029, SHF, SHH, SLC16A11, SLC16A5、SLC25A22、SLCO3A1、SMTN、SPDYA、SPINK2、SPOCK2、SPON1、SQSTM1_4156、ST8SIA1、TAF4B、TAF7、TEAD3 、TERC、TIAM1、TLE4、TMEM101、TMEM106A、TRIM9、TRPC3、TSC22D4、TSPAN2、TSPAN5、TTC14、UBB_4001、UBB_4646、 UST、VAMP5、VIM、VSTM2B、ZBTB7B、ZEB2、ZFP3、ZFP36L2、ZIC2、ZMIZ1、ZNF14、ZNF211、ZNF280B、ZNF302、ZNF382、 ZNF480、ZNF483、ZNF491、ZNF569、ZNF610、ZNF702P、ZNF709、ZNF773、ZNF845、ZNF91、CDH4、LRRC34、MAX.chr10.The novel DMR(s) comprise one or more CpG sites in 4460, NBPF24, OBSCN, SEPT9, ZNF323, ZNF506, and / or ZNF90 (Table 1), including any combination thereof. In some embodiments, the novel DMR(s) are derived from any gene or region selected from Table 1, including any combination thereof. While each novel DMR alone can distinguish one or more gynecological cancers from control samples, combining two or more of the novel DMRs may improve sensitivity. Thus, combinations of two or more novel DMRs selected from Table 1 are provided.
[0007] Embodiments of the present disclosure also include novel differentially methylated regions (DMRs) that can individually distinguish gynecological cancer from benign gynecological tissue samples, and these DMRs are prevalent in all three types of gynecological cancers (i.e., endometrial cancer (EC), ovarian cancer (OC), and cervical cancer (CC)).According to these embodiments, these novel DMRs include ACSF2, AJAP1, ARL10, ARL5C, ASCL4, ATP6V1B1, BARHL1, BEND4_2963, C17orf64, C1QL3, C2orf55, C4orf48, CA3, CDO1, CELF2, CLEC14A, CSDAP1, CYTH2_4197, DLGAP1, DSCR6, EPS8L1_2819, EPS8L1_8496, FAIM2, FGF12, GATA2, HIST1H2BE, IRF4, IRX4, ITGA5, KCNA1, LECT1, LHX1, LOC440925, LPHN1, LINC02767, MAX.chr1.2533, SOX1-OT, MAX.chr13.3357, MAX.chr14.2093, MAX.c hr17.2455, MAX.chr18.4390, MAX.chr19.2732, MAX.chr19.4467, PANTR1, MAX.chr2.0490, MAX.chr2.8148, MAX.chr2.3137, RIPOR3, S CRG1, MAX.chr4.4210, HMX1, CTC-359M8.1, MAX.chr5.0931, MAX.chr5.9924, LIN28B, MAX.chr6.9522, TTLL2, RNA5SP243, DLGAP2, MEX 3B, MNX1, NEFL, NETO1, PAX2, PDX1, psiTPTE22, RASGEF1A, SALL3_9136, SALL3_0615, SEZ6L2, SHANK2, SHANK3, SKI, SLC35D3, SORCS3_03 In some embodiments, the novel DMR(s) comprise one or more CpG sites in any gene or region selected from Table 2, including any combination thereof.Each novel DMR alone can distinguish one or more gynecological cancers from control samples, and combining two or more of the novel DMRs may improve sensitivity. Thus, combinations of two or more novel DMRs selected from Table 2 are provided.
[0008] Embodiments of the present disclosure also include novel differentially methylated regions (DMRs), each of which can individually distinguish specific subtypes of gynecological cancer (i.e., serous ovarian cancer, clear cell ovarian cancer, endometrioid ovarian cancer, mucinous ovarian cancer, adenocarcinoma cervical cancer, squamous cervical cancer, or endometrioid endometrial cancer) from benign gynecological tissue samples. According to these embodiments, the novel DMRs comprise one or more CpG sites in AIM1, AK5, c18orf18, CDO1, DLGAP1, ELMOD1, FKBP11, FLOT1, GAL3ST2, LRRC41, LYPLAL1, MAX.chr11.3750, MLH1_4513, NR3C1_2223, PISD.RABC3, RAI1, TERC, TRPC3, ZIC2, ZMIZ1, ZNF480, ZNF491, ZNF610, and / or ZNF91 (Table 3), including any combination thereof. In some embodiments, the novel DMR(s) are derived from any gene or region selected from Table 3, including any combination thereof. While each novel DMR alone can distinguish one or more gynecological cancers from control samples, combining two or more of the novel DMRs may improve sensitivity. Accordingly, combinations of two or more novel DMRs selected from Table 3 are provided.
[0009] Embodiments of the present disclosure also include novel differentially methylated regions (DMRs), each of which can individually distinguish between specific subtypes of gynecological cancer (i.e., serous ovarian cancer, clear cell ovarian cancer, endometrioid ovarian cancer, mucinous ovarian cancer, cervical adenocarcinoma, cervical squamous cell carcinoma, or endometrioid endometrioid carcinoma) and benign gynecological tissue samples. According to these embodiments, these novel DMRs comprise one or more CpG sites in LBX2, SPDYA, TERC, ZSCAN12, CYP26C1, and / or GYPC (Table 4), including any combination thereof. In some embodiments, the novel DMR(s) are derived from any gene or region selected from Table 4, including any combination thereof. Each novel DMR alone can distinguish between one or more gynecological cancer and control samples, and combining two or more of the novel DMRs may improve sensitivity. Accordingly, combinations of two or more novel DMRs selected from Table 4 are provided.
[0010] Embodiments of the present disclosure also include novel differentially methylated regions (DMRs), each of which can individually distinguish specific subtypes of gynecological cancer (i.e., serous ovarian cancer, clear cell ovarian cancer, endometrioid ovarian cancer, mucinous ovarian cancer, cervical adenocarcinoma, cervical squamous cell carcinoma, or endometrioid endometrioid carcinoma) from benign gynecological tissue samples. According to these embodiments, the novel DMRs comprise one or more CpG sites in KRT86, CDH4, c17orf64, EMX2OS, NBPF24, SFMBT2_0970, JSRP1, DIDO1, MAX.chr10.4460, MPZ, ZNF506, GATA2_6370, VILL, LINC02323, CYTH2_4043, LRRC8D_8831, LYPLAL1, SMPD5, SQSTM1_3864, ZNF323, OBSCN, ZNF90, LRRC34, GDF7, MDFI, EEF1A2, LRRC41, and / or SEPT9 (Table 8), including any combination thereof. In some embodiments, the novel DMR(s) are derived from any gene or region selected from Table 8, including any combination thereof. Each novel DMR alone can distinguish one or more gynecological cancers from control samples, and combining two or more of the novel DMRs may improve sensitivity. Thus, combinations of two or more novel DMRs selected from Table 8 are provided.
[0011] Embodiments of the present disclosure include methods for characterizing a biological sample, comprising determining a methylation profile in at least one differentially methylated region (DMR) of a DNA sample obtained from a subject having or suspected of having gynecological cancer by treating the sample with a reagent that modifies DNA in a methylation-specific manner.
[0012] In some embodiments, the methylation profile in at least one DMR indicates that the subject has or is suspected of having at least one of ovarian cancer (OC), cervical cancer (CC), and endometrial cancer (EC).
[0013] In some embodiments, at least one DMR is selected from the group consisting of ADAM8, ADHFE1, AES, AGBL2, AIM1, AK5, ALKBH3, ARAP1, ARHGAP20, ASCL2, BCAT1, BEGAIN, BEND4_3696, BMP6, C12orf68, C13orf18, C14orf169_7694, C14orf169_8382, C18orf18, C1orf61, C20orf195, C4orf31, C5orf52, C6orf147, C7orf51, CD14, CELF2, CHCHD5, CHMP2A, CHS T10, CLIC6, CLIP4, COL13A1, COL19A1, COL6A2, COPZ2, CREB3L1, CXCL2, CXXC5, CYTH2, DAB2IP, DGKZ, DLGAP3, DNASE2, DSCAML1, EBF1, EDARADD, EGR2, EI F5A2, ELMO1, ELMOD1, ELOVL4, EME2, EML6, EPSTI1, FADS2, FAM109B, FAM126A, FAM174B, FGF18, FKBP11, FLI1, FLOT1, FOXD3, FYN, GAL3ST2, GALR3, GAS7, GATA2_5878, GLT25D2, GNB2, HDAC7, HIC1, HLA-F, HNRNPF, HPDL, HS3ST4, HSPA1A, IDUA, IGSF9B, IL12RB2, IRAK3, IRF7, IRF8, ITPKA, KCNA2, KCNC3_6487 , KCNC3_7105, KCNC4, KCNH8, KDM2B, LBX2, LCMT2, LOC100129726, LOC100287216, LOC255130, LOC339290, LOC729678, LPPR3, LRRC41, LRRC8D_8856, LTB P2, LYPLAL1, MAST4, MAX.chr1.2152, HIVEP3, GRAMD1B, MAX.chr11.0394, MAX.chr11.3750, FAT3, SLC16A7, MTUS2, LINC02323, MAX.chr14.7696, MCTP2 , LOC107984974, TRIM80P, MAX.chr19.5552, ZNF433-AS1, ZNF254, MAX.chr19.0548, B3GALT1, MAX.chr2.8918, MAX.chr2.4778, MAX.chr20.3853, MAX.chr20.2903, MAX.chr21.5011, DSCR9, MAX.chr22.5665, MAX.chr3.6408. LINC02028, LINC02084, MAX.chr5.3588, CTD-2532K18.1, HS3ST5, ARHGAP 18. GRM4, LINC01004, MAX.chr8.5938, MAX.chr9.4007, MAX.chr9.2025, T.S RPM3, MED12L, MIAT, MLH1_4513, MLH1_5193, MMP16, MRPS21, MSI1, MT1E, M X1, MYC, MYH10, MYO15B, N4BP2L1, NBR1, NDRG2, NEGR1, NEU1, NOL3, NR3C1_ 2223, NR3C1_4614, NRP2, NTN1, NTNG1, PAPL, PAQR9, PDE10A, PDE3B, PDE4A PDXK, PER1, PISD, PLEC, PLIN2, PLXND1, PPM1E, PPP1R9A, PPP2R5C, PRDM5 PTP4A3, PYCARD, RAB3C, RAI1, RARG, RASA3, RPRM, RREB1, S100A6, SAMD5. SBNO2, SDC2, SDK2, SELM, SERP2, SFMBT2_2029, SHF, SHH, SLC16A11, SLC1 A5, SLC25A22, SLCO3A1, SMTN, SPDYA, SPINK2, SPOCK2, SPON1, SQSTM1_415 6. ST8SIA1, TAF4B, TAF7, TEAD3, TERC, TIAM1, TLE4, TMEM101, TMEM106A, T RIM9, TRPC3, TSC22D4, TSPAN2, TSPAN5, TTC14, UBB_4001, UBB_4646, UST. VAMP5, VIM, VSTM2B, ZBTB7B, ZEB2, ZFP3, ZFP36L2, ZIC2, ZMIZ1, ZNF14, ZN F211, ZNF280B, ZNF302, ZNF382, ZNF480, ZNF483, ZNF491, ZNF569, ZNF610 ZNF702P ZNF709 ZNF773 ZNF845 ZNF91 CDH4 LRRC34 MAX.chr10.446 0, NBPF24, OBSCN, SEPT9, ZNF323, ZNF506, and / or ZNF90 contain 1 CpG site.
[0014] In some embodiments, at least one DMR is selected from the group consisting of ACSF2, AJAP1, ARL10, ARL5C, ASCL4, ATP6V1B1, BARHL1, BEND4_2963, C17orf64, C1QL3, C2orf55, C4orf48, CA3, CDO1, CELF2, CLEC14A, CSDAP1, CYTH2_4197, DLGAP1, DSCR6, EPS8L1_2819, EPS8L1_8496, FAIM2, FGF12, GATA2, HIST1H2BE , IRF4, IRX4, ITGA5, KCNA1, LECT1, LHX1, LOC440925, LPHN1, LINC02767, MAX.chr1.2533, SOX1-OT, MAX.chr13.3357, MAX.chr14.20 93, MAX.chr17.2455, MAX.chr18.4390, MAX.chr19.2732, MAX.chr19.4467, PANTR1, MAX.chr2.0490, MAX.chr2.8148, MAX.chr2.31 37, RIPOR3, SCRG1, MAX.chr4.4210, HMX1, CTC-359M8.1, MAX.chr5.0931, MAX.chr5.9924, LIN28B, MAX.chr6.9522, TTLL2, RNA5SP2 43, DLGAP2, MEX3B, MNX1, NEFL, NETO1, PAX2, PDX1, psiTPTE22, RASGEF1A, SALL3_9136, SALL3_0615, SEZ6L2, SHANK2, SHANK3, SKI, S Containing one or more CpG sites in LC35D3, SORCS3_0305, SORCS3_1038, SOX1, SQSTM1, TBXT, TCERG1L, TERT, TNFSF11, TUBB6, ULBP1, VAC14, VWC2, WDR69, ZBTB16, ZNF132, ZSCAN12, ZSCAN23, KRT86, CYP26C1, GYPC, DIDO1, EEF1A2, EMX2OS, GDF7, JSRP1, SMPD5, MDFI, MPZ, and / or VILL.
[0015] In some embodiments, at least one DMR comprises one or more CpG sites in AIM1, FLOT1, GAL3ST2, LRRC41, LYPLAL1, MAX.chr11.3750, PISD, RAI1, ZIC2, ZMIZ1, CDH4, ZNF506, ZNF323, OBSCN, ZNF90, and / or SEPT9, and the subject has or is suspected of having OC. In some embodiments, at least one DMR comprises one or more CpG sites in AIM1, FLOT1, GAL3ST2, LYPLAL1, and / or OBSCN, and the subject has or is suspected of having serous OC. In some embodiments, at least one DMR comprises one or more CpG sites in LRRC41, PISD, ZIC2, OBSCN, and / or SEPT9, and the subject has or is suspected of having clear cell OC. In some embodiments, at least one DMR includes one or more CpG sites in MAX.chr11.3750, and the subject has or is suspected of having endometrioid OC. In some embodiments, at least one DMR includes one or more CpG sites in RAI1 and / or ZMIZ1, and the subject has or is suspected of having mucinous OC. In some embodiments, determining the methylation profile of one or more CpG sites in AIM1, FLOT1, GAL3ST2, LRRC41, LYPLAL1, MAX.chr11.3750, PISD, RAI1, ZIC2, ZMIZ1, CDH4, ZNF506, ZNF323, OBSCN, ZNF90, and / or SEPT9 comprises comparing the methylation profile to corresponding regions of a control DNA sample obtained from a subject without OC.
[0016] In some embodiments, at least one DMR comprises one or more CpG sites in AK5, ELMOD1, RABC3, TRPC3, ZNF480, ZNF491, ZNF610, ZNF91, and / or NBPF24, and the subject has or is suspected of having CC. In some embodiments, at least one DMR comprises one or more CpG sites in AK5, ELMOD1, TRPC3, and / or ZNF480, and the subject has or is suspected of having cervical adenocarcinoma (adenocarcinoma CC). In some embodiments, at least one DMR comprises one or more CpG sites in ZNF491, ZNF610, ZNF91, and / or NBPF24, and the subject has or is suspected of having squamous cell CC. In some embodiments, determining the methylation profile of one or more CpG sites in AK5, ELMOD1, RABC3, TRPC3, ZNF480, ZNF491, ZNF610, and / or ZNF91 comprises comparing the methylation profile to corresponding regions of a control DNA sample obtained from a subject who does not have CC.
[0017] In some embodiments, at least one DMR comprises one or more CpG sites in c18orf18, FKBP11, MLH1, NR3C1, and / or TERC, and the subject has or is suspected of having EC. In some embodiments, at least one DMR comprises one or more CpG sites in MLH1 and / or SEPT9, and the subject has or is suspected of having clear cell EC. In some embodiments, at least one DMR comprises one or more CpG sites in NR3C1, and the subject has or is suspected of having endometrioid EC.
[0018] In some embodiments, determining the methylation profile of one or more CpG sites in c18orf18, FKBP11, MLH1, NR3C1, and / or TERC comprises comparing the methylation profile with corresponding regions of a control DNA sample obtained from a subject without EC.
[0019] In some embodiments, at least one DMR comprises one or more CpG sites in CDO1 and / or DLGAP1, and the subject has or is suspected of having CC, OC, or EC. In some embodiments, determining the methylation profile of the one or more CpG sites in CDO1 and / or DLGAP1 comprises comparing the methylation profile to corresponding regions of a control DNA sample obtained from a subject who does not have CC, OC, or EC.
[0020] In some embodiments, the method further comprises determining a methylation profile of one or more CpG sites in AIM1, FLOT1, GAL3ST2, LRRC41, LYPLAL1, MAX.chr11.3750, PISD, RAI1, ZIC2, and / or ZMIZ1. In some embodiments, the method further comprises determining the methylation profile of one or more CpG sites in AK5, ELMOD1, RABC3, TRPC3, ZNF480, ZNF491, ZNF610, and / or ZNF91. In some embodiments, the method further comprises determining the methylation profile of one or more CpG sites in c18orf18, FKBP11, MLH1, NR3C1, and / or TERC.
[0021] In some embodiments, at least one DMR comprises one or more CpG sites in NBPF24, and the subject has or is suspected of having CC. In some embodiments, determining the methylation profile of the one or more CpG sites in NBPF24 comprises comparing the methylation profile with a corresponding region of a control DNA sample obtained from a subject who does not have CC.
[0022] In some embodiments, at least one DMR includes one or more CpG sites in CDH4, NBPF24, MAX.chr10.4460, ZNF506, ZNF323, OBSCN, ZNF90, LRRC34, SFMBT2, LINC02323, CYTH2, LRRC8D, LYPLAL1, LRRC41, and / or SEPT9, and the subject has or is suspected of having EC. In some embodiments, determining the methylation profile of the one or more CpG sites in CDH4, NBPF24, MAX.chr10.4460, ZNF506, ZNF323, OBSCN, ZNF90, LRRC34, SFMBT2, LINC02323, CYTH2, LRRC8D, LYPLAL1, LRRC41, and / or SEPT9 comprises comparing the methylation profile to corresponding regions of a control DNA sample obtained from a subject without EC.
[0023] In some embodiments, at least one DMR includes one or more CpG sites in CDH4, ZNF506, ZNF323, OBSCN, ZNF90, SFMBT2, LINC02323, CYTH2, LRRC8D, LYPLAL1, LRRC41, and / or SEPT9, and the subject has or is suspected of having OC. In some embodiments, determining the methylation profile of the one or more CpG sites in CDH4, ZNF506, ZNF323, OBSCN, ZNF90, SFMBT2, LINC02323, CYTH2, LRRC8D, LYPLAL1, LRRC41, and / or SEPT9 includes comparing the methylation profile to corresponding regions of a control DNA sample obtained from a subject without OC.
[0024] In some embodiments, at least one DMR includes one or more CpG sites in KRT86, EMX2OS, JSRP1, DIDO1, MPZ, VILL, SMPD5, GDF7, MDFI, c17orf64, GATA2, SQSTM1, and / or EEF1A2, and the subject has or is suspected of having CC, OC, or EC. In some embodiments, determining the methylation profile of the one or more CpG sites in KRT86, EMX2OS, JSRP1, DIDO1, MPZ, VILL, SMPD5, GDF7, MDFI, c17orf64, GATA2, SQSTM1, and / or EEF1A2 includes comparing the methylation profile to corresponding regions of a control DNA sample obtained from a subject without CC, OC, or EC.
[0025] In some embodiments, at least one DMR is associated with an area under the ROC curve (AUC) of 0.8 or greater, and the ROC curve distinguishes between subjects having or suspected of having OC, CC, or EC and control samples.
[0026] In some embodiments, the biological sample is selected from a tissue sample, a blood sample, a plasma sample, a serum sample, a whole blood sample, a secretion sample, an organ secretion sample, a cerebrospinal fluid (CSF) sample, a saliva sample, a urine sample, and a stool sample. In some embodiments, the tissue sample is a gynecological tissue sample. In some embodiments, the gynecological tissue sample comprises one or more of vaginal tissue, vaginal cells, cervical tissue, cervical cells, endometrial tissue, endometrial cells, ovarian tissue, and ovarian cells. In some embodiments, the tissue sample is an ovarian tissue sample, an endometrial tissue sample, or a cervical tissue sample. In some embodiments, the secretion sample is a gynecological secretion sample. In some embodiments, the subject is a human.
[0027] In some embodiments, a biological sample is obtained from the subject, and the method further comprises extracting a DNA sample from the biological sample. In some embodiments, the biological sample is collected with a collection device having an absorbent member capable of collecting the biological sample upon contact. In some embodiments, the absorbent member is a sponge configured for insertion into an orifice. In some embodiments, the collection device is selected from a tampon, a lavage that releases liquid into the vagina and recollects the fluid, a cervical brush, a Fournier cervical self-sampling device, and a swab.
[0028] In some embodiments, the reagent that modifies DNA in a methylation-specific manner is a borane reducing agent. In some embodiments, the reagent that modifies DNA in a methylation-specific manner comprises one or more of a methylation-sensitive restriction enzyme, a methylation-dependent restriction enzyme, and a bisulfite reagent.
[0029] In some embodiments, determining the methylation profile of at least one DMR comprises amplifying at least a portion of the DMR using a set of primers.
[0030] In some embodiments, determining the methylation profile of at least one DMR comprises performing at least one of methylation-specific PCR, quantitative methylation-specific PCR, methylation-specific DNA restriction enzyme analysis, quantitative bisulfite pyrosequencing, flap endonuclease assay, PCR-flap assay, and bisulfite genomic sequencing PCR.
[0031] In some embodiments, determining the methylation profile of at least one DMR comprises determining the presence or absence of methylation at CpG sites.
[0032]
[0010] Embodiments of the present disclosure also include methods for identifying gynecological cancer. According to these embodiments, the method includes determining a methylation profile in at least one differentially methylated region (DMR) of a DNA sample obtained from a subject having or suspected of having gynecological cancer by treating the sample with a reagent that modifies DNA in a methylation-specific manner. In some embodiments, the at least one DMR includes one or more CpG sites in AIM1, FLOT1, GAL3ST2, LRRC41, LYPLAL1, MAX.chr11.3750, PISD, RAI1, ZIC2, and / or ZMIZ1, and the methylation profile indicates that the subject has ovarian cancer. In some embodiments, the method further includes treating the subject with an anti-cancer therapy.
[0033]
[0010] Embodiments of the present disclosure also include methods for identifying gynecological cancer. According to these embodiments, the method includes determining a methylation profile in at least one differentially methylated region (DMR) of a DNA sample obtained from a subject having or suspected of having gynecological cancer by treating the sample with a reagent that modifies DNA in a methylation-specific manner. In some embodiments, the at least one DMR includes one or more CpG sites in AK5, ELMOD1, RABC3, TRPC3, ZNF480, ZNF491, ZNF610, and / or ZNF91, and the methylation profile indicates that the subject has cervical cancer. In some embodiments, the method further includes treating the subject with an anti-cancer therapy.
[0034]
[0010] Embodiments of the present disclosure also include methods for identifying gynecological cancer. According to these embodiments, the method includes determining a methylation profile in at least one differentially methylated region (DMR) of a DNA sample obtained from a subject having or suspected of having gynecological cancer by treating the sample with a reagent that modifies DNA in a methylation-specific manner. In some embodiments, the at least one DMR includes one or more CpG sites in c18orf18, FKBP11, MLH1, NR3C1, and / or TERC, and the methylation profile indicates that the subject has endometrial cancer. In some embodiments, the method further includes treating the subject with an anti-cancer therapy.
[0035]
[0010] Embodiments of the present disclosure also include methods for identifying gynecological cancer. According to these embodiments, the method includes determining a methylation profile in at least one differentially methylated region (DMR) of a DNA sample obtained from a subject having or suspected of having gynecological cancer by treating the sample with a reagent that modifies DNA in a methylation-specific manner. In some embodiments, the at least one DMR includes one or more CpG sites in CDO1 and / or DLGAP1, and the methylation profile indicates that the subject has ovarian cancer, cervical cancer, or endometrial cancer. In some embodiments, the method further includes treating the subject with an anti-cancer therapy.
[0036] Embodiments of the present disclosure also include methods for identifying gynecological cancer. According to these embodiments, the method includes determining a methylation profile in at least one differentially methylated region (DMR) of a DNA sample obtained from a subject with or suspected of having gynecological cancer by treating the sample with a reagent that modifies DNA in a methylation-specific manner. In some embodiments, the at least one DMR includes one or more CpG sites in NBPF24, and the methylation profile indicates that the subject has cervical cancer. In some embodiments, the method further includes treating the subject with an anti-cancer therapy.
[0037]
[0010] Embodiments of the present disclosure also include methods for identifying gynecological cancer. According to these embodiments, the method includes determining a methylation profile in at least one differentially methylated region (DMR) of a DNA sample obtained from a subject having or suspected of having gynecological cancer by treating the sample with a reagent that modifies DNA in a methylation-specific manner. In some embodiments, the at least one DMR includes one or more CpG sites in CDH4, NBPF24, MAX.chr10.4460, ZNF506, ZNF323, OBSCN, ZNF90, LRRC34, SFMBT2, LINC02323, CYTH2, LRRC8D, LYPLAL1, LRRC41, and / or SEPT9, and the methylation profile indicates that the subject has endometrial cancer. In some embodiments, the method further includes treating the subject with an anti-cancer therapy.
[0038]
[0010] Embodiments of the present disclosure also include methods for identifying gynecological cancer. According to these embodiments, the method includes determining a methylation profile in at least one differentially methylated region (DMR) of a DNA sample obtained from a subject having or suspected of having gynecological cancer by treating the sample with a reagent that modifies DNA in a methylation-specific manner. In some embodiments, the at least one DMR includes one or more CpG sites in CDH4, ZNF506, ZNF323, OBSCN, ZNF90, SFMBT2, LINC02323, CYTH2, LRRC8D, LYPLAL1, LRRC41, and / or SEPT9, and the methylation profile indicates that the subject has ovarian cancer. In some embodiments, the method further includes treating the subject with an anti-cancer therapy.
[0039]
[0010] Embodiments of the present disclosure also include methods for identifying gynecological cancer. According to these embodiments, the method includes determining a methylation profile in at least one differentially methylated region (DMR) of a DNA sample obtained from a subject having or suspected of having gynecological cancer by treating the sample with a reagent that modifies DNA in a methylation-specific manner. In some embodiments, the at least one DMR includes one or more CpG sites in KRT86, EMX2OS, JSRP1, DIDO1, MPZ, VILL, SMPD5, GDF7, MDFI, c17orf64, GATA2, SQSTM1, and / or EEF1A2, and the methylation profile indicates that the subject has ovarian cancer, cervical cancer, or endometrial cancer. In some embodiments, the method further includes treating the subject with an anti-cancer therapy. [Brief explanation of the drawings]
[0040] [Figure 1] 1 is a representative heatmap showing the ability of candidate methylated DNA markers to discriminate between gynecological cancers and cancer subtypes (see also Table 3). [Figure 2] Figures 2A-2C show representative data corresponding to the DNA methylation marker LRRC41, including a calibration plot based on ACTB normalization (Figure 2A), as well as adjusted boxplots showing epigenetic relationships among the three major gynecological cancers (Figure 2B), among gynecological cancer subtypes (Figure 2C), and in controls (Figures 2A and 2B). [Figure 3] Figures 3A-3C show representative data corresponding to the DNA methylation marker CDO1, including a calibration plot based on ACTB normalization (Figure 3A), as well as adjusted boxplots showing epigenetic relationships among the three major gynecological cancers (Figure 3B), among gynecological cancer subtypes (Figure 3C), and in controls (Figures 3A and 3B). [Figure 4]Figures 4A-4C show representative data corresponding to the DNA methylation marker ZMIZ1, including a calibration plot based on ACTB normalization (Figure 4A), as well as adjusted boxplots showing epigenetic relationships among the three major gynecological cancers (Figure 4B), among gynecological cancer subtypes (Figure 4C), and in controls (Figures 4A and 4B). [Figure 5] Figures 5A-5C show representative data corresponding to the DNA methylation marker PISD, including a calibration plot based on ACTB normalization (Figure 5A), as well as adjusted boxplots showing epigenetic relationships among the three major gynecological cancers (Figure 5B), among gynecological cancer subtypes (Figure 5C), and in controls (Figures 5A and 5B). [Figure 6] Figures 6A-6C show representative data corresponding to the DNA methylation marker AIM1, including a calibration plot based on ACTB normalization (Figure 6A), as well as adjusted boxplots showing epigenetic relationships among the three major gynecological cancers (Figure 6B), among gynecological cancer subtypes (Figure 6C), and in controls (Figures 6A and 6B). [Figure 7] Figures 7A-7C show representative data corresponding to the DNA methylation marker AK5, including a calibration plot based on ACTB normalization (Figure 7A), as well as adjusted boxplots showing epigenetic relationships among the three major gynecological cancers (Figure 7B), among gynecological cancer subtypes (Figure 7C), and in controls (Figures 7A and 7B). [Figure 8] Figures 8A-8C show representative data corresponding to the DNA methylation marker c18orf18, including a calibration plot based on ACTB normalization (Figure 8A), as well as adjusted boxplots showing epigenetic relationships among the three major gynecological cancers (Figure 8B), among gynecological cancer subtypes (Figure 8C), and in controls (Figures 8A and 8B). [Figure 9]Figures 9A-9C show representative data corresponding to the DNA methylation marker ELMOD1, including a calibration plot based on ACTB normalization (Figure 9A) and adjusted boxplots showing epigenetic relationships among the three major gynecological cancers (Figure 9B), among gynecological cancer subtypes (Figure 9C), and in controls (Figures 9A and 9B). [Figure 10] Figures 10A-10C show representative data corresponding to the DNA methylation marker FKBP11, including calibration plots based on ACTB normalization (Figure 10A), and adjusted boxplots showing epigenetic relationships among the three major gynecological cancers (Figure 10B), among gynecological cancer subtypes (Figure 10C), and in controls (Figures 10A and 10B). [Figure 11] Figures 11A-11C show representative data corresponding to the DNA methylation marker FLOT1, including a calibration plot based on ACTB normalization (Figure 11A) and adjusted boxplots showing epigenetic relationships among the three major gynecological cancers (Figure 11B), among gynecological cancer subtypes (Figure 11C), and in controls (Figures 11A and 11B). [Figure 12] Figures 12A-12C show representative data corresponding to the DNA methylation marker GAL3ST2, including a calibration plot based on ACTB normalization (Figure 12A) and adjusted boxplots showing epigenetic relationships among the three major gynecological cancers (Figure 12B), among gynecological cancer subtypes (Figure 12C), and in controls (Figures 12A and 12B). [Figure 13] Figures 13A-13C show representative data corresponding to the DNA methylation marker MAX.chr11.593, including a calibration plot based on ACTB normalization (Figure 13A), and adjusted boxplots showing epigenetic relationships among the three major gynecological cancers (Figure 13B), among gynecological cancer subtypes (Figure 13C), and in controls (Figures 13A and 13B). [Figure 14]Figures 14A-14C show representative data corresponding to the DNA methylation marker MLH1, including a calibration plot based on ACTB normalization (Figure 14A) and adjusted boxplots showing epigenetic relationships among the three major gynecological cancers (Figure 14B), among gynecological cancer subtypes (Figure 14C), and in controls (Figures 14A and 14B). [Figure 15] Figures 15A-15C show representative data corresponding to the DNA methylation marker NR3C1, including a calibration plot based on ACTB normalization (Figure 15A) and adjusted boxplots showing epigenetic relationships among the three major gynecological cancers (Figure 15B), among gynecological cancer subtypes (Figure 15C), and in controls (Figures 15A and 15B). [Figure 16] Figures 16A-16C show representative data corresponding to the DNA methylation marker RABC3, including a calibration plot based on ACTB normalization (Figure 16A), and adjusted boxplots showing epigenetic relationships among the three major gynecological cancers (Figure 16B), among gynecological cancer subtypes (Figure 16C), and in controls (Figures 16A and 16B). [Figure 17] Figures 17A-17C show representative data corresponding to the DNA methylation marker RAI1, including a calibration plot based on ACTB normalization (Figure 17A) and adjusted boxplots showing epigenetic relationships among the three major gynecological cancers (Figure 17B), among gynecological cancer subtypes (Figure 17C), and in controls (Figures 17A and 17B). [Figure 18] Figures 18A-18C show representative data corresponding to the DNA methylation marker TERC, including a calibration plot based on ACTB normalization (Figure 18A), and adjusted boxplots showing epigenetic relationships among the three major gynecological cancers (Figure 18B), among gynecological cancer subtypes (Figure 18C), and in controls (Figures 18A and 18B). [Figure 19]Figures 19A-19C show representative data for the DNA methylation marker TRPC3, including a calibration plot based on ACTB normalization (Figure 19A) and adjusted boxplots showing epigenetic relationships among the three major gynecological cancers (Figure 19B), among gynecological cancer subtypes (Figure 19C), and in controls (Figures 19A and 19B). [Figure 20] Figures 20A-20C show representative data corresponding to the DNA methylation marker ZIC2, including a calibration plot based on ACTB normalization (Figure 20A) and adjusted boxplots showing epigenetic relationships among the three major gynecological cancers (Figure 20B), among gynecological cancer subtypes (Figure 20C), and in controls (Figures 20A and 20B). [Figure 21] Figures 21A-21C show representative data corresponding to the DNA methylation marker ZNF480, including a calibration plot based on ACTB normalization (Figure 21A) and adjusted boxplots showing epigenetic relationships among the three major gynecological cancers (Figure 21B), among gynecological cancer subtypes (Figure 21C), and in controls (Figures 21A and 21B). [Figure 22] Figures 22A-22C show representative data corresponding to the DNA methylation marker ZNF491, including a calibration plot based on ACTB normalization (Figure 22A) and adjusted boxplots showing epigenetic relationships among the three major gynecological cancers (Figure 22B), among gynecological cancer subtypes (Figure 22C), and in controls (Figures 22A and 22B). [Figure 23] Figures 23A-23C show representative data corresponding to the DNA methylation marker ZNF610, including a calibration plot based on ACTB normalization (Figure 23A), and adjusted boxplots showing epigenetic relationships among the three major gynecological cancers (Figure 23B), among gynecological cancer subtypes (Figure 23C), and in controls (Figures 23A and 23B). [Figure 24]Figures 24A-24C are representative data corresponding to the DNA methylation marker ZNF91, including a calibration plot based on ACTB normalization (Figure 24A) and adjusted boxplots showing epigenetic relationships among the three major gynecological cancers (Figure 24B), among gynecological cancer subtypes (Figure 24C), and controls (Figures 24A and 24B). [Figure 25] Figures 25A-25C show representative data corresponding to the DNA methylation marker DLGAP1, including a calibration plot based on ACTB normalization (Figure 25A) and adjusted boxplots showing epigenetic relationships among the three major gynecological cancers (Figure 25B), among gynecological cancer subtypes (Figure 25C), and controls (Figures 25A and 25B). [Figure 26] Figures 26A-26C are representative data corresponding to the DNA methylation marker LYPLAP_2, including a calibration plot based on ACTB normalization (Figure 26A), and adjusted boxplots showing epigenetic relationships among the three major gynecological cancers (Figure 26B), among gynecological cancer subtypes (Figure 26C), and controls (Figures 26A and 26B). DETAILED DESCRIPTION OF THE INVENTION
[0041] The present disclosure relates to the detection of one or more types of gynecological cancer in a biological sample from a subject. In particular, the present disclosure provides compositions and methods for detecting the presence or absence of one or more types of gynecological cancer (e.g., cervical cancer, ovarian cancer, endometrial cancer) in a biological sample from a subject having or suspected of having gynecological cancer.
[0042] The section headings used in this section and throughout this disclosure are for organizational purposes only and are not intended to be limiting.
[0043] 1.Definition Throughout the specification and claims, the following terms have the meanings expressly associated therewith, unless the context clearly dictates otherwise. As used herein, the phrase "in one embodiment" may refer to the same embodiment, but does not necessarily refer to the same embodiment. Furthermore, as used herein, the phrase "in another embodiment" may refer to a different embodiment, but does not necessarily refer to a different embodiment. Thus, as described below, various embodiments of the invention can be readily combined without departing from the scope or spirit of the invention.
[0044] Additionally, as used herein, the term "or" is an inclusive "or" operator and is synonymous with the term "and / or" unless the context clearly dictates otherwise. The term "based on" is not exclusive and acknowledges that a relationship may be based on additional unlisted factors unless the context clearly dictates otherwise. Additionally, throughout this specification, the meanings of "a," "an," and "the" include plural referents. The meaning of "in" includes "in" and "on."
[0045] The transitional phrase "consisting essentially of," when used in the claims of this application, limits the scope of the claim to certain materials or steps "and which do not materially affect the basic and novel characteristic(s)" of the claimed invention, as stated in In re Herz, 537 F.2d 549, 551-52,190 USPQ 461,463 (CCPA 1976). For example, a composition "consisting essentially of" the recited elements may contain an unrecited contaminant at a level such that, although the contaminant is present, the contaminant does not alter the function of the recited composition compared to the pure composition (i.e., a composition "consisting of" the recited components).
[0046] As used herein, the term "one or more" refers to a number greater than 1. For example, the term "one or more" includes any of the following: 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 12 or more, 13 or more, 14 or more, 15 or more, 20 or more, 50 or more, 100 or more, or even more.
[0047] The terms "one or more, but less than a larger number," "two or more, but less than a larger number," "three or more, but less than a larger number," "four or more, but less than a larger number," "five or more, but less than a larger number," "six or more, but less than a larger number," "seven or more, but less than a larger number," "eight or more, but less than a larger number," "nine or more, but less than a larger number," "ten or more, but less than a larger number," "eleven or more, but less than a larger number," "twelve or more, but less than a larger number," "thirteen or more, but less than a larger number," "fourteen or more, but less than a larger number," or "fifteen or more, but less than a larger number" are not limited to the larger number. For example, the larger number can be 10,000, 1,000, 100, 50, etc. For example, the larger number can be about 50 (e.g., 50, 49, 48, 47, 46, 45, 44, 43, 42, 41, 40, 39, 38, 37, 36, 35, 34, 33, 32, 31, 32, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or 2).
[0048] The terms "one or more methylation markers" or "one or more DMRs" or "one or more genes" or "one or more markers" or "multiple methylation markers" or "multiple markers" or "multiple genes" or "multiple DMRs" are likewise not limited to a specific numerical combination. Indeed, any numerical combination of methylation markers is contemplated (e.g., 1-2 methylation markers, 1-3, 1-4, 1-5, 1-6, 1-7, 1-8, 1-9, 1-10, 1-11, 1-12, 1-13, 1-14, 1-15, 1-16, 1-17, 1-18, 1-19, 1-20, 1-21, 1-22, 1-23, 1-24, 1-25, 1-26, 1-27, 1-28, 1-29, 1-30, 1-31, 1-32, 1-33, 1-34, 1-35, 1-36, 1-37, 1-38) (e.g., 2-3, 2-4, 2-5, 2-6, 2-7, 2-8, 2-9, 2-10, 2-11, 2-12, 2-13, 2-14, 2-15, 2-16, 2-17, 2-18, 2-19, 2-20, 2-21, 2-22, 2-23, 2-24, 2-25, 2-26, 2-27, 2-28, 2-29, 2-30, 2-31, 2-32, 2-33, 2-34, 2-35, 2-36, 2-37, 2-38) (e.g., 3-4, 3-5, 3-6, 3-7, 3-8, 3-9, 3-10) , 3-11, 3-12, 3-13, 3-14, 3-15, 3-16, 3-17, 3-18, 3-19, 3-20, 3-21, 3-22, 3-23, 3-24, 3-25, 3-26, 3-27, 3-28, 3-29, 3-30, 3-31, 3-32, 3-33, 3-34, 3-35, 3-36, 3-37, 3-38) (e.g., 4-5, 4-6, 4-7, 4-8, 4-9, 4-10, 4-11, 4-12, 4-13, 4-14, 4-15, 4-16, 4-17, 4-18, 4-19, 4-20, 4-21, 4-22, 4-23, 4-24, 4-25, 4-26, 4-27, 4-28, 4-29, 4-30, 4-31, 4-32, 4-33, 4-34, 4-35, 4-36, 4-37, 4-38) (e.g., 5-6, 5-7, 5-8, 5-9, 5-10, 5-11, 5-12, 5-13, 5-14, 5-15, 5-16, 5-17, 5-18, 5-19, 5-20, 5-21, 5-22, 5-23, 5-24, 5-25, 5-26, 5-27, 5-28, 5-29,5-30, 5-31, 5-32, 5-33, 5-34, 5-35, 5-36, 5-37, 5-38) (e.g., 6-7, 6-8, 6-9, 6-10, 6-11, 6-12, 6-13, 6-14, 6-15, 6-16, 6-17, 6-18, 6-19, 6-20, 6-21, 6- 22, 6-23, 6-24, 6-25, 6-26, 6-27, 6-28, 6-29, 6-30, 6-31, 6-32, 6-33, 6-34, 6-35, 6-36, 6-37, 6-38) (e.g., 7-8, 7-9, 7-10, 7-11, 7-12, 7-13, 7-14, 7-15 , 7-16, 7-17, 7-18, 7-19, 7-20, 7-21, 7-22, 7-23, 7-24, 7-25, 7-26, 7-27, 7-28, 7-29, 7-30, 7-31, 7-32, 7-33, 7-34, 7-35, 7-36, 7-37, 7-38) (e.g., 8-9 , 8-10, 8-11, 8-12, 8-13, 8-14, 8-15, 8-16, 8-17, 8-18, 8-19, 8-20, 8-21, 8-22, 8-23, 8-24, 8-25, 8-26, 8-27, 8-28, 8-29, 8-30, 8-31, 8-32, 8-33, 8-34 , 8-35, 8-36, 8-37, 8-38) (e.g., 9-10, 9-11, 9-12, 9-13, 9-14, 9-15, 9-16, 9-17, 9-18, 9-19, 9-20, 9-21, 9-22, 9-23, 9-24, 9-25, 9-26, 9-27, 9-28, 9-29, 9, 9-30, 9-31, 9-32, 9-33, 9-34, 9-35, 9-36, 9-37, 9-38) (e.g., 10-11, 10-12, 10-13, 10-14, 10-15, 10-16, 10-17, 10-18, 10-19, 10-20, 10-21, 10-22, 1 0-23, 10-24, 10-25, 10-26, 10-27, 10-28, 10-29, 10-30, 10-31, 10-32, 10-33, 10-34, 10-35, 10-36, 10-37, 10-38) (e.g., 11-12, 11-13, 11-14, 11-15, 1 1-16, 11-17, 11-18, 11-19, 11-20, 11-21, 11-22, 11-23, 11-24, 11-25, 11-26, 11-27, 11-28, 11-29, 11-30, 11-31, 11-32, 11-33, 11-34, 11-35, 11-36,11-37, 11-38) (e.g., 12-13, 12-14, 12-15, 12-16, 12-17, 12-18, 12-19, 12-20, 12-21, 12-22, 12-23, 12-24, 12-25, 12-26, 12-27, 12-28, 12-29, 12-30, 12-31, 12-32, 12-33, 12-34, 12-35, 12-36, 12-37, 12-38) (e.g., 13-14, 13-15, 13-16, 13-17, 13-18, 13-19, 13-20, 13-21, 13-22, 13-23, 13-24, 13-25, 13-26, 13-27, 13-28, 13-29, 13-30, 13-31, 13-32, 13-33, 13-34, 13-35, 13-36, 13-37, 13-38) (e.g., 14-15, 14-16, 14-17, 14-18, 14-19, 14-20, 14-21, 14-22, 14-23, 14-24, 14-25, 14-26, 14-27, 14-28, 14-29, 14-30, 14-31, 14-32, 14-33, 14-34, 14-35, 14-36, 14-37, 14-38) (e.g., 15-16, 15-17, 15-18, 15-19, 15-20, 15-21, 15-22, 15-23, 15-24, 15-25, 15-26, 15-27, 15-28, 15-29, 15-30, 15-31, 15-32, 15-33, 15-34, 15-35, 15-36, 15-37, 15-38) (e.g., 16-17, 16-18, 16-19, 16-20, 16-21, 16-22, 16-23, 16-24, 16-25, 16-26, 16-27, 16-28, 16-29, 16-30, 16-31, 16-32, 16-33, 16-34, 16-35, 16-36, 16-37 , 16-38) (e.g., 17-18, 17-19, 17-20, 17-21, 17-22, 17-23, 17-24, 17-25, 17-26, 17-27, 17-28, 17-29, 17-30, 17-31, 17-32, 17-33, 17-34, 17-35, 17-36 , 17-37, 17-38) (e.g., 18-19, 18-20, 18-21, 18-22, 18-23, 18-24, 18-25, 18-26, 18-27, 18-28, 18-29, 18-30, 18-31, 18-32, 18-33, 18-34, 18-35, 18-36,18-37, 18-38) (e.g., 19-20, 19-21, 19-22, 19-23, 19-24, 19-25, 19-26, 19-27, 19-28, 19-29, 19-30, 19-31, 19-32, 19-33, 19-34, 19-35, 19-36, 19-37, 19-38) (e.g., 20-21, 20-22, 20-23, 20-24, 20-25, 20-26, 20-27, 20-28, 20-29, 20-30, 20-31, 20-32, 20-33, 20-34, 20-35, 20-36, 20-37, 20-38) (e.g., 21-22, 21-23, 21-24, 21-25, 21-26, 21-27, 21-28, 21-29, 21-30, 21-31, 21-32, 21-33, 21-34, 21-35, 21-36, 21-37, 21-38) (e.g., 22-23, 22-24, 22-25 , 22-26, 22-27, 22-28, 22-29, 22-30, 22-31, 22-32, 22-33, 22-34, 22-35, 22-36, 22-37, 22-38) (e.g., 23-24, 23-25, 23-26, 23-27, 23-28, 23-29, 23-30 , 23-31, 23-32, 23-33, 23-34, 23-35, 23-36, 23-37, 23-38) (e.g., 24-25, 24-26, 24-27, 24-28, 24-29, 24-30, 24-31, 24-32, 24-33, 24-34, 24-35, 24-36, 24-37, 24-38) (e.g., 25-26, 25-27, 25-28, 25-29, 25-30, 25-31, 25-32, 25-33, 25-34, 25-35, 25-36, 25-37, 25-38) (e.g., 26-27, 26-28, 26-29, 26-30 , 26-31, 26-32, 26-33, 26-34, 26-35, 26-36, 26-37, 26-38) (e.g., 27-28, 27-29, 27-30, 27-31, 27-32, 27-33, 27-34, 27-35, 27-36, 27-37, 27-38) (e.g., 28-29, 28-30, 28-31, 28-32, 28-33, 28-34, 28-35, 28-36, 28-37, 28-38) (e.g., 29-30, 29-31, 29-32, 29-33, 29-34, 29-35, 29-36, 29-37, 29-38) (e.g.,30-31, 30-32, 30-33, 30-34, 30-35, 30-36, 30-37, 30-38) (e.g., 31-32, 31-33, 31-34, 31-35, 31-36, 31-37, 31-38) (e.g., 32-33, 32-34, 32-35, 32-36, 32-37, 32-38) (e.g., 33-34, 33-35, 33-36, 33-37, 33-38) (e.g., 34-35, 34-36, 34-37, 34-38) (e.g., 35-36, 35-37, 3 5 to 38) (e.g., 36 to 37, 36 to 38) (e.g., 37 to 38) (e.g., 38 or less, 37 or less, 36 or less, 35 or less, 34 or less, 33 or less, 32 or less, 31 or less, 30 or less, 29 or less, 28 or less, 27 or less, 26 or less, 25 or less, 24 or less, 23 or less, 22 or less, 21 or less, 20 or less, 19 or less, 18 or less, 17 or less, 16 or less, 15 or less, 14 or less, 13 or less, 12 or less, 11 or less, 10 or less, 9 or less, 8 or less, 7 or less, 6 or less, 5 or less, 4 or less, 3 or less, 2 or 1).
[0049] The terms "multiple types of cancer" or "one or more types of cancer" or "one or more subtypes of cancer" or "multiple different types or subtypes of cancer" are similarly not limited to specific numerical combinations. Any number of combinations of gynecological cancer types or subtypes can be identified using the DNA methylation markers of the present disclosure, including, but not limited to, ovarian cancer, serous ovarian cancer, clear cell ovarian cancer, endometrioid ovarian cancer, mucinous ovarian cancer, cervical cancer, cervical adenocarcinoma, cervical squamous cell carcinoma, endometrial cancer, and endometrioid carcinoma.
[0050] As used herein, "nucleic acid" or "nucleic acid molecule" generally refers to any ribonucleic acid or deoxyribonucleic acid, which may be unmodified or modified DNA or RNA. "Nucleic acid" includes, but is not limited to, single-stranded and double-stranded nucleic acids. As used herein, the term "nucleic acid" also includes DNA, as described above, containing one or more modified bases. Thus, DNA with backbone modifications for stability or other reasons is a "nucleic acid." As used herein, the term "nucleic acid" encompasses chemically, enzymatically, or metabolically modified forms of nucleic acid, as well as chemical forms of DNA characteristic of viruses and cells, including, for example, simple and complex cells.
[0051] The terms "oligonucleotide" or "polynucleotide" or "nucleotide" or "nucleic acid" refer to a molecule containing two or more deoxyribonucleotides or ribonucleotides, preferably more than three, and usually more than ten. The exact size will depend on many factors and is dependent on the ultimate function or use of the oligonucleotide. Oligonucleotides can be produced in any manner, including chemical synthesis, DNA replication, reverse transcription, or a combination thereof. Typical deoxyribonucleotides of DNA are thymine, adenine, cytosine, and guanine. Typical ribonucleotides of RNA are uracil, adenine, cytosine, and guanine.
[0052] As used herein, the term "locus" or "region" of a nucleic acid refers to a small region of nucleic acid, e.g., a gene on a chromosome, a single nucleotide, a CpG island, and the like.
[0053] The terms "complementary" and "complementarity" refer to nucleotides (e.g., a single nucleotide) or polynucleotides (e.g., a sequence of nucleotides) related by the base-pairing rules. For example, the sequence 5'-AGT-3' is complementary to the sequence 3'-TCA-5'. Complementarity can be "partial," in which only a portion of the nucleic acid bases match according to the base-pairing rules. Alternatively, there can be "complete" or "total" complementarity between nucleic acids. The degree of complementarity between nucleic acid strands affects the efficiency and strength of hybridization between nucleic acid strands. This is particularly important in amplification reactions and detection methods that depend on binding between nucleic acids.
[0054] The term "gene" refers to a nucleic acid (e.g., DNA or RNA) sequence that comprises coding sequences necessary for the production of an RNA or polypeptide or its precursor. A functional polypeptide can be encoded by a full-length coding sequence or by any portion of the coding sequence, so long as the desired activity or functional property of the polypeptide (e.g., enzymatic activity, ligand binding, signal transduction, etc.) is retained. When used in reference to a gene, the term "portion" refers to fragments of that gene. These fragments can range in size from a few nucleotides to the entire gene sequence minus one nucleotide. Thus, "nucleotides comprising at least a portion of a gene" can include fragments of a gene or the entire gene.
[0055] The term "gene" includes the coding region of a structural gene as well as sequences located adjacent to the coding region at both the 5' and 3' ends, where the gene corresponds to the length of the full-length mRNA (e.g., including coding, regulatory, structural, and other sequences). Sequences located 5' of the coding region and present on the mRNA are referred to as 5' untranslated or non-translated sequences. Sequences located 3' or downstream of the coding region and present on the mRNA are referred to as 3' untranslated or 3' non-translated sequences. The term "gene" encompasses both cDNA and genomic forms of a gene. In some organisms (e.g., eukaryotes), genomic forms or clones of a gene contain coding regions interrupted by non-coding sequences termed "introns" or "intervening regions" or "intervening sequences." Introns are segments of a gene that are transcribed into nuclear RNA (hnRNA) and may contain regulatory elements such as enhancers. Introns are removed or "spliced out" from the nuclear or primary transcript; therefore, introns are absent in the messenger RNA (mRNA) transcript. mRNA functions during translation to specify the sequence or order of amino acids in a nascent polypeptide. As will be understood by those skilled in the art based on the present disclosure, one or more CpG sites in a DMR may be located in a coding region not known to be associated with a particular gene, such as a coding region of a gene, a non-coding regulatory region of a gene, or a region containing long non-coding RNA (lncRNA). In some embodiments, sequences corresponding to these regions may be obtained using corresponding accession numbers (see, e.g., Tables 1 and 2) in genome databases (e.g., GenBank, NCBI, UniProt, etc.). In some embodiments, one or more CpG sites in a DMR may be located within an unannotated genomic region. As further provided herein, an unannotated genomic region containing one or more CpG sites in a DMR may be described using a SEQ ID NO: (see, e.g., Tables 1 and 2, SEQ ID NOs: 1-32).
[0056] As will be understood by one of skill in the art based on the present disclosure, the location of one or more CpG sites within a gene or region (e.g., a CpG island) and its association with a disease or condition can be determined using a variety of techniques, including, but not limited to, those described in Chen et al., "Methods for identifying differentially methylated regions for sequence- and array-based data," Briefings in Functional Genomics, Volume 15, Issue 6, November 2016, pp. 485-490 (incorporated herein by reference in its entirety for all purposes).
[0057] In addition to containing introns, genomic forms of a gene may also include sequences located on both the 5' and 3' end of the sequences present on the RNA transcript. These sequences are referred to as "flanking" sequences or regions (these flanking sequences are located 5' or 3' to the untranslated sequences present in the mRNA transcript). The 5' flanking region may contain regulatory sequences such as promoters and enhancers that control or influence the transcription of the gene. The 3' flanking region may contain sequences that direct the termination of transcription, post-transcriptional cleavage, and polyadenylation. These flanking regions may be non-coding and therefore may not be present in the mRNA transcript.
[0058] The term "wild-type," when used in reference to a gene, refers to a gene having the characteristics of a gene isolated from a naturally occurring source. The term "wild-type," when used in reference to a gene product, refers to a gene product having the characteristics of a gene product isolated from a naturally occurring source. The term "wild-type," when used in reference to a protein, refers to a protein having the characteristics of a naturally occurring protein. The term "naturally occurring," when applied to an object, refers to the fact that the object can be found in nature. For example, a polypeptide or polynucleotide sequence present in an organism (including a virus) that can be isolated from a natural source and has not been intentionally modified by human hands in the laboratory is naturally occurring. A wild-type gene is often that gene or allele that is most frequently observed in a population and is therefore arbitrarily referred to as the "normal" or "wild-type" form of the gene. In contrast, when referring to a gene or gene product, the terms "modified" or "mutant" refer to a gene or gene product, respectively, that exhibits modifications in sequence and / or functional properties (i.e., altered characteristics) when compared to the wild-type gene or gene product. Note that naturally occurring variants can be isolated. These are identified by the fact that they have altered properties when compared to the wild-type gene or gene product.
[0059] The term "allele" refers to a genetic variation, including, without limitation, variants and mutations, polymorphic loci and single nucleotide polymorphic loci, frameshifts, and splice variants. Alleles may occur naturally within a population or may arise during the lifetime of any particular individual in a population.
[0060] Thus, when used in reference to a nucleotide sequence, the terms "variant" and "mutant" refer to a nucleic acid sequence that differs by one or more nucleotides from another, usually related, nucleotide sequence. A "mutation" is a difference between two different nucleotide sequences, typically one sequence being a reference sequence.
[0061] The term "primer" refers to an oligonucleotide, whether naturally occurring as a nucleic acid fragment from a purified restriction digest or synthesized, that can act as a point of initiation of synthesis when placed under conditions that induce synthesis of a primer extension product complementary to a template strand of nucleic acid (e.g., in the presence of nucleotides and an inducing agent such as DNA polymerase, at a suitable temperature and pH). Primers are preferably single-stranded to maximize amplification efficiency, but may alternatively be double-stranded. If double-stranded, the primer is first treated to separate its strands before being used to prepare extension products. Preferably, the primer is an oligodeoxyribonucleotide. The primer must be long enough to prime the synthesis of an extension product in the presence of the inducing agent. The exact length of the primer will depend on many factors, including temperature, primer source, and the use of the method. In some embodiments, the primer pair is specific for a particular differentially methylated region (e.g., DMR 2 in Table 1) and specifically binds to at least a portion of the genetic region containing the DMR.
[0062] The term "probe" refers to an oligonucleotide (e.g., a series of nucleotides) capable of hybridizing to another oligonucleotide of interest, whether naturally occurring, as in a purified restriction digest, or produced synthetically, recombinantly, or by PCR amplification. Probes can be single-stranded or double-stranded. Probes are useful for the detection, identification, and isolation of specific gene sequences (e.g., "capture probes"). It is contemplated that any probe used in embodiments of the present disclosure can, in some embodiments, be labeled with any "reporter molecule" and thus be detectable in any detection system, including, but not limited to, enzymatic (e.g., ELISA and enzyme-based histochemical assays), fluorescent, radioactive, and luminescent systems. It is not intended that the various embodiments of the present disclosure be limited to any particular detection system or label.
[0063] As used herein, the term "target" refers to a nucleic acid to be sorted out from other nucleic acids, e.g., by probe binding, amplification, separation, capture, etc. For example, when used in reference to the polymerase chain reaction, "target" refers to the region of nucleic acid bounded by the primers used in the polymerase chain reaction, whereas when used in assays that do not amplify the target DNA, e.g., in some embodiments of an invasion-cleavage assay, the target includes the site where a probe and an invading oligonucleotide (e.g., an INVADER oligonucleotide) bind to form an invasion-cleavage structure, thereby allowing the presence of the target nucleic acid to be detected. A "segment" is defined as a region of nucleic acid within the target sequence.
[0064] Thus, as used herein, "non-target," when used to describe a nucleic acid such as, for example, DNA, refers to a nucleic acid that may be present in a reaction but is not the subject of detection or characterization by the reaction. In some embodiments, non-target nucleic acid can refer to a nucleic acid present in a sample that does not contain, for example, a target sequence, although in some embodiments, non-target can also refer to an exogenous nucleic acid, i.e., a nucleic acid that is not derived from a sample that contains or is suspected of containing a target nucleic acid, and that is added to a reaction, for example, to reduce variability in the performance of an enzyme (e.g., a polymerase) in the reaction, to normalize the activity of the enzyme.
[0065] As used herein, "methylation" refers to cytosine methylation at the C5 or N4 position of cytosine, the N6 position of adenine, or other types of nucleic acid methylation. In vitro amplified DNA is typically unmethylated, since typical in vitro DNA amplification methods do not preserve the methylation pattern of the amplified template. However, "unmethylated DNA" or "methylated DNA" can also refer to amplified DNA in which the original template was unmethylated or methylated, respectively.
[0066] As used herein, the term "amplification reagents" refers to the reagents needed for amplification (deoxyribonucleoside triphosphates, buffers, etc.) excluding primers, nucleic acid template, and amplification enzymes. Typically, amplification reagents are placed and contained within a reaction vessel along with other reaction components.
[0067] As used herein, the term "control," when used in reference to nucleic acid detection or analysis, refers to a nucleic acid with known characteristics (e.g., known sequence, known copy number per cell) used for comparison with an experimental target (e.g., a nucleic acid of unknown concentration, etc.). A control may be an endogenous, preferably invariant, gene to which a test or target nucleic acid in an assay can be normalized. Such normalization controls for sample-to-sample variations that may arise, for example, from sample processing, assay efficiency, etc., allowing for accurate data comparison between samples. Genes used to normalize nucleic acid detection assays in human samples include, for example, β-actin, ZDHHC1, and B3GALT6 (see, e.g., U.S. Patent Application Nos. 14 / 966,617 and 62 / 364,082, each of which is incorporated herein by reference). As used herein, "ZDHHC1" refers to a gene located on human chromosome 16 (16q22.1) that encodes a protein belonging to the DHHC palmitoyltransferase family, characterized as zinc finger DHHC-type containing 1. In some embodiments, reference genes include, but are not limited to, FNBP1, NCOR2, and S1PR4 (see Table 4).
[0068] Controls can also be external. For example, in quantitative assays such as qPCR and QuARTS, a "calibrator" or "calibration control" is a nucleic acid of known sequence, e.g., a nucleic acid with the same sequence as a portion of an experimental target nucleic acid and a known concentration or series of concentrations (e.g., a serially diluted control target for generating a standard curve for quantitative PCR). Typically, calibration controls are analyzed using the same reagents and reaction conditions as those used for the experimental DNA. In certain embodiments, the measurement of the calibrator is performed simultaneously with the experimental assay, e.g., in the same thermal cycler. In preferred embodiments, multiple calibrators can be included in a single plasmid so that different calibrator sequences can be easily provided in equimolar amounts. In particularly preferred embodiments, the plasmid calibrator is digested, e.g., with one or more restriction enzymes, to release the calibrator portion from the plasmid vector. See, e.g., WO2015 / 066695, incorporated herein by reference.
[0069] As used herein, "methylated nucleotide" or "methylated nucleotide base" refers to the presence of a methyl moiety on a nucleotide base, which is not present in recognized typical nucleotide bases.For example, cytosine does not contain a methyl moiety on its pyrimidine ring, but 5-methylcytosine contains a methyl moiety at the 5th position of its pyrimidine ring.Therefore, cytosine is not a methylated nucleotide, but 5-methylcytosine is a methylated nucleotide.In another example, thymine contains a methyl moiety at the 5th position of its pyrimidine ring, but because thymine is a typical nucleotide base of DNA, for the purposes of this specification, thymine is not considered a methylated nucleotide when present in DNA.
[0070] As used herein, a "methylated nucleic acid molecule" refers to a nucleic acid molecule that contains one or more methylated nucleotides.
[0071] As used herein, the "methylation state," "methylation profile," and "methylation status" of a nucleic acid molecule refer to the presence or absence of one or more methylated nucleotide bases in a nucleic acid molecule. For example, a nucleic acid molecule that contains a methylated cytosine is considered to be methylated (e.g., the methylation state of the nucleic acid molecule is methylated). A nucleic acid molecule that does not contain any methylated nucleotides is considered to be unmethylated.
[0072] As used herein, the term "methylation level" applied to a methylation marker refers to the amount of methylation in a particular methylation marker. Methylation level may also refer to the amount of methylation in a particular methylation marker compared to an established standard or control. Methylation level may also refer to whether one or more cytosine residues present in a CpG context have a methylation group. Methylation level may also refer to the proportion of cells in a sample that have or do not have a methylation group at such cytosine. Methylation level may also represent whether a single CpG dinucleotide is methylated.
[0073] The methylation state of a particular nucleic acid sequence (e.g., a genetic marker, or a DNA region, as described herein) can indicate the methylation state of all bases in the sequence, or it can indicate the methylation state of a subset of these bases (e.g., one or more cytosines) within the sequence, or it can indicate information about the methylation density of a region within the sequence, with or without providing precise information about the position within the sequence where methylation occurs.
[0074] The methylation state of a nucleotide locus in a nucleic acid molecule refers to the presence or absence of a methylated nucleotide at a particular locus in the nucleic acid molecule. For example, the methylation state of the cytosine at the seventh nucleotide in a nucleic acid molecule is methylated if the nucleotide present at the seventh nucleotide in the nucleic acid molecule is 5-methylcytosine. Similarly, the methylation state of the cytosine at the seventh nucleotide in a nucleic acid molecule is unmethylated if the nucleotide present at the seventh nucleotide in the nucleic acid molecule is cytosine (and not 5-methylcytosine).
[0075] Methylation status can optionally be expressed or indicated by a "methylation value" (e.g., representing a methylation frequency, fraction, proportion, percent, etc.). Methylation values can be generated, for example, by quantifying the amount of intact nucleic acid present after restriction digestion with a methylation-dependent restriction enzyme, or by comparing amplification profiles after a bisulfite reaction, or by comparing sequences of bisulfite-treated nucleic acid with sequences of untreated nucleic acid, or by comparing TET-treated nucleic acid with untreated nucleic acid. Thus, a value, e.g., a methylation value, represents methylation status and can thereby be used as a quantitative indicator of methylation status across multiple copies of a locus. This is of particular use when it is desirable to compare the methylation status of sequences in a sample to a threshold or reference value.
[0076] As used herein, "methylation frequency" or "percent (%) methylation" refers to the number of instances in which a molecule or locus is methylated relative to the number of instances in which the molecule or locus is unmethylated.
[0077] As used herein, the term "methylation score" refers to a score indicating the number of methylation events detected in a marker or panel of markers compared to the median number of methylation events for that marker or panel of markers from a randomized population of mammals (e.g., a randomized population of 10, 20, 30, 40, 50, 100, or 500 mammals) that do not have a particular tumor of interest. A high methylation score for a marker or panel of markers can be any score, as long as the score is greater than the corresponding reference score. For example, a high methylation score for a marker or panel of markers can be 0.5, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more times greater than the reference methylation score.
[0078] Thus, a methylation state refers to the methylation state of a nucleic acid (e.g., a genomic sequence). Furthermore, a methylation state refers to characteristics of a nucleic acid segment at a particular genomic locus related to methylation. Such characteristics include, but are not limited to, whether any cytosine (C) residues within the DNA sequence are methylated, the location of the methylated C residue(s), the frequency or percentage of methylated Cs throughout any particular region of the nucleic acid, and allelic differences in methylation due, for example, to differences in allelic origin. Additionally, the terms "methylation state," "methylation profile," and "methylation status" refer to the relative concentration, absolute concentration, or pattern of methylated or unmethylated Cs throughout any particular region of a nucleic acid in a biological sample. For example, if a cytosine (C) residue(s) within a nucleic acid sequence are methylated, they can be said to have "hypermethylated" or "high methylation," whereas if a cytosine (C) residue(s) within the DNA are unmethylated, they can be said to have "hypomethylated" or "low methylation." Similarly, if a cytosine (C) residue(s) in a nucleic acid sequence are methylated compared to another nucleic acid sequence (e.g., from a different region or from a different individual), the sequence is considered to be hypermethylated, or have high methylation, compared to the other nucleic acid sequence. Alternatively, if a cytosine (C) residue(s) in a DNA sequence are unmethylated compared to another nucleic acid sequence (e.g., from a different region or from a different individual), the sequence is considered to be hypomethylated, or have low methylation, compared to the other nucleic acid sequence. Additionally, as used herein, the term "methylation pattern" refers to the collection of methylated and unmethylated nucleotides across a nucleic acid region. Two nucleic acids can have the same or similar methylation frequency or methylation pattern, but have different methylation patterns when the numbers of methylated and unmethylated nucleotides are the same or similar throughout the region, but the positions of the methylated and unmethylated nucleotides are different.Sequences are referred to as having "variable methylation," "differences in methylation," or "differential methylation states" when they differ in the degree (e.g., one has increased or decreased methylation compared to the other), frequency, or pattern of methylation. The term "differential methylation" refers to the difference in the level or pattern of nucleic acid methylation in a cancer-positive sample compared to the level or pattern of nucleic acid methylation in a cancer-negative sample. The term can also refer to the difference in the level or pattern between patients who experience cancer recurrence after surgery and those who do not. Differential methylation and specific levels or patterns of DNA methylation serve as prognostic and predictive biomarkers, for example, after the correct cutoff or predictive features are defined. DMRs can be located within any region of a gene. In some embodiments, DMRs comprise, are derived from, or are located within one or more regions of a gene, including, but not limited to, coding regions, non-coding regions, regulatory regions, introns, exons, promoters, enhancers, and termination sequences, 3'UTRs, and 5'UTRs. In some embodiments, one or more CpG sites in a DMR may be located within a non-coding region, such as a region corresponding to a long non-coding RNA (lncRNA).
[0079] Methylation state frequencies can be used to describe samples from a population of individuals or a single individual. For example, a nucleotide locus with a methylation state frequency of 50% means that 50% of the instances are methylated and 50% of the instances are unmethylated. Such frequencies can be used, for example, to describe the degree to which a nucleotide locus or nucleic acid region is methylated in a population of individuals or a collection of nucleic acids. Thus, if the methylation in a first population or pool of nucleic acid molecules is different from the methylation in a second population or pool of nucleic acid molecules, the methylation state frequency of the first population or pool will be different from the methylation state frequency of the second population or pool. Such frequencies can also be used, for example, to describe the degree to which a nucleotide locus or nucleic acid region is methylated in a single individual. For example, such frequencies can be used to describe the degree to which a group of cells from a tissue sample is methylated or unmethylated at a nucleotide locus or nucleic acid region.
[0080] Typically, methylation of human DNA occurs at dinucleotide sequences containing adjacent guanines and cytosines, where the cytosine is located 5' of the guanine (also called CpG dinucleotide sequences). In the human genome, most cytosines within CpG dinucleotides are methylated, although some remain unmethylated in certain CpG dinucleotide-rich genomic regions known as CpG islands (see, e.g., Antequera, et al. (1990) Cell 62:503-514).
[0081] As used herein, "CpG island" or "cytosine-phosphate-guanine-island" refers to a G:C-rich region of genomic DNA that contains more CpG dinucleotides than the total genomic DNA. A CpG island can be at least 100, 200 base pairs long, or longer, where the G:C content of the region is at least 50% and the ratio of observed CpG frequency to expected frequency is 0.6; in some cases, a CpG island can be at least 500 base pairs long, where the G:C content of the region is at least 55% and the ratio of observed CpG frequency to expected frequency is 0.65. The observed CpG frequency to expected frequency can be calculated according to the method provided in Gardiner-Garden et al. (1987) J.Mol.Biol.196:261-281. For example, the observed CpG frequency relative to the expected frequency can be calculated according to the formula R = (A x B) / (C x D), where R is the ratio of the observed CpG frequency to the expected frequency, A is the number of CpG dinucleotides in the analyzed sequence, B is the total number of nucleotides in the analyzed sequence, C is the total number of C nucleotides in the analyzed sequence, and D is the total number of G nucleotides in the analyzed sequence. Methylation status is usually determined in CpG islands, for example, in promoter regions. However, it will be appreciated that other sequences in the human genome, such as CpA and CpT, are also subject to DNA methylation (see Ramsahoye (2000) Proc. Natl. Acad. Sci. USA 97:5237-5242; Salmon and Kaye (1970) Biochim. Biophys. Acta. 204:340-351; Grafstrom (1985) Nucleic Acids Res. 13:2827-2842; Nyce (1986) Nucleic Acids Res. 14:4353-4367; Woodcock (1987) Biochem. Biophys. Res. Commun. 145:888-894).
[0082] As used herein, a "methylation-specific reagent" refers to a reagent that modifies the nucleotides of a nucleic acid molecule as a function of the methylation state of the nucleic acid molecule, or a methylation-specific reagent refers to a compound or composition, or other agent, that is capable of altering the nucleotide sequence of a nucleic acid molecule in a manner that reflects the methylation state of the nucleic acid molecule. Methods of treating nucleic acid molecules with such reagents can include contacting the nucleic acid molecule with the reagent, optionally in combination with additional steps, to achieve a desired change in nucleotide sequence. Such methods can be applied in a manner that results in the modification of unmethylated nucleotides (e.g., each unmethylated cytosine) to a different nucleotide. For example, in some embodiments, such reagents can deaminate unmethylated cytosine nucleotides to generate deoxyuracil residues. Examples of such reagents include, but are not limited to, methylation-sensitive restriction enzymes, methylation-dependent restriction enzymes, bisulfite reagents, TET enzymes, and borane reducing agents.
[0083] Alteration of a nucleic acid nucleotide sequence with a methylation-specific reagent can also result in a nucleic acid molecule in which each methylated nucleotide is modified to a different nucleotide.
[0084] The term "methylation assay" refers to any assay for determining the methylation status of one or more CpG dinucleotide sequences within a nucleic acid sequence.
[0085] The term "MS AP-PCR" (methylation-sensitive arbitrarily primed polymerase chain reaction) refers to an art-recognized technique that uses CG-rich primers to scan the entire genome and focus on regions most likely to contain CpG dinucleotides, and is described by Gonzalgo et al. (1997) Cancer Research 57:594-599.
[0086] The term "MethyLight™" refers to the art-recognized fluorescence-based real-time PCR technology described by Eads et al. (1999) Cancer Res. 59:2302-2306.
[0087] The term "HeavyMethyl™" refers to an assay in which methylation-specific inhibitory probes (also referred to herein as inhibitors) that cover the CpG positions between or covered by the amplification primers enable methylation-specific selective amplification of a nucleic acid sample.
[0088] The term "HeavyMethyl™ MethyLight™" assay refers to the HeavyMethyl™ MethyLight™ assay, which is a variation of the MethyLight™ assay in which the MethyLight™ assay is combined with a methylation-specific blocking probe that covers the CpG positions between the amplification primers.
[0089] The term "Ms-SNuPE" (methylation-sensitive single-nucleotide primer extension) refers to the art-recognized assay described in Gonzalgo & Jones (1997) Nucleic Acids Res. 25:2529-2531.
[0090] The term "MSP" (methylation-specific PCR) refers to the art-recognized methylation assay described by Herman et al. (1996) Proc. Natl. Acad. Sci. USA 93:9821-9826 and in US Pat. No. 5,786,146.
[0091] The term "COBRA" (Combined Bisulfite Restriction Analysis) refers to an art-recognized methylation assay described in Xiong & Laird (1997) Nucleic Acids Res. 25:2532-2534.
[0092] The term "MCA" (methylated CpG island amplification) refers to the methylation assay described in Toyota et al. (1999) Cancer Res. 59:2307-12 and WO00 / 26401A1.
[0093] As used herein, a "selected nucleotide" refers to one of the four nucleotides typically occurring in a nucleic acid molecule (C, G, T, and A for DNA and C, G, U, and A for RNA), and can include methylated derivatives of a typically occurring nucleotide (e.g., when C is a selected nucleotide, both methylated and unmethylated C are included in the meaning of the selected nucleotide), but a methylated selected nucleotide specifically refers to a methylated typically occurring nucleotide, and an unmethylated selected nucleotide specifically refers to an unmethylated typically occurring nucleotide.
[0094] The term "methylation-specific restriction enzyme" refers to a restriction enzyme that selectively digests nucleic acids depending on the methylation state of its recognition site. For restriction enzymes that specifically cleave when their recognition site is unmethylated or hemimethylated (methylation-sensitive enzymes), cleavage does not occur (or occurs at a much lower efficiency) when the recognition site is methylated on one or both strands. For restriction enzymes that specifically cleave only when their recognition site is methylated (methylation-dependent enzymes), cleavage does not occur (or occurs at a much lower efficiency) when the recognition site is unmethylated. Methylation-specific restriction enzymes are preferred, and their recognition sequences contain a CG dinucleotide (e.g., a recognition sequence such as CGCG or CCCGGG). Even more preferred in some embodiments are restriction enzymes that do not cleave when the cytosine in this dinucleotide is methylated at the C5 carbon atom.
[0095] As used herein, the "sensitivity" of a particular marker (or set of markers used in combination) refers to the percentage of samples reporting DNA methylation values above a threshold that distinguishes between tumor and non-tumor samples. In some embodiments, a positive result is defined as a histologically confirmed tumor reporting a DNA methylation value above a threshold (e.g., a range associated with disease), and a false negative result is defined as a histologically confirmed tumor reporting a DNA methylation value below a threshold (e.g., a range associated with non-disease). Thus, a sensitivity value reflects the probability that a DNA methylation measurement value for a given marker from a known diseased sample will fall within the range of disease-associated measurements. As defined herein, the clinical significance of a calculated sensitivity value represents an estimate of the probability that a given marker will detect the presence of a clinical condition when applied to subjects with that condition.
[0096] As used herein, the "specificity" of a given marker (or a set of markers used together) refers to the proportion of non-neoplastic samples reporting DNA methylation values below a threshold that distinguishes between neoplastic and non-neoplastic samples. In some embodiments, a negative is defined as a histologically confirmed non-neoplastic sample reporting a DNA methylation value below the threshold (e.g., a range not associated with any disease), and a false positive is defined as a histologically confirmed non-neoplastic sample reporting a DNA methylation value above the threshold (e.g., a range associated with a disease). Thus, the specificity value reflects the probability that a DNA methylation measurement value for a given marker from a known non-neoplastic sample will fall within the range of non-disease-associated measurements. As defined herein, the clinical relevance of a calculated specificity value represents an estimate of the probability that a given marker will detect the absence of a clinical condition when applied to patients without that condition.
[0097] The term "AUC" used herein is an abbreviation for "area under the curve." In particular, AUC refers to the area under the receiver operating characteristic (ROC) curve. An ROC curve is a plot of the true positive rate against the false positive rate for different possible cut points of a diagnostic test. The ROC curve shows the trade-off between sensitivity and specificity depending on the cut point selected (any increase in sensitivity will be accompanied by a decrease in specificity). The area under the ROC curve (AUC) is a measure of the accuracy of a diagnostic test (the larger the area, the better, with 1 being optimal, and a randomized test has a ROC curve located on the diagonal, with an area of 0.5. See: J.P. Egan. (1975) Signal Detection Theory and ROC Analysis, Academic Press, New York).
[0098] As used herein, the term "tumor" refers to any new, abnormal growth of tissue. Thus, a tumor can be a pre-malignant tumor or a malignant tumor.
[0099] The term "tumor-specific marker," as used herein, refers to any biological material or element that can be used to indicate the presence of a tumor. Examples of biological materials include, but are not limited to, nucleic acids, polypeptides, carbohydrates, fatty acids, cellular components (e.g., cell membranes and mitochondria), and whole cells. In some cases, a marker is a specific nucleic acid region (e.g., a gene, a region within a gene, a specific locus, etc.). A region of a nucleic acid that is a marker may be referred to, for example, as a "marker gene," a "marker region," a "marker sequence," a "marker locus," etc.
[0100] As used herein, the term "adenoma" refers to a benign tumor of glandular origin. These growths are benign, although over time they can progress to become malignant.
[0101] The terms "precancerous" or "preneoplastic" and their equivalents refer to any cell proliferative disorder undergoing malignant transformation.
[0102] The "site" of a tumor, adenoma, cancer, etc. is the tissue, organ, cell type, anatomical region, body part, etc. in a subject in which the tumor, adenoma, cancer, etc. is located.
[0103] As used herein, the application of a "diagnostic" test includes the detection or identification of a disease state or condition in a subject, determining the likelihood that a subject will suffer from a given disease or condition, determining the likelihood that a subject with a disease or condition will respond to treatment, determining the prognosis (or likelihood of progression or regression) of a subject with a disease or condition, and determining the effectiveness of a treatment for a subject with a disease or condition. For example, a diagnostic can be used to detect the presence or likelihood of a subject suffering from a tumor, or the likelihood that such a subject will respond successfully to a compound (e.g., a pharmaceutical, e.g., a drug) or other treatment.
[0104] The term "isolated," when used with reference to a nucleic acid, such as an "isolated oligonucleotide," refers to a nucleic acid sequence that is identified and separated from at least one contaminant nucleic acid normally associated with its natural source. An isolated nucleic acid exists in a form or setting that is different from that in which it is found in nature. In contrast, non-isolated nucleic acids, such as DNA and RNA, are found in the state in which they exist in nature. Examples of non-isolated nucleic acids include a given DNA sequence (e.g., a gene) found adjacent to adjacent genes on a host cell chromosome; an RNA sequence, such as a particular mRNA sequence encoding a particular protein, that is found in a cell as a mixture with many other mRNAs encoding many proteins. However, an isolated nucleic acid encoding a particular protein includes, for example, a nucleic acid in a cell that normally expresses that protein, where the nucleic acid is in a location different from that in the chromosome of the natural cell or is otherwise flanked by different nucleic acid sequences than that in which it is found in nature. An isolated nucleic acid or oligonucleotide can exist in single-stranded or double-stranded form. When an isolated nucleic acid or oligonucleotide is used to express a protein, the oligonucleotide will minimally contain a sense or coding strand (i.e., the oligonucleotide may be single-stranded), but may also contain both a sense and an antisense strand (i.e., the oligonucleotide may be double-stranded). An isolated nucleic acid may be combined with other nucleic acids or molecules after isolation from its natural or typical environment. For example, an isolated nucleic acid may be present in a host cell, for example, for heterologous expression.
[0105] The term "purified" refers to a molecule, either a nucleic acid or an amino acid sequence, that has been removed, isolated, or separated from its natural environment. Thus, an "isolated nucleic acid sequence" can be a purified nucleic acid sequence. "Substantially purified" molecules are at least 60% free, preferably at least 75% free, and more preferably at least 90% free from other components with which they are naturally associated. As used herein, the terms "purified" or "to purify" also refer to the removal of contaminants from a sample. Removal of contaminating proteins results in an increase in the percentage of the polypeptide or nucleic acid of interest in a sample. In another example, recombinant polypeptides are expressed in plant, bacterial, yeast, or mammalian host cells, and these polypeptides are purified by removal of host cell proteins, thereby increasing the percentage of recombinant polypeptides in a sample.
[0106] The term "composition comprising" a given polynucleotide sequence or polypeptide refers broadly to any composition that includes the given polynucleotide sequence or polypeptide. Compositions can include aqueous solutions containing salts (e.g., NaCl), detergents (e.g., SDS), and other components (e.g., Denhardt's solution, milk powder, salmon sperm DNA, etc.).
[0107] The term "sample" is used in its broadest sense. In one sense, it can refer to animal cells or tissues. In another sense, it refers to specimens or cultures obtained from any source, as well as biological and environmental samples. Biological samples can be obtained from plants or animals (including humans) and can include fluids, solids, tissues, and gases. Environmental samples include environmental materials such as surface material, soil, water, and industrial samples. These examples are not to be construed as limiting the types of samples applicable to this disclosure.
[0108] As used herein, a "remote sample," as used in some contexts, refers to a sample that is indirectly collected from a site that is not the source of the cell, tissue, or organ sample. For example, if sample material derived from the pancreas is evaluated in a stool sample, the sample is a remote sample.
[0109] As used herein, the term "patient" or "subject" refers to an organism that is the subject of the various tests described herein. The term "subject" includes animals, preferably mammals, including humans. In preferred embodiments, the subject is a primate. In even more preferred embodiments, the subject is a human. Furthermore, with respect to diagnostic methods, preferred subjects are vertebrate subjects. Preferred vertebrates are warm-blooded animals, and preferred warm-blooded vertebrates are mammals. Preferred mammals are most preferably humans. As used herein, the term "subject" includes both human and animal subjects. Accordingly, veterinary uses are provided herein. Thus, the present disclosure provides for the diagnosis of mammals, such as humans, as well as mammals of endangered importance, such as the Amur tiger, mammals of economic importance, such as animals raised on farms for human consumption, and / or animals of social importance to humans, such as animals kept as pets or in zoos. Examples of such animals include, but are not limited to, carnivores such as cats and dogs; swine such as pigs, hogs, and wild boars; ruminants and / or ungulates such as cows, bulls, sheep, giraffes, deer, goats, bison, and camels; pinnipeds; and horses. Accordingly, diagnostics and treatments for livestock, including but not limited to domestic pigs, ruminants, ungulates, horses (including racehorses), and the like, are also provided. Embodiments of the present disclosure further include a system for diagnosing one or more types or subtypes of gynecological cancer in a subject. This system may be provided, for example, as a commercially available kit that can be used to screen for risk of one or more types or subtypes of gynecological cancer in a subject from whom a biological sample has been obtained, or to diagnose one or more types or subtypes of gynecological cancer. Exemplary systems provided according to various embodiments of the present disclosure include assessing the methylation status or profile of the markers described herein.
[0110] As used herein, the term "kit" refers to any delivery system for delivering materials. In the context of a reaction assay, such delivery systems include systems that allow for the storage, transport, or delivery of reaction reagents (e.g., oligonucleotides, enzymes, etc. in appropriate containers) and / or supporting materials (e.g., buffers, written instructions for conducting the assay, etc.) from one location to another. For example, a kit may include one or more enclosed members (e.g., boxes) that contain the relevant reaction reagents and / or supporting materials. As used herein, the term "fragmentation kit" refers to a delivery system that includes two or more separate containers, each containing a subportion of the overall kit components. The containers can be delivered to the intended recipient together or separately. For example, a first container may contain an enzyme for use in an assay, while a second container contains an oligonucleotide. The term "fragmentation kit" is intended to encompass, but is not limited to, a kit containing an analyte-specific reagent (ASR) as defined by Section 520(e) of the Federal Food, Drug, and Cosmetic Act. Indeed, any delivery system comprising two or more separate containers, each housing a portion of the overall kit's components, is encompassed by the term "fragmented kit." In contrast, a "combined kit" refers to a delivery system that contains all components of a reaction assay in a single container (e.g., in a single box housing each of the desired components). The term "kit" encompasses both fragmented and combined kits.
[0111] As used herein, the term "information" refers to any collection of facts or data. With respect to information stored or processed using computer system(s), including but not limited to the Internet, the term refers to any data stored in any format (e.g., analog, digital, optical, etc.). As used herein, the term "information about a subject" refers to facts or data about a subject (e.g., a human, plant, or animal). The term "genomic information" refers to information about a genome, including, but not limited to, nucleic acid sequences, genes, methylation percentages, allele frequencies, RNA expression levels, protein expression, phenotypes correlated with genotypes, etc. "Allele frequency information" refers to facts or data about allele frequencies, including, but not limited to, the identity of an allele, statistical correlations between the presence of an allele and characteristics of a subject (e.g., a human subject), the presence or absence of an allele in an individual or population, the percentage likelihood of an allele being present in an individual with one or more particular characteristics, etc.
[0112] 2. Methylation DNA markers and biomarker panels Embodiments of the present disclosure provide methods, compositions, and systems for screening for multiple types of gynecological cancer from a biological sample. According to these embodiments, the present disclosure includes, but is not limited to, methods and compositions for detecting the presence of multiple types or subtypes of gynecological cancer from a biological sample. In some embodiments, the biological sample is a tissue sample, a blood sample, a plasma sample, a serum sample, a whole blood sample, a secretion sample, an organ secretion sample, a cerebrospinal fluid (CSF) sample, a saliva sample, a urine sample, and / or a stool sample. In some embodiments, the tissue sample is a gynecological tissue sample including one or more of vaginal tissue, vaginal cells, cervical tissue, cervical cells, endometrial tissue, endometrial cells, ovarian tissue, and ovarian cells. In some embodiments, the tissue sample is an ovarian tissue sample, an endometrial tissue sample, or a cervical tissue sample. In some embodiments, the secretion sample is a secretion or discharge from any gynecological organ or tissue, including, but not limited to, vaginal tissue, cervical tissue, uterine tissue, endometrial tissue, and ovarian tissue. In some embodiments, the subject is a human.
[0113] As further described herein, embodiments of the present disclosure include novel differentially methylated regions (DMRs), each of which can independently distinguish between specific types of gynecological cancer (i.e., endometrial cancer (EC), or ovarian cancer (OC), or cervical cancer (CC)) and benign gynecological tissue samples. According to these embodiments, these novel DMRs include ADAM8, ADHFE1, AES, AGBL2, AIM1, AK5, ALKBH3, ARAP1, ARHGAP20, ASCL2, BCAT1, BEGAIN, BEND4_3696, BMP6, C12orf68, C13orf18, C14orf169_7694, C14orf169_8382, C18orf18, C1orf61, C20orf195, C4orf31, C5orf52, C6orf147, C7orf169_8382, C8orf169_9382, C9orf169_10382, C10orf169_11382, C11orf169_12382, C12orf68, C13orf18, C14orf169_13382, C14orf169_14382, C15orf169_15382, C16orf169_16382, C17orf169_17382, C18orf18, C1orf61, C20orf195, C4orf31, C5orf52, C6orf147, C7orf169_11382, C8orf169_12382, C8orf169_13382, C8orf169_14382, C8orf169_ rf51, CD14, CELF2, CHCHD5, CHMP2A, CHST10, CLIC6, CLIP4, COL13A1, COL19A1, COL6A2, COPZ2, CREB3L1, CXCL2, CXXC5, CYTH2, DAB2 IP, DGKZ, DLGAP3, DNASE2, DSCAML1, EBF1, EDARADD, EGR2, EIF5A2, ELMO1, ELMOD1, ELOVL4, EME2, EML6, EPSTI1, FADS2, FAM109B, FA M126A, FAM174B, FGF18, FKBP11, FLI1, FLOT1, FOXD3, FYN, GAL3ST2, GALR3, GAS7, GATA2_5878, GLT25D2, GNB2, HDAC7, HIC1, HLA-F, HNRNPF, HPDL, HS3ST4, HSPA1A, IDUA, IGSF9B, IL12RB2, IRAK3, IRF7, IRF8, ITPKA, KCNA2, KCNC3_6487, KCNC3_7105, KCNC4, KCNH8, KDM2B, LBX2, LCMT2, LOC100129726, LOC100287216, LOC255130, LOC339290, LOC729678, LPPR3, LRRC41, LRRC8D_8856, LTBP2, LYPL AL1, MAST4, MAX.chr1.2152, HIVEP3, GRAMD1B, MAX.chr11.0394, MAX.chr11.3750, FAT3, SLC16A7, MTUS2, LINC02323, MAX.chr14.7696、MCTP2、LOC107984974、TRIM80P、MAX.chr19.5552、ZNF433-AS1、ZNF254、MAX.chr19.0548、B3GALT1、MAX.chr2.8918、MAX.chr2.4778、MAX.chr20.3853、MAX.chr20.2903、MAX.chr21.5011、DSCR9、MAX.chr22.5665、MAX.chr3.6408、LINC02028、LINC02084、MAX.chr5.3588、CTD-2532K18.1、HS3ST5、ARHGAP18、GRM4、LINC01004、MAX.chr8.5938、MAX.chr9.4007、MAX.chr9.2025、TRPM3、MED12L、MIAT、MLH1_4513、MLH1_5193、MMP16、MRPS21、MSI1、MT1E、MX1、MYC、MYH10、MYO15B、N4BP 2L1、NBR1、NDRG2、NEGR1、NEU1、NOL3、NR3C1_2223、NR3C1_4614、NRP2、NTN1、NTNG1、PAPL、PAQR9、PDE10A、PDE3B , PDE4A, PDXK, PER1, PISD, PLEC, PLIN2, PLXND1, PPM1E, PPP1R9A, PPP2R5C, PRDM5, PTP4A3, PYCARD, RAB3C, RAI1, RARG, RASA3, RPRM, RREB1, S100A6, SAMD5, SBNO2, SDC2, SDK2, SELM, SERP2, SFMBT2_2029, SHF, SHH, SLC16A11, SLC16A5、SLC25A22、SLCO3A1、SMTN、SPDYA、SPINK2、SPOCK2、SPON1、SQSTM1_4156、ST8SIA1、TAF4B、TAF7、TEAD3 、TERC、TIAM1、TLE4、TMEM101、TMEM106A、TRIM9、TRPC3、TSC22D4、TSPAN2、TSPAN5、TTC14、UBB_4001、UBB_4646、 UST、VAMP5、VIM、VSTM2B、ZBTB7B、ZEB2、ZFP3、ZFP36L2、ZIC2、ZMIZ1、ZNF14、ZNF211、ZNF280B、ZNF302、ZNF382、 ZNF480、ZNF483、ZNF491、ZNF569、ZNF610、ZNF702P、ZNF709、ZNF773、ZNF845、ZNF91、CDH4、LRRC34、MAX.chr10.The novel DMR(s) comprise one or more CpG sites in 4460, NBPF24, OBSCN, SEPT9, ZNF323, ZNF506, and / or ZNF90 (Table 1), including any combination thereof. In some embodiments, the novel DMR(s) are derived from any gene or region selected from Table 1, including any combination thereof. While each novel DMR alone can distinguish one or more gynecological cancers from control samples, combining two or more of the novel DMRs may improve sensitivity. Thus, combinations of two or more novel DMRs selected from Table 1 are provided.
[0114] Embodiments of the present disclosure also include novel differentially methylated regions (DMRs) that can individually distinguish gynecological cancer from benign gynecological tissue samples, and these DMRs are prevalent in all three types of gynecological cancers (i.e., endometrial cancer (EC), ovarian cancer (OC), and cervical cancer (CC)).According to these embodiments, these novel DMRs include ACSF2, AJAP1, ARL10, ARL5C, ASCL4, ATP6V1B1, BARHL1, BEND4_2963, C17orf64, C1QL3, C2orf55, C4orf48, CA3, CDO1, CELF2, CLEC14A, CSDAP1, CYTH2_4197, DLGAP1, DSCR6, EPS8L1_2819, EPS8L1_8496, FAIM2, FGF12, GATA2, HIST1H2BE, IRF4, IRX4, ITGA5, KCNA1, LECT1, LHX1, LOC440925, LPHN1, LINC02767, MAX.chr1.2533, SOX1-OT, MAX.chr13.3357, MAX.chr14.2093, MAX.c hr17.2455, MAX.chr18.4390, MAX.chr19.2732, MAX.chr19.4467, PANTR1, MAX.chr2.0490, MAX.chr2.8148, MAX.chr2.3137, RIPOR3, S CRG1, MAX.chr4.4210, HMX1, CTC-359M8.1, MAX.chr5.0931, MAX.chr5.9924, LIN28B, MAX.chr6.9522, TTLL2, RNA5SP243, DLGAP2, MEX 3B, MNX1, NEFL, NETO1, PAX2, PDX1, psiTPTE22, RASGEF1A, SALL3_9136, SALL3_0615, SEZ6L2, SHANK2, SHANK3, SKI, SLC35D3, SORCS3_03 In some embodiments, the novel DMR(s) comprise one or more CpG sites in any gene or region selected from Table 2, including any combination thereof.Each novel DMR alone can distinguish one or more gynecological cancers from control samples, and combining two or more of the novel DMRs may improve sensitivity. Thus, combinations of two or more novel DMRs selected from Table 2 are provided.
[0115] Embodiments of the present disclosure also include novel differentially methylated regions (DMRs), each of which can individually distinguish specific subtypes of gynecological cancer (i.e., serous ovarian cancer, clear cell ovarian cancer, endometrioid ovarian cancer, mucinous ovarian cancer, cervical adenocarcinoma, cervical squamous cell carcinoma, or endometrioid endometrioid carcinoma) from benign gynecological tissue samples. According to these embodiments, the novel DMRs comprise one or more CpG sites in AIM1, AK5, c18orf18, CDO1, DLGAP1, ELMOD1, FKBP11, FLOT1, GAL3ST2, LRRC41, LYPLAL1, MAX.chr11.3750, MLH1_4513, NR3C1_2223, PISD.RABC3, RAI1, TERC, TRPC3, ZIC2, ZMIZ1, ZNF480, ZNF491, ZNF610, and / or ZNF91 (Table 3), including any combination thereof. In some embodiments, the novel DMR(s) are derived from any gene or region selected from Table 3, including any combination thereof. While each novel DMR alone can distinguish one or more gynecological cancers from control samples, combining two or more of the novel DMRs may improve sensitivity. Accordingly, combinations of two or more novel DMRs selected from Table 3 are provided.
[0116] Embodiments of the present disclosure also include novel differentially methylated regions (DMRs), each of which can individually distinguish between specific subtypes of gynecological cancer (i.e., serous ovarian cancer, clear cell ovarian cancer, endometrioid ovarian cancer, mucinous ovarian cancer, cervical adenocarcinoma, cervical squamous cell carcinoma, or endometrioid endometrioid carcinoma) and benign gynecological tissue samples. According to these embodiments, these novel DMRs comprise one or more CpG sites in LBX2, SPDYA, TERC, ZSCAN12, CYP26C1, and / or GYPC (Table 4), including any combination thereof. In some embodiments, the novel DMR(s) are derived from any gene or region selected from Table 4, including any combination thereof. Each novel DMR alone can distinguish between one or more gynecological cancer and control samples, and combining two or more of the novel DMRs may improve sensitivity. Accordingly, combinations of two or more novel DMRs selected from Table 4 are provided.
[0117] Embodiments of the present disclosure also include novel differentially methylated regions (DMRs), each of which can individually distinguish specific subtypes of gynecological cancer (i.e., serous ovarian cancer, clear cell ovarian cancer, endometrioid ovarian cancer, mucinous ovarian cancer, cervical adenocarcinoma, cervical squamous cell carcinoma, or endometrioid endometrioid carcinoma) from benign gynecological tissue samples. According to these embodiments, the novel DMRs comprise one or more CpG sites in KRT86, CDH4, c17orf64, EMX2OS, NBPF24, SFMBT2_0970, JSRP1, DIDO1, MAX.chr10.4460, MPZ, ZNF506, GATA2_6370, VILL, LINC02323, CYTH2_4043, LRRC8D_8831, LYPLAL1, SMPD5, SQSTM1_3864, ZNF323, OBSCN, ZNF90, LRRC34, GDF7, MDFI, EEF1A2, LRRC41, and / or SEPT9 (Table 8), including any combination thereof. In some embodiments, the novel DMR(s) are derived from any gene or region selected from Table 8, including any combination thereof. Each novel DMR alone can distinguish one or more gynecological cancers from control samples, and combining two or more of the novel DMRs may improve sensitivity. Thus, combinations of two or more novel DMRs selected from Table 8 are provided.
[0118] As described in the preceding examples, experiments were conducted to identify DMRs (also referred to herein as methylated DNA markers (MDMs)) that can distinguish between types and subtypes of gynecological cancer and controls (e.g., normal or benign samples). These experiments involved validation studies of the utility and performance of panels of methylated DNA markers and proteins for detecting one or more types or subtypes of gynecological cancer by testing independent sets of case / control samples with the refined marker panel. As a result of such experiments, MDMs were identified that are useful for simultaneously detecting the presence of multiple types of gynecological cancer (i.e., endometrial cancer (EC), ovarian cancer (OC), or cervical cancer (CC)) from benign gynecological tissue samples (e.g., stool samples, tissue samples, organ samples, secretion samples (e.g., vaginal secretion samples), CSF samples, saliva samples, blood samples, plasma samples, or urine samples).
[0119] In some embodiments, the present disclosure provides compositions and methods for identifying, determining, and / or classifying multiple types or subtypes of gynecological cancer from a biological sample (e.g., a stool sample, a tissue sample, an organ sample, a secretion sample, a CSF sample, a saliva sample, a blood sample, a plasma sample, or a urine sample). The methods generally involve determining a methylation profile of at least one methylation marker in a biological sample isolated from a subject. In some embodiments, a change in the methylation status or profile of the marker indicates the presence, class, or location of a particular type of gynecological cancer. Generally, such methods are not limited to detecting the presence or absence of a particular type or subtype of gynecological cancer. In some embodiments, cancer types and subtypes include, but are not limited to, endometrial cancer, ovarian cancer, cervical cancer, serous ovarian cancer, clear cell ovarian cancer, endometrioid ovarian cancer, mucinous ovarian cancer, cervical adenocarcinoma, cervical squamous cell carcinoma, and endometrioid endometrioid carcinoma.
[0120] In some embodiments, a method is provided that includes contacting nucleic acid (e.g., genomic DNA) in a biological sample obtained from a subject with at least one reagent or set of reagents that distinguish between methylated and unmethylated nucleotides (e.g., CpG dinucleotides) within at least one methylation marker, and detecting the presence or absence of one or more types or subtypes of gynecological cancer (e.g., provided with a sensitivity of 80% or greater and a specificity of 80% or greater).
[0121] In some embodiments, methods are provided that include measuring one or both of the methylation levels of one or more genes or methylated DNA markers in a biological sample from a human individual by treating genomic DNA in the biological sample with a reagent that modifies DNA in a methylation-specific manner, amplifying the treated genomic DNA using a set of primers for selected one or more genes or methylation markers, and determining the methylation levels of the one or more genes or methylation markers.
[0122] In some embodiments, methods are provided that include measuring the amount of one or more methylated DNA markers or genes in DNA of a biological sample, measuring the amount of at least one reference marker in the DNA, and calculating a value for the amount of the at least one methylation marker gene measured in the DNA as a percentage of the amount of the reference marker gene measured in the DNA, the value representing the amount of the at least one methylation marker DNA measured in the biological sample.
[0123] In some embodiments, methods are provided that include measuring the methylation level of CpG sites of one or more genes in a biological sample from a human individual by treating genomic DNA in the biological sample with bisulfite, a reagent that can modify DNA in a methylation-specific manner; amplifying the modified genomic DNA using a set of primers for the selected one or more genes; and determining the methylation level of the CpG sites of the selected one or more genes.
[0124] In some embodiments, the present disclosure provides a method for characterizing a biological sample, comprising measuring one or both methylation levels of CpG sites of one or more genes in a biological sample from a human individual by treating genomic DNA in the biological sample with bisulfite, amplifying the bisulfite-treated genomic DNA using a set of primers for one or more selected genes, and determining the methylation levels of the CpG sites. In some embodiments, the method comprises comparing the methylation levels of one or both of the methylation markers with the methylation levels of the corresponding set of genes in a control sample that does not have the particular type of cancer, and / or determining that the subject has one or more types or subtypes of gynecological cancer if one or both of the methylation levels measured for the one or more genes are higher than the methylation levels measured in the respective control sample.
[0125] In some embodiments, the present disclosure provides methods that include one or all of measuring the methylation level of one or more genes or markers in a biological sample by treating genomic DNA in the biological sample with bisulfite, amplifying the bisulfite-treated genomic DNA using a set of primers for one or more selected genes, and determining the methylation level of the one or more genes or markers.
[0126] In some embodiments, the present disclosure provides methods of screening for one or more types or subtypes of gynecological cancer in a sample obtained from a subject. According to these embodiments, the methods include one or both of assaying the methylation status or profile of one or more methylated DNA markers and identifying the subject as having one or more types or subtypes of gynecological cancer if the methylation status or profile of the markers differs from the methylation status or profile of the markers assayed in a subject who does not have one or more types of cancer.
[0127] In some embodiments, the present disclosure provides methods that include measuring the methylation level of one or more genes or markers in a biological sample from a human individual by treating genomic DNA in the biological sample with a reagent that modifies the DNA in a methylation-specific manner, amplifying the treated genomic DNA using a set of primers for selected one or more genes or markers, and determining the methylation level of the one or more genes or markers.
[0128] In some embodiments, the present disclosure provides a method for characterizing a biological sample, comprising measuring the amount of at least one methylated DNA marker in DNA extracted from the biological sample, treating genomic DNA in the biological sample with bisulfite, and amplifying the bisulfite-treated genomic DNA using primers specific for CpG sites of each marker. In some embodiments, the primers specific for each marker can bind to amplicons bounded by the primer sequences of the markers listed in Table 1 or 2 (the amplicons bounded by the primer sequences of the markers are at least a portion of the gene regions of the methylated markers listed in Table 1 or 2), allowing the methylation level of the CpG sites of one or more genes to be determined.
[0129] In some embodiments, the present disclosure provides a method comprising extracting genomic DNA from a biological sample of a human individual suspected of or having one or more types or subtypes of gynecological cancer, measuring the methylation level of one or more methylated DNA markers in the DNA extracted from the biological sample, treating the extracted genomic DNA with bisulfite, and amplifying the bisulfite-treated genomic DNA with primers specific for the one or more markers. In some embodiments, the primers specific for the one or more markers can bind to at least a portion of the bisulfite-treated genomic DNA of a chromosomal region of the marker (e.g., one or more markers listed in Table 1 or 2) and measure the methylation level of the one or more methylation markers.
[0130] In some embodiments, the present disclosure provides methods comprising extracting genomic DNA from a biological sample of a human individual suspected of having or having one or more types or subtypes of gynecological cancer, thereby measuring the methylation level of one or more methylated DNA markers in the DNA extracted from the biological sample; treating the extracted genomic DNA with bisulfite; amplifying the bisulfite-treated genomic DNA with primers specific for one or more markers, wherein the primers specific for the one or more markers are capable of binding to at least a portion of the bisulfite-treated genomic DNA of a chromosomal region of the markers listed in Table 2; and measuring the methylation level of the one or more methylation markers.
[0131] In some embodiments, the present disclosure provides a method comprising extracting genomic DNA from a biological sample of a human individual suspected of having or having one or more types or subtypes of gynecological cancer, thereby measuring the methylation level of one or more methylated DNA markers in the DNA extracted from the biological sample; treating the extracted genomic DNA with bisulfite; amplifying the bisulfite-treated genomic DNA with primers specific for one or more markers, wherein the primers specific for the one or more markers are capable of binding to at least a portion of the bisulfite-treated genomic DNA of a chromosomal region of the markers listed in Table 3; and measuring the methylation level of the one or more methylation markers.
[0132] In some embodiments, the present disclosure provides methods comprising extracting genomic DNA from a biological sample of a human individual suspected of having or having one or more types or subtypes of gynecological cancer, thereby measuring the methylation level of one or more methylated DNA markers in the DNA extracted from the biological sample; treating the extracted genomic DNA with bisulfite; amplifying the bisulfite-treated genomic DNA with primers specific for one or more markers, wherein the primers specific for the one or more markers are capable of binding to at least a portion of the bisulfite-treated genomic DNA of a chromosomal region of the markers listed in Table 4; and measuring the methylation level of the one or more methylation markers.
[0133] In some embodiments, the present disclosure provides methods comprising extracting genomic DNA from a biological sample of a human individual suspected of having or having one or more types or subtypes of gynecological cancer, thereby measuring the methylation level of one or more methylated DNA markers in the DNA extracted from the biological sample; treating the extracted genomic DNA with bisulfite; amplifying the bisulfite-treated genomic DNA with primers specific for one or more markers, wherein the primers specific for the one or more markers are capable of binding to at least a portion of the bisulfite-treated genomic DNA of a chromosomal region of a marker listed in Table 8; and measuring the methylation level of the one or more methylation markers.
[0134] In some embodiments, the present disclosure provides a method comprising extracting genomic DNA from a biological sample of a human individual suspected of having or having cancer, treating the extracted genomic DNA with bisulfite, amplifying the bisulfite-treated genomic DNA using separate primers specific for one or more CpG sites of methylated DNA markers, and measuring the methylation level of each CpG site of the one or more markers.
[0135] In some embodiments, the present disclosure provides methods for preparing a DNA fraction from a biological sample of a human individual, useful for analyzing one or more genetic loci involved in one or more chromosomal abnormalities. According to these embodiments, the method includes extracting genomic DNA from the biological sample of the human individual, treating the extracted genomic DNA with a reagent that modifies the DNA in a methylation-specific manner to generate a fraction of extracted genomic DNA, amplifying the bisulfite-treated genomic DNA using separate primers specific for one or more methylated DNA markers, and analyzing one or more genetic loci in the generated fraction of extracted genomic DNA by measuring the methylation level of CpG sites for each of the one or more markers.
[0136] In some embodiments, the present disclosure provides methods for preparing a DNA fraction from a biological sample of a human individual, useful for analyzing one or more DNA fragments involved in one or more chromosomal abnormalities. According to these embodiments, the method includes extracting genomic DNA from the biological sample of the human individual, treating the extracted genomic DNA with a reagent that modifies DNA in a methylation-specific manner to generate a fraction of extracted genomic DNA, amplifying the bisulfite-treated genomic DNA using separate primers specific for one or more methylated DNA markers, and analyzing one or more DNA fragments in the generated fraction of extracted genomic DNA by measuring the methylation level of each CpG site of the one or more markers.
[0137] As will be understood by those skilled in the art based on the present disclosure, the various methods described herein are not limited to the use of any one particular methylated DNA marker, methylation marker gene, methylation gene, and / or DMR. That is, one or more of the methylated DNA markers, methylation marker genes, methylation genes, and / or DMRs disclosed herein can be used to distinguish and / or identify one or more types or subtypes of gynecological cancer (including any combination thereof). Furthermore, the methylated DNA markers, methylation marker genes, methylation genes, and / or DMRs disclosed herein can include regions or subregions (e.g., genes on chromosomes, single nucleotides, CpG islands, etc.) of any of the markers described herein.
[0138] In some embodiments, at least one DMR comprises one or more CpG sites in AIM1, FLOT1, GAL3ST2, LRRC41, LYPLAL1, MAX.chr11.3750, PISD, RAI1, ZIC2, ZMIZ1, CDH4, ZNF506, ZNF323, OBSCN, ZNF90, and / or SEPT9, and the subject has or is suspected of having ovarian cancer (OC). In some embodiments, at least one DMR comprises one or more CpG sites in AIM1, FLOT1, GAL3ST2, LYPLAL1, and / or OBSCN, and the subject has or is suspected of having serous OC. In some embodiments, at least one DMR comprises one or more CpG sites in LRRC41, PISD, ZIC2, OBSCN, and / or SEPT9, and the subject has or is suspected of having clear cell OC. In some embodiments, at least one DMR includes one or more CpG sites in MAX.chr11.3750, and the subject has or is suspected of having endometrioid OC. In some embodiments, at least one DMR includes one or more CpG sites in RAI1 and / or ZMIZ1, and the subject has or is suspected of having mucinous OC. In some embodiments, determining the methylation profile of one or more CpG sites in AIM1, FLOT1, GAL3ST2, LRRC41, LYPLAL1, MAX.chr11.3750, PISD, RAI1, ZIC2, ZMIZ1, CDH4, ZNF506, ZNF323, OBSCN, ZNF90, and / or SEPT9 comprises comparing the methylation profile to corresponding regions of a control DNA sample obtained from a subject without OC.
[0139] In some embodiments, at least one DMR comprises one or more CpG sites in AK5, ELMOD1, RABC3, TRPC3, ZNF480, ZNF491, ZNF610, ZNF91, and / or NBPF24, and the subject has or is suspected of having cervical cancer (CC). In some embodiments, at least one DMR comprises one or more CpG sites in AK5, ELMOD1, TRPC3, and / or ZNF480, and the subject has or is suspected of having cervical adenocarcinoma (adenocarcinoma CC). In some embodiments, at least one DMR comprises one or more CpG sites in ZNF491, ZNF610, ZNF91, and / or NBPF24, and the subject has or is suspected of having squamous cell CC. In some embodiments, determining the methylation profile of one or more CpG sites in AK5, ELMOD1, RABC3, TRPC3, ZNF480, ZNF491, ZNF610, and / or ZNF91 comprises comparing the methylation profile to corresponding regions of a control DNA sample obtained from a subject who does not have CC.
[0140] In some embodiments, at least one DMR comprises one or more CpG sites in c18orf18, FKBP11, MLH1, NR3C1, and / or TERC, and the subject has or is suspected of having endometrial cancer (EC). In some embodiments, at least one DMR comprises one or more CpG sites in MLH1 and / or SEPT9, and the subject has or is suspected of having clear cell EC. In some embodiments, at least one DMR comprises one or more CpG sites in NR3C1, and the subject has or is suspected of having endometrioid EC. In some embodiments, determining the methylation profile of one or more CpG sites in c18orf18, FKBP11, MLH1, NR3C1, and / or TERC comprises comparing the methylation profile to corresponding regions of a control DNA sample obtained from a subject without EC.
[0141] In some embodiments, at least one DMR comprises one or more CpG sites in CDO1 and / or DLGAP1, and the subject has or is suspected of having CC, OC, or EC. In some embodiments, determining the methylation profile of the one or more CpG sites CDO1 and / or DLGAP1 comprises comparing the methylation profile to corresponding regions of a control DNA sample obtained from a subject who does not have CC, OC, or EC.
[0142] In some embodiments, the methods of the disclosure include determining a methylation profile of one or more CpG sites in AIM1, FLOT1, GAL3ST2, LRRC41, LYPLAL1, MAX.chr11.3750, PISD, RAI1, ZIC2, and / or ZMIZ1. In some embodiments, the methods include determining the methylation profile of one or more CpG sites in AK5, ELMOD1, RABC3, TRPC3, ZNF480, ZNF491, ZNF610, and / or ZNF91. In some embodiments, the methods include determining the methylation profile of one or more CpG sites in c18orf18, FKBP11, MLH1, NR3C1, and / or TERC.
[0143] In some embodiments, at least one DMR comprises NBPF24, and the subject has or is suspected of having CC. In some embodiments, determining the methylation profile of NBPF24 comprises comparing the methylation profile with a corresponding region of a control DNA sample obtained from a subject who does not have CC.
[0144] In some embodiments, at least one DMR includes one or more CpG sites in CDH4, NBPF24, MAX.chr10.4460, ZNF506, ZNF323, OBSCN, ZNF90, LRRC34, SFMBT2, LINC02323, CYTH2, LRRC8D, LYPLAL1, LRRC41, and / or SEPT9, and the subject has or is suspected of having EC. In some embodiments, determining the methylation profile of the one or more CpG sites in CDH4, NBPF24, MAX.chr10.4460, ZNF506, ZNF323, OBSCN, ZNF90, LRRC34, SFMBT2, LINC02323, CYTH2, LRRC8D, LYPLAL1, LRRC41, and / or SEPT9 comprises comparing the methylation profile to corresponding regions of a control DNA sample obtained from a subject without EC.
[0145] In some embodiments, at least one DMR includes one or more CpG sites in CDH4, ZNF506, ZNF323, OBSCN, ZNF90, SFMBT2, LINC02323, CYTH2, LRRC8D, LYPLAL1, LRRC41, and / or SEPT9, and the subject has or is suspected of having OC. In some embodiments, determining the methylation profile of the one or more CpG sites in CDH4, ZNF506, ZNF323, OBSCN, ZNF90, SFMBT2, LINC02323, CYTH2, LRRC8D, LYPLAL1, LRRC41, and / or SEPT9 includes comparing the methylation profile to corresponding regions of a control DNA sample obtained from a subject without OC.
[0146] In some embodiments, at least one DMR includes one or more CpG sites in KRT86, EMX2OS, JSRP1, DIDO1, MPZ, VILL, SMPD5, GDF7, MDFI, c17orf64, GATA2, SQSTM1, and / or EEF1A2, and the subject has or is suspected of having CC, OC, or EC. In some embodiments, determining the methylation profile of the one or more CpG sites in KRT86, EMX2OS, JSRP1, DIDO1, MPZ, VILL, SMPD5, GDF7, MDFI, c17orf64, GATA2, SQSTM1, and / or EEF1A2 includes comparing the methylation profile to corresponding regions of a control DNA sample obtained from a subject without CC, OC, or EC.
[0147] As will be appreciated by those of skill in the art based on the present disclosure, one or more types or subtypes of gynecological cancer can be predicted by various combinations of markers (e.g., identified by statistical methods related to the specificity and sensitivity of the prediction). Embodiments of the present disclosure provide methods for identifying predictive combinations and validated predictive combinations for one or more types or subtypes of gynecological cancer.
[0148] Such methods are not limited to the type of subject. In some embodiments, the subject is a mammal. In some embodiments, the subject is a human. Such methods are not limited to a particular manner or technique for measuring protein expression and / or activity. Techniques for measuring protein expression and / or activity levels are known in the art. Indeed, any known technique for measuring protein expression and / or activity levels is contemplated and incorporated herein.
[0149] Such methods are not limited to a particular modality or technique for determining the methylation characterization, measurement, or assay for one or more methylation markers, methylation marker genes, genes, DMRs, and / or DNA methylation markers. In some embodiments, such techniques are based on analysis of the methylation status of at least one marker, marker region, or marker base (e.g., CpG methylation status) that comprises a DMR.
[0150] In some embodiments, measuring the methylation state or profile of a methylated DNA marker in a sample comprises determining the methylation state of a single nucleotide base. In some embodiments, measuring the methylation state of a methylated DNA marker in a sample comprises determining the degree of methylation at multiple nucleotide bases. Further, in some embodiments, the methylation state or profile of a methylated DNA marker comprises an increase in methylation of the marker relative to the marker's normal methylation state or profile. In some embodiments, the methylation state or profile of a marker comprises a decrease in methylation of the marker relative to the marker's normal methylation state. In some embodiments, the methylation state or profile of a marker comprises a different pattern of methylation of the marker compared to the marker's normal methylation state or profile.
[0151] Further, in some embodiments, the marker is a region of 100 or fewer nucleotide bases. In some embodiments, the marker is a region of 500 or fewer nucleotide bases. In some embodiments, the marker is a region of 1000 or fewer nucleotide bases. In some embodiments, the marker is a region of 5000 or fewer nucleotide bases. In some embodiments, the marker is one nucleotide base. In some embodiments, the marker is within a high CpG density promoter region.
[0152] In certain embodiments, methods for analyzing nucleic acids for the presence of 5-methylcytosine include treating DNA with reagents that modify DNA in a methylation-specific manner, examples of such reagents include, but are not limited to, methylation-sensitive restriction enzymes, methylation-dependent restriction enzymes, bisulfite reagents, TET enzymes, and borane reducing agents.
[0153] A frequently used method for analyzing nucleic acids for the presence of 5-methylcytosine is based on the bisulfite method described by Frommer et al. for detecting 5-methylcytosine or its mutations in DNA (Frommer et al. (1992) Proc. Natl. Acad. Sci. USA 89:1827-31, expressly incorporated herein by reference in its entirety for all purposes). The bisulfite method for mapping 5-methylcytosine is based on the finding that cytosine, but not 5-methylcytosine, reacts with hydrogen sulfite ions (also known as bisulfite). The reaction is typically carried out according to the following steps: first, cytosine reacts with bisulfite to form sulfonated cytosine; second, spontaneous deamination of the sulfonated reaction intermediate results in sulfonated uracil; and finally, the sulfonated uracil is desulfonated under alkaline conditions to form uracil. Detection is possible because uracil base pairs with adenine (and therefore behaves like thymine), whereas 5-methylcytosine base pairs with guanine (and therefore behaves like cytosine). This allows methylated cytosines to be distinguished from unmethylated cytosines, for example, by bisulfite genomic sequencing (Grigg G, & Clark S, Bioessays (1994) 16:431-36; Grigg G, DNA Seq. (1996) 6:189-98), methylation-specific PCR (MSP) as disclosed in U.S. Pat. No. 5,786,146, or using assays involving sequence-specific cleavage, such as the QuARTS flap endonuclease assay (see, e.g., Zou et al. (2010) "Sensitive quantification of methylated markers with a novel methylation-specific technology" Clin Chem 56:A199; and U.S. Pat. Nos. 8,361,720, 8,715,937, 8,916,344, and 9,212,392).
[0154] In some embodiments, conventional techniques involve encapsulating the DNA to be analyzed in an agarose matrix to prevent DNA diffusion and renaturation (bisulfite reacts only with single-stranded DNA) and replacing the precipitation and purification steps with high-speed dialysis (Olek A, et al. (1996) "A modified and improved method for bisulfite-based cytosine methylation analysis" Nucleic Acids Res. 24:5064-6). Thus, it is possible to analyze individual cells for methylation status, demonstrating the utility and sensitivity of the method. An overview of conventional methods for detecting 5-methylcytosine is provided by Rein, T., et al. (1998) Nucleic Acids Res. 26:2255.
[0155] Bisulfite methods generally involve bisulfite treatment followed by amplification of short, specific fragments of known nucleic acids, followed by assay of the products by sequencing (Olek & Walter (1997) Nat Genet. 17:275-6) or primer extension reactions (Gonzalgo & Jones (1997) Nucleic Acids Res. 25:2529-31; WO 95 / 00669; U.S. Patent No. 6,251,594) to analyze individual cytosine positions. Some methods use enzymatic digestion (Xiong & Laird (1997) Nucleic Acids Res. 25:2532-4). Hybridization detection has also been described in the art (Olek et al., WO 99 / 28498). Additionally, the use of bisulfite techniques for methylation detection of individual genes has been described (Grigg & Clark (1994) Bioessays 16:431-6; Zeschnigk et al. (1997) Hum Mol Genet. 6:387-95; Feil et al. (1994) Nucleic Acids Res. 22:695; Martin et al. (1995) Gene 157:261-4; WO9746705; WO9515373).
[0156] Various methylation assay techniques can be used in conjunction with bisulfite treatment according to the techniques of the present invention. These assays allow for the determination of the methylation status of one or more CpG dinucleotides (e.g., CpG islands) within a nucleic acid sequence. Such assays involve, among other techniques, sequencing of bisulfite-treated nucleic acids, PCR (for sequence-specific amplification), Southern blot analysis, and the use of methylation-specific enzymes, e.g., methylation-sensitive or methylation-dependent enzymes.
[0157] For example, genome sequencing has been simplified for the analysis of methylation patterns and 5-methylcytosine distribution by using bisulfite treatment (Frommer et al. (1992) Proc. Natl. Acad. Sci. USA 89:1827-1831). Furthermore, restriction enzyme digestion of PCR products amplified from bisulfite-converted DNA can be used to assess methylation status, as described, for example, in Sadri & Hornsby (1997) Nucl. Acids Res. 24:5058-5059, or as implemented in a method known as COBRA (Combined Bisulfite Restriction Analysis) (Xiong & Laird (1997) Nucleic Acids Res. 25:2532-2534).
[0158] The COBRA™ analysis is a quantitative methylation assay useful for determining DNA methylation levels at specific loci in small amounts of genomic DNA (Xiong & Laird, Nucleic Acids Res. 25:2532-2534, 1997). Briefly, restriction enzyme digestion is used to reveal methylation-dependent sequence differences in PCR products of sodium bisulfite-treated DNA. Methylation-dependent sequence differences are first introduced into genomic DNA by standard bisulfite treatment according to the procedure described by Frommer et al. (Proc. Natl. Acad. Sci. USA 89:1827-1831, 1992). PCR amplification of the bisulfite-converted DNA is then performed using primers specific for the CpG island of interest, followed by restriction endonuclease digestion, gel electrophoresis, and a specifically labeled hybridization probe. Methylation levels in the original DNA sample are represented by the relative amounts of digested and undigested PCR products, providing a linear method for quantitating a wide range of DNA methylation levels. In addition, this method can be reliably applied to DNA obtained from microdissected paraffin-embedded tissue samples.
[0159] Typical reagents for COBRA™ analysis (e.g., as might be found in a typical COBRA™-based kit) can include, but are not limited to: PCR primers for specific loci (e.g., specific genes, markers, DMRs, regions of genes, regions of markers, bisulfite-treated DNA sequences, CpG islands, etc.); restriction enzymes and appropriate buffers; gene hybridization oligonucleotides; control hybridization oligonucleotides; kinase labeling kits for oligonucleotide probes; and labeled nucleotides. Additionally, bisulfite conversion reagents can include DNA denaturing buffers; sulfonation buffers; DNA recovery reagents or kits (e.g., precipitation, ultrafiltration, affinity columns); desulfonation buffers; and DNA recovery components.
[0160] Assays such as "MethyLight™" (fluorescence-based real-time PCR technology) (Eads et al., Cancer Res. 59:2302-2306, 1999), Ms-SNuPE™ (methylation-sensitive single nucleotide primer extension) reactions (Gonzalgo & Jones, Nucleic Acids Res. 25:2529-2531, 1997), methylation-specific PCR ("MSP"; Herman et al., Proc. Natl. Acad. Sci. USA 93:9821-9826, 1996; U.S. Patent No. 5,786,146), and methylated CpG island amplification ("MCA"; Toyota et al., Cancer Res. 59:2307-12, 1999) are used alone or in combination with one or more of these methods.
[0161] The "HeavyMethyl™" assay technique is a quantitative method for assessing methylation differences based on methylation-specific amplification of bisulfite-treated DNA. Methylation-specific inhibitor probes ("inhibitors") covering CpG positions between or covered by the amplification primers allow for methylation-specific selective amplification of the sample.
[0162] The term "HeavyMethyl™ MethyLight™" assay refers to the HeavyMethyl™ MethyLight™ assay, which is a variation of the MethyLight™ assay in which the MethyLight™ assay is combined with a methylation-specific blocking probe that covers the CpG positions between the amplification primers. The HeavyMethyl™ assay can also be used in combination with methylation-specific amplification primers.
[0163] Typical reagents for HeavyMethyl analysis (e.g., as found in a typical MethyLight™-based kit) include, but are not limited to, PCR primers for specific loci (e.g., specific genes, markers, gene regions, marker regions, bisulfite-treated DNA sequences, CpG islands, or bisulfite-treated DNA sequences or CpG islands, etc.); blocking oligonucleotides; optimized PCR buffers and deoxynucleotides; and Taq polymerase.
[0164] MSP (methylation-specific PCR) allows assessment of the methylation status of virtually all CpG sites within a CpG island, regardless of the use of methylation-sensitive restriction enzymes (Herman et al. Proc. Natl. Acad. Sci. USA 93:9821-9826, 1996; U.S. Patent No. 5,786,146). Briefly, DNA is modified with sodium bisulfite, which converts unmethylated cytosines, but not methylated cytosines, to uracil, and these products are then amplified with primers specific for methylated versus unmethylated DNA. MSP requires only small amounts of DNA, is sensitive to 0.1% methylated alleles of a given CpG island locus, and can be performed on DNA extracted from paraffin-embedded samples. Typical reagents for MSP analysis (e.g., as found in a typical MSP-based kit) include, but are not limited to, methylated and unmethylated PCR primers for specific loci (e.g., specific genes, markers, gene regions, marker regions, bisulfite-treated DNA sequences, CpG islands, etc.); optimized PCR buffers and deoxynucleotides, and specific probes.
[0165] The MethyLight™ assay is a high-throughput quantitative methylation assay that utilizes fluorescence-based real-time PCR (e.g., TaqMan®) and requires no further manipulation after the PCR step (Eads et al., Cancer Res. 59:2302-2306, 1999). Briefly, the MethyLight™ process begins with a mixed sample of genomic DNA that is converted into a mixed pool of methylation-dependent sequence differences in a sodium bisulfite reaction according to standard procedures (the bisulfite process converts unmethylated cytosine residues to uracil). Fluorescence-based PCR is then performed in a "biased" reaction, e.g., with PCR primers that overlap known CpG dinucleotides. Sequence discrimination occurs both at the level of the amplification process and at the level of the fluorescence detection process.
[0166] The MethyLight™ assay is used as a quantitative test for methylation patterns in nucleic acids, e.g., genomic DNA samples, in which sequence discrimination occurs at the level of probe hybridization. In the quantitative version, PCR reactions result in methylation-specific amplification in the presence of fluorescent probes that overlap specific putative methylation sites. An unbiased control for input DNA amount is provided by reactions in which neither the primers nor the probe overlap any CpG dinucleotides. Alternatively, quantitative tests for genomic methylation are achieved by probing biased PCR pools with either control oligonucleotides that do not cover known methylation sites (e.g., fluorescent-based versions of the HeavyMethyl™ and MSP techniques) or oligonucleotides that cover potential methylation sites.
[0167] The MethyLight™ process can be used with any suitable probe (e.g., TaqMan® probe, Lightcycler® probe). For example, in some applications, double-stranded genomic DNA is treated with sodium bisulfite and subjected to one of two sets of PCR reactions, using, for example, a TaqMan® probe with MSP primers and / or a HeavyMethyl inhibitor oligonucleotide and a TaqMan® probe. The TaqMan® probe is dual-labeled with fluorescent "reporter" and "quencher" molecules and is designed to be specific for relatively GC-rich regions, so that it melts during PCR cycles at a temperature approximately 10°C higher than the forward or reverse primers. This allows the TaqMan® probe to remain fully hybridized during the PCR annealing / extension step. Taq polymerase enzymatically synthesizes new strands during PCR, ultimately reaching the annealed TaqMan® probe. Taq polymerase 5' to 3' endonuclease activity then displaces the TaqMan® probe by digesting it, releasing a fluorescent reporter molecule for quantitative detection of its unquenched signal using a real-time fluorescence detection system.
[0168] Typical reagents for MethyLight™ analysis (e.g., as found in a typical MethyLight™-based kit) include, but are not limited to, PCR primers for specific loci (e.g., specific genes, markers, gene regions, marker regions, bisulfite-treated DNA sequences, CpG islands, etc.); TaqMan® or Lightcycler® probes; optimized PCR buffers and deoxynucleotides, and Taq polymerase.
[0169] The QM™ (Quantitative Methylation) Assay is an alternative quantitative test for methylation patterns in genomic DNA samples, in which sequence discrimination occurs at the level of probe hybridization. In this quantitative version, PCR reactions result in unbiased amplification in the presence of fluorescent probes that overlap specific putative methylation sites. An unbiased control for input DNA amount is provided by reactions in which neither the primers nor the probe overlap any CpG dinucleotides. Alternatively, quantitative testing for genomic methylation is achieved by probing biased PCR pools with either control oligonucleotides that do not cover known methylation sites (fluorescence-based versions of HeavyMethyl™ and MSP techniques) or oligonucleotides that cover potential methylation sites.
[0170] The QM™ process can be used with any suitable probe, such as a TaqMan® probe or a Lightcycler® probe, during the amplification process. For example, double-stranded genomic DNA is treated with sodium bisulfite and subjected to unbiased primers and a TaqMan® probe. The TaqMan® probe is dual-labeled with fluorescent reporter and quencher molecules and is designed to be specific for relatively GC-rich regions, so it melts during PCR cycles at a temperature approximately 10°C higher than the forward or reverse primers. This allows the TaqMan® probe to remain fully hybridized during the PCR annealing / extension step. Taq polymerase enzymatically synthesizes new strands during PCR, ultimately reaching the annealed TaqMan® probe. The 5' to 3' endonuclease activity of Taq polymerase then displaces the TaqMan® probe by digesting it, releasing a fluorescent reporter molecule for quantitative detection of its unquenched signal using a real-time fluorescence detection system. Typical reagents for QM™ analysis (e.g., as found in a typical QM™-based kit) include, but are not limited to, PCR primers for specific loci (e.g., specific genes, markers, gene regions, marker regions, bisulfite-treated DNA sequences, CpG islands, etc.); TaqMan® or Lightcycler® probes; optimized PCR buffers and deoxynucleotides; and Taq polymerase.
[0171] The SNuPE™ technique is a quantitative method that involves single-base primer extension, based on bisulfite treatment of DNA to assess differences in methylation at specific CpG sites (Gonzalgo & Jones, Nucleic Acids Res. 25:2529-2531, 1997). Briefly, genomic DNA is reacted with sodium bisulfite to convert unmethylated cytosines to uracil, while leaving 5-methylcytosines unchanged. Next, PCR primers specific to the bisulfite-converted DNA are used to amplify the desired target sequence, and the resulting product is isolated and used as a template for methylation analysis at the CpG sites of interest. This allows for the analysis of small amounts of DNA (e.g., microdissected pathology slices), avoiding the use of restriction enzymes to determine the methylation status at CpG sites.
[0172] Typical reagents for Ms-SNuPE™ analysis (e.g., as found in a typical Ms-SNuPE™-based kit) include, but are not limited to, PCR primers for specific loci (e.g., specific genes, markers, gene regions, marker regions, bisulfite-treated DNA sequences, CpG islands, etc.); optimized PCR buffers and deoxynucleotides; gel extraction kits; positive control primers; Ms-SNuPE™ primers for specific loci; reaction buffers (for the Ms-SNuPE reaction); and labeled nucleotides. Additionally, bisulfite conversion reagents can include DNA denaturing buffers; sulfonation buffers; DNA recovery reagents or kits (e.g., precipitation, ultrafiltration, affinity columns); desulfonation buffers; and DNA recovery components.
[0173] Reduced representation bisulfite sequencing (RRBS) begins with bisulfite treatment of nucleic acids to convert all unmethylated cytosines to uracil, followed by restriction enzyme digestion (e.g., with an enzyme that recognizes sites containing CG sequences, such as Msp1) and completes fragment sequencing after coupling to an adaptor ligand. The choice of restriction enzyme enriches for fragments in CpG-dense regions, reducing the number of redundant sequences that may map to multiple gene locations during analysis. As such, RRBS reduces the complexity of a nucleic acid sample by selecting a subset of restriction fragments for sequencing (e.g., by size selection using preparative gel electrophoresis). In contrast to whole-genome bisulfite sequencing, all fragments generated by restriction enzyme digestion contain DNA methylation information for at least one CpG dinucleotide. As such, RRBS enriches samples for promoters, CpG islands, and other genomic features with high frequency of restriction enzyme cleavage sites in these regions, providing an assay to assess the methylation status of one or more genomic loci.
[0174] A typical protocol for RRBS includes the steps of digesting a nucleic acid sample with a restriction enzyme such as MspI, filling in overhangs and A-tailing, adapter ligation, bisulfite conversion, and PCR. See, e.g., Meissner et al. (2005) "Genome-scale DNA methylation mapping of clinical samples at single-nucleotide resolution" Nat Methods 7:133-6; Meissner et al. (2005) "Reduced representation bisulfite sequencing for comparative high-resolution DNA methylation analysis" Nucleic Acids Res. 33:5868-77.
[0175] In some embodiments, quantitative allele-specific real-time target and signal amplification (QuARTS) assays are used to assess methylation status. Each QuARTS assay involves three sequential reactions: a primary reaction involving amplification (reaction 1) and target probe cleavage (reaction 2), and a secondary reaction involving FRET cleavage and fluorescent signal generation (reaction 3). When a target nucleic acid is amplified with specific primers, a specific detection probe with a flap sequence loosely binds to the amplicon. The presence of a specific invasive oligonucleotide at the target binding site allows a 5' nuclease, such as FEN-1 endonuclease, to cleave the gap between the detection probe and the flap sequence, thereby releasing the flap sequence. The flap sequence is complementary to the non-hairpin portion of the corresponding FRET cassette. Thus, the flap sequence functions as an invasive oligonucleotide on the FRET cassette, resulting in cleavage between the FRET cassette fluorophore and quencher, generating a fluorescent signal. The cleavage reaction can cleave multiple probes per target, thereby releasing multiple fluorophores per flap, resulting in exponential signal amplification. QuARTS can detect multiple targets in a single reaction well by using FRET cassettes with different dyes (see, e.g., Zou et al. (2010) "Sensitive quantification of methylated markers with a novel methylation-specific technology" Clin Chem 56:A199), and U.S. Patent Nos. 8,361,720, 8,715,937, 8,916,344, and 9,212,392, each of which is incorporated by reference herein for all purposes.
[0176] The term "bisulfite reagent" refers to a reagent containing bisulfite, disulfite, hydrogen sulfite, or a combination thereof, useful for distinguishing between methylated and unmethylated CpG dinucleotide sequences, as disclosed herein. Methods for such treatment are known in the art (e.g., PCT / EP2004 / 011715 and WO2013 / 116375, each of which is incorporated by reference in its entirety). In some embodiments, the bisulfite treatment is carried out in the presence of a denaturing solvent, such as, but not limited to, n-alkylene glycol or diethylene glycol dimethyl ether (DME), or in the presence of dioxane or a dioxane derivative. In some embodiments, the denaturing solvent is used at a concentration of 1% to 35% (v / v). In some embodiments, the bisulfite reaction is carried out in the presence of a scavenger, such as, but not limited to, a chroman derivative, e.g., 6-hydroxy-2,5,7,8-tetramethylchroman-2-carboxylic acid or trihydroxybenzoic acid, and derivatives thereof, e.g., gallic acid (see PCT / EP2004 / 011715, incorporated herein by reference in its entirety). In certain preferred embodiments, the bisulfite reaction involves treatment with ammonium bisulfite, e.g., as described in WO2013 / 116375.
[0177] In some embodiments, fragments of the treated DNA are amplified using a set of primer oligonucleotides and an amplification enzyme according to the methods and compositions described herein. Amplification of several DNA segments can be performed simultaneously in one reaction vessel, and in the same reaction vessel. Typically, amplification is performed using the polymerase chain reaction (PCR). Amplicons are typically 100-2000 base pairs in length.
[0178] In some embodiments of the method, the methylation status or profile of CpG positions within or near differentially methylated regions (e.g., Tables 1 and 2) may be detected using methylation-specific primer oligonucleotides. This technique (MSP) is described in U.S. Patent No. 6,265,171 to Herman. The use of methylation-status-specific primers for amplification of bisulfite-treated DNA allows differentiation between methylated and unmethylated nucleic acids. An MSP primer pair contains at least one primer that hybridizes to a bisulfite-treated CpG dinucleotide. Thus, the primer sequence contains at least one CpG dinucleotide. MSP primers specific for unmethylated DNA contain a "T" at the C position in the CpG.
[0179] Such methods are not limited to a particular type or kind of primer or primer pair associated with one or more methylation markers, methylation marker genes, genes, DMRs, and / or methylated DNA markers. In some embodiments, a primer or primer pair specific for each methylation marker gene can bind to an amplicon bounded by the primer sequences of a marker gene listed in Table 1 or 2, and the amplicon bounded by the primer sequences of a marker gene is at least a portion of a gene region of a methylation marker gene listed in Table 1 or 2.
[0180] In another embodiment, the present disclosure provides a method for converting oxidized 5-methylcytosine residues in cell-free DNA to dihydrouracil residues (see Liu et al., 2019, Nat Biotechnol. 37, pp. 424-429; U.S. Patent Application Publication No. 202000370114). The method comprises reacting an oxidized 5mC residue selected from 5-formylcytosine (5fC), 5-carboxymethylcytosine (5caC), and combinations thereof with a borane reducing agent. The oxidized 5mC residue may be naturally occurring or, more typically, may result from prior oxidation of a 5mC or 5hmC residue, e.g., oxidation of 5mC or 5hmC by a TET family enzyme (e.g., TET1, TET2, or TET3), or chemical oxidation of 5mC or 5hmC (see, e.g., Okamato et al. (2011) Chem. Commun. 47:11231-33), e.g., using an inorganic peroxo compound or composition such as potassium perruthenate (KRuO4) or peroxotungstate and a combination of copper(II) perchlorate / 2,2,6,6-tetramethylpiperidine-1-oxyl (TEMPO) (see, e.g., Matsushita et al. (2017) Chem. Commun. 53:5756-59).
[0181] The borane reducing agent can be characterized as a complex of borane and a nitrogen-containing compound selected from a nitrogen heterocycle and a tertiary amine. The nitrogen heterocycle can be monocyclic, bicyclic, or polycyclic, but is typically a monocyclic ring in the form of a five- or six-membered ring containing a nitrogen heteroatom and, optionally, one or more additional heteroatoms selected from N, O, and S. The nitrogen heterocycle can be aromatic or alicyclic. Preferred nitrogen heterocycles herein include 2-pyrroline, 2H-pyrrole, 1H-pyrrole, pyrazolidine, imidazolidine, 2-pyrazoline, 2-imidazoline, pyrazole, imidazole, 1,2,4-triazole, 1,2,4-triazole, pyridazine, pyrimidine, pyrazine, 1,2,4-triazine, and 1,3,5-triazine, any of which may be unsubstituted or substituted with one or more non-hydrogen substituents. Typical non-hydrogen substituents are alkyl groups, particularly lower alkyl groups such as methyl, ethyl, n-propyl, isopropyl, n-butyl, isobutyl, t-butyl and the like. Exemplary compounds include, but are not limited to, borane, pyridine borane, 2-methylpyridine borane (also known as 2-picoline borane or pic-BH), 5-ethyl-2-pyridine, sodium borohydride, sodium cyanoborohydride, sodium triacetoxyborohydride, diborane, decaborane, borane-tetrahydrofuran, borane-dimethylsulfide, borane-N,N-diisopropylethylamine, borane-2-chloropyridine, borane-aniline, N,N-dimethylamine borane, tert-butylamine borane, sodium triacetoxyborohydride, borane hydride, hydrazine or dibutylamine borane, morpholine borane, borane-ammonia complex (BHNH), dicyclohexylamine borane, morpholine borane, 4-methylmorpholine borane, alkali and tetramethylamine borane (e.g., NaBH) and other -BH-containing complexes and / or derivatives. In some embodiments, the reducing agent is pyridine borane and / or pic-BH3.
[0182] The reaction of oxidized 5mC residues in cell-free DNA with borane reducing agents is advantageous insofar as it utilizes nontoxic reagents and mild reaction conditions, eliminating the need for hydrogen sulfate or other reagents that may degrade DNA. Furthermore, the conversion of oxidized 5mC residues to dihydrolauracils with borane reducing agents can be carried out in a "one-pot" or "one-tube" reaction without the need for isolation of intermediates. This is crucial because this conversion involves multiple steps: (1) reduction of the alkene bond connecting C-4 and C-5 of oxidized 5mC, (2) deamination, and (3) either decarboxylation if the oxidized 5mC is 5caC or deformylation if the oxidized 5mC is 5fC.
[0183] In addition to a method for converting oxidized 5-methylcytosine residues in cell-free DNA to dihydrouracil residues, the present disclosure also provides a reaction mixture related to the aforementioned method. The reaction mixture includes a sample of cell-free DNA containing at least one oxidized 5-methylcytosine residue selected from 5caC, 5fC, and combinations thereof, and a borane reducing agent effective in reducing, deaminating, and decarboxylating or deformylating the at least one oxidized 5-methylcytosine residue. As explained above, the borane reducing agent is a complex of borane and a nitrogen-containing compound selected from nitrogen heterocycles and tertiary amines. In a preferred embodiment, the reaction mixture is substantially bisulfite-free, meaning that it is substantially free of bisulfite ions and bisulfite salts. Ideally, the reaction mixture is free of bisulfite salts.
[0184] In a related aspect of the present disclosure, a kit for converting 5mC residues in cell-free DNA to dihydrouracil residues is provided, the kit including a reagent for blocking 5mC residues, a reagent for oxidizing 5mC residues beyond hydroxymethylation to provide oxidized 5mC residues, and a borane reducing agent effective for reducing, deaminating, and decarboxylating or deformylating the oxidized 5mC residues. The kit may also include instructions for using the components to practice the method.
[0185] In another embodiment, a method utilizing the above-described oxidation reaction is provided, which can detect the presence and location of 5-methylcytosine residues in cell-free DNA, and includes the following steps: (a) modifying 5hmC residues in fragmented, adaptor-ligated cell-free DNA to provide affinity tags thereon (the affinity tags enable removal of the modified 5hmC-containing DNA from the cell-free DNA), (b) removing the modified 5hmC-containing DNA from the cell-free DNA to leave DNA containing unmodified 5mC residues, and (c) oxidizing the unmodified 5mC residues to form 5caC, (d) contacting the DNA containing the oxidized 5mC residues with a borane reducing agent effective to reduce, deaminate, and decarboxylate or deformylate the oxidized 5mC residues, thereby providing DNA containing dihydrouracil residues in place of the oxidized 5mC residues; (e) amplifying and sequencing the DNA containing the dihydrouracil residues; and (f) determining the 5-methylation pattern from the sequencing results of (e).
[0186] In some embodiments, the present disclosure provides a method for identifying 5-methylcytosine (5mC) or 5-hydroxymethylcytosine (5hmC) in a target nucleic acid. In some embodiments, the method includes providing a biological sample containing the target nucleic acid; modifying the target nucleic acid by converting 5mC and 5hmC in the nucleic acid sample to 5-carboxylcytosine (5caC) and / or 5-formylcytosine (5fC) to generate one or more 5caC or 5fC residues by contacting the nucleic acid sample with a TET enzyme; treating the target nucleic acid with a borane reducing agent to convert the 5caC and / or 5fC to dihydrouracil (DHU) to provide a modified nucleic acid sample containing the modified target nucleic acid; and detecting the sequence of the modified target nucleic acid, wherein a cytosine (C) to thymine (T) transition or a cytosine (C) to DHU transition in the sequence of the modified target nucleic acid indicates the location of either a 5mC or a 5hmC in the target nucleic acid, relative to the target nucleic acid. In some embodiments, the borane reducing agent is 2-picoline borane.
[0187] In some embodiments, detecting the sequence of the modified target nucleic acid comprises one or more of chain termination sequencing, microarray, high-throughput sequencing, and restriction enzyme analysis. In some embodiments, the TET enzyme is selected from the group consisting of human TET1, TET2, and TET3; mouse Tet1, Tet2, and Tet3; Naegleria TET (NgTET); and Coprinopsis cinerea (CcTET). In some embodiments, the method further comprises blocking one or more modified cytosines. In some embodiments, the blocking comprises adding a sugar to 5hmC. In some embodiments, the method further comprises amplifying the copy number of one or more nucleic acid sequences. In some embodiments, the oxidizing agent is potassium perruthenate or Cu(II) / TEMPO (2,2,6,6-tetramethylpiperidine-1-oxyl).
[0188] Cell-free DNA is typically extracted from a subject's biological sample, which may be whole blood, plasma, urine, saliva, mucosal discharge, organ secretions, sputum, stool, or tears. In some embodiments, the cell-free DNA is derived from a tumor (e.g., a gynecological tumor). In other embodiments, the cell-free DNA is derived from a patient with a disease or other pathological condition. The cell-free DNA may or may not be derived from a tumor. In some embodiments, the cell-free DNA modified at 5hmC residues is purified, fragmented, and adapter-ligated. DNA purification in this regard can be performed using any suitable method known to those skilled in the art and / or described in the relevant literature. The cell-free DNA itself may be highly fragmented, although further fragmentation may be desirable in some cases, as described, for example, in U.S. Patent Publication No. 2017 / 0253924. Cell-free DNA fragments generally range in size from about 20 nucleotides to about 500 nucleotides, more typically from about 20 nucleotides to about 250 nucleotides. The purified cell-free DNA fragments modified in step (a) are end-repaired using conventional means (e.g., restriction enzymes) to ensure that the fragments have blunt ends at each 3' and 5' end. In a preferred method, as described in WO 2017 / 176630, the blunted fragments are also provided with 3' overhangs containing a single adenine residue using a polymerase such as Taq polymerase. This facilitates subsequent ligation of a selected universal adapter, i.e., a Y adapter or hairpin adapter, which ligates to both ends of the cell-free DNA fragments and contains at least one molecular barcode. The use of adapters also allows for selective PCR enrichment of adapter-ligated DNA fragments.
[0189] In some embodiments, "purified fragmented cell-free DNA" includes adaptor-ligated DNA fragments. The 5hmC residues of these cell-free DNA fragments are modified with an affinity tag to enable subsequent removal of the modified 5hmC-containing DNA from the cell-free DNA. In one embodiment, the affinity tag comprises a biotin moiety, such as biotin, desthiobiotin, oxybiotin, 2-iminobiotin, diaminobiotin, biotin sulfoxide, or biocytin. The use of a biotin moiety as an affinity tag facilitates removal with streptavidin, e.g., streptavidin beads, magnetic streptavidin beads, or the like.
[0190] Tagging of 5hmC residues with biotin moieties or other affinity tags can be achieved by covalently attaching a chemoselective group to the 5hmC residues in the DNA fragment, which can react with a functionalized affinity tag to attach the affinity tag to the 5hmC residue. In one embodiment, the chemoselective group is UDP-glucose-6-azide, which undergoes spontaneous 1,3-cycloaddition with an alkyne-functionalized biotin moiety, as described in Robertson et al. (2011) Biochem. Biophys. Res. Comm. 411(1):40-3, U.S. Patent No. 8,741,567, and WO2017 / 176630. Thus, the addition of the alkyne-functionalized biotin moiety results in the covalent attachment of the biotin moiety to each 5hmC residue.
[0191] The affinity-tagged DNA fragments can then be removed, in one embodiment, using streptavidin in the form of streptavidin beads, magnetic streptavidin beads, etc., and saved for later analysis if desired. The supernatant remaining after removal of the affinity-tagged fragments contains DNA with unmodified 5mC residues but no 5hmC residues.
[0192] In some embodiments, unmodified 5mC residues are oxidized to yield 5caC and / or 5fC residues using any suitable means. The oxidizing agent is selected to oxidize the 5mC residue beyond hydroxymethylation, i.e., to yield 5caC and / or 5fC residues. Oxidation can be performed enzymatically using a catalytically active TET family enzyme. The term "TET family enzyme" or "TET enzyme" as used herein refers to a catalytically active "TET family protein" or "TET catalytically active fragment," as defined in U.S. Pat. No. 9,115,386, the disclosure of which is incorporated herein by reference. A preferred TET enzyme in this regard is TET2 (see Ito et al. (2011) Science 333(6047):1300-1303). Oxidation can also be performed chemically using a chemical oxidizing agent, as described in the previous section. Examples of suitable oxidizing agents include, but are not limited to, perruthenate anions in the form of inorganic or organic perruthenates, including metal perruthenates such as potassium perruthenate (KRuO), tetraalkylammonium perruthenates such as tetrapropylammonium perruthenate (TPAP) and tetrabutylammonium perruthenate (TBAP), and polymer-supported perruthenate (PSP); and inorganic peroxo compounds and compositions, such as peroxotungstate or a combination of copper(II) perchlorate / TEMPO. It is not necessary to separate the 5fC-containing fragments from the 5caC-containing fragments at this point, as long as both the 5fC and 5caC residues are converted to dihydrouracil (DHU) in the next step of the process.
[0193] In some embodiments, 5-hydroxymethylcytosine residues are blocked with β-glucosyltransferase (β3GT), while 5-methylcytosine residues are oxidized with a TET enzyme, which is effective in generating a mixture of 5-formylcytosine and 5-carboxymethylcytosine. A mixture containing both of these oxidized species can be reacted with 2-picoline borane or another borane reducing agent to yield dihydrouracil. In a variation of this embodiment, 5hmC-containing fragments are not removed. Instead, in "TET-assisted picoline borane sequencing (TAPS)," 5mC- and 5hmC-containing fragments are enzymatically oxidized together to yield 5fC- and 5caC-containing fragments. Reaction with 2-picoline borane generates DHU residues where the 5mC and 5hmC residues originally resided. In "chemically assisted picoline borane sequencing (CAPS)," 5hmC-containing fragments are selectively oxidized with potassium perruthenate, leaving the 5mC residues unchanged.
[0194] As disclosed in International PCT Application PCT / US2019 / 012627 (incorporated herein by reference in its entirety), TAPS involves the use of mild enzymatic and chemical reactions to directly and quantitatively detect 5mC and 5hmC at base resolution without affecting unmodified cytosines. In a related embodiment, the above method further includes identifying the hydroxymethylation pattern of 5hmC-containing DNA removed from cell-free DNA. This can be done using the techniques described in detail in WO 2017 / 176630. This process can be performed in a one-tube manner without intermediate removal or isolation. For example, cell-free DNA fragments, preferably adapter-ligated DNA fragments, are first functionalized with βGT-catalyzed uridine diphosphoglucose 6-azide and then biotinylated with a chemoselective azide group. This procedure covalently attaches biotin to each 5hmC site. In a next step, the biotinylated strand and the strand containing unmodified (native) 5mC are simultaneously removed for further processing. Native 5mC-containing chains are removed using anti-5mC antibodies or methyl-CpG binding domain (MBD) proteins, as known to those skilled in the art. Then, with the 5hmC residues blocked, unmodified 5mC residues are selectively oxidized using any suitable technique for converting 5mC to 5fC and / or 5caC, as described elsewhere herein.
[0195] These fragments obtained by amplification can have directly or indirectly detectable labels.In some embodiments, these labels are fluorescent labels, radionuclides, or detachable molecular fragments, which have a typical mass that can be detected in mass spectrometer.When these labels are mass labels, some embodiments provide that the labeled amplicons have a single positive or negative effective charge, which allows for better detectability in mass spectrometer.For example, this detection can be performed and visualized by matrix-assisted laser desorption / ionization mass spectrometry (MALDI) or by electron spray mass spectrometry (ESI).
[0196] Methods for isolating DNA suitable for these assay techniques are well known in the art. In particular, some embodiments involve isolating nucleic acids as described in U.S. Patent No. 13 / 470,251 ("Isolation of Nucleic Acids"), which is incorporated herein by reference in its entirety.
[0197] In some embodiments, the markers described herein are used in a QUARTS assay performed on a stool sample. In some embodiments, methods are provided for generating DNA samples, particularly DNA samples containing highly purified, low-abundance nucleic acids in small volumes (e.g., less than 100 microliters, less than 60 microliters) that are substantially and / or effectively free of substances that inhibit assays used to test the DNA sample (e.g., PCR, INVADER, QuARTS assays, etc.). Such DNA samples are utilized in diagnostic tests that qualitatively detect the presence or quantitatively measure the activity, expression, or amount of genes, genetic variants (e.g., alleles), or genetic modifications (e.g., methylation) present in a sample collected from a patient. For example, some cancers are correlated with the presence of specific mutant alleles or specific methylation states; therefore, detection and / or quantification of such mutant alleles or methylation states has predictive value in cancer diagnosis and treatment.
[0198] Many useful genetic markers are present in very small amounts in samples, and many of the events that produce such markers are rare. Therefore, even highly sensitive detection methods such as PCR require a large amount of DNA to provide targets with low abundances sufficient to meet or exceed the detection threshold of the assay. Furthermore, the presence of even a small amount of inhibitors can impair the accuracy and precision of these assays aimed at detecting such low-abundance targets. Therefore, the present specification provides a method for producing such DNA samples, providing the necessary volume and concentration control.
[0199] In some embodiments, the sample comprises stool, a tissue sample, an organ secretion, CSF, saliva, blood, or urine. In some embodiments, the subject is a human. Such samples can be obtained by any number of means known in the art, as will be apparent to those skilled in the art. Cell-free or substantially cell-free samples can be obtained by subjecting the sample to various techniques known to those skilled in the art, including, but not limited to, centrifugation and filtration. While obtaining samples without invasive techniques is generally preferred, it may still be preferable to obtain samples such as tissue homogenates, tissue sections, and biopsy specimens. The present technology is not limited by the method used to prepare the sample and provide nucleic acids for testing. For example, in some embodiments, DNA is isolated from a sample (e.g., a stool sample, a tissue sample, an organ secretion sample, a CSF sample, a saliva sample, a blood sample, a plasma sample, or a urine sample) using direct gene capture or related methods, as described in detail in, for example, US Pat. Nos. 8,808,990 and 9,169,511 and WO 2012 / 155072.
[0200] Marker analysis can be performed separately or simultaneously with additional markers within a single test sample. For example, it is possible to combine several markers in one test to efficiently process multiple samples and potentially provide higher diagnostic and / or prognostic accuracy. Furthermore, those skilled in the art will recognize the value of testing multiple samples from the same subject (e.g., at successive time points). Testing such serial samples allows for the identification of changes in the methylation status of markers over time. Changes in methylation status, and the absence of changes in methylation status, can provide useful information about disease states, including, but not limited to, identifying the subject's outcome, including the approximate time since the occurrence of this event, the presence and amount of recoverable tissue, the suitability of drug therapy, the effectiveness of various therapies, and the risk of future events.
[0201] Biomarker analysis can be performed in a variety of physical formats. For example, microtiter plates or automated applications can be used to facilitate the processing of large numbers of test samples. Alternatively, single sample formats can be developed to facilitate immediate treatment and diagnosis in a timely manner, for example, in an outpatient or emergency room setting.
[0202] Genomic DNA can be isolated by any means, including the use of commercially available kits. Briefly, if the DNA of interest is encapsulated in a cell membrane, the biological sample must be disrupted and dissolved by enzymatic, chemical, or mechanical means. Proteins and other contaminants can then be removed from the DNA solution, for example, by digestion with proteinase K. The genomic DNA is then recovered from the solution. This can be done by a variety of methods, such as salting out, organic extraction, or binding of DNA to a solid support. The choice of method is influenced by several factors, including time, cost, and the amount of DNA required. All clinical sample types containing neoplastic or pre-neoplastic material are suitable for use in this method, including cell lines, histological slides, biopsies, paraffin-embedded tissues, body fluids, feces, tissues, colonic effluent, urine, plasma, serum, whole blood, isolated blood cells, cells isolated from blood, and combinations thereof.
[0203] The present technology is not limited to the method used to prepare the sample and provide the nucleic acid for testing.For example, in some embodiments, DNA is isolated from a fecal sample, or from a blood sample, or from a plasma sample, using a direct gene capture method, such as that described in U.S. Patent Application No. 61 / 485386, or by related methods.
[0204] The genomic DNA sample is then treated with at least one reagent, or a series of reagents, that distinguishes between methylated and unmethylated CpG dinucleotides within at least one marker that comprises a DMR (e.g., a DMR (Table 1 or 2)).
[0205] In some embodiments, the reagent converts unmethylated cytosine bases at the 5' position to uracil, thymine, or another base that differs from cytosine in terms of hybridization behavior, although in some embodiments the reagent may be a methylation-sensitive restriction enzyme.
[0206] In some embodiments, genomic DNA samples are treated in such a way that cytosine bases that are not methylated at the 5' position are converted to uracil, thymine, or another base that differs from cytosine in terms of hybridization behavior. In some embodiments, this treatment is carried out with bisulfite (hydrogen sulfite), followed by alkaline hydrolysis.
[0207] The processed nucleic acid is then analyzed to determine the methylation status of the target gene sequence (at least one gene, genomic sequence, or nucleotide from a marker comprising a DMR, e.g., at least one DMR selected from the DMRs in Tables 1 or 2). Analytical methods can be selected from those known in the art, including those listed herein (e.g., QuARTS and MSP as described herein).
[0208] Such samples can be obtained by any number of means known in the art, as will be apparent to those skilled in the art. For example, urine and fecal samples are readily available, while blood, ascites, serum, or pancreatic juice samples can be obtained parenterally, for example, by using a needle and syringe. Cell-free, or substantially cell-free, samples can be obtained by subjecting the sample to a variety of techniques, including, but not limited to, centrifugation and filtration. While it is generally preferred to obtain samples without the use of invasive techniques, it may still be preferable to obtain samples such as tissue homogenates, tissue sections, and biopsy specimens.
[0209] Embodiments of the present disclosure further provide compositions. In some embodiments, the present disclosure provides compositions comprising a nucleic acid comprising a DMR and a bisulfite reagent. In some embodiments, compositions are provided comprising a nucleic acid comprising a DMR and one or more primers (e.g., a primer capable of binding to at least a portion of a region of a DMR listed in Table 1 or 2, or a primer capable of binding to an amplicon bound by a primer capable of binding to at least a portion of a region of a DMR listed in Table 1 or 2). In certain embodiments, compositions are provided comprising a nucleic acid comprising a DMR and a methylation-sensitive restriction enzyme. In certain embodiments, compositions are provided comprising a nucleic acid comprising a DMR and a polymerase.
[0210] 3. Treatment method In some embodiments, the present disclosure provides methods for treating a subject (e.g., a patient having or suspected of having one or more types or subtypes of gynecological cancer). According to these embodiments, the methods include determining the methylation status or profile of one or more methylated DNA markers provided herein and / or measuring the expression and / or activity levels of one or more protein markers, and administering a treatment to the patient based on the results of determining the methylation status and / or the expression and / or activity levels of the protein markers. The treatment can be administering a pharmaceutical compound, administering a vaccine, performing surgery, imaging the patient, or performing another test. In some embodiments, treating a subject includes methods of clinical screening, prognostic evaluation, monitoring treatment results, identifying patients most likely to respond to a particular therapeutic treatment, imaging patients or subjects, and methods for drug screening and development.
[0211] In some embodiments, a method for diagnosing a particular type of cancer in a subject is provided. As used herein, the terms "diagnosing" and "diagnosis" refer to a method by which a skilled artisan can estimate, and even determine, whether a subject is suffering from a given disease or condition, or whether a subject is likely to develop a given disease or condition in the future. Those skilled in the art often make a diagnosis based on one or more diagnostic indicators, such as one or more biomarkers (e.g., one or more methylation markers, methylation marker genes, genes, DMRs, and / or DNA methylation markers disclosed herein), the methylation status of which indicates the presence, severity, or absence of the condition and / or the expression and / or activity level of one or more protein markers.
[0212] Along with diagnosis, clinical prognosis of cancer involves determining the aggressiveness of cancer and the likelihood of tumor recurrence in order to plan the most effective treatment.If a more accurate prognosis can be made or even the potential risk of developing cancer can be assessed, appropriate therapy, and in some cases, a less harsh therapy for the patient, can be selected.Evaluating cancer biomarkers (e.g., determining methylation status) is useful for separating subjects with a good prognosis and / or a low risk of developing cancer, who do not require treatment or only limited treatment, from subjects who are more likely to develop cancer or suffer from cancer recurrence, who may benefit from more intensive treatment.
[0213] Thus, "making a diagnosis" or "diagnosing," as used herein, further includes determining the risk of developing cancer or determining a prognosis, which can be provided to predict a clinical outcome (with or without medical treatment), select an appropriate treatment (or whether a treatment is effective), or monitor a current treatment to potentially modify the treatment, based on measurements of a diagnostic biomarker (e.g., DMR) disclosed herein. Furthermore, in some embodiments of the presently disclosed subject matter, multiple determinations of biomarkers over time can be made to facilitate diagnosis and / or prognosis. Changes in biomarkers over time can be used to predict clinical outcomes, monitor the progression of cancer or cancer subtypes, and / or monitor the effectiveness of appropriate cancer-directed treatments. In such embodiments, it may be expected to ascertain, for example, changes in the methylation status and / or expression and / or activity levels of protein markers of one or more biomarkers (e.g., DMRs) disclosed herein (and potentially one or more additional biomarker(s), if monitored) in biological samples over time during the course of an effective therapy.
[0214] The presently disclosed subject matter further provides, in some embodiments, a method for determining whether to initiate or continue cancer prevention or treatment in a subject. In some embodiments, such a method includes providing a series of biological samples from a subject over a period of time; analyzing the series of biological samples to determine the methylation status or profile of at least one marker disclosed herein in each of the biological samples; and comparing measurable changes in the methylation status of the one or more biomarkers in each of the biological samples. Any changes over a period of time can be used to predict the risk of developing cancer, predict clinical outcome, determine whether to initiate or continue cancer prevention or treatment, or determine whether a current treatment is effectively treating the cancer. For example, a first time point can be selected before the start of treatment, and a second time point can be selected at a time point after the start of treatment. Methylation status can be measured in each sample taken from different time points, and qualitative and / or quantitative differences can be observed. Changes in the methylation status of biomarker levels and / or protein marker expression / activity levels from different samples can be correlated with a particular cancer risk, prognosis, treatment efficacy determination, and / or progression of cancer in the subject. In some embodiments, the disclosed methods and compositions are for the treatment or diagnosis of disease at an early stage, e.g., before disease symptoms appear, hi some embodiments, the disclosed methods and compositions are for the treatment or diagnosis of disease at a clinical stage.
[0215] In some embodiments, multiple determinations of one or more diagnostic or prognostic biomarkers can be performed, and changes in the markers over time can be used to determine a diagnosis or prognosis. For example, a diagnostic marker can be determined a first time and then again a second time. In such embodiments, an increase in a marker from the first time to the second time can be diagnostic of a particular type or severity of cancer, or a given prognosis. Similarly, a decrease in a marker from the first time to the second time can indicate a particular type or severity of cancer, or a given prognosis. Furthermore, the degree of change in one or more markers can be related to the severity of cancer and future adverse events. Those skilled in the art will understand that, in certain embodiments, comparative measurements of the same biomarker can be performed at multiple time points, but a given biomarker can also be measured at one time point and a second biomarker at a second time point, and the comparison of these markers can provide diagnostic information.
[0216] As used herein, the phrase "determining a prognosis" refers to a method by which a person skilled in the art can predict the course or outcome of a condition in a subject. The term "prognosis" does not refer to the ability to predict the course or outcome of a condition with 100% accuracy, or even the ability to predict that a given course or outcome is more or less likely to occur based on the methylation status of a biomarker (e.g., a DMR). Instead, those skilled in the art will understand that the term "prognosis" refers to a high probability that a particular course or outcome will occur, i.e., a high probability that a certain course or outcome will occur in a subject exhibiting a given condition, compared to an individual not exhibiting the condition. For example, an individual who does not exhibit symptoms (e.g., having a normal methylation status of one or more DMRs and / or expression and / or activity levels of protein markers) may have a very low probability of a given outcome (e.g., suffering from a particular type of cancer).
[0217] In some embodiments, statistical analysis correlates the prognostic indicator with a predisposition to adverse outcomes. For example, in some embodiments, a difference in methylation status and / or protein marker expression / activity level from that in a normal control sample obtained from a patient without cancer, as determined by the level of statistical significance, may signal that a subject is more likely to have cancer than a subject having a level more similar to the methylation status in the control sample. Furthermore, changes in methylation status and / or protein marker expression / activity level from baseline (e.g., "normal") levels may reflect the subject's prognosis, and the degree of change in methylation status and / or protein marker expression / activity level may be related to the severity of an adverse event. Statistical significance is often determined by comparing two or more populations and determining a confidence interval and / or p-value. See, e.g., Dowdy and Wearden, *Statistics for Research*, John Wiley & Sons, New York, 1983, incorporated herein by reference in its entirety. Exemplary confidence intervals of the present subject matter are 90%, 95%, 97.5%, 98%, 99%, 99.5%, 99.9%, and 99.99%, and exemplary p-values are 0.1, 0.05, 0.025, 0.02, 0.01, 0.005, 0.001, and 0.0001.
[0218] In other embodiments, a threshold change in methylation status and / or protein marker expression / activity level of a prognostic or diagnostic biomarker (e.g., DMR; protein marker) disclosed herein can be established, and the change in methylation status and / or protein marker expression / activity level of the biomarker in a biological sample is simply compared to the threshold change in methylation status and / or protein marker expression / activity level. Preferred threshold changes in methylation status and / or protein marker expression / activity level of the biomarkers provided herein are about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 50%, about 75%, about 100%, and about 150%. In yet other embodiments, a "nomogram" can be established whereby the methylation status and / or protein marker expression / activity level of a prognostic or diagnostic indicator (biomarker or combination of biomarkers) is directly related to a trend toward a given outcome. Those skilled in the art are familiar with the use of nomograms to relate such two values, with the understanding that the uncertainty in the measurement is the same as the uncertainty in the marker concentration because it refers to the measurement of an individual sample rather than a population average.
[0219] In some embodiments, a control sample is analyzed simultaneously with the biological sample, allowing results obtained from the biological sample to be compared to those obtained from the control sample. It is further contemplated that a standard curve can be provided and assay results for the biological sample compared to the standard curve. Such a standard curve indicates the methylation status of a biomarker and / or protein marker expression / activity level as a function of assay units, e.g., fluorescent signal intensity, when fluorescent labels are used. Samples collected from multiple donors can be used to provide a standard curve of control methylation statuses of one or more biomarkers in normal tissues and "at risk" levels of one or more biomarkers in plasma collected from donors with a particular type of cancer. In certain embodiments of the method, a subject is identified as having cancer upon identification of an abnormal methylation status of one or more DMRs and / or protein marker expression / activity levels provided herein in a biological sample obtained from the subject. In other embodiments of the method, a subject is identified as having cancer by detection of an abnormal methylation status and / or protein marker expression / activity level of one or more such biomarkers in a biological sample obtained from the subject.
[0220] Analysis of markers can be performed separately or simultaneously with additional markers within a single test sample. For example, it is possible to combine several markers in a single test to efficiently process multiple samples and potentially provide greater diagnostic and / or prognostic accuracy. Furthermore, those skilled in the art will recognize the value of testing multiple samples from the same subject (e.g., at consecutive time points). Such testing of consecutive samples can enable the identification of changes in marker methylation status and / or protein marker expression / activity levels over time. Changes in methylation status and / or protein marker expression / activity levels, as well as the absence of changes in methylation status, can provide useful information regarding disease status, including, but not limited to, identifying the approximate time from the onset of an event, the presence and amount of recoverable tissue, the appropriateness of drug therapy, the effectiveness of various therapies, and the subject's outcome, including the risk of future events.
[0221] Biomarker analysis can be performed in a variety of physical formats. For example, microtiter plates or automated applications can be used to facilitate the processing of large numbers of test samples. Alternatively, single sample formats can be developed to facilitate immediate treatment and diagnosis in a timely manner, for example, in an outpatient or emergency room setting.
[0222] In some embodiments, a subject is diagnosed with a particular type of cancer if there is a measurable difference in the methylation status and / or protein marker expression / activity level of at least one biomarker in the sample compared to a control methylation status and / or protein marker expression / activity level. Conversely, if no change in the methylation status and / or protein marker expression / activity level is identified in the biological sample, the subject can be identified as not having, not at risk for, or at low risk for a particular type of cancer. In this regard, subjects with or at risk for cancer can be distinguished from subjects with or at low risk for cancer as those who are substantially free of cancer. Subjects at risk for developing a particular type of cancer can be placed on a more intensive and / or regular screening schedule. Meanwhile, subjects at low risk to those at substantially no risk can be avoided from undergoing additional cancer risk testing (e.g., invasive procedures) until future screening, e.g., screening performed according to various embodiments of the present disclosure, indicates that the subject is at risk for cancer.
[0223] As described above, depending on the embodiment of the disclosed methods, detecting a change in the methylation state and / or protein marker expression / activity level of one or more biomarkers can be a qualitative or quantitative determination. Thus, diagnosing a subject as having or at risk of developing a particular cancer type indicates that a specific threshold measurement is made, e.g., that the methylation state and / or protein marker expression / activity level of one or more biomarkers in a biological sample changes from a predetermined control methylation state and / or control protein marker expression / activity level. In some embodiments of the methods, the control methylation state is any detectable methylation state of a biomarker. In some embodiments, the control protein marker expression / activity level is any measurable and / or protein marker expression / activity level of a protein marker. In other embodiments of the methods, in which a control sample is tested simultaneously with the biological sample, the predetermined methylation state is the methylation state in the control sample, and the predetermined protein marker expression / activity level control state is the and / or protein marker expression / activity level in the control sample. In other embodiments of the methods, the predetermined methylation state and / or predetermined protein marker expression / activity level are identified based on and / or by a standard curve. In other embodiments of the method, the predetermined methylation state and / or predetermined protein marker expression / activity level is a specific state or range of states. Thus, the predetermined methylation state and / or predetermined protein marker expression / activity level can be selected within acceptable ranges that would be apparent to one of skill in the art, based in part on the embodiment of the method being performed, the desired specificity, etc.
[0224] Furthermore, with respect to diagnostic methods, preferred subjects are vertebrate subjects. Preferred vertebrates are warm-blooded, and preferred warm-blooded vertebrates are mammals. Preferred mammals are most preferably humans. As used herein, the term "subject" includes both human and animal subjects. Accordingly, veterinary uses are provided herein. Accordingly, embodiments of the present disclosure provide for the diagnosis of mammals, such as humans, as well as mammals of endangered importance, such as the Amur tiger, mammals of economic importance, such as animals raised on farms for human consumption, and / or animals of social importance to humans, such as animals kept as pets or in zoos. Examples of such animals include, but are not limited to, carnivores, such as cats and dogs; swine, including pigs, hogs, and wild boars; ruminants and / or ungulates, such as cows, oxen, sheep, giraffes, deer, goats, bison, and camels; and horses. Thus, diagnostics and treatments for livestock, including but not limited to domestic pigs, ruminants, ungulates, horses (including racehorses), and the like, are also provided.
[0225] 4. Samples, Kits, and Controls Embodiments of the present disclosure provide techniques for screening for multiple types of gynecological cancer from a biological sample. According to these embodiments, the present disclosure includes, but is not limited to, methods and compositions for detecting the presence of multiple types and / or subtypes of gynecological cancer from a biological sample. In some embodiments, the biological sample is a tissue sample, a blood sample, a plasma sample, a serum sample, a whole blood sample, a secretion sample, an organ secretion sample, a cerebrospinal fluid (CSF) sample, a saliva sample, a urine sample, and / or a stool sample. In some embodiments, the tissue sample is a gynecological tissue sample including one or more of vaginal tissue, vaginal cells, cervical tissue, cervical cells, endometrial tissue, endometrial cells, ovarian tissue, and ovarian cells. In some embodiments, the tissue sample is an ovarian tissue sample, an endometrial tissue sample, or a cervical tissue sample. In some embodiments, the subject is a human.
[0226] In other embodiments, the terms "sample," "test sample," and "biological sample" refer to a fluid sample containing or suspected of containing a methylated DNA marker of the present disclosure. A sample may be derived from any suitable source. In some cases, a sample may include a liquid, a flowable particulate solid, or a solid particle suspension. In some cases, a sample may be processed prior to analysis as described herein. For example, a sample may be separated or purified from its source prior to analysis. In certain examples, the source is a mammalian (e.g., human) bodily substance (e.g., bodily fluid, blood such as whole blood, serum, plasma, urine, saliva, sweat, sputum, semen, mucus, tears, lymph, amniotic fluid, interstitial fluid, cerebrospinal fluid, feces, tissue, organ, one or more dried blood spots, etc.). Tissues may include, but are not limited to, gynecological tissue, oropharyngeal tissue, nasopharyngeal tissue, skeletal muscle tissue, liver tissue, lung tissue, kidney tissue, cardiac muscle tissue, brain tissue, bone marrow, cervical tissue, skin, etc. A sample may be a liquid sample or a liquid extract of a solid sample. In some embodiments, the sample source may be an organ or tissue, such as a biopsy sample and / or a secretion sample (e.g., gynecological secretions), which may be solubilized by tissue disruption / cell lysis. Additionally, the sample may be a nasopharyngeal or oropharyngeal sample obtained using one or more swabs that, once obtained, are placed into a sterile tube containing viral transport medium (VTM) or universal transport medium (UTM) for testing.
[0227] A wide range of fluid sample volumes can be analyzed. In some exemplary embodiments, the sample volume can be about 0.5 nL, about 1 nL, about 3 nL, about 0.01 μL, about 0.1 μL, about 1 μL, about 5 μL, about 10 μL, about 100 μL, about 1 mL, about 5 mL, about 10 mL, etc. In some cases, the fluid sample volume is about 0.01 μL to about 10 mL, about 0.01 μL to about 1 mL, about 0.01 μL to about 100 μL, or about 0.1 μL to about 10 μL.
[0228] In some cases, the fluid sample may be diluted before use in the assay. For example, in embodiments where the source containing the methylated DNA marker is a human body fluid (e.g., blood, serum, secretions), the body fluid may be diluted with an appropriate solvent (e.g., a buffer such as PBS buffer). The fluid sample may be diluted about 1-fold, about 2-fold, about 3-fold, about 4-fold, about 5-fold, about 6-fold, about 10-fold, about 100-fold, or more before use. In other cases, the fluid sample is not diluted before use in the assay.
[0229] In some cases, the sample may undergo pre-analysis treatment. Pre-analysis treatment may provide additional functions, such as removal of non-specific proteins and / or effective yet inexpensively implemented mixing functions. Common methods of pre-analysis treatment may include the use of electrokinetic trapping, AC electrokinetics, surface acoustic waves, isotachophoresis, dielectrophoresis, electrophoresis, or other pre-concentration techniques known in the art. In some cases, the fluid sample may be concentrated before use in the assay. For example, in embodiments where the source containing the methylated DNA marker is a human bodily fluid (e.g., blood, serum, secretions), the bodily fluid may be concentrated by precipitation, evaporation, filtration, centrifugation, or a combination thereof. The fluid sample may be concentrated about 1-fold, about 2-fold, about 3-fold, about 4-fold, about 5-fold, about 6-fold, about 10-fold, about 100-fold, or more, prior to use.
[0230] It may be desirable to include a control. The control may be analyzed simultaneously with the sample from the subject, as described above. Results obtained from the subject sample may be compared to results obtained from the control sample. A standard curve may be provided to which the assay results of the sample may be compared. Such a standard curve shows the level of one or more methylated DNA markers as a function of assay units. Samples from multiple donors may be used to provide standard curves for reference levels of methylated DNA markers in normal healthy tissue and "at risk" levels of methylated DNA markers in tissue from donors, who may have one or more characteristics of gynecological cancer.
[0231] Embodiments of the present disclosure also include kits for carrying out the methods described herein. The kits include embodiments of the compositions, devices, apparatus, etc. described herein, as well as instructions for using the kit. Such instructions describe appropriate methods for preparing an analyte from a sample, e.g., methods for collecting a sample and preparing nucleic acid from the sample. Individual components of the kit are packaged in suitable containers and packaging (e.g., vials, boxes, blister packs, ampoules, jars, bottles, tubes, etc.), and the components are packaged together in suitable containers (e.g., box(es)) for convenient storage, shipping, and / or use by the user of the kit. It is understood that liquid components (e.g., buffers) may be provided in lyophilized form to be reconstituted by the user. The kit may also include controls or references for assessing, validating, and / or ensuring the performance of the kit. For example, a kit for assaying the amount of nucleic acid present in a sample may include a control containing a known concentration of the same or another nucleic acid for comparison, and in some embodiments, a detection reagent (e.g., primers) specific for the control nucleic acid. The kit is suitable for use in a clinical setting and, in some embodiments, for use in the user's home. The components of the kit, in some embodiments, provide the functionality of a system for preparing a nucleic acid solution from a sample. In some embodiments, certain components of the system are provided by the user.
[0232] In some embodiments, the disclosure provides compositions (e.g., reaction mixtures). In some embodiments, the disclosure provides compositions comprising a nucleic acid containing a DMR and a reagent capable of modifying DNA in a methylation-specific manner (e.g., a methylation-sensitive restriction enzyme, a methylation-dependent restriction enzyme, and a bisulfite reagent) (e.g., a methylation-sensitive restriction enzyme, a methylation-dependent restriction enzyme, a ten-eleven translocation (TET) enzyme (e.g., human TET1, human TET2, human TET3, mouse TET1, mouse TET2, mouse TET3, Naegleria TET (NgTET), Coprinopsis cinerea (CcTET)), or a variant thereof), a borane reducing agent). Some embodiments provide compositions comprising a nucleic acid containing a DMR and an oligonucleotide described herein. Some embodiments provide compositions comprising a nucleic acid containing a DMR and a methylation-sensitive restriction enzyme. Some embodiments provide compositions comprising a nucleic acid containing a DMR and a polymerase.
[0233] In some embodiments, the technology described herein is associated with a programmable machine designed to perform an array of arithmetic or logical operations, such as those provided by the methods described herein. For example, some embodiments of the technology are associated with (e.g., implemented in) computer software and / or computer hardware. In one aspect, the technology relates to a computer that includes a form of memory, elements for performing arithmetic and logical operations, and a processing element (e.g., a microprocessor) for executing a set of instructions for reading, manipulating, and storing data (e.g., the methods provided herein). In some embodiments, the microprocessor is part of a system for determining methylation status (e.g., of one or more DMRs in Tables 1 or 2); comparing methylation status; generating a standard curve; determining Ct values; calculating methylation rates, frequencies, or percentages; identifying CpG islands; determining assay or marker specificity and / or sensitivity; calculating ROC curves and associated AUCs; and sequence analysis, all of which are described herein or known in the art. In some embodiments, the microprocessor is part of a system for determining the level of protein expression and / or activity (e.g., one or more protein markers described herein); comparing the level of protein marker expression or activity to standard non-cancerous levels; all of which are described herein or known in the art.In some embodiments, the microprocessor is part of a system for determining methylation status (e.g., of one or more DMRs in Tables 1 or 2); comparing methylation status; generating a standard curve; determining Ct values; calculating methylation rates, frequencies, or percentages; identifying CpG islands; determining assay or marker specificity and / or sensitivity; calculating ROC curves and associated AUCs; and sequence analysis; all of which are described herein or known in the art, and / or determining protein expression and / or activity levels (e.g., one or more protein markers described herein); and comparing protein marker expression or activity levels relative to standard non-cancerous levels.
[0234] In some embodiments, the software or hardware component receives results from multiple assays and determines and reports to a user a single value result indicative of cancer risk based on the results of the multiple assays (e.g., determining the methylation status of one or more DMRs in Table 1 or 2 and determining the expression and / or activity levels of, e.g., protein markers). Related embodiments calculate a risk factor based on a mathematical combination (e.g., weighted combination, linear combination) of results from multiple assays (e.g., determining the methylation status of one or more DMRs in Table 1 or 2 and determining the expression and / or activity levels of protein markers). In some embodiments, the methylation status of the DMRs defines a dimension and can have values in a multidimensional space, and the coordinate defined by the methylation status of the multiple DMRs is a result, e.g., related to cancer risk, for reporting to a user, e.g.,
[0235] In some embodiments, various embodiments of the present disclosure involve multiple programmable devices that work in concert to perform the methods described herein. For example, in some embodiments, multiple computers (e.g., connected by a network) can operate in parallel to collect and process data, for example, in an implementation of cluster computing or grid computing or some other distributed computing architecture that relies on complete computers (on-board CPU, storage, power, network interfaces, etc.) connected to a network (private, public, or the Internet) by traditional network interfaces such as Ethernet, fiber optics, etc., or by wireless networking technology.
[0236] For example, some embodiments provide a computer including a computer-readable medium. The embodiment includes a random access memory (RAM) coupled to a processor. The processor executes computer-executable program instructions stored in the memory. Processors such as these may include microprocessors, ASICs, state machines, or other processors, and may be any of a number of computer processors, such as processors from Intel Corporation of Santa Clara, California, or Motorola Corporation of Schaumburg, Illinois. Processors such as these may include or be in communication with a medium, such as a computer-readable medium, that stores instructions that, when executed by the processor, cause the processor to perform the steps described herein.
[0237] In some embodiments, the computer is connected to a network. The computer may also include multiple external or internal devices, such as a mouse, CD-ROM, DVD, keyboard, display, or other input or output devices. Examples of computers include personal computers, digital assistants, personal digital assistants, cellular phones, mobile phones, smartphones, pagers, digital tablets, laptop computers, Internet appliances, and other processor-based devices. Generally, these computers related to aspects of the technology provided herein can be any type of processor-based platform, running any operating system, such as Microsoft Windows, Linux, UNIX, or Mac OS X, and capable of supporting one or more programs, including the technology provided herein. Some embodiments include personal computers running other application programs (e.g., applications). Applications can be stored in memory and can include, for example, word processing applications, spreadsheet applications, email applications, instant messenger applications, presentation applications, Internet browser applications, calendar / organizer applications, and any other application executable by a client device. All such components, computers, and systems described herein as related to the technology may be logical or virtual.
[0238] In some embodiments, the present disclosure provides a system for screening for one or more types or subtypes of gynecological cancer in a sample obtained from a subject. Exemplary embodiments of the system include, for example, a system for screening for multiple types or subtypes of gynecological cancer in a sample obtained from a subject (e.g., a stool sample, a tissue sample, an organ secretion sample, a CSF sample, a saliva sample, a blood sample, a plasma sample, or a urine sample). In some embodiments, the system includes an analytical component configured to determine the methylation status of one or more methylation markers in the sample and / or determine the expression and / or activity levels of one or more protein markers in the sample, a software component configured to compare the methylation status of the one or more methylation markers in the sample and / or the expression and / or activity levels of the one or more protein markers in the sample with control or reference samples recorded in a database, and an alert component configured to alert a user of a cancer-related condition.
[0239] In some embodiments, the alert is determined by a software component that receives results from multiple assays (e.g., determining the methylation status of one or more methylation markers) (e.g., determining the expression and / or activity levels of one or more protein markers) and calculates a value or result to report based on the multiple results.
[0240] Some embodiments provide a database of weighting parameters associated with each methylation marker provided herein for use in calculating a value or result and / or alert for reporting to a user (e.g., a doctor, nurse, clinician, etc.). In some embodiments, all results from multiple assays are reported. In some embodiments, one or more results are used to provide a score, value, or result based on a combined result of one or more results from multiple assays that is indicative of a subject's risk of cancer. Such methods are not limited to particular methylation markers. In such methods and systems, the one or more methylation markers comprise bases in a DMR selected from the DMRs of Table 1 or 2.
[0241] In this detailed description of various embodiments, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the disclosed embodiments. However, those skilled in the art will understand that these various embodiments may be practiced with or without these specific details. In other instances, structures and mechanisms are shown in block diagram form. Furthermore, those skilled in the art will readily appreciate that the specific order in which the methods are presented and performed is illustrative, and that the order can be changed and still be within the spirit and scope of the various embodiments disclosed herein.
[0242] The various components of the kit may be provided in suitable containers as needed. The kit may further include a container for holding or storing a sample (e.g., a container or cartridge for a urine, whole blood, plasma, serum sample, tissue, or bodily secretion sample). If desired, the kit may also optionally contain reaction vessels, mixing vessels, and reagents or other components to facilitate preparation of the test sample. The kit may also include one or more instruments to aid in obtaining a test sample, such as a syringe, pipette, forceps, measuring spoon, etc. In some embodiments, the instrument is a collection device, including, but not limited to, a tampon, a lavage device that releases liquid into the vagina and recollects the fluid, a cervical brush, a Fournier cervical self-sampling device, and a swab. In some embodiments, a biological sample is obtained from the subject, and the method further includes extracting a DNA sample from the biological sample. In some embodiments, the biological sample is collected with a collection device having an absorbent member capable of collecting the biological sample upon contact. In some embodiments, the absorbent member is a sponge configured for insertion into an orifice. [Example]
[0243] 5. Working Example As will be readily apparent to those skilled in the art, other suitable modifications and adaptations of the methods of the present disclosure described herein are readily applicable and discernible and may be made using appropriate equivalents without departing from the scope of the present disclosure or the aspects and embodiments disclosed herein. Having described the present disclosure in detail, the same will be more clearly understood by reference to the following examples, which are intended merely to illustrate certain aspects and embodiments of the present disclosure and should not be construed as limiting the scope of the disclosure. The disclosures of all journal references, U.S. patents, and publications referenced herein are hereby incorporated by reference in their entirety.
[0244] The present disclosure has multiple aspects, illustrated by the following non-limiting examples.
[0245] Example 1 Experiments were conducted to evaluate the feasibility of a panel of methylated DNA markers (MDMs) for detecting non-specific gynecological cancers, site-specific gynecological cancers (e.g., cervical cancer, ovarian cancer, endometrial cancer), and specific subtypes of gynecological cancer.
[0246] Using a proprietary methodology of sample preparation, sequencing, analytical pipeline, and filters, we identified differentially methylated regions (DMRs) and narrowed them down to DMRs that accurately represent specific types of gynecological cancer and excel in a clinical trial setting. Cross-cancer analysis of RRBS data identified 249 hypermethylated gynecological cancer-specific (either endometrial cancer (EC), ovarian cancer (OC), or cervical cancer (CC)) DMRs (Table 1). These included one or two (at most) specific hypermethylated regions in gynecological cancers, as well as subtype-specific regions. These experiments also revealed 89 regions that are universally hypermethylated in all three gynecological cancers (Table 2). Marker features include extremely low background noise (≤0.01) in leukocytes, which reduces inflammatory signals and potentially enables plasma-based testing. Signal in BCV tissue was also low (≤0.05) for the marker, which is the predominant cell type from tampon devices. The AUC of all MDMs was greater than 0.90 in separating cancer from leukocytes and ≥ 0.85 in discriminating one cancer from another.
[0247] Table 1: Methylated regions that distinguish specific types of gynecological cancer (either endometrial cancer (EC), ovarian cancer (OC), or cervical cancer (CC)) from benign tissues (genomic coordinates can be obtained using the Human Feb.2009 (GRCh37 / hg19) Assembly). [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4] [Table 1-5] [Table 1-6] [Table 1-7] [Table 1-8] [Table 1-9] [Table 1-10] [Table 1-11]
[0248] Example 2 Following the experiments described in Example 1, a novel set of differentially methylated regions (DMRs) was identified that distinguished multiple types of gynecological cancer from non-neoplastic control DNA, as shown in Table 2.
[0249] Table 2: Universally methylated regions present in all three gynecological cancers (e.g., endometrial cancer (EC), ovarian cancer (OC), and cervical cancer (CC)) assayed from benign gynecological tissues (genomic coordinates can be obtained using the Human Feb. 2009 (GRCh37 / hg19) Assembly). [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4] [Table 2-5]
[0250] Example 3 From the two marker sets in Examples 1 and 2, 25 candidate markers were selected for validation studies using independent cases and controls (Table 3). Methylation-specific PCR assays were developed from the DMR sequences and tested on tissue samples. Short amplicon primers (<150 bp) were designed to target the most discriminatory CpGs within the DMR and tested on analytical controls to ensure stable and linear amplification of fully methylated fragments and absence of amplification of unmethylated and / or unconverted fragments. Tissue samples from 82 EC (16 serous, 18 carcinosarcoma, 7 clear cell, 17 endometrioid grade 1 / 2, and 24 endometrioid grade 3), 82 OC (36 serous, 21 clear cell, 4 mucinous, and 21 endometrioid), and 64 CC (36 squamous cell and 28 adenocarcinoma) were compared with benign epithelial controls (29 cervicovaginal, 29 fallopian tube, and 14 benign endometrial tissues). As shown in Table 3, CDO1 and DLGAP1 differentiated any cancer type from benign control tissues, but gynecological cancer specificity was evident in most MDMs.
[0251] Table 3: Candidate markers selected for validation (OC: ovarian cancer, Ser OC: serous ovarian cancer, clear cell OC: clear cell ovarian cancer, endo OC: endometrioid ovarian cancer, muc OC: mucinous ovarian cancer, CC: cervical cancer, Ad CC: cervical adenocarcinoma, sq CC: cervical squamous cell carcinoma, EC: endometrioid carcinoma, endo EC: endometrioid endometrioid carcinoma, pan gyne: non-specific gynecological cancers (e.g., OC, EC, and CC)). [Table 3]
[0252] Additionally, we provide representative data for calibration plots, adjusted boxplots, and adjusted boxplots by subtype for each of the 25 identified candidate MDMs, as shown in Figures 2-26. Taken together, whole-methylome sequencing, strict filtering criteria, and biological validation for gynecological cancers yielded candidate MDMs for site-specific and universal detection of gynecological cancers.
[0253] Example 4 DNA methylation is an early event in the development of endometrial cancer (EC) and may be useful for detecting EC. One of the most promising sample types for clinical trials is vaginal fluid from tampons or similar collection devices. Preliminary global methylome NGS testing, followed by independent validation, identified discriminative EC-associated methylated DNA marker (MDM) candidates, which were then tested in tampon samples self-collected by women with and without EC. In this example, an additional round of testing was performed on a subset of tampon samples using several novel MDMs and a panel of novel epithelial reference assays that provide a measure of total epithelial shedding.
[0254] Briefly, we reanalyzed a previous reduced representation bisulfite sequencing (RRBS) study, including DNA from frozen EC, benign endometrium (BE), benign cervicovaginal (BCV) tissue, and benign buffy coat samples, to identify epithelial reference genes and several novel EC MDMs. Candidate reference markers were selected based on receiver operating characteristic (ROC) identification, fold-change methylation levels, differential methylation, and p-values for all three epithelial tissue types relative to buffy coat (white blood cell) samples. EC MDMs were selected by the same criteria, but EC tissue was compared with the other three sample types. Several other previously identified EC MDMs were also selected for vaginal fluid testing. A quantitative methylation-specific PCR (qMSP) assay was developed and tested in 50 women aged ≥45 years with abnormal uterine bleeding (AUB) or postmenopausal bleeding (PMB), or women of any age with biopsy-proven EC, using tampon-collected vaginal fluid before clinically indicated endometrial sampling or hysterectomy. Cases included 25 women with biopsy-proven EC and 25 controls with benign biopsies.
[0255] Four candidate epithelial reference markers were selected from two regions related to previous tissue RRBS data: FNBP1, NCOR2, and S1PR4. Methylation in EC, EB, and BCV tissues was consistent, concordant (all CpGs), and robust (>50%). In contrast, methylation in leukocytes was <1%. Two new EC markers, MDM, GYPC, and CYP26C1, met the cancer-specific criteria of the RRBS data (AUC >0.85; absolute mean CpG methylation in EC >20%; methylation fold change ratio (case / control) >10; p-value <0.001). These markers, along with LBX2, SPDYA, ZSCAN12, and TERC (forward and reverse strands) from Examples 1 and 2, were tested in a 50-sample tampon pilot. The four reference gene markers were strongly positive in all 25 cases and 25 controls, with the forward strand S1PR4 assay showing the most robust and consistent methylation between cases and controls. In EC-specific MDMs, TERC (forward strand) had the highest performance (AUC = 0.88), and SPDYA had the lowest performance (AUC = 0.60).
[0256] Reanalysis of the endometrial / cervical RRBS discovery data with respect to reference epithelial markers yielded candidate MDMs for vaginal fluid samples, and most of the EC-associated MDMs tested showed promising high performance in tampon-collected vaginal fluid (Table 4 ).
[0257] Table 4. Methylated regions that can distinguish endometrial cancer (EC) from benign tissue (genomic coordinates can be obtained using the Human Feb.2009 (GRCh37 / hg19) Assembly). [Table 4]
[0258] Example 5 Early detection and treatment of endometrial cancer (EC) results in excellent prognosis, and surgery alone is often curative, especially in stage IA disease. However, presentation of EC at advanced stages almost always requires multimodality therapy and results in less favorable oncological outcomes. Therefore, we conducted a study to expand the current repertoire of candidate methylated DNA markers (MDMs) for EC using methylome sequencing discovery and independent sample validation experiments. This included both the more common endometrioid histology and the less common, more invasive EC histology. The performance of these novel EC MDMs was tested in vaginal fluid obtained via self-collected tampons from women with perimenopausal AUB, PMB, or newly diagnosed biopsy-proven EC.
[0259] This study was conducted in three phases. First, tissue-based discovery of methylated DNA markers (MDMs) was performed using reduced representation bisulfite sequencing (RRBS) on DNA extracted from frozen ECs and benign tissues. Second, biological validation of EC-specific MDMs was performed using quantitative methylation-specific PCR (qMSP) on DNA extracted from an independent group of formalin-fixed, paraffin-embedded (FFPE) ECs and benign endometrium (BE). The third phase involved clinical translation of MDM detection by qMSP in DNA extracted from vaginal fluid samples (collected via patient-administered intravaginal tampons) obtained from women with EC, atypical endometrial hyperplasia (AEH), endometrial hyperplasia without atypia, or benign endometrium (BE).
[0260] Primary fresh-frozen EC tissue was identified from the EC biorepository, which prospectively maintains >1,500 frozen samples collected from consenting patients undergoing hysterectomy for EC or AEH. ECs included in the discovery phase represented the five most common EC tissue types: grade 1 / 2 endometrioid, grade 3 endometrioid, serous, and clear cell carcinoma, and uterine carcinosarcoma. Frozen tissue blocks were required to have at least 70% tumor purity for inclusion. Benign endometrial (BE) tissue was collected from consenting patients using a Pipelle or EndoSampler immediately after clinically indicated office-based endometrial biopsy for women aged ≥45 years who participated in workup for AUB or PMB, as an additional sample for research. EC tissue and BE menstrual or atrophic endometrium were confirmed by a gynecologic pathologist. Benign cervicovaginal (BCV) squamous tissue was collected from both premenopausal and postmenopausal women undergoing hysterectomy for benign indications. BE and BCV tissues were fresh-frozen until DNA extraction. Buffy coats were collected from healthy, cancer-free control female donors currently undergoing cervical cancer screening and mammography. Women diagnosed with other cancers or who had taken chemotherapy-class drugs within the past 5 years, had previously undergone pelvic irradiation, were diagnosed with synchronous cancer at the time of EC, or had a history of solid organ or bone marrow transplantation were excluded. Clinical variables were extracted from the electronic medical records of all included subjects.
[0261] An independent set of women with newly diagnosed EC who underwent hysterectomy for initial treatment was identified for the biological validation cohort. Formalin-fixed, paraffin-embedded (FFPE) EC tissue representing the same histology as the discovery cohort was included. Additionally, FFPE BE tissue from women undergoing hysterectomy for benign indications, age-matched frequency- and FFPE atypia-free endometrial hyperplasia symptoms, and AEH tissue from women undergoing hysterectomy were obtained. A gynecologic pathologist (MES) reviewed all tissues and also selected tissue block sections for macrodissection. Eligibility criteria were the same as for the discovery set.
[0262] Vaginal fluid was collected from two groups of women via self-applied tampons (Tampon Pilot). One group included women aged ≥45 years presenting to the Mayo Clinic Gynecology Department for workup of AUB or PMB, or postmenopausal women without bleeding who were referred for evaluation of endometrial thickening (ES) by pelvic ultrasound. Women were excluded if they had not undergone clinical endometrial sampling or if they had undergone endometrial sampling within the past 3 months. Final clinical pathology diagnoses were used. EC by endometrial sampling or hysterectomy pathology was included as EC cases. All women with a final clinical diagnosis of AEH or endometrial hyperplasia without atypia were included in the exploratory analysis, and any women with benign endometrial sampling were eligible as BE controls. The other group consisted of women aged ≥18 years presenting to the Mayo Clinic Gynecological Oncology Department for biopsy-proven EC or AEH and clinically indicated hysterectomy. Eligibility criteria were the same as for the discovery and biological validation cohorts.
[0263] After verification of diagnosis and tissue block selection by one of the study gynecological pathologists (JKC, SEK), microtome sections were performed on frozen EC and BCV tissues embedded in optimal cutting temperature (OCT) compound to obtain ten 10-micron scrolls. BE whole-frozen tissue samples were used, collected via one or two in-clinic biopsies using a Pipelle or EndoSampler. Genomic DNA was purified from tissue and buffy coat specimens using the DNeasy Blood and Tissue protocol and the QIAamp DNA Blood protocol (Qiagen, Valencia, CA), respectively. DNA was repurified with AMPure XP beads (Beckman-Coulter, Brea, CA) and quantified with PicoGreen (Thermo-Fisher, Waltham, MA). DNA quality was assessed using real-time quantitative PCR. RRBS libraries were prepared. Briefly, 300 ng of DNA was 1) digested with MspI, 2) ligated to methylated sequencing adapters, 3) treated with sodium bisulfite (Epitect Bisulfite Protocol, Qiagen), 4) amplified and enriched with adapter-specific primers, and 5) size-selected (160–280 bp) using AMPure beads to remove primer dimers and larger CpG-sparse regions. The library was sequenced at the Mayo Clinic Medical Genomics Facility using an Illumina HiSeq 2500 instrument (Illumina, San Diego, CA). Candidate genomic differentially methylated regions (DMRs) were selected as described below.
[0264] Quantitative methylation-specific PCR assays were developed from the CpG methylation signatures of selected DMRs. Primers were designed using MethPrimer to target the bisulfite-modified sequences of each identified gene and a CpG-free reference region within the β-actin gene. Primer quality control was performed with 20 ng (approximately 6,250 genome equivalents) of positive and negative methylation controls. DNA was bisulfite-converted using the EZ-96 DNA Methylation Kit (Zymo Research, Irvine, CA) and amplified using SYBR Green detection on Roche 480 LightCyclers (Roche, Basel, Switzerland). A serially diluted universally methylated DNA sample served as a positive control, while negative controls included bisulfite-converted and unconverted genomic DNA from leukocytes, as well as converted whole-genome amplified (unmethylated) DNA. MDM results were normalized to β-actin. Assay performance was validated using samples from the discovery cohort. Markers that performed suboptimally compared to the RRBS results and cutoffs (see below) were not considered further.
[0265] MDMs were then tested using qMSP on DNA extracted from independent FFPE EC and BE tissues. These MDMs were also tested in AEH and non-atypia-containing endometrial hyperplasia tissues. After histological verification and selection of the most representative macrodissection site by the examining gynecological pathologist (MES), tissue blocks were macrodissected using a 1 mm or 2 mm core punch. DNA was purified using the Qiagen QIAmp FFPE DNA Tissue Kit (Part No. 56404) and bisulfite converted as described above. Samples were blinded, randomized, and assayed by qMSP as described above.
[0266] Consenting women in both groups within the Tampon Pilot self-applied regular-sized unscented Playtex® tampons to collect vaginal fluid. Subjects enrolled in groups participating in workup for AUB, PMB, or thickened ES had tampons applied in the clinic before their gynecological examination and removed the tampon before clinically indicated pelvic examination and endometrial sampling. Subjects in groups presenting with biopsy-proven EC or AEH had tampons applied in the preoperative area on the day of hysterectomy and removed the tampon in the operating room. Tampon dwell time in the vagina was recorded in both groups.
[0267] After removal, each tampon was placed in a 50 mL conical tube containing sterile PBS buffer and centrifuged through a mesh filter to separate the pellet and supernatant, which were stored at -80°C until DNA extraction. Approximately midway through prospective enrollment in the tampon pilot, 50 mM EDTA was added to the PBS buffer to facilitate DNA recovery and reduce nuclease degradation. Tampon pellet DNA was extracted using a High Pure Viral Nucleic Acid Kit (Roche, Basel, Switzerland) and quantified using a Qubit Fluorometer (Invitrogen, Walther MA). DNA was bisulfite converted and assayed by qMSP against MDMs selected as described above using β-actin as the reference gene.
[0268] For discovery, we used a previously published approach. Briefly, Mayo Clinic's in-house analytical software package, the Streamline Analysis and Annotation Pipeline for RRBS (SAAP-RRBS), was used for quality scoring, sequence alignment, annotation to the University of California, Santa Cruz reference genome, and differential analysis of DMRs. Candidate CpGs were excluded if data coverage within each sample group was <50%. CpG islands are typically biochemically defined by an observed expected CpG ratio >0.6. However, for this model, regions containing five or fewer CpGs were excluded to create DMRs based on the distance between the CpG site locations on each chromosome. DMRs were then selected so that background methylation rates in benign controls (BE, BCV, and buffy coat) were <2% and then ranked by the AUC of the EC tissue referent relative to the benign controls. Statistical significance was determined by overdispersed logistic regression of the methylation percentage per candidate DMR based on read counts. To account for the varying read depth across individual subjects, an overdispersed logistic regression model was used, and the dispersion parameter was estimated using Pearson's chi-square statistic of the residuals from the fitted model. Candidate genomic DMRs were ranked according to the significance level, AUC, and difference in fold change between EC and benign controls (BE, BCV, and buffy coat) and selected for further testing. Sample size estimation for discovery was based on the method described above.
[0269] Secondary DMR analysis was performed to identify endometrium-specific MDMs that were methylated in both EC and BE but unmethylated in BCV and buffy coat.
[0270] In the independent validation study, the sample size was selected to maximize the precision of sensitivity and specificity (minimizing the width of the 95% confidence interval (95% CI)). Assuming a specificity of 95%, a set of 29 controls provides a 95% CI within ±10%. A minimum of 84 samples was required to achieve a 95% CI within ±7% for the target sensitivity of 90%. The distribution of individual markers was examined using boxplots and marker intensity maps. Accuracy was assessed by generating AUC values for each marker. A random forest (rForest) model was used to generate the predicted probability of a sample representing an EC case. Random forests perform cross-validation of the randomized marker set using 500 unique randomized training and test sets (approximately 2 / 3 for the training set and 1 / 3 for the test set) from bootstrap selection to generate 500 models. The precision error based on out-of-bag samples was then averaged across the 500 models. Marker selection was performed using the VSURF package in R. The VSURF package uses random forests in three steps to 1) eliminate the least significant predictors, 2) select all predictors associated with the response variable, and 3) reduce redundancy in the final marker selection.
[0271] For the tampon pilot, the sample size estimate was defined to detect an AUC of 0.70 (AUC = 0.50 indicates probability). 100 EC and 92 BE had greater than 90% power to detect this difference using a one-sided test at the 5% significance level. ECs were frequency matched to control BEs by subject menopausal status and tampon collection date.
[0272] All women in the practice group who underwent evaluation for AUB, PMB, or thickened ES and were diagnosed with AEH or endometrial hyperplasia without atypia after tampon collection were included for exploratory analyses. Additionally, the following tampon pilot subanalyses were performed: 1) MDM sensitivity and specificity for EC when restricted to vaginal fluid samples collected before endometrial sampling (in the context of presumed natural endometrial DNA shedding), and 2) MDM performance specifically in PBS / EDTA-buffered vaginal fluid samples.
[0273] RRBS was performed on 69 cases of EC (16 grade 1 / 2 endometrioid, 16 grade 3 endometrioid, 11 serous, 11 clear cell carcinoma, and 15 uterine carcinosarcoma), 44 cases of BE (14 proliferative, 18 disordered proliferative, and 12 atrophic), 18 cases of BCV, and 18 buffy coat samples from healthy female donors. The clinicopathological characteristics of the discovery-stage EC cases and BE controls are detailed in Table 5.
[0274] Table 5: Clinicopathological characteristics of EC cases and BE controls in the discovery cohort. [Table 5-1] [Table 5-2]
[0275] The sequencing coverage depth across all samples ranged from approximately 40 to 50X. On average, CpGs filtered by CpG context with at least 10X coverage (our minimum requirement for inclusion) averaged approximately 1.7 million per sample. A DMR-calling algorithm applied to multiple comparisons (total EC vs. total controls, histological EC subtype vs. BE, histological EC subtype vs. BCV, etc.) yielded a total of 323 statistically significant DMRs. By imposing performance cutoffs (AUC > 0.85, absolute mean CpG methylation in EC > 20%, methylation fold change ratio (cases / controls) > 10, p-value < 0.001), the number of DMRs was reduced to 54. A targeted qMSP assay was constructed and tested for the 54 selected DMRs. Twenty-one targeted qMSP assays were subsequently discarded because they either failed QC or did not perform well compared to their respective sequencing data and / or cutoffs above.
[0276] Independent histologic testing was performed on the remaining 33 MDMs. Samples included 141 cases of EC (34 grade 1 / 2 endometrioid, 31 grade 3 endometrioid, 27 serous, 19 clear cell, and 30 uterine carcinosarcoma), 112 cases of BE (35 secretory, 30 proliferative, 19 chaotic proliferative, and 28 atrophic), 35 cases of AEH, and 24 cases of endometrial hyperplasia without atypia. See Table 6 for clinicopathological characteristics.
[0277] Table 6: Clinicopathological characteristics of biological validation cohort EC cases, benign endometrium (BE) controls, atypical endometrial hyperplasia (AEH), and endometrial hyperplasia without atypia. [Table 6-1] [Table 6-2]
[0278] Several MDMs (EMX2OS, CYTH2_4043, MPZ, NBPF24) demonstrated AUCs >0.85 when distinguishing EC (all histological types combined) from BE, and most uniquely distinguished specific EC histological subtypes from BE. For example, EEF1A2, which had an AUC of 0.59 when comparing all EC histological subtypes with BE, distinguished EC from clear cell BE with an AUC of 0.91 (range, 0.80-1). In the analysis of clear cell EC vs. BE, 14 of 33 MDMs had an AUC ≥0.85. In contrast, only 5 of 33 MDMs distinguished carcinosarcoma from BE with an AUC ≥0.85. When both EC tissue-combined MDMs and tissue-specific MDMs were evaluated, only 10 MDMs had an AUC below 0.85 across all comparisons. Thus, 23 MDMs had an AUC ≥0.85 in either the tissue-combined or tissue-specific analysis.
[0279] For the tampon pilot experiment, four MDM assays (CDH4, LYPLAL1, c17orf64, and KRT86) were developed for endometrium-specific DMRs (i.e., methylated in both EC and BE, but not BCV tissue). In addition to these four, 24 MDMs were selected from biological validation, for a total of 28 MDMs tested in the Tampon Pilot: CDH4, c17orf64, CYTH2_4043, DIDO1, EEF1A2, EMX2OS, GATA2_3670, GDF7, JSRP1, LRRC8D_8831, LRRC34, LRRC41, LYPLAL1, SMPD5, MAX.chr10.4460, MAX.chr12.52652239-52652424, LINC02323, MDFI, MPZ, NBPF24, OBSCN, SEPT9, SFMBT2_0970, SQSTM1_3864, VILL, ZNF90, ZNF323, and ZNF506.
[0280] The tampon pilot included 100 women with EC, 92 with BE, 11 with AEH, and 25 with endometrial hyperplasia without atypia. EC cases included 31 women with clinical evaluation for AUB or PMB who were diagnosed with EC after tampon collection and 69 women with biopsy-proven EC before tampon collection. All women with BE, all but three with AEH, and all but one with endometrial hyperplasia without atypia were from the group of women with perimenopausal AUB or PMB, and their tampons were collected before endometrial sampling. EC cases included 49 cases of grade 1 / 2 endometrioid, 9 cases of grade 3 endometrioid, 24 cases of serous, 4 cases of clear cell carcinoma, 9 cases of uterine carcinosarcoma, and 5 cases of mixed EC histology. The clinicopathological characteristics of the EC cases, BE controls, and the group of endometrial hyperplasia without AEH and atypia are detailed in Table 7 .
[0281] Table 7: Clinicopathological characteristics of Tampon Pilot Cohort EC cases, benign endometrium (BE) controls, endometrial hyperplasia without atypia, and atypical endometrial hyperplasia (AEH). [Table 7-1] [Table 7-2]
[0282] When comparing combined EC cases from both the AUB / PMB and biopsy-proven EC groups with BE controls, 28 MDMs individually had significant methylation fold changes compared to controls. Table 8 lists the AUC for distinguishing EC from BE for each of the 28 MDMs tested in the Tampon Pilot. The 28-MDM panel distinguished EC from BE with a specificity of 96% (95% CI 89-99%) and a sensitivity of 76% (66-84%) (AUC 0.88 [0.82-0.93]). When the number of MDMs was reduced to a 3-MDM panel in a post-hoc analysis, the combination of SFMBT2_0970, NBPF24, and MAX.chr10.4460 yielded the same AUC as the 28-MDM panel. When comparing stratified AUCs, considering age ≥64 years vs. <64 years (median age of tampon pilots) and BMI ≥30 vs. <30 kg / m², there were no statistically significant differences for any covariate. When analysis was restricted to tampon samples collected before endometrial sampling, the 28-MDM panel distinguished EC (n=31) from BE with a specified specificity of 96% (95% CI 89-99%) and a similar sensitivity of 74% (95% CI 55-88%) (AUC 0.87 [0.77-0.98]).
[0283] Exploration of the performance of 28 MDMs in tampon specimens from women subsequently diagnosed with AEH or endometrial hyperplasia without atypia revealed lower methylation intensity compared to EC and higher intensity compared to BE.
[0284] As previously mentioned, approximately midway through the prospective vaginal fluid collection study, 50 mM EDTA was added to the PBS tampon buffer to improve DNA stabilization. Of all EC cases and BE controls in the tampon pilot, 57 EC cases and 52 BE cases had tampons collected in PBS / EDTA buffer. The AUCs for each of the 28 individual MDMs in distinguishing EC from BE based on tampons collected in PBS / EDTA buffer are listed in Table 8. The combined 28-MDM panel demonstrated improved sensitivity when tested on tampon specimens collected in PBS / EDTA buffer (96% (95% CI 87-99%) specificity; 82% (70-91%) sensitivity (AUC 0.91 [0.85-0.97]) compared to the complete tampon pilot containing both PBS alone and PBS / EDTA-buffered vaginal fluid (Table 8). Furthermore, in a subanalysis of PBS / EDTA buffer with specificity set at 95%, the 28-MDM panel accurately identified: 17 of 20 endometrioid ECs (85%), 18 of 23 serous ECs (78%), all of 9 uterine carcinosarcoma (100%), 2 of 3 clear cell ECs (67%), and 1 of 2 mixed ECs (50%).
[0285] Table 8: AUCs of the 28 DMRs included in the panel tested in the tampon pilot. Analysis was performed on all samples (100 ECs, 92 BEs) including tampons collected in PBS alone + tampons collected in PBS / EDTA. The subanalysis on tampons collected in PBS / EDTA included 57 ECs and 52 BEs. AUCs are listed in descending order based on the analysis of PBS alone + PBS / EDTA. Cancer specificity is also provided (CC: cervical cancer, OC: ovarian cancer, Ser OC: serous ovarian cancer, clear cell OC: clear cell ovarian cancer, EC: endometrial cancer, clear cell EC: clear cell endometrial cancer, pan gyne: non-specific gynecological cancer (e.g., OC, EC, and CC)). [Table 8]
[0286] Rigorous tissue discovery and validation identified unique EC MDMs that were detectable in tampon-collected vaginal fluid, demonstrating their efficacy in triaging patients with perimenopausal AUB or PMB using self-collected samples. Translation to tampon-collected vaginal fluid samples demonstrated that the 28-MDM EC panel tested in the Tampon Pilot had high sensitivity and specificity in distinguishing underlying EC from BE. This high sensitivity and specificity appeared to be maintained even when a smaller, three-marker panel was evaluated. The sensitivity for detecting EC in this context also remained high in a subanalysis that included only vaginal fluid samples collected from women presenting with perimenopausal AUB or PMB before the underlying endometrial lesions were determined by endometrial sampling. These data support the conclusion that EC-associated MDMs are spontaneously shed vaginally.
[0287] 6. Materials and Methods The following materials and methods were used to identify various DNA methylation markers capable of distinguishing one or more types and / or subtypes of gynecological cancer in biological samples from subjects having or suspected of having gynecological cancer.
[0288] Samples. Tissue and blood samples were obtained from the Mayo Clinic Biospecimen Repository under institutional IRB oversight. Samples were selected in strict compliance with the study approval and inclusion / exclusion criteria. Tissues were macrodissected, and an expert GI pathologist reviewed the tissues. Samples were age- and sex-matched, randomized, and blinded. Cervical cancer (CC) subtypes included 1) adenocarcinoma and 2) squamous cell carcinoma. Controls included benign cervicovaginal (BCV) tissue and whole blood-derived leukocytes. Endometrial cancer (EC) subtypes included 1) serous EC, 2) clear cell EC, 3) carcinosarcoma EC, and 4) endometrioid EC. Controls included non-neoplastic uterine tissue and whole blood-derived leukocytes. Ovarian cancer (OC) subtypes included 1) serous OC, 2) clear cell OC, 3) mucinous OC, and 4) endometrioid OC. Controls included non-neoplastic fallopian tissue and leukocytes from whole blood. DNA from 190 frozen tissues (16 grade 1 / 2 endometrioid (G1 / 2E), 16 grade 3 endometrioid (G3E), 11 serous, 11 clear cell EC, 15 uterine carcinosarcomas, 44 benign endometrial (BE) tissues (14 proliferative, 12 atrophic, 18 chaotic proliferative), 18 serous OC, 15 clear cell OC, 6 mucinous OC, 18 endometrioid OC, 6 benign fallopian tube, 14 benign fallopian tube brushings), 88 formalin-fixed paraffin-embedded (FFPE) cervical cancers (CC), and controls (36 squamous cell, 34 adenocarcinoma, 18 BCV), as well as 36 buffy coats from cancer-free women, was analyzed using the QIAamp DNA Tissue Mini Kit (frozen tissue), the QIAamp DNA FFPE Tissue Kit (FFPE tissue), and the QIAamp DNA Blood Buffy coat samples were purified using the Qiagen Mini Kit (Valencia, CA). DNA was repurified with AMPure XP beads (Beckman-Coulter, Brea, CA) and quantified with PicoGreen (Thermo-Fisher, Waltham, MA). DNA integrity was assessed using qPCR.
[0289] Sequencing. Reduced representation bisulfite sequencing (RRBS) was performed on two batches of samples (first on endometrial and cervical samples, then on ovarian samples). To account for variability, randomly selected samples from batch 1 were also included in batch 2. Sequencing libraries were prepared according to a modified Meissner protocol (Gu et al. Nature Protocols 2011). Samples were combined in a 4-plex format and sequenced by the Mayo Genomics Facility on an Illumina HiSeq 2500 instrument (Illumina, San Diego, CA). Reads were processed through Illumina pipeline modules for image analysis and base calling. Secondary analysis was performed using SAAP-RRBS, a bioinformatics suite developed by Mayo. Briefly, reads were cleaned using Trim-Galore and aligned to the GRCh37 / hg19 reference genome constructed with BSMAP. For CpGs with coverage ≥ 10X and base quality score ≥ 20, methylation rates were determined by calculating C / (C+T) or, conversely, G / (G+A) in the case of reads mapping to the opposite strand.
[0290] Biomarker selection. We used a proprietary DMR (differentially methylated region) identification pipeline and regression package to derive DMRs based on the mean methylation values of CpGs. Differences in mean methylation percentages were compared between cancer and buffy coat controls, and tiled reading frames within 100 base pairs of each mapped CpG were used to identify DMRs with control methylation <2%. DMRs were analyzed only if the total coverage depth was an average of 10 reads per subject and the variance between subgroups was >0. Assuming a biologically relevant increase in odds ratio of >3x and a coverage depth of 10 reads, ≥18 samples per group were required to achieve 80% power at a 5% significance level with a two-sided test and a binomial variance inflation factor of 1.
[0291] After regression, DMRs were ranked by p-value, area under the receiver operating characteristic curve (AUC), and fold-change difference between cancer and buffy coat controls. AUC was required to be >0.90 and FCD >20. No false positive adjustment was performed at this stage because independent validation was planned in advance.
[0292] The three cancers were analyzed separately and by subtype as described above to generate individual lists of optimally performing DMRs. The combined sample and CpG-level data were then added to the three DMR lists. Each CpG had to be present in 80% or more of the samples being compared. For each list, DMRs were then ranked by hypermethylation ratio, i.e., the number of methylated cytosines relative to the total number of cytosines at that location at a given locus. The ratio had to be ≥ 0.20 (20%) for cancers; ≤ 0.05 (5%) for BCV tissue controls; and ≤ 0.01 (1%) for buffy coat controls. Regions that did not meet these criteria were discarded. Furthermore, the pattern of CpG methylation within the region had to be continuous or consistent. The candidate DMRs (for each cancer and subtype) were then analyzed separately against the other two cancers using a logistic regression (using the average CpG methylation). For example, serous EC regions meeting the filtering criteria were compared with ovarian cancer (total) and then with cervical cancer (total). To qualify as a site-specific DMR, in this case a DMR for serous EC, the FCR between the serous EC cancer sample and either the OC or CC sample (or both) had to be 5-fold or greater. To qualify as a universal DMR, a marker had to be represented in each of the optimal lists (above).
[0293] Biomarker validation. A subset of site-specific and universal cancer DMRs was selected for further development. The criteria were primarily logistic-derived area under the receiver operating characteristic curve (ROC) measures, which provide a performance assessment of the region's discriminatory potential. An AUC of 0.85 was selected as the cancer-versus-cancer tissue comparison cutoff. Methylation differences were also a significant factor. There was a limited feasibility of 20–30 methylated DNA markers (MDMs). This was primarily due to the limited amount of sample DNA and the amount of work required to develop a high-performance analytical assay. Quantitative methylation-specific PCR (qMSP) primers were designed for candidate regions using MethPrimer (Li LC and Dahiya R. MethPrimer: designing primers for methylation PCRs. Bioinformatics 2002 Nov;18(11):1427-31 PMID:12424112), and 20 ng (6250 equivalents) of positive and negative genomic methylation controls were QC-tested. Multiple annealing temperatures were tested for optimal discrimination. Validation was performed by qMSP on independent tissue samples.
[0294] These tissues were identified as previously described by expert clinical and pathological review. DNA purification was performed as previously described. The bisulfite conversion step used the EZ-96 DNA Methylation Kit (Zymo Research, Irvine, CA). 10 ng of converted DNA (per marker) was amplified using SYBR Green detection on a Roche 480 LightCycler (Roche, Basel, Switzerland). Serially diluted universally methylated genomic DNA (Zymo Research) was used as a quantification standard. A CpG-independent ACTB (β-actin) assay was used as the input reference and normalization control. Results were expressed as methylated copies (specific marker) / ACTB copies.
[0295] Statistics. Results were analyzed for individual MDM performance. Calibration plots were examined to confirm the suitability of ACTB normalization. Boxplots (ray and calibrated) and heat matrices were generated to demonstrate the numerous epigenetic relationships between cancers, subtypes, and controls.
[0296] Sequences. The various nucleotide sequences referred to in this disclosure are provided below.
[0297] MAX.chr10.4460:
[0298] MAX.chr1.2152: CGACTCTTCGAGCGCCCCTCTGCTTCTGTAGAGGGGTCGAGCCATGTCAAGGTAGACCCTGTGTCGGCCCGTCTCCCTCGGATCCTCCGCACCAATCACTGTTGCTGAATCCGACACCCGGCGGATCCAGTGCGGAGTCTCGAACAGCTGCGGAGCTGGGAGCTACGGGACATGAGGAGTGCGGGGGGGAAGAGAAGACGGCGGAGGAAAATCCCCCGGCGGTGCTCAACTGCGGCTTTCTCTCTCGGCTGTGAGCCGGCTCCGCCCTCCGGCTTCCAGAGCAAGTGGCTTCTGCGTTCACCGCCCCCCGCCGTTTGTGGGGCGGGGCCGATTCATAAGAATCGGTTCTCACCAATGGAGGGCTTAGCATGTTTAACCTCAGGATCATAAACAAAAGACACTGCTAGAACGGTCGGGAAAGTCATACGCTTTGCTTATCTTATATATAGATTTCTAAAATTCCAAACCGGGGACGCGTTGGTGGTGTAGTGGTGAGCACAGCTGCCTTTCAAGCAGTTAACGCGGGTTCGATTCCCGGGTAACGAAACGTTTTTGTCTTTCCTTCTACGAAAAACTTTTCTGAGCCG (SEQ ID NO: 2).
[0299] MAX.chr11.0394: CCCTTGAGGCCAGGAGTTCGAGACCAGCCTGTGCAACACAGAAGACACTATCTCTACAAAAAATTAAAAAATTAGCTAGGCATAGTGGCACATGCCTGTGGTCCCAGCTACTCCGGAGATTGAAGCAGGAGGATCACTTGAGTGAGGGAGGTGGAGGCTCCAGTGAGTCGTGATCGTGCCACTGCACTCCAGCCTGGACGACAGAGCGAGACCCTGCCCCCTCGCCAAAAAAAAAATACTGGGATGCTATACACAAAATTGCCTTGAAAACTTGAGCACGGAACACCAAACAGCTAAGCGTGCCGGTTTGGGGAGGGCGGGGGAGGAATAAGGAGCTGCAACGGTAAGAGGCCGCCACACGGTGGCGCAGTGAGGCTGGGAAACGGTGCACCCCGCGCAGGAGGGGGCACTCCCCGTCGCGGCCACCCGGGGTGGGCAGGAGGCGGCGCGGGCTGGCTGGTCTCTCCCGAGAAGGTTCTCTCCCGAGAAGGGTGCGTCTCAGGGCTTGTCAGTGGACCCCTGGAACATGGGGAAGACGCACAGACAAGGGTTTCGCTCTTTGCTCTCCTCTCTCCTTGTCAGACCTCTGTGACC(SEQ ID NO: 3).
[0300] MAX.chr11.3750:
[0301] MAX.chr14.7696: CGGCACGTGGGTGGGCGATGACGCCATTTACTGAGATTTGATCCCCACCACACGGCTCCGGGGTGAGAATTATGACATCTGGCTCAACACGGCTGCCCGGGACCCACACAGCCGAGCGGCCGAGCTCGGGCGGAGTCCCAGGGCGCCCACAGCACCCCGCCAGCGCGCCCCGTCCAGCAGGGCAGCTTTTGGGCGGAGGCGACCCCCACCGCAGGTCCCAGGACCCTGCGTGCTCTTGAGCCAGGGGTGGAGAGGCCCGACCGCGGGGGGCTGCCCCACCCCGCCGCCCTTCACCGCGAGCCGGGGCCCAGACCGCCCAGCCACGCCCAGAGCCCGCGGGCGGAGACGCCAGGGGCGGTGCCAGCGAGGTCCCGTCCCCGGTACCCCGTCCCGCCCCCCCACACGCGGTGACCTGGGGACGCCCCGCGGGAGTCGTTCTGCGGCTCCCCCTGGCGTCGGCTGGGGCCACCGCCCGGGCTCCCACCTCAACCCTGCAATATGGGGTTGGGGCAGAGTGGTCTGCTGCCCGCTGCCCGCAGCCGCTTCCGGTTAGGGAGGGAGCCTGGGCCTCTGGGTGCTCACGCTGCGCTTAACGCTGGTCCCGGCAGCAGTGAGGGTGGAAGCGGCCG (SEQ ID NO: 5).
[0302] MAX.chr19.5552: CGGACCGTAGCTCCTTCCACGCATGAACCCCGCACACGAGTCGGGATTCCCCCCATGACCCTCCCGTGGCCCCCGCACAATCTGGAGAGACGCGGGGCTGCGGGCGCGGAGCTGCCCAGAGAGGACTCCTGCCCGGGCCCGCAGTCGCCGCGAAGGGACGGGACAGGACGCCCGGGGTCCCGGCTGCCAGCCCAGCCCCACCCTGCGGCCGAGGGGACCGAGGGCCGAGCTCCGCCAGCGGTACTCCGGTCCACAGAGCCCGGAGTCGCTGGCTGGGAGGCCGGGGACCCGCCACGGCCAGTTCCAACCAGCCCCTCCTCCCGTCTCGGGATCCCTGGCCCCTCACGCTCACCATTTTCCGAATTCCTCCGTGTCCCGGGGGCCTCTCTGCGGCTCCCACGACCAGTGCAGGTCCCTGTGTGACAGAGGCTGCCGCAGACTCTCCAGAGTGCCTCTCAGCGACAGAGACAGGAGCCCAGCGAAGTGGCGTGTAGAAGACGCCGCGGGCTTTTTCAATCTCGCACCCTCTTAGCTGAAGTGCGCCTGATTGACAGTTCCCACGACCCCGCCCCACGGCCCTGATTGGATAGTGCGACAGATCCCGCCCCCTGACGACTGAGTTACAGAAGCGATCTCACG (SEQ ID NO: 6).
[0303] MAX.chr19.0548:
[0304] MAX.chr2.8918: GAAGTCAGGGCAGTGCTGCAAAACCTCCACAGTGCGGAATTCCGGGAAAATTCTTTACAGAGGTGTGGAGGTGGAGGAAAGCTTCCTGGGCAGGCCTTTGGGGTCGTCCCCACGCAGGCGCTTGCAGCCACCCCAGCTCGCGCGGGGCCGGGCTTTGGGGTGTGAGAGCTGGGACGGGAGTCGGGTGGATGCCTGGCCGGAGCCGCCAGCTCCCCTCGTCCTCTTTGCTTGTCCTTTAGCACAAGGGCGAGCAGCGTAGGACAAAGACTCGGGCGGCAGCTGCCTGGTTCGGCGCGCAGGGGCGGCCTCGGCCACCCGGGGCGCCCGCCGCCTCCACCGCCCCGCGGGGGAGGCCCGATGCCCGTCTTTGTCTGTGCCGCCGCCGTGGGCCGGGTCCGCAGGAAGCGGGCGCCATCGTGCGGCCTGAGCTGGACACTGCGCCCCCGGAGGCGCGGAGGCGCGAACCACCAAGCGTGGCTCCAAGCTCCACGGGGACGCTGGTGTCATCGTGGCCACGACTGCTTGTACTGTTGTGGTGCGTTCTCTTTTGTATACTAAGTGCTGTGTGAACACAGAACCACTTCCAGTAAATGCAACTGAGCCGTCGCCAGCAGAAAAAG (SEQ ID NO: 8).
[0305] MAX.chr2.4778: AACAGTGGCGCTCAGAGAAGACAGGACAGCGGGCGAGAGCTTGGGGGGCGATGGGAGGTGGAGAGGCACTCCAGGTCCCCAGGGGGCCAGGCGGAGCTGCGGGACAGGGCGCAGACCCCGAGGCCCAGGGAGCACCGGGTGGCCGGCGGCCTGCAGGCTGGCGAGGGCGTCGGGCGGCGCAGGGCAGGCCAGGGGGCGGGGGCGTCTGGGGCCCTGGCGTGGCGCCCGGAACACCCCGTGCCGGAAGCTCCATGTGACCGTGACTCCGCAGAAGCCGCGAGCGCAGCGAAACAAAGGGCGGCTCTGCGGCCGCCTCGAGCTCAGGCTGGCACCGAGGGCCCGGACCCCCATCCCACTCCGCACCCCCGGGCCTCCCGGCCCTTCTTGCCCTCCGACCCCGGGCTCTGGCAGGGCCGGGAGGCGCAGGAACCCCGCGGGGGATGGGGCCGGCGGACTGGCACTGAAGACACTGGGATGCAAGCGGGAGGCTGGGGGCGGGGGGCTGGGGGCGGGGGGCGGGGCTGCAGGGCGTGGACGGTCTC (SEQ ID NO:9).
[0306] MAX.chr20.3853: CGGAGCGGATATTCCCGGAGCCCCTCTGCGAGCCACGCGCCCCTCTGGGAAGCCCGCTTCCCCCTGCAGACAGGCGCTGTGACACGCTTGCGCCCCGGTCGAACAGGCGAAGAGGCCGAGGCCCAGAGCGGCGCAGGGCGAGCCTGGAGGCTGCGCCCCAGACCTGGACCAGCCACGGACGCCGCTCCCGCCGCTCCCTCCGCTCCGCTGCGCTCCGCACGCTGGCGCCGGCTCCCCGAGGCCCCGGGCGCCCCGGCCGCACGCCTGGGTAAAAGGTCCCGAGGAGTCCGCAGAAGCGCGCCCACGCCCGAGACGGCCGTTTCCGCCGGCCTGGGAAAGGGGCGGAGAAGGGGGTCGCCCGGGCCGCAGCGTGCCGGTCCCCGCCGGCCGAGCCGTGTTTGGGGCCAGTCCCCGCACCCCGCTTCTTCCCCACCTGGGGAGCGGGGCGCCGCGGTAGGGGCACTGGAGCGCACGATGCACCCCGCCAACGAGTCCTTTCTGCAGACGGGGTTCCTGTTTTCATGCAAATGCCTTTGTTAGCGCACCGGGAACCAGGGGGACGGAACTGCAGCTGACGCGGGCTGCGCGCCGCTTTTCCGCCTTCGCTTGATTCGGCCTCAACGACTTAAACGCGCCGGGAACAAAAACGCCGGCGCCGCGGAAACCCTCAGGAGCGGCACGAGAAGCGCGGCTCCGCTGGGGTCCGCGAGAAGCGGTGCGGGCGGCCG (SEQ ID NO: 10).
[0307] MAX.chr20.2903: CGCTCGCCCCCTTCTCCGCGCGGGCCCTCAGCTCAGCTCCCTCTTCGCTCCCCGTGTCCCCGCGAGCGGGAGGGAGGGGATGCTAGGACGCCCTGTCGGCGTCGTCGCCGCTTTCCGCCATTGTTTAGTCGTGATGCTCTCATTTTCTCTGAATCAACAATTTTCTGCTCGGCTCCGCGCCGACCGGCGAACGCGGGGCTTTTCCTCGCCCGCCTGATGACAGCAGAGCGGCGCGGAGCAGCTGGTCCGGAAGGAAGCGCCAGGCGCCTGCCCGGTCCCAGGCGTCCGCTGCCGCCCACCCACACCAGACCCCGCCCCCGCGCGTCAAGCCCCGCCCATCCATACCAAGTCCCGCCCCCACACCCTCACCCACACACCAGGCCGTCCCCACCCCGCCCCCAGAGCCCCGGGGCGCCCCGCCCGCTAGCCGCGCACGCGCAGTGAGCACGGCGACCCCCGGTGGTCGGGTGTCTCCGCAGGCCGAACACGCTGCTCGCCCAGCTGCGGATCATTACCGCCCTTTTGTTCTCCGTCGCGCGCTCGCCCCACGCTAGGAATGCAAACTGTAGGCGCCG(SEQ ID NO:11).
[0308] MAX.chr21.5011: TGTTTTTCCAAAAGATAATAAGCGTCAACAACAACAAAAAAATAAAAAGTCCAACTCCGCCCCAAAGCAGCATCTGGCTGGCCTGCGAGATGCCCACTGGGGAGGCGAGTCCGCAGCTTAGGACTCAAGCCCGGGGTCGGAAGCTATTGCCGAAATCCGAAACGCAGCGCTCGCAGCTGCAGTGACGCGACCTGCTCATAAGTCCCCGTGCTCACAGCATCCCGGCAACTTACGAGCTAGTGCTTCCGGGTCACCCCGGCCCAGGAAGGCGCACGCGCGAAGCATAGCGAGCTTCACTCCGCACTCTTAGGCTGCGTGTGAGGCCTGCGAGTGCTCGGGAGCTGCCGCGGTCACAAGAGAAAGCCTAGCTGTCAATGACAGCCCCAGAGCATCTGGGCGCCTTGCGATACCCGGGTGTCTGTAGGCAGCCAGGAGACACTTCCAAGCTGATCTGGAATCTTTCCTCGCCCAGCTCTGTCCCTCGCAGGGATGGCAAAGGACATTACGACCTACATCCCTTCCCGGATCTGATGGCTTCAGATTGGCAGATTGTGTTAAAGTGGAAGGCTCGTGGTGCCCCTTTGCTGAGTTTTTATGGACTTAGTTTTCCCAAGTAGTTCTAATTATCG (SEQ ID NO: 12).
[0309] MAX.chr22.5665: CGGCTGTCTTTGTCTCCCGCGAGGCAACTCTGACTCAGGCTCCAGCTGCCCGTGGGAGGGAGGGGGCGCCCGGGCTCCTGAGGTCGCCAGGGAGCGGCGGGACTGGGAGGCTCCAAAGCCCTCAGTGTACGTGCGAATCCGGAGCGGACACCGAGACCTTAGCGCGGGAACCAAGAGAGGACAGAGCTCCACGGAGGCCACAGCGCGTGCACGGGGACAGGTGCGCCCTCCCCGGCAGCCCCCCTGCTCCTCGGTCACAGTTCTGTGCGGAGGCGTCTTGCGCCCTCCCCCCTGAGCCTCGCCCTTGAGTCGGGGCCGTGGGCCGCATCCAGGCCCCCAGGGCTCGGGATGCGCGTGAGGACCCGGACTCCCGAGGGCGCAGAGGTCGGGAGCCCGAAGCAGGCGCCCTTGGCCTTGGTCCCGCCCCTTATCCGGTCCCAAGCTTTTTCCTCGCCCCTTGGCCTTGACTCCACCCCTTAGGCATGCCGCTGGCCCCGCCCCTTTCCGGCCACCTTGAGGCTTGGGGGTCCCTCAGCCCCGCCTCTCTTCTTGACCCCGCCCCTTGGCAGCACCCCCTACCCCCGCCCCACGTCCAAATCTCCCGGGGCCGGTGGTGGCCGGGGCTGACGGCGGAAGCCGCGCAGAGACTCGCTTGCCCCGAAGTCGCTGGATTCGGGCCTGGATCCCAGATTATCCGCAGCCTAGGGGAGTGGAGAGATGCCCAAGGTTCCTCTGGGTCCCGGGACCCCAGTAGCGTCCCTCCCCCCGTCCCCCACGCCAACCACTGAGCGCCCTTCGGAGTCCCGGGAGGAAAGCGTAGGGGCGGGGAACTCTGGCATCTCTCTCCTCCCGGTTGCTCCCCGACTCTGCCCCGCTATTCCGCTATTTGGGGCAGTCGTTTCTACCG (SEQ ID NO: 13).
[0310] MAX.chr3.6408: CGCCGACCCCAGACCCAGTCCTAGTCCGGCCAGAGGAGGCCGTTTACGAGCCCACACCCGTAGGTGGCGCCACAGCCGGAGAATTGGCTTTGGTTCTGTTGGAGCCGCGCCGCCTTTAAATTAGCCCCACGCATGCGCGACTTTTCTAGCCCGAGCCCGCCGTCTGCGCCTGCGATTTCGCCCATACTCCCCGGTGCCCGCCTCGTGACGTGCCGCAGTGTTACGAAGGGACACCAGGGCGCAGGCGCAGCTCGCTCCTCAGGCCTGCGAGAGGCCCGCCGGCCCAGCAGAGGGCGCCCACCAATCTGCACGCGGGCCCAGCGAGTTATCTTGATTTCGGCCAAGCTTTCTGACTGCTCCAAAAAACGAAGAAAAAGATTCAGGGAGAGTAAGAGGATGAAGAGAGCTGGTGAAGCAGCTGACCAAATGGCCCGAGGTGGTATGCAGCCGCGGTAAAGCAGGGCCCCTCCGTGAGGCACAGCCGCCCGGGGGTTCCCTAGGGGAAGCAGGGGCTGCGGCAGACGCCTCTCGGGCAGGTCAGGGTATGCACCCTCCCGCAGGGGCTCCCAAGGCCGGGCGTGCGTGAGGCCAGGTCCACGGGCACACCACTGTGAACACTGATTAAACGTGGCCTCCACGGCTTCCAACCCCCAGGACAGGACCAACCCTCTGCCCCCGGCTCAGGCCAGAGCTCGGCGAGACCGTCG (SEQ ID NO: 14).
[0311] MAX.chr5.3588: CGCGTCCGGAGGAAGGCTCACCCGGAGGCCGCCTGCAGGCGGCCAGGTGCCAGCCACTGCGGGCCTCTGGGGCCGAAGCCGGCGGATGGTGAGACGCTGGTTGGTCTGCAACACTGCCCAGACCCCGGGCACTCATGTCTAGAAAAAAGCTGACTCGTGACACCAAAGGAGCTCTTTCAAGCTTCCTGCACGCTCTTAGCGCCAGAGCACCCCAGCCGTCCTGGGAGCCCCCGAAGCCAAGCATATTCGAACTCCGAATCCGCTCGATCGCCGGGGACCTGCCATCTGGGTTCGGTTCCCCAAGGTCGCTGCCGACCTTAGACCGCGGGGGTGTGGGGCGCCGGGGAAGGAGACAGAAGGACAGGCGCCGCCCAGGGCCGCGGGGACACTTGGGGCTGCGTCCTGGGTGGGCGCGATGCTCCCCCAGAACGACTGGAGATGGAGAGTGTCGGGGAGGGAAACGGGACCCACGAATTCAGGGCGCTGAGTCGGCGGAATGCGCCCTGACTCCCCCTGGCCGAGAGCCGGCTCAGAATGAAAGAGCGCGGAGTGGGAGGTCTGGGAAATGGCAGTATTTGTATTAGGGAGAGAAGGAAACAGGAGTGGGAGCCGCACGGCTTGGGGAACGCGGGAGATCGCGGATTGCGGGGATAGCGCAGCGCGGCTGCCCGGGGCTGCTGGGAGGGGCCGGACGAGGCCAGGGCGAGCGGGGTAACTGCGGCCGGCCGGACGGCGGCGGTAACCGGCTGCACCGAGGTGGTTCCACCACCGCGCTGGGCGCTTGCGGTTCGTCTGCTCCAACGAAAGCCGCGTCCCACGCTCCCTGCCGCCGCGTGGTTTTGCCTCCTCAGAGGGGCAGCGGCGACCCAGGGGCTGGCG (SEQ ID NO: 15).
[0312] MAX.chr8.5938: CGCAGCACCGGGTGTTCCCTCCTAGCCTGGTCGCTCGGGGGGAGCGTTGGTTGGCGGGGTGCAGGTCTGGTGCTCGCTCAGGTGGGCCAGGCACCCGCGCGCCAGGTGAGGCGGGCGGGGGAACACACGCCCCTGGCCCCTGCGCCGCCGTCACGGCCGCCCACCACCCGAGGGCGGGGGTCCTGGTGGGGTGTCGATTCCGCCTCCCCGCCCACAGGCACTGGGCCCCGGGCGGCCACCGGGGTGCGGGGCTCCCAGCGTCTCGGGCTCCCACTGCTTCAGGCCTGTCCAGGGGCGGGGAGCGTCTCTGTGGCCGCGGCGGGATTGCGGCGCGGTGGCCGGGCGTCCCCTGCAGGAAGCTGTTCTCGCTCGCTGCCTCCCCCACCTGGGAGGGAAGCGCCTGGATTTTGGGTCCCGCCGCCCTCCGCGCCCTGGGCCTCCACCTGTGTTCCCAAAGCCCAGCCACGAGTCTGGGGGTGCCGGGCGTGCCGTGGTGGGCGGAGGCTTCCACAGCCCCTCCCTGCCAGGGACGGCGGGGCGGGACATGGCGGGGCCGCACGCACCGGGTGGACGACAGGGGATGCCGCGGGCTCGCGTCAGCCAGGGCG (SEQ ID NO: 16).
[0313] MAX.chr9.4007:
[0314] MAX.chr9.2025: TGATGTACGCCCTGGTGGACAAAAGCTGCTAGTGTCAGCTTGATTTGCAAGATCAACATTCATGAGTTTCACCGCTTAGAAAGGGGCATTATCTGCAAACCGGAGACTGAATGGAAGCCATAAACAAGTGATTTCACACTACCAAGCAGGAAGAATATTCTACCTCCTCAATTATTCTGAACAGCATGTTTGGCCCCTTCCAGGTTCCCACCGCTAGCGAGCCCTCCACCCCTGGTGTGTGCAAACGGTGGACCTTTCGGCGCCTAAGAAACCGGGTGCTGAGCCCGGGAGCAGCGCCTGCTTTTCTTCCCAAGATCCACTCCGGGTTTTGCCTAGCGCTGCTCCGGGAACCCATTCCAGACGCAGGTAACGCCAGGCAACGTTTTCCTTCTCACCCGCCCAAGGCCAGCCCCGAGCCGCCGGGGTTCCAGGCCCAACACAGCACAGATGCACGTTTCAAAATGTGCTCGAATATGCAGCCTGCATCAAAGGCGTTGGGAGGCTCTTTCATCCTCTCAGCTGCCTAAGAAGGGACATGCTCCCAGCTACCCTCATTTGTGGCTGGGTTTACTCTGAAATGAAGATGTACCTCTGGATGCAAAAAGAAAGGGTGGAAGGTTTTTTTCCCCCTAC (SEQ ID NO: 18).
[0315] MAX.chr1.2533:
[0316] MAX.chr13.3357: AGGGTTGACCCCAGTACCTGACTTCTCCGGGAGCTGTCAGCTCTCCTCTGTTCTTCGGGCTTGGCGCGCTCCTTTCATAATGGACAGACACCAGTGGCCTTCAAAAGGTCTGGGGTGGGGGAACGGAGGAAGTGGCCTTGGGTGCAGAGGAAGAGCAGAGCTCCTGCCAAAGCTGAACGCAGTTAGCCCTACCCAAGTGCGCGCTGGCTCGGCATATGCGCTCCAGAGCCGGCAGGACAGCCCGGCCCTGCTCACCCCGAGGAGAAATCCAACAGCGCAGCCTCCTGCACCTCCTTGCCCCAGAGACCGTCCGAGCTGGAGCCACAAGCCCTCCATTCCTCTTGGAATCTTCAACCCCAAGGTAAGGTAAGTTCACCGAGCACCGCCCAGCGATGCGCAGGATCCGGGGGGGATCACGCGCGGCGACCCTACCGAGCGCTCCGTGCGCGCCCCCATCTCTCGGATCGTGTTCCTGGCTCTGTCGAAGCTGCTGAGTCCCGCGATTCGGGAAATCCGGCACTTGTTTCTCACCCTACACCATCACGTGGAAATCATTGAAAATGGGAACCCTGGTGGAGTATCTGGGAGAGCACGCTTGTGCCGAGGGGCCTGAGCTATGGGACTTCCTCCAGGTCCCTCTGTTTCCTGCCGGCGTAGGGGACTCGTAGTGTCGGATCGCATAGTGCCAAAAAATAGTGCATGGGAAACAAACAA (SEQ ID NO: 20).
[0317] MAX.chr14.2093: GTAGAGACGGTGTTTCACCATGTTGGCCAGGATGGTCTCGAACTCCTGACCTCGTGATCTGCCCGCCTCGGTCTCCCAAAGTGCTGGGATTACAGGCGTGAGCCACTGCGCCCAGCCCCAAAATTGGGAATTATTTCAAAATAAAAAGCTGGATAAATGCATACACACAAGGCAGTATCGCGTATTTTCCACGAGTGCCTGTGCAGGCAGGTAAGGATTTAGGAAAGGTCTGGAAGGATGTGCAAAATGTTCCGCCTGCGAAGGTTCCGCGGTGGCGGGGACACTGCTCCGGCTCCGCTCCCGCCCGCCCGAGCGCTCGGATGGGGCCGCCTCTGCACTGCGTGGCCACAGGCGCGGCCCGGCTGCCCACGGGCGCCCTTTGCAGCTGCTGCCCCCTGGCGGCCGCGGGCGGCTACTAGCGGGAAAGCGAAACCCGCCCGGTCCATTCAAGCCCCGCTGCCTGGCGCCCTCTAGGGTCGTTCTTGGGAACGGGCGGACCTTTCGTCAACACTTTGCCTGCAAGATCCCCCATTGGGGGAACCGAGGAGGAAGTTAAAGGAAGATGTGTGTTTTTGAGCGCTGCTTTGTGCCAGGCTCATCTTAGGTGTGGGACGTGTACTATCTGAATTAATACCCCACCAGGCCTGTGGGACAGTCACTGTCACCATTCGCAAATTATGGATGAAGAAAGGAGGTACCAAGTGGTGGTATCACCTGTCCATAGTGAGCTGTCCCTCAGGAGGGTGGCCGCCCCAC(SEQ ID NO: 21).
[0318] MAX.chr17.2455: CGTTAACAATGTCGCGTACACGCCCGAACCGGAGGAACCCCATTCCACGCTCCTTCTGGAACCGAATTCACCTCTGAGGCTTTGGGGCTTCAGAGCCGGAGCCGCTTGGGCAAAACCAGCAGAACAGCGAGAGGGAACGGGCTGGTCTAGCCCTGCCCTGAGCATTTCTACTGAGACCCCCGGTCCTGCTTCTTCCAGCCTCTGCTGGATTTCTCTCCGACCCCTCTGGAGCGAAGCCCTTTGGCCCTGCGTTGCATGCGGCACGGTGCGGGTTCGGGCTCTGCGCTGGAGCCGGGATGCCCTCCGGCGGAGGGTGCGCGTAGGCGGCGCCTGGGCGTGAGCCCCGCCTGCAAGGCTCAGCGTCGGGGAAGCACTTTTCTCGTCGACCCGGGGTCTTTTTCCGCCAAGGAGCTCGGGGCTCAAGAACTCGGGACTGGGCTGTGGGCGGGGCATGGTTTTCCTCTCTGGGCGTCCTAATCTCCAATTTCAGGCAAATTCGCTAGGAAGAACCTTCCCGAGCGCG (SEQ ID NO: 22).
[0319] MAX.chr18.4390:
[0320] MAX.chr19.2732: CTGTTTCAAAACTGTGCCATCTGGATGTTGCAGTATACCCATTTTGTCCTTCCCATACCTGTGCCCGGCCACCTGATGCAAGATGGGCACACAGCCACTGAGGAAGCGGAGTCTGCCGCCTGCCGGCTGCAGGGTGCCCTTAGGGGTGGCCTCGATCGCCGGTGGGGTCCGCATTTCTGGGGGACCCGGCGCCTCGACCCGGAGCGGGGATGGTGGCTCTCTTGCCATAACGGAGAACAGAAGCGGTAGGGTCAGCAAGAGCAGGAAAAGAAAAATAGGGGGAGGGAGGGGGCGCCGGAGAACCCAGGGGTCGCTCAGGCTCGGGCGCGAGGAGGCCCGGGGGTTCCCGCGGCTGGTGCCCGCTGAGGTGAGGGGAGGGGGCCCATGACGCCGCGGCGGCGCGGGCACTCCCTCTGCCCAACTCTCGGCTGAGCGCGGCTCCCGGCTCAGGCCCCTCTGCCGCCGCAGCCGCGGGCCCAGTAGACACAACCCAGCCGAGGAGCAGCAGCAGCAGCAGCGGCGCCCCGCGCTCCCTGGGGCCCTCCAGAAAGTTTTTTTATGGATATCAGCAATCTAATTCTACAATTTATATGGAGAGACAAAAGACTCAGAATAACCAACACAATATTGAAGGA (SEQ ID NO: 24).
[0321] MAX.chr19.4467: CTGTTCCTCTGTGGTGGAGGAAGGGACACGCGCTTTTTTTTCCGACCTTAGGAAGGAACAAGGGAGCCGGGGTCCCCTCCCAGCCTGGGAGCCCTGGGCACAGTCCCGGCTCATTTGTCAGAGCTATCGGAGCCGTCCTCGGGCTGGTGGGAGTTCAGGGCTCTGAAAGGTTTTCTGTCAAGGCTTGAAAGGGGGCCAGGTTTTTTTCCCCCCGGAGCCGCGCAGTCTCGGGGCTGTTGTTCTCAGCAATCGCAGGGCCTCGTGTTAGCAGGAAGCACAGCCAAGTAGGGTTTCCTGCGTGTTGGAGAGAGGAAGCTCCGTAATGTTCTGGGAGGCGATGGTTAAAAATAACTCCGGTATATAAAGACAGCGGAGGGTCCCCTTGTTCGCTCACTCGGGCGCCGGCCGGCTGGACGCAGGGCCGAGCAGGTGGTTTGGGGCCTCGGGAAGGCCAAACCCCCGCCTCTGGGCCCCTGGCTGGGGAAGACACCAGCCAAGTTCAGAGCCCCAAGTCGGCCTCACTTCCACAACTCAGCGTCAGGGACACCGTGGGCGTTTCTGTTTCAAAACGCTTTTCTCCAGCAAAGAACGTAACCTCAAGCTGCTGTCAGGGTAGAGGAATCCCTGCCCCCCGCC (SEQ ID NO: 25).
[0322] MAX.chr2.0490: CGAAGGATGCGGCGCGTGGAAGGAGATGCGCTGACTTGTTCCAACCCATAACCTTTCGCTCGGGTCCCCATGTGCGGGCAGAAGAAGTCAGAGCGGAACAGCCTAGTGCACTGGCAGGGCTCATTGTCTGGGAAGACACCGAGGTCTAGGCAGCTGGGACTGCGGAGTGGAGGCAAGGCCGGAGGCGGCCGGCGGCTTTGTGGAAGTTTCGCGCCGCCAGGCCCTGCGCGCCGCACGGGGCGGTGGAGTTCTTGGGCAGCCCCCGGCGCTTGGCCCACGCCTCCGCTTCCCGCGTGTGGGAAACTCGAGCACCCTACAGGCACCAGGGTAAACTGCCTGTGCCTGGCCCGGTGAGGGTCGCTCCCCCAGGCCCCGTCTCCGCCCGAGGACTGCAGGCCTAGGCCTGCGGGGAGATCCTGAGACCGCGGTGTGCGGGCGCCGGCAGCAGGGCAAGGCAGGGACTGTGCCCAGTCCGCCCGCCAAGGAGATCGCACGCCGGCTTCGCTTCTGAAGCTGCAGACGGAGGCCGTGGTGAGCCTTAGAAAGATCCCGGGACAAAGGCG(SEQ ID NO: 26).
[0323] MAX.chr2.8148: CGCCGGGGCGCAAGGCCGAGTCATCCCAGGCGTCCGTGGGCCGTGATTCCCACTCACGCCGGGGGCCCAGGCAGGCAGAGAAGAGTTAATGAGCGCGCAAGTGCAGGCGGTCACTCCTGGGCCTGAAACTCCCGCGCTGTGCATTCAGGGCCCTCGTGGCTCTCAGAGGCGCGTCCCAGGGGCGCACACTGCACCTTGGGCTGGGCAGCTCCGCCGGGTTGTGGCGAGCGGATGAGGGAAGGACGCAGAAACCAGGGCGGAGGAGCCGCGAGGGGCAGGACGAGGCTGCATGGGCCAGCGAGGGGGTCGACACCGAGCCAGAGTGAGCGCGGGGCCTGGGGCGCAGAGCCCGCCCAGGGAGCCGGGAGACGCCGCGCAAGCTCCCCGGACAAACGCAATGACCGAGGACGCGCGGGCGAGGCCGTCCAGGGAGCCCTGGTCCCTCAGCTGCACCGGACTGAGCCGCGACCGCTCAGCACGCGCTGCTTATAAATCAGGGGTGCGCTTCCCAAGCCCCGGGTGAGGTCCCCTACGTCGGCACAGCCTTAGGAGCTGCAAAGCAGCGCGCGCCTCCGGGGCTCCTGCGCGCCCCTTGAACCCCGCCTCCCGCATCCTCCTGCAACAGCCTGGAGCTCCCTGTGCAGGACGCAGCGGGGGGCGGGGGGCGGTCTTAGGAGGCTGCGGGGCGCACTCCCACCTCCTGCCTCCCCGAGACCCCCAGCGCCTTCTCCAGGGTTTAGAGCGGAGGTGAAGGGGCCTCGTCCTGCACCGCCACTGGGCGCCTGGGCTGTTCATCATCGGTTACCGCCG (SEQ ID NO: 27).
[0324] MAX.chr2.3137: CGTCCTGAGAACCCGAGAGAGCAGGGCCCGCTGGGACAGGCAGGGGAAGGCCTCGGGAGGGACACGACGGTCCGGCAGCAGAGCCTGCGGGGCTGGAGGAGGCGCCCTCCTCTCAGCTGCTCTTCCTGCCCCTTTCGGTGGCGAAGATGGATGGGGCCCGGGGCTTTCGGCGGGGCCCGAGGGGCCGGCGAGGCTGCGGCCCTGGAGCCCCCTGCCTGGCAGCCATTTGGGCCCCAGGGAAATATCGGCGCTTTGGCTAACCGAATTATTCTTTCGGTTTGAGCCAGCTCCCCTTTTTGAGTCAGATCCGGCGGCAGGGCCAGAAAAGCGCTTTCTGAAACCCCAGCGCGGTCCTCGGTGGGGGTGGAATGGGGTGGGGTGGGGGGCGCGGCCGCGCCGCTGGGCGCCCTCCCCGCCCTCCCCCCTCCCCACCCCCAGTCCTCCCTCCGCTGCCCGCCCCCCAAGCCCGGTGTCGCCCCCTCCGCCCCCTGCCGCATCCCCGGAGCCAGTGCCCACAGGGGCCAGGCAGCCCGCAGGGGTCGCTCACGGCTGGTGTAGGGGCTTGGTCCACCACGCTAGTACTTCGGGCACCAAAATAGAAAAAGAATAACGCTTGGAAAGAATCTGATGTTTCCG (SEQ ID NO: 28).
[0325] MAX.chr4.4210: CGAAAACTACCCCGCGGAAACTAGCACAGTGTGCCTGGATGTCTGTGTCCCGGGACCTCGGGGAAGAGGGCCCGCACCGGTCTGCGAATTGCAAGGCCCGGCCTTCCCCAGCGACGCTCTGGTATCCGCTGTCCCCTCCCTGTACCTCCGCGACCCAGGGGACGCCCAGTGCACCAGGCCCTTCCCCGGGGTCAGCGGAGGCGCAGGGCGTTAGCCACATCAGAGGTGCAAATTTACCCCGGGCCCAGGGGAAAATGGCGACAGCGTTCGCGGCTCCACCCGGGGCGCGTGTCAGCGTTGGAGAGCCTGCCCGGCCTGCAGAGGGCGTAACAGGCACCGCTGGGGAGAGCCAAGCACCCCTGCGTCCAGGATCCGTAGCGCCGAGCTGCAGGCCCGACCTGCAGGGGGCGTGCCCGGCATGGGAAGCTCAGGCTACGTCTCCGAAGCTTGCGCTGAAAACACCAGAGGTAGGGAAAACGGGGAGAGCGTACTGTGCTGGGCTCTACCCTGGACACCCCAGTTTCATTCTCTGCGAAGCCACGCGCTGGCAGGGCTCTCGGGACGGCGATACCCAGGGATGATGGTACCCCTGGTCTCGGCGGGACCTCCCGGGAACTTGTCCTGGGGGAGGGAGCCCAACTGGCCACGTACTGGTAGCAGCAGTGGGTGGAGCGCACAAACTCCGAGGCCCGCG (SEQ ID NO: 29).
[0326] MAX.chr5.0931: CGCCTCCTGCCTGAGGCGGGCTGGGGGGTCGTTGTCCTCGCAGCGTTAAGGCGAGTCTGGGACAGGACCCCGGCACCCCCTCCGGATCTGTGGCATCCTCCAGGACTCCGGCGCAGGACGCGCTCCAGGAGCCGCTCCTTCAGGGCCTCCGGTGCGCGCAGTCCGGGCGCCGGACGAGCTCCTTTATCAGAAAGGGCAGCCGCAGAGCCCGCGTGTGCGCGATGTGGCTGCGGGTGGGGAGCGGGCGGCGGGCCCGGGACACCGCGGCCACTGTTCTAGCCCCGCCTGGGCCGCCTGACCGCGGCTCCGCTGCGCCGCAGCCCCGCGCCCCTCTGGCTCCTGTTCCCGGGCGCGGGGAGAAGGCGGCGGGGCGCGCCTGGGCCCGCGCGGGTGCGAACGCGAGGTCTTTCCTGGGTGCTCCCAGGTCGGAGGATTCCCAGGGCGGGGGCCATCAGGGTGGCGAGGAACCGGCAGGGACCAGCCTCCGCTAGGACCGCGCTCGTGGAGCG (SEQ ID NO: 30).
[0327] MAX.chr5.9924: CTTTGGTTTGAAACACTGGAGGTGGCCCAGGGCCGTTTTCCTCAAAGGACTGAGAATCTTGATTTGCCAAGTGCTTGGGGCCTCCGCCCAAGGTGTTTGGGGGCTGCGTGGTGAGCCGAGGCAAAGCCAGGGTACCCCGATCGTCTTCCGGGCGCATCCACCATGCGGCACCGCCCCAGCCACGGCGGCCGCGCGTGGAGACCCGGGGGCTTAACAAAGGGCTCCGCGGGGGCACGGGGGGGCGCGGCCACGTGACAGGCCCGAGCGCGACGTCGCTGTCCAGCCGCGGGGAGGGGCGGCCAGCCCGGGGGGCCGTGGGGCTTCTTGACATAGAACGTCCGGGCCTCGGGCTGGCCGCGCCGGGCCGCGCTCCGCCGGGATGAGAAGTACTTGTCTGGCTCCGCGCTGGAGAAGCCGCACCTCTCATCTCTCCGGCTCTTACTTGAAAAAGCACTTGGAAGAAACTGTGTGTGCGCTGGGAGGGCCGCGGGGTGGGCCGGGGCCGGCTGCGAGGCTGAGGGGGGCCGGCTGGTGGGTGGATGGGGAGGAGGTTGAAGAAACAGCCCCTTTCTGAGTGACAGGACCCCTTTTCAAAGGGCAAACAGAAAAAAAAAAGAAGAAGAAGAAGAAGA (SEQ ID NO: 31).
[0328] MAX.chr6.9522: CGAGGCGAGTTAAATTCCTTTTGCCGGTGCCTGGCTGCGAGGACAAACGTCCGTACTTTCGTTCGGGAGCCACGGGCAGTCCAGGGGCTTGGGTTAGAAGCAACGGCTCTCTTCCAGGGGCTGTGATCCGGGTCGGCCAGGGAGAGCGAGGCCCCGGGGTCCTCTGTGAGGTCCCCAGCGAAGAGACGCAGCTGGGGAAGGCGCCGCCCCCGGGCCCCCTGCGCCACCCTAACCGGGCCTCTCCTTAGCAAAGTTGACAAATTCTTGAGAGTGTCAGCCCAGGGCTGCGCGTGAGGGCGCTGGGACCGGGGAGGAAAGAGCACCTGCCGCGCTCAGCCCGACTTTGAATTTGTTTGTTGTTACCGTTTTTGTTTTTCCTCCCAGTTTCCATAAACGCTTAGTATTTCGAGGCACTTTGCAGGTGTTGGCGCAGGTGATGATGGGCCTCGTTGGACTCTGCCTCCCACGCATCCTTTTGTTTTCTGCGCGCCAGCCTGTCTGACTGTGTCCTGCGGGGACCCCGAGACAGTCCGGGGTCAGGGCGTAGAGACTCATGCTTGCCACTTGACCCATCCGCAACCCGGGGACCCCCTAGCCCGTCGCGGAGCTGGAGTTTGGGCTTCCGGCTCCCAGCTCTCCGCCCTGGATACAGGAAGAGGGCGGGAGAGGTCGCGCACCCGCGCCGCTCGGCGGGGATCGCTCACAGGGGCTCCGGGGCCACCGCGAGCGCGGACTGCGGCTGCTGGCGGGCTCCTTCGTCGTCCAACGCACCCCATCCTCTCCCGCCCCGCAGTGTCCCAGGGAAGGCTTCACTGAAAACAGACGCTCGACGGAAAACTGACTCTGCAGGCCCGAGCTTTCG (SEQ ID NO: 32).
[0329] All publications and patents mentioned in the above specification are incorporated herein by reference in their entirety for all purposes. Various modifications and variations of the compositions, methods, and applications of the described technology will be apparent to those skilled in the art without departing from the scope and spirit of the described technology. Although such technology has been described in connection with specific exemplary embodiments, it should be understood that the invention as claimed should not be unnecessarily limited to such specific embodiments. Indeed, various modifications of the described modes for carrying out the invention that are obvious to those skilled in pharmacology, biochemistry, medicine, or related fields are intended to be within the scope of the following claims.
Claims
1. 1. A method for characterizing a biological sample, comprising: The method comprises determining a methylation profile in at least one differentially methylated region (DMR) of a DNA sample obtained from a subject having or suspected of having gynecological cancer by treating the sample with a reagent that modifies DNA in a methylation-specific manner.
2. 10. The method of claim 1, wherein the methylation profile in the at least one DMR indicates that the subject has or is suspected of having at least one of ovarian cancer (OC), cervical cancer (CC), and endometrial cancer (EC).
3. The at least one DMR is selected from the group consisting of ADAM8, ADHFEl, Aes, AGBL2, AIM1, AK5, ALKBH3, ARAP1, ARHGAP20, ASCL2, BCAT1, BEGAIN, BENDA, BMP6, C12orf68, C13orf18, C14orf169, C14orf169, C18orf18, C1orf61, C20orf195, C4orf31, C5orf52, C6orf147, C7orf51, CD14, CELF2, CHCHD5, CHMP2A, CHST10, CLIC6, CLIP4, COL13A1, COL19A1, COL6A2, COPZ2, CREB3L1, CXCL2, CXXC5, CYTH2, DAB2IP, DGKZ, DLGAp3, DNASE2, DSCAML1, EBF1, EDARADD, EGR2, EIF5A2, ELMO1, ELMOD1, ELOVL4, EMF2, EML6, EPS11, FADS2, FAM109B, FAM126A, FAM174B, FGF18, FKBP11, FLI1, FLOT1, FOXD3, FYN, GAL3ST2, GALR3, GAS7, GATA2, GLT25D2, GNB2, HDAC7, HIC1, HLA-F, HNRNPF, HPDL, HS3ST4, HSPAlA, IDUA, IGSF9B, IL12RB2, IRAK3, IRF7, IRF8, ITPKA, KCNAl, KCNC3, KCNC3, KCNC4, KCNH8, KDM2B, LBX2, LCMT2, LOC100129726, LOC100287216, LOC255130, LOC339290, LOC729678, LPPR3, LRRC41, LRRC8D, LTBP2, LYPLAL1, MAST4, MAX.chr1.2152, HIVEP3, GRAMD1B, MAX.chr11.0394, MAX.chr11.3750, FAT3, SLC16A7, MTUS2, LINCO2323, MAX.chr14.7696, MCTP2, LOC107984974, TRIM80P, MAX.chr19.5552, ZNF433-AS1, ZNF254, MAX.chr19.0548, B3GALTI, MAX.chr2.8918, MAX.chr2.4778, MAX.chr20.3853, MAX.chr20.2903, MAX.chr21.5011, DSCR9,MAX. chr22.5665, MAX. chr3.6408, LINC02028, LINC02084, MAX. chr5.3588, CTD-2532K18.1, HS3ST5, ARHGAP18, GRM4, LINC01004, MAX. chr8.5938, MAX. chr9.4007, MAX. chr9.2025, TRPM3, MED12L, MIAT, MLH1, MLH1, MMP16, MRPS21, MSI1, MT1E, MX1, MYC, MYH10, MYO15B, N4BP 2L1, NBR1, NDRG2, NEGR1, NEU1, NOL3, NR3C1, NR3C1, NRP2, NTN1, NTNG1, PAPL, PAQR9, PDE10A, PDE3B, PDE4 A, PDXK, PER1, PISD, PLEC, PLIN2, PLXND1, PPM1E, PPP1R9A, PPP2R5C, PRDM5, PTP4A3, PYCARD, RAB3C, RAI1 , RARG, RASA3, RPRM, RREB1, S100A6, SAMD5, SBNO2, SDC2, SDK2, SELM, SERP2, SFMBT2, SHF, SHH, SLC16A11, SLC16A5, SLC25A22, SLCO3A1, SMTN, SPDYA, SPINK2, SPOCK2, SPON1, SQSTM1, ST8SIA1, TAF4B, TAF7, TEAD 3, TERC, TIAM1, TLE4, TMEM101, TMEM106A, TRIM9, TRPC3, TSC22D4, TSPAN2, TSPAN5, TTC14, UBB, UBB, UST, 3. The method of claim 1 or claim 2, comprising one or more CpG sites in VAMP5, VIM, VSTM2B, ZBTB7B, ZEB2, ZFP3, ZFP36L2, ZIC2, ZMIZ1, ZNF14, ZNF211, ZNF280B, ZNF302, ZNF382, ZNF480, ZNF483, ZNF491, ZNF569, ZNF610, ZNF702P, ZNF709, ZNF773, ZNF845, ZNF91, CDH4, LRRC34, MAX.chr10.4460, NBPF24, OBSCN, SEPT9, ZNF323, ZNF506 and / or ZNF90.
4. The at least one DMR is selected from the group consisting of ACSF2, AJAP1, ARL10, ARL5C, ASCL4, ATP6V1B1, BARHL1, BEND4, C17orf64, C1QL3, C2orf55, C4orf48, CA3, CDO1, CELF2, CLEC14A, CSDAP1, CYTH2, DLGAP1, DSCR6, EPS8L1, EPS8L1, FAIM2, FGF12, GATA2, HIST1H2BE, IRF4, IRX4, ITGA5, KCNA1, LECT1, LHX1, LOC440925, LPHN1, LINC02767, MAX.chr1.2533, SOX1-OT, MAX. chr13.3357, MAX. chr14.2093, MAX. chr17.2455, MAX. chr18.4390, MAX. chr19.2732, MAX. chr19.4467, PANTR1, MAX. chr2.0490, MAX. chr2.8148, MAX. chr2.3137, RIPOR3, SCRG1, MAX. chr4.4210, HMX1, CTC-359M8.1, MAX. chr5.0931, MAX. chr5.9924, LIN28B, MAX. chr6.9522, TTLL2, RNA5SP243, DLGAP2, MEX3B, MNX1, NEFL, NETO1, PAX2, PDX1, psiTPTE22, RASGEF1A, SAL L3, SALL3, SEZ6L2, SHANK2, SHANK3, SKI, SLC35D3, SORCS3, SORCS3, SOX1, SQSTM1, TBXT, TCERG1L, TERT, T 3. The method of claim 1 or claim 2, comprising one or more CpG sites in NFSF11, TUBB6, ULBP1, VAC14, VWC2, WDR69, ZBTB16, ZNF132, ZSCAN12, ZSCAN23, KRT86, CYP26C1, GYPC, DIDO1, EEF1A2, EMX2OS, GDF7, JSRP1, SMPD5, MDFI, MPZ, and / or VILL.
5. 3. The method of claim 1 or claim 2, wherein the at least one DMR comprises one or more CpG sites in AIM1, FLOT1, GAL3ST2, LRRC41, LYPLAL1, MAX.chr11.3750, PISD, RAI1, ZIC2, ZMIZ1, CDH4, ZNF506, ZNF323, OBSCN, ZNF90, and / or SEPT9, and the subject has or is suspected of having OC.
6. 6. The method of claim 5, wherein the at least one DMR comprises one or more CpG sites in AIM1, FLOT1, GAL3ST2, LYPLAL1, and / or OBSCN, and the subject has or is suspected of having serous OC.
7. 6. The method of claim 5, wherein the at least one DMR comprises one or more CpG sites in LRRC41, PISD, ZIC2, OBSCN, and / or SEPT9, and the subject has or is suspected of having clear cell OC.
8. 6. The method of claim 5, wherein the at least one DMR comprises one or more CpG sites at MAX.chr11.3750, and the subject has or is suspected of having endometrioid OC.
9. 6. The method of claim 5, wherein the at least one DMR comprises one or more CpG sites in RAI1 and / or ZMIZ1, and the subject has or is suspected of having mucinous OC.
10. 10. The method of any one of claims 5 to 9, wherein determining the methylation profile of one or more CpG sites AIM1, FLOT1, GAL3ST2, LRRC41, LYPLAL1, MAX.chr11.3750, PISD, RAI1, ZIC2, ZMIZ1, CDH4, ZNF506, ZNF323, OBSCN, ZNF90, and / or SEPT9 comprises comparing the methylation profile with corresponding regions of a control DNA sample obtained from a subject without OC.
11. 3. The method of claim 1 or claim 2, wherein the at least one DMR comprises one or more CpG sites in AK5, ELMOD1, RABC3, TRPC3, ZNF480, ZNF491, ZNF610, ZNF91, and / or NBPF24, and the subject has or is suspected of having CC.
12. 12. The method of claim 11, wherein the at least one DMR comprises one or more CpG sites in AK5, ELMOD1, TRPC3, and / or ZNF480, and the subject has or is suspected of having cervical adenocarcinoma (CC).
13. 12. The method of claim 11, wherein the at least one DMR comprises one or more CpG sites in ZNF491, ZNF610, and / or ZNF91, and the subject has or is suspected of having squamous cell CC.
14. 14. The method of any one of claims 11 to 13, wherein determining the methylation profile of one or more CpG sites in AK5, ELMOD1, RABC3, TRPC3, ZNF480, ZNF491, ZNF610, ZNF91, and / or NBPF24 comprises comparing the methylation profile with corresponding regions of a control DNA sample obtained from a subject without CC.
15. 3. The method of claim 1 or claim 2, wherein the at least one DMR comprises one or more CpG sites in c18orf18, FKBP11, MLH1, NR3C1, and / or TERC, and the subject has or is suspected of having EC.
16. 16. The method of claim 15, wherein the at least one DMR comprises one or more CpG sites in MLH1 and / or SEPT9, and the subject has or is suspected of having clear cell EC.
17. 16. The method of claim 15, wherein the at least one DMR comprises one or more CpG sites in NR3C1 and the subject has or is suspected of having endometrioid EC.
18. 18. The method of any one of claims 15 to 17, wherein determining the methylation profile of one or more CpG sites in c18orf18, FKBP11, MLH1, NR3C1, and / or TERC comprises comparing the methylation profile with corresponding regions of a control DNA sample obtained from a subject without EC.
19. The method of claim 1 or claim 2, wherein the at least one DMR comprises one or more CpG sites in CDO1 and / or DLGAP1, and the subject has or is suspected of having CC, OC, or EC.
20. 20. The method of claim 19, wherein determining the methylation profile of at least one CpG site in CDO1 and / or DLGAP1 comprises comparing the methylation profile with a corresponding region of a control DNA sample obtained from a subject who does not have CC, OC, or EC.
21. 21. The method of claim 20, further comprising determining the methylation profile of one or more CpG sites in AIM1, FLOT1, GAL3ST2, LRRC41, LYPLAL1, MAX.chr11.3750, PISD, RAI1, ZIC2, and / or ZMIZ1.
22. 21. The method of claim 20, further comprising determining the methylation profile of one or more CpG sites in AK5, ELMOD1, RABC3, TRPC3, ZNF480, ZNF491, ZNF610, and / or ZNF91.
23. 21. The method of claim 20, further comprising determining the methylation profile of one or more CpG sites in c18orf18, FKBP11, MLH1, NR3C1, and / or TERC.
24. 3. The method of claim 1 or claim 2, wherein the at least one DMR comprises one or more CpG sites in NBPF24, and the subject has or is suspected of having CC.
25. 25. The method of claim 24, wherein determining the methylation profile of the one or more CpG sites in NBPF24 comprises comparing the methylation profile with a corresponding region of a control DNA sample obtained from a subject who does not have CC.
26. 3. The method of claim 1 or claim 2, wherein the at least one DMR comprises one or more CpG sites in CDH4, NBPF24, MAX.chr10.4460, ZNF506, ZNF323, OBSCN, ZNF90, LRRC34, SFMBT2, LINC02323, CYTH2, LRRC8D, LYPLAL1, LRRC41, and / or SEPT9, and the subject has or is suspected of having EC.
27. 27. The method of claim 26, wherein determining the methylation profile of the one or more CpG sites in CDH4, NBPF24, MAX.chr10.4460, ZNF506, ZNF323, OBSCN, ZNF90, LRRC34, SFMBT2, LINC02323, CYTH2, LRRC8D, LYPLAL1, LRRC41, and / or SEPT9 comprises comparing the methylation profile to corresponding regions of a control DNA sample obtained from a subject without EC.
28. 3. The method of claim 1 or claim 2, wherein the at least one DMR comprises one or more CpG sites in CDH4, ZNF506, ZNF323, OBSCN, ZNF90, SFMBT2, LINC02323, CYTH2, LRRC8D, LYPLAL1, LRRC41, and / or SEPT9, and the subject has or is suspected of having OC.
29. 29. The method of claim 28, wherein determining the methylation profile of the one or more CpG sites in CDH4, ZNF506, ZNF323, OBSCN, ZNF90, SFMBT2, LINC02323, CYTH2, LRRC8D, LYPLAL1, LRRC41, and / or SEPT9 comprises comparing the methylation profile with corresponding regions of a control DNA sample obtained from a subject without OC.
30. 3. The method of claim 1 or claim 2, wherein the at least one DMR comprises one or more CpG sites in KRT86, EMX2OS, JSRP1, DIDO1, MPZ, VILL, SMPD5, GDF7, MDFI, c17orf64, GATA2, SQSTM1, and / or EEF1A2, and the subject has or is suspected of having CC, OC, or EC.
31. 31. The method of claim 30, wherein determining the methylation profile of the one or more CpG sites in KRT86, EMX2OS, JSRP1, DIDO1, MPZ, VILL, SMPD5, GDF7, MDFI, c17orf64, GATA2, SQSTM1, and / or EEF1A2 comprises comparing the methylation profile with corresponding regions of a control DNA sample obtained from a subject who does not have CC, OC, or EC.
32. 32. The method of any one of claims 1-31, wherein the at least one DMR is associated with an area under the ROC curve (AUC) of 0.8 or greater, and the ROC curve distinguishes between subjects having or suspected of having OC, CC, or EC and control samples.
33. 33. The method of any one of claims 1 to 32, wherein the biological sample is selected from a tissue sample, a blood sample, a plasma sample, a serum sample, a whole blood sample, a secretion sample, an organ secretion sample, a cerebrospinal fluid (CSF) sample, a saliva sample, a urine sample, and a stool sample.
34. 34. The method of claim 33, wherein the tissue sample is a gynecological tissue sample.
35. 35. The method of claim 34, wherein the gynecological tissue sample comprises one or more of vaginal tissue, vaginal cells, cervical tissue, cervical cells, endometrial tissue, endometrial cells, ovarian tissue, and ovarian cells.
36. 36. The method of claim 35, wherein the tissue sample is an ovarian tissue sample, an endometrial tissue sample, or a cervical tissue sample.
37. 36. The method of claim 35, wherein the secretion sample is a gynecological secretion sample.
38. The method of any one of claims 1 to 37, wherein the subject is a human.
39. 39. The method of any one of claims 1 to 38, wherein the biological sample is obtained from the subject, and the method further comprises extracting the DNA sample from the biological sample.
40. 40. The method of any one of claims 1 to 39, wherein the biological sample is collected with a collection device having an absorbent member capable of collecting the biological sample upon contact.
41. 41. The method of claim 40, wherein the absorbent member is a sponge configured for insertion into the orifice.
42. 41. The method of claim 40, wherein the collection device is selected from a tampon, a lavage that releases liquid into the vagina and recollects the fluid, a cervical brush, a Fournier cervical self-sampling device, and a swab.
43. 43. The method of any one of claims 1 to 42, wherein the reagent that modifies DNA in a methylation-specific manner is a borane reducing agent.
44. 43. The method of any one of claims 1 to 42, wherein the reagent that modifies DNA in a methylation-specific manner comprises one or more of a methylation-sensitive restriction enzyme, a methylation-dependent restriction enzyme, and a bisulfite reagent.
45. 45. The method of any one of claims 1 to 44, wherein determining the methylation profile of at least one DMR comprises amplifying at least a portion of the DMR using a set of primers.
46. 46. The method of any one of claims 1 to 45, wherein determining the methylation profile of at least one DMR comprises performing at least one of methylation-specific PCR, quantitative methylation-specific PCR, methylation-specific DNA restriction enzyme analysis, quantitative bisulfite pyrosequencing, flap endonuclease assay, PCR-flap assay, and bisulfite genomic sequencing PCR.
47. 47. The method of any one of claims 1 to 46, wherein determining the methylation profile of at least one DMR comprises determining the presence or absence of methylation at CpG sites.
48. 1. A method for identifying gynecological cancer, comprising: determining a methylation profile in at least one differentially methylated region (DMR) of a DNA sample obtained from a subject having or suspected of having gynecological cancer by treating said sample with a reagent that modifies DNA in a methylation-specific manner; the at least one DMR comprises one or more CpG sites in AIM1, FLOT1, GAL3ST2, LRRC41, LYPLAL1, MAX.chr11.3750, PISD, RAI1, ZIC2, and / or ZMIZ1; The method, wherein the methylation profile indicates that the subject has ovarian cancer.
49. 1. A method for identifying gynecological cancer, comprising: determining a methylation profile in at least one differentially methylated region (DMR) of a DNA sample obtained from a subject having or suspected of having gynecological cancer by treating said sample with a reagent that modifies DNA in a methylation-specific manner; the at least one DMR comprises one or more CpG sites in AK5, ELMOD1, RABC3, TRPC3, ZNF480, ZNF491, ZNF610, and / or ZNF91; The method, wherein the methylation profile indicates that the subject has cervical cancer.
50. 1. A method for identifying gynecological cancer, comprising: determining a methylation profile in at least one differentially methylated region (DMR) of a DNA sample obtained from a subject having or suspected of having gynecological cancer by treating said sample with a reagent that modifies DNA in a methylation-specific manner; The method, wherein the at least one DMR comprises one or more CpG sites in c18orf18, FKBP11, MLH1, NR3C1, and / or TERC, and the methylation profile indicates that the subject has endometrial cancer.
51. 1. A method for identifying gynecological cancer, comprising: determining a methylation profile in at least one differentially methylated region (DMR) of a DNA sample obtained from a subject having or suspected of having gynecological cancer by treating said sample with a reagent that modifies DNA in a methylation-specific manner; the at least one DMR comprises one or more CpG sites in CDO1 and / or DLGAP1; The method, wherein the methylation profile indicates that the subject has ovarian cancer, cervical cancer, or endometrial cancer.
52. 1. A method for identifying gynecological cancer, comprising: determining a methylation profile in at least one differentially methylated region (DMR) of a DNA sample obtained from a subject having or suspected of having gynecological cancer by treating said sample with a reagent that modifies DNA in a methylation-specific manner; the at least one DMR comprises one or more CpG sites in NBPF24; The method, wherein the methylation profile indicates that the subject has cervical cancer.
53. 1. A method for identifying gynecological cancer, comprising: determining a methylation profile in at least one differentially methylated region (DMR) of a DNA sample obtained from a subject having or suspected of having gynecological cancer by treating said sample with a reagent that modifies DNA in a methylation-specific manner; wherein the at least one DMR comprises one or more CpG sites in CDH4, NBPF24, MAX.chr10.4460, ZNF506, ZNF323, OBSCN, ZNF90, LRRC34, SFMBT2, LINC02323, CYTH2, LRRC8D, LYPLAL1, LRRC41, and / or SEPT9, and the methylation profile indicates that the subject has endometrial cancer.
54. 1. A method for identifying gynecological cancer, comprising: determining a methylation profile in at least one differentially methylated region (DMR) of a DNA sample obtained from a subject having or suspected of having gynecological cancer by treating said sample with a reagent that modifies DNA in a methylation-specific manner; wherein the at least one DMR comprises one or more CpG sites in CDH4, ZNF506, ZNF323, OBSCN, ZNF90, SFMBT2, LINC02323, CYTH2, LRRC8D, LYPLAL1, LRRC41, and / or SEPT9, and the methylation profile indicates that the subject has ovarian cancer.
55. 1. A method for identifying gynecological cancer, comprising: determining a methylation profile in at least one differentially methylated region (DMR) of a DNA sample obtained from a subject having or suspected of having gynecological cancer by treating said sample with a reagent that modifies DNA in a methylation-specific manner; The method, wherein the at least one DMR comprises one or more CpG sites in KRT86, EMX2OS, JSRP1, DIDO1, MPZ, VILL, SMPD5, GDF7, MDFI, c17orf64, GATA2, SQSTM1, and / or EEF1A2, and the methylation profile indicates that the subject has ovarian, cervical, or endometrial cancer.
56. 56. The method of any one of claims 48-55, further comprising treating the subject with an anti-cancer therapy.