Methods and compositions for selective enrichment of methylation signal
By employing allele-specific and strand-specific methyl enrichment baits, the methods provide a robust and sensitive detection of methylation patterns in DNA, addressing the need for improved sensitivity and specificity in cancer diagnostics and treatment monitoring.
Patent Information
- Application Number
- PCT/US2024/059877
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-13
- Filing Date
- 2024-12-12
- Publication Date
- 2025-06-19
Smart Images

Figure US2024059877_19062025_PF_FP_ABST
Abstract
Description
METHODS AND COMPOSITIONS FOR SELECTIVE ENRICHMENT OF METHYLATION SIGNALCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 609,810, filed December 13, 2023, which is hereby incorporated by reference in its entirety.FIELD
[0002] Provided herein are compositions, such as bait sets, for selectively enriching nucleic acid molecules corresponding to one or more hypermethylation signals or one or more hypomethylation signals. Also described are methods of using such compositions, as well as methods for designing such compositions.BACKGROUND
[0003] Many diseases and conditions are products of both the genome and the epigenome. For instance, epigenetic modifications are known to cooperate with genetic alterations to drive development of cancer. DNA methylation is a type of epigenetic modification characterized by the addition of a methyl group to the C-5 position of cytosine, one of the four nucleotide bases in DNA. 5-methyl-cytosine (5mC) and 5-hydroxymethyl-cytosine (5hmC) are the two major types of DNA methylation that occur in the mammalian genome, commonly seen in CpG contexts (cytosine nucleotide followed by a guanine nucleotide in the linear sequence of bases along its 5' to 3' direction). Many cancers have profound epigenetic dysregulation that give rise to aberrant DNA methylation patterns, including hypermethylation and hypomethylation, which are distinct from healthy samples. These features can be useful in developing cancer diagnostic assays based, e.g., on next generation sequencing (NGS) for methylation analysis (Methyl-seq). Methyl-seq relies on a chemical process (e.g., bisulfite, enzymatic, pull-down, etc.) that converts unmethylated cytosines to thymine, while leaving methylated cytosines intact. Methyl-seq is employed during library construction for NGS. As such, the DNA methylation status in the original DNA molecules in the input material can be inferred and compared to identify healthy versus cancer samples to determine or quantify the amount of residual disease molecules in a sample or monitor treatment response.
[0004] There remains a need for improved methods and systems that provide robust and sensitive detection of methylation patterns in DNA, with low background signal and increased signal-to-background ratio.BRIEF SUMMARY OF THE INVENTION
[0005] Disclosed herein are methods and systems for selectively capturing nucleic acid molecules from a sample using allele- specific methyl enrichment baits and / or strand- specific methyl enrichment baits as described herein. The bait sets described herein are designed to hybridize to nucleic acids that have been converted and amplified from methylated nucleic acids that comprise a methylation pattern of interest — either a hypomethylation pattern of interest or a hypermethylation pattern of interest. An allele- specific bait as described herein is double stranded and is either hypomethylation- specific or hypermethylation- specific. A strand-specific bait as described herein is single- stranded and is either hypomethylationspecific or hypermethylation-specific. Both the allele- specific baits and strand- specific baits described herein offer improved enrichment of methylation patterns compared to conventional bait designs. The strand- specific baits described herein may be further designed to enable further improved enrichment beyond the capabilities of the allele- specific baits by, for example, avoiding off-target guanine / thymine (G:T) mismatches and / or favoring off- target adenine / cytosine (A:C) mismatches.
[0006] In one aspect, provided herein is a method of selectively capturing nucleic acid molecules from a sample from a subject, comprising: converting unmethylated cytosine (C) in one or more of the nucleic acid molecules to uracil (U), or converting methylated C in one or more of the nucleic acid molecules to U, to form converted nucleic acid molecules; amplifying the converted nucleic acid molecules into a set of amplicons; hybridizing a set of single- stranded nucleic acid baits to the set of amplicons to form a plurality of hybrids, wherein at least a portion of the hybrids in the plurality of hybrids comprise a bait from the set of single- stranded nucleic acid baits and a nucleic acid molecule comprising a nucleic acid pattern of interest hybridized to each other without mismatches, wherein the nucleic acid pattern of interest corresponds to a methylation pattern of interest at a genomic locus in the nucleic acid molecules prior to the converting; and separating the plurality of hybrids from nucleic acid molecules that are not hybridized to a bait. In some embodiments of this aspect, the methylation pattern of interest corresponds to a hypermethylation signal indicative of adisease or condition. In some embodiments of this aspect, the methylation pattern of interest corresponds to a hypomethylation signal indicative of a disease or condition.
[0007] In one aspect, provided herein is a method of selectively capturing nucleic acid molecules from a sample from a subject, comprising converting unmethylated cytosine (C) in one or more of the nucleic acid molecules to uracil (U), or converting methylated C in one or more of the nucleic acid molecules to U, to form converted nucleic acid molecules; amplifying the converted nucleic acid molecules into a set of amplicons; hybridizing a set of nucleic acid baits to the set of amplicons to form a plurality of hybrids, wherein the set of nucleic acid baits selectively hybridizes to a nucleic acid pattern of interest, wherein the nucleic acid pattern of interest corresponds to a methylation pattern of interest at a genomic locus in the nucleic acid molecules prior to the converting, wherein the methylation pattern of interest corresponds to a hypermethylation signal indicative of a disease or condition; and separating the plurality of hybrids from nucleic acid molecules that are not hybridized to a bait. In some embodiments of this aspect, the set of nucleic acid baits does not comprise a nucleic acid bait that selectively hybridizes to a nucleic acid pattern that corresponds to a hypomethylation signal at the genomic locus.
[0008] In one aspect, provided herein is a method of selectively capturing nucleic acid molecules from a sample from a subject, comprising: converting unmethylated cytosine (C) in one or more of the nucleic acid molecules to uracil (U), or converting methylated C in one or more of the nucleic acid molecules to U, to form converted nucleic acid molecules; amplifying the converted nucleic acid molecules into a set of amplicons; hybridizing a set of nucleic acid baits to the set of amplicons to form a plurality of hybrids, wherein the set of nucleic acid baits selectively hybridizes to a nucleic acid pattern of interest, wherein the nucleic acid pattern of interest corresponds to a methylation pattern of interest at a genomic locus in the nucleic acid molecules prior to the converting, wherein the methylation pattern of interest is associated with a hypomethylation signal indicative of a disease or condition; and separating the plurality of hybrids from nucleic acid molecules that are not hybridized to a bait. In some embodiments of this aspect, the set of nucleic acid baits does not comprise a nucleic acid bait that selectively hybridizes to a nucleic acid pattern that corresponds to a hypermethylation signal at the genomic locus.
[0009] In any of the preceding embodiments, the method may include converting a plurality of methylation patterns of interest at a plurality of genomic loci in the nucleic acid moleculesfrom the sample from the subject into a plurality of nucleic acid patterns of interest of the plurality of genomic loci, and hybridizing a plurality of sets of nucleic acid baits to the set of amplicons, wherein each set of nucleic acid baits of the plurality of sets of nucleic acid baits selectively hybridizes to a nucleic acid pattern of interest of the plurality of nucleic acid patterns of interest. In some embodiments, the plurality of methylation patterns of interest comprises a first methylation pattern of interest and a second methylation pattern of interest, wherein the second methylation pattern of interest is associated with a hypermethylation signal indicative of a disease or condition. In some embodiments, the plurality of methylation patterns of interest comprises a first methylation pattern of interest and a second methylation pattern of interest, wherein the second methylation pattern of interest is associated with a hypomethylation signal indicative of a disease or condition. In some embodiments, the first methylation pattern of interest and the second methylation pattern of interest are associated with hypomethylation or hypermethylation signals indicative of different diseases or conditions. In some embodiments, the first methylation pattern of interest and the second methylation pattern of interest are associated with hypomethylation or hypermethylation signals indicative of the same disease or condition.
[0010] In some embodiments, the nucleic acid baits are double- stranded. In some embodiments, the nucleic acid baits are single-stranded.
[0011] In any of the preceding embodiments, the hybrids in the plurality of hybrids may not comprise guanine / thymine (G:T) mismatches.
[0012] In some embodiments, a second portion of hybrids in the plurality of hybrids comprises a single-stranded nucleic acid bait from the set of single- stranded nucleic acid baits and a nucleic acid molecule that does not comprise the nucleic acid pattern of interest, wherein the hybrids in the second portion of hybrids comprise one or more adenine / cytosine (A:C) mismatches. In some embodiments, the method further includes separating the portion of hybrids from the second portion of hybrids.
[0013] In any of the preceding embodiments, the method may further include attaching the hybrids to a substrate and washing the substrate, wherein nucleic acid molecules that are not hybridized to a bait are depleted. In some embodiments, the method may include attaching the hybrids to a substrate and washing the substrate, wherein nucleic acid molecules in the second portion of hybrids are depleted. In some embodiments, the substrate is washed at a temperature lower than a melting temperature (Tm) of the portion of hybrids in the pluralityof hybrids comprising a single- stranded nucleic acid bait from the set of single-stranded nucleic acid baits and a nucleic acid molecule comprising the nucleic acid pattern of interest.
[0014] In any of the preceding embodiments, the converting may include enzymatic conversion and / or bisulfite conversion. In some embodiments, the converting is enzymatic conversion. In some embodiments, the converting is bisulfite conversion. In any of the preceding embodiments, the converting may include cytosine (C) to uracil (U) conversion of unmethylated C residues. In any of the preceding embodiments, the converting may include C to U conversion of methylated C residues.
[0015] In any of the preceding embodiments, the amplifying may include uracil (U) to thymine (T) conversion.
[0016] In any of the preceding embodiments, the methylation pattern of interest may include one or more CpG sites. In any of the preceding embodiments, the methylation pattern of interest may lack one or more methylated CpG sites as compared to a different methylation pattern at the same genomic locus that is not indicative of the disease or condition. In any of the preceding embodiments, the methylation pattern of interest may include one or more 3mC, 4mC, 5mC, and / or 5hmC, sites. In some embodiments, the disease or condition comprises a tissue- specific marker, a cancer, an auto-immune disease, an infectious disease, and / or a status of a transplanted tissue or organ. In some embodiments, the disease is cancer.
[0017] In any of the preceding embodiments, the sample may include cell-free DNA (cfDNA), circulating tumor DNA (ctDNA) or combination of both. In any of the preceding embodiments, the bait may be biotinylated. In some embodiments, the separating comprises contacting the baits with streptavidin-coated beads.
[0018] In any of the preceding embodiments, the converting may include TET-assisted bisulfite treatment, TET-assisted pyridine borane treatment, oxidative bisulfite treatment, APOBEC treatment, and / or other DNA deaminase treatment. In any of the preceding embodiments, the method may include fragmenting the nucleic acid molecules from the sample. In any of the preceding embodiments, the amplifying may include a polymerase chain reaction (PCR) amplification technique, a non-PCR amplification technique, or an isothermal amplification technique.
[0019] In any of the preceding embodiments, the subject may be a bacterial, plant, fungal, or animal subject. In some embodiments, the subject is an animal subject. In someembodiments, the subject is a mammalian subject. In some embodiments, the subject is a human subject.
[0020] In any of the preceding embodiments, the method may include sequencing the nucleic acid molecules in the plurality of hybrids. In some embodiments, the method may include sequencing a plurality of nucleic acid molecules from the portion of hybrids after the nucleic acid molecules in the second portion of hybrids are depleted. In some embodiments, the method may include determining a hypermethylation score and / or a hypomethylation score.
[0021] In any of the preceding embodiments, the method may include diagnosing the subject as having a disease or a condition. In some embodiments, the disease is cancer. In any of the preceding embodiments, the subject may be suspected of having or is determined to have a cancer. In some embodiments, the cancer may include a B cell cancer (multiple myeloma), a melanoma, breast cancer, lung cancer, bronchus cancer, colorectal cancer, prostate cancer, pancreatic cancer, stomach cancer, ovarian cancer, urinary bladder cancer, brain cancer, central nervous system cancer, peripheral nervous system cancer, esophageal cancer, cervical cancer, uterine cancer, endometrial cancer, cancer of an oral cavity, cancer of a pharynx, liver cancer, kidney cancer, testicular cancer, biliary tract cancer, small bowel cancer, appendix cancer, salivary gland cancer, thyroid gland cancer, adrenal gland cancer, osteosarcoma, chondrosarcoma, a cancer of hematological tissue, an adenocarcinoma, an inflammatory myofibroblastic tumor, a gastrointestinal stromal tumor (GIST), colon cancer, multiple myeloma (MM), myelodysplastic syndrome (MDS), myeloproliferative disorder (MPD), acute lymphocytic leukemia (ALL), acute myelocytic leukemia (AML), chronic myelocytic leukemia (CML), chronic lymphocytic leukemia (CLL), polycythemia Vera, Hodgkin lymphoma, non-Hodgkin lymphoma (NHL), soft-tissue sarcoma, fibrosarcoma, myxosarcoma, liposarcoma, osteogenic sarcoma, chordoma, angiosarcoma, endotheliosarcoma, lymphangiosarcoma, lymphangioendotheliosarcoma, synovioma, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinomas, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatoma, bile duct carcinoma, choriocarcinoma, seminoma, embryonal carcinoma, Wilms' tumor, bladder carcinoma, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pinealoma, hemangioblastoma, acoustic neuroma, oligodendroglioma,meningioma, neuroblastoma, retinoblastoma, follicular lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, hepatocellular carcinoma, thyroid cancer, gastric cancer, head and neck cancer, small cell cancer, essential thrombocythemia, agnogenic myeloid metaplasia, hypereosinophilic syndrome, systemic mastocytosis, familiar hypereosinophilia, chronic eosinophilic leukemia, neuroendocrine cancers, or a carcinoid tumor.
[0022] In some embodiments, the cancer may include acute lymphoblastic leukemia (Philadelphia chromosome positive), acute lymphoblastic leukemia (precursor B-cell), acute myeloid leukemia (FLT3+), acute myeloid leukemia (with an IDH2 mutation), anaplastic large cell lymphoma, basal cell carcinoma, B-cell chronic lymphocytic leukemia, bladder cancer, breast cancer (HER2 overexpressed / amplified), breast cancer (HER2+), breast cancer (HR+, HER2-), cervical cancer, cholangiocarcinoma, chronic lymphocytic leukemia, chronic lymphocytic leukemia (with 17p deletion), chronic myelogenous leukemia, chronic myelogenous leukemia (Philadelphia chromosome positive), classical Hodgkin lymphoma, colorectal cancer, colorectal cancer (dMMR / MSI-H), colorectal cancer (KRAS wild type), cryopyrin-associated periodic syndrome, a cutaneous T-cell lymphoma, dermatofibrosarcoma protuberans, a diffuse large B-cell lymphoma, fallopian tube cancer, a follicular B-cell nonHodgkin lymphoma, a follicular lymphoma, gastric cancer, gastric cancer (HER2+), gastroesophageal junction (GEJ) adenocarcinoma, a gastrointestinal stromal tumor, a gastrointestinal stromal tumor (KIT+), a giant cell tumor of the bone, a glioblastoma, granulomatosis with polyangiitis, a head and neck squamous cell carcinoma, a hepatocellular carcinoma, Hodgkin lymphoma, juvenile idiopathic arthritis, lupus erythematosus, a mantle cell lymphoma, medullary thyroid cancer, melanoma, a melanoma with a BRAF V600 mutation, a melanoma with a BRAF V600E or V600K mutation, Merkel cell carcinoma, multicentric Castleman's disease, multiple hematologic malignancies including Philadelphia chromosome-positive ALL and CML, multiple myeloma, myelofibrosis, a non-Hodgkin’s lymphoma, a nonresectable subependymal giant cell astrocytoma associated with tuberous sclerosis, a non-small cell lung cancer, a non-small cell lung cancer (ALK+), a non-small cell lung cancer (PD-L1+), a non-small cell lung cancer (with ALK fusion or ROS1 gene alteration), a non-small cell lung cancer (with BRAF V600E mutation), a non-small cell lung cancer (with an EGFR exon 19 deletion or exon 21 substitution (L858R) mutations), a non- small cell lung cancer (with an EGFR T790M mutation), ovarian cancer, ovarian cancer (with a BRCA mutation), pancreatic cancer, a pancreatic, gastrointestinal, or lung originneuroendocrine tumor, a pediatric neuroblastoma, a peripheral T-cell lymphoma, peritoneal cancer, prostate cancer, a renal cell carcinoma, rheumatoid arthritis, a small lymphocytic lymphoma, a soft tissue sarcoma, a solid tumor (MSI-H / dMMR), a squamous cell cancer of the head and neck, a squamous non-small cell lung cancer, thyroid cancer, a thyroid carcinoma, urothelial cancer, a urothelial carcinoma, or Waldenstrom's macroglobulinemia.
[0023] In some embodiments, the cancer is glioblastoma, breast cancer, bladder cancer, cervical cancer, colorectal cancer, hepatocellular carcinoma, lung cancer, ovarian cancer, esophageal cancer, pancreatic cancer, stomach cancer or prostate cancer. In some embodiments, the method further incudes treating the subject with an anti-cancer therapy. In some embodiments, the anti-cancer therapy comprises a targeted anti-cancer therapy. In some embodiments, the targeted anti-cancer therapy comprises abemaciclib (Verzenio), abiraterone acetate (Zytiga), acalabrutinib (Calquence), ado-trastuzumab emtansine (Kadcyla), afatinib dimaleate (Gilotrif), aldesleukin (Proleukin), alectinib (Alecensa), alemtuzumab (Campath), alitretinoin (Panretin), alpelisib (Piqray), amivantamab-vmjw (Rybrevant), anastrozole (Arimidex), apalutamide (Erleada), asciminib hydrochloride (Scemblix), atezolizumab (Tecentriq), avapritinib (Ayvakit), avelumab (Bavencio), axicabtagene ciloleucel (Yescarta), axitinib (Inlyta), belantamab mafodotin-blmf (Blenrep), belimumab (Benlysta), belinostat (Beleodaq), belzutifan (Welireg), bevacizumab (Avastin), bexarotene (Targretin), binimetinib (Mektovi), blinatumomab (Blincyto), bortezomib (Velcade), bosutinib (Bosulif), brentuximab vedotin (Adcetris), brexucabtagene autoleucel (Tecartus), brigatinib (Alunbrig), cabazitaxel (Jevtana), cabozantinib (Cabometyx), cabozantinib (Cabometyx, Cometriq), canakinumab (Haris), capmatinib hydrochloride (Tabrecta), carfilzomib (Kyprolis), cemiplimab-rwlc (Libtayo), ceritinib (LDK378 / Zykadia), cetuximab (Erbitux), cobimetinib (Cotellic), copanlisib hydrochloride (Aliqopa), crizotinib (Xalkori), dabrafenib (Tafinlar), dacomitinib (Vizimpro), daratumumab (Darzalex), daratumumab and hyaluronidase-fihj (Darzalex Faspro), darolutamide (Nubeqa), dasatinib (Sprycel), denileukin diftitox (Ontak), denosumab (Xgeva), dinutuximab (Unituxin), dostarlimab-gxly (Jemperli), durvalumab (Imfinzi), duvelisib (Copiktra), elotuzumab (Empliciti), enasidenib mesylate (Idhifa), encorafenib (Braftovi), enfortumab vedotin-ejfv (Padcev), entrectinib (Rozlytrek), enzalutamide (Xtandi), erdafitinib (Balversa), erlotinib (Tarceva), everolimus (Afinitor), exemestane (Aromasin), fam-trastuzumab deruxtecan-nxki (Enhertu), fedratinib hydrochloride (Inrebic), fulvestrant (Faslodex), gefitinib (Iressa), gemtuzumab ozogamicin (Mylotarg), gilteritinib (Xospata),glasdegib maleate (Daurismo), hyaluronidase-zzxf (Phesgo), ibrutinib (Imbruvica), ibritumomab tiuxetan (Zevalin), idecabtagene vicleucel (Abecma), idelalisib (Zydelig), imatinib mesylate (Gleevec), infigratinib phosphate (Truseltiq), inotuzumab ozogamicin (Besponsa), iobenguane 1131 (Azedra), ipilimumab (Yervoy), isatuximab-irfc (Sarclisa), ivosidenib (Tibsovo), ixazomib citrate (Ninlaro), lanreotide acetate (Somatuline Depot), lapatinib (Tykerb), larotrectinib sulfate (Vitrakvi), lenvatinib mesylate (Lenvima), letrozole (Femara), lisocabtagene maraleucel (Breyanzi), loncastuximab tesirine-lpyl (Zynlonta), lorlatinib (Lorbrena), lutetium Lu 177-dotatate (Lutathera), margetuximab-cmkb (Margenza), midostaurin (Rydapt), mobocertinib succinate (Exkivity), mogamulizumab-kpkc (Poteligeo), moxetumomab pasudotox-tdfk (Lumoxiti), naxitamab-gqgk (Danyelza), necitumumab (Portrazza), neratinib maleate (Nerlynx), nilotinib (Tasigna), niraparib tosylate monohydrate (Zejula), nivolumab (Opdivo), obinutuzumab (Gazyva), ofatumumab (Arzerra), olaparib (Lynparza), olaratumab (Lartruvo), osimertinib (Tagrisso), palbociclib (Ibrance), panitumumab (Vectibix), panobinostat (Farydak), pazopanib (Votrient), pembrolizumab (Keytruda), pemigatinib (Pemazyre), pertuzumab (Perjeta), pexidartinib hydrochloride (Turalio), polatuzumab vedotin-piiq (Polivy), ponatinib hydrochloride (Iclusig), pralatrexate (Folotyn), pralsetinib (Gavreto), radium 223 dichloride (Xofigo), ramucirumab (Cyramza), regorafenib (Stivarga), ribociclib (Kisqali), ripretinib (Qinlock), rituximab (Rituxan), rituximab and hyaluronidase human (Rituxan Hycela), romidepsin (Istodax), rucaparib camsylate (Rubraca), ruxolitinib phosphate (Jakafi), sacituzumab govitecan-hziy (Trodelvy), seliciclib, selinexor (Xpovio), selpercatinib (Retevmo), selumetinib sulfate (Koselugo), siltuximab (Sylvant), sipuleucel-T (Provenge), sirolimus protein-bound particles (Fyarro), sonidegib (Odomzo), sorafenib (Nexavar), sotorasib (Lumakras), sunitinib (Sutent), tafasitamab-cxix (Monjuvi), tagraxofusp-erzs (Elzonris), talazoparib tosylate (Talzenna), tamoxifen (Nolvadex), tazemetostat hydrobromide (Tazverik), tebentafusp-tebn (Kimmtrak), temsirolimus (Torisel), tepotinib hydrochloride (Tepmetko), tisagenlecleucel (Kymriah), tisotumab vedotin-tftv (Tivdak), tocilizumab (Actemra), tofacitinib (Xeljanz), tositumomab (Bexxar), trametinib (Mekinist), trastuzumab (Herceptin), tretinoin (Vesanoid), tivozanib hydrochloride (Fotivda), toremifene (Fareston), tucatinib (Tukysa), umbralisib tosylate (Ukoniq), vandetanib (Caprelsa), vemurafenib (Zelboraf), venetoclax (Venclexta), vismodegib (Erivedge), vorinostat (Zolinza), zanubrutinib (Brukinsa), ziv-aflibercept (Zaltrap), or any combination thereof.
[0024] In any of the preceding embodiments, the method may further include obtaining the sample from the subject. In any of the preceding embodiments, the sample may include a tissue biopsy sample, a liquid biopsy sample, or a normal control. In some embodiments, the sample is a liquid biopsy sample and comprises blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva. In some embodiments, the sample is a liquid biopsy sample and comprises circulating tumor cells (CTCs). In some embodiments, the sample is a liquid biopsy sample and comprises cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), or any combination thereof.
[0025] In any of the preceding embodiments, the sample may include a mixture of tumor nucleic acid molecules and non-tumor nucleic acid molecules. In some embodiments, the tumor nucleic acid molecules are derived from a tumor portion of a heterogeneous tissue biopsy sample, and the non-tumor nucleic acid molecules are derived from a normal portion of the heterogeneous tissue biopsy sample. In some embodiments, the sample comprises a liquid biopsy sample, and wherein the tumor nucleic acid molecules are derived from a circulating tumor DNA (ctDNA) fraction of the liquid biopsy sample, and the non-tumor nucleic acid molecules are derived from a non-tumor, cell-free DNA (cfDNA) fraction of the liquid biopsy sample.
[0026] In any of the preceding embodiments, the method may further include ligating one or more adapters onto one or more nucleic acid molecules prior to or after converting the methylation pattern but prior to amplifying the converted nucleic acid molecules, wherein the one or more adapters comprise amplification primers, flow cell adaptor sequences, substrate adapter sequences, or sample index sequences.
[0027] In any of the preceding embodiments, the amplifying may be performed by polymerase chain reaction (PCR). In some embodiments, the sequencing comprises use of a massively parallel sequencing (MPS) technique, whole genome sequencing, whole exome sequencing, targeted sequencing, direct sequencing, or Sanger sequencing technique. In some embodiments, the sequencing comprises use of a MPS technique, and the MPS technique comprises next generation sequencing (NGS). In some embodiments, the sequencing comprises use of a next generation sequencer. In some embodiments, a plurality of sequence reads are produced, and wherein one or more of the plurality of sequence reads overlap one or more loci within one or more subgenomic intervals in the sample. In some embodiments, the one or more loci comprise at least 10 loci each comprising more than two CpG sites.
[0028] In some embodiments, the one or more loci comprise ABL1, ACVR1B, AKT1, AKT2, AKT3, ALK, ALOX12B, AMER1, APC, AR, ARAF, ARFRP1, ARID1A, ASXL1, ATM, ATR, ATRX, AURKA, AURKB, AXIN1, AXL, BAP1, BARD1, BCL2, BCL2L1, BCL2L2, BCL6, BCOR, BCORL1, BCR, BRAF, BRCA1, BRCA2, BRD4, BRIP1, BTG1, BTG2, BTK, CALR, CARD11, CASP8, CBFB, CBL, CCND1, CCND2, CCND3, CCNE1, CD22, CD274, CD70, CD74, CD79A, CD79B, CDC73, CDH1, CDK12, CDK4, CDK6, CDK8, CDKN1A, CDKN1B, CDKN2A, CDKN2B, CDKN2C, CEBPA, CHEK1, CHEK2, CIC, CREBBP, CRKL, CSF1R, CSF3R, CTCF, CTNNA1, CTNNB1, CUL3, CUL4A, CXCR4, CYP17A1, DAXX, DDR1, DDR2, DIS3, DNMT3A, DOT1L, EED, EGFR, EMSY (Cllorf30), EP300, EPHA3, EPHB1, EPHB4, ERBB2, ERBB3, ERBB4, ERCC4, ERG, ERRFI1, ESRI, ETV4, ETV5, ETV6, EWSR1, EZH2, EZR, FAM46C, FANCA, FANCC, FANCG, FANCL, FAS, FBXW7, FGF10, FGF12, FGF14, FGF19, FGF23, FGF3, FGF4, FGF6, FGFR1, FGFR2, FGFR3, FGFR4, FH, FLCN, FLT1, FLT3, FOXL2, FUBP1, GABRA6, GATA3, GATA4, GATA6, GID4 (C17orf39), GNA11, GNA13, GNAQ, GNAS, GRM3, GSK3B, H3F3A, HDAC1, HGF, HNF1A, HRAS, HSD3B1, ID3, IDH1, IDH2, IGF1R, IKBKE, IKZF1, INPP4B, IRF2, IRF4, IRS2, JAK1, JAK2, JAK3, JUN, KDM5A, KDM5C, KDM6A, KDR, KEAP1, KEL, KIT, KLHL6, KMT2A (MLL), KMT2D (MLL2), KRAS, LTK, LYN, MAF, MAP2K1, MAP2K2, MAP2K4, MAP3K1, MAP3K13, MAPK1, MCL1, MDM2, MDM4, MED12, MEF2B, MEN1, MERTK, MET, MITF, MKNK1, MLH1, MPL, MRE11A, MSH2, MSH3, MSH6, MST1R, MTAP, MTOR, MUTYH, MYB, MYC, MYCL, MYCN, MYD88, NBN, NF1, NF2, NFE2L2, NFKBIA, NKX2-1, NOTCH1, NOTCH2, NOTCH3, NPM1, NRAS, NT5C2, NTRK1, NTRK2, NTRK3, NUTM1, P2RY8, PALB2, PARK2, PARP1, PARP2, PARP3, PAX5, PBRM1, PDCD1, PDCD1LG2, PDGFRA, PDGFRB, PDK1, PIK3C2B, PIK3C2G, PIK3CA, PIK3CB, PIK3R1, PIM1, PMS2, POLDI, POLE, PPARG, PPP2R1A, PPP2R2A, PRDM1, PRKAR1A, PRKCI, PTCHI, PTEN, PTPN11, PTPRO, QKI, RAC1, RAD21, RAD51, RAD51B, RAD51C, RAD51D, RAD52, RAD54L, RAFI, RARA, RBI, RBM10, REL, RET, RICTOR, RNF43, ROS1, RPTOR, RSPO2, SDC4, SDHA, SDHB, SDHC, SDHD, SETD2, SF3B1, SGK1, SLC34A2, SMAD2, SMAD4, SMARCA4, SMARCB1, SMO, SNCAIP, S0CS1, SOX2, SOX9, SPEN, SPOP, SRC, STAG2, STAT3, STK11, SUFU, SYK, TBX3, TEK, TERC, TERT, TET2, TGFBR2, TIPARP, TMPRSS2, TNFAIP3, TNFRSF14, TP53, TSC1, TSC2, TYRO3, U2AF1, VEGFA, VHL, WHSCI, WHSC1L1, WT1, XPO1, XRCC2, ZNF217,ZNF703, or any combination thereof. In some embodiments, the one or more loci comprise ABL, ALK, ALL, B4GALNT1, BAFF, BCL2, BRAF, BRCA, BTK, CD19, CD20, CD3, CD30, CD319, CD38, CD52, CDK4, CDK6, CML, CRACC, CS1, CTLA-4, dMMR, EGFR, ERBB1, ERBB2, FGFR1-3, FLT3, GD2, HDAC, HER1, HER2, HR, IDH2, IL-ip, IL-6, IL- 6R, JAK1, JAK2, JAK3, KIT, KRAS, MEK, MET, MSI-H, mTOR, PARP, PD-1, PDGFR, PDGFRa, PDGFRP, PD-L1, PI3K5, PIGF, PTCH, RAF, RANKL, RET, ROS1, SLAMF7, VEGF, VEGFA, VEGFB, or any combination thereof.
[0029] In any of the preceding embodiments, the method may further include generating a report indicating the presence or absence of the methylation pattern of interest. In some embodiments, the method may further include transmitting the report to a healthcare provider. In some embodiments, the report is transmitted via a computer network or a peer- to-peer connection. In some embodiments, the sample comprises a fraction of tumor nucleic acids that is less than 1% of total nucleic acids. In some embodiments, the sample comprises a fraction of tumor nucleic acids that is less than 0.1% of total nucleic acids. In some embodiments, the sample comprises a fraction of tumor nucleic acids that is at least 0.01% of total nucleic acids. In some embodiments, the method may further include determining a prognosis of the subject. In some embodiments, the method may include predicting survival of the subject.
[0030] In any of the preceding embodiments, the method may further include predicting survival of the subject, wherein survival of the subject is predicted to be decreased compared to survival of an individual whose sample does not comprise the methylation pattern of interest associated with a hypermethylation signal indicative of a disease or condition. In any of the preceding embodiments, the method may further include predicting survival of the subject, wherein survival of the subject is predicted to be decreased compared to survival of an individual whose sample does not comprise the methylation pattern of interest associated with a hypomethylation signal indicative of a disease or condition.
[0031] In some embodiments, the subject has cancer, further comprising predicting tumor burden of the subject, wherein tumor burden of the subject is predicted to be increased compared to tumor burden of a subject whose sample does not comprise the methylation pattern of interest associated with a hypermethylation signal indicative of a disease or condition. In some embodiments in which the subject has cancer, the method may further include predicting tumor burden of the subject, wherein tumor burden of the subject ispredicted to be increased compared to tumor burden of a subject whose sample does not comprise the methylation pattern of interest associated with a hypomethylation signal indicative of a disease or condition.
[0032] In any of the preceding embodiments, the subject may have cancer, and the methylation pattern of interest may be associated with a hypermethylation or hypomethylation signal indicative of responsiveness to cancer treatment, further comprising predicting the subject’s responsiveness to cancer treatment.
[0033] In an additional aspect, provided herein is a method of monitoring a cancer in a subject, comprising: (a) quantifying the amount of a methylation pattern of interest present in a first sample from the subject by selectively capturing nucleic acid molecules according to the method of any one of claims 1-83 in the first sample; (b) quantifying the amount of the methylation pattern of interest in a second sample from the subject by selectively capturing nucleic acid molecules according to the method of any one of claims 1-83 in the second sample, wherein the second sample is obtained from the subject after the first sample; and (c) determining a difference in the amount of the methylation pattern of interest between the first and second samples, thereby monitoring the cancer in the individual.
[0034] In yet another aspect, provided herein is a method of monitoring response of a subject being treated for cancer, comprising: (a) quantifying the amount of a methylation pattern of interest present in a first sample from the subject by selectively capturing nucleic acid molecules according to the method of any one of claims 1-83 in the first sample; (b) after the first sample is obtained from the subject, administering a treatment to the subject; (c) quantifying the amount of the methylation pattern of interest in a second sample from the subject by selectively capturing nucleic acid molecules according to the method of any one of claims 1-83 in the second sample, wherein the second sample is obtained from the subject after the administration of the treatment; and (d) determining a difference in the amount of the methylation pattern of interest between the first and second samples, thereby monitoring response of the individual to the treatment.
[0035] In some embodiments comprising ligating one or more adapters onto one or more nucleic acid molecules prior to or after converting the methylation pattern but prior to amplifying the converted nucleic acid molecules, wherein the one or more adapters comprise amplification primers, flow cell adaptor sequences, substrate adapter sequences, or sample index sequences, the method further includes ligating one or more adapters onto one or morenucleic acid molecules prior to converting the methylation pattern into converted nucleic acid molecules.
[0036] In a further aspect, provided herein is a sample mixed with a bait in a vial, comprising: converted nucleic acid molecules corresponding to nucleic acid molecules from a sample from a subject, wherein the converted nucleic acid molecules comprise uracil (U) in place of unmethylated cytosine (C), or U in place of methylated C, compared to the nucleic acid molecules from the sample; a set of amplicons produced by amplifying the converted nucleic acid molecules; and a set of single- stranded nucleic acid baits; wherein the set of single- stranded nucleic acid baits hybridizes to the set of amplicons to form a plurality of hybrids, wherein at least a portion of the hybrids in the plurality of hybrids comprises a bait from the set of single- stranded nucleic acid baits and a nucleic acid molecule comprising a nucleic acid pattern of interest hybridized to each other without mismatches, wherein the nucleic acid pattern of interest corresponds to a methylation pattern of interest at a genomic locus in the nucleic acid molecules from the sample; and wherein plurality of hybrids is capable of being separated from nucleic acid molecules that are not hybridized to a bait. In some embodiments, the methylation pattern of interest corresponds to a hypermethylation signal indicative of a disease or condition. In some embodiments, the methylation pattern of interest corresponds to a hypomethylation signal indicative of a disease or condition.
[0037] In a yet another aspect, provided herein is a sample mixed with a bait in a vial, comprising: converted nucleic acid molecules corresponding to nucleic acid molecules from a sample from a subject, wherein the converted nucleic acid molecules comprise uracil (U) in place of unmethylated cytosine (C), or U in place of methylated C, compared to the nucleic acid molecules from the sample; a set of amplicons produced by amplifying the converted nucleic acid molecules; and a set of nucleic acid baits; wherein the set of nucleic acid baits hybridizes to the set of amplicons to form a plurality of hybrids, wherein the set of nucleic acid baits selectively hybridizes to a nucleic acid pattern of interest, wherein the nucleic acid pattern of interest corresponds to a methylation pattern of interest at a genomic locus in the nucleic acid molecules from the sample, wherein the methylation pattern of interest corresponds to a hypermethylation signal indicative of a disease or condition; and wherein plurality of hybrids is capable of being separated from nucleic acid molecules that are not hybridized to a bait. In some embodiments, the set of nucleic acid baits does not comprise anucleic acid bait that selectively hybridizes to a nucleic acid pattern that corresponds to a hypomethylation signal at the genomic locus.
[0038] In an additional aspect, provided herein is a sample mixed with a bait in a vial, comprising: converted nucleic acid molecules corresponding to nucleic acid molecules from a sample from a subject, wherein the converted nucleic acid molecules comprise uracil (U) in place of unmethylated cytosine (C), or U in place of methylated C, compared to the nucleic acid molecules from the sample; a set of amplicons produced by amplifying the converted nucleic acid molecules; and a set of nucleic acid baits; wherein the set of nucleic acid baits hybridizes to the set of amplicons to form a plurality of hybrids, wherein the set of nucleic acid baits selectively hybridizes to a nucleic acid pattern of interest, wherein the nucleic acid pattern of interest corresponds to a methylation pattern of interest at a genomic locus in the nucleic acid molecules from the sample, wherein the methylation pattern of interest corresponds to a hypomethylation signal indicative of a disease or condition; and wherein plurality of hybrids is capable of being separated from nucleic acid molecules that are not hybridized to a bait. In some embodiments, the set of nucleic acid baits does not comprise a nucleic acid bait that selectively hybridizes to a nucleic acid pattern that corresponds to a hypermethylation signal at the genomic locus.
[0039] In any of the preceding embodiments concerning a sample mixed with a bait in a vial, the converted nucleic acid molecules may include a plurality of nucleic acid patterns of interest converted from a plurality of methylation patterns of interest at a plurality of genomic loci in the nucleic acid molecules from the subject, and further may include a plurality of sets of nucleic acid baits, wherein the plurality of sets of nucleic acid baits hybridizes to the set of amplicons, wherein each set of nucleic acid baits of the plurality of sets of nucleic acid baits selectively hybridizes to a nucleic acid pattern of interest of the plurality of nucleic acid patterns of interest. In some embodiments, the plurality of methylation patterns of interest comprises a first methylation pattern of interest and a second methylation pattern of interest, wherein the second methylation pattern of interest is associated with a hypermethylation signal indicative of a disease or condition. In some embodiments, the plurality of methylation patterns of interest comprises a first methylation pattern of interest and a second methylation pattern of interest, wherein the second methylation pattern of interest is associated with a hypomethylation signal indicative of a disease or condition. In some embodiments, the first methylation pattern of interest and the second methylation pattern of interest are associatedwith hypomethylation or hypermethylation signals indicative of different diseases or conditions. In some embodiments, the first methylation pattern of interest and the second methylation pattern of interest are associated with hypomethylation or hypermethylation signals indicative of the same disease or condition. In some embodiments, the nucleic acid baits are double-stranded. In some embodiments, the nucleic acid baits are single- stranded. In some embodiments, the hybrids in the plurality of hybrids do not comprise guanine / thymine (G:T) mismatches. In some embodiments, a second portion of hybrids in the plurality of hybrids comprises a single- stranded nucleic acid bait from the set of single- stranded nucleic acid baits and a nucleic acid molecule that does not comprise the nucleic acid pattern of interest, wherein the hybrids in the second portion of hybrids comprise one or more adenine / cytosine (A:C) mismatches.
[0040] In another aspect, provided herein are methods of designing a bait set, comprising: identifying a hypermethylation pattern of interest at a genomic locus, wherein the hypermethylation pattern of interest comprises a forward strand and a reverse strand; determining a nucleic acid sequence for: (i) a forward converted nucleic acid corresponding to the forward strand of the hypermethylation pattern of interest; (ii) a reverse complement nucleic acid of the forward converted nucleic acid; (iii) a reverse converted nucleic acid corresponding to the reverse strand of the hypermethylation pattern of interest; and / or (iv) a reverse complement nucleic acid of the reverse converted nucleic acid; wherein the converted nucleic acid sequences comprise uracil (U) in place of unmethylated cytosine (C), or U in place of methylated C, compared to the hypermethylation pattern of interest at the genomic locus; and designing a set of double- stranded nucleic acid bait molecules, wherein each double-stranded nucleic acid bait molecule in the set of double- stranded nucleic acid bait molecules comprises two single- stranded nucleic acid molecules that are hybridized to each other without mismatches, wherein a first portion of double- stranded nucleic acid bait molecules in the set of double- stranded nucleic acid bait molecules comprises (i) a nucleic acid sequence that is complementary to the forward converted nucleic acid, and (ii) a nucleic acid sequence that is complementary to the reverse complement nucleic acid of the forward converted nucleic acid, and / or wherein a second portion of double- stranded nucleic acid bait molecules in the set of double- stranded nucleic acid bait molecules comprises (iii) a nucleic acid sequence that is complementary to the reverse converted nucleic acid, and (iv) a nucleicacid sequence that is complementary to the reverse complement nucleic acid of the reverse converted nucleic acid.
[0041] In yet another aspect, provided herein are methods of designing a bait set, comprising: identifying a hypomethylation pattern of interest at a genomic locus, wherein the hypomethylation pattern of interest comprises a forward strand and a reverse strand;
[0042] determining a nucleic acid sequence for: (i) a forward converted nucleic acid corresponding to the forward strand of the hypomethylation pattern of interest; (ii) a reverse complement nucleic acid of the forward converted nucleic acid; (iii) a reverse converted nucleic acid corresponding to the reverse strand of the hypomethylation pattern of interest; and / or (iv) a reverse complement nucleic acid of the reverse converted nucleic acid; wherein the converted nucleic acid sequences comprise uracil (U) in place of unmethylated cytosine (C), or U in place of methylated C, compared to the hypomethylation pattern of interest at the genomic locus; and designing a set of double-stranded nucleic acid bait molecules, wherein each double-stranded nucleic acid bait molecule in the set of double- stranded nucleic acid bait molecules comprises two single-stranded nucleic acid molecules that are hybridized to each other without mismatches, wherein a first portion of double-stranded nucleic acid bait molecules in the set of double- stranded nucleic acid bait molecules comprises (i) a nucleic acid sequence that is complementary to the forward converted nucleic acid, and (ii) a nucleic acid sequence that is complementary to the reverse complement nucleic acid of the forward converted nucleic acid, and / or wherein a second portion of double- stranded nucleic acid bait molecules in the set of double- stranded nucleic acid bait molecules comprises (iii) a nucleic acid sequence that is complementary to the reverse converted nucleic acid, and (iv) a nucleic acid sequence that is complementary to the reverse complement nucleic acid of the reverse converted nucleic acid.
[0043] In a further aspect, provided herein are methods of designing a bait set, comprising: identifying a hypermethylation pattern of interest at a genomic locus, wherein the hypermethylation pattern of interest comprises a forward strand and a reverse strand; determining a nucleic acid sequence for: (i) a forward converted nucleic acid corresponding to the forward strand of the hypermethylation pattern of interest; (ii) a reverse complement nucleic acid of the forward converted nucleic acid; (iii) a reverse converted nucleic acid corresponding to the reverse strand of the hypermethylation pattern of interest; and / or (iv) a reverse complement nucleic acid of the reverse converted nucleic acid; wherein the convertednucleic acid sequences comprise uracil (U) in place of unmethylated cytosine (C), or U in place of methylated C, compared to the hypermethylation pattern of interest at the genomic locus; and designing a set of single-stranded nucleic acid bait molecules, wherein a first portion of single- stranded nucleic acid bait molecules in the set of single-stranded nucleic acid bait molecules comprises (i) a nucleic acid sequence that is complementary to the forward converted nucleic acid, and / or wherein a second portion of single- stranded nucleic acid bait molecules in the set of single-stranded nucleic acid bait molecules comprises (ii) a nucleic acid sequence that is complementary to the reverse complement nucleic acid of the forward converted nucleic acid, and / or wherein a third portion of single- stranded nucleic acid bait molecules in the set of single-stranded nucleic acid bait molecules comprises (iii) a nucleic acid sequence that is complementary to the reverse converted nucleic acid, and / or wherein a fourth portion of single- stranded nucleic acid bait molecules in the set of singlestranded nucleic acid bait molecules comprises (iv) a nucleic acid sequence that is complementary to the reverse complement nucleic acid of the reverse converted nucleic acid.
[0044] In an additional aspect, provided herein are methods of designing a bait set, comprising: identifying a hypomethylation pattern of interest at a genomic locus, wherein the hypomethylation pattern of interest comprises a forward strand and a reverse strand; determining a nucleic acid sequence for: (i) a forward converted nucleic acid corresponding to the forward strand of the hypomethylation pattern of interest; (ii) a reverse complement nucleic acid of the forward converted nucleic acid; (iii) a reverse converted nucleic acid corresponding to the reverse strand of the hypomethylation pattern of interest; and / or (iv) a reverse complement nucleic acid of the reverse converted nucleic acid; wherein the converted nucleic acid sequences comprise uracil (U) in place of unmethylated cytosine (C), or U in place of methylated C, compared to the hypomethylation pattern of interest at the genomic locus; and designing a set of single-stranded nucleic acid bait molecules, wherein a first portion of single- stranded nucleic acid bait molecules in the set of single-stranded nucleic acid bait molecules comprises (i) a nucleic acid sequence that is complementary to the forward converted nucleic acid, and / or wherein a second portion of single- stranded nucleic acid bait molecules in the set of single-stranded nucleic acid bait molecules comprises (ii) a nucleic acid sequence that is complementary to the reverse complement nucleic acid of the forward converted nucleic acid, and / or wherein a third portion of single- stranded nucleic acid bait molecules in the set of single-stranded nucleic acid bait molecules comprises (iii) anucleic acid sequence that is complementary to the reverse converted nucleic acid, and / or wherein a fourth portion of single- stranded nucleic acid bait molecules in the set of singlestranded nucleic acid bait molecules comprises (iv) a nucleic acid sequence that is complementary to the reverse complement nucleic acid of the reverse converted nucleic acid.
[0045] In some embodiments of any of the methods of designing a bait set provided herein, the nucleic acid sequence of the forward converted nucleic acid may include a cytosine residue at a position corresponding to a methylated cytosine residue in the hypermethylation pattern of interest and a uracil or thymine residue at a position corresponding to an unmethylated cytosine residue in the hypermethylation pattern of interest. In some embodiments of any of the methods of designing a bait set provided herein, the nucleic acid sequence of the forward converted nucleic acid may include a cytosine residue at a position corresponding to a methylated cytosine residue in the hypomethylation pattern of interest and a uracil or thymine residue at a position corresponding to an unmethylated cytosine residue in the hypomethylation pattern of interest. In some embodiments, the set of single- stranded nucleic acid bait molecules does not comprise a bait molecule that hybridizes with a G:T mismatch to a nucleic acid molecule or complement thereof comprising a sequence of a converted nucleic acid corresponding to a hypomethylation pattern at the genomic locus. In some embodiments, the set of single- stranded nucleic acid bait molecules does not comprise a bait molecule that hybridizes with a G:T mismatch to a nucleic acid molecule or complement thereof comprising a sequence of a converted nucleic acid corresponding to a hypermethylation pattern at the genomic locus. In some embodiments, the set of singlestranded nucleic acid bait molecules comprises a bait molecule that hybridizes with an A:C mismatch to a nucleic acid molecule or complement thereof comprising a sequence of a converted nucleic acid corresponding to a hypomethylation pattern at the genomic locus. In some embodiments, the set of single- stranded nucleic acid bait molecules does comprises a bait molecule that hybridizes with an A:C mismatch to a nucleic acid molecule or complement thereof comprising a sequence of a converted nucleic acid corresponding to a hypermethylation pattern at the genomic locus.
[0046] In some embodiments of any of the methods of designing a bait set provided herein, the nucleic acid bait molecules further comprise one or more additional sequence and / or modification. In some embodiments, the additional sequence comprises an adapter sequence, a primer sequence, a barcode sequence, a unique molecular identifier sequence, and / or auniversal tail on one or both ends. In some embodiments, the modification comprises covalently attaching a compound or protein. In some embodiments, the compound is biotin.
[0047] In some embodiments, the method further includes designing the bait set of any of the methods of designing a bait set provided herein and synthesizing the designed bait molecules. A further aspect of the disclosure includes a bait set made according to the method of the preceding embodiment.
[0048] It should be appreciated that all combinations of the foregoing concepts and additional concepts discussed in greater detail below are contemplated as being part of the inventive subject matter disclosed herein. In particular, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the inventive subject matter disclosed herein.INCORPORATION BY REFERENCE
[0049] All publications, patents, patent applications, and any other references mentioned in this specification are herein incorporated by reference in their entirety to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference in its entirety. In the event of a conflict between a term herein and a term in an incorporated reference, the term herein controls.BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Various aspects of the disclosed methods, devices, and systems are set forth with particularity in the appended claims. A better understanding of the features and advantages of the disclosed methods, devices, and systems will be obtained by reference to the following detailed description of illustrative embodiments and the accompanying drawings, of which:
[0051] FIG. 1 provides a schematic of a non-limiting exemplary workflow of current methods for detection of cancer-associated methylation signals, including, for example, from minimum residual disease (MRD) / early- stage cancer patient specimens or treatment response monitoring (TRM) / advanced- stage cancer patient specimens. This exemplary workflow begins with a cell-free DNA (cfDNA) sample, followed by end repair and adaptor ligation, cytosine conversion, library PCR, target capture, sequencing, determining cancer hypermethylation and / or hypomethylation scores, and finally making cancer identity calls and determining tumor fraction.
[0052] FIG. 2 provides a schematic illustration of an exemplary process for selectively capturing nucleic acid molecules using either the single- stranded hypermethylation- specific baits or the single-stranded hypomethylation- specific baits described herein.
[0053] FIG. 3 provides a schematic illustration of an exemplary process for selectively capturing nucleic acid molecules using either the single- stranded hypermethylation- specific baits or the double- stranded hypermethylation- specific baits described herein.
[0054] FIG. 4 provides a schematic illustration of an exemplary process for selectively capturing nucleic acid molecules using either the single- stranded hypomethylation- specific baits or the double- stranded hypomethylation- specific baits described herein.
[0055] FIG. 5 provides schematics illustrating a non-limiting example of design principles for allele- specific methyl enrichment baits.
[0056] FIGS. 6A-6C provide non-limiting examples of enhanced detection of cancer- associated methylation signals using methyl allele- specific baits. FIG. 6A shows a plot of cancer hypermethylation scores via cluster consensus methylation fraction(ccmf) determined from 292 CpG clusters known to be hypermethylated in cancer included in the target capture panel. The data depicted in FIG. 6A are from the same experiment as FIG. 6C, but showing the overall signal from all 292 loci. FIG. 6B shows a plot of cancer hypomethylation scores via duplex cluster consensus unmethylated fraction (dccuf) determined from 80 CpG clusters known to be hypo-methylated in cancer included in the target capture panel. FIG. 6C shows a heatmap of methylation scores per locus for a target capture panel containing CpG clusters known to be hypermethylated in cancer, from the same experiment as FIG. 6A, but showing the enhanced enrichment from each hypermethylation- specific bait at each locus. The methylation scores are displayed in the heatmap (bottom section) as indicated by the methylation scale (from 0-0.8) at right.
[0057] FIG. 7 provides schematics illustrating a non-limiting example of types of off-target capture that can happen between double-stranded, allele- specific baits and unintended amplicons.
[0058] FIGS. 8A-8B provide non-limiting examples of data demonstrating that target-bait duplexes with A:C mismatches have lower thermostability and can be selected against in post-hybridization washes. FIG. 8A shows a box-and-whiskers plot of the determined melting temperature (Tm) of hybrids from 500 randomly selected methylation loci with different baits designed against a DNA strand that was the forward (+) strand in the originalsample. FIG. 8B shows a plot of the Tm of hybrids with different baits designed against a DNA strand that was the reverse (-) strand in the original sample.
[0059] FIGS. 9A-9B provide schematics illustrating strand- selective bait design. FIG. 9A provides a non-limiting example of single stranded (ss) hypomethylation- specific baits designed to have only A:C mismatches with amplicons corresponding to hypermethylated regions. FIG. 9B provides a non-limiting example of ss strand- specific bait designs (far right) compared to double-stranded allele- specific baits (second column from right)
[0060] FIGS. 10A-10B provide non-limiting examples of data demonstrating that strand- selective methyl enrichment bait design enables further enrichment of cancer-associated methylation signals. Each plot displays cancer methylation scores (hypermethylation scores in FIG. 10A; hypomethylation scores in FIG. 10B) determined for the bait types indicated along the x-axis.DETAILED DESCRIPTION
[0061] The present disclosure relates generally to nucleic acid baits for selective enrichment of methylation signals, as well as methods of designing and / or making such baits. These may be used to enhance detection of a methylation signal on, for instance, a cluster of CpG dinucleotides that is known to be hypomethylated or hypermethylated in association with a disease or a condition. Accordingly, methods of using the baits for selective enrichment and / or sequencing of targeted methylation signals is also described herein.
[0062] Previous methods for methylation analysis involve calculation of the methylation fraction (the fraction of potential methylation sites that are methylated, such as the proportion of methylated CpG sites out of the total number of CpG sites) at a particular locus, and require capturing converted nucleic acid molecules representing each methylation state (e.g., hypermethylated and hypomethylated) equally. In contrast, the baits described herein are each designed to capture converted nucleic acids corresponding to only one methylation state (e.g., either only hypermethylation signals or only hypermethylation signals). The baits described herein may be single- stranded or double-stranded, and are designed to selectively hybridize with a nucleic acid molecule comprising a sequence corresponding to a methylation pattern of interest at a genomic locus of interest that has been converted (by, for example, enzymatic nucleotide conversion or bisulfite nucleotide conversion) into a nucleic acid pattern of interest. Some baits of the present disclosure may be further designed to form oneor more A:C mismatches and / or no G:T mismatches when hybridized to an off-target nucleic acid molecule. For example, a single- stranded hypomethylation- specific bait of the present disclosure may form one or more A:C mismatches and / or no G:T mismatches when hybridized to a nucleic acid molecule comprising a sequence corresponding to a converted nucleic acid molecule corresponding to a hypermethylated allele of the target genomic locus, and a single- stranded hypermethylation- specific bait of the present disclosure may form one or more A:C mismatches and / or no G:T mismatches when hybridized to a nucleic acid molecule comprising a sequence corresponding to a converted nucleic acid molecule corresponding to a hypomethylated allele of the target genomic locus.
[0063] When implemented in target capture following next generation sequencing (NGS) library construction from cell-free DNA samples (see FIG. 1), the nucleic acid baits described herein may, for example, improve sensitivity of detection of hypomethylation patterns or hypermethylation patterns. This can include, but is not limited to, for instance, improved sensitivity of detection of hypomethylation patterns or hypermethylation patterns that are associated with one or more diseases or conditions, such as, for example, treatment selection, treatment response, minimum residual disease (MRD) or early- stage cancer detection methods, or treatment response monitoring in, for example, advanced disease.
[0064] Aberrant methylation is a feature of many cancers and can be detected in many different types of patient samples, including those that include cfDNA. Detection of rare cancer-driven methylation patterns is a key challenge in cancer screening and monitoring of MRD, in particular. The present disclosure describes, inter alia, methods for detecting aberrant methylation (e.g., DNA methylation in CpG dinucleotide clusters) that effectively reduce background and increase signal-to-background ratio, thus allowing for detection of very low-frequency tumor DNA in otherwise normal DNA samples, which may assist in early detection and / or monitoring of cancer. The methods described in the present disclosure may similarly assist in treatment response monitoring (TRM), such as, for example, by facilitating increased signal-to-noise when tracking the amount of a methylation signal of interest in response to a treatment. Comprehensive genomic profiles (CGPs) may also include data generated according to the methods described herein, such as, for example, data relating to a methylation signal of interest from, for example, a solid tissue sample.
[0065] The designs of the nucleic acids baits described herein allow improved sensitivity of detection of hypomethylation patterns or hypermethylation patterns while also reducing assaycosts compared to other methods, by, for instance, allowing for using a smaller number of reads (i.e., lower depth sequencing) compared to conventional bait designs and associated methods.
[0066] As described herein, by enriching aberrant molecules from normal molecules in the target capture step, prior to sequencing, one can detect higher levels of aberrant molecules with improved efficiency compared to previous methods.
[0067] Described herein are bait sets and methods of designing bait sets for use in target capture, which enable allele- specific and / or strand-specific capture of methylation signals from genomic regions that contain a methylation signal (either hypomethylation or hypermethylation) associated with a disease or condition. Accordingly, also described herein are methods of selectively capturing nucleic acid molecules using the bait sets described herein.
[0068] Various bait designs are described herein. For instance, the baits described herein may be single-stranded or double-stranded, and may be hypomethylation- specific or hypermethylation- specific. The methods of designing a bait of the present disclosure can include one or more of the following: i) identifying a specific methylation pattern of interest (e.g., a hypermethylation pattern of interest or a hypomethylation pattern of interest) at a genomic locus; ii) determining a converted nucleic acid sequence and reverse complement thereof for each strand (e.g., forward and reverse) of the methylation pattern at the genomic locus; and iii) designing a nucleic acid bait molecule that selectively hybridizes to one or more of the converted nucleic acid sequences of the methylation pattern at the genomic locus.
[0069] The methods of using a bait of the present disclosure (such as, for instance, in a method of selectively capturing nucleic acid molecules) can include one or more of the following: i) obtaining a sample of nucleic acid molecules from a subject to query for the methylation pattern of interest; ii) converting any methylation patterns present in the nucleic acid molecules from the subject into nucleic acid patterns, such that the methylation pattem(s) originally present in the sample become “encoded” or “fixed” as nucleic acid patterns (which allows the methylation pattern of interest to be queried as a nucleic acid pattern of interest); iii) amplifying the converted nucleic acid molecules into a set of amplicons; iv) hybridizing the baits to the amplicons (during which, the baits selectively hybridize to the nucleic acid pattern of interest, if present); and v) removing nucleic acids that are not hybridized to a bait. The particular designs of the baits, as described herein, provideimproved sensitivity and sequencing efficiency compared to previously available methods. This overall scheme can be multiplexed and / or modified in a variety of ways, as described herein.
[0070] Depending on the disease, condition, and / or genomic locus driving the inquiry, a methylation pattern of interest may correspond to a hypermethylation signal or a hypomethylation signal. In some instances, it is possible for a single genomic locus to be associated with multiple different types of methylation states, depending on the disease or condition (for instance, hypermethylated in one disease but hypomethylated in another disease; hypomethylated in one pattern in one disease but hypomethylated in another pattern in another disease; hypermethylated in one pattern in one disease but hypermethylated in another pattern in another disease; etc.). However, each bait or set of baits described herein is designed to enrich for just one type of methylation signal: either a single hypermethylation signal at a specific locus or a single hypomethylation signal at a specific locus, but not both a hypomethylated signal and a hypermethylated signal at the same locus. Thus, for bait sets that are designed to enrich for one or more hypermethylation signals, the baits can be further designed to not include nucleic acid sequences that would hybridize to a nucleic acid pattern that corresponds to a hypomethylation signal at the genomic locus. Correspondingly, for bait sets that are designed to enrich for one or more hypomethylation signals, the baits can be further designed to not include nucleic acid sequences that would hybridize to a nucleic acid pattern that corresponds to a hypermethylation signal at the genomic locus.
[0071] As described below, the baits are either single- or double- stranded nucleic acid molecules. Further, the baits are categorized as “allele- specific” (also described herein as “methyl allele-specific”) or “strand- selective” (also described herein as “strand-specific”). Strand- selective baits as discussed herein are single-stranded (and thus are also referred to herein synonymously as, for instance, single-stranded baits). Allele- specific baits as discussed herein in are double- stranded. Both allele- specific and strand-selective baits may be designed to enrich for a hypermethylation signal or a hypomethylation signal. Thus, a hypermethylation- specific bait may be either allele-specific or strand- specific, and either double-stranded or single-stranded, respectively. Similarly, a hypomethylation- specific bait may be either allele- specific or strand-specific, and either double-stranded or single- stranded, respectively.
[0072] In some instances, the nucleic acid baits are single- stranded. This so-called “singlestranded”, “strand-selective”, or “strand-specific” methyl enrichment bait design allows selective capture and enrichment of specific strands of DNA molecules (i.e., the strands present in the original sample in 5' to 3' orientation; e.g., the “positive” (+) strand (also known as the “sense” or “coding” strand, if in the context of a gene) or the “negative” (-) strand (i.e., the strands present in the original sample in 3' to 5' orientation; also known as the “antisense” or “non-coding” strand, if in the context of a gene) from Methyl-seq libraries that harbor hypermethylation or hypomethylation signals that distinguish, e.g., cancer or other disease from healthy samples. Use of these single- stranded baits involves hybridizing the single- stranded nucleic acid baits to the set of amplicons discussed above, such that at least a portion of the resulting hybrids include mismatch-free hybridization between 1) a whole single- stranded nucleic acid bait and 2) the nucleic acid pattern of interest.
[0073] In some instances, the nucleic acid baits are double-stranded. This so-called “methyl allele-specific” enrichment bait design allows selective capture and enrichment of alleles with hypermethylation or hypomethylation signals from mMethyl-seq libraires. These “allelespecific” double-stranded baits are similar in concept to the single- stranded baits described above, except that they do not discriminate between the + / - DNA strands from the sample: each “allele-specific” bait comprises two strands, which separate during the hybridization step such that each strand of a double-stranded bait can hybridize separately with a different nucleic acid strand. The individual strands of an allele- specific bait are complementary to each other and are each still considered to be “allele- specific” baits even when separated from the complementary bait strand, such as, for example, during and after hybridization to a nucleic acid molecule comprising the nucleic acid pattern of interest.
[0074] There are more opportunities for mismatches between a bait and an amplicon when hybridizing using double-stranded baits, meaning that the double- stranded baits provide slightly less specificity than the single-stranded baits. However, as shown herein, both versions represent improvements over previous bait designs.
[0075] The methods of designing a bait set described herein are improvements compared to traditional bait designs for Methyl-seq. In traditional bait designs, methylation signals are indiscriminately captured without selectivity for alleles or strands. The new bait designs described herein provide greater flexibility and higher selectivity in capturing allele- and strand-specific methylation signals that effectively distinguish cancer and healthy samples,which improves assay sensitivity and lowers cost in, for example, MRD / early-stage cancer diagnostic assays under development.I. General Techniques
[0076] The techniques and procedures described or referenced herein are generally well understood and commonly employed using conventional methodology by those skilled in the art, such as, for example, the widely utilized methodologies described in Sambrook et al., Molecular Cloning: A Laboratory Manual 3d edition (2001) Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y.; Current Protocols in Molecular Biology (F.M. Ausubel et al. eds., (2003)); the series Methods in Enzymology (Academic Press, Inc.): PCR 2: A Practical Approach (M.J. MacPherson, B.D. Hames and G.R. Taylor eds. (1995)), Harlow and Lane, eds. (1988) Antibodies, A Laboratory Manual, and Animal Cell Culture (R.I. Freshney, ed. (1987)); Oligonucleotide Synthesis (M.J. Gait, ed., 1984); Methods in Molecular Biology, Humana Press; Cell Biology: A Laboratory Notebook (J.E. Cellis, ed., 1998) Academic Press; Animal Cell Culture (R.I. Freshney), ed., 1987); Introduction to Cell and Tissue Culture (J.P. Mather and P.E. Roberts, 1998) Plenum Press; Cell and Tissue Culture: Laboratory Procedures (A. Doyle, J.B. Griffiths, and D.G. Newell, eds., 1993-8) J. Wiley and Sons; Handbook of Experimental Immunology (D.M. Weir and C.C. Blackwell, eds.); Gene Transfer Vectors for Mammalian Cells (J.M. Miller and M.P. Calos, eds., 1987); PCR: The Polymerase Chain Reaction, (Mullis et al., eds., 1994); Current Protocols in Immunology (J.E. Coligan et al., eds., 1991); Short Protocols in Molecular Biology (Wiley and Sons, 1999); Immunobiology (C.A. Janeway and P. Travers, 1997); Antibodies (P. Finch, 1997); Antibodies: A Practical Approach (D. Catty., ed., IRL Press, 1988-1989); Monoclonal Antibodies: A Practical Approach (P. Shepherd and C. Dean, eds., Oxford University Press, 2000); Using Antibodies: A Laboratory Manual (E. Harlow and D. Lane (Cold Spring Harbor Laboratory Press, 1999); The Antibodies (M. Zanetti and J. D. Capra, eds., Harwood Academic Publishers, 1995); and Cancer: Principles and Practice of Oncology (V.T. DeVita et al., eds., J.B. Lippincott Company, 1993).II. Definitions
[0077] Unless otherwise defined, all of the technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art in the field to which this disclosure belongs.
[0078] As used in this specification and the appended claims, the singular forms “a”, “an” and “the” include plural references unless the context clearly dictates otherwise. Thus, for example, reference to “a molecule” optionally includes a combination of two or more such molecules, and the like. Any reference to “or” herein is intended to encompass “and / or” unless otherwise stated.
[0079] The term “aberrant methylation” is used herein to refer to a pattern of methylation that is not typically present in a normal epigenetic signatures of healthy cells. For example, the term can refer to increased methylation at a site that is not normally methylated in normal cells / tissue, or decreased methylation at a site that is normally methylated in normal cells / tissue. In some embodiments, nucleic acids derived from a sample (e.g., tumor derived nucleic acids) are characterized by aberrant methylation when their pattern and / or amount of methylation at one or more genomic loci differs from what is normally present at the corresponding locus / loci in a particular type of tissue. “Aberrant methylation” as used herein includes both “hypermethylation” and “hypomethylation”. “Hypermethylation” as used herein refers to aberrant methylation in the form of increased methylation at one or more genomic loci compared to what is normally present at the corresponding locus / loci in a particular type of tissue. “Hypomethylation” as used herein refers to aberrant methylation in the form of decreased methylation at one or more genomic loci compared to what is normally present at the corresponding locus / loci in a particular type of tissue.
[0080] The term “about” as used herein refers to the usual error range for the respective value readily known to the skilled person in this technical field. Reference to “about” a value or parameter herein includes (and describes) embodiments that are directed to that value or parameter per se.
[0081] When a range of values is provided, it is to be understood that each intervening value between the upper and lower limit of that range, and any other stated or intervening value in that states range, is encompassed within the scope of the present disclosure. Where the stated range includes upper or lower limits, ranges excluding either of those included limits are also included in the present disclosure.
[0082] As used herein, “administering” is meant a method of giving a dosage of a compound (e.g., an antagonist) or a pharmaceutical composition (e.g., a pharmaceutical composition including an antagonist) to a subject (e.g., a patient). Administering can be by any suitable means, including parenteral, intrapulmonary, and intranasal, and, if desired for localtreatment, intralesional administration. Parenteral infusions include, for example, intramuscular, intravenous, intraarterial, intraperitoneal, or subcutaneous administration. Dosing can be by any suitable route, e.g., by injections, such as intravenous or subcutaneous injections, depending in part on whether the administration is brief or chronic. Various dosing schedules including but not limited to single or multiple administrations over various timepoints, bolus administration, and pulse infusion are contemplated herein.
[0083] The terms “allele frequency” and “allele fraction” are used interchangeably herein and refer to the fraction of sequence reads corresponding to a particular allele relative to the total number of sequence reads for a genomic locus.
[0084] “Amplification,” as used herein generally refers to the process of producing multiple copies of a desired sequence. “Multiple copies” mean at least two copies. A “copy” does not necessarily mean perfect sequence complementarity or identity to the template sequence. For example, copies can include nucleotide analogs such as deoxyinosine, intentional sequence alterations (such as sequence alterations introduced through a primer comprising a sequence that is hybridizable, but not complementary, to the template), and / or sequence errors that occur during amplification.
[0085] The term “bait” is used herein synonymously with “nucleic acid bait”, “bait molecule” and “target capture reagent” and refers to one or more nucleic acid molecules designed according to the principles laid out herein to hybridize with a specific nucleic acid pattern of interest present in a converted nucleic acid.
[0086] The terms “cancer” and “cancerous” refer to or describe the physiological condition in mammals that is typically characterized by unregulated cell growth. Included in this definition are benign, premalignant, and malignant cancers.
[0087] It is understood that aspects and embodiments of the invention described herein include “comprising,” “consisting,” and “consisting essentially of’ aspects and embodiments. As used herein, the terms “comprising” (and any form or variant of comprising, such as “comprise” and “comprises”), “having” (and any form or variant of having, such as “have” and “has”), “including” (and any form or variant of including, such as “includes” and “include”), or “containing” (and any form or variant of containing, such as “contains” and “contain”), are inclusive or open-ended and do not exclude additional, un-recited additives, components, integers, elements, or method steps.
[0088] The term “concurrently” is used herein to refer to administration of two or more therapeutic agents, where at least part of the administration overlaps in time. Accordingly, concurrent administration includes a dosing regimen when the administration of one or more agent(s) continues after discontinuing the administration of one or more other agent(s).
[0089] A “control sample,” “control cell,” or “control tissue,” as used herein, refers to a sample, cell, tissue, standard, or level that is used for comparison purposes.
[0090] The term “CpG dinucleotide” is used herein to refer to a region of 2 or more DNA bases in which a cytosine nucleotide is followed by a guanine nucleotide in the 5'a3' direction, e.g., 5'-C-phosphate-G-3'. In many genomes, CpG dinucleotides can often be found in “clusters” or regions of DNA containing multiple CpG dinucleotides (also termed “CpG islands”). Much or most of DNA methylation in many genomes is present in CpG dinucleotides (in which the cytosine is methylated or hydroxy methylated).
[0091] The term “detection” includes any means of detecting, including direct and indirect detection.
[0092] The term “diagnosis” is used herein to refer to the identification or classification of a molecular or pathological state, disease or condition (e.g., cancer). For example, “diagnosis” may refer to identification of a particular type of cancer. “Diagnosis” may also refer to the classification of a particular subtype of cancer, for instance, by histopathological criteria, or by molecular features (e.g., a subtype characterized by expression of one or a combination of biomarkers (e.g., particular genes or proteins encoded by said genes), or by aberrant DNA methylation level and / or pattern).
[0093] An “effective amount” refers to an amount of a therapeutic agent to treat or prevent a disease or disorder in a mammal. In the case of cancers, the therapeutically effective amount of the therapeutic agent may reduce the number of cancer cells; reduce the primary tumor size; inhibit (i.e., slow to some extent and in some embodiments stop) cancer cell infiltration into peripheral organs; inhibit (i.e., slow to some extent and in some embodiments stop) tumor metastasis; inhibit, to some extent, tumor growth; and / or relieve to some extent one or more of the symptoms associated with the disorder. To the extent the drug may prevent growth and / or kill existing cancer cells, it may be cytostatic and / or cytotoxic. For cancer therapy, efficacy in vivo can, for example, be measured by assessing the duration of survival, time to disease progression (TTP), response rates (e.g., CR and PR), duration of response, and / or quality of life.
[0094] The terms “genomic locus” and “genomic loci” as used herein include nuclear and organellar (e.g., mitochondrial) DNA present in the genome of a subject, as well as transcripts and fragments thereof.
[0095] As used herein, the terms “individual,” “patient,” or “subject” are used interchangeably and refer to any single animal, e.g., a mammal (including such non-human animals as, for example, dogs, cats, horses, rabbits, zoo animals, cows, pigs, sheep, and non- human primates) for which treatment is desired. In particular embodiments, the individual, patient, or subject herein is a human.
[0096] ‘ ‘Individual response” or “response” can be assessed using any endpoint indicating a benefit to the individual, including, without limitation, (1 ) inhibition, to some extent, of disease progression (e.g., cancer progression), including slowing down or complete arrest; (2) a reduction in tumor size; (3) inhibition (i.e., reduction, slowing down, or complete stopping) of cancer cell infiltration into adjacent peripheral organs and / or tissues; (4) inhibition (i.e. reduction, slowing down, or complete stopping) of metastasis; (5) relief, to some extent, of one or more symptoms associated with the disease or disorder (e.g., cancer); (6) increase or extension in the length of survival, including overall survival and progression free survival; and / or (7) decreased mortality at a given point of time following treatment.
[0097] “Oligonucleotide,” as used herein, generally refers to short, single stranded, polynucleotides that are, but not necessarily, less than about 250 nucleotides in length. Oligonucleotides may be synthetic. The terms “oligonucleotide” and “polynucleotide” are not mutually exclusive. The description above for polynucleotides is equally and fully applicable to oligonucleotides.
[0098] A “pharmaceutically acceptable carrier” refers to an ingredient in a pharmaceutical formulation, other than an active ingredient, which is nontoxic to a subject. A pharmaceutically acceptable carrier includes, but is not limited to, a buffer, excipient, stabilizer, or preservative.
[0099] The term “pharmaceutical formulation” refers to a preparation which is in such form as to permit the biological activity of an active ingredient contained therein to be effective, and which contains no additional components which are unacceptably toxic to a subject to which the formulation would be administered.
[0100] The technique of “polymerase chain reaction” or “PCR” as used herein generally refers to a procedure wherein minute amounts of a specific piece of nucleic acid, RNA and / orDNA, are amplified as described, for example, in U.S. Pat. No. 4,683,195. Generally, sequence information from the ends of the region of interest or beyond needs to be available, such that oligonucleotide primers can be designed; these primers will be identical or similar in sequence to opposite strands of the template to be amplified. The 5' terminal nucleotides of the two primers may coincide with the ends of the amplified material. PCR can be used to amplify specific RNA sequences, specific DNA sequences from total genomic DNA, and cDNA transcribed from total cellular RNA, bacteriophage, or plasmid sequences, etc. See generally Mullis et al., Cold Spring Harbor Symp. Quant. Biol. 51 :263 (1987) and Erlich, ed., PCR Technology (Stockton Press, NY, 1989). As used herein, PCR is considered to be one, but not the only, example of a nucleic acid polymerase reaction method for amplifying a nucleic acid test sample, comprising the use of a known nucleic acid (DNA or RNA) as a primer and utilizes a nucleic acid polymerase to amplify or generate a specific piece of nucleic acid or to amplify or generate a specific piece of nucleic acid which is complementary to a particular nucleic acid.
[0101] “Polynucleotide,” or “nucleic acid,” as used interchangeably herein, refer to polymers of nucleotides of any length, and include DNA and RNA. The nucleotides can be deoxyribonucleotides, ribonucleotides, modified nucleotides or bases, and / or their analogs, or any substrate that can be incorporated into a polymer by DNA or RNA polymerase, or by a synthetic reaction. Thus, for instance, polynucleotides as defined herein include, without limitation, single- and double- stranded DNA, DNA including single- and double-stranded regions, single- and double- stranded RNA, and RNA including single- and double- stranded regions, hybrid molecules comprising DNA and RNA that may be single- stranded or, more typically, double-stranded or include single- and double- stranded regions. In addition, the term “polynucleotide” as used herein refers to triple-stranded regions comprising RNA or DNA or both RNA and DNA. The strands in such regions may be from the same molecule or from different molecules. The regions may include all of one or more of the molecules, but more typically involve only a region of some of the molecules. One of the molecules of a triple-helical region often is an oligonucleotide. The term “polynucleotide” specifically includes cDNAs.
[0102] A polynucleotide may comprise modified nucleotides, such as methylated nucleotides and their analogs. If present, modification to the nucleotide structure may be imparted before or after assembly of the polymer. The sequence of nucleotides may be interrupted by non-nucleotide components. A polynucleotide may be further modified after synthesis, such as by conjugation with a label. Other types of modifications include, for example, “caps,” substitution of one or more of the naturally-occurring nucleotides with an analog, intemucleotide modifications such as, for example, those with uncharged linkages (e.g., methyl phosphonates, phosphotriesters, phosphoamidates, carbamates, and the like) and with charged linkages (e.g., phosphorothioates, phosphorodithioates, and the like), those containing pendant moieties, such as, for example, proteins (e.g., nucleases, toxins, antibodies, signal peptides, poly-L-lysine, and the like), those with intercalators (e.g., acridine, psoralen, and the like), those containing chelators (e.g., metals, radioactive metals, boron, oxidative metals, and the like), those containing alkylators, those with modified linkages (e.g., alpha anomeric nucleic acids), as well as unmodified forms of the polynucleotide(s). Further, any of the hydroxyl groups ordinarily present in the sugars may be replaced, for example, by phosphonate groups, phosphate groups, protected by standard protecting groups, or activated to prepare additional linkages to additional nucleotides, or may be conjugated to solid or semi-solid supports. The 5' and 3' terminal OH can be phosphorylated or substituted with amines or organic capping group moieties of from 1 to 20 carbon atoms. Other hydroxyls may also be derivatized to standard protecting groups. Polynucleotides can also contain analogous forms of ribose or deoxyribose sugars that are generally known in the art, including, for example, 2'-0-methyl-, 2'-0-allyl-, 2'-fluoro-, or 2'- azido-ribose, carbocyclic sugar analogs, a-anomeric sugars, epimeric sugars such as arabinose, xyloses or lyxoses, pyranose sugars, furanose sugars, sedoheptuloses, acyclic analogs, and abasic nucleoside analogs such as methyl riboside. One or more phosphodiester linkages may be replaced by alternative linking groups. These alternative linking groups include, but are not limited to, embodiments wherein phosphate is replaced by P(0)S ("thioate"), P(S)S ("dithioate"), "(0)NR2("amidate"), P(0)R, P(0)OR', CO orCH2("formacetal"), in which each R or R' is independently H or substituted or unsubstituted alkyl (1 -20 C) optionally containing an ether (-0-) linkage, aryl, alkenyl, cycloalkyl, cycloalkenyl or araldyl. Not all linkages in a polynucleotide need be identical. A polynucleotide can contain one or more different types of modifications as described herein and / or multiple modifications of the same type. The preceding description applies to all polynucleotides referred to herein, including RNA and DNA.
[0103] The term “sample,” as used herein, refers to a composition that is obtained or derived from a subject and / or individual of interest that contains a cellular and / or other molecular entity that is to be characterized and / or identified, for example, based on physical, biochemical, chemical, and / or physiological characteristics. For example, the phrase “disease sample” and variations thereof refers to any sample obtained from a subject of interest that would be expected or is known to contain the cellular and / or molecular entity that is to be characterized. Samples include, but are not limited to, tissue samples, primary or cultured cells or cell lines, cell supernatants, cell lysates, platelets, serum, plasma, vitreous fluid, lymph fluid, synovial fluid, follicular fluid, seminal fluid, amniotic fluid, milk, whole blood, plasma, serum, blood-derived cells, urine, cerebro- spinal fluid, saliva, sputum, tears, perspiration, mucus, tumor lysates, and tissue culture medium, tissue extracts such as homogenized tissue, tumor tissue, cellular extracts, and combinations thereof. In some instances, the sample is a whole blood sample, a plasma sample, a serum sample, or a combination thereof. In some embodiments, the sample is from a tumor (e.g., a “tumor sample”), such as from a biopsy. In some embodiments, the sample is a formalin-fixed paraffin-embedded (FFPE) sample.
[0104] The term “set” as used herein refers to one or more. The term “subset”, as in subset of a set, as used herein refers to less than a complete “set”. Note that a “subset” is not necessarily synonymous with a “portion” as used herein. A “portion” as used herein refers to less than or equal to a full set.
[0105] As used herein, the term “subgenomic interval” (or “subgenomic sequence interval”) refers to a portion of a genomic sequence.
[0106] As used herein, the term “subject interval” refers to a subgenomic interval or an expressed subgenomic interval (e.g., the transcribed sequence of a subgenomic interval).
[0107] As used herein, “treatment” (and grammatical variations thereof such as “treat” or “treating”) refers to clinical intervention in an attempt to alter the natural course of the individual being treated, and can be performed either for prophylaxis or during the course of clinical pathology. Desirable effects of treatment include, but are not limited to, preventing occurrence or recurrence of disease, alleviation of symptoms, diminishment of any direct or indirect pathological consequences of the disease, preventing metastasis, decreasing the rate of disease progression, amelioration or palliation of the disease state, and remission or improved prognosis.
[0108] The term “tumor,” as used herein, refers to all neoplastic cell growth and proliferation, whether malignant or benign, and all pre-cancerous and cancerous cells and tissues. The terms “cancer,” “cancerous,” and “tumor” are not mutually exclusive as referred to herein. These terms refer to the presence of cells possessing characteristics typical of cancer-causing cells, such as uncontrolled proliferation, immortality, metastatic potential, rapid growth and proliferation rate, and certain characteristic morphological features. Cancer cells are often in the form of a tumor, but such cells can exist alone within an animal, or can be a non- tumorigenic cancer cell, such as a leukemia cell. These terms include a solid tumor, a soft tissue tumor, or a metastatic lesion. A “tumor cell” as used herein, refers to any tumor cell present in a tumor or a sample thereof. Tumor cells may be distinguished from other cells that may be present in a tumor sample, for example, stromal cells and tumor-infiltrating immune cells, using methods known in the art and / or described herein.
[0109] As used herein, the terms “variant sequence” or “variant” are used interchangeably and refer to a modified nucleic acid sequence relative to a corresponding “normal” or “wildtype” sequence, such as, for example, a converted nucleic acid sequence corresponding to an aberrant methylation signal.
[0110] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described. The description is presented to enable one of ordinary skill in the art to make and use the invention and is provided in the context of a patent application and its requirements. Various modifications to the described embodiments will be readily apparent to those persons skilled in the art and the generic principles herein may be applied to other embodiments. Thus, the present invention is not intended to be limited to the embodiment shown but is to be accorded the widest scope consistent with the principles and features described herein.III. DNA Methylation
[0111] DNA methylation is a type of epigenetic modification characterized by the addition of a methyl group to a DNA base, such as, for instance, the addition of a methyl group to the C- 5 position of a cytosine. 5-methyl-cytosine (5mC) and 5-hydroxymethyl-cytosine (5hmC) are the two major types of DNA methylation that occur in the mammalian genome, and are commonly seen in CpG contexts (cytosine nucleotide followed by a guanine nucleotide in the linear sequence of bases along its 5 '->3' direction). However, other types of methylation, suchas 3mC and 4mC, are also possible. Methylation in mammals occurs primarily at CpG sites, but may also occur less frequently at non-CpG sites (e.g., CHG or CHH sites, in which H represents A, T, or C).
[0112] During a typical workflow for the detection of cancer-associated methylation signals from, for example, MRD / early-stage cancer patient specimens (FIG. 1), a variety of Methyl- seq library amplicons are created, each representing unique patterns of DNA methylation of individual DNA fragments in the input DNA sample, depending on the methylation status of individual CpG sites in the original DNA molecule (FIG. 5, middle panel).
[0113] The present invention relates to improved methods of detecting the methylation state of a target locus.A. Methylation Pattern(s) of Interest
[0114] A methylation state of a genomic locus being queried using the methods described herein may be referred to as a methylation pattern of interest at a genomic locus, or simply a methylation pattern of interest. The methylation pattern of interest represents the original genomic sequence of the nucleic acids in a sample along with the epigenetic methylation state of any or all C residues present in the genomic sequence. Thus, a methylation pattern of interest refers to the presence or absence of methyl groups on, e.g., the C-5 position of cytosines in original genomic context of a genetic locus in question, as present in a sample. One or more methylation patterns of interest may be queried alone or in parallel using the methods described herein.
[0115] A methylation pattern of interest may include one or more of each or any of 3mC, 4mC, 5mC, and / or 5hmC sites. A methylation pattern of interest may have one or more of only 5mC sites. In some embodiments, the methylation pattern of interest may comprise methylation on every C residue on a genomic locus being queried. In contrast, the methylation pattern of interest may not comprise any methylated residues. For instance, in some embodiments, a methylation pattern of interest may be so hypomethylated compared to a control methylation pattern that no residues are methylated in the methylation pattern of interest — that is, the phrase “methylation pattern of interest” according to the methods described herein also encompasses the lack of methylation at a genomic locus, in addition to the presence of methylation to any degree.
[0116] The methylation pattern of interest may include one or more CpG sites. In some embodiments, the methylation pattern of interest comprises one CpG site. In some embodiments, the methylation pattern of interest comprises more than one CpG site. In some embodiments, the methylation pattern of interest comprises two CpG sites. In some embodiments, the methylation pattern of interest comprises more than two CpG sites. In some embodiments, the methylation pattern of interest comprises three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, fifteen, sixteen, seventeen, eighteen, nineteen, or twenty CpG sites. In some embodiments, the methylation pattern of interest comprises more than twenty CpG sites.
[0117] The methylation pattern of interest may have one or more non-CpG sites. In some embodiments, the methylation pattern of interest comprises one CHG site. In some embodiments, the methylation pattern of interest comprises more than one CHG site. In some embodiments, the methylation pattern of interest comprises two CHG sites. In some embodiments, the methylation pattern of interest comprises more than two CHG sites. In some embodiments, the methylation pattern of interest comprises three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, fifteen, sixteen, seventeen, eighteen, nineteen, or twenty CHG sites. In some embodiments, the methylation pattern of interest comprises more than twenty CHG sites.
[0118] In some embodiments, the methylation pattern of interest comprises one CHH site. In some embodiments, the methylation pattern of interest comprises more than one CHH site. In some embodiments, the methylation pattern of interest comprises two CHH sites. In some embodiments, the methylation pattern of interest comprises more than two CHH sites. In some embodiments, the methylation pattern of interest comprises three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, fifteen, sixteen, seventeen, eighteen, nineteen, or twenty CHH sites. In some embodiments, the methylation pattern of interest comprises more than twenty CHH sites.
[0119] The methylation pattern of interest may be compared to a different or “control” methylation pattern (e.g.. not the methylation pattern of interest) at the same genomic locus. For instance, the methylation pattern of interest may be compared to a previously known methylation pattern from one or more other samples and / or one or more other subjects. This previously known “control” methylation pattern may, for instance, not be known to be indicative of a disease or condition. The “control” methylation pattern may be known to notbe indicative of a particular disease or condition in question. Alternatively, “control” methylation pattern may be known to be indicative of one or more particular disease(s) or condition(s). The methylation pattern of interest may be considered “aberrant” compared to another methylation pattern, such as any of the different or “control” methylation patterns described above. In contrast, the methylation pattern of interest may match the “control” methylation pattern, and thus not be aberrant compared to the “control” methylation pattern.B. Methylation Signal(s) of Interest
[0120] A methylation pattern of interest corresponds to a methylation signal of interest (also described herein as a methylation signature, methylation signature of interest, aberrant methylation state, variant methylation state, variant, or the like) by way of the converted nucleic acids forming a nucleic acid pattern of interest as described below. That is, a methylation signal of interest may be read or inferred from a converted nucleic acid molecule (or sequence thereof) that it itself unmethylated. In this way, a converted nucleic acid can signal the methylation state of a nucleic acid molecule in a sample from a subject. A methylation signal of interest may be a hypermethylation signal or a hypomethylation signal. One or more methylation signals of interest, each corresponding to a methylation pattern of interest, may be queried alone or in parallel using the methods described herein.C. Hypomethylation Signal
[0121] A locus is considered to be hypomethylated when it exhibits less methylation at one or more nucleic acid residues compared to a “control” methylation pattern at the same locus. Thus, in some embodiments, a methylation pattern of interest my correspond to a hypomethylation signal. In some embodiments, a hypomethylation signal at a genomic locus may be indicative of a disease or condition. Thus, a methylation pattern of interest may correspond to a hypomethylation signal indicative of a disease or condition. For instance, in some embodiments, the methylation pattern of interest may lack one methylated CpG site compared to a different methylation pattern at the same genomic locus that is not indicative of a particular disease or condition, or compared to a different methylation pattern at the same genomic locus that is indicative of one or more disease(s) or condition(s). In some embodiments, the methylation pattern of interest lacks more than one methylated CpG sites as compared to a different methylation pattern at the same genomic locus that is notindicative of a particular disease or condition, or compared to a different methylation pattern at the same genomic locus that is indicative of one or more disease(s) or condition(s).
[0122] A nucleic acid bait designed to hybridize to a nucleic acid pattern of interest converted from a methylation pattern of interest corresponding to a hypomethylation signal as described herein may be referred to synonymously in various ways, such as, for example, as a hypomethylation- specific bait, a hypomethylation- specific nucleic acid bait, a hypo bait or hypo-bait, or the like. A double-stranded hypomethylation- specific nucleic acid bait may also be referred to as, for example, a hypo-allele-specific double- stranded bait, or similar.D. Hypermethylation Signal
[0123] A locus is considered to be hypermethylated when it exhibits more methylation at one or more nucleic acid residues compared to a “control” methylation pattern at the same locus. Thus, in some embodiments, a methylation pattern of interest my correspond to a hypermethylation signal. In some embodiments, a hypermethylation signal at a genomic locus may be indicative of a disease or condition. Thus, a methylation pattern of interest may correspond to a hypermethylation signal indicative of a disease or condition. For instance, in some embodiments, the methylation pattern of interest may have an additional methylated CpG site compared to a different methylation pattern at the same genomic locus that is not indicative of a particular disease or condition, or an additional methylated CpG site compared to a different methylation pattern at the same genomic locus that is indicative of one or more disease(s) or condition(s). Similarly, in some embodiments, the methylation pattern of interest may have more than one more methylated CpG sites than a different methylation pattern at the same genomic locus that is not indicative of a particular disease or condition, or more than one more methylated CpG sites than a different methylation pattern at the same genomic locus that is indicative of one or more disease(s) or condition(s).
[0124] A nucleic acid bait designed to hybridize to a nucleic acid pattern of interest converted from a methylation pattern of interest corresponding to a hypermethylation signal as described herein may be referred to synonymously in various ways, such as, for example, as a hypermethylation- specific bait, a hypermethylation- specific nucleic acid bait, a hyper bait or hyper-bait, or the like. A double-stranded hypermethylation- specific nucleic acid bait may also be referred to as, for example, a hyper- allele- specific double- stranded bait, or similar.E. Nucleic Acid Pattern(s) of Interest
[0125] During the conversion step discussed below, a methylation pattern of interest is converted to a nucleic acid pattern of interest. The nucleic acid pattern of interest corresponds to the genomic locus the methylation state of which is queried in the methods described herein. Thus, the converted nucleic acid molecules comprise a nucleic acid pattern of interest corresponding to the same genomic locus as the methylation pattern of interest. Information about the methylation pattern of interest is embedded into the nucleic acid sequence of the converted nucleic acid molecules. That is, the original nucleic acid molecules taken from the subject may contain literal methylated residues. However, during the conversion step, unmethylated cytosine residues are converted to thymine residues, while methylated cytosine residues remain as cytosines, meaning that a converted nucleic acid molecule may or may not contain literal methylated residues, but does nevertheless retain the information about which residues were and were not methylated in the original nucleic acid in the form of which C residues are converted and which are not, respectively, when compared to the known genomic sequence. Thus, a nucleic pattern of interest may correspond to or contain a methylation signal of interest (also described herein as a methylation signature, or methylation signature of interest) while not necessarily comprising methylated residues. For example, nucleic acid pattern of interest may comprise or correspond to a hypomethylation signature or a hypermethylation signature (also described herein as a hypomethylation signal or a hypermethylation signal, respectively), while not literally displaying a methylation pattern per se. In some instances, a nucleic acid pattern of interest may correspond to two adjacent genomic loci, each comprising a different methylation state. For instance, a nucleic acid pattern of interest may comprise a sequence corresponding to a genomic locus that is hypermethylated adjacent to a genomic locus that is hypomethylated, or vice versa.
[0126] When the methylation pattern of interest is on a double-stranded genomic locus (as pictured in, e.g., the hypermethylated and hypomethylated target genomic loci shown on the left side in FIG. 9B), conversion to a nucleic acid pattern of interest will produce four different nucleic acid strands (as pictured in, e.g., the amplicons shown in the second column from the left in FIG. 9B): (i) a forward strand of the converted nucleic acid pattern of interest (labeled “(+)” in FIG. 9B), corresponding to the forward strand of the target genomic locus; (ii) a reverse complement to strand (i) (labeled “(+RC)” in FIG. 9B); (iii) a reverse strand ofthe converted nucleic acid pattern of interest (labeled in FIG. 9B), corresponding to the reverse strand of the target genomic locus; and (iv) a reverse complement to strand (iii) (labeled “(-RC)” in FIG. 9B). Each strand of an amplicon that is produced from conversion of a methylation pattern of interest at a genomic locus (e.g., each of strands (i)-(iv) listed above) is considered to correspond to the methylation pattern of interest and to correspond to the target genomic locus. Thus, each nucleic acid that is produced from conversion of a methylation pattern of interest at a genomic locus is considered part of the nucleic acid pattern of interest. Further, the reverse complement of each nucleic acid that is produced from conversion of a methylation pattern of interest at a genomic locus is also considered part of the nucleic acid pattern of interest.
[0127] When the methylation pattern of interest is on a single- stranded locus (such as, e.g., a locus on 5mC RNA), conversion to a nucleic acid pattern of interest will produce two different nucleic acid strands: (i) a converted nucleic acid pattern of interest corresponding to the single-stranded target genomic locus; and (ii) a reverse complement to strand (i).
[0128] One or more nucleic acid patterns of interest may be queried alone or in parallel using the methods described herein. For instance, if there are a plurality of different methylation patterns of interest being queried in parallel, then there may also be a corresponding plurality of different nucleic acid pattens of interest in the converted nucleic acid molecules.
[0129] The bait designs described herein allow the methylation state(s) of the original nucleic acid molecules to be inferred from the nucleic acid sequence of the converted nucleic acids with accuracy and precision.F. Cytosine Conversion
[0130] Methylation patterns in the nucleic acid molecules from the sample from the subject may then be converted into converted nucleic acids via cytosine conversion (also referred to herein as “nucleotide conversion”, “methylation conversion”, and simply “conversion”). As part of this conversion, methylation patterns of interest that may be present in the nucleic acid molecules in the sample from the subject are converted into nucleic acid patterns of interest corresponding to the genomic locus. Performing a methylation conversion reaction may comprise, e.g., selectively converting non-methylated (also referred to herein as unmethylated) cytosine residues to uracil residues, or selectively converting methylated cytosine residues to uracil residues. Cytosine residues that have been converted into uracilresidues are then further converted to thymine residues by, e.g., amplifying converted nucleic acids. Thus, amplified converted nucleic acids (that contain, e.g., thymine residues in positions in which the nucleic acid molecules in the sample from the subject had, e.g., unmethylated cytosine residues), may also be referred to as converted nucleic acids.
[0131] The conversion step converts a methylation state or a methylation pattern into a methylation signal, as the converted nucleic acids may not retain actual methylated residues but do comprise information about the methylation pattern that was present in the preconverted nucleotides. Thus, converted nucleic acids are representative of a methylation state in that they comprise patterns associated with a methylation state.
[0132] Conversion may be performed by various methods known in the art, such as, e.g., a bisulfite reaction (see, e.g., Li et al. (2011), “DNA Methylation Detection: Bisulfite Genomic Sequencing Analysis”, Methods Mol. Biol. 791:11-21), enzymatic conversion reactions, or the like. For example, enzymatic deamination of non-methylated cytosine using APOBEC to form uracil can be performed using, e.g., the Enzymatic Methyl-seq Kit from New England BioLabs (Ipswich, MA) which uses prior treatment with ten-eleven translocation methylcytosine dioxygenase 2 (TET2) to oxidize 5-mC and 5-hmC, thereby providing greater protection of the methylated cytosine from deamination by APOBEC). Liu et al. (2019) recently described a bisulfite-free and base-level-resolution sequencing-based method, TET- Assisted Pyridine borane Sequencing (TAPS), for detection of 5mC and 5hmC. The method combines ten-eleven translocation methylcytosine dioxygenase (TET)-mediated oxidation of 5mC and 5hmC to 5-carboxylcytosine (5caC) with pyridine borane reduction of 5caC to dihydrouracil (DHU). Subsequent PCR amplification converts DHU to thymine, thereby enabling conversion of methylated cytosines to thymine (Liu et al. (2019), “Bisulfite-Free Direct Detection of 5 -Methylcytosine and 5-Hydroxymethylcytosine at Base Resolution”, Nature Biotechnology, vol. 37, pp. 424-429). In some embodiments, converting comprises TET-assisted bisulfite treatment, TET-assisted pyridine borane treatment, oxidative bisulfite treatment, APOBEC treatment, and / or other DNA deaminase treatment (e.g., as discussed in Vaisvila et al. (2023), “Discovery of novel DNA cytosine deaminase activities enables a nondestructive single-enzyme methylation sequencing method for base resolution high- coverage methylome mapping of cell-free and ultra-low input DNA”, bioRxiv, 2023.06.29.547047).
[0133] In some embodiments, conversion comprises enzymatic conversion. In some embodiments, conversion comprises bisulfite conversion. In some embodiments, the conversion comprises cytosine (C) to thymine (T) conversion. In some embodiments, the conversion converts non-methylated cytosine residues into thymine residues. In some embodiments, the conversion converts methylated cytosine residues into thymine residues.
[0134] In some embodiments, the conversion converts non-methylated cytosine residues into uracil residues. In some embodiments, the conversion converts methylated cytosine residues into uracil residues. In some embodiments, the conversion converts methylated cytosine residues into dihydrouracil residues. In some embodiments, the conversion converts non- methylated cytosine residues into dihydrouracil residues.
[0135] Bisulfite conversion or treatment refers to a biochemical process for converting unmethylated cytosine residue to uracil or thymine residues (e.g., deamination to uracil, followed by amplification as thymine during PCR), whereby methylated cytosine residues (e.g., 5-methylcytosine, 5mC; or 5-hydroxymethylcytosine, 5hmC) are preserved. Reagents to convert cytosine to uracil are known to those of skill in the art and include bisulfite reagents such as sodium bisulfite, potassium bisulfite, ammonium bisulfite, magnesium bisulfite, sodium metabisulfite, potassium metabisulfite, ammonium metabisulfite, magnesium metabisulfite and the like.
[0136] In some embodiments, the methods of the present disclosure comprise treating a plurality of nucleic acids or nucleic acid fragments of the present disclosure with enzymatic digestion and bisulfite treatment. The principle of the method is that the fragmentation of DNA is not achieved by ultrasound but achieved by combined enzymatic digestion by multiple endonucleases (Msel, Tsp 5091, Nlalll and Hpy CH4V), wherein the restriction enzyme cutting sites of Msel, Tsp509I, Nlalll and Hpy CH4V are TTAA, AATT, CATG and TGCA, respectively. See, e.g., Smiraglia D J et al. Oncogene 2002; 21: 5414-5426. This is followed by bisulfite treatment, e.g., as described herein.
[0137] Enzymatic methods for cytosine conversion are also known, e.g., enzymatic methyl sequencing (EM-seq). Such approaches can be advantageous because they employ enzymes instead of bisulfite, which can damage and fragment DNA, leading to DNA loss and potentially biased sequencing. For example, TET2 (the Ten-eleven translocation (Tet) family 2 methylcytosine dioxygenase) and T4-BGT (T4 phage beta-glucosyltransferase) can be used to convert 5mC and 5hmC into products that cannot be deaminated by APOBEC3A(apolipoprotein B mRNA editing enzyme, catalytic polypeptide-like 3A), then AP0BEC3A is used to deaminate unmodified cytosines by converting them into uracils. See, e.g., Vaisvila, R. et al. (2021) Genome Res. 31:1-10.
[0138] In some embodiments, the methods of the present disclosure comprise treating a plurality of nucleic acids or nucleic acid fragments of the present disclosure with TET- assisted bisulfite (e.g., TAB-seq). In the TAB-seq approach, beta-glucosyltransferase (PGT) is used to convert 5hmC into P-glucosyl-5-hydroxymethylcytosine (5gmC), and a Tet enzyme (e.g., mTetl) is used to oxidize 5mC into 5-carboxylcytosine (5caC). Subsequently, nucleic acids can be treated with bisulfite. See, e.g., Yu, M. et al. (2018) Methods Mol. Biol. 1708:645-663.
[0139] In some embodiments, the methods of the present disclosure comprise treating a plurality of nucleic acids or nucleic acid fragments of the present disclosure with TET- assisted pyridine borane (e.g., TAPS). In the TAPS approach, a TET methylcytosine dioxygenase is used to oxidize 5mC and 5hmC into 5caC, then 5caC is reduced into dihydrouracil (DHU) via pyridine borane. DHU is converted to thymine during subsequent PCR. See, e.g., Liu, Y. et al. (2019) Nat. Biotechnol. 37:424-429.
[0140] In some embodiments, the methods of the present disclosure comprise treating a plurality of nucleic acids or nucleic acid fragments of the present disclosure with oxidative bisulfite (e.g., oxBS). In the oxBS approach, 5hmC is oxidized into 5 -formylcytosine (5fC), which can be converted to uracil under bisulfite. Sequencing results from bisulfite vs. oxidative bisulfite treatment can then be used to infer 5hmC levels from 5mC. See, e.g., Booth, M.J. et al. (2013) Nat. Protocols 8:1841-1851. This approach can be scaled on a genome-wide level in oxBS-seq; see, e.g., Kirschner, K. et al. (2018) Methods Mol. Biol. 1708:665-678.
[0141] In some embodiments, the methods of the present disclosure comprise treating a plurality of nucleic acids or nucleic acid fragments of the present disclosure with APOBEC. Enzymatic reagents to convert cytosine to uracil, i.e. cytosine deaminases, include those of the APOBEC family, such as APOBEC-seq or APOBEC3A. The APOBEC family members are cytidine deaminases that convert cytosine to uracil while maintaining 5-methyl cytosine, i.e. without altering 5-methyl cytosine. Such enzymes are described in US2013 / 0244237 and WO2018165366 and are commercially available (see, e.g., the NEBNext® Enzymatic Methyl-seq Kit, New England Biolabs). Non-limiting examples of APOBEC family proteinsinclude AP0BEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3F, APOBEC3G, APOBEC3H, APOBEC4, and Activation-induced (cytidine) deaminase.G. Association with a disease or condition
[0142] A methylation pattern of interest may be indicative of or associated with a disease or condition. Similarly, a methylation pattern of interest may correspond to a methylation signal (e.g., a hypermethylation signal or a hypomethylation signal) that is indicative of or associated with a disease or condition (e.g., cancer). For instance, a methylation pattern of interest may be a biomarker of a disease or condition, such as a cancer biomarker. The disease or condition may be any disease or condition that is known to be associated with at least one hypermethylation pattern and / or at least one hypomethylation pattern. Thus, the presence of a disease or condition in a subject may be indicated by or correlated with at least one hypomethylation signal and / or at least one hypermethylation signal. In some embodiments, a hypermethylation signal at a genomic locus us indicative of a disease or condition. In some embodiments, a hypomethylation signal at a genomic locus us indicative of a disease or condition.
[0143] Exemplary diseases or conditions include but are not limited to cancer, genetic disorders (such as, for instance, Down Syndrome and Fragile X), neurological disorders, or any other disease type where detection of aberrant methylation states, e.g., one or more hypomethylation signals and / or one or more hypermethylation signals, are relevant to diagnosing, treating, or predicting said disease. For instance, in some embodiments, the disease or condition comprises a tissue-specific marker, a cancer, an auto-immune disease, an infectious disease, and / or a status of a transplanted tissue or organ. In some embodiments, the disease or condition is a tissue-specific marker, a cancer, an auto-immune disease, or a combination thereof. In some embodiments, the disease or condition is cancer.
[0144] In some embodiments, the condition may be a predisposition, such as a genetic predisposition, to develop a disease or condition, or may indicate a likelihood of developing a disease or condition. For instance, in some embodiments, the condition is a risk of having a cancer, or is associated with a risk of having cancer. In some embodiments, the condition may be a genetic predisposition to a cancer (e.g., having a genetic mutation that increases their baseline risk for developing a cancer). In some embodiments, the condition is theprogression and / or recurrence of a disease or condition, such as, e.g., cancer or tumor progression or recurrence.
[0145] In some embodiments, the disease or condition is a hyperproliferative disease. In some embodiments, the hyperproliferative disease is a cancer. In some embodiments, the cancer is a solid tumor or a metastatic form thereof. In some embodiments, the cancer is a hematological cancer, e.g., a leukemia or lymphoma. In some embodiments, the disease or condition is not a hyperproliferative disease.
[0146] In some embodiments, the same methylation pattern may be indicative of or associated with a plurality of different diseases or conditions. In some embodiments, a plurality of different methylation patterns may be associated with the same disease or condition.
[0147] In some embodiments, a methylation pattern as described herein may be used as a prognostic or diagnostic indicator associated with the sample. For example, in some instances, the prognostic or diagnostic indicator may comprise an indicator of the presence of a disease (e.g., cancer) in the sample, an indicator of the probability that a disease (e.g., cancer) is present in the sample, an indicator of the probability that the subject from which the sample was derived will develop a disease (e.g., cancer) (i.e., a risk factor), or an indicator of the likelihood that the subject from which the sample was derived will respond to a particular therapy or treatment.
[0148] In some embodiments, a methylation pattern of interest associated with hypermethylation or hypomethylation signal as described herein is indicative of responsiveness to a treatment, such as, for instance a cancer treatment. In some embodiments in which the subject has cancer, and in which the methylation pattern of interest is associated with a hypermethylation or hypomethylation signal indicative of responsiveness to cancer treatment, the methods described herein may further comprise predicting the subject’s responsiveness to cancer treatment based at least in part on the presence or absence of a methylation signal of interest, such as a hypermethylation signal or a hypomethylation signal at a locus known to be associated with responsiveness to treatment.
[0149] In some embodiments, the disease or condition is a cancer. Many cancers have profound epigenetic dysregulation that give rise to aberrant DNA methylation patterns, including hypermethylation and hypomethylation, which are distinct from healthy samples. These features can be useful in developing cancer diagnostic assays based on NGSsequencing for methylation analysis (Methyl-seq). Aberrant methylation is widespread in cancer and can be detected in many different types of patient samples, including those that comprise cfDNA. Detection of rare cancer-driven patterns is a key challenge for many liquid biopsy applications including, for example, detection and monitoring of minimal residual disease (MRD) and treatment response monitoring (TRM). Further, epigenetic data may be a valuable component of a comprehensive -omics-level analysis of a subject or a group of subject. Accordingly, data generated according to the methods described herein may be included in, for example, a comprehensive genomic profile (CGP).
[0150] Some methylation patterns in cancer are associated with or predictive of response to particular treatment regimens or disease management strategies. For example, in glioblastoma, promoter methylation in the gene MGMT has been associated with better outcomes (Lalezari et al. (2013) Neuro Oncol 15:370-381). Methylation-based studies could lead to discovery of new predictive biomarkers to guide therapy and drug development. Many late-stage cancer patients have higher levels of cancer signal in cfDNA; however, some patients have lower levels of cancer signal in cfDNA and could benefit from ultrasensitive detection of methylation levels such as those described herein. In addition, late-stage patients with the best response to treatment (chemotherapy, immunotherapy, targeted therapy, or some combination) have dramatic reduction of cancer signal observed in successive cfDNA samples just a few weeks into treatment (see, e.g., Davis, A.A. et al. (2020) Mol. Cancer Ther. 19:1486-1496; Hrebien, S. et al. (2019) Ann. Oncol. 30:945-952). Ultrasensitive detection of methylation levels may be useful, e.g., to continually monitor this subset of patients and detect recurrence as early as possible.
[0151] Exemplary cancers include, but are not limited to, B cell cancer (e.g., multiple myeloma), melanomas, breast cancer, lung cancer (such as non-small cell lung carcinoma or NSCLC), bronchus cancer, colorectal cancer, prostate cancer, pancreatic cancer, stomach cancer, ovarian cancer, urinary bladder cancer, brain or central nervous system cancer, peripheral nervous system cancer, esophageal cancer, cervical cancer, uterine or endometrial cancer, cancer of the oral cavity or pharynx, liver cancer, kidney cancer, testicular cancer, biliary tract cancer, small bowel or appendix cancer, salivary gland cancer, thyroid gland cancer, adrenal gland cancer, osteosarcoma, chondrosarcoma, cancer of hematological tissues, adenocarcinomas, inflammatory myofibroblastic tumors, gastrointestinal stromal tumor (GIST), colon cancer, multiple myeloma (MM), myelodysplastic syndrome (MDS),myeloproliferative disorder (MPD), acute lymphocytic leukemia (ALL), acute myelocytic leukemia (AML), chronic myelocytic leukemia (CML), chronic lymphocytic leukemia (CLL), polycythemia Vera, Hodgkin lymphoma, non-Hodgkin lymphoma (NHL), soft-tissue sarcoma, fibrosarcoma, myxosarcoma, liposarcoma, osteogenic sarcoma, chordoma, angiosarcoma, endotheliosarcoma, lymphangiosarcoma, lymphangioendotheliosarcoma, synovioma, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinomas, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatoma, bile duct carcinoma, choriocarcinoma, seminoma, embryonal carcinoma, Wilms' tumor, bladder carcinoma, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pinealoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, meningioma, neuroblastoma, retinoblastoma, follicular lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, hepatocellular carcinoma, thyroid cancer, gastric cancer, head and neck cancer, small cell cancers, essential thrombocythemia, agnogenic myeloid metaplasia, hypereosinophilic syndrome, systemic mastocytosis, familiar hypereosinophilia, chronic eosinophilic leukemia, neuroendocrine cancers, carcinoid tumors, and the like.
[0152] In some embodiments of any of the methods provided herein, the cancer is a carcinoma, a sarcoma, a lymphoma, a leukemia, a myeloma, a germ cell cancer, or a blastoma. In some embodiments, the cancer is a solid tumor. In some embodiments, the cancer is a hematologic malignancy. In some embodiments, the cancer is a B cell cancer, a melanoma, breast cancer, lung cancer, bronchus cancer, colorectal cancer, prostate cancer, pancreatic cancer, stomach cancer, ovarian cancer, urinary bladder cancer, brain cancer, central nervous system cancer, peripheral nervous system cancer, esophageal cancer, cervical cancer, uterine cancer, endometrial cancer, cancer of an oral cavity, cancer of a pharynx, liver cancer, kidney cancer, testicular cancer, biliary tract cancer, small bowel cancer, appendix cancer, salivary gland cancer, thyroid gland cancer, adrenal gland cancer, osteosarcoma, chondrosarcoma, a cancer of hematological tissue, an adenocarcinoma, an inflammatory myofibroblastic tumor, a gastrointestinal stromal tumor (GIST), colon cancer, multiple myeloma (MM), myelodysplastic syndrome (MDS), myeloproliferative disorder (MPD), acute lymphocytic leukemia (ALL), acute myelocytic leukemia (AML), chronic myelocytic leukemia (CML), chronic lymphocytic leukemia (CLL), polycythemia Vera, Hodgkinlymphoma, non-Hodgkin lymphoma (NHL), soft-tissue sarcoma, fibrosarcoma, myxosarcoma, liposarcoma, osteogenic sarcoma, chordoma, angiosarcoma, endotheliosarcoma, lymphangiosarcoma, lymphangioendotheliosarcoma, synovioma, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinomas, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatoma, bile duct carcinoma, choriocarcinoma, seminoma, embryonal carcinoma, Wilms' tumor, bladder carcinoma, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pinealoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, meningioma, neuroblastoma, retinoblastoma, follicular lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, hepatocellular carcinoma, thyroid cancer, gastric cancer, head and neck cancer, small cell cancer, essential thrombocythemia, agnogenic myeloid metaplasia, hypereosinophilic syndrome, systemic mastocytosis, familiar hypereosinophilia, chronic eosinophilic leukemia, neuroendocrine cancers, or a carcinoid tumor.
[0153] In some embodiments, the cancer is appendix adenocarcinoma, bladder adenocarcinoma, bladder urothelial (transitional cell) carcinoma, breast cancer not otherwise specified (NOS), breast carcinoma NOS, breast invasive ductal carcinoma (IDC), breast invasive lobular carcinoma (ILC), cervix squamous cell carcinoma (SCC), colon adenocarcinoma (CRC), esophagus adenocarcinoma, esophagus carcinoma NOS, esophagus squamous cell carcinoma (SCC), eye intraocular melanoma, gallbladder adenocarcinoma, gastroesophageal junction adenocarcinoma, intra-hepatic cholangiocarcinoma, kidney cancer NOS, liver hepatocellular carcinoma (HCC), lung cancer NOS, lung adenocarcinoma, lung large cell carcinoma, lung non-small cell lung carcinoma (NSCLC) NOS, lung small cell undifferentiated carcinoma, lung squamous cell carcinoma (SCC), ovary cancer NOS, pancreas cancer NOS, pancreas ductal adenocarcinoma, pancreatobiliary carcinoma, prostate cancer NOS, prostate acinar adenocarcinoma, prostate ductal adenocarcinoma, rectum adenocarcinoma (CRC), skin melanoma, small intestine adenocarcinoma, soft tissue sarcoma NOS, stomach adenocarcinoma NOS, unknown primary cancer NOS, unknown primary adenocarcinoma, unknown primary carcinoma (CUP) NOS, unknown primary neuroendocrine tumor, unknown primary squamous cell carcinoma (SCC), or uterus endometrial adenocarcinoma NOS.
[0154] In some instances, the cancer comprises acute lymphoblastic leukemia (Philadelphia chromosome positive), acute lymphoblastic leukemia (precursor B-cell), acute myeloid leukemia (FLT3+), acute myeloid leukemia (with an IDH2 mutation), anaplastic large cell lymphoma, basal cell carcinoma, B-cell chronic lymphocytic leukemia, bladder cancer, breast cancer (HER2 overexpressed / amplified), breast cancer (HER2+), breast cancer (HR+, HER2- ), cervical cancer, cholangiocarcinoma, chronic lymphocytic leukemia, chronic lymphocytic leukemia (with 17p deletion), chronic myelogenous leukemia, chronic myelogenous leukemia (Philadelphia chromosome positive), classical Hodgkin lymphoma, colorectal cancer, colorectal cancer (dMMR and MSI-H), colorectal cancer (KRAS wild type), cryopyrin- associated periodic syndrome, a cutaneous T-cell lymphoma, dermatofibrosarcoma protuberans, a diffuse large B-cell lymphoma, fallopian tube cancer, a follicular B-cell nonHodgkin lymphoma, a follicular lymphoma, gastric cancer, gastric cancer (HER2+), a gastroesophageal junction (GEJ) adenocarcinoma, a gastrointestinal stromal tumor, a gastrointestinal stromal tumor (KIT+), a giant cell tumor of the bone, a glioblastoma, granulomatosis with polyangiitis, a head and neck squamous cell carcinoma, a hepatocellular carcinoma, Hodgkin lymphoma, juvenile idiopathic arthritis, lupus erythematosus, a mantle cell lymphoma, medullary thyroid cancer, melanoma, a melanoma with a BRAF V600 mutation, a melanoma with a BRAF V600E or V600K mutation, Merkel cell carcinoma, multicentric Castleman's disease, multiple hematologic malignancies including Philadelphia chromosome-positive ALL and CML, multiple myeloma, myelofibrosis, a non-Hodgkin’s lymphoma, a nonresectable subependymal giant cell astrocytoma associated with tuberous sclerosis, a non-small cell lung cancer, a non-small cell lung cancer (ALK+), a non-small cell lung cancer (PD-L1+), a non-small cell lung cancer (with ALK fusion or ROS1 gene alteration), a non-small cell lung cancer (with BRAF V600E mutation), a non-small cell lung cancer (with an EGFR exon 19 deletion or exon 21 substitution (L858R) mutations), a non- small cell lung cancer (with an EGFR T790M mutation), ovarian cancer, ovarian cancer (with a BRCA mutation), pancreatic cancer, a pancreatic, gastrointestinal, or lung origin neuroendocrine tumor, a pediatric neuroblastoma, a peripheral T-cell lymphoma, peritoneal cancer, prostate cancer, a renal cell carcinoma, rheumatoid arthritis, a small lymphocytic lymphoma, a soft tissue sarcoma, a solid tumor (MSI-H / dMMR), a squamous cell cancer of the head and neck, a squamous non-small cell lung cancer, thyroid cancer, a thyroid carcinoma, urothelial cancer, a urothelial carcinoma, or Waldenstrom's macroglobulinemia.In some embodiments, the cancer is glioblastoma, breast cancer, bladder cancer, cervical cancer, colorectal cancer, hepatocellular carcinoma, lung cancer, ovarian cancer, esophageal cancer, pancreatic cancer, stomach cancer or prostate cancer.
[0155] In some instances, the cancer is a hematologic malignancy (or premaligancy). As used herein, a hematologic malignancy refers to a tumor of the hematopoietic or lymphoid tissues, e.g., a tumor that affects blood, bone marrow, or lymph nodes. Exemplary hematologic malignancies include, but are not limited to, leukemia (e.g., acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), chronic myelogenous leukemia (CML), hairy cell leukemia, acute monocytic leukemia (AMoL), chronic myelomonocytic leukemia (CMML), juvenile myelomonocytic leukemia (JMML), or large granular lymphocytic leukemia), lymphoma (e.g., AIDS-related lymphoma, cutaneous T-cell lymphoma, Hodgkin lymphoma (e.g., classical Hodgkin lymphoma or nodular lymphocyte-predominant Hodgkin lymphoma), mycosis fungoides, non-Hodgkin lymphoma (e.g., B-cell non-Hodgkin lymphoma (e.g., Burkitt lymphoma, small lymphocytic lymphoma (CLL / SLL), diffuse large B-cell lymphoma, follicular lymphoma, immunoblastic large cell lymphoma, precursor B -lymphoblastic lymphoma, or mantle cell lymphoma) or T-cell non- Hodgkin lymphoma (mycosis fungoides, anaplastic large cell lymphoma, or precursor T- lymphoblastic lymphoma)), primary central nervous system lymphoma, Sezary syndrome, Waldenstrom macroglobulinemia), chronic myeloproliferative neoplasm, Langerhans cell histiocytosis, multiple myeloma / plasma cell neoplasm, myelodysplastic syndrome, or myelodysplastic / myeloproliferative neoplasm.
[0156] In some embodiments, the disease or condition may be associated with one or more variants in one or more of the following genes: ABL1, ACVR1B, AKT1, AKT2, AKT3, ALK, ALOX12B, AMER1, APC, AR, ARAE, ARERP1, ARID1A, ASXL1, ATM, ATR, ATRX, AURKA, AURKB, AXIN1, AXL, BAP1, BARD1, BCL2, BCL2L1, BCL2L2, BCL6, BCOR, BCORL1, BCR, BRAE, BRCA1, BRCA2, BRD4, BRIP1, BTG1, BTG2, BTK, CALR, CARD11, CASP8, CBEB, CBL, CCND1, CCND2, CCND3, CCNE1, CD22, CD274, CD70, CD74, CD79A, CD79B, CDC73, CDH1, CDK12, CDK4, CDK6, CDK8, CDKN1A, CDKN1B, CDKN2A, CDKN2B, CDKN2C, CEBPA, CHEK1, CHEK2, CIC, CREBBP, CRKL, CSE1R, CSE3R, CTCE, CTNNA1, CTNNB1, CUL3, CUL4A, CXCR4, CYP17A1, DAXX, DDR1, DDR2, DIS3, DNMT3A, DOT1L, EED, EGER, EMSY (Cllorf30), EP300, EPHA3, EPHB1, EPHB4, ERBB2, ERBB3, ERBB4, ERCC4, ERG,ERRFH, ESRI, ETV4, ETV5, ETV6, EWSR1, EZH2, EZR, FAM46C, FANCA, FANCC, FANCG, FANCL, FAS, FBXW7, FGF10, FGF12, FGF14, FGF19, FGF23, FGF3, FGF4, FGF6, FGFR1, FGFR2, FGFR3, FGFR4, FH, FECN, FET1, FET3, FOXE2, FUBP1, GABRA6, GATA3, GATA4, GATA6, GID4 (C17orf39), GNA11, GNA13, GNAQ, GNAS, GRM3, GSK3B, H3F3A, HDAC1, HGF, HNF1A, HRAS, HSD3B1, ID3, IDH1, IDH2, IGF1R, IKBKE, IKZF1, INPP4B, IRF2, IRF4, IRS2, JAK1, JAK2, JAK3, JUN, KDM5A, KDM5C, KDM6A, KDR, KEAP1, KEF, KIT, KEHE6, KMT2A (MEE), KMT2D (MEE2), KRAS, ETK, EYN, MAF, MAP2K1, MAP2K2, MAP2K4, MAP3K1, MAP3K13, MAPK1, MCE1, MDM2, MDM4, MED12, MEF2B, MEN1, MERTK, MET, MITF, MKNK1, MEH1, MPE, MRE11A, MSH2, MSH3, MSH6, MST1R, MTAP, MTOR, MUTYH, MYB, MYC, MYCL, MYCN, MYD88, NBN, NF1, NF2, NFE2L2, NFKBIA, NKX2-1, NOTCH1, NOTCH2, NOTCH3, NPM1, NRAS, NT5C2, NTRK1, NTRK2, NTRK3, NUTM1, P2RY8, PALB2, PARK2, PARP1, PARP2, PARP3, PAX5, PBRM1, PDCD1, PDCD1LG2, PDGFRA, PDGFRB, PDK1, PIK3C2B, PIK3C2G, PIK3CA, PIK3CB, PIK3R1, PIM1, PMS2, POLDI, POLE, PPARG, PPP2R1A, PPP2R2A, PRDM1, PRKAR1A, PRKCI, PTCHI, PTEN, PTPN11, PTPRO, QKI, RAC1, RAD21, RAD51, RAD51B, RAD51C, RAD51D, RAD52, RAD54L, RAFI, RARA, RBI, RBM10, REL, RET, RICTOR, RNF43, ROS1, RPTOR, RSPO2, SDC4, SDHA, SDHB, SDHC, SDHD, SETD2, SF3B1, SGK1, SLC34A2, SMAD2, SMAD4, SMARCA4, SMARCB1, SMO, SNCAIP, SOCS1, SOX2, SOX9, SPEN, SPOP, SRC, STAG2, STAT3, STK11, SUFU, SYK, TBX3, TEK, TERC, TERT, TET2, TGFBR2, TIPARP, TMPRSS2, TNFAIP3, TNFRSF14, TP53, TSC1, TSC2, TYRO3, U2AF1, VEGFA, VHL, WHSCI, WHSC1L1, WT1, XPO1, XRCC2, ZNF217, or ZNF703 gene locus, or any combination thereof.
[0157] In some embodiments, the disease or condition may be associated with one or more variants in one or more of the following genes: ABL, ALK, ALL, B4GALNT1, BAFF, BCL2, BRAF, BRCA, BTK, CD19, CD20, CD3, CD30, CD319, CD38, CD52, CDK4, CDK6, CML, CRACC, CS1, CTLA-4, dMMR, EGFR, ERBB1, ERBB2, FGFR1-3, FLT3, GD2, HD AC, HER1, HER2, HR, IDH2, IL-ip, IL-6, IL-6R, JAK1, JAK2, JAK3, KIT, KRAS, MEK, MET, MSI-H, mTOR, PARP, PD-1, PDGFR, PDGFRa, PDGFRp, PD-L1, PI3K5, PIGF, PTCH, RAF, RANKL, RET, ROS1, SLAMF7, VEGF, VEGFA, or VEGFB gene locus, or any combination thereof.
[0158] In some instances, the disclosed methods for selectively capturing nucleic acid molecules may be used to select a subject (e.g., a patient) for a clinical trial based on the methylation status detected for one or more loci (e.g., CpG sites). In some instances, patient selection for clinical trials based on, e.g., identification of aberrant methylation at one or more gene loci, may accelerate the development of targeted therapies and improve the healthcare outcomes for treatment decisions.IV. Bait Design
[0159] The present disclosure describes double- stranded and / or single-stranded nucleic acid baits that are designed according to the principles laid out herein. As described herein, a bait is designed to comprise a target- specific capture sequence that selectively hybridizes to a nucleic acid sequence that corresponds to only one methylation state (e.g., either only a hypermethylation signal or only a hypermethylation signal) of a genomic locus, such that a given bait designed according to the principles described herein, whether single- stranded or double-stranded, is designed to be both specific to a genomic locus and specific to a methylation state. In this way, only a single methylation state (either only hypermethylation or only hypomethylation) is queried at a time per locus, though multiple loci could be queried in parallel. Thus, a suitable target genomic locus comprises one or more potential methylation sites. Further, once subjected to nucleotide conversion, a suitable target genomic locus that is hypermethylated produces a converted nucleic acid sequence that is not identical to a converted nucleic acid sequence produced from a hypomethylated version of the same genomic locus. Likewise, a suitable target genomic locus that is hypomethylated produces a converted nucleic acid sequence that is not identical to a converted nucleic acid sequence produced from a hypermethylated version of the same genomic locus.
[0160] The methods described herein involve at least one nucleic acid bait designed according to the principles described herein. The exact nucleotide sequence of a nucleic acid bait (whether designed to be hypomethylation- specific or hypermethylation- specific, and whether single- or double-stranded) is dependent on the sequence of the genomic locus being queried, such that at least a portion of the bait selectively hybridizes to a nucleic acid pattern of interest. Further, baits of the present invention may be, for instance, modified and / or extended beyond the region that selectively hybridizes to a nucleic acid pattern of interest as described in subsequent sections herein.
[0161] A bait molecule may be considered to selectively hybridize to a given nucleic acid molecule when at least one target- specific capture sequence of the bait molecule hybridizes without mismatches to the nucleic acid molecule along the length of the target locus against which the bait was designed (that is, perfect complementarity across the length of the nucleic acid pattern of interest). For instance, if a bait is designed with a 100 bp long target- specific capture sequence corresponding to 100 bp of a nucleic acid pattern of interest, then the bait selectively hybridizes to a nucleic acid molecule when the entire lOObp of the target- specific capture sequence of the bait hybridizes to the nucleic acid molecule without mismatches. A bait molecule may be considered to selectively hybridize to a given nucleic acid molecule when it hybridizes with fewer mismatches to the nucleic acid molecule compared to when it hybridizes to a non-target nucleic acid molecule. For example, for a hypomethylation-specific bait, the bait may be considered to selectively hybridize to a nucleic acid molecule when it hybridizes with fewer mismatches to the nucleic acid molecule compared to when it hybridizes to a nucleic acid molecule corresponding to a hypermethylation pattern at the same genomic locus. Correspondingly, a hypermethylation- specific bait may be considered to selectively hybridize to a nucleic acid molecule when it hybridizes with fewer mismatches to the nucleic acid molecule compared to when it hybridizes to a nucleic acid molecule corresponding to a hypomethylation pattern at the same genomic locus.
[0162] A bait set may be considered to selectively hybridize to a given nucleic acid molecule when at least one bait molecule in the bait set selectively hybridizes to the nucleic acid molecule. Thus, a bait set may be considered to selectively hybridize to a given nucleic acid molecule when, for instance, at least one target-specific capture sequence of at least one bait molecule in the bait set hybridizes without mismatches to the nucleic acid molecule along the length of the target locus against which the bait was designed, as outlined above, or with fewer mismatches compared to an off-target sequence.
[0163] A bait that selectively hybridizes to a nucleic acid sequence will not effectively hybridize to other sequences, such as non-target sequences, at least under high-stringency wash conditions. Thus, non-selective hybridization events could be washed out in, for instance, a wash step as described herein. For example, a nucleic acid bait that selectively hybridizes to a nucleic acid pattern of interest comprising a hypermethylation signature at a genomic locus will not hybridize to a nucleic acid pattern comprising a hypomethylation signature or a “normal” methylation signature at the same genomic locus, or would only doso with a reduced melting temperature (Tm). Similarly, a nucleic acid bait that selectively hybridizes to a nucleic acid pattern of interest comprising a hypomethylation signature at a genomic locus will not hybridize to a nucleic acid pattern comprising a hypermethylation signature or a “normal” methylation signature at the same genomic locus, or would only do so with a reduced melting temperature (Tm).
[0164] In some embodiments, a hybrid will contain a nucleic acid pattern of interest and a bait, and in which the nucleic acid pattern of interest is perfectly hybridized to all or part of a bait, such that there are no mismatched residues within the hybridized region. For example, a hybrid containing perfect hybridization between a bait and a nucleic acid pattern of interest will have complementary (e.g., A:T and G:C) pairing over the entire length of any portion of the bait that is designed to be complementary to the nucleic acid of interest and any portion of the nucleic acid of interest that is meant to be captured by a bait, without overhangs or any mismatched bases. However, in some embodiments, other sequences may hybridize and contain one or more mismatched nucleic acids. For instance, a region of hybridization that is not between a bait and a corresponding nucleic acid pattern of interest, if present, may comprise one or more mismatches.
[0165] The baits associated with the methods described herein may be “allele-specific” (also described herein as “methyl allele- specific”) nucleic acid baits or “strand- specific” (also described herein as “single-stranded” and “strand- selective”) nucleic acid baits. Unlike conventional baits, each of the baits described herein is designed to hybridize to a nucleotide- converted version of a specific methylation pattern of interest: either a hypermethylation pattern of interest or a hypomethylation pattern of interest, but not to both a hypomethylation pattern of interest and a hypomethylation pattern of interest. That is, each bait is designed to hybridize to a nucleic acid pattern of interest as described herein.
[0166] When a bait is designed to hybridize to a nucleic acid pattern of interest in a nucleic acid molecule that is converted from a methylation pattern of interest that corresponds to a hypermethylation signal, the bait may be referred to as, for example, a hypermethylationspecific bait. In contrast, when a bait is designed to hybridize to a nucleic acid pattern of interest in a nucleic acid molecule that is converted from a methylation pattern of interest that corresponds to a hypomethylation signal, the bait may be referred to as, for example, a hypomethylation- specific bait. Thus, a hypermethylation-specific allele- specific bait will hybridize to a converted nucleic acid corresponding to either strand of a hypermethylatedallele at a genomic locus of interest; correspondingly, a hypomethylation- specific allelespecific bait will hybridize to a converted nucleic acid corresponding to either strand of a hypomethylated allele at a genomic locus of interest. However, a hypermethylation- specific strand-specific bait will only hybridize to a converted nucleic acid corresponding to a strand of an amplicon corresponding to a hypermethylated allele (e.g., corresponding to one of: the + strand of a + amplicon corresponding to the + strand of the hypermethylated genomic locus; the reverse complement (RC) of the + strand of the + amplicon; the - strand of a - amplicon corresponding to the - strand of the hypermethylated genomic locus; or the RC of the - strand of the - amplicon) at a genomic locus of interest, while a hypomethylationspecific strand-specific bait will only hybridize to a converted nucleic acid corresponding to a strand of an amplicon corresponding to a hypomethylated allele (e.g., corresponding to one of: the + strand of a + amplicon corresponding to the + strand of the hypomethylated genomic locus; the reverse complement (RC) of the + strand of the + amplicon; the - strand of a - amplicon corresponding to the - strand of the hypomethylated genomic locus; or the RC of the - strand of the - amplicon) at a genomic locus of interest.A. Design Principles for Allele- Specific Baits
[0167] The allele- specific methyl enrichment baits described herein enhance detection of methylation signals (e.g., hypermethylation signals or hypomethylation signals) associated with a disease or condition (e.g., cancer). Current minimal residual disease (MRD yearly- stage cancer detection methods based on detection of cancer-associated methylation signals have low sensitivities. As described herein, selective enrichment for these signals by the allele- specific methyl enrichment baits employed in, for instance, target capture followed by next-generation sequencing (NGS), selectively enrich for cancer-associated methylation signals. These allele- specific baits improve the sensitivity of MRD / early- stage cancer detection assays by enhancing detection of cancer-associated methylation signals from MRD / early- stage cancer patient cfDNA samples, though the baits described herein are not limited to MRD applications. For example, the baits described herein may also be used in at least, for example, TRM / advanced- stage cancer patient ctDNA samples and / or as part of a comprehensive genomic profile (CGP), among other applications.
[0168] Current methods for detection of cancer-associated methylation signals from MRD / early- stage cancer patient specimens employ a workflow beginning with, for example,a cfDNA sample, followed by end repair and adaptor ligation, cytosine conversion, library PCR, target capture, sequencing, determining cancer hypermethylation and / or hypomethylation scores, and finally making cancer identity calls and determine tumor fraction (FIG. 1). During Methyl-seq library construction, the cytosine conversion and library PCR steps create variations in library sequences depending on the methylation status of the original DNA molecule in the sample. For example, methylation on cytosines in CpG contexts in the original DNA molecule from the sample protects the cytosine base from conversion and preserves its identity in NGS sequencing; unmethylated cytosines in CpG or CH (where H is any base except G) contexts are converted to uracil and subsequently sequenced as a thymine base in NGS sequencing.
[0169] Conventional targeted methylation sequencing involves capturing both hypermethylated and hypomethylated DNA fragments from a given genomic locus with similar efficiencies. In conventional methods, for the same genomic region, baits designed to match hypermethylated sequences after conversion and baits designed for hypomethylated sequences after conversion are mixed in the hybridization reaction. In contrast, a single set of allele- specific baits described herein is either hypomethylation- specific or hypermethylationspecific, such that the baits in a single set of baits selectively hybridize to a nucleic acid pattern of interest converted from either a methylation pattern of interest that corresponds to a hypomethylation signal or a methylation pattern of interest that corresponds to a hypermethylation signal, respectively, but not both.
[0170] An allele- specific bait as described herein is double stranded, as shown in, for example, FIG. 9B. Thus, an allele- specific bait set as described herein comprises an even number of strands. As shown in the column labeled “Methyl-allele specific baits” in FIG. 9B, an allele- specific bait set designed for a methylation pattern of interest (either hypermethylation specific baits or hypomethylation specific baits) at a double- stranded genomic locus may comprise, for example, four different sequences for hybridization to an amplicon (that is, four different target- specific capture sequences within the same bait set): (i) a strand comprising a sequence corresponding to the forward strand of the converted nucleic acid pattern of interest (labeled “(+)” in FIG. 9B); (ii) a reverse complement to strand (i) (labeled “(+RC)” in FIG. 9B); (iii) a strand comprising a sequence corresponding to the reverse strand of the converted nucleic acid pattern of interest (labeled “(-)” in FIG. 9B); and (iv) a reverse complement to strand (iii) (labeled “(-RC)” in FIG. 9B), meaning thathybridization with an allele- specific bait set may comprise hybridization with, for example, baits comprising four different sequences for hybridization to a nucleic acid pattern of interest. Each of the four sequences (i)-(iv) listed above is thus designed to hybridize without mismatches to an amplicon sequence corresponding to a target (i.e., to one of the four strands of a nucleic acid pattern of interest). When at least one of sequences (i)-(iv) in a bait set hybridizes without mismatches to a nucleic acid sequence of a nucleic acid pattern of interest, the bait set may be said to selectively hybridize to the nucleic acid pattern of interest. In some embodiments, when at least one of sequences (i)-(iv) in a bait set hybridizes with one or more mismatches to a nucleic acid sequence of a nucleic acid pattern of interest, the bait set may be said to selectively hybridize to the nucleic acid pattern of interest, as long as there are fewer mismatches than are present in non-target hybrids.
[0171] Each of the different target- specific capture sequences may be present on different bait molecules within the same bait set, such that a given bait set may comprise bait molecules with different target- specific capture sequences that all correspond to the same genomic locus.
[0172] In some embodiments, sequences (i) and (ii) listed above are present in an allelespecific bait set. That is, in some embodiments, a double-stranded bait set may target just sequences corresponding to the + strand of the genomic locus of interest. In some embodiments, sequences (iii) and (iv) listed above are present in an allele- specific bait set. That is, in some embodiments, a double-stranded bait set may target just sequences corresponding to the - strand of the genomic locus of interest. In some embodiments, sequences (i), (ii), (iii), and (iv) listed above are present in an allele-specific bait set. That is, in some embodiments, a double- stranded bait set may target sequences corresponding to both the + and the - strands of the genomic locus of interest. In some embodiments, a bait set targeting just the + strand or just the - strand of the genomic locus of interest may exhibit reduced capture efficiency compared to a bait set targeting sequences corresponding to both the + and the - strands of the genomic locus of interest.
[0173] A hypermethylation-specific allele- specific bait can be used in embodiments described herein comprising, for example, selectively capturing nucleic acid molecules from a sample from a subject using methods involving: converting unmethylated cytosine (C) in one or more of the nucleic acid molecules to uracil (U), or converting methylated C in one or more of the nucleic acid molecules to U, to form converted nucleic acid molecules;amplifying the converted nucleic acid molecules into a set of amplicons; hybridizing a set of nucleic acid baits to the set of amplicons to form a plurality of hybrids, wherein the set of nucleic acid baits selectively hybridizes to a nucleic acid pattern of interest, wherein the nucleic acid pattern of interest corresponds to a methylation pattern of interest at a genomic locus in the nucleic acid molecules prior to the converting, wherein the methylation pattern of interest corresponds to a hypermethylation signal indicative of a disease or condition, and wherein the nucleic acid baits are double-stranded; and separating the plurality of hybrids from nucleic acid molecules that are not hybridized to a bait. In some such embodiments, the set of nucleic acid baits does not comprise a nucleic acid bait that selectively hybridizes to a nucleic acid pattern that corresponds to a hypomethylation signal at the genomic locus.
[0174] A hypomethylation- specific allele- specific bait could be used in embodiments described herein comprising, for example, selectively capturing nucleic acid molecules from a sample from a subject using methods involving: converting unmethylated cytosine (C) in one or more of the nucleic acid molecules to uracil (U), or converting methylated C in one or more of the nucleic acid molecules to U, to form converted nucleic acid molecules; amplifying the converted nucleic acid molecules into a set of amplicons; hybridizing a set of nucleic acid baits to the set of amplicons to form a plurality of hybrids, wherein the set of nucleic acid baits selectively hybridizes to a nucleic acid pattern of interest, wherein the nucleic acid pattern of interest corresponds to a methylation pattern of interest at a genomic locus in the nucleic acid molecules prior to the converting, wherein the methylation pattern of interest is associated with a hypomethylation signal indicative of a disease or condition, and wherein the nucleic acid baits are double-stranded; and separating the plurality of hybrids from nucleic acid molecules that are not hybridized to a bait. In some such embodiments, the set of nucleic acid baits does not comprise a nucleic acid bait that selectively hybridizes to a nucleic acid pattern that corresponds to a hypermethylation signal at the genomic locus.B. Design Principles for Single- Stranded Baits
[0175] The single- stranded (ss) strand- specific methyl enrichment baits described herein enable further enrichment of methylation signals beyond the allele- specific double-stranded baits described above. The single-stranded, strand- specific baits described herein reduce off- target binding by being designed to 1) query only a single strand at a time (such as, for instance, the baits shown in the far right column of FIG. 9B), and optionally also 2) favorpotential off-target A:C mismatches over potential off-target G:T mismatches (such as, for instance, the baits shown in FIG. 9A).
[0176] A double- stranded methyl allele- specific bait set designed for a methylation pattern of interest at a double- stranded genomic locus as described above and as shown in the column labeled “Methyl- allele specific baits” in FIG. 9B may comprise, for example, four different target- specific capture sequences within the same bait set: (i) a strand comprising a sequence corresponding to the forward strand of the converted nucleic acid pattern of interest (labeled “(+)” in FIG. 9B); (ii) a reverse complement to strand (i) (labeled “(+RC)” in FIG. 9B); (iii) a strand comprising a sequence corresponding to the reverse strand of the converted nucleic acid pattern of interest (labeled “(-)” in FIG. 9B); and (iv) a reverse complement to strand (iii) (labeled “(-RC)” in FIG. 9B). In contrast, a single- stranded or “strand- selective” bait set as described herein comprises only one of the four different target- specific capture sequences described above.
[0177] The double-stranded, allele- specific baits described above, while still representing an improvement over conventional methods, can lead to off-target capture of unintended amplicons due to off-target binding between, for example, hypermethylated targets and double-stranded hypomethylation- specific baits, or between hypomethylated targets and double-stranded hypermethylation- specific baits (see, e.g., the general scheme laid out in FIG. 7). These off-target binding events form target-bait duplexes characterized by either adenine / cytosine mismatches (A:C; also referred to herein as A / C) or guanine / thymine mismatches (G:T; also referred to as G / T), depending on the corresponding strand direction (plus or minus) of the original DNA molecule (FIG. 7).
[0178] A single-stranded or “strand-selective” bait set as described herein may be optimized to reduce off-target binding by further refining the single- stranded bait design as described below. Due to the similarity of the hyper- and hypo-methylated DNA sequences after cytosine conversion, there may be relatively few mismatches between a bait and an unintended target, meaning that targets with unintended methylation status could still be enriched (FIG. 7). For example, a bait designed to be hypomethylation- specific could be hybridized with a “mismatched” sequence from a hypermethylated target region containing a few nucleotides’ mismatches, and vice versa. A:C or G:T mismatches form when, for instance, hypo-allele-specific double-stranded baits cross-bind off-target with hypermethylated amplicons, or when hyper- allele- specific double-stranded baits cross-bindoff-target with hypomethylated amplicons (FIG. 7). The strand- specific baits disclosed herein may be further designed to avoid such mismatch situations.
[0179] To design such strand- specific baits that avoid such mismatch situations, the four different target- specific capture sequences described above can be separated into two groups based on the mismatch status of unintended hybridization: an A:C mismatch group and a G:T mismatch group (see, e.g., the strands labeled “AC” and “GT” along the right side of FIG. 9B). A single-stranded bait set optimized to avoid G:T mismatches as described herein can comprise any of the “A:C” mismatch strands. One or both of the two A:C mismatch DNA strands for each targeted region can then be produced (e.g., by synthesis) as strand- selective (single-stranded) methyl enrichment baits (see FIG. 9A).
[0180] A nucleic acid hybrid with A:C mismatches has a lower melting temperature (Tm) than a hybrid with G:T mismatches. Thus, off-target hybrids containing A:C mismatches can be preferentially depleted in post-hybridization washes. As such, the strand- selective baits described herein that are designed to form only A:C mismatches with sequences corresponding to an unintended methylation signal on the target locus can offer further improved selectivity in methyl target capture. Thus, for further optimizing strand-selective bait designs, off-target A:C mismatches are preferred over off-target G:T mismatches.
[0181] Thus, a single-stranded, strand-specific bait of the present invention may be designed to query only a single strand at a time without regard for mismatches. This could be similar to, for instance, the general scheme laid out in FIG. 7, but in which each strand of the displayed hypo baits and hyper baits is considered to be a separate bait, such that the baits are each considered to be single- stranded (as shown in, for instance, the far right column of FIG. 9B), and in which bait strands that are complementary to each other are not included together in the same reaction. For instance, a hypermethylation- specific single-stranded bait could be used in embodiments described herein involving querying a nucleic acid pattern of interest corresponding to only one strand of a genomic locus of interest. This may entail, for example, selectively capturing nucleic acid molecules from a sample from a subject using methods involving: converting unmethylated cytosine (C) in one or more of the nucleic acid molecules to uracil (U), or converting methylated C in one or more of the nucleic acid molecules to U, to form converted nucleic acid molecules; amplifying the converted nucleic acid molecules into a set of amplicons; hybridizing a set of nucleic acid baits to the set of amplicons to form a plurality of hybrids, wherein at least a portion of the hybrids in the plurality of hybridscomprise a nucleic acid molecule comprising a nucleic acid pattern of interest, wherein the nucleic acid pattern of interest corresponds to a methylation pattern of interest at a genomic locus in the nucleic acid molecules prior to the converting, wherein the methylation pattern of interest corresponds to a hypermethylation signal indicative of a disease or condition, wherein the set of nucleic acid baits selectively hybridizes to the nucleic acid pattern of interest of the genomic locus, and wherein the nucleic acid baits are single- stranded; and separating the plurality of hybrids from nucleic acid molecules that are not hybridized to a bait. In some such embodiments, the set of nucleic acid baits may not comprise a nucleic acid bait that selectively hybridizes to a nucleic acid pattern that corresponds to a hypomethylation signal at the genomic locus (e.g., off-target binding). However, in some embodiments, there may be one or more A:C mismatches and / or G:T mismatches between a bait nucleic acid and an off-target nucleic acid that may nonetheless allow a bait to hybridize with an off-target nucleic acid, which could lead to a false signal. Further, a hypermethylation- specific singlestranded bait may be designed to selectively hybridize to a nucleic acid pattern of interest of the genomic locus (e.g., on-target binding) with somewhat imperfect complementarity (e.g., one or up to a few mismatches), such that nucleic acid molecules comprising the nucleic acid pattern of interest are still selectively captured over other (e.g., non-target) sequences despite not necessarily having 100% complementarity in the region of hybridization between the bait and the nucleic acid pattern of interest.
[0182] Similarly, a hypomethylation- specific single- stranded bait of the present invention could be used in embodiments described herein involving querying a nucleic acid pattern of interest corresponding to only one strand of a genomic locus of interest. This may entail, for example, selectively capturing nucleic acid molecules from a sample from a subject using methods involving: unmethylated cytosine (C) in one or more of the nucleic acid molecules to uracil (U), or converting methylated C in one or more of the nucleic acid molecules to U, to form converted nucleic acid molecules; amplifying the converted nucleic acid molecules into a set of amplicons; hybridizing a set of nucleic acid baits to the set of amplicons to form a plurality of hybrids, wherein at least a portion of the hybrids in the plurality of hybrids comprise a nucleic acid molecule comprising a nucleic acid pattern of interest, wherein the nucleic acid pattern of interest corresponds to a methylation pattern of interest at a genomic locus in the nucleic acid molecules prior to the converting, wherein the methylation pattern of interest corresponds to a hypomethylation signal indicative of a disease or condition, whereinthe set of nucleic acid baits selectively hybridizes to the nucleic acid pattern of interest of the genomic locus and wherein the nucleic acid baits are single-stranded; and separating the plurality of hybrids from nucleic acid molecules that are not hybridized to a bait. In some such embodiments, the set of nucleic acid baits may not comprise a nucleic acid bait that selectively hybridizes to a nucleic acid pattern that corresponds to a hypermethylation signal at the genomic locus (e.g., off-target binding). However, similarly to the above, in some embodiments, there may be one or more A:C mismatches and / or G:T mismatches between a bait nucleic acid and an off-target nucleic acid that may nonetheless allow a bait to hybridize with an off-target nucleic acid, which could lead to a false signal. Further, a hypomethylation- specific single-stranded bait may be designed to selectively hybridize to a nucleic acid pattern of interest of the genomic locus (e.g., on-target binding) with somewhat imperfect complementarity (e.g., one or up to a few mismatches), such that nucleic acid molecules comprising the nucleic acid pattern of interest are still selectively captured over other (e.g., non-target) sequences despite not necessarily having 100% complementarity in the region of hybridization between the bait and the nucleic acid pattern of interest.C. A:C Mismatches, Melting Temperatures, and Bait Modifications
[0183] Alternatively, a single- stranded nucleic acid bait (whether intended to be hypermethylation- specific or hypomethylation specific) of the present invention may be designed to have perfect complementarity (z.e., hybridization without mismatches) over the nucleic acid pattern of interest, but to allow hybridization with one or more off-target A:C mismatches, while not allowing hybridization with off-target G:T mismatches (as shown in, for example, FIG. 9A, in which axemplary A:C mismatched residues are indicated along the ss baits with asterisks). In the context of the exemplary schematic shown in FIG. 9B, this could entail, for instance, using the single-strand baits labeled with “AC” on the far right, but not those labeled with “GT” on the far right. In FIG. 9B, strand- specific baits of A:C and G:T subtypes (which hybridize with A:C or G:T mismatches, respectively, to a nucleic acid sequence corresponding to a converted nucleic acid sequence of a non-target methylation pattern at the same genomic locus) are indicated along the left.
[0184] In a mixed pool of hybrids, in which some hybrids contain mismatches and others do not, hybrids containing mismatches can be preferentially depleted through washes performed based on the respective melting temperatures of the mismatched vs. not-mismatched hybrids.For instance, as shown in FIG. 8B and discussed in the Examples below, hybrids between a bait and an off-target strand with an A:C mismatch (see, for example, the second plot from the left in FIG. 8B) have a lower Tm than hybrids between a bait and an off-target strand with a G:T mismatch (see, for example, the far-right plot in FIG. 8B), both of which have lower Tms than perfectly -matched on-target hybrids (see, for example, the left-most plot and second-from-right plot in FIG. 8B; in FIG. 8B, the left-most data set corresponds to hyperspecific baits hybridized to a reverse complement of a strand corresponding to a DNA strand that was hypermethylated in the original sample, showing a perfect match with G:C binding. The second data set from the left corresponds to hyper- specific baits hybridized to a reverse complement of a strand corresponding to a DNA strand that was hypomethylated in the original sample, showing an A:C mismatch. The second data set from the right corresponds to hypo-specific baits hybridized to a reverse complement of a strand corresponding to a DNA strand that was hypomethylated in the original sample, showing a perfect match with A:T binding. The right-most data set corresponds to hypo-specific baits hybridized to a reverse complement of a strand corresponding to a DNA strand that was hypermethylated in the original sample, showing a G:T mismatch. The whiskers extend to 1.5 times the interquartile range).
[0185] Thus, although embracing off-target mismatches seems counterproductive in bait design, the intentional A:C mismatches described herein have the surprising advantage of further improving signal-to-noise by allowing depletion of off-target, A:C mismatched hybrids via washes at one or more temperatures below the Tm of the perfectly matched hybrids but above the Tm of the A:C mismatched hybrids. At such temperatures, the perfectly matched (on-target) hybrids will remain annealed, but the off-target hybrids with A:C mismatches will dehybridize, allowing them to be preferentially depleted. This has the effect of further enriching the pool of hybrids for one or more nucleic acid pattern(s) of interest.
[0186] For instance, a single- stranded nucleic acid bait of the present invention could be used in embodiments described herein involving querying a nucleic acid pattern of interest corresponding to only one strand of a genomic locus of interest. This may entail, for example, selectively capturing nucleic acid molecules using methods involving: converting a methylation pattern of interest at a genomic locus in nucleic acid molecules from a sample from a subject into a nucleic acid pattern of interest corresponding to the genomic locus to form converted nucleic acid molecules; amplifying the converted nucleic acid molecules intoa set of amplicons; hybridizing a set of single-stranded nucleic acid baits to the set of amplicons to form a plurality of hybrids. The plurality of hybrids may comprise, for example, at least a first portion of hybrids and a second portion of hybrids. In some embodiments, the first portion of hybrids in the plurality of hybrids comprises hybrids between a bait and a nucleic acid strand from an amplicon hybridized to each other without mismatches (or with fewer mismatches compared to an off-target hybrid), while the second portion of hybrids comprises hybrids between a bait and a nucleic acid strand from an amplicon hybridized to each other with A:C mismatches (but not G:T mismatches). For instance, a hybrid in the first portion of hybrids could be between: i) a single- stranded nucleic acid bait from the set of single- stranded nucleic acid baits, and ii) a nucleic acid molecule comprising the nucleic acid pattern of interest, wherein i) and ii) are hybridized to each other without mismatches (or with fewer mismatches compared to an off-target hybrid) at least over the nucleic acid pattern of interest. In contrast, a hybrid in the second portion of hybrids could be between: i) a singlestranded nucleic acid bait from the set of single- stranded nucleic acid baits, and ii) and a nucleic acid molecule that does not comprise the nucleic acid pattern of interest (e.g., from an off-target amplicon), in which i) and ii) are hybridized to each other with one or more A:C mismatches. In some embodiments, a hybrid in the second portion of hybrids comprises more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more than 20 A:C mismatches. In some embodiments, a hybrid in the second portion of hybrids comprises 5 or fewer, 10 or fewer, 15 or fewer, 20 or fewer, 25 or fewer, or 30 or fewer A:C mismatches. In some embodiments, a hybrid in the second portion of hybrids comprises 20 or fewer A:C mismatches.
[0187] When at least a portion of the hybrids in the plurality of hybrids comprise a singlestranded nucleic acid bait from the set of single- stranded nucleic acid baits and a nucleic acid molecule comprising the nucleic acid pattern of interest hybridized to each other without mismatches (or with fewer mismatches compared to an off-target hybrid; e.g., the first portion of hybrids discussed above), the plurality of hybrids may be then separated from nucleic acid molecules that are not hybridized to a bait, and the first and second portions of hybrids may also be separated from each other, via one or more wash steps. In such embodiments, the methylation pattern of interest could correspond to a hypermethylation signal indicative of a disease or condition or a hypomethylation signal indicative of a disease or condition.
[0188] A wash step to remove hybrids comprising A:C mismatches enriches the pool of hybrids for nucleic acid pattem(s) of interest that bind to a bait without mismatches (or with fewer mismatches compared to an off-target hybrid). During such a wash, nucleic acid molecules that are not hybridized to a bait and nucleic acid molecules in the second portion of hybrids discussed above are depleted from, e.g., the substrate. Thus, in some embodiments, one or more washes may comprise exposing the hybrids to one or more temperatures lower than the Tm of the hybrids comprising a bait hybridized without mismatches (or with fewer mismatches compared to an off-target hybrid) to a nucleic acid pattern of interest, but higher than the Tm of the hybrids comprising A:C mismatches, such that hybrids comprising A:C mismatches would melt while hybrids not comprising mismatches would remain intact, allowing the sequences that had been hybridized with A:C mismatches to be washed away along with any other unhybridized nucleic acids, thus removing or depleting hybrids comprising A:C mismatches.
[0189] In some embodiments, none of the hybrids in the plurality of hybrids comprises a G:T mismatch. For instance, in some embodiments, neither the first portion of hybrids nor second portion of hybrids discussed above comprises a G:T mismatch. In some embodiments, some of the hybrids (e.g., one or more portions of the hybrids) in the plurality of hybrids comprise one or more G:T mismatches. In some embodiments, a hybrid comprises more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more than 20 A:C mismatches. In some embodiments, a hybrid comprises 5 or fewer, 10 or fewer, 15 or fewer, 20 or fewer, 25 or fewer, or 30 or fewer A:C mismatches. In some embodiments, a hybrid comprises 20 or fewer A:C mismatches.
[0190] Accordingly, certain aspects of the present invention involve determining one or more melting temperatures (Tms) of hybridized nucleic acids. Various methods of determining a Tm are known in the art. For example, the methods described herein may involve determining a Tm for a first hybrid comprising a bait hybridized to a nucleic acid pattern of interest without mismatches (or with fewer mismatches compared to an off-target hybrid), such as, for example, a hyper- specific bait hybridized without mismatches (or with fewer mismatches compared to an off-target hybrid) to a reverse complement of a strand corresponding to a DNA strand that was hypermethylated in the original sample. In some embodiments, the methods described herein may further involve determining a second Tm for a second hybrid comprising a bait and an off-target nucleic acid, in which the hybridizedregion between the bait comprises one or more mismatched nucleic acids, such as, for example, one or more A:C and / or one or more G:T mismatches. For example, in some embodiments, the methods described herein may comprise determining a Tm of a hybrid comprising a hyper- specific bait hybridized with an A:C mismatch to a reverse complement of a strand corresponding to a DNA strand that was hypomethylated in the original sample, and / or determining a Tm of a hybrid comprising a hypo-specific bait hybridized to a reverse complement of a strand corresponding to a DNA strand that was hypermethylated in the original sample (as shown in, e.g., FIG. 8A, in which the left-most data set corresponds to hyper- specific baits hybridized to a reverse complement of a strand corresponding to a DNA strand that was hypermethylated in the original sample, showing a perfect match with G:C binding. The second data set from the left corresponds to hyper- specific baits hybridized to a reverse complement of a strand corresponding to a DNA strand that was hypomethylated in the original sample, showing an A:C mismatch. The second data set from the right corresponds to hypo-specific baits hybridized to a reverse complement of a strand corresponding to a DNA strand that was hypomethylated in the original sample, showing a perfect match with A:T binding. The right-most data set corresponds to hypo-specific baits hybridized to a reverse complement of a strand corresponding to a DNA strand that was hypermethylated in the original sample, showing a G:T mismatch. The whiskers extend to 1.5 times the interquartile range).
[0191] For example, in some embodiments comprising a first portion of hybrids comprising hybrids between a bait and a nucleic acid strand from an amplicon hybridized to each other without mismatches, and a second portion of hybrids comprising hybrids between a bait and a nucleic acid strand from an amplicon hybridized to each other with A:C mismatches, the method may further comprise separating the first portion of hybrids from the second portion of hybrids.
[0192] In some embodiments, separating the first portion of hybrids from the second portion of hybrids and / or separating the plurality of hybrids from nucleic acid molecules that are not hybridized to a bait comprises one or more wash steps. In some embodiments, separating the first portion of hybrids from the second portion of hybrids, and separating the plurality of hybrids from nucleic acid molecules that are not hybridized to a bait may be done simultaneously in a single wash step. In some embodiments, separating the first portion of hybrids from the second portion of hybrids, and separating the plurality of hybrids fromnucleic acid molecules that are not hybridized to a bait may be done sequentially in more than one wash step. Different wash steps may be performed at the same temperature.Alternatively, different wash steps may be performed at different temperatures. In some embodiments, the temperature of a wash step is set based on the determined Tm of a hybrid. In some embodiments, the temperature of a wash step is set to be higher than the determined Tm of a hybrid, such as, for instance, higher than the Tm of a hybrid comprising one or more A:C mismatches. In some embodiments, the temperature of a wash step is set to be lower than the determined Tm of a hybrid, such as, for instance, lower than the Tm of a hybrid between a bait and a nucleic acid pattern of interest. In some embodiments, the temperature of a wash step is set to be lower than the determined Tm of a hybrid between a single stranded bait and a nucleic acid pattern of interest hybridized to each other without mismatches. In some embodiments, the temperature of a wash step is set to be lower than the determined Tm of a hybrid between a single stranded bait and a nucleic acid pattern of interest hybridized to each other with fewer mismatches compared to a hybrid comprising a bait and an off-target sequence.
[0193] In some embodiments, hybrids are attached to a substrate before one or more wash steps, such that nucleic acid molecules that are unhybridized or become unhybridized during a wash step are depleted from the substrate. Exemplary means of attaching a hybrid to a substrate may include, for example, biotinylating the baits and coating a substrate (such as, for instance, a bead) with streptavidin.
[0194] In some embodiments, a bait may be longer than the nucleic acid pattern of interest, such that only a portion of the bait is complementary to the nucleic acid pattern of interest. A portions of a bait that is not complementary to the nucleic acid pattern of interest may comprise other sequences, such as, for instance, one or more synthetic adapters, primer recognition sequences, and / or barcodes (e.g., amplification primers, sequencing adapters, flow cell adapters, substrate adapters, sample barcodes or indexes, and / or unique molecular identifier sequences). In some embodiments, a bait molecule may further comprise a universal tail on one or both ends. Further potential characteristics of or modifications to bait molecules are described below in the section on target capture reagents.
[0195] In some embodiments, a nucleic acid bait set may comprise multiple redundant a target- specific capture sequences, such as, for instance, different sequences that are complementary to different portions of the same nucleic acid pattern of interest. For example,each individual bait in a multiple redundant bait set may comprise one of a plurality of different target- specific capture sequences present in the set. The target- specific capture sequences may be semi-overlapping across the same target locus or correspond to different sections of the same target locus. Such multiple redundant bait sets may reduce hybridization bias by increasing coverage, depending on the hybridization efficiency with the nucleic acid pattern of interest. For instance, a multiple redundant bait set may have 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 different or semi-overlapping sequences corresponding to the same target binding sequence. In some embodiments, a multiple redundant bait set may have 2, 3, or 5 different or semi-overlapping sequences corresponding to the same target binding sequence.V. Methods of Selectively Capturing Nucleic Acid Molecules
[0196] The present disclosure further includes methods of selectively capturing nucleic acid molecules using a bait molecule described herein. For instance, in some embodiments, the methods described herein involve hybridizing nucleic acid molecules with one or more nucleic acid baits designed according to the principles laid out herein. This hybridization may be done in, for instance, the target capture step of a Methyl-Seq workflow (such as, e.g., in the exemplary workflow shown in FIG. 1) under conditions suitable for hybridization, wherein a plurality of nucleic acid molecules are capable of hybridization with a bait nucleic acid, followed by separating the plurality of hybrids from nucleic acid molecules that are not hybridized to a bait.
[0197] FIG. 2 provides a non-limiting example of steps taken to selectively capture nucleic acid molecules from a sample from a subject (such as, for example, a human subject) in a process 200, according to one implementation of single-stranded nucleic acid baits designed according to the principles described herein. As illustrated in FIG. 2, one or more unmethylated cytosine residues (C) in one or more of the nucleic acid molecules, or one or more methylated C in one or more of the nucleic acid molecules, is converted, 202, into a uracil residue (U) to form converted nucleic acid molecules. In some instances, the conversion comprises enzymatic conversion and / or bisulfite conversion. In some instances, the conversion comprises TET-assisted bisulfite treatment, TET-assisted pyridine borane treatment, oxidative bisulfite treatment, APOBEC treatment, and / or other DNA deaminase treatment.
[0198] In some instances, one or more adaptors (such as, for instance, amplification primers, flow cell adaptor sequences, substrate adapter sequences, and / or sample index sequences) are ligated onto one or more nucleic acid molecules prior to 202. In some instances, one or more adaptors are ligated onto one or more nucleic acid molecules between 202 and 204.
[0199] The converted nucleic acid molecules are amplified, 204, into a set of amplicons. In some instances, the amplification comprises a polymerase chain reaction (PCR) amplification technique, a non-PCR amplification technique, or an isothermal amplification technique.
[0200] The amplicons are hybridized, 206, with a set of single- stranded nucleic acid baits to form a plurality of hybrids, wherein at least a portion of the hybrids comprise a nucleic acid pattern of interest corresponding to a methylation pattern of interest at a genomic locus in the nucleic acid molecules prior to the converting. In some instances, the methylation pattern of interest corresponds to a hypermethylation signal indicative of a disease or condition. In some instances, the methylation pattern of interest corresponds to a hypomethylation signal indicative of a disease or condition. In some instances, the disease or condition is cancer. In some instances, a first portion of the hybrids in the plurality of hybrids comprise a singlestranded nucleic acid bait from the set of single- stranded nucleic acid baits and a nucleic acid molecule comprising the nucleic acid pattern of interest hybridized to each other without mismatches. In some instances, the set of baits is hypermethylation- specific. In some instances, the set of baits is hypomethylation- specific. In some instances, the hybrids in the plurality of hybrids do not comprise a guanine / thymine (G:T) mismatch between a hypomethylation- specific bait and a nucleic acid molecule comprising a sequence corresponding to a hypermethylated version of the genomic locus. In some instances, the hybrids in the plurality of hybrids do not comprise a guanine / thymine (G:T) mismatch between a hypermethylation- specific bait and a nucleic acid molecule comprising a sequence corresponding to a hypomethylated version of the genomic locus. In some instances, a second portion of hybrids in the plurality of hybrids comprises a single-stranded nucleic acid bait from the set of single- stranded nucleic acid baits and a nucleic acid molecule that does not comprise the nucleic acid pattern of interest, wherein the hybrids in the second portion of hybrids comprise one or more adenine / cytosine (A:C) mismatches. For instance, in the case of a hypomethylation- specific bait set, an exemplary second portion of hybrids may comprise a single-stranded nucleic acid bait from the set of single- stranded nucleic acid baits hybridized with one or more A:C mismatches to a nucleic acid molecule comprising asequence corresponding to a hypermethylated version of the genomic locus. In the case of a hypermethylation- specific bait set, an exemplary second portion of hybrids may comprise a single- stranded nucleic acid bait from the set of single-stranded nucleic acid baits hybridized with one or more A:C mismatches to a nucleic acid molecule comprising a sequence corresponding to a hypomethylated version of the genomic locus.
[0201] The plurality of hybrids is separated, 208, from nucleic acid molecules that are not hybridized to a bait. In some instances, the separation comprises one or more wash steps, wherein nucleic acid molecules that are not hybridized to a bait are depleted. In some instances, hybrids comprising imperfect hybridization, such as hybrids comprising one or more mismatches, are depleted in the one or more wash steps. For instance, the second portion of hybrids may be depleted. In some instances, one or more wash steps takes place at a temperature lower than a melting temperature (Tm) of a hybrid comprising a nucleic acid bait and a nucleic acid molecule comprising the nucleic acid pattern of interest hybridized to each other without mismatches.
[0202] FIG. 3 provides a non-limiting example of steps taken to selectively capture nucleic acid molecules from a sample from a subject (such as, for example, a human subject) in a process 300, according to one implementation of hypermethylation-specific nucleic acid baits designed according to the principles described herein, with either single-stranded hypermethylation- specific baits or double-stranded hypermethylation- specific baits. As illustrated in FIG. 3, one or more unmethylated cytosine residues (C) in one or more of the nucleic acid molecules, or one or more methylated C in one or more of the nucleic acid molecules, is converted, 302, into a uracil residue (U) to form converted nucleic acid molecules.
[0203] In some instances, the conversion comprises enzymatic conversion and / or bisulfite conversion. In some instances, the conversion comprises TET-assisted bisulfite treatment, TET-assisted pyridine borane treatment, oxidative bisulfite treatment, APOBEC treatment, and / or other DNA deaminase treatment.
[0204] In some instances, one or more adaptors (such as, for instance, amplification primers, flow cell adaptor sequences, substrate adapter sequences, and / or sample index sequences) are ligated onto one or more nucleic acid molecules prior to 302. In some instances, one or more adaptors are ligated onto one or more nucleic acid molecules between 302 and 304.
[0205] The converted nucleic acid molecules are amplified, 304, into a set of amplicons. In some instances, the amplification comprises a polymerase chain reaction (PCR) amplification technique, a non-PCR amplification technique, or an isothermal amplification technique.
[0206] The amplicons are hybridized, 306, with a set of hypermethylation- specific nucleic acid baits to form a plurality of hybrids, wherein the set of hypermethylation- specific nucleic acid baits selectively hybridizes to a nucleic acid pattern of interest corresponding to a methylation pattern of interest at a genomic locus in the nucleic acid molecules prior to the converting, wherein the methylation pattern of interest corresponds to a hypermethylation signal. In some instances, the hypermethylation signal is indicative of a disease or condition, such as, for instance, cancer. In some instances, the disease or condition is cancer. In some instances, the hypermethylation- specific nucleic acid baits are single-stranded. In some instances, the hypermethylation- specific nucleic acid baits are double-stranded. In instances involving double- stranded baits, a hybrid comprising at least one strand of a double- stranded bait is considered to comprise the bait. For instance, a hybrid comprising a strand of an amplicon hybridized to a single strand from a double-stranded bait is considered to comprise the bait. In some instances, a first portion of the hybrids in the plurality of hybrids comprise a hypermethylation- specific nucleic acid bait from the set of hypermethylation- specific nucleic acid baits and a nucleic acid molecule comprising the nucleic acid pattern of interest hybridized to each other without mismatches. In some instances, a second portion of hybrids in the plurality of hybrids comprises a hypermethylation- specific nucleic acid bait from the set of hypermethylation- specific nucleic acid baits and a nucleic acid molecule that does not comprise the nucleic acid pattern of interest, wherein the hybrids in the second portion of hybrids comprise one or more mismatches. For instance, an exemplary second portion of hybrids may comprise a hypermethylation- specific nucleic acid bait from the set of hypermethylation- specific nucleic acid baits hybridized with one or more mismatches to a nucleic acid molecule comprising a sequence corresponding to a hypomethylated version of the genomic locus.
[0207] The plurality of hybrids is separated, 308, from nucleic acid molecules that are not hybridized to a bait. In some instances, the separation comprises one or more wash steps, wherein nucleic acid molecules that are not hybridized to a bait are depleted. In some instances, hybrids comprising imperfect hybridization, such as hybrids comprising one or more mismatches, are depleted in the one or more wash steps. For instance, the secondportion of hybrids may be depleted. In some instances, one or more wash steps takes place at a temperature lower than a melting temperature (Tm) of a hybrid comprising a nucleic acid bait and a nucleic acid molecule comprising the nucleic acid pattern of interest hybridized to each other without mismatches.
[0208] FIG. 4 provides a non-limiting example of steps taken to selectively capture nucleic acid molecules from a sample from a subject (such as, for example, a human subject) in a process 400, according to one implementation of hypomethylation- specific nucleic acid baits designed according to the principles described herein, with either single-stranded hypomethylation- specific baits or double- stranded hypomethylation- specific baits. As illustrated in FIG. 4, one or more unmethylated cytosine residues (C) in one or more of the nucleic acid molecules, or one or more methylated C in one or more of the nucleic acid molecules, is converted, 402, into a uracil residue (U) to form converted nucleic acid molecules.
[0209] In some instances, the conversion comprises enzymatic conversion and / or bisulfite conversion. In some instances, the conversion comprises TET-assisted bisulfite treatment, TET-assisted pyridine borane treatment, oxidative bisulfite treatment, APOBEC treatment, and / or other DNA deaminase treatment.
[0210] In some instances, one or more adaptors (such as, for instance, amplification primers, flow cell adaptor sequences, substrate adapter sequences, and / or sample index sequences) are ligated onto one or more nucleic acid molecules prior to 402. In some instances, one or more adaptors are ligated onto one or more nucleic acid molecules between 402 and 404.
[0211] The converted nucleic acid molecules are amplified, 404, into a set of amplicons. In some instances, the amplification comprises a polymerase chain reaction (PCR) amplification technique, a non-PCR amplification technique, or an isothermal amplification technique.
[0212] The amplicons are hybridized, 406, with a set of hypomethylation- specific nucleic acid baits to form a plurality of hybrids, wherein the set of hypomethylation- specific nucleic acid baits selectively hybridizes to a nucleic acid pattern of interest corresponding to a methylation pattern of interest at a genomic locus in the nucleic acid molecules prior to the converting, wherein the methylation pattern of interest corresponds to a hypomethylation signal. In some instances, the hypomethylation signal is indicative of a disease or condition, such as, for instance, cancer. In some instances, the disease or condition is cancer. In some instances, the hypomethylation-specific nucleic acid baits are single-stranded. In someinstances, the hypomethylation-specific nucleic acid baits are double- stranded. In instances involving double- stranded baits, a hybrid comprising at least one strand of a double- stranded bait is considered to comprise the bait. For instance, a hybrid comprising a strand of an amplicon hybridized to a single strand from a double-stranded bait is considered to comprise the bait. In some instances, a first portion of the hybrids in the plurality of hybrids comprise a hypomethylation-specific nucleic acid bait from the set of hypomethylation-specific nucleic acid baits and a nucleic acid molecule comprising the nucleic acid pattern of interest hybridized to each other without mismatches. In some instances, a second portion of hybrids in the plurality of hybrids comprises a hypomethylation-specific nucleic acid bait from the set of hypomethylation-specific nucleic acid baits and a nucleic acid molecule that does not comprise the nucleic acid pattern of interest, wherein the hybrids in the second portion of hybrids comprise one or more mismatches. For instance, an exemplary second portion of hybrids may comprise a hypomethylation-specific nucleic acid bait from the set of hypomethylation-specific nucleic acid baits hybridized with one or more mismatches to a nucleic acid molecule comprising a sequence corresponding to a hypermethylated version of the genomic locus.
[0213] The plurality of hybrids is separated, 408, from nucleic acid molecules that are not hybridized to a bait. In some instances, the separation comprises one or more wash steps, wherein nucleic acid molecules that are not hybridized to a bait are depleted. In some instances, hybrids comprising imperfect hybridization, such as hybrids comprising one or more mismatches, are depleted in the one or more wash steps. For instance, the second portion of hybrids may be depleted. In some instances, one or more wash steps takes place at a temperature lower than a melting temperature (Tm) of a hybrid comprising a nucleic acid bait and a nucleic acid molecule comprising the nucleic acid pattern of interest hybridized to each other without mismatches.VI. Combinations of Different Baits
[0214] Certain aspects of the present invention involve combinations of different baits. For example, in some embodiments, different baits (for example, corresponding to different nucleic acid sequences of interest, and / or baits with different features, such as different numbers of strands, different lengths, and / or other various different modifications discussed above) are combined in, for instance, the same panel and / or in the same vial. These differentbaits could, for example, target the same locus or multiple different loci, and / or be singlestranded or double-stranded, or a combination thereof.A. Different Baits Targeting the Same Locus
[0215] Different baits targeting the same locus could each be specific to, for example, a different degree or pattern of methylation across the same genomic sequence. For example, there may be multiple different possible hypermethylated alleles, or multiple different possible hypomethylated alleles for a single genomic region, such that there could be multiple potential different hypermethylation signatures or multiple potential different hypomethylation signatures in a converted nucleic acid molecule corresponding to a given genomic locus. Different baits could be used to interrogate each different possible allele.
[0216] For instance, in some embodiments, a plurality of sets of hypermethylation- specific baits, such as a second, third, fourth, or more set(s) of hypermethylation- specific baits could be used together, each of which hybridizes to a different nucleic pattern of interest converted from a different hypermethylation signature at the same genomic locus. In such embodiments, the selectivity of one set of bait within the plurality of baits for its intended target would depend on the difference(s) in sequence between the converted nucleic acid sequences corresponding to the different alleles.
[0217] For example, in a hypermethylation- specific bait context, a sample may comprise a plurality of different hypermethylated alleles at a single genomic locus. For example, a plurality of copies of the same genomic locus may be present within a single sample from a single subject, and / or within one or more samples from one or more subjects, with a plurality of different hypermethylation patterns of interest. In some embodiments, one or more of the plurality of different hypermethylation patterns of interest at the same genomic locus may, for example, correspond to a hypermethylation signal indicative of a disease or condition. In some embodiments, different hypermethylation patterns of interest at the same genomic locus may correspond to one or more hypermethylation signals indicative of one or more different diseases or conditions.
[0218] Similarly, in a hypomethylation- specific bait context, a sample may comprise a plurality of different hypomethylated alleles at a single genomic locus. For example, a plurality of copies of the same genomic locus may be present within a single sample from a single subject, and / or within one or more samples from one or more subjects, with a pluralityof different hypomethylation patterns of interest. In some embodiments, one or more of the plurality of different hypomethylation patterns of interest at the same genomic locus may, for example, correspond to a hypomethylation signal indicative of a disease or condition. In some embodiments, different hypomethylation patterns of interest at the same genomic locus may correspond to one or more hypomethylation signals indicative of one or more different diseases or conditions.
[0219] Thus, in some embodiments, a plurality of different methylation patterns of interest at a single genomic locus may be converted into a plurality of different nucleic acid patterns of interest corresponding to the same genomic locus to form a plurality of different converted nucleic acid molecules. In such embodiments, the method may further comprise amplifying the plurality of different nucleic acid patterns into a plurality of sets of amplicons, hybridizing a plurality of sets of nucleic acid baits to the plurality of sets of amplicons to form a plurality of hybrids, wherein the plurality of sets of nucleic acid baits comprises a plurality of different bait sets, wherein the different baits sets selectively hybridize to the different nucleic acid patterns of interest of the genomic locus; and separating the plurality of hybrids from nucleic acid molecules that are not hybridized to a bait.
[0220] In some embodiments, there is only one bait per genomic locus. In some embodiments, there are multiple baits per genomic locus, such as, for example, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 baits per genomic locus. In some embodiments, different baits targeting the same locus would all be hypermethylation- specific baits. In some embodiments, different baits targeting the same locus would all be hypomethylation- specific baits.B. Different Baits Targeting Different Loci
[0221] Additionally or alternatively, multiple different baits could be used to query different loci in parallel. While multiple baits targeting the same locus would either all be hypermethylation- specific or hypomethylation- specific, baits targeting different loci may be, for example, hypermethylation-specific at a first locus and hypomethylation- specific at a second locus. Further, multiple baits targeting the same locus may be mixed with one or more different baits targeting one or more different loci. For example, in some embodiments, there may be one or more hypomethylation- specific baits targeting locus “A” and one or more hypermethylation- specific baits targeting locus “B”.
[0222] Any number of different loci may be targeted in parallel with different baits so long as the sequences of the converted nucleic acid molecules corresponding to each locus are sufficiently different so as to enable selective binding between a bait ad its intended target. In some embodiments, the methylation state of only one genomic locus is queried. In some embodiments, there are multiple genomic loci queried in parallel, such as, for example, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 different genomic loci. In some embodiments, different baits targeting different loci are all hypermethylation- specific baits. In some embodiments, different baits targeting different loci are all hypomethylation- specific baits. In some embodiments, one or more loci are queried with only hypermethylation- specific baits and one or more other loci are queried with only hypomethylation- specific baits. In some embodiments, one locus is targeted by multiple different baits (such as, for instance, as outlined in the section above) and another locus is targeted by multiple other different baits, such that there are, for example, different baits targeting the same locus and other different baits targeting different loci in parallel (e.g., multiple loci, each with different possible alleles that are each targeted by different baits).
[0223] A method of the present invention that involves querying multiple loci in parallel could comprise, for example: converting a plurality of methylation patterns of interest (such as, for example, a first methylation pattern of interest and a second, third, fourth, or more methylation patterns of interest), each at, respectively, a different genomic locus among a plurality of genomic loci in nucleic acid molecules from a sample from a subject into a plurality of nucleic acid patterns of interest of the plurality of genomic loci, and hybridizing a plurality of sets of nucleic acid baits to the set of amplicons, wherein each set of nucleic acid baits of the plurality of sets of nucleic acid baits selectively hybridizes to a nucleic acid pattern of interest of the plurality of nucleic acid patterns of interest. In some embodiments, any or all of the first, second, third, fourth, or more methylation patterns of interest is associated with a hypermethylation signal indicative of a disease or condition. In some embodiments, any or all of the first, second, third, fourth, or more methylation patterns of interest is associated with a hypomethylation signal indicative of a disease or condition. In some embodiments, one or more genomic loci associated with different methylation patterns of interest are associated with a hypomethylation signal indicative of a disease or condition, and / or one or more different genomic loci associated with other different methylation patterns of interest are associated with a hypermethylation signal indicative of a disease or condition.In some embodiments, the diseases and / or conditions associated with different methylation patterns of interest are the same between loci. In some embodiments, the diseases and / or conditions associated with different methylation patterns of interest are different between loci.C. Panels
[0224] Any of the different bait designs and / or combinations of different baits described above may be combined in the form of, for instance, a panel. For example, within a panel, a first set of baits designed to capture a hypermethylation signature corresponding to an allele in one genomic locus may be used a first portion of the panel; a second set of baits designed to capture a hypomethylation signature corresponding to another allele in another genomic locus or the same genomic locus may be used in a second portion of the panel; and a third set of baits designed to capture both alleles may be used in a third portion of the panel.
[0225] Within a panel, different baits may be used to query methylation states associated with different diseases or conditions. For example, a panel using the baits of the present invention may be a screening panel for different types of cancer. In some embodiments, a panel may be designed to focus just on hypermethylation signals or just on hypomethylation signals. In some such embodiments, inclusion of a portion of the panel that queries for both hypermethylation signals and hypomethylation signals could allow for quantification and / or normalization. For instance, a disease- or condition-specific panel may comprise, for example, three regions: one with only hyper-baits, one with only hypo-baits, and a normalization control that captures both, which allows for fine-tuned quantification. The normalization control region of the panel could be further tuned, for instance, based on the genomic locus or loci of interest. Tuning may include, for example, selecting for genomic regions with varying methylation levels in non-disease conditions.VII. Target Loci
[0226] A target locus (also discussed herein as a genomic locus having a methylation pattern of interest) or multiple target loci are chosen such that a bait or baits designed according to the present disclosure are able to selectively hybridize to the nucleic acid pattem(s) of interest corresponding to the genomic locus or loci. A bait may be designed according to the present disclosure to hybridize to a nucleic acid pattern of interest corresponding to any genomiclocus with a methylation pattern of interest. In some embodiments, the genomic locus comprises no CpG (a cytosine residue and a guanine residue separated by a single phosphate group) sites. In some embodiments, the genomic locus comprises one CpG site. In some embodiments, the genomic locus comprises two CpG sites. In some embodiments, the genomic locus comprises more than two CpG sites. In some embodiments, the genomic locus comprises 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more than 20 CpG sites.
[0227] In some embodiments, a target locus comprises a coding sequence, such as, for example, a gene, or a fragment thereof. In some embodiments, a target locus comprises a non-coding sequence. In some embodiments, a target locus comprises an intragenic region or a fragment thereof. In some embodiments, a target locus comprises an intergenic region or a fragment thereof. In some embodiments, a target locus comprises a regulatory region or a fragment thereof. In some embodiments, a target locus comprises a non-coding sequence or fragment thereof (e.g., a promoter sequence, enhancer sequence, 5' untranslated region (5' UTR), 3' untranslated region (3' UTR), or a fragment thereof), an exon sequence or fragment thereof, an intron sequence or a fragment thereof, or a combination of any of the above.
[0228] The methods described herein can be used in combination with, or as part of, a method for evaluating a plurality or set of subject intervals (e.g., nucleic acid patterns of interest), e.g., from a set of loci (e.g., CpG sites), as described herein.
[0229] In some instances, the set of genomic loci evaluated by the disclosed methods comprises a plurality of, e.g., genes, which in aberrantly methylated form(s), are associated with a disease or condition, such as, for example, an effect on cell division, growth or survival, or are associated with a cancer, e.g., a cancer described herein.
[0230] In some instances, the set of loci (e.g., CpG sites) evaluated by the disclosed methods comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, or more than 100 loci.
[0231] In some instances, the selected loci (also referred to herein as target loci, target sequences, genomic loci, genomic loci of interest, genomic loci comprising a methylation pattern of interest, and the like), or fragments thereof, may include subject intervals comprising non-coding sequences, coding sequences, intragenic regions, or intergenic regions of the subject genome. For example, the subject intervals can include a non-coding sequenceor fragment thereof (e.g., a promoter sequence, enhancer sequence, 5' untranslated region (5' UTR), 3' untranslated region (3' UTR), or a fragment thereof), a coding sequence of fragment thereof, an exon sequence or fragment thereof, an intron sequence or a fragment thereof.
[0232] In some embodiments, all of the CpG sites of a genomic locus are methylated in a methylation pattern of interest. In some embodiments, none of the CpG sites of a genomic locus are methylated in a methylation pattern of interest. In some embodiments, a portion of the CpG sites of a genomic locus are methylated in a methylation pattern of interest. In some embodiments, a majority of the CpG sites of a genomic locus are methylated in a methylation pattern of interest. In some embodiments, a minority of the CpG sites of a genomic locus are methylated in a methylation pattern of interest. In some embodiments, a methylation pattern of interest comprises equal numbers of methylated CpG sites and unmethylated CpG sites.
[0233] In some embodiments, a methylation pattern of interest corresponding to the genomic locus comprises no methylated CpG sites. In some embodiments, the methylation pattern of interest comprises one methylated CpG site. In some embodiments, the methylation pattern of interest comprises two methylated CpG sites. In some embodiments, the methylation pattern of interest comprises more than two methylated CpG sites. In some embodiments, the methylation pattern of interest comprises 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more than 20 methylated CpG sites.
[0234] In some embodiments, a target locus may be within a subgenomic interval. In some embodiments, more than one target loci may be within a subgenomic interval. In some embodiments, a target locus may comprise a subgenomic interval. In some embodiments, a target locus may comprise more than one subgenomic interval.
[0235] In some embodiments, the methylation state of more than one locus is queried in parallel using one or more of the bait designs disclosed herein. In some embodiments, more than one methylation state of a single locus is queried in parallel using one or more of the bait designs disclosed herein.
[0236] In some embodiments, sequencing reads produced using the methods disclosed herein overlap with one or more loci within one or more subgenomic intervals in the sample. In some embodiments, sequencing reads produced using the methods disclosed herein overlap with ten or more loci within one or more subgenomic intervals in the sample. In some embodiments, sequencing reads produced using the methods disclosed herein overlap with 10 to 20 loci within one or more subgenomic intervals in the sample. In some embodiments,sequencing reads produced using the methods disclosed herein overlap with more than 20, more than 50, more than 100, more than 200, more than 300, more than 500, more than 1,000, more than 5,000, more than 10,000, more than 20,000, more than 50,000, or more than 100,000 loci within one or more subgenomic intervals in the sample. In some embodiments, the loci within one or more subgenomic intervals in the sample with which the sequencing reads produced using the methods disclosed herein overlap comprise more than two CpG sites.
[0237] In some embodiments, one or more target loci comprise ABL1, ACVR1B, AKT1, AKT2, AKT3, ALK, ALOX12B, AMER1, APC, AR, ARAF, ARFRP1, ARID1A, ASXL1, ATM, ATR, ATRX, AURKA, AURKB, AXIN1, AXL, BAP1, BARD1, BCL2, BCL2L1, BCL2L2, BCL6, BCOR, BCORL1, BCR, BRAF, BRCA1, BRCA2, BRD4, BRIP1, BTG1, BTG2, BTK, CALR, CARD11, CASP8, CBFB, CBL, CCND1, CCND2, CCND3, CCNE1, CD22, CD274, CD70, CD74, CD79A, CD79B, CDC73, CDH1, CDK12, CDK4, CDK6, CDK8, CDKN1A, CDKN1B, CDKN2A, CDKN2B, CDKN2C, CEBPA, CHEK1, CHEK2, CIC, CREBBP, CRKL, CSF1R, CSF3R, CTCF, CTNNA1, CTNNB1, CUL3, CUL4A, CXCR4, CYP17A1, DAXX, DDR1, DDR2, DIS3, DNMT3A, DOT1L, EED, EGFR, EMSY (Cllorf30), EP300, EPHA3, EPHB1, EPHB4, ERBB2, ERBB3, ERBB4, ERCC4, ERG, ERRFI1, ESRI, ETV4, ETV5, ETV6, EWSR1, EZH2, EZR, FAM46C, FANCA, FANCC, FANCG, FANCL, FAS, FBXW7, FGF10, FGF12, FGF14, FGF19, FGF23, FGF3, FGF4, FGF6, FGFR1, FGFR2, FGFR3, FGFR4, FH, FLCN, FLT1, FLT3, FOXL2, FUBP1, GABRA6, GATA3, GATA4, GATA6, GID4 (C17orf39), GNA11, GNA13, GNAQ, GNAS, GRM3, GSK3B, H3F3A, HDAC1, HGF, HNF1A, HRAS, HSD3B1, ID3, IDH1, IDH2, IGF1R, IKBKE, IKZF1, INPP4B, IRF2, IRF4, IRS2, JAK1, JAK2, JAK3, JUN, KDM5A, KDM5C, KDM6A, KDR, KEAP1, KEL, KIT, KLHL6, KMT2A (MLL), KMT2D (MLL2), KRAS, LTK, LYN, MAF, MAP2K1, MAP2K2, MAP2K4, MAP3K1, MAP3K13, MAPK1, MCL1, MDM2, MDM4, MED12, MEF2B, MEN1, MERTK, MET, MITF, MKNK1, MLH1, MPL, MRE11A, MSH2, MSH3, MSH6, MST1R, MTAP, MTOR, MUTYH, MYB, MYC, MYCL, MYCN, MYD88, NBN, NF1, NF2, NFE2L2, NFKBIA, NKX2-1, NOTCH1, NOTCH2, NOTCH3, NPM1, NRAS, NT5C2, NTRK1, NTRK2, NTRK3, NUTM1, P2RY8, PALB2, PARK2, PARP1, PARP2, PARP3, PAX5, PBRM1, PDCD1, PDCD1LG2, PDGFRA, PDGFRB, PDK1, PIK3C2B, PIK3C2G, PIK3CA, PIK3CB, PIK3R1, PIM1, PMS2, POLDI, POLE, PPARG, PPP2R1A, PPP2R2A, PRDM1, PRKAR1A, PRKCI,PTCHI, PTEN, PTPN11, PTPRO, QKI, RAC1, RAD21, RAD51, RAD51B, RAD51C, RAD51D, RAD52, RAD54L, RAFI, RARA, RBI, RBM10, REL, RET, RICTOR, RNF43, ROS1, RPTOR, RSPO2, SDC4, SDHA, SDHB, SDHC, SDHD, SETD2, SF3B1, SGK1, SLC34A2, SMAD2, SMAD4, SMARCA4, SMARCB1, SMO, SNCAIP, SOCS1, SOX2, SOX9, SPEN, SPOP, SRC, STAG2, STAT3, STK11, SUFU, SYK, TBX3, TEK, TERC, TERT, TET2, TGFBR2, TIPARP, TMPRSS2, TNFAIP3, TNFRSF14, TP53, TSC1, TSC2, TYRO3, U2AF1, VEGFA, VHL, WHSCI, WHSC1L1, WT1, XPO1, XRCC2, ZNF217, ZNF703, or any combination thereof.
[0238] In some embodiments, one or more target loci comprise ABL, ALK, ALL, B4GALNT1, BAFF, BCL2, BRAF, BRCA, BTK, CD19, CD20, CD3, CD30, CD319, CD38, CD52, CDK4, CDK6, CML, CRACC, CS1, CTLA-4, dMMR, EGFR, ERBB1, ERBB2, FGFR1-3, FLT3, GD2, HDAC, HER1, HER2, HR, IDH2, IL-ip, IL-6, IL-6R, JAK1, JAK2, JAK3, KIT, KRAS, MEK, MET, MSI-H, mTOR, PARP, PD-1, PDGFR, PDGFRa, PDGFRP, PD-L1, PI3K5, PIGF, PTCH, RAF, RANKL, RET, ROS1, SLAMF7, VEGF, VEGFA, VEGFB, or any combination thereof.VIII. Subjects
[0239] In some embodiments, the methods comprise obtaining a sample from a subject, such as an individual. In some embodiments, one or more samples from one or more subjects are obtained. In some embodiments, samples are obtained from one subject. In some embodiments, samples are obtained from more than one subject. In some embodiments, more than one sample from a subject is obtained. In some embodiments, a subject is a bacterial, plant, fungal, or animal subject. In some embodiments, a subject is an animal subject, such as, for instance, a mammalian subject, such as a human subject. In some embodiments, the subject is an adult. In some embodiments, the subject is a child. In some embodiments, the subject is pregnant. In some embodiments, the subject is a fetus. In some embodiments, a sample is obtained from a plurality of subjects simultaneously, such as, e.g., from a pregnant person and one or more fetuses simultaneously, such as, for instance, in a blood sample taken from a pregnant person.
[0240] In some instances, a sample is obtained (e.g., collected) from a subject (e.g., patient) with a condition or disease (e.g., a hyperproliferative disease or a non-cancer indication) or suspected of having the condition or disease. In some embodiments, a subject may have, havehad, or be suspected of having any of the diseases or conditions disclosed herein and / or any disease or condition known or suspected of being associated with one or more aberrant methylation states. In some embodiments, a subject is being monitored for a disease or condition, such as, for example, progression of or recovery from a disease or condition.
[0241] In some instances, a subject has a risk of having a disease or condition. For example, in some instances, a subject has a genetic predisposition to a disease or condition (e.g., having a genetic mutation and / or aberrant methylation state that increases his or her baseline risk for developing a disease or condition). In some instances, a subject has been exposed to an environmental perturbation (e.g., radiation or a chemical) that increases his or her risk for developing a disease or condition. In some instances, a subject is in need of being monitored for development of a disease or condition. In some instances, a subject is in need of being monitored for disease or condition progression or regression, e.g., after being treated with a therapy or treatment for the disease or condition. In some instances, a subject is in need of being monitored for relapse of a disease or condition. In some instances, a subject is in need of being monitored for minimum residual disease (MRD). In some instances, a subject is undergoing treatment response monitoring (TRM). In some instances, a subject being monitored for MRD comprises a relatively low tumor fraction, such as, for example, from early in disease progression and / or shortly after surgery. In some instances, a subject undergoing TRM comprises a relatively high TF, such as, for example, from an advanced disease. In some instances, a subject has been, or is being treated, for a disease or condition. In some instances, a subject has not been treated with a therapy or treatment for the disease or condition.
[0242] In some instances, the sample is acquired from a subject having a cancer. In some instances, the cancer is a solid tumor or a metastatic form thereof. In some instances, the cancer is a hematological cancer, e.g., a leukemia or lymphoma. In some instances, the subject has a risk of having a cancer. For example, in some instances, the subject has a genetic predisposition to a cancer (e.g., having a genetic mutation and / or aberrant methylation state that increases his or her baseline risk for developing a cancer). In some instances, the subject has been exposed to an environmental perturbation (e.g., radiation or a chemical) that increases his or her risk for developing a cancer. In some instances, the subject is in need of being monitored for development of a cancer. In some instances, the subject is in need of being monitored for cancer progression or regression, e.g., after being treated with ananti-cancer therapy (or anti-cancer treatment). In some instances, the subject is in need of being monitored for relapse of cancer. In some instances, the subject is in need of being monitored for minimum residual disease (MRD). In some instances, the subject has been, or is being treated, for cancer. In some instances, the subject is in need of undergoing treatment response monitoring (TRM). In some instances, the subject is the subject of a comprehensive genomic profile (CGP) or is in need thereof. In some instances, the subject has not been treated with an anti-cancer therapy (or anti-cancer treatment).
[0243] In some instances, a subject (e.g., a patient) is being treated, or has been previously treated, with one or more targeted therapies. In some instances, e.g., for a patient who has been previously treated with a targeted therapy, a post-targeted therapy sample (e.g., specimen) is obtained (e.g., collected). In some instances, the post-targeted therapy sample is a sample obtained after the completion of the targeted therapy.
[0244] In some instances, a patient has not been previously treated with a targeted therapy. In some instances, e.g., for a patient who has not been previously treated with a targeted therapy, the sample comprises a resection, e.g., an original resection, or a resection following recurrence (e.g., following a disease recurrence post- therapy).IX. Samples
[0245] Nucleic acid molecules as used in accordance with the methods described herein may be provided in a sample or in a plurality of samples. For example, a single sample may be taken from a single subject. Alternatively, more than one sample may be taken from one subject or from more than one subject. In some embodiments, multiple samples are taken from the same subject over time. In some embodiments, multiple samples are taken from different body parts of the same subject.
[0246] Thus, in some embodiments, the methods of the present disclosure include obtaining one or more samples from one or more subjects. The sample may be derived from one or more subjects, such as, for example, any of the subjects described above, or a combination thereof. The sample may comprise nucleic acids comprising one or more methylation pattens of interest at one or more genomic loci as described above. The nucleic acid molecules can be derived from a nucleic acid duplex molecule (e.g., a DNA duplex molecule, such as a cell- free DNA duplex molecule) or a single-stranded nucleic acid molecule (e.g., a singlestranded cfDNA molecule or an RNA molecule). The duplex or single-stranded nucleic acidmolecule may be naturally occurring and may be isolated according to the methods described herein.
[0247] The disclosed methods may be used with any of a variety of samples. For example, in some instances, the sample may comprise a tissue biopsy sample, a liquid biopsy sample, or a normal control. In some instances, the sample may be a liquid biopsy sample and may comprise blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva. In some instances, the sample may be a liquid biopsy sample and may comprise circulating tumor cells (CTCs). In some instances, the sample may be a liquid biopsy sample and may comprise cfDNA, circulating tumor DNA (ctDNA), or any combination thereof.
[0248] In some embodiments, the sample is or comprises biological tissue or fluid. The sample can contain compounds that are not naturally intermixed with the tissue in nature such as preservatives, anticoagulants, buffers, fixatives, nutrients, antibiotics or the like. In one embodiment, the sample is preserved as a frozen sample or as a formaldehyde- or paraformaldehyde-fixed paraffin-embedded (FFPE) tissue preparation. For example, the sample can be embedded in a matrix, e.g., an FFPE block or a frozen sample. In another embodiment, the sample is a blood or blood constituent sample. In yet another embodiment, the sample is a bone marrow aspirate sample. In another embodiment, the sample comprises cfDNA, e.g., tumor cfDNA or tumor cfDNA. Without wishing to be bound by theory, it is believed that in some embodiments, cfDNA is DNA from apoptosed or necrotic cells.Typically, cfDNA is bound by protein (e.g., histone) and protected by nucleases. cfDNA can be used as a biomarker, for example, for non-invasive prenatal testing (NIPT), organ transplant, cardiomyopathy, microbiome, and cancer. In another embodiment, the sample comprises circulating tumor DNA (ctDNA). Without wishing to be bound by theory, it is believed that in some embodiments, ctDNA is cfDNA with a genetic or epigenetic alteration (e.g., a somatic alteration or a methylation signature) that can discriminate it originating from a tumor cell versus a non-tumor cell. In another embodiment, the sample comprises circulating tumor cells (CTCs). Without wishing to be bound by theory, it is believed that in some embodiments, CTCs are cells shed from a primary or metastatic tumor into the circulation. In some embodiments, CTCs apoptose and are a source of ctDNA in the blood / lymph.
[0249] In some embodiments, a sample may comprise fetal DNA. (e.g., from invasive or non- invasive prenatal testing). For example, a sample may be obtained using invasiveamniocentesis, chorionic villus sampling (cVS), or fetal umbilical cord sampling techniques, or obtained using non-invasive sampling of cfDNA samples (which comprises a mix of maternal cfDNA and fetal cfDNA).
[0250] The sample may comprise cfDNA. Some cfDNA molecules are free-floating double stranded DNA molecules (dsDNA or duplex DNA) or single- stranded DNA molecules found in the blood stream, typically as the result of cell apoptosis or necrosis, particularly in the context of disease. These degraded linear DNA fragments are often approximately 50-300 base pairs in length. Most commonly, cfDNA is assayed for cancer screening at early stages in disease progression by analyzing the cfDNA sequences to identify cancer-associated mutations. cfDNA can also harbor methylated residues reflective of the methylation state, including aberrant methylation states, of their genomic contexts.
[0251] The sample may comprise circulating tumor DNA (ctDNA). In some embodiments, the sample comprises tumor cells and / or tumor nucleic acids. In some embodiments, the methods comprise extracting the mixture of polynucleotides from the sample, wherein the mixture of polynucleotides is from the tumor cells and / or tumor nucleic acids. In some embodiments, the sample further comprises non-tumor cells.
[0252] In some embodiments, the sample comprises a fraction of tumor nucleic acids that is less than 1% of total nucleic acids, less than 0.5% of total nucleic acids, less than 0.1% of total nucleic acids, or less than 0.05% of total nucleic acids. In some embodiments, the sample comprises a fraction of tumor nucleic acids that is at least 0.01%, at least 0.05%, or at least 0.1% of total nucleic acids. In some embodiments, the sample comprises a fraction of tumor nucleic acids having an upper limit of 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.4%, 0.3%, 0.2%, 0.1%, 0.09%, 0.08%, 0.07%, 0.06%, 0.05%, 0.04%, 0.03%, or 0.02% of total nucleic acids and an independently selected lower limit of 0.0001%, 0.0002%, 0.0003%, 0.0004%, 0.0005%, 0.0006%, 0.0007%, 0.0008%, 0.0009%, 0.001%, 0.002%, 0.003%, 0.004%, 0.005%, 0.006%, 0.007%, 0.008%, 0.009%, 0.01%, 0.02%, 0.03%, 0.04%, 0.05%, 0.06%, 0.07%, 0.08%, 0.09%, 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, or 1% of total nucleic acids, wherein the upper limit is greater than the lower limit. Advantageously, as demonstrated herein, the methods of the present disclosure allow for robust, ultrasensitive detection of aberrant methylation levels in slight amounts of tumor nucleic acids amongst otherwise normal nucleic acids.
[0253] In some instances, the nucleic acid molecules extracted from a sample may comprise a mixture of tumor nucleic acid molecules and non-tumor nucleic acid molecules. In some instances, the tumor nucleic acid molecules may be derived from a tumor portion of a heterogeneous tissue biopsy sample, and the non-tumor nucleic acid molecules may be derived from a normal portion of the heterogeneous tissue biopsy sample. In some instances, the sample may comprise a liquid biopsy sample, and the tumor nucleic acid molecules may be derived from a circulating tumor DNA (ctDNA) fraction of the liquid biopsy sample while the non-tumor nucleic acid molecules may be derived from a non-tumor cfDNA fraction of the liquid biopsy sample. In some embodiments, the sample comprises cfDNA, circulating tumor DNA (ctDNA) or combination of both.
[0254] In some embodiments, the nucleic acid molecules from the sample are fragmented. For instance, in some embodiments, the methods further comprise, prior to sequencing the plurality of polynucleotides or providing a plurality of sequence reads: subjecting a plurality of nucleic acids to fragmentation. A variety of DNA fragmentation techniques are used in the art prior to NGS approaches. In some embodiments, nucleic acids are fragmented by nebulization, in which compressed gas is used to mechanically shear nucleic acids through a small opening. In some embodiments, nucleic acids are fragmented by sonication, in which ultrasonic waves are used to shear nucleic acids. In some embodiments, nucleic acids are fragmented enzymatically, e.g., using one or more enzymes to digest nucleic acids into fragments. For instance, some commercially available enzyme kits contain a mixture of two enzymes: one that randomly generates dsDNA nicks, and one that recognizes nicked sites and cuts the opposite strand, generating dsDNA breaks.A. Nucleic acid extraction and processing
[0255] DNA or RNA may be extracted from tissue samples, biopsy samples, blood samples, or other bodily fluid samples using any of a variety of techniques known to those of skill in the art (see, e.g., Example 1 of International Patent Application Publication No. WO 2012 / 092426; Tan et al. (2009), “DNA, RNA, and Protein Extraction: The Past and The Present”, J. Biomed. Biotech. 2009:574398; the technical literature for the Maxwell® 16 LEV Blood DNA Kit (Promega Corporation, Madison, WI); and the Maxwell 16 Buccal Swab LEV DNA Purification Kit Technical Manual (Promega Literature #TM333, January 1, 2011, Promega Corporation, Madison, WI)). Protocols for RNA isolation are disclosed in,e.g., the Maxwell® 16 Total RNA Purification Kit Technical Bulletin (Promega Literature #TB351, August 2009, Promega Corporation, Madison, WI).
[0256] A typical DNA extraction procedure, for example, comprises (i) collection of the fluid sample, cell sample, or tissue sample from which DNA is to be extracted, (ii) disruption of cell membranes (z.e., cell lysis), if necessary, to release DNA and other cytoplasmic components, (iii) treatment of the fluid sample or lysed sample with a concentrated salt solution to precipitate proteins, lipids, and RNA, followed by centrifugation to separate out the precipitated proteins, lipids, and RNA, and (iv) purification of DNA from the supernatant to remove detergents, proteins, salts, or other reagents used during the cell membrane lysis step.
[0257] Disruption of cell membranes may be performed using a variety of mechanical shear (e.g., by passing through a French press or fine needle) or ultrasonic disruption techniques. The cell lysis step often comprises the use of detergents and surfactants to solubilize lipids the cellular and nuclear membranes. In some instances, the lysis step may further comprise use of proteases to break down protein, and / or the use of an RNase for digestion of RNA in the sample.
[0258] Examples of suitable techniques for DNA purification include, but are not limited to,(i) precipitation in ice-cold ethanol or isopropanol, followed by centrifugation (precipitation of DNA may be enhanced by increasing ionic strength, e.g., by addition of sodium acetate),(ii) phenol-chloroform extraction, followed by centrifugation to separate the aqueous phase containing the nucleic acid from the organic phase containing denatured protein, and (iii) solid phase chromatography where the nucleic acids adsorb to the solid phase (e.g., silica or other) depending on the pH and salt concentration of the buffer.
[0259] In some instances, cellular and histone proteins bound to the DNA may be removed either by adding a protease or by having precipitated the proteins with sodium or ammonium acetate, or through extraction with a phenol-chloroform mixture prior to a DNA precipitation step.
[0260] In some instances, DNA may be extracted using any of a variety of suitable commercial DNA extraction and purification kits. Examples include, but are not limited to, the QIAamp (for isolation of genomic DNA from human samples) and DNAeasy (for isolation of genomic DNA from animal or plant samples) kits from Qiagen (Germantown, MD) or the Maxwell® and ReliaPrep™ series of kits from Promega (Madison, WI).
[0261] As noted above, in some instances the sample may comprise a formalin-fixed (also known as formaldehyde-fixed, or paraformaldehyde-fixed), paraffin-embedded (FFPE) tissue preparation. For example, the FFPE sample may be a tissue sample embedded in a matrix, e.g., an FFPE block. Methods to isolate nucleic acids (e.g., DNA) from formaldehyde- or paraformaldehyde-fixed, paraffin-embedded (FFPE) tissues are disclosed in, e.g., Cronin et al., (2004) Am J Pathol. 164(l):35-42; Masuda et al., (1999) Nucleic Acids Res.27(22) : 4436-4443; Specht et al., (2001) Am J Pathol. 158(2):419-429; the Ambion RecoverAll™ Total Nucleic Acid Isolation Protocol (Ambion, Cat. No. AM1975, September 2008); the Maxwell® 16 FFPE Plus LEV DNA Purification Kit Technical Manual (Promega Literature #TM349, February 2011); the E.Z.N.A.® FFPE DNA Kit Handbook (OMEGA bio-tek, Norcross, GA, product numbers D3399-00, D3399-01, and D3399-02, June 2009); and the QIAamp® DNA FFPE Tissue Handbook (Qiagen, Cat. No. 37625, October 2007). For example, the RecoverAll™ Total Nucleic Acid Isolation Kit uses xylene at elevated temperatures to solubilize paraffin-embedded samples and a glass-fiber filter to capture nucleic acids. The Maxwell® 16 FFPE Plus LEV DNA Purification Kit is used with the Maxwell® 16 Instrument for purification of genomic DNA from 1 to 10 pm sections of FFPE tissue. DNA is purified using silica-clad paramagnetic particles (PMPs), and eluted in low elution volume. The E.Z.N.A.® FFPE DNA Kit uses a spin column and buffer system for isolation of genomic DNA. QIAamp® DNA FFPE Tissue Kit uses QIAamp® DNA Micro technology for purification of genomic and mitochondrial DNA.
[0262] In some instances, the disclosed methods may further comprise determining or acquiring a yield value for the nucleic acid extracted from the sample and comparing the determined value to a reference value. In some instances, the disclosed methods may further comprise determining or acquiring a value for the size (or average size) of nucleic acid fragments in the sample, and comparing the determined or acquired value to a reference value, e.g., a size (or average size) of at least 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 base pairs (bps). In some instances, one or more parameters described herein may be adjusted or selected in response to this determination.
[0263] After isolation, the nucleic acids are typically dissolved in a slightly alkaline buffer, e.g., Tris-EDTA (TE) buffer, or in ultra-pure water. In some instances, the isolated nucleic acids (e.g., genomic DNA) may be fragmented or sheared by using any of a variety of techniques known to those of skill in the art. For example, genomic DNA can be fragmentedby physical shearing methods, enzymatic cleavage methods, chemical cleavage methods, and other methods known to those of skill in the art. Methods for DNA shearing are described in Example 4 in International Patent Application Publication No. WO 2012 / 092426. In some instances, alternatives to DNA shearing methods can be used to avoid a ligation step during library preparation.B. Library Preparation
[0264] In some instances, the nucleic acids isolated from the sample may be used to construct a library (e.g., a nucleic acid library as described herein). Exemplary methods for library preparation may be found in, for instance, the Examples below. Additional exemplary methods for library preparation may be found in, for instance, the methods described in WO2023129965A2 (the content of which is herein incorporated by reference in its entirety), in which an additional primer extension step is added after adaptor ligation and before cytosine conversion to synthesize a conversion-resistant DNA copy which preserved the genomic sequences of original DNA molecules in the input cfDNA material.
[0265] Various types of libraries may be used with the methods described herein. For example, both Methyl-seq and Single workflow libraries may be used in hybrid capture reaction with methyl enrichment baits to enrich signals from regions that contain aberrant methylation signal in cancer.
[0266] In some instances, the nucleic acids are fragmented using any of the methods described above, optionally subjected to repair of chain end damage, and optionally ligated to synthetic adapters, primers, and / or barcodes (e.g., amplification primers, sequencing adapters, flow cell adapters, substrate adapters, sample barcodes or indexes, and / or unique molecular identifier sequences), size-selected (e.g., by preparative gel electrophoresis), and / or amplified (e.g., using PCR, a non-PCR amplification technique, or an isothermal amplification technique). In some instances, the fragmented and adapter-ligated group of nucleic acids is used without explicit size selection or amplification prior to hybridization-based selection of target sequences. In some instances, the nucleic acid is amplified by any of a variety of specific or nonspecific nucleic acid amplification methods known to those of skill in the art. In some instances, the nucleic acids are amplified, e.g., by a whole-genome amplification method such as random-primed strand-displacement amplification. Examples of nucleic acid library preparation techniques for next- generation sequencing are described in, e.g., van Dijket al. (2014), Exp. Cell Research 322:12 - 20, and Illumina’s genomic DNA sample preparation kit.
[0267] In some instances, the resulting nucleic acid library may contain all or substantially all of the complexity of the genome. The term “substantially all” in this context refers to the possibility that there can in practice be some unwanted loss of genome complexity during the initial steps of the procedure. The methods described herein also are useful in cases where the nucleic acid library comprises a portion of the genome, e.g., where the complexity of the genome is reduced by design. In some instances, any selected portion of the genome can be used with a method described herein. For example, in certain embodiments, the entire exome or a subset thereof is isolated. In some instances, the library may include at least 95%, 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, or 5% of the genomic DNA. In some instances, the library may consist of cDNA copies of genomic DNA that includes copies of at least 95%, 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, or 5% of the genomic DNA. In certain instances, the amount of nucleic acid used to generate the nucleic acid library may be less than 5 micrograms, less than 1 microgram, less than 500 ng, less than 200 ng, less than 100 ng, less than 50 ng, less than 10 ng, less than 5 ng, or less than 1 ng.
[0268] In some instances, a library (e.g., a nucleic acid library) includes a collection of nucleic acid molecules. As described herein, the nucleic acid molecules of the library can include a target nucleic acid molecule (e.g., a tumor nucleic acid molecule, a reference nucleic acid molecule and / or a control nucleic acid molecule; also referred to herein as a first, second and / or third nucleic acid molecule, respectively). The nucleic acid molecules of the library can be from a single subject or individual. In some instances, a library can comprise nucleic acid molecules derived from more than one subject e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30 or more subjects). For example, two or more libraries from different subjects can be combined to form a library having nucleic acid molecules from more than one subject (where the nucleic acid molecules derived from each subject are optionally ligated to a unique sample barcode corresponding to a specific subject). In some instances, the subject is a human having, or at risk of having, a cancer or tumor.
[0269] In some instances, the library (or a portion thereof) may comprise one or more subgenomic intervals. In some instances, a subgenomic interval can be a single nucleotide position, e.g., a nucleotide position for which an aberrant methylation state at the position is associated (positively or negatively) with a disease or condition, such as, e.g., a tumor. Insome instances, a subgenomic interval comprises more than one nucleotide position. Such instances include sequences of at least 2, 5, 10, 50, 100, 150, 250, or more than 250 nucleotide positions in length. Subgenomic intervals can comprise, e.g., one or more entire genes (or portions thereof), one or more exons or coding sequences (or portions thereof), one or more introns (or portion thereof), one or more micro satellite region (or portions thereof), or any combination thereof. A subgenomic interval can comprise all or a part of a fragment of a naturally occurring nucleic acid molecule, e.g., a genomic DNA molecule. For example, a subgenomic interval can correspond to a fragment of genomic DNA which is subjected to a sequencing reaction. In some instances, a subgenomic interval is a continuous sequence from a genomic source. In some instances, a subgenomic interval includes sequences that are not contiguous in the genome, e.g., subgenomic intervals in cDNA can include exon-exon junctions formed as a result of splicing. In some instances, the subgenomic interval comprises a tumor nucleic acid molecule. In some instances, the subgenomic interval comprises a nontumor nucleic acid molecule.C. End Repair
[0270] In some embodiments, the nucleic acid molecules in a sample may further undergo end repair as part of, for instance, a Methyl-Seq workflow. End repair of a nucleic acid molecule comprises phosphorylating the 5' end of the nucleic acid molecule. This process may be performed using a polynucleotide kinase. The polynucleotide kinase preferably lacks exonuclease activity (e.g., 5' to 3' exonuclease activity), thus avoiding any blunting of the ends of the nucleic acid molecule. T4 polynucleotide kinase and Thermo PNK are exemplary polynucleotide kinases that may be used for this process.X. Adaptor Ligation
[0271] In some embodiments, the methods describe herein further comprise ligating one or more adapters onto one or more nucleic acid molecules. Suitable adapters include, but are not limited to, e.g., amplification primers, flow cell adaptor sequences, substrate adapter sequences, or sample index sequences. Such adaptor ligation may take place prior to or after the cytosine conversion step. For example, in some embodiments, an adapter may be ligated to unconverted nucleic acid molecules (e.g., prior to the conversion step). Alternatively, in some embodiments, the adapters can be ligated to converted nucleic acid molecules. In someembodiments, the adapters are ligated to converted nucleic acid molecules prior to amplifying the converted nucleic acid molecules.XI. Amplification
[0272] In some embodiments, the methods further comprise, prior to bait hybridization, amplifying the converted nucleic acid molecules into a set of amplicons. This may be accomplished by various means known in the art, such as, for instance, by a polymerase chain reaction (PCR) amplification technique, a non-PCR amplification technique, and / or an isothermal amplification technique. In some embodiments, the amplification step is performed by PCR. A variety of PCR techniques suitable for NGS are known in the art. PCR amplification may also be used to convert uracil residues or other products of cytosine conversion into thymine residues. In some embodiments, the PCR amplification is performed using deoxyribonucleotides comprising thymine.XII. Target Capture
[0273] The baits of the present disclosure may be used in, for instance, a target capture step of Methyl-Seq library preparation. Thus, the methods described herein may further comprise one or more target capture steps. Target capture takes place prior to sequencing the plurality of polynucleotides or providing a plurality of sequence reads, and involves selectively capturing one or more nucleic acid molecules using one or more of the baits described herein. This may comprise, for instance, selectively enriching for a plurality of nucleic acids or nucleic acid fragments corresponding to a genomic locus that comprises a cluster of two or more CpG dinucleotides to produce an enriched sample by selectively capturing nucleic acid molecules using one or more of the baits described herein.
[0274] Target capture using the methods described herein involves hybridizing nucleic acid molecules with one or more nucleic acid baits as described herein. This is done under conditions suitable for hybridization, wherein a plurality of nucleic acid molecules are capable of hybridization with a bait nucleic acid. After hybridization, the plurality of hybrids are separated from nucleic acid molecules that are not hybridized to a bait. Separating can be accomplished in a variety of ways, such as, for instance, isolating a plurality of nucleic acid molecules that hybridized with a bait nucleic acid. After separation, the isolated or otherwiseseparated plurality of nucleic acid molecules that hybridized with the bait molecule are sequenced by NGS.
[0275] For example, one or more baits described herein can be used to hybridize with a genomic locus of interest or fragment thereof, e.g., comprising a cluster of two or more CpG dinucleotides. For a general discussion of previously-available bait designs, see, e.g., Graham, B.I. et al. Twist Fast Hybridization targeted methylation sequencing: a tunable target enrichment solution for methylation detection [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2021; 2021 Apr 10-15 and May 17-21. Philadelphia (PA): AACR; Cancer Res 2021;81(13_Suppl):Abstract nr 2098.
[0276] In some embodiments, a hybrid capture approach is used. Further details about this and other hybrid capture processes can be found in U.S. Pat. No. 9,340,830; Frampton, G.M. et al. (2013) Nat. Biotech. 31:1023-1031; and Montesion, M. et al., Cancer Discovery (2021) l l(2):282-92.A. Target Capture Reagents
[0277] The methods described herein may comprise contacting a nucleic acid library with a plurality of baits in order to select and capture a plurality of specific target sequences e.g., gene sequences or fragments thereof) for analysis. In some instances, a bait (i.e., a molecule which can bind to and thereby allow capture of a target molecule, also described herein as an enrichment bait) is used to select subject intervals to be analyzed. For example, a target capture reagent can be a bait molecule, e.g., a nucleic acid molecule (e.g., a single- or doublestranded DNA molecule, a single-stranded RNA molecule, or an RNA-DNA duplex) which can hybridize to (i.e., is complementary to) a target molecule, and thereby allows capture of the target nucleic acid. In some instances, the target capture reagent, e.g., a bait molecule (or bait sequence), is a capture oligonucleotide (or capture probe). In some instances, the target nucleic acid is a genomic DNA molecule, an RNA molecule, a cDNA molecule derived from an RNA molecule, a microsatellite DNA sequence, and the like. In some instances, the target capture reagent is suitable for solution-phase hybridization to the target. In some instances, the target capture reagent is suitable for solid-phase hybridization to the target. In some instances, the target capture reagent is suitable for both solution-phase and solid-phase hybridization to the target. The design and construction of conventional target capturereagents is described in more detail in, e.g., International Patent Application Publication No. WO 2020 / 236941, the entire content of which is incorporated herein by reference; the design of the baits according to the methods described herein is detailed in the sections above.
[0278] The methods described herein provide for optimized sequencing of a large number of genomic loci (e.g., genes or gene products (e.g., mRNA, such as, e.g., 5mC in RNA via RNA bisulfite sequencing), micro satellite loci, etc.) from samples (e.g., cancerous tissue specimens, liquid biopsy samples, and the like) from one or more subjects by the appropriate selection and design of target capture reagents (e.g., according to one or more of the bait design strategies described herein) to select the target nucleic acid molecules to be sequenced. In some instances, a target capture reagent may hybridize to a specific target locus, e.g., a specific target locus or fragment thereof. In some instances, a target capture reagent may hybridize to a specific group of target loci, e.g., a specific group of loci or fragments thereof. In some instances, a plurality of target capture reagents comprising a mix of target- specific and / or group- specific target capture reagents may be used.
[0279] In some instances, the number of target capture reagents (e.g., bait molecules) in the plurality of target capture reagents (e.g., a bait set) contacted with a nucleic acid library to capture a plurality of target sequences for nucleic acid sequencing is greater than 10, greater than 50, greater than 100, greater than 200, greater than 300, greater than 400, greater than 500, greater than 600, greater than 700, greater than 800, greater than 900, greater than 1,000, greater than 1,250, greater than 1,500, greater than 1,750, greater than 2,000, greater than 3,000, greater than 4,000, greater than 5,000, greater than 10,000, greater than 25,000, or greater than 50,000.
[0280] In some instances, the overall length of a bait molecule can be between about 70 nucleotides and 300 nucleotides. In one instance, the bait molecule length is between about 100 and 300 nucleotides, 110 and 200 nucleotides, or 120 and 170 nucleotides, in length. In addition to those mentioned above, intermediate bait molecule lengths of less than about 70, less than about 80, less than about 90, less than about 100, less than about 110, less than about 120, less than about 130, less than about 140, less than about 150, less than about 160, less than about 170, less than about 180, less than about 190, less than about 200, less than about 210, less than about 220, less than about 230, less than about 240, less than about 250, or less than about 300 nucleotides in length can be used in the methods described herein.
[0281] In some instances, each bait sequence can include: (i) a target- specific capture sequence (e.g., hypermethylation- specific or hypomethylation-specific bait corresponding to a locus or microsatellite locus-specific complementary sequence according to the bait designs disclosed herein related to), (ii) an adapter, primer, barcode, and / or unique molecular identifier sequence, and (iii) universal tails on one or both ends. As used herein, the term “target capture reagent” or “bait” can refer to the target- specific target capture sequence or to the entire target capture reagent oligonucleotide including the target- specific target capture sequence.
[0282] In some instances, the target- specific capture sequences in the target capture reagents are between about 40 nucleotides and 1000 nucleotides in length. In some instances, the target- specific capture sequence is between about 70 nucleotides and 300 nucleotides in length. In some instances, the target- specific sequence is between about 100 nucleotides and 200 nucleotides in length. In yet other instances, the target- specific sequence is between about 120 nucleotides and 170 nucleotides in length, typically 120 nucleotides in length. Intermediate lengths in addition to those mentioned above also can be used in the methods described herein, such as target- specific sequences of about 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 300, 400, 500, 600, 700, 800, and 900 nucleotides in length, as well as target- specific sequences of lengths between the above-mentioned lengths.
[0283] In some instances, there are multiple target- specific capture sequences corresponding to multiple adjacent genomic loci within a bait molecule. In some instances, a bait molecule comprises multiple different target- specific capture sequences, such as, e.g., each corresponding to a hypermethylation signal at different adjacent genomic loci or each corresponding to a hypomethylation signal at different adjacent genomic loci.
[0284] In some instances, the target capture reagent may be designed to select a subject interval containing one or more rearrangements, e.g., an intron containing a genomic rearrangement. In such instances, the target capture reagent is designed such that repetitive sequences are masked to increase the selection efficiency. In those instances where the rearrangement has a known juncture sequence, complementary target capture reagents can be designed to recognize the juncture sequence to increase the selection efficiency.
[0285] In some instances, the disclosed methods may comprise the use of target capture reagents designed to capture two or more different target categories, each category having adifferent target capture reagent design strategy. In some instances, the hybridization-based capture methods and target capture reagent compositions disclosed herein may provide for the capture and homogeneous coverage of a set of target sequences, while minimizing coverage of genomic sequences outside of the targeted set of sequences. In some instances, the target sequences may include the entire exome of genomic DNA or a selected subset thereof. In some instances, the target sequences may include, e.g., a large chromosomal region (e.g., a whole chromosome arm). The methods and compositions disclosed herein provide different target capture reagents for achieving different sequencing depths and patterns of coverage for complex sets of target nucleic acid sequences.
[0286] DNA molecules are typically used as target capture reagent sequences, although RNA molecules can also be used. In some instances, a DNA molecule target capture reagent can be single stranded DNA (ssDNA) or double- stranded DNA (dsDNA), as described herein. In some instances, an RNA-DNA duplex is more stable than a DNA-DNA duplex and therefore provides for potentially better capture of nucleic acids.
[0287] In some instances, the disclosed methods comprise providing a selected set of nucleic acid molecules (e.g., a library catch) captured from one or more nucleic acid libraries. For example, the method may comprise: providing one or a plurality of nucleic acid libraries, each comprising a plurality of nucleic acid molecules (e.g., a plurality of target nucleic acid molecules and / or reference nucleic acid molecules) extracted from one or more samples from one or more subjects; contacting the one or a plurality of libraries (e.g., in a solution-based hybridization reaction) with one, two, three, four, five, or more than five pluralities of target capture reagents (e.g., one or more baits designed according to the principles described herein) to form a hybridization mixture comprising a plurality of target capture reagent / nucleic acid molecule hybrids; separating the plurality of target capture reagent / nucleic acid molecule hybrids from said hybridization mixture, e.g., by contacting said hybridization mixture with a binding entity that allows for separation of said plurality of target capture reagent / nucleic acid molecule hybrids from the hybridization mixture, thereby providing a library catch (e.g., a selected or enriched subgroup of nucleic acid molecules from the one or a plurality of libraries). For example, in some embodiments, separation may be accomplished using biotinylated baits and streptavidin-coated beads.
[0288] In some instances, the disclosed methods may further comprise amplifying the library catch (e.g., by performing PCR). In other instances, the library catch is not amplified.
[0289] In some instances, the target capture reagents can be part of a kit which can optionally comprise instructions, standards, buffers or enzymes or other reagents.B. Hybridization Conditions
[0290] As noted above, the methods disclosed herein may include the step of contacting a library (e.g., a nucleic acid library) with a plurality of target capture reagents to provide a selected library target nucleic acid sequences (z.e., the library catch). The contacting step can be effected in, e.g., solution-based hybridization. In some instances, the method includes repeating the hybridization step for one or more additional rounds of solution-based hybridization. In some instances, the method further includes subjecting the library catch to one or more additional rounds of solution-based hybridization with the same or a different collection of target capture reagents.
[0291] In some instances, the contacting step is effected using a solid support, e.g., an array. Suitable solid supports for hybridization are described in, e.g., Albert, T.J. et al. (2007) Nat. Methods 4(1 l):903-5; Hodges, E. et al. (2007) Nat. Genet. 39(12): 1522-7; and Okou, D.T. et al. (2007) Nat. Methods 4(11):907-9, the contents of which are incorporated herein by reference in their entireties.
[0292] Hybridization methods that can be adapted for use in the methods herein are described in the art, e.g., as described in International Patent Application Publication No. WO 2012 / 092426. Methods for hybridizing target capture reagents to a plurality of target nucleic acids are described in more detail in, e.g., International Patent Application Publication No. WO 2020 / 236941, the entire content of which is incorporated herein by reference.
[0293] Hybridization conditions may comprise one or more conditions of various stringency. For instance, the hybridization conditions may comprise one or more temperatures lower than the temperature in a wash step. The hybridization conditions may comprise one or more temperatures higher than the temperature in a wash step.
[0294] In some instances, hybridization conditions include a pre-hybridization mix comprising one or more baits described herein, one or more blockers (for example, one or two commercially available blockers, such as, for example, human Cot-1 DNA and / or sheared salmon sperm), a blocker solution, and / or a methyl enhancer. In some instances, a pre-hybridization mix is lyophilized. In some instances, a lyophilized pre-hybridization mix is resuspended in a hybridization mix. In some instances, a resuspended hybridization mix istopped with a hybridization enhancer. In some instances, a pre-hybridization mix and / or a hybridization mix is denatured. Denaturation comprises, for incubation at a temperature higher than the Tm of the baits (e.g., at least about 70, 75, 80, 85, 90, 95, or up to about 99°C) for a length of time sufficient to denature the baits (e.g., about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60 min or longer). Denaturation may comprise, for instance, incubation at 95°C for 5 min.
[0295] In some instances, hybridization conditions include incubation with a target in a library at, e.g., about 50, 55, 60, 65, 70, or 75°C for, e.g., about 0.25, 0.5, 0.75, 1, 1.25, 1.5,1.75, 2, 2.25, 2.5, 2.75, 3, 3.25, 3.5, 3.75, 4, 4.25, 4.5, 4.75, 5, 5.25, 5.5, 5.75, 5, 5.25, 5.5,5.75, 6, 6.25, 6.5, 6.75, 7, 7.25, 7.5, 7.75, 8 or more hours.
[0296] In some instances, a wash step comprises application of a wash buffer. In some instances, the method comprises a first wash step and a second wash step. In some instances, the method comprises a third wash step. In some instances, the method comprises more than three wash steps. In some instances, a wash step comprises incubation at a temperature lower than a Tm of a hybrid comprising a bait and a target without mismatches. In some instances, a wash step comprises incubation at a temperature lower than a Tm of a hybrid comprising a bait with fewer mismatches compared to a hybrid comprising a bait and an off-target sequence. In some instances, a wash step comprises incubation at, for example, about 35, 40, 45, 50, 55, 60, 65, or 70°C. In some instances, a wash step takes place at a higher temperature than a subsequent wash step. For example, in some instances, a first wash step comprises incubation at about 65 °C a second wash step comprises incubation at about 48 °C.XIII. Sequencing
[0297] The methods described herein may further comprise sequencing the nucleic acid molecules in the plurality of hybrids. For example, in some embodiments, the methods may comprise sequencing nucleic acid molecules hybridized that hybridized to a bait. In some embodiments, the sequencing takes place after one or more wash steps wherein nucleic acid molecules not bound to a bait and / or bound to a bait with imperfect hybridization (z.e., comprising one or more mismatched residues) are depleted. In some embodiments, after hybridization with one or more nucleic acid baits, the method may further comprise: i) separating a first portion of the hybrids in the plurality of hybrids from a second portion of hybrids in the plurality of hybrids, wherein the first portion of hybrids comprises hybridsbetween a nucleic acid bait and a nucleic acid molecule comprising the nucleic acid pattern of interest, and wherein the second portion of hybrids comprises hybrids between a nucleic acid bait and a nucleic acid molecule that does not comprise the nucleic acid pattern of interest, wherein the hybrids in the second portion of hybrids comprise one or more adenine / cytosine (A:C) mismatches; ii) comprising attaching the hybrids to a substrate and washing the substrate, wherein nucleic acid molecules in the second portion of hybrids are depleted; and iii) sequencing a plurality of nucleic acid molecules from the portion of hybrids after the nucleic acid molecules in the second portion of hybrids are depleted.
[0298] In some embodiments, a plurality of sequence reads of the present disclosure is obtained from next-generation sequencing (NGS).
[0299] NGS methods are known in the art, and are described, e.g., in Metzker, M. (2010) Nature Biotechnology Reviews 11:31-46. Platforms for next-generation sequencing include, e.g., Roche / 454’s Genome Sequencer (GS) FLX System, Illumina / Solexa’s Genome Analyzer (GA), Illumina’s HiSeq 2500, HiSeq 3000, HiSeq 4000 and NovaSeq 6000 Sequencing Systems, Life / APG’s Support Oligonucleotide Ligation Detection (SOLiD) system, Polonator’s G.007 system, Helicos BioSciences’ HeliScope Gene Sequencing system, and Pacific Biosciences’ PacBio RS system. NGS technologies can include one or more of steps, e.g., template preparation, sequencing and imaging, and data analysis.Methods for template preparation can include steps such as randomly breaking nucleic acids (e.g., genomic DNA) into smaller sizes and generating sequencing templates (e.g., fragment templates or mate-pair templates). The spatially separated templates can be attached or immobilized to a solid surface or support, allowing massive amounts of sequencing reactions to be performed simultaneously. Types of templates that can be used for NGS reactions include, e.g., clonally amplified templates originating from single DNA molecules, and single DNA molecule templates. Exemplary sequencing and imaging steps for NGS include, e.g., cyclic reversible termination (CRT), sequencing by ligation (SBL), single-molecule addition (pyro sequencing), and real-time sequencing. After NGS reads have been generated, they can be aligned to a known reference sequence or assembled de novo. For example, identifying genetic variations such as single-nucleotide polymorphism and structural variants in a sample (e.g., a tumor sample) can be accomplished by aligning NGS reads to a reference sequence (e.g., a wild type sequence). Methods of sequence alignment for NGS are described e.g., in Trapnell C. and Salzberg S.L. Nature Biotech., 2009, 27:455-457. Examples of de novoassemblies are described, e.g., in Warren R. et al., Bioinformatics, 2007, 23:500-501; Butler J. et al., Genome Res., 2008, 18:810-820; and Zerbino D.R. and Birney E., Genome Res., 2008, 18:821-829. Sequence alignment or assembly can be performed using read data from one or more NGS platforms, e.g., mixing Roche / 454 and Illumina / Solexa read data. In some embodiments, NGS is performed according to the methods described in, e.g., Frampton, G.M. et al. (2013) Nat. Biotech. 31:1023-1031; and / or Montesion, M. et al., Cancer Discovery (2021) l l(2):282-92.
[0300] In some embodiments, the sequencing comprises use of a massively parallel sequencing (MPS) technique (such as, e.g., next generation sequencing (NGS) and / or use of a next generation sequencer), whole genome sequencing, whole exome sequencing, targeted sequencing, direct sequencing, or Sanger sequencing technique.
[0301] In some embodiments, the sequencing produces a plurality of sequence reads, one or more of which overlap with one or more loci (e.g., at least 10 loci and even up to or more than tens of thousands of loci) within one or more subgenomic intervals in the sample. In some embodiments, each of the one or more loci comprises at least 10 loci each comprising more than two CpG sites.
[0302] In some embodiments, a plurality of sequence reads of the present disclosure is obtained from next-generation sequencing (NGS). In some embodiments, the sequencing comprises bisulfite sequencing, whole genome bisulfite sequencing (WGBS), APOBEC-seq, methyl-CpG-binding domain (MBD) protein capture, methyl-DNA immunoprecipitation (MeDIP-seq), methylation sensitive restriction enzyme sequencing (MSRE / MRE-Seq or Methyl-Seq), oxidative bisulfite sequencing (oxBS-Seq), reduced representative bisulfite sequencing (RRBS), or Tet-assisted bisulfite sequencing (TAB-Seq).
[0303] NGS methods are known in the art, and are described, e.g., in Metzker, M. (2010) Nature Biotechnology Reviews 11:31-46. NGS technologies can include one or more of steps, e.g., template preparation, sequencing and imaging, and data analysis. Methods for template preparation can include steps such as randomly breaking nucleic acids (e.g., genomic DNA) into smaller sizes and generating sequencing templates (e.g., fragment templates or mate-pair templates). The spatially separated templates can be attached or immobilized to a solid surface or support, allowing massive amounts of sequencing reactions to be performed simultaneously. Types of templates that can be used for NGS reactions include, e.g., clonally amplified templates originating from single DNA molecules, and single DNA moleculetemplates. Exemplary sequencing and imaging steps for NGS include, e.g., cyclic reversible termination (CRT), sequencing by ligation (SBL), single-molecule addition (pyro sequencing), and real-time sequencing.
[0304] After NGS reads have been generated, they can be aligned to a known reference sequence or assembled de novo. For example, identifying genetic variations such as singlenucleotide polymorphism and structural variants and / or signals of epigenetic variations, such as signals of aberrant methylation states in a sample (in, e.g., a tumor sample) can be accomplished by aligning NGS reads to a reference sequence (e.g., a wild type sequence and / or a nucleic acid sequence in which methylation information is embedded via converted from a control or wild type sequence of known methylation state). Methods of sequence alignment for NGS are described e.g., in Trapnell C. and Salzberg S.L. Nature Biotech., 2009, 27:455-457. Examples of de novo assemblies are described, e.g., in Warren R. et al., Bioinformatics, 2007, 23:500-501; Butler J. et al., Genome Res., 2008, 18:810-820; and Zerbino D.R. and Birney E., Genome Res., 2008, 18:821-829. Sequence alignment or assembly can be performed using read data from one or more NGS platforms. In some embodiments, NGS is performed according to the methods described in, e.g., Frampton, G.M. et al. (2013) Nat. Biotech. 31:1023-1031; and / or Montesion, M. et al., Cancer Discovery (2021) l l(2):282-92.
[0305] In some embodiments, a plurality of sequence reads is obtained by performing sequencing on nucleic acids captured by hybridization with a bait molecule. In some embodiments, the plurality of sequence reads is obtained by performing whole exome sequencing on nucleic acids captured by hybridization with a bait molecule. In some embodiments, the plurality of sequence reads is obtained by performing next-generation sequencing (NGS) on nucleic acids captured by hybridization with the bait molecule.
[0306] In some embodiments, a plurality of sequence reads of the present disclosure includes paired-end sequence reads. In some embodiments, consensus methylation pattern and / or CCF are determined based on paired-end sequence reads corresponding to one or more cluster(s). In some embodiments, consensus unmethylation pattern and / or CCUF are determined based on paired-end sequence reads corresponding to one or more cluster(s). Generally, paired-end sequencing methodologies are described, e.g., in W02007 / 010252, W02007 / 091077, and WO03 / 74734. This approach utilizes pairwise sequencing of a double- stranded polynucleotide template, which results in the sequential determination of nucleotidesequences in two distinct and separate regions of the polynucleotide template. The paired-end methodology makes it possible to obtain two linked or paired reads of sequence information from each double- stranded template on a clustered array, rather than just a single sequencing read as can be obtained with other methods. Paired end sequencing technology can make special use of clustered arrays, generally formed by solid-phase amplification, for example as set forth in WO03 / 74734. Target nucleic acid molecules, fitted with adapters, are immobilized to a solid support at the 5' ends of each strand, for example, via bridge amplification, forming dense clusters of double stranded DNA. Because both strands are immobilized at their 5' ends, sequencing primers are then hybridized to the free 3' end and sequencing by synthesis is performed. Adapter sequences can be inserted in between target sequences to allow for up to four reads from each duplex, as described in W02007 / 091077. In a further adaptation of this methodology, specific strands can be cleaved in a controlled fashion as set forth in W02007 / 010252. As a result, the timing of the sequencing read for each strand can be controlled, permitting sequential determination of the nucleotide sequences in two distinct and separate regions on complementary strands of the doublestranded template. See, e.g., US Pat. No. 10,174,372.
[0307] In some embodiments, the plurality of sequence reads includes unpaired sequence reads.
[0308] In some embodiments, the methods of the present disclosure further comprise, prior to determining a consensus methylation pattern and CCF: demultiplexing sequence reads from a plurality of sequence reads. In some embodiments, the methods of the present disclosure further comprise, prior to determining a consensus methylation pattern and CCF: performing alignment of sequence reads from the plurality to a reference genome, e.g., a human reference genome. In some embodiments, the alignment is a three-letter alignment to a human reference genome. In some embodiments, the methods of the present disclosure further comprise, prior to determining a consensus methylation pattern and CCF: excluding sequencing reads from the plurality that failed to undergo cytosine conversion. In some embodiments, the methods of the present disclosure further comprise, prior to determining a consensus methylation pattern and CCF: excluding sequence reads with a base other than cytosine or thymine at a first position of at least one of the CpG dinucleotides. For example, these can be due to sequencing errors or mutations (somatic or germline). In some embodiments, the methods of the present disclosure further comprise, prior to determining aconsensus methylation pattern and CCF: excluding sequence reads with a base quality below a threshold base quality. In some embodiments, base calls at a cytosine within a CpG dinucleotide are determined using two overlapping paired-end sequence reads.A. Methyl-Seq
[0309] In some embodiments, the methods disclosed herein involve one or more steps of a Methyl-seq workflow, also known as NGS (next-generation sequencing) for methylation analysis. An exemplary methyl-seq workflow can include a chemical process (e.g., bisulfite or enzymatic) that converts unmethylated cytosines to thymine while leaving methylated cytosines intact during NGS library construction. As such, DNA methylation status in the original DNA molecules in the input material can be inferred and compared to identify aberrant methylation states, such as, for instance, healthy versus cancer samples. An exemplary Methyl-seq workflow in the context of making cancer identity calls and determining tumor fractions is shown in FIG. 1 from left to right. The methods described herein may be used to determine methylation scores and detect aberrant methylation statuses not just in association with cancer, but for any DNA sequence with a potential hypermethylation or a potential hypomethylation status, such as hypomethylation or hypermethylation signals associated with any other diseases or conditions. Accordingly, the methods described herein may also be used in other workflows that involve selective detection of a methylation pattern of interest from a sample.
[0310] As shown in FIG. 1, an exemplaryMethyl-Seq workflow begins with a sample, such as a cfDNA sample. The sample undergoes, for example, end repair and adaptor ligation, cytosine conversion, library PCR, target capture, and sequencing. Adaptors ligation may take place before or after conversion. For instance, in some embodiments, the adapter is ligated to unconverted molecules. However, in some embodiments, the adapters can alternatively be ligated to converted molecules. After sequencing, hypomethylation and / or hypermethylation scores are determined, and diagnoses are made, such as, for instance, cancer detection calls and / or tumor fraction calculations.
[0311] Some Methyl-seq methods rely upon library construction and adapter ligation, followed by standard bisulfite conversion and sequencing (e.g., WGBS). Alternatively, bisulfite treatment can be carried out prior to adaptor ligation (see, e.g., Miura, F. et al. Amplification-free whole-genome bisulfite sequencing by post-bisulfite adaptor tagging,Nucleic Acids Res., 40, el36F (2012)). More recent techniques use other cytosine conversion methods such as enzymatic approaches in order to reduce damage to DNA caused by bisulfite, e.g., as in the commercially available Methyl-seq kits. Steps of library amplification, quantification, and sequencing generally follow bisulfite conversion. In some embodiments, prior to Methyl-seq, nucleic acids are extracted from a sample. In some embodiments, prior to Methyl-seq, nucleic acids are subjected to fragmentation, repair, and adaptor ligation. As noted previously, cytosine conversion can be carried out before or after adaptor ligation. In some embodiments, DNA repair is performed after cytosine conversion. PCR amplification (generally at least two cycles) is performed after cytosine conversion to convert uracil residues (generated by formerly unmethylated cytosines) into thymine, and is accomplished using a polymerase that is able to read uracil (excluding polymerases with proofreading and repair activities). In some embodiments, prior to sequencing, fragments are enriched for desired length.
[0312] Additional exemplary details about various steps of Methyl-Seq workflows that may be used as part of the present invention are discussed elsewhere herein.XIV. Methylation Scores
[0313] In some embodiments involving one or more of the nucleic acid baits disclosed herein, methylation scores are determined following sequencing. For example, one or more hypermethylation score(s) and / or one or more hypomethylation score(s) may be determined for the genomic loci corresponding to any baits used. A methylation score of the present disclosure ties the sequencing data produced from the converted nucleic acid molecules that hybridized to a bait back to the methylation pattern present in the sample from the subject, by, for instance, comparing the pattern of C and T residues in sequences converted nucleic acids that hybridized to a bait to a known genomic sequence (e.g., non-converted genomic DNA from the same genomic locus, optionally from the same subject). Such comparisons may be done with, for example, the assistance of an alignment tool. Examples of alignment tools optimized for aligning sequence reads for converted DNA include, but are not limited to, NovoAlign (Novocraft Technologies, Selangor, Malaysia), and the Bismark tool (Krueger et al. (2011), “Bismark: A Flexible Aligner and Methylation Caller for Bisulfite-Seq Applications”, Bioinformatics 27(11): 1571-1572).
[0314] In some embodiments, determining a methylation score comprises, e.g., determining the ratio of sequencing reads corresponding to converted nucleic acid molecules that hybridized to a bait to total sequencing reads. In some embodiments, determining a methylation score comprises, e.g., quantifying and / or normalizing the hypermethylation signal and / or the hypomethylation signal, such as, for example, based on a normalization region on a panel, in which the normalization region comprises mixed baits (z.e., a mix of hy ermethylation- specific and hypomethylation- specific).
[0315] In some embodiments, determining a methylation score comprises generating (e.g., by a processor, a cluster consensus fraction (CCF) for a cluster (e.g., a cluster of two or more CpG dinucleotides), wherein the CCF represents a fraction of sequence reads corresponding to the cluster that show a consensus methylation pattern out of a total number of sequence reads from the plurality corresponding to the cluster; detecting one or more of the methylation level or the unmethylation level of the cluster based on the CCF; and generating a genomic and / or epigenetic profile for the subject based at least in part on the detected methylation level, the detected unmethylation level, or both.
[0316] In some embodiments, the methods comprise sequencing (e.g., by a sequencer) a plurality of nucleic acid fragments to obtain a plurality of sequence reads, wherein the plurality of nucleic acid fragments has undergone cytosine conversion, and wherein the plurality of nucleic acid fragments corresponds to a genomic locus comprising a cluster of two or more CpG dinucleotides; determining (e.g., by a processor) a consensus methylation pattern for the cluster, wherein the consensus methylation pattern represents each CpG dinucleotide in the cluster for which methylation was detected in at least one sequence read from the plurality of sequence reads based on the cytosine conversion; and generating (e.g., by a processor) a cluster consensus fraction (CCF) for the cluster, wherein the CCF represents a fraction of sequence reads corresponding to the cluster that show the consensus methylation pattern out of a total number of sequence reads from the plurality corresponding to the cluster.
[0317] In other embodiments, the methods comprise sequencing (e.g., by a sequencer) a plurality of nucleic acid fragments to obtain a plurality of sequence reads, wherein the plurality of nucleic acid fragments has undergone cytosine conversion, and wherein the plurality of nucleic acid fragments corresponds to a genomic locus comprising a cluster of two or more CpG dinucleotides; determining (e.g., by a processor) a consensusunmethylation pattern for the cluster, wherein the consensus unmethylation pattern represents each CpG dinucleotide in the cluster for which methylation was not detected in at least one sequence read from the plurality of sequence reads based on the cytosine conversion; and generating (e.g., by a processor) a cluster consensus unmethylation fraction (CCUF) for the cluster, wherein the CCUF represents a fraction of sequence reads corresponding to the cluster that show the consensus unmethylation pattern out of a total number of sequence reads from the plurality corresponding to the cluster. It will be appreciated by those skilled in the art that the methods disclosed herein for measuring methylation (e.g., CCMF) could also be applied to measuring un- or non-methylated sites (e.g., CCUF) as well. It will be understood that the cluster consensus methylation fraction, the cluster consensus unmethylation fraction, or both may be generally referred to as a cluster consensus fraction (CCF)
[0318] Other aspects of the present disclosure relate to methods of detecting a disease or condition (e.g., cancer) in an individual, comprising detecting methylation level (e.g., of a cluster of two or more CpG dinucleotides) according to any one of the methods of the present disclosure. In some embodiments, the methods comprise sequencing (e.g., by a sequencer) a plurality of nucleic acid fragments to obtain a plurality of sequence reads, wherein the plurality of nucleic acid fragments is obtained from a sample from the individual and has subsequently undergone cytosine conversion, and wherein the plurality of nucleic acid fragments corresponds to a genomic locus comprising a cluster of two or more CpG dinucleotides; determining (e.g., by a processor) a consensus methylation pattern for the cluster, wherein the consensus methylation pattern represents each CpG dinucleotide in the cluster for which methylation was detected in at least one sequence read from the plurality of sequence reads based on the cytosine conversion; and generating (e.g., by a processor) a cluster consensus fraction (CCF) for the cluster, wherein the CCF represents a fraction of sequence reads corresponding to the cluster that show the consensus methylation pattern out of a total number of sequence reads from the plurality corresponding to the cluster. In some embodiments, a CCF at or above a threshold or reference value indicates presence of the disease or condition in the individual and identifies the individual as having the disease or condition. In some embodiments, a CCF below a threshold or reference value does not indicate presence of the disease or condition in the individual and identifies the individual as not having the disease or condition. In some embodiments, the methods may find use, e.g., in screening for a disease or condition (e.g., a new diagnosis in an individual that has notpreviously been diagnosed with cancer, or the same type of cancer) or monitoring the individual for recurrence or minimal residual disease (e.g., in an individual that has previously been diagnosed with cancer and achieved remission).
[0319] Other aspects of the present disclosure relate to methods of screening an individual suspected of having a disease or condition (e.g., cancer), comprising detecting methylation level (e.g., of a cluster of two or more CpG dinucleotides) according to any one of the methods of the present disclosure. In some embodiments, the methods comprise sequencing (e.g., by a sequencer) a plurality of nucleic acid fragments to obtain a plurality of sequence reads, wherein the plurality of nucleic acid fragments is obtained from a sample from the individual and has subsequently undergone cytosine conversion, and wherein the plurality of nucleic acid fragments corresponds to a genomic locus comprising a cluster of two or more CpG dinucleotides; determining (e.g., by a processor) a consensus methylation pattern for the cluster, wherein the consensus methylation pattern represents each CpG dinucleotide in the cluster for which methylation was detected in at least one sequence read from the plurality of sequence reads based on the cytosine conversion; and generating (e.g., by a processor) a cluster consensus fraction (CCF) for the cluster, wherein the CCF represents a fraction of sequence reads corresponding to the cluster that show the consensus methylation pattern out of a total number of sequence reads from the plurality corresponding to the cluster. In some embodiments, a CCF at or above a threshold or reference value indicates presence of the disease or condition in the individual and identifies the individual as likely to have the disease or condition. In some embodiments, a CCF below a threshold or reference value does not indicate presence of the disease or condition in the individual and identifies the individual as likely not to have the disease or condition. In some embodiments, the methods may find use, e.g., in screening for the disease or condition (e.g., a new diagnosis in an individual that has not previously been diagnosed with cancer, or the same type of cancer) or monitoring the individual for recurrence or minimal residual disease (e.g., in an individual that has previously been diagnosed with cancer and achieved remission).
[0320] Other aspects of the present disclosure relate to methods of determining prognosis of an individual having the disease or condition (e.g., cancer), comprising detecting methylation level (e.g., of a cluster of two or more CpG dinucleotides) according to any one of the methods of the present disclosure. In some embodiments, the methods comprise sequencing (e.g., by a sequencer) a plurality of nucleic acid fragments to obtain a plurality of sequencereads, wherein the plurality of nucleic acid fragments is obtained from a sample from the individual and has subsequently undergone cytosine conversion, and wherein the plurality of nucleic acid fragments corresponds to a genomic locus comprising a cluster of two or more CpG dinucleotides; determining (e.g., by a processor) a consensus methylation pattern for the cluster, wherein the consensus methylation pattern represents each CpG dinucleotide in the cluster for which methylation was detected in at least one sequence read from the plurality of sequence reads based on the cytosine conversion; and generating (e.g., by a processor) a cluster consensus fraction (CCF) for the cluster, wherein the CCF represents a fraction of sequence reads corresponding to the cluster that show the consensus methylation pattern out of a total number of sequence reads from the plurality corresponding to the cluster. In some embodiments, a CCF at or above a threshold or reference value indicates presence of the disease or condition in the individual and determines at least in part a prognosis of the individual. In some embodiments, a CCF below a threshold or reference value does not indicate presence of the disease or condition in the individual and determines at least in part a prognosis of the individual. In some embodiments, a CCF at or above a threshold or reference value corresponds to poorer prognosis of an individual, as compared to that of an individual with a CCF below the threshold or reference value.
[0321] Other aspects of the present disclosure relate to methods of predicting survival of an individual having a disease or condition (e.g., cancer), comprising detecting methylation level (e.g., of a cluster of two or more CpG dinucleotides) according to any one of the methods of the present disclosure. In some embodiments, the methods comprise sequencing (e.g., by a sequencer) a plurality of nucleic acid fragments to obtain a plurality of sequence reads, wherein the plurality of nucleic acid fragments is obtained from a sample from the individual and has subsequently undergone cytosine conversion, and wherein the plurality of nucleic acid fragments corresponds to a genomic locus comprising a cluster of two or more CpG dinucleotides; determining (e.g., by a processor) a consensus methylation pattern for the cluster, wherein the consensus methylation pattern represents each CpG dinucleotide in the cluster for which methylation was detected in at least one sequence read from the plurality of sequence reads based on the cytosine conversion; and generating (e.g., by a processor) a cluster consensus fraction (CCF) for the cluster, wherein the CCF represents a fraction of sequence reads corresponding to the cluster that show the consensus methylation pattern out of a total number of sequence reads from the plurality corresponding to the cluster. In someembodiments, a CCF at or above a threshold or reference value indicates presence of the disease or condition in the individual and predicts at least in part the survival of the individual. In some embodiments, a CCF below a threshold or reference value does not indicate presence of the disease or condition in the individual and predicts at least in part the survival of the individual. In some embodiments, a CCF at or above a threshold or reference value corresponds to shorter survival of an individual, as compared to that of an individual with a CCF below the threshold or reference value. In some embodiments, the methylation level detected in the sample is higher than a threshold or reference value, and survival of the individual is predicted to be decreased, as compared to survival of an individual whose sample has a methylation level lower than the threshold or reference value.A. Mutation calling
[0322] The methods disclosed herein may involve mutation calling in addition to calculation of methylation scores. Base calling refers to the raw output of a sequencing device, e.g., the determined sequence of nucleotides in an oligonucleotide molecule. Mutation calling refers to the process of selecting a nucleotide value, e.g., A, G, T, or C, for a given nucleotide position being sequenced. Typically, the sequence reads (or base calling) for a position will provide more than one value, e.g., some reads will indicate a T and some will indicate a G. Mutation calling is the process of assigning a correct nucleotide value, e.g., one of those values, to the sequence. Although it is referred to as “mutation” calling, it can be applied to assign a nucleotide value to any nucleotide position, e.g., positions corresponding to converted nucleic acid sequences representing a methylation state when compared to an un-converted sequence from the same locus.
[0323] In some instances, the disclosed methods may also comprise the use of customized or tuned mutation calling algorithms or parameters thereof to optimize performance when applied to sequencing data, particularly in methods that rely on massively parallel sequencing (MPS) of a large number of diverse genetic events at a large number of diverse genomic loci (e.g., CpG loci) in samples, e.g., samples from a subject having cancer. Optimization of mutation calling is described in the art, e.g., as set out in International Patent Application Publication No. WO 2012 / 092426.
[0324] Methods for mutation calling can include one or more of the following: making independent calls based on the information at each position in the reference sequence (e.g., examining the sequence reads; examining the base calls and quality scores; determining the probability of observed bases and quality scores given a potential genotype; and assigning genotypes (e.g., using Bayes’ rule)); removing false positives (e.g., using depth thresholds to reject SNPs with read depth much lower or higher than expected; local realignment to remove false positives due to small indels); and performing linkage disequilibrium (LD) / imputation- based analysis to refine the calls.
[0325] Equations used to determine the genotype likelihood associated with a specific genotype and position are described in, e.g., Li, H. and Durbin, R. Bioinformatics, 2010; 26(5): 589-95. The prior expectation for a particular mutation in a certain cancer type can be used when evaluating samples from that cancer type. Such likelihood can be derived from public databases of cancer mutations, e.g., Catalogue of Somatic Mutation in Cancer (COSMIC), HGMD (Human Gene Mutation Database), The SNP Consortium, Breast Cancer Mutation Data Base (BIC), and Breast Cancer Gene Database (BCGD).
[0326] Examples of LD / imputation based analysis are described in, e.g., Browning, B.L. and Yu, Z. Am. J. Hum. Genet. 2009, 85(6):847-61. Examples of low-coverage SNP calling methods are described in, e.g., Li, Y. et al., Annu. Rev. Genomics Hum. Genet. 2009, 10:387- 406.
[0327] After alignment, detection of substitutions can be performed using a mutation calling method (e.g., a Bayesian mutation calling method) which is applied to each base in each of the subject intervals, e.g., exons of a gene or other locus to be evaluated, where presence of alternate alleles is observed. This method will compare the probability of observing the read data in the presence of a mutation with the probability of observing the read data in the presence of base-calling error alone. Mutations can be called if this comparison is sufficiently strongly supportive of the presence of a mutation.
[0328] An advantage of a Bayesian mutation detection approach is that the comparison of the probability of the presence of a mutation with the probability of base-calling error alone can be weighted by a prior expectation of the presence of a mutation at the site. If some reads of an alternate allele are observed at a frequently mutated site for the given cancer type, then presence of a mutation may be confidently called even if the amount of evidence of mutation does not meet the usual thresholds. This flexibility can then be used to increase detectionIl lsensitivity for even rarer mutations / lower purity samples, or to make the test more robust to decreases in read coverage. The likelihood of a random base-pair in the genome being mutated in cancer is ~le-6. The likelihood of specific mutations occurring at many sites in, for example, a typical multigenic cancer genome panel can be orders of magnitude higher. These likelihoods can be derived from public databases of cancer mutations (e.g., COSMIC).
[0329] Methods have been developed that address limited deviations from allele frequencies of 50% or 100% for the analysis of cancer DNA. (see, e.g., SNVMix -Bioinformatics. 2010 March 15; 26(6): 730-736.) Methods disclosed herein, however, allow consideration of the possibility of the presence of a mutant allele at frequencies (or allele fractions) ranging from 1% to 100% (i.e., allele fractions ranging from 0.01 to 1.0), and especially at levels lower than 50%. This approach is particularly important for the detection of mutations in, for example, low-purity FFPE samples of natural (multi-clonal) tumor DNA.
[0330] In some instances, the mutation calling method used to analyze sequence reads is not individually customized or fine-tuned for detection of different mutations at different genomic loci. In some instances, different mutation calling methods are used that are individually customized or fine-tuned for at least a subset of the different mutations detected at different genomic loci. In some instances, different mutation calling methods are used that are individually customized or fine-tuned for each different mutant detected at each different genomic loci. The customization or tuning can be based on one or more of the factors described herein, e.g., the type of cancer in a sample, the gene or locus in which the subject interval to be sequenced is located, or the variant to be sequenced. This selection or use of mutation calling methods individually customized or fine-tuned for a number of subject intervals to be sequenced allows for optimization of speed, sensitivity and specificity of mutation calling.
[0331] In some instances, a nucleotide value is assigned for a nucleotide position in each of X unique subject intervals using a unique mutation calling method, and X is at least 2, at least 3, at least 4, at least 5, at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 1000, at least 1500, at least 2000, at least 2500, at least 3000, at least 3500, at least 4000, at least 4500, at least 5000, or greater. The calling methods can differ, and thereby be unique, e.g., by relying on different Bayesian prior values.
[0332] In some instances, assigning said nucleotide value is a function of a value which is or represents the prior (e.g., literature) expectation of observing a read showing a variant, e.g., a mutation, at said nucleotide position in a tumor of type.
[0333] In some instances, the method comprises assigning a nucleotide value (e.g., calling a mutation) for at least 10, 20, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1,000 nucleotide positions, wherein each assignment is a function of a unique value (as opposed to the value for the other assignments) which is or represents the prior (e.g., literature) expectation of observing a read showing a variant, e.g., a mutation, at said nucleotide position in a tumor of type.
[0334] In some instances, assigning said nucleotide value is a function of a set of values which represent th...
Claims
CLAIMSWhat is claimed is:
1. A method of selectively capturing nucleic acid molecules from a sample from a subject, comprising: converting unmethylated cytosine (C) in one or more of the nucleic acid molecules to uracil (U), or converting methylated C in one or more of the nucleic acid molecules to U, to form converted nucleic acid molecules; amplifying the converted nucleic acid molecules into a set of amplicons; hybridizing a set of single-stranded nucleic acid baits to the set of amplicons to form a plurality of hybrids, wherein at least a portion of the hybrids in the plurality of hybrids comprise a bait from the set of single-stranded nucleic acid baits and a nucleic acid molecule comprising a nucleic acid pattern of interest hybridized to each other without mismatches, wherein the nucleic acid pattern of interest corresponds to a methylation pattern of interest at a genomic locus in the nucleic acid molecules prior to the converting; and separating the plurality of hybrids from nucleic acid molecules that are not hybridized to a bait.
2. The method of claim 1, wherein the methylation pattern of interest corresponds to a hypermethylation signal indicative of a disease or condition.
3. The method of claim 1, wherein the methylation pattern of interest corresponds to a hypomethylation signal indicative of a disease or condition.
4. A method of selectively capturing nucleic acid molecules from a sample from a subject, comprising: converting unmethylated cytosine (C) in one or more of the nucleic acid molecules to uracil (U), or converting methylated C in one or more of the nucleic acid molecules to U, to form converted nucleic acid molecules; amplifying the converted nucleic acid molecules into a set of amplicons; hybridizing a set of nucleic acid baits to the set of amplicons to form a plurality of hybrids, wherein the set of nucleic acid baits selectively hybridizes to a nucleic acid patternof interest, wherein the nucleic acid pattern of interest corresponds to a methylation pattern of interest at a genomic locus in the nucleic acid molecules prior to the converting, wherein the methylation pattern of interest corresponds to a hypermethylation signal indicative of a disease or condition; and separating the plurality of hybrids from nucleic acid molecules that are not hybridized to a bait.
5. A method of selectively capturing nucleic acid molecules from a sample from a subject, comprising: converting unmethylated cytosine (C) in one or more of the nucleic acid molecules to uracil (U), or converting methylated C in one or more of the nucleic acid molecules to U, to form converted nucleic acid molecules; amplifying the converted nucleic acid molecules into a set of amplicons; hybridizing a set of nucleic acid baits to the set of amplicons to form a plurality of hybrids, wherein the set of nucleic acid baits selectively hybridizes to a nucleic acid pattern of interest, wherein the nucleic acid pattern of interest corresponds to a methylation pattern of interest at a genomic locus in the nucleic acid molecules prior to the converting, wherein the methylation pattern of interest is associated with a hypomethylation signal indicative of a disease or condition; and separating the plurality of hybrids from nucleic acid molecules that are not hybridized to a bait.
6. The method of claim 1 comprising converting a plurality of methylation patterns of interest at a plurality of genomic loci in the nucleic acid molecules from the sample from the subject into a plurality of nucleic acid patterns of interest of the plurality of genomic loci, and hybridizing a plurality of sets of nucleic acid baits to the set of amplicons, wherein each set of nucleic acid baits of the plurality of sets of nucleic acid baits selectively hybridizes to a nucleic acid pattern of interest of the plurality of nucleic acid patterns of interest.
7. The method of claim 6, wherein the plurality of methylation patterns of interest comprises a first methylation pattern of interest and a second methylation pattern of interest, wherein thesecond methylation pattern of interest is associated with a hypermethylation signal indicative of a disease or condition.
8. The method of claim 6, wherein the plurality of methylation patterns of interest comprises a first methylation pattern of interest and a second methylation pattern of interest, wherein the second methylation pattern of interest is associated with a hypomethylation signal indicative of a disease or condition.
9. The method of claim 1, wherein a second portion of hybrids in the plurality of hybrids comprises a single-stranded nucleic acid bait from the set of single- stranded nucleic acid baits and a nucleic acid molecule that does not comprise the nucleic acid pattern of interest, wherein the hybrids in the second portion of hybrids comprise one or more adenine / cytosine (A:C) mismatches.
10. The method of claim 9, further comprising separating the portion of hybrids from the second portion of hybrids.
11. The method of claim 4 comprising converting a plurality of methylation patterns of interest at a plurality of genomic loci in the nucleic acid molecules from the sample from the subject into a plurality of nucleic acid patterns of interest of the plurality of genomic loci, and hybridizing a plurality of sets of nucleic acid baits to the set of amplicons, wherein each set of nucleic acid baits of the plurality of sets of nucleic acid baits selectively hybridizes to a nucleic acid pattern of interest of the plurality of nucleic acid patterns of interest.
12. The method of claim 11, wherein the plurality of methylation patterns of interest comprises a first methylation pattern of interest and a second methylation pattern of interest, wherein the second methylation pattern of interest is associated with a hypermethylation signal indicative of a disease or condition.
13. The method of claim 12, wherein the plurality of methylation patterns of interest comprises a first methylation pattern of interest and a second methylation pattern of interest,wherein the second methylation pattern of interest is associated with a hypomethylation signal indicative of a disease or condition.
14. The method of claim 4, wherein a second portion of hybrids in the plurality of hybrids comprises a single-stranded nucleic acid bait from the set of single- stranded nucleic acid baits and a nucleic acid molecule that does not comprise the nucleic acid pattern of interest, wherein the hybrids in the second portion of hybrids comprise one or more adenine / cytosine (A:C) mismatches.
15. The method of claim 14, further comprising separating the portion of hybrids from the second portion of hybrids.
16. The method of claim 5 comprising converting a plurality of methylation patterns of interest at a plurality of genomic loci in the nucleic acid molecules from the sample from the subject into a plurality of nucleic acid patterns of interest of the plurality of genomic loci, and hybridizing a plurality of sets of nucleic acid baits to the set of amplicons, wherein each set of nucleic acid baits of the plurality of sets of nucleic acid baits selectively hybridizes to a nucleic acid pattern of interest of the plurality of nucleic acid patterns of interest.
17. The method of claim 16, wherein the plurality of methylation patterns of interest comprises a first methylation pattern of interest and a second methylation pattern of interest, wherein the second methylation pattern of interest is associated with a hypermethylation signal indicative of a disease or condition.
18. The method of claim 17, wherein the plurality of methylation patterns of interest comprises a first methylation pattern of interest and a second methylation pattern of interest, wherein the second methylation pattern of interest is associated with a hypomethylation signal indicative of a disease or condition.
19. The method of claim 5, wherein a second portion of hybrids in the plurality of hybrids comprises a single-stranded nucleic acid bait from the set of single- stranded nucleic acid baits and a nucleic acid molecule that does not comprise the nucleic acid pattern of interest,wherein the hybrids in the second portion of hybrids comprise one or more adenine / cytosine (A:C) mismatches.
20. The method of claim 19, further comprising separating the portion of hybrids from the second portion of hybrids.
21. The method of claim 1, wherein the hybrids in the plurality of hybrids do not comprise guanine / thymine (G:T) mismatches.
Citation Information
Patent Citations
Characterizing methylated DNA, RNA, and proteins in subjects suspected of having lung neoplasia
US20220403471A1
Altered cytidine deaminases and methods of use
WO2023196572A1