Combined chemistry and computational corrections for methyl-bias

Chemical and computational methods address methylation bias in sequencing by correcting strand displacement and end-repair errors, enhancing the accuracy of methylation analysis for disease diagnosis and monitoring.

WO2026102094A1PCT designated stage Publication Date: 2026-05-15FOUNDATION MEDICINE INC
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
FOUNDATION MEDICINE INC
Filing Date
2025-11-06
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Methylation bias (mbias) introduced during end repair of DNA strands prior to methylation conversion in sequencing processes leads to inaccurate assessment of methylation status, impacting the use of methylation signatures for disease diagnosis and prognosis.

Method used

A combination of chemical and computational methods to correct methylation bias, including end repair and nick/gap repair reactions using chain termination nucleotides and ligases, followed by bisulfite conversion and computational analysis to generate mbias-corrected sequence read data.

Benefits of technology

Improves the accuracy of methylation analysis by correcting methylation bias, enabling reliable detection of methylation signatures for disease diagnosis and monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025054320_15052026_PF_FP_ABST
    Figure US2025054320_15052026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides methods for generating mbias corrected sequence fragment data comprising correcting one or more strand displacement mbias instances in one or more nucleic acid fragments from a sample obtained from a subject, performing a methylation analysis, and correcting one or more end-repair bias instances in sequence fragment data from the methylation analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No: 197102018940COMBINED CHEMISRTY AND COMPUTATIONAL CORRECTIONS FOR METHYL-BIASCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the priority benefit of United States Provisional Patent Application Serial No. 63 / 717,773, filed November 7, 2024, the contents of which are incorporated herein by reference in their entirety.FIELD

[0002] The present disclosure relates generally to methods for improving the analysis of methylation sequencing data by improving methods used to correct methyl-bias when preparing libraries for nucleic acid sequencing and in methylation analysis.BACKGROUND

[0003] Methylation sequencing data obtained using, as a non-limiting example, bisulfite conversion of non-methylated cytosines to uracil (leaving methylated cytosine bases intact) and next-generation sequencing techniques, is used to study the methylation patterns of DNA. DNA methylation patterns are epigenetic markers (e.g., modifications to DNA that do not alter the base sequence of the DNA molecule) that can impact gene expression and cell differentiation (see, e.g., Kandi, et al. (2015), “Effect of DNA Methylation in Various Diseases and the Probable Protective Role of Nutrition: A Mini-Review”, Cureus 7(8):e309).

[0004] Polymerases used during end repair of dsLC prior to methylation conversion during methylation sequencing can lead to lab induced artifacts (e.g., methylation bias). For example, extracted nucleic acid molecules or fragments thereof often have breaks (nicks / gaps) in the phosphodiester backbone that exist on just one strand, and / or ssDNA regions internal to the duplex molecule (gaps). Traditional methods for methylation sequencing fix nicks and gaps by incorporating unmethylated bases. The artificial loss of cytosine methylation during this library preparation process is known a “methylation bias” (or “mbias”), and has a negative impact on the ability to accurately assess methylation status based on sequence read data. Methylation bias also has a negative impact on the ability to use methylation status (e.g., a methylation signature for a sample from a subject) as a biomarker for diagnosis of disease and / or prognosis of healthcare outcomes. Thus, there remains a need for improved methods for correcting mbias, and for detecting disease based on determining methylation patterns in heterogeneous samples.1MF-364162710Attorney Docket No: 197102018940BRIEF SUMMARY

[0005] Methylation bias (mbias) is an error in methylation sequencing that that is often introduced during end repair of the DNA strands prior to methylation conversion. In particular, the polymerase used for end-repair can cause what would normally be methylated cytosines to be unmethylated. This happens in two scenarios: (1) the end-repair fill-in itself and (2) stand-displacement associated fill-in if there are nicks or gaps in the double- stranded template fragment. Described herein are methods for correcting mbias using two complementary processes comprising chemistries and computer processes configured to correct strand displacement mbias and end-repair mbias.

[0006] Provided herein are methods for generating mbias-corrected sequence read data, comprising: receiving a plurality of nucleic acid fragments extracted from a sample from a subject; generating a plurality of modified nucleic acid fragments by performing: an end repair reaction on at least one nucleic acid fragment in the plurality of nucleic acid fragments to repair fragmentation damage; and / or a nick / gap repair reaction on at least one nucleic acid fragment in the plurality of nucleic acid fragments to fill in single- stranded nicks or gaps; converting non-methylated cytosines in the plurality of modified nucleic acid fragments into uracil to generate methylation nucleic acid fragments; sequencing, by a sequencer, the methylation nucleic acid fragments to generate a plurality of sequence reads; receiving, at one or more processors, sequence read data corresponding to the plurality of sequence reads; and performing, using the one or more processors, a computational correction of mbias in the sequence read data to produce mbias-corrected sequence read data.

[0007] In some aspects, performing the computational correction of mbias in the sequence read data comprises: determining, using the one or more processors, a methylation status of one or more methylation sites in the sequence read data for one or more sequence reads; and for each sequence read for which the methylation status of one or more methylation sites located proximal to a 3 ’-end of the sequence read is determined to be unmethylated: comparing, using the one or more processors, the unmethylated status of the one or more methylation sites located proximal to the 3 ’-end of the respective sequence read to the methylation status of other sequence reads of the plurality of sequence reads that overlap the one or more methylation sites; performing a statistical analysis to determine a probability of a correlation between the methylation status of the one or more methylation sites located proximal to the 3 ’-end of the respective sequence read to that for other sequence reads of the plurality of sequence reads that overlap the one or more methylation sites; detecting 3 ’-end2MF-364162710Attorney Docket No: 197102018940 hypomethylation bias in the respective sequence read based on a comparison between the determined probability and a predetermined threshold; and generating a corrected sequence read to replace the respective sequence read based on the methylation status of the other sequence reads of the plurality of sequence reads that overlap the one or more methylation sites if 3 ’-end hypomethylation bias is detected.

[0008] The methods described herein, further comprise generating a methylation signature by: for each corrected sequence read, calculating, using the one or more processors, a probability that a methylation state of the corrected sequence read is significantly different from a distribution of methylation states determined for a plurality of sequence reads derived from samples from healthy individuals that map to a same genomic interval. The methods may further comprise diagnosing the subject based on the methylation signature, monitoring progression of a cancer based on the methylation signature, and / or monitoring recurrence of the cancer based on the methylation signature.

[0009] In some aspects, the methylation state for each corrected sequence read is determined based on a corrected methylation status of each of one or more sites within the corrected sequence read, wherein a site at which the corrected methylation status is determined is a methylation site. In some aspects, the methylation state comprises a methylation fraction value calculated based on the corrected methylation status of each of the one or more methylation sites within the corrected sequence read.

[0010] In some aspects, the methylation fraction value is calculated as: methylation fraction = N / M, wherein M is a total number of methylation sites located within the corrected sequence read and N is a number of methylations sites that are methylated.

[0011] In some aspects, the plurality of sequence reads align to one or more genomic intervals of interest, and wherein the one or more genomic intervals of interest are selected based on a cancer to be detected, a cancer to be diagnosed, or a likelihood of response to a treatment for a cancer.

[0012] In some aspects, the sample comprises a tissue biopsy sample or a liquid biopsy sample obtained from the subject.

[0013] In some aspects, the plurality of nucleic acid fragments comprises a mixture of tumor nucleic acid molecules and non-tumor nucleic acid molecules.

[0014] In some aspects, the end repair reaction comprises use of a chain termination mechanism. In some aspects, the chain termination mechanism comprises use of a chain termination nucleotide in the end repair reaction. In some aspects, the chain termination nucleotide comprises a 2',3'-dideoxyribonucleoside 5'-triphosphate (ddNTP). In some aspects,3MF-364162710Attorney Docket No: 197102018940 the ddNTP comprises a 2',3'-dideoxycytidine 5'-triphosphate (ddCTP), 2', 3'- dideoxy guano sine 5'-triphosphate (ddGTP), 2',3'-dideoxythymidine 5'-triphosphate (ddTTP), 2',3'-dideoxyadenosine 5'-triphosphate (ddATP), or any combination thereof.

[0015] In some aspects, the nick / gap repair reaction comprises use of a ligase. In some aspects, the ligase comprises a DNA ligase. In some aspects, the DNA ligase comprises Taq DNA ligase, T4 DNA ligase, 9°N™ DNA ligase, T3 DNA ligase, or any combination thereof.

[0016] In some aspects, converting non-methylated cytosines in the plurality of modified nucleic acid fragments comprises use of a bisulfite reaction. In some aspects, converting nonmethylated cytosines in the plurality of nucleic acid fragments to generate a methylation nucleic acid fragments comprises use of an enzymatic conversion reaction.

[0017] In some aspects, the one or more methylation sites are located proximal to a 3’-end of each sequence read of the plurality of sequence reads based on the sequence read data In some aspects, a site at which the methylation status is determined is a methylation site.

[0018] In some aspects, wherein performing the statistical analysis comprises: generating a contingency table that tabulates results of the comparisons of the methylation status of the one or more methylation sites located proximal to the 3 ’-end of the respective sequence read to that for other sequence reads of the plurality of sequence reads that overlap the one or more methylation sites; and performing the statistical analysis between two or more factors in the contingency table. In some aspects, the contingency table comprises a 2 x 2 contingency table that tabulates the results of the comparisons in terms of two factors (a distal or not distal location of the one or more methylation sites) and two outcomes (a fully unmethylated or not fully unmethylated status of the one or more methylation sites). In some aspects, the predetermined threshold is determined using a multi-test correction method to adjust the probabilities of the association determined by the statistical analysis to correct for an occurrence of false positives. In some aspects, 3’-end hypomethylation bias in the respective sequence read is detected if the determined probability is greater than the predetermined threshold. In some aspects, generating a corrected sequence read comprises truncating a 3’- end of the respective sequence read.

[0019] In some aspects, the corrected sequence read is a raw sequence read, an aligned sequence read, a merged sequence read, a computationally reconstructed sequence read and / or a consensus sequence read. In some aspects, the sequence read is a raw sequence read, an aligned sequence read, a merged sequence read, a computationally reconstructed sequence read, and / or a consensus sequence read.4MF-364162710Attorney Docket No: 197102018940

[0020] In some aspects, the one or more methylation sites comprise one or more CpG dinucleotide sites and / or one or more non-CpG dinucleotide sites.

[0021] In some aspects, the plurality of modified nucleic acid fragments align to one or more genomic intervals of interest. In some aspects, the one or more genomic intervals of interest are selected based on a Tumor and Tissue of Origin (TTOO) classification. In some aspects, the one or more genomic intervals of interest are selected based on a cancer to be detected. In some aspects, the one or more genomic intervals of interest comprise one or more compact genomic regions. In some aspects, the plurality of sequence reads derived from samples from healthy individuals align to one or more genomic intervals of interest that are the same as the one or more genomic intervals of interest to which the plurality of sequence reads obtained from the sample from the subject align.

[0022] In some aspects, the sequence read data comprises methyl-seq data.

[0023] In some aspects, the subject is suspected of having or is determined to have cancer. In some aspects, the cancer is a solid tumor. In some aspects, the cancer is a hematological cancer. In some aspects, the cancer is a B cell cancer (multiple myeloma), a melanoma, breast cancer, lung cancer, bronchus cancer, colorectal cancer, prostate cancer, pancreatic cancer, stomach cancer, ovarian cancer, urinary bladder cancer, brain cancer, central nervous system cancer, peripheral nervous system cancer, esophageal cancer, cervical cancer, uterine cancer, endometrial cancer, cancer of an oral cavity, cancer of a pharynx, liver cancer, kidney cancer, testicular cancer, biliary tract cancer, small bowel cancer, appendix cancer, salivary gland cancer, thyroid gland cancer, adrenal gland cancer, osteosarcoma, chondrosarcoma, a cancer of hematological tissue, an adenocarcinoma, an inflammatory myofibroblastic tumor, a gastrointestinal stromal tumor (GIST), colon cancer, multiple myeloma (MM), myelodysplastic syndrome (MDS), myeloproliferative disorder (MPD), acute lymphocytic leukemia (ALL), acute myelocytic leukemia (AML), chronic myelocytic leukemia (CML), chronic lymphocytic leukemia (CLL), polycythemia Vera, Hodgkin lymphoma, non-Hodgkin lymphoma (NHL), soft-tissue sarcoma, fibrosarcoma, myxosarcoma, liposarcoma, osteogenic sarcoma, chordoma, angiosarcoma, endotheliosarcoma, lymphangiosarcoma, lymphangioendotheliosarcoma, synovioma, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinomas, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatoma, bile duct carcinoma, choriocarcinoma, seminoma, embryonal carcinoma, Wilms' tumor, bladder carcinoma, epithelial carcinoma, glioma, astrocytoma, medulloblastoma,5MF-364162710Attorney Docket No: 197102018940 craniopharyngioma, ependymoma, pinealoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, meningioma, neuroblastoma, retinoblastoma, follicular lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, hepatocellular carcinoma, thyroid cancer, gastric cancer, head and neck cancer, small cell cancer, essential thrombocythemia, agnogenic myeloid metaplasia, hypereosinophilic syndrome, systemic mastocytosis, familiar hypereosinophilia, chronic eosinophilic leukemia, neuroendocrine cancers, or a carcinoid tumor. In some aspects, the cancer comprises acute lymphoblastic leukemia (Philadelphia chromosome positive), acute lymphoblastic leukemia (precursor B-cell), acute myeloid leukemia (FLT3+), acute myeloid leukemia (with an IDH2 mutation), anaplastic large cell lymphoma, basal cell carcinoma, B-cell chronic lymphocytic leukemia, bladder cancer, breast cancer (HER2 overexpressed / amplified), breast cancer (HER2+), breast cancer (HR+, HER2- ), cervical cancer, cholangiocarcinoma, chronic lymphocytic leukemia, chronic lymphocytic leukemia (with 17p deletion), chronic myelogenous leukemia, chronic myelogenous leukemia (Philadelphia chromosome positive), classical Hodgkin lymphoma, colorectal cancer, colorectal cancer (dMMR / MSI-H), colorectal cancer (KRAS wild type), cryopyrin-associated periodic syndrome, a cutaneous T-cell lymphoma, dermatofibrosarcoma protuberans, a diffuse large B-cell lymphoma, fallopian tube cancer, a follicular B-cell non-Hodgkin lymphoma, a follicular lymphoma, gastric cancer, gastric cancer (HER2+), gastroesophageal junction (GEJ) adenocarcinoma, a gastrointestinal stromal tumor, a gastrointestinal stromal tumor (KIT+), a giant cell tumor of the bone, a glioblastoma, granulomatosis with polyangiitis, a head and neck squamous cell carcinoma, a hepatocellular carcinoma, Hodgkin lymphoma, juvenile idiopathic arthritis, lupus erythematosus, a mantle cell lymphoma, medullary thyroid cancer, melanoma, a melanoma with a BRAF V600 mutation, a melanoma with a BRAF V600E or V600K mutation, Merkel cell carcinoma, multicentric Castleman's disease, multiple hematologic malignancies including Philadelphia chromosome-positive ALL and CML, multiple myeloma, myelofibrosis, a non-Hodgkin’ s lymphoma, a nonresectable subependymal giant cell astrocytoma associated with tuberous sclerosis, a nonsmall cell lung cancer, a non-small cell lung cancer (ALK+), a non-small cell lung cancer (PD-L1+), a non-small cell lung cancer (with ALK fusion or ROS1 gene alteration), a non- small cell lung cancer (with BRAF V600E mutation), a non-small cell lung cancer (with an EGFR exon 19 deletion or exon 21 substitution (L858R) mutations), a non-small cell lung cancer (with an EGFR T790M mutation), ovarian cancer, ovarian cancer (with a BRCA mutation), pancreatic cancer, a pancreatic, gastrointestinal, or lung origin neuroendocrine tumor, a pediatric neuroblastoma, a peripheral T-cell lymphoma, peritoneal cancer, prostate6MF-364162710Attorney Docket No: 197102018940 cancer, a renal cell carcinoma, rheumatoid arthritis, a small lymphocytic lymphoma, a soft tissue sarcoma, a solid tumor (MSI-H / dMMR), a squamous cell cancer of the head and neck, a squamous non- small cell lung cancer, thyroid cancer, a thyroid carcinoma, urothelial cancer, a urothelial carcinoma, or Waldenstrom's macroglobulinemia.

[0024] In some aspects, the methods further comprise treating the subject with an anti-cancer therapy. In some aspects, the anti-cancer therapy comprises a targeted anti-cancer therapy, chemotherapy, radiation therapy, immunotherapy, or surgery. In some aspects, the targeted anti-cancer therapy comprises abemaciclib (Verzenio), abiraterone acetate (Zytiga), acalabrutinib (Calquence), ado-trastuzumab emtansine (Kadcyla), afatinib dimaleate (Gilotrif), aldesleukin (Proleukin), alectinib (Alecensa), alemtuzumab (Campath), alitretinoin (Panretin), alpelisib (Piqray), amivantamab-vmjw (Rybrevant), anastrozole (Arimidex), apalutamide (Erleada), asciminib hydrochloride (Scemblix), atezolizumab (Tecentriq), avapritinib (Ayvakit), avelumab (Bavencio), axicabtagene ciloleucel (Yescarta), axitinib (Inlyta), belantamab mafodotin-blmf (Blenrep), belimumab (Benlysta), belinostat (Beleodaq), belzutifan (Welireg), bevacizumab (Avastin), bexarotene (Targretin), binimetinib (Mektovi), blinatumomab (Blincyto), bortezomib (Velcade), bosutinib (Bosulif), brentuximab vedotin (Adcetris), brexucabtagene autoleucel (Tecartus), brigatinib (Alunbrig), cabazitaxel (Jevtana), cabozantinib (Cabometyx), cabozantinib (Cabometyx, Cometriq), canakinumab (Haris), capmatinib hydrochloride (Tabrecta), carfilzomib (Kyprolis), cemiplimab-rwlc (Libtayo), ceritinib (LDK378 / Zykadia), cetuximab (Erbitux), cobimetinib (Cotellic), copanlisib hydrochloride (Aliqopa), crizotinib (Xalkori), dabrafenib (Tafinlar), dacomitinib (Vizimpro), daratumumab (Darzalex), daratumumab and hyaluronidase-fihj (Darzalex Faspro), darolutamide (Nubeqa), dasatinib (Sprycel), denileukin diftitox (Ontak), denosumab (Xgeva), dinutuximab (Unituxin), dostarlimab-gxly (Jemperli), durvalumab (Imfinzi), duvelisib (Copiktra), elotuzumab (Empliciti), enasidenib mesylate (Idhifa), encorafenib (Braftovi), enfortumab vedotin-ejfv (Padcev), entrectinib (Rozlytrek), enzalutamide (Xtandi), erdafitinib (Balversa), erlotinib (Tarceva), everolimus (Afinitor), exemestane (Aromasin), fam-trastuzumab deruxtecan-nxki (Enhertu), fedratinib hydrochloride (Inrebic), fulvestrant (Faslodex), gefitinib (Iressa), gemtuzumab ozogamicin (Mylotarg), gilteritinib (Xospata), glasdegib maleate (Daurismo), hyaluronidase-zzxf (Phesgo), ibrutinib (Imbruvica), ibritumomab tiuxetan (Zevalin), idecabtagene vicleucel (Abecma), idelalisib (Zydelig), imatinib mesylate (Gleevec), infigratinib phosphate (Truseltiq), inotuzumab ozogamicin (Besponsa), iobenguane 1131 (Azedra), ipilimumab (Yervoy), isatuximab-irfc (Sarclisa), ivosidenib (Tibsovo), ixazomib citrate (Ninlaro), lanreotide acetate (Somatuline Depot),7MF-364162710Attorney Docket No: 197102018940 lapatinib (Tykerb), larotrectinib sulfate (Vitrakvi), lenvatinib mesylate (Lenvima), letrozole (Femara), lisocabtagene maraleucel (Breyanzi), loncastuximab tesirine-lpyl (Zynlonta), lorlatinib (Lorbrena), lutetium Lu 177-dotatate (Lutathera), margetuximab-cmkb (Margenza), midostaurin (Rydapt), mobocertinib succinate (Exkivity), mogamulizumab-kpkc (Poteligeo), moxetumomab pasudotox-tdfk (Lumoxiti), naxitamab-gqgk (Danyelza), necitumumab (Portrazza), neratinib maleate (Nerlynx), nilotinib (Tasigna), niraparib tosylate monohydrate (Zejula), nivolumab (Opdivo), obinutuzumab (Gazyva), ofatumumab (Arzerra), olaparib (Lynparza), olaratumab (Lartruvo), osimertinib (Tagrisso), palbociclib (Ibrance), panitumumab (Vectibix), panobinostat (Farydak), pazopanib (Votrient), pembrolizumab (Keytruda), pemigatinib (Pemazyre), pertuzumab (Perjeta), pexidartinib hydrochloride (Turalio), polatuzumab vedotin-piiq (Polivy), ponatinib hydrochloride (Iclusig), pralatrexate (Folotyn), pralsetinib (Gavreto), radium 223 dichloride (Xofigo), ramucirumab (Cyramza), regorafenib (Stivarga), ribociclib (Kisqali), ripretinib (Qinlock), rituximab (Rituxan), rituximab and hyaluronidase human (Rituxan Hycela), romidepsin (Istodax), rucaparib camsylate (Rubraca), ruxolitinib phosphate (Jakafi), sacituzumab govitecan-hziy (Trodelvy), seliciclib, selinexor (Xpovio), selpercatinib (Retevmo), selumetinib sulfate (Koselugo), siltuximab (Sylvant), sipuleucel-T (Provenge), sirolimus protein-bound particles (Fyarro), sonidegib (Odomzo), sorafenib (Nexavar), sotorasib (Lumakras), sunitinib (Sutent), tafasitamab-cxix (Monjuvi), tagraxofusp-erzs (Elzonris), talazoparib tosylate (Talzenna), tamoxifen (Nolvadex), tazemetostat hydrobromide (Tazverik), tebentafusp-tebn (Kimmtrak), temsirolimus (Torisel), tepotinib hydrochloride (Tepmetko), tisagenlecleucel (Kymriah), tisotumab vedotin-tftv (Tivdak), tocilizumab (Actemra), tofacitinib (Xeljanz), tositumomab (Bexxar), trametinib (Mekinist), trastuzumab (Herceptin), tretinoin (Vesanoid), tivozanib hydrochloride (Fotivda), toremifene (Fareston), tucatinib (Tukysa), umbralisib tosylate (Ukoniq), vandetanib (Caprelsa), vemurafenib (Zelboraf), venetoclax (Venclexta), vismodegib (Erivedge), vorinostat (Zolinza), zanubrutinib (Brukinsa), ziv-aflibercept (Zaltrap), or any combination thereof.

[0025] In some aspects, the methods further comprise obtaining the sample from the subject. In some aspects, the sample comprises a tissue biopsy sample, a liquid biopsy sample, or a normal control. In some aspects, the sample is a liquid biopsy sample and comprises blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva. In some aspects, the sample is a liquid biopsy sample and comprises circulating tumor cells (CTCs). In some aspects, the sample is a liquid biopsy sample and comprises cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), or any combination thereof. In some aspects, the plurality of nucleic acid8MF-364162710Attorney Docket No: 197102018940 fragments comprises a mixture of tumor nucleic acid fragments and non-tumor nucleic acid fragments. In some aspects, the tumor nucleic acid fragments are derived from a tumor portion of a heterogeneous tissue biopsy sample, and the non-tumor nucleic acid fragments are derived from a normal portion of the heterogeneous tissue biopsy sample. In some aspects, the sample comprises a liquid biopsy sample, and wherein the tumor nucleic acid fragments are derived from a circulating tumor DNA (ctDNA) fraction of the liquid biopsy sample, and the non-tumor nucleic acid fragments are derived from a non-tumor, cell-free DNA (cfDNA) fraction of the liquid biopsy sample.

[0026] In some aspects, the methods further comprise preparing a sequencing library. In some aspects, the sequencing comprises a single-end sequencing method. In some aspects, the sequencing comprises a paired-end sequencing method. In some aspects, the sequencing comprises use of a massively parallel sequencing (MPS) technique, whole genome sequencing (WGS), whole exome sequencing, targeted sequencing, direct sequencing, or Sanger sequencing technique. In some aspects, the sequencing comprises massively parallel sequencing, and the massively parallel sequencing technique comprises next generation sequencing (NGS). In some aspects, the sequencer comprises a next generation sequencer.

[0027] In some aspects, one or more of the plurality of sequence reads overlap one or more genomic loci within one or more subgenomic intervals in the sample. In some aspects, the one or more genomic loci comprises between 10 and 20 loci, between 10 and 40 loci, between 10 and 60 loci, between 10 and 80 loci, between 10 and 100 loci, between 10 and 150 loci, between 10 and 200 loci, between 10 and 250 loci, between 10 and 300 loci, between 10 and 350 loci, between 10 and 400 loci, between 10 and 450 loci, between 10 and 500 loci, between 20 and 40 loci, between 20 and 60 loci, between 20 and 80 loci, between 20 and 100 loci, between 20 and 150 loci, between 20 and 200 loci, between 20 and 250 loci, between 20 and 300 loci, between 20 and 350 loci, between 20 and 400 loci, between 20 and 500 loci, between 40 and 60 loci, between 40 and 80 loci, between 40 and 100 loci, between 40 and 150 loci, between 40 and 200 loci, between 40 and 250 loci, between 40 and 300 loci, between 40 and 350 loci, between 40 and 400 loci, between 40 and 500 loci, between 60 and 80 loci, between 60 and 100 loci, between 60 and 150 loci, between 60 and 200 loci, between 60 and 250 loci, between 60 and 300 loci, between 60 and 350 loci, between 60 and 400 loci, between 60 and 500 loci, between 80 and 100 loci, between 80 and 150 loci, between 80 and 200 loci, between 80 and 250 loci, between 80 and 300 loci, between 80 and 350 loci, between 80 and 400 loci, between 80 and 500 loci, between 100 and 150 loci, between 100 and 200 loci, between 100 and 250 loci, between 100 and 300 loci, between 100 and 350 loci,9MF-364162710Attorney Docket No: 197102018940 between 100 and 400 loci, between 100 and 500 loci, between 150 and 200 loci, between 150 and 250 loci, between 150 and 300 loci, between 150 and 350 loci, between 150 and 400 loci, between 150 and 500 loci, between 200 and 250 loci, between 200 and 300 loci, between 200 and 350 loci, between 200 and 400 loci, between 200 and 500 loci, between 250 and 300 loci, between 250 and 350 loci, between 250 and 400 loci, between 250 and 500 loci, between 300 and 350 loci, between 300 and 400 loci, between 300 and 500 loci, between 350 and 400 loci, between 350 and 500 loci, or between 400 and 500 loci.

[0028] In some aspects, the one or more genomic loci comprise one or more gene loci. In some aspects, the one or more gene loci comprise ABL1, ACVR1B, AKT1, AKT2, AKT3, ALK, ALOX12B, AMER1, APC, AR, ARAF, ARFRP1, ARID1A, ASXL1, ATM, ATR, ATRX, AURKA, AURKB, AXIN1, AXL, BAP1, BARD1, BCL2, BCL2L1, BCL2L2, BCL6, BCOR, BCORL1, BCR, BRAF, BRCA1, BRCA2, BRD4, BRIP1, BTG1, BTG2, BTK, CALR, CARD11, CASP8, CBFB, CBL, CCND1, CCND2, CCND3, CCNE1, CD22, CD274, CD70, CD74, CD79A, CD79B, CDC73, CDH1, CDK12, CDK4, CDK6, CDK8, CDKN1A, CDKN1B, CDKN2A, CDKN2B, CDKN2C, CEBPA, CHEK1, CHEK2, CIC, CREBBP, CRKL, CSF1R, CSF3R, CTCF, CTNNA1, CTNNB1, CUL3, CUL4A, CXCR4, CYP17A1, DAXX, DDR1, DDR2, DIS3, DNMT3A, DOT1L, EED, EGFR, EMSY (Cl lorf30), EP300, EPHA3, EPHB1, EPHB4, ERBB2, ERBB3, ERBB4, ERCC4, ERG, ERRFI1, ESRI, ETV4, ETV5, ETV6, EWSR1, EZH2, EZR, FAM46C, FANCA, FANCC, FANCG, FANCL, FAS, FBXW7, FGF10, FGF12, FGF14, FGF19, FGF23, FGF3, FGF4, FGF6, FGFR1, FGFR2, FGFR3, FGFR4, FH, FLCN, FLT1, FLT3, FOXL2, FUBP1, GABRA6, GATA3, GATA4, GATA6, GID4 (C17orf39), GNA11, GNA13, GNAQ, GNAS, GRM3, GSK3B, H3F3A, HDAC1, HGF, HNF1A, HRAS, HSD3B1, ID3, IDH1, IDH2, IGF1R, IKBKE, IKZF1, INPP4B, IRF2, IRF4, IRS2, JAK1, JAK2, JAK3, JUN, KDM5A, KDM5C, KDM6A, KDR, KEAP1, KEL, KIT, KLHL6, KMT2A (MLL), KMT2D (MLL2), KRAS, LTK, LYN, MAF, MAP2K1, MAP2K2, MAP2K4, MAP3K1, MAP3K13, MAPK1, MCL1, MDM2, MDM4, MED12, MEF2B, MEN1, MERTK, MET, MITF, MKNK1, MLH1, MPL, MRE11A, MSH2, MSH3, MSH6, MST1R, MTAP, MTOR, MUTYH, MYB, MYC, MYCL, MYCN, MYD88, NBN, NF1, NF2, NFE2L2, NFKBIA, NKX2-1, NOTCH1, NOTCH2, NOTCH3, NPM1, NRAS, NT5C2, NTRK1, NTRK2, NTRK3, NUTM1, P2RY8, PALB2, PARK2, PARP1, PARP2, PARP3, PAX5, PBRM1, PDCD1, PDCD1LG2, PDGFRA, PDGFRB, PDK1, PIK3C2B, PIK3C2G, PIK3CA, PIK3CB, PIK3R1, PIM1, PMS2, POLDI, POLE, PPARG, PPP2R1A, PPP2R2A, PRDM1, PRKAR1A, PRKCI, PTCHI, PTEN, PTPN11, PTPRO, QKI, RAC1, RAD21, RAD51, RAD51B, RAD51C,10MF-364162710Attorney Docket No: 197102018940RAD51D, RAD52, RAD54L, RAFI, RARA, RBI, RBM10, REL, RET, RICTOR, RNF43, ROS1, RPTOR, RSPO2, SDC4, SDHA, SDHB, SDHC, SDHD, SETD2, SF3B1, SGK1, SLC34A2, SMAD2, SMAD4, SMARCA4, SMARCB1, SMO, SNCAIP, SOCS1, SOX2, SOX9, SPEN, SPOP, SRC, STAG2, STAT3, STK11, SUFU, SYK, TBX3, TEK, TERC, TERT, TET2, TGFBR2, TIPARP, TMPRSS2, TNFAIP3, TNFRSF14, TP53, TSC1, TSC2, TYRO3, U2AF1, VEGFA, VHL, WHSCI, WHSC1L1, WT1, XPO1, XRCC2, ZNF217, ZNF703, or any combination thereof. In some aspects, the one or more gene loci comprise ABL, ALK, ALL, B4GALNT1, BAFF, BCL2, BRAF, BRCA, BTK, CD19, CD20, CD3, CD30, CD319, CD38, CD52, CDK4, CDK6, CML, CRACC, CS1, CTLA-4, dMMR, EGFR, ERBB1, ERBB2, FGFR1-3, FLT3, GD2, HDAC, HER1, HER2, HR, IDH2, IL-ip, IL-6, IL- 6R, JAK1, JAK2, JAK3, KIT, KRAS, MEK, MET, MSLH, mTOR, PARP, PD-1, PDGFR, PDGFRa, PDGFRP, PD-L1, PI3K6, PIGF, PTCH, RAF, RANKL, RET, ROS1, SLAMF7, VEGF, VEGFA, VEGFB, or any combination thereof.

[0029] Also provided herein are systems comprising: an automated DNA sequencing library preparation module; a sequencer; one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: receive sequence read data corresponding to the plurality of sequence reads; and perform a computational correction of mbias in the sequence read data to produce mbias-corrected sequence read data.

[0030] In some aspects, the computational correction of mbias in the sequence read data comprises: determining, using the one or more processors, a methylation status of one or more methylation sites in the sequence read data for one or more sequence reads; and for each sequence read for which the methylation status of one or more methylation sites located proximal to a 3 ’-end of the sequence read is determined to be unmethylated: comparing, using the one or more processors, the unmethylated status of the one or more methylation sites located proximal to the 3 ’-end of the respective sequence read to the methylation status of other sequence reads of the plurality of sequence reads that overlap the one or more methylation sites; performing a statistical analysis to determine a probability of a correlation between the methylation status of the one or more methylation sites located proximal to the 3 ’-end of the respective sequence read to that for other sequence reads of the plurality of sequence reads that overlap the one or more methylation sites; detecting 3 ’-end hypomethylation bias in the respective sequence read based on a comparison between the determined probability and a predetermined threshold; and generating a corrected sequence read to replace the respective sequence read based on the methylation status of the other11MF-364162710Attorney Docket No: 197102018940 sequence reads of the plurality of sequence reads that overlap the one or more methylation sites if 3 ’-end hypomethylation bias is detected.

[0031] In some aspects, the automated DNA sequencing library preparation module is configured to perform one or more steps selected from: i) generating a plurality of modified nucleic acid fragments by performing: an end repair reaction on a plurality of nucleic acid fragments extracted from a sample from a subject to repair fragmentation damage, and / or a nick / gap repair reaction on a plurality of nucleic acid fragments extracted from a sample from a subject to fill in single- stranded nicks or gaps; and / or ii) converting non-methylated cytosines in the plurality of modified nucleic acid fragments into uracil to generate methylation nucleic acid fragments.

[0032] In some aspects, the end repair reaction comprises use of a chain termination mechanism. In some aspects, the chain termination mechanism comprises use of a chain termination nucleotide in the end repair reaction. In some aspects, the chain termination nucleotide comprises a 2',3'-dideoxyribonucleoside 5'-triphosphate (ddNTP).

[0033] In some aspects, the nick / gap repair reaction comprises use of a ligase. In some aspects, the one or more methylation sites are located proximal to a 3 ’-end of each sequence read of the plurality of sequence reads based on the sequence read data.

[0034] In some aspects, a site at which the methylation status is determined is a methylation site.

[0035] In some aspects, performing the statistical analysis comprises: generating a contingency table that tabulates results of the comparisons of the methylation status of the one or more methylation sites located proximal to the 3 ’-end of the respective sequence read to that for other sequence reads of the plurality of sequence reads that overlap the one or more methylation sites; and performing the statistical analysis between two or more factors in the contingency table. In some aspects, the contingency table comprises a 2 x 2 contingency table that tabulates the results of the comparisons in terms of two factors (a distal or not distal location of the one or more methylation sites) and two outcomes (a fully unmethylated or not fully unmethylated status of the one or more methylation sites). In some aspects, the predetermined threshold is determined using a multi-test correction method to adjust the probabilities of the association determined by the statistical test to correct for an occurrence of false positives. In some aspects, 3’-end hypomethylation bias in the respective sequence read is detected if the determined probability is greater than the predetermined threshold. In some aspects, generating a corrected sequence read comprises truncating a 3 ’-end of the respective sequence read.12MF-364162710Attorney Docket No: 197102018940

[0036] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference in their entirety to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference in its entirety. In the event of a conflict between a term herein and a term in an incorporated reference, the term herein controls.BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Various aspects of the disclosed methods, devices, and systems are set forth with particularity in the appended claims. A better understanding of the features and advantages of the disclosed methods, devices, and systems will be obtained by reference to the following detailed description of illustrative embodiments and the accompanying drawings, of which:

[0038] FIGs. 1A-1C provides a non-limiting schematic of a method for generating mbias illustrations as used herein. FIG. 1A shows exemplary reads with both methylated and unmethylated cytosines mapped to an exemplary genomic interval. FIG. IB shows the exemplary reads from FIG. 1A aligned from 3’ to 5’ under an exemplary coverage plot, wherein coverage is calculated as the total number of CG dinucleotide observations (methylated + unmethylated) at the corresponding read position on the x-axis. FIG. 1C shows an exemplary methylation fraction (mF) plot that can be generated based on the reads aligned under FIG. IB, wherein mF is calculated as the number of methylated Cs divided by the number of total Cs in CpG context at the corresponding read position on the x-axis. Methylated cytosine sites are shown in black and unmethylated cytosine sites are shown in gray.

[0039] FIGs. 2A-2B provides a non-limiting schematic of an mbias illustration as used herein. FIG. 2A provides exemplary mF values for an exemplary genomic region and is annotated to show how both end repair mbias and strand displacement mbias can be viewed on an mbias illustration plot. FIG. 2B provides the corresponding coverage for the exemplary reads used to generate the mF plot in FIG. 2A.

[0040] FIG. 3 provides a non-limiting schematic of a method of generating mbias corrected sequence read data according to embodiment described herein.

[0041] FIG. 4 provides a non-limiting schematic of a method of performing the computational correction according to an embodiment described herein.

[0042] FIG. 5 provides a non-limiting schematic illustration of a method for correcting for 3 ’-end hypomethylation bias in methylation sequencing data in accordance with one embodiment of the present disclosure.13MF-364162710Attorney Docket No: 197102018940

[0043] FIGs. 6A-6F provide examples of mbias data calculated in in two different ways according to coverage in the hypermethylated regions for the samples after combining the reads from the tumor, healthy, and noncancerous samples with the healthy high coverage reads, samples are colored according to the library preparation methods used to generate the methylation analysis data. FIG. 6A shows the standard deviation of methylation position bias (SD_methyl_position) plotted against the mean coverage in the sample for the hypermethylated regions (mean_target_coverage) for samples collected from individuals with tumors processed without a mbias chemistry correction and with either the TaqLigase or ddCTP and TaqLigase chemistry corrections. FIG. 6B shows SD_methyl_position plotted against the mean_target_coverage for samples collected from healthy individuals processed without a mbias chemistry correction and with either the TaqLigase or ddCTP and TaqLigase chemistry corrections. FIG. 6C shows SD_methyl_position plotted against the mean_target_coverage for samples collected from individuals with noncancerous conditions processed without a mbias chemistry correction and with either the TaqLigase or ddCTP and TaqLigase chemistry corrections. FIG. 6D shows slip rate for plotted against mean_target_coverage for samples collected from individuals with tumors processed without an mbias chemistry correction and with either the TaqLigase or ddCTP and TaqLigase chemistry corrections. FIG. 6E shows slip rate for plotted against mean_target_coverage for samples collected from healthy individuals processed without an mbias chemistry correction and with either the TaqLigase or ddCTP and TaqLigase chemistry corrections. FIG. 6F shows slip rate for plotted against mean_target_coverage for samples collected from individuals with noncancerous conditions processed without an mbias chemistry correction and with either the TaqLigase or ddCTP and TaqLigase chemistry corrections.

[0044] FIG. 7 shows an example of mbias data (presented in mbias illustrations) before (left side) and after (right side) computational corrections for a sample collected from a healthy individual processed with no chemistry correction combined with a high coverage sample.

[0045] FIGs. 8A-8B show examples of mbias data (presented in mbias illustrations) before computational corrections (left side) and after computational correction (rights side) for samples collected from individuals with a noncancerous condition, varying in mean unique coverage sequenced with an input of lOng. FIG. 8A shows examples of mbias data before and after computational corrections for a low coverage sample collected from an individual with a noncancerous condition processed with no chemistry correction. FIG. 8B shows examples of mbias data before and after computational corrections for a sample collected14MF-364162710Attorney Docket No: 197102018940 from an individual with a noncancerous condition processed with no chemistry correction combined with a high coverage sample.

[0046] FIGs. 9A-9B show examples of mbias data (presented in mbias illustrations) before computational corrections (left side) and after computational correction (rights side) for samples collected from individuals with a noncancerous condition, varying in the chemistry corrections applied in library preparation and sequenced with an input of 20ng. FIG. 9A shows examples of mbias data before and after computational corrections for a sample collected from an individual with a noncancerous condition processed with no chemistry correction combined with a high coverage sample. FIG. 9B shows examples of mbias data before and after computational corrections for a sample collected from an individual with a noncancerous condition processed with the ddCTP and TaqLigase chemistry corrections combined with a high coverage sample.

[0047] FIGs. 10A-10B show examples of mbias data (presented in mbias illustrations) before computational correction (left side) and after computational correction (right side) for samples with high or low methyl bias processed with no chemistry corrections. FIG. 10A shows examples of mbias data before and after the computational correction for a sample collected from a healthy individual, processed with no chemistry correction and categorized as having high methyl bias. FIG. 10B shows examples of mbias data before and after the computational correction for a sample collected from an individual with a noncancerous condition, processed with no chemistry correction and categorized as having low methyl bias.

[0048] FIG. 11 depicts an exemplary computing device or system in accordance with one embodiment of the present disclosure.

[0049] FIG. 12 depicts an exemplary computer system or computer network, in accordance with some instances of the systems described herein.DETAILED DESCRIPTION

[0050] In mammalian genomes, methylation is primarily observed in genomic regions high in cytosine and guanine content comprising CpG islands. Methylation in these regions typically occurs in both strands of the double stranded DNA (dsDNA) structure, i.e., if a cytosine on one strand of the duplex is methylated within a CpG locus, typically a nearby cytosine within a CpG locus on the opposite strand have a higher likelihood of also being methylated. The DNA processing during library preparation during methylation analysis through sequencing results in CpG loci erroneously being called as unmethylated. This misidentification of15MF-364162710Attorney Docket No: 197102018940 methylation status, often called mbias, results in loss of methylation status information and increased noise in sequence read-based methylation analysis, especially when attempting to analyze regions of the genome that, e.g., become hypomethylated during development or in disease states such as cancer. Reducing or eliminating mbias can improve signal-to-noise ratios and allow for more sensitive detection of methylation status in, for example, cancerous liquid samples exhibiting lower levels of tumor fraction (i.e., amount of circulating tumor DNA in the sample relative to the amount of cell free DNA from normal or healthy cells). Sensitive detection of regions of the genome that become hypomethylated in cancer is important for cancer detection and tissue of origin determination.

[0051] Provided herein are methods integrating novel laboratory methods and computational processing, which may be used to address different types of mbias. The laboratory methods, which make use of a ligase to repair nicks, can be more effective at reducing the stranddisplacement mode, while the computational correction can be more effective at reducing the end-repair mode. A combined approach which incorporates both strategies may provide a more complete correction of mbias errors. The methods provided herein thus correct for mbias using complementary approaches, meaning they may correct for mbias introduced into the process through multiple mechanisms. These complementary approaches, the chemistry methods and computational methods, are particularly useful for regions known to be hypermethylated. These hypermethylated regions are often important for cancer diagnostics and prognostics but high quality methylation measurements have conventionally been hard to achieve in these regions due to the noise created by incorrect measures of methylation state.

[0052] The methods described herein is based at least partially on a finding that combining chemistry methods to correct for strand displacement mbias and computational approaches for correcting for end-repair mbias improve accuracy and efficiency in generating mbias corrected sequencing data. The methods disclosed herein are based at least partially on recognizing mbias levels along a saequence read rather than relative to a specific genomic region. Before recognizing mbias distribution on a read level rather than by genomic location, the benefits of combining the chemistry methods and computational methods would not have been realized. Previous methods correct for mbias using only chemistry methods during library prep, or only computational based methods during data analysis, and even then have room for improvement when it comes to those individual methods.

[0053] The methods described herein are based at least partially on insight gained by analyzing methylation frequency at a sequence read position rather than at a genomic location. Such analysis may be performed by reviewing information about methylation that16MF-364162710Attorney Docket No: 197102018940 can be annotated as a methylated or unmethylated cytosine at a CpG site along a sequence read. (FIG. 1A). The sequence reads can be aligned as a function of a position along the length of a sequence reads from 0 (3’ end) to the length of the read (5’) end and the read coverage calculated as a function of the position along the length of the sequence reads can be plotted (FIG. IB). Methylation fraction (mF) as a fraction of position along the length of the sequence reads, wherein mF can be calculated as the number of methylated Cs divided by the number of total Cs in CpG context at the corresponding sequence read position (FIG.1C).

[0054] Strand displacement mbias and end-repair mbias bias can be visualized by graphing the methylation fraction (mF) as a function of position along the length of sequence read to illustrate mbias (FIG. 2A). Corrected mbias sequencing data can be visualized as well using mbias illustrations, corrected data is illustrated by a reduced left-hand tail (3’ end of the sequencing fragments) and a higher flatter line for the mF values across the length of the sequencing fragments. The sequencing coverage can also be visualized along the length of a sequence read in a genomic interval and the relationship of mF and coverage may be related to the effectiveness of mbias corrections as described herein (FIG. 2B).

[0055] The methods provided herein utilize a subset of chemistry methods that have been implemented in a particular order. The chemistry-methods as described may be effective at correcting for mbias that is less likely able to be solved computationally. Further, by limiting the number of chemistry modifications in favor of applying computational corrections as well, the steps of the claimed methods decrease reagent costs and the possibilities of error or sample loss that may be inherent in complex library preparation protocols.

[0056] In addition, in some exemplary embodiments, by at the same time correcting for mbias in library preparation, the effectiveness of the computational correction may increase. Specifically, correcting for mbias in library preparation decreases the number of reads that will need to be discarded after sequencing or that have remaining mbias that may not be able to be corrected for computationally. When more of the sequence reads are of higher quality, the coverage at hypermethylated genomic regions increases. As described herein, the computational correction may be more effective with higher coverage. Thus, the ordered combination of steps as described herein may increase efficiency by requiring less sequencing for the coverage necessary for accuracy in the computational methods.

[0057] Also provided herein method comprising generating a methylation signature using the mbias corrected sequence fragment data. The methylation signatures comprise obtaining methylation states and comparing the methylation states obtained from the sample to those of 17MF-364162710Attorney Docket No: 197102018940 healthy individuals at the same genomic location. Methylation signatures can be used to diagnose a subject, select an anticancer therapy, monitor progression of a cancer, or monitor recurrence of a cancer.Definitions

[0058] Unless otherwise defined, technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art in the field to which this disclosure belongs.

[0059] As used in this specification and the appended claims, the singular forms “a”, “an” and “the” include plural referents unless the content clearly dictates otherwise. Thus, for example, reference to “a molecule” optionally includes a combination of two or more such molecules, and the like.

[0060] The term “or” is used herein to mean, and is used interchangeably with, the term “and / or”, unless context clearly indicates otherwise.

[0061] The terms “about” or “approximately” as used herein refer to the usual error range for the respective value readily known to the skilled person in this technical field, for example, an acceptable degree of error or deviation for the quantity measured given the nature or precision of the measurements. Reference to “about” or “approximately” a value or parameter herein includes (and describes) embodiments that are directed to that value or parameter per se.

[0062] The term “isolated” in the context of a nucleic acid molecule or a polypeptide refers to a nucleic acid molecule or polypeptide being separated from other nucleic acid molecules or polypeptides that are present in the natural source of the nucleic acid molecule or polypeptide. In some certain embodiments, the isolated nucleic acid molecule or polypeptide is free of or substantially free of other cellular material or culture medium when produced by recombinant techniques, or free of or substantially free of chemical precursors or other chemicals when chemically synthesized.

[0063] An “individual” or “subject” is a mammal. Mammals include, but are not limited to, domesticated animals (e.g., cows, sheep, cats, dogs, and horses), primates (e.g., humans and non-human primates such as monkeys), rabbits, and rodents (e.g., mice and rats). In certain embodiments, the individual or subject is a human.

[0064] As used herein, the term “sequence read” is a computationally generated sequence generated by a sequencer to represent a sequence of bases of a single strand of a sequenced fragment. A sequence read can refer to a raw sequence read (e.g., a sequence read as obtained18MF-364162710Attorney Docket No: 197102018940 directly from a sequencing instrument), an aligned sequence read (e.g., a sequence read that has been aligned to a reference genome), a single-end sequence read, a paired-end sequence read, a merged sequence read (e.g., a sequence read based on merging a group of overlapping paired-end reads), a consensus sequence read (e.g., a sequence read based on performing error correction on a merged sequence read), a computationally reconstructed sequence read (e.g., a sequence read that has been computationally truncated at the 5’ and / or 3’ end), or any combination thereof.

[0065] The terms “cancer” and “tumor” are used interchangeably herein. These terms refer to the presence of cells possessing characteristics typical of cancer-causing cells, such as uncontrolled proliferation, immortality, metastatic potential, rapid growth and proliferation rate, and certain characteristic morphological features. Cancer cells are often in the form of a tumor, but such cells can exist alone within an animal, or can be a non-tumorigenic cancer cell, such as a leukemia cell. These terms include a solid tumor, a soft tissue tumor, or a metastatic lesion. As used herein, the term “cancer” includes premalignant, as well as malignant cancers.

[0066] As used herein, “treatment” (and grammatical variations thereof such as “treat” or “treating”) refers to clinical intervention (e.g., administration of an anti-cancer agent or anticancer therapy) in an attempt to alter the natural course of the individual being treated, and can be performed either for prophylaxis or during the course of clinical pathology. Desirable effects of treatment include, but are not limited to, preventing occurrence or recurrence of disease, alleviation of symptoms, diminishment of any direct or indirect pathological consequences of the disease, preventing metastasis, decreasing the rate of disease progression, amelioration or palliation of the disease state, and remission or improved prognosis.

[0067] As used herein, the term “sequenced fragment” refers to a fragment of a larger nucleic acid molecule that has been sequenced. Different nucleic acid sequencing methods may yield one or more sequence reads per sequenced fragment, thus data for a sequenced fragment may be derived from an analysis of one or more sequence reads, e.g., sequence reads obtained from a sample from a subject.

[0068] “Likely to” or “increased likelihood,” as used herein, refer to an increased probability that an event, item, object, thing, or person will occur. Thus, in one example, an individual that is likely to respond to treatment with an anti-cancer therapy, e.g., an anti-cancer therapy provided herein, alone or in combination, has an increased probability of responding to treatment with the anti-cancer therapy alone or in combination, relative to a reference19MF-364162710Attorney Docket No: 197102018940 individual or group of individuals. “Unlikely to” refers to a decreased probability that an event, item, object, thing, or person will occur relative to a reference individual or group of individuals. Thus, an individual that is unlikely to respond to treatment with an anti-cancer therapy, e.g., an anti-cancer therapy provided herein, alone or in combination, has a decreased probability of responding to treatment with the anti-cancer therapy, alone or in combination, relative to a reference individual or group of individuals.

[0069] When a range of values is provided, it is to be understood that each intervening value between the upper and lower limit of that range, and any other stated or intervening value in that states range, is encompassed within the scope of the present disclosure. Where the stated range includes upper or lower limits, ranges excluding either of those included limits are also included in the present disclosure.

[0070] Some of the analytical methods described herein include mapping sequences to a reference sequence, determining sequence information, and / or analyzing sequence information. It is well understood in the art that complementary sequences can be readily determined and / or analyzed, and that the description provided herein encompasses analytical methods performed in reference to a complementary sequence.

[0071] The section headings used herein are for organization purposes only and are not to be construed as limiting the subject matter described. The description is presented to enable one of ordinary skill in the art to make and use the invention and is provided in the context of a patent application and its requirements. Various modifications to the described embodiments will be readily apparent to those persons skilled in the art and the generic principles herein may be applied to other embodiments. Thus, the present invention is not intended to be limited to the embodiment shown but is to be accorded the widest scope consistent with the principles and features described herein.

[0072] The figures illustrate processes according to various embodiments. In the exemplary processes, some blocks are, optionally, combined, the order of some blocks is, optionally, changed, and some blocks are, optionally, omitted. In some examples, additional steps may be performed in combination with the exemplary processes. Accordingly, the operations as illustrated (and described in greater detail below) are exemplary by nature and, as such, should not be viewed as limiting.

[0073] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference in their entirety to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be20MF-364162710Attorney Docket No: 197102018940 incorporated by reference in its entirety. In the event of a conflict between a term herein and a term in an incorporated reference, the term herein controls.Methods for generating mbias corrected sequencing fragment data

[0074] Provided herein are methods that can be used to generate mbias corrected sequencing data. The methods comprise applying chemistry methods described herein to isolated nucleic acid fragments before methylation analysis to correct for the potential causes of strand displacement mbias. The chemistry methods comprise using either an end repair reaction to repair damage done to the nucleic acids during fragmentation and / or nick / gap repair reactions to fill in single-strand nicks or gaps in the nucleic acids. A methylation analysis may comprise applying a restriction enzyme-based, affinity enrichment-based, bisulfite conversion-based, or enzymatic conversion based methylation sequencing method. The methods further comprise correcting for end-bias mbias in the sequencing data generated from the methylation analysis using computational methods as described herein. The computational methods comprise analyzing each sequence read and comparing the methylation status at one or more methylation sites proximal to the 3’ end to the respective methylation status in other sequencing fragments that overlap the same location, as described herein. If end-repair mbias is detected using the computational methods described herein, a corrected fragment is generated to correct for the mbias.

[0075] Mbias corrected sequencing fragment data can be used to generate a more accurate methylation signature as described herein. A methylation signature may be generated by calculating a probability that a methylation state of a corrected sequencing fragment is significantly different from a distribution of methylation states determined for a plurality of sequenced fragments derived from samples from healthy individuals that map to a same genomic interval. Methylation signatures can be generated for corrected sequencing data generated from nucleic acids collected from a sample collected from an individual with cancer as described herein. Methylation signatures can be used, as described herein, monitoring progression of the cancer based on the methylation signature, and / or monitoring recurrence of the cancer based on the methylation signature. Methylation signatures generated with mbias corrected sequencing fragment data will be more representative of the true methylation status of the nucleic acids in the sample and thus will be more clinically useful monitoring cancer progression, or monitoring progression of cancer as described herein.21MF-364162710Attorney Docket No: 197102018940

[0076] FIG. 3 provides a non-limiting schematic of a method of generating mbias corrected sequence read data of the invention described herein. In some instances, the mbias corrected sequencing data can be used for generating a methylation signature according to the methods described herein. In some instances, the methylation signature can be used in methods for detecting a cancer in a subject. In some instances, the methylation signature can be used to diagnose a cancer in a subject. In some instances, the methylation signature can be used to choose a treatment based on the likelihood of a cancer to respond to the treatment. In some instances, the methylation signature can be used to determine whether a cancer is responding to treatment.

[0077] Block 300 encompasses receiving a plurality of nucleic acid fragments extracted from a sample from a subject. In some instances, the plurality nucleic acid fragments are extracted from a sample obtained from a subject using the methods described herein. In some instances, the sample comprises a tissue biopsy sample, a liquid biopsy sample, or a normal control sample as described herein. In some instances, one or more nucleic acid fragments are cell- free DNA (cfDNA), circulating tumor DNA (ctDNA), or fragmented non tumor DNA as described herein. In some instances, the one or more nucleic acid fragments are derived from tumor DNA, and / or non-tumor DNA derived from a heterogeneous samples as described herein.

[0078] The sample is collected from a subject. In some instances, the subject may have or be suspected of having a condition or disease as described herein. In some instances, the condition or disease is a cancer as described herein.

[0079] Block 302 encompasses a method for correcting for strand-displacement bias in nucleic acid fragments. In some instances, the method can be used for correcting one or more strand-displacement instances in one or more nucleic acid fragments from a sample obtained from a subject by generating a plurality of modified nucleic acid fragments from the one or more nucleic acid fragments. Strand-displacement mbias may be a result of nick / gap repair and or end repair in the one or more nucleic acid fragments. In some instances, nick / gaps and damage to the end of nucleic acid fragments are the natural results of extracting nucleic acids and generating nucleic acid fragments. The traditional methods to repair nicks / gaps and damage to the end of nucleic acid fragments result in the replacement of methylated cytosines in CpG sites with non-methylated cytosine. In some instances, the one or more nucleic acid fragments are extracted from the sample according to the extraction and processing steps described herein. In some instances, the one or more nucleic acid fragments have been subject to the methods described herein to isolate or enrich for the one or more nucleic acid22MF-364162710Attorney Docket No: 197102018940 fragments that when sequenced will map to one or more target genomic region or gene loci of interest.

[0080] With reference to FIG. 3, block 302 comprises performing one or both of the chemistry methods described in block 304 and block 306 respectively. The steps at block 302 prevent mbias before the unmethylated cytosines are converted into uracil. In other words, repairing damage to DNA fragments caused by the fractionation of the DNA strands prevents methylated regions from erroneously being detected as unmethylated after conversion and sequencing. In some instances, block 302 comprises, in some embodiments, first performing an end repair reaction to repair fragmentation damage to the one or more nucleic acid fragments. In some instances, the end-repair reaction comprises use of a chain termination mechanism, as described herein. In some instances, the block 304 comprises use of a chain termination nucleotide in the end repair reaction. In some instances, the chain termination nucleotide comprises a 3'-dideoxyribonucleoside 5'-triphosphate (ddNTP), such as but not limited to the ddNTPs described herein. In some instances, the block 302 comprises performing a nick / gap repair reaction to fill in single-stranded nicks or gaps in the one or more nucleic acid fragments (block 306). In some instances, block 306 comprises performing a nick / gap repair reaction to fill in single- stranded nicks or gaps in the one or more nucleic acid fragments. In some instances, the ligase comprises a DNA ligase as described herein. In some instances, the DNA ligase comprises a Taq DNA ligase, T4 DNA ligase, 9oNTM DNA ligase, T3 DNA ligase, or any combination thereof. In some instances, block 302, in some embodiments, comprises performing an end repair reaction to repair fragmentation damage to the one or more nucleic acid fragments as described in block 304 and performing a nick / gap repair reaction to fill in single- stranded nicks or gaps in the one or more nucleic acid fragments as described in block 306 . In some embodiments, the steps at block 304 may be performed simultaneously with the steps at block 306. In some embodiments, the steps at block 304 may be performed before the steps at block 306 .In some embodiments, the steps at block 306 may be performed before the steps at block 304. In some embodiments, performing nick / gap repair before end repair may comprise use of an end repair lacking polymerase lacking strand displacement activity.

[0081] The plurality of modified nucleic acid fragments generated according to block 302 are used at block 308. At block 308 non-methylated cytosines in the plurality of modified nucleic acid fragments are converted into uracil to generate methylation nucleic acid fragments, according to methods described herein.23MF-364162710Attorney Docket No: 197102018940

[0082] At block 310, the methylation nucleic acid fragments are sequenced, by a sequencer, to generate a plurality of sequence reads. In some instances, the sequencing comprises use of a massively parallel sequencing (MPS) technique, whole genome sequencing (WGS), whole exome sequencing, targeted sequencing, or direct sequencing as described herein. In some instances, the sequencing comprises massively parallel sequencing, and the massively parallel sequencing technique comprises next generation sequencing (NGS) as described herein.

[0083] At block 312, an exemplary system (e.g., one or more electronic devices) receives, at one or more processors, sequence read data corresponding to the plurality of sequence reads (block 310). In some instances, the plurality of sequence reads have been sequenced as described in herein. In some instances, the sequence read data are mapped to a reference genome using alignment methods described herein. In some instances, the alignment data has been optimized for processing sequence fragment data generated using a methylation analysis as described herein. In some instances, block 306 comprises detecting the coverages of the sequence fragment data corresponding the plurality of sequence reads as described herein.

[0084] At block 314, the system produces mbias-corrected sequence data by performing a computational correction of mbias as described herein. In some instances, the computational correction method is the computational correction method of FIG 4.

[0085] FIG. 4 provides a non-limiting schematic illustration of a method for computationally correcting mbais in sequence read data of the invention described herein. The computational correction method corrects one or more end-repair mbias instances in the sequence read data to produce mbias corrected sequence read data comprising a plurality of mbias corrected sequence reads, using the computational methods described herein. In some instances, correcting mbias at one or more methylation site proximal to the 3’ end of a sequence read comprises correcting one or more instance of end-repair mbias. Correcting one or more endrepair mbias instances comprises determining a methylation status for one or more methylation sites (block 400) and further performing the methods described in block 402 - block 410 if the one or more methylation sites located proximal to the 3’ end of the sequence read are unmethylated. In some instances, correcting one or more end-repair mbias instances in the sequence read data comprises performing computational methods. In some instances, the computational methods are performed for sequence read data mapping to one or more hypermethylated regions as described herein. In some instances, the method comprises correcting for end-repair mbias in one or more hypermethylated regions.

[0086] In some instances, the computational methods used for correcting end-repair mbias when the sequencing coverage in the region is greater than about 20x. In some instances, the24MF-364162710Attorney Docket No: 197102018940 computational methods used for correcting end-repair mbias when the sequencing coverage in the region is greater than about lOOx. In some instances, the computational methods used for correcting end-repair mbias when the sequencing coverage in the region is greater than about 20x, 25x, 30x, 35x, 40x, 45x, 50x, 55x, 60x, 65x, 70x, 75x, 80x, 85x, 90x, 95x or lOOx. In some instances, the sequence coverage is a mean unique coverage.

[0087] With reference to FIG. 4, at block 400, an exemplary system (e.g., comprising one or more electronic devices) determines, using the one or more processors, a methylation status of one or more methylation sites. In some instances, block 400 comprises calling the methylation status of one or more methylation sites within the sequence read data. In some instances, the method for calling methylation status may be any of those known in the art or as described herein.

[0088] With reference to FIG. 4, at block 402, the exemplary system (e.g., one or more electronic devices) performs an analysis for each sequence read for which an unmethylated status is determined for one or more sites located proximal to the 3’ end. The analysis comprises the methods described in blocks 404, 406, 408, and 410. In some instances, each sequence read may be considered a sequence read under test as described herein. In some instances, the sequence read may be a raw sequence read, an aligned sequence read, a merged sequence read, a computationally reconstructed sequence read and / or a consensus sequence read.

[0089] At block 402, an exemplary system (e.g., one or more electronic devices) makes a comparison, using the one or more processors, of the methylation status of the one or more methylation sites located proximal to a 3 ’-end of the respective sequence read to that of other sequence read of the plurality of sequence read that overlap the one or more methylation sites (block 404). In some instances, making the comparison comprises any of the methods described herein.

[0090] The method of block 402 comprises, performing a statistical analysis to determine a probability of an association in the results of the comparison between the methylation status of the one or more methylation sites located proximal to the 3 ’-end of the respective sequence read to that for the other sequence reads of the plurality of sequence reads that overlap the one or more methylation sites (block 406). The statistical analysis comprises determining a probability of an association in the results of the comparison performed in the methods according to block 404.

[0091] In some instances, performing the statistical analysis at block 406 comprises, generating a contingency table that tabulates results of the comparisons of the methylation25MF-364162710Attorney Docket No: 197102018940 status of the one or more methylation sites located proximal to the 3 ’-end of the respective sequence reads to that for the other sequence reads of the plurality of sequence reads that overlap the one or more methylation sites; and performing the statistical analysis between two or more factors in the contingency table. In some instances, the contingency table is constructed as described herein. In some instances, the contingency table comprises a 2 x 2 contingency table that tabulates the results of the comparisons in terms of two factors (a distal or not distal location of the one or more methylation sites) and two outcomes (a fully unmethylated or not fully unmethylated status of the one or more methylation sites).

[0092] In some instances, the method of block 402 comprises detecting 3’-end hypomethylation bias in the respective sequence read based on a comparison between the determined probability to a second predetermined threshold (block 408). The determined probability has been generated according to the method at block 406. In some instances, the second predetermined threshold is determined using the methods as described herein. In some instances, the second predetermined probability threshold is determined using a multi-test correction method to adjust the probabilities of the association determined by the statistical test to correct for an occurrence of false positives. In some instances, 3’-end hypomethylation bias in the respective sequence read is detected if the determined probability is greater than the second predetermined threshold.

[0093] The methods of block 402 comprises generating a corrected sequence read for the respective sequence read, if 3 ’-end hypomethylation bias is detected (block 410). 3’ hypomethylation bias may have been detected using the methods described at block 408. In some instances, generating a corrected sequence read comprises applying methods as described herein. In some instances, generating a corrected sequence read comprises truncating a 3’ end of the sequence read as described herein.

[0094] The methods described herein according to FIGs. 3 and 4 improve upon the conventional methods for generating mbias corrected sequence fragment data. In some instances, the improvements are related to an ordered combination of chemistry and computational approaches to generate mbias corrected sequence read data relative to the steps of the method overall. As described above, the chemistry approaches prevent mbias before converting methylated cytosines into uracil and the computational approaches correct for mbias that was introduced during the conversion step. In some instances, the improvements are to the efficiency in generating mbias corrected sequence data. In some instances, the improvements are to the accuracy of the methylation status of the methylation sites in the26MF-364162710Attorney Docket No: 197102018940 mbias corrected sequence data, the methylation signature derived therefrom or any of the uses of the methylation signature described herein.

[0095] The methods described herein and according to FIGs. 3-4 takes advantage of a novel finding that combining chemistry methods to correct for strand displacement mbias and computational approaches to correct for end-repair mbias improve accuracy and efficiency in generating mbias corrected sequencing data. The methods described herein improve upon previous methods that rely on only chemistry methods or only computational methods. The methods apply chemistry methods during library prep that may be effective at correcting for mbias that is less likely be solved computationally. The methods may also allow mbias that can effectively be fixed computationally to persist until computational corrections can be applied. In addition, the methods rely on limiting compositional modifications to the molecules (chemistry based correction) in favor of computational correction for some instances of mbias. In some instances, by limiting the number of modifications applied to the nucleic acid fragments before sequencing in favor of applying computational corrections as well, the steps of the claimed methods decrease reagent costs and the possibilities of error or sample loss that may be inherent in complex library preparation protocols. Correcting for mbias using chemistry methods in library preparation may increase the effectiveness of the computation corrections. Correcting for mbias in library preparation decreases the number of sequence reads that will need to be discarded after sequencing or that have remaining mbias that may not be able to be corrected for computationally. The methods are more effective and efficient than previous methods because more sequence fragments of higher quality can be used as input to the computational methods. When more of the sequenced fragments are of higher quality, the coverage at hypermethylated genomic regions increases. As described herein, the computational correction may be more effective with higher coverage. The methods as described herein thus increase efficiency by requiring less sequencing for the coverage necessary for accuracy in the computational methods.Generating a plurality of modified nucleic acid fragments using chemistry methods for correcting for strand-displacement mbias

[0096] Provided herein are methods that can be used to generate mbias corrected sequence read data. The methods comprise correcting for strand-displacement mbias in nucleic acid fragments from a sample obtained from a subject by generating a plurality of modified nucleic acid fragments from the one or more nucleic acid fragments, wherein generating the27MF-364162710Attorney Docket No: 197102018940 plurality of modified nucleic acid fragments comprises applying one or more chemistry methods. Correcting for strand-displacement mbias may comprise correcting for bias that is likely an artifact of library preparation as described herein and in US provisional application 63 / 468,742, hereby incorporated by reference in its entirety.

[0097] It was observed that strand-displacement mbias likely comes from using conventional methods for library preparations. For example, if there is a nick or a gap on one strand of a dsDNA molecule that contains CpG island regions, the strand displacement and subsequent resynthesis of the DNA strand that may occur when performing conventional end repair and nick / gap repair reactions will erroneously incorporate unmethylated cytosines into the strand. For cytosines that were originally methylated, that methylation information is lost during the end repair and nick / gap repair process.

[0098] The chemistry methods described herein comprise using alternative end repair and nick / gap repair methods to minimize strand displacement mbias when preparing nucleic acids collected from a sample for sequencing. The minimization in strand displacement is achieved by introducing a ligase to repair single stranded nicks in DNA fragments, the end repair polymerase is less likely to have strand displacement activity wherein methylated cytosines are replaced by unmethylated cytosines. The disclosed methods have the advantage of lowering or eliminating strand displacement methylation bias in library preparation when combined with the computational methods described herein to reduce end-repair mbias, minimize mbias, the disclosed methods result in better signal-to-noise in sequence read data for hypomethylated cancer biomarkers and enable better sensitivity for detection of methylation status at low tumor fraction where the amount of available circulating tumor DNA (ctDNA) in the sample is typically quite low (e.g., less than 1%).

[0099] The methods for correcting for one or more instance of strand-displacement mbias may comprise performing an end repair reaction to repair fragmentation damage to the one or more nucleic acid fragments, wherein the end repair reaction comprises use of a chain termination mechanism.

[0100] During library preparation, the end repair reaction is performed to convert fragmented nucleic acid molecules into blunt-end nucleic acid molecules containing 5'-phosphate and 3'- hydroxyl groups. The 5'^-3' polymerase activity of DNA polymerase(s) used in the reaction mixture fills in 5' overhangs, while the 3'^-5' exonuclease activity of the DNA polymerase(s) removes 3' overhangs. In some instances, the fragmented nucleic acid molecules may be, e.g., naturally- occurring fragmented DNA, such as circulating tumor DNA (ctDNA) which is typically less than 300 bps in length. In some instances, the fragmented nucleic acid28MF-364162710Attorney Docket No: 197102018940 molecules may be, e.g., DNA that has been extracted from a tissue sample and fragmented using mechanical methods e.g., sonication) and / or enzymatic methods (e.g., using a fragmentase and / or restriction enzymes).

[0101] In some instances, the end repair reaction may comprise adding 2',3'-dideoxycytidine 5'-triphosphate (ddCTP) (100% or dilute) to the end repair reaction mix. Adding 2', 3'- dideoxycytidine 5'-triphosphate (ddCTP) (100% or dilute) to the end repair reaction mix instead of, or in addition to, 2 ’-deoxycytidine 5'-triphosphate (dCTP) to eliminate all 5' overhang repaired molecules, or to limit the amount of repair allowed in each molecule. Alternatively, 2',3'-dideoxythymidine 5'-triphosphate (ddTTP), 2',3'-dideoxyguanosine 5'- triphosphate (ddGTP), 2',3'-dideoxyadenosine 5'-triphosphate (ddATP) (or any combination of ddCTP, ddTTP, ddGTP, and ddATP) can be used. In some instance, the concentration of ddNTP (or other chain terminating nucleotides) used may be varied, as different polymerases have different tolerances to non-natural nucleotides. In some instances, other chain termination moieties or chain termination mechanisms may be used as well. For example, chain termination during the end repair (blunting) reaction may also be achieved by omitting a certain dNTP (e.g., dATP, dTTP, dCTP, or dGTP) from the reaction mixture (or including only a limiting amount of a certain dNTP). In some instances, a limiting amount of a specified dNTP may correspond to a concentration that is about 1 / 200*, 1 / 100th, 1 / 90*, 1 / 80*, 1 / 70*, 1 / 60*, 1 / 50*, 1 / 40*, 1 / 30*, 1 / 20*, or 1 / 10* the concentration of the other deoxynucleotide triphosphates in the reaction mixture. In some instances, a ratio of ddNTP concentration to a corresponding dNTP concentration used to perform the end repair reaction may be less than 20x, 30x, 40x, 50x, or 60x, 70x, 80x, lOOx, 120x, or 140x. In some instances, noncanonical or modified nucleotides may be used in the end repair (blunting) process to flag regions where end repair occurs. Examples of modified nucleotides that may be used include, but are not limited to, methylated cytidine (e.g., 5-methyl-dCTP and 5- hydroxy-methyl-dCTP), deoxy uracil, and / or oxoguanine.

[0102] In some instances, the end repair reaction may comprise performing end repair using a polymerase enzyme that lacks strand displacement activity. Polymerases that lack strand displacement activity lack the ability to displace downstream duplex DNA strands encountered during synthesis. There are a variety of DNA polymerases available that have varying degrees of strand displacement activity. Examples of DNA polymerases that lack strand displacement activity include, but are not limited to, T4 DNA polymerase, Klentaq, etc.29MF-364162710Attorney Docket No: 197102018940

[0103] In some instances, the disclosed methods may comprise performing both the first end repair reaction and the second end repair reaction. In some instances, the disclosed methods may comprise performing both the tailing reaction comprising the use of a single dNTP and the nick / gap repair reaction comprising the use of a DNA ligase.

[0104] The methods for correcting for one or more instance of strand-displacement mbias may comprise performing a nick / gap repair reaction to fill in single- stranded nicks or gaps in the one or more nucleic acid fragments, wherein the nick / gap repair reaction comprises use of a ligase. In some instances, performing nick / gap repair may comprise using a DNA ligase, e.g., Taq ligase. Alternative examples of ligases that may be used for performing nick / gap repair include, but are not limited to, T4 ligase, 9°NO DNA ligase, T3 DNA ligase, E. coli ligase, etc., or any combination thereof. As described above, the use of the ligase repairs single stranded nicks in DNA fragments and as a result, the end repair polymerase is less likely to have strand displacement activity wherein methylated cytosines are replaced by unmethylated cytosines.

[0105] In some instances, the disclosed methods may comprise generating a plurality of modified nucleic acid fragments by performing both an end repair reaction and a nick / gap repair reaction. In some instances, the disclosed methods may comprise generating a plurality of modified nucleic acid fragments by performing an end repair reaction or a nick / gap repair reaction. In some instances, the disclosed methods may comprise generating a plurality of modified nucleic acid fragments by performing one or more of an end repair reaction and a nick / gap repair reaction.Converting non-methylated cytosines into uracil

[0106] In some instances, as described in more detail below, the disclosed methods converting non-methylated cytosines in the plurality of modified nucleic acid fragments into uracil to generate methylation nucleic acid fragments. In some instances, converting non- methylated cytosines comprises a restriction enzyme-based, affinity enrichment-based, bisulfite conversion-based, or enzymatic conversion-based method. In some instances, a cytosine conversion reaction is used to convert non-methylated cytosine to uracil. In some instances, the cytosine conversion reaction comprises a chemical conversion reaction, e.g., a bisulfite conversion reaction. In some instances, the cytosine conversion reaction comprises an enzymatic conversion reaction, e.g., comprising the use of a tet methylcytosine dioxygenase 2 (TET2) enzyme to oxidize 5-methyl-cytosine (5mC) or 5-hydroxymethyl-30MF-364162710Attorney Docket No: 197102018940 cytosine (5hmC) to 5-carboxycytosine (5caC). In some instances, the enzymatic conversion reaction may further comprise the use of a combination of TET2 and T4 P- glucosyltransferase (T4-PGT) enzymes to convert 5-methyl-cytosine (5mC) or 5- hydroxymethyl-cytosine (5hmC) to 5-(D-glucosyl)oxymethyl-cytosine. In some instances, the enzymatic conversion reaction may comprise the use of an Apolipoprotein B mRNA Editing Catalytic Polypeptide-like (APOBEC) enzyme.

[0107] In some instances, the methods may comprise preparing a sequencing library according to the methods described herein. In some instances, the disclosed methods may further comprise performing a ligation step to ligate one or more adapters to the plurality of modified nucleic acid fragments or methylation nucleic acid fragments. In some instances, the one or more adapters comprise one or more adapters comprising an overhanging poly-T sequence. In some instances, the one or more adapters comprise one or more sequencing adapters, for example, one or more methylated stubby adapters, flow cell adapters, read 1 sequencing adapters, read 2 sequencing adapters, or any combination thereof. In some instances, the one or more adapters comprise one or more barcodes. In some instances, the disclosed methods may further comprise performing a ligation step to add a one or more barcodes to the plurality of modified nucleic acid fragments. In some instances, the one or more barcodes may comprise, for example, a library index, a sample barcode, a cell barcode, a target-specific barcode, a nucleic acid fragment barcode, a unique molecular index, or any combination thereof.

[0108] In some instances, the disclosed methods may further comprise performing a nucleic acid amplification reaction to generate amplified nucleic acid molecules that are derived from the plurality of modified nucleic acid fragments or the methylation nucleic acid fragments. In some instances, for example, the nucleic acid amplification reaction may comprise a polymerase chain reaction (PCR). In some instances, the nucleic acid amplification reaction may comprise a rolling circle amplification (RCA) reaction.

[0109] In some instances, the disclosed methods may further comprise capturing a subset of nucleic acid molecules from the amplified nucleic acid molecules and performing nucleic acid sequencing and / or analysis on the subset.Sequencing the methylation nucleic acid fragments to generate a plurality of sequence reads

[0110] The methods provided herein comprise sequencing the methylation nucleic acid fragments to generate a plurality of sequence reads.31MF-364162710Attorney Docket No: 197102018940

[0111] . In some instances, the methods may comprise sequencing lOng of the plurality of methylation nucleic acid fragments to generate a plurality of sequence reads. In some instances, the methods may comprise sequencing lOng of the plurality of methylation nucleic acid fragments to generate a plurality of sequence reads. In some instances, the methods may comprise sequencing 5ng,10ng, 20ng, 30ng, 40ng, 50ng, 60ng, 70ng, 80ng, 90ng, or more than lOOng of the plurality of methylation nucleic acid fragments to generate a plurality of sequence reads. In some instances, the methods may comprise sequencing between 5ng and lOOng, 5ng and 90ng, 5ng and 80g, 5ng and 70ng, 5ng and 60ng, 5ng and 50ng, 5ng and 40ng, 5ng and 30ng, 5ng and 20ng, or 5ng and lOng of the plurality of methylation nucleic acid fragments to generate a plurality of sequence reads. In some instances, the methods may comprise sequencing between 5ng and lOOng, lOng and lOOng, 20ng and lOOng, 30ng and lOOng, 40ng and lOOng, 50ng and lOOng, 60ng and lOOng, 70ng and lOOng, 80ng and lOOng, or 90ng and lOOng of the plurality of methylation nucleic acid fragments to generate a plurality of sequence reads.

[0112] In some instances, the disclosed methods may further comprise performing a sequence read analysis based on the plurality of sequence reads to determine a methylation signature of the subject.

[0113] In some instances, sequence read data obtained by sequencing nucleic acid molecules derived from the plurality of modified nucleic acid fragments may exhibit reduced methylation bias compared to that obtained by sequencing a conventionally-prepared DNA sequencing library. For example, in some instances, the reduction in methylation bias may be greater than 5%, 10%, 15%, 20%, 25%, 30%, or 35% as measured by the SD methyl position bias metric described elsewhere herein.

[0114] The methods provided herein for generating mbias corrected sequence read data comprise receiving at one or more processors, sequence read data corresponding to a plurality of sequence reads. In some instances, the sequence read data corresponding to a plurality of sequence read are generated through nucleic acid sequencing using a next generation sequencer. In some instances, the sequencing may comprise using the next- generation sequencing to perform whole genome sequencing.

[0115] “Next-generation sequencing” (or “NGS”) as used herein may also be referred to as “massively parallel sequencing” (or “MPS”), and refers to any sequencing method that determines the nucleotide sequence of either individual nucleic acid molecules (e.g., as in single molecule sequencing) or clonally expanded proxies for individual nucleic acid32MF-364162710Attorney Docket No: 197102018940 molecules in a high throughput fashion (e.g., wherein greater than 103, 104, 105 or more than 105 molecules are sequenced simultaneously).

[0116] Next-generation sequencing methods are known in the art, and are described in, e.g., Metzker, M. (2010) Nature Biotechnology Reviews 11:31-46, which is incorporated herein by reference. Other examples of sequencing methods suitable for use when implementing the methods and systems disclosed herein are described in, e.g., International Patent Application Publication No. WO 2012 / 092426. In some instances, the sequencing may comprise, for example, whole genome sequencing (WGS), whole exome sequencing, targeted sequencing, or direct sequencing. In some instances, the sequencing may comprise a paired-end sequencing technique that allows both ends of a fragment to be sequenced and generates high-quality, alignable sequence data for detection of, e.g., genomic rearrangements, repetitive sequence elements, gene fusions, and novel transcripts.

[0117] The disclosed methods and systems may be implemented using sequencing platforms such as the Roche 454, Illumina Solexa, ABI-SOLiD, ION Torrent, Complete Genomics, Pacific Bioscience, Helicos, and / or the Polonator platform. In some instances, sequencing may comprise Illumina MiSeq sequencing. In some instances, sequencing may comprise Illumina HiSeq sequencing. In some instances, sequencing may comprise Illumina NovaSeq sequencing. Optimized methods for sequencing a large number of target genomic loci in nucleic acids extracted from a sample are described in more detail in, e.g., International Patent Application Publication No. WO 2020 / 236941, the entire content of which is incorporated herein by reference.

[0118] The methods provided herein comprise receiving, at one or more processors, sequence read data corresponding to the plurality of sequence reads. In some instances, the sequenced fragments data may be aligned sequencing data.

[0119] In some instances, the methods may include the use of an alignment method optimized for aligning sequence reads for nucleic acid fragments that have been converted using, e.g., a bisulfite reaction, to convert unmethylated cytosine residues to uracil (which is interpreted as a thymine in sequencing results). In some instances, sequence reads may be aligned to two genomes in silico, e.g., converted, and unconverted versions of the reference genome, using such alignment tools. Methylation occurs primarily at CpG sites, but may also occur less frequently at non-CpG sites (e.g., CHG or CHH sites).

[0120] In some instances, the sequence read data may be obtained using a nucleic acid sequencing method comprising the use of a bisulfite- or enzymatic-conversion reaction (e.g., during library preparation) to convert non-methylated cytosine to uracil (see, e.g., Li, et al.33MF-364162710Attorney Docket No: 197102018940(2011), “DNA Methylation Detection: Bisulfite Genomic Sequencing Analysis”, Methods Mol. Biol. 791:11-21) or any of the sequencing methods as described herein.

[0121] Examples of alignment tools optimized for aligning sequence reads for converted DNA include, but are not limited to, NovoAlign (Novocraft Technologies, Selangor, Malaysia), and the Bismark tool (Krueger, et al. (2011), “Bismark: A Flexible Aligner and Methylation Caller for Bisulfite-Seq Applications”, Bioinformatics 27(11): 1571-1572).

[0122] In some instances, the methods described herein may comprise the use of a methylation status calling method, e.g., to call the methylation status of the CpG sites within a sequence read based on the sequence reads (complementary pairs of forward and reverse sequence reads) derived from nucleic acid fragments that have been subjected to a chemical or enzymatic conversion reaction, e.g., to convert unmethylated cytosine residues to uracil (which is interpreted as a thymine in sequencing results). Examples of such methylation status calling tools include, but are not limited to, the Bismark tool (Krueger, et al. (2011), “Bismark: A Flexible Aligner and Methylation Caller for Bisulfite-Seq Applications”, Bioinformatics 27(11): 1571-1572), TARGOMICS (Garinet, et al. (2017), “Calling Chromosome Alterations, DNA Methylation Statuses, and Mutations in Tumors by Simple Targeted Next-Generation Sequencing - A Solution for Transferring Integrated Pangenomic Studies into Routine Practice?”, J. Molecular Diagnostics 19(5):776-787), Bicycle (Grana, et al. (2018) “Bicycle: A Bioinformatics Pipeline to Analyze Bisulfite Sequencing Data”, Bioinformatics 34(8): 1414-5), SMAP (Gao, et al. (2015), “SMAP: A Streamlined Methylation Analysis Pipeline for Bisulfite Sequencing”, Gigascience 4:29), and MeDUSA (Wilson, et al. (2016), “Computational Analysis and Integration of MeDIP-Seq Methylome Data”, in: Kulski JK, editor, Next Generation Sequencing: Advances, Applications and Challenges. Rijeka: InTech, p. 153-69). See also, Rauluseviciute, et al. (2019), “DNA Methylation Data by Sequencing: Experimental Approaches and Recommendations for Tools and Pipelines for Data Analysis”, Clinical Epigenetics 11:193.

[0123] In some instances, the method may comprise detecting the coverage of the sequence read data corresponding to the plurality of sequence reads. The coverage may be related to number of unique sequence reads that align to a genomic loci. The coverage may be related to the number of unique sequence reads that align to a methylation site. The coverage may be related to the average number of sequence reads that align to one or more methylation sites. The coverage may be related to the average number of sequence reads that align to one or more methylation sites in hypermethylated regions of the genome. In some instances, the coverage may be a mean coverage for one or more hypermethylated regions.34MF-364162710Attorney Docket No: 197102018940

[0124] In some instances, the plurality of sequence reads has a coverage of about lOx, about 20x, about 30x, about 40x, about 50x, about 60x, about 70x, about 80x, about 90x, about lOOx, about HOx, about 120x, about 150x, about 160x, about 170x, about 180x, about 190x, or about 200x. In some instances, the plurality of sequence reads has a coverage of between about lOx and about 120x, about lOx and about 11 Ox, about lOx and about lOOx, about lOx and about 90x, about lOx and about 80x, about lOx and about 70x. In some instances, the plurality of sequence reads has a coverage of between about 80x and about 120x, about 70x and about 120x, about 60x and about 120x, about 50x and about 120x, or about 50x and about 120x. In some instances, the plurality of sequence reads has a coverage of between about 70x and about 120x. In some instances, increasing the input quantity of the plurality of modified nucleic acid fragments in the sequencing increases the coverage of the plurality of sequence reads.Computationally correcting for mbias to generate mbias corrected sequence read data

[0125] The methods provided herein for generating mbias corrected sequence read data comprise performing computational methods to generate mbias corrected sequence read data. In some instances, generating mbias corrected sequencing data comprises correcting for one or more instances of end-repair end bias using computational methods. In some instances, the computational methods are described herein and in US provisional application 63 / 446,238 and PCT / US2024 / 015952, incorporated by reference in their entirety. When combined with the chemistry methods to correct for strand displacement mbias, the resulting mbias correct sequence read data that are free or substantially free from mbias that is likely an artifact of library preparation.

[0126] As described herein, the methods comprise correcting for mbias in the sequence read data to produce mbias corrected sequence read data comprising a plurality of mbias corrected sequence reads. The computational correction may comprise determining, using the one or more processors, a methylation status of one or more methylation sites; and for each sequence read for which an unmethylated status is determined for the one or more methylation sites located proximal to a 3'-end: making a comparison, using the one or more processors, of the methylation status of the one or more methylation sites located proximal to a 3'-end of the respective sequence read to that of other sequence reads of the plurality of sequence reads that overlap the one or more methylation sites; performing a statistical analysis to determine a probability of an association in the results of the comparison between35MF-364162710Attorney Docket No: 197102018940 the methylation status of the one or more methylation sites located proximal to the 3'-end of the respective sequence read to that for the other sequence reads of the plurality of sequence reads that overlap the one or more methylation sites; detecting 3'-end hypomethylation bias in the respective sequence read based a comparison between the determined probability to a second predetermined threshold; and generating a corrected sequence read for the respective sequence read, if 3'-end hypomethylation bias is detected.

[0127] In some instances, a methylation site refers to a genomic site or genomic locus, e.g., a CpG dinucleotide site, which may have a methylation status of either methylated or unmethylated (z.e., not methylated). In some instances, the one or more methylation sites comprise one or more CpG dinucleotide sites and / or one or more non-CpG dinucleotide sites.

[0128] In some instances, a methylation status of one or more methylation sites is determined by aligning the sequence read data and calling a methylation status as described herein. A methylation status may be a characterization of a methylation site as methylated or unmethylated. In some instances, the one or more methylation sites may comprise, e.g., one or more CpG dinucleotide sites, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more than 100 CpG sites.

[0129] For each sequence read for which an unmethylated status is determined for the one or more methylation sites located proximal to a 3 ’-end: the methylation status of one or more methylation sites located proximal to a 3 ’-end of the respective sequence read may be compared to that of other sequence reads of the plurality of sequence reads that overlap the one or more methylation sites.

[0130] A statistical analysis may be performed to determine a probability of an association in the results of the comparison between the methylation status of the one or more methylation sites located proximal to the 3 ’-end of the respective sequence read to that for the other sequence reads of the plurality of sequence reads that overlap the one or more methylation sites. Performing the statistical analysis may comprise generating a contingency table that tabulates results of the comparisons of the methylation status of the one or more methylation sites located proximal to the 3 ’-end of the respective sequence reads to that for the other sequence reads of the plurality of sequence reads that overlap the one or more methylation sites; and performing the statistical analysis between two or more factors in the contingency table.

[0131] In some instances, the contingency table may comprise a 2 x 2 contingency table that tabulates the results of the comparisons in terms of two factors (e.g., a distal or not distal location of the one or more methylation sites) and two outcomes (e.g., a fully unmethylated 36MF-364162710Attorney Docket No: 197102018940 or not fully unmethylated status of the one or more methylation sites). In some instances, the contingency table may comprise a 3 x 3, 4 x 4, 5 x 5, or 6 x 6 contingency table that takes into account additional factors.

[0132] In some instances, for example, the statistical analysis may comprise a Fisher’s Exact Test. Alternatively, any statistical analysis used to perform association tests may be used. Examples include, but are not limited to, Barnard's exact test or Boschloo's exact test. Another alternative is to use maximum likelihood estimates to calculate a p-value from the exact binomial or multinomial distributions, and then reject or fail to reject based on the p- value.

[0133] The methods comprise detecting 3 ’-end hypomethylation bias in the respective sequence read based a comparison between the determined probability to a second predetermined threshold. In some instances, 3’-end hypomethylation bias in the respective sequence reads is detected if the determined probability is greater than the second predetermined threshold.

[0134] In some instances, the second predetermined threshold is determined using a multitest correction method to adjust the probabilities of the association determined by the statistical analysis to correct for an occurrence of false positives. For example, the multi-test correction method may comprise a Benjamini-Hochberg multi-test correction method. In some instances, the second predetermined threshold may have a percentage value ranging from 50% to 100% (e.g., 50%, 60%, 70%, 80%, 90%, 95%, 98%, or 99%, or any value within this range), or a fractional value ranging from 0.5 to 1.0 (e.g., 0.5, 0.6, 0.7, 0.8, 0.9, 0.95, 0.98, or 0.99, or any value within this range).

[0135] The methods further comprise generating a corrected sequence read for the respective sequence read, if 3 ’-end hypomethylation bias is detected. In some instances, generating a corrected sequence read for the respective sequence read comprises truncating a 3 ’-end of the sequence read under test if 3 ’-end hypomethylation bias is detected. For example, the sequence read under test may be truncated by trimming off 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, 20, or more than 20 nucleotides (or any number of nucleotides within this range) from the 3 ’-end of the sequence read under test (e.g., the nucleotides are digitally removed from the sequence read from the sequence read data for the sequence read under test).37MF-364162710Attorney Docket No: 197102018940Methods of use

[0136] In some instances, the methods further comprise generating a methylation signature. In some instances, the methylation signature is generated by, for each corrected sequence read of the plurality of mbias corrected sequence read data, calculating a probability that a methylation state of a corrected sequence read is significantly different from a distribution of methylation states determined for a plurality of sequence reads derived from samples from healthy individuals that map to a same genomic interval.

[0137] In some instances, the methylation state for each corrected sequence read is determined based on a corrected methylation status of each of one or more sites within the corrected sequence read, wherein a site at which the corrected methylation status is determined is a methylation site.

[0138] In some instances, the methylation state comprises a methylation fraction value calculated based on the corrected methylation status of each of the one or more methylation sites within the corrected sequence read.

[0139] In some instances, the methylation fraction value is calculated as: methylation fraction = N / M, wherein M is a total number of methylation sites located within the corrected sequence read and N is a number of methylations sites that are methylated.

[0140] In some instances, the methylation signature may be used for diagnosing a subject, monitoring progression of a cancer, or monitoring recurrence of cancer as described herein.

[0141] In some instances, a methylation signature generated according to the methods described herein may be used to diagnose a subject with cancer. In some instances, the disclosed methods may be applicable to diagnosis of any of a variety of cancers as described elsewhere herein. In some instances, the methods described herein may be used to diagnose (or as part of a diagnosis of) the presence of disease or other condition (e.g., cancer, genetic disorders (birth defects), neurological disorders (imprinting disorders, cardiovascular disease or any other disease type where detection of methylation in the subject is relevant. In some instances, the disclosed methods may be applicable to diagnosis of any of a variety of cancers as described elsewhere herein. In some instances, the disclosed methods may be applicable to detection of a variety of diseases or conditions and / or determination of risks associated with a variety of disease or conditions (e.g. cancer, genomic imprinting diseases, autoimmune, neurological, aging, etc.) See, e.g., Jin, et al. (2018), “DNA Methylation in Human Diseases”, Genes & Diseases 5:1-8.38MF-364162710Attorney Docket No: 197102018940

[0142] In some instances, the methods may further comprise administration of an anti-cancer therapy. In some instances, the anti-cancer therapy may be a targeted therapy. In some instances, the targeted therapy (or anti-cancer target therapy) may comprise abemaciclib (Verzenio), abiraterone acetate (Zytiga), acalabrutinib (Calquence), ado-trastuzumab emtansine (Kadcyla), afatinib dimaleate (Gilotrif), aldesleukin (Proleukin), alectinib (Alecensa), alemtuzumab (Campath), alitretinoin (Panretin), alpelisib (Piqray), amivantamab- vmjw (Rybrevant), anastrozole (Arimidex), apalutamide (Erleada), asciminib hydrochloride (Scemblix), atezolizumab (Tecentriq), avapritinib (Ayvakit), avelumab (Bavencio), axicabtagene ciloleucel (Yescarta), axitinib (Inlyta), belantamab mafodotin-blmf (Blenrep), belimumab (Benlysta), belinostat (Beleodaq), belzutifan (Welireg), bevacizumab (Avastin), bexarotene (Targretin), binimetinib (Mektovi), blinatumomab (Blincyto), bortezomib (Velcade), bosutinib (Bosulif), brentuximab vedotin (Adcetris), brexucabtagene autoleucel (Tecartus), brigatinib (Alunbrig), cabazitaxel (Jevtana), cabozantinib (Cabometyx), cabozantinib (Cabometyx, Cometriq), canakinumab (Haris), capmatinib hydrochloride (Tabrecta), carfilzomib (Kyprolis), cemiplimab-rwlc (Libtayo), ceritinib (LDK378 / Zykadia), cetuximab (Erbitux), cobimetinib (Cotellic), copanlisib hydrochloride (Aliqopa), crizotinib (Xalkori), dabrafenib (Tafinlar), dacomitinib (Vizimpro), daratumumab (Darzalex), daratumumab and hyaluronidase-fihj (Darzalex Faspro), darolutamide (Nubeqa), dasatinib (Sprycel), denileukin diftitox (Ontak), denosumab (Xgeva), dinutuximab (Unituxin), dostarlimab-gxly (Jemperli), durvalumab (Imfinzi), duvelisib (Copiktra), elotuzumab (Empliciti), enasidenib mesylate (Idhifa), encorafenib (Braftovi), enfortumab vedotin-ejfv (Padcev), entrectinib (Rozlytrek), enzalutamide (Xtandi), erdafitinib (Balversa), erlotinib (Tarceva), everolimus (Afinitor), exemestane (Aromasin), fam-trastuzumab deruxtecan-nxki (Enhertu), fedratinib hydrochloride (Inrebic), fulvestrant (Faslodex), gefitinib (Iressa), gemtuzumab ozogamicin (Mylotarg), gilteritinib (Xospata), glasdegib maleate (Daurismo), hyaluronidase-zzxf (Phesgo), ibrutinib (Imbruvica), ibritumomab tiuxetan (Zevalin), idecabtagene vicleucel (Abecma), idelalisib (Zydelig), imatinib mesylate (Gleevec), infigratinib phosphate (Truseltiq), inotuzumab ozogamicin (Besponsa), iobenguane 1131 (Azedra), ipilimumab (Yervoy), isatuximab-irfc (Sarclisa), ivosidenib (Tibsovo), ixazomib citrate (Ninlaro), lanreotide acetate (Somatuline Depot), lapatinib (Tykerb), larotrectinib sulfate (Vitrakvi), lenvatinib mesylate (Lenvima), letrozole (Femara), lisocabtagene maraleucel (Breyanzi), loncastuximab tesirine-lpyl (Zynlonta), lorlatinib (Lorbrena), lutetium Lu 177-dotatate (Lutathera), margetuximab-cmkb (Margenza), midostaurin (Rydapt), mobocertinib succinate (Exkivity), mogamulizumab-kpkc (Poteligeo), moxetumomab39MF-364162710Attorney Docket No: 197102018940 pasudotox-tdfk (Lumoxiti), naxitamab-gqgk (Danyelza), necitumumab (Portrazza), neratinib maleate (Nerlynx), nilotinib (Tasigna), niraparib tosylate monohydrate (Zejula), nivolumab (Opdivo), obinutuzumab (Gazyva), ofatumumab (Arzerra), olaparib (Lynparza), olaratumab (Lartruvo), osimertinib (Tagrisso), palbociclib (Ibrance), panitumumab (Vectibix), panobinostat (Farydak), pazopanib (Votrient), pembrolizumab (Keytruda), pemigatinib (Pemazyre), pertuzumab (Perjeta), pexidartinib hydrochloride (Turalio), polatuzumab vedotin-piiq (Polivy), ponatinib hydrochloride (Iclusig), pralatrexate (Folotyn), pralsetinib (Gavreto), radium 223 dichloride (Xofigo), ramucirumab (Cyramza), regorafenib (Stivarga), ribociclib (Kisqali), ripretinib (Qinlock), rituximab (Rituxan), rituximab and hyaluronidase human (Rituxan Hycela), romidepsin (Istodax), rucaparib camsylate (Rubraca), ruxolitinib phosphate (Jakafi), sacituzumab govitecan-hziy (Trodelvy), seliciclib, selinexor (Xpovio), selpercatinib (Retevmo), selumetinib sulfate (Koselugo), siltuximab (Sylvant), sipuleucel-T (Provenge), sirolimus protein-bound particles (Fyarro), sonidegib (Odomzo), sorafenib (Nexavar), sotorasib (Lumakras), sunitinib (Sutent), tafasitamab-cxix (Monjuvi), tagraxofusp-erzs (Elzonris), talazoparib tosylate (Talzenna), tamoxifen (Nolvadex), tazemetostat hydrobromide (Tazverik), tebentafusp-tebn (Kimmtrak), temsirolimus (Torisel), tepotinib hydrochloride (Tepmetko), tisagenlecleucel (Kymriah), tisotumab vedotin-tftv (Tivdak), tocilizumab (Actemra), tofacitinib (Xeljanz), tositumomab (Bexxar), trametinib (Mekinist), trastuzumab (Herceptin), tretinoin (Vesanoid), tivozanib hydrochloride (Fotivda), toremifene (Fareston), tucatinib (Tukysa), umbralisib tosylate (Ukoniq), vandetanib (Caprelsa), vemurafenib (Zelboraf), venetoclax (Venclexta), vismodegib (Erivedge), vorinostat (Zolinza), zanubrutinib (Brukinsa), ziv-aflibercept (Zaltrap), or any combination thereof.

[0143] In some instances, administering the anti-cancer therapy may comprise adjusting a therapy or treatment (e.g., an anti-cancer treatment or anti-cancer therapy) for a subject, e.g., by adjusting a treatment dose and / or selecting a different treatment in response to a change in the determination of a methylation status at one or more genomic loci and / or methylation signature for the subject.

[0144] In some instances, the method can further include administering or applying a treatment or therapy (e.g., an anti-cancer agent, anti-cancer treatment, or anti-cancer therapy) to the subject based on the generated methylation signature. An anti-cancer agent or anticancer treatment may refer to a compound that is effective in the treatment of cancer cells. Examples of anti-cancer agents or anti-cancer therapies include, but not limited to, alkylating agents, antimetabolites, natural products, hormones, chemotherapy, radiation therapy,40MF-364162710Attorney Docket No: 197102018940 immunotherapy, surgery, or a therapy configured to target a defect in a specific cell signaling pathway, e.g., a defect in a DNA mismatch repair (MMR) pathway.

[0145] In some instances, a methylation signature generated according to the methods described herein may be used to monitor progression of a cancer. In some instances, a methylation signature generated according to the methods described herein may be used to monitor reoccurrence of a cancer. In some instances, the methods may be used for monitoring disease progression or recurrence (e.g., cancer or tumor progression or recurrence) in a subject. For example, in some instances, the methods may be used to determine a methylation state at a specified set of genomic loci in a first sample obtained from the subject at a first time point, and used to determine a methylation state at the specified set of genomic loci in a second sample obtained from the subject at a second time point, where comparison of the first determination of methylation state and the second determination of methylation state allows one to monitor disease progression or recurrence. In some instances, the first time point is chosen before the subject has been administered a therapy or treatment, and the second time point is chosen after the subject has been administered the therapy or treatment.

[0146] In some instances, a methylation signature may be generated as part of generating a genomic profile. In some instances, a genomic profile may comprise information on the presence of mutations (e.g., short nucleotide variants), copy number variations, epigenetic traits, proteins (or modifications thereof), and / or other biomarkers in an individual’s genome and / or proteome, as well as information on the individual’s corresponding phenotypic traits and the interaction between genetic or genomic traits, phenotypic traits, and environmental factors. In some instances, a genomic profile for the subject may comprise results from a comprehensive genomic profiling (CGP) test, a nucleic acid sequencing-based test, a gene expression profiling test, a cancer hotspot panel test, a DNA methylation test, a DNA fragmentation test, an RNA fragmentation test, or any combination thereof.Samples

[0147] The disclosed methods and systems may be used with any of a variety of samples (also referred to herein as specimens) comprising nucleic acids (e.g., DNA or RNA) that are collected from a subject (e.g., a patient). Examples of a sample include, but are not limited to, a tumor sample, a tissue sample, a biopsy sample (e.g., a tissue biopsy, a liquid biopsy, or both), a blood sample (e.g., a peripheral whole blood sample), a blood plasma sample, a blood serum sample, a lymph sample, a saliva sample, a sputum sample, a urine sample, a41MF-364162710Attorney Docket No: 197102018940 gynecological fluid sample, a circulating tumor cell (CTC) sample, a cerebral spinal fluid (CSF) sample, a pericardial fluid sample, a pleural fluid sample, an ascites (peritoneal fluid) sample, a feces (or stool) sample, or other body fluid, secretion, and / or excretion sample (or cell sample derived therefrom). In certain instances, the sample may be frozen sample or a formalin-fixed paraffin-embedded (FFPE) sample.

[0148] In some instances, the sample may be collected by tissue resection (e.g., surgical resection), needle biopsy, bone marrow biopsy, bone marrow aspiration, skin biopsy, endoscopic biopsy, fine needle aspiration, oral swab, nasal swab, vaginal swab or a cytology smear, scrapings, washings, or lavages (such as a ductal lavage or bronchoalveolar lavage), etc.

[0149] In some instances, the sample is a liquid biopsy sample, and may comprise, e.g., whole blood, blood plasma, blood serum, urine, stool, sputum, saliva, or cerebrospinal fluid. In some instances, the sample may be a liquid biopsy sample and may comprise circulating tumor cells (CTCs). In some instances, the sample may be a liquid biopsy sample and may comprise cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), or any combination thereof.

[0150] In some instances, the sample may comprise one or more premalignant or malignant cells. Premalignant, as used herein, refers to a cell or tissue that is not yet malignant but is poised to become malignant. In certain instances, the sample may be acquired from a solid tumor, a soft tissue tumor, or a metastatic lesion. In certain instances, the sample may be acquired from a hematologic malignancy or pre-malignancy. In other instances, the sample may comprise a tissue or cells from a surgical margin. In certain instances, the sample may comprise tumor-infiltrating lymphocytes. In some instances, the sample may comprise one or more non-malignant cells. In some instances, the sample may be, or is part of, a primary tumor or a metastasis (e.g., a metastasis biopsy sample). In some instances, the sample may be obtained from a site (e.g., a tumor site) with the highest percentage of tumor (e.g., tumor cells) as compared to adjacent sites (e.g., sites adjacent to the tumor). In some instances, the sample may be obtained from a site (e.g., a tumor site) with the largest tumor focus (e.g., the largest number of tumor cells as visualized under a microscope) as compared to adjacent sites (e.g., sites adjacent to the tumor).42MF-364162710Attorney Docket No: 197102018940Nucleic Acid extraction and processing

[0151] In some instances, the nucleic acids extracted from the sample may comprise deoxyribonucleic acid (DNA) molecules. Examples of DNA that may be suitable for analysis by the disclosed methods include, but are not limited to, genomic DNA or fragments thereof, mitochondrial DNA or fragments thereof, cell-free DNA (cfDNA), and circulating tumor DNA (ctDNA). Cell-free DNA (cfDNA) is comprised of fragments of DNA that are released from normal and / or cancerous cells during apoptosis and necrosis, and circulate in the blood stream and / or accumulate in other bodily fluids. Circulating tumor DNA (ctDNA) is comprised of fragments of DNA that are released from cancerous cells and tumors that circulate in the blood stream and / or accumulate in other bodily fluids.

[0152] In some instances, DNA is extracted from nucleated cells from the sample. In some instances, a sample may have a low nucleated cellularity, e.g., when the sample is comprised mainly of erythrocytes, lesional cells that contain excessive cytoplasm, or tissue with fibrosis. In some instances, a sample with low nucleated cellularity may require more, e.g., greater, tissue volume for DNA extraction.

[0153] In some instances, the nucleic acids extracted from the sample may comprise ribonucleic acid (RNA) molecules. Examples of RNA that may be suitable for analysis by the disclosed methods include, but are not limited to, total cellular RNA, total cellular RNA after depletion of certain abundant RNA sequences (e.g., ribosomal RNAs), cell-free RNA (cfRNA), messenger RNA (mRNA) or fragments thereof, the poly(A)-tailed mRNA fraction of the total RNA, ribosomal RNA (rRNA) or fragments thereof, transfer RNA (tRNA) or fragments thereof, and mitochondrial RNA or fragments thereof. In some instances, RNA may be extracted from the sample and converted to complementary DNA (cDNA) using, e.g., a reverse transcription reaction. In some instances, the cDNA is produced by random-primed cDNA synthesis methods. In other instances, the cDNA synthesis is initiated at the poly(A) tail of mature mRNAs by priming with oligo(dT)-containing oligonucleotides. Methods for depletion, poly(A) enrichment, and cDNA synthesis are well known to those of skill in the art.

[0154] In some instances, the sample may comprise a tumor content (e.g., comprising tumor cells or tumor cell nuclei), or a non-tumor content (e.g., immune cells, fibroblasts, and other non-tumor cells). In some instances, the tumor content of the sample may constitute a sample metric. In some instances, the sample may comprise a tumor content of at least 1-60%, 5- 50%, 10-40%, 15-25%, or 20-30% tumor cell nuclei. In some instances, the sample may43MF-364162710Attorney Docket No: 197102018940 comprise a tumor content of at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, or at least 50% tumor cell nuclei. In some instances, the percent tumor cell nuclei (e.g., sample fraction) is determined (e.g., calculated) by dividing the number of tumor cells in the sample by the total number of all cells within the sample that have nuclei. In some instances, for example when the sample is a liver sample comprising hepatocytes, a different tumor content calculation may be required due to the presence of hepatocytes having nuclei with twice, or more than twice, the DNA content of other, e.g., non-hepatocyte, somatic cell nuclei. In some instances, the sensitivity of detection of a genetic alteration, e.g., a variant sequence, or a determination of, e.g., micro satellite instability, may depend on the tumor content of the sample. For example, a sample having a lower tumor content can result in lower sensitivity of detection for a given size sample.

[0155] In some instances, as noted above, the sample comprises nucleic acid (e.g., DNA, RNA (or a cDNA derived from the RNA), or both), e.g., from a tumor or from normal tissue. In certain instances, the sample may further comprise a non-nucleic acid component, e.g., cells, protein, carbohydrate, or lipid, e.g., from the tumor or normal tissue.Targeting gene loci for analysis

[0156] The methods described herein can be used in combination with, or as part of, a method for evaluating a plurality or set of subject intervals (e.g., target sequences or target loci), e.g., from a set of genomic loci (e.g., specific sets of genomic loci, gene loci or fragments thereof, etc.), as described herein.

[0157] In some instances, the set of genomic loci evaluated by the disclosed methods comprises a plurality of, e.g., genes, which in mutant form, are associated with an effect on cell division, growth or survival, or are associated with a cancer, e.g., a cancer described herein.

[0158] In some instances, the set of genomic loci or gene loci evaluated by the disclosed methods comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, or more than 100 genomic loci or gene loci. In some instances, the one or more gene loci comprise between 10 and 20 loci, between 10 and 40 loci, between 10 and 60 loci, between 10 and 80 loci, between 10 and 100 loci, between 10 and 150 loci, between 10 and 200 loci, between 10 and 250 loci, between 10 and 300 loci, between 10 and 350 loci, between 10 and 400 loci, between 10 and 450 loci, between 10 and44MF-364162710Attorney Docket No: 197102018940500 loci, between 20 and 40 loci, between 20 and 60 loci, between 20 and 80 loci, between 20 and 100 loci, between 20 and 150 loci, between 20 and 200 loci, between 20 and 250 loci, between 20 and 300 loci, between 20 and 350 loci, between 20 and 400 loci, between 20 and 500 loci, between 40 and 60 loci, between 40 and 80 loci, between 40 and 100 loci, between 40 and 150 loci, between 40 and 200 loci, between 40 and 250 loci, between 40 and 300 loci, between 40 and 350 loci, between 40 and 400 loci, between 40 and 500 loci, between 60 and 80 loci, between 60 and 100 loci, between 60 and 150 loci, between 60 and 200 loci, between 60 and 250 loci, between 60 and 300 loci, between 60 and 350 loci, between 60 and 400 loci, between 60 and 500 loci, between 80 and 100 loci, between 80 and 150 loci, between 80 and 200 loci, between 80 and 250 loci, between 80 and 300 loci, between 80 and 350 loci, between 80 and 400 loci, between 80 and 500 loci, between 100 and 150 loci, between 100 and 200 loci, between 100 and 250 loci, between 100 and 300 loci, between 100 and 350 loci, between 100 and 400 loci, between 100 and 500 loci, between 150 and 200 loci, between 150 and 250 loci, between 150 and 300 loci, between 150 and 350 loci, between 150 and 400 loci, between 150 and 500 loci, between 200 and 250 loci, between 200 and 300 loci, between 200 and 350 loci, between 200 and 400 loci, between 200 and 500 loci, between 250 and 300 loci, between 250 and 350 loci, between 250 and 400 loci, between 250 and 500 loci, between 300 and 350 loci, between 300 and 400 loci, between 300 and 500 loci, between 350 and 400 loci, between 350 and 500 loci, or between 400 and 500 loci.

[0159] In some instances, the one or more gene loci comprise ABL1, ACVR1B, AKT1, AKT2, AKT3, ALK, ALOX12B, AMER1, APC, AR, ARAF, ARFRP1, ARID1A, ASXL1, ATM, ATR, ATRX, AURKA, AURKB, AXIN1, AXL, BAP1, BARD1, BCL2, BCL2L1, BCL2L2, BCL6, BCOR, BCORL1, BCR, BRAF, BRCA1, BRCA2, BRD4, BRIP1, BTG1, BTG2, BTK, CALR, CARD11, CASP8, CBFB, CBL, CCND1, CCND2, CCND3, CCNE1, CD22, CD274, CD70, CD74, CD79A, CD79B, CDC73, CDH1, CDK12, CDK4, CDK6, CDK8, CDKN1A, CDKN1B, CDKN2A, CDKN2B, CDKN2C, CEBPA, CHEK1, CHEK2, CIC, CREBBP, CRKL, CSF1R, CSF3R, CTCF, CTNNA1, CTNNB1, CUL3, CUL4A, CXCR4, CYP17A1, DAXX, DDR1, DDR2, DIS3, DNMT3A, DOT1L, EED, EGFR, EMSY (Cl lorf30), EP300, EPHA3, EPHB1, EPHB4, ERBB2, ERBB3, ERBB4, ERCC4, ERG, ERRFI1, ESRI, ETV4, ETV5, ETV6, EWSR1, EZH2, EZR, FAM46C, FANCA, FANCC, FANCG, FANCL, FAS, FBXW7, FGF10, FGF12, FGF14, FGF19, FGF23, FGF3, FGF4, FGF6, FGFR1, FGFR2, FGFR3, FGFR4, FH, FLCN, FLT1, FLT3, FOXL2, FUBP1, GABRA6, GATA3, GATA4, GATA6, GID4 (C17orf39), GNA11, GNA13, GNAQ, GNAS, GRM3, GSK3B, H3F3A, HDAC1, HGF, HNF1A, HRAS, HSD3B1, ID3, IDH1, IDH2,45MF-364162710Attorney Docket No: 197102018940IGF1R, IKBKE, IKZF1, INPP4B, IRF2, IRF4, IRS2, JAK1, JAK2, JAK3, JUN, KDM5A, KDM5C, KDM6A, KDR, KEAP1, KEL, KIT, KLHL6, KMT2A (MLL), KMT2D (MLL2), KRAS, LTK, LYN, MAF, MAP2K1, MAP2K2, MAP2K4, MAP3K1, MAP3K13, MAPK1, MCL1, MDM2, MDM4, MED12, MEF2B, MEN1, MERTK, MET, MITF, MKNK1, MLH1, MPL, MRE11A, MSH2, MSH3, MSH6, MST1R, MTAP, MTOR, MUTYH, MYB, MYC, MYCL, MYCN, MYD88, NBN, NF1, NF2, NFE2L2, NFKBIA, NKX2-1, NOTCH1, NOTCH2, NOTCH3, NPM1, NRAS, NT5C2, NTRK1, NTRK2, NTRK3, NUTM1, P2RY8, PALB2, PARK2, PARP1, PARP2, PARP3, PAX5, PBRM1, PDCD1, PDCD1LG2, PDGFRA, PDGFRB, PDK1, PIK3C2B, PIK3C2G, PIK3CA, PIK3CB, PIK3R1, PIM1, PMS2, POLDI, POLE, PPARG, PPP2R1A, PPP2R2A, PRDM1, PRKAR1A, PRKCI, PTCHI, PTEN, PTPN11, PTPRO, QKI, RAC1, RAD21, RAD51, RAD51B, RAD51C, RAD51D, RAD52, RAD54L, RAFI, RARA, RBI, RBM10, REL, RET, RICTOR, RNF43, ROS1, RPTOR, RSPO2, SDC4, SDHA, SDHB, SDHC, SDHD, SETD2, SF3B1, SGK1, SLC34A2, SMAD2, SMAD4, SMARCA4, SMARCB1, SMO, SNCAIP, SOCS1, SOX2, SOX9, SPEN, SPOP, SRC, STAG2, STAT3, STK11, SUFU, SYK, TBX3, TEK, TERC, TERT, TET2, TGFBR2, TIPARP, TMPRSS2, TNFAIP3, TNFRSF14, TP53, TSC1, TSC2, TYRO3, U2AF1, VEGFA, VHL, WHSCI, WHSC1L1, WT1, XPO1, XRCC2, ZNF217, ZNF703, or any combination thereof.

[0160] In some instances, the one or more gene loci comprise ABL, ALK, ALL, B4GALNT1, BAFF, BCL2, BRAF, BRCA, BTK, CD19, CD20, CD3, CD30, CD319, CD38, CD52, CDK4, CDK6, CML, CRACC, CS1, CTLA-4, dMMR, EGFR, ERBB1, ERBB2, FGFR1-3, FLT3, GD2, HDAC, HER1, HER2, HR, IDH2, IL-ip, IL-6, IL-6R, JAK1, JAK2, JAK3, KIT, KRAS, MEK, MET, MSLH, mTOR, PARP, PD-1, PDGFR, PDGFRa, PDGFRP, PD-L1, PI3K6, PIGF, PTCH, RAF, RANKL, RET, ROS1, SLAMF7, VEGF, VEGFA, VEGFB, or any combination thereof.

[0161] In some instances, the selected genomic loci or gene loci (also referred to herein as target loci, target gene loci, or target sequences), or fragments thereof, may include subject intervals comprising non-coding sequences, coding sequences, intragenic regions, or intergenic regions of the subject genome. For example, the subject intervals can include a non-coding sequence or fragment thereof (e.g., a promoter sequence, enhancer sequence, 5’ untranslated region (5’ UTR), 3’ untranslated region (3’ UTR), or a fragment thereof), a coding sequence of fragment thereof, an exon sequence or fragment thereof, an intron sequence or a fragment thereof.46MF-364162710Attorney Docket No: 197102018940

[0162] The methods described herein may comprise contacting a nucleic acid library with a plurality of target capture reagents in order to select and capture a plurality of specific target sequences (e.g., gene sequences or fragments thereof, methylated sequences or fragments thereof, or any combination thereof) for analysis. In some instances, a target capture reagent (i.e., a molecule which can bind to and thereby allow capture of a target molecule) is used to select the subject intervals to be analyzed. For example, a target capture reagent can be a bait molecule, e.g., a nucleic acid molecule (e.g., a DNA molecule or RNA molecule) which can hybridize to (i.e., is complementary to) a target molecule, and thereby allows capture of the target nucleic acid In some instances, the target capture reagent, e.g., a bait molecule (or bait sequence), is a capture oligonucleotide (or capture probe). In some instances, the target nucleic acid is a genomic DNA molecule, an RNA molecule, a cDNA molecule derived from an RNA molecule, a micro satellite DNA sequence, and the like. In some instances, the target capture reagent (e.g., a target capture probe) may comprise a binding agent such as a peptide or protein comprising a methyl-CpG binding domain (MDB) such that the target capture reagent (e.g., MBD-protein coupled beads) selectively binds nucleic acid sequences comprising methylated CpG sites. In some instances, the target nucleic acid is a nucleic acid sequence comprising one or more methylated CpG sites. In some instances, the target capture reagent is suitable for solution-phase hybridization to the target. In some instances, the target capture reagent is suitable for solid-phase hybridization to the target. In some instances, the target capture reagent is suitable for both solution-phase and solid-phase hybridization to the target. The design and construction of target capture reagents is described in more detail in, e.g., International Patent Application Publication No. WO 2020 / 236941, the entire content of which is incorporated herein by reference.

[0163] The methods described herein provide for optimized sequencing of a large number of genomic loci (e.g., genes or gene products (e.g., mRNA), micro satellite loci, etc.) from samples (e.g., cancerous tissue specimens, liquid biopsy samples, and the like) from one or more subjects by the appropriate selection of target capture reagents to select the target nucleic acid molecules to be sequenced. In some instances, a target capture reagent may hybridize to a specific target locus, e.g., a specific target gene locus or fragment thereof. In some instances, a target capture reagent may hybridize to a specific group of target loci, e.g., a specific group of gene loci or fragments thereof. In some instances, a plurality of target capture reagents comprising a mix of target-specific and / or group- specific target capture reagents may be used.47MF-364162710Attorney Docket No: 197102018940

[0164] In some instances, the number of target capture reagents (e.g., bait molecules) in the plurality of target capture reagents (e.g., a bait set) contacted with a nucleic acid library to capture a plurality of target sequences for nucleic acid sequencing is greater than 10, greater than 50, greater than 100, greater than 200, greater than 300, greater than 400, greater than 500, greater than 600, greater than 700, greater than 800, greater than 900, greater than 1,000, greater than 1,250, greater than 1,500, greater than 1,750, greater than 2,000, greater than 3,000, greater than 4,000, greater than 5,000, greater than 10,000, greater than 25,000, or greater than 50,000.

[0165] In some instances, the overall length of the target capture reagent sequence can be between about 70 nucleotides and 1000 nucleotides. In one instance, the target capture reagent length is between about 100 and 300 nucleotides, 110 and 200 nucleotides, or 120 and 170 nucleotides, in length. In addition to those mentioned above, intermediate oligonucleotide lengths of about 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 300, 400, 500, 600, 700, 800, and 900 nucleotides in length can be used in the methods described herein. In some embodiments, oligonucleotides of about 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, or 230 bases can be used.

[0166] In some instances, each target capture reagent sequence can include: (i) a targetspecific capture sequence (e.g., a gene locus or micro satellite locus-specific complementary sequence), (ii) an adapter, primer, barcode, and / or unique molecular identifier sequence, and (iii) universal tails on one or both ends. As used herein, the term "target capture reagent" can refer to the target- specific target capture sequence or to the entire target capture reagent oligonucleotide including the target-specific target capture sequence.

[0167] In some instances, the target-specific capture sequences in the target capture reagents are between about 40 nucleotides and 1000 nucleotides in length. In some instances, the target- specific capture sequence is between about 70 nucleotides and 300 nucleotides in length. In some instances, the target- specific sequence is between about 100 nucleotides and 200 nucleotides in length. In yet other instances, the target-specific sequence is between about 120 nucleotides and 170 nucleotides in length, typically 120 nucleotides in length. Intermediate lengths in addition to those mentioned above also can be used in the methods described herein, such as target-specific sequences of about 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 300, 400, 500, 600, 700, 800, and 900 nucleotides in length, as well as target-specific sequences of lengths between the above-mentioned lengths.48MF-364162710Attorney Docket No: 197102018940

[0168] In some instances, the target capture reagent may be designed to select a subject interval containing one or more rearrangements, e.g., an intron containing a genomic rearrangement. In such instances, the target capture reagent is designed such that repetitive sequences are masked to increase the selection efficiency. In those instances where the rearrangement has a known juncture sequence, complementary target capture reagents can be designed to recognize the juncture sequence to increase the selection efficiency.

[0169] In some instances, the disclosed methods may comprise the use of target capture reagents designed to capture two or more different target categories, each category having a different target capture reagent design strategy. In some instances, the hybridization-based capture methods and target capture reagent compositions disclosed herein may provide for the capture and homogeneous coverage of a set of target sequences, while minimizing coverage of genomic sequences outside of the targeted set of sequences. In some instances, the target sequences may include the entire exome of genomic DNA or a selected subset thereof. In some instances, the target sequences may include, e.g., a large chromosomal region (e.g., a whole chromosome arm). The methods and compositions disclosed herein provide different target capture reagents for achieving different sequencing depths and patterns of coverage for complex sets of target nucleic acid sequences.

[0170] Typically, DNA molecules are used as target capture reagent sequences, although RNA molecules can also be used. In some instances, a DNA molecule target capture reagent can be single stranded DNA (ssDNA) or double- stranded DNA (dsDNA). In some instances, an RNA-DNA duplex is more stable than a DNA-DNA duplex and therefore provides for potentially better capture of nucleic acids.

[0171] In some instances, the disclosed methods comprise providing a selected set of nucleic acid molecules (e.g., a library catch) captured from one or more nucleic acid libraries. For example, the method may comprise: providing one or a plurality of nucleic acid libraries, each comprising a plurality of nucleic acid molecules (e.g., a plurality of target nucleic acid molecules and / or reference nucleic acid molecules) extracted from one or more samples from one or more subjects; contacting the one or a plurality of libraries (e.g., in a solution-based hybridization reaction) with one, two, three, four, five, or more than five pluralities of target capture reagents (e.g., oligonucleotide target capture reagents) to form a hybridization mixture comprising a plurality of target capture reagent / nucleic acid molecule hybrids; separating the plurality of target capture reagent / nucleic acid molecule hybrids from said hybridization mixture, e.g., by contacting said hybridization mixture with a binding entity that allows for separation of said plurality of target capture reagent / nucleic acid molecule49MF-364162710Attorney Docket No: 197102018940 hybrids from the hybridization mixture, thereby providing a library catch (e.g., a selected or enriched subgroup of nucleic acid molecules from the one or a plurality of libraries).

[0172] In some instances, the disclosed methods may further comprise amplifying the library catch (e.g., by performing PCR). In other instances, the library catch is not amplified. In some instances, the target capture reagents can be part of a kit which can optionally comprise instructions, standards, buffers or enzymes or other reagents.

[0173] As noted above, the methods disclosed herein may include the step of contacting the library (e.g., the nucleic acid library) with a plurality of target capture reagents to provide a selected library target nucleic acid sequences (i.e., the library catch). The contacting step can be effected in, e.g., solution-based hybridization. In some instances, the method includes repeating the hybridization step for one or more additional rounds of solution-based hybridization. In some instances, the method further includes subjecting the library catch to one or more additional rounds of solution-based hybridization with the same or a different collection of target capture reagents.

[0174] In some instances, the contacting step is effected using a solid support, e.g., an array. Suitable solid supports for hybridization are described in, e.g., Albert, T.J. et al. (2007) Nat. Methods 4(11):903-5; Hodges, E. et al. (2007) Nat. Genet. 39(12): 1522-7; and Okou, D.T. et al. (2007) Nat. Methods 4(11):907-9, the contents of which are incorporated herein by reference in their entireties.

[0175] Hybridization methods that can be adapted for use in the methods herein are described in the art, e.g., as described in International Patent Application Publication No. WO 2012 / 092426, the entire content of which is incorporated herein by reference. Methods for hybridizing target capture reagents to a plurality of target nucleic acids are described in more detail in, e.g., International Patent Application Publication No. WO 2020 / 236941, the entire content of which is incorporated herein by reference.Subjects

[0176] In some instances, the sample is obtained (e.g., collected) from a subject (e.g., patient) with a condition or disease (e.g., a hyperproliferative disease or a non-cancer indication) or suspected of having the condition or disease. In some instances, the hyperproliferative disease is a cancer. In some instances, the cancer is a solid tumor or a metastatic form thereof. In some instances, the cancer is a hematological cancer, e.g., a leukemia or lymphoma.50MF-364162710Attorney Docket No: 197102018940

[0177] In some instances, the subject has a cancer or is at risk of having a cancer. For example, in some instances, the subject has a genetic predisposition to a cancer (e.g., having a genetic mutation that increases his or her baseline risk for developing a cancer). In some instances, the subject has been exposed to an environmental perturbation (e.g., radiation or a chemical) that increases his or her risk for developing a cancer. In some instances, the subject is in need of being monitored for development of a cancer. In some instances, the subject is in need of being monitored for cancer progression or regression, e.g., after being treated with an anti-cancer therapy (or anti-cancer treatment). In some instances, the subject is in need of being monitored for relapse of cancer. In some instances, the subject is in need of being monitored for minimum residual disease (MRD). In some instances, the subject has been, or is being treated, for cancer. In some instances, the subject has not been treated with an anticancer therapy (or anti-cancer treatment).

[0178] In some instances, the subject (e.g., a patient) is being treated, or has been previously treated, with one or more targeted therapies. In some instances, e.g., for a patient who has been previously treated with a targeted therapy, a post-targeted therapy sample (e.g., specimen) is obtained (e.g., collected). In some instances, the post-targeted therapy sample is a sample obtained after the completion of the targeted therapy.

[0179] In some instances, the patient has not been previously treated with a targeted therapy. In some instances, e.g., for a patient who has not been previously treated with a targeted therapy, the sample comprises a resection, e.g., an original resection, or a resection following recurrence (e.g., following a disease recurrence post- therapy).Cancers

[0180] In some instances, the sample is acquired from a subject having a cancer. Exemplary cancers include, but are not limited to, B cell cancer (e.g., multiple myeloma), melanomas, breast cancer, lung cancer (such as non-small cell lung carcinoma or NSCLC), bronchus cancer, colorectal cancer, prostate cancer, pancreatic cancer, stomach cancer, ovarian cancer, urinary bladder cancer, brain or central nervous system cancer, peripheral nervous system cancer, esophageal cancer, cervical cancer, uterine or endometrial cancer, cancer of the oral cavity or pharynx, liver cancer, kidney cancer, testicular cancer, biliary tract cancer, small bowel or appendix cancer, salivary gland cancer, thyroid gland cancer, adrenal gland cancer, osteosarcoma, chondrosarcoma, cancer of hematological tissues, adenocarcinomas, inflammatory myofibroblastic tumors, gastrointestinal stromal tumor (GIST), colon cancer,51MF-364162710Attorney Docket No: 197102018940 multiple myeloma (MM), myelodysplastic syndrome (MDS), myeloproliferative disorder (MPD), acute lymphocytic leukemia (ALL), acute myelocytic leukemia (AML), chronic myelocytic leukemia (CML), chronic lymphocytic leukemia (CLL), polycythemia Vera, Hodgkin lymphoma, non Hodgkin lymphoma (NHL), soft-tissue sarcoma, fibrosarcoma, myxosarcoma, liposarcoma, osteogenic sarcoma, chordoma, angiosarcoma, endotheliosarcoma, lymphangiosarcoma, lymphangioendotheliosarcoma, synovioma, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinomas, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatoma, bile duct carcinoma, choriocarcinoma, seminoma, embryonal carcinoma, Wilms' tumor, bladder carcinoma, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pinealoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, meningioma, neuroblastoma, retinoblastoma, follicular lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, hepatocellular carcinoma, thyroid cancer, gastric cancer, head and neck cancer, small cell cancers, essential thrombocythemia, agnogenic myeloid metaplasia, hypereosinophilic syndrome, systemic mastocytosis, familiar hypereosinophilia, chronic eosinophilic leukemia, neuroendocrine cancers, carcinoid tumors, and the like.

[0181] In some instances, the cancer comprises acute lymphoblastic leukemia (Philadelphia chromosome positive), acute lymphoblastic leukemia (precursor B-cell), acute myeloid leukemia (FLT3+), acute myeloid leukemia (with an IDH2 mutation), anaplastic large cell lymphoma, basal cell carcinoma, B-cell chronic lymphocytic leukemia, bladder cancer, breast cancer (HER2 overexpressed / amplified), breast cancer (HER2+), breast cancer (HR+, HER2- ), cervical cancer, cholangiocarcinoma, chronic lymphocytic leukemia, chronic lymphocytic leukemia (with 17p deletion), chronic myelogenous leukemia, chronic myelogenous leukemia (Philadelphia chromosome positive), classical Hodgkin lymphoma, colorectal cancer, colorectal cancer (dMMR and MSLH), colorectal cancer (KRAS wild type), cryopyrin- associated periodic syndrome, a cutaneous T-cell lymphoma, dermatofibrosarcoma protuberans, a diffuse large B-cell lymphoma, fallopian tube cancer, a follicular B-cell nonHodgkin lymphoma, a follicular lymphoma, gastric cancer, gastric cancer (HER2+), a gastroesophageal junction (GEJ) adenocarcinoma, a gastrointestinal stromal tumor, a gastrointestinal stromal tumor (KEr+), a giant cell tumor of the bone, a glioblastoma, granulomatosis with polyangiitis, a head and neck squamous cell carcinoma, a hepatocellular carcinoma, Hodgkin lymphoma, juvenile idiopathic arthritis, lupus erythematosus, a mantle52MF-364162710Attorney Docket No: 197102018940 cell lymphoma, medullary thyroid cancer, melanoma, a melanoma with a BRAF V600 mutation, a melanoma with a BRAF V600E or V600K mutation, Merkel cell carcinoma, multicentric Castleman's disease, multiple hematologic malignancies including Philadelphia chromosome-positive ALL and CML, multiple myeloma, myelofibrosis, a non-Hodgkin’s lymphoma, a nonresectable subependymal giant cell astrocytoma associated with tuberous sclerosis, a non-small cell lung cancer, a non-small cell lung cancer (ALK+), a non-small cell lung cancer (PD-L1+), a non-small cell lung cancer (with ALK fusion or ROS1 gene alteration), a non-small cell lung cancer (with BRAF V600E mutation), a non-small cell lung cancer (with an EGFR exon 19 deletion or exon 21 substitution (L858R) mutations), a non- small cell lung cancer (with an EGFR T790M mutation), ovarian cancer, ovarian cancer (with a BRCA mutation), pancreatic cancer, a pancreatic, gastrointestinal, or lung origin neuroendocrine tumor, a pediatric neuroblastoma, a peripheral T-cell lymphoma, peritoneal cancer, prostate cancer, a renal cell carcinoma, rheumatoid arthritis, a small lymphocytic lymphoma, a soft tissue sarcoma, a solid tumor (MSLH / dMMR), a squamous cell cancer of the head and neck, a squamous non-small cell lung cancer, thyroid cancer, a thyroid carcinoma, urothelial cancer, a urothelial carcinoma, or Waldenstrom's macroglobulinemia.

[0182] In some instances, the cancer is a hematologic malignancy (or premalignancy). As used herein, a hematologic malignancy refers to a tumor of the hematopoietic or lymphoid tissues, e.g., a tumor that affects blood, bone marrow, or lymph nodes. Exemplary hematologic malignancies include, but are not limited to, leukemia (e.g., acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), chronic myelogenous leukemia (CML), hairy cell leukemia, acute monocytic leukemia (AMoL), chronic myelomonocytic leukemia (CMML), juvenile myelomonocytic leukemia (JMML), or large granular lymphocytic leukemia), lymphoma (e.g., AIDS-related lymphoma, cutaneous T-cell lymphoma, Hodgkin lymphoma (e.g., classical Hodgkin lymphoma or nodular lymphocyte-predominant Hodgkin lymphoma), mycosis fungoides, non-Hodgkin lymphoma (e.g., B-cell non-Hodgkin lymphoma (e.g., Burkitt lymphoma, small lymphocytic lymphoma (CLL / SLL), diffuse large B-cell lymphoma, follicular lymphoma, immunoblastic large cell lymphoma, precursor B-lymphoblastic lymphoma, or mantle cell lymphoma) or T- cell non-Hodgkin lymphoma (mycosis fungoides, anaplastic large cell lymphoma, or precursor T-lymphoblastic lymphoma)), primary central nervous system lymphoma, Sezary syndrome, Waldenstrom macroglobulinemia), chronic myeloproliferative neoplasm, Langerhans cell histiocytosis, multiple myeloma / plasma cell neoplasm, myelodysplastic syndrome, or myelodysplastic / myeloproliferative neoplasm.53MF-364162710Attorney Docket No: 197102018940Systems for generating mbias corrected sequence read data

[0183] Also disclosed herein are systems designed to implement any of the disclosed methods for generating mbias corrected sequence read data. The systems may comprise, e.g., an automated DNA sequencing library preparation module, a sequencer, one or more processors, and a memory unit communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: receive sequence read data corresponding to the plurality of sequence reads; and perform a computational correction of mbias in the sequence read data to produce mbias- corrected sequence read data.

[0184] In some instances, the automated DNA sequencing library preparation module is configured to perform one or more steps selected from i) generating a plurality of modified nucleic acid fragments by performing: an end repair reaction on a plurality of nucleic acid fragments extracted from a sample from a subject to repair fragmentation damage, and / or a nick / gap repair reaction on a plurality of nucleic acid fragments extracted from a sample from a subject to fill in single- stranded nicks or gaps; and / or ii) converting non-methylated cytosines in the plurality of modified nucleic acid fragments into uracil to generate methylation nucleic acid fragments. In some instances, the automated DNA sequencing library preparation module is configured to perform any of the chemistry methods described herein.

[0185] In some instances, performing the computational correction of mbias in the sequence read data comprises determining, using the one or more processors, a methylation status of one or more methylation sites in the sequence read data for one or more sequence reads; and for each sequence read for which the methylation status of one or more methylation sites located proximal to a 3 ’-end of the sequence read is determined to be unmethylated: comparing, using the one or more processors, the unmethylated status of the one or more methylation sites located proximal to the 3 ’-end of the respective sequence read to the methylation status of other sequence reads of the plurality of sequence reads that overlap the one or more methylation sites; performing a statistical analysis to determine a probability of a correlation between the methylation status of the one or more methylation sites located proximal to the 3 ’-end of the respective sequence read to that for other sequence reads of the plurality of sequence reads that overlap the one or more methylation sites; detecting 3 ’-end hypomethylation bias in the respective sequence read based on a comparison between the54MF-364162710Attorney Docket No: 197102018940 determined probability and a predetermined threshold; and generating a corrected sequence read to replace the respective sequence read based on the methylation status of the other sequence reads of the plurality of sequence reads that overlap the one or more methylation sites if 3 ’-end hypomethylation bias is detected.

[0186] In some instances, the disclosed systems may further comprise a sequencer, e.g., a next generation sequencer (also referred to as a massively parallel sequencer). Examples of next generation (or massively parallel) sequencing platforms include, but are not limited to, Roche / 454’s Genome Sequencer (GS) FLX system, Illumina / Solexa’s Genome Analyzer (GA), Illumina’s HiSeq® 2500, HiSeq® 3000, HiSeq® 4000 and NovaSeq® 6000 sequencing systems, Life / APG’s Support Oligonucleotide Ligation Detection (SOLiD) system, Polonator’s G.007 system, Helicos BioSciences’ HeliScope Gene Sequencing system, ThermoFisher Scientific’s Ion Torrent Genexus system, or Pacific Biosciences’ PacBio® RS system.

[0187] In some instances, the disclosed systems may be used for preparing sequencing libraries and / or for performing sequencing of nucleic acid molecules (e.g., DNA molecules) extracted from any of a variety of samples as described herein (e.g., a tissue sample, biopsy sample, hematological sample, or liquid biopsy sample derived from the subject).

[0188] In some instances, the plurality of genomic loci (e.g., CpG sites) for which sequencing data is processed to determine a methylation state may comprise at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or more than 1000 genomic loci.

[0189] In some instance, the nucleic acid sequence data is acquired using a next generation sequencing technique (also referred to as a massively parallel sequencing technique) having a read-length of less than 400 bases, less than 300 bases, less than 200 bases, less than 150 bases, less than 100 bases, less than 90 bases, less than 80 bases, less than 70 bases, less than 60 bases, less than 50 bases, less than 40 bases, or less than 30 bases.

[0190] In some instances, the determination of a methylation state at one or more CpG loci, or a methylations signature for DNA extracted from a sample from a subject may be used to select, initiate, adjust, or terminate a treatment for cancer in the subject (e.g., a patient) from which the sample was derived, as described elsewhere herein.

[0191] In some instances, the disclosed systems may further comprise sample processing and library preparation workstations, microplate-handling robotics, fluid dispensing systems, temperature control modules, environmental control chambers, additional data storage modules, data communication modules (e.g., Bluetooth®, WiFi, intranet, or internet55MF-364162710Attorney Docket No: 197102018940 communication hardware and associated software), display modules, one or more local and / or cloud-based software packages (e.g., instrument / system control software packages, sequencing data analysis software packages), etc., or any combination thereof. In some instances, the systems may comprise, or be part of, a computer system or computer network as described elsewhere herein.Computer systems and networks

[0192] FIG. 11 illustrates an example of a computing device or system in accordance with one embodiment. Device 1100 can be a host computer connected to a network. Device 1100 can be a client computer or a server. As shown in FIG. 11, device 1100 can be any suitable type of microprocessor-based device, such as a personal computer, workstation, server, or handheld computing device (portable electronic device) such as a phone or tablet. The device can include, for example, one or more processor(s) 1110, input devices 1120, output devices 1130, memory or storage devices 1140, communication devices 1160, and nucleic acid sequencers 1170. The mbias correction module 1150 residing in memory or storage device 1140 may comprise, e.g., an operating system as well as software for executing the computational mbias correction methods described herein. Input device 1120 and output device 1130 can generally correspond to those described herein, and can either be connectable or integrated with the computer.

[0193] Input device 1120 can be any suitable device that provides input, such as a touch screen, keyboard or keypad, mouse, or voice-recognition device. Output device 1130 can be any suitable device that provides output, such as a touch screen, haptics device, or speaker.

[0194] Storage 1140 can be any suitable device that provides storage (e.g., an electrical, magnetic or optical memory including a RAM (volatile and non-volatile), cache, hard drive, or removable storage disk). Communication device 1160 can include any suitable device capable of transmitting and receiving signals over a network, such as a network interface chip or device. The components of the computer can be connected in any suitable manner, such as via a wired media (e.g., a physical system bus 1180, Ethernet connection, or any other wire transfer technology) or wirelessly (e.g., Bluetooth®, Wi-Fi®, or any other wireless technology).

[0195] Software module 1150, which can be stored as executable instructions in storage 1140 and executed by processor(s) 1110, can include, for example, an operating system and / or the56MF-364162710Attorney Docket No: 197102018940 processes that embody the functionality of the methods of the present disclosure (e.g., as embodied in the devices as described herein).

[0196] Software module 1150 can also be stored and / or transported within any non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device, such as those described herein, that can fetch instructions associated with the software from the instruction execution system, apparatus, or device and execute the instructions. In the context of this disclosure, a computer-readable storage medium can be any medium, such as storage 1140, that can contain or store processes for use by or in connection with an instruction execution system, apparatus, or device. Examples of computer-readable storage media may include memory units like hard drives, flash drives and distribute modules that operate as a single functional unit. Also, various processes described herein may be embodied as modules configured to operate in accordance with the embodiments and techniques described above. Further, while processes may be shown and / or described separately, those skilled in the art will appreciate that the above processes may be routines or modules within other processes.

[0197] Software module 1150 can also be propagated within any transport medium for use by or in connection with an instruction execution system, apparatus, or device, such as those described above, that can fetch instructions associated with the software from the instruction execution system, apparatus, or device and execute the instructions. In the context of this disclosure, a transport medium can be any medium that can communicate, propagate, or transport programming for use by or in connection with an instruction execution system, apparatus, or device. The transport readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, or infrared wired or wireless propagation medium.

[0198] Device 1100 may be connected to a network (e.g., network 1104, as shown in FIG. 12 and / or described below), which can be any suitable type of interconnected communication system. The network can implement any suitable communications protocol and can be secured by any suitable security protocol. The network can comprise network links of any suitable arrangement that can implement the transmission and reception of network signals, such as wireless network connections, T1 or T3 lines, cable networks, DSL, or telephone lines.

[0199] Device 1100 can be implemented using any operating system, e.g., an operating system suitable for operating on the network. Software module 1150 can be written in any suitable programming language, such as C, C++, Java, or Python. In various embodiments, 57MF-364162710Attorney Docket No: 197102018940 application software embodying the functionality of the present disclosure can be deployed in different configurations, such as in a client / server arrangement or through a Web browser as a Web-based application or Web service, for example. In some embodiments, the operating system is executed by one or more processors, e.g., processor(s) 1110.

[0200] Device 1100 can further include a sequencer 1170, which can be any suitable nucleic acid sequencing instrument.

[0201] FIG. 12 illustrates an example of a computing system in accordance with one embodiment. In system 1200, device 1100 (e.g., as described above and illustrated in FIG. 11) is connected to network 1204, which is also connected to device 1206. In some embodiments, device 1206 is a sequencer. Exemplary sequencers can include, without limitation, Roche / 454’s Genome Sequencer (GS) FLX System, Illumina / Solexa’s Genome Analyzer (GA), Illumina’s HiSeq® 2500, HiSeq® 3000, HiSeq® 4000 and NovaSeq® 6000 Sequencing Systems, Life / APG’s Support Oligonucleotide Ligation Detection (SOLiD) system, Polonator’s G.007 system, Helicos BioSciences’ HeliScope Gene Sequencing system, or Pacific Biosciences’ PacBio® RS system.

[0202] Devices 1100 and 1206 may communicate, e.g., using suitable communication interfaces via network 1204, such as a Local Area Network (LAN), Virtual Private Network (VPN), or the Internet. In some embodiments, network 1204 can be, for example, the Internet, an intranet, a virtual private network, a cloud network, a wired network, or a wireless network. Devices 1100 and 1206 may communicate, in part or in whole, via wireless or hardwired communications, such as Ethernet, IEEE 802.11b wireless, or the like.Additionally, devices 1100 and 1206 may communicate, e.g., using suitable communication interfaces, via a second network, such as a mobile / cellular network. Communication between devices 1100 and 1206 may further include or communicate with various servers such as a mail server, mobile server, media server, telephone server, and the like. In some embodiments, Devices 1100 and 1206 can communicate directly (instead of, or in addition to, communicating via network 1204), e.g., via wireless or hardwired communications, such as Ethernet, IEEE 802.11b wireless, or the like. In some embodiments, devices 1100 and 1206 communicate via communications 1208, which can be a direct connection or can occur via a network (e.g., network 1204).

[0203] One or all of devices 1100 and 1206 generally include logic (e.g., http web server logic) or are programmed to format data, accessed from local or remote databases or other sources of data and content, for providing and / or receiving information via network 1404 according to various examples described herein.58MF-364162710Attorney Docket No: 197102018940EXEMPLARY EMBODIMENTS

[0204] The following exemplary embodiments are representative of some aspects of the invention:1. A method for generating mbias-corrected sequence read data, comprising: receiving a plurality of nucleic acid fragments extracted from a sample from a subject; generating a plurality of modified nucleic acid fragments by performing: an end repair reaction on at least one nucleic acid fragment in the plurality of nucleic acid fragments to repair fragmentation damage; and / or a nick / gap repair reaction on at least one nucleic acid fragment in the plurality of nucleic acid fragments to fill in single-stranded nicks or gaps; converting non-methylated cytosines in the plurality of modified nucleic acid fragments into uracil to generate methylation nucleic acid fragments; sequencing, by a sequencer, the methylation nucleic acid fragments to generate a plurality of sequence reads; receiving, at one or more processors, sequence read data corresponding to the plurality of sequence reads; and performing, using the one or more processors, a computational correction of mbias in the sequence read data to produce mbias-corrected sequence read data.2. The method of clause 1, wherein performing the computational correction of mbias in the sequence read data comprises: determining, using the one or more processors, a methylation status of one or more methylation sites in the sequence read data for one or more sequence reads; and for each sequence read for which the methylation status of one or more methylation sites located proximal to a 3 ’-end of the sequence read is determined to be unmethylated: comparing, using the one or more processors, the unmethylated status of the one or more methylation sites located proximal to the 3 ’-end of the respective sequence read to the methylation status of other sequence reads of the plurality of sequence reads that overlap the one or more methylation sites; performing a statistical analysis to determine a probability of a correlation between the methylation status of the one or more methylation sites located proximal to the 3 ’-end of the respective sequence read to that for59MF-364162710Attorney Docket No: 197102018940 other sequence reads of the plurality of sequence reads that overlap the one or more methylation sites; detecting 3 ’-end hypomethylation bias in the respective sequence read based on a comparison between the determined probability and a predetermined threshold; and generating a corrected sequence read to replace the respective sequence read based on the methylation status of the other sequence reads of the plurality of sequence reads that overlap the one or more methylation sites if 3 ’-end hypomethylation bias is detected.3. The method of clause 1 or clause 2, further comprising generating a methylation signature by: for each corrected sequence read, calculating, using the one or more processors, a probability that a methylation state of the corrected sequence read is significantly different from a distribution of methylation states determined for a plurality of sequence reads derived from samples from healthy individuals that map to a same genomic interval.4. The method of clause 3, further comprising diagnosing the subject based on the methylation signature, monitoring progression of a cancer based on the methylation signature, and / or monitoring recurrence of the cancer based on the methylation signature.5. The method of clause 3 or clause 4, wherein the methylation state for each corrected sequence read is determined based on a corrected methylation status of each of one or more sites within the corrected sequence read, wherein a site at which the corrected methylation status is determined is a methylation site.6. The method of any of clauses 3-5, wherein the methylation state comprises a methylation fraction value calculated based on the corrected methylation status of each of the one or more methylation sites within the corrected sequence read.7. The method of clause 6, wherein the methylation fraction value is calculated as: methylation fraction = N / M, wherein M is a total number of methylation sites located within the corrected sequence read and N is a number of methylations sites that are methylated.60MF-364162710Attorney Docket No: 1971020189408. The method of any of clauses 1-7, wherein the plurality of sequence reads align to one or more genomic intervals of interest, and wherein the one or more genomic intervals of interest are selected based on a cancer to be detected, a cancer to be diagnosed, or a likelihood of response to a treatment for a cancer.9. The method of any of clauses 1-8, wherein the sample comprises a tissue biopsy sample or a liquid biopsy sample obtained from the subject.10. The method of any of clauses 1-9, wherein the plurality of nucleic acid fragments comprises a mixture of tumor nucleic acid molecules and non-tumor nucleic acid molecules.11. The method of any of clauses 1-10, wherein the end repair reaction comprises use of a chain termination mechanism.12. The method of clause 11, wherein the chain termination mechanism comprises use of a chain termination nucleotide in the end repair reaction.13. The method of clause 12, wherein the chain termination nucleotide comprises a 2', 3'- dideoxyribonucleoside 5 '-triphosphate (ddNTP).14. The method of clause 13, wherein the ddNTP comprises a 2',3'-dideoxycytidine 5'- triphosphate (ddCTP), 2',3'-dideoxyguanosine 5'-triphosphate (ddGTP), 2', 3'- dideoxythymidine 5'-triphosphate (ddTTP), 2',3'-dideoxyadenosine 5'-triphosphate (ddATP), or any combination thereof.15. The method of any of clauses 1-14, wherein the nick / gap repair reaction comprises use of a ligase.16. The method of clause 15, wherein the ligase comprises a DNA ligase.17. The method of clause 16, wherein the DNA ligase comprises Taq DNA ligase, T4 DNA ligase, 9°N™ DNA ligase, T3 DNA ligase, or any combination thereof.61MF-364162710Attorney Docket No: 19710201894018. The method of any of clauses 1-17, wherein converting non-methylated cytosines in the plurality of modified nucleic acid fragments comprises use of a bisulfite reaction.19. The method of any of clauses 1-17, wherein converting non-methylated cytosines in the plurality of nucleic acid fragments to generate a methylation nucleic acid fragments comprises use of an enzymatic conversion reaction.20. The method of any of clauses 2-19, wherein the one or more methylation sites are located proximal to a 3 ’-end of each sequence read of the plurality of sequence reads based on the sequence read data.21. The method of any of clauses 2-20, wherein a site at which the methylation status is determined is a methylation site.22. The method of any of clauses 2-21, wherein performing the statistical analysis comprises: generating a contingency table that tabulates results of the comparisons of the methylation status of the one or more methylation sites located proximal to the 3 ’-end of the respective sequence read to that for other sequence reads of the plurality of sequence reads that overlap the one or more methylation sites; and performing the statistical analysis between two or more factors in the contingency table.23. The method of clause 22, wherein the contingency table comprises a 2 x 2 contingency table that tabulates the results of the comparisons in terms of two factors (a distal or not distal location of the one or more methylation sites) and two outcomes (a fully unmethylated or not fully unmethylated status of the one or more methylation sites).24. The method of any of clauses 2-23, wherein the predetermined threshold is determined using a multi-test correction method to adjust the probabilities of the association determined by the statistical analysis to correct for an occurrence of false positives.62MF-364162710Attorney Docket No: 19710201894025. The method of clause 24, wherein 3 ’-end hypomethylation bias in the respective sequence read is detected if the determined probability is greater than the predetermined threshold.26. The method of any of clauses 2- 25, wherein generating a corrected sequence read comprises truncating a 3 ’-end of the respective sequence read.27. The method of any one of clauses 2-26, wherein the corrected sequence read is a raw sequence read, an aligned sequence read, a merged sequence read, a computationally reconstructed sequence read and / or a consensus sequence read.28. The method of any one of clauses 1-27, wherein the sequence read is a raw sequence read, an aligned sequence read, a merged sequence read, a computationally reconstructed sequence read, and / or a consensus sequence read.29. The method of any of clauses 2-28, wherein the one or more methylation sites comprise one or more CpG dinucleotide sites and / or one or more non-CpG dinucleotide sites.30. The method of any of clauses 1-29, wherein the plurality of modified nucleic acid fragments align to one or more genomic intervals of interest.31. The method of clause 30, wherein one or more genomic intervals of interest are selected based on a Tumor and Tissue of Origin (TTOO) classification.32. The method of clause 31, wherein the one or more genomic intervals of interest are selected based on a cancer to be detected.33. The method any of clauses 30-32, wherein the one or more genomic intervals of interest comprise one or more compact genomic regions.34. The method of any one of claims 3-33, wherein the plurality of sequence reads derived from samples from healthy individuals align to one or more genomic intervals of interest that are the same as the one or more genomic intervals of interest to which the plurality of sequence reads obtained from the sample from the subject align.63MF-364162710Attorney Docket No: 19710201894035. The method of any one of clauses 1 -34, wherein the sequence read data comprises methyl-seq data.36. The method of any one of clauses 1 - 35, wherein the subject is suspected of having or is determined to have cancer.37. The method of clause 36, wherein the cancer is a solid tumor.38. The method of clause 36, wherein the cancer is a hematological cancer.39. The method of clause 36, wherein the cancer is a B cell cancer (multiple myeloma), a melanoma, breast cancer, lung cancer, bronchus cancer, colorectal cancer, prostate cancer, pancreatic cancer, stomach cancer, ovarian cancer, urinary bladder cancer, brain cancer, central nervous system cancer, peripheral nervous system cancer, esophageal cancer, cervical cancer, uterine cancer, endometrial cancer, cancer of an oral cavity, cancer of a pharynx, liver cancer, kidney cancer, testicular cancer, biliary tract cancer, small bowel cancer, appendix cancer, salivary gland cancer, thyroid gland cancer, adrenal gland cancer, osteosarcoma, chondrosarcoma, a cancer of hematological tissue, an adenocarcinoma, an inflammatory myofibroblastic tumor, a gastrointestinal stromal tumor (GIST), colon cancer, multiple myeloma (MM), myelodysplastic syndrome (MDS), myeloproliferative disorder (MPD), acute lymphocytic leukemia (ALL), acute myelocytic leukemia (AML), chronic myelocytic leukemia (CML), chronic lymphocytic leukemia (CLL), polycythemia Vera, Hodgkin lymphoma, non-Hodgkin lymphoma (NHL), soft-tissue sarcoma, fibrosarcoma, myxosarcoma, liposarcoma, osteogenic sarcoma, chordoma, angiosarcoma, endotheliosarcoma, lymphangiosarcoma, lymphangioendotheliosarcoma, synovioma, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinomas, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatoma, bile duct carcinoma, choriocarcinoma, seminoma, embryonal carcinoma, Wilms' tumor, bladder carcinoma, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pinealoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, meningioma, neuroblastoma, retinoblastoma, follicular lymphoma, diffuse large B-cell64MF-364162710Attorney Docket No: 197102018940 lymphoma, mantle cell lymphoma, hepatocellular carcinoma, thyroid cancer, gastric cancer, head and neck cancer, small cell cancer, essential thrombocythemia, agnogenic myeloid metaplasia, hypereosinophilic syndrome, systemic mastocytosis, familiar hypereosinophilia, chronic eosinophilic leukemia, neuroendocrine cancers, or a carcinoid tumor.40. The method of clause 37, wherein the cancer comprises acute lymphoblastic leukemia (Philadelphia chromosome positive), acute lymphoblastic leukemia (precursor B-cell), acute myeloid leukemia (FLT3+), acute myeloid leukemia (with an IDH2 mutation), anaplastic large cell lymphoma, basal cell carcinoma, B-cell chronic lymphocytic leukemia, bladder cancer, breast cancer (HER2 overexpressed / amplified), breast cancer (HER2+), breast cancer (HR+, HER2-), cervical cancer, cholangiocarcinoma, chronic lymphocytic leukemia, chronic lymphocytic leukemia (with 17p deletion), chronic myelogenous leukemia, chronic myelogenous leukemia (Philadelphia chromosome positive), classical Hodgkin lymphoma, colorectal cancer, colorectal cancer (dMMR / MSI-H), colorectal cancer (KRAS wild type), cryopyrin-associated periodic syndrome, a cutaneous T-cell lymphoma, dermatofibrosarcoma protuberans, a diffuse large B-cell lymphoma, fallopian tube cancer, a follicular B-cell nonHodgkin lymphoma, a follicular lymphoma, gastric cancer, gastric cancer (HER2+), gastroesophageal junction (GEJ) adenocarcinoma, a gastrointestinal stromal tumor, a gastrointestinal stromal tumor (KIT+), a giant cell tumor of the bone, a glioblastoma, granulomatosis with polyangiitis, a head and neck squamous cell carcinoma, a hepatocellular carcinoma, Hodgkin lymphoma, juvenile idiopathic arthritis, lupus erythematosus, a mantle cell lymphoma, medullary thyroid cancer, melanoma, a melanoma with a BRAF V600 mutation, a melanoma with a BRAF V600E or V600K mutation, Merkel cell carcinoma, multicentric Castleman's disease, multiple hematologic malignancies including Philadelphia chromosome-positive ALL and CML, multiple myeloma, myelofibrosis, a non-Hodgkin’s lymphoma, a nonresectable subependymal giant cell astrocytoma associated with tuberous sclerosis, a non-small cell lung cancer, a non-small cell lung cancer (ALK+), a non-small cell lung cancer (PD-L1+), a non-small cell lung cancer (with ALK fusion or ROS1 gene alteration), a non-small cell lung cancer (with BRAF V600E mutation), a non-small cell lung cancer (with an EGFR exon 19 deletion or exon 21 substitution (L858R) mutations), a non- small cell lung cancer (with an EGFR T790M mutation), ovarian cancer, ovarian cancer (with a BRCA mutation), pancreatic cancer, a pancreatic, gastrointestinal, or lung origin neuroendocrine tumor, a pediatric neuroblastoma, a peripheral T-cell lymphoma, peritoneal cancer, prostate cancer, a renal cell carcinoma, rheumatoid arthritis, a small lymphocytic65MF-364162710Attorney Docket No: 197102018940 lymphoma, a soft tissue sarcoma, a solid tumor (MSI-H / dMMR), a squamous cell cancer of the head and neck, a squamous non-small cell lung cancer, thyroid cancer, a thyroid carcinoma, urothelial cancer, a urothelial carcinoma, or Waldenstrom's macroglobulinemia.41. The method of any one of clauses 36-40, further comprising treating the subject with an anti-cancer therapy.42. The method of clause 41 wherein the anti-cancer therapy comprises a targeted anticancer therapy, chemotherapy, radiation therapy, immunotherapy, or surgery.43. The method of clause 42, wherein the targeted anti-cancer therapy comprises abemaciclib (Verzenio), abiraterone acetate (Zytiga), acalabrutinib (Calquence), ado- trastuzumab emtansine (Kadcyla), afatinib dimaleate (Gilotrif), aldesleukin (Proleukin), alectinib (Alecensa), alemtuzumab (Campath), alitretinoin (Panretin), alpelisib (Piqray), amivantamab-vmjw (Rybrevant), anastrozole (Arimidex), apalutamide (Erleada), asciminib hydrochloride (Scemblix), atezolizumab (Tecentriq), avapritinib (Ayvakit), avelumab (Bavencio), axicabtagene ciloleucel (Yescarta), axitinib (Inlyta), belantamab mafodotin-blmf (Blenrep), belimumab (Benlysta), belinostat (Beleodaq), belzutifan (Welireg), bevacizumab (Avastin), bexarotene (Targretin), binimetinib (Mektovi), blinatumomab (Blincyto), bortezomib (Velcade), bosutinib (Bosulif), brentuximab vedotin (Adcetris), brexucabtagene autoleucel (Tecartus), brigatinib (Alunbrig), cabazitaxel (Jevtana), cabozantinib (Cabometyx), cabozantinib (Cabometyx, Cometriq), canakinumab (Haris), capmatinib hydrochloride (Tabrecta), carfilzomib (Kyprolis), cemiplimab-rwlc (Libtayo), ceritinib (LDK378 / Zykadia), cetuximab (Erbitux), cobimetinib (Cotellic), copanlisib hydrochloride (Aliqopa), crizotinib (Xalkori), dabrafenib (Tafinlar), dacomitinib (Vizimpro), daratumumab (Darzalex), daratumumab and hyaluronidase-fihj (Darzalex Faspro), darolutamide (Nubeqa), dasatinib (Sprycel), denileukin diftitox (Ontak), denosumab (Xgeva), dinutuximab (Unituxin), dostarlimab-gxly (Jemperli), durvalumab (Imfinzi), duvelisib (Copiktra), elotuzumab (Empliciti), enasidenib mesylate (Idhifa), encorafenib (Braftovi), enfortumab vedotin-ejfv (Padcev), entrectinib (Rozlytrek), enzalutamide (Xtandi), erdafitinib (Balversa), erlotinib (Tarceva), everolimus (Afinitor), exemestane (Aromasin), fam-trastuzumab deruxtecan-nxki (Enhertu), fedratinib hydrochloride (Inrebic), fulvestrant (Faslodex), gefitinib (Iressa), gemtuzumab ozogamicin (Mylotarg), gilteritinib (Xospata), glasdegib maleate (Daurismo), hyaluronidase-zzxf (Phesgo), ibrutinib (Imbruvica), ibritumomab66MF-364162710Attorney Docket No: 197102018940 tiuxetan (Zevalin), idecabtagene vicleucel (Abecma), idelalisib (Zydelig), imatinib mesylate (Gleevec), infigratinib phosphate (Truseltiq), inotuzumab ozogamicin (Besponsa), iobenguane 1131 (Azedra), ipilimumab (Yervoy), isatuximab-irfc (Sarclisa), ivosidenib (Tibsovo), ixazomib citrate (Ninlaro), lanreotide acetate (Somatuline Depot), lapatinib (Tykerb), larotrectinib sulfate (Vitrakvi), lenvatinib mesylate (Lenvima), letrozole (Femara), lisocabtagene maraleucel (Breyanzi), loncastuximab tesirine-lpyl (Zynlonta), lorlatinib (Lorbrena), lutetium Lu 177-dotatate (Lutathera), margetuximab-cmkb (Margenza), midostaurin (Rydapt), mobocertinib succinate (Exkivity), mogamulizumab-kpkc (Poteligeo), moxetumomab pasudotox-tdfk (Lumoxiti), naxitamab-gqgk (Danyelza), necitumumab (Portrazza), neratinib maleate (Nerlynx), nilotinib (Tasigna), niraparib tosylate monohydrate (Zejula), nivolumab (Opdivo), obinutuzumab (Gazyva), ofatumumab (Arzerra), olaparib (Lynparza), olaratumab (Lartruvo), osimertinib (Tagrisso), palbociclib (Ibrance), panitumumab (Vectibix), panobinostat (Farydak), pazopanib (Votrient), pembrolizumab (Keytruda), pemigatinib (Pemazyre), pertuzumab (Perjeta), pexidartinib hydrochloride (Turalio), polatuzumab vedotin-piiq (Polivy), ponatinib hydrochloride (Iclusig), pralatrexate (Folotyn), pralsetinib (Gavreto), radium 223 dichloride (Xofigo), ramucirumab (Cyramza), regorafenib (Stivarga), ribociclib (Kisqali), ripretinib (Qinlock), rituximab (Rituxan), rituximab and hyaluronidase human (Rituxan Hycela), romidepsin (Istodax), rucaparib camsylate (Rubraca), ruxolitinib phosphate (Jakafi), sacituzumab govitecan-hziy (Trodelvy), seliciclib, selinexor (Xpovio), selpercatinib (Retevmo), selumetinib sulfate (Koselugo), siltuximab (Sylvant), sipuleucel-T (Provenge), sirolimus protein-bound particles (Fyarro), sonidegib (Odomzo), sorafenib (Nexavar), sotorasib (Lumakras), sunitinib (Sutent), tafasitamab-cxix (Monjuvi), tagraxofusp-erzs (Elzonris), talazoparib tosylate (Talzenna), tamoxifen (Nolvadex), tazemetostat hydrobromide (Tazverik), tebentafusp-tebn (Kimmtrak), temsirolimus (Torisel), tepotinib hydrochloride (Tepmetko), tisagenlecleucel (Kymriah), tisotumab vedotin-tftv (Tivdak), tocilizumab (Actemra), tofacitinib (Xeljanz), tositumomab (Bexxar), trametinib (Mekinist), trastuzumab (Herceptin), tretinoin (Vesanoid), tivozanib hydrochloride (Fotivda), toremifene (Fareston), tucatinib (Tukysa), umbralisib tosylate (Ukoniq), vandetanib (Caprelsa), vemurafenib (Zelboraf), venetoclax (Venclexta), vismodegib (Erivedge), vorinostat (Zolinza), zanubrutinib (Brukinsa), ziv-aflibercept (Zaltrap), or any combination thereof.44. The method of any one of clauses 1 - 43, further comprising obtaining the sample from the subject.67MF-364162710Attorney Docket No: 19710201894045. The method of any one of clauses 1 - 44, wherein the sample comprises a tissue biopsy sample, a liquid biopsy sample, or a normal control.46. The method of clauses 45, wherein the sample is a liquid biopsy sample and comprises blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva.47. The method of clause 45, wherein the sample is a liquid biopsy sample and comprises circulating tumor cells (CTCs).48. The method of clause 45, wherein the sample is a liquid biopsy sample and comprises cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), or any combination thereof.49. The method of any one of clauses 1 to 48, wherein the plurality of nucleic acid fragments comprises a mixture of tumor nucleic acid fragments and non-tumor nucleic acid fragments.50. The method of clause 49, wherein the tumor nucleic acid fragments are derived from a tumor portion of a heterogeneous tissue biopsy sample, and the non-tumor nucleic acid fragments are derived from a normal portion of the heterogeneous tissue biopsy sample.51. The method of clause 49, wherein the sample comprises a liquid biopsy sample, and wherein the tumor nucleic acid fragments are derived from a circulating tumor DNA (ctDNA) fraction of the liquid biopsy sample, and the non-tumor nucleic acid fragments are derived from a non-tumor, cell-free DNA (cfDNA) fraction of the liquid biopsy sample.52. The method of any of clauses 1-51, further comprising preparing a sequencing library.53. The method of any of clauses 1-52, wherein the sequencing comprises a single-end sequencing method.54. The method of any of clauses 1-52, wherein the sequencing comprises a paired-end sequencing method.68MF-364162710Attorney Docket No: 19710201894055. The method of any of clauses 1-54, wherein the sequencing comprises use of a massively parallel sequencing (MPS) technique, whole genome sequencing (WGS), whole exome sequencing, targeted sequencing, direct sequencing, or Sanger sequencing technique.56. The method of clause 55, wherein the sequencing comprises massively parallel sequencing, and the massively parallel sequencing technique comprises next generation sequencing (NGS).57. The method of any of clauses 1-56, wherein the sequencer comprises a next generation sequencer.58. The method of any one of clauses 1 - 57, wherein one or more of the plurality of sequence reads overlap one or more genomic loci within one or more subgenomic intervals in the sample.59. The method of clause 58, wherein the one or more genomic loci comprises between 10 and 20 loci, between 10 and 40 loci, between 10 and 60 loci, between 10 and 80 loci, between 10 and 100 loci, between 10 and 150 loci, between 10 and 200 loci, between 10 and 250 loci, between 10 and 300 loci, between 10 and 350 loci, between 10 and 400 loci, between 10 and 450 loci, between 10 and 500 loci, between 20 and 40 loci, between 20 and 60 loci, between 20 and 80 loci, between 20 and 100 loci, between 20 and 150 loci, between 20 and 200 loci, between 20 and 250 loci, between 20 and 300 loci, between 20 and 350 loci, between 20 and 400 loci, between 20 and 500 loci, between 40 and 60 loci, between 40 and 80 loci, between 40 and 100 loci, between 40 and 150 loci, between 40 and 200 loci, between 40 and 250 loci, between 40 and 300 loci, between 40 and 350 loci, between 40 and 400 loci, between 40 and 500 loci, between 60 and 80 loci, between 60 and 100 loci, between 60 and 150 loci, between 60 and 200 loci, between 60 and 250 loci, between 60 and 300 loci, between 60 and 350 loci, between 60 and 400 loci, between 60 and 500 loci, between 80 and 100 loci, between 80 and 150 loci, between 80 and 200 loci, between 80 and 250 loci, between 80 and 300 loci, between 80 and 350 loci, between 80 and 400 loci, between 80 and 500 loci, between 100 and 150 loci, between 100 and 200 loci, between 100 and 250 loci, between 100 and 300 loci, between 100 and 350 loci, between 100 and 400 loci, between 100 and 500 loci, between 150 and 200 loci, between 150 and 250 loci, between 150 and 300 loci, between 150 and 350 loci, between 150 and 400 loci, between 150 and 500 loci, between 20069MF-364162710Attorney Docket No: 197102018940 and 250 loci, between 200 and 300 loci, between 200 and 350 loci, between 200 and 400 loci, between 200 and 500 loci, between 250 and 300 loci, between 250 and 350 loci, between 250 and 400 loci, between 250 and 500 loci, between 300 and 350 loci, between 300 and 400 loci, between 300 and 500 loci, between 350 and 400 loci, between 350 and 500 loci, or between 400 and 500 loci.60. The method of clause 58 or clause 59, wherein the one or more genomic loci comprise one or more gene loci.61. The method of clause 60, wherein the one or more gene loci comprise ABL1, ACVR1B, AKT1, AKT2, AKT3, ALK, ALOX12B, AMER1, APC, AR, ARAF, ARFRP1, ARID1A, ASXL1, ATM, ATR, ATRX, AURKA, AURKB, AXIN1, AXL, BAP1, BARD1, BCL2, BCL2L1, BCL2L2, BCL6, BCOR, BCORL1, BCR, BRAF, BRCA1, BRCA2, BRD4, BRIP1, BTG1, BTG2, BTK, CALR, CARD11, CASP8, CBFB, CBL, CCND1, CCND2, CCND3, CCNE1, CD22, CD274, CD70, CD74, CD79A, CD79B, CDC73, CDH1, CDK12, CDK4, CDK6, CDK8, CDKN1A, CDKN1B, CDKN2A, CDKN2B, CDKN2C, CEBPA, CHEK1, CHEK2, CIC, CREBBP, CRKL, CSF1R, CSF3R, CTCF, CTNNA1, CTNNB1, CUL3, CUL4A, CXCR4, CYP17A1, DAXX, DDR1, DDR2, DIS3, DNMT3A, DOT1L, EED, EGFR, EMSY (Cl lorf30), EP300, EPHA3, EPHB1, EPHB4, ERBB2, ERBB3, ERBB4, ERCC4, ERG, ERRFI1, ESRI, ETV4, ETV5, ETV6, EWSR1, EZH2, EZR, FAM46C, FANCA, FANCC, FANCG, FANCL, FAS, FBXW7, FGF10, FGF12, FGF14, FGF19, FGF23, FGF3, FGF4, FGF6, FGFR1, FGFR2, FGFR3, FGFR4, FH, FLCN, FLT1, FLT3, FOXL2, FUBP1, GABRA6, GATA3, GATA4, GATA6, GID4 (C17orf39), GNA11, GNA13, GNAQ, GNAS, GRM3, GSK3B, H3F3A, HDAC1, HGF, HNF1A, HRAS, HSD3B1, ID3, IDH1, IDH2, IGF1R, IKBKE, IKZF1, INPP4B, IRF2, IRF4, IRS2, JAK1, JAK2, JAK3, JUN, KDM5A, KDM5C, KDM6A, KDR, KEAP1, KEL, KIT, KLHL6, KMT2A (MLL), KMT2D (MLL2), KRAS, LTK, LYN, MAF, MAP2K1, MAP2K2, MAP2K4, MAP3K1, MAP3K13, MAPK1, MCL1, MDM2, MDM4, MED12, MEF2B, MEN1, MERTK, MET, MITF, MKNK1, MLH1, MPL, MRE11A, MSH2, MSH3, MSH6, MST1R, MTAP, MTOR, MUTYH, MYB, MYC, MYCL, MYCN, MYD88, NBN, NF1, NF2, NFE2L2, NFKBIA, NKX2-1, NOTCH1, NOTCH2, NOTCH3, NPM1, NRAS, NT5C2, NTRK1, NTRK2, NTRK3, NUTM1, P2RY8, PALB2, PARK2, PARP1, PARP2, PARP3, PAX5, PBRM1, PDCD1, PDCD1LG2, PDGFRA, PDGFRB, PDK1, PIK3C2B, PIK3C2G, PIK3CA, PIK3CB, PIK3R1, PIM1, PMS2, POLDI, POLE, PPARG, PPP2R1A, PPP2R2A,70MF-364162710Attorney Docket No: 197102018940PRDM1, PRKAR1A, PRKCI, PTCHI, PTEN, PTPN11, PTPRO, QKI, RAC1, RAD21, RAD51, RAD51B, RAD51C, RAD51D, RAD52, RAD54L, RAFI, RARA, RBI, RBM10, REL, RET, RICTOR, RNF43, ROS1, RPTOR, RSPO2, SDC4, SDHA, SDHB, SDHC, SDHD, SETD2, SF3B1, SGK1, SLC34A2, SMAD2, SMAD4, SMARCA4, SMARCB1, SMO, SNCAIP, SOCS1, SOX2, SOX9, SPEN, SPOP, SRC, STAG2, STAT3, STK11, SUFU, SYK, TBX3, TEK, TERC, TERT, TET2, TGFBR2, TIPARP, TMPRSS2, TNFAIP3, TNFRSF14, TP53, TSC1, TSC2, TYRO3, U2AF1, VEGFA, VHL, WHSCI, WHSC1L1, WT1, XPO1, XRCC2, ZNF217, ZNF703, or any combination thereof.62. The method of clause 61, wherein the one or more gene loci comprise ABL, ALK, ALL, B4GALNT1, BAFF, BCL2, BRAF, BRCA, BTK, CD19, CD20, CD3, CD30, CD319, CD38, CD52, CDK4, CDK6, CML, CRACC, CS1, CTLA-4, dMMR, EGFR, ERBB1, ERBB2, FGFR1-3, FLT3, GD2, HDAC, HER1, HER2, HR, IDH2, IL-ip, IL-6, IL-6R, JAK1, JAK2, JAK3, KIT, KRAS, MEK, MET, MSLH, mTOR, PARP, PD-1, PDGFR, PDGFRa, PDGFRP, PD-L1, PI3K6, PIGF, PTCH, RAF, RANKL, RET, ROS1, SLAMF7, VEGF, VEGFA, VEGFB, or any combination thereof.63. A system comprising: an automated DNA sequencing library preparation module; a sequencer; one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: receive sequence read data corresponding to the plurality of sequence reads; and perform a computational correction of mbias in the sequence read data to produce mbias-corrected sequence read data.64. The system of clause 63, wherein performing the computational correction of mbias in the sequence read data comprises: determining, using the one or more processors, a methylation status of one or more methylation sites in the sequence read data for one or more sequence reads; and71MF-364162710Attorney Docket No: 197102018940 for each sequence read for which the methylation status of one or more methylation sites located proximal to a 3 ’-end of the sequence read is determined to be unmethylated: comparing, using the one or more processors, the unmethylated status of the one or more methylation sites located proximal to the 3 ’-end of the respective sequence read to the methylation status of other sequence reads of the plurality of sequence reads that overlap the one or more methylation sites; performing a statistical analysis to determine a probability of a correlation between the methylation status of the one or more methylation sites located proximal to the 3 ’-end of the respective sequence read to that for other sequence reads of the plurality of sequence reads that overlap the one or more methylation sites; detecting 3 ’-end hypomethylation bias in the respective sequence read based on a comparison between the determined probability and a predetermined threshold; and generating a corrected sequence read to replace the respective sequence read based on the methylation status of the other sequence reads of the plurality of sequence reads that overlap the one or more methylation sites if 3 ’-end hypomethylation bias is detected.65. The system of clause 63 or clause 64, wherein the automated DNA sequencing library preparation module is configured to perform one or more steps selected from: i) generating a plurality of modified nucleic acid fragments by performing: an end repair reaction on a plurality of nucleic acid fragments extracted from a sample from a subject to repair fragmentation damage, and / or a nick / gap repair reaction on a plurality of nucleic acid fragments extracted from a sample from a subject to fill in single- stranded nicks or gaps; and / or ii) converting non-methylated cytosines in the plurality of modified nucleic acid fragments into uracil to generate methylation nucleic acid fragments.66. The system of clause 65, wherein the end repair reaction comprises use of a chain termination mechanism.72MF-364162710Attorney Docket No: 19710201894067. The system of clause 66, wherein the chain termination mechanism comprises use of a chain termination nucleotide in the end repair reaction.68. The system of clause 67, wherein the chain termination nucleotide comprises a 2', 3'- dideoxyribonucleoside 5 '-triphosphate (ddNTP).69. The system of any of clauses 65-68, wherein the nick / gap repair reaction comprises use of a ligase.70. The system of any of clauses 64-69, wherein the one or more methylation sites are located proximal to a 3 ’-end of each sequence read of the plurality of sequence reads based on the sequence read data.71. The system of any of clauses 64-68, wherein a site at which the methylation status is determined is a methylation site.72. The system of any of clauses 64-71, wherein performing the statistical analysis comprises: generating a contingency table that tabulates results of the comparisons of the methylation status of the one or more methylation sites located proximal to the 3 ’-end of the respective sequence read to that for other sequence reads of the plurality of sequence reads that overlap the one or more methylation sites; and performing the statistical analysis between two or more factors in the contingency table.73. The system of clause 72, wherein the contingency table comprises a 2 x 2 contingency table that tabulates the results of the comparisons in terms of two factors (a distal or not distal location of the one or more methylation sites) and two outcomes (a fully unmethylated or not fully unmethylated status of the one or more methylation sites).74. The system of any of clauses 64-73, wherein the predetermined threshold is determined using a multi-test correction method to adjust the probabilities of the association determined by the statistical test to correct for an occurrence of false positives.73MF-364162710Attorney Docket No: 19710201894075. The system of clause 74, wherein 3’-end hypomethylation bias in the respective sequence read is detected if the determined probability is greater than the predetermined threshold.76. The system of any of clauses 64-75, wherein generating a corrected sequence read comprises truncating a 3 ’-end of the respective sequence read.EXAMPLES

[0205] The invention will be more fully understood by reference to the following examples. They should not, however, be construed as limiting the scope of the invention. It is understood that the examples and embodiments described herein are for illustrative purposes only and that various modifications or changes in light thereof will be suggested to persons skilled in the art and are to be included within the spirit and purview of this application and scope of the appended claims.Example 1: Correcting mbias by combining computational and chemistry corrections

[0206] cfDNA samples from 30 individuals (2 healthy individuals, 23 individuals with tumors, and 5 individuals with a non-cancerous condition)WGBS data was prepared for each of the samples according to one of three protocols. The three protocols were applied to equal numbers of samples within the three populations.

[0207] For the first protocol (EvoPrep), the cfDNA was subjected to a methylseq conversion reaction without correcting for mbias and sequenced with NGS. For the second protocol (EvoPrep + TaqLigase), cfDNA was subjected to a methylseq conversion reaction and TaqLigase was used to perform a nick or gap repair during library preparation. The libraries were sequences with WGS. For the third protocol, (Evoprep ddCTP + TaqLigase), cfDNA was subjected to a methylseq conversion reaction then TaqLigase was used to perform a nick or gap repair and dideoxyctidine triphosphates (ddCTP) were used to perform an end repair reaction during library preparation. The libraries were sequenced with WGS using lOng as input. Two libraries, one from an individual with a tumor and one from an individual with a non-cancerous condition were sequenced using WGS using 20ng DNA as input). The mean unique coverage across all 96 sequencing libraries was between 6 and 15.74MF-364162710Attorney Docket No: 197102018940

[0208] Mbias was quantified using a read position (from 5' to 3')-based metric. The 5'- 3' position of the sequence read represents a sequenced fragment. For a given genomic interval, all sequence reads that aligned to a reference position in the interval were collected and mbias illustrations were created to visualize the magnitude of mbias in hypermethylated genomic regions. (FIG. 3A) All sequence reads aligning to the region were aligned according to the position in the fragment from the 5' end to the 3' end. As shown in FIG. 3B, a graph was created using the coordinates in the read wherein x=0 represented the 3 ’-most base of each fragment, regardless of where it mapped and regardless of orientation as the x axis and the coverage, calculated as the total coverage of all unmethylated and methylated CpG dinucleotides stacked at each x position, on the y axis. As shown in FIG. 3C, a second graph was created using the same values on the x-axis but the y axis was the mF, calculated as the number of methylated C’s divided by the total number of C’s in a CpG context at that position. The fragment lengths were typically between 162 and 172 bases long.

[0209] This unique and novel visualization allowed for decoupling the two types of mbias, end-repair mbias and strand displacement mbias. A region without mbias would appear as a constant straight line close to ImF for the full read.

[0210] As demonstrated in FIGs. 4A-3B end repair mbias appears as a drop off on mF on the 3 ’end of the reads and strand displacement mbias appears as a deviation from 1.0 mF in the remaining part of the read.

[0211] End-repair mbias was corrected for computationally in the EvoPrep, EvoPrep +TaqLigase and Evoprep ddCTP + Laqligase. The method was based on searching for 3’ distal blocks of unmethylated CpG dinucleotides in each DNA sequence read derived from an individual sample. For each sequence read comprising one or more unmethylated CpG loci near the 3’ end, it’s methylation state was compared to that of all other sequence reads from the same samples that mapped to the same genomic region i.e., that map to the same or more CpG sites). The likelihood that one or more CpG site in the sequence read being evaluated are unmethylated given the methylation status observed in the one or more CpG sites in other sequence reads mapping to the same genomic region was calculated using a Fisher’s Exact Test.

[0212] The method is illustrated in FIG. 5 For a given read (or sequence read) under test (FUT), the objective was to determine whether or not to trim (z.e., discard the data for) a distal block of one or more unmethylated CpG sites in view of the methylation status observed for the one or more CpG sites in reads (sequence reads) that map to the same location in the genome.75MF-364162710Attorney Docket No: 197102018940

[0213] Each read in the sequencing library was evaluated to identify a set of reads that contain a distal block of unmethylated CpG dinucleotides; this defined the set of reads for which distal correction would potentially be applied.

[0214] Then, all reads in the sequencing library that comprised at least one CpG dinucleotide that overlapped the distal block in a given FUT were identified to define a set, S, of comparator reads (this set included the FUT).

[0215] For each read, f, in S, define (if it exists) a distal block in f, and define a “distal_status_per_CG_dinucleotide” for each CpG dinucleotide in the fragment depending on whether the CpG dinucleotide was in a distal block (true) or not (false).

[0216] For each read, f, in S, CpG dinucleotides were identified, d, that overlap with the distal block of the FUT. If at least one CpG dinucleotide in d had a distal_status_per_CG_dinucleotide that was false, “fragment_contains_non_distal” status was set to true, otherwise this was set to false. If at least one CpG dinucleotide in d is unmethylated, “fragment_fully_unmethylated” status was set to false; otherwise this was set to true.

[0217] If the fragment_fully_unmethylated status was true, and the fragment_contains_non_distal status wad false, the “fragment_has_distal_block” status was set to true; if the fragment_fully_unmethylated status is false, the fragment_has_distal_block status was set to true if at least one CpG dinucleotide in d overlaped the CpG dinucleotides in the distal block of the FUT. The fragment_fully_unmethylated and fragment_has_distal_block status of each read, f, in S was then used to construct a 2x2 contingency table, as illustrated in the lower right corner of FIG. 5.

[0218] As noted above, the FUT was only assessed if it had a 3’ block of one or more unmethylated CpGs. If it did, the genomic interval spanning the distal block of one or more CpGs was defined and two questions were asked for each read from the sample that overlaps the genomic interval (as described above):(i) Is the read fully unmethylated in the genomic interval? If yes, then the column labeled “fully unmethylated” in FIG. 5 is checked.(ii) Does the read have distal, unmethylated bases? If yes, then the column labeled “distal” in FIG. 5 is checked.

[0219] Once all reads that have overlap the specified genomic interval have been assessed, the total counts for each of the four possible outcomes were tabulated in the 2x2 contingency table (FIG. 5, lower right insert):Fully-unmethylated and distal (upper left comer of table)76MF-364162710Attorney Docket No: 197102018940Fully-unmethylated and NOT distal (upper right corner of table) NOT fully-unmethylated and distal (lower left corner of table) NOT fully-unmethylated and NOT distal (lower right comer of table)

[0220] A Fisher’s Exact Test was then performed using the contingency table to assess the significance of the association (contingency) between the two classification categories (distal and unmethylated) and generate a p-value. If the p-value (or a multi-test corrected p-value) was less than a predetermined cut-off (e.g., less than 0.05, 0.01, 0.005, or 0.001), then the distal block of the FUT was truncated (i.e., the data for the unmethylated block is discarded). If the p-value (or a multi-test corrected p-value) was greater than or equal to the predetermined cut-off, the distal block of the FUT was left intact.Example 2: Increased coverage at hypermethylated regions improve computational correction

[0221] Hybrid capture sequencing samples were obtained from 6 healthy individuals. These libraries were sequenced to a mean unique coverage of 80-120 (high coverage samples).

[0222] The reads mapping to the 3600 hypermethylated regions intervals from the high coverage samples were interested with the Evoprep, Evoprep ddCTP and Evoprep ddCTP + TaqLigase samples to increase the coverage in the hypermethylated regions for the samples with chemistry corrections. The healthy samples had a target mean coverage of about 70x. The non-cancerous samples had a target mean coverage of about 1 lOx. The integration of the high coverage samples with samples collected for healthy individuals, non-cancerous individuals, and individuals with tumors did not distort the analysis because the integration was not done to analyze methylation at individual regions but to analyze the effect of mbias corrections on samples with a range of coverages.

[0223] FIGs. 6A-6F demonstrates mbias calculated two ways according to coverage in the hypermethylated regions for the samples after combining the reads.

[0224] First, methylation bias was quantified by calculating the standard deviation of the difference between the observed methylation percentage bias and the expected methylation percentage across the length of a sequence read. The standard deviation of methylation position bias (SD_methyl_position) was plotted against the mean coverage in the sample for the hypermethylated regions (mean_target_coverage). (FIGs. 6A-6C)

[0225] Second, the mbias introduced by strand displacement activity was quantified by calculated using a “slip rate.” Slip rate was calculated by (i) looking at methylation in a control region set, e.g., a set of genomic regions that show high levels of methylation in a77MF-364162710Attorney Docket No: 197102018940 wide variety of samples, including both cancer samples and healthy samples, (ii) analyzing the trend for methylation data (e.g., average methylation fraction (AMF) data) versus sequence read base position at the 5’ end of the DNA fragment; the level typically starts high and then shows an approximately linear decrease, and (iii) analyzing the methylation versus sequence read base position data using, for example, a regression model to determine the slope of the approximately linear decrease. Slip rate for was plotted against the mean coverage in the sample for the hypermethylated regions (mean_target_coverage). (FIGs. 6D- 6F)

[0226] FIGs. 6A-6F demonstrate that the data has a wide range of coverage and mbiases. This allowed for the testing of the mbias corrections across mbias and coverage levels in the samples.

[0227] The effect of the computational and chemistry mbias corrections were evaluated on the Evoprep, Evoprep ddCTP and Evoprep ddCTP + TaqLigase the samples before and after the additional the high coverage sample reads were integrated.

[0228] FIG. 7 shows the mbias for a sample collected from a healthy individual, after integrating the high coverage reads before the high coverage reads were integrated. The left side of the plot is before the computational correction was applied and the right side shows the mbias after the computational correction was applied. The sample was processed with Evoprep.

[0229] FIG. 8A shows mbias for a sample collected from an individual a noncancerous condition before the high coverage reads were integrated. The left side of the plot is before the computational correction was applied and the right side shows the mbias after the computational correction was applied. FIG. 8B shows the same sample after the high coverage reads were integrated.

[0230] As demonstrated by the reduction in the left-hand tail in the data on computationally corrected plots in FIGs. 7 and 8A-8B, the computational correction was more effective in samples with higher coverage.

[0231] To further show that the computational correction is effective at higher coverages. The computational correction was applied to a sample with higher coverage than the one showed in FIGs. 7 and 8A-8B. The sample had higher coverage before any of the high coverage samples were added because 20ng of DNA were used as input to WGS rather than lOng DNA. The results are shown in FIGs. 9A-9B. FIG 9A shows mbias for a sample from an individual with a noncancerous condition after the high coverage reads were integrated. The left side of the plot is before the computational correction was applied and the right side78MF-364162710Attorney Docket No: 197102018940 shows the mbias after the computational correction was applied. FIG. 9B shows the same sample processed with Evoprep ddCTP + TaqLigase. The strand displacement mbias has been corrected in FIG. 9B as demonstrated by the higher and flatter mF values.

[0232] The sample in FIG. 9B shows the least 3’ end bias and least strand displacement bias. The sample also represents the highest coverage of those in FIGs. 7, 8A-8B, and 9A-9B.These results demonstrate that increased coverage at hypermethylated regions improve the effect of the computational correction when both computational and chemistry corrections are applied during methylation data processing.Example 3: mbias level does not dictate effectiveness of mbias correction methods

[0233] The effect of mbias on the ability of the computational and chemistry methods to correct for mbias was tested. Mbias was categorized as High or Low according to a laboratory assessment. FIGs. 10A-10B show mF before and after computation corrections for 3 samples that were not subjected to any of the chemistry corrections. There is not clear difference in the magnitude of mbias correction between the High and Low mbias samples.(FIGs. 10A-10B)

[0234] Together, these results demonstrated that the level of mbias in a sample processed with or without chemistry corrections does not change the magnitude of computational mbias correction. Thus, demonstrating the complementary nature of the chemistry mbias correction for correcting for strand displacement mbias and the computational mbias correction for correcting 3’ end mbias.79MF-364162710

Claims

Attorney Docket No: 197102018940CLAIMSWhat is claimed is:

1. A method for generating mbias-corrected sequence read data, comprising: receiving a plurality of nucleic acid fragments extracted from a sample from a subject; generating a plurality of modified nucleic acid fragments by performing: an end repair reaction on at least one nucleic acid fragment in the plurality of nucleic acid fragments to repair fragmentation damage; and / or a nick / gap repair reaction on at least one nucleic acid fragment in the plurality of nucleic acid fragments to fill in single-stranded nicks or gaps; converting non-methylated cytosines in the plurality of modified nucleic acid fragments into uracil to generate methylation nucleic acid fragments; sequencing, by a sequencer, the methylation nucleic acid fragments to generate a plurality of sequence reads; receiving, at one or more processors, sequence read data corresponding to the plurality of sequence reads; and performing, using the one or more processors, a computational correction of mbias in the sequence read data to produce mbias-corrected sequence read data.

2. The method of claim 1, wherein performing the computational correction of mbias in the sequence read data comprises: determining, using the one or more processors, a methylation status of one or more methylation sites in the sequence read data for one or more sequence reads; and for each sequence read for which the methylation status of one or more methylation sites located proximal to a 3 ’-end of the sequence read is determined to be unmethylated: comparing, using the one or more processors, the unmethylated status of the one or more methylation sites located proximal to the 3 ’-end of the respective sequence read to the methylation status of other sequence reads of the plurality of sequence reads that overlap the one or more methylation sites; performing a statistical analysis to determine a probability of a correlation between the methylation status of the one or more methylation sites located proximal to the 3 ’-end of the respective sequence read to that for80MF-364162710Attorney Docket No: 197102018940 other sequence reads of the plurality of sequence reads that overlap the one or more methylation sites; detecting 3 ’-end hypomethylation bias in the respective sequence read based on a comparison between the determined probability and a predetermined threshold; and generating a corrected sequence read to replace the respective sequence read based on the methylation status of the other sequence reads of the plurality of sequence reads that overlap the one or more methylation sites if 3 ’-end hypomethylation bias is detected.

3. The method of claim 1, further comprising generating a methylation signature by: for each corrected sequence read, calculating, using the one or more processors, a probability that a methylation state of the corrected sequence read is significantly different from a distribution of methylation states determined for a plurality of sequence reads derived from samples from healthy individuals that map to a same genomic interval.

4. The method of claim 3, further comprising diagnosing the subject based on the methylation signature, monitoring progression of a cancer based on the methylation signature, and / or monitoring recurrence of the cancer based on the methylation signature.

5. The method of claim 3, wherein the methylation state for each corrected sequence read is determined based on a corrected methylation status of each of one or more sites within the corrected sequence read, wherein a site at which the corrected methylation status is determined is a methylation site.

6. The method of claim 3, wherein the methylation state comprises a methylation fraction value calculated based on the corrected methylation status of each of the one or more methylation sites within the corrected sequence read , wherein the methylation fraction value is calculated as: methylation fraction = N / M, wherein M is a total number of methylation sites located within the corrected sequence read and N is a number of methylations sites that are methylated.

7. The method of claim 1, wherein the plurality of sequence reads align to one or more genomic intervals of interest, and wherein the one or more genomic intervals of interest are81MF-364162710Attorney Docket No: 197102018940 selected based on a cancer to be detected, a cancer to be diagnosed, or a likelihood of response to a treatment for a cancer.

8. The method of claim 1, wherein the end repair reaction comprises use of a chain termination mechanism, wherein the chain termination mechanism comprises use of a chain termination nucleotide in the end repair reaction.

9. The method of claim 1, wherein the nick / gap repair reaction comprises use of a ligase, wherein the ligase comprises a DNA ligase.

10. The method of claim 1, wherein converting non-methylated cytosines in the plurality of modified nucleic acid fragments comprises use of a bisulfite reaction and / or converting non-methylated cytosines in the plurality of nucleic acid fragments to generate a methylation nucleic acid fragments comprises use of an enzymatic conversion reaction.

11. The method of claim 2, wherein the one or more methylation sites are located proximal to a 3 ’-end of each sequence read of the plurality of sequence reads based on the sequence read data.

12. The method of claim 2, wherein a site at which the methylation status is determined is a methylation site.

13. The method of claim 2, wherein performing the statistical analysis comprises: generating a contingency table that tabulates results of the comparisons of the methylation status of the one or more methylation sites located proximal to the 3 ’-end of the respective sequence read to that for other sequence reads of the plurality of sequence reads that overlap the one or more methylation sites; and performing the statistical analysis between two or more factors in the contingency table.

14. The method of claim 13, wherein the contingency table comprises a 2 x 2 contingency table that tabulates the results of the comparisons in terms of two factors (a distal or not distal location of the one or more methylation sites) and two outcomes (a fully unmethylated or not fully unmethylated status of the one or more methylation sites).82MF-364162710Attorney Docket No: 19710201894015. The method of claim 2, wherein the predetermined threshold is determined using a multi-test correction method to adjust the probabilities of the association determined by the statistical analysis to correct for an occurrence of false positives.

16. The method of claim 15, wherein 3’-end hypomethylation bias in the respective sequence read is detected if the determined probability is greater than the predetermined threshold.

17. The method of claim 2, wherein generating a corrected sequence read comprises truncating a 3 ’-end of the respective sequence read.

18. The method of claim 2, wherein the one or more methylation sites comprise one or more CpG dinucleotide sites and / or one or more non-CpG dinucleotide sites.

19. The method of claim 1, wherein the plurality of modified nucleic acid fragments align to one or more genomic intervals of interest and the one or more genomic intervals of interest are selected based on a Tumor and Tissue of Origin (TTOO) classification and / or based on a cancer to be detected.

20. The method of claim 3, wherein the plurality of sequence reads derived from samples from healthy individuals align to one or more genomic intervals of interest that are the same as the one or more genomic intervals of interest to which the plurality of sequence reads obtained from the sample from the subject align.83MF-364162710