Analysis method for cell-free nucleic acids in urine

A targeted panel of genomic regions with bait oligonucleotides enhances the detection of cancer-specific methylation patterns in urine, addressing inefficiencies in existing methods by improving the cost-effectiveness and accuracy of cancer diagnosis.

JP2026504774APending Publication Date: 2026-02-10GRAIL INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025524651
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-01-20
Filing Date
2024-01-19
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing methods for analyzing cell-free nucleic acids in urine, such as DNA methylation profiling for cancer detection, are not cost-effective and inefficient due to low differential methylation signals and abundance issues, particularly when using whole-genome bisulfite sequencing (WGBS).

Method used

A targeted panel of genomic regions with bait oligonucleotides is used to enrich for cancer-specific methylation patterns in cell-free DNA (cfDNA) fragments, combined with methods to concentrate and purify nucleic acids from urine specimens, allowing for efficient sequencing and detection of cancer types and tissue of origin.

Benefits of technology

This approach enables cost-effective, non-invasive early detection of cancer by identifying specific methylation markers, improving the efficiency and accuracy of cancer diagnosis using urine samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026504774000001_ABST
    Figure 2026504774000001_ABST
Patent Text Reader

Abstract

In various aspects, the present disclosure provides methods, compositions, reaction mixtures, kits, and systems for analyzing cell-free nucleic acid molecules (e.g., cfRNA and / or cfDNA) from urine specimens. In some embodiments, the analysis is the analysis of methylation patterns at target genomic regions among cfDNA fragments in urine specimens. In some embodiments, the composition includes multiple different bait oligonucleotides. Methods for detecting various cancer types are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 480,934, filed January 20, 2023, the disclosure of which is hereby incorporated by reference in its entirety. [Background technology]

[0002] Analysis of nucleic acids, such as circulating cell-free nucleic acids (e.g., cell-free DNA (cfDNA) and cell-free RNA (cfRNA)), using next-generation sequencing (NGS) has been recognized as a useful method for characterizing various specimen types. For example, such analysis is useful as a diagnostic tool for the detection and diagnosis of cancer. These analytes can also help improve our fundamental understanding of basic biology.

[0003] DNA methylation plays an important role in regulating gene expression. Aberrant DNA methylation has been implicated in many disease processes, including cancer. DNA methylation profiling using methylation sequencing (e.g., whole-genome bisulfite sequencing (WGBS)) is increasingly recognized as a diagnostic tool useful for cancer detection, diagnosis, and / or monitoring. For example, specific patterns of differentially methylated regions may be useful as molecular markers for various diseases.

[0004] However, WGBS is not ideally suited to assaying products because either the vast majority of genomes do not exhibit differential methylation in cancer, or the local CpG density is too low to provide a robust signal. Only a few percent of genomes are expected to be useful for classification. These issues are compounded when working with specimen types that are less abundant in cell-free nucleic acids, such as urine. Summary of the Invention [Problem to be solved by the invention]

[0005] For at least the above reasons, there remains a need for cost-effective methods and compositions for analyzing cell-free nucleic acid molecules in urine. Various aspects of the present disclosure address this need and provide other advantages as well. [Means for solving the problem]

[0006] Early detection of cancer in a subject is important because it allows for early treatment and thus increases the chances of survival. Targeted detection of cancer-specific methylation patterns using cell-free DNA (cfDNA) fragments can enable early detection of cancer by providing a cost-effective, non-invasive method for obtaining information related to the presence or absence of cancer, the tissue of origin of cancer, or the type of cancer. In this method, rather than sequencing all nucleic acids in a test specimen, also known as "whole genome sequencing," the sequencing depth of the target region can be increased by using a targeted panel of genomic regions. Including methylation markers for several different types of cancer allows for more efficient use of specimens and reagents compared to performing multiple separate assays for different types of cancer. However, it may be advantageous to limit the total coverage of genome sequences in terms of capture, sequencing, and / or computational efficiency.

[0007] To this end, the present description provides cancer assay panels (alternatively referred to as "bait sets") for detecting cancer-specific methylation patterns in targeted genomic regions, along with methods for detecting cancer, cancer types, and / or tissue of origin (TOO) using the cancer assay panels. The methods described herein further include methods for designing probes to efficiently enrich for cfDNA corresponding to or derived from selected genomic regions without pulling down excessive amounts of undesired DNA. Also provided are methods for analyzing cfDNA and other cell-free nucleic acids, particularly in urine specimens.

[0008] In one aspect, provided herein is a method for sequencing a cell-free nucleic acid molecule of interest. In some embodiments, the method includes: (a) treating a urine specimen to inhibit cell lysis; (b) separating the cell-free nucleic acid molecules in the processed urine specimen from cells in the processed urine specimen, thereby producing a purified urine specimen containing the cell-free nucleic acid molecules; (c) concentrating the cell-free nucleic acid molecules in the purified urine specimen by passing at least a portion of the purified urine specimen through a filter, where (i) the concentration produces a filtrate and a residual urine specimen, and (ii) the residual urine specimen contains an increased concentration of cell-free nucleic acid molecules; (d) isolating the cell-free nucleic acid molecules from the residual urine specimen; and (e) sequencing the isolated cell-free nucleic acid molecules.

[0009] In some embodiments, treating the urine specimen to inhibit cell lysis comprises contacting the urine specimen with one or more preservative reagents. In some embodiments, treating comprises treatment with a nuclease inhibitor, a formaldehyde quencher, or both. In some embodiments, treating comprises contacting the urine specimen with a composition comprising (i) imidazolidinyl urea, EDTA, glycine, or a combination thereof; or (ii) sodium azide, EDTA, or a combination thereof. In some embodiments, separating comprises pelleting cells in the treated urine specimen by centrifugation. In some embodiments, the filter is substantially impermeable to the passage of cell-free nucleic acids in the purified urine specimen and substantially permeable to salts. In some embodiments, the filter has a nominal molecular weight cutoff of 10 kD, 5 kD, 3 kD, or less. In some embodiments, the residual urine specimen has a concentration increased by at least 2-fold, at least 5-fold, at least 10-fold, or at least 15-fold compared to the purified urine specimen. In some embodiments, the residual urine specimen has a volume that is at least 50%, at least 75%, or at least 90% smaller than the volume of the processed urine specimen. In some embodiments, the volume of the processed urine specimen is 20 mL, 30 mL, 40 mL, 50 mL, or more. In some embodiments, the method further comprises freezing the residual urine specimen. In some embodiments, the processing is completed within 120, 60, or 30 minutes after collection of the urine specimen; and the optional separating and concentrating are completed within 7 days after collection.

[0010] In some embodiments, the method further comprises amplifying one or more of the isolated cell-free nucleic acid molecules. In some embodiments, the method further comprises capturing the isolated cell-free nucleic acid molecules, or their amplification products, by hybridization to a bait oligonucleotide. In some embodiments, the method further comprises separating cell-free nucleic acid molecules bound to the bait from unbound cell-free nucleic acid molecules.

[0011] In some embodiments, each bait oligonucleotide hybridizes to a target genomic region that is differentially methylated in cancer specimens compared to non-cancer specimens. In some embodiments, differential methylation includes at least 80% of the CpG sites in the target genomic region being methylated or unmethylated. In some embodiments, the cancer is bladder cancer, prostate cancer, or renal cancer. In some embodiments, each bait oligonucleotide hybridizes to a target genomic region that includes at least five methylation sites. In some embodiments, each bait oligonucleotide hybridizes to a target genomic region that includes a target sequence of a gene selected from Table 1, and wherein the target sequence is at least 25, at least 35, or at least 45 nucleotides in length. In some embodiments, the target genomic region includes a target sequence of a gene selected from TWIST1, EOMES, HOXA9, POU4F2, and ZNF154. In some embodiments, the bait oligonucleotides collectively hybridize to target sequences from at least 10 genes in Table 1. In some embodiments, the bait oligonucleotides collectively hybridize to target sequences from: (a) a gene in Table 2 or Table 3; (b) a gene in Table 4; or (c) a gene in Table 5.

[0012] In some embodiments, the cell-free nucleic acid molecule comprises cell-free DNA (cfDNA). In some embodiments, the method further comprises deaminating the cfDNA isolated in step (d) to produce a converted cfDNA molecule; optionally, the deaminating comprises treatment with cytosine deaminase or bisulfite.

[0013] In some embodiments, the method further comprises diagnosing the cancer in the subject. In some embodiments, the cancer is bladder cancer, prostate cancer, or renal cancer. In some embodiments, the method further comprises treating the cancer in the subject. In some embodiments, treating comprises surgical resection, radiation therapy, chemotherapy, and / or immunotherapy.

[0014] In one aspect, the present specification provides a method for detecting cancer cell in object.In some embodiments, this method includes: (a) capturing the cell-free DNA (cfDNA) fragments converted from object's urine specimen or its amplification product, wherein (i) bait oligonucleotide composition comprises a plurality of different bait oligonucleotides; (ii) each bait oligonucleotide of the plurality of different bait oligonucleotides hybridizes with the target sequence of the gene selected from table 1, and the length of the target sequence is at least 25 nucleotides; (b) separating the DNA that binds to bait and unbound DNA; (c) sequencing the separated DNA to generate sequencing reads; and (d) using a trained classifier to detect cancer cell.

[0015] In some embodiments, the bait oligonucleotide is at least 45 nucleotides in length. In some embodiments, the trained classifier detects a number of sequencing reads above a threshold for one or more target sequences identified as hypermethylated and / or hypomethylated in cfDNA fragments. In some embodiments, the trained classifier distinguishes subjects with cancer from subjects without cancer with a predetermined specificity. In some embodiments, the classifier is a mixed model classifier. In some embodiments, the predetermined specificity is 0.900 or greater. In some embodiments, application of the trained classifier further comprises a sensitivity of 30% or greater. In some embodiments, the bait oligonucleotides collectively hybridize to target sequences from at least 10 genes in Table 1. In some embodiments, at least one of the bait oligonucleotides hybridizes to a target sequence of a gene selected from TWIST1, EOMES, HOXA9, POU4F2, and ZNF154. In some embodiments, (a) the cancer cells are bladder cancer cells and the bait oligonucleotides collectively hybridize to target sequences from the genes in Table 2 or Table 3; (b) the cancer cells are prostate cancer cells and the bait oligonucleotides collectively hybridize to target sequences from the genes in Table 4; or (c) the cancer cells are renal cancer cells and the bait oligonucleotides collectively hybridize to target sequences from the genes in Table 5.

[0016] In some embodiments, the converted cfDNA molecule comprises cfDNA treated with cytosine deaminase or bisulfite. In some embodiments, each bait oligonucleotide is conjugated to a solid surface or a non-nucleotide affinity moiety. In some embodiments, differential methylation comprises hypermethylation in cancer specimens compared to non-cancer specimens. In some embodiments, each target genomic region comprises at least five methylation sites.

[0017] In some embodiments, the method further comprises obtaining the converted cfDNA fragments or their amplification products, wherein obtaining further comprises (i) treating the urine specimen to inhibit cell lysis; (ii) separating the cfDNA fragments in the processed urine specimen from cells in the processed urine specimen, thereby producing a purified urine specimen containing the cfDNA fragments; (iii) concentrating the cfDNA fragments in the purified urine specimen by passing at least a portion of the purified urine specimen through a filter, whereby a filtrate and a residual urine specimen are produced, and the residual urine specimen has an increased concentration of the cfDNA fragments; and (iv) isolating the cfDNA fragments from the residual urine specimen. In some embodiments, the method further comprises (v) amplifying one or more of the isolated cfDNA fragments.

[0018] In some embodiments, treating the urine specimen to inhibit cell lysis comprises contacting the urine specimen with one or more preservative reagents. In some embodiments, treating comprises treatment with a nuclease inhibitor, a formaldehyde quencher, or both. In some embodiments, treating comprises contacting the urine specimen with a composition comprising (a) imidazolidinyl urea, EDTA, glycine, or a combination thereof; or (b) sodium azide, EDTA, or a combination thereof. In some embodiments, separating comprises pelleting cells in the treated urine specimen by centrifugation. In some embodiments, the filter is substantially impermeable to the passage of cell-free nucleic acids in the purified urine specimen and substantially permeable to salts. In some embodiments, the filter has a nominal molecular weight cutoff of 10 kD, 5 kD, 3 kD, or less. In some embodiments, the residual urine specimen has a concentration increased by at least 2-fold, at least 5-fold, at least 10-fold, or at least 15-fold compared to the purified urine specimen. In some embodiments, the residual urine specimen has a volume that is at least 50%, at least 75%, or at least 90% smaller than the volume of the processed urine specimen. In some embodiments, the volume of the processed urine specimen is 20 mL, 30 mL, 40 mL, 50 mL, or more. In some embodiments, the method further comprises freezing the residual urine specimen. In some embodiments, the processing is completed within 120, 60, or 30 minutes after collection of the urine specimen; and the optional separating and concentrating are completed within 7 days after collection.

[0019] In some embodiments, the method further comprises diagnosing the cancer in the subject. In some embodiments, the cancer is bladder cancer, prostate cancer, or renal cancer. In some embodiments, the method further comprises treating the cancer in the subject. In some embodiments, treating comprises surgical resection, radiation therapy, chemotherapy, and / or immunotherapy.

[0020] In one aspect, the present specification provides a method for treating cancer in an object.In some embodiments, the method includes selecting an object that has cancer or is at high risk of developing cancer, and administering a treatment to the object, wherein (a) selecting includes identifying the object as a source of urine cell-free DNA (cfDNA) specimen, and the target genomic region comprises one or more target genomic regions that are methylated differently above the threshold level for the existence of cancer; (b) the one or more target genomic regions comprise one or more target sequences of one or more genes selected from Table 1; (c) each target sequence is at least 25 nucleotides long; (d) the cancer is bladder cancer, prostate cancer, or renal cancer; and (e) the treatment includes surgical resection, radiation therapy, chemotherapy, immunotherapy, or any combination thereof.

[0021] In some embodiments, the threshold level for the presence of cancer is the level in a reference sample from a subject with cancer. In some embodiments, the threshold level for the presence of cancer is determined by a classifier trained on sequencing reads of converted DNA from a subject with cancer. In some embodiments, the one or more target genomic regions comprise target sequences from at least 10 genes in Table 1. In some embodiments, the one or more target genomic regions comprise target sequences of genes selected from TWIST1, EOMES, HOXA9, POU4F2, and ZNF154. In some embodiments, (a) the cancer is bladder cancer, and the one or more target genomic regions comprise one or more target sequences of one or more genes selected from Table 2 or Table 3; (b) the cancer is prostate cancer, and the one or more target genomic regions comprise one or more target sequences of one or more genes selected from Table 4; or (c) the cancer is renal cancer, and the one or more target genomic regions comprise one or more target sequences of one or more genes selected from Table 5. In some embodiments, each target genomic region comprises at least five methylation sites.

[0022] In one aspect, provided herein is a composition comprising a plurality of different bait oligonucleotides. In some embodiments, (a) the bait oligonucleotides hybridize to converted DNA molecules derived from one or more target genomic regions; (b) the one or more target genomic regions comprise one or more target sequences of one or more genes selected from Table 1; (c) the one or more target genomic regions are differentially methylated in cancer; and (d) each bait oligonucleotide comprises a sequence of at least 25 nucleotides in length that hybridizes to one of the target sequences. In some embodiments, the one or more target genomic regions comprise target sequences from at least 10 genes selected from Table 1. In some embodiments, the one or more target genomic regions comprise target sequences of genes selected from TWIST1, EOMES, HOXA9, POU4F2, and ZNF154. In some embodiments, the one or more target genomic regions comprise target sequences from (a) a gene in Table 2 or Table 3; (b) a gene in Table 4; or (c) a gene in Table 5. In some embodiments, each target genomic region comprises at least five methylation sites. In some embodiments, differential methylation includes at least 80% of the CpG sites in the target genomic region being methylated or unmethylated.

[0023] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.

[0024] The novel features of the present disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings, as follows: [Brief explanation of the drawings]

[0025] [Figure 1]1 illustrates a urine specimen processing workflow, according to one embodiment. [Figure 2A] 1 illustrates an exemplary process for biopsy-free tumor fraction estimation from urinary cell-free DNA. [Figure 2B] Illustrated is an exemplary methylation mutation at CpGs 1528723-1528728. The methylation mutations are depicted as a series of consecutive CpGs and methylation status that distinguish cancer-derived cfDNA from non-cancer cfDNA. [Figure 3] We present data showing that urological cancers (bladder or urothelial carcinoma, renal cancer, and prostate cancer) were associated with decreased detection sensitivity in the Circulating Cell-Free Genome Atlas Study (CCGA) patient cohort. Paired bars represent early-stage (left) and late-stage (right) results, respectively. [Figure 4] 1 illustrates the process of determining methylation mutations from urinary cfDNA, according to one embodiment. [Figure 5A] We provide data showing the cfDNA fragment size distribution and yield when specimens are treated by adding a preservative immediately after collection. [Figure 5B] We provide data showing the effect of delaying the addition of preservative to urine specimens by 1 hour on cfDNA fragment size distribution and yield. [Figure 6] We provide data on a cancer-specific methylation signature of four genes detected in pre- and post-operative urinary cfDNA from subjects with stage I high-grade non-muscle-invasive bladder cancer. [Figure 7] Data are provided comparing estimated tumor fraction in urine and plasma specimens from subjects with (annotated plotted points) or without (unannotated plotted points) any stage of bladder cancer. [Figure 8] Data are provided comparing estimated tumor fraction in urine and plasma specimens from subjects with and without stage III or IV prostate cancer. Data points above the "*" correspond to results from specimens from only cancer subjects. Data points below the "*" correspond to results from specimens from cancer and non-cancer subjects. [Figure 9]Data are provided comparing estimated tumor fraction in urine and plasma specimens from subjects with (annotated data points) or without (unannotated data points) renal cancer. [Figure 10A] We provide data on the classification performance of urinary cfDNA from subjects with bladder cancer based on a subset of 15 genomic regions, two of which are represented by either of two genes within the region, demonstrating an area under the curve (AUC) of 0.99. [Figure 10B]

[0023] Figure 1 provides data on the classification performance of urinary cfDNA from subjects with bladder cancer based on genomic regions within a single gene (TWIST1), showing an AUC of 0.86. Shading represents the 95% confidence interval. [Figure 10C] 1 provides data on the classification performance of urinary cfDNA from subjects with renal cancer based on a subset of genomic regions, showing an AUC of 0.52. Shading represents the 95% confidence interval. [Figure 10D]

[0023] Figure 1 provides data on the classification performance of urinary cfDNA from subjects with prostate cancer based on a subset of genomic regions, showing an AUC of 0.82. Shading represents the 95% confidence interval. [Figure 11A] Illustrates a 2x tiling probe design in accordance with one embodiment, in which three probes target a small target region, where each base in the target region (dotted rectangular box) is covered by at least two probes. [Figure 11B] Illustrates one embodiment of a 2x tiling probe design with four or more probes targeting a larger target region, where each base in the target region (dotted rectangular box) is covered by at least two probes. [Figure 11C] 1 illustrates probe designs targeting hypomethylated and / or hypermethylated fragments in a genomic region, according to certain embodiments. [Figure 11D] 1 illustrates a section of a genome containing three target genomic regions, with probes designed to hybridize to each target and its adjacent sequences in a 2x tiling configuration, according to one embodiment. [Figure 11E] Illustrated is a single pair of probes hybridizing within the same target genomic region, each containing an overlapping sequence and a non-overlapping sequence. The overlapping sequences are complementary to the same target sequence. As shown, the non-overlapping sequences are each complementary to a different sequence within the target genomic region at a different end relative to the sequence to which the overlapping sequence is complementary. [Figure 12A] 1 is a flowchart illustrating a process for creating a data structure for a control group, according to one embodiment. [Figure 12B] 12B is a flowchart depicting additional steps for validating the data structure for the control group of FIG. 12A, according to one embodiment. [Figure 13] 1 is a flowchart depicting a process for selecting genomic regions for designing probes for a cancer assay panel, according to one embodiment. [Figure 14] 1 is an illustration of an exemplary p-value score calculation according to an embodiment. [Figure 15A] 1 is a flowchart depicting a process for training a classifier based on hypomethylated and hypermethylated fragments indicative of cancer, according to an embodiment. [Figure 15B] 1 is a flowchart illustrating a process for identifying fragments indicative of cancer as determined by a probabilistic model, according to one embodiment. [Figure 16A] 1 is a flowchart illustrating a process for sequencing fragments of cell-free (cf) DNA, according to one embodiment. [Figure 16B] 1 is an illustration of a process for sequencing fragments of cell-free (cf) DNA to obtain a methylation state vector, according to one embodiment. [Figure 17A] 1 illustrates a flow chart of an apparatus for sequencing a nucleic acid sample, according to one embodiment. [Figure 17B] 1 illustrates an analytics system for analyzing the methylation status of cfDNA, according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0026] Before describing the present invention in further detail, it is to be understood that this invention is not limited to particular embodiments described, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting, since the scope of the present invention will be limited only by the appended claims.

[0027] Unless otherwise defined herein, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. To facilitate understanding of certain terms used frequently herein, the following definitions are provided:

[0028] Where a range of values ​​is provided, it is understood that each intervening value between the upper and lower limits of that range, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, and any other stated or intervening value in that stated range, is encompassed within the scope of the invention. The upper and lower limits of these smaller ranges may independently be included in those smaller ranges that are encompassed within the scope of the invention, subject to any specifically excluded limit in the stated range.

[0029] As used herein, the term "about" refers to a range of values, inclusive of the specified value, that one of ordinary skill in the art would consider to be reasonably similar to the specified value. In embodiments, about refers to within a range of standard deviation, using measurements generally accepted in the art. In embodiments, about refers to a range extending from the specified value to ±10%. In embodiments, about includes the specified value.

[0030] The term "methylation," as used herein, refers to the process of adding a methyl group to a DNA molecule. For example, a hydrogen atom on the pyrimidine ring of a cytosine base can be converted to a methyl group, resulting in the formation of 5-methylcytosine. The term also refers to the process of adding a hydroxymethyl group to a DNA molecule, for example, by oxidation of a methyl group on the pyrimidine ring of a cytosine base. Methylation and hydroxymethylation tend to occur at cytosine and guanine dinucleotides, referred to herein as "CpG sites." The principles described herein may also be applicable to detecting methylation in non-CpG contexts, including non-cytosine methylation. In such embodiments, the wet-lab assay used to detect methylation may differ from any described herein. Furthermore, a methylation state vector may include elements that are generally vectors of methylated or unmethylated sites (even if such sites are not specifically CpG sites).

[0031] The term "methylation" can also refer to the methylation state of a CpG site. A CpG site with a 5-methylcytosine moiety is methylated. A CpG site with a hydrogen atom on the pyrimidine ring of the cytosine base is unmethylated.

[0032] The term "methylation site," as used herein, refers to a region of a DNA molecule where a methyl group can be added. While "CpG" sites are the most common methylation sites, methylation sites are not limited to CpG sites. For example, DNA methylation can occur at cytosine in the form of CHG and CHH (where H is adenine, cytosine, or thymine). Cytosine methylation in the form of 5-hydroxymethylcytosine can also be determined and characterized using the methods and procedures disclosed herein (see, e.g., U.S. Patent Application Publication Nos. 20110236894A1 and 20110301045A1, which are incorporated herein by reference).

[0033] The term "CpG site," as used herein, refers to a stretch of a DNA molecule in its 5' to 3' linear sequence of bases in which a cytosine nucleotide is followed by a guanine nucleotide. "CpG" is an abbreviation for 5'-C-phosphate-G-3', in which the cytosine and guanine are separated by a single phosphate group. Methylation of the cytosine in a CpG dinucleotide can result in the formation of 5-methylcytosine.

[0034] In some embodiments, the oligonucleotide probes described herein comprise one or more CpG detection sites. The term "CpG detection site" as used herein refers to a region in the probe that is configured to hybridize to a CpG site in a target DNA molecule. The CpG site on the target DNA molecule can comprise a cytosine and a guanine separated by a single phosphate group, where the cytosine is either methylated or unmethylated. The CpG site on the target DNA molecule can comprise a uracil and a guanine separated by a single phosphate group, where the uracil is generated by the conversion of an unmethylated cytosine.

[0035] The term "UpG" is an abbreviation for 5'-U-phosphate-G-3', where uracil and guanine are separated by only one phosphate group. UpG can be generated by converting unmethylated cytosine to uracil by bisulfite treatment of DNA. Cytosine may also be converted to uracil by other methods, such as chemical modification, synthesis, or enzymatic conversion.

[0036] The terms "hypomethylated" or "hypermethylated," as used herein, refer to the methylation state of a DNA molecule having multiple CpG sites (e.g., more than 2, 3, 4, 5, 6, 7, 8, 9, 10, etc.) in which a high percentage of CpG sites (e.g., greater than 80%, 85%, 90%, or 95%, or any other percentage within the range of 50% to 100%, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97.5% or more, 98% or more, 99% or more, 99.9% or more, or any other numerical percentage within the range of 50% to 100%, the ranges provided herein including the 50% and 100% range limits) are unmethylated or methylated, respectively. For example, a "hypomethylated" nucleic acid fragment, e.g., a cfDNA fragment, can be a fragment having a certain number of CpG sites, e.g., 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 9 or more, 10 or more, in which a certain percentage, e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, or 95% or more, or 97.5% or more, 98% or more, 99% or more, 99.9% or more of the CpG sites are unmethylated. Similarly, a "hypermethylated" nucleic acid fragment, e.g., a cfDNA fragment, can be a fragment having a certain number of CpG sites, e.g., 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 9 or more, 10 or more, where a certain percentage, e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, or 95% or more, or 97.5% or more, 98% or more, 99% or more, 99.9% or more of the CpG sites are methylated. In some embodiments, a hypomethylated DNA molecule contains multiple CpG sites, at least 80% of which are unmethylated. In some embodiments, a hypermethylated DNA molecule contains multiple CpG sites, at least 80% of which are methylated.

[0037] The term "methylation state vector" or "methylation status vector," as used herein, refers to a vector that contains multiple elements, each element indicating the methylation state of a methylation site in a DNA molecule that contains multiple methylation sites, in the order in which they appear 5' to 3' in the DNA molecule. For example, <M x ,M x+1 ,M x+2 >, <M x ,M x+1 ,U x+2 >,..., x ,U x+1 ,U x+2 > may be a methylation vector of a DNA molecule containing three methylation sites, where M represents a methylation site that is methylated and U represents an unmethylated methylation site.

[0038] ​The terms "abnormal methylation pattern" and "abnormal methylation pattern," as used herein, refer to a methylation pattern of a nucleic acid molecule, e.g., a DNA molecule such as cfDNA, or a methylation state vector, that is found and / or expected to be found in a sample at a frequency lower than would be found in a healthy, e.g., non-cancer, sample. In various embodiments, such a methylation pattern is found and / or expected to be found in a sample at a frequency lower than a threshold for a healthy, e.g., non-cancer, sample. Thus, for example, the terms "abnormal methylation" and "abnormal methylation" as used herein refer to a nucleic acid molecule, e.g., a DNA molecule such as cfDNA, or a methylation state vector, that exhibits an abnormal methylation pattern. Aspects of the subject disclosure that exhibit differential methylation can, in some variations, include aspects that are not normal methylation. Whether an aspect is differentially methylated can also be used in relation to the health of the subject from whom the subject sample was derived, as an indicator of healthy, e.g., non-cancer, as opposed to diseased, e.g., cancer, characteristics. In one embodiment provided herein, the observation and / or expectation of a specific methylation state vector in a healthy control group, including healthy individuals, is represented by a p-value. In various aspects, a low p-value score corresponds to a methylation state vector that is relatively less expected compared to other methylation state vectors in a sample from a healthy individual, such as an individual in a healthy control group. In some variations, a high p-value score corresponds to a methylation state vector that is relatively more expected compared to other methylation state vectors observed in a sample from a healthy individual, such as an individual in a healthy control group. In various embodiments, a methylation state vector having an abnormal / abnormal methylation pattern is a methylation state vector that has a p-value at and / or below a threshold (e.g., 0.1, 0.01, 0.001, 0.0001, etc.), such as a threshold corresponding to a healthy sample, e.g., a non-cancerous sample.In various embodiments, the methods include associating a methylation state vector from a specimen having a p-value at and / or below a threshold (e.g., 0.1 or less, 0.01 or less, 0.001 or less, 0.0001 or less, etc.) with a determination that the specimen is not a healthy specimen, e.g., a specimen from a subject with cancer. In various embodiments, the threshold is applied as a filter in that a lower applied threshold (e.g., 0.001, 0.0001, etc.) is associated with a higher expectation of the methylation state vector from the specimen being a specimen from a non-healthy, e.g., cancerous, individual. Various methods can be used to calculate the p-value or expectation of a methylation pattern or methylation state vector. An exemplary method provided herein involves the use of Markov chain probabilities, which assume that the methylation state of a CpG site depends on the methylation state of neighboring CpG sites. Another method provided herein calculates the expected probability of observing a specific methylation state vector in a healthy individual by utilizing a mixture model including multiple mixture components, each mixture component being an independent site model in which methylation at each CpG site is assumed to be independent of the methylation state at other CpG sites. In some variations, the subject method includes determining whether a nucleic acid molecule, e.g., a DNA molecule, or a methylation state vector is abnormally methylated. In various embodiments of the method, the generated p-value is compared, e.g., by an analytics system, to a threshold to identify vectors, e.g., nucleic acid fragments such as cfDNA, that are abnormally methylated compared to a control group, e.g., a group associated with one or more healthy samples, e.g., non-cancer samples. Additionally, abnormal methylation, e.g., cfDNA methylation, can be hypermethylated and / or hypomethylated, both of which can be indicative of an unhealthy state, e.g., a cancerous state. Thus, the methods include determining a healthy or diseased state, e.g., a non-cancer or cancer state, based at least in part on a p-value, such as a relatively low p-value, such as a p-value below a threshold, where the p-value may, in various embodiments, be indicative of abnormal methylation, e.g., hypermethylation and / or hypomethylation.A low p-value, e.g., a p-value at or below a threshold value (e.g., 0.1, 0.01, 0.001, 0.0001, etc.), can be indicative of abnormal methylation, e.g., hypermethylation and / or hypomethylation, in the specimen. In various embodiments, the methods include determining a healthy or diseased state, e.g., non-cancer or cancer state, of the specimen based on nucleic acids, e.g., nucleic acid fragments, or methylation vectors from the specimen having a low p-value, e.g., 0.1, 0.01, or 0.001 or less, and being both hypermethylated and hypomethylated, or being hypermethylated or hypomethylated. In various aspects, the methods include determining a healthy or diseased state, e.g., non-cancer or cancer state, of the specimen based, at least in part, on whether nucleic acids, e.g., nucleic acid fragments, or methylation vectors from the specimen are both hypermethylated and hypomethylated. In some variations, determining that a vector, e.g., a sample fragment, is aberrantly methylated based on the generated p-value score includes determining whether the generated score for the vector is below a threshold score, where the threshold score is a confidence level that the vector is aberrantly methylated.

[0039] The term "cancerous specimen" as used herein refers to a specimen containing genomic DNA from an individual diagnosed with cancer. Genomic DNA can be, but is not limited to, cfDNA fragments or chromosomal DNA from a subject with cancer. Genomic DNA can be sequenced, and its methylation status can be determined by various methods, such as bisulfite sequencing. When a genomic sequence is obtained from a public database (e.g., The Cancer Genome Atlas (TCGA)) or experimentally obtained by sequencing the genome of an individual diagnosed with cancer, a cancerous specimen can refer to the genomic DNA or cfDNA fragments containing the genomic sequence. The plural form of the term "cancerous specimen" refers to a specimen containing genomic DNA from multiple individuals, each of whom has been diagnosed with cancer. In various embodiments, cancerous specimens are used from more than 100, 300, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 40,000, 50,000, or more individuals diagnosed with cancer.

[0040] The term "non-cancerous specimen" or "healthy specimen," as used herein, refers to a specimen containing genomic DNA from a healthy individual or an individual not diagnosed with cancer. Genomic DNA can be, but is not limited to, cfDNA fragments or chromosomal DNA from a subject without cancer. Genomic DNA can be sequenced, and its methylation status can be determined by various methods, for example, bisulfite sequencing. When a genomic sequence is obtained from a public database (e.g., The Cancer Genome Atlas (TCGA)) or experimentally obtained by sequencing the genome of an individual without cancer, a non-cancerous specimen can refer to genomic DNA or cfDNA fragments containing the genomic sequence. The plural form of the term "non-cancerous specimen" refers to a specimen containing genomic DNA from multiple individuals, each of whom is cancer-free. In various embodiments, cancerous specimens from greater than 100, 300, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 40,000, 50,000, or more cancer-free individuals are used. In various embodiments, cancerous specimens from greater than 100, greater than 300, greater than 500, greater than 1,000, greater than 2,000, greater than 5,000, greater than 10,000, greater than 20,000, greater than 40,000, or greater than 50,000 cancer-free individuals are used.

[0041] The term "training sample," as used herein, refers to a sample used to train a classifier described herein and / or select one or more genomic regions for cancer detection or for detecting the tissue of origin or cancer cell type of cancer. A training sample may include genomic DNA or modifications thereof from one or more healthy subjects and one or more subjects with a disease state (e.g., cancer, a specific type of cancer, a specific stage of cancer, etc.). Genomic DNA may be, but is not limited to, cfDNA fragments or chromosomal DNA. Genomic DNA can be sequenced, and its methylation status can be determined by various methods, such as bisulfite sequencing. When a genomic sequence is obtained from a public database (e.g., The Cancer Genome Atlas (TCGA)) or experimentally obtained by sequencing an individual's genome, a training sample may refer to the genomic DNA or cfDNA fragments containing the genomic sequence.

[0042] The term "test specimen," as used herein, refers to a specimen from a subject whose health condition is being, has been, or will be tested using the classifiers and / or assay panels described herein. The test specimen may contain genomic DNA or modifications thereof. Genomic DNA may be, but is not limited to, cfDNA fragments or chromosomal DNA.

[0043] The term "target genomic region" as used herein refers to a region in a genome selected for analysis in a test specimen. An assay panel is generated using oligonucleotide probes designed to hybridize to (and optionally pull down) nucleic acid fragments derived from the target genomic region or fragments thereof. Oligonucleotide probes directed to a target region are also referred to herein as "bait oligonucleotides." Nucleic acid fragments derived from a target genomic region refer to nucleic acid fragments generated by degradation, cleavage, bisulfite conversion, or other processing of DNA from the target genomic region. In some embodiments, multiple different bait oligonucleotides are designed to hybridize across a single target genomic region (e.g., overlapping probes tiled across a target genomic region). Generally, when referring to multiple target genomic regions, no target genomic region in the plurality is entirely contained within another target genomic region. Different target genomic regions in a plurality of target genomic regions may overlap, but will have at least different ends. In some embodiments, each target genomic region in a plurality of target genomic regions is separate from and does not overlap with any other target genomic region in the plurality. In some embodiments, the target genomic region comprises a target sequence of a gene. In this context, the term "gene" encompasses the sequence encoding that gene, and optionally flanking sequences (e.g., regulatory sequences such as promoters, enhancers, and untranslated regions), and internal non-coding sequences (e.g., introns). In some embodiments, a gene is defined by sequences encoding transcription start and stop sites, and a further 5000 nucleotides from each of these ends. A given gene may contain multiple target sequences.

[0044] A target genomic region can be described according to its chromosomal location, such as relative to a particular reference sequence (e.g., human reference genome GRCh37 / hg19). Chromosomal DNA is double-stranded, so a target genomic region contains two DNA strands: one with a sequence corresponding to a given reference sequence, and the second with its reverse complement. Probes can be designed to hybridize to one or both sequences. Optionally, the probe hybridizes to a converted sequence obtained, for example, by treatment with sodium bisulfite.

[0045] The term "off-target genomic region" as used herein refers to a region in the genome that was not selected for analysis in the test specimen, but that has sufficient homology with the target genomic region to be potentially pulled down by a probe designed to target the target genomic region. In one embodiment, the off-target genomic region is a genomic region that aligns with the probe at least 90% match rate along at least 45 bp.

[0046] The terms "converted DNA molecule," "converted cfDNA molecule," and "modified fragment obtained by processing cfDNA molecules" refer to DNA molecules obtained by processing DNA or cfDNA molecules in a sample to distinguish between methylated and unmethylated nucleotides in the DNA or cfDNA molecules. For example, in some embodiments, the sample may be treated with bisulfite ions (e.g., using sodium bisulfite), thereby converting unmethylated cytosine ("C") to uracil ("U"). In another embodiment, the conversion of unmethylated cytosine to uracil is achieved using an enzymatic conversion reaction, for example, using cytidine deaminase (e.g., APOBEC). After processing, the converted DNA molecule or cfDNA molecule contains an additional uracil that was not present in the original cfDNA sample. When a DNA strand containing uracil is replicated by DNA polymerase, an adenine is added to the nascent complementary strand instead of the guanine that is normally added as a complement to cytosine or methylcytosine.

[0047] In general, the terms "cell-free," "circulating," and "extracellular," when applied to polynucleotides (e.g., "cell-free DNA" or "cfDNA"), are used interchangeably to refer to polynucleotides or portions thereof present in a specimen from a subject that can be isolated or otherwise manipulated without applying a lysis step to the specimen as originally collected (e.g., as in lysis for extraction from cells or viruses). Cell-free polynucleotides are thus "free," meaning that they are not enclosed in, or free from, the cells or viruses from which they originate, even before the specimen from the subject is collected. Cell-free polynucleotides can arise as a byproduct of cell death (e.g., apoptosis or necrosis) or cell shedding, resulting in the release of polynucleotides into surrounding body fluids or the circulation. Thus, cell-free nucleic acids can be isolated from the noncellular fraction of blood (e.g., serum or plasma), other body fluids (e.g., urine), or the noncellular fraction of other types of specimens. Non-limiting examples of bodily fluids that may be used in conjunction with the embodiments disclosed herein include mucus, blood, plasma, serum, serum derivatives, synovial fluid, lymphatic fluid, bile, phlegm, saliva, sweat, tears, sputum, amniotic fluid, menstrual fluid, vaginal fluid, semen, urine, cerebrospinal fluid (CSF) such as lumbar or ventricular CSF, gastric juice, a liquid specimen containing one or more materials from a nasal, pharyngeal, or oral swab, a liquid specimen containing one or more materials from a lavage procedure such as a peritoneal, gastric, breast, or breast lavage procedure, etc. In some embodiments, cfDNA refers to deoxyribonucleic acid molecules circulating in a subject's body (e.g., bloodstream) and may originate from one or more healthy cells and / or one or more cancer cells. In some embodiments, the cfDNA is cfDNA from a urine specimen. In some embodiments, the compositions (e.g., assay panels) and methods disclosed herein in the context of urinary cfDNA may be applied to or adapted for use with other specimen types (e.g., cfDNA from other bodily fluids such as blood, serum, or plasma).

[0048] The term "circulating tumor DNA" or "ctDNA" refers to nucleic acid fragments originating from tumor cells that may be released into an individual's bloodstream as a result of biological processes such as apoptosis or necrosis of dying cells, or that may be actively released by viable tumor cells.

[0049] The term "fragment" as used herein may refer to a fragment of a nucleic acid molecule. For example, in one embodiment, a fragment may refer to a cfDNA molecule in a blood or plasma sample, or a cfDNA molecule extracted from a blood or plasma sample. The amplification product of a cfDNA molecule may also be referred to as a "fragment." In another embodiment, the term "fragment" refers to a sequence read or a set of sequence reads that have been processed for subsequent analysis (e.g., for machine learning-based classification) as described herein. For example, raw sequence reads can be aligned with a reference genome, and paired-end sequence reads can be matched and assembled into longer fragments for subsequent analysis.

[0050] The terms "individual" and "subject" refer to a human individual. The term "healthy individual" refers to an individual who is suspected to be cancer or disease-free. In some embodiments, a subject is an individual whose DNA is under analysis. For example, a subject can be a test subject whose DNA is under evaluation using a targeted panel as described herein to assess whether the person has cancer or another disease. In some embodiments, a subject is part of a control group (also referred to as "reference subjects") known to have (or not have) cancer or another disease. Control groups and cancer / disease groups can be used to aid in the design or validation of targeted panels.

[0051] The term "sequence read," as used herein, refers to a string of nucleotides as determined for part or all of a nucleic acid molecule by a nucleic acid sequencing process. A sequence read can be a short (e.g., 20-150) string of nucleotides sequenced from a nucleic acid fragment, a short string of nucleotides at one or both ends of a nucleic acid fragment, or the sequencing of an entire nucleic acid fragment present in a biological specimen. Sequence reads can be obtained through various methods provided herein or by other methods known in the art.

[0052] The term "sequencing depth," as used herein, refers to the count of the number of times a given target nucleic acid in a sample has been sequenced (e.g., the count of sequence reads in a given target region), or the average number of times that is actually sequenced or expected to be sequenced based on the amount of nucleic acid subjected to sequencing and the total read length generated by a given sequencing process (e.g., the average read depth for all sequenced regions from a given sequencing run). As sequencing depth increases, the required amount of nucleic acid needed to determine a disease state (e.g., cancer or tissue of origin of the cancer) may decrease.

[0053] The term "tissue of origin" or "TOO," as used herein, refers to the organ, group of organs, body region, or cell type from which cancer develops or originates. Identifying the tissue of origin or cancer cell type typically allows for identification of optimal next steps for further diagnosis, staging, and treatment decisions in the cancer continuum of care. For example, a cancer originating in cells of the bladder may be identifiable as a bladder cell based on one or more markers associated with bladder cells, even after metastasis to another tissue, such as the kidney. In the case of metastasis, "cancer tissue" refers to the metastatic growth of cells of a particular TOO that are characteristically different from the cells of the tissue to which the cancer has metastasized. Thus, metastatic cancer tissue may be located adjacent to or within healthy tissue at the metastatic site.

[0054] "Treating" or "treatment," as used herein, includes any approach to obtaining a beneficial or desired result for a subject's condition, including a clinical result. Beneficial or desired clinical results can include, but are not limited to, alleviation or amelioration of one or more symptoms or conditions, whether partial or complete, and whether detectable or undetectable, reduction in the extent of the disease, stabilization (i.e., not worsening) of the disease state, prevention of the spread or development of the disease, delay or slowing of disease progression, improvement or palliation of the disease state, reduction in recurrence of the disease, and remission. In other words, "treatment," as used herein, includes any cure, amelioration, or prevention of the disease. Treatment may prevent the disease from occurring; inhibit the spread of the disease; alleviate the symptoms of the disease, completely or partially eliminate the underlying cause of the disease, shorten the duration of the disease, or a combination of these.

[0055] "Treating" and "treatment," as used herein, include prophylactic treatment. Treatment methods include administering a therapeutically effective amount of an active agent to a subject. The administering step may consist of a single administration or may include a series of administrations. The length of treatment depends on various factors, such as the severity of the condition, the age of the patient, the concentration of the active agent, the activity of the composition used in the treatment, or a combination thereof. It will also be understood that the effective dosage of an agent used for treatment or prevention may increase or decrease over the course of a specific treatment or prevention regime. Modifications in dosage may be evident from standard diagnostic assays known in the art. In some cases, chronic administration may be required. For example, a composition is administered to a subject in an amount and for a period sufficient to treat the patient. In embodiments, the treating or treatment is not prophylactic treatment.

[0056] The term "preventing," when referring to a disease or condition in a subject, refers to a reduction in the occurrence of one or more corresponding symptoms in a subject. As noted above, prevention may be complete (no detectable symptoms) or may be partial, with fewer symptoms observed and / or a reduced incidence compared to those that would be expected to occur in the absence of treatment.

[0057] ;Semustine;Senogenesis-derived inhibitor 1;Sense oligonucleotide;Signal transduction inhibitor;Signal transduction modulator;Single-chain antigen-binding protein;Schizofuran;Sobzoxane;Borocaptate sodium;Sodium phenylacetate;Sorberol;Somatomedin-binding protein;Sonermin;Sparfosic acid;Spicamycin D;Spiromustine;Splenopentin;Spongistatin 1;Squalamine;Stem cell inhibitor;Stem cell division inhibitor;Stipiamide;Stomatin Lomelysin inhibitors; Sulfinosine; Superactive vasoactive intestinal peptide antagonists; Suragista; Suramin; Swainsonine; Synthetic glycosaminoglycans; Talimustine; Tamoxifen methiodide; Tauromustine; Tazarotene; Tecogalan sodium; Tegafur; Terlapyrylium; Telomerase inhibitors; Temoporfin; Temozolomide; Teniposide; Tetrachlorodecaoxide; Tetrazomine; Taliblastine; Thiocoraline; Thrombopoietin thrombin; thrombopoietin mimetics; thymalfasin; thymopoietin receptor agonists; thymotrin; thyroid-stimulating hormone; ethyl etiopurpurin tin; tirapazamine; titanocene dichloride; topsentin; toremifene; totipotent stem cell factor; translation inhibitors; tretinoin; triacetyluridine; triciribine; trimetrexate; triptorelin; tropisetron; turosteride; tyrosine kinase inhibitors; tyrphostins; UBC inhibitors ;Ubenimex;Urogenital sinus-derived growth inhibitor;Urokinase receptor antagonist;Vapreotide;Variolin B;Vector systems, red blood cell gene therapy;Veraresol;Veramine;Vergins;Verteporfin;Vinorelbine;Vinxartin;Vitaxin;Vorozole;Zanotherone;Zeniplatin;Zilascorub;Zinostatin stimalamer, Adriamycin, Dactinomycin, Bleomycin, Vinblastine, Cisplatin,Acivicin; Aclarubicin; Acodazole Hydrochloride; Acronine; Adzelesin; Aldesleukin; Altretamine; Ambomycin; Amethantrone Acetate; Aminoglutethimide; Amsacrine; Anastrozole; Anthramycin; Asparaginase; Asperlin; Azacitidine; Azetepa; Azotomycin; Batimastat; Benzodepa; Bicalutamide; Bisantrene Hydrochloride; Bisnafide Dimesylate; Bizelesin; Breo sulfate Mycin; brequinar sodium; bropirimine; busulfan; cactinomycin; calsterone; caracemide; carbetimer; carboplatin; carmustine; carubicin hydrochloride; carzelesin; cedefingol; chlorambucil; ciloremycin; cladribine; crisnatol mesylate; cyclophosphamide; cytarabine; dacarbazine; daunorubicin hydrochloride; decitabine; dexormaplatin; deazaguanine; dexorumaplatin mesylate Azaguanine; Diazicon; Doxorubicin; Doxorubicin hydrochloride; Droloxifene; Droloxifene citrate; Dromostanolone propionate; Duazomycin; Edatrexate; Eflornithine hydrochloride; Elsamitrucin; Enloplatin; Enpromate; Epipropizine; Epirubicin hydrochloride; Elbrozole; Esorubicin hydrochloride; Estramustine; Estramustine sodium phosphate; Etanidazole; Etoposide; Etoposide phosphate; etopurine; fadrozole hydrochloride; fazarabine; fenretinide; floxuridine; fludarabine phosphate; fluorouracil; fluorocitabine; foskidone; fostriecin sodium; gemcitabine; gemcitabine hydrochloride; hydroxyurea; idarubicin hydrochloride; ifosfamide; ilmofosine; interleukin I1 (including recombinant interleukin II, i.e., rlL2),Interferon alpha-2a; interferon alpha-2b; interferon alpha-n1; interferon alpha-n3; interferon beta-1a; interferon gamma-1b; iproplatin; irinotecan hydrochloride; lanreotide acetate; letrozole; leuprolide acetate; liarozole hydrochloride; lometrexol sodium; lomustine; losoxantrone hydrochloride; masoprocol; maytansine; mechlorethamine hydrochloride; megestrol acetate; melengestrol acetate; melphalan; menogaril; mercaptopurine; methotrexate Methotrexate; Metoprine; Meturedepa; Mitindomide; Mitocalcin; Mitochromin; Mitogillin; Mitomarcin; Mitomycin; Mitosper; Mitotane; Mitoxantrone hydrochloride; Mycophenolic acid; Nocodazole; Nogalamycin; Ormaplatin; Oxisuran; Pegaspargase; Periomycin; Pentamustine; Peplomycin sulfate; Perfosfamide; Pipobroman; Piposulfan; Piroxantrone hydrochloride; Plicamycin; Promestane; Porfimer sodium; Por Filomycin;Prednimustine;Procarbazine hydrochloride;Puromycin;Puromycin hydrochloride;Pyrazofurin;Rivopurin;Rogletimide;Safingol;Safingol hydrochloride;Semustine;Simtrazene;Sparfosate sodium;Sparsomycin;Spirogermanium hydrochloride;Spiromustine;Spiroplatin;Streptonigrin;Streptozocin;Sulofenur;Tallysomycin;Tecogalan sodium;Tegafur;Teroxantrone hydrochloride;Temoporfin;Teniposide;Teroxilon;Testolactone;Thiamiprine;Thiaminoprine Oguanine; Thiotepa; Tiazofurin; Tirapazamine; Toremifene citrate; Trestrone acetate; Triciribine phosphate; Trimetrexate; Trimetrexate glucuronate; Triptorelin; Tubrozole hydrochloride; Uracil mustard; Uredepa; Vapreotide; Verteporfin; Vinblastine sulfate; Vincristine sulfate; Vindesine; Vindesine sulfate; Binepidine sulfate; Vingricinate sulfate; Vinleurosine sulfate; Vinorelbine tartrate; Vinrocidine sulfate; Vinzolidine sulfate; Vorozole; Zeniplatin; Zinostatin; Zorubicin hydrochloride,Agents that arrest cells in the G2-M phase and / or agents that modulate the formation or stability of microtubules (e.g., Taxol™ (i.e., paclitaxel), Taxotere™ compounds containing a taxane skeleton, elbrozole (i.e., R-55104), dolastatin 10 (i.e., DLS-10 and NSC-376128), mibobulin isethionate (i.e., as CI-980), vincristine, NSC-639829, discodermolide (i.e., as NVP-XX-A-296), ABT-751 (Abbott, i.e., E-7010), altorhyrtins (e.g., altorhyrtin A and altorhyrtin C), C)), spongistatins (e.g., spongistatin 1, spongistatin 2, spongistatin 3, spongistatin 4, spongistatin 5, spongistatin 6, spongistatin 7, spongistatin 8, and spongistatin 9), cemadotin hydrochloride (i.e., LU-103793 and NSC-D-669356), epothilones (e.g., epothilone A, epothilone B, epothilone C (i.e., desoxyepothilone A or dEpoA), epothilone D (i.e., KOS-862, dEpoB, and desoxyepothilone B), epothilone E, epothilone F, epothilone B N-oxide, epothilone A N-oxide, 16-aza-epothilone B, 21-aminoepothilone B (i.e., BMS-310705), 21-hydroxyepothilone D (i.e., desoxyepothilone F and dEpoF), 26-fluoroepothilone, auristatin PE (i.e., NSC-654663), soblidotin (i.e., TZT-1027), LS-4559-P (Pharmacia, i.e., LS-4577), LS-4578 (Pharmacia), ia, i.e., LS-477-P), LS-4477 (Pharmacia), LS-4559 (Pharmacia), RPR-112378 (Aventis), vincristine sulfate, DZ-3358 (Daiichi), FR-182877 (Fujisawa, i.e., WS-9885B), GS-164 (Takeda), GS-198 (Takeda), KAR-2 (Hungarian Academy of Sciences),BSF-223651 (BASF, i.e. ILX-651 and LU-223651), SAH-49960 (Lilly / Novartis), SDZ-268970 (Lilly / Novartis), AM-97 (Armad / Kyowa Hakko), AM-132 (Armad), AM-138 (Armad / Kyowa Hakko), IDN-5005 (Indena), cryptophycin 52 (i.e., LY-355703), AC-7739 (Ajinomoto, i.e., AVE-8063A and CS-39.HCl), AC-7700 (Ajinomoto, i.e., AVE-8062, AVE-8062A, CS-39-L-Ser.HCl, and RPR-258062A), bitilebamide, tublysin A, canadensol, centaureidin (i.e., NSC-106969), T-138067 (Tularik, i.e., T-67, TL-138067, and TI-138067), COBRA-1 (Parker Hughes Laboratories), Institute, i.e., DDE-261 and WHI-261), H10 (Kansas State University), H16 (Kansas State University), oncocidin A1 (i.e., BTO-956 and DIME), DDE-313 (Parker Hughes Institute), physianolide B, laulimalide, SPA-2 (Parker Hughes Institute), SPA-1 (Parker Hughes Institute, i.e., SPIKET-P), 3-IAABU (Cytoskeleton / Mt. Sinai School of Medicine, i.e., MF-569), narcosine (also known as NSC-5366), nascapine, D-24851 (Asta Medica), A-105972 (Abbott), hemiasterin, 3-BAABU (Cytoskeleton / Mt. Sinai School of Medicine, i.e., MF-191),TMPN (Arizona State University), vanadocene acetylacetonate, T-138026 (Tularik), monastrol, indanocine (i.e., NSC-698666), 3-IAABE (Cytoskeleton / Mt. Sinai School of Medicine), A-204197 (Abbott), T-607 (Tularik, i.e., T-900607), RPR-115781 (Aventis), eleutherobins (such as desmethyleleutherobin, desacetyleleutherobin, isoeleutherobin A, and Z-eleutherobin), carybeoside, carybeolin, halichondrin B, D-64131 (Asta Medica), D-68144 (Asta Medica), diazonamide A, A-293620 (Abbott), NPI-2350 (Nereus), taccalonolide A, TUB-245 (Aventis), A-259754 (Abbott), diozostatin, (-)-phenylahistine (i.e., NSCL-96F037), D-68, 838 (Asta Medica), D-68836 (Asta Medica), myoseverin B, D-43411 (Zentaris, i.e., D-81862), A-289099 (Abbott), A-318315 (Abbott), HTI-286 (i.e., SPA-110, trifluoroacetate) (Wyeth), D-82317 (Zentaris), D-82318 (Zentaris), SC-12983 (NCI), resverastatin sodium phosphate, BPR-OY-007 (National Health Research Institute, These include but are not limited to: steroids (e.g., dexamethasone), finasteride, aromatase inhibitors, gonadotropin-releasing hormone agonists (GnRH) such as goserelin or leuprolide, corticosteroids (e.g., prednisone), progestins (e.g., hydroxyprogesterone caproate, megestrol acetate, medroxyprogesterone acetate), estrogens (e.g., diethlystilbestrol, ethinyl estradiol), antiestrogens (e.g., tamoxifen), androgens (e.g., testosterone propionate, fluoxymesterone), antiandrogens (e.g., flutamide), immunostimulants (e.g., bacillus Calmette-Guerin (BCG), levamisole, anti-CD33 monoclonal antibody-calicheamicin conjugate, anti-CD22 monoclonal antibody-Pseudomonas exotoxin conjugate, etc.), radioimmunotherapy (e.g., anti-CD20 monoclonal antibody conjugated to 111In, 90Y, or 131I, etc.), triptolide, homoharringtonine, dactinomycin, doxorubicin, epirubicin, topotecan, itraconazole, vindesine, cerivastatin, vincristine, deoxyadenosine, sertraline, pitavastatin, irinotecan, clofazimine, 5-nonyloxytryptamine, vemurafenib, dabrafenib,Erlotinib, gefitinib, EGFR inhibitors, epidermal growth factor receptor (EGFR) targeted therapies or therapeutic agents (e.g., gefitinib (Iressa™), erlotinib (Tarceva™), cetuximab (Erbitux™), lapatinib (Tykerb™), panitumumab (Vectibix™), vandetanib (Caprelsa™), afatinib / BIBW2992, CI-1033 / canertinib, neratinib / HKI-272 , CP-724714, TAK-285, AST-1306, ARRY334543, ARRY-380, AG-1478, dacomitinib / PF299804, OSI-420 / desmethylerlotinib, AZD8931, AEE788, pelitinib / EKB-569, CUDC-101, WZ8040, WZ4002, WZ3146, AG-490, XL647, PD153035, BMS-599626), sorafenib, imatinib, sunitinib, dasatinib, etc.

[0058] In some embodiments, the anticancer drug is an epigenetic inhibitor. As used herein, "epigenetic inhibitor" refers to an inhibitor of epigenetic processes, such as DNA methylation (DNA methylation inhibitor) or histone modification (histone modification inhibitor). The epigenetic inhibitor can be a histone deacetylase (HDAC) inhibitor, a DNA methyltransferase (DNMT) inhibitor, a histone methyltransferase (HMT) inhibitor, a histone demethylase (HDM) inhibitor, or a histone acetyltransferase (HAT). Examples of HDAC inhibitors include vorinostat, romidepsin, CI-994, belinostat, panobinostat, gibinostat, entinostat, mocetinostat, SRT501, CUDC-101, JNJ-26481585, or PCI24781. Examples of DNMT inhibitors include azacitidine and decitabine. Examples of HMT inhibitors include EPZ-5676. Examples of HDM inhibitors include pargyline and tranylcypromine. Examples of HAT inhibitors include CCT077791 and garcinol.

[0059] In some embodiments, the anti-cancer agent is a multikinase inhibitor. A "multikinase inhibitor" is a small molecule inhibitor of at least one protein kinase, including tyrosine protein kinases and serine / threonine kinases. A multikinase inhibitor may include a single kinase inhibitor. A multikinase inhibitor may block phosphorylation. A multikinase inhibitor may act as a covalent modulator of a protein kinase. A multikinase inhibitor may bind to the kinase active site or to a secondary or tertiary site that inhibits protein kinase activity. The multikinase inhibitor may be an anti-cancer multikinase inhibitor. Exemplary anti-cancer multikinase inhibitors include dasatinib, sunitinib, erlotinib, bevacizumab, vatalanib, vemurafenib, vandetanib, cabozantinib, poatinib, axitinib, ruxolitinib, regorafenib, crizotinib, bosutinib, cetuximab, gefitinib, imatinib, lapatinib, lenvatinib, mubritinib, nilotinib, panitumumab, pazopanib, trastuzumab, or sorafenib.

[0060] Urine specimen assay In one aspect, the present disclosure provides a method for sequencing a cell-free nucleic acid molecule of a subject. In some embodiments, the method includes: (a) treating a urine specimen to inhibit cell lysis; (b) separating the cell-free nucleic acid molecules in the processed urine specimen from cells in the processed urine specimen, thereby producing a purified urine specimen containing the cell-free nucleic acid molecules; (c) concentrating the cell-free nucleic acid molecules in the purified urine specimen by passing at least a portion of the purified urine specimen through a filter, where (i) the concentration produces a filtrate and a residual urine specimen, and (ii) the residual urine specimen has an increased concentration of cell-free nucleic acid molecules; (d) isolating the cell-free nucleic acid molecules from the residual urine specimen; and (e) sequencing the isolated cell-free nucleic acid molecules.

[0061] Urine specimens for use in the present disclosure may be collected from a variety of sources and may exhibit a variety of characteristics. In some embodiments, the urine specimen comprises a desired minimum initial volume. For example, the urine specimen may have a volume of at least 5 mL, 10 mL, 20 mL, 30 mL, 40 mL, 50 mL, 75 mL, 100 mL, or more. In some embodiments, the volume of the processed urine specimen is 5 mL, 10 mL, 15 mL, 20 mL, 30 mL, 40 mL, 50 mL, or more. In some embodiments, the urine specimen is at least a 50 mL specimen.

[0062] Typically, a urine specimen as collected will contain cell-free nucleic acid molecules and other components, such as cells, metabolites, proteins, and salts. In some embodiments, the urine specimen is treated to inhibit cell lysis. Inhibition of cell lysis is not absolutely necessary, but will generally result in a reduced rate of cell lysis in the specimen compared to an untreated urine specimen. In some embodiments, treating the urine specimen to inhibit cell lysis includes contacting the urine specimen with one or more preservation reagents. A variety of suitable preservation reagents are available. Preservative agents that can be used to inhibit cell lysis include, but are not limited to, imidazolidinyl urea (IDU), diazolidinyl urea (DU), dimethylol urea, 2-bromo-2-nitropropane-1,3-diol, 5-hydroxymethoxymethyl-1-aza-3,7-dioxabicyclo(3.3.0)octane and 5-hydroxymethyl-1-aza-3,7-dioxabicyclo(3.3.0)octane and 5-hydroxypoly[methyleneoxy]methyl-1-aza-3,7-dioxabicyclo(3.3.0)octane, bicyclic oxazolidines (e.g., Nuosept 95), DMDM ​​hydantoin, sodium hydroxymethylglycinate, hexamethylenetetramine chloroallyl chloride (quaternium-15), biocides (such as Bioban, Preventol, and Grotan), water-soluble zinc salts, sodium azide, or any combination thereof. In some embodiments, the preservation reagent comprises imidazolidinyl urea.

[0063] In some embodiments, treating the urine specimen to inhibit cell lysis includes treatment with a combination of reagents, such as treatment with a preservative reagent, a nuclease inhibitor, a formaldehyde quencher, or a combination thereof. Treating the urine specimen to inhibit cell lysis can affect the integrity of nucleic acids in the specimen. Some cell lysis inhibitors can release formaldehyde, which, if not quenched, can damage or destroy the structural integrity of nucleic acids. Therefore, adding a formaldehyde quencher can preserve the stability of nucleic acids in the specimen. Formaldehyde quenchers that can be used include, but are not limited to, glycine, tris(hydroxymethyl)aminomethane (TRIS), urea, allantoin, sulfite, or any combination thereof. In some embodiments, the formaldehyde inhibitor is glycine.

[0064] In some embodiments, treating the urine specimen includes treatment with a nuclease inhibitor. Nuclease activity in fresh urine can rapidly hydrolyze nucleic acids (e.g., DNA). Nuclease inhibitors that can be used include, but are not limited to, ethylene glycol tetraacetic acid (EGTA), pepstatin, EDTA, phosphoramidon, leupeptin, aprotinin, bestatin, proteinase inhibitor E64 (E-64), 4-(2-aminoethyl)benzenesulfonyl fluoride hydrochloride (AEBSF), or any combination thereof. In some embodiments, the nuclease inhibitor is EDTA. Treatment with a nuclease inhibitor may also be performed with or without a formaldehyde quencher. In some embodiments, the treatment includes a nuclease inhibitor but not a formaldehyde quencher. In some embodiments, the treatment includes a nuclease inhibitor and a formaldehyde quencher.

[0065] In some embodiments, the treating comprises contacting the urine specimen with a composition comprising a cytolysis inhibitor and a nuclease inhibitor, a formaldehyde quencher, or both. In some embodiments, the composition comprises imidazolidinyl urea, EDTA, glycine, or a combination thereof. In some embodiments, the imidazolidinyl urea is added to a final concentration of 0.5% to 2.0%. In some embodiments, the imidazolidinyl urea is added to a final concentration of 0.20% to 4.0%. In some embodiments, the glycine is present at a concentration of 0.01% to 0.2%. In some embodiments, the glycine is present at a concentration of 0.01% to 0.4%. In some embodiments, the glycine is present at a concentration of at least 0.01%, at least 0.05%, at least 0.2%, or even at least 0.3%. In some embodiments, the EDTA is present at a concentration of 0.5% to 2.0%. In some embodiments, EDTA is present at a concentration of 0.50% to 3.6%. In some embodiments, EDTA is present at a concentration of at least 0.5%, at least 1%, or at least 2.5%. In some embodiments, the ratio of EDTA to imidazolidinyl urea is 5:2. In some embodiments, the ratio of EDTA to imidazolidinyl urea is 1:3 to 3:1 (e.g., 9:10, 9:5, or 9:20). In some embodiments, the ratio of EDTA to glycine is 5:1 to 100:1 (e.g., 50:1 or 9:1).

[0066] In some embodiments, the composition used in treating a urine specimen to inhibit cell lysis comprises sodium azide, EDTA, or a combination thereof. In some embodiments, the sodium azide is present at a final concentration of 0.05% to 2.0%. In some embodiments, the sodium azide is present at a concentration of 0.10% to 1.0%. In some embodiments, the EDTA is present at a concentration of 0.5% to 2.0%. In some embodiments, the EDTA is present at a concentration of 0.50% to 3.6%. In some embodiments, the EDTA is present at a concentration of at least 0.5%, at least 1%, or at least 2.5%.

[0067] In some embodiments, the composition used in treating a urine specimen to inhibit cell lysis is present in an amount of 1 to 20 volume percent of the urine specimen after contact. The preservation reagent may be combined with the urine specimen after collection (e.g., by adding it to the urine specimen or by transferring all or a portion of the urine specimen to a container containing the preservation reagent), or may be present in the container used to collect the specimen. Further non-limiting examples of compositions comprising one or more of a preservation reagent, a nuclease inhibitor, or a formaldehyde quencher are described in U.S. Patent Application Publication No. 20160257995 A1, which is incorporated herein by reference.

[0068] In some embodiments, the methods include separating cell-free nucleic acid molecules (e.g., cfDNA and / or cfRNA) in the processed urine specimen from cells in the processed urine specimen to produce a purified urine specimen containing cell-free nucleic acid molecules (and having reduced cell content). Various methods are available for separating cell-free nucleic acids from cells, including, but not limited to, fractionation, centrifugation (e.g., pelleting or density gradient centrifugation), precipitation, and flow cytometry. In some embodiments, the separating includes pelleting cells by centrifuging the processed urine specimen. The cell pellet can then be removed, or the supernatant can be transferred to a new container. In some embodiments, starting with the processed urine specimen, cell-free nucleic acids (e.g., cell-free DNA) can be separated from cells by, for example, centrifugation at 3000 rpm to 8000 rpm for 10 to 15 minutes at room temperature. The processed urine specimen may be centrifuged at at least 1000 g, 2000 g, 3000 g, 4000 g, 5000 g, 6000 g, 7000 g, 8000 g, or more. Additionally, centrifugation may be performed for at least 5, 10, 15, 20, 30, 45, 60 minutes, or more. In some embodiments, the processed urine specimen is centrifuged at 4000 g for 20 minutes. The separated cells will form a cell pellet upon centrifugation, which can be removed by decanting the supernatant, thus separating the cells from the nucleic acids. Separation may be performed one, two, three, four, or more times, or until no visible cell pellet is formed after centrifugation.

[0069] In some embodiments, a purified urine specimen is subjected to a concentration step to produce a specimen with an increased concentration of nucleic acids. Concentration may provide certain advantages, such as the ability to work with a larger input sample volume compared to other specimen types (e.g., plasma), thereby offsetting the reduced cell-free nucleic acid concentration in urine. The availability of higher nucleic acid concentrations and smaller sample volumes may reduce the amount of reagents and materials required for processing and more easily allow for parallel processing of multiple specimens. In some embodiments, the concentration of cell-free nucleic acid molecules in the purified urine specimen is achieved by passing at least a portion of the purified urine specimen through a filter. Passing the purified urine specimen through the filter produces a filtrate and a residual urine specimen. The filtrate will contain salts and other small components that are able to pass through the filter, while the residual urine specimen will contain an increased concentration of nucleic acids. In some instances, the filter is substantially impermeable to nucleic acids (e.g., cell-free nucleic acids) in the purified urine specimen and substantially permeable to salts and other small components. When a large sample volume is filtered in two or more portions, each of the two or more portions may be applied sequentially to the same filter, separately to different filters, or some combination thereof. The remaining portion of the sample after filtration may be subjected to one or more additional rounds of filtration (e.g., 2, 5, 10, 15, or more rounds), such as by repeated application to the same filter or by sequential application to different filters.

[0070] Any of a variety of types of filters made from a variety of materials may be used, with several options commercially available. For example, filters may be membrane filters, such as those made from nylon, cellulose, or nitrocellulose, or may include beads (e.g., Sepharose beads). In some embodiments, filters are characterized by a molecular cutoff. Generally, the molecular weight cutoff is provided by a specific pore size and / or coating that allows for the separation of molecules above the cutoff from molecules below the cutoff. The molecular weight cutoff (also referred to as the "molecular weight limit") is typically specified, among other properties, as a characteristic of commercially available filters for processing biological specimens. In some embodiments, the molecular weight cutoff corresponds to the molecular weight of the solute that is 90% retained by the filter. In some embodiments, the filter has a nominal molecular weight cutoff of 1 kilodalton (kD) to 50 kD. In some embodiments, the filter has a nominal molecular weight cutoff of 3 kD to 10 kD. In some embodiments, the filter has a nominal molecular weight cutoff of 10 kD, 5 kD, 3 kD, or less. In some embodiments, the filter comprises a nominal molecular weight cutoff of 3 kD or less.

[0071] After one or more concentration steps, various nucleic acid concentrations can be achieved in the residual urine specimen relative to the purified urine specimen. In some embodiments, the residual urine specimen has a concentration that is at least 2-fold increased compared to the purified urine specimen (before concentration). In some embodiments, the residual urine specimen has a concentration that is at least 2-fold to 20-fold increased. In some embodiments, the residual urine specimen has a concentration that is at least 2-fold, at least 3-fold, at least 4-fold, at least 5-fold, at least 10-fold, at least 15-fold, at least 20-fold, or more increased compared to the purified urine specimen. In some embodiments, the residual urine specimen has a concentration that is at least 2-fold, at least 5-fold, at least 10-fold, or at least 15-fold increased compared to the purified urine specimen. In some embodiments, the residual urine specimen has a concentration that is at least 5-fold increased compared to the purified urine specimen. In some embodiments, the residual urine specimen has a concentration that is at least 10-fold increased compared to the purified urine specimen.

[0072] As a result of passing at least a portion of the urine specimen through the filter, the residual urine specimen (including the portion of the specimen subjected to filtration that does not pass through the filter into the filtrate) has a smaller volume compared to the starting volume of the specimen subjected to filtration. In some embodiments, the residual urine specimen has a volume that is at least 10% to at least 90% smaller than the volume of the processed urine specimen. In some embodiments, the residual urine specimen has a volume that is at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 80%, or at least 90% smaller than the volume of the processed urine specimen. In some embodiments, the residual urine specimen has a volume that is at least 50% smaller than the volume of the processed urine specimen. In some embodiments, the residual urine specimen has a volume that is at least 75% smaller than the volume of the processed urine specimen. For example, when a 50 mL processed urine specimen is concentrated to a residual urine specimen volume of 5 mL or less, the residual urine specimen volume has been reduced by 90% or more.

[0073] The residual urine specimen can be used immediately for further processing or assay of cell-free nucleic acids. For example, the residual urine specimen may be processed within one hour after concentration. Alternatively, the residual urine specimen may be stored for later use. In some instances, the residual urine specimen may be stored at 4°C, 22°C, or 37°C for up to one week after urine specimen collection. In some embodiments, the residual urine specimen may be frozen to preserve the specimen for later analysis. In some embodiments, the residual urine specimen is stored at -20°C, -80°C, or colder. In some embodiments, the frozen specimen is stored in a frozen state for one to six months.

[0074] The steps of processing, separating, and concentrating the urine specimen may be performed at various times after collection of the urine specimen. In some embodiments, processing is completed within 10 to 180 minutes after collection of the urine specimen. In some embodiments, processing is completed within 180, 150, 120, 90, 60, 45, 30, 15, or 10 minutes after collection of the urine specimen. In some embodiments, processing is completed within 120, 60, or 30 minutes after collection of the urine specimen. In some embodiments, processing is completed within 60 minutes after collection of the urine specimen. In some embodiments, processing is completed within 30 minutes after collection of the urine specimen. In some embodiments, the processed specimen is stored before proceeding to the separation and concentration steps. For example, after processing, further separation and concentration steps may be completed within 14 days (e.g., within 2, 3, 4, 5, 6, 7, 8, 9, or 10 days) after collection. In some embodiments, separation and concentration is completed within 7 days after collection of the urine specimen.

[0075] Nucleic acids can be isolated from residual urine specimens using any of a variety of nucleic acid isolation methods. Exemplary methods include, but are not limited to, extraction, solid-phase extraction, silica-based purification, magnetic particle-based purification, phenol-chloroform extraction, chromatography, anion exchange chromatography (using anion exchange surfaces), electrophoresis, filtration, precipitation, immunoprecipitation, hybridization capture with targeted bait oligonucleotides, or any combination thereof. For example, isolation of target nucleic acids by hybridization to biotinylated probes using streptavidin-coated surfaces (e.g., streptavidin-coated beads). Commercially available methods and kits are also available, including, but not limited to, the QIAamp® Circulating Nucleic Acid Kit (QIAGEN), the Chemagic Circulating Nucleic Acid Kit (Chemagen), the NucleoSpin Plasma XS Kit (Macherey-Nagel), and the High Pure Viral Nucleic Acid Large Volume Kit (Roche).

[0076] In some embodiments, the isolated cell-free nucleic acid comprises cfDNA, and the method comprises treating the cfDNA molecule to distinguish between methylated and unmethylated nucleotides, thereby producing a converted cfDNA molecule. In some embodiments, the treatment comprises deamination, such as treating the cfDNA molecule with cytosine deaminase or bisulfite. In some embodiments, the method comprises treating the cfDNA molecule with bisulfite to produce a converted cfDNA molecule.

[0077] The methods disclosed herein may further include amplifying cell-free nucleic acids isolated from residual urine specimens. Amplification may be nonspecific (e.g., amplification using random primers) or targeted (e.g., directed to a specific target region of interest). Any suitable method known in the art may be used for amplification. Examples of nucleic acid amplification reactions include, but are not limited to, polymerase chain reaction (PCR), rolling circle amplification (RCA), ligase chain reaction (LCR), simple method for amplifying RNA targets (SMART), single primer isothermal amplification (SPIA), multiple displacement amplification (MDA), nucleic acid sequence-based amplification (NASBA), hinge-initiated primer-dependent nucleic acid amplification (HIP), nicking enzyme amplification reaction (NEAR), RT-PCR, loop-mediated amplification (LAMP), exponential amplification reaction (EXPAR), and improved multiple displacement amplification (IMDA). Primer-based amplification methods may use primers targeting any region of the genome. Alternatively, primers may be used to specifically amplify targets / biomarkers of interest, thereby enriching the sample for the desired targets / biomarkers. For example, forward and reverse primers can be prepared for each genomic region of interest and used to amplify fragments corresponding to or derived from the desired genomic region. Amplification can be thermal amplification (e.g., as in PCR) or isothermal amplification.

[0078] In some embodiments, the isolated cell-free nucleic acid, or its amplification product, is captured by hybridization to a bait oligonucleotide. In some embodiments, the bait oligonucleotide is at least 45 nucleotides in length (e.g., at least 60, 75, 80, 90, 100, 110, or 120 nucleotides in length). In some embodiments, the bait oligonucleotide is no longer than 130, 140, 150, 200, 250, or 300 bases in length. In some embodiments, the bait oligonucleotide is 45-300, 60-200, or 75-150 nucleotides in length. In some embodiments, the bait oligonucleotide is at least 50 nucleotides in length. In some embodiments, the bait oligonucleotide is at least 60 nucleotides in length. In some embodiments, the bait oligonucleotide is at least 75 nucleotides in length. In some embodiments, the designated length of the bait oligonucleotide is a length designed to be complementary to a portion of a target genome sequence or its converted DNA molecule.

[0079] In some embodiments, the bait oligonucleotide targets at least 500, 1000, 1500, 5000, 10,000, 12,500, 15,000, 17,000, 19,000, or more target genomic regions. In some embodiments, the bait oligonucleotide targets 25,000, 20,000, 17,000, 15,000, 12,500, 10,000, or fewer target genomic regions. In some embodiments, the bait oligonucleotide targets 5,000-30,000, 10,000-25,000, 12,500-20,000, or 15,000-20,000 target genomic regions. In some embodiments, the bait oligonucleotide targets 100-500, 500-1,000, 1500-5,000, or 5,000-10,000 target genomic regions. In some embodiments, the bait oligonucleotide targets at least 500 target genomic regions. In some embodiments, the bait oligonucleotide targets at least 1000 target genomic regions. In some embodiments, the bait oligonucleotide targets at least 10,000 target genomic regions. In some embodiments, the bait oligonucleotide targets at least 15,000 target genomic regions. In some embodiments, the bait oligonucleotide targets less than 20,000 target genomic regions.

[0080] In some embodiments, bait oligonucleotides are configured to hybridize to converted DNA molecules (e.g., converted cfDNA molecules) corresponding to or derived from one or more genomic regions. Thus, the bait oligonucleotides can have a sequence different from that of the targeted genomic region. For example, DNA with unmethylated CpG sites can be converted to contain UpG instead of CpG by deamination (e.g., by treatment with cytosine deaminase or bisulfite). As a result, bait oligonucleotides for such targets can be configured to hybridize to sequences containing UpG instead of naturally occurring unmethylated CpG. Thus, the sites complementary to unmethylated sites in the bait oligonucleotides can contain CpA instead of CpG, and some probes targeting hypomethylated sites where all methylated sites are unmethylated can lack a guanine (G) base. In some embodiments, at least 3%, 5%, 10%, 15%, or 20% of the probes do not contain CpG sequences. In some embodiments, at least 5% of the probes do not contain CpG sequences. In some embodiments, at least 10% of the probes do not contain CpG sequences.

[0081] In some embodiments, bait oligonucleotides are used to detect the presence or absence of cancer generally and / or provide a cancer classification, such as cancer type, cancer stage, such as I, II, III, or IV, or tissue of origin (TOO), where the cancer is believed to have originated. Bait oligonucleotides can be target genomic regions that are differentially methylated between cancerous (pan-cancer) specimens in general and non-cancerous specimens, or only in cancerous specimens of a specific cancer type (e.g., urological cancer-specific targets). For example, in some embodiments, bait oligonucleotides are designed to contain genomic regions that are differentially methylated based on converted (e.g., bisulfite) sequencing data generated from cell-free nucleic acids and / or total genomic DNA of a collection of cancer and non-cancer individuals.

[0082] In some embodiments, each of the target genomic regions is differentially methylated in a cancer specimen compared to a non-cancer specimen. In some embodiments, differential methylation includes at least 50% to at least 90% of the CpG sites in the target genomic region being methylated or unmethylated. In some embodiments, differential methylation includes at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% of the CpG sites in the target genomic region being methylated or unmethylated. In some embodiments, differential methylation includes at least 70% of the CpG sites in the target genomic region being methylated or unmethylated. In some embodiments, differential methylation includes at least 80% of the CpG sites in the target genomic region being methylated or unmethylated. In some embodiments, the cancer is a urological cancer. In some embodiments, the cancer is bladder cancer, prostate cancer, or renal cancer. In some embodiments, the target genomic regions are selected to identify the presence of, and optionally distinguish between, two or more cancer types (eg, at least 3, 4, 5, or more cancer types).

[0083] In some embodiments, the cell-free nucleic acid comprises DNA or RNA. In some embodiments, the cell-free nucleic acid is cell-free DNA (cfDNA). In some embodiments, the method comprises treating the cfDNA molecule to distinguish between methylated and unmethylated nucleotides, thereby producing a converted cfDNA molecule. In some embodiments, the treatment comprises deamination, such as treating the cfDNA molecule with cytosine deaminase or bisulfite. In some embodiments, the method comprises treating the cfDNA molecule with bisulfite to produce a converted cfDNA molecule.

[0084] In some embodiments, the method comprises separating the cell-free nucleic acid molecules bound to the bait from unbound cell-free nucleic acid molecules.In some embodiments, each bait oligonucleotide is conjugated to a solid surface (e.g., a chip or a bead such as a magnetic bead or a paramagnetic bead) or to a non-nucleotide affinity moiety (e.g., a member of a binding pair).In some embodiments, such a conjugate is used to facilitate the separation of the cfDNA molecules bound to the bait oligonucleotide from unbound cfDNA molecules.In some embodiments, the affinity moiety is biotin.

[0085] In some embodiments, genomic regions can be selected to have at least 3, 5, 7, 10, or more methylation sites. In some embodiments, each target genomic region comprises at least 5 methylation sites. In some embodiments, the selected number of methylation sites (e.g., at least 5 methylation sites) are differentially methylated in at least one cancer type to be assayed by the panel. The selected target genomic regions can be located in various locations within the genome, including, but not limited to, promoters, enhancers, exons, introns, intergenic regions, and other sites.

[0086] In some embodiments, the target genomic region is selected from the sequences of genes in Table 1. In some embodiments, the target genomic region comprises a target sequence of a gene selected from Table 1. The target sequence may include two or more target sequences from a single gene and / or may include one or more target sequences from two or more genes. In some embodiments, the one or more target genomic regions comprise one or more target sequences from at least 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, 75, or more genes selected from Table 1. In some embodiments, the target genomic region comprises one or more target sequences from at least 10 genes from Table 1. In some embodiments, the target genomic region comprises one or more target sequences from at least 15 genes from Table 1. In some embodiments, the target genomic region comprises one or more target sequences from at least 25 genes from Table 1. In some embodiments, the target genomic region comprises a target sequence of a gene selected from one or more of TWIST1, EOMES, HOXA9, POU4F2, and ZNF154. In some embodiments, the target genomic region comprises a target sequence of a gene selected from one or more of TWIST1, PXDN, RP11-259O2.3, EVX2, and KNDC1. In some embodiments, the target genomic region comprises a target sequence of the TWIST1 gene. In some embodiments, the target genomic region comprises one or more target sequences from at least 20% of the genes in Table 1. In some embodiments, the target genomic region comprises one or more target sequences from at least 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100% of the genes in Table 1. In some embodiments, the target genomic region comprises one or more target sequences from at least 50% of the genes in Table 1. In some embodiments, the target genomic region comprises target sequences from all of the genes in Table 1. In some embodiments, the target sequence is at least 10, 25, 35, 45, 50, 75, 100, 200, 300, or more nucleotides in length. In some embodiments, the target sequence is at least 25, at least 35, or at least 45 nucleotides in length.In some embodiments, the target sequence is at least 45 nucleotides in length. In some embodiments, the target genomic region has diagnostic capabilities for one or more of bladder cancer, kidney cancer, or prostate cancer.

[0087] [Table 1]

[0088] In some embodiments, the target genomic region is selected from the sequences of genes in Table 2. In some embodiments, the target genomic region comprises a target sequence of a gene selected from Table 2. The target sequence may include two or more target sequences from a single gene and / or may include one or more target sequences from two or more genes. In some embodiments, the one or more target genomic regions comprise one or more target sequences from at least 1, 2, 3, 4, 5, 10, 15, 20, or more genes selected from Table 2. In some embodiments, the target genomic region comprises one or more target sequences from at least five genes from Table 2. In some embodiments, the target genomic region comprises one or more target sequences from at least 10 genes from Table 2. In some embodiments, the target genomic region comprises one or more target sequences from at least 15 genes from Table 2. In some embodiments, the target genomic region comprises a target sequence of a gene selected from one or more of TWIST1, PXDN, RP11-259O2.3, EVX2, and KNDC1. In some embodiments, the target genomic region comprises a target sequence of the TWIST1 gene. In some embodiments, the target genomic region comprises target sequences from all of the genes in Table 2. In some embodiments, the target sequence is at least 10, 25, 35, 45, 50, 75, 100, 200, 300, or more nucleotides in length. In some embodiments, the target sequence is at least 25, at least 35, or at least 45 nucleotides in length. In some embodiments, the target sequence is at least 45 nucleotides in length. In some embodiments, the target genomic region has diagnostic capabilities for bladder cancer.

[0089] [Table 2]

[0090] In some embodiments, the target genomic region is selected from the sequences of genes in Table 3. In some embodiments, the target genomic region comprises a target sequence of a gene selected from Table 3. The target sequence may include two or more target sequences from a single gene and / or may include one or more target sequences from two or more genes. In some embodiments, the one or more target genomic regions comprise one or more target sequences from at least 1, 2, 3, 4, 5, 10, or more genes selected from Table 3. In some embodiments, the target genomic region comprises one or more target sequences from at least five genes from Table 3. In some embodiments, the target genomic region comprises one or more target sequences from at least 10 genes from Table 3. In some embodiments, the target genomic region comprises one or more target sequences from at least 15 genes from Table 3. In some embodiments, the target genomic region comprises a target sequence of a gene selected from one or more of TWIST1, SEMA6B, PXDN, KNDC1, and RP11-259O2.3. In some embodiments, the target genomic region comprises a target sequence of the TWIST1 gene. In some embodiments, the target genomic region comprises target sequences from all of the genes in Table 3. In some embodiments, the target sequence is at least 10, 25, 35, 45, 50, 75, 100, 200, 300, or more nucleotides in length. In some embodiments, the target sequence is at least 25, at least 35, or at least 45 nucleotides in length. In some embodiments, the target sequence is at least 45 nucleotides in length. In some embodiments, the target genomic region has diagnostic potential for bladder cancer.

[0091] [Table 3]

[0092] In some embodiments, the target genomic region is selected from the sequences of the genes in Table 4. In some embodiments, the target genomic region comprises a target sequence of a gene selected from Table 4. The target sequence may include two or more target sequences from a single gene and / or may include one or more target sequences from two or more genes. In some embodiments, the one or more target genomic regions comprise one or more target sequences from at least 1, 2, 3, 4, 5, 10, 15, 20, 30, or more genes selected from Table 4. In some embodiments, the target genomic region comprises one or more target sequences from at least 10 genes from Table 4. In some embodiments, the target genomic region comprises one or more target sequences from at least 15 genes from Table 4. In some embodiments, the target genomic region comprises one or more target sequences from at least 25 genes from Table 4. In some embodiments, the target genomic region comprises a target sequence of a gene selected from one or more of LBH, HIST1H4D, HIST1H2APS4, FZD2, and ADARB2. In some embodiments, the target genomic region comprises target sequences from all of the genes in Table 4. In some embodiments, the target sequence is at least 10, 25, 35, 45, 50, 75, 100, 200, 300, or more nucleotides in length. In some embodiments, the target sequence is at least 25, at least 35, or at least 45 nucleotides in length. In some embodiments, the target sequence is at least 45 nucleotides in length. In some embodiments, the target genomic region has diagnostic capabilities for prostate cancer.

[0093] [Table 4]

[0094] In some embodiments, the target genomic region is selected from the sequences of the genes in Table 5. In some embodiments, the target genomic region comprises a target sequence of a gene selected from Table 5. The target sequence may include two or more target sequences from a single gene and / or may include one or more target sequences from two or more genes. In some embodiments, the one or more target genomic regions comprise one or more target sequences from at least 1, 2, 3, 4, 5, 10, 15, or more genes selected from Table 5. In some embodiments, the target genomic region comprises one or more target sequences from at least five genes from Table 5. In some embodiments, the target genomic region comprises one or more target sequences from at least 10 genes from Table 5. In some embodiments, the target genomic region comprises one or more target sequences from at least 15 genes from Table 5. In some embodiments, the target genomic region comprises a target sequence of a gene selected from one or more of NAPA-AS1, NAPA, MFHAS1, RANGAP1, and HOXB2. In some embodiments, the target genomic region comprises target sequences from all of the genes in Table 5. In some embodiments, the target sequence is at least 10, 25, 35, 45, 50, 75, 100, 200, 300, or more nucleotides in length. In some embodiments, the target sequence is at least 25, at least 35, or at least 45 nucleotides in length. In some embodiments, the target sequence is at least 45 nucleotides in length. In some embodiments, the target genomic region has diagnostic potential for renal cancer.

[0095] [Table 5]

[0096] In some embodiments, the methods disclosed herein include sequencing isolated cell-free nucleic acid molecules. Various methods for sequencing nucleic acids are available, including, but not limited to, next-generation sequencing (NGS) techniques involving synthesis (Illumina), pyrosequencing (454 Life Sciences), ion semiconductor technology (Ion Torrent sequencing), single-molecule real-time sequencing (Pacific Biosciences), sequencing-by-ligation (SOLiD sequencing), nanopore sequencing (Oxford Nanopore Technologies), or paired-end sequencing. In some embodiments, massively parallel sequencing is performed using reversible dye terminator sequencing-by-synthesis. In some embodiments, sequencing the nucleic acid molecules includes sequencing amplicons of the isolated cell-free nucleic acid molecules.

[0097] In some embodiments, the method further includes diagnosing cancer in the subject. In some embodiments, the method further includes treating cancer in the subject. In some embodiments, the cancer is bladder cancer, prostate cancer, or renal cancer. The specific mode of treatment may depend on one or more of a variety of factors, such as the specific type of cancer detected, the TOO, location, and stage of the cancer. Non-limiting examples of treatments include surgical resection, radiation therapy, chemotherapy, and / or immunotherapy.

[0098] Cancer Assay Panel In one aspect, the present disclosure provides a cancer assay panel comprising a plurality of probes or a plurality of probe pairs. The assay panel described herein may alternatively be referred to as a bait set or a composition comprising bait oligonucleotides. The probes may be polynucleotide-containing probes specifically designed to target (e.g., by complementarity) one or more genomic regions with differential methylation between cancer and non-cancerous specimens, between different tissue of origin (TOO) types, between different cancer cell types, and / or between specimens of different cancer stages, as identified by the methods provided herein. In some embodiments, the target genomic regions are selected to maximize classification accuracy, subject to a size budget (which is determined by the sequencing budget and desired sequencing depth).

[0099] In one aspect, the present disclosure provides a composition comprising a plurality of different bait oligonucleotides. In some embodiments, (a) the bait oligonucleotides hybridize to converted DNA molecules derived from one or more target genomic regions; (b) the one or more target genomic regions include one or more target sequences of one or more genes selected from Table 1; (c) the one or more target genomic regions are differentially methylated in cancer; and (d) each bait oligonucleotide includes a sequence of at least 25 nucleotides in length that hybridizes to one of the target sequences. To design a cancer assay panel, an analytics system can collect samples corresponding to various outcomes under consideration, such as samples from subjects known to have cancer, samples from subjects considered healthy, samples from subjects with cancer of known tissue of origin, etc. The source of DNA molecules (e.g., cfDNA and / or ctDNA) used to select the target genomic regions can vary depending on the purpose of the assay. For example, different sources may be desirable for assays intended to detect cancer generally, specific types of cancer, cancer stages, or tissues of origin. These samples may be processed by whole genome bisulfite sequencing (WGBS) or obtained from public databases (e.g., TCGA). In some embodiments, the analytics system is a computing system including a computer processor and a computer-readable storage medium containing instructions for causing the computer processor to execute some or all of the operations described in this disclosure. In some embodiments, at least some of the samples subjected to analysis to identify target genomic regions are urine specimens, such as urine specimens processed by the methods described herein.

[0100] The analytics system can then select target genomic regions based on the methylation patterns of the nucleic acid fragments. One approach considers the pairwise discrimination between pairs of outcomes for a region (or more specifically, for CpG sites within the region). Another approach considers the discrimination for a region (or more specifically, for CpG sites within the region) when each outcome is considered relative to the remaining outcomes. From selected target genomic regions with high discrimination power, the analytics system can design probes to target fragments from the selected genomic regions. The analytics system can generate cancer assay panels of various sizes, where, for example, a small-sized cancer assay panel includes probes targeting the most informative genomic regions, a medium-sized cancer assay panel includes probes from the small-sized cancer assay panel and additional probes targeting genomic regions of a second tier of informativeness, and a large-sized cancer assay panel includes probes from the small-sized and medium-sized cancer assay panels along with more probes targeting genomic regions of a third tier of informativeness. Using data obtained from such cancer assay panels (e.g., methylation states of nucleic acids derived from the cancer assay panel), the analytics system can train classifiers using various classification techniques to predict the likelihood that a sample has a particular outcome or condition, such as cancer, a specific cancer type, other disorder, other disease, etc.

[0101] In one illustrative example, to design a cancer assay panel, an analytics system may collect information regarding the methylation status of CpG sites of nucleic acid fragments from samples corresponding to various outcomes under consideration, e.g., samples known to have cancer, samples believed to be healthy, samples from known TOO, etc. The methylation status of CpG sites may be determined by processing these samples (e.g., by whole genome bisulfite sequencing (WGBS)), or this information may be obtained from TCGA. In some embodiments, the analytics system is a computing system including a computer processor and a computer-readable storage medium containing instructions for causing the computer processor to perform some or all of the operations described in this disclosure.

[0102] The analytics system can then select target genomic regions based on the methylation patterns of the nucleic acid fragments. One approach considers the pairwise discrimination between pairs of outcomes for a region (or more specifically, for a CpG site). Another approach considers the discrimination for a region (or more specifically, for a CpG site) when each outcome is considered relative to the remaining outcomes. From selected target genomic regions with high discrimination power, the analytics system can design probes to target fragments from the selected genomic regions. The analytics system can generate cancer assay panels of various sizes, where, for example, a small-sized cancer assay panel includes probes targeting the most informative genomic regions, a medium-sized cancer assay panel includes probes from the small-sized cancer assay panel and additional probes targeting genomic regions of a second tier of informativeness, and a large-sized cancer assay panel includes probes from the small-sized and medium-sized cancer assay panels along with more probes targeting genomic regions of a third tier of informativeness. Using such cancer assay panels, the analytics system can train classifiers with various classification techniques to predict the likelihood that a sample has a particular outcome or condition, such as cancer, a specific cancer type, other disorder, other disease, etc.

[0103] In some embodiments, the cancer assay panel comprises a plurality of probe pairs, wherein each pair of the plurality of pairs comprises two probes that are configured to overlap each other with overlapping sequences, wherein the overlapping sequences comprise at least 30 nucleotides, and wherein each probe is configured to hybridize to the same strand of (optionally converted) DNA molecules (e.g., cfDNA molecules) corresponding to one or more genomic regions.In some embodiments, each probe of the two probes in each probe pair comprises a sequence of more than 20, 25, 30, 35, 40, 45, or 50 nucleotides that does not overlap with other probes in the probe pair. Thus, in some embodiments, the first and second probes in a probe pair may overlap by 30 nucleotides, where the first probe further comprises more than 20, 25, 30, 35, 40, 45, or 50 nucleotides that do not overlap with the second probe, and where the second probe further comprises more than 20, 25, 30, 35, 40, 45, or 50 nucleotides that do not overlap with the first probe. In some embodiments, each of the genomic regions comprises at least five methylation sites, and where the at least five methylation sites have an abnormal methylation pattern in the cancerous specimen or have a different methylation pattern between specimens of different TOO. For example, in one embodiment, the at least five methylation sites are either differentially methylated between cancerous and non-cancerous specimens, or differentially methylated between one or more pairs of specimens from cancers of different tissues of origin. In some embodiments, each probe pair comprises a first probe and a second probe, where the second probe is different from the first probe. The second probe may overlap the first probe by an overlap sequence at least 30, at least 40, at least 50, or at least 60 nucleotides in length.

[0104] In some embodiments, each of the plurality of different bait oligonucleotides in the panel is at least 45 nucleotides in length (e.g., at least 60, 75, 80, 90, 100, 110, or 120 nucleotides in length). In some embodiments, each bait oligonucleotide in the plurality of probes is no more than 130, 140, 150, 200, 250, or 300 bases in length. In some embodiments, each bait oligonucleotide is 45-300, 60-200, or 75-150 nucleotides in length. In some embodiments, a bait oligonucleotide is at least 50 nucleotides in length. In some embodiments, a bait oligonucleotide is at least 60 nucleotides in length. In some embodiments, a bait oligonucleotide is at least 75 nucleotides in length. In some embodiments, the designated length of a bait oligonucleotide is a length designed to be complementary to a portion of a target genome sequence or its converted DNA molecule.

[0105] In some embodiments, the cancer assay panel is designed to target at least 5,000, 10,000, 12,500, 15,000, 17,000, 19,000, or more target genomic regions. In some embodiments, the cancer assay panel is designed to target 25,000, 20,000, 17,000, 15,000, 12,500, 10,000, or fewer target genomic regions. In some embodiments, the cancer assay panel is designed to target 5,000-30,000, 10,000-25,000, 12,500-20,000, or 15,000-20,000 target genomic regions. In some embodiments, the cancer assay panel is designed to target at least 10,000 target genomic regions. In some embodiments, the cancer assay panel is designed to target at least 15,000 target genomic regions. In some embodiments, the cancer assay panel is designed to target fewer than 20,000 target genomic regions.

[0106] In some embodiments, bait oligonucleotides are configured to hybridize to converted DNA molecules (e.g., converted cfDNA molecules) corresponding to or derived from one or more genomic regions. Thus, the bait oligonucleotides can have a sequence different from that of the targeted genomic region. For example, DNA with unmethylated CpG sites can be converted to contain UpG instead of CpG by deamination (e.g., by treatment with cytosine deaminase or bisulfite). As a result, probes for such targets can be configured to hybridize to sequences containing UpG instead of naturally occurring unmethylated CpG. Thus, the site complementary to the unmethylated site in the probe may contain CpA instead of CpG, and some probes targeting hypomethylated sites where all methylated sites are unmethylated may lack a guanine (G) base. In some embodiments, at least 3%, 5%, 10%, 15%, or 20% of the probes do not contain CpG sequences. In some embodiments, at least 5% of the probes do not contain CpG sequences. In some embodiments, at least 10% of the probes do not contain CpG sequences.

[0107] In some embodiments, cancer assay panels are used to detect the presence or absence of cancer generally and / or provide a cancer classification, such as cancer type, cancer stage, such as I, II, III, or IV, or to provide a potential TOO of cancer origin. The panel may include probes targeting genomic regions that are differentially methylated between cancerous (pan-cancer) and non-cancerous specimens in general, or only in cancerous specimens of a specific cancer type (e.g., bladder cancer-specific targets). For example, in some embodiments, cancer assay panels are designed to include genomic regions that are differentially methylated based on converted (e.g., bisulfite) sequencing data generated from cfDNA and / or whole genome DNA of a collection of cancer and non-cancer individuals.

[0108] In some embodiments, each of the target genomic regions is differentially methylated in at least one of the plurality of cancer types. In some embodiments, the plurality of cancer types includes at least two cancer types (e.g., at least 2, 3, 4, or more cancer types). In some embodiments, the plurality of cancer types includes at least two cancer types. In some embodiments, the plurality of cancer types includes at least three cancer types. In some embodiments, the plurality of cancer types includes urological cancer. In some embodiments, the plurality of cancer types includes one or more of bladder cancer, urothelial cancer, renal cancer, or prostate cancer.

[0109] Each probe (or probe pair) can be designed to target one or more target genomic regions. The target genomic regions can be selected based on several criteria designed to increase the selective enrichment of informative cfDNA fragments (e.g., cfDNA fragments from urine specimens) while reducing noise and non-specific binding. Various filtering procedures are described herein to determine whether to include a target genomic region. In some embodiments, two or more of the filtering procedures described herein are used in combination.

[0110] In one example, a panel can include probes that selectively bind to and enrich for differentially methylated cfDNA fragments in cancerous specimens. Sequencing the enriched fragments can then provide information relevant to cancer detection. Furthermore, in some embodiments, probes (or portions thereof) are designed to target genomic regions determined to have abnormal or aberrant methylation patterns in cancer specimens, or in specimens from certain cancer types, tissue types, or cell types. In one embodiment, probes are designed to target genomic regions determined to be hypermethylated or hypomethylated in certain cancers or cancer types to provide additional selectivity and specificity of detection. In some embodiments, a panel includes probes that target hypomethylated fragments. In some embodiments, a panel includes probes that target hypermethylated fragments. In some embodiments, a panel includes both a first set of probes that target hypermethylated fragments and a second set of probes that target hypomethylated fragments. In some embodiments, a cancer assay panel includes probes designed to target a region having a first methylation state (e.g., hypomethylated) as well as probes designed to hybridize to the same target region in the opposite methylation state (e.g., hypermethylated). Targeting probes to both hypomethylated and hypermethylated fragments from the same region can be referred to as "binary" targeting (see, e.g., Figure 11C). In some embodiments, the ratio between a first set of probes targeting hypermethylated fragments and a second set of probes targeting hypomethylated fragments (high:low ratio) ranges from 0.4 to 2, 0.5 to 1.5, 0.5 to 1.0, 0.8 to 1, 0.6 to 0.8, or 0.4 to 0.6. In some embodiments, the high:low ratio is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, or more. In some embodiments, the high:low ratio is at least 5.Methods for identifying genomic regions (e.g., genomic regions that give rise to differentially methylated or aberrantly methylated DNA molecules) between cancer and non-cancerous specimens, between different tissue of origin (TOO) types, between different cancer cell types, or between specimens from different stages of cancer are provided in detail herein, as are methods for identifying aberrantly methylated DNA molecules or fragments that are identified as indicative of cancer.

[0111] In a second example, genomic regions can be selected when they give rise to aberrantly methylated DNA molecules in cancer specimens or specimens of known cancer tissue of origin (TOO) type. For example, as described herein, a Markov model trained on a collection of non-cancerous specimens can be used to identify genomic regions that give rise to aberrantly methylated DNA molecules (e.g., DNA molecules with methylation patterns below a p-value threshold).

[0112] In some embodiments, each of the probes targets a genomic region comprising at least 30 bp, 35 bp, 40 bp, 45 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 100 bp, 110 bp, 120 bp, or more. In some embodiments, each of the probes targets a genomic region comprising 120 bp. In some embodiments, the genomic region can be selected to have fewer than 30, 25, 20, 15, 12, 10, 8, or 6 methylation sites. In some embodiments, the genomic region can be selected to have at least 3, 5, or 7 methylation sites. In some embodiments, each target genomic region comprises at least 5 methylation sites. In some embodiments, the selected number of methylation sites (e.g., at least 5 methylation sites) are differentially methylated in at least one cancer type to be assayed by the panel.

[0113] In some embodiments, a genomic region is selected as a target when at least 80, 85, 90, 92, 95, or 98% of at least five methylated (e.g., CpG) sites within the region are either methylated or unmethylated in a non-cancerous specimen or a cancerous specimen, or in a cancer specimen from a tissue of origin (TOO).

[0114] In some embodiments, target genomic regions are filtered to select only those likely to be informative based on their methylation patterns, e.g., CpG sites that are differentially methylated between cancerous and non-cancerous specimens, between a cancerous specimen of one TOO and a cancerous specimen of another TOO (e.g., normal methylation or unmethylated in cancer compared to non-cancer), or CpG sites that are differentially methylated only in cancerous specimens of one TOO. For selection, calculations can be performed for each CpG or for multiple CpG sites. For example, a first count (cancer_count) can be determined, which is the number of cancer-containing specimens containing fragments overlapping with the CpG, and a second count (total) is determined, which is the total number of specimens containing fragments overlapping with the CpG site. Genomic regions can be selected based on a criterion that is positively correlated with the number of cancer-containing specimens containing cancer-indicating fragments overlapping with the CpG site (cancer_count) and inversely correlated with the total number of specimens containing cancer-indicating fragments overlapping with the CpG site (total). In one embodiment, the number of non-cancerous specimens having fragments overlapping the CpG site (n 非癌 ) and the number of cancerous specimens (n 癌 ) is counted. The probability that the sample is cancer is then calculated, say, (n 癌 +1) / (n 癌 +n 非癌 +2). This principle may be applicable to other outcomes as well.

[0115] CpG sites scored by this metric can be ranked and greedily added to the panel until the panel size budget is exhausted. The process of selecting cancer-indicative genomic regions is described in further detail herein. In some embodiments, different target regions may be selected depending on whether the assay is intended to be a pan-cancer or single-cancer assay, or depending on what kind of flexibility is desired when selecting which CpG sites to contribute to the panel. Panels for detecting specific cancer types can also be designed using a similar process. In this embodiment, for each cancer type and for each CpG site, an information gain is computed to determine whether to include a probe targeting that CpG site. The information gain may be computed for a given number of samples of a given cancer type compared to all other samples. For example, consider two random variables, "AF" and "CT." "AF" is a binary variable that indicates whether a particular sample has an abnormal fragment overlapping a particular CpG site (yes or no). "CT" is a binary random variable that indicates whether the cancer is a particular type of cancer (e.g., bladder cancer or non-bladder cancer). The mutual information for "CT" given "AF" can be computed—that is, how many bits of information are gained about the cancer type (in this example, bladder vs. non-bladder) when it is known whether there is an abnormal fragment overlapping a particular CpG site. This can be used to rank CpGs based on how bladder-specific they are. This procedure is repeated for multiple cancer types. If a particular region is commonly differentially methylated only in bladder cancer (and not in other cancer types or non-cancer), CpGs in that region may tend to have high information gain for bladder cancer. For each cancer type, CpG sites are ranked by this information gain metric and then greedily added to the panel until the size budget for that cancer type is exhausted. In some embodiments, information gain is also measured relative to the likelihood of pulling down a molecule containing the target sequence and / or the cost of sequencing that target.In some embodiments, a target region is included only if the ratio of information gain to cost of sequencing that target exceeds a threshold.

[0116] In some embodiments, filtering is performed to select probes with high specificity (i.e., high binding efficiency) for enriching nucleic acids derived from targeted genomic regions. Probes can be filtered to reduce non-specific binding (or off-target binding) to nucleic acids derived from non-target genomic regions. For example, probes can be filtered to select only probes with fewer off-target binding events than a set threshold (such as may be predicted by the number of potential binding sites across the genome with a specified degree of complementarity in a given alignment window). In one embodiment, probes can be aligned to a reference genome (e.g., a human reference genome) to select probes that align with fewer regions across the genome than a set threshold. In some embodiments, the reference genome is the human reference genome GRCh37 / hg19, the sequence of which is available from the Genome Reference Consortium and from genome browsers provided by the University of Santa Cruz Genomics Institute. For example, probes can be selected that align with fewer than 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, or 8 off-target regions across the reference genome. In other cases, filtering is performed to remove a target genomic region when the sequence of the target genomic region appears more than 5, 10, 15, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 times in the genome.Further filtering can be performed to select a target genome region when a probe sequence, or set of probe sequences, that is 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% complementary to the target genome region appears less than 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, or 8 times in the reference genome, or when a probe sequence or set of probe sequences that is 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% complementary to the target genome region appears less than 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, or 8 times in the reference genome. A probe sequence or set of probe sequences designed to enrich for a targeted genomic region that is 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% complementary to the region can be removed if the target genomic region appears more than 5, 10, 15, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 times in the reference genome to eliminate repetitive probes that may pull down unwanted off-target fragments that can affect assay efficiency.

[0117] Some experiments have demonstrated that a fragment-probe overlap of at least 45 bp is effective for achieving a significant amount of pull-down (although this number may vary). Under some conditions, a mismatch rate of more than 10% between the probe and fragment sequences in the overlap region significantly disrupts binding and is therefore sufficient for pull-down efficiency. Therefore, sequences that can align with the probe at a match rate of at least 90% along at least 45 bp can be candidates for off-target pull-down. Thus, in one embodiment, the number of such regions is scored. In some embodiments, the best probes have a score of 1, meaning that these probes match only at one location (the intended target region). Probes with intermediate scores (i.e., less than 5 or 10) may be accepted in some cases, and in some cases, any probes above a certain score are discarded. In some embodiments, a candidate probe is excluded from the assay panel if it contains a contiguous 45-bp sequence that is at least 90% complementary to more than 20 off-target sites. Other cutoff values ​​may be used for specific samples.

[0118] Once the probes hybridize to and capture DNA fragments corresponding to or derived from the target genomic region, the hybridized probe-DNA fragment intermediates are pulled down (or isolated), and the target DNA is amplified and sequenced. The sequence reads provide information relevant to cancer detection. To this end, panels can be designed to include multiple probes capable of capturing fragments that together can provide information relevant to cancer detection. In some embodiments, the panel includes at least 10, 20, 30, 40, 50, 60, 80, 100, 150, 200, 300, 500, 1,000, 2,000, 2,500, 5,000, 6,000, 7,500, 10,000, 15,000, 20,000, 25,000, or 30,000 different probe pairs. In some embodiments, the panel includes at least 20, 30, 40, 60, 80, 100, 120, 160, 200, 300, 400, 600, 1,000, 2,000, 5,000, 6,000, 10,000, 12,000, 15,000, 20,000, 30,000, 40,000, 50,000, or 60,000 different probes. The plurality of different probes may collectively include at least 2,000, 5,000, 10,000, 20,000, 40,000, 60,000, 80,000, 100,000, 200,000, 400,000, 600,000, 800,000, 1 million, 2 million, 3 million, 4 million, or 5 million nucleotides. In some embodiments, the panel includes no more than 20, 30, 40, 50, 60, 80, 100, 150, 200, 300, 500, 1,000, 2,000, 2,500, 5,000, 6,000, 7,500, 10,000, 15,000, 20,000, 25,000, or 30,000, or 40,000 different probe pairs. In other embodiments, the panel includes no more than 30, 40, 60, 80, 100, 120, 160, 200, 300, 400, 600, 1,000, 2,000, 5,000, 6,000, 10,000, 12,000, 15,000, 20,000, 30,000, 40,000, 50,000, 60,000, or 70,000 different probes.

[0119] The selected target genomic region can be located in various locations within the genome, including, but not limited to, promoters, enhancers, exons, introns, intergenic regions, and other sites. In some instances, primers can be used to specifically amplify the target / biomarker of interest (e.g., by PCR), thereby enriching the sample for the desired target / biomarker (optionally without hybridization capture). For example, forward and reverse primers can be prepared for each genomic region of interest and used to amplify fragments corresponding to or derived from the desired genomic region. Thus, while the present disclosure pays particular attention to cancer assay panels and bait sets for hybridization capture, the present disclosure is broad enough to encompass other methods for enriching cell-free nucleic acid molecules (e.g., cfDNA). Thus, those skilled in the art, with the benefit of this disclosure, will recognize that methods similar to those described herein with respect to hybridization capture may alternatively be achieved by replacing hybridization capture with some other enrichment strategy, such as PCR amplification of cell-free DNA fragments corresponding to the genomic region of interest. In some embodiments, regions of interest are enriched using bisulfite padlock probe capture, such as that described by Zhang et al. (U.S. Patent Application Publication No. 2016 / 0340740). In some embodiments, enrichment (e.g., non-targeted enrichment) uses additional or alternative methods, such as reduced representation bisulfite sequencing, methylated restriction enzyme sequencing, methylated DNA immunoprecipitation sequencing, methyl-CpG-binding domain protein sequencing, methyl-DNA capture sequencing, or microdroplet PCR.

[0120] probe The cancer assay panels (alternatively referred to as "bait collections") provided herein can be panels including a collection of hybridization probes (also referred to herein as "probes") designed to target and pull down nucleic acid fragments of interest in the assay during enrichment. In some embodiments, the probes are designed to hybridize to and enrich DNA or cfDNA molecules from cancerous specimens that have been treated to convert unmethylated cytosines (C) to uracil (U). In other embodiments, the probes are designed to hybridize to and enrich DNA or cfDNA molecules from a TOO (or multiple TOO) cancerous specimens that have been treated to convert unmethylated cytosines (C) to uracil (U). The probes can be designed to anneal (or hybridize) to a target (complementary) strand of DNA or RNA. The target strand can be the "plus" strand (e.g., the strand that is transcribed into mRNA and subsequently translated into protein) or the complementary "minus" strand. In particular embodiments, a cancer assay panel may include a set of two probes, one probe targeting the plus strand and the other probe targeting the minus strand of a target genomic region.

[0121] At least four possible probe sequences can be designed for each target genomic region. Each target region is double-stranded, so a probe or probe set can target either the "plus" or forward strand or its reverse complement (the "minus" strand). Additionally, in some embodiments, a probe or probe set is designed to enrich for DNA molecules or fragments that have been treated to convert unmethylated cytosines (C) to uracils (U). For probes or probe sets designed to enrich for DNA molecules corresponding to or derived from a converted target region, the probe sequence can be designed to enrich for DNA molecules in fragments in which unmethylated Cs have been converted to Us (by utilizing A instead of G at the site of unmethylated cytosines in DNA molecules or fragments corresponding to or derived from the target region). In one embodiment, a probe is designed to bind to or hybridize to DNA molecules or fragments from genomic regions known to have cancer-specific methylation patterns (e.g., hypermethylated or hypomethylated DNA molecules), thereby enriching for cancer-specific DNA molecules or fragments. Targeting genomic regions or cancer-specific methylation patterns can be advantageous because it allows for specific enrichment of DNA molecules or fragments identified as informative for cancer or cancer TOO, thereby reducing sequencing needs and costs. In other embodiments, two probe sequences (one for each DNA strand) can be designed per target genomic region. In still other cases, probes are designed to enrich all DNA molecules or fragments corresponding to or derived from the targeted region (i.e., regardless of strand or methylation status). This may be because the cancer methylation status is highly unmethylated or unmethylated, or because the probes are designed to target small mutations or other alterations rather than methylation changes; these other variations are also indicative of the presence or absence of cancer or the presence or absence of one or more TOOs. In that case, all four possible probe sequences can be included per target genomic region.

[0122] In some embodiments, the probes are in the 10s, 100s, 200s, or 300s of base pairs in length. The probes may comprise at least 50, 75, 100, or 120 nucleotides. The probes may comprise less than 300, 250, 200, or 150 nucleotides. In certain embodiments, the probes comprise 100-150 nucleotides. In one particular embodiment, the probes comprise 120 nucleotides.

[0123] In some embodiments, the probes are designed to cover overlapping portions of the target region in a "2x tiling" fashion. Each probe optionally at least partially overlaps with another probe in the library. In such embodiments, the panel includes multiple probe pairs, where each probe in a pair overlaps with the other by at least 25, 30, 35, 40, 45, 50, 60, 70, 75, or 100 nucleotides. In some embodiments, the overlapping sequences can be designed to be complementary to the target genomic region (or cfDNA derived therefrom) or to sequences homologous to the target region or cfDNA. Thus, in some embodiments, at least two probes contain sequences complementary to the same sequence in the target genomic region, allowing at least one of the probes to bind to and pull down nucleotide fragments corresponding to or derived from the target genomic region. For a given probe pair that includes overlapping sequences, the pair may include non-overlapping sequences complementary to the target genomic region extending from separate ends of the overlapping sequences (see, e.g., Figure 11E). Other tiling levels are possible, such as 3x tiling, 4x tiling, etc., in which each nucleotide in the target region may be bound to three or more probes.

[0124] In one embodiment, as shown in Figure 11B, exactly two probes overlap one base of the target genomic region. Probes that extend beyond the target genomic region in both directions are useful for pulling down cfDNA fragments that contain a portion of the target genomic region and DNA sequences adjacent to the target genomic region. In some instances, even a relatively small target region can be targeted with three probes (see, e.g., Figure 11A). Optionally, a probe set containing three or more probes is used to capture a larger genomic region (see, e.g., Figure 11B). In some embodiments, a subset of probes collectively spans the entire target genomic region (e.g., can be complementary to unconverted or converted fragments from the entire genomic region). Optionally, a tiled probe set includes probes that collectively include at least two probes that collectively overlap every nucleotide of the target genomic region. This is done to ensure that cfDNA containing a small portion of the target genomic region at one end has substantial overlap extending to adjacent non-target genomic regions with at least one probe, resulting in efficient capture.

[0125] In some embodiments, each target genome region is targeted by a set of probes.The set of probes is designed in a tiling manner, so that adjacent probes have overlapping sequences that hybridize to the same part of genome region (see Figure 11D).Because DNA has two strands, the set of probes can also include the overlapping probe that hybridizes to the other strand, so that a total of four probes hybridize to the same part of genome region.

[0126] In some embodiments, a set of probes configured to hybridize to a target genomic region does not span the entire region—i.e., at least some sequences within the target genomic region do not have a corresponding probe. For example, a sequence within the target genomic region may be similar or identical to many other sequences in the genome, and such a probe would hybridize to more than a threshold number of off-target regions, so a probe is not designed to target that sequence.

[0127] For example, a 100-bp cfDNA fragment containing a 30-nt target genomic region will have at least 65 bp of overlap with at least one of the overlapping probes. Other tiling levels are possible. For example, to increase target size and add more probes to the panel, probes can be designed to extend at least 70 bp, 65 bp, 60 bp, 55 bp, or 50 bp into the 30-bp target region. To capture any fragments that overlap the target region (even by as little as 1 bp), probes can be designed to extend beyond the ends of the target region by at least 50 bp, 55 bp, 60 bp, 65 bp, 70 bp, 75 bp, 80 bp, or 85 bp, etc., on either side. Probes can be designed to extend 75 bp beyond the ends of the target region on either side. In some embodiments, the presence of probes designed to extend beyond the ends of the target genomic region does not increase the size of the target genomic region (e.g., they are not included in determining the size of each target genomic region or the size of multiple target genomic regions collectively).

[0128] In some embodiments, each bait oligonucleotide is conjugated to a solid surface (e.g., a chip or a bead, such as a magnetic or paramagnetic bead) or to a non-nucleotide affinity moiety (e.g., a member of a binding pair). In some embodiments, such conjugates facilitate the separation of DNA molecules bound to the bait oligonucleotide from unbound DNA molecules. Generally, a "binding pair" refers to one of a first and a second moiety, where the first and second moieties have specific binding affinity for each other. Non-limiting examples of binding pairs include antigen / antibody; biotin / avidin (or biotin / streptavidin); calmodulin-binding protein (CBP) / calmodulin; hormone / hormone receptor; lectin / carbohydrate; peptide / cell membrane receptor; enzyme / cofactor; and enzyme / substrate. In some embodiments, the affinity moiety is biotin.

[0129] In some embodiments, the target genomic region is selected so that applying the trained classifier to the sequence of the converted DNA molecule captured by the bait oligonucleotide distinguishes subjects with cancer from subjects without cancer with a predetermined specificity and / or sensitivity. Disclosed herein are methods for selecting target regions, training classifiers, and selected specificities and sensitivities useful in detecting various cancer types, such as in connection with various aspects of the methods herein. In some embodiments, the classifier is a binary classifier, a mixed model classifier, or a multilayer perceptron model classifier. In some embodiments, the predetermined specificity for each of the multiple cancer types is 0.900 or greater (e.g., at least 0.950, 0.975, 0.980, 0.985, 0.990, 0.995, or greater). In some embodiments, applying the trained classifier includes a sensitivity of at least 30% (e.g., at least 40%, 50%, 60%, 70%, 80%, or greater) for each of the multiple cancer types.

[0130] In some embodiments, the target genomic region is selected from the sequences of the genes in Table 1. In some embodiments, the cancer assay panel comprises a plurality of probes, wherein each of the plurality of probes is configured to hybridize to a DNA molecule (e.g., a converted DNA molecule) or amplicon thereof corresponding to one or more genomic regions comprising a target sequence from one or more genes in Table 1. In some embodiments, the cancer assay panel comprises a plurality of probes, wherein each of the plurality of probes is configured to hybridize to a DNA molecule (e.g., a converted DNA molecule) or amplicon thereof corresponding to at least 1, 2, 3, 4, 5, 10, 15, 20, 25, or more target sequences of one or more genes selected from Table 1. In some embodiments, the target genomic region comprises at least 10 target sequences of one or more genes from Table 1. In some embodiments, the target genomic region comprises at least 15 target sequences of one or more genes from Table 1. The target sequences may include two or more target sequences from a single gene and / or may include one or more target sequences from two or more genes. In some embodiments, the target genomic region comprises a target sequence of a gene selected from TWIST1, EOMES, HOXA9, POU4F2, and ZNF154. In some embodiments, the target genomic region comprises a target sequence of the TWIST1 gene. In some embodiments, the target genomic region comprises a target sequence from at least 20% of the genes in Table 1. In some embodiments, the target genomic region comprises a target sequence from at least 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100% of the genes from Table 1. In some embodiments, the target genomic region comprises a sequence from at least 50% of the genes from Table 1. In some embodiments, the target genomic region comprises a target sequence from all of the genes in Table 1. In some embodiments, the target sequence is at least 10, 25, 35, 45, 50, 75, 100, 200, 300, or more nucleotides in length. In some embodiments, the target sequence is at least 25, at least 35, or at least 45 nucleotides in length. In some embodiments, the target sequence is at least 45 nucleotides in length.In some embodiments, each target genomic region comprises at least five methylation sites. In some embodiments, differential methylation comprises at least 80% of the CpG sites in the target genomic region being methylated or unmethylated.

[0131] Methods for selecting target genomic regions In one aspect, the present disclosure provides a method for selecting a target genomic region for detecting cancer and / or TOO. In some embodiments, the target genomic region exhibits abnormal or aberrant methylation in cancer and / or TOO. The targeted genomic region can be used to design and manufacture probes for a cancer assay panel. The cancer assay panel can be used to screen the methylation status of DNA or cfDNA molecules corresponding to or derived from the target genomic region. Alternative methods, such as WGBS or other methods, can also be used to detect the methylation status of DNA molecules or fragments corresponding to or derived from the target genomic region.

[0132] Specimen processing 16A is a flowchart of a process for processing a nucleic acid specimen and generating methylation state vectors for DNA fragments, according to one embodiment. The method includes, but is not limited to, the following steps. For example, any step of the method may include a quantification substep for quality control or other laboratory assay procedures known to those skilled in the art.

[0133] In step 105, a urine specimen containing nucleic acids (e.g., cfDNA) is collected from a subject. The specimen can be any subset of the human genome, including the entire genome. The urine specimen can be prepared by any of the methods for collecting nucleic acids from urine disclosed herein. FIG. 1 illustrates an exemplary method for urine specimen processing. In this detailed embodiment, urine is collected from a subject, treated with a cell lysis inhibitor, and centrifuged to precipitate and remove cellular debris. The urine specimen is then concentrated by filtration to obtain a concentrated urine specimen with increased nucleic acid concentra- tion and reduced volume compared to the untreated urine specimen. This concentrated urine specimen may then be subjected to nucleic acid extraction and library preparation, which can then be sequenced for methylation analysis and cancer (e.g., bladder, renal, or prostate cancer) detection. The extracted specimen may contain cfDNA and / or ctDNA. In healthy individuals, the human body may naturally remove cfDNA and other cellular debris. If the subject has cancer or disease, cfDNA and / or ctDNA may be present in the specimen at levels detectable for the detection of the cancer or disease.

[0134] In step 110, nucleic acids are treated to convert unmethylated cytosines to uracil. In one embodiment, the method uses bisulfite treatment of DNA to convert unmethylated cytosines to uracil without converting methylated cytosines. Bisulfite conversion can be performed using a commercially available kit, such as the EZ DNA Methylation™-Gold, EZ DNA Methylation™-Direct, or EZ DNA Methylation™-Lightning kit (available from Zymo Research Corp, Irvine, CA). In another embodiment, the conversion of unmethylated cytosine to uracil is achieved using an enzymatic reaction. For example, this conversion can be performed using a commercially available kit for converting unmethylated cytosine to uracil, such as APOBEC-Seq (NEBiolabs, Ipswich, MA).

[0135] In step 115, a sequencing library is prepared. In some embodiments, ssDNA adapters are added to the 3'-OH ends of the bisulfite-converted ssDNA molecules using an ssDNA ligation reaction. In one embodiment, the ssDNA ligation reaction uses CircLigase II (Epicentre) to ligate ssDNA adapters to the 3'-OH ends of the bisulfite-converted ssDNA molecules, where the 5'-end of the adapter is phosphorylated and the bisulfite-converted ssDNA is dephosphorylated (i.e., the 5' phosphate is removed). In another embodiment, the ssDNA ligation reaction uses Thermostable 5' AppDNA / RNA Ligase (available from New England BioLabs, Ipswich, MA) to ligate ssDNA adapters to the 3'-OH ends of the bisulfite-converted ssDNA molecules. In this example, the first adapter is adenylated at the 5'-end and blocked at the 3'-end. In another embodiment, a ssDNA ligation reaction uses T4 RNA ligase (available from New England BioLabs) to ligate a ssDNA adapter to the 3'-OH end of the bisulfite-converted ssDNA molecule. In a second step, a second-strand DNA is synthesized in an extension reaction. For example, a primer extension reaction uses an extension primer that hybridizes to a primer sequence contained in the ssDNA adapter to form a double-stranded bisulfite-converted DNA molecule. Optionally, in one embodiment, the extension reaction uses an enzyme capable of reading uracil residues in the bisulfite-converted template strand. Optionally, in a third step, a dsDNA adapter is added to the double-stranded bisulfite-converted DNA molecule. Finally, a sequencing adapter is added by amplification of the double-stranded bisulfite-converted DNA. For example, P5 and P7 sequences are added to the bisulfite-converted DNA using PCR amplification with a forward primer containing the P5 sequence and a reverse primer containing the P7 sequence.Optionally, during library preparation, nucleic acid molecules (e.g., DNA molecules) may be tagged with unique molecular identifiers (UMIs) by adapter ligation. UMIs are short nucleic acid sequences (e.g., 4-10 base pairs) that are added to the ends of DNA fragments during adapter ligation. In some embodiments, UMIs contain degenerate base pair positions that serve as unique tags that can be used to identify sequence reads originating from a particular DNA fragment. During PCR amplification after adapter ligation, UMIs replicate with the DNA fragments to which they are attached, thereby providing a means for identifying sequence reads originating from the same fragment in downstream analyses (either by the UMI alone or by using the UMI in combination with a portion of the sample nucleic acid fragment end sequence, such as the first 2-10 nucleotides).

[0136] In step 120, target DNA sequences may be enriched from the library. This is used, for example, when a target panel assay is to be performed on the specimen. During enrichment, hybridization probes (also referred to herein as "probes" or "bait oligonucleotides") are used to target and, optionally, pull down nucleic acid fragments that are informative regarding the presence or absence of cancer (or disease), cancer status, or cancer classification (e.g., cancer type or tissue of origin). In some embodiments, the probes have the characteristics specified herein, such as in connection with various other aspects described herein. For a given workflow, the probes may be designed to anneal (or hybridize) to a target (complementary) strand of DNA (e.g., a converted DNA molecule). The target strand can be the "plus" strand (e.g., the strand transcribed into mRNA and subsequently translated into protein) or the complementary "minus" strand. Probes can range in length from 10s, 100s, or 1000s of base pairs. Moreover, the probes may cover overlapping portions of the target regions as described herein.

[0137] After the hybridization step 120, the hybridized nucleic acid fragments may be concentrated (e.g., by capturing or otherwise separating from unbound nucleic acids) and amplified using PCR (enrichment 125). For example, the target sequences may be enriched to obtain enriched sequences, which may then be sequenced. In general, various methods can be used to isolate and enrich the target nucleic acid to which the probe is hybridized. For example, adding a biotin moiety to the 5'-end of the probe (i.e., biotinylation) allows for easy isolation of the target nucleic acid hybridized to the probe using a streptavidin-coated surface (e.g., streptavidin-coated beads).

[0138] In step 130, sequence reads are generated from the enriched nucleic acid fragments. Sequencing data can be obtained from the enriched DNA sequences by means known in the art. For example, such methods can include next-generation sequencing (NGS) techniques involving synthesis (Illumina), pyrosequencing (454 Life Sciences), ion semiconductor technology (Ion Torrent sequencing), single-molecule real-time sequencing (Pacific Biosciences), sequencing-by-ligation (SOLiD sequencing), nanopore sequencing (Oxford Nanopore Technologies), or paired-end sequencing. In some embodiments, massively parallel sequencing is performed using sequencing-by-synthesis with reversible dye terminators.

[0139] In step 140, a methylation state vector is generated from the sequence reads. To do so, the sequence reads are aligned to a reference genome. The reference genome helps provide context as to which location in the human genome the DNA fragment (e.g., cfDNA) originated from. In a simplified example, the sequence reads are aligned so that three CpG sites correlate with CpG sites 23, 24, and 25 (arbitrary reference identification numbers are used for ease of explanation). After alignment, there is information about both the methylation state of every CpG site on the cfDNA fragment and where those CpG sites map in the human genome. With the methylation state and location, a methylation state vector can be generated for the DNA fragment.

[0140] Alternative methods for detecting methylation patterns can include the use of oligonucleotide microarrays and selective PCR amplification. The method for detecting methylation patterns can include the use of oligonucleotide microarrays. The method can include labeling cfDNA fragments with a label. The cfDNA fragments can be from an individual. The cfDNA fragments can be bisulfite-converted ssDNA molecules or converted cfDNA fragments. The label can be a fluorescent label. The fluorescent label can be a cyanine dye. The cyanine dye can be Cy3, Cy5, DY547, or DY647. The oligonucleotide microarray for use in detection can include at least 75, 150, 300, or 1000 pairs of bait oligonucleotides. In some embodiments, each pair of bait oligonucleotides includes a first bait and a second bait. The first bait can include an overlapping sequence and a first non-overlapping sequence. The second bait can include an overlapping sequence and a second non-overlapping sequence. The overlapping sequence may comprise at least 30, 40, 50, or 60 nucleotides. The first non-overlapping sequence and the second non-overlapping sequence may comprise more than 30, 40, 50, or 60 nucleotides. The first bait and the second bait of each pair of bait oligonucleotides may be configured to hybridize to converted cfDNA fragments derived from a genomic region containing at least five methylation sites that are differentially methylated in cfDNA fragments from individuals with cancer compared to cfDNA fragments from individuals without cancer. The converted cfDNA fragments may be labeled with a first fluorescent label. In some embodiments, the reference cfDNA fragments may be labeled with a second fluorescent label. The reference cfDNA fragments may be cfDNA fragments with a known methylation pattern. The second fluorescent label may be a cyanine dye. The cyanine dye may be Cy3, Cy5, DY547, or DY647. The first fluorescent label and the second fluorescent label may be different. In some embodiments, the methylation pattern is detected by contacting labeled cfDNA fragments from an individual and labeled reference cfDNA fragments with an oligonucleotide microarray.The methylation pattern can be determined by comparing the amount of the first fluorescent label and the second fluorescent label associated with each array location on the microarray. Each array location can correspond to a bait oligonucleotide pair. In some embodiments, the methylation pattern is associated with cancer or a cancer type. The method can include applying a trained classifier to predict the likelihood of cancer. In some embodiments, applying the trained classifier to sequences of converted cfDNA molecules hybridized to the bait oligonucleotides can distinguish subjects with cancer or a cancer type from subjects without cancer with a specificity of 0.990 or 0.994 and a sensitivity of at least 40%, at least 45%, or at least 50%.

[0141] The method for detecting methylation patterns may include the use of selective PCR amplification. A composition including multiple oligonucleotide pairs is configured to hybridize to converted DNA fragments derived from target genomic regions, with each oligonucleotide pair configured for selective PCR amplification of a sequence. In some embodiments, the sequence for selective PCR amplification generates PCR products containing converted DNA derived from hypermethylated target genomic regions, while not amplifying converted cfDNA derived from hypomethylated target genomic regions. In some embodiments, the sequence for selective PCR amplification generates PCR products containing converted DNA derived from hypomethylated target genomic regions, while not amplifying converted cfDNA derived from hypermethylated target genomic regions. Each target genomic region may include at least five CpG sites. In some embodiments, the target genomic region includes at least 1, 2, 3, 4, 5, 10, 15, or more target genomic regions containing target sequences from one or more genes selected from Table 1. In some embodiments, the target genomic region includes a target sequence of a gene selected from TWIST1, EOMES, HOXA9, POU4F2, and ZNF154. Determining the likelihood of a subject having cancer or a cancer type can include amplifying the converted cfDNA fragments from the subject with oligonucleotide pairs configured for selective PCR amplification as described herein, sequencing the captured DNA fragments, and applying a trained classifier to the DNA sequence to determine likelihood.The sensitivity of the classifier in determining the likelihood of cancer or a cancer type can be at least 40% with a specificity of 0.990 or greater.

[0142] Data structure generation 12A is a flowchart illustrating a process 300 for generating a healthy control data structure, according to one embodiment. To create the healthy control data structure, an analytics system obtains information about the methylation status of multiple CpG sites on sequence reads derived from multiple DNA molecules or fragments from multiple healthy subjects. The methods provided herein for creating a healthy control data structure can similarly be performed for subjects with cancer, subjects with a specific type of cancer, subjects with a known cancer type, or subjects with another known disease state. A methylation state vector is generated for each DNA molecule or fragment, for example, by process 100.

[0143] The analytics system further divides 310 the methylation state vector of each DNA fragment into strings of CpG sites. In one embodiment, the analytics system further divides 310 the methylation state vector so that all resulting strings are less than a given length. For example, a methylation state vector of length 11 may be further divided into strings of length 3 or less, resulting in 9 strings of length 3, 10 strings of length 2, and 11 strings of length 1. In another example, a methylation state vector of length 7 may be further divided into strings of length 4 or less, resulting in 4 strings of length 4, 5 strings of length 3, 6 strings of length 2, and 7 strings of length 1. If the methylation state vector resulting from a DNA fragment is shorter than or equal to the specified string length, then the methylation state vector may be converted into a single string containing all of the CpG sites in the vector.

[0144] For each possible CpG site and methylation state possibility in the vector, the analytics system tabulates the strings by counting the number of strings present in the control group that have the specified CpG site as the first CpG site in the string and have that methylation state possibility 320. For a string length of 3 at a given CpG site, 3There are 8 possible string configurations. For each CpG site, the analytics system tallies how many times each possible methylation state vector appears in the control set 320. This involves tallying the following quantities for each starting CpG site in the reference genome: <M x ,M x+1 ,M x+2 >, <M x ,M x+1 ,U x+2 >,..., x ,U x+1 ,U x+2 The analytics system creates 330 a data structure that stores the counts for each starting CpG site and the string possibilities at each starting CpG.

[0145] Putting an upper limit on string length has several benefits. First, the size of the data structures created by analytics systems can increase dramatically depending on the maximum string length. For example, a maximum string length of 4 means that at most 2 CpGs can be counted. 4 ​This means that the number of possible methylation states is 1. Increasing the maximum string length to 5 doubles the number of possible methylation states to be tallied. Reducing the string size helps reduce the computational and data storage burden on the data structure. In some embodiments, the string size is 3. In some embodiments, the string size is 4. A second reason for limiting the maximum string length is to avoid overfitting of downstream models. Calculating probabilities based on long strings of CpG sites can be problematic if long CpG strings do not have a strong biological effect on the outcome (e.g., predicting abnormality as a predictor of the presence of cancer), because the required amount of data may not be available, resulting in a model that is too sparse for adequate performance. For example, calculating the probability of abnormality / cancer conditional on the previous 100 CpG sites may require counting strings in a data structure of length 100, some of which ideally exactly match the previous 100 methylation states. If only a sparse count of strings of length 100 is available, there will be insufficient data to determine whether a given string of length 100 in a test specimen is abnormal or not.

[0146] Data structure validation Once the data structure has been created, the analytics system may attempt to validate 340 the data structure and / or any downstream models that utilize the data structure.

[0147] This first type of validation ensures that potentially cancerous samples are removed from the healthy control group so as not to affect the purity of the control group. This type of validation checks for consistency within the data structure of the control group. For example, the healthy control group may include samples from individuals with undiagnosed cancer that contain multiple aberrant methylation fragments. The analytics system may perform various calculations to determine whether to exclude data from subjects who appear to have undiagnosed cancer.

[0148] The second type of validation checks the probability model used to calculate p-values ​​with counts from the data structure itself (i.e., from the healthy control group). The p-value calculation process is described below in conjunction with Figure 14. Once the analytics system generates p-values ​​for the methylation state vectors in the validation group, it constructs a cumulative density function (CDF) with those p-values. Using the CDF, the analytics system can perform various calculations on the CDF to validate the control group's data structure. One test uses the CDF to ensure that the CDF ideally is or falls below the identity function, i.e., CDF(x)≦x. Conversely, exceeding the identity function reveals some kind of flaw in the probability model used to structure the control group's data. For example, if 1 / 100 of the fragments has a p-value score of 1 / 1000, i.e., CDF(1 / 1000)=1 / 100>1 / 1000, the second type of validation fails, indicating a problem with the probability model. See, for example, U.S. Patent Application No. 16 / 352,602, published as U.S. Patent Application Publication No. 2019 / 0287652, which is hereby incorporated by reference in its entirety.

[0149] The third type of validation uses a healthy validation sample set separate from the one used to construct the data structure. This tests whether the data structure is properly constructed and the model works. An exemplary process for performing this type of validation is described below in conjunction with FIG. 12B. This third type of validation can quantify how well the healthy control group generalizes to the distribution of healthy samples. If the third type of validation fails, then the healthy control group does not generalize well to the healthy distribution.

[0150] The fourth type of validation is tested on samples from the unhealthy validation group. The analytics system calculates p-values ​​and constructs a CDF for the unhealthy validation group. In the unhealthy validation group, the analytics system expects to see CDF(x)>x for at least some samples, or in other words, the opposite of what was expected by the healthy control and healthy validation groups in the second and third types of validation. If the fourth type of validation fails, then it indicates that the model is not adequate to identify the abnormalities that the model was designed to identify.

[0151] 12B is a flowchart illustrating an additional step 340 of validating the data structure of the control group of FIG. 12A, according to one embodiment. In this embodiment of step 340 of validating the data structure, the analytics system performs a fourth type of validation test as described above, in which a validation group is utilized in which the subject, specimen, and / or fragment composition is presumably similar to the control group. For example, if the analytics system selects healthy subjects without cancer for the control group, then the analytics system also uses healthy subjects without cancer in the validation group.

[0152] The analytics system takes the validation set and generates a set of methylation state vectors 100, as described in FIG. 12A. The analytics system performs a p-value calculation for each methylation state vector from the validation set. The p-value calculation process is further described in conjunction with FIGS. 13-14. For each possible methylation state vector, the analytics system calculates a probability from the control set data structure. Once the probabilities have been calculated for the possible methylation state vectors, the analytics system calculates a p-value score 350 for that methylation state vector based on the calculated probabilities. The p-value score represents the expectation that that particular methylation state vector, and other possible methylation state vectors, will be found to have even lower probabilities in the control set. Thus, a low p-value score generally corresponds to a methylation state vector that is relatively less likely to be found in the control set compared to other methylation state vectors, while a high p-value score generally corresponds to a methylation state vector that is relatively more likely to be found in the control set compared to other methylation state vectors. Once the analytics system generates p-value scores for the methylation state vectors in the validation set, the analytics system constructs a cumulative density function (CDF) with the p-value scores from the validation set 360. The analytics system verifies the consistency of the CDF 370 as described above in the fourth type of validation test.

[0153] Aberrantly methylated fragments According to one embodiment, as outlined in Figure 13, aberrant methylation fragments that exhibit abnormal methylation patterns in a cancer patient specimen, a subject with a certain type of cancer, a subject with a known cancer type, or a subject with another known disease state are selected as target genomic regions. An exemplary process for selecting aberrant methylation fragments 440 is illustrated in Figure 14 and further described below in the description of Figure 4. In process 400, the analytics system generates 100 methylation state vectors from the cfDNA fragments of the specimen. The analytics system treats each methylation state vector as follows:

[0154] For a given methylation state vector, the analytics system enumerates all possibilities for methylation state vectors that have the same starting CpG site and the same length (i.e., set of CpG sites) in the methylation state vector 410. Because each methylation state is either methylated or unmethylated, there are only two possible states for each CpG site, and therefore the count of distinct possibilities for a methylation state vector depends on a power of 2, so for a methylation state vector of length n, there are 2 n The probability of a methylation state vector can be correlated.

[0155] The analytics system accesses the healthy control data structure to calculate 420 the probability that each possible methylation state vector will be observed for the identified starting CpG site / methylation state vector length. In one embodiment, the calculation of the probability that a given possibility will be observed uses Markov chain probability to model joint probability calculations (as described in more detail below in connection with FIG. 14). In other embodiments, calculation methods other than Markov chain probability are used to determine the probability that each possible methylation state vector will be observed.

[0156] The analytics system uses the calculated probability for each possibility to calculate a p-value score for the methylation state vector 430. In one embodiment, this involves identifying the calculated probability that corresponds to the likelihood of matching the methylation state vector in question. Specifically, this is the likelihood of having the same set of CpG sites as the methylation state vector, or similarly the same starting CpG site and length. The analytics system generates the p-value score by summing the calculated probabilities of any possibilities that have a probability less than or equal to the identified probability.

[0157] This p-value represents the probability of observing the methylation state vector of the fragment or other methylation state vectors that are deemed even less likely in healthy controls. Thus, a low p-value score generally corresponds to a methylation state vector that is rare in healthy subjects, and therefore the fragment would be labeled as not having normal methylation compared to the healthy control group. A high p-value score generally relates to the expectation that a methylation state vector is present in healthy subjects, in a comparative sense. If the healthy control group is a non-cancerous group, for example, a low p-value may indicate that the fragment is not having normal methylation compared to the non-cancerous group, and therefore indicate the presence of cancer in the test subject.

[0158] As described above, the analytics system calculates a p-value score for each of a plurality of methylation state vectors, each representing a cfDNA fragment in the test specimen. To identify which fragments are not normally methylated, the analytics system may filter the set of methylation state vectors based on their p-value scores 440. In one embodiment, filtering is performed by comparing the p-value scores to a threshold value and retaining only fragments that are below the threshold. This threshold p-value score can be an order of magnitude number such as 0.1, 0.01, 0.001, 0.0001, or a similar number.

[0159] P-value score calculation 14 is an illustration 500 of an exemplary p-value score calculation, according to one embodiment. To calculate a p-value score given a test methylation state vector 505, the analytics system takes the test methylation state vector 505 and enumerates 410 the possibilities of the methylation state vector. In this illustrative example, the test methylation state vector 505 is <M 23 ,M 24 ,M 25 ,U 26 Because the length of the test methylation state vector 505 is 4, the probability of the methylation state vector containing CpG sites 23 to 26 is 2 4 In a general case, the number of possibilities for the methylation state vector is 2 nwhere n is the length of the test methylation state vector, or alternatively, the length of the sliding window (described further below).

[0160] The analytics system calculates 420 the probability 515 for the enumerated possibilities of the methylation state vector. Because methylation conditionally depends on the methylation state of nearby CpG sites, one way to calculate the probability that a given methylation state vector possibility will be observed is to use a Markov chain model. Generally, <S1,S2,...,S n >, where S represents a methylation state that is either methylated (denoted by M), unmethylated (denoted by U), or indeterminate (denoted by I), P( <S1,S2,...,S n >)=P(S n |S1,...,S n-1 >)×P(S n-1 |S1,...,S n-2 >)×...×P(S2|S1)×P(S1) (1) The Markov chain model allows for a more efficient calculation of the conditional probability of each possibility. In one embodiment, the analytics system selects a Markov chain order k, which corresponds to how many CpG sites in the vector (or window) precede it that will be considered in the conditional probability calculation; thus, the conditional probability is P(S n |S1,...,S n-1 )~P(S n |S n-k-2 ,...,S n-1 )

[0161] To calculate the Markov-modeled probability for each possible methylation state vector, the analytics system accesses the control data structure, specifically the counts of various strings of CpG sites and states. P(M n |S n-k-2,...,S n-1 ), the analytics system calculates n-k-2 ,...,S n-1 ,M n > the stored count of the number of strings from the data structure that match n-k-2 ,...,S n-1 ,M n > and n-k-2 ,...,S n-1 ,U n Take the ratio of the number of strings from the data structure that match > divided by the total stored count. Thus, P(M n |S n-k-2 ,...,S n-1 )teeth,

number

[0162] The calculation may additionally implement smoothing of the counts by applying a prior. In one embodiment, the prior is a uniform prior, such as in Laplace smoothing. An example of this is adding a constant to the numerator of the above equation and another constant (e.g., twice the constant in the numerator) to the denominator. In other embodiments, algorithmic techniques such as Knesser-Ney smoothing are used.

[0163] In this example, the formula depicted above is applied to a test methylation state vector 505 covering sites 23-26. Once the calculated probabilities 515 are complete, the analytics system calculates 430 a p-value score 525 that sums the probability of a methylation state vector matching the test methylation state vector 505 being less than or equal to the probability of that methylation state vector matching the test methylation state vector 505.

[0164] ​​​In one embodiment, the computational burden of calculating probability and / or p-value scores may be further reduced by caching at least some of the calculations. For example, the analytics system may cache probability calculations for the likelihood of a methylation state vector (or a window thereof) in temporary or persistent memory. If other fragments have the same CpG sites, this caching of likelihood probabilities allows p-value scores to be calculated efficiently without having to recalculate the underlying likelihood probabilities. In other words, the analytics system may calculate a p-value score for each likelihood of a methylation state vector associated with a set of CpG sites from the vector (or a window thereof). The analytics system may cache the p-value scores for use in determining p-value scores for other fragments containing the same CpG sites. In general, the p-value scores of likelihoods of methylation state vectors with the same CpG sites may be used to determine p-value scores for different possibilities from the same set of CpG sites.

[0165] Sliding Window In one embodiment, the analytics system uses a sliding window 435 to determine the probabilities of the methylation state vector and calculate p-values. Rather than enumerating the probabilities and calculating p-values ​​for the entire methylation state vector, the analytics system enumerates the probabilities and calculates p-values ​​only for a window of contiguous CpG sites, where the window is shorter than at least some fragment (of CpG sites) in length (otherwise the window may serve no purpose). The window length may be static, user-determined, dynamic, or of some other selection.

[0166] In calculating p-values ​​for methylation state vectors larger than a window, the window identifies a contiguous set of CpG sites from the vector that fall within the window, starting with the first CpG site in the vector. The analytics system calculates a p-value score for the window that includes the first CpG site. The analytics system then "slides" the window to the second CpG site in the vector and calculates another p-value score for the second window. Thus, for a window size l and a methylation vector length m, each methylation state vector will generate m-l+1 p-value scores. After completing the p-value calculations for each portion of the vector, the lowest p-value score from all sliding windows is taken as the overall p-value score for that methylation state vector. In another embodiment, the analytics system aggregates the p-value scores for a methylation state vector to generate an overall p-value score.

[0167] The use of a sliding window helps reduce the number of possible methylation state vectors to be enumerated and the corresponding probability calculations that might otherwise need to be performed. An exemplary probability calculation is shown in Figure 14, but generally the number of possible methylation state vectors increases exponentially with the size of the methylation state vector by a factor of 2. To give a realistic example, a fragment can have more than 54 CpG sites. 54 Street (approx. 1.8 x 10 16 Instead of calculating probabilities for 2 possibilities (number of possibilities) to generate a single p-value, the analytics system can instead use a window of size 5 (for example), resulting in 50 p-value calculations for each of the 50 windows of the methylation state vector for that fragment. 5 To enumerate the possibilities of the street (32), a total of 50 x 2 5 Street (1.6 x 10 3This results in a probability calculation of 100 (number of copies). This significantly reduces the calculations that must be performed, without significantly affecting the accurate identification of the aberrant fragments. This additional step can also be applied when validating the control set with the validation set methylation state vector.

[0168] Identification of cancer-indicating fragments The analytics system identifies DNA fragments that are indicative of cancer from the filtered collection of aberrantly methylated fragments450.

[0169] Hypomethylated and hypermethylated fragments According to a first method, the analytics system can identify DNA fragments from the filtered set of aberrantly methylated fragments that are considered to be hypomethylated or hypermethylated as fragments indicative of cancer. Hypomethylated and hypermethylated fragments can be defined as fragments with a high percentage of methylated CpG sites (e.g., greater than 80%, 85%, 90%, or 95%, or any other percentage within the range of 50% to 100%) or a high percentage of unmethylated CpG sites (e.g., greater than 80%, 85%, 90%, or 95%, or any other percentage within the range of 50% to 100%) and a certain length of CpG sites (e.g., greater than 3, 4, 5, 6, 7, 8, 9, 10, etc.).

[0170] Probabilistic Models According to the methods described herein, an analytics system identifies fragments indicative of cancer using a probability model of methylation patterns tailored to each cancer type and non-cancer type. The analytics system calculates a log-likelihood ratio for a sample using DNA fragments in genomic regions that account for various cancer types in the probability model tailored to each cancer type and non-cancer type. The analytics system can determine that a DNA fragment is indicative of cancer based on whether at least one of the log-likelihood ratios considered for various cancer types exceeds a threshold.

[0171] In one embodiment of dividing a genome, the analytics system divides the genome into regions in multiple stages. In the first stage, the analytics system separates the genome into blocks of CpG sites. Each block is defined when there is a separation between two adjacent CpG sites that meets and / or exceeds some threshold, e.g., greater than 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, or 1,000 bp. From each block, the analytics system further divides each block into regions of a specific length, e.g., 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, 1,100 bp, 1,200 bp, 1,300 bp, 1,400 bp, or 1,500 bp, in the second stage. The analytics system may further have an overlap with adjacent regions of a certain percentage of length, for example, 10%, 20%, 30%, 40%, 50%, or 60%, or 10% or more, 20% or more, 30% or more, 40% or more, 50% or more, or 60% or more.

[0172] The analytics system analyzes sequence reads derived from DNA fragments by region. The analytics system can process samples from tissue and / or high-signal cfDNA. High-signal cfDNA samples can be determined by a binary classification model, by cancer stage, or by another metric.

[0173] The analytics system fits separate probability models to the fragments for each cancer type and non-cancer type. In one example, each probability model is a mixture model that includes a combination of multiple mixture components, where each mixture component is an independent site model in which methylation at each CpG site is assumed to be independent of the methylation status at other CpG sites.

[0174] In alternative embodiments, calculation is carried out for each CpG site.Specifically, a first count (cancer_count) is determined, which is the number of cancerous specimens that contain the DNA fragments of abnormal methylation that overlap with this CpG in the collection, and a second count (total count) is determined, which is the total number of specimens that contain fragments that overlap with this CpG.Based on these counts, genomic regions can be selected based on a criterion that is positively correlated with the number of cancerous specimens that contain the DNA fragments that overlap with this CpG in the collection (cancer_count) and inversely correlated with the total number of specimens that contain fragments that overlap with this CpG (total count).

[0175] In some embodiments, the various types of cancer with different TOO are selected from bladder cancer, urothelial cancer, renal cancer, and prostate cancer. In some embodiments, the cancer is bladder cancer. In some embodiments, the cancer is urothelial cancer. In some embodiments, the cancer is renal cancer. In some embodiments, the cancer is prostate cancer.

[0176] In some embodiments, various cancer types can be classified and named using classification methods available in the art, such as the International Classification of Diseases for Oncology (ICD-O-3) (codes.iarc.fr) or the Surveillance, Epidemiology, and End Results Program (SEER) (seer.cancer.gov). In other embodiments, cancer types are classified by three orthogonal codes: (i) tissue distribution code, (ii) morphology code, or (iii) behavior code. Under the behavior code, benign tumor is 0, uncertain behavior is 1, carcinoma in situ is 2, malignant primary site is 3, and malignant metastatic site is 6.

[0177] In some embodiments, the cancer TOO can be selected from a group defined by guidelines used to stage the detected cancer. For example, the reference Amin, MB, Edge, S., Greene, F., Byrd, DR, Brookland, RK, Washington, MK, Gershenwald, JE, Compton, CC, Hess, KR, Sullivan, DC, Jessup, JM, Brierley, JD, Gaspar, LE, Schilsky, RL, Balch, CM, Winchester, DP, Asare, EA, Madera, M., Gress, DM, Meyer, LR (Eds.), AJCC Cancer Staging Manual, 8th edition, Springer, 2017, identifies groups of different cancers that are staged together according to standard guidelines. Staging is typically the next step in cancer management after its detection and diagnosis.

[0178] The analytics system can further calculate a log-likelihood ratio ("R") for each cancer type and non-cancer type, or a fragment indicating the likelihood of the fragment being indicative of cancer given the various cancer types in a probability model fitted to cancer TOO. These two probabilities may be taken from a probability model fitted to each of the cancer type and non-cancer type, where the probability model is defined to calculate the likelihood of observing a methylation pattern on the fragment given each of the cancer type and non-cancer type. For example, a probability model fitted to each of the cancer type and non-cancer type may be defined.

[0179] Selection of genomic regions that direct cancer In some embodiments, the analytics system can identify genomic regions that are indicative of cancer 460. To identify these informative regions, the analytics system calculates an information gain that describes the ability to distinguish between various outcomes for each genomic region, or more specifically, for each CpG site.

[0180] The method for identifying genomic regions capable of distinguishing between cancer and non-cancerous types utilizes a trained classification model that can be applied to a collection of aberrantly methylated DNA molecules or fragments corresponding to or derived from a cancerous or non-cancerous group. The trained classification model can be trained to identify any condition of interest that can be identified from a methylation state vector.

[0181] In one embodiment, the trained classification model is a binary classifier trained based on the methylation status of cfDNA fragments or genomic sequences obtained from a cohort of subjects with cancer or cancer TOO and a cohort of healthy subjects without cancer, and is then used to classify the probability that a test subject has cancer, cancer TOO, or no cancer based on the abnormal methylation status vector. In other embodiments, various classifiers may be trained using subject cohorts known to have a particular cancer (e.g., bladder cancer, urothelial cancer, prostate cancer, renal cancer, etc.); known to have a particular TOO cancer that is thought to be the origin of cancer; or known to have various stages of a particular cancer (e.g., bladder cancer, urothelial cancer, prostate cancer, renal cancer, etc.). In these embodiments, various classifiers may be trained using sequence reads obtained from specimens enriched for tumor cells from subject cohorts known to have a particular cancer (e.g., bladder cancer, urothelial cancer, prostate cancer, renal cancer, etc.). Using the ability of each genomic region to distinguish between cancer and non-cancer types in the classification model, the genomic regions are ranked from most informative to least informative in terms of classification performance, and from the ranking, the analytics system can identify genomic regions according to their information gain in classifying between non-cancer and cancer types.

[0182] Computation of information gain from hypomethylated and hypermethylated fragments indicative of cancer The analytics system may train a classifier on the cancer-indicative fragments according to one embodiment of process 600 shown in Figure 15A. Process 600 accesses two training sets of samples—a non-cancer set and a cancer set—to obtain 605 a set of non-cancer methylation state vectors and a set of cancer methylation state vectors containing aberrant methylation fragments, e.g., by step 440 from process 400.

[0183] For each methylation state vector, the analytics system determines whether the methylation state vector is indicative of cancer 610. Here, a fragment indicative of cancer may be defined as a hypermethylated or hypomethylated fragment, determined as having at least a portion of its CpG sites in a particular state (methylated or unmethylated, respectively) and / or a threshold percentage of sites in a particular state (again, methylated or unmethylated, respectively). In one example, a cfDNA fragment is identified as hypomethylated or hypermethylated if the fragment overlaps at least five CpG sites and at least 80%, 90%, or 100% of those CpG sites are methylated or at least 80%, 90%, or 100% are unmethylated, respectively.

[0184] In an alternative embodiment, the analytics system may consider a portion of the methylation state vector to determine whether that portion is hypomethylated or hypermethylated and identify that portion as hypomethylated or hypermethylated. This alternative addresses the lack of a methylation state vector that is large in size but contains at least one densely populated hypomethylated or hypermethylated region. This process of defining hypomethylation and hypermethylation can be applied in step 450 of Figure 13. In another embodiment, the cancer-indicative fragments may be defined according to the likelihood output from a trained probabilistic model.

[0185] In one embodiment, the analytics system generates a hypomethylation score (P) for each CpG site in the genome.hypo ) and hypermethylation score (P hyper ) 620. To generate any score at a given CpG site, the classifier takes four counts at that CpG site—(1) counts of (methylation state) vectors in the cancer population that are labeled as hypomethylated and overlap the CpG site; (2) counts of vectors in the cancer population that are labeled as hypermethylated and overlap the CpG site; (3) counts of vectors in the non-cancer population that are labeled as hypomethylated and overlap the CpG site; and (4) counts of vectors in the non-cancer population that are labeled as hypermethylated and overlap the CpG site. In addition, the process may normalize these counts per group to account for differences in group size between the non-cancer and cancer groups. In alternative embodiments where cancer-indicative fragments are more commonly used, the score may be defined more broadly as the count of cancer-indicative fragments at each genomic region and / or CpG site.

[0186] In one embodiment, to generate 620 the hypomethylation score at a given CpG site, the process takes the ratio of (1) to the sum of (1) and (3). Similarly, the hypermethylation score is calculated by taking the ratio of (2) to (2) and (4). Additionally, these ratios may be calculated with additional smoothing techniques as discussed above. The hypomethylation score and hypermethylation score relate to an estimation of the probability of cancer, given the presence of hypomethylation or hypermethylation of fragments from the cancer population.

[0187] The analytics system generates an overall hypomethylation score and an overall hypermethylation score for each aberrant methylation state vector 630. The overall hypermethylation and hypomethylation scores are determined based on the hypermethylation and hypomethylation scores of the CpG sites in the methylation state vector. In one embodiment, the overall hypermethylation and hypomethylation scores are assigned as the highest hypermethylation and hypomethylation scores, respectively, among the sites in each state vector. However, in alternative embodiments, the overall score may be based on the mean, median, or other calculation using the hypermethylation scores of the sites in each vector.

[0188] The analytics system ranks 640 all of the subject's methylation state vectors by their overall hypomethylation score and their overall hypermethylation score, resulting in two rankings for each subject. The process selects an overall hypomethylation score from the hypomethylation rankings and an overall hypermethylation score from the hypermethylation rankings. With the selected scores, the classifier generates 650 a single feature vector for each subject. In one embodiment, the scores selected from either ranking are selected in a fixed order that is the same for each feature vector generated for each subject in each training set. By way of example, in one embodiment, the classifier selects the first, second, fourth, and eighth overall hypermethylation score from each ranking, as well as the overall hypomethylation score, and writes those scores into the feature vector for that subject.

[0189] The analytics system trains 660 a binary classifier to distinguish feature vectors between the cancer and non-cancer training groups. Generally, any one of a number of classification techniques may be used. In one embodiment, the classifier is a nonlinear classifier. In a specific embodiment, the classifier is a nonlinear classifier utilizing L2 regularized kernel logistic regression and a Gaussian radial basis function (RBF) kernel.

[0190] Specifically, in one embodiment, the number (n) of non-cancer specimens or one or more different cancer types having aberrantly methylated fragments overlapping the CpG site is その他 ) and the number of cancer specimens or one or more cancer types (n 癌 ) is counted. Then the probability that a sample is cancer is calculated as n 癌 Positively correlated with and n その他 This score is estimated by a score ("S") that is inversely related to the 癌 +1) / (n 癌 +n その他 +2) or (n 癌 ) / (n 癌 +n その他 ) can be used to calculate the information gain for each cancer type and each genomic region or CpG site to determine whether the genomic region or CpG site is indicative of cancer 670. The information gain is calculated for training samples with a given cancer type compared to all other samples. For example, two random variables are used: "abnormal fragment" ("AF") and "cancer type" ("CT"). In one embodiment, AF is a binary variable indicating whether there is an abnormal fragment in a given sample that overlaps with a given CpG site, as determined for the abnormal score / feature vector above. CT is a random variable indicating whether the cancer is of a particular type. The analytics system calculates the mutual information for CT given AF. That is, knowing whether there is an abnormal fragment overlapping with a particular CpG site, a number of bits of information about the cancer type is obtained.

[0191] For a given cancer type, the analytics system uses that information to rank CpG sites based on how cancer-specific they are. This procedure is repeated for all cancer types under consideration. If a particular region is commonly aberrantly methylated in training samples of a given cancer, but not aberrantly methylated in training samples of other cancer types or healthy training samples, then the CpG sites where those aberrant segments overlap will tend to have high information gain for the given cancer type. The ranked CpG sites for each cancer type are greedily added (selected) into a set of selected CpG sites based on their ranking for use in the cancer classifier.

[0192] Computing pairwise information gain from cancer-indicating fragments identified from probabilistic models Once fragments indicative of cancer are identified by the second method described herein, analytics may identify genomic regions according to process 680 of FIG. 15B. The analytics system defines feature vectors 690 for each sample, each region, and each cancer type by counts of DNA fragments whose calculated log-likelihood ratios that the fragments are indicative of cancer exceed multiple thresholds, where each count is a value in the feature vector. In one embodiment, the analytics system counts, for each cancer type, the number of fragments present in the sample in a region whose log-likelihood ratios exceed one or more possible thresholds. The analytics system defines feature vectors for each sample by counts of DNA fragments for each genomic region for each cancer type that provide the calculated log-likelihood ratios for fragments that exceed multiple thresholds, where each count is a value in the feature vector. The analytics system uses the defined feature vectors to calculate an information score for each genomic region that describes the ability of the genomic region to discriminate between each pair of cancer types. For each pair of cancer types, the analytics system ranks the regions based on their information scores. The analytics system may select regions based on ranking by information score.

[0193] The analytics system calculates an information score for each region 695, describing the region's ability to discriminate between each pair of cancer types. For each individual cancer type pair, the analytics system may designate one type as a positive type and the other as a negative type. In one embodiment, the ability of a region to discriminate between positive and negative types is based on the estimated proportion of positive and negative cfDNA specimens that can be expected to have a non-zero feature value in the final assay, i.e., the mutual information calculated using at least one fragment of that hierarchy that can be sequenced in the targeted methylation assay. These proportions are estimated using the observed ratios at which the feature appears in healthy cfDNA and in high-signal cfDNA and / or tumor specimens of each cancer type. For example, if a feature appears frequently in healthy cfDNA, then it will likely also appear frequently in cfDNA of any cancer type, resulting in a low information score. The analytics system may select a specific number of regions, e.g., 1024, for each cancer type pair from the ranking.

[0194] In a further embodiment, the analytics system further identifies predominantly hypermethylated or hypomethylated regions from the ranking of the regions. The analytics system may load a set of fragments in one or more positive types for regions identified as informative. From the loaded fragments, the analytics system evaluates whether the loaded fragments are predominantly hypermethylated or hypomethylated. If the loaded fragments are predominantly hypermethylated or hypomethylated, the analytics system may select probes corresponding to the predominant methylation pattern. If the loaded fragments are not predominantly hypermethylated or hypomethylated, the analytics system may use a mix of probes to target both hypermethylated and hypomethylated regions. The analytics system may further identify a minimal set of CpG sites that overlap with fragments above a certain percentage.

[0195] In other embodiments, the analytics system ranks the regions based on their information scores and then labels each region with the lowest information ranking for all pairs of cancer types. For example, if a region is 10th most informative for distinguishing kidney from bladder and 5th most informative for distinguishing kidney from prostate, then it might be given an overall label of "5." The analytics system might start with the lowest labeled region while designing probes to add regions to the panel, for example, until the panel's size budget is exhausted.

[0196] Off-target genomic regions In some embodiments, probes targeting selected genomic regions are further filtered based on the number of similar or identical off-target sequences in the genome 475. This is to screen for probes that pull down too many cfDNA fragments corresponding to or derived from off-target genomic regions. Removing probes with many off-target sequences in the genome can be useful because it reduces the off-target rate and increases target coverage for a given amount of sequencing.

[0197] An off-target genomic region is a genomic region that has sufficient homology with a target genomic region, such that DNA molecules or fragments derived from the off-target genomic region can be hybridized to and pulled down by a probe designed to hybridize to the target genomic region. The off-target genomic region can include a sequence (or a converted sequence of the same region) that aligns with the probe at a match rate of at least 80%, 85%, 90%, 95%, or 97% along at least 35 bp, 40 bp, 45 bp, 50 bp, 60 bp, 70 bp, or 80 bp. In one embodiment, the off-target genomic region is a genomic region that includes a sequence (or a converted sequence of the same region) that aligns with the probe at a match rate of at least 90% along at least 45 bp. Various methods can be used to screen for and remove off-target genomic regions.

[0198] It can be computationally challenging to exhaustively search a genome to find all off-target genomic regions. In one embodiment, a k-mer seeding strategy (which can tolerate one or more mismatches) is combined with local alignment at the seed position. In this case, exhaustive search for good alignments can be guaranteed based on the length of the k-mer, the number of mismatches allowed, and the number of k-mer seed hits at a particular position. This requires dynamic programming local alignment at many positions, making this approach highly optimized for the use of vector CPU instructions (e.g., AVX2, AVX512) and also parallelizable across many cores within a machine and across many machines connected by a network. Those skilled in the art will recognize that improvements and variations of this approach can be implemented to identify off-target genomic regions.

[0199] In some embodiments, probes with sequence homology to, or DNA molecules corresponding to, or derived from, off-target genomic regions comprising more than a threshold number are excluded (or filtered) from the panel, for example, probes with sequence homology to, or DNA molecules corresponding to, or derived from, off-target genomic regions from more than 30, more than 25, more than 20, more than 18, more than 15, more than 12, more than 10, or more than 5 off-target regions are excluded.

[0200] In some embodiments, probes are divided into two, three, four, five, six, or more distinct groups depending on the number of off-target regions. For example, probes with no sequence homology to the off-target regions or DNA molecules corresponding to or derived from the off-target regions are assigned to a high-quality group; probes with sequence homology to 1 to 18 off-target regions or DNA molecules corresponding to or derived from 1 to 18 off-target regions are assigned to a low-quality group; and probes with sequence homology to more than 19 off-target regions or DNA molecules corresponding to or derived from 19 off-target regions are assigned to a low-quality group. Other cutoff values ​​can be used for grouping.

[0201] In some embodiments, probes in the lowest quality group are excluded. In some embodiments, probes in groups other than the highest quality group are excluded. In some embodiments, a separate panel is created for probes in each group. In some embodiments, all probes are placed on the same panel, but separate analyses are performed based on the assigned group.

[0202] In some embodiments, a panel contains more high quality probes than probes in lower groups. In some embodiments, a panel contains fewer low quality probes than probes in other groups. In some embodiments, more than 95%, 90%, 85%, 80%, 75%, or 70% of the probes in a panel are high quality probes. In some embodiments, less than 35%, 30%, 20%, 10%, 5%, 4%, 3%, 2%, or 1% of the probes in a panel are low quality probes. In some embodiments, less than 5%, 4%, 3%, 2%, or 1% of the probes in a panel are low quality probes. In some embodiments, low quality probes are not included in a panel.

[0203] In some embodiments, probes with a percentage below 50%, below 40%, below 30%, below 20%, below 10%, or below 5% are excluded, hi some embodiments, probes with a percentage above 30%, above 40%, above 50%, above 60%, above 70%, above 80%, or above 90% are selectively included in the panel.

[0204] How to Use Cancer Assay Panels In one aspect, a method of using a cancer assay panel (alternatively referred to as a "bait set") is provided. The cancer assay panel may be any disclosed herein, including those relating to any of the various aspects. The method may include treating DNA molecules or fragments to convert unmethylated cytosines to uracil (e.g., using bisulfite treatment), applying a cancer panel (as described herein) to the converted DNA molecules or fragments, enriching a subset of the converted DNA molecules or fragments that bind to the probes in the panel, and sequencing the enriched DNA fragments. In some embodiments, comparing the sequence reads to a reference genome (e.g., a human reference genome) may enable identification of the methylation status at multiple CpG sites within the DNA molecules or fragments, thereby providing information relevant to cancer detection. In some embodiments, the DNA is cfDNA. In some embodiments, the cfDNA is from a urine specimen. In some embodiments, target genomic regions identified as particularly useful in characterizing cfDNA from urine specimens are also useful in characterizing cfDNA from other specimen types (e.g., cfDNA from biological fluids such as blood, serum, or plasma). In some embodiments, information regarding target genomic regions disclosed herein is applied to assaying cfDNA from such other specimen types for the detection of cancer.

[0205] In one aspect, the disclosure provides a method for detecting cancer cells in a subject, the method comprising: (a) capturing converted cell-free DNA (cfDNA) fragments, or amplification products thereof, from a urine specimen of the subject, wherein (i) a bait oligonucleotide composition comprises a plurality of different bait oligonucleotides, and (ii) each bait oligonucleotide of the plurality of different bait oligonucleotides hybridizes to a target sequence of a gene selected from Table 1, wherein the target sequence is at least 25 nucleotides in length; (b) separating DNA bound to the bait from unbound DNA; (c) sequencing the separated DNA to generate sequencing reads; and (d) detecting the cancer cells with a trained classifier.

[0206] Cell-free DNA fragments may be obtained from a urine specimen prepared by any of the methods disclosed herein. In some embodiments, the cell-free nucleic acids are obtained by (i) treating the urine specimen to inhibit cell lysis; (ii) separating cfDNA fragments in the treated urine specimen from cells in the treated urine specimen, thereby producing a purified urine specimen containing cfDNA fragments; (iii) concentrating the cfDNA fragments in the purified urine specimen by passing at least a portion of the purified urine specimen through a filter, whereby a filtrate and a residual urine specimen are produced, and the residual urine specimen has an increased concentration of cfDNA fragments; and (iv) isolating the cfDNA fragments from the residual urine specimen. In some embodiments, treating the urine specimen to inhibit cell lysis includes contacting the urine specimen with one or more preservative reagents, non-limiting examples of which are described herein. In some embodiments, the treating includes treatment with a nuclease inhibitor, a formaldehyde quencher, or both. Non-limiting examples of nuclease inhibitors and formaldehyde quenchers are described herein. In some embodiments, the treatment includes contacting the urine specimen with a composition comprising imidazolidinyl urea, EDTA, glycine, or a combination thereof. In some embodiments, the treatment includes contacting the urine specimen with a composition comprising sodium azide, EDTA, or a combination thereof. Isolating nucleic acids can include centrifuging the treated urine specimen to pellet cells. The filter used to prepare the residual urine specimen can be substantially impermeable to cell-free nucleic acids in the purified urine specimen but permeable to salts. In some embodiments, the filter has a nominal molecular weight cutoff of 10 kD, 5 kD, 3 kD, or less. In some embodiments, the residual urine specimen has a concentration that is at least 2-fold, at least 5-fold, at least 10-fold, or at least 15-fold increased compared to the purified urine specimen. In some instances, the residual urine specimen has a volume that is at least 50%, at least 75%, or at least 90% smaller than the volume of the treated urine specimen.The volume of the processed urine specimen can be 5 mL, 10 mL, 15 mL, 20 mL, 30 mL, 40 mL, 50 mL, or more. The method can further include freezing the remaining urine specimen. In some embodiments, processing is completed within 120, 60, or 30 minutes after collection of the urine specimen. In some embodiments, separation and concentration is completed within 1 to 14 days (e.g., within 7 days) after collection. In some embodiments, the method further includes amplifying one or more of the isolated cfDNA fragments.

[0207] The bait oligonucleotides used to capture the converted cell-free DNA molecules can be any of the bait oligonucleotides described herein, such as those described for the various compositions described herein. In some embodiments, each of the plurality of different bait oligonucleotides in the panel is at least 45 nucleotides in length (e.g., at least 60, 75, 80, 90, 100, 110, or 120 nucleotides in length). In some embodiments, each bait oligonucleotide in the plurality of probes is no longer than 130, 140, 150, 200, 250, or 300 bases in length. In some embodiments, each bait oligonucleotide is 45-300, 60-200, or 75-150 nucleotides in length. In some embodiments, the bait oligonucleotide is at least 50 nucleotides in length. In some embodiments, the bait oligonucleotide is at least 60 nucleotides in length. In some embodiments, the bait oligonucleotide is at least 75 nucleotides in length. In some embodiments, the designated length of the bait oligonucleotide is a length designed to be complementary to a portion of a target genome sequence or its converted DNA molecule.

[0208] In some embodiments, the cancer assay panel is designed to target at least 500, 1000, 1500, 5000, 10,000, 12,500, 15,000, 17,000, 19,000, or more target genomic regions. In some embodiments, the cancer assay panel is designed to target fewer than 25,000, 20,000, 17,000, 15,000, 12,500, 10,000, or fewer target genomic regions. In some embodiments, the cancer assay panel is designed to target 5,000-30,000, 10,000-25,000, 12,500-20,000, or 15,000-20,000 target genomic regions. In some embodiments, the cancer assay panel is designed to target 100 to 500, 500 to 1000, 1500 to 5000, or 5000 to 10,000 target genomic regions. In some embodiments, the cancer assay panel is designed to target at least 500 target genomic regions. In some embodiments, the cancer assay panel is designed to target at least 1,000 target genomic regions. In some embodiments, the cancer assay panel is designed to target at least 10,000 target genomic regions. In some embodiments, the cancer assay panel is designed to target at least 15,000 target genomic regions. In some embodiments, the cancer assay panel is designed to target fewer than 20,000 target genomic regions.

[0209] In some embodiments, bait oligonucleotides are configured to hybridize to converted DNA molecules (e.g., converted cfDNA molecules) corresponding to or derived from one or more genomic regions. Thus, the bait oligonucleotides can have a sequence different from that of the targeted genomic region. For example, DNA with unmethylated CpG sites can be converted to contain UpG instead of CpG by deamination (e.g., by treatment with cytosine deaminase or bisulfite). As a result, probes for such targets can be configured to hybridize to sequences containing UpG instead of naturally occurring unmethylated CpG. Thus, the site complementary to the unmethylated site in the probe may contain CpA instead of CpG, and some probes targeting hypomethylated sites where all methylated sites are unmethylated may lack a guanine (G) base. In some embodiments, at least 3%, 5%, 10%, 15%, or 20% of the probes do not contain CpG sequences. In some embodiments, at least 5% of the probes do not contain CpG sequences. In some embodiments, at least 10% of the probes do not contain CpG sequences.

[0210] In some embodiments, cancer assay panels are used to detect the presence or absence of cancer generally and / or provide a cancer classification, such as cancer type, cancer stage, such as I, II, III, or IV, or to provide a potential TOO of cancer origin. The panel may include probes targeting genomic regions that are differentially methylated between cancerous (pan-cancer) and non-cancerous specimens in general, or only in cancerous specimens of a specific cancer type (e.g., bladder cancer-specific targets). For example, in some embodiments, cancer assay panels are designed to include genomic regions that are differentially methylated based on converted (e.g., bisulfite) sequencing data generated from cfDNA and / or whole genome DNA of a collection of cancer and non-cancer individuals.

[0211] In some embodiments, each of the target genomic regions is differentially methylated in at least one of a plurality of cancer types. In some embodiments, the plurality of cancer types includes at least two cancer types (e.g., at least two, three, three, four, or more cancer types). In some embodiments, the plurality of cancer types includes urological cancer. In some embodiments, the plurality of cancer types includes one or more of bladder cancer, urothelial cancer, prostate cancer, or renal cancer.

[0212] Each probe (or probe pair) can be designed to target one or more target genome regions.Target genome regions can be selected based on several criteria designed to increase selective enrichment of informative cfDNA fragments while reducing noise and non-specific binding.Described herein are various filtering procedures for determining whether a target genome region is included.In some embodiments, two or more of the filtering procedures described herein are used in combination.

[0213] In some embodiments, cancer is detected based on sequencing data using a trained classifier. In some embodiments, the trained classifier detects a number of sequencing reads above a threshold for one or more target sequences identified as hypermethylated and / or hypomethylated in cfDNA fragments. In some embodiments, the trained classifier distinguishes subjects with cancer from subjects without cancer with a specified specificity. Thus, in some embodiments, a specimen from a subject with an unknown type of cancer (e.g., an undiagnosed cancer) can be used to identify which of various different types of cancer are likely to be present and which are likely to be absent. In some embodiments, the classifier is a binary classifier, a mixed model classifier, or a multilayer perceptron model classifier. In some embodiments, the classifier is a mixed model classifier. In some embodiments, the specified specificity for each of the multiple cancer types is 0.900 or greater (e.g., at least 0.950, 0.975, 0.980, 0.985, 0.990, 0.995, or greater). In some embodiments, application of the trained classifier comprises a sensitivity of at least 30% (e.g., at least 40%, 50%, 60%, 70%, 80%, or more) for each of the plurality of cancer types. In some embodiments, the trained classifier has a sensitivity of at least 30% in determining the likelihood of cancer with a specificity of 0.900 or greater for each of the plurality of cancer types. In some embodiments, the trained classifier has a sensitivity of at least 40% in determining the likelihood of cancer with a specificity of 0.990 or greater for each of the plurality of cancer types.

[0214] In some embodiments, the method includes treating the cfDNA molecule to distinguish between methylated and unmethylated nucleotides, thereby generating a converted cfDNA molecule. In some embodiments, the treatment includes deamination, such as treating the cfDNA molecule with cytosine deaminase or bisulfite. In some embodiments, the method includes treating the cfDNA molecule with bisulfite to generate the converted cfDNA molecule.

[0215] In some embodiments, each bait oligonucleotide is conjugated to a solid surface (e.g., a chip or a bead, such as a magnetic or paramagnetic bead) or to a non-nucleotide affinity moiety (e.g., a member of a binding pair). In some embodiments, such a conjugate facilitates separation of DNA molecules bound to the bait oligonucleotide from unbound DNA molecules. In some embodiments, the affinity moiety is biotin.

[0216] In some embodiments, the genomic regions can be selected to have at least 3, 5, or 7 methylation sites. In some embodiments, each target genomic region comprises at least 5 methylation sites. In some embodiments, the selected number of methylation sites (e.g., at least 5 methylation sites) are differentially methylated in at least one cancer type to be assayed by the panel. In some embodiments, the target genomic regions comprise at least 1, 2, 3, 4, 5, 10, 15, or more target genomic regions that comprise target sequences from one or more genes selected from Table 1. In some embodiments, the target genomic regions comprise target sequences of genes selected from TWIST1, EOMES, HOXA9, POU4F2, and ZNF154.

[0217] In some embodiments, the method further comprises diagnosing cancer in the subject. In some embodiments, the method further comprises selecting a treatment for the identified cancer type. In some embodiments, the method further comprises treating the subject for cancer. In some embodiments, the cancer type comprises bladder cancer, urothelial cancer, prostate cancer, or renal cancer. The specific treatment mode may depend on one or more of a variety of factors, such as the specific type of cancer detected, the TOO, location, and stage of the cancer. Non-limiting examples of treatments include surgical resection, radiation therapy, chemotherapy, and / or immunotherapy.

[0218] Analysis of sequence reads In some embodiments, alignment position information may be determined by aligning sequence reads with a reference genome using various methods. The alignment position information may indicate the start and end positions of a region in the reference genome that corresponds to the start and end nucleotide bases of a given sequence read. The alignment position information also includes the sequence read length, which can be determined from the start and end positions. The region in the reference genome may be associated with a gene or a segment of a gene.

[0219] In various embodiments, the sequence read includes a read pair designated R1 and R2. For example, the first read R1 may be sequenced from a first end of the nucleic acid fragment, while the second read R2 may be sequenced from a second end of the nucleic acid fragment. Thus, the nucleotide base pairs of the first read R1 and the second read R2 may consistently align (e.g., in reverse) with the nucleotide bases of the reference genome. The alignment position information derived from the read pair R1 and R2 may include the start position of the reference genome corresponding to the end of the first read (e.g., R1) and the end position of the reference genome corresponding to the end of the second read (e.g., R2). In other words, the start position and end position of the reference genome correspond to the possible positions to which the nucleic acid fragment may correspond in the reference genome. An output file having a SAM (sequence alignment map) format or a BAM (binary alignment map) format may be generated and output for further analysis.

[0220] From the sequence reads, the location and methylation state for each CpG site may be determined based on alignment with the reference genome. Furthermore, a methylation state vector may be generated for each fragment, specifying the location of the fragment in the reference genome (e.g., as specified by the location of the first CpG site in each fragment, or another similar metric), the number of CpG sites in the fragment, and the methylation state of each CpG site in the fragment, whether methylated (e.g., designated M), unmethylated (e.g., designated U), or indeterminate (e.g., designated I). The methylation state vector may be stored in temporary or persistent computer memory for later use and processing. Furthermore, duplicate reads or duplicate methylation state vectors from a single subject may be removed. In a further embodiment, it may be determined that certain fragments have one or more CpG sites exhibiting an indeterminate methylation state. Such fragments may be excluded from further processing or may be selectively included so that such uncertain methylation status is taken into account in downstream data models.

[0221] Figure 16B is an illustration of the process 100 of Figure 16A for obtaining a methylation state vector by sequencing a cfDNA fragment, according to one embodiment. As an example, the analytics system takes a cfDNA fragment 112. In this example, the cfDNA fragment 112 contains three CpG sites. As shown, the first and third CpG sites of the cfDNA fragment 112 are methylated 114. During processing step 120, the cfDNA fragment 112 is converted to generate a converted cfDNA fragment 122. During processing 120, the second CpG site, which is unmethylated, has its cytosine converted to uracil. However, the first and third CpG sites are not converted.

[0222] After conversion, a sequencing library is prepared 130 and sequenced 140 to generate sequence reads 142. The analytics system aligns 150 the sequence reads 142 to a reference genome 144, which provides context for the location in the human genome from which the fragment cfDNA originated. In this simplified example, the analytics system aligns 150 the sequence reads such that three CpG sites are correlated to CpG sites 23, 24, and 25 (arbitrary reference identification numbers are used for illustrative purposes). In this manner, the analytics system generates both information regarding the methylation status of all CpG sites on the cfDNA fragment 112 and information regarding where the CpG sites map to in the human genome. As shown, the methylated CpG sites on the sequence read 142 are read as cytosines. In this example, because cytosine appears only at the first and third CpG sites in sequence read 142, it is possible to infer that the first and third CpG sites were methylated in the original cfDNA fragment. The second CpG site is read as thymine (U is converted to T during the sequencing process), and therefore it can be inferred that the second CpG site was unmethylated in the original cfDNA fragment. Using these two pieces of information, methylation state and location, the analytics system generates 160 a methylation state vector 152 for fragment cfDNA 112. In this example, the resulting methylation state vector 152 is: <M 23 ,U 24 ,M 25 >, where M corresponds to a methylated CpG site, U corresponds to an unmethylated CpG site, and the subscript numbers correspond to the position of each CpG site in the reference genome.

[0223] Cancer detection The sequence reads obtained by the methods provided herein can be further processed by automated algorithms. For example, an analytics system is used to receive sequencing data from the sequencer and perform various aspects of the processing as described herein. The analytics system may be one of a personal computer (PC), a desktop computer, a laptop computer, a notebook, a tablet PC, or a mobile device. The computing device can be communicatively coupled to the sequencer via wireless, wired, or a combination of wireless and wired communication technologies. Generally, the computing device comprises a processor and a memory that stores computer instructions that, when executed by the processor, cause the processor to perform the steps as described in the remainder of this document. Generally, the amount of genetic data and data derived therefrom is sufficiently large, and the amount of computing power required is so great that it is impossible to perform on paper or by the human brain alone.

[0224] Clinical interpretation of the methylation status of targeted genomic regions is a process that may include classifying the clinical effect of each or a combination of methylation states and reporting the results in a manner meaningful to medical professionals. Clinical interpretation can be based on comparison of sequence reads with databases specific to cancer or non-cancer subjects, and / or based on the number and type of cfDNA fragments with cancer-specific methylation patterns identified from the specimen. In some embodiments, targeted genomic regions are ranked and classified based on their likelihood of being differentially methylated in cancer specimens, and the ranking or classification is used in the interpretation process. Ranking and classification may include (1) the type of clinical effect, (2) the strength of evidence of effect, and (3) the size of the effect. Analysis of sequence reads can employ various methods for clinical analysis and interpretation of genomic data. In some other embodiments, clinical interpretation of the methylation status of such differentially methylated regions can be based on machine learning techniques that interpret current specimens based on classification or regression methods trained using the methylation status of such differentially methylated regions from specimens of cancer and non-cancer patients with known cancer conditions, cancer types, cancer stages, TOO, etc.

[0225] Clinically meaningful information may include the presence or absence of cancer in general, the presence or absence of a particular type of cancer, the stage of cancer, or the presence or absence of other types of disease. In some embodiments, the information relates to the presence or absence of urological cancer. In some embodiments, the information relates to the presence or absence of one or more cancer types selected from the group consisting of bladder cancer, urothelial cancer, renal cancer, and prostate cancer.

[0226] cancer classifier In some examples, the assay panels described herein can be used with a cancer type classifier to predict the disease state of a sample, such as cancer or non-cancer prediction, tissue of origin prediction, and / or indeterminate prediction. In some examples, the cancer type classifier can generate features based on sequence reads by taking into account methylated or unmethylated fragments of DNA in a certain genomic region of interest. For example, if the cancer type classifier determines that the methylation pattern in a fragment is similar to that of a certain cancer type, then the cancer type classifier can set the feature for that fragment to 1; otherwise, if no such fragment exists, then the feature can be set to 0. In this way, the cancer type classifier can generate a set of binary features for each sample (for example, 30,000 features). Furthermore, in some examples, inputting all or a portion of the set of binary features for a sample into the cancer type classifier can provide a set of probability scores, such as one probability score for one cancer type class and one probability score for one non-cancer type class. Furthermore, in some instances, the cancer type classifier may incorporate or otherwise use thresholding to determine whether a specimen should be called cancer or non-cancer, and / or uncertainty thresholding to reflect the degree of confidence in a specific TOO call, as further described below.

[0227] To train a cancer type classifier, an analytics system (e.g., analytics system 800, FIG. 17B) obtains a set of training samples. In some examples, the training samples include one or more fragment files (e.g., files containing sequence read data), a label corresponding to the sample's cancer type (TOO) or non-cancerous status, and / or the sample's individual's gender. The analytics system can utilize the training set to train a cancer type classifier to predict the sample's disease status.

[0228] In some instances, for training, the analytics system divides a genome (e.g., the entire genome) or a subset of a genome (e.g., a targeted methylation region) into regions. By way of example only, a portion of the genome can be separated into CpG "blocks," such that a new block begins whenever there is a separation between nearest neighboring CpGs, with at least a minimum separation distance (e.g., at least 500 bp). Further, in some instances, each block can be divided into 1000 bp regions, positioned so that adjacent regions have a particular amount of overlap (e.g., 50% or 500 bp).

[0229] Furthermore, in some examples, the analytics system may divide the training set into K subsets or K equal parts to use for K-fold cross-validation. In some examples, the folds may be balanced with respect to cancer / non-cancerous status, tissue of origin, cancer stage, age (e.g., grouped into 10-year buckets), and / or smoking status. In some examples, the training set is divided into 5 equal parts, so that 5 separate classifiers are trained, each trained on 4 out of 5 training samples and using 1 out of 5 for validation.

[0230] While training on the training set, the analytics system can fit a probability model for each cancer type (and for healthy cfDNA) to fragments derived from specimens of that type. As used herein, a "probability model" is any mathematical model capable of assigning probabilities to sequence reads based on the methylation state at one or more sites on the read. During training, the analytics system can fit sequence reads derived from one or more specimens from subjects with known diseases and use it to determine sequence read probabilities that are indicative of disease state using methylation information or a methylation state vector. Specifically, in some cases, the analytics system determines the observed methylation rate for each CpG site in the sequence read. The methylation rate corresponds to the ratio or percentage of base pairs that are methylated within the CpG site. The trained probability model can be parameterized by the product of methylation rates. In general, any known probability model that assigns probabilities to sequence reads from a specimen can be used. For example, the probability model may be a binomial model (where every site (e.g., CpG site) on a nucleic acid fragment is assigned a probability of methylation), or an independent site model (where the methylation of each CpG is specified by a separate methylation probability, and methylation at one site on a nucleic acid fragment is assumed to be independent of methylation at one or more other sites).

[0231] In some instances, the probability model is a Markov model, where the probability of methylation at each CpG site depends on the methylation state of some number of CpG sites preceding it in the sequence read or in the nucleic acid molecule from which the sequence read is derived (see, e.g., U.S. Patent Application No. 16 / 352,602, filed March 13, 2019, entitled "Anomalous Fragment Detection and Classification," which is incorporated by reference in its entirety and can be used in various embodiments).

[0232] In some examples, the probabilistic model is a "mixture model" that is fitted using a mixture of components from an underlying model. For example, in some embodiments, the mixture components can be determined using a multiple independent site model, where methylation (e.g., methylation rate) at each CpG site is assumed to be independent of methylation at other CpG sites. Utilizing the independent site model, the probability assigned to a sequence read, or the nucleic acid molecule from which it is derived, is the product of the methylation probability at each CpG site at which the sequence read is methylated and the 1-methylation probability at each CpG site at which the sequence read is unmethylated. According to this example, the analytics system determines the methylation rate for each of the mixture components. The mixture model is parameterized by the sum of the mixture components, each associated with a product of methylation rates. A probabilistic model Pr of n mixture components can be expressed as follows:

number

number

[0233] In some examples, the analytics system fits a probabilistic model using maximum likelihood estimation to find a set of parameters {β} that maximizes the log-likelihood of all fragments from the disease state, subject to applying a regularization penalty to each methylation probability with a regularization strength r. ki ,f kThe maximized amount for the sum of N fragments can be expressed as:

number

[0234] In some cases, the analytics system performs fitting separately for each cancer type and for healthy cfDNA. According to various aspects of the subject disclosure, other means can be used to fit the probabilistic model, or to identify parameters that maximize the log-likelihood of all sequence reads derived from a reference sample. For example, in some cases, Bayesian fitting (e.g., using Markov chain Monte Carlo) is used, in which each parameter is not assigned a single value, but instead associated with a distribution. In some cases, gradient-based optimization is used, in which the gradient of the likelihood (or log-likelihood) with respect to the parameter value is used to step through the parameter space toward an optimum. Still in some cases, expectation maximization is used, in which a set of latent parameters (such as the identity of the mixture component from which each fragment originates) is set to its expected value under the previous model parameters, and then the model parameters are assigned to maximize the likelihood conditional on the assumed values ​​of those latent variables. This two-step process is then repeated until convergence occurs.

[0235] Additionally, in some examples, the analytics system can generate features for each sample in the training set. For example, for each sample (regardless of label), for each region, for each cancer type, for each fragment, the analytics system can evaluate the log-likelihood ratio R with the fitted probability model according to:

number

[0236] In some examples, the analytics system can select certain features for inclusion in the feature vector for each sample. For example, for each distinct pair of cancer types, the analytics system can designate one type as a "positive type" and the other as a "negative type" and rank the features by their ability to distinguish between the types. In some cases, the ranking is based on mutual information calculated by the analytics system. For example, mutual information can be calculated using an estimated proportion of samples of positive and negative types (e.g., cancer types A and B) that are expected to have non-zero features in the resulting assay. For example, if a feature is frequently present in healthy cfDNA, the analytics system determines that it is unlikely to be frequently present in cfDNA associated with various cancers. Consequently, the feature may be a weak measure in terms of distinguishing between disease states. In calculating mutual information I, variable X is a certain feature (e.g., binary) and variable Y corresponds to a disease state, e.g., cancer type A or B:

number

[0237] In some examples, only features corresponding to the positive type are included in the ranking if their predicted occurrence is higher in the positive type compared to the negative type. For example, if "liver" is the positive type and "breast" is the negative type, then only a "liver_x" feature is considered if its predicted occurrence in liver cfDNA is higher than its predicted occurrence in breast cfDNA. Further, in some examples, for each region and for each cancer type pair (including non-cancer as a negative type), the analytics system retains only the best-performing tier. Further, in some examples, the analytics system converts feature values ​​by binarization, so that any feature value greater than 0 is set to 1, resulting in all features being either 0 or 1.

[0238] In some examples, the analytics system trains a multinomial logistic regression classifier on the training data folds to generate predictions for the held-out data. For example, for each of the K folds, one logistic regression can be trained for each combination of hyperparameters. Such hyperparameters include an L2 penalty and / or top-K (e.g., the number of highly ranked regions to retain for each tissue type pair (including non-cancer) when ranked by the mutual information procedure outlined above). For each set of hyperparameters, performance is evaluated against predictions by cross-validation on the full training set, and the best-performing hyperparameter set is selected and retrained on the full training set. In some examples, the analytics system uses log-loss as a performance metric, where log-loss is calculated by taking the negative logarithm of the prediction for the correct label for each sample and then summing over all samples (i.e., a perfect prediction of 1.0 for the correct label would result in a log-loss of 0).

[0239] To generate predictions for new samples, feature values ​​are calculated using the same method described above, but restricted to the features (region / positive class combinations) selected under the top K values ​​of selection. The generated features are then used to make predictions using the logistic regression model trained above.

[0240] In some examples, analytics trains a two-stage classifier. For example, an analytics system trains a binary cancer classifier to distinguish between the labels cancer and non-cancer based on the feature vectors of training samples. In this case, the binary classifier outputs a prediction score indicating the likelihood of cancer or non-cancer. In another example, an analytics system trains a multi-class cancer classifier to distinguish between many cancer types. In this multi-class cancer classifier, the cancer classifier is trained to determine a cancer prediction including a predictive value for each of the classified cancer types. The predictive value may correspond to the likelihood that a given sample has each of the cancer types. For example, the cancer classifier returns a cancer prediction including a predictive value for urinary cancer and non-cancer. For example, the cancer classifier may return a cancer prediction for a test sample including a prediction score for bladder or urothelial cancer, renal cancer, prostate cancer, and / or non-cancer.

[0241] The analytics system can train the cancer classifier according to any one of several methods. As an example, a binary cancer classifier may be an L2 regularized logistic regression classifier trained using a logarithmic loss function. As another example, a multi-cancer (TOO) classifier may be multinomial logistic regression. In practice, either type of cancer classifier may be trained using other techniques. These techniques are numerous, including the possibility of using kernel methods, machine learning algorithms such as multilayer neural networks, and the like. In particular, for various embodiments, methods such as those described in PCT / US2019 / 022122 and U.S. Patent Application No. 16 / 352,602 (which are incorporated by reference herein in their entireties) may be used. Still further, in some examples, the TOO classifier is trained only on cancer samples that are successfully called cancer by the binary classifier, thereby ensuring sufficient cancer signal in the cancer samples. On the other hand, in some examples, the binary classifier is trained on training samples independent of TOO.

[0242] Exemplary Sequencer and Analytics System 17A is a flow chart of a system and apparatus for sequencing a nucleic acid sample, according to one embodiment. This schematic flow chart includes apparatus such as a sequencer 820 and an analytics system 800. The sequencer 820 and analytics system 800 can work in tandem to perform one or more steps in the processes described herein.

[0243] In various embodiments, sequencer 820 receives the enriched nucleic acid sample 810. As shown in FIG. 17A , sequencer 820 includes a graphical user interface 825 that allows a user to interact with the sequencer 820 regarding a particular task (e.g., start sequencing or stop sequencing), as well as another loading station 830 for loading a sequencing cartridge containing the enriched fragment sample and / or for loading buffers necessary to perform a sequencing assay. Thus, once sequencer 820 has provided the necessary reagents and sequencing cartridges to its loading station 830, a user can initiate sequencing by interacting with the graphical user interface 825 of sequencer 820. Once initiated, sequencer 820 performs sequencing and outputs sequence reads of the enriched fragments from nucleic acid sample 810.

[0244] In some embodiments, the sequencer 820 is communicatively coupled to the analytics system 800. The analytics system 800 includes several computing devices used to process sequence reads for various applications, such as determining the methylation status at one or more CpG sites, calling variants, or quality control. The sequencer 820 may provide the sequence reads to the analytics system 800 in BAM file format. The analytics system 800 may be communicatively coupled to the sequencer 820 via wireless, wired, or a combination of wireless and wired communication technologies. Generally, the analytics system 800 comprises a processor and a non-transitory computer-readable storage medium storing computer instructions that, when executed by the processor, cause the processor to process sequence reads or perform one or more steps of any of the methods or processes disclosed herein.

[0245] In some embodiments, alignment position information may be determined by aligning sequence reads with a reference genome using various methods. The alignment position may generally describe the start and end positions of a region in the reference genome corresponding to the start and end nucleotide bases of a given sequence read. Corresponding to methylation sequencing, the alignment position information may be generalized to indicate the first and last CpG sites included in the sequence read according to the alignment with the reference genome. The alignment position information may further indicate the methylation status and positions of all CpG sites in a given sequence read. Regions in the reference genome may be associated with genes or gene segments; thus, the analytics system 800 may label the sequence read with one or more genes that align with the sequence read. In one embodiment, the fragment length (or size) is determined from the start and end positions.

[0246] In various embodiments, for example, when a paired-end sequencing process is used, the sequence read includes a read pair designated R_1 and R_2. For example, the first read R_1 may be sequenced from a first end of a double-stranded DNA (dsDNA) molecule, while the second read R_2 may be sequenced from a second end of the double-stranded DNA (dsDNA). Thus, the nucleotide base pairs of the first read R_1 and the second read R_2 may be consistently aligned (e.g., in reverse) with the nucleotide bases of the reference genome. The alignment position information derived from the read pair R_1 and R_2 may include a start position corresponding to the end of the first read (e.g., R_1) in the reference genome and an end position corresponding to the end of the second read (e.g., R_2) in the reference genome. In other words, the start position and end position in the reference genome correspond to possible positions of nucleic acid fragments in the reference genome. In one embodiment, the read pair R_1 and R_2 can be assembled into fragments, which are used for subsequent analysis and / or classification. An output file having a SAM (Sequence Alignment Map) format or a BAM (Binary Alignment Map) format may be generated and output for further analysis.

[0247] Referring now to FIG. 17B, FIG. 17B is a block diagram of an analytics system 800 for processing DNA samples, according to one embodiment. The analytics system implements one or more computing devices used to analyze DNA samples. Analytics system 800 includes a sequence processor 840, a sequence database 845, a model database 855, a model 850, a parameter database 865, and a score engine 860. In some embodiments, analytics system 800 performs one or more steps in the process of 300 of FIG. 12A , 340 of FIG. 12B , 400 of FIG. 13 , 500 of FIG. 14 , 600 of FIG. 15A , or 680 of FIG. 15B , and other processes described herein.

[0248] The sequence processor 840 generates methylation state vectors for fragments from the specimen. For each CpG site on the fragment, the sequence processor 840 generates a methylation state vector for each fragment that specifies the location of the fragment in the reference genome, the number of CpG sites in the fragment, and the methylation state of each CpG site in the fragment, whether methylated, unmethylated, or uncertain, via process 300 of Figure 12A. The sequence processor 840 may store the methylation state vectors for the fragments in a sequence database 845. The data in the sequence database 845 may be organized so that the methylation state vectors from the specimen are related to each other.

[0249] Additionally, multiple different models 850 may be stored in model database 855 or retrieved for use with test specimens. In one example, the model is a trained cancer classifier that uses feature vectors derived from abnormal fragments to determine a cancer prediction for a test specimen. The training and use of cancer classifiers is discussed elsewhere herein. Analytics system 800 may train one or more models 850 and store various trained parameters in parameter database 865. Analytics system 800 stores models 850 along with functions in model database 855.

[0250] During inference, the score engine 860 uses one or more models 850 and returns an output. The score engine 860 accesses the models 850 in the model database 855 along with trained parameters from the parameter database 865. For each model, the score engine receives appropriate inputs for the model and calculates an output based on the received inputs, parameters, and each model's function relating the inputs and outputs. In some use cases, the score engine 860 also calculates metrics that correlate with confidence in the calculated output from the model. In other cases, the score engine 860 calculates other intermediate values ​​used in the model.

[0251] Cancer and Treatment Monitoring In some embodiments, cell-free nucleic acid samples (e.g., cfDNA from a urine sample) from a cancer patient can be obtained and analyzed at first and second time points to, for example, monitor cancer progression, determine whether the cancer is in remission (e.g., after treatment), monitor or detect residual disease or disease recurrence, or monitor the effectiveness of a treatment (e.g., a therapeutic drug). In some embodiments, the first time point in cancer monitoring is before cancer treatment (e.g., before resection surgery or therapeutic intervention), and the second time point is after cancer treatment (e.g., after resection surgery or therapeutic intervention), and the method is used to monitor the effectiveness of the treatment. For example, if the second likelihood or probability score decreases compared to the first likelihood or probability score, the treatment is considered successful. However, if the second likelihood or probability score increases compared to the first likelihood or probability score, the treatment is considered unsuccessful. In some embodiments, both the first and second time points are before cancer treatment (e.g., before resection surgery or therapeutic intervention). In yet other embodiments, both the first and second time points are after cancer treatment (e.g., before resective surgery or therapeutic intervention), and the method is used to monitor the effectiveness of the treatment or loss of effectiveness of the treatment.

[0252] Test specimens can be obtained and analyzed from cancer patients according to the methods of the present invention over any desired set of time points to monitor the patient's cancer status. In some embodiments, the first and second time points are 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 hours, 1, 2, 3, 4, 5, 10, 15, 20, 25, or 30 days, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 months, or 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 24, 24.5, 25, 25.5, 26, 26.5, 27, 27.5, 28, 28.5, 29, 29.5, or 30 years. In other embodiments, test specimens can be obtained from a patient at least once every three months, at least once every six months, at least once every year, at least once every two years, at least once every three years, at least once every four years, or at least once every five years.

[0253] treatment In some embodiments, information (e.g., likelihood or probability score) obtained from any of the methods described herein may be used to make or influence clinical decisions (e.g., diagnosing cancer, selecting treatment, determining treatment efficacy, etc.). For example, in one embodiment, if the likelihood or probability score exceeds a threshold, a physician may prescribe an appropriate treatment (e.g., resective surgery, radiation therapy, chemotherapy, and / or immunotherapy). In some embodiments, information such as the likelihood or probability score may be provided as a readout to a physician or a subject.

[0254] In one aspect, the method includes selecting a subject having or at high risk of developing a cancer type, and administering to the subject a treatment effective to treat the cancer type, wherein (a) selecting includes identifying the subject as a source of a urinary cell-free DNA (cfDNA) specimen comprising one or more target genomic regions that are differentially methylated above a threshold level for the presence of cancer; (b) the one or more target genomic regions comprise one or more target sequences of one or more genes selected from Table 1; (c) each target sequence is at least 25 nucleotides in length; (d) the cancer is bladder cancer, prostate cancer, or renal cancer; and (e) the treatment comprises surgical resection, radiation therapy, chemotherapy, immunotherapy, or any combination thereof.

[0255] A classifier (as described herein) can be used to determine a likelihood or probability score that the sample feature vector is from a subject with cancer. In one embodiment, when the likelihood or probability exceeds a threshold (e.g., the level of a reference sample from a subject with cancer), an appropriate treatment (e.g., resective surgery or a therapeutic agent) may be prescribed. For example, in one embodiment, when the likelihood or probability score is 60 or greater, one or more appropriate treatments are prescribed. In another embodiment, when the likelihood or probability score is 65 or greater, 70 or greater, 75 or greater, 80 or greater, 85 or greater, 90 or greater, or 95 or greater, one or more appropriate treatments are prescribed. In other embodiments, the cancer log-odds ratio may indicate the effectiveness of a cancer treatment. For example, an increase in the cancer log-odds ratio over time (e.g., after treatment, at a second time point) may indicate that the treatment was ineffective. Similarly, a decrease in the cancer log-odds ratio over time (e.g., after treatment, at a second time point) may indicate that the treatment was successful. In another embodiment, if the cancer log odds ratio is greater than 1, greater than 1.5, greater than 2, greater than 2.5, greater than 3, greater than 3.5, or greater than 4, one or more appropriate treatments are prescribed. In some embodiments, the threshold level for the presence of cancer is determined by a classifier trained on sequencing reads of converted DNA from subjects with cancer. Non-limiting examples of classifiers are described herein. Classification can be based on one or more target genomic regions, as described herein in connection with various aspects of the present disclosure.

[0256] In some embodiments, the therapy is one or more cancer therapeutic agents selected from the group consisting of chemotherapeutic agents, targeted cancer therapeutic agents, differentiation therapeutic agents, hormonal therapeutic agents, and immunotherapeutic agents. For example, the therapy can be one or more chemotherapeutic agents selected from the group consisting of alkylating agents, antimetabolites, anthracyclines, antitumor antibiotics, cytoskeleton-disrupting agents (taxanes), topoisomerase inhibitors, mitotic inhibitors, corticosteroids, kinase inhibitors, nucleotide analogs, platinum-based agents, and any combination thereof. In some embodiments, the therapy is one or more targeted cancer therapeutic agents selected from the group consisting of signal transduction inhibitors (e.g., tyrosine kinase and growth factor receptor inhibitors), histone deacetylase (HDAC) inhibitors, retinoic acid receptor agonists, proteosome inhibitors, angiogenesis inhibitors, and monoclonal antibody conjugates. In some embodiments, the therapy is one or more differentiation therapeutic agents, including retinoids such as tretinoin, alitretinoin, and bexarotene. In some embodiments, the treatment is one or more hormonal therapy agents selected from the group consisting of antiestrogens, aromatase inhibitors, progestins, estrogens, antiandrogens, and GnRH agonists or analogs. In one embodiment, the treatment is one or more immunotherapeutic agents selected from the group including monoclonal antibody therapy, e.g., rituximab (RITUXAN) and alemtuzumab (CAMPATH), non-specific immunotherapy and adjuvants, e.g., BCG, interleukin-2 (IL-2), and interferon alpha, immunomodulatory agents, e.g., thalidomide and lenalidomide (REVLIMID). It is within the ability of a skilled physician or oncologist to select an appropriate cancer treatment agent based on the type of tumor, the stage of the cancer, previous cancer treatment or exposure to therapeutic agents, and other characteristics of the cancer.

[0257] Computer Systems and Devices In one aspect, the present disclosure provides a computer system for implementing one or more steps of the methods disclosed herein. In another aspect, the present disclosure provides a non-transitory computer-readable medium having stored thereon computer-readable instructions for implementing one or more steps of the methods disclosed herein.

[0258] The methods of the present disclosure can be implemented using software, hardware, firmware, hardwiring, or any combination thereof. The features that implement the functionality may also be in various physical locations, including those that are distributed such that the locations of the functionality are implemented in different physical locations (e.g., an imaging device in one room and a host workstation in another room or in separate buildings, with, for example, wireless or wired connections).

[0259] Processors suitable for executing a computer program include, by way of example, both general-purpose and special-purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor receives instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to, one or more mass storage devices, such as magnetic, magneto-optical, or optical disks, for storing data therefrom, receiving data therefrom, or transferring data thereto, or both. Information carriers suitable for embodying computer program instructions and data include, by way of example, all forms of non-volatile memory, including semiconductor memory devices (e.g., EPROM, EEPROM, solid-state drives (SSD), and flash memory devices); magnetic disks (e.g., internal hard disks or removable disks); magneto-optical disks; and optical disks (e.g., CD and DVD disks). The processor and memory may be provided by, or incorporated in, special-purpose logic circuitry.

[0260] To provide for user interaction, the subject matter described herein can be implemented in a computer having I / O devices, such as a CRT, LCD, LED, or projection device for displaying information to a user, and input or output devices, such as a keyboard and pointing device (e.g., a mouse or trackball), that can enable a user to provide input to the computer. Other types of devices can be used to provide for user interaction as well. For example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input from the user can be received in any form, including auditory, speech, or tactile input.

[0261] The subject matter described herein can be implemented in a computing system that includes back-end components (e.g., a data server), middleware components (e.g., an application server), or front-end components (e.g., a client computer having a graphical user interface or web browser that allows a user to interact with an implementation of the subject matter described herein), or any combination of such back-end, middleware, and front-end components. The components of the system can be interconnected via a network by any form or medium of digital data communication, e.g., a communications network. For example, a reference set of data may be stored at a remote location, and a computer may communicate across the network and thereby access the reference dataset for comparison purposes. In other embodiments, however, the reference dataset may be stored locally within a computer, and the computer may access the reference dataset within its CPU for comparison purposes. Examples of communications networks include, but are not limited to, a cellular network (e.g., 3G or 4G), a local area network (LAN), and a wide area network (WAN), e.g., the Internet.

[0262] The subject matter described herein may be implemented as one or more computer program products, such as one or more computer programs tangibly embodied in an information carrier (e.g., a non-transitory computer-readable medium) for execution by or to control the operation of a data processing device (e.g., a programmable processor, computer, or multiple computers). The computer programs (also known as programs, software, software applications, apps, macros, or code) may be written in any form of programming language, including compiled or interpreted languages ​​(e.g., C, C++, Perl), and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. The systems and methods of the present disclosure may include instructions written in any suitable programming language known in the art, including, without limitation, C, C++, Perl, Java, ActiveX, HTML5, Visual Basic, or JavaScript.

[0263] A computer program does not necessarily correspond to a file. A program can be stored in a file or portion of a file that holds other programs or data, in a single file dedicated to the program in question, or in multiple cooperating files (e.g., a file that stores one or more modules, subprograms, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers at one site or distributed across multiple sites and interconnected by a communications network.

[0264] A file may be, for example, a digital file stored on a hard drive, SSD, CD, or other tangible, non-transitory medium. A file may be sent from one device to another over a network (e.g., as packets are sent from a server to a client, e.g., through a network interface card, modem, wireless card, etc.).

[0265] Writing a file according to the present disclosure involves transforming a tangible, non-transitory, computer-readable medium, for example, by adding, removing, or rearranging particles (e.g., particles with a net charge or dipole moment are magnetized by a read / write head) so that the pattern then corresponds to a new set of information related to an objective physical phenomenon desired and useful to the user. In some embodiments, writing involves a physical transformation of material on the tangible, non-transitory, computer-readable medium (e.g., by specific optical properties that then allow an optical read / write device to read a new, useful set of information, e.g., burning a CD-ROM). In some embodiments, writing to a file involves transforming a physical flash memory device, such as a NAND flash memory device, and storing information in an array of storage cells made of floating-gate transistors by transforming the physical elements. Methods of writing to a file are well known in the art and can be invoked manually or automatically, for example, by a program or by a save command from software or a write command from a programming language.

[0266] Suitable computing devices typically include mass memory, at least one graphical user interface, at least one display device, and typically include communication between devices. Mass memory exemplifies a type of computer-readable medium, i.e., computer storage media. Computer storage media can include volatile, nonvolatile, removable, and non-removable media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include RAM, ROM, EEPROM, flash memory, or other memory technology, CD-ROM, digital versatile disks (DVDs) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage, or other magnetic storage devices, radio frequency identification (RFID) tags or chips, or any other medium that can be used to store the desired information and that can be accessed by a computing device.

[0267] The functionality described herein may be implemented using software, hardware, firmware, hardwiring, or any combination thereof. Any of the software may be in various physical locations, including being distributed such that portions of the functionality are implemented in different physical locations.

[0268] A computer system for implementing some or all of the described methods of the present invention may include one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), or both), a main memory, and a static memory, communicating with each other via a bus, as one of ordinary skill in the art would recognize as necessary or best suited for the performance of the methods of the present disclosure.

[0269] A processor will generally include a chip, such as a single-core or multi-core chip, to provide a central processing unit (CPU). The processing may be provided by a chip from Intel or AMD.

[0270] The memory may include one or more machine-readable devices that store one or more sets of instructions (e.g., software) that, when executed by one or more processors of any one of the disclosed computers, can accomplish some or all of the methodologies or functions described herein. The software may also reside, in whole or at least partially, in main memory and / or within the processor(s) during its execution by the computer system. Preferably, each computer includes non-transitory memory, such as a solid-state drive, flash drive, disk drive, hard drive, etc.

[0271] While the machine-readable device may be a single medium in an exemplary embodiment, the term "machine-readable device" should be interpreted to include a single medium or multiple media (e.g., centralized or distributed databases and / or associated caches and servers) storing one or more sets of instructions and / or data. These terms should also be interpreted to include any one or more media capable of storing, encoding, or retaining a set of instructions for execution by a machine and causing the machine to perform any one or more of the methodologies of this disclosure. Thus, these terms should be interpreted to include, but are not limited to, one or more solid-state memories (e.g., subscriber identity module (SIM) cards, secure digital cards (SD cards), microSD cards, or solid-state drives (SSDs)), optical and magnetic media, and / or any other tangible storage medium or media.

[0272] A computer of the present disclosure will generally include one or more I / O devices, such as, for example, one or more of a video display unit (e.g., a liquid crystal display (LCD) or cathode ray tube (CRT)), an alphanumeric input device (e.g., a keyboard), a cursor control device (e.g., a mouse), a disk drive unit, a signal generating device (e.g., a speaker), a touch screen, an accelerometer, a microphone, a cellular radio frequency antenna, and a network interface device, which may be, for example, a network interface card (NIC), a Wi-Fi card, or a cellular modem.

[0273] Any of the software may be located in various physical locations, including being distributed so that portions of the functionality are implemented in different physical locations.

[0274] Additionally, the system of the present disclosure may be provided with reference data. Any suitable genomic data may be stored for use within the system. Examples include, but are not limited to, comprehensive multidimensional maps of key genomic alterations in major cancer types and subtypes from The Cancer Genome Atlas (TCGA); catalogs of abnormal genomes from the International Cancer Genome Consortium (ICGC); catalogs of somatic mutations in cancer from COSMIC; the latest builds of the human genome and other popular model organisms; the latest reference SNPs from dbSNP; gold-standard indels from the 1000 Genomes Project and the Broad Institute; exome capture kit annotations from Illumina, Agilent, Nimblegen, and Ion Torrent; transcript annotations; and small-scale test data (e.g., for new users) for validation work in the pipeline.

[0275] In some embodiments, data is made available within the context of a database included in the system. Any suitable database structure may be used, including a relational database, an object-oriented database, etc. In some embodiments, the reference data is stored in a relational database, such as a "not-only-SQL" (no-SQL) database. In various embodiments, graph databases are included within the scope of the systems of the present disclosure. It should also be understood that the term "database," as used herein, is not limited to one single database; rather, multiple databases may be included in the system. For example, according to embodiments of the present disclosure, the database may include 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or more separate databases, including any integer number of databases therein. For example, one database may include public reference data, a second database may include test data from patients, a third database may include data from healthy subjects, and a fourth database may include data from diseased subjects with known conditions or disorders. It should be understood that any other database configurations for the data included in the databases are also contemplated by the methods described herein.

[0276] Illustrative Embodiments The present disclosure provides the following exemplary embodiments.

[0277] Embodiment 1. A method of sequencing cell-free nucleic acid molecules of a subject, comprising: (a) Treating the urine specimen so that cell lysis is inhibited; (b) separating the cell-free nucleic acid molecules in the processed urine specimen from the cells in the processed urine specimen, thereby producing a purified urine specimen containing the cell-free nucleic acid molecules; (c) concentrating cell-free nucleic acid molecules in the purified urine specimen by passing at least a portion of the purified urine specimen through a filter, (i) the concentration producing a filtrate and a residual urine specimen, and (ii) the residual urine specimen having an increased concentration of cell-free nucleic acid molecules; (d) isolating cell-free nucleic acid molecules from the residual urine specimen; and (e) sequencing the isolated cell-free nucleic acid molecules; A method comprising:

[0278] Embodiment 2. The method of embodiment 1, wherein treating the urine specimen to inhibit cell lysis comprises contacting the urine specimen with one or more preservation reagents.

[0279] Embodiment 3. The method of embodiment 1 or 2, wherein treating comprises treatment with a nuclease inhibitor, a formaldehyde quencher, or both.

[0280] Embodiment 4. The method of any one of embodiments 1-3, wherein processing comprises contacting the urine specimen with a composition comprising: (i) imidazolidinyl urea, EDTA, glycine, or a combination thereof; or (ii) sodium azide, EDTA, or a combination thereof.

[0281] Embodiment 5. The method of any one of embodiments 1-4, wherein separating comprises pelleting the cells in the processed urine specimen by centrifugation.

[0282] Embodiment 6. The method of any one of embodiments 1-5, wherein the filter is substantially impermeable to the passage of cell-free nucleic acids in the purified urine specimen and substantially permeable to salts.

[0283] Embodiment 7. The method of any one of embodiments 1 to 5, wherein the filter has a nominal molecular weight cutoff of 10 kD, 5 kD, 3 kD, or less.

[0284] Embodiment 8. The method of any one of embodiments 1-7, wherein the residual urine specimen has a concentration that is increased by at least 2-fold, at least 5-fold, at least 10-fold, or at least 15-fold compared to the purified urine specimen.

[0285] Embodiment 9. The method of any one of embodiments 1-8, wherein the residual urine specimen has a volume that is at least 50%, at least 75%, or at least 90% smaller than the volume of the processed urine specimen.

[0286] Embodiment 10. The method of embodiment 9, wherein the volume of the processed urine specimen is 5 mL, 10 mL, 15 mL, 20 mL, 30 mL, 40 mL, 50 mL, or more.

[0287] Embodiment 11. The method of any one of embodiments 1 to 10, further comprising freezing the residual urine specimen.

[0288] Embodiment 12. The method of any one of embodiments 1 to 11, wherein processing is completed within 120, 60, or 30 minutes after collection of the urine specimen; and optionally, separating and concentrating is completed within 7 days after collection.

[0289] Embodiment 13. The method of any one of embodiments 1 to 12, further comprising amplifying one or more of the isolated cell-free nucleic acid molecules.

[0290] Embodiment 14. The method of any one of embodiments 1 to 13, further comprising capturing the isolated cell-free nucleic acid molecule, or an amplification product thereof, by hybridization to a bait oligonucleotide.

[0291] Embodiment 15 The method of embodiment 14, further comprising separating cell-free nucleic acid molecules bound to the bait from unbound cell-free nucleic acid molecules.

[0292] Embodiment 16 The method of embodiment 15, wherein each bait oligonucleotide hybridizes to a target genomic region that is differentially methylated in cancer specimens compared to non-cancerous specimens.

[0293] Embodiment 17. The method of embodiment 16, wherein the differential methylation comprises at least 80% of the CpG sites in the target genomic region being methylated or unmethylated.

[0294] Embodiment 18. The method of embodiment 16 or 17, wherein the cancer is bladder cancer, prostate cancer, or renal cancer.

[0295] Embodiment 19. The method of any one of embodiments 14 to 18, wherein each bait oligonucleotide hybridizes to a target genomic region containing at least five methylation sites.

[0296] Embodiment 20. The method of any one of embodiments 14 to 19, wherein each bait oligonucleotide hybridizes to a target genomic region comprising a target sequence of a gene selected from Table 1, and the target sequence is at least 25, at least 35, or at least 45 nucleotides in length.

[0297] Embodiment 21. The method of embodiment 20, wherein the target genomic region comprises a target sequence of a gene selected from TWIST1, EOMES, HOXA9, POU4F2, and ZNF154.

[0298] Embodiment 22 The method of embodiment 20, wherein the bait oligonucleotides collectively hybridize to target sequences from at least 10 genes in Table 1.

[0299] Embodiment 23. The method of embodiment 20, wherein the bait oligonucleotides jointly hybridize to target sequences from: (a) a gene in Table 2 or Table 3; (b) a gene in Table 4; or (c) a gene in Table 5.

[0300] Embodiment 24. The method of any one of embodiments 1 to 23, wherein the cell-free nucleic acid molecule comprises cell-free DNA (cfDNA).

[0301] Embodiment 25. The method of embodiment 24, further comprising deaminating the cfDNA isolated in step (d) to produce a converted cfDNA molecule; optionally, the deaminating comprises treatment with cytosine deaminase or bisulfite.

[0302] Embodiment 26 The method of any one of embodiments 1 to 25, further comprising diagnosing cancer in the subject.

[0303] Embodiment 27. The method of embodiment 26, wherein the cancer is bladder cancer, prostate cancer, or renal cancer.

[0304] Embodiment 28 The method of embodiment 26 or 27, further comprising treating cancer in the subject.

[0305] Embodiment 29. The method of embodiment 28, wherein treating comprises surgical resection, radiation therapy, chemotherapy, and / or immunotherapy.

[0306] Embodiment 30. A method for detecting cancer cells in a subject, comprising: (a) capturing converted cell-free DNA (cfDNA) fragments, or amplification products thereof, from a urine specimen of a subject, (i) the bait oligonucleotide composition comprises a plurality of different bait oligonucleotides; (ii) each bait oligonucleotide of the plurality of different bait oligonucleotides hybridizes to a target sequence of a gene selected from Table 1, and the target sequence is at least 25 nucleotides in length; (b) separating the DNA bound to the bait from the unbound DNA; (c) sequencing the isolated DNA to generate sequencing reads; and (d) Detecting cancer cells with a trained classifier A method comprising:

[0307] Embodiment 31. The method of embodiment 30, wherein the trained classifier detects a number of sequencing reads above a threshold for one or more target sequences identified as hypermethylated and / or hypomethylated in the cfDNA fragments.

[0308] Embodiment 32 The method of embodiment 30 or 31, wherein the bait oligonucleotide is at least 45 nucleotides in length.

[0309] Embodiment 33. The method of any one of embodiments 30 to 32, wherein the trained classifier distinguishes subjects with cancer from subjects without cancer with a specified specificity.

[0310] Embodiment 34. The method of embodiment 33, wherein the classifier is a mixture model classifier.

[0311] Embodiment 35. The method of embodiment 33 or 34, wherein the specified specificity is 0.900 or greater.

[0312] Embodiment 36. The method of embodiment 35, wherein applying the trained classifier further comprises a sensitivity of 30% or greater.

[0313] Embodiment 37 The method of any one of embodiments 30 to 36, wherein the bait oligonucleotides collectively hybridize to target sequences from at least 10 genes in Table 1.

[0314] Embodiment 38. The method of any one of embodiments 30 to 36, wherein at least one of the bait oligonucleotides hybridizes to a target sequence of a gene selected from TWIST1, EOMES, HOXA9, POU4F2, and ZNF154.

[0315] Embodiment 39. The method according to any one of embodiments 30 to 36, (a) the cancer cells are bladder cancer cells and the bait oligonucleotides hybridize collectively to target sequences from the genes of Table 2 or Table 3; (b) the cancer cells are prostate cancer cells and the bait oligonucleotides collectively hybridize to target sequences from the genes in Table 4; or (c) The method, wherein the cancer cells are renal cancer cells and the bait oligonucleotides hybridize collectively to target sequences from the genes in Table 5.

[0316] Embodiment 40. The method of any one of embodiments 30 to 39, wherein the converted cfDNA molecule comprises cfDNA treated with cytosine deaminase or bisulfite.

[0317] Embodiment 41 The method of any one of embodiments 30 to 40, wherein each bait oligonucleotide is conjugated to a solid surface or to a non-nucleotide affinity moiety.

[0318] Embodiment 42 The method of any one of embodiments 30-41, wherein the differential methylation comprises hypermethylation in cancer specimens compared to non-cancerous specimens.

[0319] Embodiment 43. The method of any one of embodiments 30 to 42, wherein each target genomic region comprises at least five methylation sites.

[0320] Embodiment 44. The method of any one of embodiments 30 to 43, further comprising obtaining converted cfDNA fragments or amplification products thereof, and wherein obtaining further comprises: (i) treating the urine specimen to inhibit cell lysis; (ii) separating the cfDNA fragments in the processed urine specimen from cells in the processed urine specimen, thereby producing a purified urine specimen containing the cfDNA fragments; (iii) concentrating the cfDNA fragments in the purified urine specimen by passing at least a portion of the purified urine specimen through a filter, wherein the concentration produces a filtrate and a residual urine specimen, and the residual urine specimen has an increased concentration of cfDNA fragments; and (iv) isolating the cfDNA fragments from the residual urine specimen.

[0321] Embodiment 45. The method of embodiment 44, further comprising (v) amplifying one or more of the isolated cfDNA fragments.

[0322] Embodiment 46 The method of embodiment 44 or 45, wherein treating the urine specimen to inhibit cell lysis comprises contacting the urine specimen with one or more preservation reagents.

[0323] Embodiment 47. The method of any one of embodiments 44 to 46, wherein treating comprises treatment with a nuclease inhibitor, a formaldehyde quencher, or both.

[0324] Embodiment 48. The method of any one of embodiments 44-47, wherein processing comprises contacting the urine specimen with a composition comprising: (a) imidazolidinyl urea, EDTA, glycine, or a combination thereof; or (b) sodium azide, EDTA, or a combination thereof.

[0325] Embodiment 49. The method of any one of embodiments 44-48, wherein separating comprises pelleting the cells in the processed urine specimen by centrifugation.

[0326] Embodiment 50. The method of any one of embodiments 44 to 49, wherein the filter is substantially impermeable to the passage of cell-free nucleic acids in the purified urine specimen and substantially permeable to salts.

[0327] Embodiment 51. The method of any one of embodiments 44 to 49, wherein the filter has a nominal molecular weight cutoff of 10 kD, 5 kD, 3 kD, or less.

[0328] Embodiment 52. The method of any one of embodiments 44 to 51, wherein the residual urine specimen has a concentration that is increased by at least 2-fold, at least 5-fold, at least 10-fold, or at least 15-fold compared to the purified urine specimen.

[0329] Embodiment 53. The method of any one of embodiments 44 to 52, wherein the residual urine specimen has a volume that is at least 50%, at least 75%, or at least 90% smaller than the volume of the processed urine specimen.

[0330] Embodiment 54. The method of embodiment 53, wherein the volume of the processed urine specimen is 5 mL, 10 mL, 15 mL, 20 mL, 30 mL, 40 mL, 50 mL, or more.

[0331] Embodiment 55. The method of any one of embodiments 44 to 54, further comprising freezing the residual urine specimen.

[0332] Embodiment 56. The method of any one of embodiments 44 to 55, wherein processing is completed within 120, 60, or 30 minutes after collection of the urine specimen; and optionally, separating and concentrating is completed within 7 days after collection.

[0333] Embodiment 57. The method of any one of embodiments 30 to 56, further comprising diagnosing cancer in the subject.

[0334] Embodiment 58. The method of embodiment 57, wherein the cancer is bladder cancer, prostate cancer, or renal cancer.

[0335] Embodiment 59 The method of embodiment 57 or 58, further comprising treating cancer in the subject.

[0336] Embodiment 60. The method of embodiment 59, wherein treating comprises surgical resection, radiation therapy, chemotherapy, and / or immunotherapy.

[0337] Embodiment 61. A method of treating cancer in a subject, comprising selecting a subject having or at high risk of developing cancer, and administering a treatment to the subject; (a) selecting includes identifying the subject as a source of a urinary cell-free DNA (cfDNA) specimen containing one or more target genomic regions that are differentially methylated above a threshold level for the presence of cancer; (b) the one or more target genomic regions comprise one or more target sequences of one or more genes selected from Table 1; (c) each target sequence is at least 25 nucleotides in length; (d) the cancer is bladder cancer, prostate cancer, or renal cancer; and (e) The method, wherein the treatment comprises surgical resection, radiation therapy, chemotherapy, immunotherapy, or any combination thereof.

[0338] Embodiment 62 The method of embodiment 61, wherein the threshold level for the presence of cancer is the level in a reference sample from a subject with cancer.

[0339] Embodiment 63. The method of embodiment 61 or 62, wherein the threshold level for the presence of cancer is determined by a classifier trained on sequencing reads of converted DNA from subjects with cancer.

[0340] Embodiment 64. The method of any one of embodiments 61-63, wherein the one or more target genomic regions comprise target sequences from at least 10 genes in Table 1.

[0341] Embodiment 65. The method of any one of embodiments 61 to 64, wherein the one or more target genomic regions comprise target sequences of genes selected from TWIST1, EOMES, HOXA9, POU4F2, and ZNF154.

[0342] Embodiment 66. The method according to any one of embodiments 61 to 64, (a) the cancer is bladder cancer and the one or more target genomic regions comprise one or more target sequences of one or more genes selected from Table 2 or Table 3; (b) the cancer is prostate cancer and the one or more target genomic regions comprise one or more target sequences of one or more genes selected from Table 4; or (c) the cancer is renal cancer, and the one or more target genomic regions comprise one or more target sequences of one or more genes selected from Table 5.

[0343] Embodiment 67. The method of any one of embodiments 61 to 66, wherein each target genomic region comprises at least five methylation sites.

[0344] Embodiment 68. A composition comprising a plurality of different bait oligonucleotides, (a) a bait oligonucleotide hybridizes to a converted DNA molecule derived from one or more target genomic regions; (b) the one or more target genomic regions comprise one or more target sequences of one or more genes selected from Table 1; (c) one or more target genomic regions exhibit differential methylation in cancer; and (d) A composition wherein each bait oligonucleotide comprises a sequence of at least 25 nucleotides in length that hybridizes to one of the target sequences.

[0345] Embodiment 69. The method of embodiment 68, wherein the one or more target genomic regions comprise target sequences from at least 10 genes selected from Table 1.

[0346] Embodiment 70. The method of embodiment 68 or embodiment 69, wherein the one or more target genomic regions comprise target sequences of genes selected from TWIST1, EOMES, HOXA9, POU4F2, and ZNF154.

[0347] Embodiment 71. The method of embodiment 68 or embodiment 69, wherein the one or more target genomic regions comprise a target sequence from: (a) a gene in Table 2 or Table 3; (b) a gene in Table 4; or (c) a gene in Table 5.

[0348] Embodiment 72. The method of any one of embodiments 68 to 71, wherein each target genomic region comprises at least five methylation sites.

[0349] Embodiment 73. The method of any one of embodiments 68 to 72, wherein the differential methylation comprises at least 80% of the CpG sites in the target genomic region being methylated or unmethylated. [Example]

[0350] The following examples are presented to provide those of ordinary skill in the art with a complete disclosure and description of how to make and use the present description, and are not intended to limit the scope of what the inventors regard as their description, nor are they intended to represent that the experiments below are all or the only experiments performed. Efforts have been made to ensure accuracy with respect to numbers used (e.g., amounts, temperatures, etc.), but some experimental error and deviation should be accounted for.

[0351] Example 1 - Urine specimen processing workflow In the Circulating Cell-Free Genome Atlas Study ("CCGA"; ClinicalTrial.gov identification number NCT02889978), a prospective, multicenter, case-control observational study with longitudinal follow-up, urinary tract cancers, such as prostate, bladder, and kidney cancers, had low detection sensitivity (Figure 3). This low detection rate may be due to the low tumor fraction of cfDNA in the blood of subjects with urothelial cancer. Analysis of cell-free DNA from urine may increase the detection sensitivity for urinary tract cancers. Therefore, we developed an improved method for preserving and extracting cfDNA from urine.

[0352] Figure 1 shows an exemplary urine specimen processing workflow. In this workflow, approximately 50 mL of urine is collected from a subject. A preservative is then added to the collected urine specimen. Streck Urine Preserve (Streck, Nebraska, USA) or an equivalent preservative containing at least 0.5% weight / volume (w / v) of a nuclease inhibitor, at least 0.2-4.0% w / v of a preservative, and at least 0.01% w / v of a formaldehyde quencher may be used as the preservative. Other preservatives that may be used include Urine Collection Medium and UAS (Novosantis, Belgium). Alternatively, urine specimens may be collected using a Urine Collection and Preservation Apparatus (Norgen Biotek Corp., Canada) or an equivalent cup containing 10-30% w / v of a nuclease inhibitor, such as EDTA, and 0.1-1.0% w / v of a bacteriostatic preservative, such as sodium azide.

[0353] After adding the preservative, the urine specimen is centrifuged at 4000 × g for 20 minutes to sediment and remove any cellular debris. The resulting supernatant is then concentrated approximately 15-fold. Concentration of the preserved urine specimen can be achieved using a diafiltration column, such as a filter unit with a 3 kDa cutoff regenerated cellulose membrane. For example, if the volume of the resulting supernatant is 15 mL or less, the specimen can be concentrated by spinning at 4000 × g for at least 30 minutes in an Amicon Ultra-15 filter unit (Thermo Fisher Scientific, Massachusetts, US). When working with supernatants in the 15-50 mL volume range, the specimen can alternatively be concentrated by spinning at 2500 × g for at least 40 minutes in a Centricon Plus-70 centrifugal filter unit (Thermo Fisher Scientific, Massachusetts, US) or until the specimen is concentrated to a volume of less than 4.2 mL. Concentrating the urine specimen to reduce its volume makes the specimen suitable for automated bead-based extraction methods (eg, MagMax extraction) and other available benchtop techniques.

[0354] After urine specimen concentration, the specimen may be used immediately or frozen at -80°C for later use or for batch processing of specimens. The concentrated urine specimen may then be subjected to cfDNA extraction and library preparation, which can then be sequenced for methylation analysis and urological cancer detection (Figures 2 and 4).

[0355] The workflow shown in Figure 1 was applied to urine specimens using Streck Urine Preserve. Adding the preservative within 30 minutes of urine specimen collection preserved nucleosomes and generated cfDNA fragments up to 7000 bp in size (Figure 5A). In contrast, delaying the addition of the preservative for more than one hour after collection resulted in the loss of the nucleosome peak at >700 bp, a significant decrease in yield, and a narrowing of the fragment length distribution, skewing it toward lower molecular weight fragments, indicating lysis of cells in the urine along with degradation and fragmentation of cfDNA (Figure 5B). It was also found that by concentrating the urine, specimens frozen for later processing exhibited reduced cryoprecipitate formation.

[0356] Example 2 - Analysis of tumor proportion in urine-derived cfDNA Cancer-specific methylation signatures were detected in urinary cfDNA from urine specimens processed as described in Example 1, including those corresponding to specimens from subjects with stage I high-grade non-muscle-invasive bladder cancer (Figures 6A-6B).

[0357] Further studies were conducted to evaluate the detection of cancer-specific methylation markers in urinary cfDNA compared with tumor fraction estimates in plasma. Urine and blood samples were collected from patients with bladder, renal, and prostate cancer, as well as from age- and sex-matched non-cancer patients. To generate biopsy-free estimates in urinary cfDNA, the plasma-based workflow was modified in the following steps (illustrated in Figure 2A): (1) an external reference dataset of non-cancer urinary cfDNA (N = approximately 200) was used instead of plasma for the non-cancer WGBS and TM data in the workflow; (2) the noise threshold and pseudocount were adjusted to account for the small reference dataset; and (3) WGBS data from healthy urinary tissue was used to further filter out noisy methylation mutations.

[0358] Sequencing libraries were then prepared from the resulting cfDNA for methylation analysis of a panel of urological cancer methylation markers. Tumor fraction estimation was also performed using the methylation marker panel, and specimens with estimated tumor fractions above a threshold were identified as having cancer detected based on urine-derived cfDNA. The estimated tumor fractions in urine-derived cfDNA specimens were then compared with those from the corresponding plasma-derived cfDNA fractions.

[0359] The scatter plots in Figures 7-9 show the distribution of tumor fraction by cancer type in the corresponding urine and plasma cfDNA (each point represents one patient). The filled area indicates whether the multi-cancer classifier (99% specificity) detected the plasma cfDNA for that patient. An increased tumor fraction was observed in the urine cfDNA compared to plasma for all bladder cancer patients analyzed, as well as for a subset of prostate cancer patients. Plasma tumor fraction estimates were consistent with classifier detection. For renal cancer patients, there was no increase in the urine signal over that of plasma, but a higher tumor fraction was observed in the urine cfDNA compared to non-cancer patients. With a smaller number of patients, an increased signal in plasma compared to urine cfDNA was observed.

[0360] The performance of bladder cancer detection based on a subset of genomic regions in the methylation marker panel was also evaluated. High sensitivity and specificity were observed when detecting bladder cancer status based on a subset of 15 methylation markers (Figure 10A). Moreover, high sensitivity and specificity were observed even when determining bladder cancer status based only on methylation markers within a single gene (TWIST1) (Figure 10B).

[0361] In addition, we evaluated the performance of renal cancer and prostate cancer detection based on a subset of genomic regions in the methylation marker panel. The results demonstrate the diagnostic potential of urine-derived cfDNA for both renal cancer (Figure 10C) and prostate cancer (Figure 10D). The genomic regions of the genes in Tables 2 and 3 were found to contain methylation markers for bladder cancer. The genomic regions of the genes in Table 4 were found to contain methylation markers for prostate cancer. The genomic regions of the genes in Table 5 were found to contain methylation markers for renal cancer. Table 1 presents a merger of the genes in Tables 2 through 5. In this example, the genomic region of a gene was considered to be the sequence from the (actual or putative) transcription start site to the transcription stop site, plus an additional 5,000 nucleotides from each of these ends. While target genomic regions were identified for urine cfDNA samples, such markers may be used to classify other cfDNA samples, such as cfDNA from other body fluids (e.g., blood, serum, or plasma).

[0362] In conclusion, determination of urological cancer status based on analysis of methylation markers and tumor fraction was as accurate or more accurate when analyzing urine-derived cfDNA than plasma-derived cfDNA. Concentration of urine specimens improved the processing and analysis of urine-derived cfDNA for urological cancer detection.

Claims

1. 1. A method for sequencing cell-free nucleic acid molecules of a subject, comprising: (a) treating the urine specimen to inhibit cell lysis; (b) separating cell-free nucleic acid molecules in said processed urine specimen from cells in said processed urine specimen, thereby producing a purified urine specimen comprising said cell-free nucleic acid molecules; (c) concentrating the cell-free nucleic acid molecules in the purified urine specimen by passing at least a portion of the purified urine specimen through a filter, wherein (i) the concentration produces a filtrate and a residual urine specimen, and (ii) the residual urine specimen is associated with an increased concentration of cell-free nucleic acid molecules; (d) isolating cell-free nucleic acid molecules from the residual urine specimen; and (e) sequencing the isolated cell-free nucleic acid molecule. A method comprising:

2. 10. The method of claim 1, wherein treating the urine specimen to inhibit cell lysis comprises contacting the urine specimen with one or more preservative reagents.

3. 10. The method of claim 1, wherein said treating comprises treatment with a nuclease inhibitor, a formaldehyde quencher, or both.

4. 10. The method of claim 1, wherein said treating comprises contacting said urine specimen with a composition comprising: (i) imidazolidinyl urea, EDTA, glycine, or a combination thereof; or (ii) sodium azide, EDTA, or a combination thereof.

5. 10. The method of claim 1, wherein said separating comprises pelleting cells in said processed urine specimen by centrifugation.

6. 10. The method of claim 1, wherein the filter is substantially impermeable to the passage of cell-free nucleic acids in the purified urine specimen and substantially permeable to salts.

7. 10. The method of claim 1, wherein the filter has a nominal molecular weight cutoff of 10 kD, 5 kD, 3 kD, or less.

8. 10. The method of claim 1, wherein the residual urine specimen has a concentration that is increased by at least 2-fold, at least 5-fold, at least 10-fold, or at least 15-fold compared to the purified urine specimen.

9. 10. The method of claim 1, wherein the residual urine specimen has a volume that is at least 50%, at least 75%, or at least 90% smaller than the volume of the treated urine specimen.

10. 10. The method of claim 9, wherein the volume of the processed urine specimen is 5 mL, 10 mL, 15 mL, 20 mL, 30 mL, 40 mL, 50 mL, or more.

11. 10. The method of claim 1, further comprising freezing the residual urine specimen.

12. 10. The method of claim 1, wherein said processing is completed within 120, 60, or 30 minutes after collection of said urine specimen; and optionally said separating and said concentrating are completed within 7 days after collection.

13. 10. The method of claim 1, further comprising amplifying one or more of the isolated cell-free nucleic acid molecules.

14. 10. The method of claim 1, further comprising capturing the isolated cell-free nucleic acid molecule, or an amplification product thereof, by hybridization to a bait oligonucleotide.

15. The method of claim 14, further comprising separating cell-free nucleic acid molecules bound to the bait from unbound cell-free nucleic acid molecules.

16. 16. The method of claim 15, wherein each bait oligonucleotide hybridizes to a target genomic region that is differentially methylated in cancer specimens compared to non-cancerous specimens.

17. 17. The method of claim 16, wherein the differential methylation comprises at least 80% of the CpG sites in the target genomic region being methylated or unmethylated.

18. 17. The method of claim 16, wherein the cancer is bladder cancer, prostate cancer, or renal cancer.

19. 15. The method of claim 14, wherein each bait oligonucleotide hybridizes to a target genomic region containing at least five methylation sites.

20. 15. The method of claim 14, wherein each bait oligonucleotide hybridizes to a target genomic region comprising a target sequence of a gene selected from Table 1, and the target sequence is at least 25, at least 35, or at least 45 nucleotides in length.

21. 21. The method of claim 20, wherein the target genomic region comprises a target sequence of a gene selected from TWIST1, EOMES, HOXA9, POU4F2, and ZNF154.

22. 21. The method of claim 20, wherein the bait oligonucleotides collectively hybridize to target sequences from at least 10 genes in Table 1.

23. 21. The method of claim 20, wherein the bait oligonucleotides jointly hybridize to target sequences from: (a) a gene in Table 2 or Table 3; (b) a gene in Table 4; or (c) a gene in Table 5.

24. 24. The method of any one of claims 1 to 23, wherein the cell-free nucleic acid molecule comprises cell-free DNA (cfDNA).

25. 25. The method of claim 24, further comprising deaminating the cfDNA isolated in step (d) to create a converted cfDNA molecule; optionally, the deaminating comprises treatment with cytosine deaminase or bisulfite.

26. 10. The method of claim 1, further comprising diagnosing cancer in the subject.

27. 27. The method of claim 26, wherein the cancer is bladder cancer, prostate cancer, or renal cancer.

28. 28. The method of claim 26 or 27, further comprising treating the cancer in the subject.

29. 29. The method of claim 28, wherein said treating comprises surgical resection, radiation therapy, chemotherapy, and / or immunotherapy.

30. 1. A method for detecting cancer cells in a subject, comprising: (a) capturing converted cell-free DNA (cfDNA) fragments, or amplification products thereof, from a urine specimen of the subject; (i) the bait oligonucleotide composition comprises a plurality of different bait oligonucleotides; (ii) each bait oligonucleotide of the plurality of different bait oligonucleotides hybridizes to a target sequence of a gene selected from Table 1, and the target sequence is at least 25 nucleotides in length; (b) separating the DNA bound to the bait from the unbound DNA; (c) sequencing the separated DNA to generate sequencing reads; and (d) detecting the cancer cells with a trained classifier. A method comprising:

31. 31. The method of claim 30, wherein the trained classifier detects a number of sequencing reads above a threshold for one or more of the target sequences identified as hypermethylated and / or hypomethylated in the cfDNA fragments.

32. 31. The method of claim 30, wherein the bait oligonucleotide is at least 45 nucleotides in length.

33. 31. The method of claim 30, wherein the trained classifier distinguishes subjects with cancer from subjects without cancer with a specified specificity.

34. 34. The method of claim 33, wherein the classifier is a mixture model classifier.

35. 34. The method of claim 33, wherein the defined specificity is 0.900 or greater.

36. 36. The method of claim 35, wherein applying the trained classifier further comprises a sensitivity of 30% or greater.

37. 31. The method of claim 30, wherein the bait oligonucleotides collectively hybridize to target sequences from at least 10 genes in Table 1.

38. 31. The method of claim 30, wherein at least one of the bait oligonucleotides hybridizes to a target sequence of a gene selected from TWIST1, EOMES, HOXA9, POU4F2, and ZNF154.

39. (a) the cancer cells are bladder cancer cells, and the bait oligonucleotides hybridize collectively to target sequences from the genes of Table 2 or Table 3; (b) the cancer cells are prostate cancer cells and the bait oligonucleotides collectively hybridize to target sequences from the genes in Table 4; or (c) the cancer cells are renal cancer cells, and the bait oligonucleotides collectively hybridize to target sequences from the genes in Table 5.

40. 31. The method of claim 30, wherein the converted cfDNA molecule comprises cfDNA that has been treated with cytosine deaminase or bisulfite.

41. 31. The method of claim 30, wherein each bait oligonucleotide is conjugated to a solid surface or to a non-nucleotide affinity moiety.

42. 31. The method of claim 30, wherein the differential methylation comprises hypermethylation in a cancer specimen compared to a non-cancerous specimen.

43. 31. The method of claim 30, wherein each target genomic region comprises at least five methylation sites.

44. 44. The method of any one of claims 30-43, further comprising obtaining the converted cfDNA fragments or amplification products thereof, and wherein said obtaining further comprises: (i) treating a urine specimen to inhibit cell lysis; (ii) separating the cfDNA fragments in the processed urine specimen from cells in the processed urine specimen, thereby producing a purified urine specimen comprising the cfDNA fragments; (iii) concentrating the cfDNA fragments in the purified urine specimen by passing at least a portion of the purified urine specimen through a filter, wherein said concentration produces a filtrate and a residual urine specimen, and the residual urine specimen has an increased concentration of cfDNA fragments; and (iv) isolating the cfDNA fragments from the residual urine specimen.

45. 45. The method of claim 44, further comprising (v) amplifying one or more of the isolated cfDNA fragments.

46. 45. The method of claim 44, wherein treating the urine specimen to inhibit cell lysis comprises contacting the urine specimen with one or more preservation reagents.

47. 45. The method of claim 44, wherein said treating comprises treatment with a nuclease inhibitor, a formaldehyde quencher, or both.

48. 45. The method of claim 44, wherein said treating comprises contacting said urine specimen with a composition comprising: (a) imidazolidinyl urea, EDTA, glycine, or a combination thereof; or (b) sodium azide, EDTA, or a combination thereof.

49. 45. The method of claim 44, wherein said separating comprises pelleting cells in said processed urine specimen by centrifugation.

50. 45. The method of claim 44, wherein the filter is substantially impermeable to the passage of cell-free nucleic acids in the purified urine specimen and substantially permeable to salts.

51. 45. The method of claim 44, wherein the filter has a nominal molecular weight cutoff of 10 kD, 5 kD, 3 kD, or less.

52. 45. The method of claim 44, wherein the residual urine specimen has a concentration that is increased by at least 2-fold, at least 5-fold, at least 10-fold, or at least 15-fold compared to the purified urine specimen.

53. 45. The method of claim 44, wherein the residual urine specimen has a volume that is at least 50%, at least 75%, or at least 90% smaller than the volume of the treated urine specimen.

54. 54. The method of claim 53, wherein the volume of the processed urine specimen is 5 mL, 10 mL, 15 mL, 20 mL, 30 mL, 40 mL, 50 mL, or more.

55. 45. The method of claim 44, further comprising freezing the residual urine specimen.

56. 45. The method of claim 44, wherein said processing is completed within 120, 60, or 30 minutes after collection of said urine specimen; and optionally said separating and said concentrating are completed within 7 days after collection.

57. 31. The method of claim 30, further comprising diagnosing cancer in the subject.

58. 58. The method of claim 57, wherein the cancer is bladder cancer, prostate cancer, or renal cancer.

59. 58. The method of claim 57, further comprising treating the cancer in the subject.

60. 60. The method of claim 59, wherein said treating comprises surgical resection, radiation therapy, chemotherapy, and / or immunotherapy.

61. 1. A method of treating cancer in a subject, comprising selecting a subject having or at high risk of developing cancer, and administering a treatment to said subject; (a) said selecting comprises identifying said subject as a source of a urinary cell-free DNA (cfDNA) specimen containing one or more target genomic regions that are differentially methylated above a threshold level for the presence of said cancer; (b) the one or more target genomic regions comprise one or more target sequences of one or more genes selected from Table 1; (c) each target sequence is at least 25 nucleotides in length; (d) the cancer is bladder cancer, prostate cancer, or renal cancer; and (e) The method, wherein said treatment comprises surgical resection, radiation therapy, chemotherapy, immunotherapy, or any combination thereof.

62. 62. The method of claim 61, wherein the threshold level for the presence of cancer is the level in a reference specimen from a subject with cancer.

63. 62. The method of claim 61 , wherein the threshold level for the presence of cancer is determined by a classifier trained on sequencing reads of converted DNA from subjects with the cancer.

64. 62. The method of Claim 61, wherein said one or more target genomic regions comprise target sequences from at least 10 genes of Table 1.

65. 62. The method of claim 61 , wherein the one or more target genomic regions comprise a target sequence of a gene selected from TWIST1, EOMES, HOXA9, POU4F2, and ZNF154.

66. (a) the cancer is bladder cancer, and the one or more target genomic regions comprise one or more target sequences of one or more genes selected from Table 2 or Table 3; (b) the cancer is prostate cancer and the one or more target genomic regions comprise one or more target sequences of one or more genes selected from Table 4; or (c) the cancer is renal cancer and the one or more target genomic regions comprise one or more target sequences of one or more genes selected from Table 5.

67. 67. The method of any one of claims 61 to 66, wherein each target genomic region comprises at least five methylation sites.

68. 1. A composition comprising a plurality of different bait oligonucleotides, (a) the bait oligonucleotides hybridize to converted DNA molecules derived from one or more target genomic regions; (b) the one or more target genomic regions comprise one or more target sequences of one or more genes selected from Table 1; (c) the one or more target genomic regions exhibit differential methylation in cancer; and (d) a composition wherein each bait oligonucleotide comprises a sequence of at least 25 nucleotides in length that hybridizes to one of said target sequences.

69. 69. The composition of Claim 68, wherein said one or more target genomic regions comprise target sequences from at least 10 genes selected from Table 1.

70. 69. The composition of claim 68, wherein the one or more target genomic regions comprise a target sequence of a gene selected from TWIST1, EOMES, HOXA9, POU4F2, and ZNF154.

71. 69. The composition of Claim 68, wherein said one or more target genomic regions comprise a target sequence from: (a) a gene in Table 2 or Table 3; (b) a gene in Table 4; or (c) a gene in Table 5.

72. 69. The composition of claim 68, wherein each target genomic region comprises at least five methylation sites.

73. 73. The composition of any one of claims 68 to 72, wherein said differential methylation comprises at least 80% of CpG sites in said target genomic region being methylated or unmethylated.