Method for analyzing free nucleic acid in urine

By targeting the detection of cancer-specific methylation patterns in urine, bait oligonucleotides and filtration concentration technology, the problem of low-abundance cfDNA detection efficiency in urine is solved, and efficient diagnosis and treatment support for early cancers is achieved.

CN120380342APending Publication Date: 2025-07-25GRAIL INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202480005333.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-20
Filing Date
2024-01-19
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art is difficult to effectively analyze free nucleic acids with low abundance in urine, especially because whole genome sequencing methods are unstable in cancer, resulting in ineffective detection.

Method used

By designing methods to target detection of cancer-specific methylation patterns, cfDNA in urine is enriched using bait oligonucleotides, combined with filtration and concentration techniques, improving sequencing depth, and identifying cancer-related methylation patterns through sequencing and classifiers.

Benefits of technology

It realizes efficient detection of cancer-specific methylation patterns in urine, improves detection sensitivity and specificity, and supports the diagnosis and treatment decisions of early cancers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120380342A_ABST
    Figure CN120380342A_ABST
Patent Text Reader

Abstract

In various aspects, the disclosure provides methods, compositions, reaction mixtures, kits, and systems for analyzing free nucleic acid molecules (e.g., cfRNA and / or cfDNA) from a urine sample. In some embodiments, the analysis is the analysis of a methylation pattern in a target genomic region in a cfDNA fragment in a urine sample. In some embodiments, the composition includes a plurality of different bait oligonucleotides. Methods for detecting cancers of various cancer types are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims priority to U.S. Provisional Application Serial No. 63 / 480,934, filed on January 20, 2023, the disclosure of which is hereby incorporated by reference in its entirety. Technical Field

[0003] This application pertains to the technical field of analyzing free nucleic acids in urine. Background Art

[0004] The analysis of nucleic acids, such as cell-free nucleic acids (e.g., cell-free DNA (cfDNA) and cell-free RNA (cfRNA)), using next-generation sequencing (NGS) is recognized as a valuable method for characterizing various sample types. For example, such analyses can be used as diagnostic tools for detecting and diagnosing cancer. These analytes can also be valuable in improving our fundamental understanding of basic biology.

[0005] DNA methylation plays an important role in regulating gene expression. Aberrant DNA methylation has been associated with many disease processes, including cancer. DNA methylation profiling using methylation sequencing (e.g., whole-genome bisulfite sequencing (WGBS)) is increasingly recognized as an important diagnostic tool for detecting, diagnosing, and / or monitoring cancer. For example, specific patterns of differentially methylated regions can be used as molecular markers for various diseases.

[0006] However, WGBS is not ideally suited for clinical assays. The reason is that the vast majority of the genome is either not differentially methylated in cancer or has too low a local CpG density to provide a robust signal. Only a few percent of the genome may be useful for classification. These problems are exacerbated when dealing with sample types with low enrichment of cell-free nucleic acids (such as urine). Summary of the Invention

[0007] At least for the above reasons, there is still a need for cost-effective methods and compositions for analyzing free nucleic acid molecules in urine. Various aspects of this disclosure address this need and also provide other advantages.

[0008] Early detection of cancer in a subject is important because it allows for earlier treatment and thus a greater chance of survival. Detection of cancer-specific methylation patterns using cell-free DNA (cfDNA) fragments can enable early detection of cancer by providing a cost-effective and non-invasive method for obtaining information related to the presence or absence of cancer, the tissue of origin of the cancer, or the cancer type. By using a set of target genomic regions rather than sequencing all nucleic acids in a test sample (also known as "whole-genome sequencing"), this method can increase the sequencing depth of the target regions. Including methylation markers for several different types of cancer allows for more efficient use of samples and reagents compared to performing multiple separate assays for different types of cancer. However, restricting the coverage of the total genomic sequence can be advantageous for capture, sequencing, and / or computational efficiency.

[0009] Accordingly, the present specification provides a set of cancer assays (alternatively referred to as a "bait set") for detecting cancer-specific methylation patterns in target genomic regions, and methods for using these sets of cancer assays to detect cancer, cancer type, and / or tissue of origin of the cancer (TOO). The methods described herein further include methods for designing probes to effectively enrich cfDNA corresponding to or derived from selected genomic regions without pulling down excessive amounts of unwanted DNA. In particular, methods for analyzing cfDNA and other cell-free nucleic acids in urine samples are also provided.

[0010] In one aspect, the present disclosure provides methods for sequencing cell-free nucleic acid molecules from a subject. In some embodiments, the method includes: (a) treating a urine sample to inhibit cell lysis; (b) separating the cell-free nucleic acid molecules in the treated urine sample from the cells in the treated urine sample, thereby producing a purified urine sample containing the cell-free nucleic acid molecules; (c) concentrating the cell-free nucleic acid molecules in the purified urine sample by passing at least a portion of the purified urine sample through a filter, wherein (i) the concentration produces a filtrate and a retained urine sample, and (ii) the retained urine sample contains cell-free nucleic acid molecules at an increased concentration; (d) isolating the cell-free nucleic acid molecules from the retained urine sample; and (e) sequencing the isolated cell-free nucleic acid molecules.

[0011] In some embodiments, processing a urine sample to inhibit cell lysis includes contacting the urine sample with one or more preservative reagents. In some embodiments, the processing includes treating with a nuclease inhibitor, a formaldehyde quencher, or both. In some embodiments, the processing includes contacting the urine sample with a composition comprising: (i) imidazolidinyl urea, EDTA, glycine, or a combination thereof; or (ii) sodium azide, EDTA, or a combination thereof. In some embodiments, separating includes centrifuging to pellet cells in the processed urine sample. In some embodiments, the filter is substantially impermeable to the passage of free nucleic acids and substantially permeable to salts in the purified urine sample. In some embodiments, the filter has a nominal molecular weight cut-off of 10 kD, 5 kD, 3 kD, or lower. In some embodiments, the retained urine sample has a concentration that is at least 2-fold, at least 5-fold, at least 10-fold, or at least 15-fold greater than that of the purified urine sample. In some embodiments, the retained urine sample has a volume that is at least 50%, at least 75%, or at least 90% lower than the volume of the processed urine sample. In some embodiments, the volume of the processed urine sample is 20 mL, 30 mL, 40 mL, 50 mL, or more. In some embodiments, the method further includes freezing the retained urine sample. In some embodiments, the processing is completed within 120, 60, or 30 minutes after urine sample collection; and optionally wherein the separating and the concentrating are completed within 7 days after collection.

[0012] In some embodiments, the method further includes amplifying one or more of the isolated free nucleic acid molecules. In some embodiments, the method further includes capturing the isolated free nucleic acid molecules or their amplification products by hybridization with a bait oligonucleotide. In some embodiments, the method further includes separating the bait-bound free nucleic acid molecules from the unbound free nucleic acid molecules.

[0013] In some embodiments, each bait oligonucleotide hybridizes to a target genomic region that is differentially methylated in a cancer sample relative to a non-cancer sample. In some embodiments, the differential methylation comprises at least 80% of the CpG sites in the target genomic region that are methylated or unmethylated. In some embodiments, the cancer is bladder cancer, prostate cancer, or kidney cancer. In some embodiments, each bait oligonucleotide hybridizes to a target genomic region that comprises at least five methylation sites. In some embodiments, each bait oligonucleotide hybridizes to a target genomic region that comprises a target sequence of a gene selected from Table 1, and wherein the length of the target sequence is at least 25, at least 35, or at least 45 nucleotides. In some embodiments, the target genomic region comprises a target sequence of a gene selected from the following: TWIST1, EOMES, HOXA9, POU4F2, and ZNF154. In some embodiments, the bait oligonucleotides together hybridize to target sequences of at least 10 genes from Table 1. In some embodiments, the bait oligonucleotides together hybridize to target sequences from: (a) the genes in Table 2 or Table 3; (b) the genes in Table 4; or (c) the genes in Table 5.

[0014] In some embodiments, the cell-free nucleic acid molecule comprises cell-free DNA (cfDNA). In some embodiments, the method further comprises deaminating the cfDNA isolated in step (d) to produce a converted cfDNA molecule; optionally wherein the deamination comprises treatment with a cytidine deaminase or bisulfite.

[0015] In some embodiments, the method further comprises diagnosing cancer in a subject. In some embodiments, the cancer is bladder cancer, prostate cancer, or kidney cancer. In some embodiments, the method further comprises treating cancer in a subject. In some embodiments, the treatment comprises surgical resection, radiotherapy, chemotherapy, and / or immunotherapy.

[0016] In one aspect, provided herein is a method for detecting cancer cells in a subject. In some embodiments, the method comprises (a) capturing converted cell-free DNA (cfDNA) fragments or amplification products thereof from a urine sample of the subject, wherein: (i) the bait oligonucleotide composition comprises a plurality of different bait oligonucleotides; (ii) each bait oligonucleotide of the plurality of different bait oligonucleotides hybridizes to a target sequence of a gene selected from Table 1, wherein the length of the target sequence is at least 25 nucleotides; (b) separating the bait-bound DNA from the unbound DNA; (c) sequencing the separated DNA to produce sequencing reads; and (d) detecting the cancer cells with a trained classifier.

[0017] In some embodiments, the length of the decoy oligonucleotide is at least 45 nucleotides. In some embodiments, for one or more of the target sequences identified as hypermethylated and / or hypomethylated in the cfDNA fragments, the trained classifier detects a number of sequencing reads that exceeds a threshold. In some embodiments, the trained classifier differentiates between subjects with cancer and subjects without cancer with a defined specificity. In some embodiments, the classifier is a mixture model classifier. In some embodiments, the defined specificity is 0.900 or higher. In some embodiments, the application of the trained classifier further comprises a sensitivity of 30% or higher. In some embodiments, the decoy oligonucleotides hybridize together with target sequences from at least 10 genes in Table 1. In some embodiments, at least one of the decoy oligonucleotides hybridizes with a target sequence of a gene selected from: TWIST1, EOMES, HOXA9, POU4F2, and ZNF154. In some embodiments, (a) the cancer cells are bladder cancer cells and the decoy oligonucleotides hybridize together with target sequences from the genes in Table 2 or Table 3; (b) the cancer cells are prostate cancer cells and the decoy oligonucleotides hybridize together with target sequences from the genes in Table 4; or (c) the cancer cells are kidney cancer cells and the decoy oligonucleotides hybridize together with target sequences from the genes in Table 5.

[0018] In some embodiments, the transformed cfDNA molecule comprises cfDNA treated with cytidine deaminase or bisulfite. In some embodiments, each decoy oligonucleotide is conjugated to a solid surface or a non-nucleotide affinity moiety. In some embodiments, the differential methylation comprises hypermethylation in the cancer sample relative to the non-cancer sample. In some embodiments, each target genomic region comprises at least five methylation sites.

[0019] In some embodiments, the method further comprises obtaining a transformed cfDNA fragment or an amplification product thereof, wherein the obtaining further comprises: (i) treating a urine sample to inhibit cell lysis; (ii) separating the cfDNA fragments in the treated urine sample from the cells in the treated urine sample, thereby producing a purified urine sample comprising the cfDNA fragments; (iii) concentrating the cfDNA fragments in the purified urine sample by passing at least a portion of the purified urine sample through a filter, wherein the concentration produces a filtrate and a retained urine sample, and wherein the retained urine sample comprises an increased concentration of cfDNA fragments; and (iv) isolating the cfDNA fragments from the retained urine sample. In some embodiments, the method further comprises (v) amplifying one or more of the isolated cfDNA fragments.

[0020] In some embodiments, processing a urine sample to inhibit cell lysis includes contacting the urine sample with one or more preservative reagents. In some embodiments, the processing includes treating with a nuclease inhibitor, a formaldehyde quencher, or both. In some embodiments, the processing includes contacting the urine sample with a composition comprising: (a) imidazolidinyl urea, EDTA, glycine, or a combination thereof; or (b) sodium azide, EDTA, or a combination thereof. In some embodiments, separating includes centrifuging to precipitate cells in the processed urine sample. In some embodiments, the filter is substantially impermeable to the passage of free nucleic acids and substantially permeable to salts in the purified urine sample. In some embodiments, the filter has a nominal molecular weight cut-off of 10 kD, 5 kD, 3 kD, or lower. In some embodiments, the retained urine sample has a concentration that is at least 2-fold, at least 5-fold, at least 10-fold, or at least 15-fold higher compared to the purified urine sample. In some embodiments, the retained urine sample has a volume that is at least 50%, at least 75%, or at least 90% lower compared to the volume of the processed urine sample. In some embodiments, the volume of the processed urine sample is 20 mL, 30 mL, 40 mL, 50 mL, or more. In some embodiments, the method further includes freezing the retained urine sample. In some embodiments, the processing is completed within 120, 60, or 30 minutes after urine sample collection; and optionally wherein the separating and the concentrating are completed within 7 days after collection.

[0021] In some embodiments, the method further includes diagnosing cancer in a subject. In some embodiments, the cancer is bladder cancer, prostate cancer, or kidney cancer. In some embodiments, the method further includes treating cancer in a subject. In some embodiments, the treatment includes surgical resection, radiotherapy, chemotherapy, and / or immunotherapy.

[0022] In one aspect, provided herein are methods of treating cancer in a subject. In some embodiments, the method includes selecting a subject having cancer or an increased risk of developing cancer and administering a treatment to the subject, wherein: (a) the selecting includes identifying the subject as the source of a urine cell-free DNA (cfDNA) sample that contains one or more differentially methylated target genomic regions above a threshold level indicative of the presence of the cancer; (b) the one or more target genomic regions include one or more target sequences of one or more genes selected from Table 1; (c) the length of each target sequence is at least 25 nucleotides; (d) the cancer is bladder cancer, prostate cancer, or kidney cancer; and (e) the treatment includes surgical resection, radiotherapy, chemotherapy, immunotherapy, or any combination thereof.

[0023] In some embodiments, the threshold level of cancer presence is the level of a reference sample from a subject having the cancer. In some embodiments, the threshold level of cancer presence is determined by a classifier trained on sequencing reads of converted DNA from a subject having the cancer. In some embodiments, one or more target genomic regions comprise target sequences of at least 10 genes from Table 1. In some embodiments, one or more of the target genomic regions comprise target sequences of genes selected from: TWIST1, EOMES, HOXA9, POU4F2, and ZNF154. In some embodiments, (a) the cancer is bladder cancer and one or more of the target genomic regions comprise one or more target sequences of one or more genes selected from Table 2 or Table 3; (b) the cancer is prostate cancer and the one or more target genomic regions comprise one or more target sequences of one or more genes selected from Table 4; or (c) the cancer is kidney cancer and the one or more target genomic regions comprise one or more target sequences of one or more genes selected from Table 5. In some embodiments, each target genomic region comprises at least five methylation sites.

[0024] In one aspect, provided herein are compositions comprising a plurality of different bait oligonucleotides. In some embodiments, (a) the bait oligonucleotides hybridize to converted DNA molecules derived from one or more target genomic regions; (b) the one or more target genomic regions comprise one or more target sequences of one or more genes selected from Table 1; (c) the one or more target genomic regions are differentially methylated in cancer; and (d) each bait oligonucleotide comprises a sequence of at least 25 nucleotides in length that hybridizes to one of the target sequences. In some embodiments, one or more target genomic regions comprise target sequences of at least 10 genes selected from Table 1. In some embodiments, one or more of the target genomic regions comprise target sequences of genes selected from: TWIST1, EOMES, HOXA9, POU4F2, and ZNF154. In some embodiments, one or more target genomic regions comprise target sequences from: (a) genes in Table 2 or Table 3; (b) genes in Table 4; or (c) genes in Table 5. In some embodiments, each target genomic region comprises at least five methylation sites. In some embodiments, the differential methylation comprises at least 80% of the CpG sites in the target genomic regions that are methylated or unmethylated.

[0025] Incorporation by reference

[0026] All publications, patents, and patent applications mentioned in this specification are hereby incorporated by reference in their entirety to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The novel features of this disclosure are set forth in the appended claims. A better understanding of the features and advantages of this disclosure will be obtained by reference to the following detailed description, which sets forth illustrative embodiments that make use of the principles of this disclosure, and the accompanying drawings:

[0028] Figure 1 A workflow for urine sample processing is shown according to one embodiment.

[0029] Figure 2A An exemplary process for estimating tumor fraction without biopsy based on urine cell-free DNA is shown. Figure 2B An exemplary methylation variation at CpG 1528723-1528728 is shown. This methylation variation is shown as a continuous stretch of CpGs and methylation states that distinguish cancer-derived cfDNA from non-cancer cfDNA.

[0030] Figure 3 The data provided show that in the Circulating Cell-free Genome Atlas (CCGA) patient cohort, urinary tract cancers (bladder cancer or urothelial carcinoma, kidney cancer, and prostate cancer) have lower detection sensitivities. Each pair of columns represents results for early (left) and late (right) stages.

[0031] Figure 4 A process for determining methylation variation based on urine cfDNA is shown according to one embodiment.

[0032] Figure 5A The data provided show the cfDNA fragment size distribution and yield when processing by immediately adding a preservative to the sample after collection. Figure 5B The data provided show the effect of delaying the addition of a preservative to the urine sample by one hour on the cfDNA fragment size distribution and yield.

[0033] Figures 6A - 6B Data are provided on the cancer-specific methylation signatures of four genes detected in preoperative and postoperative urine cfDNA from subjects with stage I, high-grade non-muscle-invasive bladder cancer.

[0034] Figure 7 The data provided compare the estimated tumor fractions in urine and plasma samples from subjects with (plotted points with annotations) or without (plotted points without annotations) bladder cancer at any stage.

[0035] Figure 8The data provided compares the estimated tumor fractions in urine and plasma samples from subjects with or without stage III or IV prostate cancer. Data points from “*” and above correspond to results from samples from only cancer subjects. Data points below “*” correspond to results from samples from both cancer and non-cancer subjects.

[0036] Figure 9 The data provided compares the estimated tumor fractions in urine and plasma samples from subjects with (data points with annotation) or without (data points without annotation) kidney cancer.

[0037] Figure 10A Data on the classification performance of urine cfDNA from subjects with bladder cancer based on a subset of 15 genomic regions are provided, showing an area under the curve (AUC) of 0.99. Two of the genomic regions are represented by one of two genes within the region. Figure 10B Data on the classification performance of urine cfDNA from subjects with bladder cancer based on genomic regions within a single gene (TWIST1) are provided, showing an AUC of 0.86. The shaded area represents the 95% confidence interval. Figure 10C Data on the classification performance of urine cfDNA from subjects with kidney cancer based on a subset of genomic regions are provided, showing an AUC of 0.52. The shaded area represents the 95% confidence interval. Figure 10D Data on the classification performance of urine cfDNA from subjects with prostate cancer based on a subset of genomic regions are provided, showing an AUC of 0.82. The shaded area represents the 95% confidence interval.

[0038] Figure 11A A 2x tiling probe design is shown according to an embodiment, where three probes target small target regions, and each base in the target region (within the dashed rectangular box) is covered by at least two probes. Figure 11B A 2x tiling probe design is shown according to an embodiment, where more than three probes target large target regions, and each base in the target region (within the dashed rectangular box) is covered by at least two probes. Figure 11C A probe design targeting hypomethylated and / or hypermethylated fragments in genomic regions is shown according to an embodiment. Figure 11D A genomic portion containing three target genomic regions is shown. According to an embodiment, the probes are designed to hybridize to each target and its adjacent sequences in a 2x tiling configuration. Figure 11E A single pair of probes hybridizing within the same target genomic region is shown, each probe containing overlapping and non-overlapping sequences. The overlapping sequences are complementary to the same target sequence. The non-overlapping sequences are complementary to different sequences within the target genomic region, each sequence located at a different end of the sequence complementary to the overlapping sequence, as shown.

[0039] Figure 12A is a flowchart depicting the process of creating a control group data structure according to an embodiment description. Figure 12B is according to an embodiment description for verifying Figure 12A additional steps of the control group data structure.

[0040] Figure 13 is a flowchart depicting the process of selecting genomic regions for designing a set of cancer assay probes according to an embodiment.

[0041] Figure 14 is an illustration of the p-value score calculation for an example according to an embodiment.

[0042] Figure 15A is a flowchart depicting the process of training a classifier based on hypomethylated and hypermethylated fragments indicative of cancer according to an embodiment. Figure 15B is a flowchart depicting the process of identifying fragments indicative of cancer determined by a probability model according to an embodiment.

[0043] Figure 16A is a flowchart depicting the process of sequencing cell-free (cf) DNA fragments according to an embodiment. Figure 16B is an illustration of the process of sequencing cell-free (cf) DNA fragments to obtain a methylation status vector according to an embodiment.

[0044] Figure 17A Displays a flowchart of an apparatus for sequencing a nucleic acid sample according to one embodiment. Figure 17B Displays an analysis system for analyzing the methylation status of cfDNA according to one embodiment. Detailed Description

[0045] Before describing the present invention in more detail, it should be understood that the present invention is not limited to the specific embodiments described, and of course, changes can be made to the embodiments themselves. It should also be understood that the terms used herein are for the purpose of describing specific embodiments only and are not intended to be limiting, as the scope of the present invention will be limited only by the appended claims.

[0046] Unless otherwise defined herein, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. The following definitions are provided to facilitate understanding of certain terms commonly used herein.

[0047] Where a numerical range is provided, it is understood that each intermediate value (to one tenth of the unit of the lower limit) between the upper and lower limits of the range, and any other stated value or intermediate value within the stated range, is covered by the present invention, unless the context clearly dictates otherwise. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges covered by the present invention, subject to any specifically excluded limitations within the stated range.

[0048] As used herein, the term "about" means a range of values including the specified value that would be reasonably considered by a person of ordinary skill in the art to be similar to the specified value. In embodiments, about means within the standard deviation of measurements generally accepted in the art. In embodiments, about means a range extending to + / −10% of the specified value. In embodiments, about includes the specified value.

[0049] As used herein, the term "methylation" refers to the process of adding a methyl group to a DNA molecule. For example, a hydrogen atom on the pyrimidine ring of a cytosine base can be converted to a methyl group to form 5-methylcytosine. The term also refers to the process of adding a hydroxymethyl group to a DNA molecule, for example, by oxidizing the methyl group on the pyrimidine ring of a cytosine base. Methylation and hydroxymethylation tend to occur at dinucleotides consisting of cytosine and guanine, which are referred to herein as "CpG sites". The principles described herein also apply to the detection of methylation in non-CpG contexts, including methylation of non-cytosines. In such embodiments, the laboratory wet assays for detecting methylation may differ from any of the assays described herein. Further, a methylation state vector may contain elements that are typically vectors of sites where methylation has or has not occurred (even if those sites are not specifically CpG sites).

[0050] The term "methylation" can also refer to the methylation state of a CpG site. A CpG site having a 5-methylcytosine moiety is methylated. A CpG site having a hydrogen atom on the pyrimidine ring of a cytosine base is not methylated.

[0051] As used herein, the term "methylation site" refers to a region of a DNA molecule to which a methyl group can be added. "CpG" sites are the most common methylation sites, but methylation sites are not limited to CpG sites. For example, DNA methylation can occur in cytosines in CHG and CHH, where H is adenine, cytosine, or thymine. The methylation of cytosine in the form of 5-hydroxymethylcytosine can also be evaluated using the methods and procedures disclosed herein (see, for example, US20110236894 A1 and US 20110301045A1, which are incorporated herein by reference) and its characteristics.

[0052] As used herein, the term "CpG site" refers to a region of a DNA molecule in which, in the linear sequence of bases along its 5' to 3' direction, a cytosine nucleotide is followed by a guanine nucleotide. "CpG" is shorthand for 5'-C-phosphate-G-3', i.e., cytosine and guanine are separated by only one phosphate group. The cytosine in a CpG dinucleotide can be methylated to form 5-methylcytosine.

[0053] In some embodiments, the oligonucleotide probes described herein comprise one or more CpG detection sites. As used herein, the term "CpG detection site" refers to a region in the probe that is configured to hybridize to a CpG site of a target DNA molecule. The CpG site on the target DNA molecule can comprise cytosine and guanine separated by one phosphate group, where the cytosine is methylated or unmethylated. The CpG site on the target DNA molecule can comprise uracil and guanine separated by one phosphate group, where the uracil is generated by conversion of unmethylated cytosine.

[0054] The term "UpG" is shorthand for 5'-U-phosphate-G-3', i.e., uracil and guanine are separated by only one phosphate group. UpG can be generated by bisulfite treatment of DNA, which converts unmethylated cytosine to uracil. Cytosine can be converted to uracil by other methods (e.g., chemical modification, synthesis, or enzymatic conversion).

[0055] As used herein, the terms "hypomethylated" or "hypermethylated" refer to the methylation status of a DNA molecule containing multiple CpG sites (e.g., more than 2, 3, 4, 5, 6, 7, 8, 9, 10, etc.), wherein a high percentage of the CpG sites (e.g., more than 80%, 85%, 90% or 95%, or any other percentage in the range of 50%-100%, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97.5% or more, 98% or more, 99% or more, 99.9% or more, or any other numerical percentage in the range of 50%-100% or more, where the provided ranges include the range boundaries of 50% and 100%) are unmethylated or methylated, respectively. For example, a "hypomethylated" nucleic acid (e.g., cfDNA) fragment can be a fragment having a certain number (e.g., 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 9 or more, 10 or more) of CpG sites, wherein a certain proportion (e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, or 95% or more, or 97.5% or more, 98% or more, 99% or more, 99.9% or more) of the CpG sites are unmethylated. Similarly, a "hypermethylated" nucleic acid, (e.g., cfDNA) fragment can be a fragment having a certain number (e.g., 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 9 or more, 10 or more) of CpG sites, wherein a certain proportion (e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, or 95% or more, or 97.5% or more, 98% or more, 99% or more, 99.9% or more) of the CpG sites are methylated. In some embodiments, the hypomethylated DNA molecule contains multiple CpG sites, at least 80% of which are unmethylated. In some embodiments, the hypermethylated DNA molecule contains multiple CpG sites, at least 80% of which are methylated.

[0056] As used herein, the term "methylation state vector" or "methylation status vector" refers to a vector containing multiple elements, where each element indicates the methylation status of a methylation site in a DNA molecule containing multiple methylation sites, and the multiple methylation sites are arranged in the order in which they occur in the DNA molecule from 5' to 3'. For example, <Mx, Mx+1, Mx+2>, <Mx, Mx+1, Ux+2>,..., <Ux, Ux+1, Ux+2> can be methylation vectors of a DNA molecule containing three methylation sites, where M represents a methylated methylation site and U represents an unmethylated methylation site.

[0057] As used herein, the terms "abnormal methylation pattern" and "anomalous methylation pattern" refer to the methylation pattern of a nucleic acid (e.g., DNA such as cfDNA), molecule, or methylation status vector that is found and / or expected to be found in a sample at a lower frequency than in a healthy (e.g., non-cancer) sample. In various embodiments, such methylation patterns are found and / or expected to be found in a sample at a frequency lower than a threshold in a healthy (e.g., non-cancer) sample. Thus, for example, as used herein, the terms "abnormally methylated" and "anomalously methylated" describe a nucleic acid (e.g., DNA such as cfDNA), molecule, or methylation status vector that exhibits an abnormal methylation pattern. The aspects of differential methylation according to the present disclosure may include aspects of abnormal methylation in some versions. Additionally, when referring to the health status of a subject from whom a sample is derived, whether an aspect is differentially methylated can be used as an indicator to determine health (e.g., non-cancer) versus disease (e.g., cancer). In one embodiment provided herein, a particular methylation status vector is found and / or expected to be found in a healthy control group including healthy individuals, represented by a p-value. In various aspects, a low p-value score corresponds to a methylation status vector that is relatively unexpected compared to other methylation status vectors in samples from healthy individuals (e.g., individuals in a healthy control group). In some versions, a high p-value score corresponds to a methylation status vector that is relatively more expected compared to other methylation status vectors found in samples from healthy individuals (e.g., individuals in a healthy control group). In various embodiments, a methylation status vector having an abnormal / anomalous methylation pattern is a methylation status vector having a p-value equal to and / or below a threshold (e.g., 0.1, 0.01, 0.001, 0.0001, etc.), which threshold is, for example, a threshold corresponding to a healthy (e.g., non-cancer) sample. In various embodiments, the method includes associating a methylation status vector from a sample having a p-value equal to and / or below a threshold (e.g., 0.1 or less, 0.01 or less, 0.001 or less, 0.0001 or less, etc.) with the determination that the sample is not a healthy sample (e.g., a sample from a subject with cancer). In various embodiments, the threshold is applied as a filter because the application of a smaller threshold (e.g., 0.001, 0.0001, etc.) is associated with a higher expectation that the methylation status vector is from an unhealthy sample (e.g., a sample from an individual with cancer). Various methods can be used to calculate the p-value or expectation of a methylation pattern or methylation status vector. The exemplary method provided herein involves using Markov chain probabilities, which assume that the methylation status of a CpG site depends on the methylation status of neighboring CpG sites. An alternative method provided herein calculates the expectation of observing a particular methylation status vector in a healthy individual by utilizing a mixture model including multiple mixture components, each of which is an independent site model, where it is assumed that the methylation status at each CpG site is independent of the methylation status at other CpG sites.In some versions, the methods of the present invention include determining whether a nucleic acid (e.g., DNA), molecule, or methylation status vector is abnormally methylated. In various embodiments of these methods, the generated p-value (e.g., by an analysis system) is compared to a threshold to identify vectors (e.g., nucleic acids such as cfDNA fragments) that are abnormally methylated relative to a control group (e.g., a group associated with one or more healthy (e.g., non-cancer) samples). Additionally, abnormal methylation (e.g., cfDNA methylation) can be hypermethylation and / or hypomethylation, both of which can indicate a non-healthy (e.g., cancer) state. Thus, the methods include determining a healthy or diseased (e.g., non-cancer or cancer) state at least in part based on the p-value (e.g., a relatively low p-value, e.g., a p-value below a threshold), where the p-value indicates abnormal methylation in various aspects, such as hypermethylation and / or hypomethylation. A low p-value (e.g., a p-value equal to or below a threshold (e.g., 0.1, 0.01, 0.001, 0.0001, etc.)) can indicate abnormal methylation in a sample, such as hypermethylation and / or hypomethylation. In various embodiments, the methods include determining the healthy or diseased (e.g., non-cancer or cancer) state of a sample based on nucleic acids (e.g., nucleic acid fragments) or methylation vectors from a sample having a low p-value (e.g., equal to or less than 0.1, 0.01, or 0.001) and that is both hypermethylated and hypomethylated or hypermethylated or hypomethylated. In various aspects, the methods include determining the healthy or diseased (e.g., non-cancer or cancer) state of a sample at least in part based on whether the nucleic acids (e.g., nucleic acid fragments) or methylation vectors from the sample are both hypermethylated and hypomethylated. In some variants, determining whether a vector (e.g., a sample fragment) is aberrantly methylated based on a generated p-value score includes determining whether the generated score of the vector is below a threshold score, where the threshold score is the confidence that the vector is aberrantly methylated.

[0058] As used herein, the term "cancerous sample" refers to a sample containing genomic DNA from an individual diagnosed with cancer. The genomic DNA can be (but is not limited to) cfDNA fragments or chromosomal DNA from a subject with cancer. The genomic DNA can be sequenced and its methylation status can be evaluated by various methods (e.g., bisulfite sequencing). When the genomic sequence is obtained from a public database (e.g., The Cancer Genome Atlas (TCGA)) or through a sequencing experiment on the genome of an individual diagnosed with cancer, the cancerous sample can refer to genomic DNA or cfDNA fragments with a genomic sequence. The term "cancerous samples" in the plural form refers to samples containing genomic DNA from multiple individuals, each diagnosed with cancer. In various embodiments, cancerous samples from more than 100, 300, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 40,000, 50,000 or more individuals diagnosed with cancer are used.

[0059] As used herein, the term "non-cancerous sample" or "healthy sample" refers to a sample containing genomic DNA from a healthy individual or an individual not diagnosed with cancer. The genomic DNA can be (but is not limited to) cfDNA fragments or chromosomal DNA from a subject without cancer. The genomic DNA can be sequenced and its methylation status can be evaluated by various methods (e.g., bisulfite sequencing). When the genomic sequence is obtained from a public database (e.g., The Cancer Genome Atlas (TCGA)) or through a sequencing experiment on the genome of an individual without cancer, the non-cancerous sample can refer to genomic DNA or cfDNA fragments with a genomic sequence. The term "non-cancerous samples" in the plural form refers to samples containing genomic DNA from multiple individuals, each without cancer. In various embodiments, cancerous samples from more than 100, 300, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 40,000, 50,000 or more individuals without cancer are used. In various embodiments, cancerous samples from 100 or more, 300 or more, 500 or more, 1,000 or more, 2,000 or more, 5,000 or more, 10,000 or more, 20,000 or more, 40,000 or more or 50,000 or more individuals without cancer are used.

[0060] As used herein, the term "training sample" refers to a sample used to train a classifier as described herein and / or to select one or more genomic regions for cancer detection or to detect the tissue or cancer cell type of cancer origin. A training sample can comprise genomic DNA or its modifications from one or more healthy subjects and from one or more subjects with a disease condition (e.g., cancer, a particular type of cancer, cancer at a particular stage, etc.). The genomic DNA can be, but is not limited to, cfDNA fragments or chromosomal DNA. The genomic DNA can be sequenced and its methylation status can be evaluated by various methods (e.g., bisulfite sequencing). When genomic sequences are obtained from a public database (e.g., The Cancer Genome Atlas (TCGA)) or by performing a sequencing experiment on an individual's genome, a training sample can refer to genomic DNA or cfDNA fragments having the genomic sequences.

[0061] As used herein, the term "test sample" refers to a sample from a subject whose health status has been, is being, or will be tested using a classifier and / or assay set as described herein. A test sample can comprise genomic DNA or its modifications. The genomic DNA can be, but is not limited to, cfDNA fragments or chromosomal DNA.

[0062] As used herein, the term "target genomic region" refers to a region in the genome that is selected for analysis in a test sample. An assay set is generated with oligonucleotide probes that are designed to hybridize (and optionally pull down) nucleic acid fragments derived from the target genomic region or fragments thereof. The oligonucleotide probes directed against the target region are also referred to herein as "bait oligonucleotides". Nucleic acid fragments derived from the target genomic region refer to nucleic acid fragments generated by degrading, cutting, bisulfite converting, or other processing of DNA from the target genomic region. In some embodiments, multiple different bait oligonucleotides are designed to hybridize across a single target genomic region (e.g., overlapping probes tiled across the target genomic region). Generally, when referring to multiple target genomic regions, none of the multiple target genomic regions is completely contained within another target genomic region. Different target genomic regions among the multiple target genomic regions can overlap, but will have at least different ends. In some embodiments, each of the multiple target genomic regions is separate and non-overlapping from any other target genomic region among the multiple target genomic regions. In some embodiments, the target genomic region contains a target sequence of a gene. In this context, the term "gene" encompasses the sequence encoding the gene, as well as optional flanking sequences (e.g., regulatory sequences such as promoters, enhancers, and untranslated regions) and internal non-coding sequences (e.g., introns). In some embodiments, a gene is defined by the sequence encoding the transcription start and stop sites and an additional 5000 nucleotides from each of these ends. A given gene may contain multiple target sequences.

[0063] Target genomic regions can be described according to their chromosomal location, e.g., relative to a specific reference sequence (e.g., the human reference genome GRCh37 / hg19). Chromosomal DNA is double-stranded, so a target genomic region includes two DNA strands: one having the sequence according to a given reference sequence, and the second being its reverse complement. Probes can be designed to hybridize to one or both sequences. Optionally, the probes hybridize to a converted sequence generated, e.g., by treatment with sodium bisulfite.

[0064] As used herein, the term "off-target genomic region" refers to a region in the genome that is not selected for analysis in a test sample but has sufficient homology to the target genomic region to potentially bind and be captured by a probe designed to target that target genomic region. In one embodiment, the off-target genomic region is a genomic region that aligns with the probe along at least 45 bp and has a match rate of at least 90%.

[0065] The terms "converted DNA molecule", "converted cfDNA molecule", and "modified fragment obtained from processing these cfDNA molecules" refer to DNA molecules obtained by processing DNA or cfDNA molecules in a sample, the purpose of which is to distinguish methylated nucleotides and unmethylated nucleotides in the DNA or cfDNA molecules. For example, in some embodiments, a sample can be treated with bisulfite ions (e.g., using sodium bisulfite) to convert unmethylated cytosine ("C") to uracil ("U"). In another embodiment, the conversion of unmethylated cytosine to uracil is achieved using an enzymatic conversion reaction (e.g., using a cytidine deaminase such as APOBEC). After treatment, the converted DNA molecule or cfDNA molecule includes additional uracils that were not present in the original cfDNA sample. Replication of a DNA strand containing uracil by DNA polymerase results in the addition of adenine to the nascent complementary strand, rather than guanine which is normally added as the complement of cytosine or methylcytosine.

[0066] Generally, the terms "cell-free", "circulating", and "extracellular" are used interchangeably when referring to polynucleotides (e.g., "cell-free DNA" or "cfDNA") and refer to polynucleotides or portions thereof present in a sample from a subject that can be isolated or otherwise manipulated without applying a lysis step (e.g., as in lysis used for extraction from cells or viruses) to the originally collected sample. Thus, even before collection of a subject sample, cell-free polynucleotides are not encapsulated or "free" within the cells or viruses from which they originate. Cell-free polynucleotides may be produced as a byproduct of cell death (e.g., apoptosis or necrosis) or cell shedding, releasing the polynucleotides into the surrounding body fluid or circulation. Thus, cell-free nucleic acids can be isolated from the acellular fraction of blood (e.g., serum or plasma), from other body fluids (e.g., urine), or from the acellular fraction of other types of samples. Non-limiting examples of body fluids that can be used in conjunction with the embodiments disclosed herein include mucus, blood, plasma, serum, serum derivatives, synovial fluid, lymphatic fluid, bile, sputum, saliva, sweat, tears, phlegm, amniotic fluid, menstrual fluid, vaginal secretions, semen, urine, cerebrospinal fluid (CSF) (e.g., lumbar or ventricular CSF), gastric fluid, liquid samples containing one or more substances derived from nasal, throat, or oral swabs, liquid samples containing one or more substances derived from lavage procedures (e.g., peritoneal, gastric, thoracic, or catheter lavage procedures), and the like. In some embodiments, cfDNA refers to deoxyribonucleic acid molecules that circulate in a subject (e.g., in the bloodstream) and may be derived from one or more healthy cells and / or from one or more cancer cells. In some embodiments, cfDNA is the cfDNA of a urine sample. In some embodiments, the compositions (e.g., assay sets) and methods disclosed herein related to urine cfDNA can be applied or adapted to other sample types (e.g., cfDNA from other body fluids such as blood, serum, or plasma).

[0067] The term "circulating tumor DNA" or "ctDNA" refers to nucleic acid fragments that are derived from tumor cells and that may be released into the bloodstream of an individual due to biological processes such as apoptosis or necrosis of dying cells or active release by surviving tumor cells.

[0068] As used herein, the term "fragment" can refer to a fragment of a nucleic acid molecule. For example, in one embodiment, a fragment can refer to a cfDNA molecule in a blood or plasma sample, or a cfDNA molecule that has been extracted from a blood or plasma sample. An amplification product of a cfDNA molecule can also be referred to as a "fragment". In another embodiment, the term "fragment" refers to sequence reads or sets of sequence reads that have been processed for subsequent analysis (e.g., for machine learning-based classification) as described herein. For example, raw sequence reads can be aligned to a reference genome, and the matching paired-end sequence reads can be assembled into longer fragments for subsequent analysis.

[0069] The terms "individual" and "subject" refer to a human individual. The term "healthy individual" refers to an individual who is presumably free of cancer or disease. In some embodiments, the subject is the individual whose DNA is being analyzed. For example, the subject may be a test subject whose DNA will be evaluated using a targeted panel as described herein to evaluate whether the person has cancer or other disease. In some embodiments, the subject is part of a control group (also referred to as a "reference subject") known to have (or not have) cancer or other disease. The control group and the cancer / disease group can be used to assist in the design or validation of the targeted panel.

[0070] As used herein, the term "sequence read" refers to a string of nucleotides determined to be part or all of a nucleic acid molecule by a nucleic acid sequencing process. A sequence read can be a short string of nucleotides (e.g., 20 - 150) sequenced from a nucleic acid fragment, a short string of nucleotides at one or both ends of a nucleic acid fragment, or the sequencing of an entire nucleic acid fragment present in a biological sample. Sequence reads can be obtained by the various methods provided herein or by other methods known in the art.

[0071] As used herein, the term "sequencing depth" refers to a count of the number of times a given target nucleic acid in a sample has been sequenced (e.g., a count of sequence reads at a given target region), or the average number of times, actual or expected, that a given target nucleic acid is sequenced based on the amount of nucleic acid being sequenced and the total read length generated by a given sequencing process (e.g., the average read depth across all sequenced regions in a given sequencing run). Increasing the sequencing depth can reduce the amount of nucleic acid required to assess a disease state (e.g., cancer or cancer-derived tissue).

[0072] As used herein, the term "tissue of origin" or "TOO" refers to the organ, group of organs, body region, or cell type from which cancer has occurred or originated. Identifying the tissue of origin or cancer cell type typically allows identification of the most appropriate subsequent steps in the cancer care continuum for further diagnosis, staging, and treatment decisions. For example, a cancer originating from cells in the bladder can be identified as bladder cell based on one or more markers associated with bladder cells, even after metastasis to a different tissue (e.g., kidney). In the case of metastasis, "cancerous tissue" refers to the metastatic growth of specific TOO cells, distinct from the cells of the tissue to which the cancer has metastasized. Thus, metastatic cancerous tissue may be located near or within the healthy tissue at the site of metastasis.

[0073] As used herein, "treating" or "treatment" includes any approach for obtaining a beneficial or desired result in a subject's disorder, including a clinical result. Beneficial or desired clinical results can include, but are not limited to: alleviation or improvement of one or more symptoms or disorders, reduction in the degree of disease, stabilization (i.e., non-worsening) of the disease state, prevention of the spread or dissemination of the disease, delay or slowing of disease progression, amelioration or palliation of the disease state, reduction in the recurrence of the disease, and remission (whether partial or complete, and whether detectable or undetectable). In other words, as used herein, "treatment" includes any cure, improvement, or prevention of a disease. Treatment can prevent the occurrence of a disease; inhibit the spread of a disease; alleviate the symptoms of a disease, eliminate the underlying cause of the disease, in whole or in part, shorten the duration of the disease, or achieve a combination of these things.

[0074] As used herein, "treating" or "treatment" includes prophylactic treatment. The method of treatment includes administering to a subject a therapeutically effective amount of an active agent. The administering step can consist of a single administration, or can include a series of administrations. The length of the treatment period depends on a variety of factors, such as the severity of the disorder, the age of the patient, the concentration of the active agent, the activity of the composition used in the treatment, or a combination thereof. It should also be understood that the effective dose of the agent used for treatment or prophylaxis may increase or decrease during the course of a particular treatment or prophylaxis regimen. Changes in dose can be caused and become apparent by standard diagnostic assays known in the art. In some cases, long-term administration may be required. For example, administering to a subject a composition in an amount and for a duration sufficient to treat the patient. In the examples, "treating" or "treatment" is not prophylactic treatment.

[0075] The term "prevention", when referring to a disease or disorder in a subject, means reducing the occurrence of one or more corresponding symptoms in the subject. As shown above, prevention can be complete (no detectable symptoms) or partial, such that the symptoms observed are fewer and / or occur at a lower rate than those that might occur in the absence of treatment.

[0076] TM ) Erlotinib (Tarceva TM ) Cetuximab (Erbitux TM ) Lapatinib (Tykerb TM ) Panitumumab (Vectibix TM ) Vandetanib (Caprelsa TM ) afatinib / BIBW2992, CI-1033 / canertinib, neratinib / HKI-272, CP-724714, TAK-285, AST-1306, ARRY334543, ARRY-380, AG-1478, dacomitinib / PF299804, OSI-420 / desmethyl erlotinib, AZD8931, AEE788, pelitinib / EKB-569, CUDC-101, WZ8040, WZ4002, WZ3146, AG-490, XL647, PD153035, BMS-599626), sorafenib, imatinib, sunitinib, dasatinib, etc.

[0077] In some embodiments, the anti-cancer agent is an epigenetic inhibitor. As used herein, "epigenetic inhibitor" refers to an inhibitor of an epigenetic process, such as DNA methylation (DNA methylation inhibitor) or histone modification (histone modification inhibitor). Epigenetic inhibitors can be histone deacetylase (HDAC) inhibitors, DNA methyltransferase (DNMT) inhibitors, histone methyltransferase (HMT) inhibitors, histone demethylase (HDM) inhibitors or histone acetyltransferase (HAT). Examples of HDAC inhibitors include vorinostat, romidepsin, CI-994, belinostat, panobinostat, givinostat, entinostat, mocetinostat, SRT501, CUDC-101, JNJ-26481585 or PCI24781. Examples of DNMT inhibitors include azacitidine and decitabine. Examples of HMT inhibitors include EPZ-5676. Examples of HDM inhibitors include pargyline and tranylcypromine. Examples of HAT inhibitors include CCT077791 and mangostin.

[0078] In some embodiments, the anti-cancer agent is a multi-kinase inhibitor. A "multi-kinase inhibitor" is a small molecule inhibitor of at least one protein kinase, which includes tyrosine protein kinases and serine / threonine kinases. A multi-kinase inhibitor may include a single kinase inhibitor. A multi-kinase inhibitor may block phosphorylation. A multi-kinase inhibitor may act as a covalent modifier of protein kinases. A multi-kinase inhibitor may bind to the kinase active site or to a secondary or tertiary site that inhibits protein kinase activity. A multi-kinase inhibitor may be an anti-cancer multi-kinase inhibitor. Exemplary anti-cancer multi-kinase inhibitors include dasatinib, sunitinib, erlotinib, bevacizumab, vatalanib, vemurafenib, vandetanib, cabozantinib, poatinib, axitinib, ruxolitinib, regorafenib, crizotinib, bosutinib, cetuximab, gefitinib, imatinib, lapatinib, lenvatinib, mulitinib, nilotinib, panitumumab, pazopanib, trastuzumab, or sorafenib.

[0079] Urine sample assay

[0080] In one aspect, the present disclosure provides a method for sequencing free nucleic acid molecules of a subject. In some embodiments, the method includes (a) treating a urine sample to inhibit cell lysis; (b) separating the free nucleic acid molecules in the treated urine sample from the cells in the treated urine sample, thereby producing a purified urine sample containing the free nucleic acid molecules; (c) concentrating the free nucleic acid molecules in the purified urine sample by passing at least a portion of the purified urine sample through a filter, wherein (i) the concentration produces a filtrate and a retained urine sample, and (ii) the retained urine sample contains free nucleic acid molecules with increased concentration; (d) separating the free nucleic acid molecules from the retained urine sample; and (e) sequencing the separated free nucleic acid molecules.

[0081] Urine samples for use according to the present disclosure can be collected from a variety of sources and exhibit various characteristics. In some embodiments, the urine sample contains a desired minimum initial volume. For example, the urine sample may have a volume of at least 5 mL, 10 mL, 20 mL, 30 mL, 40 mL, 50 mL, 75 mL, 100 mL or more. In some embodiments, the volume of the treated urine sample is 5 mL, 10 mL, 15 mL, 20 mL, 30 mL, 40 mL, 50 mL or more. In some embodiments, the urine sample is a sample of at least 50 mL.

[0082] Typically, a collected urine sample will contain free nucleic acid molecules and other components such as cells, metabolites, proteins, and salts. In some embodiments, the urine sample is treated to inhibit cell lysis. The inhibition of cell lysis need not be absolute, but generally reduces the rate of cell lysis in the sample compared to an untreated urine sample. In some embodiments, treating the urine sample to inhibit cell lysis includes contacting the urine sample with one or more preservative reagents. There are a variety of suitable preservative reagents available. Preservative reagents that can be used to inhibit cell lysis include, but are not limited to, imidazolidinyl urea (IDU), diazolidinyl urea (DU), dihydroxymethyl urea, 2-bromo-2-nitropropane-1,3-diol, 5-hydroxymethoxymethyl-1-aza-3,7-dioxabicyclo(3.3.0)octane, and 5-hydroxymethyl-1-aza-3,7-dioxabicyclo(3.3.0)octane and 5-hydroxypoly[methyleneoxy]methyl-1-aza-3,7-dioxabicyclo(3.3.0)octane, bicyclooxazolidine (such as Nuosept 95), DMDM hydantoin, sodium hydroxymethylglycinate, hexamethylenetetramine allyl chloride (quaternium-15), microbicides (such as Bioban, dichlorophen, and Grotan), water-soluble zinc salts, sodium azide, or any combination thereof. In some embodiments, the preservative reagent comprises imidazolidinyl urea.

[0083] In some embodiments, treating the urine sample to inhibit cell lysis includes treating with a combination of reagents, such as treating with a preservative reagent, a nuclease inhibitor, a formaldehyde quencher, or a combination of these. Treating the urine sample to inhibit cell lysis may affect the integrity of the nucleic acids in the sample. Some cell lysis inhibitors release formaldehyde, which, if not quenched, can damage or destroy the structural integrity of the nucleic acids. Thus, a formaldehyde quencher can be added to maintain the stability of the nucleic acids in the sample. Available formaldehyde quenchers include, but are not limited to, glycine, tris(hydroxymethyl)aminomethane (TRIS), urea, allantoin, sulfite, or any combination thereof. In some embodiments, the formaldehyde inhibitor is glycine.

[0084] In some embodiments, processing a urine sample includes treating it with a nuclease inhibitor. Nuclease activity in fresh urine can rapidly hydrolyze nucleic acids (e.g., DNA). Available nuclease inhibitors include, but are not limited to, ethylene glycol tetraacetic acid (EGTA), pepstatin, EDTA, phosphonodipeptide, leupeptin, aprotinin, bestatin, protease inhibitor E 64 (E-64), 4-(2-aminoethyl)benzenesulfonyl fluoride hydrochloride (AEBSF), or any combination thereof. In some embodiments, the nuclease inhibitor is EDTA. Treatment with a nuclease inhibitor can also be carried out with or without a formaldehyde quencher. In some embodiments, the treatment includes a nuclease inhibitor but does not include a formaldehyde quencher. In some embodiments, the treatment includes a nuclease inhibitor and a formaldehyde quencher.

[0085] In some embodiments, processing includes contacting the urine sample with a composition comprising a cell lysis inhibitor, and a nuclease inhibitor, a formaldehyde quencher, or both. In some embodiments, the composition comprises imidazolidinyl urea, EDTA, glycine, or a combination thereof. In some embodiments, imidazolidinyl urea is added to a final concentration of 0.5% to 2.0%. In some embodiments, imidazolidinyl urea is added to a final concentration of 0.20% to 4.0%. In some embodiments, glycine is present at a concentration of 0.01% to 0.2%. In some embodiments, glycine is present at a concentration of 0.01% to 0.4%. In some embodiments, glycine is present at a concentration of at least 0.01%, at least 0.05%, at least 0.2%, or even at least 0.3%. In some embodiments, EDTA is present at a concentration of 0.5% to 2.0%. In some embodiments, EDTA is present at a concentration of 0.50% to 3.6%. In some embodiments, EDTA is present at a concentration of at least 0.5%, at least 1%, or at least 2.5%. In some embodiments, the ratio of EDTA to imidazolidinyl urea is 5:2. In some embodiments, the ratio of EDTA to imidazolidinyl urea is 1:3 to 3:1 (e.g., 9:10, 9:5, or 9:20). In some embodiments, the ratio of EDTA to glycine is 5:1 to 100:1 (e.g., 50:1 or 9:1).

[0086] In some embodiments, the composition for treating a urine sample to inhibit cell lysis comprises sodium azide, EDTA, or a combination thereof. In some embodiments, sodium azide is present at a final concentration of 0.05% to 2.0%. In some embodiments, sodium azide is present at a concentration of 0.10% to 1.0%. In some embodiments, EDTA is present at a concentration of 0.5% to 2.0%. In some embodiments, EDTA is present at a concentration of 0.50% to 3.6%. In some embodiments, EDTA is present at a concentration of at least 0.5%, at least 1%, or at least 2.5%.

[0087] In some embodiments, the composition for treating a urine sample to inhibit cell lysis is present in an amount of 1 to 20 percent of the volume of the urine sample after contact. The preservative reagent can be combined with the urine sample after collection (e.g., by adding it to the urine sample or transferring all or part of the urine sample to a container having the preservative reagent), or can be present in the container used to collect the sample. Additional non-limiting examples of compositions comprising one or more of a preservative reagent, a nuclease inhibitor, or a formaldehyde quencher are described in U.S. Publication No. 20160257995A1, which is incorporated herein by reference.

[0088] In some embodiments, the method includes the step of separating free nucleic acid molecules (e.g., cfDNA and / or cfRNA) in the treated urine sample from the cells in the treated urine sample to produce a purified urine sample comprising these free nucleic acid molecules (and reduced cell content). There are a variety of methods available for separating free nucleic acids from cells, non-limiting examples of which include fractionation, centrifugation (e.g., pelleting or density gradient centrifugation), precipitation, and flow cytometry. In some embodiments, the separation includes centrifuging the treated urine sample to pellet the cells. The cell pellet can then be removed or the supernatant can be transferred to a new container. In some embodiments, starting with the treated urine sample, the free nucleic acids (e.g., free DNA) can be separated from the cells by centrifuging at, for example, 3000 rpm to 8000 rpm for 10 to 15 minutes at room temperature. The treated urine sample can be centrifuged at at least 1000g, 2000g, 3000g, 4000g, 5000g, 6000g, 7000g, 8000g or more. Additionally, the centrifugation can be performed for at least 5, 10, 15, 20, 30, 45, 60 or more minutes. In some embodiments, the treated urine sample is centrifuged at 4000g for 20 minutes. The separated cells form a cell pellet after centrifugation, which can be removed by transferring the supernatant to a different container, thereby separating these cells from the nucleic acids. The separation can be performed one, two, three, four or more times, or until no visible cell pellet forms upon centrifugation.

[0089] In some embodiments, the purified urine sample undergoes a concentration step to produce a sample with increased nucleic acid concentration. Concentration can offer certain advantages, such as allowing for a larger input sample volume to be processed compared to other sample types (e.g., plasma), and compensating for the reduced concentration of free nucleic acids in urine. Samples with higher nucleic acid concentration and smaller volume can be processed using a reduced amount of reagents and materials, and allow for easier parallel processing of multiple samples. In some embodiments, concentration of free nucleic acid molecules in the purified urine sample is performed by passing at least a portion of the purified urine sample through a filter. Passing the purified urine sample through the filter produces a filtrate and a retained urine sample. The filtrate will contain salts and other small components that can pass through the filter, while the retained urine sample will contain nucleic acids at increased concentration. In some cases, the filter is substantially impermeable to nucleic acids (e.g., free nucleic acids) and substantially permeable to salts and other small components in the purified urine sample. When a large sample volume is filtered in two or more portions, each of the two or more portions can be applied continuously to the same filter, applied separately to different filters, or some combination of these. The retained portion of the filtered sample can undergo one or more additional rounds of filtration (e.g., 2, 5, 10, 15 or more rounds), for example by repeated application on the same filter, or by successive application to different filters.

[0090] Any of a variety of types of filters made of various materials can be used, and several options are commercially available. For example, the filter can be a membrane filter, such as a filter made of nylon, cellulose, or nitrocellulose, or can contain beads (e.g., agarose beads). In some embodiments, the filter is characterized by a molecular cut-off. Generally, the molecular weight cut-off is determined by a specific pore size and / or coating that allows molecules above the cut-off to be separated from molecules below the cut-off. The molecular weight cut-off (also referred to as "molecular weight limit") is typically specified as one of the characteristics of commercially available filters for processing biological samples. In some embodiments, the molecular weight cut-off represents the molecular weight of the solute that is 90% retained by the filter. In some embodiments, the filter has a rated molecular weight cut-off of 1 kilodalton (kD) to 50 kD. In some embodiments, the rated molecular weight cut-off of the filter is 3 kD to 10 kD. In some embodiments, the filter has a rated molecular weight cut-off of 10 kD, 5 kD, 3 kD or lower. In some embodiments, the filter has a rated molecular weight cut-off of 3 kD or lower.

[0091] After one or more concentration steps, various nucleic acid concentrations in the retained urine sample can be achieved relative to the purified urine sample. In some embodiments, the retained urine sample has a concentration increased by at least 2-fold compared to the purified urine sample (before concentration). In some embodiments, the retained urine sample has a concentration increased by at least 2-fold to 20-fold. In some embodiments, compared to the purified urine sample, the retained urine sample has a concentration increased by at least 2-fold, at least 3-fold, at least 4-fold, at least 5-fold, at least 10-fold, at least 15-fold, at least 20-fold or more. In some embodiments, the retained urine sample has a concentration increased by at least 2-fold, at least 5-fold, at least 10-fold or at least 15-fold compared to the purified urine sample. In some embodiments, compared to the purified urine sample, the retained urine sample has a concentration increased by at least 5-fold. In some embodiments, compared to the purified urine sample, the retained urine sample has a concentration increased by at least 10-fold.

[0092] Since at least a portion of the urine sample is passed through a filter, the retained urine sample (including the sample portion that has been filtered but not passed through the filter into the filtrate) will have a smaller volume compared to the starting volume of the sample that has been filtered. In some embodiments, the retained urine sample has a volume that is at least 10% to at least 90% lower than the volume of the processed urine sample. In some embodiments, the retained urine sample has a volume that is at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 80% or at least 90% lower than the volume of the processed urine sample. In some embodiments, the retained urine sample has a volume that is at least 50% lower than the volume of the processed urine sample. In some embodiments, the retained urine sample has a volume that is at least 75% lower than the volume of the processed urine sample. For example, when a 50 mL processed urine sample is concentrated to a retained urine sample volume of 5 mL or less, the volume of the retained urine sample will be reduced by 90% or more.

[0093] The retained urine sample can be immediately used for further processing or for assaying free nucleic acids. For example, the retained urine sample can be processed within one hour or less after concentration. Alternatively, the retained urine sample can be stored for subsequent use. In some cases, the retained urine sample can be stored at 4°C, 22°C or 37°C for up to one week after urine sample collection. In some embodiments, the retained urine sample can be frozen to preserve the sample for subsequent analysis. In some embodiments, the retained urine sample is stored at -20°C, -80°C or lower. In some embodiments, the frozen sample is stored in the frozen state for one to six months.

[0094] The steps of processing, separating, and concentrating the urine sample can be performed at different times after the collection of the urine sample. In some embodiments, the processing is completed within 10 to 180 minutes after the collection of the urine sample. In some embodiments, the processing is completed within 180, 150, 120, 90, 60, 45, 30, 15, or 10 minutes after the collection of the urine sample. In some embodiments, the processing is completed within 120, 60, or 30 minutes after the collection of the urine sample. In some embodiments, the processing is completed within 60 minutes after the collection of the urine sample. In some embodiments, the processing is completed within 30 minutes after the collection of the urine sample. In some embodiments, the processed sample is stored before performing the separating and concentrating steps. For example, after processing, the further separating and concentrating steps can be completed within 1 to 14 days after collection (e.g., within 2, 3, 4, 5, 6, 7, 8, 9, or 10 days). In some embodiments, the separating and concentrating are completed within 7 days after the collection of the urine sample.

[0095] Nucleic acids can be isolated from the retained urine sample by any of a variety of nucleic acid isolation methods. Exemplary methods include, but are not limited to, extraction, solid-phase extraction, silica-based purification, magnetic particle-based purification, phenol-chloroform extraction, chromatography, anion-exchange chromatography (using an anion-exchange surface), electrophoresis, filtration, precipitation, immunoprecipitation, hybridization capture with a targeted bait oligonucleotide, or any combination thereof. For example, the target nucleic acid is isolated by hybridizing a streptavidin-coated surface (e.g., streptavidin-coated beads) with a biotinylated probe. Commercially available methods and kits are also available, non-limiting examples of which include the Circulating Nucleic Acid Kit (QIAGEN), the Chemagic Circulating NA Kit (Chemagen), the NucleoSpin Plasma XS Kit (Macherey-Nagel), the High Pure Viral Nucleic Acid Large Volume Kit (Roche).

[0096] In some embodiments, the isolated cell-free nucleic acids comprise cfDNA, and the method includes processing these cfDNA molecules to distinguish methylated nucleotides from unmethylated nucleotides, thereby generating transformed cfDNA molecules. In some embodiments, the processing includes deamination, such as with a cytidine deaminase or by treating the cfDNA molecules with bisulfite. In some embodiments, the method includes treating the cfDNA molecules with bisulfite to generate transformed cfDNA molecules.

[0097] The methods disclosed herein may further include amplification of free nucleic acids isolated from a retained urine sample. The amplification may be non-specific (e.g., amplification using random primers) or target-directed (e.g., directed against a specific target region of interest). Any suitable method known in the art may be used for amplification. Examples of nucleic acid amplification reactions include, but are not limited to, polymerase chain reaction (PCR), rolling circle amplification (RCA), ligase chain reaction (LCR), simple method for amplifying RNA targets (SMART), single primer isothermal amplification (SPIA), multiple displacement amplification (MDA), nucleic acid sequence-based amplification (NASBA), hinge-initiated primer-dependent nucleic acid amplification (HIP), nicking enzyme amplification reaction (NEAR), RT-PCR, loop-mediated amplification (LAMP), exponential amplification reaction (EXPAR), and improved multiple displacement amplification (IMDA). Primer-based amplification methods may use primers targeting any region of the genome. Alternatively, primers may be used to specifically amplify a target biomarker, thereby enriching the desired target / biomarker in the sample. For example, forward and reverse primers may be prepared for each genomic region of interest and used to amplify fragments corresponding to or derived from the desired genomic region. The amplification may be thermal amplification (e.g., as in PCR) or isothermal amplification.

[0098] In some embodiments, the isolated free nucleic acids or their amplification products are captured by hybridization to a bait oligonucleotide. In some embodiments, the length of the bait oligonucleotide is at least 45 nucleotides (e.g., the length is at least 60, 75, 80, 90, 100, 110, or 120 nucleotides). In some embodiments, the length of the bait oligonucleotide does not exceed 130, 140, 150, 200, 250, or 300 bases. In some embodiments, the length of the bait oligonucleotide is 45 to 300, 60 to 200, or 75 to 150 nucleotides. In some embodiments, the length of the bait oligonucleotide is at least 50 nucleotides. In some embodiments, the length of the bait oligonucleotide is at least 60 nucleotides. In some embodiments, the length of the bait oligonucleotide is at least 75 nucleotides. In some embodiments, the indicated length of the bait oligonucleotide is designed to be complementary to a portion of the target genomic sequence or its converted DNA molecule.

[0099] In some embodiments, the bait oligonucleotides target at least 500, 1000, 1500, 5000, 10000, 12500, 15000, 17000, 19000 or more target genomic regions. In some embodiments, the bait oligonucleotides target fewer than 25000, 20000, 17000, 15000, 12500, 10000 or fewer target genomic regions. In some embodiments, the bait oligonucleotides target 5000 to 30000, 10000 to 25000, 12500 to 20000 or 15000 to 20000 target genomic regions. In some embodiments, the bait oligonucleotides target 100 to 500, 500 to 1000, 1500 to 5000, 5000 to 10000 target genomic regions. In some embodiments, the bait oligonucleotides target at least 500 target genomic regions. In some embodiments, the bait oligonucleotides target at least 1000 target genomic regions. In some embodiments, the bait oligonucleotides target at least 10000 target genomic regions. In some embodiments, the bait oligonucleotides target at least 15000 target genomic regions. In some embodiments, the bait oligonucleotides target fewer than 20000 target genomic regions.

[0100] In some embodiments, the bait oligonucleotides are configured to hybridize to a transformed DNA molecule (e.g., a transformed cfDNA molecule) corresponding to or derived from one or more genomic regions. Thus, the bait oligonucleotides can have a sequence different from the target genomic region. For example, DNA containing an unmethylated CpG site can be transformed to include UpG instead of CpG by deamination (e.g., by treatment with a cytidine deaminase or bisulfite). Thus, bait oligonucleotides targeting such targets can be configured to hybridize to a sequence including UpG instead of the naturally occurring unmethylated CpG. Thus, the site in the bait oligonucleotide complementary to the unmethylated site can contain CpA instead of CpG, and some probes targeting hypomethylated sites where all methylated sites are unmethylated may not have a guanine (G) base. In some embodiments, at least 3%, 5%, 10%, 15% or 20% of the probes do not contain a CpG sequence. In some embodiments, at least 5% of the probes do not contain a CpG sequence. In some embodiments, at least 10% of the probes do not contain a CpG sequence.

[0101] In some embodiments, the decoy oligonucleotides are typically used to detect the presence or absence of cancer and / or provide cancer classification (e.g., cancer type), cancer staging (e.g., I, II, III, or IV), or provide the tissue of origin (TOO) considered to be the cancer origin. The decoy oligonucleotides may target differentially methylated genomic regions between generally cancerous (pan-cancer) samples and non-cancerous samples or only in cancerous samples with a specific cancer type (e.g., urinary system cancer-specific targets). For example, in some embodiments, the decoy oligonucleotides are designed to include differentially methylated genomic regions based on bisulfite sequencing data generated from cell-free nucleic acids and / or whole-genome DNA from a set of cancer and non-cancer individuals.

[0102] In some embodiments, each of the target genomic regions in the cancer sample is differentially methylated relative to the non-cancer sample. In some embodiments, the differential methylation comprises at least 50% to at least 90% of the CpG sites in the target genomic region being methylated or unmethylated. In some embodiments, the differential methylation comprises at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% of the CpG sites in the target genomic region being methylated or unmethylated. In some embodiments, the differential methylation comprises at least 70% of the CpG sites in the target genomic region being methylated or unmethylated. In some embodiments, the differential methylation comprises at least 80% of the CpG sites in the target genomic region being methylated or unmethylated. In some embodiments, the cancer is a urinary system cancer. In some embodiments, the cancer is bladder cancer, prostate cancer, or kidney cancer. In some embodiments, the target genomic regions are selected to identify the presence of and optionally distinguish between: two or more cancer types (e.g., at least 3, 4, 5, or more cancer types).

[0103] In some embodiments, the cell-free nucleic acids include DNA or RNA. In some embodiments, the cell-free nucleic acids are cell-free DNA (cfDNA). In some embodiments, the method includes processing cfDNA molecules to distinguish methylated nucleotides and unmethylated nucleotides, thereby generating transformed cfDNA molecules. In some embodiments, the processing includes deamination, such as with a cytidine deaminase or treating the cfDNA molecules with bisulfite. In some embodiments, the method includes treating the cfDNA molecules with bisulfite to generate transformed cfDNA molecules.

[0104] In some embodiments, the method includes separating bait-bound free nucleic acid molecules from unbound free nucleic acid molecules. In some embodiments, each bait oligonucleotide is conjugated to a solid surface (e.g., a chip or bead, such as a magnetic or paramagnetic bead) or to a non-nucleotide affinity moiety (e.g., a member of a binding pair). In some embodiments, such conjugation is used to facilitate the separation of cfDNA molecules bound to the bait oligonucleotide from unbound cfDNA molecules. In some embodiments, the affinity moiety is biotin.

[0105] In some embodiments, genomic regions can be selected to have at least 3, 5, 7, 10 or more methylation sites. In some embodiments, each target genomic region contains at least five methylation sites. In some embodiments, the selected number of methylation sites (e.g., at least 5 methylation sites) are methylation sites that are differentially methylated in at least one type of cancer to be assayed by the panel. The selected target genomic regions can be located at different positions in the genome, including but not limited to promoters, enhancers, exons, introns, intergenic regions, and other portions.

[0106] In some embodiments, the target genomic region is selected from the gene sequences in Table 1. In some embodiments, the target genomic region comprises a target sequence of a gene selected from Table 1. The target sequence may comprise two or more target sequences from a single gene, and / or one or more target sequences from two or more genes. In some embodiments, one or more target genomic regions comprise one or more target sequences from at least 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, 75 or more genes selected from Table 1. In some embodiments, the target genomic region comprises one or more target sequences from at least 10 genes in Table 1. In some embodiments, the target genomic region comprises one or more target sequences from at least 15 genes in Table 1. In some embodiments, the target genomic region comprises one or more target sequences from at least 25 genes in Table 1. In some embodiments, the target genomic region comprises a target sequence of a gene selected from one or more of the following: TWIST1, EOMES, HOXA9, POU4F2, and ZNF154. In some embodiments, the target genomic region comprises a target sequence of a gene selected from one or more of the following: TWIST1, PXDN, RP11-259O2.3, EVX2, and KNDC1. In some embodiments, the target genomic region comprises a target sequence of the TWIST1 gene. In some embodiments, the target genomic region comprises one or more target sequences of at least 20% of the genes in Table 1. In some embodiments, the target genomic region comprises one or more target sequences of at least 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95% or 100% of the genes in Table 1. In some embodiments, the target genomic region comprises one or more target sequences of at least 50% of the genes in Table 1. In some embodiments, the target genomic region comprises target sequences of all the genes in Table 1. In some embodiments, the length of the target sequence is at least 10, 25, 35, 45, 50, 75, 100, 200, 300 or more nucleotides. In some embodiments, the length of the target sequence is at least 25, at least 35 or at least 45 nucleotides. In some embodiments, the length of the target sequence is at least 45 nucleotides. In some embodiments, the target genomic region can be used for diagnosing one or more of bladder cancer, kidney cancer or prostate cancer.

[0107] Table 1

[0108]

[0109]

[0110] In some embodiments, the target genomic region is selected from the gene sequences in Table 2. In some embodiments, the target genomic region comprises a target sequence of a gene selected from Table 2. The target sequence may comprise two or more target sequences from a single gene, and / or one or more target sequences from two or more genes. In some embodiments, one or more target genomic regions comprise one or more target sequences from at least 1, 2, 3, 4, 5, 10, 15, 20 or more genes selected from Table 2. In some embodiments, the target genomic region comprises one or more target sequences from at least 5 genes in Table 2. In some embodiments, the target genomic region comprises one or more target sequences from at least 10 genes in Table 2. In some embodiments, the target genomic region comprises one or more target sequences from at least 15 genes in Table 2. In some embodiments, the target genomic region comprises a target sequence of a gene selected from one or more of the following: TWIST1, PXDN, RP11-259O2.3, EVX2, and KNDC1. In some embodiments, the target genomic region comprises a target sequence of the TWIST1 gene. In some embodiments, the target genomic region comprises target sequences of all genes in Table 2. In some embodiments, the length of the target sequence is at least 10, 25, 35, 45, 50, 75, 100, 200, 300 or more nucleotides. In some embodiments, the length of the target sequence is at least 25, at least 35 or at least 45 nucleotides. In some embodiments, the length of the target sequence is at least 45 nucleotides. In some embodiments, the target genomic region can be used for diagnosing bladder cancer.

[0111] Table 2

[0112] TWIST1 RP11 - 434B12.1 FANCC PTCHD3P1 PXDN CRTC1 CSNK1G2 SVIL RP11 - 259O2.3 CDT1 CPT1A ATXN7L1 EVX2 C2orf43 RP11 - 712B9.2 MAP7D1 KNDC1 HNRNPUL2 - BSCL2 PSMG4 CARS SEMA6B HNRNPUL2 SLC22A23 EPN1

[0113] In some embodiments, the target genomic region is selected from the gene sequences in Table 3. In some embodiments, the target genomic region comprises a target sequence of a gene selected from Table 3. The target sequence may comprise two or more target sequences from a single gene, and / or one or more target sequences from two or more genes. In some embodiments, one or more target genomic regions comprise one or more target sequences from at least 1, 2, 3, 4, 5, 10 or more genes selected from Table 3. In some embodiments, the target genomic region comprises one or more target sequences from at least 5 genes from Table 3. In some embodiments, the target genomic region comprises one or more target sequences from at least 10 genes from Table 3. In some embodiments, the target genomic region comprises one or more target sequences from at least 15 genes from Table 3. In some embodiments, the target genomic region comprises a target sequence of a gene selected from one or more of the following: TWIST1, SEMA6B, PXDN, KNDC1, and RP11-259O2.3. In some embodiments, the target genomic region comprises a target sequence of the TWIST1 gene. In some embodiments, the target genomic region comprises target sequences of all the genes from Table 3. In some embodiments, the length of the target sequence is at least 10, 25, 35, 45, 50, 75, 100, 200, 300 or more nucleotides. In some embodiments, the length of the target sequence is at least 25, at least 35 or at least 45 nucleotides. In some embodiments, the length of the target sequence is at least 45 nucleotides. In some embodiments, the target genomic region can be used for diagnosing bladder cancer.

[0114] Table 3

[0115] TWIST1 FBXL19 PSMG4 CPT1A SEMA6B PTCHD3P1 SLC22A23 MAP7D1 PXDN SVIL RP11 - 712B9.2 CARS KNDC1 CRTC1 ATXN7L1 EPN RP11 - 259O2.3

[0116] In some embodiments, the target genomic region is selected from the gene sequences in Table 4. In some embodiments, the target genomic region comprises a target sequence of a gene selected from Table 4. The target sequence may comprise two or more target sequences from a single gene, and / or one or more target sequences from two or more genes. In some embodiments, one or more target genomic regions comprise one or more target sequences from at least 1, 2, 3, 4, 5, 10, 15, 20, 30 or more genes selected from Table 4. In some embodiments, the target genomic region comprises one or more target sequences from at least 10 genes from Table 4. In some embodiments, the target genomic region comprises one or more target sequences from at least 15 genes from Table 4. In some embodiments, the target genomic region comprises one or more target sequences from at least 25 genes from Table 4. In some embodiments, the target genomic region comprises a target sequence of a gene selected from one or more of the following: LBH, HIST1H4D, HIST1H2APS4, FZD2, and ADARB2. In some embodiments, the target genomic region comprises the target sequences of all the genes from Table 4. In some embodiments, the length of the target sequence is at least 10, 25, 35, 45, 50, 75, 100, 200, 300 or more nucleotides. In some embodiments, the length of the target sequence is at least 25, at least 35 or at least 45 nucleotides. In some embodiments, the length of the target sequence is at least 45 nucleotides. In some embodiments, the target genomic region can be used for diagnosing prostate cancer.

[0117] Table 4

[0118] LBH KCNQ4 PON1 DTYMK HIST1H4D RFTN1 PON3 SOX2 - OT HIST1H2APS4 RP11 - 807E13.3 DENND1A PACSIN2 FZD2 CTD - 2201E18.1 TRIP13 AC008752.3 ADARB2 AC053503.11 TTC1 CASKIN1 PRDM2 ASIC4 PWWP2A FAM86C1 FOXD4L1 HOXB3 ABL1 CTD - 2313N18.5 IGF2BP3 HOXB - AS3 DPF3 AC006076.1 CAV2 HOXB5 KDM4B TSC22D2 CYTH3 HOXB6 LMAN2 HGFAC RP11 - 1085N6.6 TRAPPC9 NCOR2 RP11 - 271K21.11 AC022182.1 ZFPM1 SLC2A1 UNKL RASSF2 RP11 - 21B21.4 ASS1

[0119] In some embodiments, the target genomic region is selected from the gene sequences in Table 5. In some embodiments, the target genomic region comprises a target sequence of a gene selected from Table 5. The target sequence may comprise two or more target sequences from a single gene, and / or one or more target sequences from two or more genes. In some embodiments, one or more target genomic regions comprise one or more target sequences from at least 1, 2, 3, 4, 5, 10, 15 or more genes selected from Table 5. In some embodiments, the target genomic region comprises one or more target sequences from at least 5 genes in Table 5. In some embodiments, the target genomic region comprises one or more target sequences from at least 10 genes in Table 5. In some embodiments, the target genomic region comprises one or more target sequences from at least 15 genes in Table 5. In some embodiments, the target genomic region comprises a target sequence of a gene selected from one or more of the following: NAPA-AS1, NAPA, MFHAS1, RANGAP1, and HOXB2. In some embodiments, the target genomic region comprises target sequences of all the genes in Table 5. In some embodiments, the length of the target sequence is at least 10, 25, 35, 45, 50, 75, 100, 200, 300 or more nucleotides. In some embodiments, the length of the target sequence is at least 25, at least 35 or at least 45 nucleotides. In some embodiments, the length of the target sequence is at least 45 nucleotides. In some embodiments, the target genomic region can be used for diagnosing renal cancer.

[0120] Table 5

[0121] NAPA - AS1 HOXB - AS1 MUC5B RP11 - 770E5.1 NAPA SHARPIN RP11 - 532E4.2 HOXB3 MFHAS1 NACC2 RP11 - 445F12.1 HOXB - AS3 RANGAP1 BNIP3 FBXW7 DTYMK HOXB2 WDR37 CALM3

[0122] In some embodiments, the methods disclosed herein include sequencing isolated cell-free nucleic acid molecules. There are a variety of methods available for sequencing nucleic acids, non-limiting examples of which include next-generation sequencing (NGS) technologies, including synthesis technologies (Illumina), pyrosequencing (454 Life Sciences), ion semiconductor technology (Ion Torrent sequencing), single molecule real-time sequencing (Pacific Biosciences), ligation sequencing (SOLiD sequencing), nanopore sequencing (Oxford Nanopore Technologies), or paired-end sequencing. In some embodiments, massively parallel sequencing is performed using synthesis sequencing with reversible dye terminators. In some embodiments, sequencing of the nucleic acid molecule comprises sequencing an amplicon of the isolated cell-free nucleic acid molecule.

[0123] In some embodiments, the method further comprises diagnosing cancer in a subject. In some embodiments, the method further comprises treating cancer in a subject. In some embodiments, the cancer is bladder cancer, prostate cancer, or kidney cancer. The particular treatment modality may depend on one or more of various factors, such as the specific type of cancer detected, cancer TOO, location, and stage. Non-limiting examples of treatment include surgical resection, radiotherapy, chemotherapy, and / or immunotherapy.

[0124] Cancer assay set

[0125] In one aspect, the present disclosure provides a cancer assay set comprising a plurality of probes or a plurality of probe pairs. The assay sets described herein may alternatively be referred to as bait sets or compositions comprising bait oligonucleotides. The probes may be probes containing polynucleotides that are specifically designed to target (e.g., by complementary means) one or more genomic regions that are differentially methylated between cancer and non-cancer samples, between different cancer source tissue (TOO) types, between different cancer cell types, and / or between samples of different cancer stages, as identified by the methods provided herein. In some embodiments, the target genomic regions are selected to maximize classification accuracy, subject to a size budget (determined by the sequencing budget and desired sequencing depth).

[0126] In one aspect, the present disclosure provides a composition comprising a plurality of different bait oligonucleotides. In some embodiments, (a) the bait oligonucleotides hybridize to a transformed DNA molecule derived from one or more target genomic regions; (b) the one or more target genomic regions comprise one or more target sequences of one or more genes selected from Table 1; (c) the one or more target genomic regions are differentially methylated in cancer; and (d) each bait oligonucleotide comprises a sequence of at least 25 nucleotides in length that hybridizes to one of the target sequences. To design a cancer assay set, an analysis system can collect samples corresponding to various outcomes under consideration, e.g., samples from subjects known to have cancer, samples from subjects considered to be healthy, samples from subjects with cancer having a known tissue of origin, etc. The source of the DNA molecules (e.g., cfDNA and / or ctDNA) used to select the target genomic regions can vary depending on the assay purpose. For example, different sources may be desirable for assays aimed at detecting cancer in general, a specific type of cancer, cancer stage, or tissue of origin. These samples can be processed by whole-genome bisulfite sequencing (WGBS) or obtained from a public database (e.g., TCGA). In some embodiments, the analysis system is a computing system having a computer processor and a computer-readable storage medium having instructions for causing the computer processor to perform any or all of the operations described in the present disclosure. In some embodiments, at least some of the samples subjected to analysis to identify the target genomic regions are urine samples, e.g., urine samples processed according to the methods described herein.

[0127] The analysis system can then select target genomic regions based on the methylation patterns of nucleic acid fragments. One approach considers the pairwise discriminability between the results of regions (or more specifically, CpG sites within the regions). Another approach considers the discriminability of regions (or more specifically, CpG sites within the regions) when considering each result compared to the rest. Based on the selected target genomic regions with high discriminability capabilities, the analysis system can design probes to target fragments from these selected genomic regions. The analysis system can generate cancer assay sets of different sizes. For example, a small cancer assay set includes probes targeting the most informative genomic regions; a medium cancer assay set includes the probes from the small cancer assay set and additional probes targeting the second most informative genomic regions; and a large cancer assay set includes the probes from the small and medium cancer assay sets and even more probes targeting the third most informative genomic regions. Using data obtained from such cancer assay sets (e.g., nucleic acid methylation status derived from the cancer assay set), the analysis system can train a classifier with various classification techniques to predict the likelihood that a sample has a specific outcome or status (e.g., cancer, a specific cancer type, other disorders, other diseases, etc.).

[0128] In an illustrative example, to design a cancer assay set, the analysis system can collect information on the methylation status of CpG sites of nucleic acid fragments from samples corresponding to the various outcomes under consideration (e.g., samples known to have cancer, samples considered to be healthy, samples from known TOO, etc.). These samples can be processed (e.g., with whole-genome bisulfite sequencing (WGBS)) to determine the methylation status of CpG sites, or information can be obtained from TCGA. In some embodiments, the analysis system is a computing system having a computer processor and a computer-readable storage medium having instructions for causing the computer processor to perform any or all of the operations described in this disclosure.

[0129] The analysis system can then select target genomic regions based on the methylation patterns of nucleic acid fragments. One approach considers the pairwise distinguishability between the results of regions (or more specifically CpG sites). Another approach considers the distinguishability of regions (or more specifically CpG sites) when considering each result compared to the rest. Based on the selected target genomic regions with high distinguishability capabilities, the analysis system can design probes to target fragments from these selected genomic regions. The analysis system can generate cancer assay sets of different sizes. For example, a small cancer assay set includes probes targeting the most informative genomic regions; a medium cancer assay set includes the probes from the small cancer assay set and additional probes targeting the second most informative genomic regions; and a large cancer assay set includes the probes from the small and medium cancer assay sets and even more probes targeting the third most informative genomic regions. Using such cancer assay sets, the analysis system can train classifiers with various classification techniques to predict the likelihood that a sample has a specific outcome or status (e.g., cancer, a specific cancer type, other disorders, other diseases, etc.).

[0130] In some embodiments, the cancer assay set comprises a plurality of probe pairs, where each pair of the plurality of pairs comprises two probes that are configured to overlap each other by an overlapping sequence, where the overlapping sequence comprises at least 30 nucleotides, and where each probe is configured to hybridize to the same strand of a (optionally converted) DNA molecule (e.g., a cfDNA molecule) corresponding to one or more genomic regions. In some embodiments, each of the two probes in each probe pair comprises a non-overlapping sequence that is more than 20, 25, 30, 35, 40, 45, or 50 nucleotides different from the different probe in the probe pair. Thus, in some embodiments, the first and second probes in a probe pair can overlap by 30 nucleotides, where the first probe further comprises more than 20, 25, 30, 35, 40, 45, or 50 nucleotides that do not overlap with the second probe, and where the second probe further comprises more than 20, 25, 30, 35, 40, 45, or 50 nucleotides that do not overlap with the first probe. In some embodiments, each of the genomic regions comprises at least five methylation sites, and where the at least five methylation sites have an abnormal methylation pattern in a cancerous sample, or have different methylation patterns between samples of different TOOs. For example, in one embodiment, at least five methylation sites are differentially methylated between a cancerous sample and a non-cancerous sample, or between one or more pairs of samples from cancers of different source tissues. In some embodiments, each probe pair comprises a first probe and a second probe, where the second probe is different from the first probe. The second probe can overlap the first probe by an overlapping sequence that is at least 30, at least 40, at least 50, or at least 60 nucleotides in length.

[0131] In some embodiments, the length of each of the plurality of different decoy oligonucleotides of the set is at least 45 nucleotides (e.g., the length is at least 60, 75, 80, 90, 100, 110, or 120 nucleotides). In some embodiments, the length of each decoy oligonucleotide among the plurality of probes does not exceed 130, 140, 150, 200, 250, or 300 bases. In some embodiments, the length of each decoy oligonucleotide is 45 to 300, 60 to 200, or 75 to 150 nucleotides. In some embodiments, the length of the decoy oligonucleotide is at least 50 nucleotides. In some embodiments, the length of the decoy oligonucleotide is at least 60 nucleotides. In some embodiments, the length of the decoy oligonucleotide is at least 75 nucleotides. In some embodiments, the indicated length of the decoy oligonucleotide is designed to be complementary to a portion of the target genomic sequence or its transformed DNA molecule.

[0132] In some embodiments, the cancer assay set is designed to target at least 5000, 10000, 12500, 15000, 17000, 19000 or more target genomic regions. In some embodiments, the cancer assay set is designed to target fewer than 25000, 20000, 17000, 15000, 12500, 10000 or fewer target genomic regions. In some embodiments, the cancer assay set is designed to target 5000 to 30000, 10000 to 25000, 12500 to 20000 or 15000 to 20000 target genomic regions. In some embodiments, the cancer assay set is designed to target at least 10000 target genomic regions. In some embodiments, the cancer assay set is designed to target at least 15000 target genomic regions. In some embodiments, the cancer assay set is designed to target fewer than 20000 target genomic regions.

[0133] In some embodiments, the bait oligonucleotides are configured to hybridize to transformed DNA molecules (e.g., transformed cfDNA molecules) corresponding to or derived from one or more genomic regions. Thus, the bait oligonucleotides can have a sequence different from the target genomic region. For example, DNA containing unmethylated CpG sites can be transformed to include UpG instead of CpG by deamination (e.g., by treatment with cytidine deaminase or bisulfite). Thus, probes directed to such targets can be configured to hybridize to sequences including UpG instead of the naturally occurring unmethylated CpG. Thus, the sites in the probe complementary to the unmethylated sites can contain CpA instead of CpG, and some probes targeting hypomethylated sites where all methylated sites are unmethylated may not have a guanine (G) base. In some embodiments, at least 3%, 5%, 10%, 15%, or 20% of the probes do not contain CpG sequences. In some embodiments, at least 5% of the probes do not contain CpG sequences. In some embodiments, at least 10% of the probes do not contain CpG sequences.

[0134] In some embodiments, a cancer assay set is used to detect the presence or absence of cancer and / or provide cancer classification (e.g., cancer type), cancer stage (e.g., I, II, III, or IV), or provide a TOO considered to be the origin of the cancer. The set can include probes that target differentially methylated genomic regions between generally cancerous (pan-cancer) samples and non-cancerous samples or only in cancerous samples with a specific cancer type (e.g., bladder cancer-specific targets). For example, in some embodiments, the cancer assay set is designed to include differentially methylated genomic regions based on transformed (e.g., bisulfite) sequencing data generated from cfDNA and / or whole-genome DNA from sets of cancer and non-cancer individuals.

[0135] In some embodiments, each of the target genomic regions is differentially methylated in at least one of a plurality of cancer types. In some embodiments, the plurality of cancer types includes at least 2 cancer types (e.g., at least 2, 3, 3, 4, or more cancer types). In some embodiments, the plurality of cancer types includes at least 2 cancer types. In some embodiments, the plurality of cancer types includes at least 3 cancer types. In some embodiments, the plurality of cancer types includes urinary system cancers. In some embodiments, the plurality of cancer types includes one or more of bladder cancer, urothelial cancer, kidney cancer, or prostate cancer.

[0136] Each of the probes (or probe pairs) can be designed to target one or more target genomic regions. The target genomic regions can be selected based on several criteria aimed at enhancing the selective enrichment of informative cfDNA fragments (e.g., cfDNA fragments from urine samples) while reducing noise and non-specific binding. Various filtering procedures for determining whether to include a target genomic region are described herein. In some embodiments, two or more of the filtering procedures described herein are used in combination.

[0137] In one example, a set can include probes that can selectively bind and enrich differentially methylated cfDNA fragments in a cancerous sample. In such a case, sequencing the enriched fragments can provide information relevant to cancer detection. Additionally, in some embodiments, the probes (or portions thereof) are designed to target genomic regions determined to have abnormal or aberrant methylation patterns in cancer samples or samples from certain cancer types, tissue types, or cell types. In one embodiment, the probes are designed to target genomic regions determined to be hypermethylated or hypomethylated in certain cancers or cancer types to provide additional assay selectivity and specificity. In some embodiments, the set contains probes targeting hypomethylated fragments. In some embodiments, the set contains probes targeting hypermethylated fragments. In some embodiments, the set contains a first set of probes targeting hypermethylated fragments and a second set of probes targeting hypomethylated fragments. In some embodiments, a cancer assay set includes not only probes designed to target regions having a first methylation state (e.g., hypomethylation), but also probes designed to hybridize to the same target region having the opposite methylation state (e.g., hypermethylation). The targeting of hypomethylated and hypermethylated fragments from the same region by probe pairs can be referred to as "dual" targeting (see, e.g., Figure 11C ). In some embodiments, the ratio (hyper:hypo ratio) between the first set of probes targeting hypermethylated fragments and the second set of probes targeting hypomethylated fragments ranges between 0.4 and 2, between 0.5 and 1.5, between 0.5 and 1.0, between 0.8 and 1, between 0.6 and 0.8, or between 0.4 and 0.6. In some embodiments, the hyper:hypo ratio is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, or higher. In some embodiments, the hyper:hypo ratio is at least 5. Methods for identifying genomic regions (e.g., genomic regions yielding differentially methylated DNA molecules or aberrantly methylated DNA molecules) between cancer samples and non-cancer samples, between different cancer source tissue (TOO) types, between different cancer cell types, or between samples from different cancer stages are provided in detail herein, and methods for identifying aberrantly methylated DNA molecules or fragments identified as indicative of cancer are also provided in detail herein.

[0138] In a second instance, a genomic region may be selected when it gives rise to aberrantly methylated DNA molecules in a cancer sample or a sample having a known cancer-derived tissue of origin (TOO) type. For example, as described herein, a Markov model trained on a set of non-cancerous samples can be used to identify genomic regions that give rise to aberrantly methylated DNA molecules (e.g., DNA molecules having a methylation pattern below a p-value threshold).

[0139] In some embodiments, each of the probes targets a genomic region comprising at least 30 bp, 35 bp, 40 bp, 45 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 100 bp, 110 bp, 120 bp or more. In some embodiments, each of the probes targets a genomic region comprising 120 bp. In some embodiments, a genomic region may be selected to have fewer than 30, 25, 20, 15, 12, 10, 8 or 6 methylation sites. In some embodiments, a genomic region may be selected to have at least 3, 5 or 7 methylation sites. In some embodiments, each target genomic region comprises at least five methylation sites. In some embodiments, the selected number of methylation sites (e.g., at least 5 methylation sites) are methylation sites that are differentially methylated in at least one type of cancer to be assayed by the set.

[0140] In some embodiments, regions are selected as targets when at least 80%, 85%, 90%, 92%, 95% or 98% of at least five methylation (e.g., CpG) sites within the genomic region are methylated or unmethylated in a non-cancerous sample or a cancerous sample, or in a cancer sample from a tissue of origin (TOO).

[0141] In some embodiments, target genomic regions are filtered to select only those that are likely to provide information based on their methylation patterns, e.g., CpG sites that are differentially methylated between cancerous and non-cancerous samples (e.g., abnormally methylated or unmethylated in cancer compared to non-cancer), CpG sites that are differentially methylated between cancerous samples of one TOO and cancerous samples of a different TOO, CpG sites that are differentially methylated only in cancerous samples of a TOO. For selection, calculations can be made relative to each CpG site or multiple CpG sites. For example, a first count, i.e., the number of cancerous samples (cancer_count) that include fragments overlapping with the CpG, can be determined, and a second count, i.e., the total number of samples (total) that contain fragments overlapping with the CpG site, can be determined. Genomic regions can be selected based on the following criteria: being positively correlated with the number of cancerous samples (cancer_count) that include fragments indicating cancer overlapping with the CpG site; and being negatively correlated with the total number of samples (total) that contain fragments indicating cancer overlapping with the CpG site. In one embodiment, the number of non-cancerous samples (n_non-cancer) and the number of cancerous samples (n_cancer) having fragments overlapping with the CpG site are counted. Then the probability that a sample is cancer is estimated, e.g., (n_cancer + 1) / (n_cancer + n_non-cancer + 2). This principle can be similarly applied to other outcomes.

[0142] CpG sites scored by this metric can be sorted and greedily added to the set until the set size budget is exhausted. The process of selecting genomic regions indicative of cancer is further detailed herein. In some embodiments, different target regions can be selected depending on whether the assay is intended to be a pan-cancer assay or a single-cancer assay, or according to the kind of flexibility desired in selecting which CpG sites contribute to the set. Similar methods can be used to design sets for detecting specific cancer types. In this embodiment, for each cancer type and each CpG site, the information gain is calculated to determine whether to include a probe targeting that CpG site. The information gain can be calculated for samples with a given TOO cancer type compared to all other samples. For example, consider two random variables "AF" and "CT". "AF" is a binary variable that indicates whether there is an abnormal fragment overlapping a specific CpG site in a particular sample (yes or no). "CT" is a binary random variable that indicates whether the cancer belongs to a specific type (e.g., bladder cancer or cancer other than bladder cancer). The mutual information with respect to "CT" given "AF" can be calculated. That is, how much information about the cancer type (bladder versus non-bladder in the example) can be obtained if it is known whether there is an anomalous fragment overlapping a specific CpG site. This can be used to rank the CpGs according to their bladder specificity. This procedure is repeated for multiple cancer types. If a particular region is commonly differentially methylated only in bladder cancer (and not in other cancer types or non-cancer), the CpGs in that region will tend to have a high information gain for bladder cancer. For each cancer type, the CpG sites are ranked by this information gain metric and then greedily added to the set until the size budget for that cancer type is exhausted. In some embodiments, the information gain is also measured according to the likelihood of pulling down molecules containing the target sequence and / or the cost of sequencing the target. In some embodiments, the target region is included only if the ratio of the information gain to the cost of sequencing the target is higher than a threshold.

[0143] In some embodiments, filtering is performed to select probes that have high specificity (i.e., high binding efficiency) for enrichment of nucleic acids derived from a target genomic region. Probes can be filtered to reduce non-specific binding (or off-target binding) to nucleic acids derived from non-target genomic regions. For example, probes can be filtered to select only those probes that have fewer than a set threshold of off-target binding events (e.g., predicted by the number of potential binding sites in the genome with a specific degree of complementarity across a given alignment window). In one embodiment, probes can be aligned to a reference genome (e.g., a human reference genome) to select probes that align to fewer than a region set threshold across the genome. In some embodiments, the reference genome is the human reference genome GRCh37 / hg19, the sequence of which is available from the Genome Reference Consortium and the genome browser provided by the Santa Cruz Genomics Institute. For example, probes can be selected that align to fewer than 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, or 8 off-target regions across the reference genome. In other cases, when the sequence of a target genomic region appears more than 5, 10, 15, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 times in the genome, filtering is performed to remove the genomic region. Further filtering can be performed to select a target genomic region when a probe sequence or set of probe sequences that is 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% complementary to the target genomic region appears fewer than 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, or 8 times in the reference genome, or to remove the target genomic region when a probe sequence or set of probe sequences designed to enrich the target genomic region is 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% complementary to the target genomic region and appears more than 5, 10, 15, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 times in the reference genome. This is to exclude repetitive probes that can pull down unwanted off-target fragments and can affect assay efficiency.

[0144] In some experiments, a fragment-probe overlap of at least 45 bp has proven effective for achieving non-negligible pull-down amounts (although this number can be large). In some cases, a mismatch rate of more than 10% between the probe and the fragment sequence in the overlapping region is sufficient to greatly disrupt binding and thus disrupt pull-down efficiency. Thus, sequences that can align with the probe along at least 45 bp and have a match rate of at least 90% can be candidates for off-target pull-down. Accordingly, in one embodiment, the number of such regions is scored. In some embodiments, optimal probes have a score of 1, meaning they match at only one position (the expected target region). Probes with a medium score (e.g., less than 5 or 10) may be acceptable in some cases, while in some cases, any probe with a score higher than a particular score is discarded. In some embodiments, if a candidate probe contains a sequence of 45 consecutive bp and has at least 90% complementarity with more than 20 off-target sites, the candidate probe will be excluded from the assay set. Other cut-off values can be used for a particular sample.

[0145] Once the probes hybridize and capture DNA fragments corresponding to or derived from the target genomic region, the hybridized probe-DNA fragment intermediates are pulled down (or separated), and the target DNA is amplified and sequenced. The sequence reads provide information relevant to cancer detection. To this end, a set can be designed to include multiple probes that can capture fragments that together provide information relevant to cancer detection. In some embodiments, the set includes at least 10, 20, 30, 40, 50, 60, 80, 100, 150, 200, 300, 500, 1,000, 2,000, 2,500, 5,000, 6,000, 7,500, 10,000, 15,000, 20,000, 25,000, or 30,000 different probe pairs. In some embodiments, the set includes at least 20, 30, 40, 60, 80, 100, 120, 160, 200, 300, 400, 600, 1,000, 2,000, 5,000, 6,000, 10,000, 12,000, 15,000, 20,000, 30,000, 40,000, 50,000, or 60,000 different probes. The multiple different probes together can comprise at least 2,000, 5,000, 10,000, 20,000, 40,000, 60,000, 80,000, 100,000, 200,000, 400,000, 600,000, 800,000, 1,000,000, 2,000,000, 3,000,000, 4,000,000, or 5,000,000 nucleotides. In some embodiments, the set includes no more than 20, 30, 40, 50, 60, 80, 100, 150, 200, 300, 500, 1,000, 2,000, 2,500, 5,000, 6,000, 7,500, 10,000, 15,000, 20,000, or 25,000 or 30,000 or 40,000 different probe pairs. In other embodiments, the set includes no more than 30, 40, 60, 80, 100, 120, 160, 200, 300, 400, 600, 1,000, 2,000, 5,000, 6,000, 10,000, 12,000, 15,000, 20,000, 30,000, 40,000, 50,000, 60,000, or 70,000 different probes.

[0146] The selected target genomic regions can be located at different positions in the genome, including but not limited to promoters, enhancers, exons, introns, intergenic regions, and other portions. In some cases, primers can be used to specifically amplify the target biomarker (e.g., by PCR) to enrich the desired target / biomarker in the sample (optionally without hybridization capture). For example, forward and reverse primers can be prepared for each target genomic region and used to amplify a fragment corresponding to or derived from the desired genomic region. Thus, while the present disclosure particularly focuses on the cancer assay sets and bait sets for hybridization capture, the present disclosure is broad enough to cover other methods for enriching cell-free nucleic acid molecules (e.g., cfDNA). Thus, those skilled in the art, upon benefiting from the present disclosure, will recognize that methods similar to those described herein for hybridization capture can alternatively be achieved by replacing hybridization capture with some other enrichment strategy, such as PCR amplification of cell-free DNA fragments corresponding to the target genomic region of interest. In some embodiments, bisulfite padlock probe capture is used to enrich the target region, as described in Zhang et al. (US2016 / 0340740). In some embodiments, additional or alternative methods are used for enrichment (e.g., non-targeted enrichment), such as reduced representation bisulfite sequencing, methylation-restriction enzyme sequencing, methylated DNA immunoprecipitation sequencing, methyl-CpG binding domain protein sequencing, methylated DNA capture sequencing, or droplet PCR.

[0147] probe

[0148] The cancer assay sets provided herein (alternatively referred to as "bait sets") can be sets that include a set of hybridization probes (also referred to herein as "probes") that are designed to target and pull down the nucleic acid fragments of interest for the assay during enrichment. In some embodiments, the probes are designed to hybridize and enrich DNA or cfDNA molecules from cancerous samples that have been processed to convert unmethylated cytosine (C) to uracil (U). In other embodiments, the probes are designed to hybridize and enrich DNA or cfDNA molecules from TOO (or TOOs) cancerous samples that have been processed to convert unmethylated cytosine (C) to uracil (U). These probes can be designed to anneal (or hybridize) to the target (complementary) strand of DNA or RNA. The target strand can be the "plus" strand (e.g., the strand that is transcribed into mRNA and subsequently translated into protein) or the complementary "minus" strand. In certain embodiments, the cancer assay set can include a set having two probes, one probe targeting the plus strand of the target genomic region and the other probe targeting the minus strand of the target genomic region.

[0149] For each target genomic region, at least four possible probe sequences can be designed. Each target region is double-stranded, and thus, the probe or probe set can target the "plus" or forward strand or its reverse complementary strand ("minus" strand). Additionally, in some embodiments, the probe or probe set is designed to enrich DNA molecules or fragments that have been treated to convert unmethylated cytosine (C) to uracil (U). For a probe or probe set designed to enrich DNA molecules corresponding to or derived from the target region after conversion, the probe sequence can be designed to enrich DNA molecules of fragments in which unmethylated C has been converted to U (by substituting A for G at unmethylated cytosine sites in the DNA molecule or fragment corresponding to or derived from the target region). In one embodiment, the probe is designed to bind or hybridize to DNA molecules or fragments from genomic regions known to contain cancer-specific methylation patterns (e.g., hypermethylated or hypomethylated DNA molecules), thereby enriching cancer-specific DNA molecules or fragments. Targeting genomic regions or cancer-specific methylation patterns can be advantageous, allowing for the specific enrichment of DNA molecules or fragments identified as informative for cancer or cancer TOO, thus reducing sequencing requirements and costs. In other embodiments, two probe sequences (one for each DNA strand) can be designed for each target genomic region. In still other cases, the probe is designed to enrich all DNA molecules or fragments corresponding to or derived from the target region (i.e., regardless of strand or methylation status). This may be because the cancer methylation status is not highly methylated or unmethylated, or because the probe is designed to target small mutations or other variants rather than methylation changes, and these other variants similarly indicate the presence or absence of cancer or the presence or absence of cancer in one or more TOO. In this case, each target genomic region can include all four possible probe sequences.

[0150] In some embodiments, the probe length ranges from dozens, hundreds, more than 200, or more than 300 base pairs. The probe can comprise at least 50, 75, 100, or 120 nucleotides. The probe can comprise less than 300, 250, 200, or 150 nucleotides. In an embodiment, the probe comprises 100 - 150 nucleotides. In a specific embodiment, the probe comprises 120 nucleotides.

[0151] In some embodiments, the probes are designed in a "2× tiling" manner to cover overlapping portions of the target region. The coverage of each probe optionally at least partially overlaps with another probe in the library. In such embodiments, the set contains multiple probe pairs, where each probe in the pair overlaps with the other by at least 25, 30, 35, 40, 45, 50, 60, 70, 75, or 100 nucleotides. In some embodiments, the overlapping sequences can be designed to be complementary to the target genomic region (or cfDNA derived therefrom) or to a sequence that is homologous to the target region or cfDNA. Thus, in some embodiments, at least two probes contain sequences that are complementary to the same sequence within the target genomic region, and a nucleotide fragment corresponding to or derived from the target genomic region can be bound and captured by at least one of the probes. For a given probe pair that contains overlapping sequences, the pair can contain non-overlapping sequences that extend from different ends of the overlapping sequence and are complementary to the target genomic region (see, for example, Figure 11E ). Other tiling levels are possible, such as 3× tiling, 4× tiling, etc., where each nucleotide in the target region can bind to more than two probes.

[0152] In one embodiment, a single base in the target genomic region is exactly overlapped by two probes, as shown in Figure 11B . Probes that extend bidirectionally beyond the target genomic region can be used to pull down cfDNA fragments that contain a portion of the target genomic region and DNA sequences adjacent to the target genomic region. In some cases, even relatively small target regions can be targeted with three probes (see, for example, Figure 11A ). Probe sets that contain three or more probes are optionally used to capture larger genomic regions (see, for example, Figure 11B ). In some embodiments, a subset of the probes will collectively extend across the entire target genomic region (e.g., can be complementary to untransformed or transformed fragments in the entire genomic region). The tiling probe set optionally contains probes that collectively include at least two probes that overlap with each nucleotide in the target genomic region. This is done to ensure that cfDNA that contains a small portion of the target genomic region at one end will have substantial overlap with at least one probe (the overlap extending into the adjacent non-target genomic region) to provide efficient capture.

[0153] In some embodiments, each target genomic region is targeted by a probe set. The probe set can be designed in a tiling manner such that adjacent probes have overlapping sequences that hybridize to the same part of the genomic region (see Figure 11D ). Since DNA has two strands, the probe set can also include overlapping probes that hybridize to the other strand, for a total of four probes that hybridize to the same part of the genomic region.

[0154] In some embodiments, the probe set configured to hybridize to a target genomic region does not span the entire region - i.e., at least some sequences within the target genomic region do not have corresponding probes. For example, a sequence within a target genomic region may be similar or identical to many other sequences in the genome, and no probe is designed to target that sequence because such a probe would hybridize to more than a threshold number of off-target regions.

[0155] For example, a 100bp cfDNA fragment containing a 30nt target genomic region will have at least 65bp overlap with at least one of the overlapping probes. Other tiling levels are possible. For example, to increase the target size and add more probes to the set, the probes can be designed to extend a 30bp target region by at least 70bp, 65bp, 60bp, 55bp, or 50bp. To fully capture any fragment overlapping the target region (even if only 1bp), the probes can be designed to extend beyond the ends of the target region on either side, e.g., by at least 50bp, 55bp, 60bp, 65bp, 70bp, 75bp, 80bp, or 85bp. The probes can be designed to extend 75bp beyond the ends of the target region on either side. In some embodiments, the presence of probes designed to extend beyond the ends of the target genomic region does not increase the size of the target genomic region (e.g., is not included in determining the size of the respective target genomic region or the combined size of multiple target genomic regions).

[0156] In some embodiments, each bait oligonucleotide is conjugated to a solid surface (e.g., a chip or bead, such as a magnetic or paramagnetic bead) or to a non-nucleotide affinity moiety (e.g., a member of a binding pair). In some embodiments, such conjugation is used to facilitate separation of DNA molecules bound to the bait oligonucleotide from unbound DNA molecules. Generally, a "binding pair" refers to one of a first part and a second part, where the first part and the second part have specific binding affinity for each other. Non-limiting examples of binding pairs include antigen / antibody; biotin / avidin (or biotin / streptavidin); calmodulin-binding protein (CBP) / calmodulin; hormone / hormone receptor; lectin / carbohydrate; peptide / cell membrane receptor; enzyme / cofactor; and enzyme / substrate. In some embodiments, the affinity moiety is biotin.

[0157] In some embodiments, a target genomic region is selected such that the application of a trained classifier to the sequences of transformed DNA molecules captured by bait oligonucleotides discriminates between a subject having cancer and a subject not having cancer with defined specificity and / or sensitivity. Methods for selecting target regions, classifier training, and selected specificities and sensitivities useful in detecting various cancer types are disclosed herein, such as methods related to various aspects of the methods herein. In some embodiments, the classifier is a binary classifier, a mixture model classifier, or a multi-layer perceptron model classifier. In some embodiments, the defined specificity for each of a plurality of cancer types is 0.900 or higher (e.g., at least 0.950, 0.975, 0.980, 0.985, 0.990, 0.995, or higher). In some embodiments, the application of the trained classifier comprises at least 30% (e.g., at least 40%, 50%, 60%, 70%, 80%, or higher) sensitivity for each of a plurality of cancer types.

[0158] In some embodiments, the target genomic region is selected from the gene sequences in Table 1. In some embodiments, the cancer assay set comprises a plurality of probes, wherein each of the plurality of probes is configured to hybridize to a DNA molecule (e.g., a transformed DNA molecule) or an amplicon thereof corresponding to one or more genomic regions containing target sequences of one or more genes from Table 1. In some embodiments, the cancer assay set comprises a plurality of probes, wherein each of the plurality of probes is configured to hybridize to a DNA molecule (e.g., a transformed DNA molecule) or an amplicon thereof corresponding to at least 1, 2, 3, 4, 5, 10, 15, 20, 25, or more target sequences selected from one or more genes of Table 1. In some embodiments, the target genomic region comprises at least 10 target sequences of one or more genes from Table 1. In some embodiments, the target genomic region comprises at least 15 target sequences of one or more genes from Table 1. The target sequences may comprise two or more target sequences from a single gene, and / or one or more target sequences from two or more genes. In some embodiments, the target genomic region comprises target sequences of genes selected from: TWIST1, EOMES, HOXA9, POU4F2, and ZNF154. In some embodiments, the target genomic region comprises a target sequence of the TWIST1 gene. In some embodiments, the target genomic region comprises target sequences of at least 20% of the genes from Table 1. In some embodiments, the target genomic region comprises target sequences of at least 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100% of the genes from Table 1. In some embodiments, the target genomic region comprises sequences of at least 50% of the genes from Table 1. In some embodiments, the target genomic region comprises target sequences of all genes from Table 1. In some embodiments, the length of the target sequence is at least 10, 25, 35, 45, 50, 75, 100, 200, 300, or more nucleotides. In some embodiments, the length of the target sequence is at least 25, at least 35, or at least 45 nucleotides. In some embodiments, the length of the target sequence is at least 45 nucleotides. In some embodiments, each target genomic region comprises at least five methylation sites. In some embodiments, the differential methylation comprises at least 80% of the CpG sites in the target genomic region that are methylated or unmethylated.

[0159] Method for selecting a target genomic region

[0160] In one aspect, the present disclosure provides methods for selecting target genomic regions for detecting cancer and / or TOO. In some embodiments, the target genomic regions exhibit abnormal / aberrant methylation in cancer and / or TOO. The target genomic regions can be used to design and manufacture probes for a cancer assay panel. The cancer assay panel can be used to screen the methylation status of DNA or cfDNA molecules corresponding to or derived from the target genomic regions. Alternative methods (e.g., by WGBS or other methods) can also be implemented to detect the methylation status of DNA molecules or fragments corresponding to or derived from the target genomic regions.

[0161] Sample Processing

[0162] Figure 16A is a flow chart of a process for processing a nucleic acid sample and generating a methylation status vector of DNA fragments according to one embodiment. The method includes but is not limited to the following steps. For example, any step of the method can include a quantification sub-step for quality control or other laboratory assay procedures known to those skilled in the art.

[0163] In step 105, a urine sample containing nucleic acids (e.g., cfDNA) is collected from a subject. The sample can be any subset of the human genome, including the whole genome. The urine sample can be prepared by any of the methods for collecting nucleic acids from urine disclosed herein. Figure 1 Shows an exemplary method for urine sample processing. In this particular embodiment, urine is collected from the subject, treated with a cell lysis inhibitor, and centrifuged to pellet and remove cell debris. The urine sample is then concentrated by filtration to obtain a concentrated urine sample that has an increased nucleic acid concentration and a reduced volume relative to the untreated urine sample. The concentrated urine sample can then be subjected to nucleic acid extraction and library preparation, and then sequenced for methylation analysis and cancer (e.g., bladder cancer, kidney cancer, or prostate cancer) detection. The extracted sample may contain cfDNA and / or ctDNA. In healthy individuals, the body can naturally clear cfDNA and other cell debris. If the subject has cancer or a disease, cfDNA and / or ctDNA in the sample may be present at detectable levels for detecting the cancer or disease.

[0164] In step 110, the nucleic acid is treated to convert unmethylated cytosine to uracil. In one embodiment, the method uses bisulfite treatment of DNA, which converts unmethylated cytosine to uracil without converting methylated cytosine. For example, commercially available kits can be used for bisulfite conversion, such as the EZ DNAMethylationTM-Gold, EZ DNA MethylationTM-Direct, or EZ DNA MethylationTM-Lightning kits (from Zymo Research Corporation (Irvine, CA)). In another embodiment, the conversion of unmethylated cytosine to uracil is accomplished by an enzymatic reaction. For example, the conversion can be performed using a commercially available kit for the conversion of unmethylated cytosine to uracil, such as APOBEC-Seq (Ipswich, MA, NEBiolabs).

[0165] In step 115, a sequencing library is prepared. In some embodiments, ssDNA adapters are added to the 3'-OH termini of bisulfite-converted ssDNA molecules using an ssDNA ligation reaction. In one embodiment, the ssDNA ligation reaction uses CircLigaseII (Epicentre) to ligate an ssDNA adapter to the 3'-OH terminus of a bisulfite-converted ssDNA molecule, where the 5' terminus of the adapter is phosphorylated and the bisulfite-converted ssDNA has been dephosphorylated (i.e., the 5' phosphate has been removed). In another embodiment, the ssDNA ligation reaction uses a thermostable 5’ AppDNA / RNA ligase (obtainable from New England BioLabs (Ipswich, Massachusetts)) to ligate an ssDNA adapter to the 3'-OH terminus of a bisulfite-converted ssDNA molecule. In this example, the first adapter is adenylated at the 5' terminus and blocked at the 3' terminus. In another embodiment, the ssDNA ligation reaction uses T4 RNA ligase (obtainable from New England BioLabs) to ligate an ssDNA adapter to the 3'-OH terminus of a bisulfite-converted ssDNA molecule. In the second step, a second strand of DNA is synthesized in an extension reaction. For example, an extension primer that hybridizes to a primer sequence included in the ssDNA adapter is used in a primer extension reaction to form a double-stranded bisulfite-converted DNA molecule. Optionally, in one embodiment, the extension reaction uses an enzyme capable of reading through uracil residues in the bisulfite-converted template strand. Optionally, in the third step, dsDNA adapters are added to the double-stranded bisulfite-converted DNA molecules. Finally, the double-stranded bisulfite-converted DNA is amplified to add sequencing adapters. For example, PCR amplification using a forward primer that includes a P5 sequence and a reverse primer that includes a P7 sequence is used to add the P5 and P7 sequences to the bisulfite-converted DNA. Optionally, during library preparation, unique molecular identifiers (UMIs) can be added to nucleic acid molecules (e.g., DNA molecules) by adapter ligation. A UMI is a short nucleic acid sequence (e.g., 4-10 base pairs) that is added to the end of a DNA fragment during adapter ligation. In some embodiments, the UMI contains degenerate base pair positions that serve as unique tags that can be used to identify sequence reads derived from a particular DNA fragment. During PCR amplification following adapter ligation, the UMI is replicated along with the attached DNA fragment, which provides a method for identifying sequence reads from the same original fragment in downstream analysis (by using the UMI alone, or in combination with a portion of the end sequence of the sample nucleic acid fragment (e.g., the first 2-10 nucleotides)).

[0166] In step 120, target DNA sequences can be enriched from a library. For example, this step is used in the case of performing a targeted capture assay on a sample. During enrichment, hybridization probes (also referred to herein as "probes" or "bait oligonucleotides") are used to target and optionally pull down nucleic acid fragments that can provide information about the presence or absence of cancer (or a disease), cancer status, or cancer classification (e.g., cancer type or tissue of origin). In some embodiments, the probes have specific characteristics described herein, such as characteristics related to various other aspects described herein. For a given workflow, the probes can be designed to anneal (or hybridize) to the target (complementary) strand of DNA (e.g., a transformed DNA molecule). The target strand can be the "plus" strand (e.g., the strand that is transcribed into mRNA and subsequently translated into protein) or the complementary "minus" strand. The length of the probes can range from tens, hundreds, or thousands of base pairs. Additionally, the probes can cover overlapping portions of the target region, as described herein.

[0167] After the hybridization step 120, the hybridized nucleic acid fragments are enriched (e.g., by capture or otherwise separated from unbound nucleic acids), and can also be amplified using PCR (enrichment 125). For example, the target sequences can be enriched to obtain enriched sequences that can subsequently be sequenced. Generally, a variety of methods can be used to isolate and enrich the target nucleic acids hybridized to the probes. For example, a biotin moiety can be added to the 5' end of the probes (i.e., biotinylation) to facilitate the separation of the target nucleic acids hybridized to the probes using a streptavidin-coated surface (e.g., streptavidin-coated beads).

[0168] In step 130, sequence reads are generated from the enriched nucleic acid fragments. Sequencing data can be obtained from the enriched DNA sequences by methods known in the art. For example, the methods can include next-generation sequencing (NGS) technologies, including synthesis technologies (Illumina), pyrosequencing (454 Life Sciences), ion semiconductor technology (Ion Torrent sequencing), single molecule real-time sequencing (Pacific Biosciences), ligation sequencing (SOLiD sequencing), nanopore sequencing (Oxford Nanopore Technologies), or paired-end sequencing. In some embodiments, massively parallel sequencing is performed using sequencing by synthesis with reversible dye terminators.

[0169] In step 140, a methylation status vector is generated from the sequence reads. To do this, the sequence reads are aligned to a reference genome. The reference genome helps to provide context as to where in the human genome a DNA fragment (e.g., cfDNA) originated. In a simplified example, the sequence reads are aligned such that three CpG sites correspond to CpG sites 23, 24, and 25 (these are arbitrary reference identifiers used for ease of description). After alignment, there is information about the methylation status of all CpG sites on the cfDNA fragment, as well as information about where these CpG sites map in the human genome. Using the methylation status and location information, a methylation status vector for the DNA fragment can be generated.

[0170] Alternative methods for detecting methylation patterns include using oligonucleotide microarrays and selective PCR amplification. Methods for detecting methylation patterns can include using oligonucleotide microarrays. The methods can include labeling cfDNA fragments with a label. The cfDNA fragments can be from an individual. The cfDNA fragments can be bisulfite-converted ssDNA molecules or converted cfDNA fragments. The label can be a fluorescent label. The fluorescent label can be a cyanine dye. The cyanine dye can be Cy3 or Cy5, DY547 or DY647. The oligonucleotide microarray for use in the detection can comprise at least 75, 150, 300 or 1000 bait oligonucleotide pairs. In some embodiments, each bait oligonucleotide pair comprises a first bait and a second bait. The first bait can comprise an overlapping sequence and a first non-overlapping sequence. The second bait can comprise an overlapping sequence and a second non-overlapping sequence. The overlapping sequence can comprise at least 30, 40, 50 or 60 nucleotides. The first non-overlapping sequence and the second non-overlapping sequence can comprise more than 30, 40, 50 or 60 nucleotides. The first bait and the second bait of each pair of bait oligonucleotides can be configured to hybridize to a converted cfDNA fragment derived from a genomic region comprising at least five methylation sites that are differentially methylated in cfDNA fragments from an individual with cancer compared to cfDNA fragments from an individual without cancer. The converted cfDNA fragment can be labeled with a first fluorescent label. In some embodiments, a reference cfDNA fragment is labeled with a second fluorescent label. The reference cfDNA fragment can be a cfDNA fragment having a known methylation pattern. The second fluorescent label can be a cyanine dye. The cyanine dye can be Cy3 or Cy5, DY547 or DY647. The first fluorescent label and the second fluorescent label can be different. In some embodiments, the methylation pattern is detected by contacting the labeled cfDNA fragments from the individual and the labeled reference cfDNA fragments with the oligonucleotide microarray. The methylation pattern can be determined by comparing the amounts of the first fluorescent label and the second fluorescent label associated with each array position on the microarray. Each array position can correspond to a pair of bait oligonucleotides. In some embodiments, the methylation pattern is associated with cancer or cancer type. The method can include applying a trained classifier to predict the likelihood of cancer. In some embodiments, applying the trained classifier to the sequence of the converted cfDNA molecule hybridized to the bait oligonucleotides can distinguish a subject with cancer or cancer type from a subject without cancer, wherein the specificity is 0.990 or 0.994 and the sensitivity is at least 40%, at least 45% or at least 50%.

[0171] Methods for detecting methylation patterns can include using selective PCR amplification. A composition includes a plurality of oligonucleotide pairs configured to hybridize to transformed DNA fragments derived from a target genomic region, wherein each oligonucleotide pair is configured for selective PCR amplification of a sequence. In some embodiments, the sequence for selective PCR amplification comprises transformed DNA derived from a hypermethylated target genomic region to produce a PCR product and cannot amplify transformed cfDNA derived from a hypomethylated target genomic region. In some embodiments, the sequence for selective PCR amplification comprises transformed DNA derived from a hypomethylated target genomic region to produce a PCR product and cannot amplify transformed cfDNA derived from a hypermethylated target genomic region. Each target genomic region can include at least five CpG sites. In some embodiments, the target genomic region includes at least 1, 2, 3, 4, 5, 10, 15, or more of the target genomic regions containing a target sequence from one or more genes selected from Table 1. In some embodiments, the target genomic region includes a target sequence of a gene selected from: TWIST1, EOMES, HOXA9, POU4F2, and ZNF154. Determining the likelihood of the presence or absence of cancer or a cancer type in a subject can include amplifying transformed cfDNA fragments from the subject with the oligonucleotide pairs configured for selective PCR amplification described herein, sequencing the captured DNA fragments, and applying a trained classifier to the DNA sequences to determine the likelihood. The sensitivity of the classifier for determining the likelihood of cancer or a cancer type can be at least 40%, and the specificity can be 0.990 or higher.

[0172] Data structure generation

[0173] Figure 12A FIG. 300 is a flow chart of a process for generating a healthy control group data structure according to the embodiment description. To create a healthy control group data structure, the analysis system obtains information related to the methylation status of a plurality of CpG sites on sequence reads derived from a plurality of DNA molecules or fragments from a plurality of healthy subjects. The methods provided herein for creating a healthy control group data structure can be similarly performed on subjects with cancer, subjects with TOO cancer, subjects with a known cancer type, or subjects with other known disease states. For example, via process 100, a methylation status vector is generated for each DNA molecule or fragment.

[0174] The analysis system subdivides the methylation state vector of each DNA fragment into a string of CpG sites. In one embodiment, the analysis system subdivides the methylation state vector such that the resulting strings are all less than a given length. For example, a methylation state vector of length 11 may be subdivided into strings of length less than or equal to 3, resulting in 9 strings of length 3, 10 strings of length 2, and 11 strings of length 1. In another example, a methylation state vector of length 7 is subdivided into strings of length less than or equal to 4, resulting in 4 strings of length 4, 5 strings of length 3, 6 strings of length 2, and 7 strings of length 1. If the methylation state vector generated from a DNA fragment is shorter than or the same as a particular string length, then the methylation state vector may be converted into a single string containing all the CpG sites of the vector.

[0175] The analysis system tallies 320 these strings by computing, for each possible CpG site and methylation state possibility in the vector, the number of strings in the control group that have the designated CpG site as the first CpG site in the string and have that methylation state possibility. For a string length of three at a given CpG site, there are 2^3 or 8 possible string configurations. For each CpG site, the analysis system tallies 320 the number of occurrences of each possible methylation state vector in the control group. This may involve tallying the following quantities: for each starting CpG site in the reference genome, <Mx,Mx+1,Mx+2>, <Mx,Mx+1,Ux+2>,..., <Ux,Ux+1,Ux+2>. The analysis system creates 330 a data structure that stores the statistical counts for each starting CpG site and the string possibilities at each starting CpG.

[0176] Setting an upper limit on the string length has several benefits. First, the size of the data structures created by the analysis system can increase dramatically depending on the maximum length of the string. For example, a maximum string length of 4 means that at most 2^4 numbers need to be counted at each CpG. Increasing the maximum string length to 5 doubles the number of possible methylation states that need to be counted. Reducing the string size helps to alleviate the computational and data storage burden on the data structures. In some embodiments, the string size is 3. In some embodiments, the string size is 4. A second reason for limiting the maximum string length is to avoid overfitting of downstream models. If long CpG strings do not have a strong biological effect on the outcome (e.g., an anomalous prediction for the presence of cancer), then the calculated probabilities based on long CpG site strings can be problematic because this requires a large amount of data that may not be available and is thus too sparse for the model to execute properly. For example, calculating the probability of an anomalous / cancer condition based on the first 100 CpG sites would require counting strings of length 100 in the data structure, and ideally, some of these would exactly match the methylation states of these first 100. If only sparse counts of strings of length 100 are available, there may not be enough data to determine whether a given string of length 100 in the test sample is anomalous.

[0177] Data structure validation

[0178] Once the data structure has been created, the analysis system may attempt to validate 340 the data structure and / or any downstream models that utilize the data structure.

[0179] A first type of validation ensures that potentially cancerous samples are removed from the healthy control group so as not to affect the purity of that control group. This type of validation checks for consistency within the control group data structure. For example, a healthy control group may contain samples from individuals with undiagnosed cancer that contain multiple anomalously methylated segments. The analysis system can perform various calculations to determine whether to exclude data from subjects with apparently undiagnosed cancer.

[0180] A second type of validation checks the probability model used to calculate p-values using the counts from the data structure itself (i.e., from the healthy control group). Jointly described below Figure 14The process of p-value calculation is described. Once the analysis system generates p-values for the methylation status vectors in the validation group, the analysis system constructs a cumulative density function (CDF) using these p-values. Using the CDF, the analysis system can perform various calculations on the CDF to verify the data structure of the control group. One test uses the fact that ideally, the CDF should be equal to or below the identity function, such that CDF(x) ≤ x. Conversely, revealing above the identity function indicates some flaws within the probability model for the control group data structure. For example, if 1 / 100 of the segments have a p-value score of 1 / 1000, meaning CDF(1 / 1000) = 1 / 100 > 1 / 1000, then the second validation type fails, indicating a problem with the probability model. See, for example, U.S. Application No. 16 / 352,602, published as U.S. Publication No. 2019 / 0287652, which is hereby incorporated by reference in its entirety.

[0181] The third validation type uses a set of healthy validation samples that are separate from the samples used to construct the data structure. This tests whether the data structure is correctly constructed and whether the model is working properly. Jointly described below Figure 12B is an exemplary process for performing this validation type. The third validation type can quantify how well the healthy control group generalizes the distribution of healthy samples. If the third validation type fails, then the healthy control group does not generalize well to the healthy distribution.

[0182] The fourth validation type is tested using samples from a non-healthy validation group. The analysis system calculates p-values and constructs a CDF for the non-healthy validation group. For the non-healthy validation group, the analysis system expects to see CDF(x) > x for at least some of the samples, or in other words, the opposite of what is expected in the second and third validation types for the healthy control group and healthy validation group. If the fourth validation type fails, it indicates that the model cannot appropriately identify the anomalies it is designed to identify.

[0183] Figure 12B is a flowchart of additional step 340 for validating Figure 12A the control group data structure according to an embodiment. In an embodiment of validation data structure step 340, the analysis system performs the fourth validation type test as described above, which utilizes a validation group having subjects, samples, and / or segments assumed to be similar to the control group. For example, if the analysis system selects healthy subjects without cancer for the control group, then the analysis system also uses healthy subjects without cancer in the validation group.

[0184] The analysis system adopts the validation group and generates a set of 100 methylation status vectors, as Figure 12A described. The analysis system performs p-value calculation on each methylation status vector from the validation group. The p-value calculation process will jointly Figures 13 - 14Further description. For each possible methylation state vector, the analysis system calculates probabilities from the data structure of the control group. Once the likelihood probabilities of the methylation state vectors are calculated, the analysis system calculates a 350 p-value score for that methylation state vector based on the calculated probabilities. The p-value score represents the expectancy of finding that particular methylation state vector and other methylation state vectors that may have even lower probabilities in the control group. Thus, a low p-value score generally corresponds to a methylation state vector that is relatively unexpected compared to other methylation state vectors in the control group, while a high p-value score generally corresponds to a methylation state vector that is relatively more expected compared to other methylation state vectors found in that control group. Once the analysis system generates p-value scores for the methylation state vectors in the validation group, the analysis system constructs a 360 cumulative density function (CDF) with the p-value scores from that validation group. The analysis system verifies the consistency of the CDF as described above in a fourth validation type test.

[0185] Aberrantly methylated fragment

[0186] According to Figure 13 the embodiments outlined in, aberrantly methylated fragments with abnormal methylation patterns are selected as target genomic regions in cancer patient samples, subjects with TOO cancer, subjects with known cancer types, or subjects with other known disease states. Figure 14 An exemplary process of the selected aberrantly methylated fragments 440 is visually shown in and is further described under the description of Figure 4 In process 400, the analysis system generates 100 methylation state vectors from the cfDNA fragments of the sample. The analysis system processes each methylation state vector as follows.

[0187] For a given methylation state vector, the analysis system enumerates 410 all possible methylation state vectors that have the same starting CpG site and the same length (i.e., set of CpG sites) as that given methylation state vector. Since each methylation state can be either methylated or unmethylated, there are only two possible states at each CpG site, so the number of different possibilities of the methylation state vector depends on the power of 2, such that a methylation state vector of length n will be associated with 2n methylation state vector possibilities.

[0188] The analysis system calculates 420 the probability of observing each methylation state vector possibility for the identified starting CpG site / methylation state vector length by accessing the healthy control group data structure. In one embodiment, calculating the probability of observing a given possibility uses Markov chain probability to model the joint probability calculation, which will be described below with respect to Figure 14described in more detail. In other embodiments, computational methods other than Markov chain probabilities are used to determine the probability of observing each methylation state vector possibility.

[0189] The analysis system uses the calculated probability of each possibility to calculate a 430p value score for the methylation state vector. In one embodiment, this includes identifying the calculated probability corresponding to the possibility that matches the methylation state vector under discussion. In particular, this is the possibility that has the same set of CpG sites as the methylation state vector or, similarly, has the same starting CpG site and length as the methylation state vector. The analysis system sums the calculated probabilities of any possibility whose probability is less than or equal to the identified probability to generate a p value score.

[0190] This p value represents the probability of observing the methylation state vector of the fragment or other more unlikely methylation state vectors in the healthy control group. Thus, a low p value score generally corresponds to a methylation state vector that is rare in healthy subjects and causes the fragment to be labeled as abnormally methylated relative to the healthy control group. A high p value score is generally associated with a methylation state vector that is expected to be present in healthy subjects in a relative sense. For example, if the healthy control group is a non-cancer group, a low p value indicates that the fragment is abnormally methylated relative to that non-cancer group and thus may indicate the presence of cancer in the test subject.

[0191] As described above, the analysis system calculates a p value score for each of the multiple methylation state vectors, where each methylation state vector represents a cfDNA fragment in the test sample. To identify which fragments are abnormally methylated, the analysis system may filter the set of 440 methylation state vectors based on their p value scores. In one embodiment, the filtering is performed by comparing the p value scores with a threshold and retaining only those fragments that are below the threshold. This threshold p value score can be on the order of 0.1, 0.01, 0.001, 0.0001, or similar.

[0192] P value score calculation

[0193] Figure 14 is an illustration 500 of an example p value score calculation according to an embodiment. To calculate the p value score for a given test methylation state vector 505, the analysis system takes the test methylation state vector 505 and enumerates the possibilities of 410 methylation state vectors. In this illustrative example, the test methylation state vector 505 is <M23,M24,M25,U26>. Since the length of the test methylation state vector 505 is 4, there are 2^4 possibilities for methylation state vectors covering CpG sites 23–26. In a general example, the number of possibilities for a methylation state vector is 2^n, where n is the length of the test methylation state vector or, alternatively, the length of the sliding window (described further below).

[0194] The analysis system calculates the probability 515 of 420 for the possibilities of the enumerated methylation state vectors. Since methylation is conditionally dependent on the methylation states of neighboring CpG sites, one way to calculate the probability of observing a given methylation state vector possibility is to use a Markov chain model. Generally, a methylation state vector such as <S1, S2, …, S n >, where S represents the methylation state, whether methylated (denoted as M), unmethylated (denoted as U), or undetermined (denoted as I), has a joint probability that can be expanded using the probability chain rule as follows:

[0195] P(<S1, S2,..., S n >) = P(S n |S1,..., S n-1 >)*P(S n-1 |S1,..., S n-2 >)*...*P(S2|S1)*P(S1) (1)

[0196] The Markov chain model can be used to make the calculation of the conditional probability of each possibility more efficient. In one embodiment, the analysis system selects a Markov chain order k, which corresponds to the number of previous CpG sites in the vector (or window) to be considered in the conditional probability calculation, such that the conditional probability is modeled as P(S n |S1,…,S n-1 ) ~ P(S n |S n-k-2 ,…,S n-1 ).

[0197] To calculate each Markov-modeled probability of the methylation state vector possibility, the analysis system accesses the data structure of the control group, particularly the counts of various CpG sites and state strings. To calculate P(M n |S n-k-2 ,..., S n-1 ), the analysis system employs the following ratio: the count of the number of strings stored from the data structure matching <S n-k-2 ,..., S n-1 , M n > is divided by the sum of the count of the number of strings stored from the data structure matching <S n-k-2 ,..., S n-1 , M n > and <S n-k-2 ,..., S n-1 , U n >. Thus, P(M n |S n-k-2 ,..., S n-1) is a calculated ratio having the following form:

[0198]

[0199] Count smoothing can also be achieved by applying a prior distribution. In one embodiment, the prior distribution is a uniform prior as in Laplace smoothing. As an example, a constant is added to the numerator of the above equation, and another constant (e.g., twice the constant in the numerator) is added to the denominator of the above equation. In other embodiments, algorithmic techniques such as Knesser-Ney smoothing are used.

[0200] In the illustration, the equation represented above is applied to the test methylation status vector 505 covering sites 23-26. Once the probability 515 is calculated, the analysis system will calculate the p-value score 525, which will be less than or equal to the sum of the probabilities of the methylation status vectors that match the test methylation status vector 505.

[0201] In one embodiment, the computational load of calculating the probability and / or the p-value score may be further reduced by caching at least some of the computations. For example, the analysis system may cache the probability calculations of the methylation status vector (or its window) likelihoods in transient or permanent memory. If other segments have the same CpG sites, caching the probabilities of these likelihoods allows for efficient calculation of the p-value score without having to recalculate the probabilities of the underlying likelihoods. Equivalently, the analysis system may calculate the p-value score for each likelihood of the methylation status vector associated with a set of CpG sites in the vector (or its window). The analysis system may cache these p-value scores for use in determining the p-value scores of other segments that include the same CpG sites. Overall, the p-value scores calculated for multiple likelihoods with the same CpG sites in the methylation status vector can be used to determine the p-value score of another one of these multiple likelihoods under the same set of CpG sites.

[0202] Sliding window

[0203] In one embodiment, the analysis system uses a 435 sliding window to determine the likelihood of the methylation status vector and calculate the p-value. The analysis system only needs to enumerate the likelihoods and calculate the p-value for windows that contain consecutive CpG sites, where the length of the window (in CpG sites) is shorter than at least some of the segments (otherwise, the window would be meaningless), rather than enumerating the likelihoods and calculating the p-value for all methylation status vectors. The window length may be static, user-determined, dynamic, or otherwise selected.

[0204] When calculating the p-value of a methylation status vector larger than the window, the window starts identifying a set of consecutive CpG sites within the vector from the first CpG site in the vector. The analysis system calculates a p-value score for the window including the first CpG site. Then, the analysis system "slides" the window to the second CpG site in the vector and calculates another p-value score for the second window. Thus, for a window size l and a methylation vector length m, each methylation status vector will generate m – l + 1 p-value scores. After completing the p-value calculation for each part of the vector, the lowest p-value score among all the sliding windows is taken as the overall p-value score for that methylation status vector. In another embodiment, the analysis system aggregates the p-value scores of the methylation status vector to generate an overall p-value score.

[0205] Using a sliding window helps reduce the number of methylation status vector possibilities that need to be enumerated and their corresponding probability calculations that would otherwise have to be performed. An example probability calculation is shown in Figure 14 but generally, the number of methylation status vector possibilities increases exponentially with a factor of 2 with the size of the methylation status vector. As a practical example, a fragment may have up to 54 CpG sites. The analysis system can use a window of size 5 (for example) and perform 50 p-value calculations for each of the 50 windows of the methylation status vector of the fragment, instead of calculating the probabilities of 2^54 (about 1.8 × 10^16) possibilities to generate a single p-value. Each of these 50 calculations enumerates 2^5 (32) methylation status vector possibilities, for a total of 50 × 2^5 (1.6 × 10^3) probability calculations. This greatly reduces the calculations that need to be performed, with no significant impact on the accurate identification of aberrant fragments. This additional step can also be applied when validating 340 control groups with the methylation status vectors of the validation group.

[0206] Identifying fragments indicative of cancer

[0207] The analysis system identifies 450 DNA fragments indicative of cancer from the filtered set of aberrant methylation fragments.

[0208] Hypomethylated and hypermethylated fragments

[0209] According to the first method, the analysis system can identify DNA fragments that are considered hypomethylated or hypermethylated from the filtered set of aberrant methylation fragments as fragments indicative of cancer. Hypomethylated and hypermethylated fragments can be defined as fragments having a certain number of CpG sites (e.g., more than 3, 4, 5, 6, 7, 8, 9, 10, etc.) of a certain length, where the percentage of methylated CpG sites is high (e.g., more than 80%, 85%, 90%, or 95%, or any other percentage within the range of 50%-100%), or the percentage of unmethylated CpG sites is high (e.g., more than 80%, 85%, 90%, or 95%, or any other percentage within the range of 50%-100%).

[0210] Probability model

[0211] According to the methods described herein, the analysis system uses a probability model suitable for the methylation patterns of each cancer type and non-cancer type to identify fragments indicative of cancer. The analysis system uses a fitted probability model for each cancer type and non-cancer type to consider various cancer types and calculates the log-likelihood ratio of the sample using DNA fragments in genomic regions. The analysis system can determine DNA fragments indicative of cancer based on whether at least one of the log-likelihood ratios considered for various cancer types is higher than a threshold.

[0212] In one embodiment of partitioning the genome, the analysis system partitions the genome into multiple regions in multiple stages. In the first stage, the analysis system divides the genome into CpG site blocks. Each block is defined when the interval between two adjacent CpG sites reaches and / or exceeds a certain threshold (e.g., greater than 200bp, 300bp, 400bp, 500bp, 600bp, 700bp, 800bp, 900bp, or 1,000bp). Starting from each block, in the second stage, the analysis system subdivides each block into regions of a certain length (e.g., 500bp, 600bp, 700bp, 800bp, 900bp, 1,000bp, 1,100bp, 1,200bp, 1,300bp, 1,400bp, or 1,500bp). The analysis system can further overlap adjacent regions by a percentage of length (e.g., 10%, 20%, 30%, 40%, 50%, or 60%, or 10% or more, 20% or more, 30% or more, 40% or more, 50% or more, or 60% or more).

[0213] The analysis system analyzes sequence reads of DNA fragments derived from each region. The analysis system can process samples from tissues and / or high-signal cfDNA. High-signal cfDNA samples can be determined by a binary classification model, by cancer stage, or by other metrics.

[0214] For each cancer type and non-cancer type, the analysis system fits a separate probability model to the segments. In one instance, each probability model is a mixture model that includes a combination of multiple mixture components, each of which is an independent site model, where it is assumed that methylation at each CpG site is independent of the methylation status at other CpG sites.

[0215] In an alternative embodiment, calculations are performed with respect to each CpG site. Specifically, a first count is determined, which is the number of cancerous samples that include an aberrantly methylated DNA segment that overlaps the CpG (cancer_count), and a second count is determined, which is the total number of samples in the set that contain a segment that overlaps the CpG (total_count). Genomic regions can be selected based on these numbers, for example, based on criteria that are positively correlated with the number of cancerous samples that include an aberrantly methylated DNA segment that overlaps the CpG (cancer_count), and negatively correlated with the total number of samples in the set that contain a segment that overlaps the CpG (total_count).

[0216] In some embodiments, the various types of cancer having different TOOs are selected from bladder cancer, urothelial cancer, kidney cancer, and prostate cancer. In some embodiments, the cancer is bladder cancer. In some embodiments, the cancer is urothelial cancer. In some embodiments, the cancer is kidney cancer. In some embodiments, the cancer is prostate cancer.

[0217] In some embodiments, the various cancer types can be classified and labeled using classification methods available in the art, such as the International Classification of Diseases for Oncology, 3rd Edition (ICD-O-3) (codes.iarc.fr) or the Surveillance, Epidemiology, and End Results Program (SEER) (seer.cancer.gov). In other embodiments, the cancer types are divided into three orthogonal codes: (i) anatomic site code, (ii) morphology code, or (iii) behavior code. Under the behavior code, a benign tumor is 0, uncertain behavior is 1, carcinoma in situ is 2, malignant primary site is 3, and malignant metastatic site is 6.

[0218] In some embodiments, the cancer TOO can be selected from the groups defined by a guide that will be used to stage the detected cancer. For example, the reference Amin, M.B., Edge, S., Greene, F., Byrd, D.R., Brookland, R.K., Washington, M.K., Gershenwald, J.E., Compton, C.C., Hess, K.R., Sullivan, D.C., Jessup, J.M., Brierley, J.D., Gaspar, L.E., Schilsky, R.L., Balch, C.M., Winchester, D.P., Asare, E.A., Madera, M., Gress, D.M., Meyer, L.R. (editors), AJCC Cancer Staging Manual [AJCC Cancer Staging Manual], 8th edition, Springer, 2017, identifies groups of different cancers that are staged together according to standard guidelines. Staging is typically the next step in cancer management after detection and diagnosis.

[0219] The analysis system can further calculate the log-likelihood ratio ("R") of the fragment for various cancer types using a fitted probability model for each cancer type and non-cancer type or for the cancer TOO, which log-likelihood ratio indicates the likelihood that the fragment indicates cancer. These two probabilities can be obtained from the probability models fitted for each of the cancer type and non-cancer type, which probability models are defined to calculate the likelihood of observing the methylation pattern on the fragment for each of the cancer type and non-cancer type. For example, probability models can be defined for each of the cancer type and non-cancer type.

[0220] Selection of genomic regions indicative of cancer

[0221] In some embodiments, the analysis system can identify 460 genomic regions indicative of cancer. To identify these informative regions, the analysis system calculates the information gain for each genomic region or more specifically for each CpG site, which information gain describes the ability to distinguish various outcomes.

[0222] The method for identifying genomic regions capable of distinguishing cancer types and non-cancer types utilizes a trained classification model that can be applied to a set of aberrantly methylated DNA molecules or fragments corresponding to or derived from cancerous or non-cancerous groups. The trained classification model can be trained to identify any target condition that can be identified from the methylation state vector.

[0223] In one embodiment, the trained classification model is a binary classifier (the binary classifier is trained based on the methylation status of cfDNA fragments or genomic sequences obtained from a cohort of subjects with cancer or cancer TOO and a cohort of healthy subjects without cancer), and is then used to classify the probability that a test subject has cancer, cancer TOO, or does not have cancer based on an anomalous methylation status vector. In other embodiments, different classifiers can be trained using the following subject cohorts: known to have a specific cancer (e.g., bladder cancer, urothelial cancer, prostate cancer, kidney cancer, etc.); known to have a cancer with a specific TOO considered to be the origin of the cancer; or known to have different stages of a specific cancer (e.g., bladder cancer, urothelial cancer, prostate cancer, kidney cancer, etc.). In these embodiments, sequence reads obtained from samples enriched with tumor cells from a cohort of subjects known to have a specific cancer (e.g., bladder cancer, urothelial cancer, prostate cancer, kidney cancer, etc.) can be used to train different classifiers. The ability of each genomic region in the classification model to distinguish between cancer types and non-cancer types is used to rank the genomic regions from most informative to least informative in terms of classification performance. The analysis system can identify genomic regions from the ranking based on the information gain in the classification between non-cancer types and cancer types.

[0224] Calculate the information gain of fragments indicating hypomethylation and hypermethylation of cancer

[0225] According to an embodiment, using fragments indicating cancer, the analysis system can, according to Figure 15A the process 600 shown in. The process 600 accesses two sample training groups - a non-cancer group and a cancer group - and obtains 605 a set of non-cancer methylation status vectors and a set of cancer methylation status vectors (including anomalous methylated fragments), for example via step 440 in process 400.

[0226] For each methylation status vector, the analysis system determines 610 whether the methylation status vector indicates cancer. Here, fragments indicating cancer can be defined as hypomethylated or hypermethylated fragments determined as follows: whether at least a certain number of CpG sites have a specific status (methylated or unmethylated respectively) and / or a threshold percentage of sites with a specific status (again, methylated or unmethylated respectively). In one example, if a fragment overlaps with at least 5 CpG sites and at least 80%, 90%, or 100% of its CpG sites are methylated, or at least 80%, 90%, or 100% are unmethylated, the cfDNA fragment is identified as hypomethylated or hypermethylated respectively.

[0227] In an alternative embodiment, the analysis system considers portions of the methylation state vectors and determines whether the portions are hypomethylated or hypermethylated, and can distinguish whether the portions are hypomethylated or hypermethylated. This alternative addresses missing methylation state vectors that, while large in size, contain at least one dense hypomethylated or hypermethylated region. This process of defining hypomethylation and hypermethylation can be applied to Figure 13 step 450. In another embodiment, a segment indicative of cancer can be defined based on the likelihood output by a trained probability model.

[0228] In one embodiment, the analysis system generates a hypomethylation score (P low) and a hypermethylation score (P high) for each CpG site in the genome 620. To generate either score at a given CpG site, the classifier makes four counts at that CpG site — (1) the (methylation state) vector count of the cancer set marked as hypomethylated that overlaps the CpG site; (2) the vector count of the cancer set marked as hypermethylated that overlaps the CpG site; (3) the vector count of the non-cancer set marked as hypomethylated that overlaps the CpG site; and (4) the vector count of the non-cancer set marked as hypermethylated that overlaps the CpG site. Additionally, the process can normalize these counts for each group to account for differences in group size between the non-cancer and cancer groups. In an alternative embodiment, where segments indicative of cancer are more common, the scores can be more broadly defined as the count of segments indicative of cancer at each genomic region and / or CpG site.

[0229] In one embodiment, to generate the hypomethylation score at a given CpG site 620, the process takes the ratio of (1) compared to the sum of (1) and (3). Similarly, the hypermethylation score is calculated by taking the ratio of (2) compared to the sum of (2) and (4). Additionally, these ratios can be calculated with additional smoothing techniques as discussed above. The hypomethylation score and hypermethylation score are related to the estimation of cancer probability, assuming hypomethylation or hypermethylation of segments from the cancer set.

[0230] The analysis system generates 630 an overall hypomethylation score and an overall hypermethylation score for each aberrant methylation state vector. The overall hypermethylation score and hypomethylation score are determined based on the hypermethylation scores and hypomethylation scores of the CpG sites in the methylation state vector. In one embodiment, the overall hypermethylation score and hypomethylation score are respectively designated as the maximum hypermethylation score and hypomethylation score of the sites in each state vector. However, in an alternative embodiment, the overall scores can be based on the mean, median, or other calculations that use the hyper / hypomethylation scores of the sites in each vector.

[0231] The analysis system sorts all of the methylation status vectors of these subjects 640 according to the overall hypomethylation score of the subject and the overall hypermethylation score of the subject, thereby generating two sorts for each subject. The process selects the overall hypomethylation score from the hypomethylation sort and the overall hypermethylation score from the hypermethylation sort. Using the selected scores, the classifier generates 650 a single feature vector for each subject. In one embodiment, the scores selected from either sort are selected in a fixed order that is the same for each generated feature vector of each subject in each of the training groups. For example, in one embodiment, the classifier selects the first, second, fourth, and eighth overall hypermethylation scores from each sort, and similarly selects each overall hypomethylation score, and writes these scores into the feature vector of the subject.

[0232] The analysis system trains 660 a binary classifier to distinguish the feature vectors between the cancer and non-cancer training groups. Generally, any one of a variety of classification techniques can be used. In one embodiment, the classifier is a non-linear classifier. In a particular embodiment, the classifier is a non-linear classifier that employs L2-regularized kernel logistic regression with a Gaussian radial basis function (RBF) kernel.

[0233] Specifically, in one embodiment, the number of non-cancer samples or one or more different cancer types (n_other) and the number of cancer samples or one or more cancer types having anomalous methylation overlapping with the CpG site (n_cancer) are counted. Then the probability that the sample is cancer is estimated by a score ("S") that is positively correlated with n_cancer and negatively correlated with n_other. The score can be calculated using the following equation: (n_cancer + 1) / (n_cancer + n_other + 2) or (n_cancer) / (n_cancer + n_other). The analysis system calculates 670 the information gain for each cancer type and each genomic region or CpG site to determine whether the genomic region or CpG site indicates cancer. The information gain is calculated for the training samples, and these training samples are given the cancer type compared to all other training samples. For example, two random variables "anomalous fragment" ("AF") and "cancer type" ("CT") are used. In one embodiment, AF is a binary variable indicating whether there is an anomalous fragment overlapping with a given CpG site in a given sample (as determined for the anomalous score / feature vector above). CT is a random variable indicating whether the cancer belongs to a certain specific type. The analysis system calculates the mutual information of CT with respect to AF given. That is, how much information about the cancer type can be obtained if it is known whether there is an anomalous fragment overlapping with a specific CpG site.

[0234] For a given cancer type, the analysis system uses this information to rank these CpG sites according to their cancer specificity. This procedure is repeated for all cancer types under consideration. If a particular region is typically aberrantly methylated in the training samples of a given cancer, but not in the training samples of other cancer types or healthy training samples, then for that given cancer type, the CpG sites overlapping those aberrant segments will tend to have high information gain. The CpG sites for each cancer type, after being ranked, are added (selected) to the set of selected CpG sites by a greedy algorithm according to their ranking for use in a cancer classifier.

[0235] Calculate the pairwise information gain from the segments indicative of cancer identified from the probabilistic model

[0236] Using the segments indicative of cancer identified according to the second method described herein, the analysis can identify genomic regions according to the Figure 15B process 680 in. The analysis system defines a feature vector 690 for each sample, each region, and each cancer type by counting DNA segments having a calculated log-likelihood ratio above a plurality of thresholds indicative of cancer, where each count is a value in the feature vector. In one embodiment, the analysis system counts the number of segments present at a region of each cancer type in a sample that have a log-likelihood ratio above one or more possible thresholds. The analysis system defines a feature vector for each sample by counting the number of DNA segments in each genomic region of each cancer type, the feature vector providing the calculated log-likelihood ratio for segments above a plurality of thresholds, where each count is a value in the feature vector. The analysis system calculates an information content score for each genomic region using the defined feature vector, the score describing the ability of that genomic region to distinguish each pair of cancer types. For each pair of cancer types, the analysis system ranks the regions based on the information content score. The analysis system can select regions based on the ranking according to the information content score.

[0237] The analysis system calculates an information score for each region, which describes the ability of the region to distinguish each pair of cancer types. For each different pair of cancer types, the analysis system can designate one type as the positive type and the other type as the negative type. In one embodiment, the ability of the region to distinguish the positive type and the negative type is based on mutual information, which is calculated using the estimated scores of cfDNA samples of the positive type and the negative type, for which it is expected that the features will be non-zero in the final assay (i.e., at least one fragment of that level will be sequenced in the targeted methylation assay). These scores are estimated using the ratios at which the features appear in healthy cfDNA and in the highly-signal cfDNA and / or tumor samples of each cancer type. For example, if a feature appears frequently in healthy cfDNA, then it is estimated that it will also appear frequently in the cfDNA of any cancer type, and this may result in a low information score. The analysis system can select a certain number of regions, such as 1024, from the ranking for each pair of cancer types.

[0238] In additional embodiments, the analysis system further identifies regions that are predominantly hypermethylated or hypomethylated from the region ranking. The analysis system can load a set of one or more positive type fragments of the regions identified as informative. The analysis system evaluates whether the loaded fragments are predominantly hypermethylated or hypomethylated based on the loaded fragments. If the loaded fragments are predominantly hypermethylated or hypomethylated, the analysis system can select the probes corresponding to the predominant methylation pattern. If the loaded fragments are not predominantly hypermethylated or hypomethylated, the analysis system can use a mixed probe to target hypermethylation and hypomethylation. The analysis system can further identify the minimum set of CpG sites that overlap more than a certain percentage of the fragments.

[0239] In other embodiments, after ranking the regions based on the information score, the analysis system marks each region that has the lowest information across all cancer type pair rankings. For example, if a region is the 10th most informative region for distinguishing kidney and bladder and the 5th most informative region for distinguishing kidney and prostate, then the region will be given an overall label of "5". The analysis system can start designing probes from the lowest-labeled regions while adding regions to the set, such as until the size budget of the set has been exhausted.

[0240] Off-target genomic regions

[0241] In some embodiments, probes targeting selected genomic regions are further filtered 475 based on the number of similar or identical off-target sequences in the genome. This is to screen out probes that pull down an excessive number of cfDNA fragments corresponding to or derived from off-target genomic regions. Excluding probes with many off-target sequences in the genome can be valuable by reducing the off-target rate and increasing the target coverage for a given sequencing volume.

[0242] An off-target genomic region is a genomic region that has sufficient homology to the target genomic region such that a DNA molecule or fragment derived from the off-target genomic region hybridizes to and is pulled down by a probe designed to hybridize to the target genomic region. The off-target genomic region can contain a sequence (or a transformed sequence of that same region) that aligns with the probe along at least 35 bp, 40 bp, 45 bp, 50 bp, 60 bp, 70 bp, or 80 bp and has a match rate of at least 80%, 85%, 90%, 95%, or 97%. In one embodiment, the off-target genomic region is a genomic region that contains a sequence (or a transformed sequence of that same region) that aligns with the probe along at least 45 bp and has a match rate of at least 90%. Various methods can be employed to screen for off-target genomic regions.

[0243] Performing an exhaustive search of the genome to find all off-target genomic regions can be computationally challenging. In one embodiment, a k-mer seeding strategy (which can allow for one or more mismatches) is combined with a local alignment at the seed positions. In this case, an exhaustive search for a good alignment can be guaranteed based on the k-mer length, the number of allowed mismatches, and the number of k-mer seed hits at a particular position. This requires dynamic programming local alignments at a large number of positions, so this approach is highly optimized to use vector CPU instructions (e.g., AVX2, AVX512), and can also be parallelized across multiple cores within a machine and also across multiple machines connected by a network. One of ordinary skill in the art will recognize that modifications and variations of this approach can be implemented for the purpose of identifying off-target genomic regions.

[0244] In some embodiments, probes that have sequence homology to off-target genomic regions, or DNA molecules corresponding to or derived from off-target genomic regions that contain more than a threshold number of off-target genomic regions are excluded (or filtered) from the set. For example, probes that have sequence homology to off-target genomic regions, or DNA molecules corresponding to or derived from off-target genomic regions that have more than 30, more than 25, more than 20, more than 18, more than 15, more than 12, more than 10, or more than 5 off-target regions are excluded.

[0245] In some embodiments, the probes are divided into 2, 3, 4, 5, 6 or more separate groups according to the number of off-target regions. For example, probes having sequence homology to non-off-target regions or DNA molecules corresponding to or derived from off-target regions are assigned to a high-quality group; probes having sequence homology to 1-18 off-target regions or DNA molecules corresponding to or derived from 1-18 off-target regions are assigned to a low-quality group; and probes having sequence homology to more than 19 off-target regions or DNA molecules corresponding to or derived from 19 off-target regions are assigned to a poor-quality group. Other cut-off values can be used for grouping.

[0246] In some embodiments, the probes in the lowest-quality group are excluded. In some embodiments, the probes in groups other than the highest-quality group are excluded. In some embodiments, separate sets are made for the probes in each set. In some embodiments, all the probes are placed on the same set, but separate analyses are performed based on the assigned groups.

[0247] In some embodiments, the set contains a greater number of high-quality probes than the number of probes in the lower groups. In some embodiments, the set contains a smaller number of poor-quality probes than the number of probes in the other groups. In some embodiments, more than 95%, 90%, 85%, 80%, 75% or 70% of the probes in the set are high-quality probes. In some embodiments, less than 35%, 30%, 20%, 10%, 5%, 4%, 3%, 2% or 1% of the probes in the set are low-quality probes. In some embodiments, less than 5%, 4%, 3%, 2% or 1% of the probes in the set are poor-quality probes. In some embodiments, no poor-quality probes are included in the set.

[0248] In some embodiments, probes having less than 50%, less than 40%, less than 30%, less than 20%, less than 10% or less than 5% are excluded. In some embodiments, probes having more than 30%, more than 40%, more than 50%, more than 60%, more than 70%, more than 80% or more than 90% are selectively included in the set.

[0249] Method of using a cancer assay set

[0250] In one aspect, methods are provided that use a cancer assay set (alternatively referred to as a “bait set”). The cancer assay set can be any of those disclosed herein, such as any of those in the various aspects. The methods can include the steps of: processing DNA molecules or fragments to convert unmethylated cytosine to uracil (e.g., using bisulfite treatment); applying a cancer set (as described herein) to the processed DNA molecules or fragments; enriching a subset of the processed DNA molecules or fragments that bind to the probes in the set; and sequencing the enriched DNA fragments. In some embodiments, sequence reads can be compared to a reference genome (e.g., a human reference genome), thereby allowing identification of the methylation status at multiple CpG sites within the DNA molecules or fragments and thus providing information relevant to cancer detection. In some embodiments, the DNA is cfDNA. In some embodiments, the cfDNA is from a urine sample. In some embodiments, target genomic regions identified as being particularly useful in characterizing cfDNA from urine samples can also be used to characterize cfDNA in other sample types (e.g., cfDNA from a biological fluid such as blood, serum, or plasma). In some embodiments, information regarding the target genomic regions disclosed herein is applied to assays of cfDNA from such other sample types for cancer detection.

[0251] In one aspect, the present disclosure provides a method of detecting cancer cells in a subject, the method comprising: (a) capturing processed cell-free DNA (cfDNA) fragments or amplification products thereof from a urine sample of the subject, wherein (i) the bait oligonucleotide composition comprises a plurality of different bait oligonucleotides and (ii) each of the plurality of different bait oligonucleotides hybridizes to a target sequence of a gene selected from Table 1, wherein the length of the target sequence is at least 25 nucleotides; (b) separating the bait-bound DNA from the unbound DNA; (c) sequencing the separated DNA to generate sequencing reads; and (d) detecting the cancer cells with a trained classifier.

[0252] Cell-free DNA fragments can be obtained from urine samples prepared by any of the methods disclosed herein. In some embodiments, the cell-free nucleic acids are obtained by: (i) treating the urine sample to inhibit cell lysis; (ii) separating cfDNA fragments in the treated urine sample from the cells in the treated urine sample, thereby producing a purified urine sample containing these cfDNA fragments; (iii) concentrating the cfDNA fragments in the purified urine sample by passing at least a portion of the purified urine sample through a filter, wherein the concentration produces a filtrate and a retained urine sample, and wherein the retained urine sample contains an increased concentration of cfDNA fragments; and (iv) isolating cfDNA fragments from the retained urine sample. In some embodiments, treating the urine sample to inhibit cell lysis includes contacting the urine sample with one or more preservative reagents, non-limiting examples of which are described herein. In some embodiments, the treatment includes treating with a nuclease inhibitor, a formaldehyde quencher, or both. Non-limiting examples of nuclease inhibitors and formaldehyde quenchers are described herein. In some embodiments, the treatment includes contacting the urine sample with a composition comprising imidazolidinyl urea, EDTA, glycine, or a combination thereof. In some embodiments, the treatment includes contacting the urine sample with a composition comprising sodium azide, EDTA, or a combination thereof. Separating the nucleic acids can include centrifuging the treated urine sample to pellet the cells. The filter used to prepare the retained urine sample may be substantially impermeable to cell-free nucleic acids but permeable to salts in the purified urine sample. In some embodiments, the filter has a nominal molecular weight cut-off of 10 kD, 5 kD, 3 kD, or lower. In some embodiments, the retained urine sample has a concentration that is increased by at least 2-fold, at least 5-fold, at least 10-fold, or at least 15-fold compared to the purified urine sample. In some cases, the retained urine sample has a volume that is at least 50%, at least 75%, or at least 90% lower than the volume of the treated urine sample. The volume of the treated urine sample can be 5 mL, 10 mL, 15 mL, 20 mL, 30 mL, 40 mL, 50 mL, or more. The method can further include freezing the retained urine sample. In some embodiments, the treatment is completed within 120, 60, or 30 minutes after urine sample collection. In some embodiments, the separation and concentration are completed within 1 to 14 days (e.g., within 7 days) after collection. In some embodiments, the method further includes amplifying one or more of the isolated cfDNA fragments.

[0253] The bait oligonucleotides for capturing the converted free DNA molecules can be any of the bait oligonucleotides described herein, such as for the various compositions described herein. In some embodiments, the length of each of the plurality of different bait oligonucleotides in the set is at least 45 nucleotides (e.g., the length is at least 60, 75, 80, 90, 100, 110, or 120 nucleotides). In some embodiments, the length of each bait oligonucleotide among the plurality of probes does not exceed 130, 140, 150, 200, 250, or 300 bases. In some embodiments, the length of each bait oligonucleotide is 45 to 300, 60 to 200, or 75 to 150 nucleotides. In some embodiments, the length of the bait oligonucleotide is at least 50 nucleotides. In some embodiments, the length of the bait oligonucleotide is at least 60 nucleotides. In some embodiments, the length of the bait oligonucleotide is at least 75 nucleotides. In some embodiments, the indicated length of the bait oligonucleotide is designed to be complementary to a portion of the target genomic sequence or its converted DNA molecule.

[0254] In some embodiments, the cancer assay set is designed to target at least 500, 1000, 1500, 5000, 10000, 12500, 15000, 17000, 19000 or more target genomic regions. In some embodiments, the cancer assay set is designed to target fewer than 25000, 20000, 17000, 15000, 12500, 10000 or fewer target genomic regions. In some embodiments, the cancer assay set is designed to target 5000 to 30000, 10000 to 25000, 12500 to 20000 or 15000 to 20000 target genomic regions. In some embodiments, the cancer assay set is designed to target 100 to 500, 500 to 1000, 1500 to 5000, 5000 to 10000 target genomic regions. In some embodiments, the cancer assay set is designed to target at least 500 target genomic regions. In some embodiments, the cancer assay set is designed to target at least 1000 target genomic regions. In some embodiments, the cancer assay set is designed to target at least 10000 target genomic regions. In some embodiments, the cancer assay set is designed to target at least 15000 target genomic regions. In some embodiments, the cancer assay set is designed to target fewer than 20000 target genomic regions.

[0255] In some embodiments, the bait oligonucleotides are configured to hybridize to transformed DNA molecules (e.g., transformed cfDNA molecules) corresponding to or derived from one or more genomic regions. Thus, the bait oligonucleotides can have sequences different from the target genomic regions. For example, DNA containing unmethylated CpG sites can be transformed to include UpG instead of CpG by deamination (e.g., by treatment with cytidine deaminase or bisulfite). Thus, probes directed against such targets can be configured to hybridize to sequences including UpG instead of the naturally occurring unmethylated CpG. Thus, the sites in the probe complementary to the unmethylated sites can contain CpA instead of CpG, and some probes targeting hypomethylated sites where all methylated sites are unmethylated may not have a guanine (G) base. In some embodiments, at least 3%, 5%, 10%, 15%, or 20% of the probes do not contain a CpG sequence. In some embodiments, at least 5% of the probes do not contain a CpG sequence. In some embodiments, at least 10% of the probes do not contain a CpG sequence.

[0256] In some embodiments, a cancer assay set is used to detect the presence or absence of cancer and / or to provide cancer classification (e.g., cancer type), cancer staging (e.g., I, II, III, or IV), or to provide a TOO that is considered the origin of the cancer. The set can include probes that target differentially methylated genomic regions between generally cancerous (pan-cancer) samples and non-cancerous samples or only among cancerous samples having a specific cancer type (e.g., bladder cancer-specific targets). For example, in some embodiments, a cancer assay set is designed to include differentially methylated genomic regions based on transformed (e.g., bisulfite) sequencing data generated from cfDNA and / or whole-genome DNA from sets of cancer and non-cancer individuals.

[0257] In some embodiments, each of the target genomic regions is differentially methylated in at least one of a plurality of cancer types. In some embodiments, the plurality of cancer types includes at least 2 cancer types (e.g., at least 2, 3, 3, 4, or more cancer types). In some embodiments, the plurality of cancer types includes urinary system cancers. In some embodiments, the plurality of cancer types includes one or more of bladder cancer, urothelial cancer, prostate cancer, or kidney cancer.

[0258] Each of the probes (or probe pairs) can be designed to target one or more target genomic regions. The target genomic regions can be selected based on several criteria that are intended to enhance the selective enrichment of informative cfDNA fragments while reducing noise and non-specific binding. Various filtering procedures for determining whether to include a target genomic region are described herein. In some embodiments, two or more of the filtering procedures described herein are used in combination.

[0259] In some embodiments, a trained classifier is used to detect cancer based on sequencing data. In some embodiments, for one or more of the target sequences identified as hypermethylated and / or hypomethylated in cfDNA fragments, the trained classifier detects that the number of sequencing reads exceeds a threshold. In some embodiments, the trained classifier differentiates subjects with cancer from subjects without cancer with a defined specificity. Thus, in some embodiments, samples from patients with an unknown type of cancer (e.g., undiagnosed cancer) can be used to identify which of a variety of different types of cancer may be present and which are not. In some embodiments, the classifier is a binary classifier, a mixture model classifier, or a multi-layer perceptron model classifier. In some embodiments, the classifier is a mixture model classifier. In some embodiments, the defined specificity for each of a variety of cancer types is 0.900 or higher (e.g., at least 0.950, 0.975, 0.980, 0.985, 0.990, 0.995, or higher). In some embodiments, the application of the trained classifier includes at least 30% (e.g., at least 40%, 50%, 60%, 70%, 80%, or higher) sensitivity for each of a variety of cancer types. In some embodiments, the trained classifier has at least 30% sensitivity for determining cancer likelihood and 0.900 or higher specificity for each of a variety of cancer types. In some embodiments, the trained classifier has at least 40% sensitivity for determining cancer likelihood and 0.990 or higher specificity for each of a variety of cancer types.

[0260] In some embodiments, the method includes processing cfDNA molecules to distinguish methylated nucleotides from unmethylated nucleotides, thereby generating transformed cfDNA molecules. In some embodiments, the processing includes deamination, such as treating the cfDNA molecules with a cytidine deaminase or with bisulfite. In some embodiments, the method includes treating the cfDNA molecules with bisulfite to generate transformed cfDNA molecules.

[0261] In some embodiments, each bait oligonucleotide is conjugated to a solid surface (e.g., a chip or bead, such as a magnetic or paramagnetic bead) or to a non-nucleotide affinity moiety (e.g., a member of a binding pair). In some embodiments, such conjugation is used to facilitate separation of DNA molecules bound to the bait oligonucleotide from unbound DNA molecules. In some embodiments, the affinity moiety is biotin.

[0262] In some embodiments, the selectable genomic region has at least 3, 5, or 7 methylation sites. In some embodiments, each target genomic region contains at least five methylation sites. In some embodiments, the selected number of methylation sites (e.g., at least 5 methylation sites) are methylation sites that are differentially methylated in at least one type of cancer to be assayed by the panel. In some embodiments, the target genomic region contains at least 1, 2, 3, 4, 5, 10, 15, or more of the target genomic regions containing target sequences from one or more genes selected from Table 1. In some embodiments, the target genomic region contains target sequences of genes selected from: TWIST1, EOMES, HOXA9, POU4F2, and ZNF154.

[0263] In some embodiments, the method further comprises diagnosing cancer in a subject. In some embodiments, the method further comprises selecting a treatment for the identified cancer type. In some embodiments, the method further comprises treating the cancer of the subject. In some embodiments, the cancer type includes bladder cancer, urothelial cancer, prostate cancer, or kidney cancer. The specific mode of treatment may depend on one or more of various factors, such as the specific type of cancer detected, cancer TOO, location, and stage. Non-limiting examples of treatment include surgical resection, radiotherapy, chemotherapy, and / or immunotherapy.

[0264] Sequence read analysis

[0265] In some embodiments, various methods can be used to align sequence reads with a reference genome to determine alignment position information. The alignment position information can indicate the start position and end position of a region in the reference genome that corresponds to the start nucleotide base and end nucleotide base of a given sequence read. The alignment position information can also include the sequence read length, which can be determined by the start position and end position. The region in the reference genome may be associated with a gene or a fragment of a gene.

[0266] In various embodiments, the sequence reads are composed of read pairs represented as and. For example, the first read can be sequenced from the first end of the nucleic acid fragment, and the second read is sequenced from the second end of the nucleic acid fragment. Therefore, the nucleotide base pairs of the first read and the second read can be aligned in a consistent manner (for example, in the opposite direction) with the nucleotide bases of the reference genome. The alignment position information derived from the read pairs and may include a starting position in the reference genome, which corresponds to one end of the first read (for example,), and an end position in the reference genome, which corresponds to one end of the second read (for example,). In other words, the starting position and end position in the reference genome represent positions that the nucleic acid fragment may correspond to in the reference genome. Output files in SAM (sequence alignment map) format or BAM (binary alignment map) format can be generated and output for further analysis.

[0267] According to the sequence reads, the position and methylation state of each of the CpG sites can be determined based on the comparison with the reference genome. Further, a methylation state vector for each fragment can be generated, which specifies the position of the fragment in the reference genome (e.g., as specified by the position of the first CpG site in each fragment or another similar indicator), the number of CpG sites in the fragment, and the methylation state of each CpG site in the fragment (methylation (e.g., expressed as M), unmethylation (e.g., expressed as U), or uncertain (e.g., expressed as I)). The methylation state vector may be stored in a transient or permanent computer memory for subsequent use and processing. Further, repeated reads or repeated methylation state vectors from a single subject can be removed. In another embodiment, it can be determined that a certain fragment has one or more CpG sites with uncertain methylation states. Such fragments can be excluded from subsequent processing, or selectively included in the case where the downstream data model explains such uncertain methylation states.

[0268] Figure 16B According to the embodiment Figure 16A 100 for sequencing cfDNA fragments to obtain a methylation state vector. For example, the analysis system uses cfDNA fragment 112. In this example, cfDNA fragment 112 contains three CpG sites. As shown, the first and third CpG sites of cfDNA fragment 112 are methylated 114. During processing step 120, cfDNA fragment 112 is converted to generate converted cfDNA fragment 122. During processing 120, the unmethylated second CpG site, its cytosine is converted to uracil. However, the first and third CpG sites are not converted.

[0269] After conversion, a sequencing library 130 is prepared and sequencing 140 is performed to generate sequence reads 142. An analysis system aligns 150 the sequence reads 142 with a reference genome 144. The reference genome 144 provides context as to where in the human genome the cfDNA fragment originated. In this simplified example, the analysis system performs alignment 150 of the sequence reads such that those three CpG sites correspond to CpG sites 23, 24, and 25 (these are arbitrary reference identifiers used for ease of description). Thereby, the analysis system generates information regarding the methylation status of all CpG sites on the cfDNA fragment 112 and the locations in the human genome where these CpG sites map. As shown, the methylated CpG sites on the sequence reads 142 are read as cytosine. In this example, cytosine appears only at the first and third CpG sites of the sequence reads 142, from which it can be inferred that the first and third CpG sites in the original cfDNA fragment are methylated. The second CpG site is read as thymine (U is converted to T during the sequencing process), from which it can be inferred that the second CpG site in the original cfDNA fragment is unmethylated. With these two pieces of information, the methylation status and the location, the analysis system generates 160 a methylation status vector 152 for the cfDNA fragment 112. In this example, the resulting methylation status vector 152 is <M23,U24,M25>, where M corresponds to a methylated CpG site, U corresponds to an unmethylated CpG site, and the subscript numbers correspond to the location of each CpG site in the reference genome.

[0270] Cancer detection

[0271] The sequence reads obtained by the methods provided herein can be further processed by automated algorithms. For example, an analysis system is used to receive sequencing data from a sequencer and perform various aspects of the processing as described herein. The analysis system can be one of a personal computer (PC), a desktop computer, a laptop computer, a notebook, a tablet PC, a mobile device. The computing device can be communicatively coupled to the sequencer by way of wireless, wired, or a combination of wireless and wired communication technologies. Generally, the computing device is configured with a processor and a memory storing computer instructions that, when executed by the processor, cause the processor to perform the steps described in the remainder of this document. Generally, the amount of genetic data and the data derived therefrom is large enough and the amount of computing power required is so great that it is not possible to perform on paper or by the human brain alone.

[0272] The clinical interpretation of the methylation status of a target genomic region is a process that can include classifying the clinical effects of each or combinations of methylation statuses and reporting the results in a manner meaningful to medical professionals. The clinical interpretation can be based on comparing sequence reads to databases specific to cancer or non-cancer subjects, and / or based on the number and type of cfDNA fragments identified in the sample that have cancer-specific methylation patterns. In some embodiments, the target genomic regions are ranked or classified based on their similarity in differential methylation in cancer samples, and these rankings or classifications are used during the interpretation process. The ranking and classification can include: (1) the type of clinical effect; (2) the strength of evidence for the effect; and (3) the magnitude of the effect. A variety of methods for clinical analysis and genomic data interpretation can be employed to analyze the sequence reads. In some other embodiments, the clinical interpretation of the methylation status of such differentially methylated regions can be based on machine learning methods that interpret the current sample based on classification or regression methods that are trained using the methylation status of such differentially methylated regions from samples of cancer patients and non-cancer patients with known cancer status, cancer type, cancer stage, TOO, etc.

[0273] Clinical significance information can include the presence or absence of cancer in general, the presence or absence of certain types of cancer, cancer stage, or the presence or absence of other types of disease. In some embodiments, the information pertains to the presence or absence of urinary system cancer. In some embodiments, the information pertains to the presence or absence of one or more cancer types selected from the group consisting of bladder cancer, urothelial cancer, kidney cancer, and prostate cancer.

[0274] Cancer classifier

[0275] In some instances, the assay sets described herein can be used with a cancer type classifier that can predict the disease state of a sample, such as a cancer or non-cancer prediction, a source tissue prediction, and / or an indeterminate state prediction. In some instances, the cancer type classifier can generate features based on sequence reads by considering methylated or unmethylated DNA fragments in certain genomic regions of interest. For example, if the cancer type classifier determines that the methylation pattern of a fragment is similar to the methylation pattern of a certain cancer type, the cancer type classifier can set the feature of that fragment to 1, otherwise, if no such fragment exists, the feature can be set to 0. In this way, the cancer type classifier can generate a binary feature set (e.g., 30,000 features for illustration only) for each sample. Further, in some instances, all or part of the binary feature set of a sample can be input into the cancer type classifier to provide a set of probability scores, such as one probability score for each cancer type category and non-cancer type category. Additionally, in some instances, the cancer type classifier can be combined with a threshold setting or otherwise used in conjunction with a threshold setting to determine whether a sample is determined to be cancer or non-cancer, and / or combined with an uncertainty threshold setting or otherwise used in conjunction with an uncertainty threshold setting to reflect the confidence in a particular TOO determination. Such methods will be described further below.

[0276] To train the cancer type classifier, an analysis system (e.g., analysis system 800, Figure 17B ) can obtain a set of training samples. In some instances, each training sample includes one or more fragment files (e.g., files containing sequence read data), a label corresponding to the cancer type (TOO) or non-cancer status of the sample, and / or the gender of the sample individual. The analysis system can use the training set to train the cancer type classifier to predict the disease state of the sample.

[0277] In some instances, for training, the analysis system divides the genome (e.g., the whole genome) or a subset of the genome (e.g., targeted methylation regions) into multiple regions. For illustration only, a portion of the genome can be divided into CpG "blocks", and a new block starts whenever the interval between the nearest neighbor CpGs is at least a minimum interval distance (e.g., at least 500 bp). Further, in some instances, each block can be divided into 1000 bp regions and positioned such that neighboring regions have a certain amount (e.g., 50% or 500 bp) of overlap.

[0278] In addition, in some instances, the analysis system can divide the training set into K subsets or folds for use in K-fold cross-validation. In some instances, the folds can be balanced with respect to cancer / non-cancer status, source tissue, cancer stage, age (e.g., grouped in 10-year intervals), and / or smoking status. In some instances, the training set is divided into 5 folds, such that 5 independent classifiers are trained, each time on 4 / 5 of the training samples and using the remaining 1 / 5 for validation.

[0279] During training using the training set, the analysis system can fit a probability model to the fragments derived from samples of each cancer type (and for healthy cfDNA). As used herein, a "probability model" is any mathematical model capable of assigning a probability to a sequence read based on the methylation status at one or more positions on the read. During training, the analysis system fits sequence reads derived from one or more samples from subjects with a known disease and can be used to determine the sequence read probability indicative of the disease state using methylation information or a methylation status vector. In particular, in some cases, the analysis system determines the methylation rate of each CpG site observed in the sequence read. The methylation rate represents the proportion or percentage of base pairs within the CpG site that are methylated. The trained probability model can be parameterized by the product of the methylation rates. Generally, any known probability model for assigning probabilities to sequence reads from samples can be used. For example, the probability model can be a binomial model where each position on the nucleic acid fragment (e.g., CpG site) is assigned a methylation probability; or it can be an independent sites model where the methylation of each CpG is specified by a different methylation probability, where methylation at one position is assumed to be independent of methylation at one or more other positions on the nucleic acid fragment.

[0280] In some instances, the probability model is a Markov model where the methylation probability at each CpG site depends on the methylation status of a certain number of preceding CpG sites in the sequence read or the nucleic acid molecule from which the sequence read is derived. See, for example, U.S. Patent Application No. 16 / 352,602, titled "Anomalous Fragment Detection and Classification", and filed on March 13, 2019, which is incorporated herein by reference in its entirety and can be used in various embodiments.

[0281] In some instances, the probability model is a "mixture model" fit using mixture components from a base model. For example, in some embodiments, the mixture components can be determined using multiple independent site models, where it is assumed that methylation (e.g., methylation rate) at each CpG site is independent of methylation at other CpG sites. Using the independent site model, the probability assigned to a sequence read or the nucleic acid molecule from which it is derived is the product of the methylation probability at each CpG site where the sequence read is methylated and one minus the methylation probability at each CpG site where the sequence read is not methylated. According to this instance, the analysis system determines the methylation rate for each of the mixture components. The mixture model is parameterized by the sum of the mixture components each associated with a product of the methylation rate. The probability model Pr for n mixture components can be expressed as:

[0282]

[0283] For an input segment, m i ∈ {0, 1} represents the methylation state observed at position i in the reference genome for the segment, where 0 represents unmethylated and 1 represents methylated. For f k ≥ 0 and , the fractional assignment for each mixture component k is f k . The methylation probability at position i in the CpG sites of mixture component k is β ki . Thus, the probability of being unmethylated is 1 - β ki . The number of mixture components n can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, etc.

[0284] In some instances, the analysis system fits the probability model using maximum likelihood estimation to identify the parameter set {β ki , f k} that maximizes the log likelihood of all segments derived from the disease state and imposes a regularization penalty on each methylation probability with a regularization strength r. The maximized number for N total segments can be expressed as:

[0285]

[0286] In some instances, the analysis system performs a fit for each cancer type and healthy cfDNA separately. According to various aspects of the present disclosure, other means can be used to fit a probability model or identify parameters that maximize the log-likelihood of all sequence reads derived from a reference sample. For example, in some instances, Bayesian fitting (using, e.g., Markov Chain Monte Carlo) is used, where each parameter is not assigned a single value but is associated with a distribution. In some instances, gradient-based optimization is used, where the gradient of the likelihood (or log-likelihood) with respect to the parameter values is used to step through the parameter space towards the optimal solution. In still other instances, expectation maximization is used, where a set of latent parameters (e.g., the identity of the mixture components from which each fragment is derived) is set to the expected values under the previous model parameters, and then the parameters of the model are assigned to maximize the likelihood conditional on the assumed values of these latent variables. This two-step process is then repeated until convergence.

[0287] Further, in some instances, the analysis system can generate features for each sample in the training set. For example, for each sample (regardless of label), in each region, for each cancer type, for each fragment, the analysis system can evaluate the log-likelihood ratio R using the fitted probability model as follows:

[0288]

[0289] Next, for each sample, for each region, for each cancer type, for each of the values in the "tier" set, the analysis system can count the number of fragments where Rcancer type > tier and assign these counts as non-negative integer-valued features. For example, the tiers include thresholds of 1, 2, 3, 4, 5, 6, 7, 8, and 9, resulting in 9 features for each region for each cancer type.

[0290] In some instances, the analysis system can select certain features for inclusion in the feature vector for each sample. For example, for each pair of different cancer types, the analysis system can designate one type as the "positive type" and the other type as the "negative type" and rank the features according to their ability to distinguish between these types. In some cases, the ranking is based on the mutual information calculated by the analysis system. For example, the mutual information can be calculated using the estimated scores for samples of the positive type and negative type (e.g., cancer types A and B), for which the feature is expected to be non-zero in the outcome assay. For example, if a feature occurs frequently in healthy cfDNA, the analysis system determines that the feature is unlikely to occur frequently in cfDNA associated with various types of cancer. Thus, the feature may be a weaker indicator in distinguishing the disease state. When calculating the mutual information I, the variable X is a certain feature (e.g., binary), and the variable Y represents the disease state, e.g., cancer type A or B:

[0291]

[0292] p(1|A) = f A +f H -f H f A

[0293] The joint probability mass function of X and Y is p(x, y), and the marginal probability mass functions are p(x) and p(y). The analysis system can assume that missing features are not informative, and either disease state is equally likely a priori, e.g., p(Y = A) = p(Y = B) = 0.5. The probability of a given binary feature for cancer type A, observed (e.g., in cfDNA), is denoted by p(1|A), where f A is the probability of observing the feature in tumor ctDNA samples (or high-signal cfDNA samples) associated with cancer type A, and f H is the probability of observing the feature in healthy or non-cancer cfDNA samples.

[0294] In some instances, only features corresponding to positive types are included in the ranking, and only if their predicted incidence in the positive type is higher than in the negative type. For example, if "liver" is the positive type and "breast" is the negative type, only "liver_x" features are considered, and only if their estimated incidence in liver cfDNA is higher than in breast cfDNA. Further, in some instances, for each region, for each pair of cancer types (including non-cancer as the negative type), the analysis system retains only the best-performing tier. Further, in some instances, the analysis system binarizes the feature values, thereby setting any feature value greater than 0 to 1, such that all features are either 0 or 1.

[0295] In some instances, the analysis system trains a multinomial logistic regression classifier on the folded training data and generates predictions for the held-out data. For example, for each of the K folds, a logistic regression can be trained for each combination of hyperparameters. Such hyperparameters can include L2 penalty and / or topK (e.g., the number of highly ranked regions to retain for each pair of tissue types (including non-cancer), as ranked by the mutual information procedure described above). For each set of hyperparameters, the performance is evaluated based on cross-validation predictions on the full training set, and the set of hyperparameters with the best performance is selected for retraining on the full training set. In some instances, the analysis system uses log loss as the performance metric, thereby calculating the log loss as follows: taking the negative logarithm of the correct-label prediction for each sample and then summing over the samples (i.e., for a perfect prediction with the correct label of 1.0, the log loss is 0).

[0296] To generate predictions for new samples, eigenvalue calculations are performed using the same method as above, but limited to the features (region / positive class combinations) selected at the chosen topK value. The resulting features are then used to create predictions using the logistic regression model trained above.

[0297] In some instances, the analysis system trains a two-level classifier. For example, the analysis system trains a binary cancer classifier based on the feature vectors of training samples to distinguish labels (cancer and non-cancer). In this case, the binary classifier outputs a prediction score indicating the likelihood of the presence or absence of cancer. In another instance, the analysis system trains a multi-class cancer classifier to distinguish between multiple cancer types. In this multi-class cancer classifier, the cancer classifier is trained to determine a cancer prediction that includes a prediction value for each of the cancer types being classified. The prediction value can correspond to the likelihood that a given sample has each of the cancer types. For example, the cancer classifier returns a cancer prediction that includes prediction values for urinary system cancer and non-cancer. For example, the cancer classifier can return a cancer prediction for a test sample that includes prediction scores for bladder cancer or urothelial cancer, kidney cancer, prostate cancer, and / or non-cancer.

[0298] The analysis system may train the cancer classifier according to any of a number of methods. For example, the binary cancer classifier may be an L2-regularized logistic regression classifier trained using a log-loss function. As another example, the multi-cancer (TOO) classifier may be a multinomial logistic regression classifier. In practice, other techniques may be used to train either type of cancer classifier. There are many such techniques, including kernel methods, machine learning algorithms (such as multi-layer neural networks), etc., that may be used. In particular, the methods described in PCT / US2019 / 022122 and U.S. Patent Application No. 16 / 352,602, which are incorporated herein by reference in their entirety, may be used in various embodiments. Further, in some instances, the TOO classifier is trained only on cancer samples that have been successfully determined to be cancer by the binary classifier, thus ensuring sufficient cancer signal in the cancer samples. On the other hand, in some instances, the binary classifier is trained on the training samples regardless of the TOO.

[0299] Exemplary Sequencer and Analysis System

[0300] Figure 17A is a flowchart of a system and apparatus for sequencing a nucleic acid sample according to one embodiment. The illustrative flowchart includes, for example, apparatus such as sequencer 820 and analysis system 800. Sequencer 820 and analysis system 800 can work in series to perform one or more steps of the processes described herein.

[0301] In various embodiments, a sequencer 820 receives an enriched nucleic acid sample 810. As Figure 17A shown, the sequencer 820 can include a graphical user interface 825 that enables a user to interact with specific tasks (e.g., initiate sequencing or terminate sequencing) and one or more loading stations 830 for loading sequencing cartridges that include these enriched fragment samples and / or for loading buffers necessary to perform a sequencing assay. Thus, once a user of the sequencer 820 has provided the necessary reagents and sequencing cartridges to the loading stations 830 of the sequencer 820, the user can initiate sequencing by interacting with the graphical user interface 825 of the sequencer 820. Once initiated, the sequencer 820 performs sequencing and outputs sequence reads of the enriched fragments in the nucleic acid sample 810.

[0302] In some embodiments, the sequencer 820 is communicatively coupled to an analysis system 800. The analysis system 800 includes a number of computing devices for processing the sequence reads to meet various application requirements, such as assessing the methylation status of one or more CpG sites, variant identification, or quality control. The sequencer 820 can provide the sequence reads to the analysis system 800 in a BAM file format. The analysis system 800 can be communicatively coupled to the sequencer 820 via wireless, wired, or a combination of wireless and wired communication technologies. Generally, the analysis system 800 is configured with a processor and a non-transitory computer-readable storage medium storing computer instructions that, when executed by the processor, cause the processor to process the sequence reads or perform one or more steps of any method or process disclosed herein.

[0303] In some embodiments, various methods can be used to align the sequence reads with a reference genome to determine alignment position information. The alignment position can generally describe the start position and end position of a region in the reference genome that corresponds to the start nucleotide base and end nucleotide base of a given sequence read. For methylation sequencing, the concept of alignment position information can be generalized to indicate the first CpG site and the last CpG site included in a sequence read according to the alignment with the reference genome. The alignment position information can further indicate the methylation status and positions of all CpG sites in a given sequence read. A region in the reference genome may be associated with a gene or a fragment of a gene; thus, the analysis system 800 can label the sequence read with one or more genes that the sequence read aligns to. In one embodiment, the fragment length (or size) is determined by the start position and the end position.

[0304] In various embodiments, for example, when using a paired-end sequencing process, sequence reads consist of a pair of reads, denoted as R_1 and R_2. For example, the first read R_1 can be sequenced from the first end of a double-stranded DNA (dsDNA) molecule, while the second read R_2 is sequenced from the second end of the dsDNA. Thus, the nucleotide base pairs of the first read R_1 and the second read R_2 can be aligned in a consistent manner (e.g., in opposite directions) with the nucleotide bases of the reference genome. The alignment position information derived from the read pair R_1 and R_2 may include a starting position in the reference genome that corresponds to one end of the first read (e.g., R_1), and a termination position in the reference genome that corresponds to one end of the second read (e.g., R_2). In other words, the starting and termination positions in the reference genome represent the possible positions of the nucleic acid fragment in the reference genome. In one embodiment, the read pair R_1 and R_2 can be assembled into a fragment, and the fragment is used for subsequent analysis and / or classification. An output file in SAM (Sequence Alignment Map) format or BAM (binary) format can be generated and output for further analysis.

[0305] Now referring to Figure 17B , Figure 17B which is a block diagram of an analysis system 800 for processing DNA samples according to one embodiment. The analysis system employs one or more computing devices for analyzing DNA samples. The analysis system 800 includes a sequence processor 840, a sequence database 845, a model database 855, a model 850, a parameter database 865, and a scoring engine 860. In some embodiments, the analysis system 800 performs Figure 12A process 300, Figure 12B process 340, Figure 13 process 400, Figure 14 process 500, Figure 15A process 600, or Figure 15B process 680 and one or more steps of other processes described herein.

[0306] The sequence processor 840 generates a methylation status vector for fragments from a sample. For each CpG site on a fragment, the sequence processor 840 generates a methylation status vector for each fragment that indicates the position of the fragment in the reference genome, the number of CpG sites within the fragment, and the methylation status (methylated, unmethylated, or indeterminate) of each CpG site in the fragment, and this process is implemented by Figure 12A process 300 in

[0307] Further, multiple different models 850 can be stored in or retrieved from the model database 855 for testing samples. In one instance, the model is a trained cancer classifier that uses feature vectors derived from aberrant segments to determine cancer predictions for test samples. The training and use of the cancer classifier are discussed elsewhere in this document. The analysis system 800 can train one or more models 850 and store various trained parameters in the parameter database 865. The analysis system 800 stores the model 850 together with associated functions in the model database 855.

[0308] During inference, the scoring engine 860 uses one or more models 850 to return an output. The scoring engine 860 accesses the model 850 in the model database 855 and the trained parameters from the parameter database 865. For each model, the scoring engine receives inputs appropriate for that model and calculates an output based on the received inputs, parameters, and functions associated with the inputs and outputs in each model. In some use cases, the scoring engine 860 further calculates metrics related to the confidence in the output calculated by the model. In other use cases, the scoring engine 860 calculates other intermediate values used in the model.

[0309] Cancer and Treatment Monitoring

[0310] In some embodiments, cell-free nucleic acid samples (e.g., cfDNA from a urine sample) may be obtained and analyzed from a cancer patient at first and second time points, e.g., to monitor cancer progression, to determine whether the cancer is in remission (e.g., after treatment), to monitor or detect residual disease or disease recurrence, or to monitor the effectiveness of a treatment (e.g., a therapeutic). In some embodiments, the first time point in cancer monitoring is before cancer treatment (e.g., before resection surgery or a treatment intervention), the second time point is after cancer treatment (e.g., after resection surgery or a treatment intervention), and the method is used to monitor the effectiveness of the treatment. For example, if the second likelihood or probability score decreases compared to the first likelihood or probability score, the treatment is considered successful. However, if the second likelihood or probability score increases compared to the first likelihood or probability score, the treatment is considered a failure. In some embodiments, both the first and second time points are before cancer treatment (e.g., before resection surgery or a treatment intervention). In still other embodiments, both the first and second time points are after cancer treatment (e.g., before resection surgery or a treatment intervention), and the method is used to monitor the effectiveness of the treatment or the loss of treatment effectiveness.

[0311] Test samples can be obtained from a cancer patient at any desired set of time points and analyzed according to the methods of the present invention to monitor the cancer status of the patient. In some embodiments, the duration between the first time point and the second time point ranges from 15 minutes to 30 years, such as 30 minutes, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 hours, such as 1, 2, 3, 4, 5, 10, 15, 20, 25, or 30 days, or such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 months, or such as about 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, 10, 10.5, 11, 11.5, 12, 12.5, 13, 13.5, 14, 14.5, 15, 15.5, 16, 16.5, 17, 17.5, 18, 18.5, 19, 19.5, 20, 20.5, 21, 21.5, 22, 22.5, 23, 23.5, 24, 24.5, 25, 25.5, 26, 26.5, 27, 27.5, 28, 28.5, 29, 29.5, or 30 years. In other embodiments, the frequency of obtaining test samples from the patient can be: at least once every 3 months, at least once every 6 months, at least once a year, at least once every 2 years, at least once every 3 years, at least once every 4 years, or at least once every 5 years.

[0312] Process

[0313] In some embodiments, information obtained from any of the methods described herein (e.g., likelihood or probability scores) can be used to make or influence clinical decisions (e.g., cancer diagnosis, treatment selection, assessment of treatment efficacy, etc.). For example, in one embodiment, if the likelihood or probability score exceeds a threshold, a doctor can prescribe an appropriate treatment (e.g., resection surgery, radiotherapy, chemotherapy, and / or immunotherapy). In some embodiments, information such as likelihood or probability scores can be provided as a readout to a doctor or a subject.

[0314] In one aspect, a method includes selecting a subject having a cancer type or at increased risk of developing a cancer type, and administering to the subject a treatment effective to treat the cancer type, wherein: (a) the selection includes identifying the subject as the source of a cell-free DNA (cfDNA) sample that contains one or more differentially methylated target genomic regions at a threshold level above the presence of the cancer; (b) the one or more target genomic regions include one or more target sequences of one or more genes selected from Table 1; (c) the length of each target sequence is at least 25 nucleotides; (d) the cancer is bladder cancer, prostate cancer, or kidney cancer; and (e) the treatment includes surgical resection, radiation therapy, chemotherapy, immunotherapy, or any combination thereof.

[0315] A classifier (as described herein) can be used to determine the likelihood or probability score that a sample feature vector is from a subject having a cancer. In one embodiment, when the likelihood or probability exceeds a threshold (e.g., the level of a reference sample from a subject having a cancer), an appropriate treatment (e.g., a resection surgery or a therapeutic agent) is prescribed. For example, in one embodiment, if the likelihood or probability score is greater than or equal to 60, one or more appropriate treatments are prescribed. In other embodiments, if the likelihood or probability score is greater than or equal to 65, greater than or equal to 70, greater than or equal to 75, greater than or equal to 80, greater than or equal to 85, greater than or equal to 90, or greater than or equal to 95, one or more appropriate treatments are prescribed. In other embodiments, the cancer log odds ratio can indicate the effectiveness of cancer treatment. For example, over time (e.g., at one second after treatment), an increase in the cancer log odds ratio can indicate ineffective treatment. Similarly, over time (e.g., at one second after treatment), a decrease in the cancer log odds ratio can indicate successful treatment. In another embodiment, if the cancer log odds ratio is greater than 1, greater than 1.5, greater than 2, greater than 2.5, greater than 3, greater than 3.5, or greater than 4, one or more appropriate treatments are prescribed. In some embodiments, the threshold level of cancer presence is determined by a classifier trained on sequencing reads of converted DNA from subjects having a cancer. Non-limiting examples of classifiers are described herein. Classification can be based on one or more target genomic regions, as described herein with respect to various aspects of the disclosure.

[0316] In some embodiments, the treatment is one or more cancer therapeutic agents selected from the group consisting of: chemotherapeutic agents, targeted cancer therapeutics, differentiation therapeutics, hormonal therapeutics, and immunotherapeutics. For example, the treatment can be one or more chemotherapeutic agents selected from the group consisting of: alkylating agents, antimetabolites, anthracyclines, antitumor antibiotics, cytoskeletal disruptors (taxanes), topoisomerase inhibitors, mitotic inhibitors, corticosteroids, kinase inhibitors, nucleotide analogs, platinum-based agents, and any combination thereof. In some embodiments, the treatment is one or more targeted cancer therapeutics selected from the group consisting of: signal transduction inhibitors (e.g., tyrosine kinase and growth factor receptor inhibitors), histone deacetylase (HDAC) inhibitors, retinoic acid receptor agonists, proteasome inhibitors, angiogenesis inhibitors, and monoclonal antibody conjugates. In some embodiments, the treatment is one or more differentiation therapeutics comprising retinoids, such as tretinoin, alitretinoin, and bexarotene. In some embodiments, the treatment is one or more hormonal therapeutics selected from the group consisting of: antiestrogens, aromatase inhibitors, progesterones, estrogens, antiandrogens, and GnRH agonists or analogs. In one embodiment, the treatment is one or more immunotherapeutics selected from the group comprising: monoclonal antibody therapies (such as rituximab (RITUXAN) and alemtuzumab (CAMPATH)), nonspecific immunotherapies and adjuvants (such as BCG, interleukin-2 (IL-2), and interferon-α), immunomodulatory drugs (e.g., thalidomide and lenalidomide (REVLIMID)). An experienced physician or oncologist can select an appropriate cancer therapeutic agent based on a variety of characteristics, such as tumor type, cancer stage, prior exposure to cancer treatments or therapeutic agents, and other characteristics of the cancer.

[0317] Computer systems and devices

[0318] In one aspect, the present disclosure provides a computer system for performing one or more steps of the methods disclosed herein. In another aspect, the present disclosure provides a non-transitory computer-readable medium having stored thereon computer-readable instructions for performing one or more steps of the methods disclosed herein.

[0319] The methods of the present disclosure can be performed using software, hardware, firmware, hardwiring, or any combination thereof. The features that implement the functionality can also be physically located at various positions, including distributed, such that parts of the functionality are implemented at different physical locations (e.g., an imaging device in one room and a host workstation in another room, or in different buildings, e.g., via a wireless or wired connection).

[0320] Processors suitable for executing computer programs include, by way of example, general and special-purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for executing the instructions and one or more memory devices for storing the instructions and data. Generally, a computer will also include or be operatively coupled to receive data from or transfer data to or both: one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks. Information carriers suitable for embodying computer program instructions and data include all forms of non-volatile memory, by way of example including semiconductor storage devices (such as EPROM, EEPROM, solid state drives (SSD) and flash memory devices); magnetic disks (such as internal hard disks or removable disks); magneto-optical disks; and optical disks (such as CD and DVD disks). The processor and the memory may be supplemented by, or incorporated in, special purpose logic circuitry.

[0321] For providing interaction with a user, the subject matter described herein may be implemented on a computer having I / O devices such as a CRT, LCD, LED, or projection device for displaying information to the user, and input or output devices such as a keyboard and a pointing device (such as a mouse or a trackball) by which the user may provide input to the computer. Other kinds of devices may also be used to provide interaction with the user. For example, feedback provided to the user may be any form of sensory feedback (such as visual feedback, auditory feedback, or tactile feedback), and input received from the user may be in any form, including acoustic, speech, or tactile input.

[0322] The subject matter described herein may be implemented in a computing system that includes backend components (such as data servers), middleware components (such as application servers), or frontend components (such as client computers having a graphical user interface or a web browser through which users may interact with an implementation of the subject matter described herein), or any combination of such backend, middleware, and frontend components. The components of the system may be interconnected by any form or medium of a digital data communication network (such as a communication network). For example, a reference data set may be stored at a remote location and a computer may communicate across the network to access the reference data set for comparison purposes. However, in other embodiments, the reference data set may be stored locally on the computer and the computer accesses the reference data set within the CPU for comparison purposes. Examples of communication networks include, but are not limited to, cellular networks (such as 3G or 4G), local area networks (LAN), and wide area networks (WAN), such as the Internet.

[0323] The subject matter described herein can be implemented as one or more computer program products, e.g., one or more computer programs tangibly embodied in an information carrier (e.g., in a non-transitory computer-readable medium) for execution by or to control the operation of a data processing apparatus (e.g., a programmable processor, a computer, or multiple computers). A computer program (also referred to as a program, software, software application, application program, macro, or code) can be written in any form of programming language, including compiled or interpreted languages (e.g., C, C++, Perl), and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. The systems and methods of the present disclosure can include instructions written in any suitable programming language known in the art, including but not limited to C, C++, Perl, Java, ActiveX, HTML5, VisualBasic, or JavaScript.

[0324] A computer program does not necessarily correspond to a file. A program can be stored in a file or part of a file that holds other programs or data, in a single file dedicated to the program, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or multiple computers at a site, or distributed across multiple sites and interconnected via a communication network.

[0325] A file can be a digital file, e.g., stored on a hard disk, SSD, CD, or other tangible, non-transitory medium. A file can be sent from one device to another via a network (e.g., as a data packet from a server to a client, e.g., via a network interface card, modem, wireless network card, or the like).

[0326] Writing a file according to the present disclosure involves transforming a tangible, non-transitory computer-readable medium, e.g., by adding, removing, or rearranging particles (e.g., converting particles with a net charge or dipole moment into a magnetization pattern by a read / write head), and these patterns subsequently represent a new combination of information about objective physical phenomena that is desired by and useful to the user. In some embodiments, writing involves physically transforming the material in a tangible, non-transitory computer-readable medium (e.g., having certain optical properties so that an optical read / write device can read the new and useful combination of information, e.g., burning a CD-ROM). In some embodiments, writing a file includes transforming a physical flash device (e.g., a NAND flash device) and storing information by transforming physical elements in an array of memory cells made of floating-gate transistors. Methods of writing files are well known in the art and can be invoked manually or automatically, e.g., by a program or by a save command of software or a write command of a programming language.

[0327] Suitable computing devices typically include a mass storage, at least one graphical user interface, at least one display device, and typically include communication between devices. The mass storage is a type of computer-readable medium, i.e., a computer storage medium. The computer storage medium may include volatile, non-volatile, removable, and non-removable media, which are implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include RAM, ROM, EEPROM, flash memory or other storage technologies, CD-ROM, digital versatile disc (DVD) or other optical memory, magnetic tape cartridges, tapes, magnetic disk storage or other magnetic storage devices, radio frequency identification (RFID) tags or chips, or any other medium that can be used to store the desired information and can be accessed by a computing device.

[0328] The functions described herein can be implemented using software, hardware, firmware, hardwiring, or any combination thereof. Any of the software can be physically located in various positions, including distributed, such that portions of the functionality are implemented in different physical locations.

[0329] As will be recognized by those skilled in the art as necessary or most appropriate for the execution of the disclosed methods, a computer system for implementing some or all of the inventive methods may include one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), or both), a main memory, and a static memory, which communicate with each other via a bus.

[0330] The processor will typically include a chip, such as a single-core or multi-core chip, to provide a central processing unit (CPU). The processor can be provided by a chip from Intel or AMD.

[0331] The memory may include one or more machine-readable devices on which one or more instruction sets (e.g., software) are stored, which, when executed by one or more processors of any of the disclosed computers, can implement some or all of the methods or functions described herein. The software may also reside entirely or at least partially in the main memory and / or the processor during execution by the computer system. Preferably, each computer includes non-transitory memory, such as a solid-state drive, a flash drive, a disk drive, a hard disk drive, etc.

[0332] While in the exemplary embodiments, the machine-readable device may be a single medium, the term "machine-readable device" should be understood to include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) storing one or more instructions and / or data sets. These terms should also be understood to include any one or more media capable of storing, encoding, or preserving a set of instructions executable by a machine and causing the machine to perform any one or more of the methods disclosed herein. Thus, these terms should be understood to include, but not be limited to, one or more solid-state memories (e.g., a subscriber identity module (SIM) card, a secure digital card (SD card), a micro SD card, or a solid-state drive (SSD)), optical and magnetic media, and / or any other tangible storage medium.

[0333] The computers disclosed herein will generally include one or more I / O devices, such as, for example, one or more video display units (e.g., a liquid crystal display (LCD) or a cathode ray tube (CRT)), alphanumeric input devices (e.g., a keyboard), cursor control devices (e.g., a mouse), disk drive units, signal generation devices (e.g., a speaker), touchscreens, accelerometers, microphones, cellular radio frequency antennas, and network interface devices (which may be, for example, a network interface card (NIC), a Wi-Fi card, or a cellular modem).

[0334] Any of the software may be physically located in various places, including being distributed, such that portions of the functionality are implemented in different physical locations.

[0335] In addition, the systems disclosed herein may be provided to include reference data. Any suitable genomic data may be stored in the system for use. Examples include, but are not limited to: comprehensive, multi-dimensional maps of key genomic changes in the major types and subtypes of cancers from The Cancer Genome Atlas (TCGA); catalogs of genomic abnormalities from the International Cancer Genome Consortium (ICGC); catalogs of somatic mutations in cancers from COSMIC; the latest versions of the human genome and other common model organisms; the latest reference SNPs from dbSNP; the gold standard insertions and deletions from the 1000 Genomes Project and the Broad Institute; exon capture kit annotations from Illumina, Inc., Agilent Technologies, Inc., Nimblegen, and Ion Torrent; transcript annotations; small test data for pipeline experiments (e.g., for new users).

[0336] In some embodiments, data is available in the context of databases included in the system. Any suitable database structure can be used, including relational databases, object-oriented databases, and others. In some embodiments, reference data is stored in a relational database (e.g., a "Not Only SQL" (NoSQL) database). In various embodiments, a graph database is included in the system disclosed herein. It should also be understood that the term "database" as used herein is not limited to a single database; rather, multiple databases can be included in the system. For example, according to embodiments of the present disclosure, the database can include two, three, four, five, six, seven, eight, nine, ten, fifteen, twenty, or more individual databases, including any integer number of databases therebetween. For example, one database can contain common reference data, a second database can contain test data from patients, a third database can contain data from healthy subjects, and a fourth database can contain data from diseased subjects suffering from a known disease or disorder. It should be understood that any other database configuration regarding the data contained therein is also contemplated by the methods described herein.

[0337] Exemplary embodiments

[0338] The present disclosure provides the following exemplary embodiments:

[0339] Example 1. A method for sequencing free nucleic acid molecules of a subject, the method comprising:

[0340] Processing a urine sample to inhibit cell lysis;

[0341] Separating free nucleic acid molecules in the processed urine sample from the cells in the processed urine sample, thereby producing a purified urine sample containing these free nucleic acid molecules;

[0342] Concentrating the free nucleic acid molecules in the purified urine sample by passing at least a portion of the purified urine sample through a filter, wherein (i) the concentration produces a filtrate and a retained urine sample, and (ii) the retained urine sample contains free nucleic acid molecules with increased concentration;

[0343] Isolating free nucleic acid molecules from the retained urine sample; and

[0344] Sequencing the isolated free nucleic acid molecules.

[0345] Example 2. The method according to Example 1, wherein processing the urine sample to inhibit cell lysis includes contacting the urine sample with one or more preservative reagents.

[0346] Example 3. The method according to Example 1 or 2, wherein the processing includes treating with a nuclease inhibitor, a formaldehyde quencher, or both.

[0347] Example 4. The method according to any one of Examples 1-3, wherein the treatment comprises contacting the urine sample with a composition comprising: (i) imidazolidinyl urea, EDTA, glycine, or a combination thereof; or (ii) sodium azide, EDTA, or a combination thereof.

[0348] Example 5. The method according to any one of Examples 1-4, wherein the separation comprises centrifuging to precipitate cells in the treated urine sample.

[0349] Example 6. The method according to any one of Examples 1-5, wherein the filter is substantially impermeable to the passage of free nucleic acids and substantially permeable to salts in the purified urine sample.

[0350] Example 7. The method according to any one of Examples 1-5, wherein the filter has a nominal molecular weight cut-off of 10 kD, 5 kD, 3 kD, or lower.

[0351] Example 8. The method according to any one of Examples 1-7, wherein the retained urine sample has a concentration that is at least 2-fold, at least 5-fold, at least 10-fold, or at least 15-fold higher compared to the purified urine sample.

[0352] Example 9. The method according to any one of Examples 1-8, wherein the retained urine sample has a volume that is at least 50%, at least 75%, or at least 90% lower compared to the volume of the treated urine sample.

[0353] Example 10. The method according to Example 9, wherein the volume of the treated urine sample is 5 mL, 10 mL, 15 mL, 20 mL, 30 mL, 40 mL, 50 mL, or more.

[0354] Example 11. The method according to any one of Examples 1-10, further comprising freezing the retained urine sample.

[0355] Example 12. The method according to any one of Examples 1-11, wherein the treatment is completed within 120, 60, or 30 minutes after collection of the urine sample; and optionally wherein the separation and the concentration are completed within 7 days after collection.

[0356] Example 13. The method according to any one of Examples 1-12, further comprising amplifying one or more of the isolated free nucleic acid molecules.

[0357] Example 14. The method according to any one of Examples 1-13, further comprising capturing the isolated free nucleic acid molecules or their amplification products by hybridization with a bait oligonucleotide.

[0358] Example 15. The method according to Example 14, further comprising separating the free nucleic acid molecules bound to the bait from the unbound free nucleic acid molecules.

[0359] Example 16. The method according to Example 15, wherein each bait oligonucleotide hybridizes to a target genomic region that is differentially methylated in a cancer sample relative to a non-cancer sample.

[0360] Example 17. The method according to Example 16, wherein the differential methylation comprises at least 80% of the CpG sites in the target genomic region being methylated or unmethylated.

[0361] Example 18. The method according to Example 16 or 17, wherein the cancer is bladder cancer, prostate cancer, or kidney cancer.

[0362] Example 19. The method according to any one of Examples 14-18, wherein each bait oligonucleotide hybridizes to a target genomic region comprising at least five methylation sites.

[0363] Example 20. The method according to any one of Examples 14-19, wherein each bait oligonucleotide hybridizes to a target genomic region comprising a target sequence of a gene selected from Table 1, and wherein the length of the target sequence is at least 25, at least 35, or at least 45 nucleotides.

[0364] Example 21. The method according to Example 20, wherein the target genomic region comprises a target sequence of a gene selected from the following: TWIST1, EOMES, HOXA9, POU4F2, and ZNF154.

[0365] Example 22. The method according to Example 20, wherein the bait oligonucleotides together hybridize to target sequences of at least 10 genes from Table 1.

[0366] Example 23. The method according to Example 20, wherein the bait oligonucleotides together hybridize to target sequences from: (a) the genes in Table 2 or Table 3; (b) the genes in Table 4; or (c) the genes in Table 5.

[0367] Example 24. The method according to any one of Examples 1-23, wherein the free nucleic acid molecules comprise cell-free DNA (cfDNA).

[0368] Example 25. The method according to Example 24, wherein the method further comprises deaminating the cfDNA isolated in step (d) to produce converted cfDNA molecules; optionally wherein the deamination comprises treatment with a cytidine deaminase or bisulfite.

[0369] Example 26. The method according to any one of Examples 1-25, wherein the method further comprises diagnosing cancer in the subject.

[0370] Example 27. The method according to Example 26, wherein the cancer is bladder cancer, prostate cancer or kidney cancer.

[0371] Example 28. The method according to Example 26 or 27, wherein the method further comprises treating the cancer in the subject.

[0372] Example 29. The method according to Example 28, wherein the treatment comprises surgical resection, radiotherapy, chemotherapy and / or immunotherapy.

[0373] Example 30. A method for detecting cancer cells in a subject, the method comprising:

[0374] Capturing transformed cell-free DNA (cfDNA) fragments or amplification products thereof from a urine sample of the subject, wherein:

[0375] The bait oligonucleotide composition comprises a plurality of different bait oligonucleotides;

[0376] Each bait oligonucleotide of the plurality of different bait oligonucleotides hybridizes to a target sequence of a gene selected from Table 1, wherein the length of the target sequence is at least 25 nucleotides;

[0377] Separating the bait-bound DNA from the unbound DNA;

[0378] Sequencing the separated DNA to generate sequencing reads; and

[0379] Detecting these cancer cells with a trained classifier,

[0380] Example 31. The method according to Example 30, wherein for one or more of the target sequences identified as hypermethylated and / or hypomethylated in these cfDNA fragments, the trained classifier detects that the number of sequencing reads exceeds a threshold.

[0381] Example 32. The method according to Example 30 or 31, wherein the length of these bait oligonucleotides is at least 45 nucleotides.

[0382] Example 33. The method according to any one of Examples 30-32, wherein the trained classifier differentiates between subjects with cancer and subjects without cancer with a defined specificity.

[0383] Example 34. The method according to Example 33, wherein the classifier is a hybrid model classifier.

[0384] Example 35. The method according to Example 33 or 34, wherein the defined specificity is 0.900 or higher.

[0385] Example 36. The method according to Example 35, wherein the application of the trained classifier further comprises a sensitivity of 30% or higher.

[0386] Example 37. The method according to any one of Examples 30 - 36, wherein these decoy oligonucleotides hybridize in common with target sequences from at least 10 genes in Table 1.

[0387] Example 38. The method according to any one of Examples 30 - 36, wherein at least one of these decoy oligonucleotides hybridizes with a target sequence of a gene selected from: TWIST1, EOMES, HOXA9, POU4F2, and ZNF154.

[0388] Example 39. The method according to any one of Examples 30 - 36, wherein:

[0389] these cancer cells are bladder cancer cells, and these decoy oligonucleotides hybridize in common with target sequences of genes from Table 2 or Table 3;

[0390] these cancer cells are prostate cancer cells, and these decoy oligonucleotides hybridize in common with target sequences of genes from Table 4; or

[0391] these cancer cells are kidney cancer cells, and these decoy oligonucleotides hybridize in common with target sequences of genes from Table 5.

[0392] Example 40. The method according to any one of Examples 30 - 39, wherein these transformed cfDNA molecules comprise cfDNA treated with cytidine deaminase or bisulfite.

[0393] Example 41. The method according to any one of Examples 30 - 40, wherein each decoy oligonucleotide is conjugated to a solid surface or a non - nucleotide affinity moiety.

[0394] Example 42. The method according to any one of Examples 30 - 41, wherein the differential methylation comprises hypermethylation in the cancer sample relative to the non - cancer sample.

[0395] Example 43. The method according to any one of Examples 30 - 42, wherein each target genomic region comprises at least five methylation sites.

[0396] Example 44. The method according to any one of Examples 30 - 43, wherein the method further comprises obtaining the transformed cfDNA fragments or their amplification products, and wherein the obtaining further comprises: (i) treating the urine sample to inhibit cell lysis; (ii) separating the cfDNA fragments in the treated urine sample from the cells in the treated urine sample, thereby producing a purified urine sample containing the cfDNA fragments; (iii) concentrating the cfDNA fragments in the purified urine sample by passing at least a portion of the purified urine sample through a filter, wherein the concentration produces a filtrate and a retained urine sample, and wherein the retained urine sample contains an increased concentration of cfDNA fragments; and (iv) isolating the cfDNA fragments from the retained urine sample.

[0397] Example 45. The method according to Example 44, further comprising (v) amplifying one or more of the isolated cfDNA fragments.

[0398] Example 46. The method according to Example 44 or 45, wherein treating the urine sample to inhibit cell lysis comprises contacting the urine sample with one or more anti - preservative reagents.

[0399] Example 47. The method according to any one of Examples 44 - 46, wherein the treatment comprises treating with a nuclease inhibitor, a formaldehyde quencher, or both.

[0400] Example 48. The method according to any one of Examples 44 - 47, wherein the treatment comprises contacting the urine sample with a composition comprising: (a) imidazolidinyl urea, EDTA, glycine, or a combination thereof; or (b) sodium azide, EDTA, or a combination thereof.

[0401] Example 49. The method according to any one of Examples 44 - 48, wherein the separation comprises centrifuging to precipitate the cells in the treated urine sample.

[0402] Example 50. The method according to any one of Examples 44 - 49, wherein the filter is substantially impermeable to the passage of free nucleic acids and substantially permeable to salts in the purified urine sample.

[0403] Example 51. The method according to any one of Examples 44 - 49, wherein the filter has a nominal molecular weight cut - off of 10 kD, 5 kD, 3 kD, or lower.

[0404] Example 52. The method according to any one of Examples 44 - 51, wherein the retained urine sample has a concentration that is at least 2 - fold, at least 5 - fold, at least 10 - fold, or at least 15 - fold increased compared to the purified urine sample.

[0405] Example 53. The method according to any one of Examples 44-52, wherein the retained urine sample has a volume that is at least 50%, at least 75%, or at least 90% lower than the volume of the processed urine sample.

[0406] Example 54. The method according to Example 53, wherein the volume of the processed urine sample is 5 mL, 10 mL, 15 mL, 20 mL, 30 mL, 40 mL, 50 mL or more.

[0407] Example 55. The method according to any one of Examples 44-54, further comprising freezing the retained urine sample.

[0408] Example 56. The method according to any one of Examples 44-55, wherein the treatment is completed within 120, 60, or 30 minutes after collection of the urine sample; and optionally wherein the separation and concentration are completed within 7 days after collection.

[0409] Example 57. The method according to any one of Examples 30-56, wherein the method further comprises diagnosing cancer in the subject.

[0410] Example 58. The method according to Example 57, wherein the cancer is bladder cancer, prostate cancer, or kidney cancer.

[0411] Example 59. The method according to Example 57 or 58, wherein the method further comprises treating the cancer in the subject.

[0412] Example 60. The method according to Example 59, wherein the treatment comprises surgical resection, radiotherapy, chemotherapy, and / or immunotherapy.

[0413] Example 61. A method of treating cancer in a subject, the method comprising selecting a subject having cancer or an increased risk of developing cancer and administering a treatment to the subject, wherein:

[0414] the selection comprises identifying the subject as the source of a urine cell-free DNA (cfDNA) sample that contains one or more differentially methylated target genomic regions above a threshold level for the presence of the cancer;

[0415] the one or more target genomic regions comprise one or more target sequences of one or more genes selected from Table 1;

[0416] the length of each target sequence is at least 25 nucleotides;

[0417] the cancer is bladder cancer, prostate cancer, or kidney cancer; and

[0418] The treatment includes surgical resection, radiotherapy, chemotherapy, immunotherapy, or any combination thereof.

[0419] Example 62. The method according to Example 61, wherein the threshold level at which the cancer is present is the level of a reference sample from a subject having the cancer.

[0420] Example 63. The method according to Example 61 or 62, wherein the threshold level at which the cancer is present is determined by a classifier trained on sequencing reads of the converted DNA from a subject having the cancer

[0421] Example 64. The method according to any one of Examples 61-63, wherein the one or more target genomic regions comprise target sequences from at least 10 genes in Table 1.

[0422] Example 65. The method according to any one of Examples 61-64, wherein the one or more target genomic regions comprise target sequences of genes selected from: TWIST1, EOMES, HOXA9, POU4F2, and ZNF154.

[0423] Example 66. The method according to any one of Examples 61-64, wherein:

[0424] the cancer is bladder cancer, and the one or more target genomic regions comprise one or more target sequences of one or more genes selected from Table 2 or Table 3;

[0425] the cancer is prostate cancer, and the one or more target genomic regions comprise one or more target sequences of one or more genes selected from Table 4; or

[0426] the cancer is kidney cancer, and the one or more target genomic regions comprise one or more target sequences of one or more genes selected from Table 5.

[0427] Example 67. The method according to any one of Examples 61-66, wherein each target genomic region comprises at least five methylation sites.

[0428] Example 68. A composition comprising a plurality of different bait oligonucleotides, wherein:

[0429] the bait oligonucleotides hybridize to converted DNA molecules derived from one or more target genomic regions;

[0430] the one or more target genomic regions comprise one or more target sequences of one or more genes selected from Table 1;

[0431] the one or more target genomic regions are differentially methylated in cancer; and

[0432] Each decoy oligonucleotide comprises a sequence of at least 25 nucleotides in length that hybridizes to one of these target sequences.

[0433] Example 69. The method according to Example 68, wherein the one or more target genomic regions comprise target sequences from at least 10 genes selected from Table 1.

[0434] Example 70. The method according to Example 68 or Example 69, wherein the one or more target genomic regions comprise target sequences of genes selected from: TWIST1, EOMES, HOXA9, POU4F2, and ZNF154.

[0435] Example 71. The method according to Example 68 or Example 69, wherein the one or more target genomic regions comprise target sequences from: (a) the genes in Table 2 or Table 3; (b) the genes in Table 4; or (c) the genes in Table 5.

[0436] Example 72. The method according to any one of Examples 68 - 71, wherein each target genomic region comprises at least five methylation sites.

[0437] Example 73. The method according to any one of Examples 68 - 72, wherein the differential methylation comprises at least 80% of the CpG sites in the target genomic region being methylated or unmethylated.

[0438] Example

[0439] The following examples are presented to provide a complete disclosure and description of how to make and use the present specification to those of ordinary skill in the art, and are not intended to limit the scope of what the inventors regard as their specification, nor are they intended to represent that the following experiments are all or the only experiments conducted. Efforts have been made to ensure the accuracy of the numbers used (e.g., amounts, temperatures, etc.), but some experimental error and deviation should be accounted for.

[0440] Example 1 - Urine Sample Processing Workflow

[0441] Urological cancers such as prostate cancer, bladder cancer, and kidney cancer have low detection sensitivity in the Circulating Cell-free Genome Atlas study ("CCGA"; ClinicalTrials.gov identifier NCT02889978), which is a prospective, multi-center, case-control, observational study with longitudinal follow-up ( Figure 3 ). This low detection rate may be due to the low tumor fraction of cfDNA in the blood of subjects with urothelial cancer. Analyzing cell-free DNA from urine can improve the detection sensitivity of urological cancers. Therefore, improved methods for preserving and extracting cfDNA from urine have been developed.

[0442] Figure 1 Shows an exemplary urine sample processing workflow. In this workflow, approximately 50 mL of urine is collected from a subject. After collection, a preservative is added to the urine sample. Streck urine preservative (Streck, Nebraska, USA) or an equivalent preservative can be used as the preservative, and the equivalent preservative contains at least 0.5% weight / volume (w / v) nuclease inhibitor, at least 0.2%-4.0% w / v preservative agent, and at least 0.01% w / v formaldehyde quencher. Other available preservatives include urine collection media and UAS (Novosantis, Belgium). Alternatively, a urine collection and preservation device (Norgen Biotek Corp., Canada) or an equivalent cup can be used to collect the urine sample, and the equivalent cup contains 10%-30% w / v nuclease inhibitor (such as EDTA) and 0.1%-1.0% w / v antibacterial preservative agent (such as sodium azide).

[0443] After adding the preservative, the urine sample is centrifuged at 4000 x g for 20 minutes to precipitate and remove any cell debris. Then the resulting supernatant is concentrated approximately 15-fold. Concentration of the preserved urine sample can be achieved using a filtration column (e.g., a filtration device with a regenerated cellulose membrane having a 3 kDa cut-off). For example, if the resulting supernatant has a volume of 15 mL or less, the sample can be concentrated by spinning at 4000 x g for at least 30 minutes in an Amicon Ultra-15 filtration device (Thermo Fisher Scientific, Massachusetts, USA). If a supernatant with a volume range of 15 mL to 50 mL is used, the sample can be alternatively concentrated by spinning at 2500 x g for at least 40 minutes or until the sample has been concentrated to a volume below 4.2 mL in a Centricon Plus-70 centrifugal filtration device (Thermo Fisher Scientific, Massachusetts, USA). Concentrating the urine sample to a lower volume makes the sample suitable for automated bead-based extraction methods (e.g., MagMax extraction) and other bench-top feasible techniques.

[0444] After concentrating the urine sample, the sample can be used immediately or frozen at -80 °C for subsequent use or for batch processing of samples. Then cfDNA extraction and library preparation can be performed on the concentrated urine sample, and then it can be sequenced for methylation analysis and urinary system cancer detection (Figure 2 and Figure 4 )

[0445] Using Streck urine preservative, Figure 1The workflow shown in Figure 5A was applied to urine samples. Adding preservatives within 30 minutes after urine sample collection can preserve nucleosomes and generate cfDNA fragments up to 7000 bp in size( Figure 5B ). In contrast, delaying the addition of preservatives until one hour or more after collection results in the loss of the nucleosome peak at >700 bp, a significant decrease in yield, and a narrower distribution of fragment lengths biased towards low molecular weight fragments, indicating cell lysis in urine, accompanied by degradation and fragmentation of cfDNA(

[0446] Example 2 – Analysis of tumor fraction in urine-derived cfDNA

[0447] Cancer-specific methylation signatures were detected in urine cfDNA, such as those corresponding to samples from subjects with stage I high-grade non-muscle-invasive bladder cancer, and this urine cfDNA was processed from urine samples as described in Example 1( Figures 6A - 6B ).

[0448] Further studies were conducted to evaluate the detection of cancer-specific methylation markers in urine-derived cfDNA, as compared to plasma tumor fraction estimates. Urine and blood were collected from patients with bladder, kidney, and prostate cancer, as well as age- and sex-matched non-cancer patients. To generate biopsy-free estimates in urine cfDNA, the plasma-based workflow was modified as follows (shown in Figure 2A ): (1) an external reference dataset of non-cancer urine cfDNA (N = ~200) was used in the workflow instead of plasma for non-cancer WGBS and TM data; (2) the noise threshold and pseudocounts were adjusted to accommodate the smaller reference dataset; and (3) WGBS data from healthy urinary tract tissue was used to further filter out noise methylation variants.

[0449] Sequencing libraries were then prepared from the resulting cfDNA for methylation analysis of a panel of urinary tract cancer methylation markers. The panel of methylation markers was also used for tumor fraction estimation, and samples with an estimated tumor fraction above the threshold were identified as detecting cancer based on urine-derived cfDNA. The estimated tumor fraction in urine-derived cfDNA samples was then compared to the estimated tumor fraction from the corresponding plasma-derived cfDNA fraction.

[0450] Figures 7 - 9The scatter plot shows the distribution of tumor fractions for each cancer type in matched urine and plasma cfDNA (each point is an individual patient). The fill indicates whether the plasma cfDNA of that patient was detected by the multi-cancer classifier (99% specificity). An increase in the tumor fraction in urine cfDNA relative to plasma was observed in all bladder cancer patients analyzed and in a subset of prostate cancer patients. The plasma tumor fraction estimates were consistent with classifier detection. In renal cancer patients, although the signal in urine was not increased relative to that in plasma, a higher tumor fraction in urine cfDNA was observed relative to non-cancer patients. In a few patients, an increase in the signal in plasma relative to urine cfDNA was observed.

[0451] The performance of bladder cancer detection based on subsets of genomic regions within the methylation marker set was also evaluated. High sensitivity and specificity were observed when detecting bladder cancer status based on a subset of 15 methylation markers ( Figure 10A ). Additionally, high sensitivity and specificity were observed even when determining bladder cancer status based only on methylation markers within a single gene (TWIST1) ( Figure 10B ).

[0452] Furthermore, the performance of renal cancer and prostate cancer detection based on subsets of genomic regions within the methylation marker set was evaluated. The results showed that urine-derived cfDNA diagnosed renal cancer ( Figure 10C ) and prostate cancer ( Figure 10D ). Genomic regions of the genes in Tables 2 and 3 were found to contain methylation markers for bladder cancer. Genomic regions of the genes in Table 4 were found to contain methylation markers for prostate cancer. Genomic regions of the genes in Table 5 were found to contain methylation markers for renal cancer. Table 1 presents the union of the genes in Tables 2 - 5. In this example, the genomic region of a gene is considered to be the sequence from the (actual or putative) transcription start site to the transcription termination site, and an additional 5000 nucleotides from each of these ends. Although the target genomic regions of urine cfDNA samples have been identified, such markers can also be used for the classification of other cfDNA samples (e.g., cfDNA from other body fluids such as blood, serum, or plasma).

[0453] In summary, determining urinary system cancer status based on methylation marker and tumor fraction analysis is comparable or more accurate when analyzing urine-derived cfDNA compared to plasma-derived cfDNA. Concentrating urine samples can improve the handling and analysis of urine-derived cfDNA for urinary system cancer detection.

Claims

1. A method for sequencing free nucleic acid molecules of a subject, characterized in that: The method includes: (a) Processing a urine sample to inhibit cell lysis; (b) Separating free nucleic acid molecules in the processed urine sample from the cells in the processed urine sample, thereby producing a purified urine sample containing these free nucleic acid molecules; (c) Concentrating these free nucleic acid molecules in the purified urine sample by passing at least a portion of the purified urine sample through a filter, wherein (i) the concentration produces a filtrate and a retained urine sample, and (ii) the retained urine sample contains free nucleic acid molecules with increased concentration; (d) Isolating free nucleic acid molecules from the retained urine sample; and (e) Sequencing the isolated free nucleic acid molecules.

2. The method according to claim 1, characterized in that: Processing the urine sample to inhibit cell lysis includes contacting the urine sample with one or more preservative reagents.

3. The method according to claim 1, characterized in that: The processing includes treating with a nuclease inhibitor, a formaldehyde quencher, or both.

4. The method according to claim 1, characterized in that: The processing includes contacting the urine sample with a composition comprising: (i) imidazolidinyl urea, EDTA, glycine, or a combination thereof; or (ii) sodium azide, EDTA, or a combination thereof.

5. The method according to claim 1, characterized in that: The separation includes centrifuging to precipitate the cells in the processed urine sample.

6. The method according to claim 1, wherein: The filter is substantially impermeable to the passage of free nucleic acids and substantially permeable to salts in the purified urine sample.

7. The method according to claim 1, wherein: The filter has a nominal molecular weight cut-off of 10 kD, 5 kD, 3 kD, or lower.

8. The method according to claim 1, characterized in that: The retained urine sample has a concentration increased by at least 2-fold, at least 5-fold, at least 10-fold, or at least 15-fold compared to the purified urine sample.

9. The method according to claim 1, characterized in that: The retained urine sample has a volume that is at least 50%, at least 75%, or at least 90% lower than the volume of the processed urine sample.

10. The method according to claim 9, wherein: The volume of the processed urine sample is 5 mL, 10 mL, 15 mL, 20 mL, 30 mL, 40 mL, 50 mL, or more.

11. The method according to claim 1, characterized in that: The method further includes freezing the retained urine sample.

12. The method according to claim 1, wherein: The processing is completed within 120, 60, or 30 minutes after collection of the urine sample; and optionally wherein the separation and the concentration are completed within 7 days after collection.

13. The method according to claim 1, characterized in that: The method further includes amplifying one or more of the isolated free nucleic acid molecules.

14. The method according to claim 1, characterized in that: The method further includes capturing the isolated free nucleic acid molecules or their amplification products by hybridization with bait oligonucleotides.

15. The method according to claim 14, characterized in that: The method further includes separating the bait-bound free nucleic acid molecules from the unbound free nucleic acid molecules.

16. The method according to claim 15, wherein: Each bait oligonucleotide hybridizes to a target genomic region that is differentially methylated in a cancer sample relative to a non-cancer sample.

17. The method according to claim 16, characterized in that: The differential methylation includes at least 80% of the CpG sites in the target genomic region being methylated or unmethylated.

18. The method according to claim 16, characterized in that: The cancer is bladder cancer, prostate cancer, or kidney cancer.

19. The method according to claim 14, characterized in that: Each bait oligonucleotide hybridizes to a target genomic region containing at least five methylation sites.

20. The method according to claim 14, wherein: Each bait oligonucleotide hybridizes to a target genomic region containing a target sequence of a gene selected from Table 1, and wherein the length of the target sequence is at least 25, at least 35, or at least 45 nucleotides.

21. The method according to claim 20, wherein: The target genomic region contains target sequences of genes selected from: TWIST1, EOMES, HOXA9, POU4F2, and ZNF154.

22. The method according to claim 20, wherein: These bait oligonucleotides hybridize in common with target sequences of at least 10 genes from Table 1.

23. The method according to claim 20, wherein: These bait oligonucleotides hybridize in common with target sequences from: (a) genes in Table 2 or Table 3; (b) genes in Table 4; or (c) genes in Table 5.

24. The method according to any one of claims 1-23, characterized in that: These free nucleic acid molecules contain cell-free DNA (cfDNA).

25. The method according to claim 24, wherein: The method further includes deaminating the cfDNA isolated in step (d) to produce converted cfDNA molecules; optionally wherein the deamination includes treatment with a cytidine deaminase or bisulfite.

26. The method according to claim 1, characterized in that: The method further includes diagnosing cancer in the subject.

27. The method according to claim 26, wherein: The cancer is bladder cancer, prostate cancer, or kidney cancer.

28. The method according to claim 26 or 27, characterized in that: The method further includes treating the cancer in the subject.

29. The method according to claim 28, characterized in that: The treatment includes surgical resection, radiotherapy, chemotherapy, and / or immunotherapy.

30. A method for detecting cancer cells in a subject, characterized in that: The method includes: (a) Capturing converted cell-free DNA (cfDNA) fragments or their amplification products from a urine sample of the subject, wherein: (i) The bait oligonucleotide composition contains a plurality of different bait oligonucleotides; (ii) Each of the plurality of different bait oligonucleotides hybridizes with a target sequence of a gene selected from Table 1, wherein the length of the target sequence is at least 25 nucleotides; (b) Separating the bait-bound DNA from the unbound DNA; (c) Sequencing the separated DNA to produce sequencing reads; and (d) Detecting these cancer cells with a trained classifier.

31. The method according to claim 30, wherein: For one or more of the target sequences identified as hypermethylated and / or hypomethylated among these cfDNA fragments, the trained classifier detects that the number of sequencing reads exceeds a threshold.

32. The method according to claim 30, wherein: The length of these bait oligonucleotides is at least 45 nucleotides.

33. The method according to claim 30, wherein: The trained classifier distinguishes subjects with cancer from subjects without cancer with a defined specificity.

34. The method according to claim 33, wherein: The classifier is a hybrid model classifier.

35. The method according to claim 33, wherein: The defined specificity is 0.900 or higher.

36. The method according to claim 35, characterized in that: The application of the trained classifier further includes a sensitivity of 30% or higher.

37. The method according to claim 30, characterized in that: These bait oligonucleotides hybridize in common with target sequences of at least 10 genes from Table 1.

38. The method according to claim 30, wherein: At least one of these bait oligonucleotides hybridizes with a target sequence of a gene selected from: TWIST1, EOMES, HOXA9, POU4F2, and ZNF154.

39. The method according to claim 30, wherein: (a) These cancer cells are bladder cancer cells, and these bait oligonucleotides hybridize in common with target sequences of genes from Table 2 or Table 3; (b) These cancer cells are prostate cancer cells, and these bait oligonucleotides hybridize in common with target sequences of genes from Table 4; or (c) These cancer cells are kidney cancer cells, and these bait oligonucleotides hybridize in common with target sequences of genes from Table 5.

40. The method according to claim 30, characterized in that: These converted cfDNA molecules contain cfDNA treated with a cytidine deaminase or bisulfite.

41. The method according to claim 30, wherein: Each bait oligonucleotide is conjugated to a solid surface or a non-nucleotide affinity moiety.

42. The method according to claim 30, wherein: Differential methylation includes hypermethylation in a cancer sample relative to a non-cancer sample.

43. The method according to claim 30, characterized in that: Each target genomic region includes at least five methylation sites.

44. The method according to any one of claims 30-43, characterized in that: The method further includes obtaining these transformed cfDNA fragments or their amplification products, and wherein the obtaining further includes: (i) treating a urine sample to inhibit cell lysis; (ii) separating cfDNA fragments in the treated urine sample from cells in the treated urine sample, thereby producing a purified urine sample containing these cfDNA fragments; (iii) concentrating these cfDNA fragments in the purified urine sample by passing at least a portion of the purified urine sample through a filter, wherein the concentration produces a filtrate and a retained urine sample, and wherein the retained urine sample contains an increased concentration of cfDNA fragments; and (iv) isolating cfDNA fragments from the retained urine sample.

45. The method according to claim 44, characterized in that: The method further includes (v) amplifying one or more of the isolated cfDNA fragments.

46. The method according to claim 44, wherein: Treating the urine sample to inhibit cell lysis includes contacting the urine sample with one or more preservative reagents.

47. The method according to claim 44, characterized in that: The treatment includes treating with a nuclease inhibitor, a formaldehyde quencher, or both.

48. The method according to claim 44, characterized in that: The treatment includes contacting the urine sample with a composition comprising: (a) imidazolidinyl urea, EDTA, glycine, or a combination thereof; or (b) sodium azide, EDTA, or a combination thereof.

49. The method according to claim 44, wherein: The separating includes centrifuging to pellet cells in the treated urine sample.

50. The method according to claim 44, characterized in that: The filter is substantially impermeable to the passage of free nucleic acids and substantially permeable to salts in the purified urine sample.

51. The method according to claim 44, characterized in that: The filter has a nominal molecular weight cut-off of 10 kD, 5 kD, 3 kD, or lower.

52. The method according to claim 44, characterized in that: The retained urine sample has a concentration that is at least 2-fold, at least 5-fold, at least 10-fold, or at least 15-fold increased compared to the purified urine sample.

53. The method according to claim 44, characterized in that: The retained urine sample has a volume that is at least 50%, at least 75%, or at least 90% lower than the volume of the treated urine sample.

54. The method according to claim 53, wherein: The volume of the treated urine sample is 5 mL, 10 mL, 15 mL, 20 mL, 30 mL, 40 mL, 50 mL, or more.

55. The method according to claim 44, wherein: The method further includes freezing the retained urine sample.

56. The method according to claim 44, characterized in that: The treatment is completed within 120, 60, or 30 minutes after collection of the urine sample; and optionally wherein the separating and the concentrating are completed within 7 days after collection.

57. The method according to claim 30, characterized in that: The method further includes diagnosing cancer in the subject.

58. The method according to claim 57, characterized in that: The cancer is bladder cancer, prostate cancer, or kidney cancer.

59. The method according to claim 57, wherein: The method further includes treating the cancer in the subject.

60. The method according to claim 59, characterized in that: The treatment includes surgical resection, radiotherapy, chemotherapy, and / or immunotherapy.

61. A method for treating cancer in a subject, characterized in that: The method includes selecting a subject having cancer or an increased risk of developing cancer and administering a treatment to the subject, wherein: (a) the selecting includes identifying the subject as the source of a urine cell-free DNA (cfDNA) sample that contains one or more differentially methylated target genomic regions above a threshold level for the presence of cancer; (b) The one or more target genomic regions comprise one or more target sequences of one or more genes selected from Table 1; (c) The length of each target sequence is at least 25 nucleotides; (d) The cancer is bladder cancer, prostate cancer or kidney cancer; and (e) The treatment comprises surgical resection, radiotherapy, chemotherapy, immunotherapy or any combination thereof.

62. The method according to claim 61, wherein: The threshold level of the presence of the cancer is the level of a reference sample from a subject having the cancer.

63. The method according to claim 61, wherein: The threshold level of the presence of the cancer is determined by a classifier trained on sequencing reads of the converted DNA from a subject having the cancer.

64. The method according to claim 61, characterized in that: The one or more target genomic regions comprise target sequences from at least 10 genes in Table 1.

65. The method according to claim 61, wherein: The one or more target genomic regions comprise target sequences of genes selected from: TWIST1, EOMES, HOXA9, POU4F2 and ZNF154.

66. The method according to claim 61, wherein: (a) The cancer is bladder cancer, and the one or more target genomic regions comprise one or more target sequences of one or more genes selected from Table 2 or Table 3; (b) The cancer is prostate cancer, and the one or more target genomic regions comprise one or more target sequences of one or more genes selected from Table 4; or (c) The cancer is kidney cancer, and the one or more target genomic regions comprise one or more target sequences of one or more genes selected from Table 5.

67. The method according to any one of claims 61-66, characterized in that: Each target genomic region comprises at least five methylation sites.

68. A composition comprising a plurality of different bait oligonucleotides, wherein: (a) The bait oligonucleotides hybridize to converted DNA molecules derived from one or more target genomic regions; (b) The one or more target genomic regions comprise one or more target sequences of one or more genes selected from Table 1; (c) The one or more target genomic regions are differentially methylated in the cancer; and (d) Each bait oligonucleotide comprises a sequence of at least 25 nucleotides in length that hybridizes to one of the target sequences.

69. The composition according to claim 68, characterized in that: The one or more target genomic regions comprise target sequences from at least 10 genes in Table 1.

70. The composition according to claim 68, wherein: The one or more target genomic regions comprise target sequences of genes selected from: TWIST1, EOMES, HOXA9, POU4F2 and ZNF154.

71. The composition according to claim 68, characterized in that: The one or more target genomic regions comprise target sequences from: (a) genes in Table 2 or Table 3; (b) genes in Table 4; or (c) genes in Table 5.

72. The composition according to claim 68, characterized in that: Each target genomic region comprises at least five methylation sites.

73. The composition according to any one of claims 68 - 72, characterized in that: The differential methylation comprises at least 80% of the CpG sites in the target genomic region being methylated or unmethylated.

Citation Information

Patent Citations

  • Anomalous fragment detection and classification

    US12027237B2

  • Selective oxidation of 5-methylcytosine by TET-family proteins

    US20110236894A1

  • Composition and Methods Related to Modification of 5-Hydroxymethylcytosine (5-hmC)

    US20110301045A1

  • Stabilization of nucleic acids in urine

    US20160257995A1

  • Methylation haplotyping for non-invasive diagnosis (MONOD)

    US20160340740A1