Systems and methods for multi-analyte detection of cancer

By analyzing urine-derived cfDNA through NGS for genetic biomarkers, the method enhances cancer detection sensitivity and specificity, offering a non-invasive alternative to tissue biopsies and improving treatment strategies.

CN120322567APending Publication Date: 2025-07-15HUIDU MEDICAL CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380083753.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-02-13
Filing Date
2023-10-04
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently detect cancers, especially urogenital cancers such as bladder cancer, with tissue biopsy at risk of pain and complications and cannot be monitored frequently.

Method used

By extracting cell-free DNA (cfDNA) from urine and using next-generation sequencing technology (NGS) for assays, genetic changes such as single nucleotide variation, copy number variation and DNA rearrangement are identified, combined with bioinformatics analysis, early detection and monitoring of cancer can be achieved.

Benefits of technology

It provides a non-invasive, cost-effective method that improves the sensitivity and accuracy of cancer detection, can frequently monitor cancer progress and guide personalized treatments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120322567A_ABST
    Figure CN120322567A_ABST
Patent Text Reader

Abstract

Provided herein are methods and systems for detecting cancer. The method may include using a nucleic acid from a urine sample. The method may include determining nucleic acids in urine to detect a panel of biomarkers from a sample. The method may include treating a set of biomarkers to determine the presence of cancer or a cancer parameter. The processing may be performed by an algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference

[0002] This application claims the benefit of U.S. Application No. 63 / 413,433, filed on October 5, 2022, U.S. Application No. 63 / 445,201, filed on February 13, 2023, U.S. Application No. 63 / 445,145, filed on February 13, 2023, and U.S. Application No. 63 / 445,150, filed on February 13, 2023, each of which is incorporated herein by reference in its entirety. BACKGROUND OF THE INVENTION

[0003] Cancer is a leading cause of death worldwide. Detection of cancer in an individual can be crucial for providing treatment and improving patient outcomes. Cancer can be caused by genetic aberrations, which can lead to unregulated cell growth. Detection of genetic aberrations can be important for cancer detection. Sequencing nucleic acids in a sample from a patient can be used to detect genetic aberrations. SUMMARY OF THE INVENTION

[0004] Systems and methods for detecting the presence or absence of cancer in a subject are provided herein. The systems and methods provided herein include assaying polynucleotides to identify biomarkers of cancer in the subject. Detecting a particular biomarker for one type of cancer or a given cancer can allow for effective treatment to be provided to an individual and can result in improved outcomes. For multiple types of cancer, specific biomarkers indicative of a particular cancer type (or subtype) can be used to identify the prognosis of an individual having that cancer. To provide accurate detection and prognosis of cancer, multiple analytes can be examined. By analyzing an increasing number of analytes (and biomarker sets derived from the analytes), detection of cancer (or cancer parameters) can be improved and effective treatment can be recommended, and can also allow for a more accurate prognosis.

[0005] In one aspect, the present disclosure provides a method for detecting the presence or absence of cancer in a subject, comprising: (a) assaying cell-free deoxyribonucleic acid (cfDNA) molecules from a biological sample obtained or derived from the subject, wherein the biological sample comprises a urine sample; (b) detecting a set of biomarkers from the cfDNA molecules, wherein the set of biomarkers comprises differentially expressed markers or variants; (c) computationally processing the set of biomarkers to detect the presence or absence of the cancer in the subject. The method of claim 1, wherein the biological sample is obtained or derived from the subject using an ethylenediaminetetraacetic acid (EDTA) collection tube, a cell-free RNA collection tube, or a cell-free deoxyribonucleic acid (DNA) collection tube, other blood collection tubes, and CTC collection tubes.

[0006] In some embodiments, (a) comprises subjecting the biological sample to conditions sufficient to isolate, enrich, or extract the cfDNA molecules. In some embodiments, at least one of the cfDNA molecules is assayed using nucleic acid sequencing to generate nucleic acid sequencing reads. In some embodiments, the method further comprises filtering at least a subset of the nucleic acid sequencing reads based on a quality score. In some embodiments, the method further comprises error correction of the nucleic acid sequencing reads using a sample barcode or a molecular barcode attached to at least one of the cfDNA molecules. In some embodiments, the method further comprises performing at least one of single-stranded consensus sequence identification and double-stranded consensus sequence identification on the nucleic acid sequencing reads, thereby suppressing sequencing and PCR errors in the nucleic acid sequencing reads.

[0007] In some embodiments, the cfDNA molecules are assayed using DNA sequencing. In some embodiments, the DNA sequencing is selected from: next-generation sequencing, whole-genome sequencing, low-throughput sequencing, targeted sequencing, whole-exome sequencing, methylation-sensitive sequencing, bisulfite sequencing, and combinations thereof. In some embodiments, the DNA sequencing comprises targeted sequencing. In some embodiments, the nucleic acid sequencing comprises nucleic acid amplification. In some embodiments, the nucleic acid amplification comprises polymerase chain reaction (PCR) or isothermal amplification. In some embodiments, at least one of the cfDNA molecules is assayed by polymerase chain reaction (PCR), microarray, or isothermal amplification.

[0008] In some embodiments, the cancer is selected from: breast cancer, lung cancer, prostate cancer, colorectal cancer, melanoma, bladder cancer, non-Hodgkin lymphoma, kidney cancer, endometrial cancer, leukemia, pancreatic cancer, thyroid cancer, and liver cancer, and any combination thereof. In some embodiments, the cancer comprises bladder cancer. In some embodiments, the cancer comprises kidney cancer. In some embodiments, the bladder cancer comprises non-muscle invasive bladder cancer. In some embodiments, the subject is asymptomatic for the cancer.

[0009] In some embodiments, (b) comprises detecting the presence or absence of the cancer in the subject with at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99% accuracy. In some embodiments, (b) comprises detecting the presence or absence of the cancer in the subject with at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99% sensitivity. In some embodiments, (b) comprises detecting the presence or absence of the cancer in the subject with at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99% specificity. In some embodiments, (b) comprises detecting the presence or absence of the cancer in the subject with at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99% positive predictive value. In some embodiments, (b) comprises detecting the presence or absence of the cancer in the subject with at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99% negative predictive value.

[0010] In some embodiments, the biological sample is obtained or derived from the subject before the subject receives therapy for the cancer. In some embodiments, the biological sample is obtained or derived from the subject during therapy for the cancer. In some embodiments, the biological sample is obtained or derived from the subject after receiving therapy for the cancer. In some embodiments, the therapy is selected from: surgical resection, chemotherapy, radiotherapy, immunotherapy, cell therapy, adjuvant therapy, neoadjuvant therapy, androgen deprivation therapy, and combinations thereof. In some embodiments, the method further comprises identifying a clinical intervention for the subject based at least in part on the detected presence or absence of the cancer. In some embodiments, the clinical intervention is selected from a plurality of clinical interventions. In some embodiments, the clinical intervention is selected from: surgical resection, chemotherapy, radiotherapy, immunotherapy, adjuvant therapy, neoadjuvant therapy, androgen deprivation therapy, and combinations thereof. In some embodiments, the method further comprises administering the clinical intervention to the subject.

[0011] In some embodiments, the biomarker panel includes quantitative measurements of a cancer-related genomic locus panel. In some embodiments, the cancer-related genomic locus panel includes one or more members selected from the genes listed in Table 1. In some embodiments, the cancer-related genomic locus panel includes 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 members selected from the genes listed in Table 1. In some embodiments, the cancer-related genomic locus panel includes PTEN, TP53, or RB1. In some embodiments, the cancer-related genomic locus panel includes PTEN. In some embodiments, the cancer-related genomic locus panel includes FGFR3 or ERBB2. In some embodiments, the cancer-related genomic locus panel includes one or more members selected from the genes listed in Table 2. In some embodiments, the cancer-related genomic locus panel includes 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 members selected from the genes listed in Table 2.

[0012] In some embodiments, the method further includes using a probe configured to selectively enrich nucleic acid molecules corresponding to the genomic locus panel in the biological sample. In some embodiments, the probe includes nucleic acid primers. In some embodiments, the probe includes nucleic acid capture probes. In some embodiments, the probe has sequence complementarity with at least a portion of the nucleic acid sequence of the genomic locus panel. In some embodiments, the probe includes at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 different probes.

[0013] In some embodiments, the method further comprises determining the likelihood of the determination of the presence or absence of the cancer in the subject. In some embodiments, the method further comprises monitoring the presence or absence of the cancer in the subject, wherein the monitoring comprises assessing the presence or absence of the cancer in the subject at each of a plurality of time points. In some embodiments, a difference in the assessment of the presence or absence of the cancer in the subject between the plurality of time points indicates one or more clinical indicators selected from: (i) diagnosis of the cancer, (ii) prognosis of the cancer, and (iii) effectiveness or ineffectiveness of a treatment course for treating the cancer in the subject.

[0014] In some embodiments, the prognosis comprises progression-free survival (PFS) or overall survival (OS). In some embodiments, the biomarker set from the cfDNA molecules comprises tumor-related alterations selected from: copy number alterations (CNA), copy number losses (CNL), single nucleotide variants (SNV), insertions or deletions (indels), and rearrangements. In some embodiments, the biomarker set from the cfDNA molecules comprises copy number variations. In some embodiments, the biomarker set from the cfDNA molecules comprises copy number losses. In some embodiments, the biomarker set from the cfDNA molecules comprises single nucleotide variants.

[0015] In some embodiments, the method further comprises determining the mutant allele frequency of a set of somatic mutations in the biomarker set. In some embodiments, the method further comprises determining a blood copy number burden based on copy number alterations or copy number losses in the biomarker set. In some embodiments, the method further comprises determining a circulating tumor DNA (ctDNA) fraction of the cancer in the subject at least in part based on the set of mutant allele frequencies. In some embodiments, the method further comprises determining a tumor mutation burden (TMB) of the cancer in the subject at least in part based on the set of mutant allele frequencies. In some embodiments, the method further comprises determining a tumor mutation burden (TMB) of the cancer in the subject at least in part based on the set of mutant allele frequencies including microsatellites. In some embodiments, the method further comprises determining an aberration score of the cancer in the subject at least in part based on the set of mutant allele frequencies.

[0016] In one aspect, the present disclosure provides a method for detecting the presence or absence of cancer in an object, comprising: (a) providing cell-free deoxyribonucleic acid (cfDNA) molecules from a biological sample obtained or derived from the object, wherein the biological sample comprises a urine sample; (b) hybridizing the cfDNA molecules or derivatives thereof with a plurality of nucleic acid capture probes to produce a plurality of enriched cfDNA molecules; (c) sequencing the nucleic acids of the plurality of enriched cfDNA molecules to produce sequencing data; (d) computationally processing the sequencing data to detect a biomarker panel; and (e) detecting the presence or absence of the cancer in the object based at least on the presence of the biomarker panel.

[0017] In one aspect, the present disclosure provides a method for detecting the presence or absence of cancer in an object, comprising: (a) providing cell-free deoxyribonucleic acid (cfDNA) molecules from a biological sample obtained or derived from the object; (b) performing a first sequencing assay on the cfDNA or derivatives thereof to produce copy number data for at least one region of the object's genome, wherein the sequencing is performed at a depth of no more than 10x; (c) performing a second sequencing assay on the cfDNA molecules or derivatives thereof, wherein the second sequencing assay comprises a whole exome sequencing assay or a methylation-sensitive sequencing assay, to produce sequencing data; (d) computationally processing the copy number data and the sequencing data to detect a biomarker panel; and (e) detecting the presence or absence of the cancer in the object based at least on the presence of the biomarker panel.

[0018] In some embodiments, the biological sample is obtained or derived from the object using an ethylenediaminetetraacetic acid (EDTA) collection tube, a cell-free RNA collection tube, a cell-free deoxyribonucleic acid (DNA) collection tube, other blood collection tubes, and a CTC collection tube. In some embodiments, (a) comprises subjecting the biological sample to conditions sufficient to separate, enrich, or extract the cfDNA molecules.

[0019] In some embodiments, the biological sample comprises a urine, blood, or cerebrospinal sample. In some embodiments, the method further comprises filtering at least one subset of the nucleic acid sequencing reads based on a quality score. In some embodiments, the method further comprises error correction of the nucleic acid sequencing reads using a sample barcode or a molecular barcode attached to at least one of the cfDNA molecules. In some embodiments, the method further comprises performing at least one of single-stranded consensus sequence identification and double-stranded consensus sequence identification on the nucleic acid sequencing reads, thereby suppressing sequencing and PCR errors in the nucleic acid sequencing reads. In some embodiments, (b) or (c) comprises nucleic acid amplification. In some embodiments, the nucleic acid amplification comprises polymerase chain reaction (PCR) or isothermal amplification. In some embodiments, at least one of the cfDNA molecules is assayed by polymerase chain reaction (PCR), microarray, or isothermal amplification. In some embodiments, the cancer is selected from: breast cancer, lung cancer, prostate cancer, colorectal cancer, melanoma, bladder cancer, non-Hodgkin lymphoma, kidney cancer, endometrial cancer, leukemia, pancreatic cancer, thyroid cancer, and liver cancer, and any combination thereof. In some embodiments, the cancer comprises bladder cancer. In some embodiments, the bladder cancer comprises non-muscle invasive bladder cancer. In some embodiments, the subject is asymptomatic for the cancer.

[0020] In some embodiments, (d) comprises detecting the presence or absence of the cancer in the subject with at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98% or at least about 99% accuracy. In some embodiments, (d) comprises detecting the presence or absence of the cancer in the subject with at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98% or at least about 99% sensitivity. In some embodiments, (d) comprises detecting the presence or absence of the cancer in the subject with at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98% or at least about 99% specificity. In some embodiments, (d) comprises detecting the presence or absence of the cancer in the subject with at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98% or at least about 99% positive predictive value. In some embodiments, (d) comprises detecting the presence or absence of the cancer in the subject with at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98% or at least about 99% negative predictive value.

[0021] In some embodiments, the biological sample is obtained or derived from the subject before the subject receives therapy for the cancer. In some embodiments, the biological sample is obtained or derived from the subject during therapy for the cancer. In some embodiments, the biological sample is obtained or derived from the subject after receiving therapy for the cancer. In some embodiments, the treatment is selected from: surgical resection, chemotherapy, radiotherapy, immunotherapy, cell therapy, adjuvant therapy, neoadjuvant therapy, androgen deprivation therapy, and combinations thereof.

[0022] In some embodiments, the method further includes identifying a clinical intervention for the subject based at least in part on the detected presence or absence of the cancer. In some embodiments, the clinical intervention is selected from a plurality of clinical interventions. In some embodiments, the clinical intervention is selected from: surgical resection, chemotherapy, radiotherapy, immunotherapy, adjuvant therapy, neoadjuvant therapy, androgen deprivation therapy, and combinations thereof. In some embodiments, the method further includes administering the clinical intervention to the subject. In some embodiments, the biomarker panel includes quantitative measurements of a cancer-related genomic locus panel. In some embodiments, the cancer-related genomic locus panel includes one or more members selected from the genes listed in Table 1. In some embodiments, the cancer-related genomic locus panel includes 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 members selected from the genes listed in Table 1. In some embodiments, the cancer-related genomic locus panel includes PTEN, TP53, or RB1. In some embodiments, the cancer-related genomic locus panel includes PTEN. In some embodiments, the cancer-related genomic locus panel includes FGFR3 or ERBB2. In some embodiments, the cancer-related genomic locus panel includes one or more members selected from the genes listed in Table 2. In some embodiments, the cancer-related genomic locus panel includes 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 members selected from the genes listed in Table 2. In some embodiments, the cancer-related genomic locus panel includes one or more members selected from the genes listed in Table 3. In some embodiments, the cancer-related genomic locus panel includes 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 members selected from the genes listed in Table 3.In some embodiments, the method further comprises using probes configured to selectively enrich nucleic acid molecules corresponding to a genomic locus set in the biological sample. In some embodiments, the probes comprise nucleic acid primers. In some embodiments, the probes comprise nucleic acid capture probes. In some embodiments, the probes have sequence complementarity with at least a portion of the nucleic acid sequences of the genomic locus set. In some embodiments, the probes have sequence complementarity with at least a portion of the nucleic acid sequences of genes selected from the genes listed in Table 1, Table 2, or Table 3. In some embodiments, the probes comprise at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 different probes. In some embodiments, the method further comprises determining the likelihood of the determination of the presence or absence of the cancer in the subject.

[0023] In some embodiments, the method further comprises determining the likelihood of the determination of the presence or absence of the cancer in the subject. In some embodiments, the method further comprises monitoring the presence or absence of the cancer in the subject, wherein the monitoring comprises assessing the presence or absence of the cancer in the subject at each of a plurality of time points. In some embodiments, a difference in the assessment of the presence or absence of the cancer in the subject between the plurality of time points indicates one or more clinical indications selected from: (i) diagnosis of the cancer, (ii) prognosis of the cancer, and (iii) effectiveness or ineffectiveness of a treatment course for treating the cancer in the subject.

[0024] In some embodiments, the prognosis comprises progression-free survival (PFS) or overall survival (OS). In some embodiments, the biomarker set from the cfDNA molecules comprises tumor-related alterations selected from: copy number alterations (CNA), copy number losses (CNL), single nucleotide variants (SNV), insertions or deletions (indels), and rearrangements. In some embodiments, the biomarker set from the cfDNA molecules comprises copy number variations. In some embodiments, the biomarker set from the cfDNA molecules comprises copy number losses. In some embodiments, the biomarker set from the cfDNA molecules comprises single nucleotide variants.

[0025] In some embodiments, the method further comprises determining a mutant allele frequency of a somatic mutation set in the biomarker set. In some embodiments, the method further comprises determining a blood copy number burden based on a copy number alteration or a copy number loss of the biomarker set. In some embodiments, the method further comprises determining a circulating tumor DNA (ctDNA) fraction of the cancer of the subject at least in part based on the set of mutant allele frequencies. In some embodiments, the method further comprises determining a tumor mutation burden (TMB) of the cancer of the subject at least in part based on the set of mutant allele frequencies. In some embodiments, the method further comprises determining a tumor mutation burden (TMB) of the cancer of the subject at least in part based on the set of mutant allele frequencies including microsatellites. In some embodiments, the method further comprises determining an aberration score of the cancer of the subject at least in part based on the set of mutant allele frequencies.

[0026] Another aspect of the present disclosure provides a non - transitory computer - readable medium comprising machine - executable code that, when executed by one or more computer processors, implements any of the methods described above or elsewhere herein.

[0027] Another aspect of the present disclosure provides a system comprising one or more computer processors and a computer memory coupled thereto. The computer memory comprises machine - executable code that, when executed by one or more computer processors, implements any of the methods described above or elsewhere herein.

[0028] Those skilled in the art will readily appreciate other aspects and advantages of the present disclosure from the following detailed description, which shows and describes only illustrative embodiments of the present disclosure. As will be recognized, the present disclosure is capable of other and different embodiments, and several details can be modified in various obvious aspects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature and not restrictive.

[0029] Incorporated by reference

[0030] All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference to the extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. If a publication, patent, or patent application incorporated by reference conflicts with the present disclosure contained in this specification, this specification is intended to supersede and / or take precedence over any such conflicting material. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The novel features of the present invention are set forth in the appended claims. The features and advantages of the present invention will be better understood by reference to the following detailed description of illustrative embodiments in which the principles of the invention are utilized, and to the accompanying drawings (also referred to herein as "figures"), wherein:

[0032] Figure 1 An exemplary workflow for identifying somatic mutations in cell-free DNA of urine is shown.

[0033] Figure 2 An exemplary workflow for detecting variants is shown.

[0034] Figure 3 A graph showing the sensitivity of the assay and the mutant allele frequency is shown.

[0035] Figure 4 The agreement between the expected MAF and the MAF detected by the PredicineCARE assay is shown.

[0036] Figure 5 The agreement between genomic alterations detected by the PredicineCARE assay in tissue and urine is shown.

[0037] Figure 6 A schematic diagram of an exemplary workflow is shown.

[0038] Figure 7 A schematic diagram of an exemplary workflow is shown.

[0039] Figure 8 A schematic diagram of an exemplary workflow is shown.

[0040] Figure 9 A schematic diagram of an exemplary workflow is shown.

[0041] Figure 10 A schematic diagram of an exemplary workflow of a bioinformatic pipeline is shown.

[0042] Figure 11 A schematic diagram of an exemplary study design is shown.

[0043] Figure 12 A heatmap showing matched urine NGS and FFPE tissue RT-PCR is shown.

[0044] Figures 13A-13B A scatter plot showing the variant allele frequency (VAF) between matched FFPE and urine variants is shown.

[0045] Figures 14A-14D A schematic diagram of an exemplary assay workflow is shown. Figure 14BShows an illustration of whole-genome CNV detection and CNB calculation. Figure 14C -D shows the LP-WGS CNV profile.

[0046] Figures 15A-15D Shows a chart related to the analytical evaluation of PredicineCNB on clinical plasma samples.

[0047] Figure 16A Shows a heatmap of the LP-WGS CNV profiles of 14 non-muscle-invasive bladder cancer patient samples.

[0048] Figure 16B Shows a heatmap of the LP-WGS CNV profiles of 33 muscle-invasive bladder cancer patient samples.

[0049] Figure 17A Shows a comparison of CNB scores of FFPE, plasma, and urine samples between non-muscle-invasive and muscle-invasive / non-organ-confined bladder cancer patients.

[0050] Figure 17B Shows the LP-WGS CNV profiles of two non-invasive bladder cancer patients before and after TURBT surgery.

[0051] Figure 17C Shows the LP-WGS gene copy numbers of key bladder cancer genes before and after TURBT.

[0052] Figure 18 Shows a schematic diagram of an example workflow.

[0053] Figure 19 Shows the PredicineEPIC library with 1 ng, 2.5 ng, and 10 ng inputs compared to the standard 50 ng DNA input for a standard whole-genome bisulfite sequencing library.

[0054] Figure 20 Shows a figure displaying Uniform Manifold Approximation and Projection (UMAP).

[0055] Figure 21 Shows an example heatmap of the significance scores (color scale) of aberrant methylation fragments.

[0056] Figure 22 Shows a chart related to the methylation aberration score.

[0057] Figure 23 Shows the mutation profiles of urine and tissue tumor DNA from MIBC.

[0058] Figure 24 A-24D shows a graph related to the consistency and number of mutations detected from tDNA and utDNA in the WES region and the ATLAS region.

[0059] Figure 25 A graph showing the tumor fractions in tissue and urine is presented.

[0060] Figures 26A-26C A schematic diagram of a study showing an exemplary workflow, sample collection, and assay timeline and operations is presented.

[0061] Figures 27A-27B Data from an exemplary PredicineWES+ assay is presented.

[0062] Figure 28 A series of bCNB scores for multiple patients during treatment is presented.

[0063] Figure 29 Copy number gains and losses determined by PredicineCNB are presented.

[0064] Figure 30 The mutational landscape determined by PredicineWES+ is presented.

[0065] Figure 31 A computer system programmed or otherwise configured to implement the methods provided herein is presented. Detailed Description

[0066] Although various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that these embodiments are provided by way of example only. Many variations, changes, and substitutions will occur to those skilled in the art without departing from the present invention. It should be understood that various alternatives to the embodiments of the present invention described herein may be employed.

[0067] Systems and methods for detecting the presence or absence of cancer in a subject are provided herein. The systems and methods provided herein include assaying polynucleotides to identify biomarkers of cancer in the subject. The biomarkers can be processed to identify the presence or absence of cancer. The methods described herein can process an analyte to determine the presence or absence of cancer. The analyte can include cfDNA or other analytes that can be provided by non-invasive methods. By analyzing analytes obtained by non-invasive methods, these methods can allow for improved or similar detection or prognosis determination compared to methods using tumor or tissue biopsies.

[0068] Human urine can contain fragmented DNA known as urine cell-free DNA (ucfDNA), which is derived from dead cells in the urogenital tract or circulating DNA filtered through the glomeruli. Due to the direct accessibility of the urinary tract, urine is a viable source for detecting cfDNA biomarkers, and ucfDNA can improve the diagnostic sensitivity of current liquid biopsies for urogenital cancers. Many gene variants can be identified in ucfDNA from cancer patients, especially from patients with bladder cancer. Using urine for liquid biopsy provides a completely non-invasive method for detecting genomic biomarkers to guide cancer treatment.

[0069] Over the past 10+ years, next-generation sequencing technologies have revolutionized cancer genome research. NGS technologies are commercially used to guide cancer patient treatment plans and have been FDA approved or cleared for processing DNA from patient tissue or blood samples. Performing next-generation sequencing (NGS) assays on cfDNA allows for the precise detection of genomic alterations, including single nucleotide variants (SNVs), insertions and deletions (indels), copy number variants (CNVs), and DNA rearrangements. Urine cfDNA is directly derived from dead cells shed in urine and, due to tumor heterogeneity, can be considered more representative of the tumor than a tissue biopsy, which only reflects mutations found in a specific region of the tumor. In addition, urine can contain fewer contaminating proteins than blood, and cfDNA levels in urine can be higher than in the bloodstream. Urine cfDNA can be sequenced to allow for the detection of cancer and cancer-related genetic alterations from a subject's urine. The methods and assays described in this disclosure represent applications of NGS technologies and provide a non-invasive, cost-effective, and potentially more sensitive sampling method for patients with cancer, such as urogenital cancers, including bladder cancer.

[0070] In addition to staging and grading a patient's tumor, tissue biopsies are generally considered the gold standard for guiding cancer patient treatment, including eligibility for clinical trials or FDA-approved treatment options, such as the QIAGEN Therascreen FGFR RGQ RT-PCR Kit. However, depending on the tumor location or patient condition, tumor biopsies can be painful, and patients can be at risk of complications, which can make medical treatment expensive. In some cases, a tissue biopsy may not be feasible. For the treatment of patients with bladder cancer, less invasive sampling methods remain an unmet clinical need. Urine cfDNA assays and liquid biopsy (e.g., from urine) options can help fill this gap. In addition, compared to tissue biopsies, patients have the opportunity to undergo multiple tests via urine liquid biopsy. Thus, urine liquid biopsy represents a non-invasive and cost-effective method for obtaining patient samples to determine molecular fitness for specific treatment strategies.

[0071] Urine liquid biopsy samples can also improve patient care. In cases where tissue biopsy is not available for NGS testing in bladder cancer patients, urine liquid biopsy can be used to monitor bladder cancer patients. This assay can identify the FGFR molecular fitness of specific therapeutic products. In addition, this assay can also detect other gene mutations in urine samples from bladder cancer patients, including but not limited to alterations in CDKN2A, HRAS / KRAS, KDM6A, PIK3CA, TERT, TP53, and TSC1, and if these alterations are identified, they can help guide patient treatment.

[0072] Bladder cancer is the tenth most common malignancy globally, with an estimated 550,000 new cases and 200,000 deaths reported in 2018. The majority of bladder cancer cases are non-muscle-invasive bladder cancer (NMIBC), which requires frequent monitoring, local resection (transurethral resection of bladder tumor [TURBT]), and an intensified regimen of intravesical therapy to reduce the risk of recurrence and progressive disease. Despite these efforts, within 5 years, 45% to 60% of patients will develop recurrent disease, and nearly 20% of high-risk disease patients will progress to muscle-invasive tumors, requiring radical cystectomy (RC). The natural course of high-risk NMIBC is unpredictable; the recurrence rate varies from 15% to 78%, and the rate of progression to muscle invasion and metastasis varies from <1% to 45%. Long-term outcomes indicate that approximately 20% to 25% of high-risk NMIBC patients will ultimately die of bladder cancer.

[0073] Analytes available for tumor diagnosis in urine liquid biopsy include cfDNA, non-coding RNA, shed tumor cells, and proteins. During tumor destruction therapy or during apoptosis and necrosis, both healthy and diseased cells can release cfDNA fragments that are typically 100 - 200 base pairs in length. In non-diseased patients, phagocytes engulf cell debris and necrotic cells, so cfDNA levels are very low. In the case of diseased patients, phagocytosis is impaired, DNA digestion is minimal, and the DNA fragments have random sizes that can exceed 10,000 base pairs. Therefore, cfDNA levels are usually elevated in diseased patients.

[0074] A prospective study was conducted that included blood, urine, and tumor tissue samples from 16 patients with bladder cancer presenting with hematuria to identify a panel of DNA markers using NGS assays. The study showed that there was a higher concordance of gene mutations in urine supernatant and pellet compared to plasma with cancer tissue. A 48-gene panel was used, and the analysis showed that two gene combinations for genetic diagnostic modeling of DNA in bladder cancer patients were: TERT, FGFR3, TP53, PIK3CA, and KRAS for urine supernatant, and TERT, FGFR3, TP53, HRAS, PIK3CA, KRAS, and ERBB2. The accuracy of the five-gene panel and seven-gene panel produced AUCs of 0.94 (95% CI 0.91 - 0.97) and 0.91 (95% CI 0.86 - 0.96), respectively. The study concluded that urine cfDNA has great potential in identifying bladder cancer in patients with hematuria.

[0075] Fibroblast growth factor receptors FGFR1 to FGFR4 are tyrosine kinases that are present in various types of endothelial and tumor cells and have been shown to play important roles in tumor cell growth, survival, and migration, as well as in maintaining tumor angiogenesis (Turner 2010). FGFR-related tumorigenic mechanisms are diverse and include gene amplification, mutation, and fusion. FGFR3 genetic alterations were found in 60% to 70% of early NMIBCs. The frequency of FGFR3 alterations observed in bladder cancer varies with tumor stage and grade. For example, in a prospective cohort of 772 NMIBC patients, TaG1 (158 / 257) and TaG2 (139 / 239) tumors showed similar mutation frequencies of 61.5% and 58.1%, respectively. FGFR3 mutation is an independent predictor of recurrence in patients with low-grade Ta tumors. TaG3 (30 / 88; 34.1%), T1G2 (7 / 26; 26.9%), and T1G3 tumors (20 / 119; 17%) had lower mutation frequencies.

[0076] A recent study of a pooled dataset with matched clinical and genomic data from 263 pT1 patients showed that FGFR alterations were frequent in high-risk patients (39% mutations, 6% fusions, not mutually exclusive). Additionally, different from previous reports, the study showed that the prognosis of patients with FGFR alterations was not different from that of patients without FGFR alterations (Breyer 2020). The incidence of FGFR alterations in MIBC seems to be slightly lower than in NMIBC, with FGFR3 mutation rates of approximately 15% reported in multiple studies.

[0077] Urine cfDNA (ucfDNA) can be extracted from the urine of an object, and various reactions can be performed on the ucfDNA to allow ucfDNA sequencing. Library construction of ucfDNA can include amplification, ligation of adapters or other sequences, and / or barcoding to generate a sequencing library. In addition, capture probes or amplification primers can be used to enrich ucfDNA or the ucfDNA library to enrich specific target sequences from cfDNA. Then, a sequencing reaction can be performed on the library to generate sequencing data.

[0078] Figure 1 A general workflow for performing an exemplary assay using urine is shown. A urine sample can be isolated and collected from an individual. After collection, extraction of urine cfDNA can be performed. Then, library construction of the extracted urine cfDNA can be carried out, followed by enrichment of specific targets. Once enrichment has been performed, the cfDNA can be sequenced, and then the sequencing data can be processed.

[0079] Genetic alterations, such as single nucleotide variants (SNVs), copy number gains and losses, DNA rearrangements, and copy number variations (CNVs), can be identified by performing bioinformatics analysis on the sequencing data. Generally, a bioinformatics pipeline can utilize the raw sequencing data (e.g., BCL files) and output mutational calls. The pipeline can perform various tasks to analyze the sequencing data, such as adapter trimming, barcode checking, or error correction. Aligning tools (e.g., the BWA aligning tool) can be used to align the cleaned paired files (e.g., FASTQ files) to the human reference genome. Then, consensus sequences can be derived by merging paired-end reads that originate from the same molecule as the single-stranded fragment. Single-stranded fragments from the same double-stranded DNA molecule can be further merged into double strands. These processes can allow for the correction of sequencing and PCR errors. Figure 2 An exemplary bioinformatics analysis workflow is shown.

[0080] Figure 6 – Figure 9 The workflow of an exemplary urine cfDNA assay (e.g., PredicineCARE) is shown. The blue ovals provide connectivity between the individual figures. Figure 10 The workflow of a bioinformatics pipeline (e.g., DeepSEA) is shown. As Figure 6 shown, a sample can be provided, and then cfDNA can be extracted and isolated from the sample. Then, the cfDNA can be verified and quality controlled. A sequencing library can be generated from the cfDNA by end repair and A-tail adapter ligation. Then, the library can be amplified and quantified.

[0081] As Figure 7As shown, once the sample library is generated, the library can be hybridized with capture probes in solution. The capture probes can be biotinylated. The probes can then be incubated with magnetic beads (e.g., streptavidin magnetic beads) and then washed to remove any contaminants. Then, the captured DNA can be eluted from the beads and can be further amplified, normalized, and / or pooled and a capture library can be formed.

[0082] As Figure 8 shown, the capture library can be sequenced using a next-generation sequencer (e.g., Illumina NovaSeq) and raw sequencing data can be generated.

[0083] Figure 9 A schematic diagram of the raw sequencing process is shown. The data is input into a bioinformatics pipeline (e.g., DeepSea pipeline) for variant calls. Once variant calls are made and the quality is checked, the variants can be classified and analyzed to determine the presence of specific variants and information about the presence of cancer and specific mutations associated with cancer can be provided. Then, this can generate a report and the report can then be sent to a doctor.

[0084] Figure 10 A schematic diagram of a variant calling pipeline (e.g., DeepSea pipeline) is shown. The pipeline can demultiplex and extract unique molecular identifiers (UMIs) to generate FASTq files with UMIs. These files can then be aligned and then consensus sequences can be generated. Various errors can be corrected at least in part based on the analysis of the UMIs or consensus sequences. The sequencing with these errors corrected can then be input into various variant callers to generate variant calling results. These files can then be annotated and filtered to generate variant calling results.

[0085] An object can be suspected of having cancer. The cancer can be specific or originate from an organ or other area of the object. For example, the cancer can be breast cancer, lung cancer, prostate cancer, colorectal cancer, melanoma, bladder cancer, non-Hodgkin lymphoma, kidney cancer, endometrial cancer, leukemia, pancreatic cancer, thyroid cancer, and liver cancer, and any combination thereof. The cancer can be hormone-sensitive prostate cancer (HSPC), castration-resistant prostate cancer (CRPC), metastatic prostate cancer, and combinations thereof. The cancer can include biomarkers specific to a particular cancer. The specific biomarkers can indicate the presence of a particular cancer. For example, the biomarker can indicate the presence of castration-resistant prostate cancer. Identifying the presence of the cancer type can allow for the determination of treatment options or recommendations.

[0086] In some cases, an object may not have cancer symptoms. For example, cancer may not exhibit any symptoms, and the object may be unaware of the presence of cancer. The methods described herein may allow for the identification of cancer at an earlier stage than other methods. Identifying the presence of cancer at an earlier stage may allow for the determination of treatment options or recommendations at an earlier stage and may allow the object to have an improved prognosis.

[0087] A biological sample may include nucleic acids. The biological sample may be a cell-free deoxyribonucleic acid (cfDNA) sample or a cell-free ribonucleic acid (cfRNA) sample. The biological sample may include genomic DNA or germline DNA (gDNA). The nucleic acid may be DNA (e.g., double-stranded DNA, single-stranded DNA, single-stranded DNA hairpin, cDNA, genomic DNA, germline DNA, circulating tumor DNA (ctDNA), cell-free DNA (cfDNA)), RNA (e.g., cfRNA, mRNA, cRNA, miRNA, siRNA, miRNA, snoRNA, piRNA, tiRNA, snRNA) or a DNA / RNA hybrid. The biological sample may be derived from or contain a biological fluid. For example, the biological sample may be a plasma sample, a serum sample, a buffy coat sample, a peripheral blood mononuclear cell (PBMC) sample, a red blood cell sample, a urine sample, a saliva sample or other body fluid sample. The biological sample may include or be a pleural fluid sample, a peritoneal fluid sample, an amniotic fluid sample, a cerebrospinal fluid sample, a lymph fluid sample, a sweat sample, a tear sample, a semen sample or any combination of biological fluids.

[0088] A biological sample may be collected, obtained or derived from an object using a collection tube. The collection tube may be an ethylenediaminetetraacetic acid (EDTA) collection tube, a cell-free RNA collection tube, a cell-free deoxyribonucleic acid (DNA) collection tube, and a CTC collection tube, or other blood collection tubes. The collection tube may include additional reagents for stabilizing nucleic acid molecules or blood cells. The collection tube may stabilize the nucleic acid or blood cells to minimize degradation of the biological sample prior to assay. The additional reagents may include buffer salts or chelating agents.

[0089] Biological samples can be obtained or derived from a subject at different times. Biological samples can be obtained or derived from a subject before the subject receives therapy for cancer. Biological samples can be obtained or derived from a subject during the subject's receipt of therapy for cancer. Biological samples can be obtained or derived from a subject after the subject has received therapy for cancer. Biological samples can be collected at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 time points. The time points can occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 hours or longer. The time points can occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 days or longer. The time points can occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 weeks or longer. The time points can occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 months or longer. The time points can occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 years or longer.

[0090] In various aspects described herein, a clinical intervention or therapy can be identified based at least in part on the identification of the presence of cancer or the presence of cancer parameters. The clinical intervention can be a variety of clinical interventions. The clinical intervention can be selected from a variety of clinical interventions. The clinical intervention can be surgical resection, chemotherapy, radiotherapy, immunotherapy, adjuvant therapy, neoadjuvant therapy, androgen deprivation therapy, or a combination thereof. In some cases, a clinical intervention can be administered to a subject. After administering the clinical intervention, a sample can be obtained or derived from the subject to monitor the cancer or cancer parameters. Thus, the methods and systems herein can be implemented iteratively so that cancer monitoring can be implemented. Additionally, by implementing the method or system iteratively, the therapy or clinical intervention can be updated based on the results of the method. Monitoring of cancer can include an assessment and an assessment difference from a previously generated assessment. A difference in the cancer assessment of a subject among multiple time points (or samples) can indicate one or more clinical indications, such as the diagnosis of cancer, the prognosis of cancer, or the effectiveness or ineffectiveness of a treatment course for treating the subject's cancer. The prognosis can include progression-free survival (PFS), overall survival (OS), or other metrics related to cancer severity or viability.

[0091] The biological sample can be subjected to additional reactions or conditions prior to the assay. For example, the biological sample can be subjected to conditions sufficient to isolate, enrich, or extract nucleic acids (such as cfDNA molecules).

[0092] The methods disclosed herein may include performing one or more enrichment reactions on one or more nucleic acid molecules in a sample. The enrichment reactions may include contacting the sample with one or more beads or bead sets. The enrichment reactions may include one or more hybridization reactions. For example, the enrichment reactions may include contacting the sample with one or more capture probes or bait molecules that hybridize to nucleic acid molecules of a biological sample. The enrichment reactions may include differential amplification of a set of nucleic acid molecules. The enrichment reactions may enrich multiple genetic loci or sequences corresponding to genetic loci. The enrichment reactions may include using primers or probes that may be complementary to the sequence (or upstream or downstream sequences) of the sequence to be enriched. For example, the capture probes may include sequences complementary to a set of genomic loci and allow enrichment of the genomic loci. The enrichment reactions may include multiple probes or primers. For example, the capture probes may include sequences complementary to the selected genes in Table 1, Table 2, or Table 3. The multiple probes may include 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 205, 210, 215, 220, 225, 230, 235, 240, 245, 250, 255, 260, 265, 270, 275, 280, 285, 290, 295, 300, 305, 310, 315, 320, 325, 330, 335, 340, 345, 350, 355, 360, 365, 370, 375, 380, 385, 390, 395, 400, 405, 410, 415, 420, 425, 430, 435, 440, 445, 450, 455, 460, 465, 470, 475, 480, 485, 490, 495, 500, 505, 510, 515, 520, 525, 530, 535, 540, 545, 550, 555, 560, 565, 570, 575, 580, 585, 590, 595 or 600 different probes.

[0093] The methods disclosed herein can include performing one or more separation or purification reactions on one or more nucleic acid molecules in a sample. The separation or purification reaction can include contacting the sample with one or more beads or bead sets. The separation or purification reaction can include one or more hybridization reactions, enrichment reactions, amplification reactions, sequencing reactions, or combinations thereof. The separation or purification reaction can include using one or more separators. One or more separators can include a magnetic separator. The separation or purification reaction can include separating nucleic acid molecules bound to beads from bead-free nucleic acid molecules. The separation or purification reaction can include separating nucleic acid molecules hybridized to capture probes from capture-probe-free nucleic acid molecules. The separation reaction can include removing or separating a set of nucleic acid molecules from another set of nucleic acids.

[0094] The methods disclosed herein can include performing an extraction reaction on one or more nucleic acids in a biological sample. The extraction reaction can lyse cells or disrupt the interaction between nucleic acids and cells, thereby separating, purifying, enriching nucleic acids, or subjecting them to other reactions.

[0095] The methods disclosed herein can include amplification or extension reactions. The amplification reaction can include polymerase chain reaction. The amplification reaction can include PCR-based amplification, non-PCR-based amplification, or combinations thereof. One or more PCR-based amplifications can include PCR, qPCR, nested PCR, linear amplification, or combinations thereof. One or more non-PCR-based amplifications can include multiple displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, rolling circle amplification, loop-to-loop amplification, or combinations thereof. The amplification reaction can include isothermal amplification.

[0096] The methods disclosed herein may include a barcoding reaction. The barcoding reaction may include adding a barcode or tag to a nucleic acid. The barcode may be a molecular barcode or a sample barcode. For example, the barcoded nucleic acid may include a barcode sequence, which may be a degenerate n-mer. The sequence may be randomly generated or generated to synthesize a specific barcode sequence. The barcoded nucleic acid may be added to a sample to label nucleic acid molecules in the sample. The barcode may be sample-specific. For example, multiple barcoded nucleic acids may be added to a sample with the same barcode sequence. After barcoding the nucleic acid, nucleic acids from the same sample may have the same barcode sequence and may allow the nucleic acids to be identified as belonging to a particular or given sample. Molecular barcodes may also be used such that each molecule (or molecules) of the same volume has a different molecular barcode. The barcode may undergo amplification such that all amplicons derived from the molecule have the same barcode. In this way, molecules derived from the same molecule can be identified. The sequence reads may be processed based on the barcode sequence. For example, the processing may reduce errors or allow tracing of the molecules. The barcode sequence may be appended or otherwise added or incorporated into the sequence through various reactions such as amplification, extension, or ligation reactions, and may be performed enzymatically using a nucleic acid polymerase or ligase. The ligation may be an overhang or blunt-end ligation, and the barcode may include complementarity to the nucleic acid to be barcoded. The complementarity may be a sequence derived from a sample of an object or may be a constant sequence generated by reacting nucleic acids in the sample.

[0097] In some cases, a biological sample may include multiple components. For example, the biological sample may be a whole blood sample. The biological sample may undergo a reaction to separate or fractionate the biological sample. For example, the whole blood sample may be fractionated and cell-free nucleic acids may be obtained. The whole blood sample may be fractionated using centrifugation to separate the blood cells from the plasma (which may contain cell-free nucleic acids). The sample may undergo multiple rounds of separation or fractionation.

[0098] In various aspects described in the present disclosure, nucleic acids can undergo a sequencing reaction. The sequencing reaction can be used for DNA, RNA, or other nucleic acid molecules. Examples of sequencing reactions that can be used include capillary sequencing, next-generation sequencing, Sanger sequencing, sequencing by synthesis, single-molecule nanopore sequencing, sequencing by ligation, hybridization sequencing, nanopore current-limited sequencing, or combinations thereof. Sequencing by synthesis can include reversible terminator sequencing, progressive single-molecule sequencing, sequential nucleotide flow sequencing, or combinations thereof. Sequential nucleotide flow sequencing can include pyrosequencing, pH-mediated sequencing, semiconductor sequencing, or combinations thereof. The sequencing reaction can include whole-genome sequencing, whole-exome sequencing, low-throughput whole-genome sequencing, targeted sequencing, methylation-sensitive sequencing, enzymatic methylation sequencing, bisulfite methylation sequencing. The sequencing reaction can be transcriptome sequencing, mRNA-seq, totalRNA-seq, smallRNA-seq, exosome sequencing, or combinations thereof. Combinations of sequencing reactions can be used in the methods described elsewhere herein. For example, a sample can undergo whole-genome sequencing and whole-transcriptome sequencing. Since a sample can include multiple types of nucleic acids (e.g., RNA and DNA), sequencing reactions specific for DNA or RNA can be used to obtain sequence reads related to the nucleic acid type.

[0099] Sequencing reactions can be performed at different sequencing depths. The sequencing depth of the sequencing reaction can be selected or adjusted. The sequencing reaction can include sequencing a region at a depth of at least 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 11x, 12x, 13x, 14x, 15x, 16x, 17x, 18x, 19x, 20x, 25x, 30x, 35x, 40x, 45x, 50x, 60x, 70x, 80x, 90x, 100x, 200x, 300x, 400x, 500x, 600x, 700x, 800x, 900x, 1000x, 2000x, 3000x, 4000x, 5000x, 6000x, 7000x, 8000x, 9000x, 10,000x, 20,000x, 30,000x, 40,000x, 50,000x, 60,000x, 70,000x, 80,000x, 90,000, 100,000x or higher. The sequencing reaction can include sequencing a region at a depth of no more than 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 11x, 12x, 13x, 14x, 15x, 16x, 17x, 18x, 19x, 20x, 25x, 30x, 35x, 40x, 45x, 50x, 60x, 70x, 80x, 90x, 100x, 200x, 300x, 400x, 500x, 600x, 700x, 800x, 900x, 1000x, 2000x, 3000x, 4000x, 5000x, 6000x, 7000x, 8000x, 9000x, 10,000x, 20,000x, 30,000x, 40,000x, 50,000x, 60,000x, 70,000x, 80,000x, 90,000, 100,000x or lower.

[0100] In various embodiments, low-throughput whole-genome sequencing is used to sequence nucleic acids. Low-throughput whole-genome sequencing can be performed at an average sequencing depth of at least 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x or higher. Low-throughput whole-genome sequencing can be performed at an average sequencing depth of no more than 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x or lower. Low-throughput whole-genome sequencing can be performed at an average depth between 1x and 2x.

[0101] In various embodiments, a set of personalized or customized probes can be used for a sequencing reaction. A sequencing reaction using a set of personalized or customized probes can be a deep sequencing reaction or an ultra-deep sequencing reaction. For example, a sequencing reaction using a set of personalized or customized probes can be performed at a depth of 50x, 60x, 70x, 80x, 90x, 100x, 200x, 300x, 400x, 500x, 600x, 700x, 800x, 900x, 1000x, 2000x, 3000x, 4000x, 5000x, 6000x, 7000x, 8000x, 9000x, 10,000x, 20,000x, 30,000x, 40,000x, 50,000x, 60,000x, 70,000x, 80,000x, 90,000, 100,000x or higher.

[0102] In various embodiments, whole exome sequencing is used to sequence the nucleic acids of an object. Whole exome sequencing can be performed at a non-uniform depth. For example, certain regions of the exome can be enhanced or otherwise sequenced at a greater depth than other regions, or at a depth greater than the average depth of whole exome sequencing. By sequencing certain regions at a higher depth, genes or regions of greater interest can be analyzed with higher sensitivity, accuracy, and / or precision. Genes or regions associated with or related to cancer can be sequenced at a greater depth. For example, at least 100, 200, 300, 400, 500, 600, 700, 800, 900 or more genes can be sequenced at a depth higher than the rest of the exome (e.g., the average depth of whole exome sequencing).

[0103] Nucleic acid sequencing can generate sequencing read data. The sequencing reads can be processed to generate quality-improved data. Sequencing reads with quality scores can be generated. The quality scores can indicate the accuracy of the sequence reads or the level above the noise threshold for the identification of a given base or the signal. The quality scores can be used to filter the sequencing reads. For example, sequencing reads that do not meet a specific quality score threshold can be removed. The sequencing reads can be processed to generate a consensus sequence or consensus base identification. A given nucleic acid (or nucleic acid fragment) can be sequenced, and due to reactions before or during sequencing, errors can occur in the sequence. For example, amplification or PCR can introduce errors in the amplicon, such that the sequence is not identical to the parental sequence. Using sample barcodes or molecular barcodes, error correction can be performed. Error correction can include identifying sequence reads that cannot be corroborated with other sequences from the same sample or the same original parental molecule. The use of barcodes can allow the identification of the same parental or sample. In addition, the sequence reads can be processed by performing single-stranded consensus sequence identification or double-stranded consensus sequence identification to reduce or suppress errors.

[0104] The methods disclosed herein can include determining allele frequencies or other cancer - related metrics. The method can include the mutant allele frequencies of a somatic mutation set in a biomarker panel. The mutant allele frequencies can be used to determine the circulating tumor DNA (ctDNA) fraction of an object's cancer. The plasma tumor mutation burden (pTMB) of an object's cancer can be determined at least in part based on the set of mutant allele frequencies. Microsatellite instability detection can also be used to determine the presence or absence of cancer or a cancer metric. The methylation status can be determined using the methods described herein and can be used to identify the presence of cancer or a cancer parameter.

[0105] In various aspects, a biomarker panel is processed and data corresponding to the biomarker is generated. The biomarker panel can include quantitative measurements from a set of cancer - related genomic loci. The cancer - related genomic loci can correspond to a genome. The cancer - related genomic loci can include one or more genes selected from Table 1. In some cases, the set of cancer - related genomic loci includes 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 members selected from the genes listed in Table 1. The cancer - related genomic loci can include one or more genes selected from Table 2. In some cases, the set of cancer - related genomic loci includes 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 members selected from the genes listed in Table 2.

[0106] Table 1: List of Genes

[0107]

[0108] Table 2: List of Genes in the PredicineCARE Panel

[0109] List of Genes

[0110]

[0111] *The indicated genes are covered by the complete coding region

[0112] Cancer-related genomic loci may include one or more genes selected from Table 3. In some cases, the cancer-related genomic loci group includes 2 to 600 genes selected from Table 3. In some cases, the cancer-related genomic loci group includes 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 205, 210, 215, 220, 225, 230, 235, 240, 245, 250, 255, 260, 265, 270, 275, 280, 285, 290, 295, 300, 305, 310, 315, 320, 325, 330, 335, 340, 345, 350, 355, 360, 365, 370, 375, 380, 385, 390, 395, 400, 405, 410, 415, 420, 425, 430, 435, 440, 445, 450, 455, 460, 465, 470, 475, 480, 485, 490, 495, 500, 505, 510, 515, 520, 525, 530, 535, 540, 545, 550, 555, 560, 565, 570, 575, 580, 585, 590, 595 or 600 members of the genes listed in Table 3.

[0113] Table 3: List of Genes in PredicineATLAS

[0114]

[0115]

[0116] The biomarker group may correspond to genetic aberrations of genetic loci. The genetic aberrations can be tumor-related alterations. The genetic aberrations can include copy number alterations (CNA), copy number losses (CNL), single nucleotide variants (SNV), insertions or deletions (indels), and / or rearrangements. The biomarker group can be identified in a variety of nucleic acid types. For example, tumor-related alterations can be identified in cfDNA. Tumor-related alterations can include changes in allele expression or gene expression. The methods and systems disclosed herein can allow for gene expression profiling and identification of changes in gene expression levels.

[0117] In various aspects, the method can include identifying the presence of cancer or a cancer parameter. The method can include determining the probability or likelihood of the presence of cancer or a cancer parameter. For example, an output can be generated that indicates the probability that a subject has cancer, rather than a binary output indicating presence or absence. The probability can be determined based on algorithms described elsewhere herein. Similarly, the probability or likelihood of response to a particular treatment or the probability of recurrence can be output.

[0118] In various aspects, an algorithm is used to process a biomarker panel. The algorithm can be a trained algorithm. The trained algorithm can use the biomarker panel as input and generate an output regarding the presence or absence of cancer. The output can be specific to a cancer type or cancer subtype. For example, the output can indicate the presence of bladder cancer.

[0119] The trained algorithm can be trained on multiple samples. For example, at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000 or more independent training samples can be used to train the trained algorithm. No more than 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000 or fewer independent training samples can be used to train the trained algorithm. The training samples can be related to the presence or absence of cancer. The training samples can be related to cancer recurrence. The training samples can be related to cancer that is resistant to a particular drug or treatment. A single training sample can be positive for a particular cancer. A single training sample can be negative for a particular cancer. By using the training samples, the trained algorithm can detect cancer, determine the probability of cancer recurrence or relapse, or determine whether the cancer includes a biomarker group that is resistant to treatment. The training samples can be related to additional clinical health data of the subject. For example, the additional clinical health data can include the subject's gender, weight, height, or levels of metabolites or antibodies. The additional clinical health data can include indications of other diseases, disorders, or medical conditions.

[0120] The trained algorithm can be trained using multiple sets of training samples. These sets can include the training samples described elsewhere in this document. For example, a first set of independent training samples related to the presence of cancer and a second set of independent training samples related to the absence of cancer can be used for training. Similarly, the first set can be related to recurrence, and the second set of samples can be related to the absence of recurrence.

[0121] The trained algorithm can also process additional clinical health data of the subject. For example, the additional clinical health data can include the subject's gender, weight, height, or levels of metabolites or antibodies. The additional clinical health data can include indications of other diseases, disorders, or medical conditions that the subject may have. By using the additional clinical health data, in combination with biomarkers, the trained algorithm can output the presence or absence of cancer, the probability of recurrence, or drug treatment resistance, which can be different from the output of an algorithm that does not process additional clinical health.

[0122] The trained algorithm can be an unsupervised machine learning algorithm. For example, an unsupervised machine learning algorithm can utilize cluster analysis to identify attributes of interest. The trained algorithm can be a supervised machine learning algorithm. For example, the trained algorithm can be trained with training data to generate an expected or desired output. Supervised learning algorithms can include deep learning algorithms, support vector machines (SVMs), neural networks, or random forests. Through the machine learning algorithm, the trained algorithm can identify the relationship between biomarkers and a specific cancer prognosis or diagnosis. Without a trained algorithm, it may be difficult to identify the relationship of biomarkers to accurately identify the presence of cancer or other parameters related to cancer.

[0123] In various aspects, the systems and methods can include an accuracy, sensitivity, or specificity for detecting cancer or a cancer parameter. For example, a method or system can include detecting the presence or absence of cancer (or the presence of a cancer parameter such as recurrence, relapse, or drug resistance) in a subject with at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99% accuracy. A method or system can include detecting the presence or absence of cancer (or the presence of a cancer parameter such as recurrence, relapse, or drug resistance) in a subject with at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99% sensitivity. A method or system can include detecting the presence or absence of cancer (or the presence of a cancer parameter such as recurrence, relapse, or drug resistance) in a subject with at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99% specificity. A method or system can include detecting the presence or absence of cancer (or the presence of a cancer parameter such as recurrence, relapse, or drug resistance) in a subject with at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99% positive predictive value. A method or system can include detecting the presence or absence of cancer (or the presence of a cancer parameter such as recurrence, relapse, or drug resistance) in a subject with at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99% negative predictive value.

[0124] Computer control system

[0125] The present disclosure provides a computer system programmed to implement the methods of the present disclosure. Figure 31 Shown is a computer system 3101 programmed or otherwise configured to perform the analysis or operations of the method, such as determining the likelihood of cancer presence or running an algorithm based on an individual's biomarker panel. The computer system 3101 can regulate various aspects of the methods and systems of the present disclosure, such as, for example, executing an algorithm, inputting training data, analyzing a biomarker panel, or outputting results regarding the presence or absence of cancer to a user. The computer system 3101 can be the user's electronic device or a computer system placed remotely relative to the electronic device. The electronic device can be a mobile electronic device.

[0126] The computer system 3101 includes a central processing unit (CPU, also referred to herein as "processor" and "computer processor") 3105, which can be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system 3101 also includes a memory or storage location 3110 (such as random access memory, read-only memory, flash memory), an electronic storage unit 3115 (such as a hard disk), a communication interface 3120 for communicating with one or more other systems (such as a network adapter), and peripheral devices 3125, such as caches, other memories, data memories, and / or electronic display adapters. The memory 3110, storage unit 3115, interface 3120, and peripheral devices 3125 communicate with the CPU 3105 via a communication bus (solid lines), such as a motherboard. The storage unit 3115 can be a data storage unit (or data repository) for storing data. The computer system 3101 can be operably coupled to a computer network ("network") 3130 via the communication interface 3120. The network 3130 can be the Internet, an intranet, and / or an extranet, or an intranet and / or extranet that communicates with the Internet. In some cases, the network 3130 is a telecommunications and / or data network. The network 3130 can include one or more computer servers, which can implement distributed computing, such as cloud computing. In some cases, with the help of the computer system 3101, the network 3130 can implement a peer-to-peer network, which can enable devices coupled to the computer system 3101 to act as clients or servers.

[0127] The CPU 3105 can execute a series of machine-readable instructions, which can be embodied in a program or software. The instructions can be stored in a memory location such as the memory 3110. The instructions can be directed to the CPU 3105 and subsequently program or otherwise configure the CPU 3105 to implement the methods of the present disclosure. Examples of operations performed by the CPU 3105 can include fetching, decoding, executing, and writing back.

[0128] The CPU 3105 can be part of a circuit, such as an integrated circuit. One or more other components of the system 3101 can be included in the circuit. In some cases, the circuit is an application-specific integrated circuit (ASIC).

[0129] The storage unit 3115 can store files, such as drivers, libraries, and saved programs. The storage unit 3115 can store user data, such as user preferences and user programs. In some cases, the computer system 3101 can include one or more additional data storage units located outside the computer system 3101, such as on a remote server that communicates with the computer system 3101 via an intranet or the Internet.

[0130] The computer system 3101 can communicate with one or more remote computer systems via the network 3130. For example, the computer system 3101 can communicate with a remote computer system of a user (such as a medical professional or a patient). Examples of remote computer systems include personal computers (such as portable PCs), tablets or tablet computers (such as iPad, Galaxy Tab), telephones, smartphones (such as iPhone, Android - enabled devices, ) or personal digital assistants. A user can access the computer system 3101 via the network 3130.

[0131] The methods described herein can be implemented by machine (e.g., computer processor) - executable code stored at an electronic storage location of the computer system 3101, such as, for example, the memory 3110 or the electronic storage unit 3115. The machine - executable or machine - readable code can be provided in the form of software. In use, the processor 3105 can execute the code. In some cases, the code can be retrieved from the storage unit 3115 and stored in the memory 3110 for ready access by the processor 3105. In some cases, the electronic storage unit 3115 can be removed and the machine - executable instructions can be stored in the memory 3110.

[0132] The code can be pre - compiled and configured to be used with a machine having a processor suitable for executing the code, or it can be compiled at runtime. The code can be provided in a programming language that can be chosen such that the code can be executed in a pre - compiled or compiled manner.

[0133] Aspects of the systems and methods provided herein, such as computer system 3101, may be embodied in programming. Various aspects of the technology may be considered a "product" or "manufacture", which typically exists in the form of machine (or processor) executable code and / or associated data, carried on or embodied in one type of machine-readable medium. The machine executable code may be stored in an electronic storage unit such as a memory (e.g., read-only memory, random access memory, flash memory) or a hard disk. A "storage" type medium may include any or all of the tangible memories of a computer, processor, etc., or associated modules thereof, such as various semiconductor memories, tape drives, disk drives, etc., which can provide non-transitory storage for software programming at any time. All or part of the software can sometimes be communicated via the Internet or various other telecommunications networks. For example, such communication may load the software from one computer or processor to another, such as from a management server or a main computer to the computer platform of an application server. Thus, another type of medium that can carry software elements includes light waves, electrical waves, and electromagnetic waves, such as those used via a physical interface between local devices, via wired and fiber optic landline networks, and via various air links. Physical elements that carry such waves (such as wired or wireless links, fiber optic links, etc.) may also be considered a medium carrying software. As used herein, unless restricted to non-transitory, tangible "storage" media, the term such as computer or machine "readable medium" refers to any medium that participates in providing instructions to a processor for execution.

[0134] Thus, a machine-readable medium, such as computer-executable code, can take many forms, including but not limited to tangible storage media, carrier media, or physical transmission media. Non-volatile storage media include, for example, optical or magnetic disks, such as any storage device in any computer or the like, such as may be used to implement a database shown in the figures. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that make up the buses within a computer system. Carrier transmission media can take the form of electrical or electromagnetic signals or acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example: floppy disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punch cards, paper tapes, any other physical storage media with hole patterns, RAM, ROM, PROM, and EPROM, FLASH-EPROM, any other memory chip or cartridge, a carrier wave that transports data or instructions, a cable or link that transports such a carrier wave, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer-readable media may involve transporting one or more sequences of one or more instructions to a processor for execution.

[0135] Computer system 3101 may include or communicate with an electronic display 3135 that includes a user interface (UI) 3140 for providing, for example, input of biomarker or sequencing data or visual output related to detection, diagnosis, or prognosis. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.

[0136] The methods and systems of the present disclosure can be implemented by one or more algorithms. The algorithms can be implemented by software executed by a central processing unit 3105. For example, the algorithm can determine the presence or absence of cancer or cancer parameters based on a set of input sequencing data from a sample derived from an object.

[0137] Embodiments

[0138] Example 1: Analysis of cell-free DNA for cancer detection

[0139] Using the methods and systems of the present disclosure, circulating tumor DNA in urine is analyzed to identify genomic alterations in bladder cancer patients. Up to 90 mL of urine is collected into a tube containing a urine preservative buffer. Upon receipt of the sample, the sample is logged and the urine supernatant is immediately isolated and stored in an -80°C freezer. The ucfDNA is extracted and purified with beads and then quantified using Qubit, Agilent Bioanalyzer, or Fragment Analyzer. Up to 15 ng of size-selected ucfDNA is used for this assay. To construct the library, the extracted ucfDNA is tagged with unique molecular barcodes. The ligated sequencing library is PCR amplified with a high-fidelity polymerase and quantified with a Bioanalyzer. Target enrichment and hybridization capture are performed. During enrichment, the sequencing library is blocked with adaptors of specific blocking oligonucleotides and then hybridized with the PredicineCARE panel (such as Table 2). The library is captured with beads, amplified, and quantified by a Bioanalyzer. To perform sequencing, the enriched library is normalized, pooled, and loaded onto an Illumina platform for 2X 150bp paired-end sequencing. The library is sequenced to a median depth >20,000X.

[0140] The PredicineCARE assay identifies SNVs, copy number gains and losses, CNVs, and DNA rearrangements in a single workflow using a mature bioinformatics pipeline for high accuracy in variant detection. The NGS data is analyzed using the DeepSEA NGS analysis pipeline, which starts from the raw sequencing data (BCL files) and outputs the final mutation identification. Figure 2 A schematic diagram of the DeepSEA pipeline is shown. The pipeline performs adaptor trimming, barcode checking, and error correction. The cleaned paired FASTQ files are aligned to the human reference genome version hg19 using the BWA alignment tool. Then, the consensus sequence BAM file is derived by merging the paired-end reads originating from the same molecule as the single-stranded fragment. During this process, sequencing and PCR errors are corrected.

[0141] Next, SNV and copy number gain and loss detection are performed using a variant filter. The variants are filtered based on the variant background from a pool of normal control samples and other historical samples. Other metrics (such as base quality, log odds ratio, and distance to the fragment end) are used to remove variants with lower confidence. A detected identification refers to a variant with at least 4 unique supporting fragments and one of them should be double-stranded.

[0142] Variants are identified after filtering out low base quality and low mapping score reads. The detected variants are further filtered based on variant context (defined by normal plasma samples and historical data), repetitive regions, and other quality metrics. Benign and likely benign SNPs are excluded from the variant identification list.

[0143] If a matching normal sample is not available, a germline variant filter is performed. In practice, most liquid biopsy assays do not have a matching normal sample. When there is a copy number change at the variant position, the variant allele frequency is adjusted according to the copy number change. Variants with relatively high population allele frequencies annotated in public germline databases such as 1000 Genomes are also filtered out.

[0144] Copy number variations (CNVs) are estimated at the gene level. This process calculates the on-target unique fragment coverage based on the consensus sequence BAM file and then adjusts for GC bias coverage. Using the unique fragment coverage profiles of a set of normal samples as a reference, the copy number changes are normalized and z-scores are estimated. Before performing the consensus sequence operation, rearrangements are detected by identifying alignment breakpoints based on the BAM file. Suspect alignments are filtered based on repetitive regions, local entropy calculations, and the similarity between the reference alignment and alternative alignments. Greater than or equal to 2 unique alignments are used to report rearrangement identification.

[0145] To obtain the sample, urine was collected using a urine collection kit. The components of the urine collection kit are listed in Table 4. The kit contains one reagent, namely urine preservative. Streck Inc. (Omaha, NE, USA) is the supplier of the urine preservative in this kit.

[0146] Table 4. Components of the urine sample collection kit

[0147] Item Kit components 1 Centrifuge tube, 50 mL 2 Streck urine preservative, 5 mL 3 Disposable urine cup, 40 mL 4 Urine collection instruction manual / Instructions for use 5 Foam lining; general 6 Biohazard bag; general

[0148] To evaluate the PredicineCARE panel and assay for somatic alteration detection in ucfDNA, various samples with known mutation alterations (including SNVs, gains and losses, DNA rearrangements, copy number increases, and copy number losses) were used for this analytical validation. In this example, the limit of detection (LoD) is defined as the lowest mutant allele frequency at which the PredicineCARE assay can reliably detect 95% of the variants in all replicate experiments for a certain variant type. The percent positive agreement (PPA) is used to demonstrate assay sensitivity. A 15 ng ucfDNA sample input was used in the validation study.

[0149] To evaluate the LoD for single nucleotide variant (SNV) detection in urine specimens, ucfDNA with unique SNPs from healthy male donors was incorporated into ucfDNA from healthy female donors, generating a series of test materials with variable mutant allele frequencies (MAF) ranging from 0.125% to 10%. Among 90 expected mutations, 89 mutations were detected at 0.5% MAF, achieving a 98.89% assay sensitivity (95% CI, 94 - 100%) and 100% PPV (95% CI, 95.9 - 100%) (Table 5 and Figure 3 ).

[0150] Table 5. Analytical sensitivity of the PredicineCARE assay for SNV detection.

[0151]

[0152] To evaluate the LOD for copy number gains and losses, CNVs, and fusion detection in ucfDNA, we fragmented HD753 gDNA into ucfDNA-sized fragments and incorporated them into ucfDNA from healthy donors, generating samples with a predefined allele frequency (AF). HD753 reference gDNA from Horizon Discovery (Cambridge, UK) was used, which contains known copy number gains and losses, CNVs, and fusions. For copy number gains and losses, a total of 35 variants were detected at 0.5% to 1% AF, achieving a 97.22% assay sensitivity (95% CI, 85.5% - 99.9%) and 100% PPV (95% CI, 90% - 100%) (Table 6).

[0153] Table 6. Analytical sensitivity of the PredicineCARE assay for copy number gains and losses detection

[0154]

[0155] For copy number variant (CNV) detection, samples with titers between 2.375 and 3.125 copies were evaluated. All CNV variants were detected at 2.375 copy number, achieving 100% assay sensitivity (95% CI, 69.2 - 100%) and 100% PPV (95% CI, 69.2 - 100%) (Table 7).

[0156] Table 7. Analytical sensitivity of the PredicineCARE assay for CNV detection.

[0157]

[0158] For DNA rearrangement, samples with titration levels between 0% AF and 0.825% AF were analyzed. As shown in Table 5, the LoD for DNA rearrangement detection was 0.25% to 0.55% AF, achieving 100% detection sensitivity (95% CI, 78.2 - 100%) and 100% PPV (95% CI, 78.2 - 100%) (Table 8).

[0159] Table 8. Analytical sensitivity of the PredicineCARE assay for fusion detection.

[0160]

[0161] To evaluate assay specificity and ensure that "blank" samples do not generate analytical signals, 18 ucfDNA samples from healthy donors were tested. Analytical specificity was estimated based on the number of false - positive mutations in the target panel. Analytical specificity = 100 * (1 - number of false - positives / panel size). Although pathogenic mutations with low variant frequencies, such as CHIP mutations, may exist in healthy donors, we considered all variants that met the NGS analysis workflow detection criteria as false - positives. The analytical specificity was 99.9998%. Precision was measured by the variation in the estimated variant frequencies between replicate experiments. All samples for precision determination were evaluated from library construction operations to sequencing analysis.

[0162] The repeatability test included within - run performance (samples processed under the same conditions). Each sample was replicated three times under the same conditions, and the results were compared. Reproducibility was evaluated based on six samples independently processed under different operating conditions. Variant consistency identified from replicate experiments was used to evaluate within - assay and between - assay precision.

[0163] 100% consistency between replicate experiments was detected in both within - run and between - run precision studies. Additionally, a high degree of consistency was observed between the expected MAF and the detected MAF in the samples of the precision study (Table 9 and Figure 4 ).

[0164] Table 9. 100% consistency between replicate experiments detected in within - run and between - run precision studies

[0165]

[0166] Using DNA from 43 paired urine and tumor tissues of bladder cancer patients, the performance of urine-based variant detection was evaluated in clinical samples. Mutations detected in tissue samples were used as references for mutations detected in ucfDNA. Agreement was determined by comparing variants detected in urine with those in paired tissues. The agreement of mutations detected in urine and tissues was 81.0% (95% CI: 77.2 - 84.4%) (Table 10). Additionally, for five commonly mutated genes in bladder cancer (TERT, TP53, KDM6A, PIK3CA, and FGFR3)( Figure 5 ), the PredicineCARE assay achieved 94.6% agreement (95% CI: 87.9 - 98.2%).

[0167] Table 10. The PredicineCARE assay detected highly consistent genomic alterations in paired urine and tissue samples from bladder cancer patients.

[0168]

[0169] The PredicineCARE assay described in this example analyzed somatic variants of 152 cancer-related genes from liquid biopsies. The assay underwent rigorous analytical validation testing and showed stable and reproducible results in plasma cfDNA. As shown in this study, the PredicineCARE assay can detect gene variants in plasma cfDNA with high sensitivity and specificity (Table 11), making the assay a valuable tool for analyzing somatic variants in various body fluids. Combining completely non-invasive sample collection and the ease of urine-based liquid biopsy, the PredcineCARE assay has great potential in implementing real-time genomic profiling to detect cancer and enable more frequent monitoring.

[0170] Table 11. Summary of the performance of the PredicineCARE assay for variant detection in urine cfDNA

[0171]

[0172] Example 2: Detection of bladder cancer using urine

[0173] This example used PredicineCare urine cfDNA assays for research to evaluate the concordance between tissue tumor DNA analysis and urine cfDNA or circulating tumor DNA (ctDNA). This study prospectively enrolled 59 patients with pathologically confirmed bladder cancer and provided matched tissue / urine paired samples. Baseline peripheral blood mononuclear cells (PBMCs) and plasma specimens were collected during the visit (Zhang, et al. 2021). Urine, tissue, PBMC, and plasma samples were processed using the PredicineCARE assay and analyzed using the DeepSEA bioinformatics pipeline. Concordance analysis was performed using tissue tumor DNA as the reference. Urine cfDNA achieved a specificity of 99.3%, a sensitivity of 86.7%, and a diagnostic accuracy of 99.1%. FGFR3 alterations and ERBB2 amplifications were identified in urine cfDNA. Quantitative metrics including cancer cell fraction, variant allele frequency, and tumor mutation burden were concordant between tumor tissue DNA and urine cfDNA. Due to ctDNA aberrations caused by clonal hematopoiesis, the concordance between plasma cfDNA and tumor tissue DNA was low.

[0174] Example 3: Detection of bladder cancer using urine-based NGS assay

[0175] The clinical performance of the urine-based NGS assay was evaluated by comparing the results with those of an FDA-approved tissue-based PCR CDx assay that detects key alterations in the FGFR gene. Figure 11 A schematic of the study design is provided. Matched urine and tissue samples were collected from 107 (muscle- and non-muscle-invasive) bladder cancer patients in the German Bladder BRIDGister clinical trial. Tissue specimens were analyzed using the FDA-approved Qiagen therascreen FGFR RGQ RT-PCR kit, while matched urine samples were processed using the PredicineCARE TM urine (cell-free DNA) cfDNA NGS assay with a detection sensitivity of 0.3% (0.1% for hotspot mutations). The results of urine cfDNA NGS and tissue RT-PCR for 107 paired bladder cancers were analyzed to determine the concordance (PPA and NPA) between the two assays. In addition, tissue NGS was compared with therascreen RT-PCR and tissue NGS with urine cfDNA NGS using a smaller sample set. In addition, the discordant mutation subset was verified by ddPCR.

[0176] Consistency analysis between PredicineCARE urine cfDNA NGS and the therascreen FGFR RGQ RT-PCR kit for 107 samples showed a PPA of 100% (20 / 20, 95% CI: 83.2 - 100) and an NPA of 94.1% (48 / 51, 95% CI: 83.8 - 98.8). The PPA and NPA between PredicineCARE urine cfDNA and tissue NGS were 100% (19 / 19, 95% CI: 82.4 - 100) and 94.2% (49 / 52, 95% CI: 84.1 - 98.8), respectively. The PPA and NPA between PredicineCARE tissue NGS and the therascreen FGFR RGQ RT-PCR kit were both > 95%.

[0177] Figure 12 Heatmaps showing matched urine NGS and FFPE tissue RT-PCR for identified genetic alterations are presented. Each column represents a sample, adjacent columns are matched FFPE and urine, and each row represents a gene. There is a high degree of concordance between urine NGS and FFPE RT-PCR results. Most samples showed genetic alterations in both urine NGS and FFPE samples.

[0178] Figures 13A-13B A scatter plot of variant allele frequencies (VAF) between matched FFPE and urine variants is presented. Figure 13A A scatter plot of all identified genetic alterations (including somatic variants and germline variants) is presented. Figure 13B A scatter plot of somatic FGFR3 alterations is presented. The X-axis is the VAF of urine NGS, and the Y-axis is the VAF of FFPE RT-PCR.

[0179] Table 12. Summary of assay performance

[0180] WT* Mut* Failure / Invalid Assay success rate (%) PredicineCARE urine cfDNA 63 27 17 84.1%(90 / 107)** Therascreen tissue RT-PCR 63 21 23 78.5%(84 / 107)** PredicineCARE tissue gDNA 66 23 6 93.6%(89 / 95)***

[0181] *WT and Mut: Based on 4 SNVs and 5 fusions defined in the Therascreen RT-PCR assay.

[0182] **Samples (107) were selected from a set of matched urine and tissue samples for measuring assay performance.

[0183] ***This is from the comparison between Therascreen RT-PCR and PredicineCARE tissue gDNA NGS.

[0184] Allele frequency (AF) cut-off values for the reportable range of urine NGS (SNV and copy number gains and losses): 0.3% (hotspot: 0.1%)

[0185] Allele frequency (AF) cut-off values for the reportable range of FFPE NGS (SNV and copy number gains and losses): 5% (hotspot: 2%)

[0186] Allele frequency (AF) cut-off values for the reportable range of fusions: urine (0.1%) and tissue (1%)

[0187] Table 13. PredicineCARE TM Consistency results between urine NGS and the QIAGEN therascreen FGFR RGQ RT-PCR kit

[0188]

[0189] Note: FGFR+ and FGFR- are based on 4 previously defined SNVs and 5 fusions included in the Qiagen therascreen FGFR RGQ RT-PCR kit.

[0190] Three RT-PCR-tested tissue FGFR-negative samples were detected as FGFR-positive by urine NGS testing, while 0 urine NGS FGFR-negative samples were detected as FGFR-positive by tissue RT-PCR (Table 13). Samples with discrepancies (tissue FGFR WT or invalid but urine FGFR mutation positive) were further analyzed and confirmed as positive by independent orthogonal ddPCR (Bio-Rad ddPCR mutation detection assay), indicating that the inconsistency between urine and tissue results is usually due to the reduced sensitivity of the tissue FGFR RT-PCR test.

[0191] The high consistency between FGFR alterations detected using an FDA-approved tissue companion diagnostic (CDx) assay and urine cfDNA NGS assay indicates that the PredicineCARE urine cfDNA NGS assay represents a novel, accurate, and non-invasive clinical application for molecular diagnostic testing to identify biomarkers in bladder cancer.

[0192] Example 4: Determination of copy number burden using low-throughput whole-genome sequencing

[0193] Copy number variation (CNV) is an important feature of the cancer genome. Low-pass whole-genome sequencing (LP-WGS) based on blood / urine has been increasingly used to identify copy number variations in large regions of the cancer genome. In this study, we developed PredicineCNB TM, an accompanying LP-WGS assay to robustly estimate the whole-genome copy number burden (CNB) from plasma and urine clinical samples, thereby providing cancer / normal classification and longitudinal therapy monitoring in a cost-effective manner (e.g., using PredicineSCORE TM liquid biopsy assay (Predicine, Inc., Hayward, CA)).

[0194] Based on the analysis and evaluation of clinical plasma titration samples, PredicineCNB TM has been demonstrated to have high cancer detection sensitivity (LOD of 1%) and high specificity, with a DNA input as low as 1 ng. Figure 14A Shows the PredicineCNB assay workflow. Blood (or urine sample) is obtained. For blood samples, plasma is separated and DNA is extracted. Then a library is prepared and sequenced at 3x depth using low-throughput whole-genome sequencing, followed by analysis. Figure 14B Shows an illustration of whole-genome CNV detection, and for as little as 0.5 ng of plasma input, CNB calculation remains robust, showing 5 ng of input plasma, CNB score = 11.7 (upper figure); 0.5 ng of input plasma, CNB score = 11.7 (lower figure). Figure 14C Shows that the LP-WGS CNV profiles of 5 ng and 0.5 ng of input plasma are highly consistent in the 1 Mb region, and Figure 14D shows that the LP-WGS CNV copy numbers between 5 ng and 0.5 ng of input plasma are highly consistent at the chromosomal arm level, with a correlation coefficient even higher than that in the 1 Mb region.

[0195] Figures 15A-15D Shows the analysis and evaluation of PredicineCNB on clinical plasma samples and its application in 688 prostate cancer patient samples. As Figure 15A shown, the PredicineCNB LOD is 1%, and its specificity > 97.6% (41 / 42). CNB scores were generated for cancer plasma titration samples and corresponding normal plasma baseline samples, as Figure 15B shown. As Figure 15C shown, the CNB score distribution of 688 clinical prostate cancer patients was generated, among which 430 samples were classified as high-risk and 258 samples were classified as low-risk. Figure 15D Shows a heatmap of the LP-WGS CNV profiles of 430 prostate cancer patient samples classified as high-risk, where each row represents the copy number deviation of 1 Mb bins on all chromosomes of each clinical sample.

[0196] The assay was also performed on patients with non-muscle-invasive and muscle-invasive bladder cancer. Figures 16A-16BShows an overview of the LP-WGS CNV profile heatmap of patient samples with non-muscle-invasive and muscle-invasive bladder cancer. Figure 16A Shows the LP-WGS CNV profile heatmap of 14 patient samples with non-muscle-invasive bladder cancer. Figure 16B Shows the LP-WGS CNV profile heatmap of 33 patient samples with muscle-invasive bladder cancer.

[0197] Tested the effectiveness of PredicineCNB for longitudinal therapy monitoring of bladder cancer patients. Figure 17A Shows the comparison of CNB scores of FFPE, plasma, and urine samples between patients with non-muscle-invasive bladder cancer and patients with muscle-invasive / non-organ-confined bladder cancer. Figure 17B Shows the LP-WGS CNV profiles of two patients with non-muscle-invasive bladder cancer before and after TURBT surgery, showing that the CNB scores decreased in both patients after TURBT surgery. Figure 17C Shows the LP-WGS gene copy numbers of key bladder cancer genes before and after TURBT, demonstrating the monitoring of different cancer genes before and after surgery.

[0198] The examples show the algorithm analysis process of LP-WGS determination. The copy number burden (CNB) LOD is 1% tumor fraction, with a plasma input of 1 ng. PredicineCNB TM Demonstrates its high sensitivity for cancer detection and good clinical application prospects of whole-genome copy number changes based on urine and blood in therapy monitoring.

[0199] References

[0200] [1] Davis AA, Luo J, Zheng T, Dai C, et al. Genomic complexity predicts resistance to endocrine therapy and CDK4 / 6 inhibition in hormone receptor-positive (HR+) / HER2-negative metastatic breast cancer. Clin Cancer Res. 2023 Jan 24; CCR-22-2177. doi: 10.1158 / 1078-0432.CCR-22-2177, which is incorporated herein by reference in its entirety.

[0201] Example 5: Detection and monitoring using methylation

[0202] Using the methods and systems of the present disclosure, circulating tumor DNA from a biological sample was analyzed. A combined MRD assay was performed using a targeted panel (Predicine WES+) that covers hot spot mutations and important genes, an accompanying LP-WGS assay for copy number burden (Predicine CNB), and a whole genome methylation assay (e.g., Predicine EPIC).

[0203] Figure 18 An example schematic diagram of the workflow is shown. A sample (blood, urine, or tissue) is obtained from an object. DNA is extracted and a library is generated. The library is methylated. After library construction is completed, next-generation sequencing is performed, and mutation, CNV, and methylation data are analyzed.

[0204] The analysis and evaluation were based on titration of real-world clinical patient samples. The DNA input amount ranged from as low as 1 ng to 30 ng. The clinical evaluation was based on longitudinal samples from patients with different cancer indications (including bladder cancer, mCRPC, CRC, breast cancer, and NSCLC).

[0205] Figure 19 A comparison of Predicine EPIC libraries with 1 ng, 2.5 ng, and 10 ng input to a standard 50 ng DNA input for a standard whole genome bisulfite sequencing library is shown. Specifically, compared to the standard 50 ng whole genome bisulfite sequencing library, the Predicine EPIC libraries with 1 ng, 2.5 ng, and 10 ng provided similar coverage and data.

[0206] 97 cancer and normal tissue types were analyzed using fragment-level DNA methylation analysis. These assays were able to identify fragments with increased or decreased methylation compared to a background model. These assays were able to group different cancers according to tissue type and whether the sample was cancerous or non-cancerous. Figure 20 A graph showing Uniform Manifold Approximation and Projection (UMAP) is shown. Different populations were visualized based on different UMAP scores.

[0207] For each sample, a methylation aberration score can be calculated based on the methylation data. Figure 21 An example heatmap of the significance scores (color scale) of aberrant methylation fragments at 28.4K of the most variable CpG sites (columns) out of 142K genome-wide covered sites for 35 bladder cancer samples (rows) from patients at different stages and different sources (row-annotated) is shown.

[0208] Combined with the methylation aberration score, a copy number burden aberration score can also be calculated based on low-throughput sequencing data. As Figure 22As shown, these two abnormal scores are shown to be highly correlated and can be used jointly to identify whether an object has cancer. In addition, as Figure 22 shown, a DNA methylation score can be used to calculate the disease progression of a patient. Three patients are shown, where T0 is before treatment, and T1 and T2 are after treatment, and the abnormal scores decreased significantly after treatment.

[0209] Example 6: Enhancement of whole exome sequencing for muscle-invasive bladder cancer

[0210] Urine tumor DNA profiling can be used for the diagnosis, monitoring, and treatment stratification of bladder cancer. However, previous studies mainly used targeted next-generation sequencing (NGS) panel methods, which were limited to predefined genes and thus lacked comprehensiveness. Here, this example demonstrates the use of enhanced whole-exome sequencing (WES) for urine and tissue tumor DNA of muscle-invasive bladder cancer (MIBC) to comprehensively compare the mutation spectra in matched urine and tissue samples.

[0211] Matched tumor tissue, urine, and peripheral blood mononuclear cell (PBMC) samples were collected from 20 MIBC patients. Nineteen tumor tissue samples, nineteen urine samples, and twenty PBMC samples that passed sample quality control were processed for NGS. PredicineWES+ is an NGS assay from the PredicineATLAS panel with whole-exome coverage and enhanced coverage of 600 cancer-related genes, which was applied to the matched tumor, urine, and PBMC samples for variant spectrum analysis. The mutation spectra of tumor tissue and urine DNA were analyzed and compared.

[0212] Figure 23 The mutation spectra of urine and tissue tumor DNA from MIBC are shown. The mutation spectra of urine and tissue tumor DNA are highly consistent among patients, and the incidence rates of common mutant genes (TERT, TP53, ARID1A, KMT2D, KDMSA, PIK3CA, etc.) are comparable. Two tissue samples that failed sequencing QC are not shown.

[0213] Consistency of mutations detected from tissue tumor DNA (tDNA) and urine tumor DNA (utDNA) by PredicineWES+. The number of mutations detected from tDNA and utDNA in the WES region ( Figure 24 A - B) and the ATLAS region ( Figure 24 C - D) was compared. Most tDNA mutations (67.5% in the WES region and 80.1% in the ATLAS region) A) were also detected in utDNA. However, less than half of the utDNA mutations (42.1% in the WES region and 39.9% in the ATLAS region) were detected in tDNA.

[0214] The tumor fraction (TF) was inferred from paired urine (2 - 52%) and tumor tissue (17 - 68%), which also showed a significant difference (a, p = 0.05). Figure 25 Graph showing tumor fractions in tissue and urine. Although the TF in urine was relatively low, more somatic mutations were detected in urine than in tumor tissue (b, p < 0.05). The tumor mutation burden (TMB) of tDNA and utDNA was also calculated and compared, showing a high correlation (R = 0.84).

[0215] Overall, the results demonstrated the effectiveness of urine tumor DNA as a tissue surrogate for whole - exome - scale MIBC mutation profiling, supporting the application of urine - based non - invasive molecular profiling in precision medicine for bladder cancer patients.

[0216] PredicineWES+ TM (Enhanced whole - exome sequencing) identified 1,493 somatic variants in 11 CSF samples, of which 97 variants had previously been reported as likely pathogenic by public clinical databases. For NSCLC - specific biomarkers, 7 out of 11 patients carried EGFR variants, including 3 exon 19 deletions, 1 exon 20 insertion, 1 L858R mutation, and other gain - of - function mutations. In addition, an EML - ALK mutation was detected in 1 patient.

[0217] Example 7: Monitoring of breast cancer using blood samples

[0218] We performed two comprehensive NGS assays to profile somatic mutations and copy number variations in blood samples collected from HR+ / HER2 - negative metastatic breast cancer patients at baseline and during treatment with endocrine therapy and CIDK4 / 6 inhibition (ET+CDK4 / 6i). Specifically, blood samples from a phase II study of palbociclib plus letrozole or fulvestrant, with a weekly schedule of 5 days on / 2 days off and a 28 - day cycle, were evaluated as first - line or second - line treatment.

[0219] The first assay used was PredicineWES+, an enhanced whole - exome sequencing (WES) assay that combines WES with deep coverage of 600 cancer genes targeted by the PredicineATLAS panel to generate a whole - exome genomic profile of somatic single - nucleotide variants (SNVs), indels, and copy number variations (CNVs), and to determine a blood tumor mutation burden (bTMB) score reflecting the number of DNA mutations.

[0220] The second assay is PredicineCNB, a low-pass whole-genome sequencing (LP-WGS) assay used to generate a blood copy number burden (bCNB) score, which represents a comprehensive measure of copy number variation, including amplifications and deletions on all chromosomal arms (e.g., using PredicineSCORE TM liquid biopsy assay (Predicine, Inc., Hayward, CA)).

[0221] Figures 26A-26C Shows the study schematic, sample collection, and assay timeline and procedures. Figure 26A Shows the study schematic. Figure 26B Shows the timeline of sampling and sequencing. Sample collection was performed at baseline (BL), day 15 of cycle 1 during treatment (C1D15), day 1 of cycle 2 (C2D1), Q3 monthly staging scan with no progressive disease (PD), and at the time of radiographic detection of PD. A total of 216 consecutive blood samples were collected from 51 patients at baseline and during treatment. In addition, germline DNA samples were collected from each patient. After QC procedures, 78 blood samples were sequenced using PredicineWES+ and 218 blood samples were sequenced using PredicineCNB. Similarly, 49 germline samples were sequenced using PredicineWES+. Figure 26C Shows the NGS profiling using PredicineWES+ and PredicineCNB. Blood samples were separated into plasma samples containing cfDNA, buffy coats containing gDNA, and red blood cells (RBC). DNA was extracted from the samples and libraries were constructed. Two different sequencing workflows were performed on the libraries. For a portion of the samples, low-pass whole-genome sequencing at 5x depth was performed and reads were analyzed to determine the blood copy number burden. For the second portion of the samples, targeted enrichment was performed for whole exome sequencing (2500x depth) and additional enrichment of 600 specific target genes of Predicine ATLAS (20,000x depth). Then, these reads were used to identify SNVs, gains and losses, CNVs, fusions, and to determine the tumor mutation burden.

[0222] Figures 27A-27B Shows the data of the PredicineWES+ assay. Figure 27A Shows a heatmap of the genes most significantly altered in baseline and progression samples. The number and type of genomic alterations and the frequency of alterations in specific genes in baseline and progression samples were also analyzed. Based on these data, enrichment of stage-specific variants could be identified and a decrease in the total number of SNVs at progression was observed. Figure 27BShows the status related to changes in the baseline as compared to the advanced stage. However, an increase in the total number of CNVs (mainly copy loss events) was observed in the progression samples. When observing the blood tumor mutation burden in the baseline and progression samples, the median level did not change.

[0223] When observing the blood copy number burden (bCNB), the medians for the baseline and progression were also similar. However, bCNB was able to track treatment progression. Figure 28 A series of bCNBs during treatment are shown, and in 12 out of 18 patients (66.7%) analyzed for staged blood samples, bCNB decreased at C1D15 and / or C2D1 and then increased prior to radiographic detection of disease progression. Thus, bCNB was able to track treatment and predict progression prior to any radiographic detection.

[0224] As shown in the examples, WES in plasma is a highly sensitive and comprehensive NGS method that can be used to detect individual variants at baseline and during treatment, some of which are significantly enriched during the disease progression phase. However, NGS assays designed around specific variants for monitoring disease progression are costly. Therefore, a shallow LP-WGS assay can be used to detect the dynamic changes in CNVs during treatment prior to radiographic detection of recurrence. This method is a promising and cost-effective approach for continuously monitoring early signs of metastatic disease progression during treatment.

[0225] Example 8: Application of ctDNA in cerebrospinal fluid for cancer detection

[0226] More than 3% of patients diagnosed with non-small cell lung cancer (NSCLC) will develop leptomeningeal metastases throughout the course of the disease, leading to poor clinical outcomes and limited therapies. Cerebrospinal fluid (CSF) is a direct liquid biopsy for the pathological diagnosis of leptomeningeal metastases. However, traditional clinical methods for detecting tumor cells in CSF show limited sensitivity. At the same time, the unique genomic aberrations of leptomeningeal metastases remain unclear. We hereby report a prospective clinical study aimed at identifying the genomic aberrations carried by NSCLC patients with leptomeningeal metastases through circulating tumor DNA (ctDNA) in CSF.

[0227] Thirteen patients were included in this study, and CSF samples were collected after diagnosis of metastases. Among them, PBMC samples were collected from 11 patients as germline control materials. A PredicineCNB (low-pass whole-genome sequencing (LP-WGS)) assay was performed to identify copy number variations and tumor fractions in the CSF samples from all 13 patients. In addition, PredicineWES+ (enhanced whole-exome sequencing assay) was also performed on paired cerebrospinal fluid and PBMC samples from 11 patients.

[0228] By PredicineCNB TM determination (e.g., using PredicineSCORE TM liquid biopsy assay (Predicine, Inc., Hayward, CA)) identified ctDNA fractions in all 13 CSF samples. Gene copy variants associated with NSCLC were also detected, such as increased copy numbers of EGFR (7 pts), BRAF (5 pts), MET (5 pts), KRAS (2 pts), ERBB2 (2 pts), ROS1 (2 pts), ALK (1 pts), and decreased copy numbers of RB1 (4 pts), PTEN (2 pts), TP53 (1 pts). Figure 29 Shows copy number increases and losses determined by PredicineCNB.

[0229] PredicineWES+ TM (Enhanced whole exome sequencing) identified 1,493 somatic variants in 11 CSF samples, of which 97 variants had previously been reported as likely pathogenic by public clinical databases. For NSCLC-specific biomarkers, 7 out of 11 patients carried EGFR variants, including 3 exon 19 deletion events, 1 exon 20 insertion event, 1 L858R mutation, and other gain-of-function mutations. Figure 30 Shows the mutation landscape determined by PredicineWES+.

[0230] Example 9: Analysis of TMB and MSI using liquid biopsy

[0231] Tumor mutational burden (TMB) and microsatellite instability (MSI) are emerging biomarkers associated with response to immunotherapy. PredicineATLAS is a proprietary NGS-based assay capable of robust measurement of TMB and MSI in cell-free circulating DNA (cfDNA) extracted from blood samples. This report summarizes the analytical validation of the PredicineATLAS assay, including accuracy, specificity, limit of detection - lowest tumor content (LOD), and precision (repeatability and reproducibility).

[0232] The PredicineATLAS assay uses a 600-genome panel and is designed to measure TMB and MSI in liquid biopsy samples collected from cancer patients. cfDNA is extracted, labeled with unique molecular barcodes during library construction, then enriched using the PredicineATLAS panel, and then sequenced on an Illumina platform using paired-end reads.

[0233] For blood samples, 10 mL of peripheral venous blood was collected in Streck cell-free DNA BCT. Upon receipt of the sample, it was immediately processed into plasma and stored at -80 °C. cfDNA was extracted using the QIAamp Circulating Nucleic Acid Kit and quantified using Qubit. Genomic DNA (gDNA) samples from cell lines were enzymatically digested and serially size selected to mimic the plasma cfDNA profile. The extracted cfDNA was labeled with unique molecular barcodes, and the ligated sequencing libraries were PCR amplified using a high-fidelity polymerase pair and quantified using a Bioanalyzer. For enrichment, the sequencing libraries were blocked with adaptors containing specific blocking oligonucleotides and hybridized to the PredicineATLAS panel. The bead-bound capture libraries were amplified and quantified using a Bioanalyzer. The enriched libraries were normalized, pooled, and loaded onto an Illumina platform for 2X150 bp paired-end sequencing. The libraries were sequenced to a median depth >20,000X.

[0234] To achieve accurate and robust TMB estimation, only highly confident somatic single nucleotide variant (SNV) mutations in the targeted coding regions were considered in the TMB calculation.

[0235] The NGS data was analyzed using Predicine's DeepSEA NGS analysis pipeline, which starts from the raw sequencing data (BCL files) and outputs the final variant calls. The pipeline first performs adapter trimming, barcode checking, and error correction. The cleaned paired FASTQ files were aligned to the human reference genome version hg19 using the BWA alignment tool. Then, the consensus sequence bam file was derived by merging the paired-end reads originating from the same molecule of the single-stranded fragment. Subsequently, the single-stranded fragments from the same double-stranded DNA molecule were further merged into double-stranded. During this process, both sequencing and PCR errors were corrected.

[0236] After using the DeepSEA variant caller, the variants were filtered based on the variant background from a pool of normal control samples and other historical samples. Other metrics (such as base quality, log odds ratio, and distance to the fragment end) were used to remove variants with lower confidence. The detected calls refer to variants with at least 4 unique supporting fragments and one of them should be double-stranded.

[0237] When no matched normal sample is available, the germline variant filter can be executed. In many applications, liquid biopsy assays may not include a matched normal sample. The germline variant filter can assume that the variant allele frequencies of somatic mutations originating from tumors are much lower than those of heterozygous germline variants. It can be assumed that variants with high allele frequencies are derived from the germline. When copy number changes occur at the variant locus, the variant allele frequencies are adjusted according to the copy number changes. Variants annotated with relatively high population allele frequencies in public germline databases are also filtered out.

[0238] The TMB score is estimated based on the number of somatic SNVs in the coding region and is normalized by the total coding region size covered by the panel. If the MSAF (maximum somatic allele frequency) is less than the threshold, the TMB score is not estimated.

[0239] The MSI score of a tumor sample is evaluated by counting the number of instability markers. For MSI detection, the PredicineATLAS assay analyzes 50 MSI markers in the panel, which are short tandem repeat regions in the reference genome. If the MSI score is greater than the threshold defined in the titration experiment ( Figure 6 ), the tumor sample is predicted to be MSI-high (MSI-H). For each MSI marker, the z-score is calculated by comparing the repeat length distribution of the tumor sample and the normal background constructed from a batch of normal plasma samples. If the z-score is higher than the threshold estimated from the validation data, the marker is considered unstable. Before constructing the repeat length distribution, the DeepSEA algorithm is used to suppress PCR or sequencing noise through error correction.

[0240] The PredicineATLAS panel contains 600 cancer-related genes, covering a genomic region of 2.4 Mb and a coding region of 1.36 Mb.

[0241] To evaluate the PredicineATLAS panel for TMB measurement, the TMB score obtained from the assay was compared with the TMB score calculated from publicly available WES data. 7116 tissue samples were used in the analysis, which covered >30 tumor types from publicly available TCGA data downloaded from Broad GDAC Firehose (https: / / gdac.broadinstitute.org / ). This in silico analysis showed that the TMB score based on the PredicineATLAS panel was highly correlated with the TMB score based on WES (R = 0.98, P < 0.001).

[0242] Performance metrics of Predicine NGS panel

[0243] Based on a standard 15 ng DNA input, Table 14 summarizes the satisfactory assay performance for identifying base substitutions, indels, rearrangements, copy number gains (CNG), and copy number losses (CNL).

[0244] Table 14. Summary of performance metrics

[0245]

[0246]

[0247] Eight cell lines with WES data available from the COSMIC database (https: / / cancer.sanger.ac.uk / cosmic / ) were tested by the PredicineATLAS assay for TMB measurement. The TMB scores of the PredicineATLAS assay were shown to be highly correlated with the TMB from publicly available WES data (R = 0.97, P < 0.001).

[0248] LoD of the assay

[0249] For TMB, the LoD was defined as the lowest tumor content required to achieve at least 90% concordance between the detected TMB status and the expected TMB status. Four cell lines with different TMBs were titrated to five different tumor content levels (20%, 10%, 1%, 0.5%, and 0.25%), and each level was replicated at least twice. The tumor content of the resulting diluted samples ranged from 0.25% to 20%. All samples were fragmented and sized appropriately to mimic the size of plasma cfDNA. Since these cancer cell lines had copy number variations across the genome, the MAF (mutant allele frequency) ranged from 10% to 90%. To facilitate TMB LoD assessment, TMB LoD assessment only included cell line mutations with MAF > 35% and no CNV regions. A high correlation was observed between the TMB scores of the PredicineATLAS panel and the WES data from COSMIC.

[0250] To evaluate the LoD, the concordance between the detected TMB status and the expected TMB status was evaluated at different TMB cutoffs (5, 10, 15, 20, 30 mutations / Mb). At LoD 1%, the assay achieved at least 90% concordance.

[0251] For MSI, the LoD is defined as the lowest titration level at which at least 90% of MSI-high samples can be detected as MSI-H samples. To determine the LoD of MSI, MSI-H cell lines were serially diluted in MSS cell lines targeting multiple titration levels (20%, 10%, 1%, 0.5%, 0.25%). Samples with >1% tumor cells were 100% detected as MSI-H, and at the 1% titration level, 90.9% of the samples were identified as MSI-H. Thus, the MSI assay achieved 90.9% concordance even with a tumor content as low as 1%.

[0252] Accuracy was determined by comparing the TMB scores calculated from a series of titrated cell lines with the expected TMB scores at a TMB cut-off of 10 mutations / Mb. A total of 39 samples with a tumor content ≥1% were used to evaluate TMB accuracy. All samples with high TMB (TMB score ≥10) were classified as high, and all samples with low TMB (TMB score <10) were classified as low. No samples were misclassified, and thus the PPV of the PredicineATLAS TMB assay was determined to be 100%.

[0253] For MSI, accuracy was determined by comparing the MSI status calculated from a series of titrated cell lines with the expected MSI status. A total of 21 MSI samples and 9 MSS samples with a tumor content ≥1% were used to calculate MSI accuracy. The MSI accuracy was determined to be 96.7% (95% CI: 82.8 - 99.9%).

[0254] To evaluate the performance of the PredicineATLAS assay and ensure that "blank" samples do not generate analytical signals, the TMB and MSI scores of cfDNA from healthy donors were evaluated using the PredicineATLAS assay for specificity. The TMB scores of all samples from healthy donors were 0 (AF cut-off threshold = 0.5%) and were predicted to be MSS( Figure 6 ). For MSI, 100% (95% CI: [69.2 - 100%]) of the samples were classified as TMB low at a cut-off of 10 mutations / Mb and MSS at a cut-off of 14. These baseline data indicate that the PredicineATLAS assay has 100% specificity for TMB and MSI measurements.

[0255] Repeatability (intra-assay precision)

[0256] To evaluate the concordance between repeated tests of the same sample under the same operating conditions, TMB and MSI analyses were performed in 8 replicate groups (with at least 2 replicate experiments per group) under the same operating conditions. High similarity was observed for all replicate samples.

[0257] To evaluate the consistency between measurement results when operating conditions change, the reproducibility among different reagent batches, operators, and sequencers was evaluated and compared. Groups of 6 samples with high or low TMB were analyzed under different conditions, and the TMB scores between replicate experiments were compared. Groups of 8 samples with MSI-H status were analyzed under different conditions, and the MSI status between replicate experiments was compared. High consistency was observed in all replicate samples used for TMB and MSI studies. Over a seven-month period, cell lines at 1% titration levels were processed multiple times using the PredicineATLAS assay. The measured TMB scores (sorted by processing time) are as Figure 7 shown. The difference between the measured TMB scores and the expected TMB score (157.36) was between -5.6% and 6.5%, indicating that the PredicineATLAS assay has high reproducibility.

[0258] PredicineATLAS analyzes 600 cancer-related genes in a single assay to provide an assessment of immunotherapy biomarkers (TMB and MSI). The assay underwent rigorous analytical validation testing using a low cfDNA input volume of 15 ng. The validation study showed high consistency between PredicineATLAS and WES in the assessment of TMB accuracy. In addition, the assessment of MSI status also showed high consistency with cell line samples. The development of the cfDNA-based PredicineATLAS assay provides a complementary method to tissue-based TMB and MSI assays for patients with solid tumors. PredicineATLAS is a robust and efficient assay that allows for the simultaneous measurement of TMB and MSI in liquid biopsy samples.

[0259] Example 10: Blood cancer detection

[0260] Blood cancers mainly originate from the bone marrow and account for approximately 10% of all newly diagnosed cancer patients. As an alternative to the current invasive bone marrow biopsy used to monitor blood cancers, comprehensive genomic profiling of cell-free circulating DNA (cfDNA) or peripheral blood mononuclear cells (PBMC) in the blood provides a minimally invasive and clinically convenient solution for detecting genomic biomarkers to guide clinical decision-making in oncology.

[0261] PredicineHEME TMIt is a capture-based targeted next-generation sequencing (NGS) assay that can accurately detect genomic alterations in plasma cfDNA or genomic DNA (gDNA) extracted from peripheral blood mononuclear cells (PBMCs) or bone marrow aspirates (BMAs) of patients with hematological malignancies, including small nucleotide variants (SNVs), insertions and deletions (indels), copy number variants (CNVs), and DNA rearrangements.

[0262] PredicineHEME TM is unique in its ability to detect variant allele frequencies as low as 0.1% in plasma cfDNA or blood / BMA gDNA. This example summarizes the analytical validation of the PredicineHEME TM assay, including accuracy, specificity, sensitivity, and precision. PredicineHEME TM analyzed potential genomic biomarkers of 106 key blood cancer-related genes in patients with blood cancers. The capture-based NGS assay is designed to detect SNVs, indels, copy number variations, and gene rearrangements in plasma, blood, or BMA samples collected from patients with blood cancers.

[0263] During library construction, the extracted plasma cfDNA or enzymatically fragmented gDNA is labeled with unique molecular barcodes and then enriched using the PredicineHEME TM panel, followed by paired-end sequencing using the Illumina platform. If plasma separation is required, peripheral venous blood is collected in Streck cell-free DNA BCT; otherwise, whole blood samples are collected using EDTA tubes. After receiving the samples, they are registered and immediately processed into plasma or buffy coat and stored at -80 °C.

[0264] Plasma cfDNA, as well as PBMC or BMA genomic DNA (gDNA), is extracted using the QIAamp Circulating Nucleic Acid Kit. Samples are extracted using the DNeasy Blood and Tissue Kit. Genomic DNA is further enzymatically digested and purified before DNA quantification and qualification. The extracted cfDNA or fragmented gDNA is labeled with unique molecular barcodes, and the ligated sequencing library is PCR amplified using high-fidelity polymerase and quantified using a Bioanalyzer. For enrichment, the sequencing library is blocked with adaptors containing specific blocking oligonucleotides and hybridized with the PredicineHEME TM panel. The bead-bound capture library is amplified and quantified using a Bioanalyzer. The enriched library is normalized, pooled, and loaded onto the Illumina platform for 2X150bp paired-end sequencing. The library is sequenced to a median depth >20,000X.

[0265] Combined with the in-house developed DeepSEA variant identification software for high-accuracy variant detection, PredicineHEME TM Evaluates SNVs, indels, rearrangements, and CNVs in a workflow. NGS data is analyzed using Predicine's proprietary DeepSEA NGS analysis pipeline, which starts from raw sequencing data (BCL files) and outputs final mutation identification. The pipeline performs adapter trimming, barcode checking, and error correction. Then the cleaned paired FASTQ files are aligned to the human reference genome version hg19 using the BWA alignment tool. Then a consensus sequence BAM file is obtained by merging paired-end reads originating from the same molecule as the single-stranded fragment. Then single-stranded fragments from the same double-stranded DNA molecule are further merged into double-stranded. During this process, both sequencing and PCR errors are corrected.

[0266] After the DeepSEA variant caller, variants are filtered based on the variant background from a pool of normal control samples and other historical samples. Other metrics (such as base quality, log odds ratio, entropy, and distance to fragment end) are used to remove variants with lower confidence.

[0267] When a matching normal sample is not available, a germline variant filter can be performed. When there is a copy number change at the variant locus, the variant allele frequency is adjusted according to the copy number change. It is assumed that variants with an adjusted allele frequency close to 50% or higher than 90% are from the germline. Variants annotated in public germline databases (such as the 1000 Genomes Project) with relatively high population allele frequencies are also filtered out.

[0268] After filtering reads with low base quality and low mapping scores, SNV and indel variants are identified. The detected variants are further filtered based on the variant background (defined by normal plasma samples and historical data), repeat regions, and other quality metrics. Benign and likely benign SNPs are excluded from the variant identification list. Copy number variations (CNVs) are estimated at the gene level. The pipeline first calculates the on-target unique fragment coverage based on the consensus sequence bam file, and then adjusts for GC bias coverage. Using the unique fragment coverage profile of a set of normal samples as a reference, copy number changes are normalized and z-scores are estimated. DNA rearrangements are detected by identifying alignment breakpoints based on the bam file before performing consensus sequence operations. Suspect alignments are filtered based on repeat regions, local entropy calculations, and similarity between the reference alignment and alternative alignments. Two or more unique alignments are used to report DNA fusions.

[0269] To assess the accuracy of the assay, reference samples with known genomic alterations present in multiple genes being assayed were tested. These samples included cell lines, commercial reference samples, and patient samples tested by a third-party laboratory. PredicineHEME TM The assay analyzes a maximum of 30 ng of cfDNA or fragmented gDNA input.

[0270] Agreement was determined by comparing the expected variants of the reference samples with the variants detected by the assay and used as a performance metric to demonstrate accuracy. The panel assay was expected to detect a total of 102 mutations in 41 samples, and all mutations were confirmed with 100% PPA [96.6% - 100%]. In this validation study, the limit of detection (LoD) was defined as PredicineHEME TM the lowest mutant allele frequency at which the PredicineHEME assay reliably detected a particular variant type in 95% of replicate experiments. Studies were conducted to demonstrate the putative sensitivity values for each variant type (including SNV, copy number gains and losses, DNA rearrangements, and CNG). Reference cfDNA samples with defined MAFs at 4 different levels (0.1, 0.25, 0.375, 0.5%) were evaluated to define the LOD for each variant type. Both PPA and PPV were calculated for each mutation level and used to define the LoD.

[0271] For SNV, a total of 80 expected variants with an expected AF of 0.25% were evaluated, and the sensitivity achieved by the assay was 98.75% (95% CI, 93.2 - 100%), and the PPV was 97.53% (95% CI, 91.4 - 99.7%). For copy number gains and losses, a total of 20 expected variants with an expected AF of 0.375% were evaluated, and the sensitivity achieved by the assay was 100% (95% CI, 83.2 - 100%), and the PPV was 100% (95% CI, 83.2 - 100%). For DNA rearrangements, a total of 20 expected variants with an expected AF of 0.375% were evaluated, and the sensitivity achieved by the assay was 95% (95% CI, 75.1 - 99.9%), and the PPV was 100% (95% CI, 82.4 - 100%). For CNG, a total of 20 expected variants with an expected 2.23 copies were evaluated, and the sensitivity achieved by the assay was 100% (95% CI, 83.2 - 100%), and the PPV was 100% (95% CI, 83.2 - 100%). Based on the analytical validation data, the SNV LoD for the MYC gene was 0.25% MAF, the copy number gains and losses LoD was 0.375% MAF, the DNA rearrangement LoD was 0.375% MAF, and the CNG LoD was 2.23 copies.

[0272] To evaluate the specificity of the assay and ensure that “blank” samples do not generate analytical signals, 24 samples from healthy donors, including buffy coat, plasma, and BMA, were tested. Analytical specificity was estimated based on the number of false-positive mutations in the target panel. Analytical specificity = 100*(1 - number of false positives / panel size).

[0273] Although pathogenic mutations with low variant frequencies can also be present in healthy donors, such as CHIP mutations. We considered all variants passing the NGS analysis workflow detection criteria as false positives, including CHIP mutations. The analytical specificity was 99.9999%.

[0274] Precision was measured by the variation in variant frequencies estimated between replicate experiments. All samples in the precision study were evaluated for the entire workflow from library construction to sequencing analysis. A total of 18 sets of identical sample aliquots were used for the precision study, and each set included a reference standard and NTC (no-template control) with a defined MAF above the LoD.

[0275] The repeatability test included within-run performance (samples processed under the same conditions). Each sample was run in triplicate under the same conditions, and the results were compared. The reproducibility assessment was based on samples independently processed by at least 2 operators over 5 days. 100% consistency between replicate experiments was detected for both within-run and between-run precision studies.

[0276] Make PredicineHEME TM The PredicineHEME assay analyzed 102 clinical gDNA samples from patients with hematological indications. A total of 714 genomic alterations, including SNVs, indels, and copy number variations, were detected in 98 out of 102 total samples. Copy number variations were detected in 13 genes.

[0277] While the preferred embodiments of the invention have been shown and described herein, it will be apparent to those skilled in the art that these embodiments are provided by way of example only. The invention is not intended to be limited by the specific examples provided in the specification. While the invention has been described with reference to the foregoing specification, the description and illustration of the embodiments herein are not intended to be construed in a limiting sense. Many variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. In addition, it should be understood that all aspects of the invention are not limited to the specific descriptions, configurations, or relative proportions described herein, which depend on various conditions and variables. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. Accordingly, the invention should also cover any such alternatives, modifications, variations, or equivalents. The following claims are intended to define the scope of the invention and thereby cover the methods and structures within the scope of these claims and their equivalents.

Claims

1. A method for detecting the presence or absence of cancer in an object, the method comprising: (a) Assaying cell-free deoxyribonucleic acid (cfDNA) molecules from a biological sample obtained or derived from the object, wherein the biological sample comprises a urine sample; (b) Detecting a biomarker panel from the cfDNA molecules, wherein the biomarker panel comprises differentially expressed markers or variants; (c) Computer processing the biomarker panel to detect the presence or absence of the cancer in the object.

2. The method according to claim 1, wherein the biological sample is obtained or derived from the object using an ethylenediaminetetraacetic acid (EDTA) collection tube, a cell-free RNA collection tube, a cell-free deoxyribonucleic acid (DNA) collection tube, other blood collection tubes, and a CTC collection tube.

3. The method according to any one of claims 1-2, wherein (a) comprises subjecting the biological sample to conditions sufficient to separate, enrich, or extract the cfDNA molecules.

4. The method according to any one of claims 1-3, wherein at least one of the cfDNA molecules is assayed using nucleic acid sequencing to generate nucleic acid sequencing reads.

5. The method according to claim 4, further comprising filtering at least one subset of the nucleic acid sequencing reads based on a quality score.

6. The method according to claim 4 or 5, further comprising error correction of the nucleic acid sequencing reads using a sample barcode or a molecular barcode attached to at least one of the cfDNA molecules.

7. The method according to any one of claims 4-6, further comprising performing at least one of single-stranded consensus sequence identification and double-stranded consensus sequence identification on the nucleic acid sequencing reads, thereby suppressing sequencing and PCR errors in the nucleic acid sequencing reads.

8. The method according to any one of claims 4-7, wherein the cfDNA molecules are assayed using DNA sequencing.

9. The method according to claim 8, wherein the DNA sequencing is selected from: next-generation sequencing, whole-genome sequencing, low-throughput sequencing, targeted sequencing, whole-exome sequencing, methylation-sensitive sequencing, bisulfite sequencing, and combinations thereof.

10. The method according to claim 8, wherein the DNA sequencing comprises targeted sequencing.

11. The method according to claim 4, wherein the nucleic acid sequencing comprises nucleic acid amplification.

12. The method according to claim 11, wherein the nucleic acid amplification comprises polymerase chain reaction (PCR) or isothermal amplification.

13. The method according to any one of claims 1-12, wherein at least one of the cfDNA molecules is assayed by polymerase chain reaction (PCR), microarray, or isothermal amplification.

14. The method according to any one of claims 1-13, wherein the cancer is selected from: breast cancer, lung cancer, prostate cancer, colorectal cancer, melanoma, bladder cancer, non-Hodgkin lymphoma, kidney cancer, endometrial cancer, leukemia, pancreatic cancer, thyroid cancer, and liver cancer, and any combination thereof.

15. The method according to claim 14, wherein the cancer includes renal cancer.

16. The method according to claim 14, wherein the cancer includes bladder cancer.

17. The method according to claim 16, wherein the bladder cancer includes non-muscle invasive bladder cancer.

18. The method according to any one of claims 1-17, wherein the subject is asymptomatic for the cancer.

19. The method according to any one of claims 1-18, wherein (b) includes detecting the presence or absence of the cancer in the subject with an accuracy of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.

20. The method according to any one of claims 1-19, wherein (b) includes detecting the presence or absence of the cancer in the subject with a sensitivity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.

21. The method according to any one of claims 1-20, wherein (b) includes detecting the presence or absence of the cancer in the subject with a specificity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.

22. The method according to any one of claims 1-21, wherein (b) includes detecting the presence or absence of the cancer in the subject with a positive predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.

23. The method according to any one of claims 1-22, wherein (b) includes detecting the presence or absence of the cancer in the subject with a negative predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.

24. The method according to any one of claims 1-23, wherein the biological sample is obtained or derived from the subject before the subject receives therapy for the cancer.

25. The method according to any one of claims 1-23, wherein the biological sample is obtained or derived from the subject during therapy for the cancer.

26. The method according to any one of claims 1-23, wherein the biological sample is obtained or derived from the subject after receiving therapy for the cancer.

27. The method according to any one of claims 24-26, wherein the therapy is selected from: surgical resection, chemotherapy, radiotherapy, immunotherapy, cell therapy, adjuvant therapy, neoadjuvant therapy, androgen deprivation therapy, and combinations thereof.

28. The method according to any one of claims 1-27 further comprises identifying a clinical intervention for the subject, at least in part based on the detected presence or absence of the cancer.

29. The method according to claim 28, wherein the clinical intervention is selected from a plurality of clinical interventions.

30. The method according to claim 28, wherein the clinical intervention is selected from: surgical resection, chemotherapy, radiotherapy, immunotherapy, adjuvant therapy, neoadjuvant therapy, androgen deprivation therapy, and combinations thereof.

31. The method according to claim 28 further comprises administering the clinical intervention to the subject.

32. The method according to any one of claims 1-31, wherein the biomarker panel comprises quantitative measurements of a cancer-associated genomic locus panel.

33. The method according to claim 32, wherein the cancer-associated genomic locus panel comprises one or more members selected from the genes listed in Table 1.

34. The method according to claim 33, wherein the cancer-associated genomic locus panel comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 members selected from the genes listed in Table 1.

35. The method according to claim 32, wherein the cancer-associated genomic locus panel comprises PTEN, TP53, or RB1.

36. The method according to claim 32, wherein the cancer-associated genomic locus panel comprises PTEN.

37. The method according to claim 32, wherein the cancer-associated genomic locus panel comprises FGFR3 or ERBB2.

38. The method according to claim 32, wherein the cancer-associated genomic locus panel comprises one or more members selected from the genes listed in Table 2.

39. The method according to claim 38, wherein the cancer-associated genomic locus panel comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 members selected from the genes listed in Table 2.

40. The method according to any one of claims 32-39 further comprises using a probe configured to selectively enrich nucleic acid molecules corresponding to the genomic locus panel in the biological sample.

41. The method according to claim 40, wherein the probe comprises nucleic acid primers.

42. The method according to claim 40, wherein the probe comprises nucleic acid capture probes.

43. The method according to any one of claims 40 - 42, wherein the probe has sequence complementarity with at least a portion of the nucleic acid sequence of the genomic locus group.

44. The method according to any one of claims 40 - 42, wherein the probe has sequence complementarity with at least a portion of the nucleic acid sequence of a gene selected from the genes listed in Tables 1 and 2.

45. The method according to any one of claims 40 - 43, wherein the probe comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175 or 180 different probes.

46. The method according to any one of claims 1 - 45, further comprising determining the likelihood of the determination of the presence or absence of the cancer in the subject.

47. The method according to any one of claims 1 - 46, further comprising monitoring the presence or absence of the cancer in the subject, wherein the monitoring comprises assessing the presence or absence of the cancer in the subject at each of a plurality of time points.

48. The method according to claim 47, wherein a difference in the assessment of the presence or absence of the cancer in the subject between the plurality of time points indicates one or more clinical indicators selected from: (i) diagnosis of the cancer, (ii) prognosis of the cancer, and (iii) effectiveness or ineffectiveness of a treatment course for treating the cancer in the subject.

49. The method according to claim 48, wherein the prognosis comprises progression - free survival (PFS) or overall survival (OS).

50. The method according to any one of claims 1 - 49, wherein the biomarker group from the cfDNA molecules comprises tumor - associated alterations selected from: copy number alterations (CNA), copy number losses (CNL), single - nucleotide variants (SNV), insertions or deletions (indels), and rearrangements.

51. The method according to any one of claims 1 - 50, wherein the biomarker group from the cfDNA molecules comprises copy number variations.

52. The method according to any one of claims 1 - 50, wherein the biomarker group from the cfDNA molecules comprises copy number losses.

53. The method according to any one of claims 1 - 50, wherein the biomarker group from the cfDNA molecules comprises single - nucleotide variants.

54. The method according to any one of claims 1 - 53, further comprising determining the mutant allele frequency of the somatic mutation group in the biomarker group.

55. The method according to any one of claims 1-54 further comprises determining a blood copy number burden based on a copy number alteration or copy number loss of the biomarker panel.

56. The method according to claim 54 further comprises determining a circulating tumor DNA (ctDNA) fraction of the cancer of the subject at least in part based on the mutant allele frequency panel.

57. The method according to any one of claims 54-56 further comprises determining a tumor mutation burden (TMB) of the cancer of the subject at least in part based on the mutant allele frequency panel.

58. The method according to any one of claims 54-57 further comprises determining a tumor mutation burden (TMB) of the cancer of the subject at least in part based on the mutant allele frequency panel including microsatellites.

59. The method according to any one of claims 54-58 further comprises determining an aberration score of the cancer of the subject at least in part based on the mutant allele frequency panel.

60. A method for detecting the presence or absence of cancer in a subject, comprising: (a) providing cell-free deoxyribonucleic acid (cfDNA) molecules from a biological sample obtained or derived from the subject, wherein the biological sample comprises a urine sample; (b) hybridizing the cfDNA molecules or derivatives thereof with a plurality of nucleic acid capture probes to produce a plurality of enriched cfDNA molecules; (c) sequencing the nucleic acids of the plurality of enriched cfDNA molecules to produce sequencing data; (d) computationally processing the sequencing data to detect a biomarker panel; and (e) detecting the presence or absence of the cancer in the subject based at least on the presence of the biomarker panel.

61. A method for detecting the presence or absence of cancer in a subject, comprising: (a) providing cell-free deoxyribonucleic acid (cfDNA) molecules from a biological sample obtained or derived from the subject; (b) performing a first sequencing assay on the cfDNA or derivatives thereof to produce copy number data for at least one region of the subject's genome, wherein the sequencing is performed at a depth of no more than 10x; (c) performing a second sequencing assay on the cfDNA molecules or derivatives thereof, wherein the second sequencing assay comprises a whole exome sequencing assay or a methylation-sensitive sequencing assay, to produce sequencing data; (d) computationally processing the copy number data and the sequencing data to detect a biomarker panel; and (e) detecting the presence or absence of the cancer in the subject based at least on the presence of the biomarker panel.

62. The method according to claim 61, wherein the biological sample is obtained or derived from the subject using an ethylenediaminetetraacetic acid (EDTA) collection tube, a cell-free RNA collection tube, or a cell-free deoxyribonucleic acid (DNA) collection tube, other blood collection tubes, and CTC collection tubes.

63. The method according to any one of claims 61-62, wherein (a) comprises subjecting the biological sample to conditions sufficient to separate, enrich, or extract the cfDNA molecules.

64. The method according to any one of claims 61-63, wherein the biological sample comprises a urine, blood, or cerebrospinal sample.

65. The method according to claim 64, further comprising filtering at least one subset of the nucleic acid sequencing reads based on a quality score.

66. The method according to claim 64 or 65, further comprising error correction of the nucleic acid sequencing reads using a sample barcode or a molecular barcode attached to at least one of the cfDNA molecules.

67. The method according to any one of claims 64-66, further comprising performing at least one of single-stranded consensus sequence identification and double-stranded consensus sequence identification on the nucleic acid sequencing reads, thereby suppressing sequencing and PCR errors in the nucleic acid sequencing reads.

68. The method according to any one of claims 61-67, wherein (b) or (c) comprises nucleic acid amplification.

69. The method according to claim 68, wherein the nucleic acid amplification comprises polymerase chain reaction (PCR) or isothermal amplification.

70. The method according to any one of claims 61-69, wherein at least one of the cfDNA molecules is assayed by polymerase chain reaction (PCR), microarray, or isothermal amplification.

71. The method according to any one of claims 61-70, wherein the cancer is selected from: breast cancer, lung cancer, prostate cancer, colorectal cancer, melanoma, bladder cancer, non-Hodgkin lymphoma, kidney cancer, endometrial cancer, leukemia, pancreatic cancer, thyroid cancer, and liver cancer, and any combination thereof.

72. The method according to claim 71, wherein the cancer comprises bladder cancer.

73. The method according to claim 72, wherein the bladder cancer comprises non-muscle invasive bladder cancer.

74. The method according to any one of claims 61-73, wherein the subject is asymptomatic for the cancer.

75. The method according to any one of claims 61-74, wherein (d) comprises detecting the presence or absence of the cancer in the subject with an accuracy of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.

76. The method according to any one of claims 61-75, wherein (d) comprises detecting the presence or absence of the cancer in the subject with a sensitivity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.

77. The method according to any one of claims 61 - 76, wherein (d) comprises detecting the presence or absence of said cancer in said subject with a specificity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98% or at least about 99%.

78. The method according to any one of claims 61 - 77, wherein (d) comprises detecting the presence or absence of said cancer in said subject with a positive predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98% or at least about 99%.

79. The method according to any one of claims 61 - 78, wherein (d) comprises detecting the presence or absence of said cancer in said subject with a negative predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98% or at least about 99%.

80. The method according to any one of claims 61 - 79, wherein said biological sample is obtained or derived from said subject prior to said subject receiving therapy for said cancer.

81. The method according to any one of claims 61 - 79, wherein said biological sample is obtained or derived from said subject during therapy for said cancer.

82. The method according to any one of claims 61 - 79, wherein said biological sample is obtained or derived from said subject after receiving therapy for said cancer.

83. The method according to any one of claims 80 - 82, wherein said treatment is selected from: surgical resection, chemotherapy, radiotherapy, immunotherapy, cell therapy, adjuvant therapy, neoadjuvant therapy, androgen deprivation therapy, and combinations thereof.

84. The method according to any one of claims 61 - 83, further comprising identifying a clinical intervention for said subject based at least in part on the detected presence or absence of said cancer.

85. The method according to claim 84, wherein said clinical intervention is selected from a plurality of clinical interventions.

86. The method according to claim 84, wherein said clinical intervention is selected from: surgical resection, chemotherapy, radiotherapy, immunotherapy, adjuvant therapy, neoadjuvant therapy, androgen deprivation therapy, and combinations thereof.

87. The method according to claim 84, further comprising administering said clinical intervention to said subject.

88. The method according to any one of claims 61 - 87, wherein said biomarker panel comprises quantitative measurements of a cancer - associated genomic locus panel.

89. The method according to claim 88, wherein said cancer - associated genomic locus panel comprises one or more members selected from the genes listed in Table 1.

90. The method according to claim 89, wherein the cancer-related genomic locus group comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175 or 180 members selected from the genes listed in Table 1.

91. The method according to claim 88, wherein the cancer-related genomic locus group comprises PTEN, TP53 or RB1.

92. The method according to claim 88, wherein the cancer-related genomic locus group comprises PTEN.

93. The method according to claim 88, wherein the cancer-related genomic locus group comprises FGFR3 or ERBB2.

94. The method according to claim 88, wherein the cancer-related genomic locus group comprises one or more members selected from the genes listed in Table 2.

95. The method according to claim 88, wherein the cancer-related genomic locus group comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175 or 180 members selected from the genes listed in Table 2.

96. The method according to claim 88, wherein the cancer-related genomic locus group comprises one or more members selected from the genes listed in Table 3.

97. The method according to claim 88, wherein the cancer-related genomic locus group comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175 or 180 members selected from the genes listed in Table 3.

98. The method according to any one of claims 88-97, further comprising using a probe configured to selectively enrich nucleic acid molecules corresponding to the genomic locus group in the biological sample.

99. The method according to claim 98, wherein the probe comprises nucleic acid primers.

100. The method according to claim 98, wherein the probe comprises nucleic acid capture probes.

101. The method according to any one of claims 98-100, wherein the probe has sequence complementarity with at least a portion of the nucleic acid sequence of the genomic locus group.

102. The method according to any one of claims 98 - 100, wherein the probe has sequence complementarity with at least a portion of the nucleic acid sequence of a gene selected from the genes listed in Table 1, Table 2, or Table 3.

103. The method according to any one of claims 98 - 102, wherein the probe comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 different probes.

104. The method according to any one of claims 61 - 103, further comprising determining the likelihood of the determination of the presence or absence of the cancer in the subject.

105. The method according to any one of claims 61 - 104, further comprising monitoring the presence or absence of the cancer in the subject, wherein the monitoring comprises evaluating the presence or absence of the cancer in the subject at each of a plurality of time points.

106. The method according to claim 105, wherein a difference in the evaluation of the presence or absence of the cancer in the subject between the plurality of time points indicates one or more clinical indicators selected from: (i) diagnosis of the cancer, (ii) prognosis of the cancer, and (iii) effectiveness or ineffectiveness of a treatment course for treating the cancer in the subject.

107. The method according to claim 106, wherein the prognosis comprises progression - free survival (PFS) or overall survival (OS).

108. The method according to any one of claims 61 - 107, wherein the biomarker set from the cfDNA molecules comprises tumor - associated alterations selected from: copy number alterations (CNA), copy number losses (CNL), single - nucleotide variants (SNV), insertions or deletions (indels), and rearrangements.

109. The method according to any one of claims 61 - 108, wherein the biomarker set from the cfDNA molecules comprises copy number variations.

110. The method according to any one of claims 61 - 108, wherein the biomarker set from the cfDNA molecules comprises copy number losses.

111. The method according to any one of claims 61 - 108, wherein the biomarker set from the cfDNA molecules comprises single - nucleotide variants.

112. The method according to any one of claims 61 - 111, further comprising determining the mutant allele frequency of the somatic mutation set in the biomarker set.

113. The method according to any one of claims 61 - 112, further comprising determining a blood copy number load based on the copy number alteration or copy number loss of the biomarker set.

114. The method according to claim 112, further comprising determining the circulating tumor DNA (ctDNA) fraction of the cancer of the subject at least in part based on the set of mutant allele frequencies.

115. The method according to any one of claims 112-114, further comprising determining the tumor mutation burden (TMB) of the cancer of the subject at least in part based on the set of mutant allele frequencies.

116. The method according to any one of claims 112-115, further comprising determining the tumor mutation burden (TMB) of the cancer of the subject at least in part based on the set of mutant allele frequencies including microsatellites.

117. The method according to any one of claims 112-116, further comprising determining an aberration score of the cancer of the subject at least in part based on the set of mutant allele frequencies.