Systems and methods for multi-analyte detection of cancer - Patents.com

Non-invasive urine-derived cfDNA analysis via NGS for cancer detection addresses the limitations of invasive biopsies by enhancing diagnostic accuracy and treatment guidance for genitourinary cancers.

JP2025535077APending Publication Date: 2025-10-22PREDICINE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025519927
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-02-13
Filing Date
2023-10-04
Publication Date
2025-10-22

AI Technical Summary

Technical Problem

Current cancer detection methods, particularly for genitourinary cancers like bladder cancer, are invasive and may not accurately represent tumor heterogeneity, posing challenges in diagnosis and treatment guidance.

Method used

A non-invasive method using urine-derived cell-free DNA (cfDNA) analysis through next-generation sequencing (NGS) to identify specific biomarkers, such as FGFR3 and TP53, for accurate cancer detection and prognosis, enabling multiple analyte testing and bioinformatics processing to enhance diagnostic sensitivity and treatment recommendations.

Benefits of technology

Provides accurate and non-invasive cancer detection with high sensitivity and specificity, allowing for effective treatment strategies and monitoring, filling the gap left by invasive tissue biopsies and improving patient outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025535077000001_ABST
    Figure 2025535077000001_ABST
Patent Text Reader

Abstract

Methods and systems for detecting cancer are provided herein. The methods may include using nucleic acids from a urine sample. The methods may include assaying the nucleic acids in the urine to detect a set of biomarkers from the sample. The methods may include processing the set of biomarkers to determine the presence of cancer or a cancer parameter. The processing may be performed by an algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] cross reference This application claims the benefit of U.S. Application No. 63 / 413,433, filed October 5, 2022, U.S. Application No. 63 / 445,201, filed February 13, 2023, U.S. Application No. 63 / 445,145, filed February 13, 2023, and U.S. Application No. 63 / 445,150, filed February 13, 2023, each of which is incorporated by reference in its entirety. [Background technology]

[0002] Cancer is a leading cause of death worldwide. Detecting cancer in an individual can be important in providing treatment and improving patient outcomes. Cancer can be caused by genetic abnormalities, which can result in uncontrolled cell growth. Detecting genetic abnormalities can be important in detecting cancer. Sequencing of nucleic acids in a patient-derived sample can be used to detect genetic abnormalities. Summary of the Invention

[0003] Systems and methods are provided herein for detecting the presence or absence of cancer in a subject. The systems and methods provided herein involve assaying polynucleotides to identify biomarkers for the subject's cancer. Detecting specific biomarkers for a cancer type or a given cancer can provide effective treatment to the individual and improve outcomes. In the case of multiple types of cancer, specific biomarkers indicative of a specific cancer type (or subtype) can be used to identify the prognosis of an individual suffering from cancer. Multiple analytes can be tested to provide accurate cancer detection and prognosis. Increasing the number of analytes (and sets of biomarkers from the analytes) analyzed can improve cancer (or cancer parameters) detection, allow for effective treatment recommendations, and more accurate prognosis.

[0004] In one aspect, the present disclosure provides a method for detecting the presence or absence of cancer in a subject, the method comprising: (a) assaying cell-free deoxyribonucleic acid (cfDNA) molecules from a biological sample obtained from or derived from the subject, wherein the biological sample comprises a urine sample; (b) detecting a set of biomarkers from the cfDNA molecules, wherein the set of biomarkers comprises differentially expressed markers or variants; and (c) computer-processing the set of biomarkers to detect the presence or absence of cancer in the subject. The method of claim 1, wherein the biological sample is obtained from or derived from the subject using an ethylenediaminetetraacetic acid (EDTA) collection tube, a cell-free RNA collection tube, or a cell-free deoxyribonucleic acid (DNA), other blood collection tube, and CTC collection tube.

[0005] In some embodiments, (a) comprises subjecting the biological sample to conditions sufficient to isolate, enrich, or extract cfDNA molecules. In some embodiments, at least one of the cfDNA molecules is assayed using nucleic acid sequencing to generate nucleic acid sequencing reads. In some embodiments, the method further comprises filtering at least a subset of the nucleic acid sequencing reads based on quality scores. In some embodiments, the method further comprises performing error correction on the nucleic acid sequencing reads using a sample barcode or molecular barcode attached to at least one of the cfDNA molecules. In some embodiments, the method further comprises performing at least one of single-strand consensus calling and double-strand consensus calling on the nucleic acid sequencing reads, thereby suppressing sequencing and PCR errors in the nucleic acid sequencing reads.

[0006] In some embodiments, the cfDNA molecules are assayed using DNA sequencing. In some embodiments, the DNA sequencing is selected from the group consisting of next-generation sequencing, whole genome sequencing, low-pass sequencing, targeted sequencing, whole exome sequencing, methylation-recognition sequencing, bisulfite sequencing, and combinations thereof. In some embodiments, the DNA sequencing comprises targeted sequencing. In some embodiments, the nucleic acid sequencing comprises nucleic acid amplification. In some embodiments, the nucleic acid amplification comprises polymerase chain reaction (PCR) or isothermal amplification. In some embodiments, at least one of the cfDNA molecules is assayed using a polymerase chain reaction (PCR) assay, a microarray, or isothermal amplification.

[0007] In some embodiments, the cancer is selected from the group consisting of breast cancer, lung cancer, prostate cancer, colon cancer, melanoma, bladder cancer, non-Hodgkin's lymphoma, kidney cancer, endometrial cancer, leukemia, pancreatic cancer, thyroid cancer, and liver cancer, and any combination thereof. In some embodiments, the cancer comprises bladder cancer. In some embodiments, the cancer comprises kidney cancer. In some embodiments, the bladder cancer comprises non-muscle-invasive bladder cancer. In some embodiments, the subject is asymptomatic for the cancer.

[0008] In some embodiments, (b) comprises detecting the presence or absence of cancer in the subject with an accuracy of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. In some embodiments, (b) comprises detecting the presence or absence of cancer in the subject with a sensitivity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. In some embodiments, (b) comprises detecting the presence or absence of cancer in the subject with a specificity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. In some embodiments, (b) comprises detecting the presence or absence of cancer in the subject with a positive predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. In some embodiments, (b) comprises detecting the presence or absence of cancer in the subject with a negative predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.

[0009] In some embodiments, the biological sample is obtained from or derived from a subject before the subject receives treatment for cancer. In some embodiments, the biological sample is obtained from or derived from a subject during treatment for cancer. In some embodiments, the biological sample is obtained from or derived from a subject after receiving treatment for cancer. In some embodiments, the treatment is selected from the group consisting of surgical resection, chemotherapy, radiation therapy, immunotherapy, cell therapy, adjuvant therapy, neoadjuvant therapy, androgen deprivation therapy, and combinations thereof. In some embodiments, the method further comprises identifying a clinical intervention for the subject based at least in part on the presence or absence of detected cancer. In some embodiments, the clinical intervention is selected from a plurality of clinical interventions. In some embodiments, the clinical intervention is selected from the group consisting of surgical resection, chemotherapy, radiation therapy, immunotherapy, adjuvant therapy, neoadjuvant therapy, androgen deprivation therapy, and combinations thereof. In some embodiments, the method further comprises administering the clinical intervention to the subject.

[0010] In some embodiments, the set of biomarkers comprises quantitative measurements of a set of genomic loci associated with cancer. In some embodiments, the set of genomic loci associated with cancer comprises one or more members selected from the group consisting of the genes listed in Table 1. In some embodiments, the set of genomic loci associated with cancer comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 members selected from the group consisting of the genes listed in Table 1. In some embodiments, the set of genomic loci associated with cancer comprises PTEN, TP53, or RB1. In some embodiments, the set of genomic loci associated with cancer comprises PTEN. In some embodiments, the set of genomic loci associated with cancer comprises FGFR3 or ERBB2. In some embodiments, the set of genomic loci associated with cancer comprises one or more members selected from the group consisting of the genes listed in Table 2. In some embodiments, the set of genomic loci associated with cancer comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 members selected from the group consisting of the genes listed in Table 2.

[0011] In some embodiments, the method further includes using a probe configured to selectively enrich the biological sample for nucleic acid molecules corresponding to the set of loci. In some embodiments, the probe comprises a nucleic acid primer. In some embodiments, the probe comprises a nucleic acid capture probe. In some embodiments, the probe has sequence complementarity to at least a portion of the nucleic acid sequence of the set of genomic loci. In some embodiments, the probe comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 different probes.

[0012] In some embodiments, the method further comprises determining the likelihood of determining the presence or absence of cancer in the subject. In some embodiments, the method further comprises monitoring the presence or absence of cancer in the subject, comprising assessing the presence or absence of cancer in the subject at each of a plurality of time points. In some embodiments, the difference in the assessment of the presence or absence of cancer in the subject between the plurality of time points indicates one or more clinical indications selected from the group consisting of: (i) diagnosis of cancer, (ii) prognosis of cancer, and (iii) the effectiveness or ineffectiveness of a treatment course for treating the subject's cancer.

[0013] In some embodiments, the prognosis includes a prediction of expected progression-free survival (PFS) or overall survival (OS). In some embodiments, the set of biomarkers from cfDNA molecules includes tumor-associated alterations selected from the group consisting of copy number alterations (CNAs), copy number losses (CNLs), single nucleotide variants (SNVs), insertions or deletions (indels), and rearrangements. In some embodiments, the set of biomarkers from cfDNA molecules includes copy number variations. In some embodiments, the set of biomarkers from cfDNA molecules includes copy number losses. In some embodiments, the set of biomarkers from cfDNA molecules includes single nucleotide variants.

[0014] In some embodiments, the method further comprises determining mutant allele frequencies of a set of somatic mutations among the set of biomarkers. In some embodiments, the method further comprises determining blood copy number burden based on copy number alterations or copy number losses of the set of biomarkers. In some embodiments, the method further comprises determining a circulating tumor DNA (ctDNA) fraction of the subject's cancer based at least in part on the set of mutant allele frequencies. In some embodiments, the method further comprises determining a tumor mutation burden (TMB) of the subject's cancer based at least in part on the set of mutant allele frequencies. In some embodiments, the method further comprises determining a plasma tumor mutation burden (TMB) of the subject's cancer based at least in part on the set of mutant allele frequencies that include microsatellites. In some embodiments, the method further comprises determining an aberration score for the subject's cancer based at least in part on the set of mutant allele frequencies.

[0015] In one aspect, the present disclosure provides a method for detecting the presence or absence of cancer in a subject, the method comprising: (a) providing cell-free deoxyribonucleic acid (cfDNA) molecules from a biological sample obtained from or derived from the subject, wherein the biological sample comprises a urine sample; (b) hybridizing the cfDNA molecules or derivatives thereof to a plurality of nucleic acid capture probes to generate a plurality of enriched cfDNA molecules; (c) sequencing the nucleic acids of the plurality of enriched cfDNA molecules to generate sequencing data; (d) computationally processing the sequencing data to detect a set of biomarkers; and (e) detecting the presence or absence of cancer in the subject based at least on the presence of the set of biomarkers.

[0016] In one aspect, the present disclosure provides a method for detecting the presence or absence of cancer in a subject, the method comprising: (a) providing cell-free deoxyribonucleic acid (cfDNA) from a biological sample obtained from or derived from the subject; (b) performing a first sequencing assay on the cfDNA or a derivative thereof to generate copy number data for at least one region of the subject's genome, wherein the sequencing is performed at a depth of 10x or less; (c) performing a second sequencing assay on the cfDNA molecule or a derivative thereof to generate sequencing data, wherein the second sequencing assay comprises a whole-exome sequencing assay or a methylation-aware sequencing assay; (d) computationally processing the copy number data and the sequencing data to detect a set of biomarkers; and (e) detecting the presence or absence of cancer in the subject based at least on the presence of the set of biomarkers.

[0017] In some embodiments, the biological sample is obtained from or derived from a subject using ethylenediaminetetraacetic acid (EDTA) collection tube, cell-free RNA collection tube, or cell-free deoxyribonucleic acid (DNA) collection tube, other blood collection tube, and CTC collection tube.In some embodiments, (a) comprises subjecting the biological sample to conditions sufficient to isolate, enrich, or extract cfDNA molecules.

[0018] In some embodiments, the biological sample comprises a urine, blood, or cerebrospinal fluid sample. In some embodiments, the method further comprises filtering at least a subset of the nucleic acid sequencing reads based on a quality score. In some embodiments, the method further comprises performing error correction on the nucleic acid sequencing reads using a sample barcode or molecular barcode attached to at least one of the cfDNA molecules. In some embodiments, the method further comprises performing at least one of single-strand consensus calling and double-strand consensus calling on the nucleic acid sequencing reads, thereby suppressing sequencing and PCR errors in the nucleic acid sequencing reads. In some embodiments, (b) or (c) comprises nucleic acid amplification. In some embodiments, the nucleic acid amplification comprises polymerase chain reaction (PCR) or isothermal amplification. In some embodiments, at least one of the cfDNA molecules is assayed using a polymerase chain reaction (PCR) assay, a microarray, or isothermal amplification. In some embodiments, the cancer is selected from the group consisting of breast cancer, lung cancer, prostate cancer, colon cancer, melanoma, bladder cancer, non-Hodgkin's lymphoma, kidney cancer, endometrial cancer, leukemia, pancreatic cancer, thyroid cancer, and liver cancer, and any combination thereof. In some embodiments, the cancer comprises bladder cancer. In some embodiments, the bladder cancer comprises non-muscle-invasive bladder cancer. In some embodiments, the subject is asymptomatic for the cancer.

[0019] In some embodiments, (d) comprises detecting the presence or absence of cancer in the subject with an accuracy of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. In some embodiments, (d) comprises detecting the presence or absence of cancer in the subject with a sensitivity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. In some embodiments, (d) comprises detecting the presence or absence of cancer in the subject with a specificity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. In some embodiments, (d) comprises detecting the presence or absence of cancer in the subject with a positive predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. In some embodiments, (d) comprises detecting the presence or absence of cancer in the subject with a negative predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.

[0020] In some embodiments, the biological sample is obtained from or derived from a subject before the subject receives treatment for cancer. In some embodiments, the biological sample is obtained from or derived from a subject during treatment for cancer. In some embodiments, the biological sample is obtained from or derived from a subject after receiving treatment for cancer. In some embodiments, the treatment is selected from the group consisting of surgical resection, chemotherapy, radiation therapy, immunotherapy, cellular therapy, adjuvant therapy, neoadjuvant therapy, androgen deprivation therapy, and combinations thereof.

[0021] In some embodiments, the method further comprises identifying a clinical intervention for the subject based at least in part on the presence or absence of detected cancer. In some embodiments, the clinical intervention is selected from a plurality of clinical interventions. In some embodiments, the clinical intervention is selected from the group consisting of surgical resection, chemotherapy, radiation therapy, immunotherapy, adjuvant therapy, neoadjuvant therapy, androgen deprivation therapy, and combinations thereof. In some embodiments, the method further comprises administering the clinical intervention to the subject. In some embodiments, the set of biomarkers comprises quantitative measurements of a set of genomic loci associated with cancer. In some embodiments, the set of genomic loci associated with cancer comprises one or more members selected from the group consisting of the genes listed in Table 1. In some embodiments, the set of genomic loci associated with cancer comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 members selected from the group consisting of genes listed in Table 1. In some embodiments, the set of genomic loci associated with cancer comprises PTEN, TP53, or RB1. In some embodiments, the set of genomic loci associated with cancer comprises PTEN. In some embodiments, the set of genomic loci associated with cancer comprises FGFR3 or ERBB2. In some embodiments, the set of genomic loci associated with cancer comprises one or more members selected from the group consisting of the genes listed in Table 2. In some embodiments, the set of genomic loci associated with cancer comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 members selected from the group consisting of the genes listed in Table 2.In some embodiments, the set of genomic loci associated with cancer comprises one or more members selected from the group consisting of the genes listed in Table 3. In some embodiments, the set of genomic loci associated with cancer comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 members selected from the group consisting of the genes listed in Table 3. In some embodiments, the method further comprises using a probe configured to selectively enrich the biological sample for nucleic acid molecules corresponding to the set of loci. In some embodiments, the probe comprises a nucleic acid primer. In some embodiments, the probe comprises a nucleic acid capture probe. In some embodiments, the probes have sequence complementarity to at least a portion of the nucleic acid sequence of a set of genomic loci. In some embodiments, the probes have sequence complementarity to at least a portion of the nucleic acid sequence of a gene selected from the genes listed in Table 1, Table 2, or Table 3. In some embodiments, the probes comprise at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 different probes. In some embodiments, the method further comprises determining the likelihood of determining the presence or absence of cancer in the subject.

[0022] In some embodiments, the method further comprises determining the likelihood of determining the presence or absence of cancer in the subject. In some embodiments, the method further comprises monitoring the presence or absence of cancer in the subject, comprising assessing the presence or absence of cancer in the subject at each of a plurality of time points. In some embodiments, the difference in the assessment of the presence or absence of cancer in the subject between the plurality of time points indicates one or more clinical indications selected from the group consisting of: (i) diagnosis of cancer, (ii) prognosis of cancer, and (iii) the effectiveness or ineffectiveness of a treatment course for treating the subject's cancer.

[0023] In some embodiments, the prognosis includes a prediction of expected progression-free survival (PFS) or overall survival (OS). In some embodiments, the set of biomarkers from cfDNA molecules includes tumor-associated alterations selected from the group consisting of copy number alterations (CNAs), copy number losses (CNLs), single nucleotide variants (SNVs), insertions or deletions (indels), and rearrangements. In some embodiments, the set of biomarkers from cfDNA molecules includes copy number variations. In some embodiments, the set of biomarkers from cfDNA molecules includes copy number losses. In some embodiments, the set of biomarkers from cfDNA molecules includes single nucleotide variants.

[0024] In some embodiments, the method further comprises determining mutant allele frequencies of a set of somatic mutations among the set of biomarkers. In some embodiments, the method further comprises determining blood copy number burden based on copy number alterations or copy number losses of the set of biomarkers. In some embodiments, the method further comprises determining a circulating tumor DNA (ctDNA) fraction of the subject's cancer based at least in part on the set of mutant allele frequencies. In some embodiments, the method further comprises determining a tumor mutation burden (TMB) of the subject's cancer based at least in part on the set of mutant allele frequencies. In some embodiments, the method further comprises determining a plasma tumor mutation burden (TMB) of the subject's cancer based at least in part on the set of mutant allele frequencies that include microsatellites. In some embodiments, the method further comprises determining an aberration score for the subject's cancer based at least in part on the set of mutant allele frequencies.

[0025] Another aspect of the present disclosure provides a non-transitory computer-readable medium comprising machine-executable code that, when executed by one or more computer processors, implements any of the methods described above or elsewhere herein.

[0026] Another aspect of the present disclosure provides a system comprising one or more computer processors and a computer memory coupled thereto, the computer memory comprising machine-executable code that, when executed by the one or more computer processors, implements any of the methods described above or elsewhere herein.

[0027]

[0013] Further aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.

[0028] Incorporation by Reference All publications, patents, and patent applications mentioned herein are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. In the event that the publications and patents or patent applications incorporated by reference conflict with the disclosure contained herein, the present specification is intended to supersede and / or supersede any such conflicting material. [Brief explanation of the drawings]

[0029] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also referred to herein as "figure" and "FIG.").

[0030] [Figure 1] An example workflow for the identification of somatic mutations in urinary cell-free DNA is shown. [Figure 2] An example workflow for variant detection is shown. [Figure 3] Graphs of assay sensitivity and mutant allele frequency are shown. [Figure 4] The agreement between the predicted MAF and the MAF detected by PredicineCARE is shown. [Figure 5]Matched tissue and urine for genomic alterations detected by PredicineCARE are shown. [Figure 6] 1 shows a schematic diagram of an example workflow. [Figure 7] 1 shows a schematic diagram of an example workflow. [Figure 8] 1 shows a schematic diagram of an example workflow. [Figure 9] 1 shows a schematic diagram of an example workflow. [Figure 10] 1 shows a schematic diagram of an example bioinformatics pipeline workflow. [Figure 11] A schematic diagram of an example study design is provided. [Figure 12] Heatmaps of matched urine NGS and FFPE tissue RT-PCR are shown. [Figure 13A] Scatter plots for variant allelic variation (VAF) between matched FFPE and urine variants are shown. [Figure 13B] Scatter plots for variant allelic variation (VAF) between matched FFPE and urine variants are shown. [Figure 14A] 1 shows a schematic diagram of an example assay workflow. [Figure 14B] A schematic diagram of an example assay workflow is shown. Figure 14B shows an illustration of genome-wide CNV detection and CNB calculation. [Figure 14C] A schematic diagram of an example assay workflow is shown. Figures 14C-14D show LP-WGS CNV profiles. [Figure 14D] A schematic diagram of an example assay workflow is shown. Figures 14C-14D show LP-WGS CNV profiles. [Figure 15A] 1 shows a chart for analytical evaluation of Predicine CNB on clinical plasma samples. [Figure 15B] 1 shows a chart for analytical evaluation of Predicine CNB on clinical plasma samples. [Figure 15C] 1 shows a chart for analytical evaluation of Predicine CNB on clinical plasma samples. [Figure 15D] 1 shows a chart for analytical evaluation of Predicine CNB on clinical plasma samples. [Figure 16A] LP-WGS CNV profile heatmap of 14 non-muscle invasive bladder cancer patient samples is shown. [Figure 16B] LP-WGS CNV profile heatmap of 33 muscle-invasive bladder cancer patient samples is shown. [Figure 17A] 1 shows a comparison of CNB scores for FFPE, plasma, and urine samples between patients with non-muscle invasive bladder cancer and patients with invasive / non-organ confined bladder cancer. [Figure 17B] LP-WGS CNV profiles of two non-invasive bladder cancer patients before and after TURBT surgery are shown. [Figure 17C] LP-WGS gene copy numbers of major bladder cancer genes before and after TURBT are shown. [Figure 18] 1 shows a schematic diagram of an example workflow. [Figure 19] Shown are PredicineEPIC libraries with 1 ng, 2.5 ng, and 10 ng input compared to the standard 50 ng DNA input used for standard whole genome bisulfite sequencing libraries. [Figure 20] A graph showing uniform manifold approximation and projection (UMAP) is shown. [Figure 21] An example heatmap of significance scores (color scale) of aberrantly methylated fragments is shown. [Figure 22] A chart showing the methylation aberration score is shown. [Figure 23] Mutational profiles of tumor DNA in urine and tissues from MIBC are shown. [Figure 24] Figures 24A-D show charts of the concordance and number of mutations detected in tDNA and utDNA within the WES and ATLAS regions. [Figure 25] Plots of tumor fraction in tissue and urine are shown. [Figure 26A]1 shows a study schematic, sample collection, and assay timeline and operation for an exemplary workflow. [Figure 26B] 1 shows a study schematic, sample collection, and assay timeline and operation for an exemplary workflow. [Figure 26C] 1 shows a study schematic, sample collection, and assay timeline and operation for an exemplary workflow. [Figure 27A] Data from an example Predicine WES+ assay is shown. [Figure 27B] Data from an example Predicine WES+ assay is shown. [Figure 28] Serial bCNB scores during treatment for multiple patients are shown. [Figure 29] Copy number gains and losses determined via PredicineCNB are shown. [Figure 30] The mutation landscape determined via PredicineWES+ is shown. [Figure 31] 1 illustrates a computer system programmed or configured to implement the methods provided herein. DETAILED DESCRIPTION OF THE INVENTION

[0031] While various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It will be understood that various alternatives to the embodiments of the invention described herein may be employed.

[0032] The present invention provides a system and method for detecting the presence or absence of cancer in a subject. The system and method provided herein include assaying polynucleotides to identify biomarkers for cancer in a subject. The biomarkers can be processed to identify the presence or absence of cancer. The methods described herein can also process analytes to determine the presence or absence of cancer. The analytes can include cfDNA or other analytes that can be provided by non-invasive methods. By analyzing analytes obtained by non-invasive methods, the present method allows for improved or similar detection or determination of prognosis compared to methods using tumor or tissue biopsy.

[0033] Human urine may contain fragmented DNA, known as urinary cell-free DNA (ucfDNA), derived from dying cells in the urogenital tract or from circulating DNA that has passed through glomerular filtration. Given direct access to the urinary tract, urine is a viable source for detecting cfDNA biomarkers, and ucfDNA can improve the current diagnostic sensitivity of liquid biopsies in urogenital cancers. Many genetic variants can be identified in ucfDNA from cancer patients, particularly those with bladder cancer. Using urine for liquid biopsies offers a completely noninvasive approach for the detection of genomic biomarkers to guide cancer treatment.

[0034] Next-generation sequencing (NGS) has revolutionized cancer genomics research over the past decade. NGS technologies are commercially available, either FDA-cleared or approved for use in processing DNA obtained from patient tissue or blood samples to guide cancer patient treatment plans. Next-generation sequencing (NGS) assays on cfDNA can enable accurate detection of genomic alterations, including single-nucleotide variants (SNVs), insertions and deletions (indels), copy number variations (CNVs), and DNA rearrangements. Because urinary cfDNA is derived directly from dying cells exfoliated in urine, whereas tissue biopsies only account for mutations found in specific regions of the tumor, it can be considered more representative of the tumor than tissue biopsies due to tumor heterogeneity. In addition, urine may contain fewer contaminating proteins than blood, and cfDNA levels in urine may be higher than those in the bloodstream. Urinary cfDNA can be subjected to sequencing, enabling the detection of cancer and cancer-associated genetic alterations in a subject's urine. The methods and assays described in this disclosure represent an application of NGS technology to provide a non-invasive, cost-effective, and potentially more sensitive method of sample collection for patients with cancers such as genitourinary cancers, including bladder cancer.

[0035] In addition to staging and grading a patient's tumor, tissue biopsy often represents the gold standard in guiding cancer patients' treatment, including eligibility for clinical trials or FDA-approved treatment options (e.g., QIAGEN Therascreen FGFR RGQ RT-PCR Kit). However, depending on the tumor location or patient condition, tumor biopsy can be painful and expose patients to potential complications, which can increase the cost of medical procedures. In some cases, tissue biopsy may not be feasible. Minimally invasive sample collection methods remain an unmet clinical need for the treatment of bladder cancer patients. Urine cfDNA assays and liquid biopsy options (e.g., from urine) may help fill this void. Furthermore, in contrast to tissue biopsies, there is the opportunity for patients to be tested multiple times with urine biopsies. Therefore, urine liquid biopsies represent a noninvasive and cost-effective method for obtaining patient samples to determine molecular eligibility for specific treatment strategies.

[0036] Urine biopsy samples can also improve patient care. Bladder cancer patients can be monitored using urine biopsies when tissue biopsies are not feasible for NGS testing. The assay can identify FGFR molecular eligibility for specific therapeutic products. Additionally, the assay can detect other genetic mutations in urine samples from bladder cancer patients, including, but not limited to, alterations in CDKN2A, HRAS / KRAS, KDM6A, PIK3CA, TERT, TP53, and TSC1, which, if identified, can help inform patient treatment.

[0037] Bladder cancer is the 10th most common malignancy worldwide, with an estimated 550,000 new cases and 200,000 deaths reported in 2018. The majority of bladder cancer cases are non-muscle-invasive bladder cancer (NMIBC), which requires frequent monitoring, local resection (transurethral resection of bladder tumor [TURBT]), and an intensive regimen of intravesical therapy to reduce the risk of both recurrent and progressive disease. Despite these efforts, within 5 years, 45%–60% of patients experience recurrent disease, and nearly 20% of patients with high-risk disease progress to muscle-invasive tumors requiring radical cystectomy (RC). The natural history of high-risk NMIBC is unpredictable, with recurrence rates ranging from 15%–78% and progression to muscle invasion and metastasis ranging from 1%–45%. Long-term outcomes suggest that approximately 20%–25% of high-risk NMIBC patients ultimately die from bladder cancer.

[0038] Analytes that can be used for tumor diagnosis from urine biopsies include cfDNA, non-coding RNA, exfoliated tumor cells, and proteins. During tumor-destructive therapy or during apoptosis and necrosis processes, both healthy and diseased cells can release cfDNA fragments, typically 100–200 base pairs in length. In disease-free patients, phagocytes engulf cellular debris and necrotic cells, resulting in very low levels of cfDNA. In diseased patients, phagocytosis is impaired, DNA digestion is minimal, and DNA fragments have random dimensions that can exceed 10,000 base pairs. Therefore, cfDNA levels in diseased patients are often elevated.

[0039] To identify a panel of DNA markers using an NGS assay, we conducted a prospective registry study of blood, urine, and tumor tissue samples from 16 bladder cancer patients with hematuria. This study demonstrated that gene mutations in urine supernatant and sediment had better concordance with cancer tissue compared to plasma. Using a 48-gene panel, analysis suggested that two combinations of genes for genetic diagnostic modeling of DNA from bladder cancer patients were TERT, FGFR3, TP53, PIK3CA, and KRAS for urine supernatant, and TERT, FGFR3, TP53, HRAS, PIK3CA, KRAS, and ERBB2 for urinary supernatant. The accuracy of the 5-gene and 7-gene panels yielded AUCs of 0.94 (95% CI 0.91-0.97) and 0.91 (95% CI 0.86-0.96), respectively. This study concluded that urine-derived cfDNA has great diagnostic potential for identifying bladder cancer in patients with hematuria.

[0040] Fibroblast growth factor receptors FGFR1-FGFR4 are tyrosine kinases present in many types of endothelial cells and tumor cells and have been shown to play important roles in tumor cell growth, survival, and migration as well as in maintaining tumor angiogenesis (Turner 2010). Various mechanisms exist for FGFR-related tumorigenesis, including gene amplification, mutation, and fusion. FGFR3 gene alterations are found in 60%-70% of early-stage NMIBC cases. The observed frequency of FGFR3 alterations in bladder cancer varies with tumor stage and grade. For example, in a prospective cohort of 772 patients with NMIBC, TaG1 (158 / 257) and TaG2 (139 / 239) tumors showed similar mutation frequencies of 61.5% and 58.1%, respectively. FGFR3 mutations were an independent predictor of recurrence in patients with low-grade Ta tumors. Mutation frequencies were low among TaG3 (30 / 88; 34.1%), T1G2 (7 / 26; 26.9%), and T1G3 tumors (20 / 119; 17%).

[0041] A more recent study using a pooled dataset of matched clinical and genomic data for 263 patients with stage pT1 disease demonstrated that FGFR alterations were frequent in high-risk patients (39% mutations, 6% fusions, not mutually exclusive). Furthermore, this study demonstrated that, unlike previous reports, the prognosis of patients with FGFR alterations did not differ from that of patients without FGFR alterations (Breyer 2020). The prevalence of FGFR alterations appears to be somewhat lower in MIBC compared with NMIBC, with an FGFR3 mutation rate of approximately 15% reported across several studies.

[0042] Urine cfDNA (ucfDNA) can be extracted from the urine of a subject, and the ucfDNA can be subjected to various reactions to allow sequencing of the ucfDNA.The construction of a ucfDNA library can include amplification, adapter ligation, or labeling with barcodes to generate additional sequences and / or sequencing libraries.Furthermore, the cfDNA or ucfDNA library can be subjected to enrichment using capture probes or amplification primers to enrich specific sequences of interest from cfDNA.The library can then be subjected to sequencing reactions to generate sequencing data.

[0043] Figure 1 shows the general workflow of an example urine-based assay. A urine sample may be isolated and collected from an individual. After collection, urinary cfDNA can be extracted. Library construction can then be performed on the extracted urinary cfDNA, followed by enrichment for specific targets. After enrichment, the cfDNA can be sequenced, and the sequencing data can then be processed.

[0044] Genetic alterations such as single-nucleotide variations (SNVs), indels, DNA rearrangements, and copy number variations (CNVs) can be identified through bioinformatics analysis of sequencing data. Generally, bioinformatics pipelines utilize raw sequencing data (e.g., BCL files) and can output variant calls. The pipeline can perform various tasks to analyze sequencing data, such as adapter trimming, barcode checking, or error correction. Pairs of clean files (e.g., FASTQ files) can be aligned to the human reference genome using an alignment tool such as the BWA alignment tool. A consensus sequence can then be obtained by merging single-stranded fragments and paired-end reads derived from the same molecule. Single-stranded fragments derived from the same double-stranded DNA molecule can be further merged as double strands. These processes can allow for the correction of sequencing and PCR errors. Figure 2 shows an example workflow for bioinformatics analysis.

[0045] Figures 6-9 show an example process workflow for an exemplary urine cfDNA assay (e.g., PredicineCARE). Blue ellipses provide connectivity between diagrams. Figure 10 shows the workflow of a bioinformatics pipeline (e.g., DeepSEA). As shown in Figure 6, a sample is provided, and then cfDNA can be extracted and isolated from the sample. The cfDNA can then be validated and quality controlled. A sequencing library can be generated from the cfDNA via end-repair and A-tailing adapter ligation. The library can then be amplified and quantified.

[0046] Once this sample library is generated, the library can be subjected to in-solution hybridization with capture probes, as shown in Figure 7. The capture probes can be biotinylated. The probes can then be incubated with magnetic beads (e.g., streptavidin magnetic beads) and then subjected to washing to remove any contaminants. The captured DNA can then be eluted from the beads and subjected to further amplification, normalization, and / or pooling to form a capture library.

[0047] As shown in Figure 8, the captured library can be sequenced using a next-generation sequencer (e.g., Illumina NovaSeq) to generate raw sequencing data.

[0048] Figure 9 shows a schematic of the raw sequencing process. The data is fed into a bioinformatics pipeline (e.g., the DeepSea pipeline) for variant calling. After variant calling and quality checking, the variants can be classified and analyzed to determine the presence of specific variants, which can provide information about the presence of cancer and specific mutations associated with cancer. This can then be generated into a report, which can then be sent to a physician.

[0049] Figure 10 shows a schematic diagram of a variant calling pipeline (e.g., the DeepSea pipeline). The pipeline can demultiplex and extract unique molecular identifiers (UMIs) to generate FASTq files with UMIs. These files can then be aligned, and a consensus sequence can then be generated. Based at least in part on the analysis of the UMIs or consensus sequence, various errors can be corrected. These error-corrected sequences can then be input into various variant callers to generate variant calling results. These files can then be annotated and filtered to generate variant calling results.

[0050] The subject may be a subject suspected of having cancer.Cancer may be specific to or originate from an organ or other region of the subject.For example, cancer may be breast cancer, lung cancer, prostate cancer, colon cancer, melanoma, bladder cancer, non-Hodgkin's lymphoma, kidney cancer, endometrial cancer, leukemia, pancreatic cancer, thyroid cancer, and liver cancer, and any combination thereof.Cancer may be hormone-sensitive prostate cancer (HSPC), castration-resistant prostate cancer (CRPC), metastatic prostate cancer, and combinations thereof.Cancer may include biomarkers specific to specific cancers.Specific biomarkers may indicate the presence of specific cancers.For example, biomarkers may indicate the presence of castration-resistant prostate cancer.Identifying the presence of cancer types may allow for the determination of treatment options or recommendations.

[0051] In some cases, the subject may be asymptomatic about cancer.For example, cancer may not show any symptoms, and the subject may not be aware of the existence of cancer.The method described herein can allow cancer to be identified at an earlier stage than otherwise.Identifying the existence of cancer at an early stage can allow treatment options or recommendations to be determined at an early stage, and can allow the subject to have an improved prognosis.

[0052] The biological sample may contain nucleic acids. The biological sample may be a cell-free deoxyribonucleic acid (cfDNA) sample or a cell-free ribonucleic acid (cfRNA) sample. The biological sample may contain genomic DNA or germline DNA (gDNA). The nucleic acid may be DNA (e.g., double-stranded DNA, single-stranded DNA, single-stranded DNA hairpin, cDNA, genomic DNA, germline DNA, circulating tumor DNA (ctDNA), cell-free DNA (cfDNA)), RNA (e.g., cfRNA, mRNA, cRNA, miRNA, siRNA, miRNA, snoRNA, piRNA, tiRNA, snRNA), or a DNA / RNA hybrid. The biological sample may be derived from or contain a biological fluid. For example, the biological sample may be a plasma sample, a serum sample, a buffy coat sample, a peripheral blood mononuclear cell (PBMC) sample, a red blood cell sample, a urine sample, a saliva sample, or other bodily fluid sample. The biological sample may include or be a pleural fluid sample, an ascites sample, an amniotic fluid sample, a cerebrospinal fluid sample, a lymphatic fluid sample, a sweat sample, a tear sample, a semen sample, or any combination of biological fluids.

[0053] The biological sample can be collected, obtained, or derived from a subject using a collection tube. The collection tube can be an ethylenediaminetetraacetic acid (EDTA) collection tube, a cell-free RNA collection tube, a cell-free deoxyribonucleic acid (DNA) collection tube, a CTC collection tube, or other blood collection tube. The collection tube can contain additional reagents to stabilize nucleic acid molecules or blood cells. The collection tube can allow the nucleic acids or blood cells to stabilize to minimize degradation of the biological sample before assay. The additional reagents can include buffer salts or chelating agents.

[0054] Biological samples may be obtained from or derived from a subject at various time points. Biological samples may be obtained from or derived from a subject before the subject is treated for cancer. Biological samples may be obtained from or derived from a subject while the subject is undergoing treatment. Biological samples may be obtained from or derived from a subject after the subject is treated for cancer. Biological samples may be collected at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or over a period of time. The time points can occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 hours or more. The time points can occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 days or more. The time points can occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 weeks or more. The time points can occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 months or more. The time points can occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 years or more.

[0055] In various embodiments described herein, a clinical intervention or treatment may be identified based at least in part on the identification of the presence of cancer or the presence of a cancer parameter. The clinical intervention may be multiple clinical interventions. The clinical intervention may be selected from multiple clinical interventions. The clinical intervention may be surgical resection, chemotherapy, radiation therapy, immunotherapy, adjuvant therapy, neoadjuvant therapy, androgen deprivation therapy, or a combination thereof. In some cases, the clinical intervention may be administered to the subject. After administration of the clinical intervention, a sample may be obtained from or derived from the subject to monitor the cancer or cancer parameters. Thus, the methods and systems disclosed herein may be repeatedly performed so that cancer monitoring can be performed. Furthermore, by repeatedly performing the method or system, a treatment or clinical intervention can be updated based on the results of the method. Cancer monitoring may include not only an assessment but also a difference between the assessment and a previously generated assessment. Differences in the assessment of a subject's cancer between multiple time points (or samples) may indicate one or more clinical signs, such as a diagnosis of cancer, a prognosis of cancer, or the effectiveness or ineffectiveness of a treatment course for treating the subject's cancer. Prognosis may include expected progression-free survival (PFS), overall survival (OS), or other metrics related to cancer severity or survival.

[0056] The biological sample may be subjected to additional reactions or conditions prior to the assay. For example, the biological sample may be subjected to conditions sufficient to isolate, enrich, or extract nucleic acids, such as cfDNA molecules.

[0057] The methods disclosed herein may include performing one or more enrichment reactions on one or more nucleic acid molecules in a sample. The enrichment reaction may include contacting the sample with one or more beads or bead sets. The enrichment reaction may include one or more hybridization reactions. For example, the enrichment reaction may include contacting the sample with one or more capture probes or bait molecules that hybridize to nucleic acid molecules in the biological sample. The enrichment reaction may include differential amplification of a set of nucleic acid molecules. The enrichment reaction may enrich multiple loci or sequences corresponding to loci. The enrichment reaction may include the use of primers or probes that may be complementary to the sequence (or upstream or downstream sequence) of the sequence to be enriched. For example, the capture probe may include sequence complementarity to a set of genomic loci, allowing for enrichment of genomic loci. The enrichment reaction may include multiple probes or primers. For example, the capture probe may include sequence complementarity to a gene selected from Table 1, Table 2, or Table 3. Multiple probes are available at 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155 5, 160, 165, 170, 175, 180, 185, 190, 195, 200, 205, 210, 215, 220, 225, 230, 235, 240, 245, 250, 255, 260, 265, 270, 275, 280, 285, 290, 295, 300, 305, 310 , 315, 320, 325, 330, 335, 340, 345, 350, 355, 360, 365, 370, 375, 380, 385, 390, 395, 400, 405, 410, 415, 420, 425, 430, 435, 440, 445, 450, 455, 460, 465, 470, 475, 480, 485, 490, 495, 500, 505, 510, 515, 520, 525, 530, 535, 540, 545, 550, 555, 560, 565, 570, 575, 580, 585, 590, 595, or 600 different probes.

[0058] The methods disclosed herein may include performing one or more isolation or purification reactions on one or more nucleic acid molecules in a sample. The isolation or purification reaction may include contacting the sample with one or more beads or bead sets. The isolation or purification reaction may include one or more hybridization reactions, enrichment reactions, amplification reactions, sequencing reactions, or combinations thereof. The isolation or purification reaction may include the use of one or more separators. The one or more separators may comprise magnetic separators. The isolation or purification reaction may include separating bead-bound nucleic acid molecules from bead-free nucleic acid molecules. The isolation or purification reaction may include separating capture probe-hybridized nucleic acid molecules from capture probe-free nucleic acid molecules. The isolation reaction may include removing or separating a group of nucleic acid molecules from another group of nucleic acids.

[0059] The methods disclosed herein can include conducting an extraction reaction on one or more nucleic acids in a biological sample. The extraction reaction can lyse cells or disrupt nucleic acid interactions with cells so that the nucleic acids can be isolated, purified, concentrated, or subjected to other reactions.

[0060] The methods disclosed herein may include an amplification or extension reaction. The amplification reaction may include a polymerase chain reaction. The amplification reaction may include a PCR-based amplification, a non-PCR-based amplification, or a combination thereof. The one or more PCR-based amplifications may include PCR, qPCR, nested PCR, linear amplification, or a combination thereof. The one or more non-PCR-based amplifications may include multiplex displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, rolling circle amplification, circle-to-circle amplification, or a combination thereof. The amplification reaction may include isothermal amplification.

[0061] The methods disclosed herein may include a barcoding reaction. The barcoding reaction may include adding a barcode or tag to a nucleic acid. The barcode may be a molecular barcode or a sample barcode. For example, the barcode nucleic acid may include a barcode sequence, which may be a degenerate n-mer. The sequence may be randomly generated or may be generated to synthesize a specific barcode sequence. The barcode nucleic acid may be added to a sample to label nucleic acid molecules in the sample. The barcode may be specific to the sample. For example, multiple barcode nucleic acids may be added to a sample with the same barcode sequence. When barcoding nucleic acids, those derived from the same sample may have the same barcode sequence, allowing the nucleic acid to be identified as belonging to a specific or given sample. Molecular barcodes may also be used so that each molecule (or multiple molecules) in the same volume has a different molecular barcode. This barcode may be subjected to amplification so that all amplicons derived from the molecule have the same barcode. In this way, molecules derived from the same molecule can be identified. Sequence reads may be processed based on the barcode sequence. For example, the processing may reduce errors or allow molecules to be tracked. The barcode sequence may be added or otherwise appended or incorporated into a sequence by various reactions, such as amplification, extension, or ligation reactions, or may be performed enzymatically using a nucleic acid polymerase or ligase. The ligation may be overhang or blunt-end ligation, and the barcode may comprise complementarity to the nucleic acid being barcoded. This complementarity may be a sequence derived from a sample from a subject, or may be a constant sequence generated via a reaction performed on the nucleic acid in the sample.

[0062] In some cases, a biological sample may contain multiple components. For example, the biological sample may be a whole blood sample. The biological sample may be subjected to a reaction, such as separating or fractionating the biological sample. For example, a whole blood sample may be fractionated to obtain cell-free nucleic acids. A whole blood sample may be fractionated using centrifugation so that blood cells can be separated from plasma (which may contain cell-free nucleic acids). The sample may be subjected to multiple separations or fractionations.

[0063] In various embodiments described throughout this disclosure, nucleic acids may be subjected to a sequencing reaction. Sequencing reactions may be used for DNA, RNA, or other nucleic acid molecules. Examples of sequencing reactions that may be used include capillary sequencing, next-generation sequencing, Sanger sequencing, sequencing-by-synthesis, single-molecule nanopore sequencing, sequencing-by-ligation, sequencing-by-hybridization, nanopore current-limiting sequencing, or a combination thereof. Sequencing-by-synthesis may include reversible terminator sequencing, processive single-molecule sequencing, sequential nucleotide flow sequencing, or a combination thereof. Sequential nucleotide flow sequencing may include pyrosequencing, pH-mediated sequencing, semiconductor sequencing, or a combination thereof. Sequencing reactions may include whole genome sequencing, whole exome sequencing, low-pass whole genome sequencing, targeted sequencing, methylation-recognition sequencing, enzymatic methylation sequencing, and bisulfite sequencing. The sequencing reaction can be transcriptome sequencing, mRNA-seq, total RNA-seq, small RNA-seq, exosome sequencing, or a combination thereof.A combination of sequencing reactions can be used in the methods described elsewhere herein.For example, a sample can be subjected to whole genome sequencing and whole transcriptome sequencing.Because a sample can contain multiple types of nucleic acids (such as RNA and DNA), DNA or RNA specific sequencing reactions can be used to obtain sequence readings related to nucleic acid type.

[0064] The sequencing reaction can be performed at various sequencing depths. The sequencing depth of the sequencing reaction can be selected or adjusted. The sequencing reaction can be performed at least 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 11x, 12x, 13x, 14x, 15x, 16x, 17x, 18x, 19x, 20x, 25x, 30x, 35x, 40x, 45x, 50x, 60x, 70x, 80x, 90x, 100x, 200x, 300x, 400x, 500x, 600x, 700x, 800x, 900x, 1000x, 1100x, 1200x, 1300x, 1400x, 1500x, 1600x, 1700x, 1800x, 1900x, 2000x, 2500x, 3000x, 4000x, 5000x, 6000x, 7000x, 8000x, 9000x, 1000x, 1100x, 1200x, 1300x, 1400x, 1500x, 1600x, 1700x, 1800x, 1900x, 2000x, 2500x, 3000x, 3500x, 4000x, 5000x, 6000x, 7000x, 8000x, 9000x, 1000x, 11 The present invention may include sequencing in an area of ​​00x, 1000x, 2000x, 3000x, 4000x, 5000x, 6000x, 7000x, 8000x, 9000x, 10,000x, 20,000x, 30,000x, 40,000x, 50,000x, 60,000x, 70,000x, 80,000x, 90,000, 100,000x, or more depth. Sequencing reactions are 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 11x, 12x, 13x, 14x, 15x, 16x, 17x, 18x, 19x, 20x, 25 x, 30x, 35x, 40x, 45x, 50x, 60x, 70x, 80x, 90x, 100x, 200x, 300x, 400x, 500x, 600x, 700x, 800x, 900x , 1000x, 2000x, 3000x, 4000x, 5000x, 6000x, 7000x, 8000x, 9000x, 10,000x, 20,000x, 30,000x, 40,000x, 50,000x, 60,000x, 70,000x, 80,000x, 90,000, 100,000x, or less depth.

[0065] In various embodiments, nucleic acids are sequenced using low-pass whole genome sequencing.Low-pass whole genome sequencing can be performed at an average sequencing depth of at least 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, or more.Low-pass whole genome sequencing can be performed at an average sequencing depth of 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, or less.Low-pass whole genome sequencing can be performed at an average sequencing depth of 1x to 2x.

[0066] In various embodiments, the sequencing reaction may be performed using an individualized or customized set of probes. The sequencing reaction using an individualized or customized set of probes may be a deep or ultra-deep sequencing reaction. For example, sequencing reactions using individualized or customized probe sets can be performed at a sequencing depth of 50x, 60x, 70x, 80x, 90x, 100x, 200x, 300x, 400x, 500x, 600x, 700x, 800x, 900x, 1000x, 2000x, 3000x, 4000x, 5000x, 6000x, 7000x, 8000x, 9000x, 10,000x, 20,000x, 30,000x, 40,000x, 50,000x, 60,000x, 70,000x, 80,000x, 90,000, 100,000x, or more.

[0067] In various embodiments, whole exome sequencing is used to sequence the nucleic acid of a subject. Whole exome sequencing can be performed at uneven depths. For example, some regions of the exome can be boosted or otherwise sequenced at a deeper depth than other regions, or at a deeper depth than the average depth of whole exome sequencing. By sequencing specific regions at a higher depth, more interesting genes or regions can be analyzed with higher sensitivity, accuracy, and / or precision. Genes or regions related to or associated with cancer can be sequenced at a deeper depth. For example, at least 100, 200, 300, 400, 500, 600, 700, 800, 900 or more genes can be sequenced at a higher depth than the rest of the exome (for example, the average depth of whole exome sequencing).

[0068] Sequencing of nucleic acids can generate sequencing read data. Sequencing reads can be processed to generate data of improved quality. Sequencing reads can be generated using quality scores. The quality score can indicate the accuracy of the sequence read for a given base call, or a level or signal above a threshold. The quality score can be used to filter sequencing reads. For example, sequencing reads that do not meet a certain quality score threshold can be removed. Sequencing reads can be processed to generate a consensus sequence or consensus base call. A given nucleic acid (or nucleic acid fragment) can be sequenced, and errors in the sequence can be generated due to reactions before or during sequencing. For example, amplification or PCR can generate errors in the amplicon so that the sequence is not identical to the parent sequence. Error correction can be performed using sample barcodes or molecular barcodes. Error correction can include identifying sequence reads that do not corroborate with other sequences from the same sample or the same original parent molecule. The use of barcodes can enable identification of the same parent or sample. Additionally, sequence reads can be processed by performing single-stranded or double-stranded consensus calling, thereby reducing or suppressing errors.

[0069] The methods disclosed herein may include determining allele frequencies or other cancer-related metrics. The methods may include determining the mutant allele frequencies of a set of somatic mutations among a set of biomarkers. The mutant allele frequencies may be used to determine the circulating tumor DNA (ctDNA) fraction of a subject's cancer. The plasma tumor mutation burden (pTMB) of a subject's cancer may be determined based at least in part on the set of mutant allele frequencies. Detection of microsatellite instability may also be used to determine the presence or absence of cancer or cancer metrics. Methylation status may be determined using the methods described herein and used to identify the presence of cancer or cancer parameters.

[0070] In various embodiments, a set of biomarkers is processed to generate data corresponding to the biomarkers. The set of biomarkers may include quantitative measurements from a set of genomic loci associated with cancer. The genomic loci associated with cancer may correspond to a set of genes. The genomic loci associated with cancer may include one or more genes selected from Table 1. In some cases, the set of genomic loci associated with cancer includes 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 members selected from the group consisting of the genes listed in Table 1. The genomic loci associated with cancer can include one or more genes selected from Table 2. In some cases, the set of genomic loci associated with cancer includes 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 members selected from the group consisting of the genes listed in Table 2.

[0071] [Table 1]

[0072] [Table 2]

[0073] The cancer-associated genomic loci can include one or more genes selected from Table 3. In some cases, the cancer-associated genomic loci include between 2 and 600 genes selected from Table 3. In some cases, the set of cancer-associated genomic loci is selected from the group consisting of genes listed in Table 3, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 205, 210, 215, 220, 225, 230, 235, 240, 245, 250, 255, 260, 265, 270, 275, 280, 285, 290, 305, 310, 315, 320, 325, 330, 335, 340, 345, 350, 355, 360, 365, 370, 375, 380, 385, 390, 405, 410, 415, 420, 425, 430, 435, 440, 445, 450, 455, 460, 465, 470, 475, 480, 485, 490, 505, 5 5, 290, 295, 300, 305, 310, 315, 320, 325, 330, 335, 340, 345, 350, 355, 360, 365, 370, 375, 380, 385, 390, 395, 400, 405, 410, 415, 420, 425, 430, 435, 440, 445, 45 Contains 0, 455, 460, 465, 470, 475, 480, 485, 490, 495, 500, 505, 510, 515, 520, 525, 530, 535, 540, 545, 550, 555, 560, 565, 570, 575, 580, 585, 590, 595, or 600 members.

[0074] [Table 3-1]

[0075] [Table 3-2]

[0076] The set of biomarkers can correspond to genetic abnormalities of gene loci. The genetic abnormalities can be tumor-related changes. The genetic abnormalities can include copy number alterations (CNAs), copy number losses (CNLs), single nucleotide variants (SNVs), insertions or deletions (indels), and / or rearrangements. The set of biomarkers can be identified in various nucleic acid types. For example, tumor-related changes can be identified in cfDNA. The tumor-related changes can include changes in allele expression or gene expression. The methods and systems disclosed herein can enable gene expression profiling and the identification of changes in gene expression levels.

[0077] In various embodiments, the method may include identifying the presence of cancer or a cancer parameter. The method may include determining the probability or likelihood of the presence of cancer or a cancer parameter. For example, instead of a binary output indicating presence or absence, an output may be generated indicating the probability that the subject has cancer. This probability may be determined based on an algorithm described elsewhere herein. Similarly, the probability or likelihood of response to a particular treatment, or the probability of recurrence, may be output.

[0078] In various aspects, the set of biomarkers is processed using an algorithm. The algorithm may be a trained algorithm. The trained algorithm may use the set of biomarkers as input and generate an output regarding the presence or absence of cancer. The output may be specific to a type of cancer or a subtype of cancer. For example, the output may indicate the presence of bladder cancer.

[0079] The trained algorithm may be trained on a plurality of samples. For example, the trained algorithm may be trained on at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, It may be trained using 190, 195, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or more independent training samples. The trained algorithms were: 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 205, 210, 215, 220, 225, 230, 235, 240, 245, 250, 255, 260, 265, 270, 275, 280, 285, 290, 300, 305, 310, 315, 320, 325, 330, 335, 340, 345, 350, 355, 360, 365, 370, 375, 380, 385, 390, 410, 420, 430, 440, 450, 460, 470, 480, 490, 510, 520, 530, 540, 550, 560, 570, 580, 590, 610, 620, 630, 640, 650, 66 The algorithm may be trained using up to 95, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, or fewer independent training samples. The training samples may be associated with the presence or absence of cancer. The training samples may be associated with cancer recurrence. The training samples may be associated with cancer that is resistant to a particular drug or treatment. Individual training samples may be positive for a particular cancer. Individual training samples may be negative for a particular cancer. By using the training samples, the trained algorithm may be able to detect cancer, determine the recurrence or probability of recurrence of cancer, or determine whether a cancer contains a set of biomarkers. The training samples may be associated with additional clinical health data of the subject, for example, the additional clinical health data may include gender, weight, height, or levels of metabolites or antibodies in the subject.The additional clinical health data may include indicators of other diseases, disorders, or disease states.

[0080] The trained algorithm can be trained using multiple sets of training samples.Sets can include the training samples as described elsewhere herein.For example, training can be carried out using a first set of independent training samples that are related to the presence of cancer and a second set of independent training samples that are related to the absence of cancer.Similarly, the first set can be associated with recurrence, and the second set can be associated with the absence of recurrence.

[0081] The trained algorithm may also process additional clinical health data of the subject. For example, the additional clinical health data may include the gender, weight, height, or levels of metabolites or antibodies in the subject. The additional clinical health data may include indicators of other diseases, disorders, or disease states that the subject may suffer from. By using the additional clinical health data in conjunction with the biomarkers, the trained algorithm may output the presence or absence of cancer, the probability of recurrence, or resistance to drug treatment, which may differ from the output of an algorithm that does not process the additional clinical health data.

[0082] The trained algorithm may be an unsupervised machine learning algorithm. For example, the unsupervised machine learning algorithm may utilize cluster analysis to identify attributes of interest. The trained algorithm may be a supervised machine learning algorithm. For example, the trained algorithm may be trained using training data to generate expected or desired outputs. Supervised learning algorithms may include deep learning algorithms, support vector machines (SVMs), neural networks, or random forests. Through the machine learning algorithm, the trained algorithm may be able to identify the relationship of biomarkers to a particular cancer prognosis or diagnosis. Without the trained algorithm, it may otherwise be difficult to identify the relationship of biomarkers to accurately identify the presence of cancer or other parameters associated with cancer.

[0083] In various embodiments, the systems and methods may include accuracy, sensitivity, or specificity of detecting cancer or a parameter of cancer. For example, the method or system may include detecting the presence or absence of cancer (or the presence of a parameter of cancer, e.g., recurrence, relapse, or drug resistance) in a subject with an accuracy of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. The method or system may include detecting the presence or absence of cancer (or the presence of a parameter of cancer, e.g., recurrence or drug resistance) in a subject with a sensitivity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. The method or system may include detecting the presence or absence of cancer (or the presence of a parameter of cancer, e.g., recurrence, relapse, or drug resistance) in a subject with a specificity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. The method or system may include detecting the presence or absence of cancer (or the presence of a parameter of cancer, e.g., recurrence, relapse, or drug resistance) in a subject with a positive predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. The method or system may include detecting the presence or absence of cancer (or the presence of a parameter of cancer, e.g., recurrence or drug resistance) in a subject with a negative predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.

[0084] Computer Control System The present disclosure provides a computer system programmed to implement the methods of the present disclosure. Figure 31 shows a computer system (3101) programmed or otherwise configured to perform the analyses or operations of the present disclosure, for example, to determine the likelihood of the presence of cancer based on a set of biomarkers for an individual, or to execute an algorithm. The computer system (3101) can coordinate various aspects of the methods and systems of the present disclosure, such as, for example, executing an algorithm, inputting training data, analyzing a set of biomarkers, or outputting a result to a user regarding the presence or absence of cancer. The computer system (3101) may be an electronic device of a user or a computer system, and may be located remotely relative to the electronic device. The electronic device may be a mobile electronic device.

[0085] The computer system (3101) includes a central processing unit (CPU, also referred to herein as a "processor" and a "computer processor") (3105), which may be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system (3101) also includes memory or storage locations (3110) (e.g., random access memory, read-only memory, flash memory), electronic storage devices (3115) (e.g., hard disks), communication interfaces (3120) (e.g., network adapters) for communicating with one or more other systems, and peripherals (3125), such as cache, other memory, data storage devices, and / or electronic display adapters. The memory (3110), storage devices (3115), interfaces (3120), and peripherals (3125) communicate with the CPU (3105) via a communication bus (solid lines), such as a motherboard. The storage devices (3115) may be a data storage device (or data repository) for storing data. The computer system 3101 may be operatively connected to a computer network ("network") 3130 with the aid of a communication interface 3120. The network 3130 may be the Internet and / or an extranet, an intranet and / or an extranet in communication with the Internet. In some cases, the network 3130 is a telecommunications and / or data network. The network 3130 may include one or more computer servers, which may enable distributed computing, such as cloud computing. In some cases, the network 3130 may implement a peer-to-peer network with the aid of the computer system 3101, thereby enabling devices coupled to the computer system 3101 to act as clients or servers.

[0086] The CPU (3105) can execute a series of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a storage location, such as the memory (3110). The instructions may be directed to the CPU (3105), which may then program or otherwise configure the CPU (3105) to implement the methods of the present disclosure. Examples of operations performed by the CPU (3105) include fetch, decode, execute, and writeback.

[0087] The CPU 3105 may be part of a circuit, such as an integrated circuit. One or more other components of the system 3101 may also be included in the circuit. In some cases, the circuit is an application-specific integrated circuit (ASIC).

[0088] The storage device (3115) can store files such as drivers, libraries, and saved programs. The storage device (3115) can store user data, such as user preferences and user programs. The computer system (3101) may optionally include one or more additional data storage devices external to the computer system (3101), such as located on a remote server in communication with the computer system (3101) via an intranet or the Internet.

[0089] The computer system (3101) can communicate with one or more remote computer systems via a network (3130). For example, the computer system (3101) can communicate with a remote computer system of a user (e.g., a service provider or a patient). Examples of remote computer systems include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., an Apple® iPad, a Samsung® Galaxy Tab), a telephone, a smartphone (e.g., an Apple® iPhone, an Android-enabled device, a Blackberry®), or a personal digital assistant. A user can access the computer system (3101) via the network (3130).

[0090] Methods as described herein can be performed by machine (e.g., computer processor) executable code stored in an electronic storage location of the computer system (3101), such as, for example, on memory (3110) or electronic storage (3115). The machine-executable or machine-readable code can be provided in the form of software. During use, the code can be executed by the processor (3105). In some cases, the code can be retrieved from storage (3115) and stored in memory (3110) for immediate access by the processor (3105). In some situations, the electronic storage (3115) can be omitted, and machine-executable instructions are stored in memory (3110).

[0091] The code may be pre-compiled and configured for use with a machine having a suitable processor to execute the code, or may be compiled at run-time. The code may be provided in a programming language that can be selected to render the code executable in a pre-compiled or as-compiled fashion.

[0092] Aspects of the systems and methods provided herein, such as the computer system (3101), can be integrated into programming. Various aspects of this technology may be considered as an "article of manufacture" or "article of manufacture," typically in the form of machine (or processor) executable code and / or associated data carried on or embedded in a type of machine-readable medium. The machine-executable code may be stored in electronic storage, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. "Storage" type media may include any or all of the tangible memory of a computer or processor, or its associated modules, such as various semiconductor memories, tape drives, disk drives, etc., which may provide non-transitory recording media at any time for programming the software. All or portions of the software are sometimes communicated via the Internet or various other telecommunications networks. Such communication may enable, for example, loading of the software from one computer or processor to another, e.g., from a management server or host computer to an application server computer platform. Thus, other types of media that may carry software elements include optical, electrical, and electromagnetic waves, such as those used over wired and terrestrial optical communication networks between local devices and various air links. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., may also be considered media that carry software. As used herein, unless limited to non-transitory, tangible "storage" media, terms such as computer or machine "readable medium" refer to media that participate in providing instructions to a processor for execution.

[0093] Thus, a machine-readable medium such as a computer-executable code may take many forms, including, but not limited to, a tangible storage medium, a carrier wave medium, or a physical transmission medium. Non-volatile storage media include optical or magnetic disks, such as any of the storage devices in any computer(s), such as those that may be used to implement the databases shown in the figures. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible storage media include coaxial cables, copper wire, and fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, other magnetic media, CD-ROMs, DVDs or DVD-ROMs, other optical media, punch cards, paper tape, other physical storage media with patterns of holes, RAM, ROM, PROMs and EPROMs, FLASH-EPROMs, other memory chips or cartridges, carrier waves carrying data or instructions, cables or links which transmit such carrier waves, or other media from which a computer can read programming code and / or data. Many of these forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.

[0094] The computer system (3101) may include or be in communication with an electronic display (3135) that includes a user interface (UI) (3140) for providing visual output regarding, for example, biomarkers of sequencing data, or detection, diagnosis, or prognosis. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.

[0095] The method and system of the present disclosure can be implemented by one or more algorithms.The algorithm can be implemented by software after being executed by a central processing unit (3105).For example, the algorithm can determine the presence or absence of cancer or cancer parameters based on a set of input sequencing data from a sample derived from a subject. [Example]

[0096] Example 1: Analysis of cell-free DNA for cancer detection

[0097] Using the disclosed method and system, circulating tumor DNA in urine was analyzed to identify genomic alterations in subjects with bladder cancer. Up to 90 mL of urine was collected in tubes containing urine storage buffer. Upon receipt, samples were deposited, immediately processed into urine supernatant, and stored in a -80°C freezer. ucfDNA was extracted, purified by beads, and then quantified using a Qubit, Agilent Bioanalyzer, or Fragment Analyzer. Up to 15 ng of size-selected ucfDNA can be used in this assay. To construct libraries, extracted ucfDNA was labeled with unique molecular barcodes. Ligated sequencing libraries were PCR-amplified using a high-fidelity polymerase and quantified by a Bioanalyzer. Target enrichment and hybrid capture were performed. During enrichment, the sequencing libraries were blocked with adapter-specific blocking oligonucleotides and then hybridized with the PredicineCARE panel (e.g., Table 2). Library capture was performed by beads, amplified, and quantified by a Bioanalyzer. To perform sequencing, the enriched libraries were normalized, pooled, and loaded onto an Illumina platform for 2×150 bp paired-end sequencing. Libraries were sequenced to a median depth of over 20,000×.

[0098] The PredicineCARE assay identified SNVs, indels, CNVs, and DNA rearrangements in a single workflow by using a bioinformatics pipeline developed for high accuracy in variant detection. NGS data were analyzed using the DeepSEA NGS analysis pipeline, which starts with raw sequencing data (BCL files) and outputs final variant calls. Figure 2 shows a schematic diagram of the DeepSEA pipeline. The pipeline performs adapter trimming, barcode checking, and error correction. Pairs of clean FASTQ files were aligned to the human reference genome build hg19 using the BWA alignment tool. A consensus BAM file was then obtained by merging single-stranded fragments and paired-end reads derived from the same molecule. Single-stranded fragments derived from the same double-stranded DNA molecule were further merged as double strands. Both sequencing and PCR errors were corrected during this process.

[0099] Next, a variant filter for SNV and indel detection was used. Variants were filtered based on variant background from a pool of normal control samples and other historical samples. Other metrics, such as base quality, log odds ratio, and distance to fragment ends, were used to remove variants with low confidence. Detected calls were variants with at least four unique supporting fragments, one of which was double-stranded.

[0100] Variants were called after filtering out low base quality and low mapping score reads. Detected variants were further filtered based on variant background (defined by normal plasma samples and historical data), repetitive regions, and other quality metrics. Benign and likely benign SNPs were excluded from the variant call list.

[0101] When matched normal samples are not available, a germline variant filter is implemented. In practical applications, most liquid biopsy assays are not matched to normal samples. The variant allele frequency was adjusted by the copy number change occurring at the variant position. Variants annotated in public germline databases, such as 1000 Genomes, that have relatively high population allele frequencies were also excluded.

[0102] Copy number variation (CNV) was estimated at the gene level. The pipeline calculated on-target unique fragment coverage based on the consensus BAM file and then adjusted for GC bias coverage. Unique fragment coverage profiles from a group of normal samples were used as a reference to normalize and estimate z-scores of copy number changes. Rearrangements were detected by identifying alignment breakpoints based on the BAM file before the consensus operation. Suspicious alignments were filtered based on repetitive regions, local entropy calculations, and similarity between the reference and alternative alignments. Rearrangement calls were reported using two or more unique alignments.

[0103] To obtain the samples, urine was collected using a urine collection kit. The components of the urine collection kit are listed in Table 4. The kit contains one reagent, Streck® Urine Preserve. Streck Inc. (Omaha, Nebraska, USA) is the supplier of the Streck® Urine Preserve reagent in the kit.

[0104] [Table 4]

[0105] To evaluate the PredicineCARE panel and assay for somatic alteration detection in ucfDNA, various samples with known mutational alterations, including SNVs, indels, DNA rearrangements, copy number gains, and copy number losses, were used in this assay validation. In this example, the limit of detection (LoD) was defined as the lowest variant allele frequency at which 95% of variants across all replicates for a variant type could be reliably detected by the PredicineCARE assay. The positive percent agreement (PPA) was used to indicate assay sensitivity. 15 ng of ucfDNA sample input was used in the validation study.

[0106] To evaluate the LoD of single-nucleotide variation (SNV) detection in urine specimens, we spiked ucfDNA from healthy male donors carrying unique SNPs into ucfDNA from healthy female donors to generate a series of test materials with variable variant allele frequencies (MAFs) ranging from 0.125 to 10%. We detected 89 of 90 predicted mutations with an MAF of 0.5%, achieving an assay sensitivity of 98.89% (95% CI, 94-100%) with a PPV of 100% (95% CI, 95.9-100%) (Table 5 and Figure 3).

[0107] [Table 5]

[0108] To evaluate the LOD of indel, CNV, and fusion detection in ucfDNA, we fragmented HD753 gDNA to the ucfDNA size and spiked them into ucfDNA from healthy donors to generate samples with predefined allele frequencies (AF). We used HD753 reference gDNA from Horizon Discovery (Cambridge, UK), which contains known indels, CNVs, and fusions. For indels, a total of 35 variants were detected with an AF of 0.5–1%, achieving an assay sensitivity of 97.22% (95% CI, 85.5–99.9%) and a PPV of 100% (95% CI, 90–100%) (Table 6).

[0109] [Table 6]

[0110] For copy number variation (CNV) detection, samples with titer levels between 2.375 and 3.125 copies were evaluated. All CNV variants were detected at 2.375 copies, achieving 100% assay sensitivity (95% CI, 69.2–100%) and 100% PPV (95% CI, 69.2–100%) (Table 7).

[0111] [Table 7]

[0112] For DNA rearrangements, samples with titer levels ranging from 0% AF to 0.825% AF were analyzed. As shown in Table 5, the LoD for DNA rearrangement detection is 0.25-0.55% AF, with an assay sensitivity of 100% (95% CI, 78.2-100%) and a PPV of 100% (95% CI, 78.2-100%) (Table 8).

[0113] [Table 8]

[0114] To assess assay specificity and ensure that "blank" samples did not generate analytical signals, 18 ucfDNA samples from healthy donors were tested. Analytical specificity was estimated based on the number of false-positive mutations in the target panel. Analytical specificity = 100 * (1 - number of false positives / panel size). Although healthy donors may have pathogenic mutations, such as CHIP mutations, with low variant frequencies, we considered all variants that passed the detection criteria of the NGS analysis pipeline as false positives. Analytical specificity is 99.9998%. Accuracy was measured by the variation in estimated variant frequencies between replicates. All samples in the accuracy assay were evaluated from the library construction stage to sequencing analysis.

[0115] Repeatability testing included intra-run performance (samples processed under the same conditions). Three replicates of each sample were performed under the same conditions, and the results were compared. Reproducibility was assessed based on six samples processed independently under different operating conditions. Agreement of called variants from replicates is used to assess intra- and inter-assay precision.

[0116] 100% agreement between replicates was detected for both intra- and inter-run precision studies. Furthermore, high agreement was observed between the expected and detected MAFs in precision study samples (Table 9 and Figure 4).

[0117] [Table 9]

[0118] We evaluated the performance of urine-based variant detection in clinical samples using DNA from 43 paired urine and tumor tissue samples from bladder cancer patients. Mutations detected in tissue samples were used as a reference for mutations detected in ucfDNA. Concordance was determined by comparing variants detected in urine with those in paired tissue samples. The concordance between mutations detected in urine and tissue samples was 81.0% (95% CI: 77.2–84.4%) (Table 10). Furthermore, the PredicineCARE assay achieved 94.6% concordance (95% CI, 87.9–98.2%) for five highly mutated genes in bladder cancer (TERT, TP53, KDM6A, PIK3CA, and FGFR3) (Figure 5).

[0119] [Table 10]

[0120] The PredicineCARE assay described in this example analyzed somatic variants in 152 cancer-associated genes from liquid biopsies. The assay underwent rigorous analytical validation testing and demonstrated robust and reproducible results in plasma cfDNA. As demonstrated in this study, the PredicineCARE assay can detect genetic variants in ucfDNA with high sensitivity and specificity (Table 11), making the assay a valuable tool for analyzing somatic variants in various body fluids. Combined with the ease of use offered by completely non-invasive sample collection and urine-based liquid biopsies, the PredicineCARE assay offers great potential for real-time genomic profiling for cancer detection, enabling more frequent monitoring.

[0121] [Table 11]

[0122] Example 2: Detection of bladder cancer using urine

[0123] In this example, we conducted a study to evaluate the concordance between tissue tumor DNA profiling and urinary cfDNA or circulating tumor DNA (ctDNA) using the PredicineCare urine cfDNA assay. This study prospectively enrolled 59 cases of bladder cancer with pathologically confirmed disease and matched paired tissue / urine samples. Baseline peripheral blood mononuclear cells (PBMCs) and plasma specimens were collected during clinic visits (Zhang et al., 2021). Urine, tissue, PBMCs, and plasma samples were processed with the PredicineCARE assay and analyzed with the DeepSEA bioinformatics pipeline. Concordance analysis was performed using tissue tumor DNA as a reference. Urine cfDNA achieved 99.3% specificity, 86.7% sensitivity, and 99.1% diagnostic accuracy. FGFR3 alterations and ERBB2 amplification were identified in the urine cfDNA. Quantitative metrics, including cancer cell fraction, variant allele frequency, and tumor mutation burden, were concordant between tumor tissue DNA and urine cfDNA. Plasma cfDNA was poorly matched with tumor tissue DNA due to ctDNA abnormalities arising from clonal hematopoiesis.

[0124] Example 3: Detection of Bladder Cancer Using a Urine-Based NGS Assay

[0125] The clinical performance of the urine-based NGS assay was evaluated by comparing its results with those from an FDA-approved tissue-based PCR CDx assay that detects significant alterations in FGFR genes. Figure 11 provides a schematic diagram of the study design. Paired urine and tissue samples were collected from 107 bladder cancer patients (muscle-invasive and non-muscle-invasive) from the German Bladder BRIDGister clinical trial. Tissue specimens were analyzed using the FDA-approved Qiagen therascreen FGFR RGQ RT-PCR kit, and matched urine samples were processed using the PredicineCARE™ Urine (cell-free DNA) cfDNA NGS assay, which has a detection sensitivity of 0.3% (0.1% for hotspot mutations). 107 paired bladder cancer urine cfDNA NGS and tissue RT-PCR results were analyzed to determine concordance (PPA and NPA) between these two assays. Smaller sample sets were also used to compare tissue NGS with therascreen RT-PCR and tissue NGS with urine cfDNA NGS. A subset of discordant mutations was further validated by ddPCR.

[0126] For concordance analysis between PredicineCARE urine cfDNA NGS and therascreen FGFR RGQ RT-PCR kit (107 samples), PPA was 100% (20 / 20, 95% CI: 83.2-100) and NPA was 94.1% (48 / 51, 95% CI: 83.8-98.8). PPA and NPA between PredicineCARE urine cfDNA and tissue NGS were 100% (19 / 19, 95% CI: 82.4-100) and 94.2% (49 / 52, 95% CI: 84.1-98.8), respectively. PPA and NPA between PredicineCARE tissue NGS and therascreen FGFR RGQ RT-PCR kit were both >95%.

[0127] Figure 12 shows a heatmap of matched urine NGS and FFPE tissue RT-PCR for identified gene alterations. Each column represents a sample, adjacent columns represent matched FFPE and urine, and rows represent genes. There was high concordance between urine NGS and FFPE RT-PCR results. Most samples showed gene alterations in both urine NGS and FFPE samples.

[0128] Figures 13A-13B show scatter plots of variant allele frequencies (VAF) between matched FFPE and urinary variants. Figure 13A shows a scatter plot of all identified genetic alterations, including somatic and germline variants. Figure 13B shows a scatter plot of somatic FGFR3 alterations. The X-axis is the VAF of urine NGS, while the Y-axis is the VAF of FFPE RT-PCR.

[0129] [Table 12]

[0130] [Table 13]

[0131] Three tissue FGFR-negative samples by RT-PCR were FGFR-positive by urine NGS testing, but no urine NGS FGFR-negative samples were FGFR-positive by tissue RT-PCR (Table 13). Discrepant samples (tissue FGFR WT or invalid but positive urine for FGFR mutations) were further analyzed and confirmed as positive by independent orthogonal ddPCR (Bio-Rad ddPCR Mutation Detection Assay), suggesting that discrepancies in results between urine and tissue are often caused by the reduced sensitivity of tissue FGFR RT-PCR testing.

[0132] The high concordance between FGFR alterations detected with the FDA-approved tissue companion diagnostic (CDx) assay and the urine cfDNA NGS assay demonstrates that the PredicineCARE urine cfDNA NGS assay may represent a novel, accurate, and non-invasive clinical application for molecular diagnostic testing to identify biomarkers in bladder cancer.

[0133] Example 4: Determining copy number burden using low-pass whole genome sequencing

[0134] Copy number variation (CNV) is a key feature in cancer genomes. Blood / urine-based low-pass whole-genome sequencing (LP-WGS) is increasingly being used to identify large genome-wide copy number variations in cancer. In this study, we developed PredicineCNB™, a companion LP-WGS assay for robustly estimating genome-wide copy number burden (CNB) from plasma and urine clinical samples, which can provide cancer / normal classification and long-term treatment monitoring in a cost-effective manner (e.g., using the PredicineSCORE™ liquid biopsy assay (Predicine, Inc., Hayward, CA)).

[0135] Analytical evaluation was performed based on clinical plasma titration samples, and PredicineCNB™ demonstrated high cancer detection sensitivity (LOD is 1%) with high specificity using DNA input as low as 1 ng. Figure 14A shows the workflow of the PredicineCNB assay. Blood (or urine) samples are obtained. For blood samples, plasma is isolated and DNA is extracted. Libraries are then prepared and sequenced using low-pass whole-genome sequencing at 3x depth, followed by analysis. Figure 14B shows an example of genome-wide CNV detection, demonstrating that the CNB calculation is robust to plasma inputs as low as 0.5 ng: 5 ng input plasma, CNB score = 11.7 (top); 0.5 ng input plasma, CNB score = 11.7 (bottom). Figure 14C shows that the LP-WGS CNV profiles of 5 ng and 0.5 ng input plasma are highly consistent over a 1 Mb region, and Figure 14D shows that the LP-WGS copy numbers between 5 ng and 0.5 ng input plasma are highly consistent at the chromosome arm level, with an even higher correlation coefficient than the 1 Mb region.

[0136] Figures 15A-15D show analytical evaluation of PredicineCNB on clinical plasma samples and its application to 688 prostate cancer patient samples. As shown in Figure 15A, the LOD of PredicineCNB is 1%, and the specificity is greater than 97.6% (41 / 42). CNB scores are generated for cancer plasma titration samples and corresponding normal plasma baseline samples, as shown in Figure 15B. As shown in Figure 15C, the CNB score distribution of 688 clinical prostate cancer patients is generated using 430 samples classified as high-risk and 258 samples classified as low-risk. Figure 15D shows an LP-WGS CNV profile heatmap of 430 prostate cancer patient samples classified as high-risk, with each row representing copy number deviations across all chromosomes in 1-Mb bins for each clinical sample.

[0137] The assay was also performed on non-muscle invasive and muscle invasive bladder cancer. Figures 16A-16B show an overview of the LP-WGS CNV profile heatmaps of non-muscle invasive and muscle invasive bladder cancer patient samples. Figure 16A shows the LP-WGS CNV profile heatmaps of 14 non-muscle invasive bladder cancer patient samples. Figure 16B shows the LP-WGS CNV profile heatmaps of 33 muscle invasive bladder cancer patient samples.

[0138] Predicine CNB was tested for its usefulness in long-term bladder cancer patient treatment monitoring. Figure 17A shows a comparison of CNB scores of FFPE, plasma, and urine samples between patients with non-muscle invasive bladder cancer and patients with muscle invasive / non-organ confined bladder cancer. Figure 17B shows the LP-WGS CNV profiles of two non-invasive bladder cancer patients before and after TURBT surgery, demonstrating a decrease in CNB scores for both patients after TURBT surgery. Figure 17C shows the LP-WGS gene copy number of major bladder cancer genes before and after TURBT, demonstrating the monitoring of different cancer genes before and after surgery.

[0139] This example demonstrates an algorithmic analysis pipeline for the LP-WGS assay. The copy number burden (CNB) LOD is 1% tumor fraction with 1 ng of plasma input. PredicineCNB™ demonstrates the promising clinical application of urine- and blood-based genome-wide copy number alterations in highly sensitive cancer detection and treatment monitoring.

[0140] reference

[0141] [1] Davis AA, Luo J, Zheng T, Dai C, et al. Genomic complexity predicts resistance to endocrine therapy and CDK4 / 6 inhibition in hormone receptor-positive (HR+) / HER2-negative metastatic breast cancer. Clin Cancer Res. 2023 Jan 24; CCR-22-2177. doi: 10.1158 / 1078-0432. CCR-22-2177, which is incorporated herein by reference in its entirety.

[0142] Example 5: Detection and monitoring using methylation

[0143] Using the disclosed method and system, circulating tumor DNA from biological samples was analyzed. A combined MRD assay was performed using a targeted panel covering hotspot mutations and critical genes (Predicine WES+), a companion LP-WGS assay for copy number burden (Predicine CNB), and a whole-genome methylation assay (e.g., Predicine EPIC).

[0144] Figure 18 shows an example workflow schematic. A sample (blood, urine, or tissue) was obtained from a subject. DNA was extracted and a library was generated. The library was methylation treated. Once the library was constructed, next-generation sequencing was performed to analyze mutations, CNVs, and methylation data.

[0145] Analytical evaluation was based on titration of real-world clinical patient samples. DNA input ranged from 30 ng to 1 ng. Clinical evaluation was based on longitudinal samples from patients with different cancer indications, including bladder cancer, mCRPC, CRC, breast cancer, and NSCLC.

[0146] Figure 19 shows Predicine EPIC libraries with 1 ng, 2.5 ng, and 10 ng inputs compared to the standard 50 ng DNA input used for standard whole-genome bisulfite sequencing libraries. Specifically, Predicine EPIC libraries with 1 ng, 2.5 ng, and 10 ng provide similar coverage and data as split into a standard 50 ng whole-genome bicarbonate sequencing library.

[0147] Ninety-seven cancer and normal tissue types were analyzed using fragment-level DNA methylation analysis. These assays were able to identify fragments with increased or decreased methylation compared to background models. These assays were able to classify different cancers by tissue type and whether the sample was cancerous or non-cancerous. Figure 20 shows a graph depicting uniform manifold approximation and projection (UMAP). Different populations are visualized based on different UMAP scores.

[0148] For each sample, a methylation aberration score can be calculated based on the methylation data. Figure 21 shows an exemplary heatmap of the significance scores (color scale) of aberrantly methylated fragments at 28.4K of the most variable CpG sites of a total of 142K genome-wide covered sites (columns) for 35 bladder cancer samples (rows) from patients of different stages and different sources (row annotation).

[0149] In addition to the methylation aberration score, a copy number load aberration score can also be calculated on the low-pass sequencing data. These two aberration scores have been shown to be highly correlated, as shown in Figure 22, and can be used in combination to identify subjects with cancer. In addition, the disease progression of patients can be calculated using the DNA methylation score as shown in Figure 22. Three patients are shown, where T0 is before treatment, and T1 and T2 are after treatment, with a significant decrease in the aberration score after treatment.

[0150] Example 6: Boosted whole-exome sequencing to prevent muscle-invasive bladder cancer

[0151] Urinary tumor DNA profiling can be used for the diagnosis, monitoring, and treatment stratification of bladder cancer.However, previous studies have mainly used targeted next-generation sequencing (NGS) panel approaches, which are limited to certain genes and therefore lack comprehensiveness.Here, this example demonstrates that boosted whole exome sequencing (WES) is used for urine and tissue tumor DNA in muscle-invasive bladder cancer (MIBC) to comprehensively compare the mutation profiles in matched urine and tissue samples.

[0152] Matched tumor tissue, urine, and peripheral blood mononuclear cell (PBMC) samples were collected from 20 MIBC patients. Nineteen tumor tissue samples, 19 urine samples, and 20 PBMC samples that passed sample quality control were processed for next-generation sequencing (NGS). Predicine WES+, an NGS assay with whole-exome and boosted coverage in 600 cancer-associated genes from the Predicine ATLAS panel, was applied to matched tumor, urine, and PBMC samples for variant profiling. Mutational profiles of tumor tissue and urine DNA were analyzed and compared.

[0153] Figure 23 shows the mutation profiles of urine and tissue tumor DNA from MIBC. The mutation profiles of urine and tissue tumor DNA were highly consistent between patients, with hypermutated genes (such as TERT.TP53, ARID1A, KMT2D, KDMSA, and PIK3CA) showing comparable prevalence. No two tissue samples failed sequencing QC.

[0154] Concordance of mutations detected by Predicine WES+ in tissue tumor DNA (tDNA) and urinary tumor DNA (utDNA). We compared the number of mutations detected in tDNA and utDNA in the WES region (Figures 24A-24B) and the ATLAS region (Figures 24C-24D). The majority of tDNA mutations in the ATLAS region (67.5% and 80.1% in the WES region) were also detected in utDNA. However, less than half of the utDNA mutations were detected in tDNA (42.1% in the WES region and 39.9% in the ATLAS region).

[0155] Tumor fraction (TF) was also estimated from pairs of urine (2-52%) and tumor tissue (17-68%) samples, showing significant differences (a, p=0.05). Figure 25 shows plots of tumor fraction in tissue and urine. Although TF in urine was relatively low, more somatic mutations were detected in urine than in tumor tissue (b, p<0.05). TMB was also calculated and compared from tDNA and utDNA, showing a high correlation (R=0.84).

[0156] Overall, the results demonstrate the validity of urinary tumor DNA as a tissue surrogate for mutational profiling in MIBC at the whole-exome scale and support urine-based noninvasive molecular profiling in precision medicine for bladder cancer patients.

[0157] PredicineWES+™, a boosted whole-exome sequencing assay, identified 1,493 somatic variants in 11 CSF samples, 97 of which had previously been reported as likely pathogenic in public clinical databases. Regarding NSCLC-specific biomarkers, 7 of 11 patients harbored EGFR variants, including 3 Exon19del events, 1 Exon20ins event, 1 L858R mutation, and other gain-of-function mutations. Additionally, an EML-ALK mutation was detected in 1 patient.

[0158] Example 7: Monitoring breast cancer using blood samples

[0159] Two comprehensive NGS assays were performed to profile somatic mutations and copy number variations in blood samples collected from patients with HR+ / HER2-negative metastatic breast cancer at baseline and during treatment with endocrine therapy and CIDK4 / 6 inhibition (ET+CDK4 / 6i). Specifically, blood samples were evaluated from a phase II trial of palbociclib plus letrozole or fulvestrant on a 5-day-on / 2-day-off weekly schedule in a 28-day cycle as first- or second-line treatment.

[0160] The first assay used was PredicineWES+, a boosted whole exome sequencing (WES) assay that uses WES combined with the coverage depth of 600 cancer genes targeted by the PredicineATLAS panel to generate an exome-wide genomic profile of somatic single nucleotide variants (SNVs), indels, and copy number variations (CNVs) and determine a hematologic tumor mutation burden (bTMB) score, which reflects the number of mutations in the DNA.

[0161] The second assay was PredicineCNB, a low-pass whole genome sequencing (LP-WGS) assay, used to generate a blood copy number burden (bCNB) score that represents a comprehensive measure of copy number variation, including amplifications and deletions, across all chromosomal arms (e.g., using the PredicineSCORE™ Liquid Biopsy Assay (Predicine, Inc., Hayward, CA)).

[0162] Figures 26A-26C show the study schematic, sample collection, and assay timeline and procedures. Figure 26A shows the study schematic. Figure 26B shows the sample collection and sequencing timeline. Sample collection at baseline (BL), and on-treatment at cycle 1 day 15 (C1D15), cycle 2 day 1 (C2D1), Q3-month staging scans without progressive disease (PD), and imaging detection of PD. 216 serial blood samples were collected from 51 patients at baseline and during treatment. In addition, germline DNA samples were collected from each patient. Following QC procedures, 78 blood samples were sequenced using PredicineWES+ and 218 samples were sequenced using PredicineCNB. Similarly, 49 germline samples were sequenced using PredicineWES+. Figure 26C shows NGS profiling using PredicineWES+ and PredicineCNB. Blood samples were separated into plasma samples containing cfDNA, buffy coats containing gDNA, and red blood cells (RBCs). DNA was extracted from the samples and libraries were generated. The libraries were subjected to two different sequencing workflows. One portion of the sample underwent low-pass whole-genome sequencing at 5x depth, and the reads were analyzed to determine the blood copy number burden. The second portion of the sample underwent targeted enrichment, with whole-exome sequencing (at 2500x depth) and further enrichment for 600 specific target genes from Predicine ATLAS (at 20,000x depth). These reads were then used to identify SNVs, indels, CNVs, and fusions and determine the tumor mutation burden.

[0163] Figures 27A-27B show data from the Predicine WES+ assay. Figure 27A shows a heatmap of the top altered genes in baseline and progression samples. The number and type of genomic alterations, as well as the frequency of alterations for specific genes in baseline and progression samples, were also analyzed. Based on this data, enrichment of specific variants could be identified at progression, and a decrease in the total number of SNVs was observed at progression. Figure 27B shows the status of baseline alterations compared to progression. However, an increase in the total number of CNVs, primarily copy number loss events, was observed in progression samples. When observing blood tumor mutation burden between baseline and progression samples, there was no change in the median level.

[0164] When looking at blood copy number burden (bCNB), the median values ​​at baseline and progression were also similar. However, bCNB was able to track progression during treatment. Figure 28 shows a series of bCNB during treatment, showing a decrease in C1D15 and / or C2D1, followed by an increase in bCNB, which preceded imaging detection of progressive disease in 12 / 18 (66.7%) patients whose staging blood samples were analyzed. Thus, bCNB was able to track treatment and predict progression before any radiographic detection.

[0165] As demonstrated in the Examples, plasma WES is a highly sensitive and comprehensive NGS approach for detecting individual variants at baseline and during treatment, some of which are significantly enriched at progression. However, NGS assays designed around specific variants to monitor disease progression are expensive. Therefore, dynamic changes in CNVs during treatment can be detected before radiographic detection of recurrence using shallow LP-WGS assays. This approach constitutes a promising, cost-effective method for serially monitoring early signs of metastatic disease progression during treatment.

[0166] Example 8: Use of ctDNA in cerebrospinal fluid for cancer detection

[0167] Leptomeningeal metastasis occurs in more than 3% of patients diagnosed with non-small cell lung cancer (NSCLC) throughout the disease course, resulting in poor clinical outcomes and limiting treatment options. Cerebrospinal fluid (CSF) is a direct liquid biopsy for the pathological diagnosis of leptomeningeal metastasis. However, traditional clinical methods for detecting tumor cells in CSF have shown limited sensitivity. Meanwhile, the intrinsic genomic abnormalities of leptomeningeal metastases remain unknown. Here, we report a prospective clinical study aimed at identifying genomic abnormalities harbored by NSCLC patients with leptomeningeal metastases via circulating tumor DNA (ctDNA) in CSF.

[0168] Thirteen patients were enrolled in this study, and CSF samples were collected after the diagnosis of metastasis. PBMC samples were collected from 11 patients as germline controls. A low-pass whole-genome sequencing (LP-WGS) assay, Predicine CNB, was performed to identify copy number variations and tumor fractions in CSF samples from all 13 patients. Additionally, a boosted whole-exome sequencing assay, Predicine WES+, was performed on paired CSF and PBMC samples from 11 patients.

[0169] The ctDNA fraction was identified in all 13 CSF samples by the PredicineCNB™ assay (e.g., using the PredicineSCORE™ liquid biopsy assay (Predicine, Inc., Hayward, CA)). Gene copy variants related to NSCLC were also detected, including copy gains of EGFR (7 pts), BRAF (5 pts), MET (5 pts), KRAS (2 pts), ERBB2 (2 pts), ROS1 (2 pts), and ALK (1 pt), and copy losses of RB1 (4 pts), PTEN (2 pts), and TP53 (1 pt). Figure 29 shows the copy number gains and losses determined via PredicineCNB.

[0170] PredicineWES+™, a boosted whole-exome sequencing platform, identified 1,493 somatic variants in 11 CSF samples, 97 of which had previously been reported as likely pathogenic by public clinical databases. Regarding NSCLC-specific biomarkers, 7 of 11 patients carried EGFR variants, including 3 Exon19del events, 1 Exon20ins event, 1 L858R mutation, and other gain-of-function mutations. Figure 30 shows the mutation landscape determined via PredicineWES+.

[0171] Example 9: Analysis of TMB and MSI using liquid biopsies

[0172] Tumor mutation burden (TMB) and microsatellite instability (MSI) are emerging biomarkers that correlate with response to immunotherapy. Predicine ATLAS is a proprietary NGS-based assay that enables robust measurement of TMB and MSI in cell-free circulating DNA (cfDNA) extracted from blood samples. This report summarizes the analytical validation of the Predicine ATLAS assay, including its accuracy, specificity, limit of detection (LOD), and precision (repeatability and reproducibility).

[0173] The Predicine ATLAS assay uses a 600-gene panel designed to measure TMB and MSI in liquid biopsy samples from cancer patients. cfDNA was extracted and labeled with unique molecular barcodes during library construction, then enriched using the Predicine ATLAS panel and subsequently subjected to paired-end sequencing using the Illumina platform.

[0174] For blood samples, 10 mL of peripheral venous blood was collected in a Streck Cell-Free DNA BCT. Upon receipt, samples were immediately processed to plasma and stored at -80°C. cfDNA was extracted using the QIAamp Circulating Nucleic Acid Kit and quantified using Qubit. Cell line genomic DNA (gDNA) samples were enzymatically digested and sequentially size-selected to mimic the plasma cfDNA profile. Extracted cfDNA was labeled with unique molecular barcodes, and ligated sequencing libraries were PCR-amplified with high-fidelity polymerases and quantified using Bioanalyzer. For enrichment, sequencing libraries were blocked with adapter-specific blocking oligonucleotides and hybridized with the Predicine ATLAS panel. Bead-bound captured libraries were amplified and quantified using Bioanalyzer. Enriched libraries were normalized, pooled, and loaded onto an Illumina platform for 2X 150bp paired-end sequencing. Libraries were sequenced to a median depth of over 20,000X.

[0175] To achieve accurate and robust TMB estimation, only highly reliable somatic single nucleotide variant (SNV) mutations in the target coding region were taken into account in the TMB calculation.

[0176] NGS data were analyzed using Predicine's DeepSEA NGS analysis pipeline, which starts with raw sequencing data (BCL files) and outputs final variant calls. The pipeline first performs adapter trimming, barcode checking, and error correction. Pairs of clean FASTQ files were aligned to the human reference genome build hg19 using the BWA alignment tool. A consensus bam file was then obtained by merging single-stranded fragments and paired-end reads derived from the same molecule. Single-stranded fragments derived from the same double-stranded DNA molecule were further merged as double strands. Both sequencing and PCR errors were corrected during this process.

[0177] Following the DeepSEA variant caller, variants were filtered based on variant background from a pool of normal control samples and other historical samples. Other metrics, such as base quality, log odds ratio, and distance to fragment ends, were used to remove variants with low confidence. Detected calls were variants with at least four unique supporting fragments, one of which was double-stranded.

[0178] When a matched normal sample is not available, a germline variant filter can be implemented. In many applications, liquid biopsy assays may not include a matched normal. The germline variant filter can use the assumption that tumor-derived somatic mutations have a much lower variant allele frequency than heterozygous germline variants. This allows the assumption that variants with high allele frequencies are germline derived. The variant allele frequency is adjusted by the copy number change occurring at the variant position. Variants annotated in public germline databases with relatively high population allele frequencies are also excluded.

[0179] The TMB score was estimated based on the number of somatic SNVs in the coding region and normalized by the total coding region size covered by the panel. If the MSAF (maximum somatic allele frequency) was below a threshold, the TMB score was not estimated.

[0180] MSI scores were assessed from tumor samples by counting the number of unstable markers. For MSI detection, the Predicine ATLAS assay analyzes 50 MSI markers in a panel, which are short tandem repeat regions in the reference genome. A tumor sample is predicted as MSI-high (MSI-H) if its MSI score is greater than a threshold defined in a titration experiment (Figure 6). For each MSI marker, a z-score was calculated by comparing the repeat length distribution of the tumor sample with that of a normal background constructed from a batch of normal plasma samples. If a marker is considered unstable, a z-value threshold is estimated from validation data. PCR or sequencing noise is suppressed by error correction using the DeepSEA algorithm before constructing the repeat length distribution.

[0181] The Predicine ATLAS panel contains 600 cancer-associated genes covering a 2.4 Mb genomic region and 1.36 Mb coding region.

[0182] To evaluate the Predicine ATLAS panel for TMB measurement, assay-derived TMB scores were compared with those calculated from publicly available WES data. 7,116 tissue samples spanning over 30 tumor types from publicly available TCGA data downloaded from the Broad GDAC Firehose (https: / / gdac.broadinstitute.org / ) were used for analysis. This in silico analysis showed that the Predicine ATLAS panel-based TMB scores were highly correlated with WES-based TMB scores (R = 0.98, P < 0.001).

[0183] Predicine NGS Panel Performance Metrics

[0184] Based on a standard 15 ng DNA input, satisfactory assay performance for base substitution calling, indels, rearrangements, copy number gains (CNGs), and copy number losses (CNLs) is summarized in Table 14.

[0185] [Table 14]

[0186] Eight cell lines with available WES data from the COSMIC database (https: / / cancer.sanger.ac.uk / cosmic / ) were tested by the Predicine ATLAS assay for TMB measurement. TMB scores from the Predicine ATLAS assay were demonstrated to be highly correlated with TMB from published WES data (R = 0.97, P < 0.001).

[0187] Assay LoD

[0188] For TMB, the LoD is defined as the minimum tumor content required to obtain at least 90% concordance between the detected and predicted TMB status. Four cell lines with different TMB were titrated to five different levels of tumor content (20%, 10%, 1%, 0.5%, and 0.25%), with each level having a minimum of two replicates. The resulting diluted samples had tumor content ranging from 0.25% to 20%. All samples were fragmented, with sizes selected to mimic those of plasma cfDNA. Because these cancer cell lines have copy number alterations across the entire genome, the MAF (mutant allele frequency) ranged from 10% to 90%. To facilitate TMB LoD assessment, only cell line mutations with an MAF greater than 35% and no CNV regions were included in the TMB LoD assessment. A high correlation of TMB scores between the Predicine ATLAS panel and WES data from COSMIC was observed.

[0189] To assess the LoD, the agreement between the detected and predicted TMB status was evaluated for different TMB cutoffs (5, 10, 15, 20, and 30 Muts / Mb), respectively. The assay reached at least 90% agreement at an LoD of 1%.

[0190] For MSI, the LoD is defined as the lowest titration level at which at least 90% of MSI-high samples can be detected as MSI-high. To determine the LoD for MSI, MSI-H cell lines were serially diluted with target MSS cell lines at multiple titration levels (20%, 10%, 1%, 0.5%, and 0.25%). 100% of samples with less than 1% tumor cells were detected as MSI-H, and 90.9% of samples at a titration level of 1% were identified as MSI-H. Thus, the MSI assay achieved 90.9% concordance at tumor content as low as 1%.

[0191] Accuracy was established by comparing the calculated TMB score from a series of titrated cell lines with the expected TMB score at a TMB cutoff of 10 Muts / Mb. A total of 39 samples with ≥1% tumor content were used for TMB accuracy. All samples with high TMB (TMB score ≥10) were classified as high, and all samples with low TMB (TMB score <10) were classified as low. Since none of the samples were misclassified, the PPV of the Predicine ATLAS TMB assay is established as 100%.

[0192] For MSI, accuracy was established by comparing the calculated MSI status from a series of titrated cell lines with the expected MSI status. In total, 21 MSI samples and 9 MSS samples with tumor content ≥ 1% were used for MSI accuracy. The MSI accuracy was established as 96.7% (95% CI: 82.8–99.9%).

[0193] To evaluate the performance of the Predicine ATLAS assay and ensure that "blank" samples did not generate analytical signals, TMB and MSI scores of cfDNA from healthy donors were evaluated for specificity using the Predicine ATLAS assay. All samples from healthy donors had a TMB score of 0 (AF cutoff threshold = 0.5%) and were predicted as MSS (Figure 6). 100% (95% CI: [69.2-100%]) of the samples were classified as TMB low, with a cutoff of 10 Mut / Mb, and MSS, with an MSI cutoff of 14. These baseline data suggest that the Predicine ATLAS assay has 100% specificity for TMB and MSI measurements.

[0194] Repeatability (intra-assay precision)

[0195] To assess the closeness of agreement between replicates of the same sample under the same operating conditions, eight replicate groups with at least two replicates per group were each run under the same operating conditions for TMB and MSI analysis. High similarity was observed for all replicates.

[0196] To assess the closeness of agreement between assay results when operating conditions were varied, reproducibility was assessed and compared across different reagent lots, operators, and sequencers. A set of six samples with either high or low TMB was run under different conditions, and TMB scores were compared between replicates. A set of eight samples with MSI-H status was run under different conditions, and MSI status was compared between replicates. High agreement was observed across all replicates for both TMB and MSI testing. Cell lines at 1% titration were processed multiple times over a 7-month period using the Predicine ATLAS assay. Measured TMB scores (sorted by processing time) are shown in Figure 7. The difference between the measured and expected TMB scores (157.36) ranged from -5.6% to 6.5%, demonstrating the high reproducibility of the Predicine ATLAS assay.

[0197] The Predicine ATLAS analyzes 600 cancer-related genes in a single assay to provide assessment of immunotherapy biomarkers (TMB and MSI). The assay underwent rigorous analytical validation testing using a low cfDNA input of 15 ng. Validation studies demonstrated that the Predicine ATLAS assay had high concordance with WES for accurate assessment of TMB. Furthermore, assessment of MSI status showed high concordance with cell line samples. The development of the cfDNA-based Predicine ATLAS assay provides a complementary approach to tissue-based TMB and MSI assays for patients with solid tumors. The Predicine ATLAS is a robust, high-performance assay that enables simultaneous measurement of TMB and MSI in liquid biopsy samples.

[0198] Example 10: Detection of Hematological Cancer

[0199] Hematological cancers primarily begin in the bone marrow and account for approximately 10% of all newly diagnosed cancer cases. As an alternative to the invasive bone marrow biopsy currently used to monitor hematological cancers, comprehensive genomic profiling of cell-free circulating DNA (cfDNA) in blood or peripheral blood mononuclear cells (PBMCs) offers a minimally invasive and clinically convenient solution for detecting genomic biomarkers that guide clinical decisions in oncology.

[0200] PredicineHEME™ is a capture-based targeted next-generation sequencing (NGS) assay that enables accurate detection of genomic alterations, including small nucleotide variants (SNVs), insertions and deletions (indels), copy number variants (CNVs), and DNA rearrangements, in plasma cfDNA or genomic DNA (gDNA) extracted from PBMCs or bone marrow aspirates (BMA) of patients with hematological malignancies.

[0201] PredicineHEME™ is unique in its ability to detect variant allele frequencies down to 0.1% in plasma cfDNA or blood / BMA gDNA. This example presents an overview of the analytical validation of the 106-gene PredicineHEME™ assay, including accuracy, specificity, sensitivity, and precision. PredicineHEME™ analyzes 106 key blood cancer-associated genes for potential genomic biomarkers in blood cancer patients. The capture-based NGS assay is designed to detect SNVs, indels, copy number variations, and gene rearrangements in plasma, blood, or BMA samples collected from patients with blood cancer.

[0202] Extracted plasma cfDNA or enzymatically fragmented gDNA is labeled with unique molecular barcodes during library construction, subsequently enriched using the PredicineHEME™ panel, and then paired-end sequenced using the Illumina platform. If plasma isolation is required, peripheral venous blood is collected in a Streck Cell-Free DNA BCT; otherwise, EDTA tubes are used for whole blood sampling. Upon receipt, samples are deposited and immediately processed into plasma or buffy coat and stored at -80°C.

[0203] Plasma cfDNA is extracted using the QIAamp Circulating Nucleic Acid Kit, and PBMC or BMA genomic DNA (gDNA) samples are extracted using the DNeasy Blood & Tissue Kit. Genomic DNA is further enzymatically digested and purified, followed by DNA quantification and weighing. The extracted cfDNA or fragmented gDNA is labeled with unique molecular barcodes, and ligated sequencing libraries are PCR-amplified with high-fidelity polymerases and quantified by Bioanalyzer. For enrichment, the sequencing libraries are blocked with adapter-specific blocking oligonucleotides and hybridized with Predicine HEME™ panels. The bead-bound captured libraries are amplified and quantified by Bioanalyzer. The enriched libraries are normalized, pooled, and loaded onto an Illumina platform for 2X 150bp paired-end sequencing. Libraries are sequenced to a median depth of over 20,000X.

[0204] Combined with our in-house developed DeepSEA variant calling software for high accuracy in variant detection, PredicineHEME™ evaluates SNVs, indels, rearrangements, and CNVs in a single workflow. NGS data is analyzed using Predicine's proprietary DeepSEA NGS analysis pipeline, which starts with raw sequencing data (BCL files) and outputs final variant calls. The pipeline performs adapter trimming, barcode checking, and error correction. Clean FASTQ file pairs are then aligned to the human reference genome build hg19 using the BWA alignment tool. A consensus BAM file is then obtained by merging single-stranded fragments and paired-end reads derived from the same molecule. Single-stranded fragments derived from the same double-stranded DNA molecule are then further merged as double strands. Both sequencing and PCR errors are corrected during this process.

[0205] Following the DeepSEA variant caller, variants are filtered based on variant background from a pool of normal controls and other historical samples. Other metrics, such as base quality, log-odds ratio, entropy, and distance to fragment ends, are used to remove variants with low confidence.

[0206] When no matched normal sample is available, germline variant filter can be implemented.Variant allele frequency is adjusted by the copy number change when occurring at variant position.This assumes that the adjusted variant with an allele frequency close to 50% or more than 90% is derived from germline.Also filter out the variants that have relatively high population allele frequency and are annotated in public germline databases (for example, 1000 Genomes Project).

[0207] SNVs and indel variants are called after filtering out low-base-quality and low-mapping-score reads. Detected variants are further filtered based on variant background (defined by normal plasma samples and historical data), repetitive regions, and other quality metrics. Benign and likely benign SNPs are excluded from the variant call list. Copy number variations (CNVs) are estimated at the gene level. The pipeline first calculates on-target unique fragment coverage based on the consensus bam file and then adjusts for GC bias coverage. The unique fragment coverage profile from the normal sample group is used as a reference to normalize and estimate the z-score of copy number changes. DNA rearrangements are detected by identifying alignment breakpoints based on the bam file before the consensus operation. Suspicious alignments are filtered based on repetitive regions, local entropy calculations, and the similarity between the reference alignment and alternative alignments. DNA fusions are reported using two or more unique alignments.

[0208] To assess assay accuracy, reference samples with known genomic alterations in the assayed gene variants were tested, including cell lines, commercially available reference samples, and patient samples tested by a third-party laboratory. Up to 30 ng of cfDNA or fragmented gDNA input was analyzed by the Predicine HEME™ assay.

[0209] Concordance was determined by comparing the expected variants from the reference sample with the variants detected by the assay and was used as a performance indicator to demonstrate accuracy. In total, 102 mutations from 41 samples were predicted to be detected by this panel assay, and all mutations were confirmed with a 100% PPA (96.6-100%). In this validation study, the limit of detection (LoD) was defined as the lowest variant allele frequency at which the Predicine HEME™ assay could reliably detect the variant in 95% of replicates. Studies were conducted to demonstrate estimated sensitivity values ​​for each variant type, including SNVs, indels, DNA rearrangements, and CNGs. Reference cfDNA samples with MAFs defined at four different levels (0.1, 0.25, 0.375, and 0.5%) were evaluated to define the LOD for each variant type. Both PPA and PPV for each mutation level were calculated and used to define the LoD.

[0210] For SNVs, a total of 80 expected variants were evaluated with an expected AF of 0.25%, and the assay achieved 98.75% sensitivity (95% CI, 93.2–100%) with a PPV of 98.53% (95% CI, 91.4–99.7%). For indels, a total of 20 expected variants were evaluated with an expected AF of 0.375%, and the assay achieved 100% sensitivity (95% CI, 83.2–100%) with a PPV of 100% (95% CI, 83.2–100%). For DNA rearrangements, a total of 20 expected variants were evaluated with an expected AF of 0.375%, and the assay achieved 95% sensitivity (95% CI, 75.1–100%) with a PPV of 100% (95% CI, 82.4–99.9%). For CNG, a total of 20 predicted variants were evaluated at an expected 2.23 copies, and the assay achieved 100% sensitivity (95% CI, 83.2-100%) with 100% PPV (95% CI, 83.2-100%). Based on assay validation data, the LoD for SNVs was 0.25% MAF, the LoD for indels was 0.375% MAF, and the LoD for DNA rearrangements was 0.375% MAF. The LoD for CNG is 2.23 copies for the MYC gene.

[0211] To assess assay specificity and ensure that "blank" samples did not generate analytical signals, 24 samples were tested, including buffy coat, plasma, and BMA from healthy donors. Analytical specificity was estimated based on the number of false-positive mutations in the target panel. Analytical specificity = 100 * (1 - number of false-positives / panel size).

[0212] However, healthy donors may also carry pathogenic mutations with low variant frequencies, such as CHIP mutations. We consider all variants that pass the detection criteria of the NGS analysis pipeline to be false positives, including CHIP mutations. The analytical specificity is 99.9999%.

[0213] Accuracy was measured by the variation in estimated variant frequency between replicates. All samples in the precision study were evaluated throughout the entire workflow, from library to sequencing analysis. In total, 18 sets of identical sample aliquots were used in the precision study, and each set contained a reference standard with a MAF defined above the LoD and NTC (no template control).

[0214] Repeatability testing included intra-run performance (samples processed under the same conditions). Three replicates of each sample were performed under the same conditions, and the results were compared. Reproducibility was assessed based on samples processed independently by at least two operators on five days. 100% agreement between replicates was detected for both intra-run and inter-run precision testing.

[0215] 102 clinical gDNA samples from patients with hematological indications were analyzed using the Predicine HEME™ assay. A total of 714 genomic alterations, including SNVs, indels, and copy number variations, were detected in 98 of the 102 samples. Copy number variations were detected in 13 genes.

[0216] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. The present invention is not intended to be limited by the specific examples provided herein. While the present invention has been described with reference to the foregoing specification, the descriptions and illustrations of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will occur to those skilled in the art without departing from the invention. Furthermore, it should be understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative proportions set forth herein, which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the present invention described herein may be employed in practicing the invention. Accordingly, it is contemplated that the present invention also encompasses any such alternatives, modifications, variations, or equivalents. The following claims define the scope of the invention, and it is intended that methods and structures within the scope of these claims and their equivalents be covered thereby.

Claims

1. 1. A method for detecting the presence or absence of cancer in a subject, comprising: (a) assaying cell-free deoxyribonucleic acid (cfDNA) molecules from a biological sample obtained from or derived from the subject, wherein the biological sample comprises a urine sample; (b) detecting a set of biomarkers from the cfDNA molecules, wherein the set of biomarkers comprises differentially expressed markers or variants; (c) computerizing the set of biomarkers to detect the presence or absence of the cancer in the subject; A method comprising:

2. 10. The method of claim 1, wherein the biological sample is obtained from or derived from the subject using ethylenediaminetetraacetic acid (EDTA) collection tubes, cell-free RNA collection tubes, or cell-free deoxyribonucleic acid (DNA) collection tubes, other blood collection tubes, and CTC collection tubes.

3. 3. The method of any one of claims 1-2, wherein (a) comprises subjecting the biological sample to conditions sufficient to isolate, enrich, or extract the cfDNA molecules.

4. 4. The method of claim 1, wherein at least one of the cfDNA molecules is assayed using nucleic acid sequencing to generate nucleic acid sequencing reads.

5. 5. The method of claim 4, further comprising filtering at least a subset of the nucleic acid sequencing reads based on a quality score.

6. 6. The method of Claim 4 or 5, further comprising performing error correction on the nucleic acid sequencing reads using a sample barcode or molecular barcode attached to at least one of the cfDNA molecules.

7. 7. The method of claim 4, further comprising performing at least one of single-strand consensus calling and double-strand consensus calling on the nucleic acid sequencing reads, thereby suppressing sequencing and PCR errors in the nucleic acid sequencing reads.

8. 8. The method of any one of claims 4 to 7, wherein the cfDNA molecules are assayed using DNA sequencing.

9. 9. The method of claim 8, wherein the DNA sequencing is selected from the group consisting of next-generation sequencing, whole genome sequencing, low-pass sequencing, targeted sequencing, whole exome sequencing, methylation-aware sequencing, bisulfite sequencing, and combinations thereof.

10. The method of claim 8 , wherein the DNA sequencing comprises targeted sequencing.

11. The method of claim 4 , wherein the nucleic acid sequencing comprises nucleic acid amplification.

12. 12. The method of claim 11, wherein the nucleic acid amplification comprises polymerase chain reaction (PCR) or isothermal amplification.

13. 13. The method of any one of claims 1 to 12, wherein at least one of the cfDNA molecules is assayed using a polymerase chain reaction (PCR) assay, a microarray, or isothermal amplification.

14. 14. The method of any one of claims 1 to 13, wherein the cancer is selected from the group consisting of breast cancer, lung cancer, prostate cancer, colon cancer, melanoma, bladder cancer, non-Hodgkin's lymphoma, kidney cancer, endometrial cancer, leukemia, pancreatic cancer, thyroid cancer, and liver cancer, and any combination thereof.

15. 15. The method of claim 14, wherein the cancer comprises kidney cancer.

16. 15. The method of claim 14, wherein the cancer comprises bladder cancer.

17. 17. The method of claim 16, wherein the bladder cancer comprises non-muscle invasive bladder cancer.

18. The method of any one of claims 1 to 17, wherein the subject is asymptomatic for the cancer.

19. 19. The method of any one of claims 1-18, wherein (b) comprises detecting the presence or absence of cancer in the subject with an accuracy of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.

20. 20. The method of any one of claims 1-19, wherein (b) comprises detecting the presence or absence of cancer in the subject with a sensitivity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.

21. 21. The method of any one of claims 1-20, wherein (b) comprises detecting the presence or absence of cancer in the subject with a specificity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.

22. 22. The method of any one of claims 1-21, wherein (b) comprises detecting the presence or absence of cancer in the subject with a positive predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.

23. 23. The method of any one of claims 1-22, wherein (b) comprises detecting the presence or absence of cancer in the subject with a negative predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.

24. 24. The method of any one of claims 1 to 23, wherein the biological sample is obtained or derived from the subject before the subject receives treatment for the cancer.

25. The method of any one of claims 1 to 23, wherein the biological sample is obtained or derived from the subject during treatment for the cancer.

26. The method of any one of claims 1 to 23, wherein the biological sample is obtained or derived from the subject after undergoing treatment for the cancer.

27. 27. The method of any one of claims 24 to 26, wherein the treatment is selected from the group consisting of surgical resection, chemotherapy, radiation therapy, immunotherapy, cell therapy, adjuvant therapy, neoadjuvant therapy, androgen deprivation therapy, and combinations thereof.

28. 28. The method of any one of claims 1 to 27, further comprising identifying a clinical intervention for said subject based at least in part on said presence or absence of said detected cancer.

29. 29. The method of claim 28, wherein the clinical intervention is selected from a plurality of clinical interventions.

30. 29. The method of claim 28, wherein the clinical intervention is selected from the group consisting of surgical resection, chemotherapy, radiation therapy, immunotherapy, adjuvant therapy, neoadjuvant therapy, androgen deprivation therapy, and combinations thereof.

31. 30. The method of claim 28, further comprising administering the clinical intervention to the subject.

32. 32. The method of any one of claims 1 to 31, wherein the set of biomarkers comprises quantitative measurements of a set of genomic loci associated with cancer.

33. 33. The method of claim 32, wherein the set of cancer-associated genomic loci comprises one or more members selected from the group consisting of the genes listed in Table 1.

34. 34. The method of claim 33, wherein the set of cancer-associated genomic loci comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 members selected from the group consisting of the genes listed in Table 1.

35. 33. The method of claim 32, wherein the set of cancer-associated genomic loci comprises PTEN, TP53, or RB1.

36. 33. The method of claim 32, wherein the set of cancer-associated genomic loci comprises PTEN.

37. 33. The method of claim 32, wherein the set of cancer-associated genomic loci comprises FGFR3 or ERBB2.

38. 33. The method of Claim 32, wherein the set of cancer-associated genomic loci comprises one or more members selected from the group consisting of the genes listed in Table 2.

39. 39. The method of claim 38, wherein the set of cancer-associated genomic loci comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 members selected from the group consisting of the genes listed in Table 2.

40. 40. The method of any one of claims 32-39, further comprising using probes configured to selectively enrich the biological sample in nucleic acid molecules corresponding to a set of loci.

41. 41. The method of claim 40, wherein the probe comprises a nucleic acid primer.

42. 41. The method of claim 40, wherein the probe comprises a nucleic acid capture probe.

43. 43. The method of any one of claims 40 to 42, wherein the probes have sequence complementarity to at least a portion of the nucleic acid sequences of the set of loci.

44. 43. The method of any one of claims 40 to 42, wherein the probe has sequence complementarity to at least a portion of a nucleic acid sequence of a gene selected from the genes listed in Tables 1 and 2.

45. 44. The method of any one of claims 40-43, wherein the probes comprise at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 different probes.

46. 46. ​​The method of any one of claims 1 to 45, further comprising determining the likelihood of determining the presence or absence of cancer in the subject.

47. 47. The method of any one of claims 1-46, further comprising monitoring the subject for the presence or absence of cancer, comprising assessing the subject for the presence or absence of cancer at each of a plurality of time points.

48. 48. The method of claim 47, wherein a difference in the assessment of the presence or absence of cancer in the subject between the multiple time points is indicative of one or more clinical indications selected from the group consisting of: (i) a diagnosis of the cancer, (ii) a prognosis of the cancer, and (iii) the effectiveness or ineffectiveness of a course of treatment for treating the cancer in the subject.

49. 49. The method of claim 48, wherein the prognosis comprises expected progression-free survival (PFS) or overall survival (OS).

50. 50. The method of any one of claims 1-49, wherein the set of biomarkers from the cfDNA molecule comprises tumor-associated alterations selected from the group consisting of copy number alterations (CNAs), copy number losses (CNLs), single nucleotide variants (SNVs), insertions or deletions (indels), and rearrangements.

51. 51. The method of any one of claims 1 to 50, wherein the set of biomarkers from the cfDNA molecule comprises copy number variation.

52. 51. The method of any one of claims 1-50, wherein the set of biomarkers from the cfDNA molecule comprises a copy number reduction.

53. 51. The method of any one of claims 1-50, wherein the set of biomarkers from the cfDNA molecule comprises single base variants.

54. 54. The method of any one of claims 1 to 53, further comprising determining mutant allele frequencies of a set of somatic mutations among said set of biomarkers.

55. 55. The method of any one of claims 1 to 54, further comprising determining the blood copy number burden based on copy number alterations or copy number losses of said set of biomarkers.

56. 55. The method of claim 54, further comprising determining a circulating tumor DNA (ctDNA) fraction of said cancer in said subject based at least in part on the set of variant allele frequencies.

57. 57. The method of any one of claims 54-56, further comprising determining a tumor mutational burden (TMB) of said cancer in said subject based at least in part on the set of variant allele frequencies.

58. 58. The method of any one of claims 54-57, further comprising determining the tumor mutational burden (TMB) of said cancer in said subject based at least in part on a set of variant allele frequencies comprising microsatellites.

59. 59. The method of any one of claims 54-58, further comprising determining an aberration score for said cancer in said subject based at least in part on the set of variant allele frequencies.

60. 1. A method for detecting the presence or absence of cancer in a subject, comprising: (a) providing cell-free deoxyribonucleic acid (cfDNA) molecules from a biological sample obtained from or derived from the subject, wherein the biological sample comprises a urine sample; (b) hybridizing the cfDNA molecules or derivatives thereof to a plurality of nucleic acid capture probes to generate a plurality of enriched cfDNA molecules; (c) sequencing the nucleic acids of the plurality of enriched cfDNA molecules to generate sequencing data; (d) computationally processing the sequencing data to detect a set of biomarkers; (e) detecting the presence or absence of the cancer in the subject based at least on the presence of the set of biomarkers; A method comprising:

61. 1. A method for detecting the presence or absence of cancer in a subject, comprising: (a) providing cell-free deoxyribonucleic acid (cfDNA) from a biological sample obtained from or derived from the subject; (b) performing a first sequencing assay on the cfDNA or derivative thereof to generate copy number data for at least one region of the genome of a subject, wherein the sequencing is performed to a depth of 10x or less; (c) performing a second sequencing assay on the cfDNA molecule or derivative thereof, wherein the second sequencing assay comprises a whole-exome sequencing assay or a methylation-aware sequencing assay to generate sequencing data; (d) computationally processing the copy number data and the sequencing data to detect a set of biomarkers; (e) detecting the presence or absence of the cancer in the subject based at least on the presence of the set of biomarkers; A method comprising:

62. 62. The method of claim 61, wherein the biological sample is obtained from or derived from the subject using ethylenediaminetetraacetic acid (EDTA) collection tubes, cell-free RNA collection tubes, or cell-free deoxyribonucleic acid (DNA) collection tubes, other blood collection tubes, and CTC collection tubes.

63. 63. The method of any one of claims 61-62, wherein (a) comprises subjecting the biological sample to conditions sufficient to isolate, enrich, or extract the cfDNA molecules.

64. 64. The method of any one of claims 61 to 63, wherein the biological sample comprises a urine, blood, or cerebrospinal sample.

65. 65. The method of Claim 64, further comprising filtering at least a subset of the nucleic acid sequencing reads based on a quality score.

66. 66. The method of Claim 64 or 65, further comprising performing error correction on nucleic acid sequencing reads using a sample barcode or molecular barcode attached to at least one of the cfDNA molecules.

67. 67. The method of any one of claims 64 to 66, further comprising performing at least one of single-stranded consensus calling and double-stranded consensus calling on nucleic acid sequencing reads, thereby suppressing sequencing and PCR errors in said nucleic acid sequencing reads.

68. 68. The method of any one of claims 61 to 67, wherein (b) or (c) comprises nucleic acid amplification.

69. 69. The method of claim 68, wherein the nucleic acid amplification comprises polymerase chain reaction (PCR) or isothermal amplification.

70. 70. The method of any one of claims 61-69, wherein at least one of the cfDNA molecules is assayed using a polymerase chain reaction (PCR) assay, a microarray, or isothermal amplification.

71. 71. The method of any one of claims 61 to 70, wherein the cancer is selected from the group consisting of breast cancer, lung cancer, prostate cancer, colon cancer, melanoma, bladder cancer, non-Hodgkin's lymphoma, kidney cancer, endometrial cancer, leukemia, pancreatic cancer, thyroid cancer, and liver cancer, and any combination thereof.

72. 72. The method of claim 71, wherein the cancer comprises bladder cancer.

73. 73. The method of claim 72, wherein the bladder cancer comprises non-muscle invasive bladder cancer.

74. 74. The method of any one of claims 61 to 73, wherein the subject is asymptomatic for the cancer.

75. 75. The method of any one of claims 61-74, wherein (d) comprises detecting the presence or absence of cancer in the subject with an accuracy of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.

76. 76. The method of any one of claims 61-75, wherein (d) comprises detecting the presence or absence of cancer in the subject with a sensitivity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.

77. 77. The method of any one of claims 61-76, wherein (d) comprises detecting the presence or absence of cancer in the subject with a specificity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.

78. 78. The method of any one of claims 61-77, wherein (d) comprises detecting the presence or absence of cancer in the subject with a positive predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.

79. 79. The method of any one of claims 61-78, wherein (d) comprises detecting the presence or absence of cancer in the subject with a negative predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.

80. 80. The method of any one of claims 61 to 79, wherein the biological sample is obtained or derived from the subject before the subject receives treatment for the cancer.

81. 80. The method of any one of claims 61 to 79, wherein the biological sample is obtained or derived from the subject during treatment for the cancer.

82. 80. The method of any one of claims 61 to 79, wherein the biological sample is obtained or derived from the subject after undergoing treatment for the cancer.

83. 83. The method of any one of claims 80-82, wherein the treatment is selected from the group consisting of surgical resection, chemotherapy, radiation therapy, immunotherapy, cell therapy, adjuvant therapy, neoadjuvant therapy, androgen deprivation therapy, and combinations thereof.

84. 84. The method of any one of claims 61-83, further comprising identifying a clinical intervention for said subject based at least in part on said presence or absence of said detected cancer.

85. 85. The method of claim 84, wherein the clinical intervention is selected from a plurality of clinical interventions.

86. 85. The method of claim 84, wherein the clinical intervention is selected from the group consisting of surgical resection, chemotherapy, radiation therapy, immunotherapy, adjuvant therapy, neoadjuvant therapy, androgen deprivation therapy, and combinations thereof.

87. 85. The method of claim 84, further comprising administering the clinical intervention to the subject.

88. 88. The method of any one of claims 61-87, wherein the set of biomarkers comprises quantitative measurements of a set of genomic loci associated with cancer.

89. 89. The method of Claim 88, wherein said set of cancer-associated genomic loci comprises one or more members selected from the group consisting of the genes listed in Table 1.

90. 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 members selected from the group consisting of the genes listed in Table 1.

91. 89. The method of claim 88, wherein the set of cancer-associated genomic loci comprises PTEN, TP53, or RB1.

92. 89. The method of claim 88, wherein the set of cancer-associated genomic loci comprises PTEN.

93. 89. The method of claim 88, wherein the set of cancer-associated genomic loci comprises FGFR3 or ERBB2.

94. 89. The method of Claim 88, wherein said set of cancer-associated genomic loci comprises one or more members selected from the group consisting of the genes listed in Table 2.

95. 89. The method of claim 88, wherein the set of cancer-associated genomic loci comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 members selected from the group consisting of the genes listed in Table 2.

96. 89. The method of Claim 88, wherein said set of cancer-associated genomic loci comprises one or more members selected from the group consisting of the genes listed in Table 3.

97. 89. The method of claim 88, wherein the set of cancer-associated genomic loci comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 members selected from the group consisting of the genes listed in Table 3.

98. 98. The method of any one of claims 88-97, further comprising using probes configured to selectively enrich the biological sample in nucleic acid molecules corresponding to a set of loci.

99. 99. The method of claim 98, wherein the probe comprises a nucleic acid primer.

100. 99. The method of claim 98, wherein the probe comprises a nucleic acid capture probe.

101. 101. The method of any one of claims 98 to 100, wherein the probes have sequence complementarity to at least a portion of the nucleic acid sequences of the set of loci.

102. 101. The method of any one of claims 98-100, wherein the probe has sequence complementarity to at least a portion of a nucleic acid sequence of a gene selected from the genes listed in Table 1, Table 2, or Table 3.

103. 103. The method of any one of claims 98-102, wherein the probes comprise at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 different probes.

104. 104. The method of any one of claims 61 to 103, further comprising determining the likelihood of determining the presence or absence of cancer in the subject.

105. 105. The method of any one of claims 61-104, further comprising monitoring the subject for the presence or absence of cancer, comprising assessing the subject for the presence or absence of cancer at each of a plurality of time points.

106. 106. The method of claim 105, wherein a difference in the assessment of the presence or absence of cancer in the subject between the multiple time points indicates one or more clinical indications selected from the group consisting of: (i) a diagnosis of the cancer, (ii) a prognosis of the cancer, and (iii) the effectiveness or ineffectiveness of a course of treatment for the treatment of the cancer in the subject.

107. 107. The method of claim 106, wherein the prognosis comprises expected progression-free survival (PFS) or overall survival (OS).

108. 108. The method of any one of claims 61-107, wherein the set of biomarkers from the cfDNA molecule comprises tumor-associated alterations selected from the group consisting of copy number alterations (CNAs), copy number losses (CNLs), single nucleotide variants (SNVs), insertions or deletions (indels), and rearrangements.

109. 109. The method of any one of claims 61-108, wherein the set of biomarkers from the cfDNA molecule comprises copy number variation.

110. 109. The method of any one of claims 61-108, wherein the set of biomarkers from the cfDNA molecule comprises a copy number reduction.

111. 109. The method of any one of claims 61-108, wherein the set of biomarkers from the cfDNA molecule comprises single base variants.

112. 112. The method of any one of claims 61 to 111, further comprising determining mutant allele frequencies of a set of somatic mutations among said set of biomarkers.

113. 113. The method of any one of claims 61 to 112, further comprising determining the blood copy number burden based on copy number alterations or copy number losses of said set of biomarkers.

114. 113. The method of claim 112, further comprising determining a circulating tumor DNA (ctDNA) fraction of said cancer in said subject based at least in part on the set of variant allele frequencies.

115. 115. The method of any one of claims 112-114, further comprising determining a tumor mutational burden (TMB) of said cancer in said subject based at least in part on the set of variant allele frequencies.

116. 116. The method of any one of claims 112-115, further comprising determining the tumor mutational burden (TMB) of said cancer in said subject based at least in part on a set of variant allele frequencies comprising microsatellites.

117. 117. The method of any one of claims 112-116, further comprising determining an aberration score for said cancer in said subject based at least in part on the set of variant allele frequencies.