Methods and systems for porting tissue-based classifiers into liquid biopsy samples

A computer-implemented method using integrative correlation coefficients and linear discriminant analysis addresses the limitations of genomic alteration testing and cfDNA analysis by identifying gene expression signatures and classifiers, improving cancer subtype classification and therapeutic response prediction.

WO2025250678A1PCT designated stage Publication Date: 2025-12-04GENECENTRIC THERAPEUTICS INC +5
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/031248
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-28
Filing Date
2025-05-28
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Current genomic alteration testing and cfDNA analysis in liquid biopsies are limited by low signal-to-noise ratios, contamination from normal tissues, low ctDNA frequency, and the challenge of differentiating tumor-specific features, which hinder accurate cancer diagnosis and treatment decision-making.

Method used

A computer-implemented method using integrative correlation coefficients and linear discriminant analysis to identify gene expression signatures and generate classifiers from tissue and bodily fluid samples, enhancing the analysis of ctDNA to improve cancer subtype classification and therapeutic response prediction.

Benefits of technology

The method provides accurate gene expression signatures and classifiers that enhance cancer subtype identification and therapeutic response prediction, addressing the limitations of existing genomic alteration testing and cfDNA analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025031248_04122025_PF_FP_ABST
    Figure US2025031248_04122025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure discloses methods for assigning gene expression signature labels to samples with liquid circulating tumor DNA (ctDNA) fragmentomics profiles and using said profiles to uncover genes whose expression patterns are reproducible between tissue samples and bodily fluid samples comprising ctDNA. Also provided herein are methods and systems for generating classifiers for a desired feature of a cancer of interest in a subject suffering from said cancer of interest.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No. GNCN-026 / 01WO 320289-2157 METHODS AND SYSTEMS FOR PORTING TISSUE-BASED CLASSIFIERS INTO LIQUID BIOPSY SAMPLES CROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority from U.S. Provisional Application No. 63 / 652,432 filed May 28, 2024, which is incorporated by reference herein in its entirety for all purposes. FIELD

[0002] The present disclosure is directed to methods for assigning gene expression signature labels to samples with liquid circulating tumor DNA (ctDNA) fragmentomics profiles. In some embodiments, the methods presented herein utilize the integrative correlation coefficient (ICC) to uncover genes whose expression patterns are reproducible between tissue samples and bodily fluid samples comprising ctDNA and uses linear discriminant analysis (LDA) methods (e.g., Classification to Nearest Centroids (ClaNC)) on said reproducible genes to generate classifiers for desired features of a cancer of interest in a subject suffering from said cancer of interest. BACKGROUND

[0003] Accurate cancer diagnosis and subtype classification are critical for guiding clinical care and precision oncology. Moreover, tumor subtypes are often characterized by distinct transcriptional regulation, which can change during treatment resistance, leading to different clinical tumor phenotypes. Therefore, accurate subtype classification and identification of transcriptional patterns underlying emergent clinical phenotype during therapy has critical implications for studying mechanisms of resistance and informing treatment decisions.

[0004] Genomic alteration testing is a mature technology and its utility and limitations in diagnostic testing and precision medicine are understood. For example, genomic alteration testing can identify pathogenic or actionable alterations as well as screening and early detection, monitoring for minimal residual disease (MRD) and therapeutic response, detecting and classifying emerging resistance, and monitoring for progressive disease. However, a chief limitation of genomic alteration testing is its silence on the phenotype or clinical impact of pathogenic / actionable alterations, which is important because many of the most effective targeted and immuno-oncology therapies have overall response rates of 50% or less. Knowing the likelihood of therapeutic failure, effectively monitoring therapeutic response, and spotting theAttorney Docket No. GNCN-026 / 01WO 320289-2157 first signs of emerging resistance can be critical for healthcare providers and patients to make timely, informed treatment decisions. Penetrance of pathogenic / actionable alterations, gene or pathway activity, and gene expression and epigenetic regulation are facets of phenotype not addressed by current DNA sequencing based genomic alteration testing. Accordingly, there is a need for additional technologies and biomarkers to complement the deficits of alteration testing.

[0005] Liquid biopsies are clinical tests that analyze a patient’s blood, urine, or other bodily fluid to detect cancer cells or tumor-related molecules. Due to its non-invasive quality, liquid biopsy has become an increasingly attractive diagnostic tool in oncology. The non-invasive assay permits serial sample collection to discover a range of information and monitor the evolution of many types of cancers.

[0006] Most cancers release a milieu of liquid biopsy analytes, such as tumor cells, proteins, extracellular vesicles, and / or cell free nucleic acids (cfNAs). Arguably the most clinically actionable information to date is derived from analyzing cell free DNA (cfDNA) comprised of DNA from healthy, and tumor cells alike to capture genetic and epigenetic differences within circulating tumor DNA (ctDNA) shed specifically from tumor cells. Current applications of cfDNA analysis support the identification of personalized treatment strategies by assessing commonly mutated oncogenes such as EGFR, BRCA1, and TP53 mutations as well as offer insight into copy number alterations which could inform treatment strategies and prognosis. Additionally, some of the current research and clinical efforts have focused on the detection of genetic alterations in ctDNA and how said alterations from ctDNA can be used to help distinguish molecular subsets of tumors (see Wyatt, A. W. et al. Concordance of circulating tumor DNA and matched metastatic tissue biopsy in prostate cancer. J. Natl Cancer Inst.110, 78–86 (2018) and Viswanathan, S. R. et al. Structural alterations driving castration-resistant prostate cancer revealed by linked-read genome sequencing. Cell 174, 433–447.e19 (2018)). So while cfDNA counterparts for DNA alteration test are established, these cfDNA counterparts are devoid of gene expression data.

[0007] Fundamental challenges associated with cfDNA analysis remain, which greatly limit its utility in research and clinical settings. These fundamental challenges can include the low signal- to-noise ratio due to, in part, the inherent contamination of cfDNA shed from normal tissue and circulating immune cells coupled with the relative low frequency of ctDNA presence and, like genomic alteration testing in general, the finding that the presence of a mutation may notAttorney Docket No. GNCN-026 / 01WO 320289-2157 associate with targeted treatment response or penetrance of that mutation. Additionally, even when ctDNA is present in the liquid biopsy sample, differentiating features unique to the tumor phenotype represent the proverbial needle in the haystack obstacle that requires extensive analysis to identify.

[0008] The methods, kits and systems provided herein address the limitations and challenges of genomic alteration testing as well as the analysis of cfDNA. SUMMARY

[0009] In one aspect, provided herein is a computer-implemented method for determining a gene expression signature in disparate samples from subjects suffering from a cancer of interest, the method comprising: (a) receiving in a computer system, a first set of nucleic acid expression data for each of a plurality of tissue samples obtained from a first population of subjects suffering from or suspected of suffering from the cancer of interest and a second set of nucleic acid expression data for each of a plurality of bodily fluid samples obtained from a second population of subjects suffering from or suspected of suffering from the cancer of interest, wherein each of the plurality of bodily fluid samples comprises cell-free DNA (cfDNA); (b) determining by the computer system, dependence relationships between each gene within the first set of nucleic acid expression data to generate a first set of dependence relationships and dependence relationships between each gene within the second set of nucleic acid expression data to generate a second set of dependence relationships; and (c) selecting by the computer system, each gene for which the dependence relationships for a respective gene from the first set of dependence relationships is substantially similar to the dependence relationships for the respective gene in the second set of dependence relationships, thereby generating a gene signature that comprises each gene selected by the computer system to possess substantially similar dependence relationships between the first set of nucleic acid expression data for the plurality of tissue samples and the second set of nucleic acid expression data for the plurality of bodily fluid samples. In some cases, the dependence relationships between each gene within the first set of nucleic acid expression data comprises all possible pairwise correlations between each gene within the first set of nucleic acid expression data and each other gene within the first set of nucleic acid expression data and the dependence relationships between each gene within the second set of nucleic acid expression data comprises all possible pairwise correlations between each gene within the second set of nucleic acidAttorney Docket No. GNCN-026 / 01WO 320289-2157 expression data and each other gene within the second set of nucleic acid expression data. In some cases, the substantial similarity in (c) is evidenced by an integrative correlation coefficient (ICC) for the respective gene that is above a desired threshold, wherein the ICC is a correlation of the pairwise correlations determined for the respective gene within the first set of nucleic acid expression data and the pairwise correlations determined for the respective gene within the second set of nucleic acid expression data. In some cases, the desired threshold is the 99thpercentile of a null distribution of integrative correlations. In some cases, prior to (a), the second set of nucleic acid expression data is subjected to a method comprising (a) extracting gene expression information across the dataset by (i) mapping the sequence reads from the dataset to a reference human genome, thereby generating read count data across the mapped genome; and (ii) applying fast Fourier transformation (FFT) to the read count data in nucleosome occupancy windows across the mapped genome to determine an FFT signal at each nucleosome occupancy window, wherein an increased FFT signal is indicative of nucleosomal depletion and a decreased FFT signal is indicative of nucleosomal presence; and (b) performing quality control of the cfDNA dataset comprising removal of samples from the dataset that possess a low or inconsistent FFT signal as determined in (b)(i)-(b)(ii) and / or a measured circulating tumor DNA (ctDNA) content below a dataset-specific threshold, thereby generating a quality-controlled cfDNA dataset comprised of features that are reflective of gene activity or expression.

[0010] In another aspect, provided herein is a computer-implemented method for generating a classifier for a desired feature of a cancer of interest in a subject suffering from the cancer of interest or suspected of suffering from the cancer of interest, the method comprising: (a) receiving in a computer system, a first set of nucleic acid expression data for each of a plurality of tissue samples obtained from a first population of subjects suffering from or suspected of suffering from the cancer of interest and a second set of nucleic acid expression data for each of a plurality of bodily fluid samples obtained from a second population of subjects suffering from or suspected of suffering from the cancer of interest, wherein each of the plurality of bodily fluid samples comprises cell-free DNA (cfDNA); (b) determining by the computer system, dependence relationships between each gene within the first set of nucleic acid expression data to generate a first set of dependence relationships and dependence relationships between each gene within the second set of nucleic acid expression data to generate a second set of dependence relationships; (c) selecting by the computer system, each gene for which the dependence relationships for aAttorney Docket No. GNCN-026 / 01WO 320289-2157 respective gene from the first set of dependence relationships is substantially similar to the dependence relationships for the respective gene in the second set of dependence relationships, thereby generating a gene signature that comprises each gene selected by the computer system to possess substantially similar dependence relationships between the first set of nucleic acid expression data for the plurality of tissue samples and the second set of nucleic acid expression data for the plurality of bodily fluid samples; (d) inputting nucleic acid expression data for each gene in the gene signature from (c) from at least two training sets, wherein one of the at least two training sets comprises nucleic acid expression data for each gene in the gene signature from (c) from each of a plurality of tissue samples from subjects that are indicative of the presence a desired feature for the cancer of interest, while another of the at least two training sets comprises nucleic acid expression data for each gene in the gene signature from (c) from each of a plurality of tissue samples from subjects that are indicative of the absence of the desired feature for the cancer of interest; and (e) conducting on the computer system a linear discriminate analysis (LDA) comprising feature selection to generate a classifier for the desired feature of the cancer of interest, wherein the feature selection comprises: (i) calculating a test statistic (t-statistic) for each gene in the gene signature from (c) from the at least two training sets, wherein the t-statistic for each gene indicates each gene's ability to distinguish between the presence or the absence of the desired feature; (ii) ranking each gene based on each gene’s t-statistic; and (iii) selecting each gene whose t-statistic is above a desired threshold for distinguishing between the presence or the absence of the desired feature, thereby generating a classifier for the desired feature comprise each of the selected genes. In some cases, the method, further comprises (f) classifying on the computer system one or more test samples obtained from an independent population of subjects suffering the cancer of interest as possessing or not possessing the desired feature using the classifier from (e)(iii) on the one or more test samples; (g) comparing on the computer system, the classification of the one or more test samples to classification of the one or more samples for the desired feature as determined using a control classifier of the desired feature for the cancer of interest; and (h) validating on the computer system the classifier from (e)(iii) if the comparing indicates that the classifications of the one or more test samples is substantially similar to classification of the one or more test samples determined using the control classifier. In some cases, the dependence relationships between each gene within the first set of nucleic acid expression data comprises all possible pairwise correlations between each gene within the first set of nucleic acid expressionAttorney Docket No. GNCN-026 / 01WO 320289-2157 data and each other gene within the first set of nucleic acid expression data and the dependence relationships between each gene within the second set of nucleic acid expression data comprises all possible pairwise correlations between each gene within the second set of nucleic acid expression data and each other gene within the second set of nucleic acid expression data. In some cases, the substantial similarity in (c) is evidenced by an integrative correlation coefficient (ICC) for the respective gene that is above a desired threshold, wherein the ICC is a correlation of the pairwise correlations determined for the respective gene within the first set of nucleic acid expression data and the pairwise correlations determined for the respective gene within the second set of nucleic acid expression data. In some cases, the desired threshold is the 99thpercentile of a null distribution of integrative correlations. In some cases, prior to (a), the second set of nucleic acid expression data is subjected to a method comprising (a) extracting gene expression information across the dataset by (i) mapping the sequence reads from the dataset to a reference human genome, thereby generating read count data across the mapped genome; and (ii) applying fast Fourier transformation (FFT) to the read count data in nucleosome occupancy windows across the mapped genome to determine an FFT signal at each nucleosome occupancy window, wherein an increased FFT signal is indicative of nucleosomal depletion and a decreased FFT signal is indicative of nucleosomal presence; and (b) performing quality control of the cfDNA dataset comprising removal of samples from the dataset that possess a low or inconsistent FFT signal as determined in (b)(i)-(b)(ii) and / or a measured circulating tumor DNA (ctDNA) content below a dataset-specific threshold, thereby generating a quality-controlled cfDNA dataset comprised of features that are reflective of gene activity or expression. In some cases, the desired feature is selected from the group consisting of a subtype of the cancer of interest, microsatellite instability, mismatch repair proficiency, minimal residual disease, prognostic risk, response to a therapeutic agent, development of resistance to a therapeutic agent. In some cases, the cancer of interest is COAD, the desired feature is microsatellite instability and the classifier generated comprises, consists essentially of or consists of the genes in Table 2. In some cases, the cancer of interest is PAAD, the desired feature is a PAAD subtype, and the classifier generated comprises, consists essentially of or consists of the genes in Table 4. In some cases, the subtype is basal or classical. In some cases, the cancer of interest is BLCA, the desired feature is an FGFR activation signature, and the classifier generated comprises, consists essentially of or consists of the genes in Table 6. In some cases, the FGFR activation signature is a FGFR3 activation signature.Attorney Docket No. GNCN-026 / 01WO 320289-2157

[0011] In another aspect, provided herein is a computer-implemented method for determining a gene expression signature in disparate samples from subjects suffering from a cancer of interest, the method comprising: (a) receiving in a computer system, a first set of nucleic acid expression data for each of a plurality of tissue samples obtained from a first population of subjects suffering from or suspected of suffering from the cancer of interest and a second set of nucleic acid expression data for each of a plurality of bodily fluid samples obtained from a second population of subjects suffering from or suspected of suffering from the cancer of interest, wherein each of the plurality of bodily fluid samples comprises cell-free DNA (cfDNA); (b) determining by the computer system, an integrative correlation coefficient for each gene from the first and second set of nucleic acid expression data, wherein the integrative correlation coefficient is a numerical representation of how each gene from the first set of nucleic acid expression data correlate with each other compared to the rank of the same gene from the second set of nucleic acid expression data; (c) ranking by the computer system, the integrative correlation coefficient of each gene from step (b); and (d) selecting, by the computer system, genes that have an integrative correlation coefficient above a desired threshold to generate a gene signature, wherein the gene signature comprises genes whose expression patterns are substantially similar between the tissue sample and the bodily fluid sample. In some cases, the ICC is above the desired threshold of the 99thpercentile of a null distribution of integrative correlations. In some cases, prior to (a), the second set of nucleic acid expression data is subjected to a method comprising (a) extracting gene expression information across the dataset by (i) mapping the sequence reads from the dataset to a reference human genome, thereby generating read count data across the mapped genome; and (ii) applying fast Fourier transformation (FFT) to the read count data in nucleosome occupancy windows across the mapped genome to determine an FFT signal at each nucleosome occupancy window, wherein an increased FFT signal is indicative of nucleosomal depletion and a decreased FFT signal is indicative of nucleosomal presence; and (b) performing quality control of the cfDNA dataset comprising removal of samples from the dataset that possess a low or inconsistent FFT signal as determined in (b)(i)-(b)(ii) and / or a measured circulating tumor DNA (ctDNA) content below a dataset-specific threshold, thereby generating a quality-controlled cfDNA dataset comprised of features that are reflective of gene activity or expression.

[0012] In yet another aspect, provided herein is a computer-implemented method for generating a classifier for determining a desired feature of a cancer of interest in a subject suffering from theAttorney Docket No. GNCN-026 / 01WO 320289-2157 cancer of interest or suspected of suffering from the cancer of interest, the method comprising: (a) receiving in a computer system, a first set of nucleic acid expression data for each of a plurality of tissue samples obtained from a first population of subjects suffering from or suspected of suffering from the cancer of interest and a second set of nucleic acid expression data from each of a plurality of bodily fluid samples obtained from a second population of subjects suffering from or suspected of suffering from the cancer of interest, wherein each of the plurality of bodily fluid samples comprises cell-free DNA (cfDNA); (b) determining by the computer system, an integrative correlation coefficient for each gene from the first and second set of nucleic acid expression data, wherein the integrative correlation coefficient is a numerical representation of how each gene from the first set of nucleic acid expression data correlate with each other compared to the rank of the same gene from the second set of nucleic acid expression data; (c) ranking by the computer system, the integrative correlation coefficient of each gene from step (b); (d) selecting, by the computer system, genes that have an integrative correlation coefficient above a desired threshold to generate a gene signature, wherein the gene signature comprises genes whose expression patterns are substantially similar between the tissue sample and the bodily fluid sample; (e) inputting nucleic acid expression data for each gene in the gene signature from (d) from at least two training sets, wherein one of the at least two training sets comprises nucleic acid expression data for each gene in the gene signature from (d) from each of a plurality of tissue samples from subjects that are indicative of the presence a desired feature for the cancer of interest, while another of the at least two training sets comprises nucleic acid expression data for each gene in the gene signature from (d) from each of a plurality of tissue samples from subjects that are indicative of the absence of the desired feature for the cancer of interest; and (f) conducting on the computer system a linear discriminate analysis (LDA) comprising feature selection to generate a classifier for the desired feature of the cancer of interest, wherein the feature selection comprises: (i) calculating a test statistic (t-statistic) for each gene in the gene signature from (d) from the at least two training sets, wherein the t-statistic for each gene indicates each gene's ability to distinguish between the presence or the absence of the desired feature; (ii) ranking each gene based on each gene’s t- statistic; and (iii) selecting each gene whose t-statistic is above a desired threshold for distinguishing between the presence or the absence of the desired feature, thereby generating a classifier for the desired feature comprise each of the selected genes. In some cases, the method further comprises (g) classifying on the computer system one or more test samples obtained fromAttorney Docket No. GNCN-026 / 01WO 320289-2157 an independent population of subjects suffering the cancer of interest as possessing or not possessing the desired feature using the classifier from (f)(iii) on the one or more test samples; (h) comparing on the computer system, the classification of the one or more test samples to classification of the one or more samples for the desired feature as determined using a control classifier of the desired feature for the cancer of interest; and (i) validating on the computer system the classifier from (f)(iii) if the comparing indicates that the classifications of the one or more test samples is substantially similar to classification of the one or more test samples determined using the control classifier. In some cases, the ICC is above the desired threshold of the 99thpercentile of a null distribution of integrative correlations. In some cases, prior to (a), the second set of nucleic acid expression data is subjected to a method comprising (a) extracting gene expression information across the dataset by (i) mapping the sequence reads from the dataset to a reference human genome, thereby generating read count data across the mapped genome; and (ii) applying fast Fourier transformation (FFT) to the read count data in nucleosome occupancy windows across the mapped genome to determine an FFT signal at each nucleosome occupancy window, wherein an increased FFT signal is indicative of nucleosomal depletion and a decreased FFT signal is indicative of nucleosomal presence; and (b) performing quality control of the cfDNA dataset comprising removal of samples from the dataset that possess a low or inconsistent FFT signal as determined in (b)(i)-(b)(ii) and / or a measured circulating tumor DNA (ctDNA) content below a dataset-specific threshold, thereby generating a quality-controlled cfDNA dataset comprised of features that are reflective of gene activity or expression. In some cases, the desired feature is selected from the group consisting of a subtype of the cancer of interest, microsatellite instability, mismatch repair proficiency, minimal residual disease, prognostic risk, response to a therapeutic agent, development of resistance to a therapeutic agent. In some cases, the cancer of interest is COAD, the desired feature is microsatellite instability and the classifier generated comprises, consists essentially of or consists of the genes in Table 2. In some cases, the cancer of interest is PAAD, the desired feature is a PAAD subtype, and the classifier generated comprises, consists essentially of or consists of the genes in Table 4. In some cases, the subtype is basal or classical. In some cases, the cancer of interest is BLCA, the desired feature is an FGFR activation signature, and the classifier generated comprises, consists essentially of or consists of the genes in Table 6. In some cases, the FGFR activation signature is a FGFR3 activation signature.Attorney Docket No. GNCN-026 / 01WO 320289-2157

[0013] Further to any of the above aspects, in some cases, the first and the second population of subjects consist of the same subjects.

[0014] Further to any of the above aspects, in some cases, the tissue samples are tumor tissue samples selected from the group consisting of a formalin-fixed, paraffin-embedded (FFPE) tissue sample, a fresh tissue sample and a frozen tissue sample. In some cases, the bodily fluid is selected from the group consisting of whole blood, plasma, serum, an exosome, wash fluids, urine, saliva, cerebrospinal fluid or sputum.

[0015] Further to any of the above aspects, in some cases, the first set of nucleic acid expression data is nucleic acid sequencing data or methylome sequencing data obtained from nucleic acid extracted from each tissue sample from the plurality of tissue samples. In some cases, the nucleic acid sequencing data is DNA sequencing data or RNA sequencing data.

[0016] Further to any of the above aspects, in some cases, the second set of nucleic acid expression data is nucleic acid sequencing data or methylome sequencing data obtained from nucleic acid extracted from each bodily fluid sample from the plurality of bodily fluid samples. In some cases, the nucleic acid sequencing data is DNA sequencing data or RNA sequencing data.

[0017] Further to any of the above aspects, in some cases, the second set of nucleic acid expression data comprises a fast Fourier transform (FFT) magnitude matrix for each of a plurality of genomic regions of interest. In some cases, the FFT magnitude matrix for each of the plurality of genomic regions of interest is generated by the computer system is a method that comprises: (i) inputting into the computer system, nucleic acid sequencing data generated from nucleic acid extracted from a each of the bodily fluid samples, wherein the nucleic acid sequencing data includes a plurality of fragment reads, wherein each fragment read has a fragment length and a GC content indicating a percentage of bases in the fragment read that are G or C; (ii) determining by the computing system, GC bias values for each fragment read based on the fragment length and the GC content of the fragment read; (iii) generating by the computing system, a genomic coverage distribution that is adjusted for GC bias using the sequence read data and the GC bias values; (iv) calculating by the computing system mean sequence read counts for a sliding window across a defined window in each of the plurality of genomic regions of interest from the genomic coverage distribution to generate smoothed mean read counts; and (v) performing a fast Fourier transform (FFT) on the smoothed mead read counts to generate the FFT magnitude matrix for each of the plurality of genomic regions of interest. In some cases, the sliding window has a width of at least, at most orAttorney Docket No. GNCN-026 / 01WO 320289-2157 exactly 5, 10, 15, 20, 25, 30, 35, 40, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95 or 100 nucleotides across the defined window. In some cases, the sliding window has a width of 15 nucleotides across the defined window. In some cases, the defined window has a width of 2000 base pairs.

[0018] Further to any of the above aspects, in some cases, the cancer of interest is selected from the group consisting of kidney renal papillary cell carcinoma (KIRP); breast invasive carcinoma (BRCA); thyroid cancer (THCA); bladder urothelial carcinoma (BLCA); prostate adenocarcinoma (PRAD); kidney chromophobe (KICH); cervical squamous cell carcinoma and endocervical adenocarcinoma (CESC); kidney renal clear cell carcinoma (KIRC); liver hepatocellular carcinoma (LIHC); low grade glioma (LGG); sarcoma (SARC); lung adenocarcinoma (LUAD); colon adenocarcinoma (COAD); head and neck squamous cell carcinoma (HNSC); uterine corpus endometrial carcinoma (UCEC); glioblastoma multiforme (GBM); esophageal carcinoma (ESCA); stomach adenocarcinoma (STAD); ovarian serous cystadenocarcinoma (OV); rectum adenocarcinoma (READ); adrenocortical carcinoma (ACC); uveal melanoma (UVM); mesothelioma (MESO); pheochromocytoma and paraganglioma (PCPG); skin cutaneous melanoma (SKCM); uterine carcinosarcoma (UCS); lung squamous cell carcinoma (LUSC); testicular germ cell tumors (TGCT); cholangiocarcinoma (CHOL); pancreatic adenocarcinoma (PAAD); thymoma (THYM); or Lymphoid Neoplasm Diffuse Large B-cell Lymphoma (DLBC). In some cases, the cancer of interest is COAD. In some cases, the first set and the second set of nucleic acid expression data is nucleic acid sequencing data and the gene signature generated comprises, consists essentially of or consists of the genes in Table 1. In some cases, the cancer of interest is PAAD. In some cases, the first set and the second set of nucleic acid expression data is nucleic acid sequencing data and the gene signature comprises, consists essentially of or consists of the genes in Table 3. In some cases, the cancer of interest is BLCA. In some cases, the first set and the second set of nucleic acid expression data is nucleic acid sequencing data and the gene signature comprises, consists essentially of or consists of the genes in Table 5.

[0019] In another aspect, provided herein is a method of assaying a sample obtained from a subject suffering from COAD, the method comprising measuring the expression level of a plurality of biomarkers selected from Table 1 or Table 2 using a sequencing assay. In some cases, the sequencing assay is a DNA sequencing assay or RNA sequencing assay. In some cases, the plurality of biomarkers selected from Table 1 comprises at least 50 biomarkers, at least 100 biomarkers, at least 150 biomarkers, at least 200 biomarkers, at least 250 classifier biomarkers, atAttorney Docket No. GNCN-026 / 01WO 320289-2157 least 300 biomarkers, at least 350 biomarkers, at least 400 biomarkers, at least 450 biomarkers, at least 500 biomarkers, at least 550 biomarkers, at least 600 biomarkers, at least 650 biomarkers, at least 700 biomarkers, at least 750 biomarkers, at least 800 biomarkers, at least 850 biomarkers, at least 900 biomarkers or at least 950 biomarkers from Table 1. In some cases, the plurality of classifiers selected from Table 1 comprises at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% of the classifiers from Table 1. In some cases, the plurality of classifiers selected from Table 1 consists of all the classifiers of Table 1. In some cases, the plurality of classifiers selected from Table 2 comprises at least 2 classifiers, at least 4 classifiers, at least 6 classifiers, at least 8 classifiers, at least 10 classifier biomarkers, at least 12 classifiers, at least 14 classifiers, at least 16 classifiers, at least 18 classifiers, at least 20 classifiers, at least 22 classifiers, at least 24 classifiers, at least 28 classifiers, at least 30 classifiers, at least 32 classifiers, at least 34 classifiers, at least 36 classifiers, at least 38 classifiers, at least 40 classifiers, at least 42 classifiers, at least 44 classifiers, at least 46 classifiers, at least 48 classifiers, at least 50 classifiers, at least 52 classifiers, at least 54 classifiers, at least 56 classifiers, at least 58 classifiers, at least 60 classifiers, at least 62 classifiers, at least 64 classifiers, at least 66 classifiers or at least 68 classifiers from Table 2. In some cases, the plurality of classifiers selected from Table 2 comprises at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% of the classifiers from Table 2. In some cases, the plurality of classifiers selected from Table 2 consists of all the classifiers of Table 2. In some cases, the sample is a formalin-fixed, paraffin-embedded (FFPE) tissue sample from the colon of the subject, fresh or a frozen tissue sample from the colon of the subject, an exosome, wash fluids, cell pellets, or a bodily fluid obtained from the subject. In some cases, the bodily fluid is whole blood, plasma, serum, an exosome, wash fluids, urine, saliva, cerebrospinal fluid or sputum.

[0020] In one aspect, provided herein is a method of treating COAD in a subject, the method comprising: determining microsatellite instability of COAD of a subject suffering from COAD by measuring a nucleic acid expression level of a plurality of classifier biomarkers in a sample obtained from a subject suffering from or suspected of suffering from COAD, wherein the plurality of classifier biomarkers is selected from Table 2, wherein the nucleic acid expression level of theAttorney Docket No. GNCN-026 / 01WO 320289-2157 plurality of classifier biomarkers indicates the presence of microsatellite instability (MSI); and administering a therapeutic intervention based on the presence or absence of MSI, wherein the therapeutic intervention is immune checkpoint inhibitor therapy (ICI) when MSI is present (MSI positive) or an immuno-oncology (IO) treatment and / or chemotherapy if MSI is absent (MSI negative). In some cases, the determining step further comprises comparing the nucleic acid expression levels of the plurality of classifier biomarkers from Table 2 to the nucleic acid expression levels of the plurality of classifier biomarkers from Table 2 in at least one sample training set(s), wherein the at least one sample training set comprises nucleic acid expression level data of the plurality of classifier biomarkers from Table 2 from a reference MSI positive sample, nucleic acid expression level data of the plurality of classifier biomarkers from Table 2 from a reference MSI negative sample or a combination thereof; and classifying the sample obtained from the subject as MSI positive or MSI negative based on the results of the comparing step. In some cases, the comparing step comprises applying a statistical algorithm which comprises determining a correlation between the expression data obtained from the sample obtained from the subject and the expression data from the at least one training set(s); and classifying the sample obtained from the subject as MSI positive or MSI negative based on the results of the statistical algorithm. In some cases, the plurality of classifiers selected from Table 2 comprises at least 2 classifiers, at least 4 classifiers, at least 6 classifiers, at least 8 classifiers, at least 10 classifier biomarkers, at least 12 classifiers, at least 14 classifiers, at least 16 classifiers, at least 18 classifiers, at least 20 classifiers, at least 22 classifiers, at least 24 classifiers, at least 28 classifiers, at least 30 classifiers, at least 32 classifiers, at least 34 classifiers, at least 36 classifiers, at least 38 classifiers, at least 40 classifiers, at least 42 classifiers, at least 44 classifiers, at least 46 classifiers, at least 48 classifiers, at least 50 classifiers, at least 52 classifiers, at least 54 classifiers, at least 56 classifiers, at least 58 classifiers, at least 60 classifiers, at least 62 classifiers, at least 64 classifiers, at least 66 classifiers or at least 68 classifiers from Table 2. In some cases, the plurality of classifiers selected from Table 2 comprises at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% of the classifiers from Table 2. In some cases, the plurality of classifiers selected from Table 2 consists of all the classifiers of Table 2. In some cases, the sample is a formalin-fixed, paraffin-embedded (FFPE) tissue sample from the colon of the subject, fresh or a frozen tissue sample from the colon of theAttorney Docket No. GNCN-026 / 01WO 320289-2157 subject, an exosome, wash fluids, cell pellets, or a bodily fluid obtained from the subject. In some cases, the bodily fluid is whole blood, plasma, serum, an exosome, wash fluids, urine, saliva, cerebrospinal fluid or sputum.

[0021] In another aspect, provided herein is a method of assaying a sample obtained from a subject suffering from PAAD, the method comprising measuring the expression In some cases, the sequencing assay is a DNA sequencing assay or RNA sequencing assay. In some cases, the plurality of biomarkers selected from Table 3 comprises at least 50 biomarkers, at least 100 biomarkers, at least 150 biomarkers, at least 200 biomarkers, at least 250 classifier biomarkers, at least 300 biomarkers, at least 350 biomarkers, at least 400 biomarkers, at least 450 biomarkers, at least 500 biomarkers, at least 550 biomarkers, at least 600 biomarkers, at least 650 biomarkers, at least 700 biomarkers, at least 750 biomarkers, at least 800 biomarkers, at least 850 biomarkers, at least 900 biomarkers or at least 950 biomarkers from Table 3. In some cases, the plurality of classifiers selected from Table 3 comprises at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% of the classifiers from Table 3. In some cases, the plurality of classifiers selected from Table 3 consists of all the classifiers of Table 3. In some cases, the plurality of classifiers selected from Table 4 comprises at least 2 classifiers, at least 4 classifiers, at least 6 classifiers, at least 8 classifiers, at least 10 classifier biomarkers, at least 12 classifiers, at least 14 classifiers, at least 16 classifiers, at least 18 classifiers, at least 20 classifiers, at least 22 classifiers, at least 24 classifiers, at least 28 classifiers, at least 30 classifiers, at least 32 classifiers, at least 34 classifiers, at least 36 classifiers, at least 38 classifiers, at least 40 classifiers, at least 42 classifiers, at least 44 classifiers, at least 46 classifiers, at least 48 classifiers, at least 50 classifiers, at least 52 classifiers, at least 54 classifiers, at least 56 classifiers, at least 58 classifiers, at least 60 classifiers, at least 62 classifiers, at least 64 classifiers, at least 66 classifiers, at least 68 classifiers, at least 70 classifiers, at least 72 classifiers, at least 74 classifiers, at least 76 classifiers, at least 78 classifiers or at least 80 classifiers from Table 4. In some cases, the plurality of classifiers selected from Table 4 comprises at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% of the classifiers from Table 4. In some cases, the plurality of classifiers selected from Table 4 consists of all the classifiers of Table 4. In someAttorney Docket No. GNCN-026 / 01WO 320289-2157 cases, the sample is a formalin-fixed, paraffin-embedded (FFPE) tissue sample from the pancreas of the subject, fresh or a frozen tissue sample from the pancreas of the subject, an exosome, wash fluids, cell pellets, or a bodily fluid obtained from the subject. In some cases, the bodily fluid is whole blood, plasma, serum, an exosome, wash fluids, urine, saliva, cerebrospinal fluid or sputum.

[0022] In one aspect, provided herein is a method of treating PAAD in a subject, the method comprising: determining subtype of PAAD of a subject suffering from PAAD by measuring a nucleic acid expression level of a plurality of classifier biomarkers in a sample obtained from a subject suffering from or suspected of suffering from PAAD, wherein the plurality of classifier biomarkers is selected from Table 4, wherein the nucleic acid expression level of the plurality of classifier biomarkers indicates the subtype of PAAD as being classical or basal and administering a therapeutic intervention based on the subtype of PAAD, wherein the therapeutic intervention is selected from: (i) agents listed for the classical subtype in Table 7, 5-flourouracil and platinum- based therapy if the subject is classified as having the classical subtype of PAAD, or (ii) agents listed for the basal subtype in Table 7, cisplatin- or oxaliplatin-based therapies and gemcitabine if the subject is classified as having the basal subtype of PAAD. In some cases, the determining step further comprises comparing the nucleic acid expression levels of the plurality of classifier biomarkers from Table 4 to the nucleic acid expression levels of the plurality of classifier biomarkers from Table 4 in at least one sample training set(s), wherein the at least one sample training set comprises nucleic acid expression level data of the plurality of classifier biomarkers from Table 4 from a reference PAAD classical sample, nucleic acid expression level data of the plurality of classifier biomarkers from Table 4 from a reference PAAD basal sample or a combination thereof; and classifying the sample obtained from the subject as PAAD classical or PAAD basal based on the results of the comparing step. In some cases, the comparing step comprises applying a statistical algorithm which comprises determining a correlation between the expression data obtained from the sample obtained from the subject and the expression data from the at least one training set(s); and classifying the sample obtained from the subject as PAAD classical or PAAD basal based on the results of the statistical algorithm. In some cases, the plurality of classifiers selected from Table 4 comprises at least 2 classifiers, at least 4 classifiers, at least 6 classifiers, at least 8 classifiers, at least 10 classifier biomarkers, at least 12 classifiers, at least 14 classifiers, at least 16 classifiers, at least 18 classifiers, at least 20 classifiers, at least 22 classifiers, at least 24 classifiers, at least 28 classifiers, at least 30 classifiers, at least 32 classifiers, at least 34Attorney Docket No. GNCN-026 / 01WO 320289-2157 classifiers, at least 36 classifiers, at least 38 classifiers, at least 40 classifiers, at least 42 classifiers, at least 44 classifiers, at least 46 classifiers, at least 48 classifiers, at least 50 classifiers, at least 52 classifiers, at least 54 classifiers, at least 56 classifiers, at least 58 classifiers, at least 60 classifiers, at least 62 classifiers, at least 64 classifiers, at least 66 classifiers, at least 68 classifiers, at least 70 classifiers, at least 72 classifiers, at least 74 classifiers, at least 76 classifiers, at least 78 classifiers or at least 80 classifiers from Table 4. In some cases, the plurality of classifiers selected from Table 4 comprises at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% of the classifiers from Table 4. In some cases, the plurality of classifiers selected from Table 4 consists of all the classifiers of Table 4. In some cases, the sample is a formalin-fixed, paraffin-embedded (FFPE) tissue sample from the pancreas of the subject, fresh or a frozen tissue sample from the pancreas of the subject, an exosome, wash fluids, cell pellets, or a bodily fluid obtained from the subject. In some cases, the bodily fluid is whole blood, plasma, serum, an exosome, wash fluids, urine, saliva, cerebrospinal fluid or sputum.

[0023] In another aspect, provided herein is a method of assaying a sample obtained from a subject suffering from bladder cancer (BLCA), the method comprising measuring the expression level of a plurality of classifiers selected from Table 5 or Table 6 using a sequencing assay. In some cases, the sequencing assay is a DNA sequencing assay or RNA sequencing assay. In some cases, the plurality of biomarkers selected from Table 5 comprises at least 50 biomarkers, at least 100 biomarkers, at least 150 biomarkers, at least 200 biomarkers, at least 250 classifier biomarkers, at least 300 biomarkers, at least 350 biomarkers, at least 400 biomarkers, at least 450 biomarkers, at least 500 biomarkers, at least 550 biomarkers, at least 600 biomarkers, at least 650 biomarkers, at least 700 biomarkers, at least 750 biomarkers, at least 800 biomarkers, at least 850 biomarkers, at least 900 biomarkers or at least 950 biomarkers from Table 5. In some cases, the plurality of classifiers selected from Table 5 comprises at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% of the classifiers from Table 5. In some cases, the plurality of classifiers selected from Table 5 consists of all the classifiers of Table 5. In some cases, the plurality of classifiers selected from Table 6 comprises at least 2 classifiers, at least 4 classifiers, at least 6 classifiers, at least 8Attorney Docket No. GNCN-026 / 01WO 320289-2157 classifiers, at least 10 classifier biomarkers, at least 12 classifiers, at least 14 classifiers, at least 16 classifiers, at least 18 classifiers, at least 20 classifiers, at least 22 classifiers, at least 24 classifiers, at least 28 classifiers, at least 30 classifiers, at least 32 classifiers, at least 34 classifiers, at least 36 classifiers, at least 38 classifiers, at least 40 classifiers, at least 42 classifiers, at least 44 classifiers, at least 46 classifiers, at least 48 classifiers or at least 50 classifiers from Table 6. In some cases, the plurality of classifiers selected from Table 6 comprises at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% of the classifiers from Table 6. In some cases, the plurality of classifiers selected from Table 6 consists of all the classifiers of Table 6. In some cases, the sample is a formalin-fixed, paraffin-embedded (FFPE) tissue sample from the bladder of the subject, fresh or a frozen tissue sample from the bladder of the subject, an exosome, wash fluids, cell pellets, or a bodily fluid obtained from the subject. In some cases, the bodily fluid is whole blood, plasma, serum, an exosome, wash fluids, urine, saliva, cerebrospinal fluid or sputum. In some cases, the bodily fluid is urine.

[0024] In one aspect, provided herein is a method of treating BLCA in a subject, the method comprising: measuring a nucleic acid expression level of a plurality of classifiers in a sample obtained from a subject suffering from or suspected of suffering from BLCA, wherein the plurality of classifiers is selected from Table 6, wherein the measured nucleic acid expression levels of the plurality of classifiers provide an FGFR3 activation signature for the sample; and administering an FGFR inhibitor based on presence of a positive FGFR3 activation signature, wherein the positive FGFR3 activation signature is indicative of presence of one or more FGFR3 mutations. In some cases, the determining step further comprises comparing the nucleic acid expression levels of the plurality of classifiers from Table 6 to the nucleic acid expression levels of the plurality of classifier from Table 6 in at least one sample training set(s), wherein the at least one sample training set is from a reference FGFR3 mutation-containing BLCA sample, or is from a reference FGFR3 mutation-free BLCA sample; and classifying the tumor sample as having a positive FGFR3 activation signature based on the results of the comparing step. In some cases, the at least one training set is from a reference FGFR3 mutation-containing cancer sample and the sample is classified as possessing the positive FGFR3 activation signature if the nucleic acid expression levels of the plurality of classifiers of Table 6 correlate with the nucleic acid expression levels ofAttorney Docket No. GNCN-026 / 01WO 320289-2157 the plurality of classifiers of Table 6 from the reference FGFR3 mutation-containing cancer sample. In some cases, the at least one training set is from a reference FGFR3 mutation-containing cancer sample and from a reference FGFR3 mutation-free cancer sample and the sample is classified as possessing the positive FGFR3 activation signature if the expression levels of the plurality of classifiers of Table 6 correlate with the expression levels of the plurality of classifiers of Table 6 from the reference FGFR3 mutation-containing cancer sample. In some cases, the comparing step comprises applying a statistical algorithm which comprises determining a correlation between the expression data obtained from the sample obtained from the subject and the expression data from the at least one training set(s); and classifying the sample obtained from the subject as possessing a positive FGFR3 activation signature on the results of the statistical algorithm. In some cases, the plurality of classifiers selected from Table 6 comprises at least 2 classifiers, at least 4 classifiers, at least 6 classifiers, at least 8 classifiers, at least 10 classifier biomarkers, at least 12 classifiers, at least 14 classifiers, at least 16 classifiers, at least 18 classifiers, at least 20 classifiers, at least 22 classifiers, at least 24 classifiers, at least 28 classifiers, at least 30 classifiers, at least 32 classifiers, at least 34 classifiers, at least 36 classifiers, at least 38 classifiers, at least 40 classifiers, at least 42 classifiers, at least 44 classifiers, at least 46 classifiers, at least 48 classifiers or at least 50 classifiers from Table 6. In some cases, the plurality of classifiers selected from Table 6 comprises at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% of the classifiers from Table 6. In some cases, the plurality of classifiers selected from Table 6 consists of all the classifiers of Table 6. In some cases, the sample is a formalin-fixed, paraffin-embedded (FFPE) tissue sample from the bladder of the subject, fresh or a frozen tissue sample from the bladder of the subject, an exosome, wash fluids, cell pellets, or a bodily fluid obtained from the subject. In some cases, the bodily fluid is whole blood, plasma, serum, an exosome, wash fluids, urine, saliva, cerebrospinal fluid or sputum. In some cases, the bodily fluid is urine. In some cases, the FGFR inhibitor shows inhibitory activity toward fibroblast growth factor receptor-3 (FGFR3). In some cases, the FGFR inhibitor is a tyrosine kinase inhibitor. In some cases, the FGFR inhibitor is a selective tyrosine kinase inhibitor. In some cases, the FGFR inhibitor is a non-selective tyrosine kinase inhibitor. In some cases, the FGFR inhibitor is selected from the group consisting of erdafitinib (JNJ 42756493), infigratinib (BGJ1398), Rogaritinib (BAY 1163877), AZD4547,Attorney Docket No. GNCN-026 / 01WO 320289-2157 Pemigatinib (INCB54828), TAS-120, LY2874455, DEBIO 1347, PD173074, BLU9931, pazopanib, brivanib, ponatinib (AP24534), regorafenib (BAY 73-4506), lenvatinib (E7080), dovitinib (TKI258), lucitanib (E3810), nintedanib (BIBF 1120), Foretinib, and any combination thereof. In some cases, the FGFR inhibitor is nintedanib (BIBF 1120). In some cases, the FGFR inhibitor is an antibody or antibody-conjugate. In some cases, the FGFR inhibitor is B-701 or MFGR1877S. In some cases, the FGFR inhibitor is LY3076226.

[0025] In one aspect, provided herein is a computer-implemented method for identifying regulatory regions for one or more genes that correlate with gene expression of the one or more genes, the method comprising: (a) inputting into a computer system, a first dataset comprising RNA sequencing data and a second dataset comprising chromatin accessibility data; (b) mapping by the computer system, sequence reads obtained from the first dataset to a reference genome sequence; (c) mapping by the computer system, sequence reads obtained from the second dataset to the reference genome sequence; (d) annotating by the computer system, peaks of the sequence reads over random or low-level noise in the chromatin accessibility data from the second dataset wherein each peak represents a putative chromatin accessible region;(e) linking by the computer system, the annotated peaks from (d) to a nearest gene found from the mapped sequence reads to the first dataset using machine learning, thereby generating a plurality of genes and chromatin accessibility regions associated with each gene from the plurality of genes; (f) measuring by the computer system, a correlation between a level of expression of each gene in the plurality of genes to each chromatin accessibility region associated with each gene in the plurality of genes, wherein the level of expression of each gene in the plurality of genes is evidenced by the number of sequence reads from the first dataset that mapped to each gene in the plurality of genes; (g) ranking by the computer system, the correlations obtained in (f); (h) selecting by the computer system, chromatin accessibility regions whose correlations coefficients were greater than a desired threshold, thereby generating a list of chromatin accessibility regions that are highly correlated with RNA expression levels; and (i) selecting by the computer system, regulatory regions from the chromatin accessibility regions selected in (h), thereby identifying regulatory regions for one or more genes that highly correlate with RNA expression of the one or more genes. In some cases, the method further comprises (j) designing by the computer, a set of probes, wherein each probe in the set comprises sequence complementary to one of the regulatory regions selected in (i). In some cases, the method further comprises (j) designing by the computer, a set of primer pairs,Attorney Docket No. GNCN-026 / 01WO 320289-2157 wherein each primer pair in the set comprises sequence complementary to one of the regulatory regions selected in (i). In some cases, the correlation is a Pearson correlation or a Spearman correlation. In some cases, the correlation is a Pearson and a Spearman correlation. In some cases, the correlation is a Spearman correlation. In some cases, the desired threshold can be a correlation coefficient greater than 0.6, 0.7, 0.8 or 0.9. In some cases, the desired threshold is a correlation coefficient greater than 0.8. In some cases, the chromatin accessibility data from the second dataset is obtained from a chromatin accessibility sequencing (CA-seq) assay. In some cases, the CA-seq assay is selected from the group consisting of micrococcal nuclease digestion with deep sequencing (MNase-seq), DNase I hypersensitive sites sequencing (DNase-seq), Formaldehyde- Assisted Isolation of Regulatory Elements sequencing (FAIRE-seq) and Assay for Transposase- Accessible Chromatin using sequencing (ATAC-seq). In some cases, the CA-seq assay is ATAC- seq and the chromatin accessibility regions are ATAC-seq regions. In some cases, the regulatory regions are selected from the group consisting of the 5’UTR, the 3’ UTR, a promoter, a proximal enhancer, a distal enhancer and any combination thereof. In some cases, the first dataset and the second dataset are each obtained from tissue samples or tissue biopsies. In some cases, the tissue samples are tissue matched samples between the first and second datasets. In some cases, the first dataset and the second dataset are obtained from a population of subjects that each have matched chromatin accessibility sequencing data and RNA expression sequencing data. In some cases, the first and second datasets are obtained from one or more subjects that suffer from or are suspected of suffering from a desired feature of a cancer. In some cases, the desired feature is a type of cancer, a subtype of cancer or the presence or absence of microsatellite instability. In some cases, prior to step (a), the second dataset comprising chromatin accessibility data is generated by inputting into the computer system, chromatin accessibility sequencing data for a first plurality of subjects suffering from one type of cancer or subtype thereof and chromatin accessibility sequencing (e.g., ATAC-seq) data for a second plurality of subjects suffering from a second type of cancer or subtype thereof; and filtering out by the computer system, chromatin accessibility regions present in the chromatin accessibility sequencing data for the first and second plurality of subjects, thereby generating a second dataset comprising chromatin accessibility data specific to the first and second type of cancer or subtype thereof. In some cases, the machine learning comprises a tool from Algorithms for Calculating Microarray Enrichment (ACME). In some cases, the tool from ACME is a findClosestGene function. In some cases, the reference genome is aAttorney Docket No. GNCN-026 / 01WO 320289-2157 reference human genome. In some cases, the reference human genome is a GRCh38 (hg38) assembly.

[0026] In one aspect, provided herein is a method for sequencing target regulatory regions in cfDNA, the method comprising: (a) obtaining a sample comprising cfDNA from a subject suffering from or suspected of suffering from cancer; and (b) performing a target enrichment sequencing assay on the cfDNA sample using a panel comprising a plurality of probes, wherein each probe in the plurality comprises sequence complementary to regulatory regions for one or more genes that correlate with gene expression of the one or more genes, wherein the panel is generated using any method provided herein, thereby generating a gene activity matrix for the cfDNA for the subject. In some cases, the target enrichment sequencing assay comprises a hybrid- capture based enrichment step followed by a sequencing assay. In some cases, the sequencing assay comprises next-generation sequencing (NGS). In some cases, the sample comprising cfDNA is a bodily fluid sample or liquid biopsy sample.

[0027] In one aspect, provided herein is a computer-implemented method comprising: (a) inputting into a computer system, paired-end DNA sequencing data obtained from a target enrichment sequencing assay performed on a sample comprising cfDNA obtained from a subject, wherein the target enrichment sequencing assay is performed using a panel comprising a plurality of probes generated using any method provided herein; (b) calculating by the computer system, fragment sequence lengths for sequences from the pair-end DNA sequencing data that comprise overlapping sequence for a first regulatory region from (a)); (c) calculating a probability of finding each fragment size by the computer system, using counts of each unique fragment sequence fragment length for the first regulatory region; (d) calculating Shannon entropy using the probabilities from (c), thereby generating a Shannon entropy value for the first regulatory region; and (e) repeating (b)-(d) for each additional regulatory region from the target enrichment sequencing assay, thereby generating a Shannon entropy matrix for each regulatory region obtained from the target enrichment sequencing assay for the sample comprising cfDNA obtained from the subject. In some cases, the target enrichment sequencing assay is a hybrid capture enrichment sequencing assay, wherein each probe in the panel binds to and pulls down a regulatory region associated with a target gene from the plurality of target genes prior to being subjected to a sequencing assay. In some cases, the subject suffers from a desired feature of cancer. In some cases, the Shannon entropy matrix for each regulatory region obtained from the target enrichmentAttorney Docket No. GNCN-026 / 01WO 320289-2157 sequencing assay from cfDNA for the subject represents a sample positive for the desired feature of cancer. In some cases, (a)-(e) are repeated on samples comprising cfDNA obtained from a plurality of subjects that each suffer from the desired feature of cancer. In some cases, the Shannon entropy matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for each of the plurality of subjects that suffer from the desired feature of cancer are combined by computer system. In some cases, the subject does not suffer from the desired feature of cancer. In some cases, the Shannon entropy matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for the subject represents a sample negative for the desired feature of cancer. In some cases, (a)-(e) are repeated on samples comprising cfDNA obtained from a plurality of subjects that each do not suffer from the desired feature of cancer. In some cases, the Shannon entropy matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for each of the plurality of subjects that do not suffer from the desired feature of cancer are combined by the computer system. In some cases, the desired feature of the cancer is selected from the group consisting of a type of cancer, a subtype of the cancer of interest, microsatellite instability, mismatch repair proficiency, minimal residual disease, prognostic risk, response to a therapeutic agent, development of resistance to a therapeutic agent cancer.

[0028] Ine one aspect, provided herein is a computer-implemented method for generating a classifier for determining a desired feature of cancer, the method comprising: (a) receiving in a computer system, a first set of Shannon entropy indices obtained from cfDNA comprising samples obtained from a first population of subjects suffering from the desired feature of the cancer and a second set of Shannon entropy indices obtained from cfDNA comprising samples obtained from a second population of subjects that do not suffer from the desired feature of the cancer; (b) selecting by the computer system, regulatory regions for model training and validation using the first set and second set of Shannon entropy matrices; and (c) training and validating by the computer system a machine learning algorithm for determining the presence or absence of the desired feature of the cancer in a cfDNA comprising sample, thereby generating a classifier for determining a desired feature of a cancer. In some cases, the method further comprises inputting target enrichment sequencing data obtained for cfDNA comprising sample obtained from a test subject suspected of suffering from the desired feature of the cancer into the classifier generated by (a)- (c), thereby determining the presence or absence of the desired feature of the cancer in the testAttorney Docket No. GNCN-026 / 01WO 320289-2157 subject. In some cases, the cfDNA comprising samples are a bodily fluid samples or liquid biopsy samples. In some cases, the desired feature is selected from the group consisting of a type of cancer, a subtype of the cancer of interest, microsatellite instability, mismatch repair proficiency, minimal residual disease, prognostic risk, response to a therapeutic agent, development of resistance to a therapeutic agent cancer. In some cases, the Shannon entropy matrices for the first and the second populations are generated using the computer-implemented method comprising: (a) inputting into a computer system, paired-end DNA sequencing data obtained from a target enrichment sequencing assay performed on a sample comprising cfDNA obtained from a subject, wherein the target enrichment sequencing assay is performed using a panel comprising a plurality of probes generated using any method provided herein; (b) calculating by the computer system, fragment sequence lengths for sequences from the pair-end DNA sequencing data that comprise overlapping sequence for a first regulatory region from (a)); (c) calculating a probability of finding each fragment size by the computer system, using counts of each unique fragment sequence fragment length for the first regulatory region; (d) calculating Shannon entropy using the probabilities from (c), thereby generating a Shannon entropy value for the first regulatory region; and (e) repeating (b)-(d) for each additional regulatory region from the target enrichment sequencing assay, thereby generating a Shannon entropy matrix for each regulatory region obtained from the target enrichment sequencing assay for the sample comprising cfDNA obtained from the subject. In some cases, the target enrichment sequencing assay is a hybrid capture enrichment sequencing assay, wherein each probe in the panel binds to and pulls down a regulatory region associated with a target gene from the plurality of target genes prior to being subjected to a sequencing assay. In some cases, the subject suffers from a desired feature of cancer. In some cases, the Shannon entropy matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for the subject represents a sample positive for the desired feature of cancer. In some cases, (a)-(e) are repeated on samples comprising cfDNA obtained from a plurality of subjects that each suffer from the desired feature of cancer. In some cases, the Shannon entropy matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for each of the plurality of subjects that suffer from the desired feature of cancer are combined by computer system. In some cases, the subject does not suffer from the desired feature of cancer. In some cases, the Shannon entropy matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for the subject represents a sample negative for the desired featureAttorney Docket No. GNCN-026 / 01WO 320289-2157 of cancer. In some cases, (a)-(e) are repeated on samples comprising cfDNA obtained from a plurality of subjects that each do not suffer from the desired feature of cancer. In some cases, the Shannon entropy matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for each of the plurality of subjects that do not suffer from the desired feature of cancer are combined by the computer system. In some cases, the desired feature of the cancer is selected from the group consisting of a type of cancer, a subtype of the cancer of interest, microsatellite instability, mismatch repair proficiency, minimal residual disease, prognostic risk, response to a therapeutic agent, development of resistance to a therapeutic agent cancer.

[0029] In one aspect, provided herein is a computer-implemented method comprising: (a) inputting into a computer system, paired-end DNA sequencing data obtained from a target enrichment sequencing assay performed on a sample comprising cfDNA obtained from a subject, wherein the target enrichment sequencing assay is performed using a panel comprising a plurality of probes generated using any method provided herein; and (b) calculating by the computer system one or more sets of fragment characteristics for sequences from the pair-end DNA sequencing data, thereby generating a fragment characteristic matrix for each regulatory region obtained from the target enrichment sequencing assay for the sample comprising cfDNA obtained from the subject. In some cases, the target enrichment sequencing assay is a hybrid capture enrichment sequencing assay, wherein each probe in the panel binds to and pulls down a regulatory region associated with a target gene from the plurality of target genes prior to being subjected to a sequencing assay. In some cases, the subject suffers from a desired feature of cancer. In some cases, the fragment characteristic matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for the subject represents a sample positive for the desired feature of cancer. In some cases, (a)-(e) are repeated on samples comprising cfDNA obtained from a plurality of subjects that each suffer from the desired feature of cancer. In some cases, the fragment characteristic matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for each of the plurality of subjects that suffer from the desired feature of cancer are combined by computer system. In some cases, the subject does not suffer from the desired feature of cancer. In some cases, the fragment characteristic matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for the subject represents a sample negative for the desired feature of cancer. In some cases, (a)-(e) are repeated on samples comprising cfDNA obtained from a plurality of subjects that each do notAttorney Docket No. GNCN-026 / 01WO 320289-2157 suffer from the desired feature of cancer. In some cases, the fragment characteristic matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for each of the plurality of subjects that do not suffer from the desired feature of cancer are combined by the computer system. In some cases, the desired feature of the cancer is selected from the group consisting of a type of cancer, a subtype of the cancer of interest, microsatellite instability, mismatch repair proficiency, minimal residual disease, prognostic risk, response to a therapeutic agent, development of resistance to a therapeutic agent cancer. In some cases, the fragment characteristics are selected from the group consisting of fragment sequence length, end motif frequency, end motif sequence, jagged end length, fragment diversity, FFT amplitude magnitude and any combination thereof.

[0030] In one aspect, provided herein is a computer-implemented method for generating a classifier for determining a desired feature of cancer, the method comprising: (a) receiving in a computer system, a first set of fragment characteristic indices obtained from cfDNA comprising samples obtained from a first population of subjects suffering from the desired feature of the cancer and a second set of fragment characteristic indices obtained from cfDNA comprising samples obtained from a second population of subjects that do not suffer from the desired feature of the cancer; (b) selecting by the computer system, regulatory regions for model training and validation using the first set and second set of fragment characteristic matrices; and (c) training and validating by the computer system a machine learning algorithm for determining the presence or absence of the desired feature of the cancer in a cfDNA comprising sample, thereby generating a classifier for determining a desired feature of a cancer. In some cases, the method further comprises inputting target enrichment sequencing data obtained for cfDNA comprising sample obtained from a test subject suspected of suffering from the desired feature of the cancer into the classifier generated by (a)-(c), thereby determining the presence or absence of the desired feature of the cancer in the test subject. In some cases, the cfDNA comprising samples are a bodily fluid samples or liquid biopsy samples. In some cases, the desired feature is selected from the group consisting of a type of cancer, a subtype of the cancer of interest, microsatellite instability, mismatch repair proficiency, minimal residual disease, prognostic risk, response to a therapeutic agent, development of resistance to a therapeutic agent cancer. In some cases, wherein the fragment characteristic matrices for the first and the second populations are generated using the computer- implemented method comprising: (a) inputting into a computer system, paired-end DNAAttorney Docket No. GNCN-026 / 01WO 320289-2157 sequencing data obtained from a target enrichment sequencing assay performed on a sample comprising cfDNA obtained from a subject, wherein the target enrichment sequencing assay is performed using a panel comprising a plurality of probes generated using the method of embodiment 102; and (b) calculating by the computer system one or more sets of fragment characteristics for sequences from the pair-end DNA sequencing data, thereby generating a fragment characteristic matrix for each regulatory region obtained from the target enrichment sequencing assay for the sample comprising cfDNA obtained from the subject. In some cases, the target enrichment sequencing assay is a hybrid capture enrichment sequencing assay, wherein each probe in the panel binds to and pulls down a regulatory region associated with a target gene from the plurality of target genes prior to being subjected to a sequencing assay. In some cases, the subject suffers from a desired feature of cancer. In some cases, the fragment characteristic matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for the subject represents a sample positive for the desired feature of cancer. In some cases, (a)-(e) are repeated on samples comprising cfDNA obtained from a plurality of subjects that each suffer from the desired feature of cancer. In some cases, the fragment characteristic matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for each of the plurality of subjects that suffer from the desired feature of cancer are combined by computer system. In some cases, the subject does not suffer from the desired feature of cancer. In some cases, the fragment characteristic matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for the subject represents a sample negative for the desired feature of cancer. In some cases, (a)-(e) are repeated on samples comprising cfDNA obtained from a plurality of subjects that each do not suffer from the desired feature of cancer. In some cases, the fragment characteristic matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for each of the plurality of subjects that do not suffer from the desired feature of cancer are combined by the computer system. In some cases, the desired feature of the cancer is selected from the group consisting of a type of cancer, a subtype of the cancer of interest, microsatellite instability, mismatch repair proficiency, minimal residual disease, prognostic risk, response to a therapeutic agent, development of resistance to a therapeutic agent cancer. In some cases, the fragment characteristics are selected from the group consisting of fragment sequence length, end motif frequency, end motif sequence, jagged end length, fragment diversity, FFT amplitude magnitude and any combination thereof.Attorney Docket No. GNCN-026 / 01WO 320289-2157 BRIEF DESCRIPTION OF THE FIGURES

[0031] FIG. 1 shows genes ranked by cross-platform integrative correlation coefficient (gene reproducibility score; y-axis) across quartiles. Genes with larger positive values behave similarly across the two platforms in that the within-platform correlation coefficients between a given gene and all other genes from the first platform are themselves highly positively correlated with those from the second.

[0032] FIG. 2 illustrates gene-gene correlation coefficients using random subsets of genes from the 1stquartile.

[0033] FIG. 3 illustrates gene-gene correlation coefficients using random subsets of genes from the 2ndquartile.

[0034] FIG. 4 illustrates gene-gene correlation coefficients using random subsets of genes from the 3rdquartile.

[0035] FIG. 5 illustrates gene-gene correlation coefficients using random subsets of genes from the 4thquartile.

[0036] FIG.6 illustrates gene-gene correlation coefficients using the top-ranked 1K genes.

[0037] FIG. 7 illustrates cross-validation curves for generating a methylome data ctDNA-based Colorectal (COAD) microsatellite instability (MSI) yes / no classifier using a >0.4 training approach from top 1K genes (Table 1).

[0038] FIG. 8 illustrates the process for developing a surrogate tissue trained microsatellite stability predictive response signature (MSS-PRS) for liquid (ctDNA) performance that entails using an established tumor tissue-based signature to train a surrogate tumor signature (TCGA) using the top 100 genes of the ~1k of qualified tumor-ctDNA features (left panel in FIG. 8) followed by qualifying the resultant prototype signature by projecting genes in the new signature onto COAD cfDNA (right panel in FIG.8).

[0039] FIG.9 illustrates cross-validation curves for generating a ctDNA-based pancreatic cancer classifier (Basal vs Classical) from top 1K genes (Table 3).

[0040] FIG. 10 illustrates the process for developing a surrogate tissue trained basal / classical pancreatic cancer (PuRIST) subtyper for liquid (ctDNA) performance for a 2ndcohort that entails using an established tumor tissue-based signature (left panel in FIG.10) to train a surrogate tumor signature (TCGA) using ~1k of qualified tumor-ctDNA features (middle panel in FIG. 10)Attorney Docket No. GNCN-026 / 01WO 320289-2157 followed by qualifying the resultant prototype signature by projecting genes in the new signature onto pancreatic cancer cfDNA (right panel in FIG.10).

[0041] FIG. 11 shows genes ranked by cross-platform integrative correlation coefficient (gene reproducibility score; y-axis) across quartiles for a number (i.e., 4) of permutations under the null hypothesis of no RNAseq gene being linked to its fragmentomics partner.

[0042] FIGs 12A-12D illustrates gene-gene correlation plots for a number (i.e., 4) of permutations under the null hypothesis of no RNAseq gene being linked to its fragmentomics partner.

[0043] FIG. 13 illustrates five-fold cross-validation curves for generating a FGFR3 activation signature for determining FGFR3 alteration / activation status (yes / no mutation classifier) from top 1K genes selected for BLCA (Table 5).

[0044] FIG.14 illustrates the process for developing a surrogate tissue trained FGFR3 activation signature (FAS), which can also be referred to as an FGFR predictive response signature (FGFR- PRS),for liquid (ctDNA) performance that entails using an established tumor tissue-based signature to train a surrogate tumor signature (TCGA) using the ~1k of qualified tumor-ctDNA features (left panel in FIG. 14) followed by qualifying the resultant prototype signature by projecting genes in the new signature onto COAD cfDNA (right panel in FIG.14).

[0045] FIG.15 displays the results of a bioinformatic down-sampling experiment where pooled low-pass whole genome sequencing (LP-WGS) data from n = 5 COAD, n = 1 READ, n = 3 BRCA and n = 1 PAAD cancer samples whose DNA was sequenced to an average depth of 8x and was randomly permuted to achieve down-sampling dilutions equivalent to 6x, 4x, 2x, 0.8x, and 0.4x. Following down-sampling, the correlation structure of the “diluted” samples was compared to the original undiluted sample (designated 100%) the Pearson correlation method.

[0046] FIG. 16 shows the three datasets used to collectively demonstrate (1) utilization of any liquid cfDNA dataset by the methods provided herein to generate clinically actionable classifiers and (2) mounting of methods provided herein onto existing assays to expand their clinical utility.

[0047] FIGs 17A-17B shows how nucleosome depletion should associate with increased FFT signal. FIG.17A depicts a cartoon schematic of a “closed” promoter region where nucleosomes across an approximately 800 bp region are present. Whole genome sequencing of cfDNA at this region would result in a wave-like read pile up where majority of reads will map to nucleosome occupied regions. Linker regions will have fewer reads associated with them. FIG.17B depicts a cartoon schematic of an “open” promoter region where a region is depleted of a nucleosome acrossAttorney Docket No. GNCN-026 / 01WO 320289-2157 approximately 800 bp. Whole genome sequencing of cfDNA at this region would result in a wave- like read pile up where the majority of reads will map to nucleosome occupied regions. Linker regions will have fewer reads associated with them and nucleosome depleted regions will resemble linker regions or may have even fewer reads. Amplitude is a measurement of the frequency peak from the DC component (the mean signal strength). The FFT signal will be higher in FIG. 17B than in FIG.17A because the absence of nucleosomes changes the overall signal strength of the wave, decreasing the DC component and thereby increasing the amplitude of the frequency.

[0048] FIG.18 shows demonstrate high variance at the assessed regulatory elements from Wei et al. 2020. “Genome-Wide Profiling of Circulating Tumor DNA Depicts Landscape of Copy Number Alterations in Pancreatic Cancer with Liver Metastasis.” Molecular Oncology 14 (9): 1966–77 samples. Boxplot of the variance of FFT signal at the regulatory elements of 10,904 PDAC-expressed genes per sample. Only n = 6 samples show markedly inconsistent FFT signal across these regulatory elements and were removed from downstream analysis that leverages this dataset.

[0049] FIGs 19A-19B shows tumor content of vLP-WGS dataset as determined by ichorCNA. FIG. 19A shows boxplots of the tumor fraction of samples by sample stage. Stage IV breast samples had significantly higher tumor fraction than stages I – III. FIG. 19B shows boxplots of the tumor fraction of samples by tumor type. The dotted line shows the 75thpercentile of tumor fraction in the NSCLC samples. This 75thpercentile is at a tumor fraction of 0.07 as measured by ichorCNA. This tumor fraction threshold was also applied to the breast samples (dotted line). Samples above the dotted were used in downstream analysis.

[0050] FIGs 20A-20C shows the FFT signal at regulatory elements with high fidelity to their associated genes that were identified and used in downstream analysis. FIG.20A shows boxplots of the quartiles of the FFT signal fidelity score. FIG. 20B shows heatmap of the pairwise correlation of the BRCA RNA expression of the top 100 ranked genes by reproducibility score. FIG. 20C shows a heatmap pairwise correlation of the FFT signal at the regulatory elements of the top 100 ranked genes by reproducibility score. The same genes are listed in FIGs 20B and 20C.

[0051] FIGs 21A-21B shows logistic regression of PCA identifies which highly reproducible features and genes associate with differentiating breast from lung. FIG. 21A shows table of samples included in Logistic regression PCA comparing BRCA vs NSCLC. Degrees of freedomAttorney Docket No. GNCN-026 / 01WO 320289-2157 and significance in association with differentiating breast from lung are shown for PC1 – 10. FIG. 21B shows a heatmap of the FFT signal for the top 50 genes regions (right) and the tissue RNA- seq expression from BRCA, lung adenocarcinoma (LUAD), and lung squamous carcinoma (LUSC) for the same 50 genes (left).

[0052] FIGs 22A-22C shows FFT signal in n = 90 PDAC cfDNA samples at regulatory elements with high fidelity to their associated genes were identified and used in downstream analysis. FIG. 22A show boxplots of the quartiles of the FFT signal fidelity score. Features linked to genes in the PurIST classifier are highlighted in orange. FIG.22B shows heatmaps of the pairwise correlation of the PurIST genes in PDAC RNA expression (left) and cfDNA FFT signal at regulatory elements (right). FIG.22C shows heatmaps of the pairwise correlation of the PDAC RNA expression (left) and FFT signal at the regulatory elements of the top 100 ranked genes by reproducibility score in cfDNA (right).

[0053] FIGs 23A-23C shows training a new basal vs classical PDAC signature using PurIST labels and m = 1000 fidelitous features. FIG. 23A shows cross validation agreement plot of PurIST subtyper with m = 1000 genes with high integrative correlation with FFT signal and using published PurIST calls on PDAC samples. FIG.23B shows a heatmap of gene expression from PDAC samples of genes selected in 5-fold cross validation. Two label bars are shown at the top of the slide heatmap. The top label bar shows the sample subtype according to published PurIST calls; the bottom label bar shows the sample subtype according to the newly trained PurIST subtyper. FIG.23C shows a heatmap of feature strength at regulatory elements of genes selected in 5-fold cross validation. The top label bar shows putative calls for basal and classical based on feature signal and their relative to PDAC gene expression.

[0054] FIG.24 provides a schematic overview of how open chromatin regions in DNA identified by ATAC-seq and RNA transcripts identified by RNA-seq can be correlated and the integrative correlation methods provided herein can distill cfDNA fragments to those mirroring the levels and variance of RNA transcripts in a highly correlated manner.

[0055] FIG.25 describes an approach for identifying promoter and cis-regulatory regions against which to design hybrid capture probes for pull down and library enrichment of fragments reflecting expression of targeted genes and describes the subsequent use of fragmentomics data obtained by targeting the specified promoter and cis-regulatory regions to train an “algorithmic expressionAttorney Docket No. GNCN-026 / 01WO 320289-2157 cluster” as a fragmentomics equivalent to immunohistochemistry for a specific protein target such as HER2.

[0056] FIG.26 depicts a method by which the “algorithmic expression cluster” of FIG.25 could be reported as an “algorithmic expression index” (e.g., HER2 low, medium, or high) for a protein target of therapeutic significance.

[0057] FIG. 27 describes how Simultaneous capture and reporting of gene variant and gene activity information can increase the information content and value of genomic alteration testing on multiple levels.

[0058] FIG.28 depicts how GenomicsNext cfDNA Assay Platform Spans Multiple Applications

[0059] FIG.29 depicts how cfDNA platform provided herein (i.e., GenomicsNextTM) combines DNA alteration testing with a second dimension of gene expression

[0060] FIG.30 depicts technical paths for leveraging the cfDNA platform provided herein (i.e., GenomicsNextTM). DETAILED DESCRIPTION Definitions

[0061] While the following terms are believed to be well understood by one of ordinary skill in the art, the following definitions are set forth to facilitate explanation of the presently disclosed subject matter.

[0062] As used herein, the term “a” or “an” can refer to one or more of that entity, i.e., can refer to a plural referents. As such, the terms “a” or “an”, “one or more” and “at least one” can be used interchangeably herein. In addition, reference to “an element” by the indefinite article “a” or “an” does not exclude the possibility that more than one of the elements is present, unless the context clearly requires that there is one and only one of the elements.

[0063] Unless the context requires otherwise, throughout the present specification and claims, the word “comprise” and variations thereof, such as, “comprises” and “comprising” are to be construed in an open, inclusive sense that is as “including, but not limited to”.

[0064] Reference throughout this specification to “one embodiment” or “an embodiment” means that a particular feature, structure or characteristic described in connection with the embodiment may be included in at least one embodiment of the present disclosure. Thus, the appearances of the phrases “in one embodiment” or “in an embodiment” in various places throughout this specification may not necessarily all referring to the same embodiment. It is appreciated thatAttorney Docket No. GNCN-026 / 01WO 320289-2157 certain features of the disclosure, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the disclosure, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable sub-combination.

[0065] Throughout this disclosure, various aspects of the methods and compositions provided herein can be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the invention. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range.

[0066] Unless otherwise indicated, the methods and compositions provided herein can utilize conventional techniques and descriptions of organic chemistry, polymer technology, molecular biology (including recombinant techniques), cell biology, biochemistry, and immunology, which are within the skill of the art. Such conventional techniques include polymer array synthesis, hybridization, ligation, and detection of hybridization using a label. Specific illustrations of suitable techniques can be had by reference to the example herein below. However, other equivalent conventional procedures can, of course, also be used. Such conventional techniques and descriptions can be found in standard laboratory manuals such as Genome Analysis: A Laboratory Manual Series (Vols. I-IV), Using Antibodies: A Laboratory Manual, Cells: A Laboratory Manual, PCR Primer: A Laboratory Manual, and Molecular Cloning: A Laboratory Manual (all from Cold Spring Harbor Laboratory Press), Gait, "Oligonucleotide Synthesis: A Practical Approach" 1984, IRL Press, London, Nelson and Cox (2000), Lehninger et al., (2008) Principles of Biochemistry 5th Ed., W.H. Freeman Pub., New York, N.Y. and Berg et al. (2006) Biochemistry, 6.sup.th Ed., W.H. Freeman Pub., New York, N.Y., all of which are herein incorporated in their entirety by reference for all purposes.

[0067] Conventional software and systems may also be used in the methods and compositions provided herein. Computer software products of the invention typically include computer readable medium having computer-executable instructions for performing the logic steps of theAttorney Docket No. GNCN-026 / 01WO 320289-2157 method of the invention. Suitable computer readable medium include floppy disk, CD- ROM / DVD / DVD-ROM, hard-disk drive, flash memory, ROM / RAM, magnetic tapes, etc. The computer-executable instructions may be written in a suitable computer language or combination of several languages. Basic computational biology methods are described in, for example, Setubal and Meidanis et al., Introduction to Computational Biology Methods (PWS Publishing Company, Boston, 1997); Salzberg, Searles, Kasif, (Ed.), Computational Methods in Molecular Biology, (Elsevier, Amsterdam, 1998); Rashidi and Buehler, Bioinformatics Basics: Application in Biological Science and Medicine (CRC Press, London, 2000) and Ouelette and Bzevanis Bioinformatics: A Practical Guide for Analysis of Gene and Proteins (Wiley & Sons, Inc., 2.sup.nd ed., 2001). See U.S. Pat. No.6,420,108.

[0068] The methods and compositions provided herein may also make use of various computer program products and software for a variety of purposes, such as probe design, management of data, analysis, and instrument operation. See, U.S. Pat. Nos.5,593,839, 5,795,716, 5,733,729, 5,974,164, 6,066,454, 6,090,555, 6,185,561, 6,188,783, 6,223,127, 6,229,911 and 6,308,170. Computer methods related to genotyping using high density microarray analysis may also be used in the present methods, see, for example, US Patent Pub. Nos.20050250151, 20050244883, 20050108197, 20050079536 and 20050042654.

[0069] Additionally, the present disclosure may have preferred embodiments that include methods for providing genetic information over networks such as the Internet as shown in U.S. Patent Pub. Nos. 20030097222, 20020183936, 20030100995, 20030120432, 20040002818, 20040126840, and 20040049354.

[0070] As used herein, the terms “individual,” “patient,” and “subject” can refer to any single animal, more preferably a mammal (including such non-human animals as, for example, dogs, cats, horses, rabbits, zoo animals, cows, pigs, sheep, and non-human primates) for which treatment is desired. In particular embodiments, the individual or patient herein is a human.

[0071] It will be appreciated that the term "healthy" as used herein, is relative to cancer status, as the term "healthy" cannot be defined to correspond to any absolute evaluation or status. Thus, an individual defined as healthy with reference to any specified disease or disease criterion, can in fact be diagnosed with any other one or more diseases, or exhibit any other one or more disease criterion, including one or more other cancers.Attorney Docket No. GNCN-026 / 01WO 320289-2157

[0072] The term “tumor,” as used herein, can refer to all neoplastic cell growth and proliferation, whether malignant or benign, and all pre-cancerous and cancerous cells and tissues. The terms “cancer,” “cancerous,” and “tumor” are not mutually exclusive and can be used interchangeably.

[0073] The term “detection” can include any means of detecting, including direct and indirect detection.

[0074] The terms “substantially” or “substantial” as used herein can mean substantially similar in level (e.g., expression level), function or capability or otherwise competitive to the products, items (e.g., type of cancer, nucleic acid complement), services or methods recited herein. Substantially similar products, items (e.g., type of cancer, nucleic acid complement), services or methods are at least 80%, 81%, 82%, 83%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% similar or the same as a product, item (e.g., type of cancer, nucleic acid complement), service or method recited herein. Substantially similar level (e.g., type of expression level) are at least 80%, 81%, 82%, 83%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% similar or the same as a level (e.g., type of expression level) recited herein.

[0075] As used herein, the term “FGFR mutation” or “FGFR mutations” can refer to any mutation known in the art in an fgfr gene and / or the protein encoded thereby. Likewise, the term “FGFR3 mutation” or “FGFR3 mutations” can refer to any mutation known in the art in an fgfr3 gene and / or the protein encoded thereby.

[0076] As used herein, an “expression profile” or an “expression pattern” or a “biomarker profile” or a “gene signature” can comprise one or more values corresponding to a measurement of the relative abundance, level, presence, or absence of expression of a discriminative or classifier biomarker or biomarker. An expression profile can be derived from a subject prior to or subsequent to a diagnosis of a type of cancer, can be derived from a biological sample collected from a subject at one or more time points prior to or following treatment or therapy, can be derived from a biological sample collected from a subject at one or more time points during which there is no treatment or therapy (e.g., to monitor progression of disease or to assess development of disease in a subject diagnosed with or at risk for a type of cancer), or can be collected from a healthy subject. The term subject can be used interchangeably with patient. The patient can be a human patient. The one or a plurality of classifier biomarkers that can make up an expression profile asAttorney Docket No. GNCN-026 / 01WO 320289-2157 provided herein can be selected from one or more biomarkers of Tables 1-6 and / or any additional set of biomarker classifiers disclosed herein.

[0077] As used interchangeably herein, a “liquid biopsy” or “fluid biopsy” or “fluid phase biopsy” can comprise the sampling and analysis of non-solid biological tissue. The aforementioned terms can refer to the molecular analysis in biological fluids of nucleic acids, subcellular structures, especially exosomes, and, in the context of cancer, circulating tumor cells. The non-biological tissue can be a bodily fluid. The bodily fluid can be any bodily fluid known in the art such as, for example, blood or fractions thereof, urine, saliva, amniotic fluid, cerebrospinal fluid (CSF) or any combination thereof.

[0078] Quantitation (and grammatical variants thereof) as used herein can refer to characterizing the fragmentation size distributions with a quantitative value. The quantitative value can be an absolute or relative value and can be, without limitation, a number, a statistical value (e.g., frequency, mean, median, standard deviation, or quantile), or a degree or a relative quantity (e.g., high, medium, and low). A quantitative value can be a ratio of two quantitative values. A quantitative value can be a linear combination of quantitative values. A quantitative value may be a normalized value. Overview

[0079] Provided herein are methods, kits and systems for extracting features reflective of gene expression from cell-free DNA (cfDNA) datasets. The cfDNA datasets can comprise sequencing data. The sequencing data can be obtained from any sequencing method used on cfDNA containing samples known in the art. The sequencing data can be obtained from a sequencing method selected from the group consisting of whole genome sequencing (WGS), whole exome sequencing (WES), next generation sequencing (NGS) panel assays, target hybridization capture NGS, whole genome bisulfate sequencing (WGBS) and any combination thereof.

[0080] In one embodiment, provided herein is a method for identifying or extracting features reflective of or corresponding to gene activity or gene expression from cfDNA datasets. The method can comprise (a) obtaining a dataset comprising sequence reads from cfDNA; (b) extracting gene expression information across the dataset by (i) mapping the sequence reads from the dataset to a reference human genome, thereby generating read count data across the mapped genome; and (ii) applying fast Fourier transformation (FFT) to the read count data in nucleosome occupancy windows across the mapped genome to determine an FFT signal at each nucleosomeAttorney Docket No. GNCN-026 / 01WO 320289-2157 occupancy window, wherein an increased FFT signal is indicative of nucleosomal depletion and a decreased FFT signal is indicative of nucleosomal presence; and (c) performing quality control of the cfDNA dataset comprising removal of samples from the dataset that possess a low or inconsistent FFT signal as determined in (b)(i)-(b)(ii) and / or a measured circulating tumor DNA (ctDNA) content below a dataset-specific threshold, thereby generating a quality-controlled cfDNA dataset comprised of features that are reflective of gene activity or expression. The nucleosome occupancy windows can be determined by incorporating published sequencing data from a chromatin accessibility sequencing (CA-seq) method. The CA-seq method can be selected from the group consisting of micrococcal nuclease digestion with deep sequencing (MNase-seq), DNase I hypersensitive sites sequencing (DNase-seq), Formaldehyde-Assisted Isolation of Regulatory Elements sequencing (FAIRE-seq) and Assay for Transposase-Accessible Chromatin using sequencing (ATAC-seq). In one embodiment, the nucleosome occupancy windows are determined from published ATAC-seq data. In one embodiment, the CA-seq method is ARAC- seq can be from tumor tissue(s). Nucleosome occupancy can exhibit an average frequency of one nucleosome approximately every 190 base pairs (bp) (~145 bp of DNA wraps around the nucleosome and another ~45 bp linker DNA). As a consequence, endogenous DNAses in circulation may preferentially degrade DNA that is not protected by nucleosomes. The short linker DNA is unprotected as are any active promoter or cis-regulatory element that are commonly depleted for nucleosomes. Accordingly, these unprotected DNA segments are generally not captured and sequenced by current methods as readily as protected regions. Therefore, mapping the sequence reads obtained from cfDNA across the genome can result in peaks and troughs of read pileups every 190 bps as DNA is protected from nucleosomes (peaks) or subject to DNAse activity (linker). Even more DNA degradation by DNAses may occur resulting in very few mapped reads at nucleosome depleted regions. The result of this preferential degradation at linkers and depleted regions can result in a wave-like signal where peaks of the wave are observed at nucleosome occupied sites and troughs are observed at depleted regions or at linker sites. The FFT signal can be an amplitude of a frequency of wave-like data over each nucleosome occupancy window. The amplitude can be the distance from the DC component, or mean value of the waveform, to the peaks of that frequency. In some cases, the cfDNA datasets are obtained from bodily fluid samples or liquid biopsies obtained from individuals suffering from or suspected of suffering from a cancer. In some cases, ctDNA content, which is reflective of the tumor fraction,Attorney Docket No. GNCN-026 / 01WO 320289-2157 can be estimated, determined or assessed using machine learning. The machine learning can comprise a probabilistic model that uses a hidden Markov model. For example, tumor fraction can be assessed using ichorCNA software or Fragle, which is an ultra-fast deep learning-based method (see Zhu G, et al. A deep-learning model for quantifying circulating tumour DNA from the density distribution of DNA-fragment lengths. Nat Biomed Eng. 2025 Mar;9(3):307-319. doi: 10.1038 / s41551-025-01370-3. PMID: 40055581, which is incorporated by reference in its entirety). Alternatively, if matching tumor tissue mutation data to the cfDNA dataset is available, then the tumor content can be determined or assessed using the ratio of variant allele frequency in the cfDNA vs matched tumor DNA (or median ratio if multiple variants in cfDNA and tumor DNA are present from the same subject).

[0081] In one embodiment, the aforementioned method for identifying or extracting features reflective of or corresponding to gene activity or gene expression from cfDNA datasets is a computer-implemented method, wherein each step of said method is performed by a computer system. In one embodiment, the aforementioned method for identifying or extracting features reflective of or corresponding to gene activity or gene expression from cfDNA datasets may be performed by a system comprising (a) one or more processors; (b) one or more memories or computer-readable medium operatively coupled to at least one of the one or more processors and having instructions stored thereon that, when executed by at least one of the one or more processors, cause the system to performing the aforementioned method for identifying or extracting features reflective of or corresponding to gene activity or gene expression from cfDNA datasets.

[0082] In another embodiment, provided herein is a system for identifying or extracting features reflective of or corresponding to gene activity or gene expression from cfDNA datasets. In some cases, the system comprises (a) one or more processors; (b) one or more memories or computer-readable medium operatively coupled to at least one of the one or more processors and having instructions stored thereon that, when executed by at least one of the one or more processors, cause the system to perform the aforementioned method for identifying or extracting features reflective of or corresponding to gene activity or gene expression from cfDNA datasets.

[0083] Also provided herein are methods, kits and systems for identifying features in cfDNA datasets that reflect gene expression data obtained from tissue samples. The cfDNA datasets can be obtained from bodily fluid samples (e.g., liquid biopsies) obtained from individuals. TheAttorney Docket No. GNCN-026 / 01WO 320289-2157 individuals may suffer from or be suspected of suffering from cancer. The cancer can be any cancer known in the art and / or provided herein. In one embodiment, provided herein are kits that comprise primer sets or probe sets, wherein each primer pair in the primer set or each probe in the probe set comprises sequence complementary to a regulatory region identified using a method provided herein. The kit may further comprise instructions to use said primer or probe sets on samples that comprise cfDNA in a NGS panel assay or target enrichment NGS assay and / or access to means for analyzing data generated from use of said primer or probe sets. In one embodiment, provided herein a system that utilizes primer sets or probe sets, wherein each primer pair in the primer set or each probe in the probe set comprises sequence complementary to a regulatory region identified using a method provided herein. In some cases, the system comprises (a) one or more processors; (b) one or more memories or computer-readable medium operatively coupled to at least one of the one or more processors and having instructions stored thereon that, when executed by at least one of the one or more processors, cause the system to perform a NGS panel assay or target enrichment NGS assay on a sample comprising cfDNA obtained from a subject using the primer set or probe set and transmit and / or analyze the sequencing data generated from the NGS panel assay or target enrichment NGS assay. The subject can be suffering from a cancer. The analysis can entail classifying the subject as possessing a desired feature of the cancer. The desired feature can be a subtype, predictive response signature or microsatellite stability.

[0084] Provided herein are methods, kits and systems for porting or adapting a tissue-based phenotyping test to analytes obtained from a bodily fluid sample (e.g., liquid biopsy) obtained from one or more subjects. The methods, kits and systems provided herein can capture tissue gene expression variance and represent said variance in bodily fluid or liquid biopsy features such as, for example, cfDNA. In some cases, the aforementioned methods, kits and systems can permit the development of cfDNA signatures that are reflective of tumor gene activity in the presence of or independent of a large cohort of matched tissue and cfDNA samples. In one embodiment, the methods, kits and systems for porting or adapting a tissue-based phenotyping test to analytes obtained from a bodily fluid sample (e.g., liquid biopsy) obtained from a subject comprises determining one or a plurality of genes whose expression patterns relative to the expression pattern of each other gene in a first sample obtained from the subject is substantially similar to the expression pattern for the one or plurality of genes in a second sample obtained from the subject relative to the expression pattern of each other gene in the second sample. The one or plurality ofAttorney Docket No. GNCN-026 / 01WO 320289-2157 genes whose expression pattern is determined to be substantially similar in the first and second sample relative to the expression pattern of each other gene in the respective samples can be considered to be portable between the first and second sample. In some cases, the first sample is a tissue sample obtained from the subject. In some cases, the second sample is a bodily fluid sample obtained from the subject. In one embodiment, the first sample obtained from the subject is a tissue sample, while the second subject is a bodily fluid sample obtained from the subject. The tissue sample can be any tissue sample known in the art and / or provided herein. The bodily fluid sample can be any bodily fluid sample known in the art and / or provided herein. In one embodiment, the subject suffers from or is suspected of suffering from a cancer. The cancer can be any cancer known in the art and / or provided herein. In one embodiment, the second sample is a bodily fluid that comprises cell-free DNA (cfDNA) or circulating tumor DNA (ctDNA). In one embodiment, provided herein is a computer-implemented method for determining a gene expression signature in disparate samples from subjects suffering from a cancer, the method comprising: (a) receiving in a computer system, a first set of nucleic acid expression data for each of a plurality of tissue samples obtained from a first population of subjects suffering from or suspected of suffering from the cancer and a second set of nucleic acid expression data for each of a plurality of bodily fluid samples obtained from a second population of subjects suffering from or suspected of suffering from the cancer; (b) determining by the computer system, dependence relationships between each gene within the first set of nucleic acid expression data to generate a first set of dependence relationships and dependence relationships between each gene within the second set of nucleic acid expression data to generate a second set of dependence relationships; and (c) selecting by the computer system, each gene for which the dependence relationships for a respective gene from the first set of dependence relationships is substantially similar to the dependence relationships for the respective gene in the second set of dependence relationships, thereby generating a gene signature that comprises each gene selected by the computer system to possess substantially similar dependence relationships between the first set of nucleic acid expression data for the plurality of tissue samples and the second set of nucleic acid expression data for the plurality of bodily fluid samples. In some cases, each of the plurality of bodily fluid samples comprises cell-free DNA (cfDNA) and, thus, the second set of nucleic acid expression data is a cfDNA dataset. In some cases, the cfDNA dataset is subjected to any method provided herein for identifying or extracting features reflective of or corresponding to gene activity or gene expression from cfDNA datasetsAttorney Docket No. GNCN-026 / 01WO 320289-2157 prior to being subjected to the aforementioned methods, kits and systems for porting or adapting a tissue-based phenotyping test to analytes obtained from a bodily fluid sample (e.g., liquid biopsy) obtained from one or more subjects. For example, prior to being subjected to the aforementioned porting or adapting method, the cfDNA dataset is subjected to a method comprising (a) extracting gene expression information across the dataset by (i) mapping the sequence reads from the dataset to a reference human genome, thereby generating read count data across the mapped genome; and (ii) applying fast Fourier transformation (FFT) to the read count data in nucleosome occupancy windows across the mapped genome to determine an FFT signal at each nucleosome occupancy window, wherein an increased FFT signal is indicative of nucleosomal depletion and a decreased FFT signal is indicative of nucleosomal presence; and (b) performing quality control of the cfDNA dataset comprising removal of samples from the dataset that possess a low or inconsistent FFT signal as determined in (b)(i)-(b)(ii) and / or a measured circulating tumor DNA (ctDNA) content below a dataset-specific threshold, thereby generating a quality-controlled cfDNA dataset comprised of features that are reflective of gene activity or expression.

[0085] The tissue sample can be any tissue sample known in the art and / or provided herein. The bodily fluid sample can be any bodily fluid sample known in the art and / or provided herein. The cancer can be any cancer known in the art and / or provided herein.

[0086] In one embodiment, the dependence relationships between each gene within the first set of nucleic acid expression data comprises all possible pairwise correlations between each gene within the first set of nucleic acid expression data and each other gene within the first set of nucleic acid expression data. In one embodiment, the dependence relationships between each gene within the second set of nucleic acid expression data comprises all possible pairwise correlations between each gene within the second set of nucleic acid expression data and each other gene within the second set of nucleic acid expression data. In one embodiment, the dependence relationships between each gene within the first set of nucleic acid expression data comprises all possible pairwise correlations between each gene within the first set of nucleic acid expression data and each other gene within the first set of nucleic acid expression data and the dependence relationships between each gene within the second set of nucleic acid expression data comprises all possible pairwise correlations between each gene within the second set of nucleic acid expression data and each other gene within the second set of nucleic acid expression data.Attorney Docket No. GNCN-026 / 01WO 320289-2157

[0087] Further to the above embodiments, the substantial similarity in (c) can be evidenced by an integrative correlation coefficient (ICC) for the respective gene that is above a desired threshold. The desired threshold can be the 99thpercentile of a null distribution of integrative correlations. The ICC can be as described in Parmigiani G, Garrett-Mayer ES, Anbazhagan R, Gabrielson E. Cross-study comparison of gene expression data sets for the molecular classification of lung cancer. Clin. Cancer Res. 2004;10:2922, which is incorporated by reference for all purposes. In some cases, the ICC can be a correlation of the pairwise correlations determined for the respective gene within the first set of nucleic acid expression data and the pairwise correlations determined for the respective gene within the second set of nucleic acid expression data.

[0088] In another aspect, provided herein is a computer-implemented method for determining a gene expression signature in disparate samples from subjects suffering from a cancer of interest, the method comprising: (a) receiving in a computer system, a first set of nucleic acid expression data for each of a plurality of tissue samples obtained from a first population of subjects suffering from or suspected of suffering from the cancer of interest and a second set of nucleic acid expression data for each of a plurality of bodily fluid samples obtained from a second population of subjects suffering from or suspected of suffering from the cancer of interest; (b) determining by the computer system, an integrative correlation coefficient for each gene from the first and second set of nucleic acid expression data, wherein the integrative correlation coefficient is a numerical representation of how each gene from the first set of nucleic acid expression data correlate with each other compared to the rank of the same gene from the second set of nucleic acid expression data; (c) ranking by the computer system, the integrative correlation coefficient of each gene from step (b); and (d) selecting, by the computer system, genes that have an integrative correlation coefficient above a desired threshold to generate a gene signature, wherein the gene signature comprises genes whose expression patterns are substantially similar between the tissue sample and the bodily fluid sample. In some cases, each of the plurality of bodily fluid samples comprises cell- free DNA (cfDNA) and, thus, the second set of nucleic acid expression data is a cfDNA dataset. In some cases, the cfDNA dataset is subjected to any method provided herein for identifying or extracting features reflective of or corresponding to gene activity or gene expression from cfDNA datasets prior to being subjected to the aforementioned methods, kits and systems for porting or adapting a tissue-based phenotyping test to analytes obtained from a bodily fluid sample (e.g., liquid biopsy) obtained from one or more subjects. For example, prior to being subjected to theAttorney Docket No. GNCN-026 / 01WO 320289-2157 aforementioned porting or adapting method, the cfDNA dataset is subjected to a method comprising (a) extracting gene expression information across the dataset by (i) mapping the sequence reads from the dataset to a reference human genome, thereby generating read count data across the mapped genome; and (ii) applying fast Fourier transformation (FFT) to the read count data in nucleosome occupancy windows across the mapped genome to determine an FFT signal at each nucleosome occupancy window, wherein an increased FFT signal is indicative of nucleosomal depletion and a decreased FFT signal is indicative of nucleosomal presence; and (b) performing quality control of the cfDNA dataset comprising removal of samples from the dataset that possess a low or inconsistent FFT signal as determined in (b)(i)-(b)(ii) and / or a measured circulating tumor DNA (ctDNA) content below a dataset-specific threshold, thereby generating a quality-controlled cfDNA dataset comprised of features that are reflective of gene activity or expression. The desired threshold can be the 99thpercentile of a null distribution of integrative correlations.

[0089] In another aspect, provided herein is a computer-implemented method for generating a classifier for determining a desired feature of the cancer of interest in a subject suffering from the cancer of interest or suspected of suffering from the cancer of interest that comprises performing the steps as described herein for determining a gene expression signature in disparate samples from subjects suffering from a cancer of interest followed by: inputting nucleic acid expression data for each gene from the gene signature generated from the steps as described herein for determining a gene expression signature in disparate samples from subjects suffering from the cancer of interest from at least two training sets; and conducting on the computer system a linear discriminate analysis (LDA) comprising feature selection to generate a classifier for the desired feature of the cancer of interest. In some cases, one of the at least two training sets comprises nucleic acid expression data for each gene in the gene signature from the previous steps from each of a plurality of tissue samples from subjects that are indicative of the presence a desired feature for the cancer of interest, while another of the at least two training sets comprises nucleic acid expression data for each gene in the gene signature from the previous steps from each of a plurality of tissue samples from subjects that are indicative of the absence of the desired feature for the cancer of interest. In some cases, the feature selection comprises: (i) calculating a test statistic (t-statistic) for each gene in the gene signature from the previous steps from the at least two training sets, wherein the t-statistic for each gene indicates each gene's ability to distinguish between theAttorney Docket No. GNCN-026 / 01WO 320289-2157 presence or the absence of the desired feature; (ii) ranking each gene based on each gene’s t- statistic; and (iii) selecting each gene whose t-statistic is above a desired threshold for distinguishing between the presence or the absence of the desired feature, thereby generating a classifier for the desired feature comprise each of the selected genes. In some cases, the LDA that comprises feature selection that is used in a method, kit or system provided herein is the classification to nearest centroids (ClaNC) method. In some cases, the ClaNC method is a traditional nearest centroid classifier (see Dabney, Alan R. "ClaNC: point-and-click software for classifying microarrays to nearest centroids." Bioinformatics22, no. 1 (2005): 122-123) to determine a set of genes for use in the classifier. Generation of the classifier using training data sets and CLaNC and use of the ordinary nearest centroid classifier to generate a fitted classifier can be as described in the Examples provided herein and / or Dabney A.R.. Classification of microarrays to nearest centroids, Bioinformatics, 2005, vol. 21 (pg. 4148-4154) or Parker JS, Mullins M, Cheang MC, Leung S, Voduc D, Vickery T, Davies S, Fauron C, He X, Hu Z, Quackenbush JF, Stijleman IJ, Palazzo J, Marron JS, Nobel AB, Mardis E, Nielsen TO, Ellis MJ, Perou CM, Bernard PS. Supervised risk predictor of breast cancer based on intrinsic subtypes. J Clin Oncol.2009 Mar 10;27(8):1160-7.

[0090] In some cases, the method for generating a classifier for determining a desired feature of the cancer of interest further comprises: classifying on the computer system one or more test samples obtained from an independent population of subjects suffering the cancer of interest as possessing or not possessing the desired feature using the classifier from (iii) on the one or more test samples; comparing on the computer system, the classification of the one or more test samples to classification of the one or more samples for the desired feature as determined using a control classifier of the desired feature for the cancer of interest; and validating on the computer system the classifier from (iii) if the comparing indicates that the classifications of the one or more test samples is substantially similar to classification of the one or more test samples determined using the control classifier.

[0091] In some cases, the desired feature of the cancer of interest can be selected from the group consisting of a subtype of the cancer of interest, microsatellite instability, mismatch repair proficiency, minimal residual disease, prognostic risk, response to a therapeutic agent, development of resistance to a therapeutic agent and any combination thereof.Attorney Docket No. GNCN-026 / 01WO 320289-2157

[0092] In some cases, the first and the second population of subjects in any method, kits or systems for determining a gene expression signature in disparate samples provided herein consist of the same subjects.

[0093] In some cases, the first and the second population of subjects in any method, kits or systems for generating a classifier for determining a desired feature of the cancer of interest as provided herein consist of the same subjects.

[0094] In some cases, the first set of nucleic acid expression data in any method, kits or systems for determining a gene expression signature in disparate samples provided herein can be nucleic acid sequencing data or methylome sequencing data obtained from nucleic acid extracted from each tissue sample from the plurality of tissue samples. In some cases, the second set of nucleic acid expression data in any method, kits or systems for determining a gene expression signature in disparate samples provided herein can be nucleic acid sequencing data or methylome sequencing data obtained from nucleic acid extracted from each tissue sample from the plurality of tissue samples. The nucleic acid sequencing data can be DNA sequencing data or RNA sequencing data.

[0095] In some cases, the first set of nucleic acid expression data in any method, kits or systems for generating a classifier for determining a desired feature of the cancer of interest provided herein can be nucleic acid sequencing data or methylome sequencing data obtained from nucleic acid extracted from each tissue sample from the plurality of tissue samples. In some cases, the second set of nucleic acid expression data in any method, kits or systems for generating a classifier for determining a desired feature of the cancer of interest provided herein can be nucleic acid sequencing data or methylome sequencing data obtained from nucleic acid extracted from each tissue sample from the plurality of tissue samples. The nucleic acid sequencing data can be DNA sequencing data or RNA sequencing data.

[0096] In some cases, nucleic acid expression data obtained from bodily fluid samples in any method, systems or kits provided herein comprises a fast Fourier transform (FFT) magnitude matrix for each of a plurality of genomic regions of interest. The FFT magnitude matrix for each of the plurality of genomic regions of interest in any method, systems or kits provided herein can be generated by the computer system in any method, systems or kits provided herein that comprises: (i) inputting into the computer system, nucleic acid sequencing data generated from nucleic acid extracted from each of the bodily fluid samples; (ii) determining by the computing system, GC bias values for each fragment read based on the fragment length and the GC contentAttorney Docket No. GNCN-026 / 01WO 320289-2157 of the fragment read; (iii) generating by the computing system, a genomic coverage distribution that is adjusted for GC bias using the sequence read data and the GC bias values; (iv) calculating by the computing system mean sequence read counts for a sliding window across a defined window in each of the plurality of genomic regions of interest from the genomic coverage distribution to generate smoothed mean read counts; and (e) performing a fast Fourier transform (FFT) on the smoothed mead read counts to generate the FFT magnitude matrix for each of the plurality of genomic regions of interest. In some cases, the nucleic acid sequencing data can include a plurality of fragment reads, wherein each fragment read has a fragment length and a GC content indicating a percentage of bases in the fragment read that are G or C.

[0097] In some cases, the sliding window has a width of at least 5, 10, 15, 20, 25, 30, 35, 40, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95 or 100 nucleotides across the defined window. In some cases, the sliding window has a width of exactly 5, 10, 15, 20, 25, 30, 35, 40, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95 or 100 nucleotides across the defined window. In some cases, the sliding window has a width of at most 5, 10, 15, 20, 25, 30, 35, 40, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95 or 100 nucleotides across the defined window. In some cases, the sliding window has a width of 15 nucleotides across the defined window. In some cases, the defined window has a width of 2000 base pairs.

[0098] In one embodiment, the cancer of interest in any method, kits or system for determining a gene expression signature in disparate samples provided herein is colorectal cancer (COAD). Further to this embodiment, can be nucleic acid sequencing data or methylome sequencing data obtained from nucleic acid extracted from each tissue sample from the plurality of tissue samples. In some cases, the first set and the second set of nucleic acid expression data is nucleic acid sequencing data and the gene signature generated comprises the genes in Table 1. In some cases, the first set and the second set of nucleic acid expression data is nucleic acid sequencing data and the gene signature generated consists essentially of the genes in Table 1. In some cases, the first set and the second set of nucleic acid expression data is nucleic acid sequencing data and the gene signature generated consists of the genes in Table 1.

[0099] In one embodiment, the cancer of interest in any method, kits or system for determining a gene expression signature in disparate samples provided herein is pancreatic adenocarcinoma (PAAD). Further to this embodiment, can be nucleic acid sequencing data or methylome sequencing data obtained from nucleic acid extracted from each tissue sample from the plurality of tissue samples. In some cases, the first set and the second set of nucleic acid expression data isAttorney Docket No. GNCN-026 / 01WO 320289-2157 nucleic acid sequencing data and the gene signature generated comprises the genes in Table 3. In some cases, the first set and the second set of nucleic acid expression data is nucleic acid sequencing data and the gene signature generated consists essentially of the genes in Table 3. In some cases, the first set and the second set of nucleic acid expression data is nucleic acid sequencing data and the gene signature generated consists of the genes in Table 3.

[0100] In one embodiment, the cancer of interest in any method, kits or system for determining a gene expression signature in disparate samples provided herein is bladder cancer (BLCA). Further to this embodiment, can be nucleic acid sequencing data or methylome sequencing data obtained from nucleic acid extracted from each tissue sample from the plurality of tissue samples. In some cases, the first set and the second set of nucleic acid expression data is nucleic acid sequencing data and the gene signature generated comprises the genes in Table 5. In some cases, the first set and the second set of nucleic acid expression data is nucleic acid sequencing data and the gene signature generated consists essentially of the genes in Table 5. In some cases, the first set and the second set of nucleic acid expression data is nucleic acid sequencing data and the gene signature generated consists of the genes in Table 5.

[0101] In one embodiment, the cancer of interest in a method for generating a classifier for determining a desired feature of the cancer of interest is COAD and the desired feature is microsatellite instability. In some cases, the classifier generated comprises the genes found in Table 2. In some cases, the classifier generated consists of the genes found in Table 2. In some cases, the classifier generated consists essentially of the genes found in Table 2.

[0102] In one embodiment, the cancer of interest in a method for generating a classifier for determining a desired feature of the cancer of interest is PAAD and the desired feature is a PAAD basal or classical subtype. In some cases, the classifier generated comprises the genes found in Table 4. In some cases, the classifier generated consists of the genes found in Table 4. In some cases, the classifier generated consists essentially of the genes found in Table 4.

[0103] In one embodiment, the cancer of interest in a method for generating a classifier for determining a desired feature of the cancer of interest is BLCA and the desired feature is an FGFR activation signature. The FGFR activation signature can also be referred to as an FGFR predictive response (FGFR-PRS) signature. In some cases, the classifier generated comprises the genes found in Table 6. In some cases, the classifier generated consists of the genes found in Table 6. In some cases, the classifier generated consists essentially of the genes found in Table 6.Attorney Docket No. GNCN-026 / 01WO 320289-2157

[0104] In some cases, any method or system for generating a classifier for determining a desired feature of the cancer of interest provided herein can employ the use of machine learning. The machine learning can be used to perform one or more steps in any of the methods provided herein for generating a classifier for determining a desired feature of the cancer of interest provided herein. In some cases, any method or system provided herein for identifying features in cfDNA datasets that reflect gene expression data obtained from tissue samples can employ the use of machine learning. Machine learning for use in any method or system provided herein can employ algorithms, executed by computer, that automate analytical model building, e.g., for clustering, classification or pattern recognition. The machine learning algorithms for use in any method provided herein may be supervised or unsupervised. Machine learning algorithms for use in any method provided herein can include, for example, artificial neural networks (e.g., back propagation networks), discriminant analyses (e.g., Bayesian classifier or Fischer analysis), support vector machines, decision trees (e.g., recursive partitioning processes such as CART - classification and regression trees, or random forests), linear classifiers (e.g., multiple linear regression (MLR), partial least squares (PLS) regression, and principal components regression), hierarchical clustering, and cluster analysis.

[0105] In one embodiment, a specified statistical confidence level may be determined in order to provide a confidence level regarding any of the classifiers provided herein. For example, it may be determined that a confidence level of greater than 90% may be a useful predictor of any classifier provided herein. In other embodiments, more or less stringent confidence levels may be chosen. For example, a confidence level of about or at least about 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, 99.5%, or 99.9% may be chosen. The confidence level provided may in some cases be related to the quality of the sample, the quality of the data, the quality of the analysis, the specific methods used, and / or the number of gene expression values (i.e., the number of genes or regulatory regions thereof) analyzed. The specified confidence level for providing the likelihood of response may be chosen on the basis of the expected number of false positives or false negatives. Methods for choosing parameters for achieving a specified confidence level or for identifying markers with diagnostic power include but are not limited to Receiver Operating Characteristic (ROC) curve analysis, binormal ROC, principal component analysis, odds ratio analysis, partial least squares analysis, singular value decomposition, least absolute shrinkage andAttorney Docket No. GNCN-026 / 01WO 320289-2157 selection operator analysis, least angle regression, and the threshold gradient directed regularization method.

[0106] Determining any classifier provided herein in some cases can be improved through the application of algorithms designed to normalize and or improve the reliability of the gene expression data. In some embodiments of the present invention, the data analysis utilizes a computer or other device, machine or apparatus for application of the various algorithms described herein due to the large number of individual data points that are processed. A “machine learning algorithm” refers to a computational-based prediction methodology, also known to persons skilled in the art as a “classifier,” employed for characterizing a gene expression profile or profiles, e.g., to determine any classifier provided herein. The biomarker levels, determined by, e.g., microarray- based hybridization assays, sequencing assays, NanoString assays, etc., are in one embodiment subjected to the algorithm in order to classify the profile. In embodiments related to assessing any classifier provided herein, supervised learning generally involves “training” a classifier to recognize the distinctions among a sample positive for a desired feature or a sample negative for a desired feature, and then “testing” the accuracy of the classifier on an independent test set. Therefore, for new, unknown samples the classifier can be used to predict, for example, the class in which the samples belong. The desired feature can be any desired feature described herein. The desired feature can be selected from the group consisting of a type of cancer, a subtype of the cancer of interest, microsatellite instability, mismatch repair proficiency, minimal residual disease, prognostic risk, response to a therapeutic agent, development of resistance to a therapeutic agent cancer.

[0107] Various other software programs may be implemented for any method or system provided herein. In certain methods, feature selection and model estimation may be performed by logistic regression with lasso penalty using glmnet (Friedman et al. (2010). Journal of statistical software 33(1): 1-22, incorporated by reference in its entirety). Raw reads may be aligned using TopHat (Trapnell et al. (2009). Bioinformatics 25(9): 1105-11, incorporated by reference in its entirety). In methods, top features (N ranging from 10 to 200) are used to train a linear support vector machine (SVM) (Suykens JAK, Vandewalle J. Least Squares Support Vector Machine Classifiers. Neural Processing Letters 1999; 9(3): 293-300, incorporated by reference in its entirety) using the e1071 library (Meyer D. Support vector machines: the interface to libsvm in package e1071. 2014, incorporated by reference in its entirety). Confidence intervals, in oneAttorney Docket No. GNCN-026 / 01WO 320289-2157 embodiment, are computed using the pROC package (Robin X, Turck N, Hainard A, et al. pROC: an open-source package for R and S+ to analyze and compare ROC curves. BMC bioinformatics 2011; 12: 77, incorporated by reference in its entirety).

[0108] In addition, data may be filtered to remove data that may be considered suspect.

[0109] In some embodiments of the present invention, data from probe-sets obtained using any of the methods provided herein may be excluded from analysis if they are not identified at a detectable level (above background).

[0110] In some embodiments of the present disclosure, probe-sets obtained using any of the methods provided herein that exhibit no, or low variance may be excluded from further analysis. Low-variance probe-sets are excluded from the analysis via a Chi-Square test. In one embodiment, a probe-set is considered to be low-variance if its transformed variance is to the left of the 99 percent confidence interval of the Chi-Squared distribution with (N-l) degrees of freedom. (N-l)*Probe-set Variance / (Gene Probe-set Variance). Chi-Sq(N-l) where N is the number of input CEL files, (N-l) is the degrees of freedom for the Chi-Squared distribution, and the “probe-set variance for the gene” is the average of probe-set variances across the gene. In some embodiments of the present invention, probe-sets for a given gene or regulatory region thereof may be excluded from further analysis if they contain less than a minimum number of probes that pass through filter steps for GC content, reliability, variance and the like.

[0111] In some cases, accuracy of any classifier generated using the methods provided herein may be determined by tracking the subject over time to determine the accuracy of the original diagnosis or classification. In other cases, accuracy may be established in a deterministic manner or using statistical methods. For example, receiver operator characteristic (ROC) analysis may be used to determine the optimal assay parameters to achieve a specific level of accuracy, specificity, positive predictive value, negative predictive value, and / or false discovery rate.

[0112] In some cases, the results of the classifier assay as provided herein, are entered into a database for access by representatives or agents of a molecular profiling business, the individual, a medical provider, or insurance provider. In some cases, assay results include sample classification, identification, or diagnosis by a representative, agent or consultant of the business, such as a medical professional. In other cases, a computer or algorithmic analysis of the data is provided automatically. In some cases, the molecular profiling business may bill the individual, insurance provider, medical provider, researcher, or government entity for one or more of theAttorney Docket No. GNCN-026 / 01WO 320289-2157 following: molecular profiling assays performed, consulting services, data analysis, reporting of results, or database access.

[0113] In some embodiments of the present invention, the results of the classifier assay provided herein are presented as a report on a computer screen or as a paper record. In some embodiments, the report may include, but is not limited to, such information as one or more of the following: the levels of classifier biomarkers as compared to the reference sample or reference value(s); the likelihood the subject will respond to a particular therapy and / or will be resistant or non-responsive to a particular therapy, based on the classifier biomarker values and proposed therapies.

[0114] In some embodiments of the present invention, results (e.g., from a hybrid-capture enrichment sequencing assay on a cfDNA containing sample as provided herein) are classified using a trained algorithm. Trained algorithms of the present invention can include algorithms that have been developed using a reference set of samples from subjects known to possess a desired feature of a cancer and / or normal samples from subjects known not to possess a desired feature of a cancer.

[0115] Algorithms suitable for categorization of samples include but are not limited to k- nearest neighbor algorithms, k-top scoring pairs (TSPs), top scoring pairs (TSPs), support vector machines, linear discriminant analysis, CLaNC, diagonal linear discriminant analysis, updown, naive Bayesian algorithms, neural network algorithms, hidden Markov model algorithms, genetic algorithms, or any combination thereof.

[0116] When a binary classifier is compared with actual true values (e.g., values from a biological sample), there are typically four possible outcomes. If the outcome from a prediction is p (where “p” is a positive classifier output, such as the presence of a deletion or duplication syndrome) and the actual value is also p, then it is called a true positive (TP); however, if the actual value is n then it is said to be a false positive (FP). Conversely, a true negative has occurred when both the prediction outcome and the actual value are n (where “n” is a negative classifier output, such as no deletion or duplication syndrome), and false negative is when the prediction outcome is n while the actual value is p. In one embodiment, consider a test that seeks to determine whether a person is likely or unlikely to respond to a target gene (e.g., ERBB2) inhibitor therapy. A false positive in this case occurs when the person tests positive, but actually does respond. A false negative, on the other hand, occurs when the person tests negative, suggesting they are unlikely toAttorney Docket No. GNCN-026 / 01WO 320289-2157 respond, when they actually are likely to respond. The same holds true for classifying any one of or a combination of classifiers for a desired feature provided herein.

[0117] The positive predictive value (PPV), or precision rate, or post-test probability of disease, is the proportion of subjects with positive test results who are correctly diagnosed possessing a desired feature of a cancer or not. It reflects the probability that a positive test reflects the underlying condition being tested for. Its value does however depend on the prevalence of the disease, which may vary. In one example the following characteristics are provided: FP (false positive); TN (true negative); TP (true positive); FN (false negative). False positive rate( )=FP / (FP+TN)-specificity; False negative rate ( )=FN / (TP+FN)-sensitivity; Power = sensitivity= 1- ; Likelihood-ratio positive=sensitivity / (l-specificity); Likelihood-ratio negative = (1 - sensitivity) / specificity. The negative predictive value (NPV) is the proportion of subjects with negative test results who are correctly diagnosed.

[0118] In some embodiments, the results of any classifier analysis provided herein can provide a statistical confidence level that a given diagnosis is correct. In some embodiments, such statistical confidence level is at least about, or more than about 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% 99.5%, or more.

[0119] In some embodiments, the methods described herein that employ a classifier developed using any of the method provided herein are capable of classifying a subject as possessing a desired feature of a cancer with a predictive success or accuracy of at least about 70%, at least about 71%, at least about 72%, about 73%, about 74%, about 75%, about 76%, about 77%, about 78%, about 79%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, up to 100%, and all values in between. The cancer can be any cancer known in the art and / or provided herein. The desired feature can be any desired feature described herein. The desired feature can be selected from the group consisting of a type of cancer, a subtype of the cancer of interest, microsatellite instability, mismatch repair proficiency, minimal residual disease, prognostic risk, response to a therapeutic agent, development of resistance to a therapeutic agent cancer. The sample obtained from the subject for use in any method described herein that employs a classifier developed using any of the method provided herein can comprise cfDNA. Moreover, the cfDNA can have a fraction of ctDNA. The ctDNA fraction can be from about 0.000001 to about 0.01, such as about 0.000005Attorney Docket No. GNCN-026 / 01WO 320289-2157 to about 0.01, about 0.00001 to about 0.01, about 0.00005 to about 0.01, about 0.0001 to about 0.01, about 0.0005 to about 0.01, 0.000001 to about 0.005, about 0.000005 to about 0.005, about 0.00005 to about 0.005, about 0.00005 to about 0.005, about 0.0001 to about 0.005, about 0.0005 to about 0.005, 0.000001 to about 0.001, about 0.000005 to about 0.001, about 0.00001 to about 0.001, about 0.00005 to about 0.001, about 0.0001 to about 0.001, about 0.0005 to about 0.001, about 0.000001 to about 0.0005, about 0.000005 to about 0.0005, about 0.00001 to about 0.0005, about 0.00005 to about 0.0005, about 0.0001 to about 0.0005, about 0.0005 to about 0.0005, about 0.000001 to about 0.0001, about 0.000005 to about 0.0001, about 0.00001 to about 0.0001, about 0.00005 to about 0.0001, about 0.000001 to about 0.00005, about 0.000005 to about 0.00005, about 0.00001 to about 0.00005, about 0.000001 to about 0.00001, about 0.000005 to about 0.00001, or about 0.000001 to about 0.000005. In some cases, the ct fraction can be from about 0.001 to about 0.25, such as from about 0.005 to about 0.25, from about 0.01 to about 0.25, from about 0.05 to about 0.25, from about 0.1 to about 0.25, from about 0.001 to about 0.1, from about 0.005 to about 0.1, from about 0.01 to about 0.1, from about 0.05 to about 0.1, from about 0.001 to about 0.05, from about 0.005 to about 0.05, from about 0.01 to about 0.05, from about 0.001 to about 0.01, from about 0.005 to about 0.01, or from about 0.001 to about 0.005.

[0120] In some embodiments, any classifier method provided herein can include classifying the sample as being positive or negative for a desired feature based on the comparison of classifier biomarker levels in the sample and reference biomarker levels, for example present in at least one training set. In some embodiments, the sample is classified as being positive or negative if the results of the comparison meet one or more criterion such as, for example, a minimum percent agreement, a value of a statistic calculated based on the percentage agreement such as (for example) a kappa statistic, a minimum correlation (e.g., Pearson’s correlation) and / or the like.

[0121] It is intended that the methods described herein can be performed by software (stored in memory and / or executed on hardware), hardware, or a combination thereof. Hardware modules may include, for example, a general-purpose processor, a field programmable gate array (FPGA), and / or an application specific integrated circuit (ASIC). Software modules (executed on hardware) can be expressed in a variety of software languages (e.g., computer code), including Unix utilities, C, C++, Java™, Ruby, SQL, SAS®, the R programming language / software environment, Visual Basic™, and other object-oriented, procedural, or other programmingAttorney Docket No. GNCN-026 / 01WO 320289-2157 language and development tools. Examples of computer code include, but are not limited to, micro-code or micro-instructions, machine instructions, such as produced by a compiler, code used to produce a web service, and files containing higher-level instructions that are executed by a computer using an interpreter. Additional examples of computer code include, but are not limited to, control signals, encrypted code, and compressed code.

[0122] Some embodiments described herein relate to devices with a non-transitory computer-readable medium (also can be referred to as a non-transitory processor-readable medium or memory) having instructions or computer code thereon for performing various computer- implemented operations and / or methods disclosed herein. The computer-readable medium (or processor-readable medium) is non-transitory in the sense that it does not include transitory propagating signals per se (e.g., a propagating electromagnetic wave carrying information on a transmission medium such as space or a cable). The media and computer code (also can be referred to as code) may be those designed and constructed for the specific purpose or purposes. Examples of non-transitory computer-readable media include, but are not limited to: magnetic storage media such as hard disks, floppy disks, and magnetic tape; optical storage media such as Compact Disc / Digital Video Discs (CD / DVDs), Compact Disc-Read Only Memories (CD-ROMs), and holographic devices; magneto-optical storage media such as optical disks; carrier wave signal processing modules; and hardware devices that are specially configured to store and execute program code, such as Application-Specific Integrated Circuits (ASICs), Programmable Logic Devices (PLDs), Read-Only Memory (ROM) and Random-Access Memory (RAM) devices. Other embodiments described herein relate to a computer program product, which can include, for example, the instructions and / or computer code discussed herein. Prognostic Uses

[0123] In one aspect, provided herein is a method for determining a disease outcome in a subject suffering from or suspected of suffering from cancer. The cancer can be any cancer known in the art and / or provided herein. The prognostic information that can be obtained by the methods provided herein can comprise a number of possible endpoints, which can be selected from time from surgery to distant metastases (distant recurrence-free survival), time of disease-free survival (recurrence free survival), time of progression-free survival (progression free survival) and time of overall survival. In some cases, Kaplan-Meier plots (Kaplan and Meier. J Am Stat Assoc 53:Attorney Docket No. GNCN-026 / 01WO 320289-2157 457-481 (1958)) can be used to display time-to-event curves for any or all of these three endpoints. In some cases, a cox regression (or proportional hazards regression) can be performed in order to determine a hazard ratio for any or all of these three endpoints. In one embodiment, a cox regression (or proportional hazards regression) is used to assess the prognostic performance in terms of overall survival of sample positive or negative for a desired feature of a cancer as determined using the methods provided herein. The Cox Proportional Hazards analysis is a regression method for survival data that provides an estimate of the hazard ratio and its confidence interval. The Cox model is a well-recognized statistical technique for exploring the relationship between the survival of a subject and particular variables. This statistical method permits estimation of the hazard (i.e., risk) of individuals given their prognostic variables (e.g., target gene (e.g., ERBB2) activation status with or without other additional clinical factors, as described herein). The "hazard ratio" is the risk of death at any given time point for patients displaying particular prognostic variables. See generally Spruance et al., Antimicrob. Agents & Chemo. 48:2787-92 (2004). The additional clinical factors can include age, sex, tumor diameter, tumor stage and smoking history. A relevant time interval or time point can be at least 1 year, at least two years, at least three years, at least five years, or at least ten years. Sample Types

[0124] The sample can be any sample known in the art and / or provided herein. In one embodiment, the tissue sample for any method, kit or system provided herein is obtained from an individual and comprises formalin-fixed paraffin-embedded (FFPE) tissue. However, other tissue and sample types are amenable for use herein. In one embodiment, the other tissue and sample types can be fresh frozen tissue or cell pellets, or the like. In one embodiment, the bodily fluid sample for any method, kit or system provided herein can be a bodily fluid or a liquid biopsy obtained from the individual. The bodily fluid can be blood or fractions thereof (e.g., serum, plasma), urine, sputum, saliva, wash fluids or cerebrospinal fluid (CSF). A biomarker nucleic acid as provided herein can be extracted from a cell or can be cell free or extracted from an extracellular vesicular entity such as an exosome.

[0125] In one embodiment, the tissue sample for any method, kit or system provided herein comprises cells harvested from a tissue sample. Cells can be harvested from a biological sample using standard techniques known in the art. For example, in one embodiment, cells are harvestedAttorney Docket No. GNCN-026 / 01WO 320289-2157 by centrifuging a cell sample and resuspending the pelleted cells. The cells can be resuspended in a buffered solution such as phosphate-buffered saline (PBS). After centrifuging the cell suspension to obtain a cell pellet, the cells can be lysed to extract nucleic acid, e.g., messenger RNA. All samples obtained from a subject, including those subjected to any sort of further processing, are considered to be obtained from the subject.

[0126] Methods are known in the art for the isolation of RNA from FFPE tissue. In one embodiment, total RNA can be isolated from FFPE tissues as described by Bibikova et al. (2004) American Journal of Pathology 165:1799-1807, herein incorporated by reference. Likewise, the High Pure RNA Paraffin Kit (Roche) can be used. Paraffin is removed by xylene extraction followed by ethanol wash. RNA can be isolated from sectioned tissue blocks using the MasterPure Purification kit (Epicenter, Madison, Wis.); a DNase I treatment step is included. RNA can be extracted from frozen samples using Trizol reagent according to the supplier's instructions (Invitrogen Life Technologies, Carlsbad, Calif.). Samples with measurable residual genomic DNA can be resubjected to DNaseI treatment and assayed for DNA contamination. All purification, DNase treatment, and other steps can be performed according to the manufacturer's protocol. After total RNA isolation, samples can be stored at -80 ºC until use.

[0127] General methods for mRNA extraction from tissue and non-tissue-based sources are well known in the art and are disclosed in standard textbooks of molecular biology, including Ausubel et al., ed., Current Protocols in Molecular Biology, John Wiley & Sons, New York 1987- 1999. Methods for RNA extraction from paraffin embedded tissues are disclosed, for example, in Rupp and Locker (Lab Invest.56:A67, 1987) and De Andres et al. (Biotechniques 18:42-44, 1995). In particular, RNA isolation can be performed using a purification kit, a buffer set and protease from commercial manufacturers, such as Qiagen (Valencia, Calif.), according to the manufacturer's instructions. For example, total RNA from cells in culture can be isolated using Qiagen RNeasy mini-columns. Other commercially available RNA isolation kits include MasterPureTM. Complete DNA and RNA Purification Kit (Epicentre, Madison, Wis.) and Paraffin Block RNA Isolation Kit (Ambion, Austin, Tex.). Total RNA from tissue samples can be isolated, for example, using RNA Stat-60 (Tel-Test, Friendswood, Tex.). RNA prepared from a tumor can be isolated, for example, by cesium chloride density gradient centrifugation. Additionally, large numbers of tissue samples can readily be processed using techniques well known to those of skillAttorney Docket No. GNCN-026 / 01WO 320289-2157 in the art, such as, for example, the single-step RNA isolation process of Chomczynski (U.S. Pat. No.4,843,155, incorporated by reference in its entirety for all purposes).

[0128] The sample, in one embodiment, is further processed before the detection of the biomarker levels any of the combination of biomarkers set forth herein. For example, mRNA in a cell, exosome or tissue sample can be separated from other components of the sample. The sample can be concentrated and / or purified to isolate mRNA in its non-natural state, as the mRNA is not in its natural environment. For example, studies have indicated that the higher order structure of mRNA in vivo differs from the in vitro structure of the same sequence (see, e.g., Rouskin et al. (2014). Nature 505, pp.701-705, incorporated herein in its entirety for all purposes).

[0129] mRNA from the sample in one embodiment, is hybridized to a synthetic DNA probe, which in some embodiments, includes a detection moiety (e.g., detectable label, capture sequence, barcode reporting sequence). Accordingly, in these embodiments, a non-natural mRNA-cDNA complex is ultimately made and used for detection of the biomarker. In another embodiment, mRNA from the sample is directly labeled with a detectable label, e.g., a fluorophore. In a further embodiment, the non-natural labeled-mRNA molecule is hybridized to a cDNA probe and the complex is detected.

[0130] In one embodiment, once the mRNA is obtained from a sample, it is converted to complementary DNA (cDNA) prior to the hybridization reaction or is used in a hybridization reaction together with one or more cDNA probes. cDNA does not exist in vivo and therefore is a non-natural molecule. Furthermore, cDNA-mRNA hybrids are synthetic and do not exist in vivo. Besides cDNA not existing in vivo, cDNA is necessarily different than mRNA, as it includes deoxyribonucleic acid and not ribonucleic acid. The cDNA is then amplified, for example, by the polymerase chain reaction (PCR) or other amplification method known to those of ordinary skill in the art. For example, other amplification methods that may be employed include the ligase chain reaction (LCR) (Wu and Wallace, Genomics, 4:560 (1989), Landegren et al., Science, 241:1077 (1988), incorporated by reference in its entirety for all purposes, transcription amplification (Kwoh et al., Proc. Natl. Acad. Sci. USA, 86:1173 (1989), incorporated by reference in its entirety for all purposes), self-sustained sequence replication (Guatelli et al., Proc. Nat. Acad. Sci. USA, 87:1874 (1990), incorporated by reference in its entirety for all purposes), incorporated by reference in its entirety for all purposes, and nucleic acid based sequence amplification (NASBA). Guidelines for selecting primers for PCR amplification are known to those of ordinaryAttorney Docket No. GNCN-026 / 01WO 320289-2157 skill in the art. See, e.g., McPherson et al., PCR Basics: From Background to Bench, Springer- Verlag, 2000, incorporated by reference in its entirety for all purposes. The product of this amplification reaction, i.e., amplified cDNA is also necessarily a non-natural product. First, as mentioned above, cDNA is a non-natural molecule. Second, in the case of PCR, the amplification process serves to create hundreds of millions of cDNA copies for every individual cDNA molecule of starting material. The numbers of copies generated are far removed from the number of copies of mRNA that are present in vivo.

[0131] In one embodiment, cDNA is amplified with primers that introduce an additional DNA sequence (e.g., adapter, reporter, capture sequence or moiety, barcode) onto the fragments (e.g., with the use of adapter-specific primers), or mRNA or cDNA biomarker sequences are hybridized directly to a cDNA probe comprising the additional sequence (e.g., adapter, reporter, capture sequence or moiety, barcode). Amplification and / or hybridization of mRNA to a cDNA probe therefore serves to create non-natural double stranded molecules from the non-natural single stranded cDNA, or the mRNA, by introducing additional sequences and forming non-natural hybrids. Further, as known to those of ordinary skill in the art, amplification procedures have error rates associated with them. Therefore, amplification introduces further modifications into the cDNA molecules. In one embodiment, during amplification with the adapter-specific primers, a detectable label, e.g., a fluorophore, is added to single strand cDNA molecules. Amplification therefore also serves to create DNA complexes that do not occur in nature, at least because (i) cDNA does not exist in vivo, (i) adapter sequences are added to the ends of cDNA molecules to make DNA sequences that do not exist in vivo, (ii) the error rate associated with amplification further creates DNA sequences that do not exist in vivo, (iii) the disparate structure of the cDNA molecules as compared to what exists in nature, and (iv) the chemical addition of a detectable label to the cDNA molecules.

[0132] In some embodiments, the expression of a biomarker of interest is detected at the nucleic acid level via detection of non-natural cDNA molecules.

[0133] In one embodiment, the bodily fluid sample provided in any method, kit or system provided herein has nucleic extracted therefrom as part of the method, kit or system. The nucleic acid can be DNA, RNA or cDNA. In some cases, the nucleic acid is cell-free DNA (cfDNA) and / or circulating tumor DNA (ctDNA). In some cases, the nucleic acid extracted from the bodily fluid sample is subjected to nucleic acid amplification (e.g., PCR) and / or sequencing (e.g., nextAttorney Docket No. GNCN-026 / 01WO 320289-2157 generation sequencing, whole genome sequencing (WGS) or whole exome sequencing (WES) or sequencing data from an NGS-panel assay) as part of the method, kit or system provided herein. In some cases, the methods, kits and / or systems provided here can entail running the Griffin pipeline on sequencing data (e.g., whole genome sequencing (WGS) or whole exome sequencing (WES) data) obtained from nucleic acid (e.g., cfDNA or ctDNA) extracted from the sample obtained from the subject as outlined in Doebley, AL., Ko, M., Liao, H. et al. Author Correction: A framework for clinical cancer subtyping from nucleosome profiling of cell-free DNA. Nat Commun 14, 403 (2023). Following the generation of GC corrected coverage data around gene regions of interest, the method, kits and / or systems can entail calculating the mean read counts for a sliding window in each region, applying fast Fourier transform (FFT) to the mean data and extracting the highest magnitude.

[0134] The FFT step can entail taking the 132 count values in a 2KB region of interest (-990 ~ +975, 15bp interval); performing fast Fourier transform with fft() in R; and extracting the first element (max magnitude) from the FFT output using Mod(). The amplitude can be encoded as the magnitude of the complex number. The desired signal can be the one with smallest frequency, and highest magnitude. Cancer Types

[0135] The cancer from which a subject may be suffering from or suspected of suffering from in a method, kit or system provided herein can be any cancer known in the art and / or provided herein. The cancer can include, but is not limited to, carcinoma, lymphoma, blastoma (including medulloblastoma and retinoblastoma), sarcoma (including liposarcoma and synovial cell sarcoma), neuroendocrine tumors (including carcinoid tumors, gastrinoma, and islet cell cancer), mesothelioma, schwannoma (including acoustic neuroma), meningioma, adenocarcinoma, melanoma, and leukemia or lymphoid malignancies. Examples of a cancer can also include, but is not limited to, a lung cancer (e.g., a non-small cell lung cancer (NSCLC) or small cell lung cancer), a kidney cancer (e.g., a kidney urothelial carcinoma or RCC), a bladder cancer (e.g., a bladder urothelial (transitional cell) carcinoma (e.g., locally advanced or metastatic urothelial cancer, including 1L or 2L+ locally advanced or metastatic urothelial carcinoma)), a breast cancer, a colorectal cancer (e.g., a colon adenocarcinoma), an ovarian cancer, a pancreatic cancer, a gastric carcinoma, an esophageal cancer, a mesothelioma, a melanoma (e.g., a skin melanoma), a headAttorney Docket No. GNCN-026 / 01WO 320289-2157 and neck cancer (e.g., a head and neck squamous cell carcinoma (HNSCC)), a thyroid cancer, a sarcoma (e.g., a soft-tissue sarcoma, a fibrosarcoma, a myxosarcoma, a liposarcoma, an osteogenic sarcoma, an osteosarcoma, a chondrosarcoma, an angiosarcoma, an endotheliosarcoma, a lymphangiosarcoma, a lymphangioendotheliosarcoma, a leiomyosarcoma, or a rhabdomyosarcoma), a prostate cancer, a glioblastoma, a cervical cancer, a thymic carcinoma, a leukemia (e.g., an acute lymphocytic leukemia (ALL), an acute myelocytic leukemia (AML), a chronic myelocytic leukemia (CML), a chronic eosinophilic leukemia, or a chronic lymphocytic leukemia (CLL)), a lymphoma (e.g., a Hodgkin lymphoma or a non-Hodgkin lymphoma (NHL)), a myeloma (e.g., a multiple myeloma (MM)), a mycosis fungoides, a Merkel cell cancer, a hematologic malignancy, a cancer of hematological tissues, a B cell cancer, a bronchus cancer, a stomach cancer, a brain or central nervous system cancer, a peripheral nervous system cancer, a uterine or endometrial cancer, a cancer of the oral cavity or pharynx, a liver cancer, a testicular cancer, a biliary tract cancer, a small bowel or appendix cancer, a salivary gland cancer, an adrenal gland cancer, an adenocarcinoma, an inflammatory myofibroblastic tumor, a gastrointestinal stromal tumor (GIST), a colon cancer, a myelodysplastic syndrome (MDS), a myeloproliferative disorder (MPD), a polycythemia Vera, a chordoma, a synovioma, a Ewing’s tumor, a squamous cell carcinoma, a basal cell carcinoma, a sweat gland carcinoma, a sebaceous gland carcinoma, a papillary carcinoma, a papillary adenocarcinoma, a medullary carcinoma, a bronchogenic carcinoma, a renal cell carcinoma, a hepatoma, a bile duct carcinoma, a choriocarcinoma, a seminoma, an embryonal carcinoma, a Wilms' tumor, a bladder carcinoma, an epithelial carcinoma, a glioma, an astrocytoma, a medulloblastoma, a craniopharyngioma, an ependymoma, a pinealoma, a hemangioblastoma, an acoustic neuroma, an oligodendroglioma, a meningioma, a neuroblastoma, a retinoblastoma, a follicular lymphoma, a diffuse large B-cell lymphoma, a mantle cell lymphoma, a hepatocellular carcinoma, a thyroid cancer, a small cell cancer, an essential thrombocythemia, an agnogenic myeloid metaplasia, a hypereosinophilic syndrome, a systemic mastocytosis, a familiar hypereosinophilia, a neuroendocrine cancer, or a carcinoid tumor.

[0136] In one embodiment, the cancer is selected from kidney renal papillary cell carcinoma (KIRP); breast invasive carcinoma (BRCA); thyroid cancer (THCA); bladder urothelial carcinoma (BLCA); prostate adenocarcinoma (PRAD); kidney chromophobe (KICH); cervical squamous cell carcinoma and endocervical adenocarcinoma (CESC); kidney renal clear cellAttorney Docket No. GNCN-026 / 01WO 320289-2157 carcinoma (KIRC); liver hepatocellular carcinoma (LIHC); low grade glioma (LGG); sarcoma (SARC); lung adenocarcinoma (LUAD); colon adenocarcinoma (COAD); head and neck squamous cell carcinoma (HNSC); uterine corpus endometrial carcinoma (UCEC); glioblastoma multiforme (GBM); esophageal carcinoma (ESCA); stomach adenocarcinoma (STAD); ovarian serous cystadenocarcinoma (OV); rectum adenocarcinoma (READ); adrenocortical carcinoma (ACC); uveal melanoma (UVM); mesothelioma (MESO); pheochromocytoma and paraganglioma (PCPG); skin cutaneous melanoma (SKCM); uterine carcinosarcoma (UCS); lung squamous cell carcinoma (LUSC); testicular germ cell tumors (TGCT); cholangiocarcinoma (CHOL); pancreatic adenocarcinoma (PAAD); thymoma (THYM); Lymphoid Neoplasm Diffuse Large B-cell Lymphoma (DLBC); and Acute Myeloid Leukemia [LAML]. In another embodiment, the cancer is selected from kidney renal papillary cell carcinoma (KIRP); breast invasive carcinoma (BRCA); thyroid cancer (THCA); bladder urothelial carcinoma (BLCA); prostate adenocarcinoma (PRAD); kidney chromophobe (KICH); cervical squamous cell carcinoma and endocervical adenocarcinoma (CESC); kidney renal clear cell carcinoma (KIRC); liver hepatocellular carcinoma (LIHC); low grade glioma (LGG); sarcoma (SARC); lung adenocarcinoma (LUAD); colon adenocarcinoma (COAD); head and neck squamous cell carcinoma (HNSC); uterine corpus endometrial carcinoma (UCEC); glioblastoma multiforme (GBM); esophageal carcinoma (ESCA); stomach adenocarcinoma (STAD); ovarian serous cystadenocarcinoma (OV); rectum adenocarcinoma (READ); adrenocortical carcinoma (ACC); uveal melanoma (UVM); mesothelioma (MESO); pheochromocytoma and paraganglioma (PCPG); skin cutaneous melanoma (SKCM); uterine carcinsarcoma (UCS); lung squamous cell carcinoma (LUSC); testicular germ cell tumors (TGCT); cholangiocarcinoma (CHOL); pancreatic adenocarcinoma (PAAD); thymoma (THYM); and Lymphoid Neoplasm Diffuse Large B-cell Lymphoma (DLBC). Clinical / Therapeutic Uses

[0137] In one embodiment, a method is provided herein for determining a disease outcome or prognosis for a patient suffering from a cancer. The cancer can be any cancer known in the art and / or provided herein. The disease outcome or prognosis can be measured by examining the overall survival for a period of time or intervals (e.g., 0 to 36 months or 0 to 60 months). In one embodiment, survival is analyzed as a function of the desired feature. The desired feature can be selected from the group consisting of a type of cancer, a subtype of the cancer of interest,Attorney Docket No. GNCN-026 / 01WO 320289-2157 microsatellite instability, mismatch repair proficiency, minimal residual disease, prognostic risk, response to a therapeutic agent, development of resistance to a therapeutic agent cancer. Relapse- free and overall survival can be assessed using standard Kaplan-Meier plots as well as Cox proportional hazards modeling.

[0138] In one embodiment, the methods, kits and systems as provided herein for determining a desired feature of a patient suffering or suspected of suffering from a cancer of interest is used to determine whether or not said patient is a candidate for treatment with a specific type or types of cancer therapy.

[0139] In one embodiment, upon determining the desired feature of the cancer of interest, the patient is selected for suitable therapy, for example, radiotherapy (radiation therapy), surgical intervention, target therapy, chemotherapy or drug therapy with an angiogenesis inhibitor or immunotherapy or combinations thereof. In some embodiments, the suitable treatment can be any treatment or therapeutic method that can be used for a patient suffering from the cancer of interest. In one embodiment, upon determining a patient’s desired feature of the cancer of interest, the patient is administered a suitable therapeutic agent, for example chemotherapeutic agent(s) or an angiogenesis inhibitor or immunotherapeutic agent(s). In one embodiment, the therapy is immunotherapy, and the immunotherapeutic agent is a checkpoint inhibitor, monoclonal antibody, biological response modifier, therapeutic vaccine or cellular immunotherapy. In some embodiments, the determination of a suitable treatment can identify treatment responders. In some embodiments, the determination of a suitable treatment can identify treatment non responders. In some embodiments, upon determining a patient’s desired feature of the cancer of interest, the cancer patient can be selected for any combination of suitable therapies. For example, chemotherapy or drug therapy with a radiotherapy, a tumor dissection with an immunotherapy or a chemotherapeutic agent with a radiotherapy. In some embodiments, immunotherapy, or immunotherapeutic agent can be a checkpoint inhibitor, monoclonal antibody, biological response modifier, therapeutic vaccine or cellular immunotherapy.

[0140] The methods of present invention are also useful for evaluating clinical response to therapy, as well as for endpoints in clinical trials for efficacy of new therapies. The extent to which sequential diagnostic expression profiles move towards normal can be used as one measure of the efficacy of the candidate therapy. ImmunotherapyAttorney Docket No. GNCN-026 / 01WO 320289-2157

[0141] In one embodiment, provided herein is a method for determining whether a cancer patient is likely to respond to immunotherapy by determining the desired feature of the cancer of interest in a sample or samples obtained from the patient and, based on the desired feature ascertained for the sample, assessing whether the patient is likely to respond to immunotherapy. The immunotherapy can be any immunotherapy provided herein. In one embodiment, the immunotherapy comprises administering one or more checkpoint inhibitors. The checkpoint inhibitors can be any checkpoint inhibitor provided herein such as, for example, a checkpoint inhibitor that targets PD-l, PD-LI or CTLA4.

[0142] In one embodiment, upon determining a patient’s desired feature of the cancer of interest using any of the methods and classifier biomarkers panels or subsets thereof as provided herein, the patient is selected for treatment with or administered an immunotherapeutic agent. The immunotherapeutic agent can be a checkpoint inhibitor, monoclonal antibody, biological response modifiers, therapeutic vaccine or cellular immunotherapy.

[0143] In another embodiment, the immunotherapeutic agent is a checkpoint inhibitor. In some cases, a method for determining the likelihood of response to one or more checkpoint inhibitors is provided. In one embodiment, the checkpoint inhibitor is a PD-l / PD-LI checkpoint inhibitor. The PD-l / PD-LI checkpoint inhibitor can be nivolumab, pembrolizumab, atezolizumab, durvalumab, lambrolizumab, or avelumab. In one embodiment, the checkpoint inhibitor is a CTLA-4 checkpoint inhibitor. The CTLA-4 checkpoint inhibitor can be ipilimumab or tremelimumab. In one embodiment, the checkpoint inhibitor is a combination of checkpoint inhibitors such as, for example, a combination of one or more PD-l / PD-LI checkpoint inhibitors used in combination with one or more CTLA-4 checkpoint inhibitors.

[0144] In one embodiment, the immunotherapeutic agent is a monoclonal antibody. In some cases, a method for determining the likelihood of response to one or more monoclonal antibodies is provided. The monoclonal antibody can be directed against tumor cells or directed against tumor products. The monoclonal antibody can be panitumumab, matuzumab, necitumunab, trastuzumab, amatuximab, bevacizumab, ramucirumab, bavituximab, patritumab, rilotumumab, cetuximab, immu-l32, or demcizumab.

[0145] In yet another embodiment, the immunotherapeutic agent is a therapeutic vaccine. In some cases, a method for determining the likelihood of response to one or more therapeutic vaccines is provided. The therapeutic vaccine can be a peptide or tumor cell vaccine. The vaccineAttorney Docket No. GNCN-026 / 01WO 320289-2157 can target MAGE-3 antigens, NY-ESO-l antigens, p53 antigens, survivin antigens, or MUC1 antigens. The therapeutic cancer vaccine can be GVAX (GM- CSF gene-transfected tumor cell vaccine), belagenpumatucel-L (allogeneic tumor cell vaccine made with four irradiated NSCLC cell lines modified with TGF-beta2 antisense plasmid), MAGE- A3 vaccine (composed of MAGE- A3 protein and adjuvant AS 15), (l)-BLP- 25 anti-MUC-l (targets MUC-l expressed on tumor cells), CimaVax EGF (vaccine composed of human recombinant Epidermal Growth Factor (EGF) conjugated to a carrier protein), WT1 peptide vaccine (composed of four Wilms’ tumor suppressor gene analogue peptides), CRS-207 (live-attenuated Listeria monocytogenes vector encoding human mesothelin), Bec2 / BCG (induces anti-GD3 antibodies), GV1001 (targets the human telomerase reverse transcriptase), TG4010 (targets the MUC1 antigen), racotumomab (anti- idiotypic antibody which mimicks the NGcGM3 gangboside that is expressed on multiple human cancers), tecemotide (liposomal BLP25; liposome-based vaccine made from tandem repeat region of MUC1) or DRibbles (a vaccine made from nine cancer antigens plus TLR adjuvants).

[0146] In one embodiment, the immunotherapeutic agent is a biological response modifier. In some cases, a method for determining the likelihood of response to one or more biological response modifiers is provided. The biological response modifier can trigger inflammation such as, for example, PF-3512676 (CpG 7909) (a toll-like receptor 9 agonist), CpG-ODN 2006 (downregulates Tregs), Bacillus Calmette-Guerin (BCG), mycobacterium vaccae (SRL172) (nonspecific immune stimulants now often tested as adjuvants). The biological response modifier can be cytokine therapy such as, for example, IL-2+ tumor necrosis factor alpha (TNF-alpha) or interferon alpha (induces T-cell proliferation), interferon gamma (induces tumor cell apoptosis), or Mda-7 (IL-24) (Mda-7 / IL-24 induces tumor cell apoptosis and inhibits tumor angiogenesis). The biological response modifier can be a colony-stimulating factor such as, for example granulocyte colony-stimulating factor. The biological response modifier can be a multi-modal effector such as, for example, multi-target

[0147] VEGFR: thalidomide and analogues such as lenalidomide and pomalidomide, cyclophosphamide, cyclosporine, denileukin diftitox, talactoferrin, trabecetedin or all-trans- retinmoic acid.

[0148] In one embodiment, the immunotherapy is cellular immunotherapy. In some cases, a method for determining the likelihood of response to one or more cellular therapeutic agents. The cellular immunotherapeutic agent can be dendritic cells (DCs) (ex vivo generated DC-Attorney Docket No. GNCN-026 / 01WO 320289-2157 vaccines loaded with tumor antigens), T-cells (ex vivo generated lymphokine-activated killer cells; cytokine-induce killer cells; activated T-cells; gamma delta T-cells), or natural killer cells. Angiogenesis Inhibitors

[0149] In general, methods of determining whether a patient is likely to respond to angiogenesis inhibitor therapy, or methods of selecting a patient for angiogenesis inhibitor therapy are provided herein. In one embodiment, upon determining a patient’s or subject’s desired feature of a cancer of interest, the patient is selected for drug therapy with an angiogenesis inhibitor.

[0150] In one embodiment, the angiogenesis inhibitor is a vascular endothelial growth factor (VEGF) inhibitor, a VEGF receptor inhibitor, a platelet derived growth factor (PDGF) inhibitor or a PDGF receptor inhibitor.

[0151] In one embodiment, angiogenesis inhibitor treatments include, but are not limited to an integrin antagonist, a selectin antagonist, an adhesion molecule antagonist, an antagonist of intercellular adhesion molecule (ICAM)-l, IC AM-2, IC AM-3, platelet endothelial adhesion molecule (PC AM), vascular cell adhesion molecule (VCAM)), lymphocyte function-associated antigen 1 (LFA-l), a basic fibroblast growth factor antagonist, a vascular endothelial growth factor (VEGF) modulator, a platelet derived growth factor (PDGF) modulator (e.g., a PDGF antagonist).

[0152] In one embodiment of determining whether a subject is likely to respond to an integrin antagonist, the integrin antagonist is a small molecule integrin antagonist, for example, an antagonist described by Paolillo et al. (Mini Rev Med Chem, 2009, volume 12, pp. 1439-1446, incorporated by reference in its entirety), or a leukocyte adhesion-inducing cytokine or growth factor antagonist (e.g., tumor necrosis factor-a (TNF-a), interleukin- 1b (IL- 1 b). monocyte chemotactic protein-l (MCP-l) and a vascular endothelial growth factor (VEGF)), as described in U.S. Patent No.6,524,581, incorporated by reference in its entirety herein.

[0153] The methods provided herein are also useful for determining whether a subject is likely to respond to one or more of the following angiogenesis inhibitors: interferon gamma 1b, interferon gamma 1b (Actimmune®) with pirfenidone, ACUHTR028, anb5, aminobenzoate potassium, amyloid P, ANG1122, ANG1170, ANG3062, ANG3281, ANG3298, ANG4011, anti- CTGF RNAi, Aplidin, astragalus membranaceus extract with salvia and schisandra chinensis, atherosclerotic plaque blocker, Azol, AZX100, BB3, connective tissue growth factor antibody, CT140, danazol, Esbriet, EXC001, EXC002, EXC003, EXC004, EXC005, F647, FG3019, Fibrocorin, Follistatin, FT011, a galectin-3 inhibitor, GKT137831, GMCT01, GMCT02,Attorney Docket No. GNCN-026 / 01WO 320289-2157 GRMD01, GRMD02, GRN510, Heberon Alfa R, interferon a-2b, ITMN520, JKB119, JKB121, JKB122, KRX168, LPA1 receptor antagonist, MGN4220, MIA2, microRNA 29a oligonucleotide, MMI0100, noscapine, PBI4050, PBI4419, PDGFR inhibitor, PF-06473871, PGN0052, Pirespa, Pirfenex, pirfenidone, plitidepsin, PRM151, Pxl02, PYN17, PYN22 with PYN17, Relivergen, rhPTX2 fusion protein, RXI109, secretin, STX100, TGF-b Inhibitor, transforming growth factor, b- receptor 2 oligonucleotide, VA999260, XV615 or a combination thereof.

[0154] In another embodiment, a method is provided for determining whether a subject is likely to respond to one or more endogenous angiogenesis inhibitors. In a further embodiment, the endogenous angiogenesis inhibitor is endostatin, a 20 kDa C-terminal fragment derived from type XVIII collagen, angiostatin (a 38 kDa fragment of plasmin), a member of the thrombospondin (TSP) family of proteins. In a further embodiment, the angiogenesis inhibitor is a TSP-l, TSP-2, TSP-3, TSP-4 and TSP-5. Methods for determining the likelihood of response to one or more of the following angiogenesis inhibitors are also provided a soluble VEGF receptor, e.g., soluble VEGFR-l and neuropilin 1 (NPR1), angiopoietin-l, angiopoietin-2, vasostatin, calreticulin, platelet factor-4, a tissue inhibitor of metalloproteinase (TIMP) (e.g., TIMP1, TIMP2, TIMP3, TIMP4), cartilage- derived angiogenesis inhibitor (e.g., peptide troponin I and chrondomodulin I), a disintegrin and metalloproteinase with thrombospondin motif 1, an interferon (IFN), (e.g., IFN-a, IFN-b, IFN-g), a chemokine, e.g., a chemokine having the C-X-C motif (e.g., CXCL10, also known as interferon gamma-induced protein 10 or small inducible cytokine B10), an interleukin cytokine (e.g, IL-4, IL-12, IL-18), prothrombin, antithrombin III fragment, prolactin, the protein encoded by the TNFSF15 gene, osteopontin, maspin, canstatin, proliferin-related protein.

[0155] In one embodiment, a method for determining the likelihood of response to one or more of the following angiogenesis inhibitors is provided is angiopoietin-l, angiopoietin-2, angiostatin, endostatin, vasostatin, thrombospondin, calreticulin, platelet factor-4, TIMP, CDAI, interferon a, interferon b, vascular endothelial growth factor inhibitor (VEGI) meth-l, meth-2, prolactin, VEGI, SPARC, osteopontin, maspin, canstatin, proliferin-related protein (PRP), restin, TSP-l, TSP-2, interferon gamma 1b, ACUHTR028, anb5, aminobenzoate potassium, amyloid P, ANG1122, ANG1170, ANG3062, ANG3281, ANG3298, ANG4011, anti-CTGF RNAi, Aplidin, astragalus membranaceus extract with salvia and schisandra chinensis, atherosclerotic plaque blocker, Azol, AZX100, BB3, connective tissue growth factor antibody, CT140, danazol, Esbriet, EXC001, EXC002, EXC003, EXC004, EXC005, F647, FG3019, Fibrocorin, Follistatin, FT011, aAttorney Docket No. GNCN-026 / 01WO 320289-2157 galectin-3 inhibitor, GKT137831, GMCT01, GMCT02, GRMD01, GRMD02, GRN510, Heberon Alfa R, interferon a-2b, ITMN520, JKB119, JKB121, JKB122, KRX168, LPA1 receptor antagonist, MGN4220, MIA2, microRNA 29a oligonucleotide, MMI0100, noscapine, PBI4050, PBI4419, PDGFR inhibitor, PF-06473871, PGN0052, Pirespa, Pirfenex, pirfenidone, plitidepsin, PRM151, Pxl02, PYN17, PYN22 with PYN17, Relivergen, rhPTX2 fusion protein, RXI109, secretin, STX100, TGF-b Inhibitor, transforming growth factor, b-receptor 2 oligonucleotide, VA999260, XV615 or a combination thereof.

[0156] In yet another embodiment, the angiogenesis inhibitor can include pazopanib (Votrient), sunitinib (Sutent), sorafenib (Nexavar), axitinib (Inlyta), ponatinib (Iclusig), vandetanib (Caprelsa), cabozantinib (Cometrig), ramucirumab (Cyramza), regorafenib (Stivarga), ziv-aflibercept (Zaltrap), motesanib, or a combination thereof. In another embodiment, the angiogenesis inhibitor is a VEGF inhibitor. In a further embodiment, the VEGF inhibitor is axitinib, cabozantinib, aflibercept, brivanib, tivozanib, ramucirumab or motesanib. In yet a further embodiment, the angiogenesis inhibitor is motesanib.

[0157] In one embodiment, the methods provided herein relate to determining a subject’s likelihood of response to an antagonist of a member of the platelet derived growth factor (PDGF) family, for example, a drug that inhibits, reduces or modulates the signaling and / or activity of PDGF -receptors (PDGFR). For example, the PDGF antagonist, in one embodiment, is an anti- PDGF aptamer, an anti-PDGF antibody or fragment thereof, an anti- PDGFR antibody or fragment thereof, or a small molecule antagonist. In one embodiment, the PDGF antagonist is an antagonist of the PDGFR-a or PDGFR-b. In one embodiment, the PDGF antagonist is the anti-PDGF-b aptamer El 0030, sunitinib, axitinib, sorefenib, imatinib, imatinib mesylate, nintedanib, pazopanib HC1, ponatinib, MK-2461, dovitinib, pazopanib, crenolanib, PP-121, telatinib, imatinib, KRN 633, CP 673451, TSU-68, Ki875l, amuvatinib, tivozanib, masitinib, motesanib diphosphate, dovitinib dilactic acid, bnifanib (ABT-869).

[0158] Upon making a determination of whether a patient is likely to respond to angiogenesis inhibitor therapy, or selecting a patient for angiogenesis inhibitor therapy, in one embodiment, the patient is administered the angiogenesis inhibitor. The angiogenesis in inhibitor can be any of the angiogenesis inhibitors described herein. RadiotherapyAttorney Docket No. GNCN-026 / 01WO 320289-2157

[0159] In one embodiment, provided herein is a method for determining whether a patient is likely to respond to radiotherapy by determining the desired feature of interest for a cancer of interest of a sample obtained from the patient and, based on the desired feature of interest for a cancer of interest assessing whether the patient is likely to respond to or benefit from radiotherapy. In another embodiment, provided herein is a method of selecting a patient suffering from cancer for radiotherapy by determining a desired feature of interest for a cancer of interest of a sample from the patient and, based on the desired feature of interest for a cancer of interest selecting the patient for radiotherapy.

[0160] In some embodiments, the radiotherapy can include but are not limited to proton therapy and external-beam radiation therapy. In some embodiments, the radiotherapy can include any types or forms of treatment that is suitable for patients with specific types of cancer.

[0161] In some embodiments, a patient with a specific type of cancer can have or display resistance to radiotherapy. Radiotherapy resistance in any cancer or subtype thereof can be determined by measuring or detecting the expression levels of one or more genes known in the art and / or provided herein associated with or related to the presence of radiotherapy resistance. Genes associated with radiotherapy resistance can include NFE2L2, KEAP1 and CUL3. In some embodiments, radiotherapy resistance can be associated with the alterations of KEAP1 (Kelch-like ECH-associated protein l) / NRF2 (nuclear factor E2-related factor 2) pathway. Association of a particular gene to radiotherapy resistance can be determined by examining expression of said gene in one or more patients known to be radiotherapy non responders and comparing expression of said gene in one or more patients known to be radiotherapy responders. Surgical Intervention

[0162] In one embodiment, provided herein is a method for determining whether a cancer patient is likely to respond to surgical intervention by determining the patient’s desired feature of a cancer of interest from a sample or samples obtained from the patient and, based on the patient’s desired feature of the cancer of interest, assessing whether the patient is likely to respond to or benefit from surgery. In another embodiment, provided herein is a method of selecting a patient suffering from cancer for surgery by determining the patient’s desired feature of a cancer of interest from a sample or samples from the patient and, based on the patient’s desired feature of a cancer of interest, selecting the patient for surgery. In some embodiments, the surgery can include laser technology, excision, dissection, and reconstructive surgery.Attorney Docket No. GNCN-026 / 01WO 320289-2157 Microsatellite Instability and Colorectal Cancer

[0163] In one embodiment, the desired feature is microsatellite instability (MSI) or conversely, microsatellite stability (MSS) in a cancer of interest in a method, kit or system as provided herein for determining a desired feature of a patient suffering or suspected of suffering from the cancer of interest. Further to this embodiment, provided herein are methods, kits and systems for determining the presence or absence of microsatellite instability (MSI) or conversely, microsatellite stability (MSS) in subjects suffering from the cancer of interest. In one embodiment, the presence or absence of MSI or MSS is determined by determining a MSS predictive response signature (MSS-PRS) of a subject suffering from the cancer. In some cases, the cancer is colon adenocarcinoma (COAD) and the MSS-PRS is the expression profile of a plurality of biomarkers selected from the biomarkers listed in Table 2. The MSS-PRS was developed to combine both MSI and mismatch repair (MMR) to capture tumors, independent of mutation status and overt MSI defects that may harbor MMR defects that facilitate an immuno-oncology (IO) or other modality response. In some cases, the MSS-PRS provided herein can determine a subtype of a subject suffering from cancer (e.g., COAD) that phenocopies MSI and / or MMR deficient (i.e., dMMR). In some cases, determining a subject’s MSS-PRS gene signature can inform whether or not the subject is a candidate or a potential responder for immuno-oncology (IO) treatment with or without combo treatments with therapeutics in clinical trials if the subject’s MSS-PRS signature is MSS- PRS positive or is a candidate or a potential responder for 5-FU Chemotherapy treatment if the subject’s MSS-PRS signature is MSS-PRS negative. In one embodiment, provided herein is a method for treating a subject suffering from cancer comprising determining the subject’s MSS- PRS gene signature and administering a therapeutic agent to the subject based on the subject’s MSS-PRS signature. The determining step comprises measuring the expression of a plurality of biomarkers selected from the biomarkers in Table 2 in a sample obtained from the subject. The expression pattern of said plurality of biomarkers selected from the biomarkers in Table 2 reflects the subject’s COAD MSS-PRS. If the subject’s MSS-PRS signature is MSS-PRS positive, the subject is administered immune checkpoint inhibitor (ICI) therapy. If the subject’s MSS-PRS signature is MSS-PRS positive, the subject is administered immuno-oncology (IO) treatment with or without combo treatments with therapeutics in clinical trials. If the subject’s MSS-PRS signature is MSS-PRS negative, the subject is administered a chemotherapy agent such as, for example, 5-FU.Attorney Docket No. GNCN-026 / 01WO 320289-2157

[0164] Exemplary immune checkpoint inhibitors for use in any method provided herein can be any checkpoint inhibitor known in the art and / or provided herein. The checkpoint inhibitors can be any checkpoint inhibitor provided herein such as, for example, a checkpoint inhibitor that targets PD-1, PD-LI or CTLA4.

[0165] In some cases, the determining step further comprises comparing the nucleic acid expression levels of the plurality of classifier biomarkers from Table 2 to the nucleic acid expression levels of the plurality of classifier biomarkers from Table 2 in at least one sample training set(s), and classifying the sample obtained from the subject as MSI positive or MSI negative based on the results of the comparing step. The at least one sample training set can comprise nucleic acid expression level data of the plurality of classifier biomarkers from Table 2 from a reference MSI positive sample, nucleic acid expression level data of the plurality of classifier biomarkers from Table 2 from a reference MSI negative sample or a combination thereof. In one embodiment, the comparing step comprises applying a statistical algorithm which comprises determining a correlation between the expression data obtained from the sample obtained from the subject and the expression data from the at least one training set(s); and classifying the sample obtained from the subject as MSI positive or MSI negative based on the results of the statistical algorithm. The statistical algorithm can be any statistical algorithm known in the art and / or provided herein for performing the necessary correlation. In some cases, the statistical algorithm is performed by a computer system.

[0166] In some cases, the plurality of classifiers selected from Table 2 comprises at least 2 classifiers, at least 4 classifiers, at least 6 classifiers, at least 8 classifiers, at least 10 classifier biomarkers, at least 12 classifiers, at least 14 classifiers, at least 16 classifiers, at least 18 classifiers, at least 20 classifiers, at least 22 classifiers, at least 24 classifiers, at least 28 classifiers, at least 30 classifiers, at least 32 classifiers, at least 34 classifiers, at least 36 classifiers, at least 38 classifiers, at least 40 classifiers, at least 42 classifiers, at least 44 classifiers, at least 46 classifiers, at least 48 classifiers, at least 50 classifiers, at least 52 classifiers, at least 54 classifiers, at least 56 classifiers, at least 58 classifiers, at least 60 classifiers, at least 62 classifiers, at least 64 classifiers, at least 66 classifiers or at least 68 classifiers from Table 2.

[0167] In some cases, the plurality of classifiers selected from Table 2 comprises at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%,Attorney Docket No. GNCN-026 / 01WO 320289-2157 at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% of the classifiers from Table 2.

[0168] In some cases, the plurality of classifiers selected from Table 2 consists of all the classifiers of Table 2. Pancreatic Adenocarcinoma (PAAD) Subtyping

[0169] In one embodiment, the cancer of interest is PAAD and the desired feature is a subtype of PAAD. The subtype of PAAD can be a basal or classical subtype. Further to this embodiment, provided herein are methods, kits and systems for determining the subtype of PAAD for a patient suffering from PAAD and, based on the determined subtype of PAAD, administering a specific treatment to the patient.

[0170] In one embodiment, provided herein is a method of treating PAAD in a subject, the method comprising: determining a subtype of PAAD of a subject suffering from PAAD by measuring a nucleic acid expression level of a plurality of classifier biomarkers in a sample obtained from a subject suffering from or suspected of suffering from PAAD, wherein the plurality of classifier biomarkers is selected from Table 4, wherein the nucleic acid expression level of the plurality of classifier biomarkers indicates the subtype of PAAD as being classical or basal and administering a therapeutic intervention based on the subtype of PAAD. In some cases, the subject is determined to possess a classical subtype of PAAD and is administered 5-flourouracil and platinum-based therapy. In some cases, the subject is determined to possess a classical subtype of PAAD and is administered one or more therapeutic agents listed for the classical in Table 7. In some cases, the subject is determined to possess a basal subtype of PAAD and is administered cisplatin- or oxaliplatin-based therapies and gemcitabine. In some cases, the subject is determined to possess a basal subtype of PAAD and is administered one or more therapeutic agents listed for the basal subtype in Table 7.

[0171] Table 7. Exemplary Chemotherapeutics Applicable To Specific PAAD SubtypesAttorney Docket No. GNCN-026 / 01WO 320289-2157Attorney Docket No. GNCN-026 / 01WO 320289-2157

[0172] In some cases, the determining step further comprises comparing the nucleic acid expression levels of the plurality of classifier biomarkers from Table 4 to the nucleic acid expression levels of the plurality of classifier biomarkers from Table 4 in at least one sample training set(s), and classifying the sample obtained from the subject as PAAD classical or PAAD basal based on the results of the comparing step. The at least one sample training set can comprise nucleic acid expression level data of the plurality of classifier biomarkers from Table 4 from a reference PAAD classical sample, nucleic acid expression level data of the plurality of classifier biomarkers from Table 4 from a reference PAAD basal sample or a combination thereof. In some cases, the comparing step comprises applying a statistical algorithm which comprises determining a correlation between the expression data obtained from the sample obtained from the subject and the expression data from the at least one training set(s); and classifying the sample obtained from the subject as PAAD classical or PAAD basal based on the results of the statistical algorithm. The statistical algorithm can be any statistical algorithm known in the art and / or provided herein for performing the necessary correlation. In some cases, the statistical algorithm is performed by a computer system.

[0173] In some cases, the plurality of classifiers selected from Table 4 comprises at least 2 classifiers, at least 4 classifiers, at least 6 classifiers, at least 8 classifiers, at least 10 classifierAttorney Docket No. GNCN-026 / 01WO 320289-2157 biomarkers, at least 12 classifiers, at least 14 classifiers, at least 16 classifiers, at least 18 classifiers, at least 20 classifiers, at least 22 classifiers, at least 24 classifiers, at least 28 classifiers, at least 30 classifiers, at least 32 classifiers, at least 34 classifiers, at least 36 classifiers, at least 38 classifiers, at least 40 classifiers, at least 42 classifiers, at least 44 classifiers, at least 46 classifiers, at least 48 classifiers, at least 50 classifiers, at least 52 classifiers, at least 54 classifiers, at least 56 classifiers, at least 58 classifiers, at least 60 classifiers, at least 62 classifiers, at least 64 classifiers, at least 66 classifiers, at least 68 classifiers, at least 70 classifiers, at least 72 classifiers, at least 74 classifiers, at least 76 classifiers, at least 78 classifiers or at least 80 classifiers from Table 4.

[0174] In some cases, the plurality of classifiers selected from Table 4 comprises at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% of the classifiers from Table 4.

[0175] In some cases, the plurality of classifiers selected from Table 4 consists of all the classifiers of Table 4. FGFR Predictive Response Signature

[0176] In one embodiment, the cancer of interest is bladder cancer (BLCA) and the desired feature is an FGFR activation signature that is indicative of the presence of FGFR3 mutations. Further to this embodiment, provided herein are methods, kits and systems for determining the FGFR activation signature of BLCA for a patient suffering from BLCA and, based on the determined subtype of BLCA, administering a specific treatment to the patient.

[0177] In one embodiment, a method is provided herein for determining a disease outcome or prognosis for a patient suffering from bladder cancer. In some cases, the bladder cancer is Muscle Invasive Bladder Cancer (MIBC). The disease outcome or prognosis can be measured by examining the overall survival for a period of time or intervals (e.g., 0 to 36 months or 0 to 60 months). In one embodiment, survival is analyzed as a function of bladder cancer FGFR activation signature. In one embodiment, survival is analyzed as a function of FGFR activation signature. The bladder cancer FGFR activation signature can be determined using the methods provided herein such as, for example, determining the expression of all or subsets of the genes in Table 6. Relapse-free and overall survival can be assessed using standard Kaplan-Meier plots as well as Cox proportional hazards modeling.Attorney Docket No. GNCN-026 / 01WO 320289-2157

[0178] In one embodiment, upon determining a patient’s bladder cancer FGFR activation signature (e.g., by measuring the expression of all or subsets of the genes in Table 6), the patient is selected for suitable therapy, for example, radiotherapy (radiation therapy), surgical intervention, target therapy, chemotherapy or drug therapy with an angiogenesis inhibitor or immunotherapy or combinations thereof. In some embodiments, the suitable treatment can be any treatment or therapeutic method that can be used for a bladder cancer patient. In one embodiment, upon determining a patient’s bladder cancer FGFR activation signature is negative, the patient is administered a suitable therapeutic agent, for example chemotherapeutic agent(s) or an angiogenesis inhibitor or immunotherapeutic agent(s). In one embodiment, the therapy is immunotherapy, and the immunotherapeutic agent is a checkpoint inhibitor, monoclonal antibody, biological response modifier, therapeutic vaccine or cellular immunotherapy. In some embodiments, the determination of a suitable treatment can identify treatment responders. In some embodiments, the determination of a suitable treatment can identify treatment non responders. In some embodiments, upon determining a patient’s bladder cancer FGFR activation signature, the bladder cancer patient can be selected for any combination of suitable therapies. For example, chemotherapy or drug therapy with a radiotherapy, a tumor dissection with an immunotherapy or a chemotherapeutic agent with a radiotherapy. In some embodiments, immunotherapy, or immunotherapeutic agent can be a checkpoint inhibitor, monoclonal antibody, biological response modifier, therapeutic vaccine or cellular immunotherapy.

[0179] In one embodiment, provided herein is a method of treating BLCA in a subject, the method comprising: measuring a nucleic acid expression level of a plurality of classifiers in a sample obtained from a subject suffering from or suspected of suffering from BLCA, wherein the plurality of classifiers is selected from Table 6, wherein the measured nucleic acid expression levels of the plurality of classifiers provide an FGFR activation signature (e.g., FGFR3 activation signature) for the sample; and administering an FGFR inhibitor based on presence of a positive FGFR activation signature (e.g., FGFR3 activation signature), wherein the positive FGFR activation signature (e.g., FGFR3 activation signature) is indicative of presence of one or more FGFR3 mutations. In one embodiment, absence of a positive FGFR activation signature (FAS), i.e., a negative FAS of a sample obtained from a subject suffering from or suspected of suffering from a cancer indicates that the subject may be responsive to a therapeutic agent or defined set of therapeutic agents other than those that exhibit(s) inhibitory activity toward a fibroblast growthAttorney Docket No. GNCN-026 / 01WO 320289-2157 factor receptor (FGFR) generally and / or a fibroblast growth factor receptor-3 (FGFR3), specifically such as one or more therapeutic agents or modalities known in the art and / or as described herein. In some cases, the sample is a bodily fluid sample. In some cases, the bodily fluid sample is urine.

[0180] In some cases, the determining step further comprises comparing the nucleic acid expression levels of the plurality of classifiers from Table 6 to the nucleic acid expression levels of the plurality of classifier from Table 6 in at least one sample training set(s), and classifying the tumor sample as having a positive FGFR activation signature (FAS (+)) or negative FGFR activation signature (FAS (-)) based on the results of the comparing step. In some cases, the at least one sample training set is from a reference FGFR3 mutation-containing BLCA sample and / or is from a reference FGFR3 mutation-free BLCA sample.

[0181] In some cases, the comparing step comprises applying a statistical algorithm which comprises determining a correlation between the expression data obtained from the sample obtained from the subject and the expression data from the at least one training set(s); and classifying the sample obtained from the subject as being FAS (+) or FAS (-) based on the results of the statistical algorithm. The statistical algorithm can be any statistical algorithm known in the art and / or provided herein for performing the necessary correlation. In some cases, the statistical algorithm is performed by a computer system. In some cases, the at least one training set is from a reference FGFR3 mutation-containing cancer sample and the sample is classified as possessing the positive FGFR3 activation signature if the nucleic acid expression levels of the plurality of classifiers of Table 6 correlate with the nucleic acid expression levels of the plurality of classifiers of Table 6 from the reference FGFR3 mutation-containing cancer sample. In some cases, the at least one training set is from a reference FGFR3 mutation-containing cancer sample and from a reference FGFR3 mutation-free cancer sample and the sample is classified as possessing the positive FGFR3 activation signature if the expression levels of the plurality of classifiers of Table 6 correlate with the expression levels of the plurality of classifiers of Table 6 from the reference FGFR3 mutation-containing cancer sample.

[0182] In some cases, the plurality of classifiers selected from Table 6 comprises at least 2 classifiers, at least 4 classifiers, at least 6 classifiers, at least 8 classifiers, at least 10 classifier biomarkers, at least 12 classifiers, at least 14 classifiers, at least 16 classifiers, at least 18 classifiers, at least 20 classifiers, at least 22 classifiers, at least 24 classifiers, at least 28 classifiers, at least 30Attorney Docket No. GNCN-026 / 01WO 320289-2157 classifiers, at least 32 classifiers, at least 34 classifiers, at least 36 classifiers, at least 38 classifiers, at least 40 classifiers, at least 42 classifiers, at least 44 classifiers, at least 46 classifiers, at least 48 classifiers or at least 50 classifiers from Table 6.

[0183] In some cases, the plurality of classifiers selected from Table 6 comprises at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% of the classifiers from Table 6.

[0184] In some cases, the plurality of classifiers selected from Table 6 consists of all the classifiers of Table 6.

[0185] The methods of present invention are also useful for evaluating clinical response to therapy, as well as for endpoints in clinical trials for efficacy of new therapies. The extent to which sequential diagnostic expression profiles move towards normal can be used as one measure of the efficacy of the candidate therapy.

[0186] In some cases, the FGFR inhibitor exhibits inhibitory activity toward a fibroblast growth factor receptor (FGFR) generally and / or a fibroblast growth factor receptor-3 (FGFR3), specifically. In some cases, the FGFR inhibitor shows inhibitory activity toward fibroblast growth factor receptor-3 (FGFR3). In some cases, the FGFR inhibitor is a tyrosine kinase inhibitor. In some cases, the FGFR inhibitor is a selective tyrosine kinase inhibitor. In some cases, the FGFR inhibitor is a non-selective tyrosine kinase inhibitor. In some cases, the FGFR inhibitor is selected from the group consisting of erdafitinib (JNJ 42756493), infigratinib (BGJ1398), Rogaritinib (BAY 1163877), AZD4547, Pemigatinib (INCB54828), TAS-120, LY2874455, DEBIO 1347, PD173074, BLU9931, pazopanib, brivanib, ponatinib (AP24534), regorafenib (BAY 73-4506), lenvatinib (E7080), dovitinib (TKI258), lucitanib (E3810), nintedanib (BIBF 1120), Foretinib, and any combination thereof. In some cases, the FGFR inhibitor is nintedanib (BIBF 1120). In some cases, the FGFR inhibitor is an antibody or antibody-conjugate. In some cases, the FGFR inhibitor is B-701 or MFGR1877S. In some cases, the FGFR inhibitor is LY3076226.

[0187] In one embodiment, an agent for use in any of the diagnostic and / or therapeutic methods provided herein is an agent that shows or exhibits inhibitory activity towards a fibroblast growth factor receptor (FGFR). In one embodiment, the detection of a positive FAS in a sample obtained from a patient using any of the FGFR activation signatures provided herein (e.g., Table 6) indicates that the patient is a responder to an agent that shows or exhibits inhibitory activityAttorney Docket No. GNCN-026 / 01WO 320289-2157 towards an FGFR. The agent that shows or exhibits inhibitory activity towards an FGFR can be administered to a responder (patient with a positive FAS) alone or in combination with an additional therapy or therapies. The additional therapy or therapies can be selected from the group consisting of a chemotherapeutic agent, an angiogenesis inhibitor, immunotherapy, radiotherapy, surgical intervention and any combination thereof.

[0188] In one embodiment, the detection of a negative FAS in a sample obtained from a patient using any of the FGFR activation signatures provided herein (e.g., Table 6) indicates that the patient is a non-responder to an agent that shows or exhibits inhibitory activity towards an FGFR. The agent that shows or exhibits inhibitory activity towards an FGFR can thusly, not be administered to a non-responder (patient with a negative FAS). Instead, a patient determined to be a non-responder using any of the diagnostic or detection methods provided herein (e.g., through the use of one or more FGFR3 activation signatures provided herein, i.e., Table 6) is administered a non-FGFR inhibitor therapy or therapies. The additional therapy or therapies can be selected from the group consisting of a chemotherapeutic agent, an angiogenesis inhibitor, immunotherapy, radiotherapy, surgical intervention and any combination thereof.

[0189] The present disclosure provides methods for predicting overall survival rate for a bladder cancer patient. In some embodiments, the prediction of overall survival rate can involve obtaining a bladder tissue sample or bodily fluid (e.g., urine) from a bladder cancer patient. In some embodiments, the bladder cancer patients can have various stages of cancers. In some embodiments, the overall survival rate can be determined by detecting the expression level of at least one subtype classifier of a publicly available bladder cancer database or dataset. In some embodiments, an overall survival rate can be determined by detecting the expression level (e.g., protein and / or nucleic acid) of any subtype classifiers that are relevant to bladder cancer. In one embodiment, the subtype classifiers can be all or a subset of classifiers from Table 6.

[0190] In some embodiments, the present disclosure further provides methods of predicting overall survival in bladder cancer from specific areas of the bladder. In some embodiments, the prediction includes detecting an expression level of at least one gene from a bladder cancer dataset (e.g., Table 6) in a bladder tissue sample or a bodily fluid (e.g., urine) obtained from a bladder cancer patient. In some embodiments, the detection of the expression level of a subtype classifier from a bladder cancer dataset (e.g., Table 6) using the methods providedAttorney Docket No. GNCN-026 / 01WO 320289-2157 herein specifically identifies a FAS (+) or FAS (-). In some embodiments, the identification of the FAS is indicative of the overall survival in the patient. Detection Methods

[0191] Also provided herein are methods for assaying samples for the presence, absence or expression level of a plurality of biomarkers from a set of biomarkers. The set of biomarkers can represent a gene signature and can be ascertained using any of the methods, kits or systems provided herein. The expression level can be a nucleic acid expression level. The samples can be obtained from a subject suffering from a cancer of interest. The cancer of interest is any type of cancer known in the art and / or provided herein. The method for assaying the sample can comprise measuring the expression level of the plurality of biomarkers. The measuring the expression level of the plurality of biomarkers can be using any methods for measuring nucleic acid expression levels known in the art and / or provided herein. The measuring the plurality of biomarkers can be using a sequencing assay. The sequencing assay can be a DNA sequencing assay or RNA sequencing assay.

[0192] In one embodiment, provided herein is a method of assaying a sample obtained from a subject suffering from COAD, the method comprising measuring the expression level of a plurality of biomarkers selected from Table 1 or Table 2 using the sequencing assay.

[0193] In some cases, the plurality of biomarkers selected from Table 1 comprises at least 50 biomarkers, at least 100 biomarkers, at least 150 biomarkers, at least 200 biomarkers, at least 250 classifier biomarkers, at least 300 biomarkers, at least 350 biomarkers, at least 400 biomarkers, at least 450 biomarkers, at least 500 biomarkers, at least 550 biomarkers, at least 600 biomarkers, at least 650 biomarkers, at least 700 biomarkers, at least 750 biomarkers, at least 800 biomarkers, at least 850 biomarkers, at least 900 biomarkers or at least 950 biomarkers from Table 1.

[0194] In some cases, the plurality of classifiers selected from Table 1 comprises at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% of the classifiers from Table 1.

[0195] In some cases, the plurality of classifiers selected from Table 1 consists of all the classifiers of Table 1.

[0196] In some cases, the plurality of classifiers selected from Table 2 comprises at least 2 classifiers, at least 4 classifiers, at least 6 classifiers, at least 8 classifiers, at least 10 classifierAttorney Docket No. GNCN-026 / 01WO 320289-2157 biomarkers, at least 12 classifiers, at least 14 classifiers, at least 16 classifiers, at least 18 classifiers, at least 20 classifiers, at least 22 classifiers, at least 24 classifiers, at least 28 classifiers, at least 30 classifiers, at least 32 classifiers, at least 34 classifiers, at least 36 classifiers, at least 38 classifiers, at least 40 classifiers, at least 42 classifiers, at least 44 classifiers, at least 46 classifiers, at least 48 classifiers, at least 50 classifiers, at least 52 classifiers, at least 54 classifiers, at least 56 classifiers, at least 58 classifiers, at least 60 classifiers, at least 62 classifiers, at least 64 classifiers, at least 66 classifiers or at least 68 classifiers from Table 2.

[0197] In some cases, the plurality of classifiers selected from Table 2 comprises at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% of the classifiers from Table 2.

[0198] In some cases, the plurality of classifiers selected from Table 2 consists of all the classifiers of Table 2.

[0199] In one embodiment, provided herein is a method of assaying a sample obtained from a subject suffering from PAAD, the method comprising measuring the expression level of a plurality of biomarkers selected from Table 3 or Table 4 using the sequencing assay.

[0200] In some cases, the plurality of biomarkers selected from Table 3 comprises at least 50 biomarkers, at least 100 biomarkers, at least 150 biomarkers, at least 200 biomarkers, at least 250 classifier biomarkers, at least 300 biomarkers, at least 350 biomarkers, at least 400 biomarkers, at least 450 biomarkers, at least 500 biomarkers, at least 550 biomarkers, at least 600 biomarkers, at least 650 biomarkers, at least 700 biomarkers, at least 750 biomarkers, at least 800 biomarkers, at least 850 biomarkers, at least 900 biomarkers or at least 950 biomarkers from Table 3.

[0201] In some cases, the plurality of classifiers selected from Table 3 comprises at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% of the classifiers from Table 3.

[0202] In some cases, the plurality of classifiers selected from Table 3 consists of all the classifiers of Table 3.

[0203] In some cases, the plurality of classifiers selected from Table 4 comprises at least 2 classifiers, at least 4 classifiers, at least 6 classifiers, at least 8 classifiers, at least 10 classifier biomarkers, at least 12 classifiers, at least 14 classifiers, at least 16 classifiers, at least 18 classifiers,Attorney Docket No. GNCN-026 / 01WO 320289-2157 at least 20 classifiers, at least 22 classifiers, at least 24 classifiers, at least 28 classifiers, at least 30 classifiers, at least 32 classifiers, at least 34 classifiers, at least 36 classifiers, at least 38 classifiers, at least 40 classifiers, at least 42 classifiers, at least 44 classifiers, at least 46 classifiers, at least 48 classifiers, at least 50 classifiers, at least 52 classifiers, at least 54 classifiers, at least 56 classifiers, at least 58 classifiers, at least 60 classifiers, at least 62 classifiers, at least 64 classifiers, at least 66 classifiers, at least 68 classifiers, at least 70 classifiers, at least 72 classifiers, at least 74 classifiers, at least 76 classifiers, at least 78 classifiers or at least 80 classifiers from Table 4.

[0204] In some cases, the plurality of classifiers selected from Table 4 comprises at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% of the classifiers from Table 4.

[0205] In some cases, the plurality of classifiers selected from Table 4 consists of all the classifiers of Table 4.

[0206] In one embodiment, provided herein is a method of assaying a sample obtained from a subject suffering from BLCA, the method comprising measuring the expression level of a plurality of biomarkers selected from Table 5 or Table 6 using the sequencing assay.

[0207] In some cases, the plurality of biomarkers selected from Table 5 comprises at least 50 biomarkers, at least 100 biomarkers, at least 150 biomarkers, at least 200 biomarkers, at least 250 classifier biomarkers, at least 300 biomarkers, at least 350 biomarkers, at least 400 biomarkers, at least 450 biomarkers, at least 500 biomarkers, at least 550 biomarkers, at least 600 biomarkers, at least 650 biomarkers, at least 700 biomarkers, at least 750 biomarkers, at least 800 biomarkers, at least 850 biomarkers, at least 900 biomarkers or at least 950 biomarkers from Table 5.

[0208] In some cases, the plurality of classifiers selected from Table 5 comprises at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% of the classifiers from Table 5.

[0209] In some cases, the plurality of classifiers selected from Table 5 consists of all the classifiers of Table 5.

[0210] In some cases, the plurality of classifiers selected from Table 6 comprises at least 2 classifiers, at least 4 classifiers, at least 6 classifiers, at least 8 classifiers, at least 10 classifier biomarkers, at least 12 classifiers, at least 14 classifiers, at least 16 classifiers, at least 18 classifiers,Attorney Docket No. GNCN-026 / 01WO 320289-2157 at least 20 classifiers, at least 22 classifiers, at least 24 classifiers, at least 28 classifiers, at least 30 classifiers, at least 32 classifiers, at least 34 classifiers, at least 36 classifiers, at least 38 classifiers, at least 40 classifiers, at least 42 classifiers, at least 44 classifiers, at least 46 classifiers, at least 48 classifiers or at least 50 classifiers from Table 6.

[0211] In some cases, the plurality of classifiers selected from Table 6 comprises at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% of the classifiers from Table 6.

[0212] In some cases, the plurality of classifiers selected from Table 6 consists of all the classifiers of Table 6. Selection of regulatory regions to target

[0213] Chromatin accessibility sequencing (CA-seq) methods such as those selected from the group consisting of micrococcal nuclease digestion with deep sequencing (MNase-seq), DNase I hypersensitive sites sequencing (DNase-seq), Formaldehyde-Assisted Isolation of Regulatory Elements sequencing (FAIRE-seq) and Assay for Transposase-Accessible Chromatin using sequencing (ATAC-seq) can yield information about chromatin accessibility. Gene expression can be measured directly by transcriptomics using RNA-seq or indirectly inferred using a CA-seq method known in the art (e.g., ATAC-seq). However, correlation between the levels of gene expression measured by a CA-seq method (e.g., ATAC-seq) and RNA-seq can vary from no correlation to very highly correlated. Provided herein are methods or systems for identifying features (i.e., genes or regulatory regions thereof) with high correlation between CA-seq (e.g., ATAC-seq) data and RNA-seq data. Once identified, said features can then be used to generate panels for use in a NGS panel assay or in a computer implemented method provided herein for extracting features in cfDNA datasets that translate to or reflect tissue based gene expression.

[0214] In one embodiment, provided herein is a computer-implemented method for identifying regulatory regions for one or more genes that correlate with gene expression of the one or more genes. The computer-implemented method can comprise: (a) inputting into a computer system, a first dataset comprising RNA sequencing data and a second dataset comprising chromatin accessibility data; (b) mapping by the computer system, sequence reads obtained from the first dataset to a reference genome sequence; (c) mapping by the computer system, sequence reads obtained from the second dataset to the reference genome sequence; (d) annotating by theAttorney Docket No. GNCN-026 / 01WO 320289-2157 computer system, peaks of the sequence reads from the second dataset over random or low-level noise in the chromatin accessibility data; wherein each peak represents a putative chromatin accessible region; (e) mapping by the computer system, sequence reads from the second dataset to a nearest gene found from the mapped sequence reads from the first dataset using machine learning, thereby generating a plurality of genes and chromatin accessibility regions associated with each gene from the plurality of genes; (f) measuring by the computer system, a correlation between a level of expression of each gene in the plurality of genes to each chromatin accessibility region associated with each gene in the plurality of genes, wherein the level of expression of each gene in the plurality of genes is evidenced by the number of sequence reads from the first dataset that mapped to each gene in the plurality of genes; (g) ranking by the computer system, the correlations obtained in (f); (h) selecting by the computer system, chromatin accessibility regions whose correlations coefficients were greater than a desired threshold, thereby generating a list of chromatin accessibility regions that are highly correlated with RNA expression levels; and (i) selecting by the computer system, regulatory regions from the chromatin accessibility regions selected in (h), thereby identifying regulatory regions for one or more genes that highly correlate with RNA expression of the one or more genes. In one embodiment, the method further comprises (j) designing by the computer, a set of probes, wherein each probe in the set comprises sequence complementary to one of the regulatory regions selected in (i). In one embodiment, the method further comprises (j) designing by the computer, a set of primer pairs, wherein each primer pair in the set comprises sequence complementary to one of the regulatory regions selected in (i). The correlation can be a Pearson correlation and / or a Spearman correlation. In one embodiment, the correlation is a Pearson and a Spearman correlation. In one embodiment, the correlation is a Spearman correlation. The desired threshold can be a correlation coefficient greater than 0.6, 0.7, 0.8 or 0.9. In one embodiment, the desired threshold is a correlation coefficient greater than 0.8. The chromatin accessibility data from the second dataset can be obtained from a CA-seq assay. The CA-seq assay can be selected from the group consisting of micrococcal nuclease digestion with deep sequencing (MNase-seq), DNase I hypersensitive sites sequencing (DNase-seq), Formaldehyde-Assisted Isolation of Regulatory Elements sequencing (FAIRE-seq) and Assay for Transposase-Accessible Chromatin using sequencing (ATAC-seq). In one embodiment, the CA- seq assay is ATAC-seq and the chromatin accessibility regions are ATAC-seq regions. The machine learning can employ algorithms, executed by computer, that automate analytical modelAttorney Docket No. GNCN-026 / 01WO 320289-2157 building, e.g., for clustering, classification or pattern recognition. The machine learning algorithms can be supervised or unsupervised. The machine learning algorithms can include, for example, artificial neural networks (e.g., back propagation networks), discriminant analyses (e.g., Bayesian classifier or Fischer analysis), support vector machines, decision trees (e.g., recursive partitioning processes such as CART - classification and regression trees, or random forests), linear classifiers (e.g., multiple linear regression (MLR), partial least squares (PLS) regression, and principal components regression), hierarchical clustering, and cluster analysis. In some cases, the machine learning is the “findClosestGene” function from the ACME package. In some cases, the “findClosestGene” function from the ACME package can be used to map chromatin accessibility (e.g., ATAC) regions to the nearest gene using refseq transcripts listed at genome.ucsc.edu. For the chromatin accessibility regions that may be found to have multiple equal matches to multiple genes (e.g., refseq transcripts), which can be due to alternate transcripts, the first listed gene (e.g., refseq gene) can be selected. The regulatory regions can also be referred to as regulatory elements (REs) or cis-regulatory elements (CREs). The CREs associated with a gene can be selected from the group consisting of the 5’UTR, the 3’ UTR, a promoter, a proximal enhancer, a distal enhancer and any combination thereof. In some cases, the first dataset and the second dataset can each be obtained from tissue samples or tissue biopsies. The tissue samples can be tissue matched samples between the first and second datasets. In some cases, the first dataset and the second dataset can be obtained from a population of subjects that each have matched chromatin accessibility sequencing (e.g., ATAC-seq) data and RNA expression sequencing (e.g., RNA-seq) data. The reference genome can be haploid. The reference genome can represent a mosaic of the genomes of several individuals of species. The reference genome can be publicly available or a privately available reference genome. The reference genome can be a reference human genome. The reference human genome can be any human genome assembly known in the art, such as, for example, hg19, hg38 or CHM13 assemblies. In some cases, the reference human genome can be any version of the GRCh38 (hg38) assembly. The reference genome can be accessed through any web browser known in the art. The random or low-level noise in the chromatin accessibility data can be random or low-level noise in the sequencing data (i.e., not enriched for reads) obtained from the CA-seq assay. This background can be associated with non-accessible regions (e.g., DNA wrapped around nucleosomes).Attorney Docket No. GNCN-026 / 01WO 320289-2157

[0215] In one embodiment, the first and second datasets used in a method for identifying regulatory regions for one or more genes that correlate with gene expression of the one or more genes as provided herein can be obtained from one or more subjects that suffer from or are suspected of suffering from a desired feature of a cancer. The cancer can be any cancer known in the art and / or provided herein. Further to this embodiment, the regulatory regions thereby identified by said method can represent regulatory regions expressed when a subject suffers from the desired feature of the cancer. The desired feature can be any desired feature of cancer provided herein. In some cases, the desired feature can be a type of cancer, a subtype of cancer or the presence or absence of microsatellite instability. A panel of probes or set of primers can be designed to target the regulatory regions identified to be representative of the desired feature of cancer.

[0216] In one embodiment, prior to step (a) in the aforementioned computer-implemented method from identifying regulatory regions for one or more genes that correlate with gene expression of the one or more genes, the second dataset comprising chromatin accessibility data can be generated by inputting into the computer system, chromatin accessibility sequencing (e.g., ATAC-seq) data for a first plurality of subjects suffering from one type of cancer or subtype thereof and chromatin accessibility sequencing (e.g., ATAC-seq) data for a second plurality of subjects suffering from a second type of cancer or subtype thereof; and filtering out by the computer system, chromatin accessibility regions present in the chromatin accessibility sequencing data for the first and second plurality of subjects, thereby generating a second dataset comprising chromatin accessibility data specific to the first and second type of cancer or subtype thereof. The cancer can be any cancer known in the art and / or provided herein.

[0217] In one embodiment, provided herein is a method for generating a panel of regulatory regions for one or more genes differential for one or more desired features of cancer. In one embodiment, the panel is generated by a method that comprises (a) inputting into a computer system, a first dataset comprising RNA sequencing data and a second dataset comprising chromatin accessibility data, wherein sequencing data in the first dataset and the second dataset are obtained from tissue samples from a plurality of subjects that suffer from a cancer; (b) mapping by the computer system, sequence reads obtained from the first dataset to a reference genome sequence; (c) mapping by the computer system, sequence reads from the second dataset to a reference genome and regions with higher than background read coverage, or “peaks” areAttorney Docket No. GNCN-026 / 01WO 320289-2157 annotated to the nearest gene found from the mapped sequence reads to the first dataset using machine learning, thereby generating a plurality of genes and chromatin accessibility regions associated with each gene from the plurality of genes; (d) measuring by the computer system, a correlation between a level of expression of each gene in the plurality of genes to each chromatin accessibility region associated with each gene in the plurality of genes, wherein the level of expression of each gene in the plurality of genes is evidenced by the number of sequence reads from the first dataset that mapped to each gene in the plurality of genes; (e) ranking by the computer system, the correlations obtained in (d); (f) selecting by the computer system, chromatin accessibility regions whose correlation coefficients were greater than a desired threshold, thereby generating a list of chromatin accessibility regions that are highly correlated with RNA expression levels; (g) selecting by the computer system, regulatory regions from the chromatin accessibility regions selected in (f), thereby identifying regulatory regions for one or more genes that highly correlate with RNA expression of the one or more genes; (h) inputting into the computer system, gene expression based signatures for a desired feature of the cancer; and (i) selecting by the computer system, regulatory regions for one or more genes from (g) that correspond to one or more genes from the gene expression based signature inputted in (h). In one embodiment, the method further comprises (j) designing by the computer system, a set of probes, wherein each probe in the set comprises sequence complementary to one of the regulatory regions selected in (i). In one embodiment, the method further comprises (j) designing by the computer system, a set of primer pairs, wherein each primer pair in the set comprises sequence complementary to one of the regulatory regions selected in (i). The correlation can be a Pearson correlation and / or a Spearman correlation. In one embodiment, the correlation is a Pearson and a Spearman correlation. In one embodiment, the correlation is a Spearman correlation. The desired threshold can be a correlation coefficient greater than 0.6, 0.7, 0.8 or 0.9. In one embodiment, the desired threshold is a correlation coefficient greater than 0.8. The chromatin accessibility data from the second dataset can be obtained from a CA-seq assay. The CA-seq assay can be selected from the group consisting of micrococcal nuclease digestion with deep sequencing (MNase-seq), DNase I hypersensitive sites sequencing (DNase-seq), Formaldehyde-Assisted Isolation of Regulatory Elements sequencing (FAIRE-seq) and Assay for Transposase-Accessible Chromatin using sequencing (ATAC-seq). In one embodiment, the CA-seq assay is ATAC-seq and the chromatin accessibility regions are ATAC-seq regions. The machine learning can employ algorithms, executed byAttorney Docket No. GNCN-026 / 01WO 320289-2157 computer, that automate analytical model building, e.g., for clustering, classification or pattern recognition. The machine learning algorithms can be supervised or unsupervised. The machine learning algorithms can include, for example, artificial neural networks (e.g., back propagation networks), discriminant analyses (e.g., Bayesian classifier or Fischer analysis), support vector machines, decision trees (e.g., recursive partitioning processes such as CART - classification and regression trees, or random forests), linear classifiers (e.g., multiple linear regression (MLR), partial least squares (PLS) regression, and principal components regression), hierarchical clustering, and cluster analysis. In some cases, the machine learning is the “findClosestGene” function from the ACME package. In some cases, the “findClosestGene” function from the ACME package can be used to map chromatin accessibility (e.g., ATAC) regions to the nearest gene using refseq transcripts listed at genome.ucsc.edu. For the chromatin accessibility regions that may be found to have multiple equal matches to multiple genes (e.g., refseq transcripts), which can be due to alternate transcripts, the first listed gene (e.g., refseq gene) can be selected. The regulatory regions can also be referred to as regulatory elements (REs) or cis-regulatory elements (CREs). The CREs associated with a gene can be selected from the group consisting of the 5’UTR, the 3’ UTR, a promoter, a proximal enhancer, a distal enhancer and any combination thereof. In some cases, the sequencing data in the first dataset and the second dataset are obtained from matched tissue samples from a plurality of subjects that suffer from a cancer or a desired feature of a cancer. In some cases, the first dataset and the second dataset can be obtained from a population of subjects that each have matched chromatin accessibility sequencing (e.g., ATAC-seq) data and RNA expression sequencing (e.g., RNA-seq) data. In some cases, the desired feature of cancer is a subtype of the cancer. The desired feature can be any desired feature provided herein. In some cases, the desired feature is a microsatellite stability of the cancer. The cancer can be any cancer known in the art and / or provided herein. The reference genome can be haploid. The reference genome can represent a mosaic of the genomes of several individuals of species. The reference genome can be publicly available or a privately available reference genome. The reference genome can be a reference human genome. The reference human genome can be any human genome assembly known in the art, such as, for example, hg19, hg38 or CHM13 assemblies. In some cases, the reference human genome can be any version of the GRCh38 (hg38) assembly. The reference genome can be accessed through any web browser known in the art. In some cases, each gene that has a probe directed against (i.e., sequence complementary to) a regulatory region of said gene inAttorney Docket No. GNCN-026 / 01WO 320289-2157 a panel of probes generated using a method provided herein comprises a probe directed against one or a plurality of regulatory regions of said gene. In some cases, the plurality of regulatory regions of said gene against which a probe is directed in a panel can be at least 2, 3, 4, 5, 6, 7, 8, 9 or 10 regulatory regions of said gene.

[0218] In one embodiment, provided herein is a method for sequencing target regulatory regions in cfDNA. In some cases, the method for sequencing target regulatory regions in cfDNA can comprise (a) obtaining a cfDNA comprising sample (e.g., a liquid biopsy) from a subject (e.g., a s subject suffering from or suspected of suffering from cancer); and (b) performing a target enrichment sequencing assay on the cfDNA comprising sample using a set of primer pairs, wherein each primer pair in the set comprises sequence complementary to one of the regulatory regions generated using a method provided herein, thereby generating a gene activity matrix for the cfDNA for the subject. The target enrichment sequencing assay can comprise a hybrid-capture based enrichment step followed by a sequencing assay (e.g., next-generation sequencing (NGS) assay). In one embodiment, the method for sequencing target regulatory regions in cfDNA can comprise (a) obtaining a cfDNA comprising sample (e.g., a liquid biopsy) from a subject (e.g., a subject suffering from or suspected of suffering from cancer); and (b) performing a target enrichment sequencing assay on the cfDNA comprising sample using a panel comprising a plurality of probes that comprises sequence complementary to the regulatory regions generated using a method provided herein, thereby generating a gene activity matrix for the cfDNA for the subject. The target enrichment sequencing assay can be any hybrid-capture based enrichment sequencing assay (e.g., next-generation sequencing (NGS) assay) known in the art. In one embodiment, the hybrid-capture based enrichment sequencing assay of (b) can entail (i) isolating target cfDNA using the panel comprising the plurality of probes that comprises sequence complementary to the regulatory regions generated using a method provided herein, thereby generating a pool of target cfDNA pulled down via hybridization with probes from the panel; and (ii) performing a sequencing (e.g., NGS) assay on the pool of target cfDNA. The sequencing assay can be any sequencing assay known in the art and / or provided herein. The cancer can be any cancer known in the art and / or provided herein. The cfDNA sample can be a bodily fluid sample or liquid biopsy sample obtained from the subject. The bodily fluid sample can be any type of bodily fluid known in the art and / or provided herein. The regulatory regions can also be referred to as regulatory elements (REs) or cis-regulatory elements (CREs). The CREs associated with a gene can be selected from the groupAttorney Docket No. GNCN-026 / 01WO 320289-2157 consisting of the 5’UTR, the 3’ UTR, a promoter, a proximal enhancer, a distal enhancer and any combination thereof. In some cases, the gene activity matrix generated by the target enrichment sequencing assay (e.g., hybrid-capture enrichment)) comprises fragmentation patterns of the cfDNA sequenced from the cfDNA comprising sample obtained from the subject.

[0219] In one embodiment, provided herein is a method for detecting a desired feature of a cancer in a subject using fragmentation patterns of cfDNA. The fragmentation patterns of cfDNA can be generated using any target enrichment sequencing assay provided herein. The cancer can be any type of cancer known in the art and / or provided herein. In one embodiment, the method comprises: (a) determining fragmentation patterns of sequencing data for a plurality of regulatory regions associated with a plurality of target genes, wherein the sequencing data is obtained from a sample comprising cell-free deoxyribonucleic acid (cfDNA) obtained from the subject; and (b) classifying the fragmentation patterns to identify the subject as being negative or positive for desired feature of the subject. In some cases, the sequencing data for the plurality of regulatory regions associated with one or more genes is obtained by performing a target enrichment sequencing assay (e.g., hybrid-capture enrichment) using a panel of probes, wherein each probe in the panel comprises sequence complementary to a regulatory region associated with a target gene from the plurality of target genes. The regulatory regions for inclusion in the panel can be identified using any method provided herein for identifying regulatory regions that correlate with target gene expression of one or target genes to which the regulatory regions are associated. In some cases, each target gene from the plurality of target genes comprises probes comprising complementary sequence to at least 2, 3, 4, 5, 6, 7, 8, 9 or 10 regulatory regions associated with the target gene. In some cases, each target gene from the plurality of target genes comprises probes comprising sequence complementary to 2 or 3 regulatory regions associated with the target gene. The regulatory regions can also be referred to as regulatory elements (REs) or cis-regulatory elements (CREs). The CREs associated with a gene can be selected from the group consisting of the 5’UTR, the 3’ UTR, a promoter, a proximal enhancer, a distal enhancer and any combination thereof. In some cases, the determining the fragmentation patterns comprises determining a fragment size distribution of the plurality of regulatory regions associated with a plurality of target genes.

[0220] In some cases, the determining the fragmentation patterns comprises quantitating each fragment size distribution. Any of a number of distribution quantitations can be used. These include but are not limited to quantitation of entropy, sum, minimum, maximum, interquartileAttorney Docket No. GNCN-026 / 01WO 320289-2157 range, mean, median, mode, variance, standard deviation, kurtosis, diversity, depth of sequencing, bins, and / or Kolmogorov-Smirnov statistic. In some cases, the quantitating comprises quantitating an entropy value for each fragment size distribution. The distribution quantitation used in a method provided herein can be an entropy quantitation as described in Roach TNF. Use and Abuse of Entropy in Biology: A Case for Caliber. Entropy (Basel). 2020 Nov 25;22(12):1335). In one embodiment, the entropy value is a Shannon entropy value. The Shannon entropy value for use in a method provided herein can be a Shannon entropy quantitation as described in Shannon, Claude E. (July 1948). “A Mathematical Theory of Communication”. Bell System Technical Journal.27 (3): 379-423) (Shannon, Claude E. (October 1948). “A Mathematical Theory of Communication”. Bell System Technical Journal. 27 (4): 623-656). Other suitable entropy quantitations include Rényi entropy (Rényi, Alfréd (1961). “On measures of information and entropy” (PDF). Proceedings of the fourth Berkeley Symposium on Mathematics, Statistics and Probability 1960. pp.547-561) and Tsallis entropy (Tsallis, C. (1988). “Possible generalization of Boltzmann-Gibbs statistics”. Journal of Statistical Physics.52 (1-2): 47-487), among others.

[0221] In some cases, the classifying the fragmentation patterns to identify the subject as being negative or positive for desired feature of the subject comprises inputting the fragmentation patterns from (a) into classifier trained for detecting the desired feature of the cancer. The cancer can be any cancer known in the art and / or provided herein. The desired feature can be selected from the group consisting of a type of cancer, a subtype of the cancer of interest, microsatellite instability, mismatch repair proficiency, minimal residual disease, prognostic risk, response to a therapeutic agent, development of resistance to a therapeutic agent cancer. The classifier can identify or diagnose the subject as being negative or positive for the desired feature of a cancer by determining a probability or numerical score from the fragmentation patterns and classifying the subject based on certain thresholds thereof. Classifying the fragmentation patterns can be performed with any classifier known in the art and / or provided herein. In one embodiment, the classifier is an algorithm or machine learning model that was generated by receiving, as input, test data and producing, as output, a classification of the input data as belonging to one or another class (e.g., being positive or negative for the desired feature of a cancer). Further to this embodiment, the classifier can be trained for the purposes herein (e.g., identifying or diagnosing a subject as possessing a desired feature of a cancer) by determining the fragmentation patterns of the regulatory regions sequenced using a target enrichment sequencing assay (e.g., hybrid-captureAttorney Docket No. GNCN-026 / 01WO 320289-2157 enrichment) as described herein from samples comprising cfDNA obtained from test subjects having a desired feature of a cancer (e.g., cancer or particular subtypes of cancer) and control subjects not having the desired feature, or model samples representative of the same. Machine learning can then be used to distinguish the fragmentation patterns of regulatory regions sequenced from the cfDNA comprising samples obtained from the subjects having the desired feature of interest from cfDNA from subjects not having the desired feature of interest. The machine learning algorithms may be supervised or unsupervised. The machine learning algorithms can include, for example, artificial neural networks (e.g., back propagation networks), discriminant analyses (e.g., Bayesian classifier or Fischer analysis), support vector machines, decision trees (e.g., recursive partitioning processes such as CART - classification and regression trees, or random forests), linear classifiers (e.g., multiple linear regression (MLR), partial least squares (PLS) regression, and principal components regression), hierarchical clustering, and cluster analysis.

[0222] In one embodiment, the identification of the desired feature in a subject using the aforementioned method can then allow for determining a prognosis of the subject and / or selecting a therapeutic intervention for the subject. In some cases, the subject is selected for a therapeutic intervention, which can be any therapeutic intervention known in the art and / or provided herein for treating a subject with the desired feature of cancer.

[0223] In one embodiment, provided herein is a computer-implemented method comprising: (a) inputting into a computer system, paired-end DNA sequencing data obtained from a target enrichment sequencing assay performed on a sample comprising cfDNA obtained from a subject, wherein the target enrichment sequencing assay is performed using a panel comprising a plurality of probes, wherein each probe in the plurality of probes comprises sequence directed to or complementary to a regulatory region associated with a target gene from a plurality of target genes as identified using any method provided herein sand (b) calculating by the computer system one or more sets of fragment characteristics for sequences from the pair-end DNA sequencing data, thereby generating a fragment characteristic matrix for each regulatory region obtained from the target enrichment sequencing assay for the sample comprising cfDNA obtained from the subject. The target enrichment sequencing assay can be a hybrid capture enrichment sequencing assay, wherein each probe in the panel binds to and pulls down a regulatory region associated with a target gene from the plurality of target genes prior to being subjected to a sequencing assay. In some cases, the subject suffers from a desired feature of cancer. In some cases, the fragmentAttorney Docket No. GNCN-026 / 01WO 320289-2157 characteristic matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for the subject represents a sample positive for the desired feature of cancer. In some cases, steps (a)-(e) are repeated on samples comprising cfDNA obtained from a plurality of subjects that each suffer from the desired feature of cancer. In some cases, the fragment characteristic matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for each of the plurality of subjects that suffer from the desired feature of cancer are combined by computer system. In some cases, the subject does not suffer from the desired feature of cancer. In some cases, the fragment characteristic matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for the subject represents a sample negative for the desired feature of cancer. In some cases, steps (a)-(e) are repeated on samples comprising cfDNA obtained from a plurality of subjects that each do not suffer from the desired feature of cancer. In some cases, the fragment characteristic matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for each of the plurality of subjects that do not suffer from the desired feature of cancer are combined by the computer system. In some cases, the desired feature of the cancer is selected from the group consisting of a type of cancer, a subtype of the cancer of interest, microsatellite instability, mismatch repair proficiency, minimal residual disease, prognostic risk, response to a therapeutic agent, development of resistance to a therapeutic agent cancer. In some cases, the fragment characteristics are selected from the group consisting of fragment sequence length, end motif frequency, end motif sequence, jagged end length, fragment diversity, FFT amplitude magnitude and any combination thereof. In some cases, the panel comprising a plurality of probes further comprise one or more probes that each comprise sequence complementary to at least one exon and / or intron of a target gene. The target gene can be any gene known in the art to be present on a NGS panel for assessing gene alteration.

[0224] In another embodiment, provided herein is a computer-implemented method for generating a classifier for determining a desired feature of a cancer, the method comprising: (a) receiving in a computer system, a first set of fragment characteristic indices obtained from cfDNA comprising samples obtained from a first population of subjects suffering from the desired feature of the cancer and a second set of fragment characteristic indices obtained from cfDNA comprising samples obtained from a second population of subjects that do not suffer from the desired feature of the cancer; (b) selecting by the computer system, regulatory regions for model training andAttorney Docket No. GNCN-026 / 01WO 320289-2157 validation using the first set and second set of fragment characteristic matrices; and (c) training and validating by the computer system a machine learning algorithm for determining the presence or absence of the desired feature of the cancer in a cfDNA comprising sample, thereby generating a classifier for determining a desired feature of a cancer. In some cases, the method further comprises inputting target enrichment sequencing data obtained for cfDNA comprising sample obtained from a test subject suspected of suffering from the desired feature of the cancer into the classifier generated by (a)-(c), thereby determining the presence or absence of the desired feature of the cancer in the test subject. In some cases, the cfDNA comprising samples are a bodily fluid samples or liquid biopsy samples. In some cases, the desired feature is selected from the group consisting of a type of cancer, a subtype of the cancer of interest, microsatellite instability, mismatch repair proficiency, minimal residual disease, prognostic risk, response to a therapeutic agent, development of resistance to a therapeutic agent cancer. In some cases, the fragment characteristic matrices for the first and the second populations are generated using the method of any of the methods provided herein.

[0225] In one embodiment, a Shannon Entropy (SE) based fragmentomics method can be applied to a gene activity matrix generated from a target enrichment sequencing assay (e.g., hybrid- capture enrichment) to quantify the variability of cfDNA fragment lengths from the target enrichment sequencing assay (e.g., hybrid-capture enrichment). The target enrichment sequencing assay can be performed using a targeted DNA sequencing panel comprising a plurality of probes, wherein each probe in the plurality of probes comprises sequence directed to or complementary to a regulatory region. The regulatory region can be identified using any method provided herein for identifying regulatory regions for one or more genes that correlate with gene expression of the one or more genes. In some cases, the SE based fragmentomics method can be used to quantify the variability of cfDNA sequence fragment lengths found in the targeted DNA sequencing panel data. Low SE values can be associated with a more repressed nucleosomal DNA state, indicating a more protected and less transcriptionally active region of DNA, while the opposite can be expected for high SE values (i.e., a high SE value can be associated with a more open nucleosomal DNA state, indicating a less protected and more transcriptionally active region of DNA). Low vs high SE at a given region can be relative to the rest of the dataset and the length of the region queried. In general, a region can have low SE if the SE is below the median SE value of all SEs with the same region length. Likewise, a region can have high SE if the SE is above the medianAttorney Docket No. GNCN-026 / 01WO 320289-2157 SE value of all SEs with the same region length. Comparing SEs across regions of different lengths is not recommended. Paired-end DNA sequences are used to calculate fragment sequence lengths that overlap “features”, specific genomic regions of interest. Counts of each unique sequence fragment length in each feature are used to calculate the probabilities of finding each fragment size. Shannon entropy is then calculated using those probabilities: -sum(probabilities * log(probabilities)). A Shannon Entropy matrix can then be compiled for each sample and feature targeted in the panel.

[0226] In some cases, a Shannon Entropy (SE) based fragmentomics method as provided herein comprises: (a) inputting into a computer system, paired-end DNA sequencing data obtained from a target enrichment sequencing assay (e.g., hybrid-capture enrichment) performed on a sample comprising cfDNA obtained from a subject, wherein the target enrichment sequencing assay (e.g., hybrid-capture enrichment) assay is performed using a panel comprising a plurality of probes, wherein each probe in the plurality of probes comprises sequence directed to or complementary to a regulatory region associated with a target gene from a plurality of target genes; (b) calculating by the computer system, fragment sequence lengths for sequences from the pair- end DNA sequencing data that comprise overlapping sequence for a first regulatory region from (a)); (c) calculating a probability of finding each fragment size by the computer system, using counts of each unique fragment sequence fragment length for the first regulatory region; (d) calculating Shannon entropy using the probabilities from (c), thereby generating a Shannon entropy value for the first regulatory region; and (e) repeating (b)-(d) for each additional regulatory region from the target enrichment sequencing assay, thereby generating a Shannon entropy matrix for each regulatory region obtained from the target enrichment sequencing assay for the sample comprising cfDNA obtained from the subject. In some cases, the target enrichment sequencing assay is a hybrid capture enrichment sequencing assay, wherein each probe in the panel binds to and pulls down a regulatory region associated with a target gene from the plurality of target genes prior to being subjected to a sequencing assay. In one embodiment, the subject suffers from a desired feature of cancer. Further to this embodiment, the aforementioned method generates a Shannon entropy matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for a subject positive for the desired feature of a cancer. In some cases, the aforementioned method is repeated on samples comprising cfDNA obtained from a plurality of subjects that each suffer from the desired feature of a cancer. In some cases, the Shannon entropyAttorney Docket No. GNCN-026 / 01WO 320289-2157 matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for each of a plurality of subjects positive for the desired feature of a cancer can be combined by computer system. In one embodiment, the subject does not suffer from the desired feature of cancer. Further to this embodiment, the aforementioned method generates a Shannon entropy matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for a subject negative for the desired feature of a cancer. In some cases, the aforementioned method is repeated on samples comprising cfDNA obtained from a plurality of subjects that each do not suffer from the desired feature of a cancer. In some cases, the Shannon entropy matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for each of a plurality of subjects negative for the desired feature of a cancer can be combined by the computer system. The desired feature can be selected from the group consisting of a type of cancer, a subtype of the cancer of interest, microsatellite instability, mismatch repair proficiency, minimal residual disease, prognostic risk, response to a therapeutic agent, development of resistance to a therapeutic agent cancer. In some cases, the panel comprising a plurality of probes further comprise one or more probes that each comprise sequence complementary to at least one exon and / or intron of a target gene. The target gene can be any gene known in the art to be present on a NGS panel for assessing gene alteration.

[0227] In one embodiment, provided herein is a computer-implemented method for generating a classifier for determining a desired feature of a cancer, the method comprising: (a) receiving in a computer system, a first set of Shannon entropy indices obtained from cfDNA comprising samples (e.g., liquid biopsies) obtained from a first population of subjects suffering from the desired feature of the cancer and a second set of Shannon entropy indices obtained from cfDNA comprising samples (e.g., liquid biopsies) obtained from a second population of subjects that do not suffer from the desired feature of the cancer; (b) selecting by the computer system, regulatory regions for model training and validation using the first set and second set of Shannon entropy matrices; and (c) training and validating by the computer system a machine learning algorithm for determining the presence or absence of the desired feature of the cancer in a cfDNA comprising sample, thereby generating a classifier for determining a desired feature of a cancer. In some cases, the method further comprises inputting target enrichment sequencing data obtained for cfDNA comprising sample obtained from a subject suspected of suffering from the desired feature of the cancer into the classifier generated by (a)-(c), thereby determining the presence orAttorney Docket No. GNCN-026 / 01WO 320289-2157 absence of the desired feature of the cancer in the subject. The desired feature can be selected from the group consisting of a type of cancer, a subtype of the cancer of interest, microsatellite instability, mismatch repair proficiency, minimal residual disease, prognostic risk, response to a therapeutic agent, development of resistance to a therapeutic agent cancer. The Shannon entropy matrices can be determined using any method provided herein. The cancer can be any cancer known in the art and / or provided herein. In some cases, (b) comprises (i) applying by the computer system, labels for each sample in the first set of Shannon entropy indices and the second set of Shannon entropy indices as possessing the desired feature of the cancer or not; and (ii) performing feature selection on the regulatory regions to identify sets of features (e.g., regulatory regions) that differentiate samples within the first and second sets as possessing the desired feature of the cancer or not. The sets of feature identified in (b)(ii) can then be used in (c). The number of features used in (c) can serve to minimizes error rates in cross-validation on the training set.

[0228] An exemplary computer implemented method for generating an algorithmic expression index or classifier for predicting HER2 or ERBB2 status in a subject. The computer implemented method cancomprise: (a) performing sequencing on a liquid panel that comprises all or some combination of the red, orange, and yellow cis-regulatory elements (CREs) around a target gene (e.g., the ERBB2 gene as shown in FIG.25) as well as all possible regions / CREs that were found to be highly correlated with gene expression for genes whose expression is known to be associated with a phenotype of interest (i.e., HER2 / ERBB2 status) on a plurality of samples obtained from subjects suffering from cancer; (b) calculating the Shannon Entropy for each CRE from the liquid panel sequencing data obtained from each sample; (c) applying immunohistochemistry (IHC) labels (high, medium, and low) or other clinical (e.g. responder non- responder) or molecular labels (e.g., PAM50 subtypes) to each sample; (d) performing feature selection on CREs and promoters from the target genes (e.g., ERBB2 and other related genes) using high and low samples; and (e) training a nearest centroid classifier (i.e., an algorithmic expression cluster) to bin samples as high, medium, and low.

[0229] This example is for developing an algorithmic expression cluster corresponding to IHC, clinical status, or other molecular subtype; but, with the use of matched tissue RNA-seq, computer learning models such as pamr, Lasso or Ridge regression models could be built, and samples could be rank ordered from the CRE Shannon Entropy results. To do this, machine learning models such as pamr or glmnet can be used to select cfDNA CRE Shannon EntropyAttorney Docket No. GNCN-026 / 01WO 320289-2157 features that predict the matched tissue RNA-expression of ERBB2. Samples can then be ranked by those with highest CRE signal corresponding to ERBB2 expression.

[0230] Mapping of sequences in any of the methods or system provided herein can be performed by alignment of sequences using an alignment algorithm, for example, Needleman- Wunsch algorithm (see e.g., the EMBOSS Needle aligner available at the URL ebi.ac.uk / Tools / psa / emboss_needle / nucleotide.html, optionally with default settings), the BLAST algorithm (see e.g., the BLAST alignment tool available at the URL blast.ncbi.nlm.nih.gov / Blast.cgi, optionally with default settings), or the Smith-Waterman algorithm (see e.g., the EMBOSS Water aligner available at the URL ebi.ac.uk / Tools / psa / emboss_water / nucleotide.html, optionally with default settings). Optimal alignment may be assessed using any suitable parameters of a chosen algorithm, including default parameters.

[0231] The isolating of cfDNA in any method provided herein can utilize methods of isolating targeted cfDNA that are known in the art. See, e.g., US 2019 / 0287645 A1, US 2022 / 0259647 A1, and US 2022 / 0090207 A1, which are incorporated herein by reference in their entireties.

[0232] The sequencing performed in any method provided herein can be performed using a first-generation sequencing method, such as Maxam-Gilbert or Sanger sequencing, or a high- throughput sequencing (e.g., next-generation sequencing or NGS) method. Sequencing methods may include, but are not limited to: pyrosequencing, sequencing-by-synthesis, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing-by-ligation, sequencing-by-hybridization, Digital Gene Expression (Helicos), massively parallel sequencing, e.g., Helicos, Clonal Single Molecule Array (Solexa / Illumina), sequencing using PacBio, SOLID, Ion Torrent, or Nanopore platforms. Single stream cfDNA alteration and fragmentation analysis Assay (GenomicsNextTMAssay)

[0233] In one embodiment, provided herein are methods, systems or kits for generating DNA alteration data as well as gene activity or expression data from a single sample that comprises cfDNA obtained from a subject suffering from a disease or disorder. This type of method, system or kit can be referred to as a GenomicsNextTMassay. In some cases, the disease or disorder is a cancer or subtype thereof. The cancer or subtype thereof can be any subtype of cancer or subtypeAttorney Docket No. GNCN-026 / 01WO 320289-2157 thereof known in the art and / or provided herein. The single sample that comprises cfDNA can be a bodily fluid sample or liquid biopsy sample. In some cases, the methods, systems or kits for generating DNA alteration data as well as gene activity or expression data from a single sample that comprises cfDNA obtained from a subject suffering from cancer or subtype thereof and a single DNA sequencing run (i.e., GenomicsNextTMassay) comprises gene alteration testing by next generation sequencing (NGS) oncology panels with additional probes to gene regulatory regions to enable a hybrid-capture based enrichment step followed by sequencing (e.g., next- generation sequencing (NGS)) and subsequent gene expression analysis by fragmentomics. The gene regulatory regions can be selected by a method provided herein for identifying regulatory regions for one or more genes that correlate with gene expression of the one or more genes. The gene expression analysis by fragmentomics can be any gene analysis method provided herein. A methods, systems or kits for generating DNA alteration data as well as gene activity or expression data from a single sample that comprises cfDNA obtained from a subject suffering from a disease or disorder provided herein can provide phenotype information that can be applied to obtain a deeper understanding of gene alteration results (FIG. 27) by straightforward means such as calculating gene penetrance based on fragment size distributions or reporting the expression levels of target genes of interest as well as through creating complex algorithmic signatures to identify subtypes or predict outcomes. The algorithmic signature to identify subtypes or predict outcomes can be accomplished by applying a liquid biopsy classifier generated using a method provided herein to the sequencing data obtained from the single sample. In some cases, the GenomicsNextTMassay can generate a gene activity matrix. The gene activity matrix can be used to discover biomarkers orthogonal to gene alterations, increase the interpretability of alterations, and aidclinical decision- In this way, the GenomicsNextTMassay can provide multimodal clinical-genomic data. The multimodal clinical-genomic data can be housed in a database. In some cases, any database generated by the GenomicsNextTMassay can be accessed with tools that enable real-world data access, exploration, and discovery. Accordingly, use of the GenomicsNextTMassay can provide new patient, biomarker and drug development insights.

[0234] Because gene alteration and fragmentomics information can be complementary, the second dimension added by DNA fragment profiling can enable DNA sequencing to extend across the full range of liquid biopsy applications being pursued today (FIG.28). On one hand, the secondAttorney Docket No. GNCN-026 / 01WO 320289-2157 dimension of information can improve quality, accuracy, and interpretation of existing uses of genomic testing such as screening, identifying actionable alterations, monitoring MRD and therapy response, and detecting resistance. On the other hand, fragment profiling can open new applications for DNA sequencing such as, for example, determining tissue of origin, subtype, germline / somatic origin of ctDNA, and primary / metastatic origin of ctDNA. With the overlay of software analysis and visualization tools, this new dimension of phenotype information can be

[0235] In one embodiment, the GenomicsNextTMassay includes genomic alteration test results that include pathogenic / actionable alterations covering a broad range of important cancer hotspot genes and tumor suppressor genes in ctDNA major variant classes (e.g., SNVs, INDELs, CNGs and gene rearrangements). In some cases, fragmentomics data from the same cfDNA sample and DNA sequencing run are used to deliver algorithmic expression index (i.e., “gene activity”) scores, a gene activity matrix (panel genes + other targeted regions), fragment distribution trends,and custom gene activity biomarkers (FIG.29 some cases, the hybrid capture enrichment stepwith targeted panels generated using the methods provided herein coupled with whole exome sequencing, can be used to gene expression / phenotype information from LP-WGS data (FIG.30). EXAMPLES

[0236] The present disclosure is further illustrated by reference to the following Examples. However, it should be noted that these Examples, like the embodiments described above, are illustrative and are not to be construed as restricting the scope of the invention in any way. Example 1- Development of ExpressCT Prospector (also called ExpressCT Decoder): Backbone discovery engine that efficiently ports signatures to ctDNA across tumor types. Introduction

[0237] The prevalence of short fragments of DNA in circulation and the subsequent realization that these fragments correspond to regions of open chromatin accessible to DNAses in blood led to the understanding that differences in the short DNA fragment size distributions and normalized amounts could be used to differentiate between different states, such as blood from people with and without cancer.Attorney Docket No. GNCN-026 / 01WO 320289-2157

[0238] Transcriptomics data can be used in different ways to measure gene expression: by quantifying the level of transcripts for an individual gene-of-interest or by training an algorithmic model to identify patterns among a set of transcripts to delineate a qualitative trait or phenotype, such as molecular phenotype. Figuring out how to translate established tissue gene expression algorithms into liquid biopsy equivalents is a multidimensional challenge as there are multiple liquid biopsy technologies (e.g., cfDNA-methylomics, cfDNA-seq, cfRNA-seq) that can serve as potential counterparts to tissue RNA-seq data. Translation of a tissue gene expression algorithm may occur under ideal circumstances such as where tumor and cfRNA collected from the same patient at the same time point are available for transcriptomic analysis. However, non-ideal constraints are likely and could include unmatched tumor and blood sample from different patients with a shared tumor type and different sequencing technologies (e.g., tumor RNA-seq and cfDNA- seq). The complexities and messiness of real-world conditions likely to be encountered when translating existing tumor gene expression signatures to liquid biopsy equivalents indicates a clear need for approaches capable of identifying signals in cfDNA that track gene expression variance in tissue with high fidelity. The methods provided throughout this specification provide a suite of methods for identifying cfDNA and ctDNA features that reflect tissue gene expression levels with high fidelity. Objective

[0239] Provide a proof of concept for a method for identifying a pool of genes analytically and biologically robust between tumor RNA expression data and ctDNA. Materials and Methods Representative Method for Selecting the set of genes whose expression patterns are highly likely to be mirrored in tissue and liquid biopsies for any specific cancer type.

[0240] For each of two data sets (i.e., TCGA CRC cohort (COAD; n=268) and liquid biopsy samples with methylation sequencing data from liquid samples and matching clinical diagnosis for CRC), the integrative correlation coefficient was calculated for each gene in accordance with Parmigiani G et al., “A cross-study comparison of gene expression studies for the molecular classification of lung cancer”. Clin Cancer Res. 2004 May 1;10(9):2922-7. doi:Attorney Docket No. GNCN-026 / 01WO 320289-2157 10.1158 / 1078-0432.ccr-03-0490. PMID: 15131026. In summary, within each of the two data sets, the correlation between each gene (x) and each other gene (y) was calculated. Subsequently, the cross-study correlation of correlations over all genes (y) for each gene (x) was then calculated, which represents the integrative correlation coefficient (ICC) for gene (x). Once the ICC was calculated for each gene in each of the two data sets, the genes were ranked by cross-platform integrative correlation coefficients for each gene and divided into quartiles (see FIG. 1). The genes in the top quartile by ICC value were selected as the top genes that behaved similarly across the two data sets (see Table 1). Results and Conclusions

[0241] Genes with larger, positive ICC values behaved similarly across the two data sets in that the within-data set correlation coefficients between a given gene and all other genes from the first data set were themselves highly-positively correlated with those from the second. FIGs 2-6 show images of gene-gene correlation coefficients using random subsets of genes from each quartile, as well as an image using the top-ranked 1K genes. Reproducible correlation structure emerged (from TCGA to liquid) in quartile 4, and in the top 1K genes the gene-gene correlation coefficients were remarkably similar across platforms. These top 1k genes were selected as the tumor specific genes that behaved the most similarly across the data sets and thus had the greatest potential for developing classifiers that would be predicted to behave similarly in tumor tissue samples as well as bodily fluid samples or liquid biopsies obtained from a subject suffering from the cancer.Attorney Docket No. GNCN-026 / 01WO 320289-2157 .recnaClatceroloCrof CCInodesabsenemg u84.843 7 4272779249 99640 483 8 3 9 4 9 7 0 9 0 9 0 9 7 2 4 0 .88100060 7 6 1 5 2 7 6 0 5 1 9 0 6 41dNe n917089602031408051704169049705310051162054037 8 3 9 0 2 578360602022401057knoisas0re_001 0 0 100 100 0 0 0 1 0 0 0 0 0 1 1 0 0 0 0 0 0 00cM C_N BM_NM_NM 0N_ _M0N_ _M_ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _CNMNMNMNMNMNM M M M M M M M M M M MB0cAMNMN N N N N N N N N N N N0N01poTl.ob1 m31A2A7N3el yAS0P 31A D3C632 1PMSBP 3TMP 0 5A5 72 SP3 2P226 IC G1 P1 3E1 MFTBKD AR K 3PLK1KT 2GT36 G 811 AFOLAS9NF QGR AaT MA ASR S CILEMHCFTI 3 Rb enL TAYLBZALCBAFNZMPCC S EHI URRMLS RHC P 1LI ATeGMT S PM r]e2b4m 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 721 1 1 1 1 1 1 1 1 1 2 2 2 2 2 20u2 20N[Att D kt N GNCN026 / 01WO 3202892157 4.35.0221200300_0M_NMN2A1GPTRIFS4956963.790264507410_10M 0N _MN47N1THNDCCP1623634.34.7077 0 4.8.822404328 6 2 6845181210 0307951451233570925815128508260.45 7 6 397 3 1 76 896 03 0 0 1 2 0 7 3 0813591704042617 1 7 7 5 2 0 417 1 273 31 0 0 0 0 0 1 0 1 0 1 0 2 0 0 000301010105120211 0 110 0_M_M_M_ 0MC_M_ _ 00_ _ _ _ _ _ _ _ _ _ _ _ _ _0_ 10_0_0_ 10_0_NMNMN_MM M M M M M M M C BN N N N BM M M M M M MBM M MAM MNN N N N N N N N N N N N N N N N N N N N642 3228 2 2 81B2A1L5O D2P PN18LF91TL 2DHF275480GKG5 GP DOP3 3N NA2 TNP 4T R 2HKD3N5PCDDP 8 16 A1P4OA D ELCX7 B PC6QG IF AOP 1L NS T GBUA PNCMRZQSMRID FSXPI S RF G 1MSBTLENH1PBCR SWNCSAR SLPPO PRG HMAOR B BPTG AA TNF AG NDC S82920313233343536373839304142434445464748494051525354555657585Attorney Docket No. GNCN-026 / 01WO 320289-2157320.049123744272606141741909762 24232910829197209 7 3 6 3 7 84 6 1 6 6 4 5 2 7 2 6 8 6 214 1 5 5 2 8 0 987377 2 0 6 01 1 1 1 0 0 1 3 1 1 0 2 7 040 0 0 3 5 5 0 3 1 041305 3 90 0 5 0 0 0 0 0 0 0 0 1 1 020 0 0 1 1 0 0 1 0 0 0 000001_M_ 0NM C_ _ _ _ _ _ _ _ 0 _ _0_ _ _ _ 0 _ _ _ _ _ _ _ _ _ _0_N BMNMNMNMNMNMNMNM 0N_MM M R M M M 0_M M M M M M M M M M MNN N N N N NMN N N N N N N N N N N N11 3A4 B3LL L1 8 3 52 K3 P12113 4 1 151 95261TA5 BG PAKL E8SS ALP DPBPU NCTSD ABRDLG 1FPP 5 C D4BA 6FD SD KP 1frCOE4M4VRR 1C1LU 3HCSAOP APNENDPTTGOSFAF5MKOH4DPNETF R URA oC S CMNZRAPR DHYLT LRM1C SFZOPDC A0616263646566676869607172737475767778797081828384858687888983.4736 .4 52 00 21 60 10 0_ _M MN N25CFNIGLDP64.74.7730490200_0M_NMNB1LTQPCCER3244245.24.15.779.81.41.65. . .354216314 .9.63376 .117.. . .339713984 .5.1 05 .0.1.3.5 13 .6.91.744 4 3 8 1 5 3 5 0 0 3 4 0 8 9 1 7 52868 297464812 47840 403003 2 2 4 8 3 2 312 3 209 815 0 0 832 5 103 2 8 5 310 5 73_00_05_10_00_07_11_00_0_10 005_1_10 10_70 0 390_0_0_10 200_0_10 000 1 2_0_0 0 10 8041 30M M M M M M M M M 0_M M 0_M Y M M M 0_M M 0_M M_M_M0_ _M_CN N N N N N N N NMN N NMN N A N N NMN N NM MBN N N N NMN N N55 0S7GF142 SPG252IRTK2 D11H SARQY3 M I 8L423ODD 83 7SN 1N2F412911P1 1NAA CGDCDOUCRBR SAGA5 S2RP C PAR NAPCH EPDIM OC NNE RN CAB ME XA OUT P RFIP2 U RS LBDTHBTCS GCTRL I TBF H P S SNC P PUTLS D U A0 1 2 3 4 5 6 7 8 900102030405060708090011 2 3 4 5 6 7 8 99 9 9 9 9 9 9 9 9 9 1 1 1 1 1 1 1 1 1 1 111111111111111111196285283119425536429131.9698352667030709716230.677087.1772 4464 3 1 3 8 3 4 3 8 7 8 1 0 4 5 5 9 9 4 0 9 0 0 8 3 3 256 42200 2 1 1 3 2 0 3 7 1 0 7 0 2 0 1 1 2 3 1 0 8 7 7 0_0 12 0 870M_0_0_0_0_0_0_1_1_0_0_010 0 0 0 0 0 0 0 0 0 7C_ _ _ _ _ _ _ _ _ _ 11U_ 00 0_0_300C_NMNMNMNMNMNMNMNMNMNMNMNBMNMNMNMNMNMNMNMNMNMNKMN_MM MBMNN N N221 439B A5 R812A 1 O31371 9 N42IH 1L D T OG31P 69 L A5I A2 41 1L6 XR TLM641 2 3D8 B L TL1B 6CCMT BA52 TSLHLCDCULGAS 1CBD OP KNGHV TX CGEAH OF GATMPR G AM J 8CE A SPNY3P MLMR CLMK C ABT AS R PS D N IBR3B AM NDE D A TS S C A T122232425262728292031323334353637383930414243444546474849 0 11 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 141512.03.5423439110_0M_NMNA408 OF STNCZ7188184.797311253401_10M 0N _MNBX9A4HfProXC4854846.54.0962212638210458231474306242017418221985571286439360728231980790922132014908438315 1 1 0 31 2 7 2 5 5 2 8 3 1 6 1 1 01 0 0 0 01000300000109 0 5 1 3 8 2 0 1 0 1 0 1 9 0 0 1 0 8 2_ _ _ _ _ _ _ _ _ _ _1_0_1_0_1_1_0 0 0 0 0 0 0 1 0 0 0 0 1 0M M M M M M M M M M M M M M M M M_M_R_M_M_M_M_M_ _ _ _ _ _ _N N N N N N N N N N N N N N N N N N N N N N N NMNMNMNMNMNMNMN A951 A73 1 01C 1B 7L 7 14K 812A 333F2A 3 6BR 27P4FMINN LGIN 1 CAM 2PO7S13E1SO1D D 2YR K3ARZ R P T ET B PXCK H AJP PMR2K9 D1 1PRF AOYLG GMHB 4FM BTDAAEPIMA UMU GBF DC RPTPS RSAOPGATM L3HIE R GM RTAF C Y R FNS P SMK S T G1525354555657585950616263646566676869607172737475 6 7 8 9 0 11 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1717171717181815.26.5120365010_0M_NMN3D FWGRHB8498483.05.1646012100_0M_NMN1A3 21 CP RT IA B5165155.84.03071651511324160120 21132230669687567065282904933004 3 8 1 5 1 2 3 2 6 4 9163836655825205230118308 3027687 52110398410 0 0 0 1 0 0200010201020000 0 019 1 7 3 2 0 01010 1 0_ _ _ _ 0 _ _ _ _ _ _ _ _ _0_0_0_ 01_0_1_0 0 0 00000 0 0M M M M 0_M M M M M M M M M M M M 0_M M M_M_M_M_0_ _0_ _ _ _N N N NMN N N N N N N N N N N N NMN N N N N N NMNM MNNM MNNMNMN112L141N43 AN 1K XPHB4R L A1 4113 5P45 -01 12 41 XZ 1LDBWNAI 1 2 2 21FZ C MLKCGF F 1 A C TYR S1IRS6111 C 2LCPN 3 2AOH RDL X XNLENLOD 1BEL 1BV ID MB N TGZ ELSRPLPPMDAEONAK2P FGRRDCTPI HGEDS D PHGEGFNT A CR DZWD R UK2838485868788898091929394959697989990010203040506 7 8 9 0 1 21 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 2 2 2 2 2 2020202021212123.946200_MNGC3KIP0881.702560CB2D0V6PTA7454.11 0 2 2 .15 7 4 3 0 9 6 5 7 4 8 6 6 645 9 5 264 8 1 9 2 8 0 4 1 0 0 0 6 93085244 0 3 3 5 2 9 62 2 2 193 6 7 4 0 2 6 6 9 5 1 9 7 1 4 3 29624324 1 4 2 08 0 8 203 0 0 1 0 0 0 1 3 1 0 9 1 0 1 7 1 0 0 30061104 31_0 1 061 0 0 0 0 0 0 0 1 0 101 0 0 0 1 0 0 0 0 0 0 00000M_ _ _8L _ _ _ _ _ _ _ _ _ _0NMNMNMN OMNMNMNMNMNMNMNMNMNMN_ _MM_M_M_M_M_M_M_M_M_M_M_M_M_RNN N N N N N N N N N N N N N2N 2545113 A RA1LN1 4312I 1 B 1 BL202FA2D7L 11T 61R4 1 2BNPCCA NDRH THEV TFSZPTNBAA EPF ACSG BRD VNEAA P GK FD1 LRRROCAAMT 2 UDRT E UAT FAY FKE TMSPMHSC MFTLNIYYMTDNAE EGGRS KPMP H P P EMR P41516171819102122232425262728292031323334353637383930 1 2 32 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2424242423.812.37560801_70M CN B2S 2SA1HPA0119194.28.3416122500_0M_NMNBC 13PAPRPFG7785757.24.0793564893949645690118454172 1 6 2 6.506 3 1 9 7 4 3 8 3 5 6 324 0 2 4 366 1 3 2 2 4 3 6 4 6 8 9389221061718 7 5 3 4 6 4 6 6 3 6 753 7 4 3 7 7 2 3 1 6 0 20 1 0 0 0 03110000010002 0 0 0 0 080 0 9 1 0 1 5 0 0 0 8 0_ _ _ _ _ _ _ _ _ _ _ _0_0_0_0_0_0_ 90 0 1 0 0 0 1 0 0 0 0 0M M M M M M M M M M M M M M M M M MM_M_M_M_M_M_M_ _ _ _ _ _N N N N N N N N N N N N N N N N N NON N N N N NMNMNMNMNMNMNA635 91 88 S 3 8 6D 1 1 1L1R N2 L1A21 23O1429f 2 K2 311 FI PA5 DP EZ VL3 NYLF3D 3HSLOLE 3FPS GBMPG R7LULNB F5OC ZL 7FBDT3 Fr AC Ao TXL 3DT SBRS HNP JA EOP CP YLSM LZKN PTT HLE RAML T S NBA BNZBZA 2SU 12ACPS PXE FPS R445464748494051525354555657585950616263646566676869 0 1 2 3 42 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 26272727272721.3.759296171100C_BMN01R1TNCLATI14246.64.5746811100_0M_NMN2FPAAEIX8096065.85.103982513800941217030511 51900944307540349848581805 8 9 5 6 2 8 4 8 1 90214733572217299699.175740 3.8 0 423 4 5774 1 37 2 354 4 20 0 0 0 0 1 00091100000514 2 911 0 0 3 0 0 112 321 9 0_ _ _ _ _ 0 _ _ _ _ _ _ _1_0_1_ 00_0_0_0 0 0 000 000 1 0M M M M M 0_M M M M M M M M M M 0_M M M_G_M_M_MC_M_C_ _ _N N N N NMN N N N N N N N N N NMN BMBM M MNN N N N N N N N N N N21 L 75 L8 214EB A L 1 1219T111P 6LM 7 3L 32 2 1P1 AC 1KBDG 83 C G D 1 H E C 1RBN7N N LI CAN 1RD M P B D BD 2DL X PK D R A NRD P AP PRMI -1PDKLS1F2B R I U C SI N NCC T Z FTSCCRPMS NCPNLDHILNO G MRT EAMC ERPMNFTRBDANTMTBA CED ASM N 5767778797081828384858687888980919293949596979899 0 1 2 3 4 52 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2920303030303033.392000_MNBKHP3794.928000_MN4AIRG04651..9044.63.05.27.53.74.03.73.654..023.4.4.4.2.7.3. 1.4.4.3.4.5.4.3.9 8 4 7 1 2 2 2 9 563 955005121612705 8211 1 8 0 1 8101072 096 3 8 9 6 0 5 6 5371422002615 9 3 00 0662642723059429689588675311694 8 3 3 7 2 30 1 1 0 0 0 0 0101031 100000000512 120 1 0 0 1 0 7_ 0M 0_ _ _ _ _ _ _ _ _ _ 0 _ _ _ _ _0_0_ 10_0_0_0_0_0_1_MM M M M M M M M M 0_M M CNN N N N N N N N N NM M M M M MBM M M M M M MNN N N N N N N N N N N N N NBH 4 1 1 7 2 2510 1 ES 58FK1B AL2GZPSKI AUTTPNT 1T 46P DQ8 LB NPAFS PNALS LDJC 4E N2FAF5S1PA R1C1K HOEHMNDPS NR UE BTG BOGT BH FJ GE S AA 0LG 8GAR AC PMO M OG V Z CMJAL P T P P ZRA TNDCOHNI7080900111213141516171819102122232425 6 7 8 9 0 1 2 33 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3232323232333333333Attorney Docket No. GNCN-026 / 01WO 320289-2157 *Each GenBank Accession Number is a representative or exemplary GenBank Accession Number for the listed gene and is herein incorporated by reference in its entirety for all purposes. Further, each listed representative or exemplary accession number should not be construed to limit the claims to the specific accession number.Attorney Docket No. GNCN-026 / 01WO 320289-2157 Example 2- Use of ExpressCT Prospector (ExpressCT Decoder) selected gene sets for developing microsatellite stable (MSS) predictive response signature (MSS-PRS) that is portable to classifying liquid biopsy samples from colorectal cancer patients. Objective

[0243] To use the ~1k tumor specific gene set developed for Example 1 for colorectal cancer to develop a classifier for determining whether a subject suffering from colorectal cancer possesses microsatellite instability in either tissue or liquid biopsy samples obtained from the subject. Materials & Methods

[0244] The final gene set was determined by building a nearest centroid classifier, using Clanc with cross-validation, predicting known expression subtype (yes / no) in TCGA with feature selection from the top 100 of the 1K tumor specific gene set developed in Example 1 (i.e., Table 1; FIG. 7). Intrinsic subtypes were trained on TCGA tumor tissue with the nearest centroid approach (Dabney AR. ClaNC: point-and-click software for classifying microarrays to nearest centroids. Bioinformatics.2006 Jan 1;22(1):122-3. doi: 10.1093 / bioinformatics / bti756. Epub 2005 Nov 2. PMID: 16269418) to find genes that had the highest agreement with gold standard classifier. In essence, the top 100 of qualified tumor-ctDNA methylome features were used to train a new model using ClaNC that was based on an established tumor tissue classifier (see FIG. 8, left panel). Following training, the prototype signature was qualified by projecting genes in the new signature onto colon cancer cfDNA samples. Expression subtypes were assigned to samples in the fragmentomics matrix using the final gene set (see Table 2). Visual inspection of matching profiles across data sets was conducted with attention to subtype prevalence to assign subtypes to the samples in the fragmentomics matrix. Results and Conclusions

[0245] As shown in the right panel of FIG.8, the profile and prevalence obtained from the cfDNA samples was as expected. As such, this Example shows the development of a surrogate tissue trained MSS-PRS for liquid (ctDNA) performance using methylome features (i.e., Table 2).Attorney Docket No. GNCN-026 / 01WO 320289-2157

[0246] Table 2. Surrogate tissue trained MSS-PRS for liquid (ctDNA) performance using methylome features.*Each GenBank Accession Number is a representative or exemplary GenBank Accession Number for the listed gene and is herein incorporated by reference in its entirety for all purposes. Further, each listed representative or exemplary accession number should not be construed to limit the claims to the specific accession number. Example 3- Use of ExpressCT Prospector (ExpressCT Decoder) selected gene sets for developing basal / classical subtyper that is portable to classifying liquid biopsy samples from pancreatic cancer patients. Objective

[0247] To use the methods described in Examples 1 and 2 to determine ~1k tumor specific gene set of genes whose expression patterns are highly similar in tissue biopsies and liquid biopsies obtained from patients suffering from pancreatic cancer and using the ~1k tumor specific gene set to develop a classifier for determining whether a subject suffering from pancreatic cancer possesses a basal or classical subtype.Attorney Docket No. GNCN-026 / 01WO 320289-2157 Materials & Methods

[0248] The methods described in Example 1 (e.g., use of integrative correlation coefficient) were used on each of two data sets (i.e., TCGA cohort (PAAD; n=150) and liquid biopsy samples with sequencing data from liquid samples and matching clinical diagnosis for PAAD to generate the top genes that behaved similarly across the two data sets for PAAD. In summary, within each of the two data sets, the correlation between each gene (x) and each other gene (y) was calculated. Subsequently, the cross-study correlation of correlations over all genes(y) for each gene (x) was then calculated, which represents the integrative correlation coefficient(ICC) for gene (x). Once the ICC was calculated for each gene in each of the two data sets, the genes were ranked by cross-platform integrative correlation coefficients for each gene and divided into quartiles. The genes with ICC values above X were selected as the top genes that behaved similarly across the two data sets (see Table 3).

[0249] Like in Example 2, the final gene set was determined by building a nearest centroid classifier, using Clanc with cross-validation, predicting known expression subtype (basal / classical) in TCGA with feature selection from the 1K tumor specific gene set developed in above (i.e., Table 3; FIG.9). Intrinsic subtypes were trained on TCGA PAAD tumor tissue with the nearest centroid approach (Dabney AR. ClaNC: point-and-click software for classifying microarrays to nearest centroids. Bioinformatics. 2006 Jan 1;22(1):122-3. doi: 10.1093 / bioinformatics / bti756. Epub 2005 Nov 2. PMID: 16269418) to find genes that had the highest agreement with gold standard classifier. In essence, the top 1000 of qualified tumor- ctDNA features were used to train a new model using ClaNC that was based on an established tumor tissue classifier (see FIG. 10, left and middle panel). Following training, the prototype signature was qualified by projecting genes in the new signature onto pancreatic cancer cfDNA samples (see FIG. 10, right panel). Expression subtypes were assigned to samples in the fragmentomics matrix using the final gene set (see Table 4). Visual inspection of matching profiles across data sets was conducted with attention to subtype prevalence to assign subtypes to the samples in the fragmentomics matrix. Results and ConclusionsAttorney Docket No. GNCN-026 / 01WO 320289-2157

[0250] As shown in the right panel of FIG. 10, the profile and prevalence obtained from the cfDNA samples was as expected. As such, this Example shows the development of a surrogate tissue trained Pancreatic cancer subtypes (PurIST) for liquid (ctDNA) performance (i.e., Table 4).Attorney Docket No. GNCN-026 / 01WO 320289-2157333333333343434343434343434343535353E UEUEUEUEUEUEUEUE E E E E E E E E ER RURU URU U URU U U UTRTRT TRTRTRTRT TRT TRTRT TRTRTRTRT5.15.45.96.25.83.14.85.4.575.3.6.5.4.4.764.5 6 8 16 1 2 2 5 0 9 7 3 9 3452 8 349803822961 1710650914283 880255160132 4 2 2013. 4 9 5 1 1 3013. 00 1 0 0 00000000009 00100000300001 0_ _ _ _ _ _ _ _ _0_ _ _0 0M M M M M M M M M_ _ _ _ _ _N N N N N N N N NMNMNMNMNMNMNMNMNMN71 A5 2 736 I1 R1 25 4 FB1M1 R 2 E2L933D D 2SFADK S HDC EOSCS CRPG MGAL UJO T GPC HHMST RARB1FDPHAC FIPLSMN POP C LMNT YMI2 3 4 5 6 7 8 90111213141516171819136240100_MN1FAHDS6864.503400_MN1NIB353EUEUEUEUEUEUEUEUEUEUEUE E E E E E E E E E E E E ER R R R R R R R R R RURURURURURURURURURURURU URUT T T T T T T T T T T T T T T T T T T T T TRT TRT5.64.4.5.3.3.4.5.4.5. 11. 24.5.3.4.2.4.344.5.4.864.4291459260598184008.454.80 2 2 3 7 5 9 2 5 8 6 2153 6 1 7 4 6 6 3 5 6 9 2553403833888906049734132051500050714 3 3 4 2152375 7 3 2 0 7 14. 0 1 0 12. 50 0 0 0 0 000000210 910 100100000001 06 2 3 2 05 2_ _ _ _ _ _ _ _ _ _ 1 00 0 0 0 0 0 0M M M M B_C_ _ _ _ _ _ _ _ _ _ _ _N N N NMNMNMNMNMNMNAMNBMNMNMNMNMNMNMNMNMNMNMNMN1T1RC 6 1 8221K D 8 N2C P1A 316 ARR 3401488 4ADSPFASN TPHMPIA AX 3XA2R 5R CVPIZANI 1CK A N1TMBPERITATDB MP E CML G 3SA C T FLIOT RMRFLU2C GAPAEESGA K P T GCHLSST NMT R021222324252627282920313233343536373839304142434441.5.743922755441R_CMN2LH0R9PRDDAW1 2173.84.9119112201_0M_NMN1PS 6D TT AC N879373EUEUEUEUEUEUEUEUEUEUEUEUEUEUEUEUEUEUEUEUEUEUE ER R R R R R R R R R R R R R R RTRTRU UT T T T T T T T T T T T T T T TRTRTRTRTRTRT3.42.75.33.74.94.45.23.14.64.45.44.83.44.04.33.24.24.43.05.34.2.3.7.3.2.0 3 2 8 4 7 5 8 8 0 9 0 9 85 9 4 4 1 93 3 2 5 9 9 2 9 34 5 4 9 9 0 9 3 8 2 4 03 4 6 0 2 0 1 0 31281648052960546156824346823662 51 0 1 3 0 3 2 3 3 2 0 2 0 3 0 0 0 01 40 0 0 0 0 0 0 0 0 0 0 03 0 0 3 5 0 2 2_ _ _ _ _ _0 0 0 0 0 0 1 0 0 1 0 0 0 0M_ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _NMNMNMNMNMNMNMNMNMNMNMNMNMNMNMNMNMNMNMNMNMNMNMNMNMN1C 2 3 2T2H NINI1NI8 B1L 1E1 2 1 BB1T 1211N1D3HGDSSLCAL PGPRR 1AB1 L1ABAAFCIK LI B7S1X3CP P3ICNNY N LK N ZINZOAP4BMHS R RS E HLEHO AARBEIR MB 2L KIF Y 1WFR MP C N S R D C GKP DCBWMZ L54647484940515253545556575859506162636465666768696071.5.65.24.92.34.35.3.3. 1.4.5.5.4.4.5.54 2.5.4.3.3.3.4 0 0 5 4 7875534 38 7 1 7 7 5 151 9 1 4 83 7 7 2 9 3 4 9 4 8622663910 7 894 2 4 0 37 4 6 2 4 3 6 0 3 3 5 2 8 567802 287 4 2 0 08 2 0 3 0 1 0 0 0 6 2 3 3 1 1 81 .72023523 2100 0 0 0 0 0 0 0 5 0 0 1 0 0 0030 0 0 05030C_ _ _ _ _ _ _ _ 4R_ _ _ _ _ _0_ 0C_ _ _ _ _BMNMNMNMNMNMNMNMNCMNMNMNMNMNMNMNBMNMNMNMNMN 2 D8 6 4 8 1 1 212 H6D6AFL2 DPRWCHL2R7E1BA1 4SM1PM2AS 12 3 8C R CBID D 12 6 LNS NX11BSMAND H S G TRE BZNPSSA 1P B HWMDDZAGTPBZ AMT SRL PPM P MB KTA PZLBUNNGPS273747576777879708182838485868788898091929394964080100_MNGPSA1674.669410_MN03XHD8245.64.23.54.01.4.67.96.22.5.4.4.6.4.3.4.5.3.4.4.3.795.3.8 9 6 4 3 4 1 921918933670827244 3 7 0 9 9 5 6987 2 0 0 4 3 8 9 3 9 9 1 6 9 020119771732 54131233530522 1 9 0 6 8 1 6 0 4 4 7 3 7 1 3113. 0 80 0 1 0 30000102000102 1 2 1 2 0 3 1 2 1 09 0 1_ _ _ _ 0 _ _ _ _ _ _0_0_0_0_0_0_0_0_0 0 0 0 0M MNMNM CNMNMNMNMNM_ _ _ _ _N B NMNMNMNMNMNMNMNMNMNMNMNMNMNMN1C5 21G H 1 8 LT43B 3L2A4CP K3P1E5 C T 50K2P P 41W1PIWII FL5NY2F R 1TM2 256 D 4 71 162HTMC B S PE 1KLSBZ ELASNA N P YDXABMDSC CE E FT EHPMSCNILF C F S TMMD CTNZ BYSTS HLK5 6 7 8 900102030405060708090011 2 3 4 5 6 7 89 9 9 9 9 1 1 1 1 1 1 1 1 1 1 111111111111111113.940810_MN1JHKELP58798900100_MN92fro21C2543.04.44.63.43.4 025.34.94.1 824.5.5.4.4.74144.3.4.444.1.5.5.7 0 4 3 058 4 78 98541052999 0 6856586 028 135328 6 4 2013603 3 3 12.0838442013.022545039340313.114.0377155211.78542094620 000023007 103 2 02 2 4 1 0 4 06 03 0 1 4 00 0 7 2 0_ _ _ _ _0_1_0_0 0_1_0_0_1_0 0 0_0_1_0 0 1 0 0MNMNMNMNM_N NMNM_ _ _ _ _ _N N N N N N N N N N N N N N NR_ _MM M M M M M M M M M M M M MNMNMN4133 21 D 477 B4DZFRK L5 2 5C851I 41PAL 7 L54 3 1 C 2K4f192 4E2NAGR FM3 3 2H31 5ICAG C LRN P CL 2Pro C C 1RRLZTS DEIMFLRAN5ZPTDM CEC PT VED3 RA HHSR PN AL NC QA 9 PSNS PM1CRRARLRPOP9102122232425262728292031323334 5 6 7 8 9 0 1 2 31 1 1 1 1 1 1 1 1 1 1 1 1 1 1313131313131414141414.733410_MN2LIPP0184.378710_MN6BSA7744.34.74.65.52.93.83.55.11.7.94.54.04. 1.2. 1.5.3.4.5.1.5.5.2.4.4.8 2 4 7 2 4 6 8 6 4 5 778 871 4407659208 98576 1 1030 3 8 2 3 2 1 3 7 4 48 6 0 205070570010212181 5 012 0 8 2 4932 62 1 6 5 7 7 7 4 8 444 8 0 3 0 2 5 8 3 30 0 0 1 0 0 0 0 420203100 400 40050207 6 3 0 3 0 3_M_ _ _ _ _ _ _ 1C_ _ _ _ 1 _ 0 _ _ _1_1_0_0_1_0_0_NMNMNMNMNMNMNMNBMNMNMNM CNBM CNBMNMNMNMNRNMNMNMNMNMN2C2PR 33071132 6 2 0 3 71A3RF 5XT KD 8IA PD P8BB 6C1231 ANK R 830NPGR RF NPAQJ 1G2EO3 8fXSS r 1 S 3oP BR8FA1A2XSIGPOSCLABDH K BMBR F AAFPC GGLSR FARG IAPBUERRP 7 R1CAPOSNZPGNS4454647484940515253545556575859506162 3 4 5 6 7 8 91 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 16161616161616161050505050505 5 5 5 5 5 5 5 5 5 5 5 5 5 5 5 5 5 52.5.44.16.44.32.3.83.24.24.13.64.15.55.8 395.4.6.3.5.4.4.3.5.4 0 0 7 4 8 2 2 4 1 9 2 2 52 434320375906212 71 3 1 1 6 5 3 8 4 1 1 01 0 22 84 9 03 8 3 9 5 8 4 1 6 5413008161 110228 6 0 3 6 1 1 . 2 8 4 0 0 2 4 6 21 0 0 0 0 1 0 031102000000004 2 3 1 3 2 0 2 1 30 _ _ _ _ 0 _ _ _ _ _ _ _ _0 0_1_0_0_0_0_0_0 0C M MNMNM CNMNMNMNMNMNM_ _ _B N B NMNMNMNMNMNMNMNMNMNMNMNMN1S 146A H84711 62PHAB 1 7 1B1P 2 21 22 A 7Pf 12 PPD1f 01TA AODY P 7CA F T B CL C OL B T UMroADSAAMAM rR PMNZBATC K DC E AXGSLSYT 6W OGC T1C V JLDHT ARU oGM9CCCGZ17273747576777879708182838485868788 9 0 1 2 3 41 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1818191919191914.97.74.25.4 363.9 8 21 93.65.24.5 305.92. 1.3.3. 2.4.3.4.4.1 6 8 9597 78 1 72010 75998 981821 6512 0 983 87010 6 7 473 4 3 4 1 5 8 4 2196307000050102._0_0 0 08251103.08102.01376106001001. 403 030180141042370212255110 0 000 0 1 0M M_M_M_ _M_ _ _M_M_M_ _ _F_ _C_ _ _ _N N N NMNNMNMNN N ...

Claims

Attorney Docket No. GNCN-026 / 01WO 320289-2157 CLAIMS What is claimed:

1. A computer-implemented method for determining a gene expression signature in disparatesamples from subjects suffering from a cancer of interest, the method comprising:(a) receiving in a computer system, a first set of nucleic acid expression data for each of a pluralityof tissue samples obtained from a first population of subjects suffering from or suspected of suffering from the cancer of interest and a second set of nucleic acid expression data for each of a plurality of bodily fluid samples obtained from a second population of subjects suffering from or suspected of suffering from the cancer of interest, wherein each of the plurality of bodily fluid samples comprises cell-free DNA (cfDNA);(b) determining by the computer system, dependence relationships between each gene within thefirst set of nucleic acid expression data to generate a first set of dependence relationships and dependence relationships between each gene within the second set of nucleic acid expression data to generate a second set of dependence relationships; and(c) selecting by the computer system, each gene for which the dependence relationships for arespective gene from the first set of dependence relationships is substantially similar to the dependence relationships for the respective gene in the second set of dependence relationships, thereby generating a gene signature that comprises each gene selected by the computer system to possess substantially similar dependence relationships between the first set of nucleic acid expression data for the plurality of tissue samples and the second set of nucleic acid expression data for the plurality of bodily fluid samples.

2. A computer-implemented method for generating a classifier for a desired feature of a cancer ofinterest in a subject suffering from the cancer of interest or suspected of suffering from the cancer of interest, the method comprising:(a) receiving in a computer system, a first set of nucleic acid expression data for each of a pluralityof tissue samples obtained from a first population of subjects suffering from or suspected of suffering from the cancer of interest and a second set of nucleic acid expression data for each of a plurality of bodily fluid samples obtained from a second population of subjects suffering from or suspected of suffering from the cancer of interest, wherein each of the plurality of bodily fluid samples comprises cell-free DNA (cfDNA);Attorney Docket No. GNCN-026 / 01WO 320289-2157 (b) determining by the computer system, dependence relationships between each gene within the first set of nucleic acid expression data to generate a first set of dependence relationships and dependence relationships between each gene within the second set of nucleic acid expression data to generate a second set of dependence relationships; (c) selecting by the computer system, each gene for which the dependence relationships for a respective gene from the first set of dependence relationships is substantially similar to the dependence relationships for the respective gene in the second set of dependence relationships, thereby generating a gene signature that comprises each gene selected by the computer system to possess substantially similar dependence relationships between the first set of nucleic acid expression data for the plurality of tissue samples and the second set of nucleic acid expression data for the plurality of bodily fluid samples; (d) inputting nucleic acid expression data for each gene in the gene signature from (c) from at least two training sets, wherein one of the at least two training sets comprises nucleic acid expression data for each gene in the gene signature from (c) from each of a plurality of tissue samples from subjects that are indicative of the presence a desired feature for the cancer of interest, while another of the at least two training sets comprises nucleic acid expression data for each gene in the gene signature from (c) from each of a plurality of tissue samples from subjects that are indicative of the absence of the desired feature for the cancer of interest; and (e) conducting on the computer system a linear discriminate analysis (LDA) comprising feature selection to generate a classifier for the desired feature of the cancer of interest, wherein the feature selection comprises: (i) calculating a test statistic (t-statistic) for each gene in the gene signature from (c) from the at least two training sets, wherein the t-statistic for each gene indicates each gene's ability to distinguish between the presence or the absence of the desired feature; (ii) ranking each gene based on each gene’s t-statistic; and (iii) selecting each gene whose t-statistic is above a desired threshold for distinguishing between the presence or the absence of the desired feature, thereby generating a classifier for the desired feature comprise each of the selected genes.

3. The method of claim 2, further comprising: (f) classifying on the computer system one or more test samples obtained from an independent population of subjects suffering the cancer of interest as possessing or notAttorney Docket No. GNCN-026 / 01WO 320289-2157 possessing the desired feature using the classifier from (e)(iii) on the one or more test samples; (g) comparing on the computer system, the classification of the one or more test samples to classification of the one or more samples for the desired feature as determined using a control classifier of the desired feature for the cancer of interest; and (h) validating on the computer system the classifier from (e)(iii) if the comparing indicates that the classifications of the one or more test samples is substantially similar to classification of the one or more test samples determined using the control classifier.

4. The method of any one of the above claims, wherein the dependence relationships between each gene within the first set of nucleic acid expression data comprises all possible pairwise correlations between each gene within the first set of nucleic acid expression data and each other gene within the first set of nucleic acid expression data and the dependence relationships between each gene within the second set of nucleic acid expression data comprises all possible pairwise correlations between each gene within the second set of nucleic acid expression data and each other gene within the second set of nucleic acid expression data.

5. The method of claim 4, wherein the substantial similarity in (c) is evidenced by an integrative correlation coefficient (ICC) for the respective gene that is above a desired threshold, wherein the ICC is a correlation of the pairwise correlations determined for the respective gene within the first set of nucleic acid expression data and the pairwise correlations determined for the respective gene within the second set of nucleic acid expression data.

6. The method of claim 5, wherein the desired threshold is the 99thpercentile of a null distribution of integrative correlations.

7. A computer-implemented method for determining a gene expression signature in disparate samples from subjects suffering from a cancer of interest, the method comprising: (a) receiving in a computer system, a first set of nucleic acid expression data for each of a plurality of tissue samples obtained from a first population of subjects suffering from or suspected of suffering from the cancer of interest and a second set of nucleic acid expression data for each of a plurality of bodily fluid samples obtained from a second population of subjects suffering from orAttorney Docket No. GNCN-026 / 01WO 320289-2157 suspected of suffering from the cancer of interest, wherein each of the plurality of bodily fluid samples comprises cell-free DNA (cfDNA); (b) determining by the computer system, an integrative correlation coefficient for each gene from the first and second set of nucleic acid expression data, wherein the integrative correlation coefficient is a numerical representation of how each gene from the first set of nucleic acid expression data correlate with each other compared to the rank of the same gene from the second set of nucleic acid expression data; (c) ranking by the computer system, the integrative correlation coefficient of each gene from step (b); and (d) selecting, by the computer system, genes that have an integrative correlation coefficient above a desired threshold to generate a gene signature, wherein the gene signature comprises genes whose expression patterns are substantially similar between the tissue sample and the bodily fluid sample.

8. A computer-implemented method for generating a classifier for determining a desired feature of a cancer of interest in a subject suffering from the cancer of interest or suspected of suffering from the cancer of interest, the method comprising: (a) receiving in a computer system, a first set of nucleic acid expression data for each of a plurality of tissue samples obtained from a first population of subjects suffering from or suspected of suffering from the cancer of interest and a second set of nucleic acid expression data from each of a plurality of bodily fluid samples obtained from a second population of subjects suffering from or suspected of suffering from the cancer of interest, wherein each of the plurality of bodily fluid samples comprises cell-free DNA (cfDNA); (b) determining by the computer system, an integrative correlation coefficient for each gene from the first and second set of nucleic acid expression data, wherein the integrative correlation coefficient is a numerical representation of how each gene from the first set of nucleic acid expression data correlate with each other compared to the rank of the same gene from the second set of nucleic acid expression data; (c) ranking by the computer system, the integrative correlation coefficient of each gene from step (b); (d) selecting, by the computer system, genes that have an integrative correlation coefficient above a desired threshold to generate a gene signature, wherein the gene signature comprises genes whose expression patterns are substantially similar between the tissue sample and the bodily fluid sample;Attorney Docket No. GNCN-026 / 01WO 320289-2157 (e) inputting nucleic acid expression data for each gene in the gene signature from (d) from at least two training sets, wherein one of the at least two training sets comprises nucleic acid expression data for each gene in the gene signature from (d) from each of a plurality of tissue samples from subjects that are indicative of the presence a desired feature for the cancer of interest, while another of the at least two training sets comprises nucleic acid expression data for each gene in the gene signature from (d) from each of a plurality of tissue samples from subjects that are indicative of the absence of the desired feature for the cancer of interest; and (f) conducting on the computer system a linear discriminate analysis (LDA) comprising feature selection to generate a classifier for the desired feature of the cancer of interest, wherein the feature selection comprises: (i) calculating a test statistic (t-statistic) for each gene in the gene signature from (d) from the at least two training sets, wherein the t-statistic for each gene indicates each gene's ability to distinguish between the presence or the absence of the desired feature; (ii) ranking each gene based on each gene’s t-statistic; and (iii) selecting each gene whose t-statistic is above a desired threshold for distinguishing between the presence or the absence of the desired feature, thereby generating a classifier for the desired feature comprise each of the selected genes.

9. The method of claim 8, further comprising: (g) classifying on the computer system one or more test samples obtained from an independent population of subjects suffering the cancer of interest as possessing or not possessing the desired feature using the classifier from (f)(iii) on the one or more test samples; (h) comparing on the computer system, the classification of the one or more test samples to classification of the one or more samples for the desired feature as determined using a control classifier of the desired feature for the cancer of interest; and (i) validating on the computer system the classifier from (f)(iii) if the comparing indicates that the classifications of the one or more test samples is substantially similar to classification of the one or more test samples determined using the control classifier.

10. The method of claim 7 or 8, wherein the ICC is above the desired threshold of the 99thpercentile of a null distribution of integrative correlations.Attorney Docket No. GNCN-026 / 01WO 320289-2157 11. The computer-implemented method of any one of claims 2-3 or claims 8-9, wherein the desired feature is selected from the group consisting of a subtype of the cancer of interest, microsatellite instability, mismatch repair proficiency, minimal residual disease, prognostic risk, response to a therapeutic agent, development of resistance to a therapeutic agent.

12. The computer-implemented method of any one of claims 2-3 or claims 8-9, wherein the first and the second population of subjects consist of the same subjects.

13. The computer-implemented method of any one of claims 2-3 or claims 8-9, wherein the tissue samples are tumor tissue samples selected from the group consisting of a formalin-fixed, paraffin- embedded (FFPE) tissue sample, a fresh tissue sample and a frozen tissue sample.

14. The computer-implemented method of any one of claims 2-3 or claims 8-9, wherein the bodily fluid is selected from the group consisting of whole blood, plasma, serum, an exosome, wash fluids, urine, saliva, cerebrospinal fluid or sputum.

15. The computer-implemented method of any one of claims 2-3 or claims 8-9, wherein the first set of nucleic acid expression data is nucleic acid sequencing data or methylome sequencing data obtained from nucleic acid extracted from each tissue sample from the plurality of tissue samples.

16. The computer-implemented method of any one of claims 2-3 or claims 8-9, wherein the second set of nucleic acid expression data is nucleic acid sequencing data or methylome sequencing data obtained from nucleic acid extracted from each bodily fluid sample from the plurality of bodily fluid samples.

17. The computer-implemented method of claims 15, wherein the nucleic acid sequencing data is DNA sequencing data or RNA sequencing data.

18. The computer implemented method of any one of claims 2-3 or claims 8-9, wherein the second set of nucleic acid expression data comprises a fast Fourier transform (FFT) magnitude matrix for each of a plurality of genomic regions of interest.

19. The computer-implemented method of claim 18, wherein the FFT magnitude matrix for each of the plurality of genomic regions of interest is generated by the computer system is a method that comprises:Attorney Docket No. GNCN-026 / 01WO 320289-2157 (i) inputting into the computer system, nucleic acid sequencing data generated from nucleic acid extracted from a each of the bodily fluid samples, wherein the nucleic acid sequencing data includes a plurality of fragment reads, wherein each fragment read has a fragment length and a GC content indicating a percentage of bases in the fragment read that are G or C; (ii) determining by the computing system, GC bias values for each fragment read based on the fragment length and the GC content of the fragment read; (iii) generating by the computing system, a genomic coverage distribution that is adjusted for GC bias using the sequence read data and the GC bias values; (iv) calculating by the computing system mean sequence read counts for a sliding window across a defined window in each of the plurality of genomic regions of interest from the genomic coverage distribution to generate smoothed mean read counts; and (v) performing a fast Fourier transform (FFT) on the smoothed mead read counts to generate the FFT magnitude matrix for each of the plurality of genomic regions of interest.

20. The method of claim 19, wherein the sliding window has a width of at least, at most or exactly 5, 10, 15, 20, 25, 30, 35, 40, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95 or 100 nucleotides across the defined window.

21. The method of claim 19, wherein the sliding window has a width of 15 nucleotides across the defined window.

22. The method of claim 19, wherein the defined window has a width of 2000 base pairs.

23. The method of claim 1, wherein prior to (a), the second set of nucleic acid expression data is subjected to a method comprising (a) extracting gene expression information across the dataset by (i) mapping the sequence reads from the dataset to a reference human genome, thereby generating read count data across the mapped genome; and (ii) applying fast Fourier transformation (FFT) to the read count data in nucleosome occupancy windows across the mapped genome to determine an FFT signal at each nucleosome occupancy window, wherein an increased FFT signal is indicative of nucleosomal depletion and a decreased FFT signal is indicative of nucleosomal presence; and (b) performing quality control of the cfDNA dataset comprising removal of samplesAttorney Docket No. GNCN-026 / 01WO 320289-2157 from the dataset that possess a low or inconsistent FFT signal as determined in (b)(i)-(b)(ii) and / or a measured circulating tumor DNA (ctDNA) content below a dataset-specific threshold, thereby generating a quality-controlled cfDNA dataset comprised of features that are reflective of gene activity or expression.

24. The method of any one of claims 2-3 or claims 8-9, wherein the cancer of interest is selected from the group consisting of kidney renal papillary cell carcinoma (KIRP); breast invasive carcinoma (BRCA); thyroid cancer (THCA); bladder urothelial carcinoma (BLCA); prostate adenocarcinoma (PRAD); kidney chromophobe (KICH); cervical squamous cell carcinoma and endocervical adenocarcinoma (CESC); kidney renal clear cell carcinoma (KIRC); liver hepatocellular carcinoma (LIHC); low grade glioma (LGG); sarcoma (SARC); lung adenocarcinoma (LUAD); colon adenocarcinoma (COAD); head and neck squamous cell carcinoma (HNSC); uterine corpus endometrial carcinoma (UCEC); glioblastoma multiforme (GBM); esophageal carcinoma (ESCA); stomach adenocarcinoma (STAD); ovarian serous cystadenocarcinoma (OV); rectum adenocarcinoma (READ); adrenocortical carcinoma (ACC); uveal melanoma (UVM); mesothelioma (MESO); pheochromocytoma and paraganglioma (PCPG); skin cutaneous melanoma (SKCM); uterine carcinosarcoma (UCS); lung squamous cell carcinoma (LUSC); testicular germ cell tumors (TGCT); cholangiocarcinoma (CHOL); pancreatic adenocarcinoma (PAAD); thymoma (THYM); or Lymphoid Neoplasm Diffuse Large B-cell Lymphoma (DLBC).

25. The computer-implemented method of any one of claims 2-3 or claims 8-9, wherein the cancer of interest is COAD.

26. The computer-implemented method of claim 25, wherein the first set and the second set of nucleic acid expression data is nucleic acid sequencing data and the gene signature generated comprises, consists essentially of or consists of the genes in Table 1.

27. The computer-implemented method of claim 1, wherein the cancer of interest is PAAD.Attorney Docket No. GNCN-026 / 01WO 320289-2157 28. The computer-implemented method of claim 27, wherein the first set and the second set of nucleic acid expression data is nucleic acid sequencing data and the gene signature comprises, consists essentially of or consists of the genes in Table 3.

29. The computer-implemented method of claim 1, wherein the cancer of interest is BLCA.

30. The computer-implemented method of claim 29, wherein the first set and the second set of nucleic acid expression data is nucleic acid sequencing data and the gene signature comprises, consists essentially of or consists of the genes in Table 5.

31. The computer-implemented method of any one of claims 2-3 or claims 8-9, wherein the cancer of interest is COAD, the desired feature is microsatellite instability and the classifier generated comprises, consists essentially of or consists of the genes in Table 2.

32. The computer-implemented method of any one of claims 2-3 or claims 8-9, wherein the cancer of interest is PAAD, the desired feature is a PAAD subtype, and the classifier generated comprises, consists essentially of or consists of the genes in Table 4.

33. The computer-implemented method of any one of claims 2-3 or claims 8-9, wherein the cancer of interest is BLCA, the desired feature is an FGFR activation signature, and the classifier generated comprises, consists essentially of or consists of the genes in Table 6.

34. A method of assaying a sample obtained from a subject suffering from COAD, the method comprising measuring the expression level of a plurality of biomarkers selected from Table 1 or Table 2 using a sequencing assay.

35. The method of claim 34, wherein the sequencing assay is a DNA sequencing assay or RNA sequencing assay.

36. The method of claim 34, wherein the plurality of biomarkers selected from Table 1 comprises at least 50 biomarkers, at least 100 biomarkers, at least 150 biomarkers, at least 200 biomarkers,Attorney Docket No. GNCN-026 / 01WO 320289-2157 at least 250 classifier biomarkers, at least 300 biomarkers, at least 350 biomarkers, at least 400 biomarkers, at least 450 biomarkers, at least 500 biomarkers, at least 550 biomarkers, at least 600 biomarkers, at least 650 biomarkers, at least 700 biomarkers, at least 750 biomarkers, at least 800 biomarkers, at least 850 biomarkers, at least 900 biomarkers or at least 950 biomarkers from Table 1.

37. The method of claim 34, wherein the plurality of classifiers selected from Table 1 comprises at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% of the classifiers from Table 1.

38. The method of claim 34, wherein the plurality of classifiers selected from Table 1 consists of all the classifiers of Table 1.

39. The method of claim 34, wherein the plurality of classifiers selected from Table 2 comprises at least 2 classifiers, at least 4 classifiers, at least 6 classifiers, at least 8 classifiers, at least 10 classifier biomarkers, at least 12 classifiers, at least 14 classifiers, at least 16 classifiers, at least 18 classifiers, at least 20 classifiers, at least 22 classifiers, at least 24 classifiers, at least 28 classifiers, at least 30 classifiers, at least 32 classifiers, at least 34 classifiers, at least 36 classifiers, at least 38 classifiers, at least 40 classifiers, at least 42 classifiers, at least 44 classifiers, at least 46 classifiers, at least 48 classifiers, at least 50 classifiers, at least 52 classifiers, at least 54 classifiers, at least 56 classifiers, at least 58 classifiers, at least 60 classifiers, at least 62 classifiers, at least 64 classifiers, at least 66 classifiers or at least 68 classifiers from Table 2.

40. The method of claim 34, wherein the plurality of classifiers selected from Table 2 comprises at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% of the classifiers from Table 2.

41. The method of claim 34, wherein the plurality of classifiers selected from Table 2 consists of all the classifiers of Table 2.Attorney Docket No. GNCN-026 / 01WO 320289-2157 42. The method of claim 34, wherein the sample is a formalin-fixed, paraffin-embedded (FFPE) tissue sample from the colon of the subject, fresh or a frozen tissue sample from the colon of the subject, an exosome, wash fluids, cell pellets, or a bodily fluid obtained from the subject.

43. The method of claim 39, wherein the bodily fluid is whole blood, plasma, serum, an exosome, wash fluids, urine, saliva, cerebrospinal fluid or sputum.

44. A method of treating COAD in a subject, the method comprising: determining microsatellite instability of COAD of a subject suffering from COAD by measuring a nucleic acid expression level of a plurality of classifier biomarkers in a sample obtained from a subject suffering from or suspected of suffering from COAD, wherein the plurality of classifier biomarkers is selected from Table 2, wherein the nucleic acid expression level of the plurality of classifier biomarkers indicates the presence of microsatellite instability (MSI); and administering a therapeutic intervention based on the presence or absence of MSI, wherein the therapeutic intervention is immune checkpoint inhibitor therapy (ICI) when MSI is present (MSI positive) or an immuno-oncology (IO) treatment and / or chemotherapy if MSI is absent (MSI negative).

45. The method of claim 44, wherein the determining step further comprises comparing the nucleic acid expression levels of the plurality of classifier biomarkers from Table 2 to the nucleic acid expression levels of the plurality of classifier biomarkers from Table 2 in at least one sample training set(s), wherein the at least one sample training set comprises nucleic acid expression level data of the plurality of classifier biomarkers from Table 2 from a reference MSI positive sample, nucleic acid expression level data of the plurality of classifier biomarkers from Table 2 from a reference MSI negative sample or a combination thereof; and classifying the sample obtained from the subject as MSI positive or MSI negative based on the results of the comparing step.

46. The method of claim 45, wherein the comparing step comprises applying a statistical algorithm which comprises determining a correlation between the expression data obtained from the sample obtained from the subject and the expression data from the at least one training set(s); andAttorney Docket No. GNCN-026 / 01WO 320289-2157 classifying the sample obtained from the subject as MSI positive or MSI negative based on the results of the statistical algorithm.

47. The method of claim 44, wherein the plurality of classifiers selected from Table 2 comprises at least 2 classifiers, at least 4 classifiers, at least 6 classifiers, at least 8 classifiers, at least 10 classifier biomarkers, at least 12 classifiers, at least 14 classifiers, at least 16 classifiers, at least 18 classifiers, at least 20 classifiers, at least 22 classifiers, at least 24 classifiers, at least 28 classifiers, at least 30 classifiers, at least 32 classifiers, at least 34 classifiers, at least 36 classifiers, at least 38 classifiers, at least 40 classifiers, at least 42 classifiers, at least 44 classifiers, at least 46 classifiers, at least 48 classifiers, at least 50 classifiers, at least 52 classifiers, at least 54 classifiers, at least 56 classifiers, at least 58 classifiers, at least 60 classifiers, at least 62 classifiers, at least 64 classifiers, at least 66 classifiers or at least 68 classifiers from Table 2.

48. The method of claim 44, wherein the plurality of classifiers selected from Table 2 comprises at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% of the classifiers from Table 2.

49. The method of claim 44, wherein the plurality of classifiers selected from Table 2 consists of all the classifiers of Table 2.

50. The method of claim 44, wherein the sample is a formalin-fixed, paraffin-embedded (FFPE) tissue sample from the colon of the subject, fresh or a frozen tissue sample from the colon of the subject, an exosome, wash fluids, cell pellets, or a bodily fluid obtained from the subject.

51. The method of claim 50, wherein the bodily fluid is whole blood, plasma, serum, an exosome, wash fluids, urine, saliva, cerebrospinal fluid or sputum.

52. A method of assaying a sample obtained from a subject suffering from PAAD, the method comprising measuring the expression level of a plurality of classifiers selected from Table 3 or Table 4 using a sequencing assay.Attorney Docket No. GNCN-026 / 01WO 320289-2157 53. The method of claim 52, wherein the sequencing assay is a DNA sequencing assay or RNA sequencing assay.

54. The method of claim 52, wherein the plurality of biomarkers selected from Table 3 comprises at least 50 biomarkers, at least 100 biomarkers, at least 150 biomarkers, at least 200 biomarkers, at least 250 classifier biomarkers, at least 300 biomarkers, at least 350 biomarkers, at least 400 biomarkers, at least 450 biomarkers, at least 500 biomarkers, at least 550 biomarkers, at least 600 biomarkers, at least 650 biomarkers, at least 700 biomarkers, at least 750 biomarkers, at least 800 biomarkers, at least 850 biomarkers, at least 900 biomarkers or at least 950 biomarkers from Table 3.

55. The method of claim 52, wherein the plurality of classifiers selected from Table 3 comprises at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% of the classifiers from Table 3.

56. The method of claim 52 , wherein the plurality of classifiers selected from Table 3 consists of all the classifiers of Table 3.

57. The method of claim 52, wherein the plurality of classifiers selected from Table 4 comprises at least 2 classifiers, at least 4 classifiers, at least 6 classifiers, at least 8 classifiers, at least 10 classifier biomarkers, at least 12 classifiers, at least 14 classifiers, at least 16 classifiers, at least 18 classifiers, at least 20 classifiers, at least 22 classifiers, at least 24 classifiers, at least 28 classifiers, at least 30 classifiers, at least 32 classifiers, at least 34 classifiers, at least 36 classifiers, at least 38 classifiers, at least 40 classifiers, at least 42 classifiers, at least 44 classifiers, at least 46 classifiers, at least 48 classifiers, at least 50 classifiers, at least 52 classifiers, at least 54 classifiers, at least 56 classifiers, at least 58 classifiers, at least 60 classifiers, at least 62 classifiers, at least 64 classifiers, at least 66 classifiers, at least 68 classifiers, at least 70 classifiers, at least 72 classifiers, at least 74 classifiers, at least 76 classifiers, at least 78 classifiers or at least 80 classifiers from Table 4.Attorney Docket No. GNCN-026 / 01WO 320289-2157 58. The method of claim 52, wherein the plurality of classifiers selected from Table 4 comprises at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% of the classifiers from Table 4.

59. The method of claim 52, wherein the plurality of classifiers selected from Table 4 consists of all the classifiers of Table 4.

60. The method of claim 52, wherein the sample is a formalin-fixed, paraffin-embedded (FFPE) tissue sample from the pancreas of the subject, fresh or a frozen tissue sample from the pancreas of the subject, an exosome, wash fluids, cell pellets, or a bodily fluid obtained from the subject.

61. The method of claim 60, wherein the bodily fluid is whole blood, plasma, serum, an exosome, wash fluids, urine, saliva, cerebrospinal fluid or sputum.

62. A method of treating PAAD in a subject, the method comprising: determining subtype of PAAD of a subject suffering from PAAD by measuring a nucleic acid expression level of a plurality of classifier biomarkers in a sample obtained from a subject suffering from or suspected of suffering from PAAD, wherein the plurality of classifier biomarkers is selected from Table 4, wherein the nucleic acid expression level of the plurality of classifier biomarkers indicates the subtype of PAAD as being classical or basal and administering a therapeutic intervention based on the subtype of PAAD, wherein the therapeutic intervention is selected from: (i) agents listed for the classical subtype in Table 7, 5-flourouracil and platinum-based therapy if the subject is classified as having the classical subtype of PAAD, or (ii) agents listed for the basal subtype in Table 7, cisplatin- or oxaliplatin-based therapies and gemcitabine if the subject is classified as having the basal subtype of PAAD.

63. The method of claim 62, wherein the determining step further comprises comparing the nucleic acid expression levels of the plurality of classifier biomarkers from Table 4 to the nucleic acid expression levels of the plurality of classifier biomarkers from Table 4 in at least one sample training set(s), wherein the at least one sample training set comprises nucleic acid expression levelAttorney Docket No. GNCN-026 / 01WO 320289-2157 data of the plurality of classifier biomarkers from Table 4 from a reference PAAD classical sample, nucleic acid expression level data of the plurality of classifier biomarkers from Table 4 from a reference PAAD basal sample or a combination thereof; and classifying the sample obtained from the subject as PAAD classical or PAAD basal based on the results of the comparing step.

64. The method of claim 63, wherein the comparing step comprises applying a statistical algorithm which comprises determining a correlation between the expression data obtained from the sample obtained from the subject and the expression data from the at least one training set(s); and classifying the sample obtained from the subject as PAAD classical or PAAD basal based on the results of the statistical algorithm.

65. The method of claim 62, wherein the plurality of classifiers selected from Table 4 comprises at least 2 classifiers, at least 4 classifiers, at least 6 classifiers, at least 8 classifiers, at least 10 classifier biomarkers, at least 12 classifiers, at least 14 classifiers, at least 16 classifiers, at least 18 classifiers, at least 20 classifiers, at least 22 classifiers, at least 24 classifiers, at least 28 classifiers, at least 30 classifiers, at least 32 classifiers, at least 34 classifiers, at least 36 classifiers, at least 38 classifiers, at least 40 classifiers, at least 42 classifiers, at least 44 classifiers, at least 46 classifiers, at least 48 classifiers, at least 50 classifiers, at least 52 classifiers, at least 54 classifiers, at least 56 classifiers, at least 58 classifiers, at least 60 classifiers, at least 62 classifiers, at least 64 classifiers, at least 66 classifiers, at least 68 classifiers, at least 70 classifiers, at least 72 classifiers, at least 74 classifiers, at least 76 classifiers, at least 78 classifiers or at least 80 classifiers from Table 4.

66. The method of claim 62, wherein the plurality of classifiers selected from Table 4 comprises at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% of the classifiers from Table 4.

67. The method of claim 62, wherein the plurality of classifiers selected from Table 4 consists of all the classifiers of Table 4.Attorney Docket No. GNCN-026 / 01WO 320289-2157 68. The method of claim 62, wherein the sample is a formalin-fixed, paraffin-embedded (FFPE) tissue sample from the pancreas of the subject, fresh or a frozen tissue sample from the pancreas of the subject, an exosome, wash fluids, cell pellets, or a bodily fluid obtained from the subject.

69. The method of claim 68, wherein the bodily fluid is whole blood, plasma, serum, an exosome, wash fluids, urine, saliva, cerebrospinal fluid or sputum.

70. A method of assaying a sample obtained from a subject suffering from bladder cancer (BLCA), the method comprising measuring the expression level of a plurality of classifiers selected from Table 5 or Table 6 using a sequencing assay.

71. The method of claim 70, wherein the sequencing assay is a DNA sequencing assay or RNA sequencing assay.

72. The method of claim 70, wherein the plurality of biomarkers selected from Table 5 comprises at least 50 biomarkers, at least 100 biomarkers, at least 150 biomarkers, at least 200 biomarkers, at least 250 classifier biomarkers, at least 300 biomarkers, at least 350 biomarkers, at least 400 biomarkers, at least 450 biomarkers, at least 500 biomarkers, at least 550 biomarkers, at least 600 biomarkers, at least 650 biomarkers, at least 700 biomarkers, at least 750 biomarkers, at least 800 biomarkers, at least 850 biomarkers, at least 900 biomarkers or at least 950 biomarkers from Table 5.

73. The method of claim 70, wherein the plurality of classifiers selected from Table 5 comprises at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% of the classifiers from Table 5.

74. The method of claim 70, wherein the plurality of classifiers selected from Table 5 consists of all the classifiers of Table 5.Attorney Docket No. GNCN-026 / 01WO 320289-2157 75. The method of claim 70, wherein the plurality of classifiers selected from Table 6 comprises at least 2 classifiers, at least 4 classifiers, at least 6 classifiers, at least 8 classifiers, at least 10 classifier biomarkers, at least 12 classifiers, at least 14 classifiers, at least 16 classifiers, at least 18 classifiers, at least 20 classifiers, at least 22 classifiers, at least 24 classifiers, at least 28 classifiers, at least 30 classifiers, at least 32 classifiers, at least 34 classifiers, at least 36 classifiers, at least 38 classifiers, at least 40 classifiers, at least 42 classifiers, at least 44 classifiers, at least 46 classifiers, at least 48 classifiers or at least 50 classifiers from Table 6.

76. The method of claim 70, wherein the plurality of classifiers selected from Table 6 comprises at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% of the classifiers from Table 6.

77. The method of claim 70 , wherein the plurality of classifiers selected from Table 6 consists of all the classifiers of Table 6.

78. The method of claim 70, wherein the sample is a formalin-fixed, paraffin-embedded (FFPE) tissue sample from the bladder of the subject, fresh or a frozen tissue sample from the bladder of the subject, an exosome, wash fluids, cell pellets, or a bodily fluid obtained from the subject.

79. The method of claim 78, wherein the bodily fluid is whole blood, plasma, serum, an exosome, wash fluids, urine, saliva, cerebrospinal fluid or sputum.

80. The method of claim 78, wherein the bodily fluid is urine.

81. A method of treating BLCA in a subject, the method comprising: measuring a nucleic acid expression level of a plurality of classifiers in a sample obtained from a subject suffering from or suspected of suffering from BLCA, wherein the plurality of classifiers is selected from Table 6, wherein the measured nucleic acid expression levels of the plurality of classifiers provide an FGFR3 activation signature for the sample; and administering an FGFR inhibitor based onAttorney Docket No. GNCN-026 / 01WO 320289-2157 presence of a positive FGFR3 activation signature, wherein the positive FGFR3 activation signature is indicative of presence of one or more FGFR3 mutations.

82. The method of claim 81, wherein the determining step further comprises comparing the nucleic acid expression levels of the plurality of classifiers from Table 6 to the nucleic acid expression levels of the plurality of classifier from Table 6 in at least one sample training set(s), wherein the at least one sample training set is from a reference FGFR3 mutation-containing BLCA sample, or is from a reference FGFR3 mutation-free BLCA sample; and classifying the tumor sample as having a positive FGFR3 activation signature based on the results of the comparing step.

83. The method of claim 82, wherein the at least one training set is from a reference FGFR3 mutation-containing cancer sample and the sample is classified as possessing the positive FGFR3 activation signature if the nucleic acid expression levels of the plurality of classifiers of Table 6 correlate with the nucleic acid expression levels of the plurality of classifiers of Table 6 from the reference FGFR3 mutation-containing cancer sample.

84. The method of claim 82, wherein the at least one training set is from a reference FGFR3 mutation-containing cancer sample and from a reference FGFR3 mutation-free cancer sample and the sample is classified as possessing the positive FGFR3 activation signature if the expression levels of the plurality of classifiers of Table 6 correlate with the expression levels of the plurality of classifiers of Table 6 from the reference FGFR3 mutation-containing cancer sample.

85. The method of claim 82, wherein the comparing step comprises applying a statistical algorithm which comprises determining a correlation between the expression data obtained from the sample obtained from the subject and the expression data from the at least one training set(s); and classifying the sample obtained from the subject as possessing a positive FGFR3 activation signature on the results of the statistical algorithm.

86. The method of claim 81, wherein the plurality of classifiers selected from Table 6 comprises at least 2 classifiers, at least 4 classifiers, at least 6 classifiers, at least 8 classifiers, at least 10 classifier biomarkers, at least 12 classifiers, at least 14 classifiers, at least 16 classifiers, at least 18Attorney Docket No. GNCN-026 / 01WO 320289-2157 classifiers, at least 20 classifiers, at least 22 classifiers, at least 24 classifiers, at least 28 classifiers, at least 30 classifiers, at least 32 classifiers, at least 34 classifiers, at least 36 classifiers, at least 38 classifiers, at least 40 classifiers, at least 42 classifiers, at least 44 classifiers, at least 46 classifiers, at least 48 classifiers or at least 50 classifiers from Table 6.

87. The method of claim 81, wherein the plurality of classifiers selected from Table 6 comprises at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% of the classifiers from Table 6.

88. The method of claim 81, wherein the plurality of classifiers selected from Table 6 consists of all the classifiers of Table 6.

89. The method of claim 81, wherein the sample is a formalin-fixed, paraffin-embedded (FFPE) tissue sample from the bladder of the subject, fresh or a frozen tissue sample from the bladder of the subject, an exosome, wash fluids, cell pellets, or a bodily fluid obtained from the subject.

90. The method of claim 89, wherein the bodily fluid is whole blood, plasma, serum, an exosome, wash fluids, urine, saliva, cerebrospinal fluid or sputum.

91. The method of claim 90, wherein the bodily fluid is urine.

92. The method of claim 81, wherein the FGFR inhibitor shows inhibitory activity toward fibroblast growth factor receptor-3 (FGFR3).

93. The method of any claim 92, wherein the FGFR inhibitor is a tyrosine kinase inhibitor.

94. The method of claim 93, wherein the FGFR inhibitor is a selective tyrosine kinase inhibitor.

95. The method of claim 93, wherein the FGFR inhibitor is a non-selective tyrosine kinase inhibitor.Attorney Docket No. GNCN-026 / 01WO 320289-2157 96. The method of claim 92, wherein the FGFR inhibitor is selected from the group consisting of erdafitinib (JNJ 42756493), infigratinib (BGJ1398), Rogaritinib (BAY 1163877), AZD4547, Pemigatinib (INCB54828), TAS-120, LY2874455, DEBIO 1347, PD173074, BLU9931, pazopanib, brivanib, ponatinib (AP24534), regorafenib (BAY 73-4506), lenvatinib (E7080), dovitinib (TKI258), lucitanib (E3810), nintedanib (BIBF 1120), Foretinib, and any combination thereof.

97. The method of claim 92, wherein the FGFR inhibitor is nintedanib (BIBF 1120).

98. The method of claim 92, wherein the FGFR inhibitor is an antibody or antibody-conjugate.

99. The method of claim 98, wherein the FGFR inhibitor is B-701 or MFGR1877S.

100. The method of claim 98, wherein the FGFR inhibitor is LY3076226.

101. A computer-implemented method for identifying regulatory regions for one or more genes that correlate with gene expression of the one or more genes, the method comprising: (a) inputting into a computer system, a first dataset comprising RNA sequencing data and a second dataset comprising chromatin accessibility data; (b) mapping by the computer system, sequence reads obtained from the first dataset to a reference genome sequence; (c) mapping by the computer system, sequence reads obtained from the second dataset to the reference genome sequence; (d) annotating by the computer system, peaks of the sequence reads over random or low- level noise in the chromatin accessibility data from the second dataset wherein each peak represents a putative chromatin accessible region; (e) linking by the computer system, the annotated peaks from (d) to a nearest gene found from the mapped sequence reads to the first dataset using machine learning, thereby generating a plurality of genes and chromatin accessibility regions associated with each gene from the plurality of genes;Attorney Docket No. GNCN-026 / 01WO 320289-2157 (f) measuring by the computer system, a correlation between a level of expression of each gene in the plurality of genes to each chromatin accessibility region associated with each gene in the plurality of genes, wherein the level of expression of each gene in the plurality of genes is evidenced by the number of sequence reads from the first dataset that mapped to each gene in the plurality of genes; (g) ranking by the computer system, the correlations obtained in (f); (h) selecting by the computer system, chromatin accessibility regions whose correlations coefficients were greater than a desired threshold, thereby generating a list of chromatin accessibility regions that are highly correlated with RNA expression levels; and (i) selecting by the computer system, regulatory regions from the chromatin accessibility regions selected in (h), thereby identifying regulatory regions for one or more genes that highly correlate with RNA expression of the one or more genes.

102. The computer-implemented method of claim 101, further comprising (j) designing by the computer, a set of probes, wherein each probe in the set comprises sequence complementary to one of the regulatory regions selected in (i).

103. The computer-implemented method of claim 101, further comprising (j) designing by the computer, a set of primer pairs, wherein each primer pair in the set comprises sequence complementary to one of the regulatory regions selected in (i).

104. The computer-implemented method of claim 101, wherein the correlation is a Pearson correlation or a Spearman correlation.

105. The computer-implemented method of claim 101, wherein the correlation is a Pearson and a Spearman correlation.

106. The computer-implemented method of claim 101, wherein the correlation is a Spearman correlation.Attorney Docket No. GNCN-026 / 01WO 320289-2157 107. The computer-implemented method of claim 101, wherein the desired threshold can be a correlation coefficient greater than 0.6, 0.7, 0.8 or 0.

9.

108. The computer-implemented method of claim 101, the desired threshold is a correlation coefficient greater than 0.

8.

109. The computer-implemented method of claim 101, wherein the chromatin accessibility data from the second dataset is obtained from a chromatin accessibility sequencing (CA-seq) assay.

110. The computer-implemented method of claim 109, wherein the CA-seq assay is selected from the group consisting of micrococcal nuclease digestion with deep sequencing (MNase-seq), DNase I hypersensitive sites sequencing (DNase-seq), Formaldehyde-Assisted Isolation of Regulatory Elements sequencing (FAIRE-seq) and Assay for Transposase-Accessible Chromatin using sequencing (ATAC-seq).

111. The computer-implemented method of claim 109, wherein the CA-seq assay is ATAC-seq and the chromatin accessibility regions are ATAC-seq regions.

112. The computer-implemented method of claim 101, wherein the regulatory regions are selected from the group consisting of the 5’UTR, the 3’ UTR, a promoter, a proximal enhancer, a distal enhancer and any combination thereof.

113. The computer-implemented method of claim 101, wherein the first dataset and the second dataset are each be obtained from tissue samples or tissue biopsies.

114. The computer-implemented method of claim 113, wherein the tissue samples are tissue matched samples between the first and second datasets.

115. The computer-implemented method of claim 113, wherein the first dataset and the second dataset are obtained from a population of subjects that each have matched chromatin accessibility sequencing data and RNA expression sequencing data.Attorney Docket No. GNCN-026 / 01WO 320289-2157 116. The computer-implemented method of claim 101, wherein the first and second datasets are obtained from one or more subjects that suffer from or are suspected of suffering from a desired feature of a cancer.

117. The computer-implemented method of claim 116, wherein the desired feature is a type of cancer, a subtype of cancer or the presence or absence of microsatellite instability.

118. The computer-implemented method of claim 101, wherein, prior to step (a), the second dataset comprising chromatin accessibility data is generated by inputting into the computer system, chromatin accessibility sequencing data for a first plurality of subjects suffering from one type of cancer or subtype thereof and chromatin accessibility sequencing (e.g., ATAC-seq) data for a second plurality of subjects suffering from a second type of cancer or subtype thereof; and filtering out by the computer system, chromatin accessibility regions present in the chromatin accessibility sequencing data for the first and second plurality of subjects, thereby generating a second dataset comprising chromatin accessibility data specific to the first and second type of cancer or subtype thereof.

119. The computer-implemented method of claim 101, wherein the machine learning comprises a tool from Algorithms for Calculating Microarray Enrichment (ACME).

120. The computer-implemented method of claim 119, wherein the tool from ACME is a findClosestGene function.

121. The computer-implemented method of claim 101, wherein the reference genome is a reference human genome.

122. The computer-implemented method of claim 121, wherein the reference human genome is a GRCh38 (hg38) assembly.

123. A method for sequencing target regulatory regions in cfDNA, the method comprising: (a) obtaining a sample comprising cfDNA from a subject suffering from or suspected of suffering from cancer; andAttorney Docket No. GNCN-026 / 01WO 320289-2157 (b) performing a target enrichment sequencing assay on the cfDNA sample using a panel comprising a plurality of probes, wherein each probe in the plurality comprises sequence complementary to regulatory regions for one or more genes that correlate with gene expression of the one or more genes, wherein the panel is generated using the method of claim 102, thereby generating a gene activity matrix for the cfDNA for the subject.

124. The method of claim 123, wherein the target enrichment sequencing assay comprises a hybrid-capture based enrichment step followed by a sequencing assay.

125. The method of claim 124, wherein the sequencing assay comprises next-generation sequencing (NGS).

126. The method of claim 125, wherein the sample comprising cfDNA is a bodily fluid sample or liquid biopsy sample.

127. A computer-implemented method comprising: (a) inputting into a computer system, paired- end DNA sequencing data obtained from a target enrichment sequencing assay performed on a sample comprising cfDNA obtained from a subject, wherein the target enrichment sequencing assay is performed using a panel comprising a plurality of probes generated using the method of claim 102; (b) calculating by the computer system, fragment sequence lengths for sequences from the pair-end DNA sequencing data that comprise overlapping sequence for a first regulatory region from (a)); (c) calculating a probability of finding each fragment size by the computer system, using counts of each unique fragment sequence fragment length for the first regulatory region; (d) calculating Shannon entropy using the probabilities from (c), thereby generating a Shannon entropy value for the first regulatory region; and (e) repeating (b)-(d) for each additional regulatory region from the target enrichment sequencing assay, thereby generating a Shannon entropy matrix for each regulatory region obtained from the target enrichment sequencing assay for the sample comprising cfDNA obtained from the subject.

128. The computer-implemented method of claim 127, wherein the target enrichment sequencing assay is a hybrid capture enrichment sequencing assay, wherein each probe in the panel binds toAttorney Docket No. GNCN-026 / 01WO 320289-2157 and pulls down a regulatory region associated with a target gene from the plurality of target genes prior to being subjected to a sequencing assay.

129. The computer-implemented method of claim 127 or 128, wherein the subject suffers from a desired feature of cancer.

130. The computer-implemented method of claim 129, wherein the Shannon entropy matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for the subject represents a sample positive for the desired feature of cancer.

131. The computer-implemented method of claim 129 or 130, wherein (a)-(e) are repeated on samples comprising cfDNA obtained from a plurality of subjects that each suffer from the desired feature of cancer.

132. The computer-implemented method of claim 131, wherein the Shannon entropy matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for each of the plurality of subjects that suffer from the desired feature of cancer are combined by computer system.

133. The computer-implemented method of claim 127 or 128, wherein the subject does not suffer from the desired feature of cancer.

134. The computer-implemented method of claim 133, wherein the Shannon entropy matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for the subject represents a sample negative for the desired feature of cancer.

135. The computer-implemented method of claim 133 or 134, wherein (a)-(e) are repeated on samples comprising cfDNA obtained from a plurality of subjects that each do not suffer from the desired feature of cancer.Attorney Docket No. GNCN-026 / 01WO 320289-2157 136. The computer-implemented method of claim 135, wherein the Shannon entropy matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for each of the plurality of subjects that do not suffer from the desired feature of cancer are combined by the computer system.

137. The computer-implemented method of claim 127, wherein the desired feature of the cancer is selected from the group consisting of a type of cancer, a subtype of the cancer of interest, microsatellite instability, mismatch repair proficiency, minimal residual disease, prognostic risk, response to a therapeutic agent, development of resistance to a therapeutic agent cancer.

138. A computer-implemented method for generating a classifier for determining a desired feature of cancer, the method comprising: (a) receiving in a computer system, a first set of Shannon entropy indices obtained from cfDNA comprising samples obtained from a first population of subjects suffering from the desired feature of the cancer and a second set of Shannon entropy indices obtained from cfDNA comprising samples obtained from a second population of subjects that do not suffer from the desired feature of the cancer; (b) selecting by the computer system, regulatory regions for model training and validation using the first set and second set of Shannon entropy matrices; and (c) training and validating by the computer system a machine learning algorithm for determining the presence or absence of the desired feature of the cancer in a cfDNA comprising sample, thereby generating a classifier for determining a desired feature of a cancer.

139. The computer-implemented method of claim 138, further comprises inputting target enrichment sequencing data obtained for cfDNA comprising sample obtained from a test subject suspected of suffering from the desired feature of the cancer into the classifier generated by (a)- (c), thereby determining the presence or absence of the desired feature of the cancer in the test subject.

140. The computer-implemented method of claim 138, wherein the cfDNA comprising samples are a bodily fluid samples or liquid biopsy samples.Attorney Docket No. GNCN-026 / 01WO 320289-2157 141. The computer-implemented method of claim 138, wherein the desired feature is selected from the group consisting of a type of cancer, a subtype of the cancer of interest, microsatellite instability, mismatch repair proficiency, minimal residual disease, prognostic risk, response to a therapeutic agent, development of resistance to a therapeutic agent cancer.

142. The computer-implemented method of claim 138, wherein the Shannon entropy matrices for the first and the second populations are generated using the method of any one of claims 127-137.

143. A computer-implemented method comprising: (a) inputting into a computer system, paired- end DNA sequencing data obtained from a target enrichment sequencing assay performed on a sample comprising cfDNA obtained from a subject, wherein the target enrichment sequencing assay is performed using a panel comprising a plurality of probes generated using the method of claim 102; and (b) calculating by the computer system one or more sets of fragment characteristics for sequences from the pair-end DNA sequencing data, thereby generating a fragment characteristic matrix for each regulatory region obtained from the target enrichment sequencing assay for the sample comprising cfDNA obtained from the subject.

144. The computer-implemented method of claim 143, wherein the target enrichment sequencing assay is a hybrid capture enrichment sequencing assay, wherein each probe in the panel binds to and pulls down a regulatory region associated with a target gene from the plurality of target genes prior to being subjected to a sequencing assay.

145. The computer-implemented method of claim 143 or 144, wherein the subject suffers from a desired feature of cancer.

146. The computer-implemented method of claim 143, wherein the fragment characteristic matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for the subject represents a sample positive for the desired feature of cancer.Attorney Docket No. GNCN-026 / 01WO 320289-2157 147. The computer-implemented method of claim 145 or 146, wherein (a)-(e) are repeated on samples comprising cfDNA obtained from a plurality of subjects that each suffer from the desired feature of cancer.

148. The computer-implemented method of claim 147, wherein the fragment characteristic matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for each of the plurality of subjects that suffer from the desired feature of cancer are combined by computer system.

149. The computer-implemented method of claim 143 or 144, wherein the subject does not suffer from the desired feature of cancer.

150. The computer-implemented method of claim 149, wherein the fragment characteristic matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for the subject represents a sample negative for the desired feature of cancer.

151. The computer-implemented method of claim 149 or 150, wherein (a)-(e) are repeated on samples comprising cfDNA obtained from a plurality of subjects that each do not suffer from the desired feature of cancer.

152. The computer-implemented method of claim 151, wherein the fragment characteristic matrix for each regulatory region obtained from the target enrichment sequencing assay from cfDNA for each of the plurality of subjects that do not suffer from the desired feature of cancer are combined by the computer system.

153. The computer-implemented method of claim 143, wherein the desired feature of the cancer is selected from the group consisting of a type of cancer, a subtype of the cancer of interest, microsatellite instability, mismatch repair proficiency, minimal residual disease, prognostic risk, response to a therapeutic agent, development of resistance to a therapeutic agent cancer.Attorney Docket No. GNCN-026 / 01WO 320289-2157 154. The computer-implemented method of claim 143, wherein the fragment characteristics are selected from the group consisting of fragment sequence length, end motif frequency, end motif sequence, jagged end length, fragment diversity, FFT amplitude magnitude and any combination thereof.

155. A computer-implemented method for generating a classifier for determining a desired feature of cancer, the method comprising: (a) receiving in a computer system, a first set of fragment characteristic indices obtained from cfDNA comprising samples obtained from a first population of subjects suffering from the desired feature of the cancer and a second set of fragment characteristic indices obtained from cfDNA comprising samples obtained from a second population of subjects that do not suffer from the desired feature of the cancer; (b) selecting by the computer system, regulatory regions for model training and validation using the first set and second set of fragment characteristic matrices; and (c) training and validating by the computer system a machine learning algorithm for determining the presence or absence of the desired feature of the cancer in a cfDNA comprising sample, thereby generating a classifier for determining a desired feature of a cancer.

156. The computer-implemented method of claim 155, further comprising inputting target enrichment sequencing data obtained for cfDNA comprising sample obtained from a test subject suspected of suffering from the desired feature of the cancer into the classifier generated by (a)- (c), thereby determining the presence or absence of the desired feature of the cancer in the test subject.

157. The computer-implemented method of claim 155 or 156, wherein the cfDNA comprising samples are a bodily fluid samples or liquid biopsy samples.

158. The computer-implemented method of any one of claims 155-157, wherein the desired feature is selected from the group consisting of a type of cancer, a subtype of the cancer of interest, microsatellite instability, mismatch repair proficiency, minimal residual disease, prognostic risk, response to a therapeutic agent, development of resistance to a therapeutic agent cancer.Attorney Docket No. GNCN-026 / 01WO 320289-2157 159. The computer-implemented method of any one of claims 155-158, wherein the fragment characteristic matrices for the first and the second populations are generated using the method of claim 143.

Citation Information

Patent Citations

  • Biological status determination using cell-free nucleic acids

    US20200115762A1

  • Cell-free DNA sequence data analysis method to examine nucleosome protection and chromatin accessibility

    WO2022217096A2

  • Systems and methods for gene expression and tissue of origin inference from cell-free DNA

    WO2023091517A2