Method for using circulating tumor DNA for cancer subtyping

The FFT-based cancer subtyping method using ctDNA addresses noise issues in existing technologies by generating a precise cancer classifier, enhancing cancer subtype identification and treatment selection accuracy.

WO2025144650A1PCT designated stage expired Publication Date: 2025-07-03GENECENTRIC THERAPEUTICS INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/060700
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-28
Filing Date
2024-12-18
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Current methods for cancer subtyping using circulating tumor DNA (ctDNA) face challenges such as high noise levels and reduced relevance in gene set enrichment analyses, failing to fully explain treatment failure or identify therapeutic targets, and genomic alterations may not always accurately reflect treatment responses.

Method used

A method utilizing Fast Fourier Transform (FFT) magnitude as a gene expression surrogate marker, combined with a Classifying arrays to Nearest Centroid (CLaNC) model, to generate a cancer type classifier from cell-free DNA (cfDNA) data, adjusting for GC bias and calculating smoothed mean read counts in defined genomic regions.

Benefits of technology

The method provides accurate cancer subtyping by identifying key genes with high mean and variance, improving signal-to-noise ratio and enabling precise cancer subtype classification and treatment selection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024060700_03072025_PF_FP_ABST
    Figure US2024060700_03072025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure discloses methods for using cell-free DNA (cfDNA) or circulating tumor DNA (ctDNA) from whole genome or whole exome sequencing data from tumor samples obtained from subjects suffering from cancer in combination with fast Fourier transformations on said sequencing data to determine cancer type or cancer subtype. The present disclosure also discloses methods for selecting a treatment and / or treating a subject based on the subject's determined cancer type or cancer subtype.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No. GNCN-025 / 01WO 320289-2155 METHOD FOR USING CIRCULATING TUMOR DNA FOR CANCER SUBTYPING CROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority from U.S. Provisional Application No. 63 / 615,390 filed December 28, 2023, which is incorporated by reference herein in its entirety for all purposes. FIELD

[0002] The present disclosure is directed to methods of determining a subtype of cancer that a subject is suffering from or suspected of suffering from using cell-free DNA (cfDNA) found in a biological sample (e.g., liquid biopsy) obtained from the subject. In some embodiments, the methods presented herein utilize a mathematical transformation (e.g., Fast Fourier Transformation (FFT)) of whole genome or whole exome sequencing data obtained from cfDNA found in a biological sample obtained from a subject suffering from or suspected of suffering from cancer to determine the subject’s cancer subtype. In some embodiments, the methods pf the present disclosure can be prognostic and / or select the subject for specific treatments based on the subject’s cancer subtype. BACKGROUND

[0003] Accurate cancer diagnosis and subtype classification are critical for guiding clinical care and precision oncology. Moreover, tumor subtypes are often characterized by distinct transcriptional regulation, which can change during treatment resistance, leading to different clinical tumor phenotypes. Therefore, accurate subtype classification and identification of transcriptional patterns underlying emergent clinical phenotype during therapy has critical implications for studying mechanisms of resistance and informing treatment decisions.

[0004] In patients with cancer, cell-free DNA (cfDNA) can be released from tumor cells, called circulating tumor DNA (ctDNA). Recent studies have shown the analysis of ctDNA can address the challenges in tissue accessibility and have demonstrated great potential for clinical utility (Doebley, AL., Ko, M., Liao, H. et al. Author Correction: A framework for clinical cancer subtyping from nucleosome profiling of cell-free DNA. Nat Commun 14, 403 (2023). Much of the current research and clinical efforts have focused on the detection of genetic alterations in ctDNA and how said alterations from ctDNA can be used to help distinguish molecular subsetsof tumors (see Wyatt, A. W. et al. Concordance of circulating tumor DNA and matched metastatic tissue biopsy in prostate cancer. J. Natl Cancer Inst.110, 78–86 (2018) and Viswanathan, S. R. et al. Structural alterations driving castration-resistant prostate cancer revealed by linked-read genome sequencing. Cell 174, 433–447.e19 (2018). However, these genomic alterations, including somatic mutations, may not always fully explain treatment failure or identify therapeutic targets, exemplifying a major limitation of cancer precision medicine. Moreover, methods used in some of the studies utilizing ctDNA for cancer subtyping show high levels of noise as well as reduced relevant biology in gene set enrichment analyses (GSEA).

[0005] The methods provided herein address these issues and provide a cancer subtyping method that utilizes fast Fourier transformation (FFT) magnitude as a gene expression surrogate marker. SUMMARY

[0006] In one aspect, provided herein is a computer implemented method for generating a classifier for determining a cancer type of a subject suffering from a cancer or suspected of suffering from a cancer; the method comprising: (a) receiving in a computer system, sequence read data generated from nucleic acid extracted from a sample obtained from each subject from a plurality of subjects suffering from a cancer or suspected of suffering from a cancer, wherein the sample comprises cell-free DNA (cfDNA), wherein the sequence read data includes a plurality of fragment reads, wherein each fragment read has a fragment length and a GC content indicating a percentage of bases in the fragment read that are G or C; (b) determining by the computing system, GC bias values for each fragment read based on the fragment length and the GC content of the fragment read; (c) generating by the computing system, a genomic coverage distribution that is adjusted for GC bias using the sequence read data and the GC bias values; (d) calculating by the computing system mean sequence read counts for a sliding window across a defined window in each of a plurality of genomic regions of interest from the genomic coverage distribution to generate smoothed mean read counts, wherein each of the plurality of genomic regions of interest are cancer type informative sites; (e) performing a fast Fourier transform (FFT) on the smoothed mead read counts to generate an FFT magnitude matrix for each of the plurality of genomic regions of interest; (f) selecting genes from the FFT magnitude matrix for each of the plurality of genomic regions of interest that have a high mean and high variance; (g) inputting the selected genes into a Classifying arrays to Nearest Centroid (CLaNC) model on the computing system to determine aset of genes for the cancer type classifier; and (h) fitting an ordinary nearest centroid classifier using the set of genes determined in step (g) to generate the cancer type classifier. In some cases, the sliding window has a width of at least, at most or exactly 5, 10, 15, 20, 25, 30, 35, 40, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95 or 100 nucleotides across the defined window. In some cases, the sliding window has a width of 15 nucleotides across the defined window. In some cases, the defined window has a width of 2000 base pairs. In some cases, the cancer-type informative sites are sites that have differential expression between a first cancer type and a second cancer type. In some cases, the cancer type is a cancer subtype of a type of cancer. In some cases, the cancer is selected from the group consisting of kidney renal papillary cell carcinoma (KIRP); breast invasive carcinoma (BRCA); thyroid cancer (THCA); bladder urothelial carcinoma (BLCA); prostate adenocarcinoma (PRAD); kidney chromophobe (KICH); cervical squamous cell carcinoma and endocervical adenocarcinoma (CESC); kidney renal clear cell carcinoma (KIRC); liver hepatocellular carcinoma (LIHC); low grade glioma (LGG); sarcoma (SARC); lung adenocarcinoma (LUAD); colon adenocarcinoma (COAD); head and neck squamous cell carcinoma (HNSC); uterine corpus endometrial carcinoma (UCEC); glioblastoma multiforme (GBM); esophageal carcinoma (ESCA); stomach adenocarcinoma (STAD); ovarian serous cystadenocarcinoma (OV); rectum adenocarcinoma (READ); adrenocortical carcinoma (ACC); uveal melanoma (UVM); mesothelioma (MESO); pheochromocytoma and paraganglioma (PCPG); skin cutaneous melanoma (SKCM); uterine carcinsarcoma (UCS); lung squamous cell carcinoma (LUSC); testicular germ cell tumors (TGCT); cholangiocarcinoma (CHOL); pancreatic adenocarcinoma (PAAD); thymoma (THYM); or Lymphoid Neoplasm Diffuse Large B-cell Lymphoma (DLBC). In some cases, the sample is a formalin-fixed, paraffin-embedded (FFPE) tissue sample, a fresh or a frozen tissue sample, an exosome, wash fluids, cell pellets, or a bodily fluid obtained from the subject. In some cases, the bodily fluid is blood or fractions thereof, urine, saliva, or sputum. In some cases, the nucleic acid is DNA, RNA or cDNA. In some cases, the sequence read data is whole genome sequencing (WGS) data. In some cases, the sequence read data is whole exome sequencing (WES) data.

[0007] In another aspect, provided herein is a method for determining a cancer type in a subject suffering from or suspected of suffering from a cancer comprising inputting sequence read data generated from a sample obtained from the subject into the cancer type classifier provided herein.

[0008] In still another aspect, provided herein is a method for treating cancer in a subject suffering from or suspected of suffering from a cancer comprising: (a) determining a cancer type of the subject by inputting sequence read data generated from a sample obtained from the subject into the cancer type classifier provided herein; and (b) administering a therapeutic specific for the cancer type of the subject determined in step (a), thereby treating the cancer.

[0009] In still another aspect, provided herein is a method for selecting a cancer treatment for a subject suffering from or suspected of suffering from a cancer comprising: (a) determining a cancer type of the subject by inputting sequence read data generated from a sample obtained from the subject into the cancer type classifier provided herein; and (b) selecting a therapeutic specific the cancer type of the subject determined in step (a). BRIEF DESCRIPTION OF THE FIGURES

[0010] FIG.1 illustrates the histogram of a magnitude matrix.

[0011] FIG.2 illustrates cross-validation curves.

[0012] FIG. 3 illustrates the results from the training sets and testing sets using the classifier developed in Example 1.

[0013] FIG.4 illustrates heatmap using the classifier genes (126 of 184 symbols directly present) for clustering.

[0014] FIG. 5 illustrates the results from permutation tests to evaluate how well a random set of genes classifies the BRCA vs CRC cohorts.

[0015] FIG. 6 illustrates a comparison of the Sample-Sample Correlation of the FFT method provided herein (i.e., GC FFT Method) vs. Griffin FFT Method. The method provided herein has more range while the Griffin FFT method is close to zero.

[0016] FIG.7 illustrates a comparison of the top genes correlated with healthy vs. malignant status of the FFT method provided herein (i.e., GC FFT Method) vs. Griffin FFT Method. DETAILED DESCRIPTION Definitions

[0017] While the following terms are believed to be well understood by one of ordinary skill in the art, the following definitions are set forth to facilitate explanation of the presently disclosed subject matter.

[0018] As used herein, the term “a” or “an” can refer to one or more of that entity, i.e. can refer to a plural referents. As such, the terms “a” or “an”, “one or more” and “at least one” can be used interchangeably herein. In addition, reference to “an element” by the indefinite article “a” or “an” does not exclude the possibility that more than one of the elements is present, unless the context clearly requires that there is one and only one of the elements.

[0019] Unless the context requires otherwise, throughout the present specification and claims, the word “comprise” and variations thereof, such as, “comprises” and “comprising” are to be construed in an open, inclusive sense that is as “including, but not limited to”.

[0020] Reference throughout this specification to “one embodiment” or “an embodiment” means that a particular feature, structure or characteristic described in connection with the embodiment may be included in at least one embodiment of the present disclosure. Thus, the appearances of the phrases “in one embodiment” or “in an embodiment” in various places throughout this specification may not necessarily all referring to the same embodiment. It is appreciated that certain features of the disclosure, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the disclosure, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable sub-combination.

[0021] Throughout this disclosure, various aspects of the methods and compositions provided herein can be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the invention. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range.

[0022] Unless otherwise indicated, the methods and compositions provided herein can utilize conventional techniques and descriptions of organic chemistry, polymer technology, molecular biology (including recombinant techniques), cell biology, biochemistry, and immunology, which are within the skill of the art. Such conventional techniques include polymer array synthesis, hybridization, ligation, and detection of hybridization using a label. Specific illustrations ofsuitable techniques can be had by reference to the example herein below. However, other equivalent conventional procedures can, of course, also be used. Such conventional techniques and descriptions can be found in standard laboratory manuals such as Genome Analysis: A Laboratory Manual Series (Vols. I-IV), Using Antibodies: A Laboratory Manual, Cells: A Laboratory Manual, PCR Primer: A Laboratory Manual, and Molecular Cloning: A Laboratory Manual (all from Cold Spring Harbor Laboratory Press), Gait, "Oligonucleotide Synthesis: A Practical Approach" 1984, IRL Press, London, Nelson and Cox (2000), Lehninger et al., (2008) Principles of Biochemistry 5th Ed., W.H. Freeman Pub., New York, N.Y. and Berg et al. (2006) Biochemistry, 6.sup.th Ed., W.H. Freeman Pub., New York, N.Y., all of which are herein incorporated in their entirety by reference for all purposes.

[0023] Conventional software and systems may also be used in the methods and compositions provided herein. Computer software products of the invention typically include computer readable medium having computer-executable instructions for performing the logic steps of the method of the invention. Suitable computer readable medium include floppy disk, CD- ROM / DVD / DVD-ROM, hard-disk drive, flash memory, ROM / RAM, magnetic tapes, etc. The computer-executable instructions may be written in a suitable computer language or combination of several languages. Basic computational biology methods are described in, for example, Setubal and Meidanis et al., Introduction to Computational Biology Methods (PWS Publishing Company, Boston, 1997); Salzberg, Searles, Kasif, (Ed.), Computational Methods in Molecular Biology, (Elsevier, Amsterdam, 1998); Rashidi and Buehler, Bioinformatics Basics: Application in Biological Science and Medicine (CRC Press, London, 2000) and Ouelette and Bzevanis Bioinformatics: A Practical Guide for Analysis of Gene and Proteins (Wiley & Sons, Inc., 2.sup.nd ed., 2001). See U.S. Pat. No.6,420,108.

[0024] The methods and compositions provided herein may also make use of various computer program products and software for a variety of purposes, such as probe design, management of data, analysis, and instrument operation. See, U.S. Pat. Nos. 5,593,839, 5,795,716, 5,733,729, 5,974,164, 6,066,454, 6,090,555, 6,185,561, 6,188,783, 6,223,127, 6,229,911 and 6,308,170. Computer methods related to genotyping using high density microarray analysis may also be used in the present methods, see, for example, US Patent Pub. Nos. 20050250151, 20050244883, 20050108197, 20050079536 and 20050042654.

[0025] Additionally, the present disclosure may have preferred embodiments that include methods for providing genetic information over networks such as the Internet as shown in U.S. Patent Pub. Nos. 20030097222, 20020183936, 20030100995, 20030120432, 20040002818, 20040126840, and 20040049354.

[0026] As used herein, the terms “individual,” “patient,” and “subject” can refer to any single animal, more preferably a mammal (including such non-human animals as, for example, dogs, cats, horses, rabbits, zoo animals, cows, pigs, sheep, and non-human primates) for which treatment is desired. In particular embodiments, the individual or patient herein is a human. Overview

[0027] Provided herein are methods, kits and compositions for determining a cancer type or cancer subtype from nucleic acid expression data measured in a sample obtained from a subject suffering from cancer or suspected of suffering from cancer. The nucleic acid can be DNA, RNA or cDNA. The nucleic acid can be extracted and / or manipulated by any methods known in the art and / or provided herein. The sample can be any sample known in the art and / or provided herein. The cancer can be any cancer known in the art and / or provided herein. The methods provided here entail running the Griffin pipeline on sequencing data (e.g., whole genome sequencing (WGS) or whole exome sequencing (WES) data) obtained from nucleic acid (e.g., cfDNA or ctDNA) extracted from the sample obtained from the subject as outlined in Doebley, AL., Ko, M., Liao, H. et al. Author Correction: A framework for clinical cancer subtyping from nucleosome profiling of cell-free DNA. Nat Commun 14, 403 (2023). Following the generation of GC corrected coverage data around gene regions of interest, the method entails calculating the mean read counts for a sliding window in each region, applying fast Fourier transform (FFT) to the mean data and extracting the highest magnitude.

[0028] The FFT step can entail taking the 132 count values in a 2KB region of interest (-990 ~ +975, 15bp interval); performing fast Fourier transform with fft() in R; and extracting the first element (max magnitude) from the FFT output using Mod(). The amplitude can be encoded as the magnitude of the complex number. The desired signal can be the one with smallest frequency, and highest magnitude (see FIG. 1).

[0029] The magnitude can then be utilized in a traditional nearest centroid classifier (see Dabney, Alan R. "ClaNC: point-and-click software for classifying microarrays to nearestcentroids." Bioinformatics22, no. 1 (2005): 122-123) to determine a set of genes for use in the classifier. Generation of the classifier using training data sets and CLaNC and use of the ordinary nearest centroid classifier to generate a fitted classifier can be as described in the Examples provided herein and / or Dabney A.R.. Classification of microarrays to nearest centroids, Bioinformatics, 2005, vol. 21 (pg. 4148-4154) or Parker JS, Mullins M, Cheang MC, Leung S, Voduc D, Vickery T, Davies S, Fauron C, He X, Hu Z, Quackenbush JF, Stijleman IJ, Palazzo J, Marron JS, Nobel AB, Mardis E, Nielsen TO, Ellis MJ, Perou CM, Bernard PS. Supervised risk predictor of breast cancer based on intrinsic subtypes. J Clin Oncol.2009 Mar 10;27(8):1160-7.

[0030] In one embodiment, the sample used herein is obtained from an individual, and comprises formalin-fixed paraffin-embedded (FFPE) tissue. However, other tissue and sample types are amenable for use herein. In one embodiment, the other tissue and sample types can be fresh frozen tissue, wash fluids, or cell pellets, or the like. In one embodiment, the sample can be a bodily fluid or a liquid biopsy obtained from the individual. The bodily fluid can be blood or fractions thereof (e.g., serum, plasma), urine, sputum, saliva or cerebrospinal fluid (CSF). A biomarker nucleic acid as provided herein can be extracted from a cell or can be cell free or extracted from an extracellular vesicular entity such as an exosome.

[0031] Methods are known in the art for the isolation of RNA from FFPE tissue. In one embodiment, total RNA can be isolated from FFPE tissues as described by Bibikova et al. (2004) American Journal of Pathology 165:1799-1807, herein incorporated by reference. Likewise, the High Pure RNA Paraffin Kit (Roche) can be used. Paraffin is removed by xylene extraction followed by ethanol wash. RNA can be isolated from sectioned tissue blocks using the MasterPure Purification kit (Epicenter, Madison, Wis.); a DNase I treatment step is included. RNA can be extracted from frozen samples using Trizol reagent according to the supplier's instructions (Invitrogen Life Technologies, Carlsbad, Calif.). Samples with measurable residual genomic DNA can be resubjected to DNaseI treatment and assayed for DNA contamination. All purification, DNase treatment, and other steps can be performed according to the manufacturer's protocol. After total RNA isolation, samples can be stored at -80 ºC until use.

[0032] General methods for mRNA extraction are well known in the art and are disclosed in standard textbooks of molecular biology, including Ausubel et al., ed., Current Protocols in Molecular Biology, John Wiley & Sons, New York 1987-1999. Methods for RNA extraction from paraffin embedded tissues are disclosed, for example, in Rupp and Locker (Lab Invest. 56:A67,1987) and De Andres et al. (Biotechniques 18:42-44, 1995). In particular, RNA isolation can be performed using a purification kit, a buffer set and protease from commercial manufacturers, such as Qiagen (Valencia, Calif.), according to the manufacturer's instructions. For example, total RNA from cells in culture can be isolated using Qiagen RNeasy mini-columns. Other commercially available RNA isolation kits include MasterPureTM. Complete DNA and RNA Purification Kit (Epicentre, Madison, Wis.) and Paraffin Block RNA Isolation Kit (Ambion, Austin, Tex.). Total RNA from tissue samples can be isolated, for example, using RNA Stat-60 (Tel-Test, Friendswood, Tex.). RNA prepared from a tumor can be isolated, for example, by cesium chloride density gradient centrifugation. Additionally, large numbers of tissue samples can readily be processed using techniques well known to those of skill in the art, such as, for example, the single- step RNA isolation process of Chomczynski (U.S. Pat. No. 4,843,155, incorporated by reference in its entirety for all purposes).

[0033] In one embodiment, a sample comprises cells harvested from a tissue sample. Cells can be harvested from a biological sample using standard techniques known in the art. For example, in one embodiment, cells are harvested by centrifuging a cell sample and resuspending the pelleted cells. The cells can be resuspended in a buffered solution such as phosphate-buffered saline (PBS). After centrifuging the cell suspension to obtain a cell pellet, the cells can be lysed to extract nucleic acid, e.g, messenger RNA. All samples obtained from a subject, including those subjected to any sort of further processing, are considered to be obtained from the subject.

[0034] The sample, in one embodiment, is further processed before the detection of the biomarker levels of the combination of biomarkers set forth herein. For example, mRNA in a cell or tissue sample can be separated from other components of the sample. The sample can be concentrated and / or purified to isolate mRNA in its non-natural state, as the mRNA is not in its natural environment. For example, studies have indicated that the higher order structure of mRNA in vivo differs from the in vitro structure of the same sequence (see, e.g., Rouskin et al. (2014). Nature 505, pp.701-705, incorporated herein in its entirety for all purposes).

[0035] mRNA from the sample in one embodiment, is hybridized to a synthetic DNA probe, which in some embodiments, includes a detection moiety (e.g., detectable label, capture sequence, barcode reporting sequence). Accordingly, in these embodiments, a non-natural mRNA-cDNA complex is ultimately made and used for detection of the biomarker. In another embodiment, mRNA from the sample is directly labeled with a detectable label, e.g., a fluorophore. In a furtherembodiment, the non-natural labeled-mRNA molecule is hybridized to a cDNA probe and the complex is detected.

[0036] In one embodiment, once the mRNA is obtained from a sample, it is converted to complementary DNA (cDNA) prior to the hybridization reaction or is used in a hybridization reaction together with one or more cDNA probes. cDNA does not exist in vivo and therefore is a non-natural molecule. Furthermore, cDNA-mRNA hybrids are synthetic and do not exist in vivo. Besides cDNA not existing in vivo, cDNA is necessarily different than mRNA, as it includes deoxyribonucleic acid and not ribonucleic acid. The cDNA is then amplified, for example, by the polymerase chain reaction (PCR) or other amplification method known to those of ordinary skill in the art. For example, other amplification methods that may be employed include the ligase chain reaction (LCR) (Wu and Wallace, Genomics, 4:560 (1989), Landegren et al., Science, 241:1077 (1988), incorporated by reference in its entirety for all purposes, transcription amplification (Kwoh et al., Proc. Natl. Acad. Sci. USA, 86:1173 (1989), incorporated by reference in its entirety for all purposes), self-sustained sequence replication (Guatelli et al., Proc. Nat. Acad. Sci. USA, 87:1874 (1990), incorporated by reference in its entirety for all purposes), incorporated by reference in its entirety for all purposes, and nucleic acid based sequence amplification (NASBA). Guidelines for selecting primers for PCR amplification are known to those of ordinary skill in the art. See, e.g., McPherson et al., PCR Basics: From Background to Bench, Springer- Verlag, 2000, incorporated by reference in its entirety for all purposes. The product of this amplification reaction, i.e., amplified cDNA is also necessarily a non-natural product. First, as mentioned above, cDNA is a non-natural molecule. Second, in the case of PCR, the amplification process serves to create hundreds of millions of cDNA copies for every individual cDNA molecule of starting material. The numbers of copies generated are far removed from the number of copies of mRNA that are present in vivo.

[0037] In one embodiment, cDNA is amplified with primers that introduce an additional DNA sequence (e.g., adapter, reporter, capture sequence or moiety, barcode) onto the fragments (e.g., with the use of adapter-specific primers), or mRNA or cDNA biomarker sequences are hybridized directly to a cDNA probe comprising the additional sequence (e.g., adapter, reporter, capture sequence or moiety, barcode). Amplification and / or hybridization of mRNA to a cDNA probe therefore serves to create non-natural double stranded molecules from the non-natural single stranded cDNA, or the mRNA, by introducing additional sequences and forming non-naturalhybrids. Further, as known to those of ordinary skill in the art, amplification procedures have error rates associated with them. Therefore, amplification introduces further modifications into the cDNA molecules. In one embodiment, during amplification with the adapter-specific primers, a detectable label, e.g., a fluorophore, is added to single strand cDNA molecules. Amplification therefore also serves to create DNA complexes that do not occur in nature, at least because (i) cDNA does not exist in vivo, (i) adapter sequences are added to the ends of cDNA molecules to make DNA sequences that do not exist in vivo, (ii) the error rate associated with amplification further creates DNA sequences that do not exist in vivo, (iii) the disparate structure of the cDNA molecules as compared to what exists in nature, and (iv) the chemical addition of a detectable label to the cDNA molecules.

[0038] In some embodiments, the expression of a biomarker of interest is detected at the nucleic acid level via detection of non-natural cDNA molecules. EXAMPLES

[0039] The present disclosure is further illustrated by reference to the following Examples. However, it should be noted that these Examples, like the embodiments described above, are illustrative and are not to be construed as restricting the scope of the invention in any way. Example 1- Fast Fourier Transformation (FFT) magnitude as gene expression surrogate marker Objective

[0040] Develop a “expression-like” surrogate features from liquid fragmentomics data. Methods

[0041] Run the Griffin pipeline on cfDNA whole genome sequencing (WGS) data as outlined in Doebley, AL., Ko, M., Liao, H. et al. Author Correction: A framework for clinical cancer subtyping from nucleosome profiling of cell-free DNA. Nat Commun 14, 403 (2023). The WGS data was in measured samples obtained from subjects suffering from cancer.

[0042] Take GC corrected coverage data around gene regions of interest.

[0043] Calculate the mean read counts for a sliding window in each region.

[0044] Apply fast Fourier transform (FFT) to the mean data and extract the highest magnitude as follows:

[0045] Take the 132 count values in the 2KB region (-990 ~ +975, 15bp interval)

[0046] Perform fast Fourier transform with fft() in R

[0047] Amplitude is encoded as the magnitude of the complex number.

[0048] Signal we wanted was the one with smallest frequency, and highest magnitude (see FIG. 1).

[0049] Extract the first element (max magnitude) from the FFT output using Mod().

[0050] The magnitude was utilized then in a traditional nearest centroid classifier (see Dabney, Alan R. "ClaNC: point-and-click software for classifying microarrays to nearest centroids." Bioinformatics22, no.1 (2005): 122-123). Training Nearest Centroid Classifier

[0051] Training for BRCA vs CRC classifier used 7 / 10 of the n=62 liquid biopsy samples with whole genome sequencing of the liquid samples and matching clinical diagnosis.

[0052] Derive gene regions of interest. Based on Pan cancer chromatin accessibility peak calls (Corces et al. The chromatin accessibility landscape of primary human cancers. Science.2018 Oct 26;362(6413)), we filtered for rows which have the value ‘Promoter’ in the Annotation column. Data table was then sorted in descending order for the ‘Start’ column corresponding to the starting base position for that promoter site. First occurrence of promoter sites for each gene were retained to use as “region of interest” for corresponding genes. Nearest centroid classifiers

[0053] For each gene in the FFT magnitude matrix (one promoter region mapped to each gene as described above, ~15K genes total) using the training samples, we derived mean and variance across all training samples. We kept the high mean high variance genes (~5,800 genes) for feature selection. The number of genes to include in the classifier was determined using ClaNC softwareand 5-fold cross-validation as shown in FIG. 2 (see Dabney, Alan R. "ClaNC: point-and-click software for classifying microarrays to nearest centroids." Bioinformatics22, no. 1 (2005): 122- 123). ClaNC was then run on the entire training set to determine the set of genes for the classifier. An ordinary nearest centroid classifier was then fit using the selected genes. The fitted classifier was then tested in the testing samples to evaluate performance (see FIG. 3). Classifier Genes in TCGA gene expression

[0054] We evaluated whether the 184 genes segregating BRCA vs CRC were indeed surrogate of gene expression markers. We plotted a subset of these genes whose expression value were available in the TCGA set (126 of 184 genes).

[0055] We plotted this for the combined BRCA, READ, and COAD cohorts. The classifier genes, derived algorithmically from FFT magnitude, separate cancer types in TCGA based on their expression profile, suggesting that FFT magnitude derived using this method is a good surrogate to the gene expression markers (see FIG. 4).

[0056] To further evaluate the significance of this finding, we performed permutation tests to evaluate how well a random set of genes from a classifier derived elsewhere classifies the BRCA vs CRC cohorts (see FIG. 5).

[0057] The heatmap using the classifier genes (126 of 184 symbols directly present) for clustering appeared to separate the tumor types well (see FIG.4). But how well? Is it better than chance?

[0058] The bar may be different (higher?) here (to eclipse the random gene sets compared to the first classifier) because the gene set is twice as big. Nonetheless, this is great evidence to support the CV in the cfDNA data. Conclusions

[0059] Genes that were chosen were those with FFT magnitude differences that stratified tumor types in a similar pattern as with TCGA gene expression data. WES data, depending on the capture method, may also be applicable to this approach.

[0060] Provide more dynamic range when examining sample-sample correlation. This is usually a sign of better control of noise (see FIG.6).

[0061] Show more relevance biology in GSEA analyses of top genes associated with healthy vs malignant comparison. This again is usually indicative of better signal to noise ratio (see FIG.7). Further Numbered Embodiments of the Disclosure

[0062] Other subject matter contemplated by the present disclosure is set out in the following numbered embodiments:

[0063] 1. A computer implemented method for generating a classifier for determining a cancer type of a subject suffering from a cancer or suspected of suffering from a cancer; the method comprising: (a) receiving in a computer system, sequence read data generated from nucleic acid extracted from a sample obtained from each subject from a plurality of subjects suffering from a cancer or suspected of suffering from a cancer, wherein the sample comprises cell-free DNA (cfDNA), wherein the sequence read data includes a plurality of fragment reads, wherein each fragment read has a fragment length and a GC content indicating a percentage of bases in the fragment read that are G or C; (b) determining by the computing system, GC bias values for each fragment read based on the fragment length and the GC content of the fragment read; (c) generating by the computing system, a genomic coverage distribution that is adjusted for GC bias using the sequence read data and the GC bias values; (d) calculating by the computing system mean sequence read counts for a sliding window across a defined window in each of a plurality of genomic regions of interest from the genomic coverage distribution to generate smoothed mean read counts, wherein each of the plurality of genomic regions of interest are cancer type informative sites; (e) performing a fast Fourier transform (FFT) on the smoothed mead read counts to generate an FFT magnitude matrix for each of the plurality of genomic regions of interest; (f) selecting genes from the FFT magnitude matrix for each of the plurality of genomic regions of interest that have a high mean and high variance; (g) inputting the selected genes into a Classifying arrays to Nearest Centroid (CLaNC) model on the computing system to determine a set of genes for the cancer type classifier; and (h) fitting an ordinary nearest centroid classifier using the set of genes determined in step (g) to generate the cancer type classifier.

[0064] 2. The method of embodiment 1, wherein the sliding window has a width of at least, at most or exactly 5, 10, 15, 20, 25, 30, 35, 40, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95 or 100 nucleotides across the defined window.

[0065] 3. The method of embodiment 1 or 2, wherein the sliding window has a width of 15 nucleotides across the defined window.

[0066] 4. The method of any one of the above embodiments, wherein the defined window has a width of 2000 base pairs.

[0067] 5. The method of any one of the above embodiments, wherein the cancer-type informative sites are sites that have differential expression between a first cancer type and a second cancer type.

[0068] 6. The method of any one of the above embodiments, wherein the cancer type is a cancer subtype of a type of cancer.

[0069] 7. The method of any one of the above embodiments, wherein the cancer is selected from the group consisting of kidney renal papillary cell carcinoma (KIRP); breast invasive carcinoma (BRCA); thyroid cancer (THCA); bladder urothelial carcinoma (BLCA); prostate adenocarcinoma (PRAD); kidney chromophobe (KICH); cervical squamous cell carcinoma and endocervical adenocarcinoma (CESC); kidney renal clear cell carcinoma (KIRC); liver hepatocellular carcinoma (LIHC); low grade glioma (LGG); sarcoma (SARC); lung adenocarcinoma (LUAD); colon adenocarcinoma (COAD); head and neck squamous cell carcinoma (HNSC); uterine corpus endometrial carcinoma (UCEC); glioblastoma multiforme (GBM); esophageal carcinoma (ESCA); stomach adenocarcinoma (STAD); ovarian serous cystadenocarcinoma (OV); rectum adenocarcinoma (READ); adrenocortical carcinoma (ACC); uveal melanoma (UVM); mesothelioma (MESO); pheochromocytoma and paraganglioma (PCPG); skin cutaneous melanoma (SKCM); uterine carcinosarcoma (UCS); lung squamous cell carcinoma (LUSC); testicular germ cell tumors (TGCT); cholangiocarcinoma (CHOL); pancreatic adenocarcinoma (PAAD); thymoma (THYM); or Lymphoid Neoplasm Diffuse Large B-cell Lymphoma (DLBC).

[0070] 8. The method of any one of the above embodiments, wherein the sample is a formalin- fixed, paraffin-embedded (FFPE) tissue sample, a fresh or a frozen tissue sample, an exosome, wash fluids, cell pellets, or a bodily fluid obtained from the subject.

[0071] 9. The method of embodiment 8, wherein the bodily fluid is blood or fractions thereof, urine, saliva, or sputum.

[0072] 10. The method of any one of the above embodiments, wherein the nucleic acid is DNA, RNA or cDNA.

[0073] 11. The method of any one of the above embodiments, wherein the sequence read data is whole genome sequencing (WGS) data.

[0074] 12. The method of any one of embodiments 1-10, wherein the sequence read data is whole exome sequencing (WES) data.

[0075] 13. A method for determining a cancer type in a subject suffering from or suspected of suffering from a cancer comprising inputting sequence read data generated from a sample obtained from the subject into the cancer type classifier developed by computer-implemented method of embodiments 1-12.

[0076] 14. A method for treating cancer in a subject suffering from or suspected of suffering from a cancer comprising: (a) determining a cancer type of the subject by inputting sequence read data generated from a sample obtained from the subject into the cancer type classifier developed by computer-implemented method of embodiments 1-12; and (b) administering a therapeutic specific for the cancer type of the subject determined in step (a), thereby treating the cancer.

[0077] 15. A method for selecting a cancer treatment for a subject suffering from or suspected of suffering from a cancer comprising: (a) determining a cancer type of the subject by inputting sequence read data generated from a sample obtained from the subject into the cancer type classifier developed by computer-implemented method of embodiments 1-12; and (b) selecting a therapeutic specific the cancer type of the subject determined in step (a). * * * * * * *

[0078] The various embodiments described above can be combined to provide further embodiments. All of the U.S. patents, U.S. patent application publications, U.S. patent application, foreign patents, foreign patent application and non-patent publications referred to in this specification and / or listed in the Application Data Sheet are incorporated herein by reference, in their entirety. Aspects of the embodiments can be modified, if necessary to employ concepts of the various patents, application and publications to provide yet further embodiments.

[0079] These and other changes can be made to the embodiments in light of the above-detailed description. In general, in the following claims, the terms used should not be construed to limit the claims to the specific embodiments disclosed in the specification and the claims, but should beconstrued to include all possible embodiments along with the full scope of equivalents to which such claims are entitled. Accordingly, the claims are not limited by the disclosure. INCORPORATION BY REFERENCE

[0080] All references, articles, publications, patents, patent publications, and patent applications cited herein are incorporated by reference in their entireties for all purposes. However, mention of any reference, article, publication, patent, patent publication, and patent application cited herein is not, and should not be taken as an acknowledgment or any form of suggestion that they constitute valid prior art or form part of the common general knowledge in any country in the world.

Claims

CLAIMS What is claimed:

1. A computer implemented method for generating a classifier for determining a cancer type of a subject suffering from a cancer or suspected of suffering from a cancer; the method comprising: (a) receiving in a computer system, sequence read data generated from nucleic acid extracted from a sample obtained from each subject from a plurality of subjects suffering from a cancer or suspected of suffering from a cancer, wherein the sample comprises cell-free DNA (cfDNA), wherein the sequence read data includes a plurality of fragment reads, wherein each fragment read has a fragment length and a GC content indicating a percentage of bases in the fragment read that are G or C; (b) determining by the computing system, GC bias values for each fragment read based on the fragment length and the GC content of the fragment read; (c) generating by the computing system, a genomic coverage distribution that is adjusted for GC bias using the sequence read data and the GC bias values; (d) calculating by the computing system mean sequence read counts for a sliding window across a defined window in each of a plurality of genomic regions of interest from the genomic coverage distribution to generate smoothed mean read counts, wherein each of the plurality of genomic regions of interest are cancer type informative sites; (e) performing a fast Fourier transform (FFT) on the smoothed mead read counts to generate an FFT magnitude matrix for each of the plurality of genomic regions of interest; (f) selecting genes from the FFT magnitude matrix for each of the plurality of genomic regions of interest that have a high mean and high variance; (g) inputting the selected genes into a Classifying arrays to Nearest Centroid (CLaNC) model on the computing system to determine a set of genes for the cancer type classifier; and (h) fitting an ordinary nearest centroid classifier using the set of genes determined in step (g) to generate the cancer type classifier.

2. The method of claim 1, wherein the sliding window has a width of at least, at most or exactly 5, 10, 15, 20, 25, 30, 35, 40, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95 or 100 nucleotides across the defined window.

3. The method of claim 1, wherein the sliding window has a width of 15 nucleotides across the defined window.

4. The method of claim 1, wherein the defined window has a width of 2000 base pairs.

5. The method of claim 1, wherein the cancer-type informative sites are sites that have differential expression between a first cancer type and a second cancer type.

6. The method of claim 1, wherein the cancer type is a cancer subtype of a type of cancer.

7. The method of claim 1, wherein the cancer is selected from the group consisting of kidney renal papillary cell carcinoma (KIRP); breast invasive carcinoma (BRCA); thyroid cancer (THCA); bladder urothelial carcinoma (BLCA); prostate adenocarcinoma (PRAD); kidney chromophobe (KICH); cervical squamous cell carcinoma and endocervical adenocarcinoma (CESC); kidney renal clear cell carcinoma (KIRC); liver hepatocellular carcinoma (LIHC); low grade glioma (LGG); sarcoma (SARC); lung adenocarcinoma (LUAD); colon adenocarcinoma (COAD); head and neck squamous cell carcinoma (HNSC); uterine corpus endometrial carcinoma (UCEC); glioblastoma multiforme (GBM); esophageal carcinoma (ESCA); stomach adenocarcinoma (STAD); ovarian serous cystadenocarcinoma (OV); rectum adenocarcinoma (READ); adrenocortical carcinoma (ACC); uveal melanoma (UVM); mesothelioma (MESO); pheochromocytoma and paraganglioma (PCPG); skin cutaneous melanoma (SKCM); uterine carcinosarcoma (UCS); lung squamous cell carcinoma (LUSC); testicular germ cell tumors (TGCT); cholangiocarcinoma (CHOL); pancreatic adenocarcinoma (PAAD); thymoma (THYM); or Lymphoid Neoplasm Diffuse Large B-cell Lymphoma (DLBC).

8. The method of claim 1, wherein the sample is a formalin-fixed, paraffin-embedded (FFPE) tissue sample, a fresh or a frozen tissue sample, an exosome, wash fluids, cell pellets, or a bodily fluid obtained from the subject.

9. The method of claim 8, wherein the bodily fluid is blood or fractions thereof, urine, saliva, or sputum.

10. The method of claim 1, wherein the nucleic acid is DNA, RNA or cDNA.

11. The method of claim 1, wherein the sequence read data is whole genome sequencing (WGS) data.

12. The method of claim 1, wherein the sequence read data is whole exome sequencing (WES) data.

13. A method for determining a cancer type in a subject suffering from or suspected of suffering from a cancer comprising inputting sequence read data generated from a sample obtained from the subject into the cancer type classifier developed by computer-implemented method of claim 1.

14. A method for treating cancer in a subject suffering from or suspected of suffering from a cancer comprising: (a) determining a cancer type of the subject by inputting sequence read data generated from a sample obtained from the subject into the cancer type classifier developed by computer- implemented method of claim 1; and (b) administering a therapeutic specific for the cancer type of the subject determined in step (a), thereby treating the cancer.

15. A method for selecting a cancer treatment for a subject suffering from or suspected of suffering from a cancer comprising: (a) determining a cancer type of the subject by inputting sequence read data generated from a sample obtained from the subject into the cancer type classifier developed by computer-implemented method of claim 1; and (b) selecting a therapeutic specific the cancer type of the subject determined in step (a).

Citation Information

Patent Citations

  • Methods for fragmentome profiling of cell-free nucleic acids

    US20190287645A1

  • Detecting cancer cell of origin

    US20230037765A1