Methods and systems for determining the circulating tumor DNA fraction in a patient sample
By utilizing somatic short variants and historical data, the method accurately estimates tumor DNA fraction in liquid biopsy samples, addressing inaccuracies in existing methods and improving diagnostic and monitoring capabilities.
Patent Information
- Application Number
- JP2025501526
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-15
- Filing Date
- 2023-07-14
- Publication Date
- 2025-07-30
AI Technical Summary
The prior art is difficult to accurately estimate the content of circulating tumor DNA (ctDNA) in liquid biopsy samples, resulting in insufficient diagnostic, prediction and disease monitoring performance, affecting the widespread application of liquid biopsy technology.
By analyzing the genomic variation frequency and tumor ploidy status, combining historical patient genomic data, the probability density model is used to estimate ctDNA content, considering the impact of copy number and tumor ploidy on the variation frequency.
Improve the accuracy estimation of ctDNA content and enhance the diagnostic, predictive and disease monitoring capabilities of liquid biopsy technology.
Smart Images

Figure 2025524649000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 389,733, filed on July 15, 2022, the content of which is hereby incorporated by reference in its entirety.
[0002] Field of the Invention The present disclosure generally relates to methods and systems for analyzing genomic profiling data, and more specifically, to methods and systems for using genomic profiling data to determine the fraction of tumor DNA in a sample obtained from a subject (e.g., a patient).
Background Art
[0003] Background Analysis of genomic profile data obtained by sequencing DNA extracted from patient samples for disease diagnosis or prognosis often requires estimating the proportion of tumor DNA present in the sample. Particularly in liquid biopsy - based technologies, accurate assessment of the amount of circulating tumor DNA (ctDNA) in patient plasma (or other liquid biopsy samples) can be essential. Plasma ctDNA content serves as an important quality control (QC) metric for assays and also as a biomarker for disease diagnosis, disease prognosis, and disease monitoring purposes. The fraction of ctDNA present in the total amount of cell - free DNA (cfDNA) isolated from a sample can be estimated, for example, by detecting the level of tumor aneuploidy indicated by cfDNA. However, due to the characteristics of cfDNA present in liquid biopsy samples and the inherently low levels of cfDNA that are often present in liquid biopsy samples, evidence of tumor aneuploidy may not be present or may be below detectable levels for many samples. Therefore, there is a need for improved methods for determining the ctDNA fraction in liquid biopsy samples in order to improve the diagnostic, prognostic, and disease monitoring performance of liquid biopsy - based technologies and thereby accelerate their adoption rate and impact on improved healthcare outcomes.
Summary of the Invention
[0004] Brief Summary of the Invention This specification discloses methods and systems that can generally provide a more accurate assessment of the tumor DNA content in a biopsy sample, and in particular, a more accurate assessment of the circulating tumor DNA content in a cell-free DNA sample collected from a patient's plasma or other liquid biopsy sample. The methods described herein overcome the problems encountered by existing tumor aneuploidy-based methods for estimating the tumor DNA fraction when no detectable aneuploidy signal is present, by leveraging alternative signals, such as signals associated with the presence of somatic short variants, to infer the tumor DNA fraction (e.g., the ctDNA fraction). This approach is based on the following. (i) The derived physical relationship between the allele frequency of a variant (e.g., a somatic short variant) and the tumor content determined by the copy number state of the genomic position of the variant and the tumor average ploidy, and (ii) access to a broad database collection of patient genomic profile data.
[0005] Briefly, variants (e.g., somatic short variants) are identified in a sample (e.g., a biopsy sample, a liquid biopsy sample, or a plasma sample) based on sequence read data derived from the sample, and their allele frequencies are quantified. Next, an empirical distribution of all possible sample tumor DNA fraction values considering the observed somatic short variant allele frequencies can be generated by (i) using an equation that describes the physical relationship between the somatic variant allele frequency and the tumor DNA fraction for a given copy number state of the variant genomic position and a given tumor average ploidy, and (ii) leveraging historical patient data that includes known copy number and tumor ploidy profiles. Then, a model, such as a probability density model (e.g., a non-parametric probability density model), can be fit to the empirical distribution of the tumor DNA fraction values, and an estimated value of the tumor DNA fraction of the sample can be derived from the model along with the upper and lower limits of the estimated values based on the desired confidence interval.
[0006] The disclosed method achieves a more accurate assessment of tumor DNA content by considering the effects of copy number and tumor ploidy on variant allele frequency (VAF), and thus provides an estimate that more closely reflects the true tumor DNA fraction of the sample.
[0007] Providing a plurality of nucleic acid molecules obtained from a sample from a subject; ligating one or more adapters onto one or more of the nucleic acid molecules from the plurality of nucleic acid molecules; amplifying one or more ligated nucleic acid molecules from the plurality of nucleic acid molecules; capturing the amplified nucleic acid molecules from the amplified nucleic acid molecules; sequencing the captured nucleic acid molecules by a sequencer to obtain a plurality of sequence reads representing the captured nucleic acid molecules; receiving, in one or more processors, sequence read data in the plurality of sequence reads; using one or more processors to determine a variant allele frequency (VAF) in one or more variants detected in the sample based on the sequence read data; using one or more processors to generate an empirical distribution of tumor DNA fraction values according to the determined VAF in the one or more variants; using one or more processors to fit a model to the empirical distribution of tumor DNA fraction values; and determining a tumor DNA fraction in the sample based on the model. A method is disclosed herein that includes these steps.
[0008] In some embodiments, the method further includes determining a confidence interval in the tumor DNA fraction based on the model. In some embodiments, the one or more variants include one or more somatic short variants. In some embodiments, the one or more somatic short variants are known not to be associated with clonal hematopoiesis of indeterminate potential (CHIP).
[0009] In some embodiments, generating an empirical distribution of tumor DNA fraction values involves calculating a tumor DNA fraction value based on a known copy number in one or more variants, a determined VAF in one or more variants, and a corresponding known tumor average ploidy in a plurality of historical subject samples having a known VAF in one or more variants that is substantially the same as the determined VAF in one or more variants. In some embodiments, generating an empirical distribution of tumor DNA fraction values involves pre-calculating a tumor DNA fraction value based on a known copy number in one or more variants and a corresponding known tumor average ploidy in a plurality of historical subject samples having a range of VAF values in one or more variants, and selecting a subset of the pre-calculated tumor DNA fraction values corresponding to samples having a known VAF in one or more variants that is substantially the same as the determined VAF in one or more variants.
[0010] In some embodiments, the tumor DNA fraction value is calculated or pre-calculated by solving a set of equations representing the relationships between (i) tumor DNA fraction, tumor purity, and tumor average ploidy, and (ii) somatic VAF, tumor purity, copy number at the genomic location of one or more variants, and the number of variant alleles in each of the one or more variants, based on a known VAF in one or more variants, a known copy number in one or more variants, and a corresponding known tumor average ploidy in a plurality of historical subject samples, thereby deriving a relationship in tumor DNA fraction as a function of somatic VAF, tumor average ploidy, copy number at the genomic location of one or more variants, and the number of variant alleles for one or more variants, excluding tumor purity.
[0011] In some embodiments, the plurality of historical subject samples include solid biopsy samples, liquid biopsy samples, or any combination thereof. In some embodiments, the plurality of historical subject samples include cancer samples. In some embodiments, the plurality of historical subject samples include samples in a single type of cancer. In some embodiments, the plurality of historical subject samples include samples in multiple types of cancer.
[0012] In some embodiments, the plurality of historical subject samples are acute lymphoblastic leukemia (Philadelphia chromosome positive) samples, acute lymphoblastic leukemia (precursor B cell) samples, acute myeloid leukemia (FLT3+) samples, acute myeloid leukemia (with IDH2 mutation) samples, anaplastic large cell lymphoma samples, basal cell carcinoma samples, B-cell chronic lymphocytic leukemia samples, bladder cancer samples, breast cancer (HER2 overexpression / amplification) samples, breast cancer (HER2+) samples, breast cancer (HR+, HER2-) samples, cervical cancer samples, cholangiocarcinoma samples, chronic lymphocytic leukemia samples, chronic lymphocytic leukemia (with 17p deletion) samples, chronic myeloid leukemia samples, chronic myeloid leukemia (Philadelphia chromosome positive) samples, classical Hodgkin lymphoma samples, colorectal cancer samples, colorectal cancer (dMMR and MSI-H) samples, colorectal cancer (KRAS wild type) samples, cryopyrin-associated periodic syndrome samples, cutaneous T-cell lymphoma samples, dermatofibrosarcoma protuberans samples, diffuse large B-cell lymphoma samples, fallopian tube cancer samples, follicular B-cell non-Hodgkin lymphoma samples, follicular lymphoma samples, gastric cancer samples, gastric cancer (HER2+) samples, gastroesophageal junction (GEJ) adenocarcinoma samples, gastrointestinal stromal tumor samples, gastrointestinal stromal tumor samples, gastrointestinal stromal tumor (KIT+) samples, giant cell tumor of bone samples, glioblastoma samples, polyangiitis granulomatosa samples, head and neck squamous cell carcinoma samples, hepatocellular carcinoma samples, Hodgkin lymphoma samples, juvenile idiopathic arthritis samples, lupus erythematosus samples, mantle cell lymphoma samples, medullary thyroid cancer samples, melanoma samples, melanoma with BRAF V600 mutation samples, melanoma with BRAF V600E or V600K mutation samples, Merkel cell carcinoma samples, Castleman disease samples, multiple blood samples including Philadelphia chromosome positive ALL and CML, multiple myeloma samples, myelofibrosis samples, non-Hodgkin lymphoma samples, inoperable subependymal giant cell astrocytoma samples associated with tuberous sclerosis, non-small cell lung cancer samples, non-small cell lung cancer (ALK+) samples, non-small cell lung cancer (PD-L1+) samples, non-small cell lung cancer (with ALK fusion or ROS1 gene alteration) samples, non-small cell lung cancer (with BRAF V600E mutation) samples, non-small cell lung cancer samples (with EGFR exon 19 deletion or exon 21 substitution (L858R) mutation), non-small cell lung cancer samples (EGFR(with T790M mutation), ovarian cancer samples, ovarian cancer samples (with BRCA mutation), pancreatic cancer samples, pancreatic cancer samples, gastrointestinal cancer samples, lung-derived neuroendocrine tumor samples, pediatric neuroblastoma samples, peripheral T-cell lymphoma samples, prostate cancer samples, peritoneal cancer samples, renal cell cancer samples, rheumatoid arthritis samples, small lymphocyte lymphoma samples, soft tissue sarcoma samples, solid tumor (MSI-H / dMMR) samples, head and neck squamous cell carcinoma samples, squamous non-small cell lung cancer samples, thyroid cancer samples, thyroid cancer samples, urothelial cancer samples, urothelial cancer samples, Waldenström macroglobulinemia samples, or any combination thereof.
[0013] In some embodiments, the model is a parametric probability density model. In some embodiments, the model is a non-parametric probability density model. In some embodiments, the determined tumor DNA fraction in the sample is the most likely tumor DNA fraction. In some embodiments, the determined tumor DNA fraction in the cfDNA sample is the mean, median, or mode of the dominant peak in the empirical distribution of the tumor DNA fraction values.
[0014] In some embodiments, the sample comprises DNA extracted from a blood sample, plasma sample, cerebrospinal fluid sample, pleural effusion sample, sputum sample, fecal sample, urine sample or saliva sample.
[0015] In some embodiments, the subject is suspected of having cancer or is determined to have cancer. In some embodiments, the cancer is acute lymphoblastic leukemia (Philadelphia chromosome positive), acute lymphoblastic leukemia (precursor B cell), acute myeloid leukemia (FLT3+), acute myeloid leukemia (with IDH2 mutation), anaplastic large cell lymphoma, basal cell carcinoma, B-cell chronic lymphocytic leukemia, bladder cancer, breast cancer (HER2 overexpression / amplification), breast cancer (HER2+), breast cancer (HR+, HER2-), cervical cancer, cholangiocarcinoma, chronic lymphocytic leukemia, chronic lymphocytic leukemia (with 17p deletion), chronic myelogenous leukemia, chronic myelogenous leukemia (Philadelphia chromosome positive), classical Hodgkin lymphoma, colorectal cancer, colorectal cancer (dMMR / MSI-H), colorectal cancer (KRAS wild type), cryopyrin-associated periodic syndrome, cutaneous T-cell lymphoma, precursor cutaneous fibrosarcoma, diffuse large B-cell lymphoma, fallopian tube cancer, follicular B-cell non-Hodgkin lymphoma, follicular lymphoma, gastric cancer, gastric cancer (HER2+), gastroesophageal junction (GEJ) adenocarcinoma, gastrointestinal stromal tumor, gastrointestinal stromal tumor (KIT+), giant cell tumor of bone, glioblastoma, polyangiitis granulomatosa, head and neck squamous cell carcinoma, hepatocellular carcinoma, Hodgkin lymphoma, juvenile idiopathic arthritis, lupus erythematosus, mantle cell lymphoma, medullary thyroid carcinoma, melanoma, melanoma with BRAF V600 mutation, melanoma with BRAF V600E or V600K mutation, Merkel cell carcinoma, multicentric Castleman disease, multiple hematologic malignancies including Philadelphia chromosome positive ALL and CML, multiple myeloma, myelofibrosis, non-Hodgkin lymphoma, unresectable subependymal giant cell astrocytoma associated with tuberous sclerosis, non-small cell lung cancer, non-small cell lung cancer (ALK+), non-small cell lung cancer (PD-L1+), non-small cell lung cancer (with ALK fusion or ROS1 gene change), non-small cell lung cancer (with BRAF V600E mutation), non-small cell lung cancer (with EGFR exon 19 deletion or exon 21 substitution (L858R) mutation), non-small cell lung cancer (EGFRaccompanied by the T790M mutation), ovarian cancer, ovarian cancer (accompanied by the BRCA mutation), pancreatic cancer, pancreatic cancer, gastrointestinal cancer or lung-derived neuroendocrine tumor, pediatric neuroblastoma, peripheral T-cell lymphoma, peritoneal cancer, prostate cancer, renal cell carcinoma, rheumatoid arthritis, small lymphocyte lymphoma, soft tissue sarcoma, solid tumor (MSI-H / dMMR), head and neck squamous cell carcinoma, non-small cell lung cancer, squamous cell carcinoma, thyroid cancer, thyroid cancer, urothelial cancer, urothelial cancer, or Waldenström macroglobulinemia.
[0016] In some embodiments, the method further comprises treating the subject with an anti-cancer therapy. In some embodiments, the anti-cancer therapy comprises a targeted anti-cancer therapy. In some embodiments, the targeted anti-cancer therapy is abemaciclib (Verzenio), abiraterone acetate (Zytiga), acalabrutinib (Calquence), trastuzumab emtansine (Kadcyla), afatinib dimaleate (Gilotrif), aldesleukin (Proleukin), alectinib (Alecensa), alemtuzumab (Campath), alitretinoin (Panretin), alpelisib (Piqray), amivantamab-vmjw (Rybrevant), anastrozole (Arimidex), apalutamide (Erleada), asimicinib hydrochloride (Scemblix), atezolizumab (Tecentriq), avapritinib (Ayvakit), avelumab (Bavencio), axicabtagene ciloleucel (Yescarta), axitinib (Inlyta), belantamab mafodotin-blmf (Blenrep), belimumab (Benlysta), belinostat (Beleodaq), belzutifan (Welireg), bevacizumab (Avastin), bexarotene (Targretin), binimetinib (Mektovi), blinatumomab (Blincyto), bortezomib (Velcade), bosutinib (Bosulif), brentuximab vedotin (Adcetris), brexucabtagene autoleucel (Tecartus), brigatinib (Alunbrig), cabazitaxel (Jevtana), cabozantinib (Cabometyx), cabozantinib (Cabometyx, Cometriq), canakinumab (Ilaris), capmatinib hydrochloride (Tabrecta), carfilzomib (Kyprolis), cemiplimab-rwlc (Libtayo), ceritinib (LDK378 / Zykadia), cetuximab (Erbitux), cobimetinib (Cotellic), copanlisib hydrochloride (Aliqopa), crizotinib (Xalkori), dabrafenib (Tafinlar), dacomitinib (Vizimpro), daratumumab (Darzalex),Daratumumab and hyaluronidase-fihj (Darzalex Faspro), darolutamide (Nubeqa), dasatinib (Sprycel), denileukin diftitox (Ontak), denosumab (Xgeva), dinutuximab (Unituxin), dostarlimab-gxly (Jemperli), durvalumab (Imfinzi), duvelisib (Copiktra), elotuzumab (Empliciti), enasidenib mesylate (Idhifa), encorafenib (Braftovi), enfortumab vedotin-ejfv (Padcev), entrectinib (Rozlytrek), enzalutamide (Xtandi), erdafitinib (Balversa), erlotinib (Tarceva), everolimus (Afinitor), exemestane (Aromasin), fam-trastuzumab deruxtecan-nxki (Enhertu), fedratinib hydrochloride (Inrebic), fulvestrant (Faslodex), gefitinib (Iressa), gemtuzumab ozogamicin (Mylotarg), gilteritinib (Xospata), glasdegib maleate (Daurismo), hyaluronidase-zzxf (Phesgo), ibrutinib (Imbruvica), ibritumomab tiuxetan (Zevalin), idecabtagene vicleucel (Abecma), idelalisib (Zydelig), imatinib mesylate (Gleevec), infigratinib phosphate (Truseltiq), inotuzumab ozogamicin (Besponsa), iobenguane I131 (Azedra), ipilimumab (Yervoy), isatuximab-irfc (Sarclisa),ivosidenib (Tibsovo), ixazomib citrate (Ninlaro), lanreotide acetate (Somatuline Depot), lapatinib (Tykerb), larotrectinib sulfate (Vitrakvi), lenvatinib mesylate (Lenvima), letrozole (Femara), isocabtagene autoleucel (Breyanzi), loncastuximab tesirine-lpyl (Zynlonta), lorlatinib (Lorbrena), lutetium Lu177-dotatate (Lutathera), margetuximab-cmkb (Margenza)Midostaurin (Rydapt), mobocertinib succinate (Exkivity), mogamulizumab-kpkc (Poteligeo), moxetumomab pasudotox-tdfk (Lumoxiti), naxitamab-gqgk (Danyelza), necitumumab (Portrazza), neratinib maleate (Nerlynx), nilotinib (Tasigna), niraparib tosylate monohydrate (Zejula), nivolumab (Opdivo), obinutuzumab (Gazyva), ofatumumab (Arzerra), olaparib (Lynparza), olaratumab (Lartruvo), osimertinib (Tagrisso), palbociclib (Ibrance), panitumumab (Vectibix), panobinostat (Farydak), pazopanib (Votrient), pembrolizumab (Keytruda), pemigatinib (Pemazyre), pertuzumab (Perjeta), peqidartinib hydrochloride (Turalio), polatuzumab vedotin-piiq (Polivy), ponatinib hydrochloride (Iclusig), pralatrexate (Folotyn), pralsetinib (Gavreto), radium 223 dichloride (Xofigo), ramucirumab (Cyramza), regorafenib (Stivarga), ribociclib (Kisqali), ripretinib (Qinlock), rituximab (Rituxan), rituximab and hyaluronidase human (Rituxan Hycela), romidepsin (Istodax), rucaparib (Rubraca), ruxolitinib phosphate (Jakafi), sacituzumab govitecan-fuji (Trodelvy), selinexor, selpercatinib (Retevmo), serdemetanib sulfate (Koselugo), siltuximab (Sylvant), sipuleucel-T (Provenge), sirolimus protein-bound particles (Fyarro), sonidegib (Odomzo), sorafenib (Nexavar), sotorasib (Lumakras), sunitinib (Sutent), tafasitamab-cxix (Monjuvi), tagraxofusp-erzs (Elzonris), talazoparib tosylate (Talzenna), tamoxifen (Nolvadex),Tazemetostat Hydrobromide (Tazverik), Tebentafusp-teb (Kimmtrak), Temsirolimus (Torisel), Tepotinib Hydrochloride (Tepmetko), tisagenlecleucel (Kymriah), tisotumab vedotin-tftv (Tivdak), tocilizumab (Actemra), tofacitinib (Xeljanz), tositumomab (Bexxar), trametinib (Mekinist), trastuzumab (Herceptin), tretinoin (Vesanoid), tivozanib hydrochloride (Fotivda), tamoxifen (Fareston), tucatinib (Tukysa), umbralisib tosylate (Ukoniq), vandetanib (Caprelsa), vemurafenib (Zelboraf), venetoclax (Venclexta), vismodegib (Erivedge), vorinostat (Zolinza), zanubrutinib (Brukinsa), aflibercept (Zaltrap), or any combination thereof.
[0017] In some embodiments, the method further comprises obtaining a tissue sample from a subject. In some embodiments, the sample comprises a tissue biopsy sample, a liquid biopsy sample, or a normal control. In some embodiments, the sample is a liquid biopsy sample and comprises blood, plasma, cerebrospinal fluid, sputum, feces, urine, or saliva. In some embodiments, the sample is a liquid biopsy sample and comprises circulating tumor cells (CTCs). In some embodiments, the sample is a liquid biopsy sample and comprises cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), or any combination thereof. In some embodiments, the plurality of nucleic acid molecules comprises a mixture of tumor nucleic acid molecules and non-tumor nucleic acid molecules. In some embodiments, the tumor nucleic acid molecules are derived from the tumor portion of a heterogeneous tissue biopsy sample and the non-tumor nucleic acid molecules are derived from the normal portion of the heterogeneous tissue biopsy sample. In some embodiments, the sample comprises a liquid biopsy sample, the tumor nucleic acid molecules are derived from the circulating tumor DNA (ctDNA) fraction of the liquid biopsy sample, and the non-tumor nucleic acid molecules are derived from the non-tumor cell-free DNA (cfDNA) fraction of the liquid biopsy sample.
[0018] In some embodiments, one or more adapters include amplification primers, flow cell adapter arrays, substrate adapter arrays, or sample index arrays. In some embodiments, captured nucleic acid molecules are captured from amplified nucleic acid molecules by hybridization to one or more bait molecules. In some embodiments, one or more bait molecules include one or more nucleic acid molecules, each nucleic acid molecule including a region complementary to a region of the captured nucleic acid molecule. In some embodiments, amplifying the nucleic acid molecule includes performing polymerase chain reaction (PCR) amplification techniques, non-PCR amplification techniques, or isothermal amplification techniques. In some embodiments, sequencing includes the use of massively parallel sequencing (MPS) techniques, whole genome sequencing (WGS), whole exome sequencing, targeted sequencing, direct sequencing, or Sanger sequencing techniques. In some embodiments, sequencing includes massively parallel sequencing, and massively parallel sequencing techniques include next generation sequencing (NGS). In some embodiments, sequencing includes next generation sequencers.
[0019] In some embodiments, one or more of the plurality of array determination reads overlap with one or more loci within one or more sub-genomic intervals in the sample. In some embodiments, the one or more loci include from 10 to 20 loci, from 10 to 40 loci, from 10 to 60 loci, from 10 to 80 loci, from 10 to 100 loci, from 10 to 150 loci, from 10 to 200 loci, from 10 to 250 loci, from 10 to 300 loci, from 10 to 350 loci, from 10 to 400 loci, from 10 to 450 loci, from 10 to 500 loci, from 20 to 40 loci, from 20 to 60 loci, from 20 to 80 loci, from 20 to 100 loci, from 20 to 150 loci, from 20 to 200 loci, from 20 to 250 loci, from 20 to 300 loci, from 20 to 350 loci, from 20 to 400 loci, from 20 to 500 loci, from 40 to 60 loci, from 40 to 80 loci, from 40 to 100 loci, from 40 to 150 loci, from 40 to 200 loci, from 40 to 250 loci, from 40 to 300 loci, from 40 to 350 loci, from 40 to 400 loci, from 40 to 500 loci, from 60 to 80 loci, from 60 to 100 loci, from 60 to 150 loci, from 60 to 200 loci, from 60 to 250 loci, from 60 to 300 loci, from 60 to 350 loci, from 60 to 400 loci, from 60 to 500 loci, from 80 to 100 loci, from 80 to 150 loci, from 80 to 200 loci, from 80 to 250 loci, from 80 to 300 loci, from 80 to 350 loci, from 80 to 400 loci, from 80 to 500 loci, from 100 to 150 loci, from 100 to 200 loci, from 100 to 250 loci, from 100 to 300 loci, from 100 to 350 loci, from 100 to 400 loci, from 100 to 500 loci, from 150 to 200 loci, from 150 to 250 loci, from 150 to 300 loci, from 150 to 350 loci, from 150 to 400 loci, from 150 to 500 loci, from 200 to 250 loci, from 200 to 300 loci, from 200 to 350 loci, from 200 to 400 loci, from 200 to 500 loci, from 250 to 300 loci, from 250 to 350 loci, from 250 to 400 loci, from 250 to 500 loci, from 300 to 350 loci, from 300 to 400 loci, from 300 to 500 loci, from 350 to 400 loci, from 350 to 500 loci, or from 400 to 500 loci.
[0020] In some embodiments, one or more loci are ABL1, ACVR1B, AKT1, AKT2, AKT3, ALK, ALOX12B, AMER1, APC, AR, ARAF, ARFRP1, ARID1A, ASXL1, ATM, ATR, ATRX, AURKA, AURKB, AXL, BAP1, BARD1, BCL2, BCL2L1, BCL2L2, BCL6, BCOR, BCORL1, BCR, BRAF, BRCA1, BRCA2, BRD4, BRIP1, BTG1, BTG2, BTK, CALR, CARD11, CASP8, CBFB, CBL, CCND1, CCND2, CCND3, CCNE1, CD22, CD274, CD70, CD74, CD79A, CD79B, CDC73, CDH1, CDK12, CDK4, CDK6, CDK8, CDKN1A, CDKN1B, CDKN2A, CDKN2B, CDKN2C, CEBPA, CHEK1, CHEK2, CIC, CREBBP, CRKL, CSF1R, CSF3R, CTCF, CTNNA1, CTNNB1, CUL3, CUL4A, CXCR4, CYP17A1, DAXX, DDR1, DDR2, DIS3, DNMT3A, DOT1L, EED, EGFR, EMSY (C11orf30), EP300, EPHA3, EPHB1, EPHB4, ERBB2, ERBB3, ERBB4, ERCC4, ERG, ERRFI1, ESR1, ETV4, ETV5, ETV6, EWSR1, EZH2, EZR, FAM46C, FANCA, FANCC, FANCG, FANCL, FAS, FBXW7, FGF10, FGF12, FGF14, FGF19, FGF23, FGF3, FGF4, FGF6, FGFR1, FGFR2, FGFR3, FGFR4, FH, FLCN, FLT1, FLT3, FOXL2, FUBP1, GABRA6, GATA3, GATA4, GATA6, GID4 (C17orf39), GNA11, GNA13, GNAQ, GNAS, GRM3, GSK3B, H3F3A, HDAC1, HGF, HNF1A, HRAS, HSD3B1, ID3, IDH1, IDH2, IGF1R, IKBKE, IKZF1, INPP4B, IRF2, IRF4, IRS2, JAK1, JAK2, JAK3, JUN, KDM5A, KDM5C, KDM6A, KDR, KEAP1, KEL, KIT, KLHL6, KMT2A (MLL),KMT2D (MLL2), KRAS, LTK, LYN, MAF, MAP2K1, MAP2K2, MAP2K4, MAP3K1, MAP3K13, MAPK1, MCL1, MDM2, MDM4, MED12, MEF2B, MEN1, MERTK, MET, MITF, MKNK1, MLH1, MPL, MRE11A, MSH2, MSH3, MSH6, MST1R, MTAP, MTOR, MUTYH, MYB, MYC, MYCL, MYCN, MYD88, NBN, NF1, NF2, NFE2L2, NFKBIA, NKX2-1, NOTCH1, NOTCH2, NOTCH3, NPM1, NRAS, NT5C2, NTRK1, NTRK2, NTRK3, NUTM1, P2RY8, PALB2, PARK2, PARP1, PARP2, PARP3, PAX5, PBRM1, PDCD1, PDCD1LG2, PDGFRA, PDGFRB, PDK1, PIK3C2B, PIK3C2G, PIK3CA, PIK3CB, PIK3R1, PIM1, PMS2, POLD1, POLE, PPARG, PPP2R1A, PPP2R2A, PRDM1, PRKAR1A, PRKCI, PTCH1, PTEN, PTPN11, PTPRO, QKI, RAC1, RAD21, RAD51, RAD51B, RAD51C, RAD51D, RAD52, RAD54L, RAF1, RARA, RB1, RBM10, REL, RET, RICTOR, RNF43, ROS1, RPTOR, RSPO2, SDC4, SDHA, SDHB, SDHC, SDHD, SETD2, SF3B1, SGK1, SLC34A2, SMAD2, SMAD4, SMARCA4, SMARCB1, SMO, SNCAIP, SOCS1, SOX2, SOX9, SPEN, SPOP, SRC, STAG2, STAT3, STK11, SUFU, SYK, TBX3, TEK, TERC, TERT, TET2, TGFBR2, TIPARP, TMPRSS2, TNFAIP3, TNFRSF14, TP53, TSC1, TSC2, TYRO3, U2AF1, VEGFA, VHL, WHSC1, WHSC1L1, WT1, XPO1, XRCC2, ZNF217, ZNF703, or any combination thereof.,
[0021] In some embodiments, one or more loci include ABL, ALK, ALL, B4GALNT1, BAFF, BCL2, BRAF, BRCA, BTK, CD19, CD20, CD3, CD30, CD319, CD38, CD52, CDK4, CDK6, CML, CRACC, CS1, CTLA-4, dMMR, EGFR, ERBB1, ERBB2, FGFR1-3, FLT3, GD2, HDAC, HER1, HER2, HR, IDH2, IL-1β, IL-6, IL-6R, JAK1, JAK2, JAK3, KIT, KRAS, MEK, MET, MSI-H, mTOR, PARP, PD-1, PDGFR, PDGFRα, PDGFRβ, PD-L1, PI3Kδ, PIGF, PTCH, RAF, RANKL, RET, ROS1, SLAMF7, VEGF, VEGFA, VEGFB, or any combination thereof.
[0022] In some embodiments, the method further includes generating, by one or more processors, a report indicating the determined tumor DNA fraction of the sample. In some embodiments, the method further includes transmitting the report to a healthcare provider. In some embodiments, the report is transmitted via a computer network or a peer-to-peer connection.
[0023] Disclosed herein is a method for determining a tumor DNA fraction in a cell-free DNA sample from a subject, the method comprising receiving, in one or more processors, sequence read data in a plurality of sequence reads derived from a cell-free DNA (cfDNA) sample from the subject; using one or more processors to determine, based on the sequence read data, the variant allele frequency (VAF) in one or more variants detected in the cfDNA sample; using one or more processors to generate an empirical distribution of tumor DNA fraction values in response to the determined VAFs in the one or more variants; using one or more processors to fit a model to the empirical distribution of tumor DNA fraction values; and determining, based on the model, the tumor DNA fraction in the cfDNA sample.
[0024] In some embodiments, the method further includes determining a confidence interval in the tumor DNA fraction based on the model. In some embodiments, one or more variants include one or more short variants. In some embodiments, one or more short variants include one or more somatic short variants. In some embodiments, one or more of the somatic short variants are known not to be associated with clonal hematopoiesis of indeterminate potential (CHIP).
[0025] In some embodiments, generating an empirical distribution of tumor DNA fraction values involves calculating the tumor DNA fraction values based on the known copy number in one or more variants, the determined VAF in one or more variants, and the corresponding known tumor average ploidy in a plurality of historical subject samples having known VAFs in one or more variants that are substantially the same as the determined VAF in one or more variants. In some embodiments, generating an empirical distribution of tumor DNA fraction values involves pre-calculating the tumor DNA fraction values based on the known copy number in one or more variants and the corresponding known tumor average ploidy in a plurality of historical subject samples having a range of VAF values in one or more variants, and selecting a subset of the pre-calculated tumor DNA fraction values corresponding to samples having known VAFs in one or more variants that are substantially the same as the determined VAF in one or more variants. In some embodiments, the tumor DNA fraction value is calculated or selected with respect to the variant showing the highest VAF in the cfDNA sample from the subject. In some embodiments, the tumor DNA fraction value is calculated or selected with respect to the ranked set of two or more variants showing the highest ranked VAFs in the cfDNA sample from the subject. In some embodiments, the tumor DNA fraction value is calculated or selected with respect to a predetermined set of two or more variants detected in the cfDNA sample from the subject. In some embodiments, the tumor DNA fraction value is calculated or selected with respect to a predetermined set of two or more variants detected in the cfDNA sample from the subject that includes known driver mutations. In some embodiments, the tumor DNA fraction value is calculated or selected with respect to all variants detected in the cfDNA sample from the subject.
[0026] In some embodiments, the tumor DNA fraction value is calculated or pre-calculated based on the known VAF in one or more variants, the known copy number in one or more variants, and the corresponding known tumor average ploidy in a plurality of historical subject samples.
[0027] In some embodiments, the tumor DNA fraction value is based on the known VAF in one or more variants, the known copy number in one or more variants, and the corresponding known tumor average ploidy in a plurality of historical subject samples, and by solving a set of equations representing (i) the relationship among tumor DNA fraction, tumor purity, and tumor average ploidy, and (ii) the relationship among somatic VAF, tumor purity, copy number at the genomic position of one or more variants, and the number of variant alleles in each of the one or more variants, thereby deriving the relationship in tumor DNA fraction as a function of somatic VAF, tumor average ploidy, copy number at the genomic position of one or more variants, and the number of variant alleles for one or more variants, excluding tumor purity.
[0028] In some embodiments, the tumor DNA fraction value is calculated or pre-calculated by solving a set of equations including a first equation that equates the tumor DNA fraction to the product of tumor purity and tumor average ploidy divided by the sum of the product of tumor purity and tumor average ploidy and the product of 2 and the magnitude obtained by subtracting tumor purity from 1, and a second equation that equates the somatic VAF to the product of tumor purity and the number of variant alleles in each of the one or more variants divided by the sum of the product of tumor purity and copy number at the genomic position of the one or more variants and the product of 2 and the magnitude obtained by subtracting tumor purity from 1, thereby excluding tumor purity and deriving a relationship that equates the tumor DNA fraction to the magnitude equal to dividing the tumor average ploidy by the sum of the ratio of the number of variant alleles in the one or more variants and the somatic VAF in each of the one or more variants to the magnitude obtained by subtracting the copy number at the genomic position of the one or more variants from the tumor average ploidy.
[0029] In some embodiments, the tumor DNA fraction value is based on the known VAF in one or more variants, the known copy number in one or more variants, and the corresponding known tumor average ploidy in a plurality of historical subject samples, and is given by the following set of equations, namely, is calculated or pre-calculated by solving TIFF2025524649000002.tif30170, thereby eliminating ρ, TIFF2025524649000003.tif20170where ρ is the tumor purity, ψ is the tumor average ploidy, C is the copy number at the genomic position of one or more variants, and V is the number of variant alleles in each of one or more variants.
[0030] In some embodiments, the plurality of historical subject samples include solid biopsy samples, liquid biopsy samples, or any combination thereof. In some embodiments, the plurality of historical subject samples include cancer samples. In some embodiments, the plurality of historical subject samples include samples in a single type of cancer. In some embodiments, the plurality of historical subject samples include samples in multiple types of cancer.
[0031] In some embodiments, the plurality of historical subject samples include bladder cancer samples, breast cancer samples, colorectal cancer samples, endometrial cancer samples, kidney cancer samples, leukemia samples, liver cancer samples, lung cancer samples, melanoma samples, non-Hodgkin lymphoma samples, pancreatic cancer samples, prostate cancer samples, thyroid cancer samples, or any combination thereof.
[0032] In some embodiments, the plurality of historical subject samples are acute lymphoblastic leukemia (Philadelphia chromosome positive) samples, acute lymphoblastic leukemia (precursor B cell) samples, acute myeloid leukemia (FLT3+) samples, acute myeloid leukemia (with IDH2 mutation) samples, anaplastic large cell lymphoma samples, basal cell carcinoma samples, B-cell chronic lymphocytic leukemia samples, bladder cancer samples, breast cancer (HER2 overexpression / amplification) samples, breast cancer (HER2+) samples, breast cancer (HR+, HER2-) samples, cervical cancer samples, cholangiocarcinoma samples, chronic lymphocytic leukemia samples, chronic lymphocytic leukemia (with 17p deletion) samples, chronic myelogenous leukemia samples, chronic myelogenous leukemia (Philadelphia chromosome positive) samples, classical Hodgkin lymphoma samples, colorectal cancer samples, colorectal cancer (dMMR and MSI-H) samples, colorectal cancer (KRAS wild type) samples, cryopyrin-associated periodic syndrome samples, cutaneous T-cell lymphoma samples, dermatofibrosarcoma protuberans samples, diffuse large B-cell lymphoma samples, fallopian tube cancer samples, follicular B-cell non-Hodgkin lymphoma samples, follicular lymphoma samples, gastric cancer samples, gastric cancer (HER2+) samples, gastroesophageal junction (GEJ) adenocarcinoma samples, gastrointestinal stromal tumor samples, gastrointestinal stromal tumor samples, gastrointestinal stromal tumor (KIT+) samples, giant cell tumor of bone samples, glioblastoma samples, polyangiitis granulomatosa samples, head and neck squamous cell carcinoma samples, hepatocellular carcinoma samples, Hodgkin lymphoma samples, juvenile idiopathic arthritis samples, lupus erythematosus samples, mantle cell lymphoma samples, medullary thyroid cancer samples, melanoma samples, melanoma with BRAF V600 mutation samples, melanoma with BRAF V600E or V600K mutation samples, Merkel cell carcinoma samples, Castleman disease samples, multiple blood samples including Philadelphia chromosome positive ALL and CML, multiple myeloma samples, myelofibrosis samples, non-Hodgkin lymphoma samples, inoperable subependymal giant cell astrocytoma samples associated with tuberous sclerosis, non-small cell lung cancer samples, non-small cell lung cancer (ALK+) samples, non-small cell lung cancer (PD-L1+) samples, non-small cell lung cancer (with ALK fusion or ROS1 gene alteration) samples, non-small cell lung cancer (with BRAF V600E mutation) samples, non-small cell lung cancer samples (with EGFR exon 19 deletion or exon 21 substitution (L858R) mutation), non-small cell lung cancer samples (EGFR(with T790M mutation), ovarian cancer sample, ovarian cancer sample (with BRCA mutation), pancreatic cancer sample, pancreatic cancer sample, gastrointestinal cancer sample, lung-origin neuroendocrine tumor sample, pediatric neuroblastoma sample, peripheral T-cell lymphoma sample, prostate cancer sample, peritoneal cancer sample, renal cell cancer sample, rheumatoid arthritis sample, small lymphocyte lymphoma sample, soft tissue sarcoma sample, solid tumor (MSI-H / dMMR) sample, head and neck squamous cell carcinoma sample, squamous non-small cell lung cancer sample, thyroid cancer sample, thyroid cancer sample, urothelial cancer sample, urothelial cancer sample, Waldenström macroglobulinemia sample, or any combination thereof.
[0033] In some embodiments, the model is a parametric probability density model. In some embodiments, the model is a non-parametric probability density model. In some embodiments, the determined tumor DNA fraction in the sample is the most likely tumor DNA fraction. In some embodiments, the determined tumor DNA fraction in the cfDNA sample is the mean, median, or mode of the dominant peak in the empirical distribution of the tumor DNA fraction values.
[0034] In some embodiments, the cfDNA sample comprises DNA extracted from a blood sample, plasma sample, cerebrospinal fluid sample, pleural fluid sample, sputum sample, fecal sample, urine sample, or saliva sample.
[0035] Disclosed herein is a method for determining the tumor DNA fraction in a sample from a subject, the method comprising, in one or more processors, receiving sequence read data in a plurality of sequence reads derived from a sample from the subject, using one or more processors to determine, based on the sequence read data, the variant allele frequency (VAF) in one or more variants detected in the sample, using one or more processors to generate an empirical distribution of tumor DNA fraction values in response to the determined VAF in the one or more variants, using one or more processors to fit a model to the empirical distribution of tumor DNA fraction values, and determining the tumor DNA fraction in the sample based on the model.
[0036] In some embodiments, the method further includes determining a confidence interval in the tumor DNA fraction based on the model. In some embodiments, one or more variants include one or more short variants. In some embodiments, one or more short variants include one or more somatic short variants. In some embodiments, one or more of the somatic short variants are known not to be associated with clonal hematopoiesis of indeterminate potential (CHIP).
[0037] In some embodiments, generating an empirical distribution of tumor DNA fraction values involves calculating a tumor DNA fraction value based on a known copy number in one or more variants, a determined VAF in one or more variants, and a corresponding known tumor average ploidy in a plurality of historical subject samples having a known VAF in one or more variants that is substantially the same as the determined VAF in one or more variants. In some embodiments, generating an empirical distribution of tumor DNA fraction values involves pre-calculating a tumor DNA fraction value based on a known copy number in one or more variants and a corresponding known tumor average ploidy in a plurality of historical subject samples having a range of VAF values in one or more variants, and selecting a subset of the pre-calculated tumor DNA fraction values corresponding to samples having a known VAF in one or more variants that is substantially the same as the determined VAF in one or more variants. In some embodiments, the tumor DNA fraction value is calculated or selected with respect to a variant showing the highest VAF in a sample from a subject. In some embodiments, the tumor DNA fraction value is calculated or selected with respect to a ranked set of two or more variants showing the highest ranked VAFs in a sample from a subject. In some embodiments, the tumor DNA fraction value is calculated or selected with respect to a predetermined set of two or more variants detected in a sample from a subject. In some embodiments, the tumor DNA fraction value is calculated or selected with respect to a predetermined set of two or more variants detected in a sample from a subject that includes a known driver mutation. In some embodiments, the tumor DNA fraction value is calculated or selected with respect to all variants detected in a cfDNA sample from a subject.
[0038] In some embodiments, the tumor DNA fraction value is based on the known VAF in one or more variants, the known copy number in one or more variants, and the corresponding known tumor average ploidy in a plurality of historical subject samples, and is calculated or pre-calculated by solving a set of equations representing (i) the relationship between tumor DNA fraction, tumor purity, and tumor average ploidy, and (ii) the relationship between somatic VAF, tumor purity, copy number at the genomic position of one or more variants, and the number of variant alleles in each of the one or more variants, thereby deriving the relationship in tumor DNA fraction as a function of somatic VAF, tumor average ploidy, copy number at the genomic position of one or more variants, and the number of variant alleles for the one or more variants, excluding tumor purity.
[0039] In some embodiments, the tumor DNA fraction value is calculated or pre-calculated by solving a set of equations including a first equation that equates the tumor DNA fraction to the product of tumor purity and tumor average ploidy divided by the sum of the product of tumor purity and tumor average ploidy and the product of 2 and the magnitude obtained by subtracting tumor purity from 1, and a second equation that equates the somatic VAF to the product of tumor purity and the number of variant alleles in each of the one or more variants divided by the sum of the product of tumor purity and copy number at the genomic position of the one or more variants and the product of 2 and the magnitude obtained by subtracting tumor purity from 1, thereby deriving a relationship that equates the tumor DNA fraction to the magnitude obtained by dividing the tumor average ploidy by the sum of the ratio of the number of variant alleles in the one or more variants and the somatic VAF in each of the one or more variants to the magnitude obtained by subtracting the copy number at the genomic position of the one or more variants from the tumor average ploidy, excluding tumor purity.
[0040] In some embodiments, the tumor DNA fraction value is based on the known VAF in one or more variants, the known copy number in one or more variants, and the corresponding known tumor average ploidy in a plurality of historical subject samples, and is calculated or pre-calculated by solving the following set of equations, namely, by solving TIFF2025524649000004.tif30170, thereby eliminating ρ, TIFF2025524649000005.tif20170where ρ is the tumor purity, ψ is the tumor average ploidy, C is the copy number at the genomic position of one or more variants, and V is the number of variant alleles in each of one or more variants.
[0041] In some embodiments, the plurality of historical subject samples include solid biopsy samples, liquid biopsy samples, or any combination thereof. In some embodiments, the plurality of historical subject samples include cancer samples. In some embodiments, the plurality of historical subject samples include samples in a single type of cancer. In some embodiments, the plurality of historical subject samples include samples in multiple types of cancer.
[0042] In some embodiments, the plurality of historical subject samples include bladder cancer samples, breast cancer samples, colorectal cancer samples, endometrial cancer samples, kidney cancer samples, leukemia samples, liver cancer samples, lung cancer samples, melanoma samples, non-Hodgkin lymphoma samples, pancreatic cancer samples, prostate cancer samples, thyroid cancer samples, or any combination thereof.
[0043] In some embodiments, the plurality of historical subject samples are acute lymphoblastic leukemia (Philadelphia chromosome positive) samples, acute lymphoblastic leukemia (precursor B cell) samples, acute myeloid leukemia (FLT3+) samples, acute myeloid leukemia (with IDH2 mutation) samples, anaplastic large cell lymphoma samples, basal cell carcinoma samples, B-cell chronic lymphocytic leukemia samples, bladder cancer samples, breast cancer (HER2 overexpression / amplification) samples, breast cancer (HER2+) samples, breast cancer (HR+, HER2-) samples, cervical cancer samples, cholangiocarcinoma samples, chronic lymphocytic leukemia samples, chronic lymphocytic leukemia (with 17p deletion) samples, chronic myelogenous leukemia samples, chronic myelogenous leukemia (Philadelphia chromosome positive) samples, classical Hodgkin lymphoma samples, colorectal cancer samples, colorectal cancer (dMMR and MSI-H) samples, colorectal cancer (KRAS wild type) samples, cryopyrin-associated periodic syndrome samples, cutaneous T-cell lymphoma samples, dermatofibrosarcoma protuberans samples, diffuse large B-cell lymphoma samples, fallopian tube cancer samples, follicular B-cell non-Hodgkin lymphoma samples, follicular lymphoma samples, gastric cancer samples, gastric cancer (HER2+) samples, gastroesophageal junction (GEJ) adenocarcinoma samples, gastrointestinal stromal tumor samples, gastrointestinal stromal tumor samples, gastrointestinal stromal tumor (KIT+) samples, giant cell tumor of bone samples, glioblastoma samples, polyangiitis granulomatosa samples, head and neck squamous cell carcinoma samples, hepatocellular carcinoma samples, Hodgkin lymphoma samples, juvenile idiopathic arthritis samples, lupus erythematosus samples, mantle cell lymphoma samples, medullary thyroid cancer samples, melanoma samples, melanoma with BRAF V600 mutation samples, melanoma with BRAF V600E or V600K mutation samples, Merkel cell carcinoma samples, Castleman disease samples, multiple blood samples including Philadelphia chromosome positive ALL and CML, multiple myeloma samples, myelofibrosis samples, non-Hodgkin lymphoma samples, inoperable subependymal giant cell astrocytoma samples associated with tuberous sclerosis, non-small cell lung cancer samples, non-small cell lung cancer (ALK+) samples, non-small cell lung cancer (PD-L1+) samples, non-small cell lung cancer (with ALK fusion or ROS1 gene alteration) samples, non-small cell lung cancer (with BRAF V600E mutation) samples, non-small cell lung cancer samples (with EGFR exon 19 deletion or exon 21 substitution (L858R) mutation), non-small cell lung cancer samples (EGFRincluding ovarian cancer samples (with T790M mutations), ovarian cancer samples (with BRCA mutations), pancreatic cancer samples, pancreatic cancer samples, gastrointestinal cancer samples, lung-origin neuroendocrine tumor samples, pediatric neuroblastoma samples, peripheral T-cell lymphoma samples, prostate cancer samples, peritoneal cancer samples, renal cell cancer samples, rheumatoid arthritis samples, small lymphocyte lymphoma samples, soft tissue sarcoma samples, solid tumors (MSI-H / dMMR) samples, head and neck squamous cell carcinoma samples, squamous non-small cell lung cancer samples, thyroid cancer samples, thyroid cancer samples, urothelial cancer samples, urothelial cancer samples, Waldenström macroglobulinemia samples, or any combination thereof.
[0044] In some embodiments, the model is a parametric probability density model. In some embodiments, the model is a non-parametric probability density model. In some embodiments, the determined tumor DNA fraction in the sample is the most likely tumor DNA fraction. In some embodiments, the determined tumor DNA fraction in the cfDNA sample is the mean, median, or mode of the dominant peak in the empirical distribution of tumor DNA fraction values.
[0045] In some embodiments, the sample comprises DNA extracted from a blood sample, plasma sample, cerebrospinal fluid sample, pleural fluid sample, sputum sample, fecal sample, urine sample or saliva sample.
[0046] In some embodiments, the determination of the tumor DNA fraction is used for the diagnosis or confirmation of a diagnosis of a disease in a subject. In some embodiments, the disease is cancer. In some embodiments, the cancer is acute lymphoblastic leukemia (Philadelphia chromosome positive), acute lymphoblastic leukemia (precursor B cell), acute myeloid leukemia (FLT3+), acute myeloid leukemia (with IDH2 mutation), anaplastic large cell lymphoma, basal cell carcinoma, B-cell chronic lymphocytic leukemia, bladder cancer, breast cancer (HER2 overexpression / amplification), breast cancer (HER2+), breast cancer (HR+, HER2-), cervical cancer, cholangiocarcinoma, chronic lymphocytic leukemia, chronic lymphocytic leukemia (with 17p deletion), chronic myelogenous leukemia, chronic myelogenous leukemia (Philadelphia chromosome positive), classical Hodgkin lymphoma, colorectal cancer, colorectal cancer (dMMR / MSI-H), colorectal cancer (KRAS wild type), cryopyrin-associated periodic syndrome, cutaneous T-cell lymphoma, precursor cutaneous fibrosarcoma, diffuse large B-cell lymphoma, fallopian tube cancer, follicular B-cell non-Hodgkin lymphoma, follicular lymphoma, gastric cancer, gastric cancer (HER2+), gastroesophageal junction (GEJ) adenocarcinoma, gastrointestinal stromal tumor, gastrointestinal stromal tumor (KIT+), giant cell tumor of bone, glioblastoma, polyangiitis nodosa, head and neck squamous cell carcinoma, hepatocellular carcinoma, Hodgkin lymphoma, juvenile idiopathic arthritis, lupus erythematosus, mantle cell lymphoma, medullary thyroid carcinoma, melanoma, melanoma with BRAF V600 mutation, melanoma with BRAF V600E or V600K mutation, Merkel cell carcinoma, multicentric Castleman disease, multiple hematologic malignancies including Philadelphia chromosome positive ALL and CML, multiple myeloma, myelofibrosis, non-Hodgkin lymphoma, unresectable subependymal giant cell astrocytoma associated with tuberous sclerosis, non-small cell lung cancer, non-small cell lung cancer (ALK+), non-small cell lung cancer (PD-L1+), non-small cell lung cancer (with ALK fusion or ROS1 gene alteration), non-small cell lung cancer (with BRAF V600E mutation), non-small cell lung cancer (with EGFR exon 19 deletion or exon 21 substitution (L858R) mutation), non-small cell lung cancer (EGFR(with the T790M mutation), ovarian cancer, ovarian cancer (with BRCA mutation), pancreatic cancer, pancreatic cancer, gastrointestinal cancer or lung-derived neuroendocrine tumor, pediatric neuroblastoma, peripheral T-cell lymphoma, peritoneal cancer, prostate cancer, renal cell cancer, rheumatoid arthritis, small lymphocyte lymphoma, soft tissue sarcoma, solid tumor (MSI-H / dMMR), head and neck squamous cell carcinoma, non-small cell lung cancer, squamous cell carcinoma, thyroid cancer, thyroid cancer, urothelial cancer, urothelial cancer, or Waldenström macroglobulinemia.
[0047] In some embodiments, the method further comprises selecting an anti-cancer therapy for administration to the subject based on determination of the tumor DNA fraction. In some embodiments, the method further comprises determining an effective amount of an anti-cancer therapy for administration to the subject based on determination of the tumor DNA fraction. In some embodiments, the method further comprises administering an anti-cancer therapy to the subject based on determination of the tumor DNA fraction. In some embodiments, the anti-cancer therapy includes chemotherapy, radiation therapy, immunotherapy, targeted therapy, or surgery.
[0048] Disclosed herein is a method for diagnosing a disease, the method comprising diagnosing that a subject has the disease based on determination of a tumor DNA fraction in a sample from the subject, wherein the tumor DNA fraction is determined according to any of the methods described herein.
[0049] Disclosed herein is a method for selecting an anti-cancer therapy, the method comprising selecting an anti-cancer therapy for a subject in response to determining a tumor DNA fraction in a sample from the subject, wherein the tumor DNA fraction is determined according to any of the methods described herein.
[0050] Disclosed herein is a method for treating cancer in a subject, the method comprising administering an effective amount of an anti-cancer therapy to the subject in response to determining a tumor DNA fraction in a sample from the subject, wherein the tumor DNA fraction is determined according to any of the methods described herein.
[0051] Disclosed herein is a method for determining the prognosis of a subject having cancer, the method comprising determining a tumor DNA fraction in a sample from the subject and determining the prognosis of the subject based on the tumor DNA fraction, wherein the tumor DNA fraction is determined according to any of the methods described herein.
[0052] Disclosed herein is a method for assessing minimal residual disease (MRD), the method comprising determining a tumor DNA fraction in a sample from the subject and assessing minimal residual disease (MRD) in the subject based on the tumor DNA fraction, wherein the tumor DNA fraction is determined according to any of the methods described herein.
[0053] Disclosed herein is a method for monitoring the progression or recurrence of cancer in a subject, the method comprising determining a first tumor DNA fraction in a first sample obtained from the subject at a first time point according to any of the methods described herein, determining a second tumor DNA fraction in a second sample obtained from the subject at a second time point, and comparing the first tumor DNA fraction with the second tumor DNA fraction to thereby monitor the progression or recurrence of cancer. In some embodiments, the second tumor DNA fraction in the second sample is determined according to any of the methods described herein. In some embodiments, the method further comprises selecting an anti-cancer therapy for the subject in response to cancer progression. In some embodiments, the method further comprises administering an anti-cancer therapy to the subject in response to cancer progression. In some embodiments, the method further comprises adjusting an anti-cancer therapy for the subject in response to cancer progression. In some embodiments, the method further comprises adjusting the dosage of the anti-cancer therapy or selecting a different anti-cancer therapy in response to the progression of cancer. In some embodiments, the method further comprises administering the adjusted anti-cancer treatment to the subject. In some embodiments, the first time point is before the subject is administered an anti-cancer treatment, and the second time point is after the subject is administered an anti-cancer treatment. In some embodiments, the subject has cancer, is at risk of having cancer, is routinely screened for cancer, or is suspected of having cancer. In some embodiments, the cancer is a solid tumor. In some embodiments, the cancer is a blood cancer. In some embodiments, the anti-cancer therapy comprises chemotherapy, radiation therapy, immunotherapy, targeted therapy, or surgery.
[0054] In some embodiments, the method further comprises determining, identifying, or applying a value of the tumor DNA fraction in the sample as a diagnostic value associated with the sample. In some embodiments, the method further comprises generating a genomic profile for a subject based on the determination of the tumor DNA fraction. In some embodiments, the genomic profile of the subject further comprises results from a comprehensive genomic profiling (CGP) test, a gene expression profiling test, a cancer hot spot panel test, a DNA methylation test, a DNA fractionation test, an RNA fractionation test, or any combination thereof. In some embodiments, the genomic profile of the subject further comprises results from a test based on nucleic acid sequencing. In some embodiments, the method further comprises selecting, administering, or applying an anti-cancer therapy to the subject based on the generated genomic profile. In some embodiments, the determination of the tumor DNA fraction in the sample is used in making a treatment decision proposed for the subject. In some embodiments, the determination of the tumor DNA fraction in the sample is used in applying or administering treatment to the subject.
[0055] Disclosed herein is a system comprising one or more processors and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to receive sequence read data in a plurality of sequence reads derived from a sample of a subject, determine a variant allele frequency (VAF) in one or more variants detected in the sample based on the sequence read data, generate an empirical distribution of tumor DNA fraction values in response to the determined VAFs in the one or more variants, fit a model to the empirical distribution of tumor DNA fraction values, and determine the tumor DNA fraction in the sample based on the model.
[0056] In some embodiments, the system further includes instructions for determining a confidence interval in the tumor DNA fraction based on a model. In some embodiments, one or more variants include one or more short variants. In some embodiments, one or more short variants include one or more somatic short variants. In some embodiments, one or more somatic short variants are known not to be associated with clonal hematopoiesis of indeterminate potential (CHIP).
[0057] In some embodiments, generating an empirical distribution of tumor DNA fraction values includes calculating the tumor DNA fraction values based on known copy numbers in one or more variants, determined VAFs in one or more variants, and corresponding known tumor average ploidies in a plurality of historical subject samples having known VAFs in one or more variants that are substantially the same as the determined VAFs in one or more variants. In some embodiments, generating an empirical distribution of tumor DNA fraction values includes pre-calculating the tumor DNA fraction values based on known copy numbers in one or more variants and corresponding known tumor average ploidies in a plurality of historical subject samples having a range of VAF values in one or more variants, and selecting a subset of the pre-calculated tumor DNA fraction values corresponding to samples having known VAFs in one or more variants that are substantially the same as the determined VAFs in one or more variants. In some embodiments, the tumor DNA fraction values are calculated or selected with respect to the variant showing the highest VAF in a sample from a subject. In some embodiments, the tumor DNA fraction values are calculated or selected with respect to a ranked set of two or more variants showing the highest ranked VAFs in a sample from a subject. In some embodiments, the tumor DNA fraction values are calculated or selected with respect to a predetermined set of two or more variants detected in a sample from a subject. In some embodiments, the tumor DNA fraction values are calculated or selected with respect to a predetermined set of two or more variants detected in a sample from a subject that includes known driver mutations. In some embodiments, the tumor DNA fraction values are calculated or selected with respect to all variants detected in a cfDNA sample from a subject.
[0058] In some embodiments, the tumor DNA fraction value is calculated or pre-calculated by solving a set of equations representing (i) the relationship between tumor DNA fraction, tumor purity, and tumor average ploidy, and (ii) the relationship between somatic VAF, tumor purity, copy number at the genomic position of one or more variants, and the number of variant alleles in each of the one or more variants, based on the known VAF in one or more variants, the known copy number in one or more variants, and the corresponding known tumor average ploidy in a plurality of historical subject samples, thereby deriving the relationship in tumor DNA fraction as a function of somatic VAF, tumor average ploidy, copy number at the genomic position of one or more variants, and the number of variant alleles for one or more variants, excluding tumor purity.
[0059] In some embodiments, the model is a parametric probability density model. In some embodiments, the model is a non-parametric probability density model. In some embodiments, the determined tumor DNA fraction in a sample is the most likely tumor DNA fraction. In some embodiments, the determined tumor DNA fraction in a cfDNA sample is the mean, median, or mode of the dominant peak in the empirical distribution of tumor DNA fraction values.
[0060] Disclosed herein is a non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a system, cause the system to receive sequence read data in a plurality of sequence reads derived from a sample from a subject, determine a variant allele frequency (VAF) in one or more variants detected in the sample based on the sequence read data, generate an empirical distribution of tumor DNA fraction values according to the determined VAF in the one or more variants, fit a model to the empirical distribution of tumor DNA fraction values, and determine a tumor DNA fraction in the sample based on the model.
[0061] In some embodiments, the non-transitory computer-readable storage medium further includes instructions for determining a confidence interval in a tumor DNA fraction based on a model. In some embodiments, one or more variants include one or more short variants. In some embodiments, one or more short variants include one or more somatic short variants. In some embodiments, one or more of the somatic short variants are known not to be associated with clonal hematopoiesis of indeterminate potential (CHIP).
[0062] In some embodiments, generating an empirical distribution of tumor DNA fraction values involves calculating a tumor DNA fraction value based on the known copy number in one or more variants, the determined VAF in one or more variants, and the corresponding known tumor average ploidy in a plurality of historical subject samples having a known VAF in one or more variants that is substantially the same as the determined VAF in one or more variants. In some embodiments, generating an empirical distribution of tumor DNA fraction values involves pre-calculating a tumor DNA fraction value based on the known copy number in one or more variants and the corresponding known tumor average ploidy in a plurality of historical subject samples having a range of VAF values in one or more variants, and selecting a subset of the pre-calculated tumor DNA fraction values corresponding to samples having a known VAF in one or more variants that is substantially the same as the determined VAF in one or more variants. In some embodiments, the tumor DNA fraction value is calculated or selected with respect to the variant showing the highest VAF in a sample from the subject. In some embodiments, the tumor DNA fraction value is calculated or selected with respect to a ranked set of two or more variants showing the highest ranked VAFs in a sample from the subject. In some embodiments, the tumor DNA fraction value is calculated or selected with respect to a predetermined set of two or more variants detected in a sample from the subject. In some embodiments, the tumor DNA fraction value is calculated or selected with respect to a predetermined set of two or more variants detected in a sample from a subject comprising a known driver mutation. In some embodiments, the tumor DNA fraction value is calculated or selected with respect to all variants detected in a cfDNA sample from the subject.
[0063] In some embodiments, the tumor DNA fraction value is calculated or pre-calculated by solving a set of equations representing (i) the relationship between tumor DNA fraction, tumor purity, and tumor average ploidy, and (ii) the relationship between somatic VAF, tumor purity, copy number at the genomic location of one or more variants, and the number of variant alleles at each of the one or more variants, based on the known VAF in one or more variants, the known copy number in one or more variants, and the corresponding known tumor average ploidy in a plurality of historical subject samples, thereby deriving the relationship in tumor DNA fraction as a function of somatic VAF, tumor average ploidy, copy number at the genomic location of one or more variants, and the number of variant alleles for one or more variants, excluding tumor purity.
[0064] In some embodiments, the model is a parametric probability density model. In some embodiments, the model is a non-parametric probability density model. In some embodiments, the determined tumor DNA fraction in a sample is the most likely tumor DNA fraction. In some embodiments, the determined tumor DNA fraction in a cfDNA sample is the mean, median, or mode of the dominant peak in the empirical distribution of the tumor DNA fraction values.
[0065] Incorporation by reference All publications, patents, and patent applications mentioned in this specification are hereby incorporated by reference in their entirety, to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference in its entirety. In case of conflict between the terms of this specification and the incorporated references, the terms of this specification shall govern. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] The various aspects of the disclosed methods, devices, and systems are described in detail in the appended claims. A better understanding of the features and advantages of the disclosed methods, devices, and systems will be obtained by reference to the following detailed description of the exemplary embodiments and the accompanying drawings.
[0067]
Figure 1
[0068]
Figure 2
[0069]
Figure 3
[0070]
Figure 4
[0071]
Figure 5
[0072]
Figure 6
[0073]
Figure 7
[0074]
Figure 8
[0075] Detailed Description This specification discloses methods and systems that can generally provide a more accurate assessment of the tumor DNA content in a biopsy sample, particularly a more accurate assessment of the circulating tumor DNA content in a cell-free DNA sample collected from a patient's plasma or other liquid biopsy sample. The methods described herein overcome the problems encountered by existing tumor aneuploidy-based methods for estimating the tumor DNA fraction when no detectable aneuploidy signal is present, by leveraging alternative signals, such as signals associated with the presence of somatic short variants, to infer the tumor DNA fraction (e.g., the ctDNA fraction). This approach is based on the following: (i) the derived physical relationship between the allele frequency of a variant (e.g., a somatic short variant) and the tumor content determined by the copy number state at the genomic position of the variant and the tumor average ploidy, and (ii) access to a broad database collection of patient genomic profile data.
[0076] Briefly, variants (e.g., somatic short variants) are identified in a sample (e.g., a biopsy sample, a liquid biopsy sample, or a plasma sample) based on sequence read data derived from the sample, and their allele frequencies are quantified. Next, the empirical distribution of all possible sample tumor DNA fraction values considering the observed somatic short variant allele frequencies can be generated by (i) using an equation that describes the physical relationship between the somatic variant allele frequency and the tumor DNA fraction for a given copy number state at the variant genomic position and a given tumor average ploidy, and (ii) leveraging historical patient data that includes known copy number and tumor ploidy profiles. Then, a model, such as a probability density model (e.g., a non-parametric probability density model), can be fit to the empirical distribution of the tumor DNA fraction values, and an estimated value of the tumor DNA fraction in the sample can be derived from the model along with the upper and lower bounds of the estimated values based on the desired confidence interval.
[0077] In some cases, for example, receiving array read data in a plurality of array reads derived from a sample from a subject, determining a variant allele frequency (VAF) in one or more variants detected in the sample based on the array read data, generating an empirical distribution of tumor DNA fraction values according to the determined VAF in the one or more variants, fitting a model to the empirical distribution of tumor DNA fraction values, and determining a tumor DNA fraction in the sample based on the model are described. In some cases, the disclosed method is a computer-implemented method.
[0078] In some cases, the method further includes determining a confidence interval of the tumor DNA fraction based on the model. In some cases, the one or more variants include one or more short variants (e.g., one or more somatic short variants). In some cases, the one or more short variants are known not to be associated with clonal hematopoiesis of indeterminate potential (CHIP).
[0079] In some cases, generating the empirical distribution of tumor DNA fraction values includes calculating tumor DNA fraction values based on the known copy number in the one or more variants, the determined VAF in the one or more variants, and the corresponding known tumor average ploidy in a plurality of historical subject samples having known VAFs in the one or more variants that are substantially the same as the determined VAF in the one or more variants. In some cases, generating the empirical distribution of tumor DNA fraction values includes pre-calculating tumor DNA fraction values based on the known copy number in the one or more variants and the corresponding known tumor average ploidy in a plurality of historical subject samples having a range of VAF values in the one or more variants, and selecting a subset of the pre-calculated tumor DNA fraction values corresponding to samples having known VAFs in the one or more variants that are substantially the same as the determined VAF in the one or more variants.
[0080] In some cases, the tumor DNA fraction value is calculated or selected for the variant showing the highest VAF in the sample from the subject. In some cases, the tumor DNA fraction value is calculated or selected for the set of ranks of two or more variants showing the highest ranked VAFs in the sample from the subject. In some cases, the tumor DNA fraction value is calculated or selected for a predetermined set of two or more variants detected in the sample from the subject. In some cases, the tumor DNA fraction value is calculated or selected for a predetermined set of two or more variants detected in the sample from the subject that includes known driver mutations. In some cases, the tumor DNA fraction value is calculated or selected for all variants detected in the cfDNA sample from the subject. In some cases, the tumor DNA fraction value is calculated or pre-calculated by solving a set of equations representing the relationships between (i) tumor DNA fraction, tumor purity, and tumor average ploidy, and (ii) somatic VAF, tumor purity, copy number at the genomic location of one or more variants, and the number of variant alleles at each of one or more variants, based on the known VAFs at one or more variants, the known copy numbers at one or more variants, and the corresponding known tumor average ploidies in a plurality of historical subject samples, thereby deriving the relationship in tumor DNA fraction as a function of somatic VAF, tumor average ploidy, copy number at the genomic location of one or more variants, and the number of variant alleles for one or more variants, excluding tumor purity.
[0081] In some cases, the model is a non-parametric probability density model. In some cases, the tumor DNA fraction determined for the sample is the most likely tumor DNA fraction. In some cases, the tumor DNA fraction determined for the cfDNA sample is the mean, median, or mode of the dominant peak in the empirical distribution of the tumor DNA fraction values.
[0082] The disclosed method achieves a more accurate assessment of tumor DNA content by considering the effects of copy number and tumor average ploidy on variant allele frequency (VAF), and thus provides an estimate that more precisely reflects the true tumor DNA fraction of a sample.
[0083] Definitions Unless otherwise defined, all technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0084] As used in this specification and the appended claims, the singular forms "a", "an", and "the" include plural references unless the context clearly dictates otherwise. Any reference to "or" herein is intended to include "and / or" unless specifically stated otherwise.
[0085] "About" and "approximately" generally mean an acceptable degree of error of the measured quantity, taking into account the nature or precision of the measurement. Exemplary degrees of error are within 20 percent (%) of a given value or range of values, typically within 10%, more typically within 5%.
[0086] As used herein, the terms "comprising" (and any form or variation of comprising such as "comprise" and "comprises"), "having" (and any form or variation of having such as "have" and "has"), "including" (and any form or variation including such as "includes" and "include"), or "containing" (and any form or variation of containing such as "contains" and "contain") are inclusive or open-ended and do not exclude additional, unrecited additives, components, integers, elements, or method steps.
[0087] As used herein, the terms "individual," "patient," or "subject" are used interchangeably and refer to any single animal for which treatment is desired, e.g., a mammal (including non-human animals such as dogs, cats, horses, rabbits, zoo animals, cows, pigs, sheep, and non-human primates). In certain embodiments, the individual, patient, or subject herein is a human.
[0088] The terms "cancer" and "tumor" are used interchangeably herein. These terms refer to the presence of cells having characteristics typical of cells that cause cancer, such as uncontrolled proliferation, immortality, metastatic potential, rapid growth and proliferation rate, and certain characteristic morphological features. Cancer cells are often in the form of a tumor, but such cells can exist alone within an animal or can be non-tumorigenic cancer cells such as leukemia cells. These terms include solid tumors, soft tissue tumors, or metastatic lesions. As used herein, the term "cancer" includes pre-cancerous as well as malignant cancers.
[0089] As used herein, "treatment" (and its grammatical variations such as "treat" or "treating") refers to a clinical intervention in an attempt to alter the natural course of a treated individual (e.g., administration of an anti-cancer agent or anti-cancer agent treatment), which can be done for prophylaxis or during the course of clinical pathology. Desired effects of treatment include, but are not limited to, preventing the occurrence or recurrence of a disease, alleviating symptoms, reducing any direct or indirect pathological consequence of the disease, preventing metastasis, decreasing the rate of disease progression, improving or alleviating the disease state, and remission or improved prognosis.
[0090] As used herein, the term "subgenomic interval" (or "subgenomic sequence interval") refers to a portion of a genomic sequence.
[0091] As used herein, the term "target interval" refers to a subgenomic interval or an expressed subgenomic interval (e.g., the transcribed sequence of a subgenomic interval).
[0092] As used herein, the terms "variant sequence" or "variant" are used interchangeably and refer to a nucleic acid sequence that has been modified relative to the corresponding "normal" or "wild-type" sequence. In some cases, a variant sequence can be a "short variant sequence" (or "short variant"), i.e., a variant sequence that is less than about 50 base pairs in length.
[0093] The terms "allele frequency" and "allele fraction" are used interchangeably herein and refer to the fraction of sequence reads corresponding to a particular allele relative to the total number of sequence reads for a genomic locus.
[0094] The terms "variant allele frequency" and "variant allele fraction" are used interchangeably herein and refer to the fraction of sequence reads corresponding to a particular variant allele relative to the total number of sequence reads for a genomic locus.
[0095] Any headings of items used herein are for purposes of organization only and should not be construed as limiting the subject matter described.
[0096] Method for determining tumor DNA fraction The methods described herein can generally provide a more accurate assessment of the tumor DNA content in a biopsy sample, particularly a more accurate assessment of the circulating tumor DNA content in a cell-free DNA sample collected from a patient's plasma or other liquid biopsy sample. Briefly, variants (e.g., somatic short variants) are identified in a sample (e.g., a biopsy sample, a liquid biopsy sample, or a plasma sample) based on sequence read data derived from the sample, and their allele frequencies are quantified. An empirical distribution of all possible sample tumor DNA fraction values taking into account the observed somatic short variant allele frequencies can then be generated by (i) using an equation that describes the physical relationship between the somatic variant allele frequency and the tumor DNA fraction for a given copy number state at the variant genomic location and a given tumor ploidy, and (ii) leveraging historical patient data that includes known copy number and tumor ploidy profiles. A model, such as a probability density model (e.g., a non-parametric probability density model), can then be fit to the empirical distribution of the tumor DNA fraction values, and an estimated value of the tumor DNA fraction in the sample can be derived from the model along with upper and lower bounds of the estimated value based on a desired confidence interval.
[0097] FIG. 1 provides a non-limiting example of a flowchart of process 100 for determining the tumor DNA fraction (e.g., ctDNA fraction) in a sample. Process 100 can be implemented using one or more electronic devices that implement, for example, a software platform. In some examples, process 100 is implemented using a client-server system, and the blocks of process 100 are divided between the server and the client device in any manner. In other examples, the blocks of process 100 are divided between the server and multiple client devices. Thus, while a portion of process 100 is described herein as being implemented by a particular device of a client-server system, it should be understood that process 100 is not so limited. In other examples, process 100 is implemented using only client devices, or only multiple client devices. In process 100, some blocks are optionally combined, the order of some blocks is optionally changed, and some blocks are optionally omitted. In some cases, additional steps can be implemented in combination with process 100. Thus, the operations illustrated (and described in more detail below) are exemplary in nature and should not be considered limiting.
[0098] In step 102 of FIG. 1, nucleic acid (e.g., DNA) is extracted from the sample, and sequence read data is received for a plurality of sequence reads obtained by sequencing the nucleic acid. In some cases, the sample can include a tissue sample, a solid biopsy sample, a liquid biopsy sample, or any combination thereof. In some cases, the sample can include DNA extracted from a blood sample, a plasma sample, a cerebrospinal fluid sample, a pleural fluid sample, a sputum sample, a fecal sample, a urine sample, or a saliva sample. In some cases, the sample can include cell-free DNA (cfDNA). In some cases, the sample can include circulating tumor DNA (ctDNA).
[0099] In step 104 of FIG. 1, based on the array read data, the variant allele frequency (VAF) is determined for one or more variants detected in the sample. In some cases, the one or more variants may include one or more short variants, one or more rearrangement events, or any combination thereof. In some cases, the one or more short variants include one or more somatic short variants. In some cases, the one or more short variants (e.g., one or more short somatic variants) may be variants known not to be associated with clonal hematopoiesis of indeterminate potential (CHIP).
[0100] In some cases, one or more variants are ABL1, ACVR1B, AKT1, AKT2, AKT3, ALK, ALOX12B, AMER1, APC, AR, ARAF, ARFRP1, ARID1A, ASXL1, ATM, ATR, ATRX, AURKA, AURKB, AXIN1, AXL, BAP1, BARD1, BCL2, BCL2L1, BCL2L2, BCL6, BCOR, BCORL1, BCR, BRAF, BRCA1, BRCA2, BRD4, BRIP1, BTG1, BTG2, BTK, CALR, CARD11, CASP8, CBFB, CBL, CCND1, CCND2, CCND3, CCNE1, CD22, CD274, CD70, CD74, CD79A, CD79B, CDC73, CDH1, CDK12, CDK4, CDK6, CDK8, CDKN1A, CDKN1B, CDKN2A, CDKN2B, CDKN2C, CEBPA, CHEK1, CHEK2, CIC, CREBBP, CRKL, CSF1R, CSF3R, CTCF, CTNNA1, CTNNB1, CUL3, CUL4A, CXCR4, CYP17A1, DAXX, DDR1, DDR2, DIS3, DNMT3A, DOT1L, EED, EGFR, EMSY (C11orf30), EP300, EPHA3, EPHB1, EPHB4, ERBB2, ERBB3, ERBB4, ERCC4, ERG, ERRFI1, ESR1, ETV4, ETV5, ETV6, EWSR1, EZH2, EZR, FAM46C, FANCA, FANCC, FANCG, FANCL, FAS, FBXW7, FGF10, FGF12, FGF14, FGF19, FGF23, FGF3, FGF4, FGF6, FGFR1, FGFR2, FGFR3, FGFR4, FH, FLCN, FLT1, FLT3, FOXL2, FUBP1, GABRA6, GATA3, GATA4, GATA6, GID4 (C17orf39), GNA11, GNA13, GNAQ, GNAS, GRM3, GSK3B, H3F3A, HDAC1, HGF, HNF1A, HRAS, HSD3B1, ID3, IDH1, IDH2, IGF1R, IKBKE, IKZF1, INPP4B, IRF2, IRF4, IRS2, JAK1, JAK2, JAK3, JUN, KDM5A, KDM5C, KDM6A, KDR, KEAP1, KEL, KIT, KLHL6, KMT2A (MLL),Variants in the KMT2D (MLL2), KRAS, LTK, LYN, MAF, MAP2K1, MAP2K2, MAP2K4, MAP3K1, MAP3K13, MAPK1, MCL1, MDM2, MDM4, MED12, MEF2B, MEN1, MERTK, MET, MITF, MKNK1, MLH1, MPL, MRE11A, MSH2, MSH3, MSH6, MST1R, MTAP, MTOR, MUTYH, MYB, MYC, MYCL, MYCN, MYD88, NBN, NF1, NF2, NFE2L2, NFKBIA, NKX2-1, NOTCH1, NOTCH2, NOTCH3, NPM1, NRAS, NT5C2, NTRK1, NTRK2, NTRK3, NUTM1, P2RY8, PALB2, PARK2, PARP1, PARP2, PARP3, PAX5, PBRM1, PDCD1, PDCD1LG2, PDGFRA, PDGFRB, PDK1, PIK3C2B, PIK3C2G, PIK3CA, PIK3CB, PIK3R1, PIM1, PMS2, POLD1, POLE, PPARG, PPP2R1A, PPP2R2A, PRDM1, PRKAR1A, PRKCI, PTCH1, PTEN, PTPN11, PTPRO, QKI, RAC1, RAD21, RAD51, RAD51B, RAD51C, RAD51D, RAD52, RAD54L, RAF1, RARA, RB1, RBM10, REL, RET, RICTOR, RNF43, ROS1, RPTOR, RSPO2, SDC4, SDHA, SDHB, SDHC, SDHD, SETD2, SF3B1, SGK1, SLC34A2, SMAD2, SMAD4, SMARCA4, SMARCB1, SMO, SNCAIP, SOCS1, SOX2, SOX9, SPEN, SPOP, SRC, STAG2, STAT3, STK11, SUFU, SYK, TBX3, TEK, TERC, TERT, TET2, TGFBR2, TIPARP, TMPRSS2, TNFAIP3, TNFRSF14, TP53, TSC1, TSC2, TYRO3, U2AF1, VEGFA, VHL, WHSC1, WHSC1L1, WT1, XPO1, XRCC2, ZNF217 or ZNF703 loci, or any combination thereof may be included.,
[0101] In some cases, one or more variants may include variants at the ABL, ALK, ALL, B4GALNT1, BAFF, BCL2, BRAF, BRCA, BTK, CD19, CD20, CD3, CD30, CD319, CD38, CD52, CDK4, CDK6, CML, CRACC, CS1, CTLA-4, dMMR, EGFR, ERBB1, ERBB2, FGFR1-3, FLT3, GD2, HDAC, HER1, HER2, HR, IDH2, IL-1β, IL-6, IL-6R, JAK1, JAK2, JAK3, KIT, KRAS, MEK, MET, MSI-H, mTOR, PARP, PD-1, PDGFR, PDGFRα, PDGFRβ, PD-L1, PI3Kδ, PIGF, PTCH, RAF, RANKL, RET, ROS1, SLAMF7, VEGF, VEGFA or VEGFB loci, or any combination thereof.
[0102] In step 106 of FIG. 1, an empirical distribution of tumor DNA fraction values (e.g., ctDNA fraction values) as a function of the VAF determined for one or more variants is generated based on historical data for a plurality of subjects (e.g., patients) including, for example, known copy number and tumor ploidy profiles.
[0103] In some cases, generating an empirical distribution of tumor DNA fraction values may include calculating the tumor DNA fraction values based on the known copy number in one or more variants, the determined VAF in one or more variants, and the corresponding known tumor average ploidy in a plurality of historical subject samples having known VAFs in one or more variants that are substantially the same as the determined VAF in one or more variants. In some cases, generating an empirical distribution of tumor DNA fraction values may include pre-calculating the tumor DNA fraction values based on the known copy number in one or more variants and the corresponding known tumor average ploidy in a plurality of historical subject samples having a range of VAF values in one or more variants, and selecting a subset of the pre-calculated tumor DNA fraction values corresponding to samples having known VAFs in one or more variants that are substantially the same as the determined VAF in one or more variants. The term "substantially the same" in this context refers to situations where the known and determined values of the VAF differ by less than about 10%, less than about 8%, less than about 6%, less than about 4%, less than about 2%, less than about 1%, or less than about 0.5%.
[0104] In some cases, the tumor DNA fraction value can be calculated or selected for the variant showing the highest VAF in the sample from the subject. In some examples, the tumor DNA fraction value can be calculated or selected for the ranked set of two or more variants showing the highest ranked VAFs in the sample from the subject (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 variants showing the highest ranked VAF in the sample). In some cases, the tumor DNA fraction value is calculated or selected for a predetermined set of two or more variants detected in the sample from the subject (e.g., a predetermined set including 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 variants). In some cases, the tumor DNA fraction value is calculated or selected for a predetermined set of two or more variants detected in the sample from the subject that includes known driver mutations (e.g., a predetermined set including 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 variants including known driver mutations). In some cases, the tumor DNA fraction value is calculated or selected for all variants detected in the cfDNA sample from the subject.
[0105] In some cases, for example, if two or more variants are used to determine the tumor DNA fraction, the variants can be clustered according to their VAFs using an unsupervised clustering algorithm such as, for example, an unsupervised K - means clustering algorithm, a hierarchical clustering algorithm, or a Gaussian mixture model clustering algorithm. The center of each identified cluster (e.g., the average of all VAF values within the cluster) can represent the VAF corresponding to a particular copy number. If only one cluster is identified, the center of the cluster can be used as the value of the VAF used to generate the possible distribution of the tumor DNA fraction. If two or more clusters are identified, a set of rules can be applied to select one of the clusters that is likely to represent the diploid (i.e., copy number = 2) state, and then the center of the diploid cluster can be used as the value of the VAF used to generate the possible distribution of the tumor DNA fraction.
[0106] In some cases, the tumor DNA fraction value is calculated or pre-calculated by solving a set of equations representing (i) the relationship between tumor DNA fraction, tumor purity, and tumor average ploidy, and (ii) the relationship between somatic VAF, tumor purity, copy number at the genomic location of one or more variants, and the number of variant alleles in each of the one or more variants, based on the known VAF in one or more variants, the known copy number in one or more variants, and the corresponding known tumor average ploidy in a plurality of historical subject samples, thereby eliminating tumor purity and deriving the relationship in tumor DNA fraction as a function of somatic VAF, tumor average ploidy, copy number at the genomic location of one or more variants, and the number of variant alleles for one or more variants.
[0107] In some embodiments, the tumor DNA fraction value is such that the product of tumor purity and tumor average ploidy is divided by the sum of the product of tumor purity and tumor average ploidy and the product of 2 and the magnitude obtained by subtracting tumor purity from 1, and the tumor DNA fraction is made equal in a first equation, and the product of tumor purity and the number of variant alleles in each of the one or more variants is divided by the sum of the product of tumor purity and copy number at the genomic location of the one or more variants and the product of 2 and the magnitude obtained by subtracting tumor purity from 1, and the somatic VAF is made equal in a second equation, based on the known VAF in one or more variants, the known copy number in one or more variants, and the corresponding known tumor average ploidy in a plurality of historical subject samples, and may be calculated or pre-calculated by solving the set of equations including the first and second equations, thereby eliminating tumor purity and deriving a relationship where the tumor DNA fraction is equal to the magnitude obtained by dividing the tumor average ploidy by the sum of the ratio of the number of variant alleles in the one or more variants and the somatic VAF in each of the one or more variants to the magnitude obtained by subtracting the copy number at the genomic location of the one or more variants from the tumor average ploidy.
[0108] In some cases, the tumor DNA fraction value is based on the known VAF in one or more variants, the known copy number in one or more variants, and the corresponding known tumor average ploidy in a plurality of historical subject samples, and may be calculated or pre-calculated by solving a set of equations consisting of TIFF2025524649000006.tif34170, thereby eliminating ρ, TIFF2025524649000007.tif20170where ρ is the tumor purity, ψ is the tumor average ploidy, C is the copy number at the genomic position of one or more variants, and V is the number of variant alleles in each of one or more variants.
[0109] In some cases, the plurality of historical subject samples (e.g., patient samples) may include solid biopsy samples, liquid biopsy samples, or any combination thereof. In some cases, the plurality of historical subject samples may include cancer samples. In some cases, the plurality of historical subject samples may include samples of a single type of cancer. In some cases, the plurality of historical subject samples may include samples of multiple types of cancer.
[0110] In some cases, the plurality of historical subject samples (e.g., patient samples) may include bladder cancer samples, breast cancer samples, colorectal cancer samples, endometrial cancer samples, kidney cancer samples, leukemia samples, liver cancer samples, lung cancer samples, melanoma samples, non-Hodgkin lymphoma samples, pancreatic cancer samples, prostate cancer samples, thyroid cancer samples, or any combination thereof.
[0111] A plurality of historical target samples (e.g., patient samples) include acute lymphoblastic leukemia (Philadelphia chromosome positive) samples, acute lymphoblastic leukemia (precursor B cell) samples, acute myeloid leukemia (FLT3+) samples, acute myeloid leukemia (with IDH2 mutation) samples, anaplastic large cell lymphoma samples, basal cell carcinoma samples, B-cell chronic lymphocytic leukemia samples, bladder cancer samples, breast cancer (HER2 overexpression / amplification) samples, breast cancer (HER2+) samples, breast cancer (HR+, HER2-) samples, cervical cancer samples, cholangiocarcinoma samples, chronic lymphocytic leukemia samples, chronic lymphocytic leukemia (with 17p deletion) samples, chronic myelogenous leukemia samples, chronic myelogenous leukemia (Philadelphia chromosome positive) samples, classical Hodgkin lymphoma samples, colorectal cancer samples, colorectal cancer (dMMR and MSI-H) samples, colorectal cancer (KRAS wild-type) samples, cryopyrin-associated periodic syndrome samples, cutaneous T-cell lymphoma samples, dermatofibrosarcoma protuberans samples, diffuse large B-cell lymphoma samples, fallopian tube cancer samples, follicular B-cell non-Hodgkin lymphoma samples, follicular lymphoma samples, gastric cancer samples, gastric cancer (HER2+) samples, gastroesophageal junction (GEJ) adenocarcinoma samples, gastrointestinal stromal tumor samples, gastrointestinal stromal tumor samples, gastrointestinal stromal tumor (KIT+) samples, giant cell tumor of bone samples, glioblastoma samples, polyangiitis granulomatosa samples, head and neck squamous cell carcinoma samples, hepatocellular carcinoma samples, Hodgkin lymphoma samples, juvenile idiopathic arthritis samples, lupus erythematosus samples, mantle cell lymphoma samples, medullary thyroid cancer samples, melanoma samples, melanoma with BRAF V600 mutation samples, melanoma with BRAF V600E or V600K mutation samples, Merkel cell carcinoma samples, Castleman disease samples, multiple blood samples including Philadelphia chromosome positive ALL and CML, multiple myeloma samples, myelofibrosis samples, non-Hodgkin lymphoma samples, inoperable subependymal giant cell astrocytoma samples associated with tuberous sclerosis, non-small cell lung cancer samples, non-small cell lung cancer (ALK+) samples, non-small cell lung cancer (PD-L1+) samples, non-small cell lung cancer (with ALK fusion or ROS1 gene alteration) samples, non-small cell lung cancer (with BRAF V600E mutation) samples, non-small cell lung cancer samples (with EGFR exon 19 deletion or exon 21 substitution (L858R) mutation), non-small cell lung cancer samples (EGFR(with T790M mutation), ovarian cancer samples, ovarian cancer samples (with BRCA mutation), pancreatic cancer samples, pancreatic cancer samples, gastrointestinal cancer samples, lung-derived neuroendocrine tumor samples, pediatric neuroblastoma samples, peripheral T-cell lymphoma samples, prostate cancer samples, peritoneal cancer samples, renal cell cancer samples, rheumatoid arthritis samples, small lymphocyte lymphoma samples, soft tissue sarcoma samples, solid tumor (MSI-H / dMMR) samples, head and neck squamous cell carcinoma samples, squamous non-small cell lung cancer samples, thyroid cancer samples, thyroid cancer samples, urothelial cancer samples, urothelial cancer samples, Waldenström macroglobulinemia samples, or any combination thereof.
[0112] In step 108 of FIG. 1, the model is fitted to the empirical distribution of the tumor DNA fraction values. In some cases, the model may be a probability density model. In some cases, the model may be a parametric probability density model. In some cases, the model may be a non-parametric probability density model.
[0113] In step 110 of FIG. 1, a value of the tumor DNA fraction (e.g., ctDNA fraction) for the sample is determined based on the model. In some cases, the method may further include the step of determining a confidence interval for the tumor DNA fraction based on the model. In some cases, the tumor DNA fraction determined for the sample may be the most likely tumor DNA fraction. In some cases, the tumor DNA fraction determined for the cfDNA sample may be the mean, median, or mode of the dominant peak in the empirical distribution of the tumor DNA fraction values.
[0114] In some cases, the confidence interval can be determined for the tumor DNA fraction (e.g., ctDNA fraction) based on the model. For example, the confidence interval can be calculated as the determined average value (e.g., estimated value) of the tumor DNA fraction plus or minus a measure of the variability of that estimated value (e.g., a value proportional to the standard deviation), providing a range of tumor DNA fraction values within which the determined value of the tumor DNA fraction is expected to fall at a specified confidence level if the determination is repeated. As a non-limiting example, the confidence interval for data following a normal distribution is given by: TIFF2025524649000008.tif15170where CI is the confidence interval (e.g., 95%), TIFF2025524649000009.tif6170is the average value of the tumor DNA fraction determined based on a plurality of historical samples, Z is the critical value of the z-distribution (e.g., Z = 1.96 for a 95% confidence level), s is the standard deviation of the tumor DNA fraction determined based on a plurality of historical samples, and N is the number of samples in the plurality of historical samples. In some cases, the confidence interval can be determined for a confidence level of, for example, 90%, 95%, 98%, or 99%.
[0115] Method of Use In some cases, the disclosed method may further include one or more of: (i) obtaining a sample from a subject (e.g., a subject suspected of having cancer or determined to have cancer); (ii) extracting nucleic acid molecules (e.g., a mixture of tumor nucleic acid molecules and non-tumor nucleic acid molecules) from the sample; (iii) ligating one or more adapters (e.g., one or more amplification primers, flow cell adapter sequences, substrate adapter sequences, or sample index sequences) to the nucleic acid molecules extracted from the sample; (iv) amplifying the nucleic acid molecules (e.g., using polymerase chain reaction (PCR) amplification techniques, non-PCR amplification techniques, or isothermal amplification techniques); (v) capturing nucleic acid molecules from the amplified nucleic acid molecules (e.g., by hybridization to one or more bait molecules each containing a region complementary to a region of the captured nucleic acid molecules); (vi) sequencing the nucleic acid molecules extracted from the sample (or a library proxy derived therefrom) using, for example, a next-generation (e.g., massively parallel) sequencer and, for example, next-generation (massively parallel) sequencing techniques, whole-genome sequencing (WGS) techniques, whole-exome sequencing techniques, targeted sequencing techniques, direct sequencing techniques, or Sanger sequencing techniques; and (vii) generating, displaying, transmitting, and / or delivering a report (e.g., an electronic report, a web-based report, or a paper report) to a subject (or patient), caregiver, healthcare provider, physician, oncologist, electronic health record system, hospital, clinic, medical practice, third-party payer, insurance company, or government agency. In some cases, the report includes an output from the methods described herein. In some cases, all or part of the report can be displayed on the graphical user interface of an online or web-based healthcare portal. In some cases, the report is transmitted via a computer network or a peer-to-peer connection.
[0116] The disclosed method can be used with any of a variety of samples. For example, in some cases, the sample can include a tissue biopsy sample, a liquid biopsy sample, or a normal control. In some cases, the sample can be a liquid biopsy sample and can include blood, plasma, cerebrospinal fluid, sputum, feces, urine, or saliva. In some cases, the sample can be a liquid biopsy sample and can include circulating tumor cells (CTCs). In some cases, the sample can be a liquid biopsy sample and can include cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), or any combination thereof.
[0117] In some cases, the nucleic acid molecules extracted from the sample can include a mixture of tumor nucleic acid molecules and non-tumor nucleic acid molecules. In some cases, the tumor nucleic acid molecules can be derived from the tumor portion of a heterogeneous tissue biopsy sample, and the non-tumor nucleic acid molecules can be derived from the normal portion of the heterogeneous tissue biopsy sample. In some cases, the sample can include a liquid biopsy sample, the tumor nucleic acid molecules can be derived from the circulating tumor DNA (ctDNA) fraction of the liquid biopsy sample, and the non-tumor nucleic acid molecules can be derived from the non-tumor cell-free DNA (cfDNA) fraction of the liquid biopsy sample.
[0118] In some cases, the disclosed method for determining a tumor DNA fraction (e.g., a ctDNA fraction) can be used to diagnose (or as part of a diagnosis) the presence of a disease or other condition (e.g., cancer, a genetic disorder (such as Down syndrome and Fragile X), a neurological disorder, or a variant, such as a copy number change) in a subject (e.g., a patient) where the detection of said disease is relevant to the diagnosis, treatment, or prediction of any other disease type. In some cases, the disclosed method can be applicable to the diagnosis of any of a variety of cancers, as described elsewhere herein.
[0119] In some cases, the disclosed methods for determining a tumor DNA fraction (e.g., a ctDNA fraction) can be used to select subjects (e.g., patients) for a clinical trial based on the determined tumor DNA fraction. In some cases, for example, patient selection for a clinical trial based on the determination of a tumor DNA fraction can accelerate the development of targeted therapies and improve healthcare outcomes for treatment decisions.
[0120] In some cases, the disclosed methods for determining a tumor DNA fraction (e.g., a ctDNA fraction) can be used to select an appropriate treatment or therapy (e.g., an anti-cancer therapy or treatment) for a subject. In some cases, for example, the anti-cancer therapy or treatment can include the use of poly(ADP-ribose) polymerase inhibitors (PARPis), platinum compounds, chemotherapy, radiation therapy, targeted therapy (e.g., immunotherapy), surgery, or any combination thereof.
[0121] In some cases, targeted therapy (or targeted cancer therapy) includes abemaciclib (Verzenio), abiraterone acetate (Zytiga), acalabrutinib (Calquence), trastuzumab emtansine (Kadcyla), afatinib dimaleate (Gilotrif), aldesleukin (Proleukin), alectinib (Alecensa), alemtuzumab (Campath), alitretinoin (Panretin), alpelisib (Piqray), amivantamab-vmjw (Rybrevant), anastrozole (Arimidex), apalutamide (Erleada), asimicinib hydrochloride (Scemblix), atezolizumab (Tecentriq), avapritinib (Ayvakit), avelumab (Bavencio), axicabtagene ciloleucel (Yescarta), axitinib (Inlyta), belantamab mafodotin-blmf (Blenrep), belimumab (Benlysta), belinostat (Beleodaq), belzutifan (Welireg), bevacizumab (Avastin), bexarotene (Targretin), binimetinib (Mektovi), blinatumomab (Blincyto), bortezomib (Velcade), bosutinib (Bosulif), brentuximab vedotin (Adcetris), brexucabtagene autoleucel (Tecartus), brigatinib (Alunbrig), cabazitaxel (Jevtana), cabozantinib (Cabometyx), cabozantinib (Cabometyx, Cometriq), canakinumab (Ilaris), capmatinib hydrochloride (Tabrecta), carfilzomib (Kyprolis), cemiplimab-rwlc (Libtayo), ceritinib (LDK378 / Zykadia), cetuximab (Erbitux), cobimetinib (Cotellic), copanlisib hydrochloride (Aliqopa), crizotinib (Xalkori), dabrafenib (Tafinlar), dacomitinib (Vizimpro), daratumumab (Darzalex), daratumumab and hyaluronidase-fihj (DarzalexFaspro, darolutamide (Nubeqa), dasatinib (Sprycel), denileukin diftitox (Ontak), denosumab (Xgeva), dinutuximab (Unituxin), dostarlimab-gxly (Jemperli), durvalumab (Imfinzi), duvelisib (Copiktra), elotuzumab (Empliciti), enasidenib mesylate (Idhifa), encorafenib (Braftovi), enfortumab vedotin-ejfv (Padcev), entrectinib (Rozlytrek), enzalutamide (Xtandi), erdafitinib (Balversa), erlotinib (Tarceva), everolimus (Afinitor), exemestane (Aromasin), fam-trastuzumab deruxtecan-nxki (Enhertu), fedratinib hydrochloride (Inrebic), fulvestrant (Faslodex), gefitinib (Iressa), gemtuzumab ozogamicin (Mylotarg), gilteritinib (Xospata), glasdegib maleate (Daurismo), hyaluronidase-zzxf (Phesgo), ibrutinib (Imbruvica), ibritumomab tiuxetan (Zevalin), idecabtagene vicleucel (Abecma), idelalisib (Zydelig), imatinib mesylate (Gleevec), infigratinib phosphate (Truseltiq), inotuzumab ozogamicin (Besponsa), iobenguane I131 (Azedra), ipilimumab (Yervoy), isatuximab-irfc (Sarclisa),ivosidenib (Tibsovo), ixazomib citrate (Ninlaro), lanreotide acetate (SomatulineDepot), Lapatinib (Tykerb), Larotrectinib Sulfate (Vitrakvi), Lenvatinib Mesylate (Lenvima), Letrozole (Femara), Isocabtagene Autoleucel (Breyanzi), Loncastuximab Tesirine-lpyl (Zynlonta), Lorlatinib (Lorbrena), Lutetium Lu177-dotatate (Lutathera), Margetuximab-cmkb (Margenza), Midostaurin (Rydapt), Mobocertinib Succinate (Exkivity), Mogamulizumab-kpkc (Poteligeo), Moxetumomab Pasudotox-tdfk (Lumoxiti), Naxitamab-gqgk (Danyelza), Necitumumab (Portrazza), Neratinib Maleate (Nerlynx), Nilotinib (Tasigna), Niraparib Tosylate Monohydrate (Zejula), Nivolumab (Opdivo), Obinutuzumab (Gazyva), Ofatumumab (Arzerra), Olaparib (Lynparza), Olaratumab (Lartruvo), Osimertinib (Tagrisso), Palbociclib (Ibrance), Panitumumab (Vectibix), Panobinostat (Farydak), Pazopanib (Votrient), Pembrolizumab (Keytruda), Pemigatinib (Pemazyre), Pertuzumab (Perjeta), Pegaspargase Hydrochloride (Turalio), Polatuzumab Vedotin-piiq (Polivy), Ponatinib Hydrochloride (Iclusig), Pralatrexate (Folotyn), Pralsetinib (Gavreto), Radium 223 Dichloride (Xofigo), Ramucirumab (Cyramza), Regorafenib (Stivarga), Ribociclib (Kisqali), Repotrectinib (Qinlock), Rituximab (Rituxan), Rituximab and Hyaluronidase Human (RituxanHycela), romidepsin (Istodax), rucaparib (Rubraca), ruxolitinib phosphate (Jakafi), sacituzumab govitecan-fujii (Trodelvy), selinexor, selinexor (Xpovio), selpercatinib (Retevmo), selumetinib sulfate (Koselugo), siltuximab (Sylvant), sipuleucel-T (Provenge), sirolimus protein-bound particles (Fyarro), sonidegib (Odomzo), sorafenib (Nexavar), sotorasib (Lumakras), sunitinib (Sutent), tafasitamab-cxix (Monjuvi), tagraxofusp-erzs (Elzonris), talazoparib tosylate (Talzenna), tamoxifen (Nolvadex), tazemetostat hydrobromide (Tazverik), tebentafusp-teb (Kimmtrak), temsirolimus (Torisel), tepotinib hydrochloride (Tepmetko), tisagenlecleucel (Kymriah), tisotumab vedotin-tftv (Tivdak), tocilizumab (Actemra), tofacitinib (Xeljanz), tositumomab (Bexxar), trametinib (Mekinist), trastuzumab (Herceptin), tretinoin (Vesanoid), tivozanib hydrochloride (Fotivda), toremifene (Fareston), tucatinib (Tukysa), umbralisib tosylate (Ukoniq), vandetanib (Caprelsa), vemurafenib (Zelboraf), venetoclax (Venclexta), vismodegib (Erivedge), vorinostat (Zolinza), zanubrutinib (Brukinsa), aflibercept (Zaltrap), or any combination thereof may be included.
[0122] In some cases, the disclosed methods for determining tumor DNA fraction (e.g., ctDNA fraction) can be used in treating a subject's disease (e.g., cancer). For example, in response to determining the tumor DNA fraction of a sample from a subject using any of the methods disclosed herein, an effective amount of an anti-cancer therapy or treatment can be administered to the subject.
[0123] In some cases, the disclosed methods for determining tumor DNA fraction (e.g., ctDNA fraction) can be used to monitor disease progression or recurrence in a subject (e.g., cancer or tumor progression or recurrence). For example, in some cases, the method can be used to determine the tumor DNA fraction in a first sample obtained from a subject at a first time point and to determine the tumor DNA fraction in a second sample obtained from the subject at a second time point, and the progression or recurrence of the disease can be monitored by comparing the first determination of the tumor DNA fraction with the second determination of the tumor DNA fraction. In some cases, the first time point is selected before the subject is administered a therapy or treatment, and the second time point is selected after the subject has been administered a therapy or treatment.
[0124] In some cases, the disclosed methods can be used to adjust a subject's therapy or treatment (e.g., anti-cancer treatment or anti-cancer therapy), such as by adjusting the therapeutic dose and / or selecting a different treatment in response to a change in the determination of the tumor DNA fraction (e.g., ctDNA fraction).
[0125] In some cases, the value of the tumor DNA fraction (e.g., ctDNA fraction) determined using the disclosed method can be used as a prognostic or diagnostic indicator related to the sample. For example, in some cases, the prognostic or diagnostic indicator can include an indicator of the presence of a disease (e.g., cancer) in the sample, an indicator of the probability that the disease (e.g., cancer) is present in the sample, an indicator of the probability that the subject from whom the sample is derived will develop the disease (e.g., cancer) (i.e., a risk factor), or an indicator of the likelihood that the subject from whom the sample is derived will respond to a particular therapy or treatment.
[0126] In some cases, the disclosed methods for determining tumor DNA fraction (e.g., ctDNA fraction) can be part of a process of genomic profiling that includes identification of the presence of variant sequences at one or more loci in a sample from a subject as part of detecting a particular disease, such as cancer, monitoring, predicting risk factors, or selecting a treatment. In some cases, the variant panel selected for genomic profiling can include detection of variant sequences at a selected set of loci. In some cases, the variant panel selected for genomic profiling can include detection of variant sequences at several loci via comprehensive genomic profiling (CGP), which is a next-generation sequencing (NGS) approach used to evaluate hundreds of genes (including relevant cancer biomarkers) in a single assay. Including the disclosed methods for determining tumor DNA fraction as part of the genomic profiling step (or including the output from the disclosed methods for determining tumor DNA fraction as part of a subject's genomic profile) can be done based on the genomic profile, for example, by independently verifying the tumor DNA fraction in a given patient sample, and can improve, for example, the validity of disease detection calls and treatment decisions.
[0127] In some cases, a genomic profile can include information regarding the presence of genes (or variant sequences thereof), copy number variations, epigenetic traits, proteins (or modifications thereof), and / or other biomarkers in an individual's genome and / or proteome, as well as information regarding the individual's corresponding phenotypic traits and the interactions between genetic or genomic traits, phenotypic traits, and environmental factors.
[0128] In some cases, a subject's genomic profile can include results from a comprehensive genomic profiling (CGP) test, a nucleic acid sequencing-based test, a gene expression profiling test, a cancer hot spot panel test, a DNA methylation test, a DNA fractionation test, an RNA fractionation test, or any combination thereof.
[0129] In some cases, the method may further include administering or applying a treatment or therapy (e.g., an anti-cancer agent, anti-cancer treatment, or anti-cancer therapy) to a subject based on the generated genomic profile. The anti-cancer agent or anti-cancer treatment may refer to a compound that is effective in treating cancer cells. Examples of anti-cancer agents or anti-cancer therapies include, but are not limited to, alkylating agents, antimetabolites, natural products, hormones, chemotherapy, radiation therapy, immunotherapy, surgery, or a therapy configured to target a defect in a specific cell signaling pathway, such as a defect in the DNA mismatch repair (MMR) pathway.
[0130] Sample The disclosed methods and systems can be used with any of a variety of samples (also referred to herein as specimens) that contain nucleic acids (e.g., DNA or RNA) collected from a subject (e.g., a patient). Examples of samples include, but are not limited to, tumor samples, tissue samples, biopsy samples (e.g., tissue biopsies, liquid biopsies, or both), blood samples (e.g., peripheral whole blood samples), plasma samples, serum samples, lymph samples, saliva samples, sputum samples, urine samples, gynecological fluid samples, circulating tumor cell (CTC) samples, cerebrospinal fluid (CSF) samples, pericardial fluid samples, pleural effusion samples, ascites (peritoneal fluid) samples, fecal (or stool) samples, or other body fluids, secretions, and / or excretory samples (or cell samples derived therefrom). In certain cases, the sample can be a frozen sample or a formalin-fixed paraffin-embedded (FFPE) sample.
[0131] In some cases, the sample can be collected by tissue resection (e.g., surgical resection), needle biopsy, bone marrow biopsy, bone marrow aspiration, skin biopsy, endoscopic biopsy, fine needle aspiration, oral swab, nasal swab, vaginal swab, or cytological smear, abrasion, washing or washings (such as a lumen washing or a bronchoalveolar lavage fluid), etc.
[0132] In some cases, the sample is a liquid biopsy sample and can include, for example, whole blood, plasma, serum, urine, feces, sputum, saliva, or cerebrospinal fluid. In some cases, the sample can be a liquid biopsy sample and can include circulating tumor cells (CTCs). In some cases, the sample can be a liquid biopsy sample and can include cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), or any combination thereof.
[0133] In some cases, the sample can include one or more pre-malignant or malignant cells. As used herein, a pre-malignant tumor refers to cells or tissues that are not yet malignant but are ready to become malignant. In certain examples, the sample can be obtained from a solid tumor, soft tissue tumor, or metastatic lesion. In certain examples, the sample can be obtained from a hematologic malignancy or pre-malignant tumor. In other examples, the sample can include tissue or cells from a surgical margin. In certain examples, the sample can include tumor-infiltrating lymphocytes. In some cases, the sample can include one or more non-malignant cells. In some cases, the sample can be or can be a part of a primary tumor or metastasis (e.g., a metastatic biopsy sample). In some cases, the sample can be obtained from a site where the percentage of the tumor (e.g., tumor cells) is the highest compared to an adjacent site (e.g., a site adjacent to the tumor) (e.g., the tumor site). In some cases, the sample can be obtained from a site that has the largest tumor lesion (e.g., the largest number of tumor cells when visualized microscopically) compared to an adjacent site (e.g., a site adjacent to the tumor) (e.g., the tumor site).
[0134] In some cases, the disclosed method may further include analyzing a primary control (e.g., a normal tissue sample). In some cases, the disclosed method may further include determining whether a primary control is available and, if available, isolating a control nucleic acid (e.g., DNA) from the primary control. In some cases, the sample may include any normal control (e.g., normal adjacent tissue (NAT)) if a primary control is not available. In some cases, the sample may be or include histologically normal tissue. In some cases, the method includes evaluating a sample, e.g., a histologically normal sample (e.g., from a surgical tissue margin), using the methods described herein. In some cases, the disclosed method may further include obtaining a sub-sample enriched in non-tumor cells by macrodissecting non-tumor tissue from the NAT in a sample without a primary control, for example. In some cases, the disclosed method may further include determining that a primary control and NAT are not available and marking the sample for analysis without a matched control.
[0135] In some cases, a sample obtained from histologically normal tissue (e.g., a histologically normal tissue margin otherwise) may still contain genetic changes such as the variant sequences described herein. Thus, the method may further include reclassifying the sample based on the presence of the detected genetic changes. In some cases, multiple samples (e.g., from different subjects) are processed simultaneously.
[0136] The disclosed methods and systems can be applied to the analysis of nucleic acids extracted from various tissue samples (or their disease states), such as solid tissue samples, soft tissue samples, metastatic lesions, or liquid biopsy samples. Examples of tissues include, but are not limited to, connective tissue, muscle tissue, nervous system tissue, epithelial tissue, and blood. Tissue samples can be collected from any of the organs in an animal or human body. Examples of human organs include, but are not limited to, the brain, heart, lungs, liver, kidneys, pancreas, spleen, thyroid, breast, uterus, prostate, large intestine, small intestine, bladder, bone, skin, etc.
[0137] In some cases, the nucleic acids extracted from the sample can contain deoxyribonucleic acid (DNA) molecules. Examples of DNA that may be suitable for analysis by the disclosed methods include, but are not limited to, genomic DNA or fractions thereof, mitochondrial DNA or fractions thereof, cell-free DNA (cfDNA), and circulating tumor DNA (ctDNA). Cell-free DNA (cfDNA) is composed of fractions of DNA that are released from normal and / or cancer cells during apoptosis and necrosis, circulate in the bloodstream, and / or accumulate in other body fluids. Circulating tumor DNA (ctDNA) is composed of fractions of DNA that are released from cancer cells and tumors that circulate in the bloodstream and / or accumulate in other body fluids.
[0138] In some cases, DNA is extracted from nucleated cells in the sample. In some cases, the sample has a low nucleated cell enrichment, for example, when the sample consists mainly of red blood cells, diseased cells containing excessive cytoplasm, or tissue with fibrosis. In some cases, samples with low nucleated cell enrichment may require more, for example, a larger tissue volume for DNA extraction.
[0139] In some cases, the nucleic acid extracted from a sample can contain ribonucleic acid (RNA) molecules. Examples of RNA that may be suitable for analysis by the disclosed methods include, but are not limited to, total cellular RNA, total cellular RNA after depletion of a specific abundance of RNA sequences (e.g., ribosomal RNA), cell-free RNA (cfRNA), messenger RNA (mRNA) or fractions thereof, poly(A)-tailed mRNA fractions of total RNA, ribosomal RNA (rRNA) or fractions thereof, transfer RNA (tRNA) or fractions thereof, and mitochondrial RNA or fractions thereof. In some cases, the RNA is extracted from the sample and can be converted to complementary DNA, for example, using a reverse transcription reaction. In some cases, the cDNA is produced by a random-prime cDNA synthesis method. In other examples, cDNA synthesis is initiated at the poly(A) tail of mature mRNA by priming with an oligo(dT)-containing oligonucleotide. Methods for depletion, poly(A) enrichment, and cDNA synthesis are well known to those of skill in the art.
[0140] In some cases, the sample can contain tumor content (e.g., containing tumor cells or tumor cell nuclei) or non-tumor content (e.g., immune cells, fibroblasts, and other non-tumor cells). In some cases, the tumor content of the sample can constitute a sample metric. In some cases, the sample can contain a tumor content having at least 5-50%, 10-40%, 15-25%, or 20-30% tumor cell nuclei. In some cases, the sample can contain a tumor content of at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, or at least 50% tumor cell nuclei. In some cases, the percentage of tumor cell nuclei (e.g., a sample fraction) is determined (e.g., calculated) by dividing the number of tumor cells in the sample by the total number of all cells in the sample having nuclei. In some cases, for example, when the sample is a liver sample containing hepatocytes, different tumor content calculations may be required due to the presence of hepatocytes having two or more than two nuclei, the presence of other DNA content, e.g., non-hepatocytes, somatic cell nuclei. In some cases, the sensitivity of detecting genetic changes, e.g., variant sequences, or, e.g., determining microsatellite instability, can depend on the tumor content of the sample. For example, a sample having a lower tumor content can result in a lower sensitivity of detection for a given size of sample.
[0141] In some cases, as described above, the sample can contain nucleic acids (e.g., DNA, RNA (or cDNA derived from RNA), or both), for example, from a tumor or from normal tissue. In certain examples, the sample can further contain non-nucleic acid components, for example, cells, proteins, carbohydrates, or lipids, for example, derived from a tumor or normal tissue.
[0142] Subject In some cases, the sample is obtained (e.g., collected) from a subject (e.g., a patient) having a certain condition or disease (e.g., a proliferative disorder or a non-cancerous marker), or suspected of having a certain condition or disease. In some cases, the proliferative disorder is cancer. In some cases, the cancer is a solid tumor or its metastatic form. In some cases, the cancer is a blood cancer, e.g., leukemia or lymphoma.
[0143] In some cases, the subject has cancer or is at risk of having cancer. For example, in some cases, the subject has a genetic predisposition to cancer (e.g., having a genetic mutation that increases the baseline risk of developing cancer). In some cases, the subject is exposed to environmental variations (e.g., radiation or chemicals) that increase the risk of developing cancer. In some cases, the subject needs to be monitored for the development of cancer. In some cases, the subject needs to be monitored for the progression or regression of cancer, e.g., after being treated with an anti-cancer therapy (or anti-cancer treatment). In some cases, the subject needs to be monitored for cancer recurrence. In some cases, the subject needs to be monitored for minimal residual disease (MRD). In some cases, the subject has been or is being treated for cancer. In some cases, the subject has not been treated with an anti-cancer therapy (or anti-cancer treatment).
[0144] In some cases, the subject (e.g., a patient) is being or has previously been treated with one or more targeted therapies. In some cases, for example, a post-targeted-therapy sample (e.g., a specimen) is obtained (e.g., collected) from a patient who has previously been treated with a targeted therapy. In some cases, the post-targeted-therapy sample is a sample obtained after completion of the targeted therapy.
[0145] In some cases, the patient has not been previously treated with a targeted therapy. In some cases, for example, for patients who have not been previously treated with a targeted therapy, the sample includes an excision, for example, an initial excision, or an excision after recurrence (for example, after disease recurrence after therapy).
[0146] cancer In some cases, the sample is obtained from a subject having cancer. Exemplary cancers include, but are not limited to, B-cell cancers (e.g., multiple myeloma), melanoma, breast cancer, lung cancer (such as non-small cell lung cancer or NSCLC), bronchial cancer, colorectal cancer, prostate cancer, pancreatic cancer, stomach cancer, ovarian cancer, bladder cancer, brain or central nervous system cancer, peripheral nervous system cancer, esophageal cancer, cervical cancer, uterine or endometrial cancer, oral or pharyngeal cancer, liver cancer, kidney cancer, testicular cancer, biliary tract cancer, small intestine or accessory organ cancer, salivary gland cancer, thyroid cancer, adrenal gland cancer, osteosarcoma, chondrosarcoma, cancers of the blood tissue, adenocarcinoma, inflammatory fibroblastic tumor, gastrointestinal stromal tumor (GIST), colon cancer, multiple myeloma (MM), myelodysplastic syndrome (MDS), myeloproliferative disorder (MPD), acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic myeloid leukemia (CML), chronic lymphocytic leukemia (CLL), polycythemia vera, Hodgkin lymphoma, non-Hodgkin lymphoma (NHL), soft tissue sarcoma, fibrosarcoma, myxosarcoma, liposarcoma, osteogenic sarcoma, chordoma, angiosarcoma, endotheliosarcoma synovial sarcoma, mesothelioma, Ewing tumor, leiomyosarcoma, rhabdomyosarcoma, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinoma, medullary carcinoma, bronchiogenic carcinoma, renal cell carcinoma, hepatoma, cholangiocarcinoma, choriocarcinoma, seminoma, embryonal carcinoma tumor, Wilms tumor, bladder cancer, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pinealoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, meningioma, neuroblastoma, retinoblastoma, diffuse large B-cell lymphoma, mantle cell lymphoma, hepatocellular carcinoma, thyroid cancer, stomach cancer, head and neck cancer, small cell carcinoma, essential thrombocythemia, agnogenic myeloid metaplasia, hypereosinophilic syndrome, systemic mastocytosis, familial hypereosinophilia, chronic eosinophilic leukemia, neuroendocrine cancer, cancer-like tumor, and the like.
[0147] In some cases, the cancer is acute lymphoblastic leukemia (Philadelphia chromosome positive), acute lymphoblastic leukemia (precursor B cell), acute myeloid leukemia (FLT3+), acute myeloid leukemia (with IDH2 mutation), anaplastic large cell lymphoma, basal cell carcinoma, B-cell chronic lymphocytic leukemia, bladder cancer, breast cancer (HER2 overexpression / amplification), breast cancer (HER2+), breast cancer (HR+, HER2-), cervical cancer, cholangiocarcinoma, chronic lymphocytic leukemia, chronic lymphocytic leukemia (with 17p deletion), chronic myeloid leukemia, chronic myeloid leukemia (Philadelphia chromosome positive), classical Hodgkin lymphoma, colorectal cancer, colorectal cancer (dMMR / MSI-H), colorectal cancer (KRAS wild type), cryopyrin-associated periodic syndrome, cutaneous T-cell lymphoma, precursor cutaneous fibrosarcoma, diffuse large B-cell lymphoma, fallopian tube cancer, follicular B-cell non-Hodgkin lymphoma, follicular lymphoma, gastric cancer, gastric cancer (HER2+), gastroesophageal junction (GEJ) adenocarcinoma, gastrointestinal stromal tumor, gastrointestinal stromal tumor (KIT+), giant cell tumor of bone, glioblastoma, granulomatosis with polyangiitis, head and neck squamous cell carcinoma, hepatocellular carcinoma, Hodgkin lymphoma, juvenile idiopathic arthritis, lupus erythematosus, mantle cell lymphoma, medullary thyroid cancer, melanoma, melanoma with BRAF V600 mutation, melanoma with BRAF V600E or V600K mutation, Merkel cell carcinoma, multicentric Castleman disease, multiple hematologic malignancies including Philadelphia chromosome positive ALL and CML, multiple myeloma, myelofibrosis, non-Hodgkin lymphoma, unresectable subependymal giant cell astrocytoma associated with tuberous sclerosis, non-small cell lung cancer, non-small cell lung cancer (ALK+), non-small cell lung cancer (PD-L1+), non-small cell lung cancer (with ALK fusion or ROS1 gene alteration), non-small cell lung cancer (with BRAF V600E mutation), non-small cell lung cancer (with EGFR exon 19 deletion or exon 21 substitution (L858R) mutation), non-small cell lung cancer (EGFR(with T790M mutation), ovarian cancer, ovarian cancer (with BRCA mutation), pancreatic cancer, pancreatic cancer, gastrointestinal cancer or lung-derived neuroendocrine tumor, pediatric neuroblastoma, peripheral T-cell lymphoma, peritoneal cancer, prostate cancer, renal cell cancer, rheumatoid arthritis, small lymphocyte lymphoma, soft tissue sarcoma, solid tumor (MSI-H / dMMR), head and neck squamous cell carcinoma, non-small cell lung cancer, squamous cell carcinoma, thyroid cancer, thyroid cancer, urothelial cancer, urothelial cancer, or Waldenström macroglobulinemia.
[0148] In some cases, the cancer is a hematologic malignancy (or pre-malignancy). As used herein, a hematologic malignancy refers to a tumor of hematopoietic or lymphoid tissue, e.g., a tumor affecting the blood, bone marrow, or lymph nodes. Exemplary hematologic malignancies include leukemia (e.g., acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), hairy cell leukemia, acute monocytic leukemia (AMoL), chronic myelomonocytic leukemia (CMML), juvenile myelomonocytic leukemia (JMML), or large granular lymphocytic leukemia), lymphoma (e.g., AIDS-related lymphoma, cutaneous T-cell lymphoma, Hodgkin lymphoma (e.g., classical Hodgkin lymphoma or nodular lymphocyte-predominant Hodgkin lymphoma), mycosis fungoides, non-Hodgkin lymphoma (e.g., B-cell non-Hodgkin lymphoma (e.g., Burkitt lymphoma, small lymphocytic lymphoma (CLL / SLL), diffuse large B-cell lymphoma, follicular lymphoma, immunoblastic large cell lymphoma, precursor B-lymphoblastic lymphoma, or mantle cell lymphoma) or T-cell non-Hodgkin lymphoma (mycosis fungoides, anaplastic large cell lymphoma, or precursor T-lymphoblastic lymphoma)), primary central nervous system, but are not limited thereto.
[0149] Nucleic Acid Extraction and Processing DNA or RNA may be extracted from tissue samples, biopsy samples, blood samples, or other body fluid samples using any of a variety of techniques known to those of skill in the art (e.g., Example 1 of International Patent Application Publication No. WO 2012 / 092426; Tan, et al. (2009), "DNA, RNA and Protein Extraction: Past and Present", J. Biomed. Biotech. 2009:574398; technical literature for the Maxwell® 16 LEV Blood DNA Kit (Promega Corporation, Madison, WI); and the Maxwell 16 Buccal Swab LEV DNA Purification Kit Technical Manual (Promega Literature #TM333, January 1, 2011, Promega Corporation, Madison, WI)). Protocols for RNA isolation are disclosed, for example, in the Maxwell® 16 Total RNA Purification Kit Technical Bulletin (Promega Literature #TB351, August 2009, Promega Corporation, Madison, WI).
[0150] Typical DNA extraction procedures include, for example, (i) collection of a fluid sample, cell sample, or tissue sample from which DNA is to be extracted, (ii) disruption of cell membranes (i.e., cell lysis), if necessary, to release DNA and other cytoplasmic components, (iii) treatment of the liquid sample or lysed sample with a concentrated salt solution to precipitate proteins, lipids, and RNA, followed by centrifugation to separate the precipitated proteins, lipids, and RNA, and (iv) purification of DNA from the supernatant to remove detergents, proteins, salts, or other reagents used during the cell lysis step.
[0151] Disruption of the cell membrane can be carried out using various mechanical shearing (e.g., French press or fine needle) or sonication techniques. The cell lysis step often involves the use of detergents and surfactants to lyse lipids, cells, and nuclear membranes. In some cases, the lysis step may further include the use of proteases to break down proteins and / or the use of RNases to digest RNA in the sample.
[0152] Examples of suitable techniques for DNA purification include, but are not limited to, (i) precipitation in ice-cold ethanol or isopropanol followed by centrifugation (e.g., DNA precipitation that can be enhanced by increasing the ionic strength by adding sodium acetate), (ii) phenol-chloroform extraction followed by centrifugation to separate the aqueous phase containing nucleic acids from the organic phase containing denatured proteins, and (iii) solid-phase chromatography where nucleic acids adsorb to a solid phase (e.g., silica or others) depending on the pH and salt concentration of the buffer.
[0153] In some cases, cell and histone proteins bound to DNA can be removed by adding proteases, or by precipitating proteins with sodium acetate or ammonium acetate, or through extraction with a phenol-chloroform mixture prior to the DNA precipitation step.
[0154] In some cases, DNA can be extracted using any of various suitable commercially available DNA extraction and purification kits. Examples include, but are not limited to, the QIAamp (for isolation of genomic DNA from human samples) and DNAeasy (for isolation of genomic DNA from animal or plant samples) kits from Qiagen (Germantown, MD), or the Maxwell® and ReliaPrep™ series from Promega (Madison, WI).
[0155] As described above, in some cases, the sample may include formalin fixation (formaldehyde fixation or paraformaldehyde fixation), paraffin embedding (FFPE) tissue preparation. For example, an FFPE sample can be a substrate, e.g., a tissue sample embedded in an FFPE block. Methods for isolating nucleic acids (e.g., DNA) from formaldehyde- or paraformaldehyde-fixed paraffin-embedded (FFPE) tissue are described, for example, in Cronin et al., (2004) Am J Pathol. 164(1):35-42; Masuda et al., (1999) Nucleic Acids Res. 27(22):4436-4443; Specht et al., (2001) Am J Pathol. 158(2):419-429; Ambion RecoverAll (trademark) Total Nucleic Acid Isolation Protocol (Ambion, Catalog No. AM1975, September 2008); Maxwell (registered trademark) 16 FFPE Plus LEV DNA Purification Kit Technical Manual (Promega Literature#TM349, February 2011); E.Z.N.A. (登録商標)It is disclosed in FFPE DNA Kit Handbook (OMEGA bio-tek, Norcross, GA, product numbers D3399-00, D3399-01, D3399-02, June 2009); and QIAamp® DNA FFPE Tissue Handbook (Qiagen, catalog No. 37625, October 2007). For example, the RecoverAll™ Total Nucleic Acid Isolation Kit solubilizes paraffin-embedded samples using xylene at high temperature and captures nucleic acids by passing them through a glass fiber filter. The Maxwell® 16 FFPE Plus LEV DNA Purification Kit, together with the Maxwell® 16 Instrument, is used to purify genomic DNA from 1-10 μm sections of FFPE tissue. DNA is purified using silica-coated paramagnetic particles (PMPs) and eluted with a low elution volume. The E.Z.N.A.® FFPE DNA Kit uses a spin column and buffer system for the isolation of genomic DNA. The QIAamp® DNA FFPE Tissue Kit uses QIAamp® DNA Micro technology for the purification of genomic and mitochondrial DNA.
[0156] In some cases, the disclosed methods may further include determining or obtaining a yield value of nucleic acids extracted from a sample and comparing the determined value to a reference value. For example, if the determined or obtained value is less than the reference value, the nucleic acids may be amplified prior to proceeding with library construction. In some cases, the disclosed methods may further include determining or obtaining a value for the size (or average size) of the nucleic acid fraction in a sample and comparing the determined or obtained value to a reference value, e.g., a size (or average size) of at least 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 base pairs (bps). In some cases, one or more parameters described herein may be adjusted or selected in response to this determination.
[0157] After isolation, the nucleic acid is typically dissolved in a slightly alkaline buffer, such as Tris-EDTA (TE) buffer, or in ultrapure water. In some cases, the isolated nucleic acid (e.g., genomic DNA) can be fractionated or sheared by using any of a variety of techniques known to those of skill in the art. For example, genomic DNA can be fractionated by physical shearing methods, enzymatic cleavage methods, chemical cleavage methods, and other methods well-known to those of skill in the art. Methods for DNA shearing are described, for example, in Example 4 of International Patent Application Publication No. 2012 / 092426. In some cases, alternative methods to DNA shearing can be used to avoid the ligation step during library preparation.
[0158] Library preparation In some cases, nucleic acids isolated from a sample can be used to construct a library (e.g., a nucleic acid library as described herein). In some cases, the nucleic acids are fractionated using any of the methods described above, optionally subjected to repair of strand-end damage, and optionally ligated to synthesize adapters, primers, and / or barcodes (e.g., amplification primers, sequencing adapters, flow cell adapters, substrate adapters, sample barcodes or indexes, and / or unique molecular identifier sequences), size selected (e.g., by preparative gel electrophoresis), and / or amplified (e.g., using PCR, non-PCR amplification techniques, or isothermal amplification techniques). In some cases, the fractionated and adapter-ligated nucleic acid population is used without explicit size selection or amplification prior to hybridization-based selection of target sequences. In some cases, the nucleic acids are amplified by any of a variety of specific or non-specific nucleic acid amplification methods well known to those of skill in the art. In some cases, the nucleic acids are amplified by a whole genome amplification method such as, for example, random primer strand displacement amplification. Examples of nucleic acid library preparation techniques for next generation sequencing are described, for example, in van Dijk, et al. (2014), Exp. Cell Research 322:12-20, and in the Illumina genomic DNA sample preparation kit.
[0159] In some cases, the resulting nucleic acid library can contain all or substantially all of the genomic complexity. The term "substantially all" in this context actually refers to the possibility that there may be some undesirable loss of genomic complexity during the initial steps of the procedure. The methods described herein are also useful when the nucleic acid library is a part of the genome, for example, when the genomic complexity is reduced by design. In some cases, any selected portion of the genome can be used with the methods described herein. For example, in certain embodiments, the entire exome or a subset thereof is isolated. In some cases, the library can contain at least 95%, 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, or 5% of the genomic DNA. In some cases, the library can consist of cDNA copies of genomic DNA that contain at least 95%, 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, or 5% of the genomic DNA. In a particular example, the amount of nucleic acid used to generate the nucleic acid library can be less than 5 micrograms, less than 1 microgram, less than 500 ng, less than 200 ng, less than 100 ng, less than 50 ng, less than 10 ng, less than 5 ng, or less than 1 ng.
[0160] In some cases, a library (e.g., a nucleic acid library) comprises a collection of nucleic acid molecules. As described herein, the nucleic acid molecules of the library can include target nucleic acid molecules (e.g., tumor nucleic acid molecules, reference nucleic acid molecules, and / or control nucleic acid molecules, also referred to herein as the first, second, and / or third nucleic acid molecules, respectively). The nucleic acid molecules of the library can be derived from a single subject or individual. In some cases, the library can include nucleic acid molecules derived from two or more subjects (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30 or more subjects). For example, two or more libraries from different subjects can be combined to form a library having nucleic acid molecules from two or more subjects (the nucleic acid molecules from each subject are optionally ligated to a unique sample barcode corresponding to a particular subject). In some cases, the subject is a human having or at risk of having cancer or a tumor.
[0161] In some cases, a library (or a portion thereof) may include one or more sub-genomic intervals. In some cases, a sub-genomic interval can be a single nucleotide position, e.g., a nucleotide position where a variant at that position is associated (either positively or negatively) with a tumor phenotype. In some cases, a sub-genomic interval includes two or more nucleotide positions. Such examples include sequences of nucleotide positions that are at least 2, 5, 10, 50, 100, 150, 250, or more than 250 in length. A sub-genomic interval can include, for example, one or more entire genes (or portions thereof), one or more exons or coding sequences (or portions thereof), one or more introns (or portions thereof), one or more microsatellite regions (or portions thereof), or any combination thereof. A sub-genomic interval can include all or a portion of a fraction of a naturally occurring nucleic acid molecule, e.g., a genomic DNA molecule. For example, a sub-genomic interval can correspond to a fraction of genomic DNA that is subjected to a sequencing reaction. In some cases, a sub-genomic interval is a contiguous sequence from a genomic source. In some cases, a sub-genomic interval includes sequences that are not contiguous in the genome, e.g., a sub-genomic interval in cDNA can include exon-exon junctions formed as a result of splicing. In some cases, a sub-genomic interval includes a tumor nucleic acid molecule. In some cases, a sub-genomic interval includes a non-tumor nucleic acid molecule.
[0162] Targeting of Loci for Analysis The methods described herein can be used, as described herein, in combination with, or as part of, a method for evaluating a set of target intervals (e.g., target sequences) from a set of genomic loci (e.g., loci or fractions thereof), for example.
[0163] In some cases, the set of genomic loci evaluated by the disclosed methods includes multiple, e.g., genes, in mutant form that are associated with an effect on cell division, proliferation, or survival, or are associated with cancer, e.g., the cancers described herein.
[0164] In some cases, the set of loci evaluated by the disclosed method comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, or more than 100 loci.
[0165] In some cases, the selected locus (also referred to herein as the target locus or target sequence) or a fraction thereof may comprise a target interval that includes a non-coding sequence, a coding sequence, an intronic region, or an intergenic region of the subject genome. For example, the target interval may include a non-coding sequence or a fraction thereof (e.g., a promoter sequence, an enhancer sequence, a 5' untranslated region (5'UTR), a 3' untranslated region (3'UTR), or a fraction thereof), a coding sequence of the fraction, an exon sequence or a fraction thereof, or an intron sequence or a fraction thereof.
[0166] Target capture reagent The methods described herein may include contacting a nucleic acid library with a plurality of target capture reagents to select and capture a plurality of specific target sequences (e.g., gene sequences or fractions thereof) for analysis. In some cases, a target capture reagent (i.e., a molecule that binds to a target molecule, thereby enabling capture of the target molecule) is used to select the region of interest to be analyzed. For example, the target capture reagent can be a bait molecule, such as a nucleic acid molecule (e.g., a DNA molecule or an RNA molecule), that hybridizes to (i.e., is complementary to) the target molecule, thereby enabling capture of the target nucleic acid. In some cases, the target capture reagent, such as a bait molecule (or bait sequence), is a capture oligonucleotide (or capture probe). In some cases, the target nucleic acid is a genomic DNA molecule, an RNA molecule, a cDNA molecule derived from an RNA molecule, a microsatellite DNA sequence, etc. In some cases, the target capture reagent is suitable for solution-phase hybridization to the target. In some cases, the target capture reagent is suitable for solid-phase hybridization to the target. In some cases, the target capture reagent is suitable for both solution-phase hybridization and solid-phase hybridization to the target. The design and construction of the target capture reagent are described in more detail, for example, in International Patent Application Publication No. 2020 / 236941, the entire content of which is incorporated herein by reference.
[0167] The methods described herein provide optimized sequencing of a number of genomic loci (e.g., genes or gene products (e.g., mRNA), microsatellite loci, etc.) from a sample (e.g., a cancer tissue specimen, a liquid biopsy sample, etc.) from one or more subjects by appropriate selection of a target capture reagent for selecting a target nucleic acid molecule to be sequenced. In some cases, the target capture reagent can hybridize to a specific target locus, e.g., a specific target locus or a fraction thereof. In some cases, the target capture reagent can hybridize to a specific group of target loci, e.g., a specific group of loci or a fraction thereof. In some cases, a plurality of target capture reagents can be used, including a mixture of target-specific and / or group-specific target capture reagents.
[0168] In some cases, the number of target capture reagents (e.g., bait molecules) in a plurality of target capture reagents (e.g., a bait set) contacted with a nucleic acid library to capture a plurality of target sequences for nucleic acid sequencing is greater than 10, greater than 50, greater than 100, greater than 200, greater than 300, greater than 400, greater than 500, greater than 600, greater than 700, greater than 800, greater than 900, greater than 1,000, greater than 1,250, greater than 1,500, greater than 1,750, greater than 2,000, greater than 3,000, greater than 4,000, greater than 5,000, greater than 10,000, greater than 25,000, or greater than 50,000.
[0169] In some cases, the full length of the target capture reagent sequence can be from about 70 nucleotides to 1000 nucleotides. In one example, the length of the target capture reagent is about 100 - 300 nucleotides, 110 - 200 nucleotides, or 120 - 170 nucleotides in length. In addition to the above, intermediate oligonucleotide lengths of about 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 300, 400, 500, 600, 700, 800 and 900 nucleotides in length can be used in the methods described herein. In some embodiments, oligonucleotides of about 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220 or 230 bases can be used.
[0170] In some cases, each target capture reagent sequence can include (i) a target-specific capture sequence (e.g., a locus- or microsatellite locus-specific complementary sequence), (ii) an adapter, primer, barcode, and / or unique molecular identifier sequence, and (iii) a universal tail at one or both ends. As used herein, the term "target capture reagent" can refer to the entire target capture reagent oligonucleotide including the target-specific target capture sequence or a target capture reagent oligonucleotide containing the target-specific target capture sequence.
[0171] In some cases, the target-specific capture sequence in the target capture reagent is from about 40 nucleotides to 1000 nucleotides in length. In some cases, the target-specific capture sequence is from about 70 nucleotides to 300 nucleotides in length. In some cases, the target-specific sequence is from about 100 nucleotides to 200 nucleotides in length. In yet other examples, the target-specific sequence is from about 120 nucleotides to 170 nucleotides in length, typically 120 nucleotides in length. In addition to the above, intermediate lengths, for example, target-specific sequences of about 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 300, 400, 500, 600, 700, 800, and 900 nucleotides in length, as well as target-specific sequences of lengths between the above lengths, can also be used in the methods described herein.
[0172] In some cases, the target capture reagent can be designed to select a target region that includes one or more rearrangements, for example, an intron that includes a genomic rearrangement. In such examples, the target capture reagent is designed such that repetitive sequences are masked to enhance the selection efficiency. In these examples where the rearrangement has a known junction sequence, a complementary target capture reagent can be designed to the junction sequence to enhance the selection efficiency.
[0173] In some cases, the disclosed methods can include the use of target capture reagents designed to capture two or more different target categories, each category having a different target capture reagent design strategy. In some cases, the hybridization-based capture methods and target capture reagent compositions disclosed herein provide capture and uniform coverage of a target sequence set while minimizing coverage of genomic sequences outside the targeted sequence set. In some cases, the target sequences can include the entire exome of genomic DNA or a selected subset thereof. In some cases, the target sequences can include, for example, large chromosomal regions (e.g., entire chromosomal arms). The methods and compositions disclosed herein provide different target capture reagents for achieving different sequencing depths and coverage patterns for complex target nucleic acid sequence sets.
[0174] Typically, DNA molecules are used as target capture reagent sequences, although RNA molecules can also be used. In some cases, the DNA molecule target capture reagent can be single-stranded DNA (ssDNA) or double-stranded DNA (dsDNA). In some cases, RNA-DNA duplexes are more stable than DNA-DNA duplexes and thus potentially provide better nucleic acid capture.
[0175] In some cases, the disclosed method includes providing a selected set of nucleic acid molecules (e.g., library catch) captured from one or more nucleic acid libraries. For example, the method includes providing one or more nucleic acid libraries, each containing a plurality of nucleic acid molecules (e.g., a plurality of target nucleic acid molecules and / or reference nucleic acid molecules) extracted from one or more samples from one or more subjects; contacting one or more libraries (e.g., in a solution-based hybridization reaction) with a plurality of target capture reagents (e.g., oligonucleotide target capture reagents), such as 1, 2, 3, 4, 5, or more than 5 target capture reagents, to form a hybridization mixture containing a plurality of target capture reagent / nucleic acid molecule hybrids; and separating the plurality of target capture reagent / nucleic acid molecule hybrids from the hybridization mixture, for example, by contacting the hybridization mixture with a binding entity that allows separation of the plurality of target capture reagent / nucleic acid molecule hybrids from the hybridization mixture, thereby providing a library catch (e.g., a subset of selected or enriched nucleic acid molecules from one or more libraries).
[0176] In some cases, the disclosed method may further include amplifying the library catch (e.g., by performing PCR). In other examples, the library catch is not amplified.
[0177] In some cases, the target capture reagent may be part of a kit that may include instructions, standards, buffers, or enzymes or other reagents as needed.
[0178] Hybridization Conditions As described above, the methods disclosed herein may include contacting a library (e.g., a nucleic acid library) with a plurality of target capture reagents to contact a selected library target nucleic acid sequence (i.e., a library catch). The contacting step may be performed, for example, by solution-based hybridization. In some cases, the method includes repeating the hybridization step with respect to one or more additional solution-based hybridizations. In some cases, the method further includes subjecting the library catch to one or more additional solution-based hybridizations with the same or a different set of target capture reagents.
[0179] In some cases, the contacting step is performed using a solid support, e.g., an array. Solid supports suitable for hybridization are disclosed, for example, in Albert, T.J. et al. (2007) Nat. Methods 4(11):903-5; Hodges, E. et al. (2007) Nat. Genet. 39(12):1522-7; and Okou, D.T. et al. (2007) Nat. Methods 4(11):907-9, the contents of which are incorporated herein by reference in their entirety.
[0180] Hybridization methods that can be adapted for use in the methods herein are described in the art, for example, as described in International Patent Application Publication No. 2012 / 092426. Methods for hybridizing target capture reagents to a plurality of target nucleic acids are described in more detail, for example, in International Patent Application Publication No. 2020 / 236941, the entire contents of which are incorporated herein by reference.
[0181] Sequencing method The methods and systems disclosed herein are used in combination with, or as part of, a method or system for sequencing nucleic acids (e.g., a next-generation sequencing system) to generate a plurality of sequence reads that overlap one or more loci within a sub-genomic interval in a sample, thereby, for example, determining gene alleles at a plurality of loci. As used herein, "next-generation sequencing" (or "NGS"), which may also be referred to as "ultra-parallel sequencing," determines the nucleotide sequences of individual nucleic acid molecules (e.g., as in single-molecule sequencing) or clonally amplified proxies of individual nucleic acid molecules in a high-throughput manner (e.g., where more than 10 3 , 10 4 , 10 5 , or 10 5 supramolecular molecules are sequenced simultaneously).
[0182] Next-generation sequencing methods are known in the art and are described, for example, in Metzker, M. (2010) Nature Biotechnology Reviews 11:31-46, which is incorporated herein by reference. Other examples of sequencing methods suitable for use when implementing the methods and systems disclosed herein are described, for example, in International Patent Application Publication No. 2012 / 092426. In some cases, sequencing may include, for example, whole-genome sequencing (WGS), whole-exome sequencing, targeted sequencing, or direct sequencing. In some cases, sequencing may be performed using, for example, Sanger sequencing. In some cases, sequencing may include paired-end sequencing techniques that enable both ends of a fraction to be sequenced and that generate high-quality alignable sequence data, for example, for the detection of genomic rearrangements, repetitive sequence elements, gene fusions, and novel transcripts.
[0183] The disclosed methods and systems can be implemented using sequencing platforms such as the Roche454, Illumina Solexa, ABI-SOLiD, ION Torrent, Complete Genomics, Pacific Bioscience, Helicos, and / or Polonator platforms. In some cases, sequencing can include Illumina MiSeq sequencing. In some cases, sequencing can include Illumina HiSeq sequencing. In some cases, sequencing can include Illumina NovaSeq sequencing. Optimized methods for sequencing a number of target genomic loci in nucleic acids extracted from a sample are described in detail, for example, in International Patent Application Publication No. 2020 / 236941, the entire content of which is incorporated herein by reference.
[0184] In certain instances, the disclosed method includes: (a) obtaining from a sample a library comprising a plurality of normal and / or tumor nucleic acid molecules; (b) contacting the library simultaneously or sequentially with one, two, three, four, five, or more than five target capture reagents under conditions that allow hybridization of the target capture reagents to the target nucleic acid molecules, thereby providing a selected set of captured normal and / or tumor nucleic acid molecules (i.e., a library catch); (c) separating a selected subset of nucleic acid molecules (e.g., the library catch) from the hybridization mixture, for example, by contacting the hybridization mixture with a binding entity that allows separation of the target capture reagent / nucleic acid molecule hybrids from the hybridization mixture; (d) sequencing the library catch to obtain a plurality of reads (e.g., sequence reads) that overlap one or more target intervals (e.g., one or more target sequences) from a library catch that may contain mutations (or variations), for example, variant sequences that include somatic mutations or germline mutations; (e) aligning the sequence reads using an alignment method described elsewhere herein; and / or (f) assigning nucleotide values to nucleotide positions within the target interval from one or more of the plurality of sequence reads (e.g., calling mutations using the Bayesian method or other methods described herein).
[0185] In some instances, obtaining array reads for one or more target intervals can include sequencing at least 1, at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1,000, at least 1,250, at least 1,500, at least 1,750, at least 2,000, at least 2,250, at least 2,500, at least 2,750, at least 3,000, at least 3,500, at least 4,000, at least 4,500, or at least 5,000 loci, such as genomic loci, loci, microsatellite loci, etc. In some cases, obtaining array reads for one or more target intervals can include sequencing the target interval for any number of loci within the ranges described in this paragraph, e.g., for at least 2,850 loci.
[0186] In some cases, obtaining sequence reads for one or more target regions involves sequencing the target regions using a sequencing method that provides a sequence read length (or average sequence read length) of at least 20 bases, at least 30 bases, at least 40 bases, at least 50 bases, at least 60 bases, at least 70 bases, at least 80 bases, at least 90 bases, at least 100 bases, at least 120 bases, at least 140 bases, at least 160 bases, at least 180 bases, at least 200 bases, at least 220 bases, at least 240 bases, at least 260 bases, at least 280 bases, at least 300 bases, at least 320 bases, at least 340 bases, at least 360 bases, at least 380 bases, or at least 400 bases. In some cases, obtaining sequence reads for one or more target regions can involve sequencing the target regions using a sequencing method that provides a sequence read length (or average sequence read length) of any number of bases within the ranges described in this paragraph, for example, a sequence read length (or average sequence read length) of 56 bases.
[0187] In some cases, obtaining sequence reads for one or more target regions can involve sequencing at an average coverage (or depth) of at least 100× or more. In some cases, obtaining sequence reads for one or more target regions involves sequencing at an average coverage (or depth) of at least 100×, at least 150×, at least 200×, at least 250×, at least 500×, at least 750×, at least 1,000×, at least 1,500×, at least 2,000×, at least 2,500×, at least 3,000×, at least 3,500×, at least 4,000×, at least 4,500×, at least 5,000×, at least 5,500×, or at least 6,000× or more. In some cases, obtaining sequence reads for one or more target regions can involve sequencing at an average coverage (or depth) having any value within the range of values described in this paragraph, for example, at least 160×.
[0188] In some cases, obtaining array reads for one or more target intervals involves sequencing at an average sequencing depth having any value in the range of at least 100× to at least 6,000× for at least about 90%, 92%, 94%, 95%, 96%, 97%, 98%, or greater than 99% of the sequenced loci. For example, in some cases, obtaining reads for a target interval involves sequencing at an average sequencing depth of at least 125× for at least 99% of the sequenced loci. As another example, in some cases, obtaining reads for a target interval involves sequencing at an average sequencing depth of at least 4,100× for at least 95% of the sequenced loci.
[0189] In some cases, the relative abundance of nucleic acid species in a library can be estimated by counting the relative number of appearances of their homologous sequences (e.g., the number of sequence reads for a given homologous sequence) in the data generated by a sequencing experiment.
[0190] In some cases, the disclosed methods and systems provide nucleotide sequences for a set of target intervals (e.g., loci), as described herein. In certain instances, the sequences are provided without using methods that include a matched normal control (e.g., wild-type control) and / or a matched tumor control (e.g., primary to metastatic).
[0191] In some cases, as used herein, the level of sequencing depth (e.g., the X-fold level of sequencing depth) refers to the number of reads (e.g., unique reads) obtained after detection and removal of duplicate reads (e.g., PCR duplicate reads). In other examples, duplicate reads are evaluated, for example, to assist in the detection of copy number variations (CNAs).
[0192] Alignment Alignment is the process of matching reads to a location, such as a genomic location or locus. In some cases, NGS reads can be aligned to a known reference sequence (e.g., a wild-type sequence). In some cases, NGS reads can be de novo assembled. Methods for sequence alignment of NGS reads are described, for example, in Trapnell, C. and Salzberg, S.L. Nature Biotech., 2009, 27:455-457. Examples of de novo sequence assembly are described, for example, in Warren R. et al., Bioinformatics, 2007, 23:500-501, Butler J. et al., Genome Res., 2008, 18:810-820, and Zerbino D.R. and Birney E., Genome Res., 2008, 18:821-829. Optimization of sequence alignment is described in the art, as described, for example, in International Patent Application Publication No. 2012 / 092426. Additional explanations of sequence alignment methods are described in more detail, for example, by International Patent Application Publication No. 2020 / 236941, the entire contents of which are incorporated herein by reference.
[0193] Misalignment (e.g., the placement of base pairs from short reads at inaccurate locations within the genome), such as the reads of alternative alleles being shifted from the histogram peaks of alternative allele reads, can lead to a decrease in the sensitivity of mutation detection due to the sequence context (e.g., the presence of repetitive sequences) around the actual cancer mutation. Other examples of sequence contexts that can cause misalignment include short tandem repeats, interspersed repetitive sequences, low complexity regions, insertions - deletions (indels), and paralogs. If a problematic sequence situation occurs when there is no actual mutation, misalignment can introduce reads of "mutated" allele artifacts by placing reads of the actual reference genome base sequence in the wrong location. Since mutation calling algorithms for multi - gene analysis must be sensitive even to low - abundance mutations, sequence misalignment can increase the false positive discovery rate and / or decrease specificity.
[0194] In some cases, the methods and systems disclosed herein may integrate the use of multiple individually tuned alignment methods or algorithms to optimize base calling performance in sequencing methods, particularly methods that rely on massively parallel sequencing (MPS) of a large number of diverse genetic events at a large number of diverse genomic loci. In some cases, the disclosed methods and systems may include the use of one or more global alignment algorithms. In some cases, the disclosed methods and systems may include the use of one or more local alignment algorithms. Examples of alignment algorithms that may be used include, but are not limited to, the Burrows-Wheeler Alignment (BWA) software bundle (e.g., Li, et al. (2009), "Fast and accurate short read alignment with Burrows-Wheeler Transform," Bioinformatics 25:1754-60; Li et al. (2010), "Fast and accurate long read alignment with Burrows-Wheeler Transform," Bioinformatics epub. PMID:20080505), the Smith-Waterman algorithm (e.g., see Smith et al. (1981), "Identification of common molecular subsequences," J. Molecular Biology 147(1):195-197), the Striped Smith-Waterman algorithm (e.g., see Farrar (2007), "Striped Smith-Waterman speeds database searches six times compared to other SIMD implementations," Bioinformatics 23(2):156-161), the Needleman-Wunsch algorithm (Needleman et al. (1970) "A general method applicable to the search for similarities in the amino acid sequences of two proteins," J. Molecular Biology 48(3):443-53), or any combination thereof.
[0195] In some cases, the methods and systems disclosed herein may also include the use of sequence assembly algorithms, such as the Arachne sequence assembly algorithm (see, e.g., Batzoglou et al. (2002), "ARACHNE: A Whole Genome Shotgun Assembler", Genome Res., 12:177-189).
[0196] In some cases, the alignment methods used to analyze sequence reads are not individually customized or adjusted for the detection of different variants (e.g., point mutations, insertions, deletions, etc.) at different genomic loci. In some cases, different alignment methods that are individually customized or adjusted for the detection of at least a subset of different variants detected at different genomic loci are used to analyze the reads. In some cases, different alignment methods that are individually customized or adjusted for detecting each different variant at different genomic loci are used to analyze the reads. In some cases, the adjustment can be a function of one or more of (i) the locus to be sequenced (e.g., a locus, a microsatellite locus, or other target interval), (ii) the tumor type associated with the sample, (iii) the variant to be sequenced, or (iv) a characteristic of the sample or subject. The selection or use of alignment conditions that are individually adjusted for some specific target intervals to be sequenced allows for the optimization of speed, sensitivity, and specificity. This method is particularly effective when the alignment of reads for a relatively large number of diverse target intervals is optimized. In some cases, the method includes the combined use of an alignment method optimized for rearrangements and other alignment methods optimized for target intervals not associated with rearrangements.
[0197] In some cases, the methods disclosed herein further comprise selecting or using an alignment method for analyzing, e.g., aligning, array reads, the alignment method being a function of, selected according to, or optimized for one or more of: (i) the tumor type, e.g., the tumor type in the sample, (ii) the location of the target interval to be sequenced (e.g., locus), (iii) the type of variant within the target interval to be sequenced (e.g., point mutation, insertion, deletion, substitution, copy number variation (CNV), rearrangement, or fusion), (iv) the site being analyzed (e.g., nucleotide position), (v) the type of sample (e.g., the samples described herein), and / or (vi) the adjacent sequences within or near the target interval being evaluated (e.g., according to its expected tendency to misalign the target interval due to the presence of repetitive sequences within or near the target interval).
[0198] In some cases, the methods disclosed herein enable rapid and efficient alignment of difficult reads, e.g., reads having rearrangements. Thus, in some cases where the reads for a target interval include nucleotide positions associated with a rearrangement, e.g., a translocation, the method may be appropriately adjusted and may include using an alignment method that includes: (i) selecting a rearrangement reference sequence for alignment with the reads, the rearrangement reference sequence being aligned with a rearrangement (in some cases, the reference sequence is not identical to a genomic rearrangement), and (ii) comparing, e.g., aligning, the reads with the rearrangement reference sequence.
[0199] In some cases, alternative methods can be used to align problematic reads. These methods are particularly effective when the alignment of reads to a relatively large number of diverse target intervals is optimized. As an example, a method for analyzing a sample can include: (i) performing a comparison of reads using a first parameter set (e.g., an alignment comparison), such as using a first mapping algorithm or by comparison with a first reference sequence, to determine whether the read meets a first alignment criterion (e.g., the read can be aligned with the first reference sequence with, for example, less than a certain number of mismatches); (ii) if the read does not meet the first alignment criterion, performing a second alignment comparison using a second parameter set (e.g., using a second mapping algorithm or by comparison with a second reference sequence); and (iii) optionally, determining whether the read meets a second criterion (e.g., the read can be aligned with the second reference sequence with, for example, less than a certain number of mismatches), where the second parameter set includes the use of a second reference sequence that is more likely to result in an alignment (e.g., rearrangement, insertion, deletion, or translocation) of the read with respect to the variant compared to the first parameter set.
[0200] In some cases, the alignment of array reads in the disclosed method can be combined with the variant calling methods described elsewhere in this specification. As discussed herein, a decrease in sensitivity to detect actual variants can be addressed by evaluating (either manually or in an automated fashion) the quality of the alignment around the predicted variant sites of the gene or genomic locus being analyzed (e.g., locus). In some cases, the sites to be evaluated can be obtained from a database of the human genome (e.g., HG19 human reference genome) or cancer mutations (e.g., COSMIC). Regions identified as problematic can be repaired using an algorithm selected to provide better performance in the relevant sequence context by alignment optimization (or realignment) using a slower but more accurate alignment algorithm such as the Smith-Waterman alignment. If the general alignment algorithm cannot improve the problem, a customized alignment approach can be created, for example, by adjusting the maximum different mismatch penalty parameter for genes with a high likelihood of containing substitutions, by adjusting specific mismatch penalty parameters based on specific variant types common to a particular tumor type (e.g., C→T in melanoma), or by adjusting specific mismatch penalty parameters based on specific variant types common to a particular sample type (e.g., substitutions common to FFPE).
[0201] A decrease in the specificity of the evaluated target interval (increase in false positive rate) due to misalignment can be evaluated by manual or automated inspection of all variant calls in the sequencing data. Regions found to be prone to false variant calls due to misalignment can be subjected to the alignment improvement discussed above. If algorithmic improvements are not possible, "variants" from the problematic regions can be classified or screened from a panel of target loci.
[0202] Variant calling Base calling refers to the raw output of a sequencing device, e.g., the determined sequence of nucleotides in an oligonucleotide molecule. Variant calling refers to the process of selecting a nucleotide value, e.g., A, G, T, or C, for a given nucleotide position being sequenced. Typically, a sequence read (or base call) for a position provides two or more values, e.g., some reads indicate T and some indicate G. Variant calling is the process of assigning the correct nucleotide value, e.g., one of those values, to the sequence. Although called "variant" calling, it can be applied to assign nucleotide values to any nucleotide position, e.g., a position corresponding to a variant allele, a wild-type allele, an allele not characterized as variant or wild-type, or a position not characterized by variability.
[0203] In some cases, the disclosed methods can include the use of customized or adjusted variant calling algorithms or parameters to optimize performance when applied to sequencing data, particularly in methods that rely on massively parallel sequencing (MPS) of multiple diverse genetic events at multiple diverse genomic loci (e.g., loci, microsatellite regions, etc.) in a sample, e.g., a sample from a subject having cancer. Optimization of variant calling is described in the art, e.g., as described in International Patent Application Publication No. 2012 / 092426.
[0204] Methods for variant calling can include one or more of the following: making independent calls based on information at each position in a reference sequence (e.g., examining sequence reads; examining base calls and quality scores; calculating the probability of the observed bases and quality scores given a potential genotype; and assigning a genotype (e.g., using Bayes' rule)); removing false positives (e.g., using a depth threshold to reject SNPs with a read depth much lower or higher than expected; local realignment to remove false positives due to small indels); performing an analysis based on linkage disequilibrium (LD) / complementation to improve the calls.
[0205] The equations used to calculate the genotype likelihoods related to specific genotypes and positions are described, for example, in Li H. and Durbin R., Bioinformatics, 2010; 26(5):589-95. The prior prediction for a specific mutation in a specific cancer type can be used when evaluating samples from that cancer type. Such likelihoods can be obtained from public databases of cancer mutations, such as the Catalogue of Somatic Mutation in Cancer (COSMIC), HGMD (Human Gene Mutation Database), The SNP Consortium, Breast Cancer Mutation Data Base (BIC), and Breast Cancer Gene Database (BCGD).
[0206] Examples of LD / imputation-based analysis are described, for example, in Browning, B.L. and Yu, Z., Am. J. Hum. Genet. 2009, 85(6):847-61. Examples of low-coverage SNP calling methods are described, for example, in Li, Y., et al., Annu. Rev. Genomics Hum. Genet. 2009, 10:387-406.
[0207] After alignment, detection of substitutions can be performed using a calling method (e.g., a Bayesian mutation calling method), which is applied to each base of each target interval, e.g., the exons of the gene being evaluated or other loci, and the presence of alternative alleles is observed. This method compares the probability of observing read data in the presence of a mutation to the probability of observing read data in the presence of only base calling errors. If this comparison strongly enough supports the presence of a mutation, the mutation can be called.
[0208] The advantage of the Bayesian variant detection method is that the comparison between the probability of the presence of a mutation and only the probability of a base calling error can be weighted by the prior expectation of the presence of a mutation at that site. If several reads of an alternative allele are observed at a site that frequently mutates for a given cancer type, the presence of a mutation can be reliably called even if the amount of evidence for the mutation does not meet the normal threshold. This flexibility can then be used to increase the detection sensitivity for rarer mutations / lower purity samples, or to make the test more robust to a decrease in read coverage. The likelihood that a random base pair in the genome is mutated in cancer is about 1e-6. For example, the likelihood of a specific mutation occurring at many sites in a typical multi-gene cancer genome panel can be orders of magnitude higher. These likelihoods can be derived from public databases of cancer mutations (e.g., COSMIC).
[0209] Indel calling is the process of finding bases in sequencing data that differ from the reference sequence by an insertion or deletion, typically including a related confidence score or statistical evidence metric. Methods for indel calling can include steps of identifying candidate indels, calculating genotype likelihoods by local realignment, and performing LD-based genotype inference and calling. Typically, the Bayesian method is used to obtain potential indel candidates, which are then tested with the reference sequence within a Bayesian framework.
[0210] Algorithms for generating candidate indels are described, for example, in McKenna, A. et al., Genome Res. 2010;20(9):1297-303; Ye, K., et al., Bioinformatics, 2009;25(21):2865-71; Lunter, G. and Goodson, M., Genome Res., 2011;21(6):936-9; and Li, H., et al. (2009), Bioinformatics 25(16):2078-9.
[0211] Examples of methods for making indel calls and generating individual-level genotype likelihoods include, for example, the Dindel algorithm (Albers, C. A., et al., Genome Res. 2011; 21(6): 961-73). For example, the Bayesian EM algorithm can be used to analyze reads, make initial indel calls, generate genotype likelihoods for each candidate indel, and subsequently perform genotype attribution using, for example, QCALL (Le S. Q. and Durbin R. Genome Res. 2011; 21(6): 952-60). Parameters such as prior expectations for observing indels can be adjusted (e.g., increased or decreased) based on the size or position of the indel.
[0212] Methods have been developed to address limited deviations from 50% or 100% allele frequencies for the analysis of cancer DNA. (See, for example, SNVMix - Bioinformatics. 2010 March 15; 26(6): 730-736). However, the methods disclosed herein allow for the consideration of frequencies (or allele fractions) in the range of 1% to 100% (i.e., allele fractions in the range of 0.01 to 1.0), and in particular, the possibility of the presence of mutant alleles at levels below 50%. This approach is particularly important, for example, for the detection of mutations in low-purity FFPE samples of native (multiclonal) tumor DNA.
[0213] In some cases, the variant calling methods used to analyze array reads are not customized or adjusted individually for the detection of different variants at different genomic loci. In some cases, different variant calling methods that are customized or fine-tuned individually for at least a subset of different variants detected at different genomic loci are used. In some cases, different variant calling methods that are customized or fine-tuned individually for each different variant detected at each different genomic locus are used. The customization or adjustment can be based on one or more of the factors described herein, such as the type of cancer in the sample, the gene or locus in which the sequenced target interval is located, or the variant being sequenced. This selection or use of variant calling methods customized or fine-tuned individually for the number of sequenced target intervals enables optimization of the speed, sensitivity, and specificity of variant calling.
[0214] In some cases, nucleotide values are assigned to the nucleotide positions of each of X unique target intervals using a unique variant calling method, where X is at least 2, at least 3, at least 4, at least 5, at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 1000, at least 1500, at least 2000, at least 2500, at least 3000, at least 3500, at least 4000, at least 4500, at least 5000 or more. The calling methods are different and can be unique, for example, by depending on different Bayesian prior values.
[0215] In some cases, assigning the nucleotide value is a function of the expected value or a value representing a variant, e.g., a read indicating a mutation, at that nucleotide position in that type of tumor, e.g., prior to observation (e.g., in the literature).
[0216] In some cases, the method involves assigning nucleotide values (e.g., variant calls) for at least 10, 20, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1,000 nucleotide positions, where each assignment is a function of an expected value (e.g., in the literature) prior to observing a read indicating a variant, e.g., a mutation, at that nucleotide position in the type of tumor or a unique value (as contrasted with values of other assignments) representing it.
[0217] In some cases, assigning the nucleotide value is a function of a set of values representing the probability of observing a read indicating the variant at that nucleotide position when the variant is present in the sample at a particular frequency (e.g., 1%, 5%, 10%, etc.) and / or when the variant is not present (e.g., observed in a read due to base calling error only).
[0218] In some cases, the variant calling method described herein involves: (a) for each nucleotide position in each of the X target intervals, obtaining: (i) a first value that is an expected value (e.g., in the literature) prior to observing a read indicating a variant, e.g., a mutation, at that nucleotide position in a tumor of type X or a value representing it, and (ii) a set of second values representing the likelihood of observing a read indicating the variant at that nucleotide position when the variant is present in the sample at a certain frequency (e.g., 1%, 5%, 10%, etc.) and / or when the variant is not present (e.g., observed in a read due to base calling error alone); and (b) in response to the values, for example, by weighting a comparison between values in a second set that uses the first value, such as by the Bayesian method described herein, assigning a nucleotide value (e.g., calling a mutation) from the read to each of the nucleotide positions, thereby analyzing the sample.
[0219] Additional explanations of the variant calling method are described in detail, for example, in International Patent Application Publication No. 2020 / 236941, the entire content of which is incorporated herein by reference.
[0220] System Also disclosed herein is a system designed to perform any of the disclosed methods for determining the tumor DNA fraction (e.g., ctDNA fraction) in a sample from a subject. The system may include, for example, one or more processors and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to receive sequence read data in a plurality of sequence reads derived from a sample from a subject, determine a variant allele frequency (VAF) in one or more variants detected in the sample based on the sequence read data, generate an empirical distribution of tumor DNA fraction values in response to the determined VAFs in the one or more variants, fit a model to the empirical distribution of tumor DNA fraction values, and determine the tumor DNA fraction in the sample based on the model.
[0221] In some cases, the disclosed system may further include a sequencer, e.g., a next-generation sequencer (also referred to as a massively parallel sequencer). Examples of next-generation (or massively parallel) sequencing platforms include, but are not limited to, Roche / 454's Genome Sequencer (GS) FLX system, Illumina / Solexa's Genome Analyzer (GA), Illumina's HiSeq® 2500, HiSeq® 3000, HiSeq® 4000, and NovaSeq® 6000 sequencing systems, Life / APG's Support Oligonucleotide Ligation Detection (SOLiD) system, Polonator's G.007 system, Helicos BioSciences' HeliScope Gene Sequencing system, ThermoFisher Scientific's Ion Torrent Genexus system, or Pacific Biosciences' PacBio® RS system.
[0222] In some cases, the disclosed system can be used to determine the tumor DNA fraction in any of the various samples described herein (e.g., tissue samples, biopsy samples, hematological samples, or liquid biopsy samples) derived from a subject.
[0223] In some cases, the plurality of loci for which sequencing data is processed to determine a tumor DNA fraction (e.g., a ctDNA fraction) includes 10 to 20 loci, 10 to 40 loci, 10 to 60 loci, 10 to 80 loci, 10 to 100 loci, 10 to 150 loci, 10 to 200 loci, 10 to 250 loci, 10 to 300 loci, 10 to 350 loci, 10 to 400 loci, 10 to 450 loci, 10 to 500 loci, 20 to 40 loci, 20 to 60 loci, 20 to 80 loci, 20 to 100 loci, 20 to 150 loci, 20 to 200 loci, 20 to 250 loci, 20 to 300 loci, 20 to 350 loci, 20 to 400 loci, 20 to 500 loci, 40 to 60 loci, 40 to 80 loci, 40 to 100 loci, 40 to 150 loci, 40 to 200 loci, 40 to 250 loci, 40 to 300 loci, 40 to 350 loci, 40 to 400 loci, 40 to 500 loci, 60 to 80 loci, 60 to 100 loci, 60 to 150 loci, 60 to 200 loci, 60 to 250 loci, 60 to 300 loci, 60 to 350 loci, 60 to 400 loci, 60 to 500 loci, 80 to 100 loci, 80 to 150 loci, 80 to 200 loci, 80 to 250 loci, 80 to 300 loci, 80 to 350 loci, 80 to 400 loci, 80 to 500 loci, 100 to 150 loci, 100 to 200 loci, 100 to 250 loci, 100 to 300 loci, 100 to 350 loci, 100 to 400 loci, 100 to 500 loci, 150 to 200 loci, 150 to 250 loci, 150 to 300 loci, 150 to 350 loci, 150 to 400 loci, 150 to 500 loci, 200 to 250 loci, 200 to 300 loci, 200 to 350 loci, 200 to 400 loci, 200 to 500 loci, 250 to 300 loci, 250 to 350 loci, 250 to 400 loci, 250 to 500 loci, 300 to 350 loci, 300 to 400 loci, 300 to 500 loci, 350 to 400 loci, 350 to 500 loci, or 400 to 500 loci.
[0224] In some cases, nucleic acid sequence data is obtained using next-generation sequencing techniques (also referred to as massively parallel sequencing techniques) having read lengths of less than 400 bases, less than 300 bases, less than 200 bases, less than 150 bases, less than 100 bases, less than 90 bases, less than 80 bases, less than 70 bases, less than 60 bases, less than 50 bases, less than 40 bases, or less than 30 bases.
[0225] In some cases, determination of the tumor DNA fraction (e.g., ctDNA fraction) is used to select, initiate, adjust, or terminate treatment of cancer in a subject (e.g., a patient) from whom the sample is derived, as described elsewhere herein.
[0226] In some cases, the disclosed system may further include a sample processing and library preparation workstation, a microplate handling robot, a fluid dispensing system, a temperature control module, an environmental control chamber, an additional data storage module, a data communication module (e.g., Bluetooth®, WiFi, intranet, or Internet communication hardware and related software), a display module, one or more local and / or cloud-based software packages (e.g., an instrument / system control software package, a sequencing data analysis software package), etc., or any combination thereof. In some instances, the system may include or be part of a computer system or computer network as described elsewhere herein.
[0227] Computer Systems and Networks FIG. 2 shows an example of a computing device or system according to one embodiment. Device 200 can be a host computer connected to a network. Device 200 can be a client computer or a server. As shown in FIG. 2, device 200 can be any suitable type of microprocessor-based device such as a personal computer, a workstation, a server, or a handheld computing device (portable electronic device) such as a phone or a tablet. The device can include, for example, one or more processors 210, an input device 220, an output device 230, a memory or storage device 240, a communication device 260, and a nucleic acid sequencer 270. Software 250 resident in the memory or storage device 240 can include, for example, an operating system and software for implementing the methods described herein. The input device 220 and the output device 230 can generally correspond to those described herein, be connectable to the computer, or be integrated with the computer.
[0228] The input device 220 can be any suitable device that provides input such as a touch screen, a keyboard or keypad, a mouse, or a voice recognition device. The output device 230 can be any suitable device that provides output such as a touch screen, a tactile device, or a speaker.
[0229] Storage 240 can be any suitable device that provides storage (e.g., electrical, magnetic, or optical memory including RAM (volatile and non-volatile), cache, hard drive, or removable storage disk). Communication device 260 can include any suitable device that can transmit and receive signals over a network such as a network interface chip or device. The components of the computer can be connected in any suitable manner, for example, via wired media (e.g., physical system bus 280, Ethernet connection, or any other wired transfer technology) or wirelessly (e.g., Bluetooth®, Wi-Fi®, or any other wireless technology).
[0230] Software module 250 is stored as executable instructions in storage 240 and can be executed by processor 210 and can include, for example, an operating system and / or a process that embodies the functionality of the methods of the present disclosure (e.g., embodied in the devices described above).
[0231] Software module 250 can also be stored and / or transferred in any non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device (e.g., those described herein), and can fetch instructions associated with the software from the instruction execution system, apparatus, or device and execute the instructions. In the context of the present disclosure, the computer-readable storage medium can be any medium such as storage 240, and can include or store a process for use by or in connection with an instruction execution system, apparatus, or device. Examples of computer-readable storage media can include hard drives, flash drives, and memory units such as distribution modules that operate as a single functional unit. Also, the various processes described herein can be embodied as modules configured to operate according to the above-described embodiments and techniques. Further, although the processes can be shown and / or described separately, those skilled in the art will understand that the above processes can be routines or modules within other processes.
[0232] Software module 250 can also be propagated in any transmission medium for use by or in connection with an instruction execution system, apparatus, or device such as those described above, and can fetch instructions associated with the software from the instruction execution system, apparatus, or device and execute the instructions. In the context of the present disclosure, the transmission medium can be any medium and can communicate, propagate, or transmit transmission programming for use by or in connection with an instruction execution system, apparatus, or device. The transmission-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, or infrared wired or wireless propagation media.
[0233] Device 200 may be connected to a network (e.g., network 304 as shown in FIG. 3 and / or described below) which can be any suitable type of interconnected communication system. The network may implement any suitable communication protocol and may be protected by any suitable security protocol. The network may include network links of any suitable arrangement that can implement the transmission and reception of network signals, such as a wireless network connection (T1 or T3 line), a cable network, DSL, or a telephone line.
[0234] Device 200 may be implemented using any operating system, e.g., an operating system suitable for operating on a network. Software module 250 can be written in any suitable programming language such as C, C++, Java, or Python. In various embodiments, the application software embodying the functions of the present disclosure can be deployed in different configurations (e.g., in a client / server arrangement or via a web browser as a web-based application or web service). In some embodiments, the operating system is executed by one or more processors, e.g., processor 210.
[0235] Device 200 can further include a sequencer 270 which can be any suitable nucleic acid sequencing device.
[0236] Figure 3 shows an example of a computing system according to one embodiment. In system 300, device 200 (e.g., as shown in FIG. 2 described above) is connected to a network 304 that is also connected to device 306. In some embodiments, device 306 is a sequencer. Exemplary sequencers include, but are not limited to, the Roche / 454 Genome Sequencer (GS) FLX System, the Illumina / Solexa Genome Analyzer (GA), the Illumina HiSeq® 2500, HiSeq® 3000, HiSeq® 4000, and NovaSeq® 6000 Sequencing Systems, the Life / APG Support Oligonucleotide Ligation Detection (SOLiD) system, the Polonator G.007 system, the Helicos BioSciences HeliScope Gene Sequencing system, or the Pacific Biosciences PacBio® RS system.
[0237] Devices 200 and 306 can communicate using an appropriate communication interface via a network 304 such as, for example, a local area network (LAN), a virtual private network (VPN), or the Internet. In some embodiments, network 304 can be, for example, the Internet, an intranet, a virtual private network, a cloud network, a wired network, or a wireless network. Devices 200 and 306 can communicate partially or entirely via wireless or wired communication such as Ethernet, IEEE802.11b wireless, etc. Additionally, devices 200 and 306 can communicate via a second network such as a mobile / cellular network using, for example, a suitable communication interface. Communication between device 200 and 306 can further include or communicate with various servers such as a mail server, a mobile server, a media server, a telephone server, etc. In some embodiments, devices 200 and 306 can communicate directly (instead of or in addition to communication via network 304) via wireless or wired communication such as Ethernet, IEEE802.11b wireless, etc. In some embodiments, devices 200 and 306 can be directly connected or communicate via communication 308 that can occur via a network (e.g., network 304).
[0238] One or all of devices 200 and 306 generally includes logic (e.g., http web server logic) accessed from a local or remote database or other source of data and content to provide and / or receive information via network 304 according to various examples described herein, or is programmed to format data.
Example
[0239] Example 1 - Simulation of VAF vs ctDNA Fraction and Comparison with Actual Patient Data As described above, the methods disclosed herein may include the following concepts.
[0240] I. The fraction of the cfDNA sample consisting of ctDNA can be represented using the following equation derived based on the physical characteristics of the cfDNA sample. TIFF2025524649000010.tif16170Where ρ is the tumor purity (as used herein, the terms "tumor purity" and "cellular purity" are equivalent), and ψ is the tumor average ploidy (i.e., the average ploidy across the genome in tumor cells).
[0241] II. The variant allele frequency (VAF) of a variant (e.g., a somatic variant, or a somatic short variant) can be represented using the following equation derived based on the physical characteristics of the somatic variant. TIFF2025524649000011.tif16170Where ρ is the tumor purity, C is the copy number at the genomic position of the variant, and V is the number of variant alleles.
[0242] III. As shown below, by solving equations (5) and (6), the relationship between the ctDNA fraction and the variant allele frequency, the copy number state of the genomic position containing a given variant, and the tumor average ploidy can be derived as functions of each other. ctDNA_fraction = f(somatic VAF, C, V, ψ) (7) somatic VAF = f(ctDNA_fraction, C, V, ψ) (8)
[0243] IV. Equations (7) and (8) show that for any given sample with known C and V in a somatic variant and the tumor average ploidy ψ in the sample, the ctDNA fraction can be derived from the somatic variant allele frequency (somatic VAF). Alternatively, the somatic VAF can be derived from the ctDNA fraction. Equations (7) and (8) can be solved to derive an explicit relationship between the ctDNA fraction and the somatic VAF given by the following. TIFF2025524649000012.tif21170
[0244] This conclusion is verified by the following analysis. A set of approximately 2200 samples with known tumor average ploidy ψ, and known C and V in the top somatic variants (e.g., somatic variants with the highest allele frequency) in each sample were randomly selected from the genomics database. Using a random number generator, ctDNA fractions in the range of 0.04 to 0.96 were simulated. Each of the simulated ctDNA fraction values was substituted into Equation (4) together with a set of values of C, V, and ψ randomly selected from one of the approximately 2200 samples to calculate the expected somatic VAF for a given value of ctDNA fraction, C, V, and ψ. The results of the simulated somatic VAF versus ctDNA fraction data are plotted in FIGS. 4 (where the traces for variants with different copy number values of 1, 2, 3, or 4 are overlaid) and 5 (showing simulated data for variants with copy number 1 in the upper left panel, simulated data for variants with copy number 2 in the upper right panel, simulated data for variants with copy number 3 in the lower left panel, and simulated data for variants with copy number 4 in the lower right panel). The solid lines in FIGS. 4 to 7 indicate the case where VAF is equal to the ctDNA fraction. The dashed lines in FIGS. 4 to 7 indicate the case where VAF is equal to 1 / 2 of the ctDNA fraction. As can be seen from the figures, in most cases, the VAF of the top variant (the variant with the highest VAF) is equal to either the tumor DNA fraction or half of the tumor fraction. This explains the two peaks evident in the probability density plot shown in FIG. 8 and the "branching" observed in FIGS. 4 to 7. These variants that follow the relationship VAF = ctDNA fraction typically have variant copy allele numbers close to the tumor average ploidy. Variants that follow the relationship VAF = 1 / 2(ctDNA fraction) typically have variant allele numbers equal to approximately half of the tumor average ploidy (e.g., if the tumor average ploidy is 2, but only one allele has a mutation).
[0245] The simulation results were found to closely match what was observed for actual patient data, as shown in FIGS. 6 and 7. FIG. 6 provides a plot of the somatic VAF measured for patient samples against the ctDNA fraction determined during copy number modeling, where again, the traces of variants with different copy number values of 1, 2, 3, or 4 are overlaid. FIG. 7 shows the same patient data plotted for variants with a copy number of 1 in the upper left panel, patient data for variants with a copy number of 2 in the upper right panel, patient data for variants with a copy number of 3 in the lower left panel, and patient data for variants with a copy number of 4 in the lower right panel.
[0246] The agreement between the simulated and actual results demonstrates that the ctDNA fraction, somatic allele frequencies, C, V, and ψ, as measured by sequencing-based assays, actually follow the principles described by the above equations, and by leveraging data on the known values of C, V, and ψ in a large number of patient samples included in existing genomic databases, it is shown that the ctDNA fraction can be accurately inferred based on the measured somatic variant allele frequencies.
[0247] FIG. 8 provides a non-limiting example of an output probability density plot as a function of the possible ctDNA fraction values in a sample for a given value of the observed maximum somatic VAF. A reference table containing sets of (C, V, ψ) was constructed for somatic variants showing the maximum VAF in a plurality of tissue samples for which good copy number alteration (CNA) modeling data was available. To estimate the tumor DNA fraction of the liquid sample (i.e., the cfDNA fraction), the somatic variant showing the maximum VAF was identified for the sample, and Equation (9) was used to calculate the ctDNA fraction value based on the tissue reference table of (C, V, ψ), thereby forming an empirical distribution of all possible ctDNA fraction values. The empirical distribution was fitted to a non-parametric probability density model and used to output the most likely ctDNA fraction value as well as the upper and lower limits of the ctDNA fraction (indicated by the vertical dashed lines in FIG. 8).
[0248] Example 2 - Exemplary Process Workflow One non - limiting example of a process workflow for implementing the disclosed method can include the following steps.
[0249] 1. Select a representative set of patient samples for which reliable CNA modeling data is available from the variant database. Non - limiting examples of selection criteria can include, but are not limited to, disease ontology, tissue type, clinical data, and other considerations.
[0250] 2. For the somatic variants with the highest allele frequencies in the selected sample set, construct a reference table of sets of (C, V, ψ). An exemplary incomplete table is shown in Table 1 (for illustrative purposes only).
[0251] 3. To estimate the ctDNA fraction of a liquid biopsy sample, identify all somatic variants in the sample and find the highest allele frequency among these somatic variants. This is the somatic VAF value used to infer the ctDNA fraction.
[0252] 4. Use the above formula (9) to calculate all possible ctDNA fraction values based on the reference table of (C, V, ψ) values, thereby forming an empirical distribution of possible ctDNA fractions.
[0253] 5. Fit a model, such as a non - parametric probability density model, to the distribution. Identify the most likely ctDNA fraction based on the non - parametric probability density model and identify the upper and lower bounds of the ctDNA fraction based on the desired confidence interval. [Table 1]
[0254] Exemplary Embodiments Exemplary embodiments of the methods and systems described herein can include the following. 1. Prepare a plurality of nucleic acid molecules obtained from a sample from a subject; Ligate one or more adapters onto one or more of the nucleic acid molecules from the plurality of nucleic acid molecules; Amplify one or more ligated nucleic acid molecules from the plurality of nucleic acid molecules; Capture the amplified nucleic acid molecules from the amplified nucleic acid molecules; Sequence the captured nucleic acid molecules with a sequencer to obtain a plurality of sequence reads representing the captured nucleic acid molecules; In one or more processors, receive sequence read data regarding the plurality of sequence reads; Using one or more processors, determine a variant allele frequency (VAF) regarding one or more variants detected in the sample based on the sequence read data; Using one or more processors, generate an empirical distribution of tumor DNA fraction values according to the determined VAFs in the one or more variants; Using one or more processors, fit a model to the empirical distribution of tumor DNA fraction values; Determine the tumor DNA fraction in the sample based on the model; A method comprising the above steps. 2. The method of item 1, further comprising determining a confidence interval in the tumor DNA fraction based on the model. 3. The method of item 1 or 2, wherein the one or more variants include one or more somatic short variants. 4. The method of item 3, wherein the one or more somatic short variants are known not to be associated with clonal hematopoiesis of indeterminate potential (CHIP). 5. Generating the empirical distribution of tumor DNA fraction values includes calculating tumor DNA fraction values based on known copy numbers in the one or more variants, the determined VAFs in the one or more variants, and corresponding known tumor average ploidies in a plurality of historical subject samples having known VAFs in the one or more variants that are substantially the same as the determined VAFs in the one or more variants. The method according to any one of items 1 to 4. 6. Generating an empirical distribution of tumor DNA fraction values includes pre-calculating tumor DNA fraction values based on known copy numbers in one or more variants and corresponding known tumor average ploidies in a plurality of historical subject samples having a range of VAF values in one or more variants, and selecting a subset of the pre-calculated tumor DNA fraction values corresponding to samples having known VAFs in one or more variants that are substantially the same as the determined VAFs in one or more variants, the method of any one of items 1 to 4. 7. Tumor DNA fraction values are calculated or pre-calculated by solving a set of equations representing the relationships between (i) tumor DNA fraction, tumor purity, and tumor average ploidy, and (ii) somatic VAF, tumor purity, copy number at the genomic location of one or more variants, and the number of variant alleles at each of the one or more variants, based on known VAFs in one or more variants, known copy numbers in one or more variants, and corresponding known tumor average ploidies in a plurality of historical subject samples, thereby deriving the relationship in tumor DNA fraction as a function of somatic VAF, tumor average ploidy, copy number at the genomic location of one or more variants, and the number of variant alleles for one or more variants, excluding tumor purity, the method of item 5 or item 6. 8. The plurality of historical subject samples include solid biopsy samples, liquid biopsy samples, or any combination thereof, the method of any one of items 5 to 7. 9. The plurality of historical subject samples include cancer samples, the method of any one of items 5 to 8. 10. The plurality of historical subject samples include samples of a single type of cancer, the method of item 9. 11. The plurality of historical subject samples include samples of multiple types of cancer, the method of item 9. 12. The plurality of historical target samples include acute lymphoblastic leukemia (Philadelphia chromosome positive) samples, acute lymphoblastic leukemia (precursor B cell) samples, acute myeloid leukemia (FLT3+) samples, acute myeloid leukemia (with IDH2 mutation) samples, anaplastic large cell lymphoma samples, basal cell carcinoma samples, B-cell chronic lymphocytic leukemia samples, bladder cancer samples, breast cancer (HER2 overexpression / amplification) samples, breast cancer (HER2+) samples, breast cancer (HR+, HER2-) samples, cervical cancer samples, cholangiocarcinoma samples, chronic lymphocytic leukemia samples, chronic lymphocytic leukemia (with 17p deletion) samples, chronic myeloid leukemia samples, chronic myeloid leukemia (Philadelphia chromosome positive) samples, classical Hodgkin lymphoma samples, colorectal cancer samples, colorectal cancer (dMMR and MSI-H) samples, colorectal cancer (KRAS wild type) samples, cryopyrin-associated periodic syndrome samples, cutaneous T-cell lymphoma samples, dermatofibrosarcoma protuberans samples, diffuse large B-cell lymphoma samples, fallopian tube cancer samples, follicular B-cell non-Hodgkin lymphoma samples, follicular lymphoma samples, gastric cancer samples, gastric cancer (HER2+) samples, gastroesophageal junction (GEJ) adenocarcinoma samples, gastrointestinal stromal tumor samples, gastrointestinal stromal tumor samples, gastrointestinal stromal tumor (KIT+) samples, giant cell tumor of bone samples, glioblastoma samples, polyangiitis granulomatosa samples, head and neck squamous cell carcinoma samples, hepatocellular carcinoma samples, Hodgkin lymphoma samples, juvenile idiopathic arthritis samples, lupus erythematosus samples, mantle cell lymphoma samples, medullary thyroid cancer samples, melanoma samples, melanoma with BRAF V600 mutation samples, melanoma with BRAF V600E or V600K mutation samples, Merkel cell carcinoma samples, Castleman disease samples, multiple blood samples including Philadelphia chromosome positive ALL and CML, multiple myeloma samples, myelofibrosis samples, non-Hodgkin lymphoma samples, inoperable subependymal giant cell astrocytoma samples associated with tuberous sclerosis, non-small cell lung cancer samples, non-small cell lung cancer (ALK+) samples, non-small cell lung cancer (PD-L1+) samples, non-small cell lung cancer (with ALK fusion or ROS1 gene alteration) samples, non-small cell lung cancer (with BRAF V600E mutation) samples, non-small cell lung cancer samples (with EGFR exon 19 deletion or exon 21 substitution (L858R) mutation), non-small cell lung cancer samples (EGFRThe method according to any one of items 5 to 11, comprising a T790M mutation), an ovarian cancer sample, an ovarian cancer sample (with a BRCA mutation), a pancreatic cancer sample, a pancreatic cancer sample, a gastrointestinal cancer sample, a lung-derived neuroendocrine tumor sample, a pediatric neuroblastoma sample, a peripheral T cell lymphoma sample, a prostate cancer sample, a peritoneal cancer sample, a renal cell cancer sample, a rheumatoid arthritis sample, a small lymphocyte lymphoma sample, a soft tissue sarcoma sample, a solid tumor (MSI-H / dMMR) sample, a squamous cell carcinoma sample of the head and neck, a squamous non-small cell lung cancer sample, a thyroid cancer sample, a thyroid cancer sample, a urothelial cancer sample, a urothelial cancer sample, a Waldenström macroglobulinemia sample, or any combination thereof. 13. The method according to any one of items 1 to 12, wherein the model is a non-parametric probability density model. 14. The method according to any one of items 1 to 13, wherein the determined tumor DNA fraction in the sample is the most likely tumor DNA fraction. 15. The method according to any one of items 1 to 14, wherein the determined tumor DNA fraction in the cfDNA sample is the mean, median, or mode of the dominant peak in the empirical distribution of tumor DNA fraction values. 16. The method according to any one of items 1 to 15, wherein the sample comprises DNA extracted from a blood sample, a plasma sample, a cerebrospinal fluid sample, a pleural effusion sample, a sputum sample, a fecal sample, a urine sample, or a saliva sample. 17. The method according to any one of items 1 to 16, wherein the subject is suspected of having cancer or is determined to have cancer. 18. Cancer includes acute lymphoblastic leukemia (Philadelphia chromosome positive), acute lymphoblastic leukemia (precursor B cell), acute myeloid leukemia (FLT3+), acute myeloid leukemia (with IDH2 mutation), anaplastic large cell lymphoma, basal cell carcinoma, B-cell chronic lymphocytic leukemia, bladder cancer, breast cancer (HER2 overexpression / amplification), breast cancer (HER2+), breast cancer (HR+, HER2-), cervical cancer, cholangiocarcinoma, chronic lymphocytic leukemia, chronic lymphocytic leukemia (with 17p deletion), chronic myelogenous leukemia, chronic myelogenous leukemia (Philadelphia chromosome positive), classical Hodgkin lymphoma, colorectal cancer, colorectal cancer (dMMR / MSI-H), colorectal cancer (KRAS wild type), cryopyrin-associated periodic syndrome, cutaneous T-cell lymphoma, precursor cutaneous fibrosarcoma, diffuse large B-cell lymphoma, fallopian tube cancer, follicular B-cell non-Hodgkin lymphoma, follicular lymphoma, gastric cancer, gastric cancer (HER2+), gastroesophageal junction (GEJ) adenocarcinoma, gastrointestinal stromal tumor, gastrointestinal stromal tumor (KIT+), giant cell tumor of bone, glioblastoma, granulomatosis with polyangiitis, head and neck squamous cell carcinoma, hepatocellular carcinoma, Hodgkin lymphoma, juvenile idiopathic arthritis, lupus erythematosus, mantle cell lymphoma, medullary thyroid carcinoma, melanoma, melanoma with BRAF V600 mutation, melanoma with BRAF V600E or V600K mutation, Merkel cell carcinoma, multicentric Castleman disease, multiple hematological malignancies including Philadelphia chromosome positive ALL and CML, multiple myeloma, myelofibrosis, non-Hodgkin lymphoma, unresectable subependymal giant cell astrocytoma associated with tuberous sclerosis, non-small cell lung cancer, non-small cell lung cancer (ALK+), non-small cell lung cancer (PD-L1+), non-small cell lung cancer (with ALK fusion or ROS1 gene alteration), non-small cell lung cancer (with BRAF V600E mutation), non-small cell lung cancer (with EGFR exon 19 deletion or exon 21 substitution (L858R) mutation), non-small cell lung cancer (EGFRThe method of item 17, which is for ovarian cancer (with T790M mutation), ovarian cancer (with BRCA mutation), pancreatic cancer, pancreatic cancer, gastrointestinal cancer or lung-derived neuroendocrine tumor, pediatric neuroblastoma, peripheral T-cell lymphoma, peritoneal cancer, prostate cancer, renal cell cancer, rheumatoid arthritis, small lymphocyte lymphoma, soft tissue sarcoma, solid tumor (MSI-H / dMMR), head and neck squamous cell carcinoma, non-small cell lung cancer, squamous cell carcinoma, thyroid cancer, thyroid cancer, urothelial cancer, urothelial cancer, or Waldenström macroglobulinemia. 19. The method of item 18, further comprising treating the subject with an anti-cancer therapy. 20. The method of item 19, wherein the anti-cancer therapy comprises a targeted anti-cancer therapy. 21. Targeted cancer therapies include abemaciclib (Verzenio), abiraterone acetate (Zytiga), acalabrutinib (Calquence), trastuzumab emtansine (Kadcyla), afatinib dimaleate (Gilotrif), aldesleukin (Proleukin), alectinib (Alecensa), alemtuzumab (Campath), alitretinoin (Panretin), alpelisib (Piqray), amivantamab-vmjw (Rybrevant), anastrozole (Arimidex), apalutamide (Erleada), asimicinib hydrochloride (Scemblix), atezolizumab (Tecentriq), avapritinib (Ayvakit), avelumab (Bavencio), axicabtagene ciloleucel (Yescarta), axitinib (Inlyta), belantamab mafodotin-blmf (Blenrep), belimumab (Benlysta), belinostat (Beleodaq), belzutifan (Welireg), bevacizumab (Avastin), bexarotene (Targretin), binimetinib (Mektovi), blinatumomab (Blincyto), bortezomib (Velcade), bosutinib (Bosulif), brentuximab vedotin (Adcetris), brexucabtagene autoleucel (Tecartus), brigatinib (Alunbrig), cabazitaxel (Jevtana), cabozantinib (Cabometyx), cabozantinib (Cabometyx, Cometriq), canakinumab (Ilaris), capmatinib hydrochloride (Tabrecta), carfilzomib (Kyprolis), cemiplimab-rwlc (Libtayo), ceritinib (LDK378 / Zykadia), cetuximab (Erbitux), cobimetinib (Cotellic), copanlisib hydrochloride (Aliqopa), crizotinib (Xalkori), dabrafenib (Tafinlar), dacomitinib (Vizimpro), daratumumab (Darzalex), daratumumab and hyaluronidase-fihj (DarzalexFaspro, darolutamide (Nubeqa), dasatinib (Sprycel), denileukin diftitox (Ontak), denosumab (Xgeva), dinutuximab (Unituxin), dostarlimab-gxly (Jemperli), durvalumab (Imfinzi), duvelisib (Copiktra), elotuzumab (Empliciti), enasidenib mesylate (Idhifa), encorafenib (Braftovi), enfortumab vedotin-ejfv (Padcev), entrectinib (Rozlytrek), enzalutamide (Xtandi), erdafitinib (Balversa), erlotinib (Tarceva), everolimus (Afinitor), exemestane (Aromasin), fam-trastuzumab deruxtecan-nxki (Enhertu), fedratinib hydrochloride (Inrebic), fulvestrant (Faslodex), gefitinib (Iressa), gemtuzumab ozogamicin (Mylotarg), gilteritinib (Xospata), glasdegib maleate (Daurismo), hyaluronidase-zzxf (Phesgo), ibrutinib (Imbruvica), ibritumomab tiuxetan (Zevalin), idecabtagene vicleucel (Abecma), idelalisib (Zydelig), imatinib mesylate (Gleevec), infigratinib phosphate (Truseltiq), inotuzumab ozogamicin (Besponsa), iobenguane I131 (Azedra), ipilimumab (Yervoy), isatuximab-irfc (Sarclisa),ivosidenib (Tibsovo), ixazomib citrate (Ninlaro), lanreotide acetate (SomatulineDepot), Lapatinib (Tykerb), Larotrectinib Sulfate (Vitrakvi), Lenvatinib Mesylate (Lenvima), Letrozole (Femara), Isocabtagene Autoleucel (Breyanzi), Loncastuximab Tesirine-lpyl (Zynlonta), Lorlatinib (Lorbrena), Lutetium Lu177-dotatate (Lutathera), Margetuximab-cmkb (Margenza), Midostaurin (Rydapt), Mobocertinib Succinate (Exkivity), Mogamulizumab-kpkc (Poteligeo), Moxetumomab Pasudotox-tdfk (Lumoxiti), Naxitamab-gqgk (Danyelza), Necitumumab (Portrazza), Neratinib Maleate (Nerlynx), Nilotinib (Tasigna), Niraparib Tosylate Monohydrate (Zejula), Nivolumab (Opdivo), Obinutuzumab (Gazyva), Ofatumumab (Arzerra), Olaparib (Lynparza), Olalatumab (Lartruvo), Osimertinib (Tagrisso), Palbociclib (Ibrance), Panitumumab (Vectibix), Panobinostat (Farydak), Pazopanib (Votrient), Pembrolizumab (Keytruda), Pemigatinib (Pemazyre), Pertuzumab (Perjeta), Pegaspargase Hydrochloride (Turalio), Polatuzumab Vedotin-piiq (Polivy), Ponatinib Hydrochloride (Iclusig), Pralatrexate (Folotyn), Pralsetinib (Gavreto), Radium 223 Dichloride (Xofigo), Ramucirumab (Cyramza), Regorafenib (Stivarga), Ribociclib (Kisqali), Ripretinib (Qinlock), Rituxan (Rituxan), Rituximab and Hyaluronidase Human (RituxanThe method of item 20, comprising Hycela, romidepsin (Istodax), rucaparib (Rubraca), ruxolitinib phosphate (Jakafi), sacituzumab govitecan-fujii (Trodelvy), selinexor, selinexor (Xpovio), selpercatinib (Retevmo), selumetinib sulfate (Koselugo), siltuximab (Sylvant), sipuleucel-T (Provenge), sirolimus protein-bound particles (Fyarro), sonidegib (Odomzo), sorafenib (Nexavar), sotorasib (Lumakras), sunitinib (Sutent), tafasitamab-cxix (Monjuvi), tagraxofusp-erzs (Elzonris), talazoparib tosylate (Talzenna), tamoxifen (Nolvadex), tazemetostat hydrobromide (Tazverik), tebentafusp-teb (Kimmtrak), temsirolimus (Torisel), tepotinib hydrochloride (Tepmetko), tisagenlecleucel (Kymriah), tisotumab vedotin-tftv (Tivdak), tocilizumab (Actemra), tofacitinib (Xeljanz), tositumomab (Bexxar), trametinib (Mekinist), trastuzumab (Herceptin), tretinoin (Vesanoid), tivozanib hydrochloride (Fotivda), tremifen (Fareston), tucatinib (Tukysa), umbralisib tosylate (Ukoniq), vandetanib (Caprelsa), vemurafenib (Zelboraf), venetoclax (Venclexta), vismodegib (Erivedge), vorinostat (Zolinza), zanubrutinib (Brukinsa), ziv-aflibercept (Zaltrap), or any combination thereof. 22. The method of any one of items 1 to 21, further comprising obtaining a sample from a subject. 23. The method of any one of items 1 to 22, wherein the sample comprises a tissue biopsy sample, a liquid biopsy sample, or a normal control. 24. The method of item 23, wherein the sample is a liquid biopsy sample and comprises blood, plasma, cerebrospinal fluid, sputum, feces, urine, or saliva. 25. The sample is a liquid biopsy sample and contains circulating tumor cells (CTC), and the method of item 23. 26. The sample is a liquid biopsy sample and contains cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), or any combination thereof, and the method of item 23. 27. The plurality of nucleic acid molecules includes a mixture of tumor nucleic acid molecules and non-tumor nucleic acid molecules, and the method of any one of items 1 to 26. 28. The tumor nucleic acid molecules are derived from the tumor part of a heterogeneous tissue biopsy sample, and the non-tumor nucleic acid molecules are derived from the normal part of the heterogeneous tissue biopsy sample, and the method of item 27. 29. The sample includes a liquid biopsy sample, the tumor nucleic acid molecules are derived from the circulating tumor DNA (ctDNA) fraction of the liquid biopsy sample, and the non-tumor nucleic acid molecules are derived from the non-tumor cell-free DNA (cfDNA) fraction of the liquid biopsy sample, and the method of item 27. 30. The one or more adapters include amplification primers, flow cell adapter sequences, substrate adapter sequences, or sample index sequences, and the method of any one of items 1 to 29. 31. The captured nucleic acid molecules are captured from nucleic acid molecules amplified by hybridization to one or more bait molecules, and the method of any one of items 1 to 30. 32. The one or more bait molecules include one or more nucleic acid molecules, and each nucleic acid molecule includes a region complementary to a region of the captured nucleic acid molecules, and the method of item 31. 33. Amplifying the nucleic acid molecules includes performing polymerase chain reaction (PCR) amplification techniques, non-PCR amplification techniques, or isothermal amplification techniques, and the method of any one of items 1 to 32. 34. Sequencing includes the use of massively parallel sequencing (MPS) techniques, whole genome sequencing (WGS), whole exome sequencing, targeted sequencing, direct sequencing, or Sanger sequencing techniques, and the method of any one of items 1 to 33. 35. Sequencing includes massively parallel sequencing, and the massively parallel sequencing technique includes next-generation sequencing (NGS), and the method of item 34. 36. The method of any one of items 1 to 35, wherein the sequencer includes a next-generation sequencer. 37. The method of any one of claims 1 to 36, wherein one or more of the plurality of array determination reads overlap with one or more loci within one or more sub-genomic intervals in the sample. 38. The method of claim 37, wherein the one or more loci comprise 10 to 20 loci, 10 to 40 loci, 10 to 60 loci, 10 to 80 loci, 10 to 100 loci, 10 to 150 loci, 10 to 200 loci, 10 to 250 loci, 10 to 300 loci, 10 to 350 loci, 10 to 400 loci, 10 to 450 loci, 10 to 500 loci, 20 to 40 loci, 20 to 60 loci, 20 to 80 loci, 20 to 100 loci, 20 to 150 loci, 20 to 200 loci, 20 to 250 loci, 20 to 300 loci, 20 to 350 loci, 20 to 400 loci, 20 to 500 loci, 40 to 60 loci, 40 to 80 loci, 40 to 100 loci, 40 to 150 loci, 40 to 200 loci, 40 to 250 loci, 40 to 300 loci, 40 to 350 loci, 40 to 400 loci, 40 to 500 loci, 60 to 80 loci, 60 to 100 loci, 60 to 150 loci, 60 to 200 loci, 60 to 250 loci, 60 to 300 loci, 60 to 350 loci, 60 to 400 loci, 60 to 500 loci, 80 to 100 loci, 80 to 150 loci, 80 to 200 loci, 80 to 250 loci, 80 to 300 loci, 80 to 350 loci, 80 to 400 loci, 80 to 500 loci, 100 to 150 loci, 100 to 200 loci, 100 to 250 loci, 100 to 300 loci, 100 to 350 loci, 100 to 400 loci, 100 to 500 loci, 150 to 200 loci, 150 to 250 loci, 150 to 300 loci, 150 to 350 loci, 150 to 400 loci, 150 to 500 loci, 200 to 250 loci, 200 to 300 loci, 200 to 350 loci, 200 to 400 loci, 200 to 500 loci, 250 to 300 loci, 250 to 350 loci, 250 to 400 loci, 250 to 500 loci, 300 to 350 loci, 300 to 400 loci, 300 to 500 loci, 350 to 400 loci, 350 to 500 loci, or 400 to 500 loci.39. One or more loci are ABL1, ACVR1B, AKT1, AKT2, AKT3, ALK, ALOX12B, AMER1, APC, AR, ARAF, ARFRP1, ARID1A, ASXL1, ATM, ATR, ATRX, AURKA, AURKB, AXL1, AXL, BAP1, BARD1, BCL2, BCL2L1, BCL2L2, BCL6, BCOR, BCORL1, BCR, BRAF, BRCA1, BRCA2, BRD4, BRIP1, BTG1, BTG2, BTK, CALR, CARD11, CASP8, CBFB, CBL, CCND1, CCND2, CCND3, CCNE1, CD22, CD274, CD70, CD74, CD79A, CD79B, CDC73, CDH1, CDK12, CDK4, CDK6, CDK8, CDKN1A, CDKN1B, CDKN2A, CDKN2B, CDKN2C, CEBPA, CHEK1, CHEK2, CIC, CREBBP, CRKL, CSF1R, CSF3R, CTCF, CTNNA1, CTNNB1, CUL3, CUL4A, CXCR4, CYP17A1, DAXX, DDR1, DDR2, DIS3, DNMT3A, DOT1L, EED, EGFR, EMSY (C11orf30), EP300, EPHA3, EPHB1, EPHB4, ERBB2, ERBB3, ERBB4, ERCC4, ERG, ERRFI1, ESR1, ETV4, ETV5, ETV6, EWSR1, EZH2, EZR, FAM46C, FANCA, FANCC, FANCG, FANCL, FAS, FBXW7, FGF10, FGF12, FGF14, FGF19, FGF23, FGF3, FGF4, FGF6, FGFR1, FGFR2, FGFR3, FGFR4, FH, FLCN, FLT1, FLT3, FOXL2, FUBP1, GABRA6, GATA3, GATA4, GATA6, GID4 (C17orf39), GNA11, GNA13, GNAQ, GNAS, GRM3, GSK3B, H3F3A, HDAC1, HGF, HNF1A, HRAS, HSD3B1, ID3, IDH1, IDH2, IGF1R, IKBKE, IKZF1, INPP4B, IRF2, IRF4, IRS2, JAK1, JAK2, JAK3, JUN, KDM5A, KDM5C, KDM6A, KDR, KEAP1, KEL, KIT, KLHL6, KMT2A (MLL), KMT2D (MLL2), KRAS,The method of claim 37 or claim 38, comprising LTK, LYN, MAF, MAP2K1, MAP2K2, MAP2K4, MAP3K1, MAP3K13, MAPK1, MCL1, MDM2, MDM4, MED12, MEF2B, MEN1, MERTK, MET, MITF, MKNK1, MLH1, MPL, MRE11A, MSH2, MSH3, MSH6, MST1R, MTAP, MTOR, MUTYH, MYB, MYC, MYCL, MYCN, MYD88, NBN, NF1, NF2, NFE2L2, NFKBIA, NKX2-1, NOTCH1, NOTCH2, NOTCH3, NPM1, NRAS, NT5C2, NTRK1, NTRK2, NTRK3, NUTM1, P2RY8, PALB2, PARK2, PARP1, PARP2, PARP3, PAX5, PBRM1, PDCD1, PDCD1LG2, PDGFRA, PDGFRB, PDK1, PIK3C2B, PIK3C2G, PIK3CA, PIK3CB, PIK3R1, PIM1, PMS2, POLD1, POLE, PPARG, PPP2R1A, PPP2R2A, PRDM1, PRKAR1A, PRKCI, PTCH1, PTEN, PTPN11, PTPRO, QKI, RAC1, RAD21, RAD51, RAD51B, RAD51C, RAD51D, RAD52, RAD54L, RAF1, RARA, RB1, RBM10, REL, RET, RICTOR, RNF43, ROS1, RPTOR, RSPO2, SDC4, SDHA, SDHB, SDHC, SDHD, SETD2, SF3B1, SGK1, SLC34A2, SMAD2, SMAD4, SMARCA4, SMARCB1, SMO, SNCAIP, SOCS1, SOX2, SOX9, SPEN, SPOP, SRC, STAG2, STAT3, STK11, SUFU, SYK, TBX3, TEK, TERC, TERT, TET2, TGFBR2, TIPARP, TMPRSS2, TNFAIP3, TNFRSF14, TP53, TSC1, TSC2, TYRO3, U2AF1, VEGFA, VHL, WHSC1, WHSC1L1, WT1, XPO1, XRCC2, ZNF217, ZNF703, or any combination thereof., 40. The method of claim 37 or claim 38, wherein one or more loci comprise ABL, ALK, ALL, B4GALNT1, BAFF, BCL2, BRAF, BRCA, BTK, CD19, CD20, CD3, CD30, CD319, CD38, CD52, CDK4, CDK6, CML, CRACC, CS1, CTLA-4, dMMR, EGFR, ERBB1, ERBB2, FGFR1-3, FLT3, GD2, HDAC, HER1, HER2, HR, IDH2, IL-1β, IL-6, IL-6R, JAK1, JAK2, JAK3, KIT, KRAS, MEK, MET, MSI-H, mTOR, PARP, PD-1, PDGFR, PDGFRα, PDGFRβ, PD-L1, PI3Kδ, PIGF, PTCH, RAF, RANKL, RET, ROS1, SLAMF7, VEGF, VEGFA, VEGFB, or any combination thereof. 41. The method of any one of claims 1 to 40, further comprising generating, by one or more processors, a report indicating the determined tumor DNA fraction of a sample. 42. The method of claim 41, further comprising transmitting the report to a healthcare provider. 43. The method of claim 42, wherein the report is transmitted via a computer network or a peer-to-peer connection. 44. A method for determining the tumor DNA fraction in a cell-free DNA sample from a subject, comprising: receiving, in one or more processors, sequence read data in a plurality of sequence reads derived from a cell-free DNA (cfDNA) sample from the subject; using one or more processors to determine, based on the sequence read data, the variant allele frequency (VAF) in one or more variants detected in the cfDNA sample; using one or more processors to generate an empirical distribution of tumor DNA fraction values according to the VAF determined for one or more variants; using one or more processors to fit a model to the empirical distribution of tumor DNA fraction values; determining the tumor DNA fraction in the cfDNA sample based on the model; and comprising. The method of claim 44, further comprising determining a confidence interval in the tumor DNA fraction based on a model. The method of claim 44 or 45, wherein one or more variants comprise one or more short variants. The method of claim 46, wherein one or more short variants comprise one or more somatic short variants. The method of claim 47, wherein it is known that one or more somatic short variants are not associated with clonal hematopoiesis of indeterminate potential (CHIP). The method of any one of claims 44 to 48, wherein generating an empirical distribution of tumor DNA fraction values comprises calculating tumor DNA fraction values based on known copy numbers in one or more variants, the determined VAF in one or more variants, and corresponding known tumor average ploidies in a plurality of historical subject samples having known VAFs in one or more variants that are substantially the same as the determined VAF in one or more variants. The method of any one of claims 44 to 48, wherein generating an empirical distribution of tumor DNA fraction values comprises pre-calculating tumor DNA fraction values based on known copy numbers in one or more variants and corresponding known tumor average ploidies in a plurality of historical subject samples having a range of VAF values in one or more variants, and selecting a subset of the pre-calculated tumor DNA fraction values corresponding to samples having known VAFs in one or more variants that are substantially the same as the determined VAF in one or more variants. The method of claim 49 or 50, wherein the tumor DNA fraction value is calculated or selected with respect to the variant showing the highest VAF in the cfDNA sample from the subject. The method of claim 49 or 50, wherein the tumor DNA fraction value is calculated or selected with respect to a ranked set of two or more variants showing the highest ranked VAFs in the cfDNA sample from the subject. The method of claim 49 or 50, wherein the tumor DNA fraction value is calculated or selected with respect to a predetermined set of two or more variants detected in the cfDNA sample from the subject. 54. The tumor DNA fraction value is the method of item 49 or item 50, calculated or selected with respect to a predetermined set of two or more variants detected in a cfDNA sample from a subject containing known driver mutations. 55. The tumor DNA fraction value is the method of item 49 or item 50, calculated or selected with respect to all variants detected in a cfDNA sample from a subject. 56. The tumor DNA fraction value is the method of any one of items 49 to 55, calculated or pre-calculated based on the known VAF in one or more variants, the known copy number in one or more variants, and the corresponding known tumor average ploidy in a plurality of historical subject samples. 57. The tumor DNA fraction value is calculated or pre-calculated based on the known VAF in one or more variants, the known copy number in one or more variants, and the corresponding known tumor average ploidy in a plurality of historical subject samples, by solving a set of equations representing (i) the relationship between tumor DNA fraction, tumor purity, and tumor average ploidy, and (ii) the relationship between somatic VAF, tumor purity, copy number at the genomic location of one or more variants, and the number of variant alleles in each of one or more variants, thereby eliminating tumor purity and deriving the relationship in tumor DNA fraction as a function of somatic VAF, tumor average ploidy, copy number at the genomic location of one or more variants, and the number of variant alleles for one or more variants. The method of any one of items 49 to 56. 58. The tumor DNA fraction value is based on the known VAF in one or more variants, the known copy number in one or more variants, and the corresponding known tumor average ploidy in a plurality of historical subject samples. Making the tumor DNA fraction equal to the product of tumor purity and tumor average ploidy divided by the sum of the product of tumor purity and tumor average ploidy and the product of 2 and the magnitude obtained by subtracting tumor purity from 1, the first equation; Divide the product of the tumor purity and the number of variant alleles in each of one or more variants by the sum of the product of the tumor purity and the copy number at the genomic position of the one or more variants and the product of two and the magnitude obtained by subtracting the tumor purity from one, and make the somatic VAF equal, with a second equation, Calculated or pre-calculated by solving a set of equations including Excluding tumor purity, divide the tumor average ploidy by the magnitude equal to the sum of the ratio of the number of variant alleles in one or more variants to the somatic VAF in each of the one or more variants subtracted from the copy number at the genomic position of the one or more variants from the tumor average ploidy, and derive a relationship that makes the tumor DNA fraction equal, The method according to any one of items 49 to 57. 59. The tumor DNA fraction value is based on the known VAF in one or more variants, the known copy number in one or more variants, and the corresponding known tumor average ploidy in a plurality of historical subject samples, that is, Calculated or pre-calculated by solving TIFF2025524649000014.tif29170, thereby excluding ρ, TIFF2025524649000015.tif20170 where ρ is the tumor purity, ψ is the tumor average ploidy, C is the copy number at the genomic position of one or more variants, and V is the number of variant alleles in each of the one or more variants, The method according to any one of items 49 to 58. 60. The method according to any one of items 49 to 59, wherein the plurality of historical subject samples include solid biopsy samples, liquid biopsy samples, or any combination thereof. 61. The method according to any one of items 49 to 60, wherein the plurality of historical subject samples include cancer samples. 62. The method according to item 61, wherein the plurality of historical subject samples include samples of a single type of cancer. 63. The method according to item 62, wherein the plurality of historical subject samples include samples of a plurality of types of cancer. 64. The plurality of historical subject samples are from the method of any one of items 49 to 63, including a bladder cancer sample, a breast cancer sample, a colorectal cancer sample, an endometrial cancer sample, a kidney cancer sample, a leukemia sample, a liver cancer sample, a lung cancer sample, a melanoma sample, a non-Hodgkin lymphoma sample, a pancreatic cancer sample, a prostate cancer sample, a thyroid cancer sample, or any combination thereof. 65. The plurality of historical subject samples include acute lymphoblastic leukemia (Philadelphia chromosome positive) samples, acute lymphoblastic leukemia (precursor B cell) samples, acute myeloid leukemia (FLT3+) samples, acute myeloid leukemia (with IDH2 mutation) samples, anaplastic large cell lymphoma samples, basal cell carcinoma samples, B-cell chronic lymphocytic leukemia samples, bladder cancer samples, breast cancer (HER2 overexpression / amplification) samples, breast cancer (HER2+) samples, breast cancer (HR+, HER2-) samples, cervical cancer samples, cholangiocarcinoma samples, chronic lymphocytic leukemia samples, chronic lymphocytic leukemia (with 17p deletion) samples, chronic myeloid leukemia samples, chronic myeloid leukemia (Philadelphia chromosome positive) samples, classical Hodgkin lymphoma samples, colorectal cancer samples, colorectal cancer (dMMR and MSI-H) samples, colorectal cancer (KRAS wild type) samples, cryopyrin-associated periodic syndrome samples, cutaneous T-cell lymphoma samples, dermatofibrosarcoma protuberans samples, diffuse large B-cell lymphoma samples, fallopian tube cancer samples, follicular B-cell non-Hodgkin lymphoma samples, follicular lymphoma samples, gastric cancer samples, gastric cancer (HER2+) samples, gastroesophageal junction (GEJ) adenocarcinoma samples, gastrointestinal stromal tumor samples, gastrointestinal stromal tumor samples, gastrointestinal stromal tumor (KIT+) samples, giant cell tumor of bone samples, glioblastoma samples, polyangiitis granulomatosa samples, head and neck squamous cell carcinoma samples, hepatocellular carcinoma samples, Hodgkin lymphoma samples, juvenile idiopathic arthritis samples, lupus erythematosus samples, mantle cell lymphoma samples, medullary thyroid cancer samples, melanoma samples, melanoma with BRAF V600 mutation samples, melanoma with BRAF V600E or V600K mutation samples, Merkel cell carcinoma samples, Castleman disease samples, multiple hematological samples including Philadelphia chromosome positive ALL and CML, multiple myeloma samples, myelofibrosis samples, non-Hodgkin lymphoma samples, inoperable subependymal giant cell astrocytoma samples associated with tuberous sclerosis, non-small cell lung cancer samples, non-small cell lung cancer (ALK+) samples, non-small cell lung cancer (PD-L1+) samples, non-small cell lung cancer (with ALK fusion or ROS1 gene alteration) samples, non-small cell lung cancer (with BRAF V600E mutation) samples, non-small cell lung cancer samples (with EGFR exon 19 deletion or exon 21 substitution (L858R) mutation), non-small cell lung cancer samples (EGFRThe method of any one of items 49 to 64, comprising a T790M mutation), an ovarian cancer sample, an ovarian cancer sample (with a BRCA mutation), a pancreatic cancer sample, a pancreatic cancer sample, a gastrointestinal cancer sample, a lung-derived neuroendocrine tumor sample, a pediatric neuroblastoma sample, a peripheral T cell lymphoma sample, a prostate cancer sample, a peritoneal cancer sample, a renal cell cancer sample, a rheumatoid arthritis sample, a small lymphocyte lymphoma sample, a soft tissue sarcoma sample, a solid tumor (MSI-H / dMMR) sample, a squamous cell carcinoma sample of the head and neck, a squamous non-small cell lung cancer sample, a thyroid cancer sample, a thyroid cancer sample, a urothelial cancer sample, a urothelial cancer sample, a Waldenström macroglobulinemia sample, or any combination thereof. 66. The method of any one of items 44 to 65, wherein the model is a non-parametric probability density model. 67. The method of any one of items 44 to 66, wherein the determined tumor DNA fraction in the sample is the most likely tumor DNA fraction. 68. The method of any one of items 44 to 67, wherein the determined tumor DNA fraction in the cfDNA sample is the mean, median, or mode of the dominant peak in the empirical distribution of tumor DNA fraction values. 69. The method of any one of items 44 to 68, wherein the cfDNA sample comprises DNA extracted from a blood sample, a plasma sample, a cerebrospinal fluid sample, a pleural effusion sample, a sputum sample, a fecal sample, a urine sample, or a saliva sample. 70. A method for determining the tumor DNA fraction in a sample from a subject, comprising: receiving, in one or more processors, sequence read data in a plurality of sequence reads derived from a sample from the subject; using one or more processors to determine a variant allele frequency (VAF) for one or more variants detected in the sample based on the sequence read data; using one or more processors to generate an empirical distribution of tumor DNA fraction values according to the VAF determined for one or more variants; using one or more processors to fit a model to the empirical distribution of tumor DNA fraction values; Determining the tumor DNA fraction in a sample based on a model, and A method comprising. 71. The method of item 70, further comprising determining a confidence interval in the tumor DNA fraction based on a model. 72. The method of item 70 or 71, wherein one or more variants comprise one or more short variants. 73. The method of item 72, wherein one or more short variants comprise one or more somatic short variants. 74. The method of item 73, wherein one or more somatic short variants are known not to be associated with clonal hematopoiesis of indeterminate potential (CHIP). 75. Generating an empirical distribution of tumor DNA fraction values comprises calculating tumor DNA fraction values based on known copy numbers in one or more variants, determined VAFs in one or more variants, and corresponding known tumor average ploidies in a plurality of historical subject samples having known VAFs in one or more variants that are substantially the same as the determined VAFs in one or more variants. The method according to any one of items 70 to 74. 76. Generating an empirical distribution of tumor DNA fraction values comprises pre-calculating tumor DNA fraction values based on known copy numbers in one or more variants and corresponding known tumor average ploidies in a plurality of historical subject samples having a range of VAF values in one or more variants, and selecting a subset of the pre-calculated tumor DNA fraction values corresponding to samples having known VAFs in one or more variants that are substantially the same as the determined VAFs in one or more variants. The method according to any one of items 70 to 75. 77. The method of item 75 or 76, wherein the tumor DNA fraction value is calculated or selected with respect to the variant showing the highest VAF in the sample from the subject. 78. The method of item 75 or 76, wherein the tumor DNA fraction value is calculated or selected with respect to a ranked set of two or more variants showing the highest ranked VAFs in the sample from the subject. 79. The method of item 75 or 76, wherein the tumor DNA fraction value is calculated or selected with respect to a predetermined set of two or more variants detected in the sample from the subject. 80. The tumor DNA fraction value is calculated or selected with respect to a predetermined set of two or more variants detected in a sample from a subject containing known driver mutations, by the method of paragraph 75 or paragraph 76. 81. The tumor DNA fraction value is calculated or selected with respect to all variants detected in a cfDNA sample from a subject, by the method of paragraph 75 or paragraph 76. 82. The tumor DNA fraction value is calculated or pre-calculated by solving a set of equations representing (i) the relationship between tumor DNA fraction, tumor purity, and tumor average ploidy, and (ii) the relationship between somatic VAF, tumor purity, copy number at the genomic position of one or more variants, and the number of variant alleles at each of the one or more variants, based on the known VAF at one or more variants, the known copy number at one or more variants, and the corresponding known tumor average ploidy in a plurality of historical subject samples, thereby deriving the relationship in tumor DNA fraction as a function of somatic VAF, tumor average ploidy, copy number at the genomic position of one or more variants, and the number of variant alleles for one or more variants, excluding tumor purity. The method of any one of paragraphs 75 to 81. 83. The tumor DNA fraction value is based on the known VAF at one or more variants, the known copy number at one or more variants, and the corresponding known tumor average ploidy in a plurality of historical subject samples. The first equation that makes the tumor DNA fraction equal to the product of tumor purity and tumor average ploidy divided by the sum of the product of tumor purity and tumor average ploidy and the product of 2 and the magnitude obtained by subtracting tumor purity from 1. The second equation that makes the somatic VAF equal to the product of tumor purity and the number of variant alleles at each of the one or more variants divided by the sum of the product of tumor purity and copy number at the genomic position of the one or more variants and the product of 2 and the magnitude obtained by subtracting tumor purity from 1. Calculated or pre-calculated by solving a set of equations including the above, thereby Derive a relationship that equalizes the tumor DNA fraction to a magnitude obtained by dividing the tumor average ploidy by a value equal to the sum of the ratio of the number of variant alleles at the genomic positions of one or more variants subtracted from the tumor average ploidy excluding tumor purity and the somatic VAF in each of the one or more variants. The method according to any one of items 75 to 82. 84. The tumor DNA fraction value is based on the known VAFs in one or more variants, the known copy numbers in one or more variants, and the corresponding known tumor average ploidies in a plurality of historical subject samples, and is a set of the following equations, namely, Calculated or pre-calculated by solving TIFF2025524649000016.tif30170, whereby ρ is excluded, TIFF2025524649000017.tif20170 where ρ is the tumor purity, ψ is the tumor average ploidy, C is the copy number at the genomic positions of one or more variants, and V is the number of variant alleles in each of one or more variants. The method according to any one of items 75 to 83. 85. The method according to any one of items 75 to 84, wherein the plurality of historical subject samples include solid biopsy samples, liquid biopsy samples, or any combination thereof. 86. The method according to any one of items 75 to 85, wherein the plurality of historical subject samples include cancer samples. 87. The method according to item 86, wherein the plurality of historical subject samples include samples of a single type of cancer. 88. The method according to item 86, wherein the plurality of historical subject samples include samples of multiple types of cancer. 89. The method according to any one of items 75 to 88, wherein the plurality of historical subject samples include bladder cancer samples, breast cancer samples, colorectal cancer samples, endometrial cancer samples, kidney cancer samples, leukemia samples, liver cancer samples, lung cancer samples, melanoma samples, non-Hodgkin lymphoma samples, pancreatic cancer samples, prostate cancer samples, thyroid cancer samples, or any combination thereof. 90. A plurality of historical target samples include acute lymphoblastic leukemia (Philadelphia chromosome positive) samples, acute lymphoblastic leukemia (precursor B cell) samples, acute myeloid leukemia (FLT3+) samples, acute myeloid leukemia (with IDH2 mutation) samples, anaplastic large cell lymphoma samples, basal cell carcinoma samples, B-cell chronic lymphocytic leukemia samples, bladder cancer samples, breast cancer (HER2 overexpression / amplification) samples, breast cancer (HER2+) samples, breast cancer (HR+, HER2-) samples, cervical cancer samples, cholangiocarcinoma samples, chronic lymphocytic leukemia samples, chronic lymphocytic leukemia (with 17p deletion) samples, chronic myeloid leukemia samples, chronic myeloid leukemia (Philadelphia chromosome positive) samples, classical Hodgkin lymphoma samples, colorectal cancer samples, colorectal cancer (dMMR and MSI-H) samples, colorectal cancer (KRAS wild type) samples, cryopyrin-associated periodic syndrome samples, cutaneous T-cell lymphoma samples, dermatofibrosarcoma protuberans samples, diffuse large B-cell lymphoma samples, fallopian tube cancer samples, follicular B-cell non-Hodgkin lymphoma samples, follicular lymphoma samples, gastric cancer samples, gastric cancer (HER2+) samples, gastroesophageal junction (GEJ) adenocarcinoma samples, gastrointestinal stromal tumor samples, gastrointestinal stromal tumor samples, gastrointestinal stromal tumor (KIT+) samples, giant cell tumor of bone samples, glioblastoma samples, polyangiitis granulomatosa samples, head and neck squamous cell carcinoma samples, hepatocellular carcinoma samples, Hodgkin lymphoma samples, juvenile idiopathic arthritis samples, lupus erythematosus samples, mantle cell lymphoma samples, medullary thyroid cancer samples, melanoma samples, melanoma with BRAF V600 mutation samples, melanoma with BRAF V600E or V600K mutation samples, Merkel cell carcinoma samples, Castleman disease samples, multiple blood samples including Philadelphia chromosome positive ALL and CML, multiple myeloma samples, myelofibrosis samples, non-Hodgkin lymphoma samples, inoperable subependymal giant cell astrocytoma samples associated with tuberous sclerosis, non-small cell lung cancer samples, non-small cell lung cancer (ALK+) samples, non-small cell lung cancer (PD-L1+) samples, non-small cell lung cancer (with ALK fusion or ROS1 gene alteration) samples, non-small cell lung cancer (with BRAF V600E mutation) samples, non-small cell lung cancer samples (with EGFR exon 19 deletion or exon 21 substitution (L858R) mutation), non-small cell lung cancer samples (EGFRThe method according to any one of claims 75 to 89, comprising an ovarian cancer sample, an ovarian cancer sample (with a BRCA mutation), a pancreatic cancer sample, a pancreatic cancer sample, a gastrointestinal cancer sample, a lung-derived neuroendocrine tumor sample, a pediatric neuroblastoma sample, a peripheral T-cell lymphoma sample, a prostate cancer sample, a peritoneal cancer sample, a renal cell cancer sample, a rheumatoid arthritis sample, a small lymphocyte lymphoma sample, a soft tissue sarcoma sample, a solid tumor (MSI-H / dMMR) sample, a head and neck squamous cell carcinoma sample, a squamous non-small cell lung cancer sample, a thyroid cancer sample, a thyroid cancer sample, a urothelial cancer sample, a urothelial cancer sample, a Waldenström macroglobulinemia sample, or any combination thereof. 91. The method according to any one of claims 70 to 90, wherein the model is a non-parametric probability density model. 92. The method according to any one of claims 70 to 91, wherein the determined tumor DNA fraction in the sample is the most likely tumor DNA fraction. 93. The method according to any one of claims 70 to 92, wherein the determined tumor DNA fraction in the cfDNA sample is the mean, median, or mode of the dominant peak in the empirical distribution of tumor DNA fraction values. 94. The method according to any one of claims 70 to 93, wherein the sample comprises DNA extracted from a blood sample, a plasma sample, a cerebrospinal fluid sample, a pleural effusion sample, a sputum sample, a fecal sample, a urine sample, or a saliva sample. 95. The method according to any one of claims 44 to 94, wherein the determination of the tumor DNA fraction is used for the diagnosis or confirmation of the diagnosis of the disease of the subject. 96. The method according to claim 95, wherein the disease is cancer. 97. Cancer includes acute lymphoblastic leukemia (Philadelphia chromosome positive), acute lymphoblastic leukemia (precursor B cell), acute myeloid leukemia (FLT3+), acute myeloid leukemia (with IDH2 mutation), anaplastic large cell lymphoma, basal cell carcinoma, B-cell chronic lymphocytic leukemia, bladder cancer, breast cancer (HER2 overexpression / amplification), breast cancer (HER2+), breast cancer (HR+, HER2-), cervical cancer, cholangiocarcinoma, chronic lymphocytic leukemia, chronic lymphocytic leukemia (with 17p deletion), chronic myeloid leukemia, chronic myeloid leukemia (Philadelphia chromosome positive), classical Hodgkin lymphoma, colorectal cancer, colorectal cancer (dMMR / MSI-H), colorectal cancer (KRAS wild-type), cryopyrin-associated periodic syndrome, cutaneous T-cell lymphoma, precursor cutaneous fibrosarcoma, diffuse large B-cell lymphoma, fallopian tube cancer, follicular B-cell non-Hodgkin lymphoma, follicular lymphoma, gastric cancer, gastric cancer (HER2+), gastroesophageal junction (GEJ) adenocarcinoma, gastrointestinal stromal tumor, gastrointestinal stromal tumor (KIT+), giant cell tumor of bone, glioblastoma, polyangiitis granulomatosa, head and neck squamous cell carcinoma, hepatocellular carcinoma, Hodgkin lymphoma, juvenile idiopathic arthritis, lupus erythematosus, mantle cell lymphoma, medullary thyroid carcinoma, melanoma, melanoma with BRAF V600 mutation, melanoma with BRAF V600E or V600K mutation, Merkel cell carcinoma, multicentric Castleman disease, multiple hematologic malignancies including Philadelphia chromosome positive ALL and CML, multiple myeloma, myelofibrosis, non-Hodgkin lymphoma, unresectable subependymal giant cell astrocytoma associated with tuberous sclerosis, non-small cell lung cancer, non-small cell lung cancer (ALK+), non-small cell lung cancer (PD-L1+), non-small cell lung cancer (with ALK fusion or ROS1 gene alteration), non-small cell lung cancer (with BRAF V600E mutation), non-small cell lung cancer (with EGFR exon 19 deletion or exon 21 substitution (L858R) mutation), non-small cell lung cancer (EGFRThe method of item 96, which is ovarian cancer (with T790M mutation), ovarian cancer (with BRCA mutation), pancreatic cancer, pancreatic cancer, gastrointestinal cancer or lung-derived neuroendocrine tumor, pediatric neuroblastoma, peripheral T-cell lymphoma, peritoneal cancer, prostate cancer, renal cell cancer, rheumatoid arthritis, small lymphocyte lymphoma, soft tissue sarcoma, solid tumor (MSI-H / dMMR), head and neck squamous cell carcinoma, non-small cell lung cancer, squamous cell carcinoma, thyroid cancer, thyroid cancer, urothelial cancer, urothelial cancer, or Waldenström macroglobulinemia. 98. The method of any one of items 44 to 97, further comprising selecting an anticancer therapy to be administered to a subject based on determination of the tumor DNA fraction. 99. The method of any one of items 44 to 98, further comprising determining an effective amount of an anticancer therapy to be administered to a subject based on determination of the tumor DNA fraction. 100. The method of item 98 or item 99, further comprising administering an anticancer therapy to a subject based on determination of the tumor DNA fraction. 101. The method of any one of items 98 to 100, wherein the anticancer therapy includes chemotherapy, radiotherapy, immunotherapy, targeted therapy, or surgery. 102. A method for diagnosing a disease, comprising diagnosing that a subject has the disease based on determination of the tumor DNA fraction in a sample from the subject, wherein the tumor DNA fraction is determined according to the method of any one of items 44 to 94. 103. A method for selecting an anticancer therapy, comprising selecting an anticancer therapy for a subject according to determination of the tumor DNA fraction in a sample from the subject, wherein the tumor DNA fraction is determined according to the method of any one of items 44 to 94. 104. A method for treating cancer in a subject, comprising administering an effective amount of a cancer therapy to the subject according to determination of the tumor DNA fraction in a sample from the subject, wherein the tumor DNA fraction is determined according to the method of any one of items 44 to 94. 105. A method for determining the prognosis in a subject having cancer, Determining a tumor DNA fraction in a sample from a subject and determining a prognosis in the subject based on the tumor DNA fraction, wherein the tumor DNA fraction is determined according to the method of any one of items 44 to 94. 106. A method for evaluating minimal residual disease (MRD), comprising: Determining a tumor DNA fraction in a sample from a subject and evaluating minimal residual disease (MRD) in the subject based on the tumor DNA fraction, wherein the tumor DNA fraction is determined according to the method of any one of items 44 to 94. 107. A method for monitoring cancer progression or recurrence in a subject, comprising: Determining a first tumor DNA fraction in a first sample obtained from the subject at a first time point according to the method of any one of items 44 to 94; Determining a second tumor DNA fraction in a second sample obtained from the subject at a second time point, and comparing the first tumor DNA fraction with the second tumor DNA fraction to thereby monitor cancer progression or recurrence. 108. The method of item 107, wherein the second tumor DNA fraction in the second sample is determined according to the method of any one of items 44 to 94. 109. The method of item 107 or item 108, further comprising selecting an anti-cancer therapy for the subject according to cancer progression. 110. The method of item 107 or item 108, further comprising administering an anti-cancer therapy to the subject according to cancer progression. 111. The method of item 107 or item 108, further comprising adjusting an anti-cancer therapy for the subject according to cancer progression. 112. The method of any one of items 109 to 111, further comprising adjusting the dosage of an anti-cancer therapy according to cancer progression or selecting a different anti-cancer treatment. 113. The method of item 112, further comprising administering the adjusted anti-cancer therapy to the subject. 114. The method of any one of items 107 to 113, wherein the first time point is before the subject is administered anti-cancer treatment, and the second time point is after the subject is administered anti-cancer treatment. 115. The subject is a method according to any one of items 107 to 114, who has cancer, is at risk of having cancer, is routinely examined for cancer, or is suspected of having cancer. 116. A method according to any one of items 107 to 115, wherein the cancer is a solid tumor. 117. A method according to any one of items 107 to 116, wherein the cancer is a blood cancer. 118. A method according to any one of items 107 to 117, wherein the anti-cancer therapy includes chemotherapy, radiotherapy, immunotherapy, targeted therapy, or surgery. 119. A method according to any one of items 44 to 94, further comprising determining, identifying, or applying a value of a tumor DNA fraction in a sample as a diagnostic value associated with the sample. 120. A method according to any one of items 44 to 94, further comprising generating a genomic profile in a subject based on the determination of the tumor DNA fraction. 121. The method of item 120, wherein the genomic profile of the subject further includes results from a comprehensive genomic profiling (CGP) test, a gene expression profiling test, a cancer hot spot panel test, a DNA methylation test, a DNA fraction test, an RNA fraction test, or any combination thereof. 122. The method of item 120 or item 121, wherein the genomic profile of the subject further includes results from a test based on nucleic acid sequencing. 123. A method according to any one of items 120 to 122, further comprising selecting an anti-cancer therapy, administering an anti-cancer therapy, or applying an anti-cancer therapy to the subject based on the generated genomic profile. 124. A method according to any one of items 44 to 94, wherein the determination of the tumor DNA fraction in the sample is used when making a treatment decision proposed for the subject. 125. A method according to any one of items 44 to 94, wherein the determination of the tumor DNA fraction in the sample is used when applying or administering treatment to the subject. 126. One or more processors, A memory communicably coupled to one or more processors and configured to store instructions, wherein when the instructions are executed by the one or more processors, the system is caused to receive array read data in a plurality of array reads derived from a sample from a subject, determine a variant allele frequency (VAF) in one or more variants detected in the sample based on the array read data, generate an empirical distribution of tumor DNA fraction values according to the determined VAF in one or more variants, fit a model to the empirical distribution of tumor DNA fraction values, determine the tumor DNA fraction of the sample based on the model, a memory, A system comprising 127. The system of item 126, further comprising instructions for determining a confidence interval in the tumor DNA fraction based on the model. 128. The system of item 126 or item 127, wherein one or more variants include one or more short variants. 129. The system of item 128, wherein one or more short variants include one or more somatic short variants. 130. The system of item 129, wherein one or more somatic short variants are known not to be associated with clonal hematopoiesis of indeterminate potential (CHIP). 131. Generating the empirical distribution of tumor DNA fraction values includes calculating tumor DNA fraction values based on known copy numbers in one or more variants, the determined VAF in one or more variants, and corresponding known tumor average ploidy in a plurality of historical subject samples having known VAFs in one or more variants that are substantially the same as the determined VAF in one or more variants. The system of any one of items 126 to 130. 132. Generating an empirical distribution of tumor DNA fraction values involves pre - calculating tumor DNA fraction values based on known copy numbers in one or more variants and corresponding known tumor ploidies in a plurality of historical subject samples having a range of VAF values in one or more variants, and selecting a subset of the pre - calculated tumor DNA fraction values corresponding to samples having known VAFs in one or more variants that are substantially the same as the determined VAF in one or more variants. The system of any one of claims 126 to 130 includes these steps. 133. The tumor DNA fraction value is calculated or selected with respect to the variant showing the highest VAF in the sample from the subject. The system of claim 131 or claim 132. 134. The tumor DNA fraction value is calculated or selected with respect to an ordered set of two or more variants showing the highest ranked VAFs in the sample from the subject. The system of claim 131 or claim 132. 135. The tumor DNA fraction value is calculated or selected with respect to a predetermined set of two or more variants detected in the sample from the subject. The system of claim 131 or claim 132. 136. The tumor DNA fraction value is calculated or selected with respect to a predetermined set of two or more variants detected in the sample from the subject that includes known driver mutations. The system of claim 131 or claim 132. 137. The tumor DNA fraction value is calculated or selected with respect to all variants detected in the cfDNA sample from the subject. The system of claim 131 or claim 132. The tumor DNA fraction value is calculated or pre-calculated by solving a set of equations representing (i) the relationships among tumor DNA fraction, tumor purity, and tumor average ploidy, and (ii) the relationships among somatic VAF, tumor purity, copy number at the genomic location of one or more variants, and the number of variant alleles in each of the one or more variants, based on the known VAFs in one or more variants, the known copy numbers in one or more variants, and the corresponding known tumor average ploidies in a plurality of historical subject samples, thereby eliminating tumor purity and deriving the relationship in tumor DNA fraction as a function of somatic VAF, tumor average ploidy, copy number at the genomic location of one or more variants, and the number of variant alleles for one or more variants. The system of any one of items 131 to 137. The system of any one of items 126 to 138, wherein the model is a non-parametric probability density model. The system of any one of items 126 to 139, wherein the determined tumor DNA fraction in the sample is the most likely tumor DNA fraction. The system of any one of items 126 to 140, wherein the determined tumor DNA fraction in the cfDNA sample is the mean, median, or mode of the dominant peak in the empirical distribution of tumor DNA fraction values. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of the system, cause the system to receive sequence read data in a plurality of sequence reads derived from a sample of a subject, determine the variant allele frequency (VAF) in one or more variants detected in the sample based on the sequence read data, generate an empirical distribution of tumor DNA fraction values according to the determined VAFs in the one or more variants, fit a model to the empirical distribution of tumor DNA fraction values, and determine the tumor DNA fraction of the sample based on the model. Non-transitory computer-readable storage medium. 143. The non-transitory computer-readable storage medium of item 142, further comprising instructions for determining a confidence interval in a tumor DNA fraction based on a model. 144. The non-transitory computer-readable storage medium of item 142 or item 143, wherein one or more variants include one or more short variants. 145. The non-transitory computer-readable storage medium of item 144, wherein one or more short variants include one or more somatic short variants. 146. The non-transitory computer-readable storage medium of item 145, wherein it is known that one or more somatic short variants are not related to clonal hematopoiesis of indeterminate potential (CHIP). 147. Generating an empirical distribution of tumor DNA fraction values includes calculating the tumor DNA fraction values based on the known copy number in one or more variants, the determined VAF in one or more variants, and the corresponding known tumor average ploidy in a plurality of historical subject samples having known VAFs in one or more variants that are substantially the same as the determined VAF in one or more variants. The non-transitory computer-readable storage medium of any one of items 142 to 146. 148. Generating an empirical distribution of tumor DNA fraction values includes pre-calculating the tumor DNA fraction values based on the known copy number in one or more variants and the corresponding known tumor average ploidy in a plurality of historical subject samples having a range of VAF values in one or more variants, and selecting a subset of the pre-calculated tumor DNA fraction values corresponding to samples having known VAFs in one or more variants that are substantially the same as the determined VAF in one or more variants. The non-transitory computer-readable storage medium of any one of items 142 to 147. 149. The tumor DNA fraction value is calculated or selected with respect to the variant showing the highest VAF in the sample from the subject. The non-transitory computer-readable storage medium of item 147 or item 148. 150. The tumor DNA fraction value is calculated or selected with respect to a set of ranks of two or more variants showing the highest rank order of VAF in a sample from a subject, the non-transitory computer-readable storage medium of claim 147 or claim 148. 151. The tumor DNA fraction value is calculated or selected with respect to a predetermined set of two or more variants detected in a sample from a subject, the non-transitory computer-readable storage medium of claim 147 or claim 148. 152. The tumor DNA fraction value is calculated or selected with respect to a predetermined set of two or more variants detected in a sample from a subject including known driver mutations, the non-transitory computer-readable storage medium of claim 147 or claim 148. 153. The tumor DNA fraction value is calculated or selected with respect to all variants detected in a cfDNA sample from a subject, the non-transitory computer-readable storage medium of claim 147 or claim 148. 154. The tumor DNA fraction value is calculated or pre-calculated by solving a set of equations representing (i) the relationships among tumor DNA fraction, tumor purity, and tumor average ploidy, and (ii) the relationships among somatic VAF, tumor purity, copy number at the genomic positions of one or more variants, and the number of variant alleles at each of the one or more variants, based on the known VAF in one or more variants, the known copy number in one or more variants, and the corresponding known tumor average ploidy in a plurality of historical subject samples, thereby deriving the relationship in tumor DNA fraction as a function of somatic VAF, tumor average ploidy, copy number at the genomic positions of one or more variants, and the number of variant alleles for one or more variants, excluding tumor purity. The non-transitory computer-readable storage medium of any one of claims 147 to 153. 155. The non-transitory computer-readable storage medium of any one of claims 142 to 154, wherein the model is a non-parametric probability density model. 156. The non-transitory computer-readable storage medium of any one of claims 142 to 155, wherein the determined tumor DNA fraction in the sample is the most likely tumor DNA fraction. The non-transitory computer-readable storage medium of any one of claims 142 to 156, wherein the determined tumor DNA fraction in the 157.cfDNA sample is the mean, median, or mode of the dominant peak in the empirical distribution of the tumor DNA fraction values.
[0255] From the above, specific embodiments of the disclosed methods and systems have been illustrated and described, but various modifications can be made to them, and it should be understood that this is contemplated herein. It is also not intended that the present invention be limited by the specific examples provided herein. Although the present invention has been described with reference to the foregoing specification, the description and illustration of the preferred embodiments herein are not meant to be construed in a limiting sense. Furthermore, it should be understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative proportions described herein, which depend on various conditions and variables. Various modifications in the form and details of the embodiments of the present invention will be apparent to those skilled in the art. Accordingly, the present invention is also intended to encompass any such modifications, variations, and equivalents.
Claims
1. A method for determining a tumor DNA fraction in a cell-free DNA sample from a subject, comprising: receiving, in one or more processors, sequence read data in a plurality of sequence reads derived from a cell-free DNA (cfDNA) sample from the subject; using the one or more processors to determine, based on the sequence read data, a variant allele frequency (VAF) in one or more variants detected in the cfDNA sample; using the one or more processors to generate an empirical distribution of tumor DNA fraction values according to the determined VAF in the one or more variants; using the one or more processors to fit a model to the empirical distribution of tumor DNA fraction values; determining, based on the model, the tumor DNA fraction in the cfDNA sample; A method comprising the above steps.
2. The method according to claim 1, further comprising determining a confidence interval in the tumor DNA fraction based on the model.
3. The method according to claim 1, wherein the one or more variants comprise one or more somatic short variants known not to be associated with clonal hematopoiesis of indeterminate potential (CHIP).
4. Generating the empirical distribution of tumor DNA fraction values comprises: (i) calculating tumor DNA fraction values based on known copy numbers in the one or more variants, the determined VAF in the one or more variants, and corresponding known tumor average ploidies in a plurality of historical subject samples having known VAFs in the one or more variants that are substantially the same as the determined VAF in the one or more variants; (ii) pre-calculating tumor DNA fraction values based on known copy numbers in the one or more variants and corresponding known tumor average ploidies in a plurality of historical subject samples having a range of VAF values in the one or more variants, and selecting a subset of the pre-calculated tumor DNA fraction values corresponding to samples having known VAFs in the one or more variants that are substantially the same as the determined VAF in the one or more variants; The method according to claim 1, comprising the above steps.
5. The method according to claim 4, wherein the tumor DNA fraction value is calculated or selected with respect to the variant showing the highest VAF in the cfDNA sample from the subject.
6. The method according to claim 4, wherein the tumor DNA fraction value is calculated or selected with respect to a set of ranks of two or more variants indicating the highest-ranked VAF in the cfDNA sample from the subject.
7. The method according to claim 4, wherein the tumor DNA fraction value is calculated or selected with respect to a predetermined set of two or more variants detected in the cfDNA sample from the subject.
8. The method according to claim 4, wherein the tumor DNA fraction value is calculated or selected with respect to a predetermined set of two or more variants detected in the cfDNA sample from a subject containing a known driver mutation.
9. The method according to claim 4, wherein the tumor DNA fraction value is calculated or selected with respect to all variants detected in the cfDNA sample from the subject.
10. The method according to claim 4, wherein the tumor DNA fraction value is calculated or pre-calculated based on the known VAF in the one or more variants, the known copy number in the one or more variants, and the corresponding known tumor ploidy in the plurality of historical subject samples.
11. The method according to claim 4, wherein the plurality of historical subject samples includes solid biopsy samples, liquid biopsy samples, or any combination thereof.
12. The method according to claim 4, wherein the plurality of historical subject samples includes cancer samples.
13. The method according to claim 12, wherein the plurality of historical subject samples includes samples from a single type of cancer.
14. The method according to claim 13, wherein the plurality of historical subject samples includes samples from multiple types of cancer.
15. The method according to claim 4, wherein the plurality of historical subject samples includes bladder cancer samples, breast cancer samples, colorectal cancer samples, endometrial cancer samples, kidney cancer samples, leukemia samples, liver cancer samples, lung cancer samples, melanoma samples, non-Hodgkin lymphoma samples, pancreatic cancer samples, prostate cancer samples, thyroid cancer samples, or any combination thereof.
16. The method according to claim 1, wherein the model is a non-parametric probability density model.
17. The method according to claim 1, wherein the determined tumor DNA fraction in the sample is the most likely tumor DNA fraction.
18. The method according to claim 1, wherein the determined tumor DNA fraction in the cfDNA sample is the mean, median, or mode of the dominant peak in the empirical distribution of the tumor DNA fraction values.
19. The method according to claim 1, wherein the cfDNA sample comprises DNA extracted from a blood sample, a plasma sample, a cerebrospinal fluid sample, a pleural fluid sample, a sputum sample, a fecal sample, a urine sample, or a saliva sample.
20. The method according to claim 1, wherein the determination of the tumor DNA fraction is used for the diagnosis or confirmation of cancer in a subject.
21. A system comprising: one or more processors; a memory communicably coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to perform the method according to any one of claims 1 to 20; A system comprising the above.
22. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a system, cause the system to perform the method according to any one of claims 1 to 20.