Methods and systems for estimation of molecular residual disease, tumor mutational burden, and clonal hematopoiesis
By analyzing cell-free DNA through genomic segment processing and machine learning, the method addresses the limitations of current cancer detection methods, offering early and reliable residual cancer detection.
Patent Information
- Application Number
- PCT/US2025/031796
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-31
- Filing Date
- 2025-05-30
- Publication Date
- 2025-12-04
AI Technical Summary
Current methods for detecting residual cancer cells post-treatment are invasive, expensive, and unreliable, often failing to detect very small quantities of cancer genetic material due to subtle genetic differences and read errors in sequencing techniques, leading to a high risk of recurrence undetected by conventional means.
A method involving the analysis of cell-free DNA (cfDNA) derived from a subject to generate genomic data, process it into segments, and apply an embedding algorithm to estimate the frequency of genomic sequences originating from cancer cells, using machine learning to determine tumor mutational burden and clonal hematopoiesis, enabling early detection of residual disease.
Provides affordable, minimally invasive early detection of residual cancer by accurately quantifying ctDNA fractions, overcoming challenges of small sample sizes and sequencing errors, thus reducing the risk of undetected recurrence.
Smart Images

Figure US2025031796_04122025_PF_FP_ABST
Abstract
Description
[0001]Attorney Docket No.36024-790601 METHODS AND SYSTEMS FOR ESTIMATION OF MOLECULAR RESIDUAL DISEASE, TUMOR MUTATIONAL BURDEN, AND CLONAL HEMATOPOIESIS CROSS-REFERENCE This application claims the benefit of U.S. Provisional Application No.63 / 654,281, filed May 31, 2024, which is incorporated herein by reference in its entirety. BACKGROUND Cancer recurrence is of foremost concern for patients and medical practitioners following initial cancer treatment. Following interventions such as chemotherapy, radiotherapy, targeted drugs, immunotherapies, surgical excision, etc., a patient’s cancer may become clinically undetectable and nominally “cured,” after which the patient is shifted to post-treatment monitoring. Unfortunately, cancers frequently do recur. Recurrence of a treated cancer may be diagnosed weeks, months, or even many years after the cancer became clinically undetectable using conventional detection methods. The recurrence rate varies widely according to many factors, especially cancer type and subtype—for example, glioblastomas seem to recur in about 100 percent of cases, whereas estimated risk of childhood acute lymphoblastic leukemia is 15–20 percent. SUMMARY The continuing risk of cancer recurrence even many years after initial treatment presents many clinical challenges. For example, clinical monitoring for recurrence may be haphazard, imperfect, invasive and require frequent expensive and inconvenient patient visits. Detection may rely on imaging, cell count, or biochemical cues that are only detectable after a recurrence has matured. Even if recurrence never happens, the costs, inconveniences, and anxieties relating to recurrence risk are significant. An estimated 7 percent of post-treatment cancer patients experience debilitating fear of recurrence. These problems may be addressed using genomic detection methods. For example, using a simple serum test configured to detect genetic material and / or markers that signal lingering or recurring cancer, physicians may be able to detect recurrence early using low-cost minimally invasive techniques. Unfortunately, the residual cancer cells in a treated patient’s body (a) represent a very small fraction of the overall genetic material present in that patient’s body; and (b) the genetic material of cancer cells may only differ very slightly from healthy cells. Consequently, there is a long-felt unmet need for methods and systems to detect, identify, and quantify cancer genetic material (if any) circulating in a patient. Attorney Docket No.36024-790601 In an aspect, the present disclosure provides a method comprising: assaying a sample comprising cell-free DNA (cfDNA) derived from a subject to produce genomic data; processing the genomic data into a set of genomic segments; producing an embedding based at least in part on the set of genomic segments; and based at least in part on the embedding, generating an estimated frequency that a genomic sequence from the genomic data arises from a source. In some embodiments, the generating the estimated frequency is based at least in part on a frequency vector. In some embodiments, the frequency vector is a plasma frequency vector. In some embodiments, the set of genomic segments comprises one or more sequences indicative of the source. In some embodiments, the one or more sequences indicative of the source are determined based at least in part on assaying a diseased sample derived from the subject. In some embodiments, the diseased sample comprises a plasma sample. In some embodiments, the diseased sample comprises a tumor sample. In some embodiments, the diseased sample does not comprise a tumor sample. In some embodiments, the diseased sample derived from the subject is assayed to produce a diseased sequencing data and the one or more sequences indicative of the source are determined based at least in part on the diseased sequencing data. In some embodiments, the sample comprising cfDNA is obtained after the diseased sample is obtained from the subject. In some embodiments, the methods further comprise estimating, based at least in part on the set of genomic segments, a major allelic fraction and / or a minor allelic fraction for at least one segment in the set of genomic segments. In some embodiments, the methods further comprise, estimating a measure of ploidy for at least one segment in the set of genomic segments. In some embodiments, the methods further comprise, estimating a measure of ploidy for at least one segment, and correcting the major allelic fraction or the minor allelic fraction, based at least in part on the measure of ploidy. In some embodiments, the major allelic fraction or the minor allelic fraction is based at least in part on a copy number factor (CNF). In some embodiments, the embedding is based at least in part on the major allelic fraction or the minor allelic fraction for a genomic segment in the set of genomic segments that is associated with a diseased sample. In some embodiments, the embedding is based at least in part on filtering the major allelic fraction or the minor allelic fraction for a genomic segment in the set of genomic segments that is associated with a reference sample. In some embodiments, the method further comprised determining a measure of tumor mutational burden. In some embodiments, the measure of tumor mutational burden is based at least in part on one or more of a log likelihood function, the minor allelic fraction, the major Attorney Docket No.36024-790601 allelic fraction, or an observed frequency. In some embodiments, the method further comprised determining a measure of clonal hematopoiesis of indeterminate potential (ChIP). In some embodiments, the measure of ChIP is based at least in part on one or more of a log likelihood function, the minor allelic fraction, the major allelic fraction, or an observed frequency. In some embodiments, the method further comprised determining a tumor fraction. In some embodiments, the generating an estimated frequency comprises applying an algorithm to the embedding to produce a plasma frequency. In some embodiments, the algorithm is a machine learning algorithm. In some embodiments, the source is a cancer cell. In some embodiments, the source is a tissue type. In some embodiments, the source is a tumor associated with a cancer. In some embodiments, the cancer is a metastatic cancer. In some embodiments, the cancer is selected from the group consisting of childhood lymphoblastic leukemia, leukemia, lymphoma, multiple myeloma, adrenocortical carcinoma, bladder cancer, bladder urothelial carcinoma, bone cancer, brain lower grade glioma, breast cancer, breast invasive carcinoma, cervical squamous cell carcinoma, endocervical adenocarcinoma, cholangiocarcinoma, chronic lymphocytic leukemia, chronic myeloid disorders, colon adenocarcinoma, colorectal cancer, early onset prostate cancer, esophageal adenocarcinoma, esophageal carcinoma, gallbladder cancer, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney cancer, kidney chromophobe, kidney renal clear cell carcinoma, kidney renal papillary cell carcinoma, liver cancer, liver hepatocellular carcinoma lower grade glioma, lung cancer, lung adenocarcinoma, lung squamous cell carcinoma, lymphoid neoplasm diffuse large b-cell lymphoma, malignant lymphoma, mesothelioma, neuroblastoma, oral cancer, ovarian serous cystadenocarcinoma, pancreatic cancer endocrine neoplasms, pancreatic adenocarcinoma, pediatric brain cancer, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectum adenocarcinoma, renal cancer, sarcoma, secretory cancer, skin cutaneous melanoma, stomach adenocarcinoma, testicular germ cell tumor, testicular germ cell tumors, thymoma, thyroid carcinoma, uterine carcinosarcoma, uterine corpus endometrial carcinoma, uveal melanoma, adrenocortical carcinoma. In some embodiments, a primary tumor has not been identified in relation to the metastatic cancer. In some embodiments, a primary tumor was removed prior to development of the metastatic cancer. In some embodiments, the sample is obtained or derived from a blood sample, a urine sample, or a saliva sample. In some embodiments, the sample is obtained or derived from a blood plasma sample. In some embodiments, the set of genomic segments is determined by a split criterion. In some embodiments, the split criterion is based at least in part on a sliding window. In some Attorney Docket No.36024-790601 embodiments, the major allelic fraction or the minor allelic fraction is represented as a matrix of base pairs, comprising a frequency of occurrence of A, T, C, G, an insertion, or a deletion. In some embodiments, the genomic sequence comprises a mutation. In some embodiments, the genomic sequence comprises a genomic loci. In another aspect the present disclosure provides a method comprising: assaying a diseased sample derived from a subject to produce a diseased sequencing data, wherein the diseased sample is other than a solid tumor sample; determining one or more sequences indicative of a source based at least in part on the diseased sequencing data; and assaying a sample comprising cell-free DNA (cfDNA) derived from the subject to produce genomic data; generating an estimated frequency that a genomic sequence from the genomic data arises from the source based at least in part on the one or more sequences indicative of the source. In some embodiments, the assaying comprises assaying at least two samples comprising cfDNA obtained at different time points. In some embodiments, the source is a cancer cell. In some embodiments, the source is a tissue type. In some embodiments, the source is a tumor. In some embodiments, the diseased sample comprises a bodily fluid. In some embodiments, the bodily fluid comprises urine. In some embodiments, the bodily fluid comprises blood. In some embodiments, the diseased sample is processed to generate a cell-free sample and / or comprises a cell sample. In some embodiments, the diseased sample comprises a plasma sample. In some embodiments, the assaying the diseased sample comprises sequencing cell-free nucleic acids. In some embodiments, the assaying the diseased sample further comprises sequencing genomic DNA. In some embodiments, the determining the one or more sequences indicative of the cancer comprises comparing sequences present in the genomic DNA of the subject and sequences present in cell-free nucleic acids of the subject. In some embodiments, the diseased sample comprises a buffy coat portion and a plasma portion, and wherein the sequences present in the genomic DNA of the subject are determined by assaying the buffy coat portion and the sequences present in cell-free nucleic acids of the subject are determined by assaying the plasma portion. In some embodiments, the diseased sample comprises ctDNA. In some embodiments, the diseased sample is derived from the subject before the sample comprising cfDNA is derived from the subject. In some embodiments, the one or more sequences indicative of a cancer make up at least a portion of a tumor signature. In some embodiments, the one or more sequences indicative of the cancer comprises a set of genetic variants. In some embodiments, the estimated frequency comprises a tumor fraction. In some embodiments, the method further comprises determining the Attorney Docket No.36024-790601 presence or absence of a cancer, a recurrence of a cancer, and / or a metastasis of a cancer. In some embodiments, the method further comprises determining the presence or absence of a residual disease from a cancer. Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive. INCORPORATION BY REFERENCE All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material. BRIEF DESCRIPTION OF THE DRAWINGS A better understanding of the features and advantages of certain embodiments will be obtained by reference to the following description that sets forth illustrative embodiments and the accompanying drawings (also “figure” and “FIG.” herein), of which: Fig.1 depicts an example illustration showing genome data broken down and saved as a matrix. The in-set figure is an example zoomed-in sample showing how the matrix records base pair, insert, and / or deletion frequency at genomic loci. Fig.2 depicts an example three-dimension vector representing an estimate of tumor fraction from a plasma sample from a subject. Fig.3 depicts an example alignment between a hypothetical “normal” somatic DNA sequence and corresponding mutant tumor sequence for the same loci. As shown in the labels, mutant differences may be single-nucleotide variant (SNV) and / or copy-number variant (CNV). Figs.4A–4B depict two example three-dimension vectors representing estimates of tumor fraction from a plasma sample from a subject. Fig.4A depicts an example involving SNV type mutations; and Fig.4B depicts an example involving CNV type mutations. Figs.5A–5B depict block diagrams illustrating an example of the flow of genomic data and operations of a method for estimating tumor fraction DNA in a patient sample. Attorney Docket No.36024-790601 Fig.6 depicts a block diagram of an example of a modeling flow of data and operations. Fig.7 depicts a block diagram of an example modeling operation. Figs.8A–8D depict four alternate block diagrams of example modeling operations, depending on the materials and information available to a clinician. Fig.8A depicts an example of a scenario in which clinicians have access to tumor DNA sequence and matched-locus normal DNA sequence; Fig.8B depicts an example “tumor informed” scenario without a “matched normal”; Fig.8C depicts an example “tumor naïve” scenario; and Fig.8D depicts an example “plasma only” scenario. Figs.9A–9F depict an example of test results in three test subjects. Figs.9A–9B represent example test results from a tumor sample and plasma sample from patient 5084; Figs.9C–9D represent example test results from a tumor sample and plasma sample from patient 5234; and Figs.9E–9F represent example test results from a tumor sample and plasma sample from patient 4874. Figs.10A–10F depict an example of test results in test subjects 5084. Fig.10A represents test results from a tumor sample, and Figs.10B–10F represent example test results from five plasma samples from patient 5084. Figs.11A–11B depict an example of estimated tumor DNA fraction (Y-axis) against in- vitro synthetic data (X-axis) in two patients. Figs.12A–12C depict an example of estimated tumor DNA fraction (Y-axis) against in- vitro synthetic data (X-axis) in 10 control samples and 77 test samples. Fig.13 depicts an example graph showing score estimation of tumor fraction by cancer type. Fig.14 depicts an example analytical performance evaluation showing relation between fMSE and coverage. Fig.15 depicts an example analysis of overlap of variants for various samples. DETAILED DESCRIPTION While various embodiments of the invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Variations, changes, and substitutions may occur to those skilled in the art. It should be understood that various alternatives to the embodiments described herein may be employed. As used in the specification and claims, the singular form “a”, “an”, and “the” include plural references unless the context clearly dictates otherwise. For example, the term “a nucleic acid” includes a plurality of nucleic acids, including mixtures thereof. Attorney Docket No.36024-790601 As used herein, the term “subject,” generally refers to an entity or a medium that has testable or detectable genetic information. A subject can be a person, individual, or patient. A subject can be a vertebrate, such as, for example, a mammal. Non-limiting examples of mammals include humans, simians, farm animals, sport animals, rodents, and pets. The subject may have cancer or be suspected of having cancer. The subject may be displaying a symptom(s) indicative of cancer. As an alternative, the subject can be asymptomatic with respect to such cancer. As used herein, the term “nucleic acid” generally refers to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof. Nucleic acids may have any three-dimensional structure, and may perform any function. Non- limiting examples of nucleic acids include deoxyribonucleic acid (DNA), ribonucleic acid (RNA), messenger RNA (mRNA), transfer RNA, ribosomal RNA, short interfering RNA (siRNA), short-hairpin RNA (shRNA), micro-RNA (miRNA), ribozymes, cDNA, recombinant nucleic acids, branched nucleic acids, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. A nucleic acid may comprise one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure may be made before or after assembly of the nucleic acid. The sequence of nucleotides of a nucleic acid may be interrupted by non-nucleotide components. A nucleic acid may be further modified after polymerization, such as by conjugation or binding with a reporter agent. As used herein, the terms “amplifying” and “amplification” generally refer to increasing the size or quantity of a nucleic acid molecule. The nucleic acid molecule may be single-stranded or double-stranded. Amplification may include generating one or more copies or “amplified product” of the nucleic acid molecule. Amplification may be performed, for example, by extension (e.g., primer extension) or ligation. Amplification may include performing a primer extension reaction to generate a strand complementary to a single-stranded nucleic acid molecule, and in some cases generate one or more copies of the strand and / or the single-stranded nucleic acid molecule. The term “DNA amplification” generally refers to generating one or more copies of a DNA molecule or “amplified DNA product.” The term “reverse transcription amplification” generally refers to the generation of deoxyribonucleic acid (DNA) from a ribonucleic acid (RNA) template via the action of a reverse transcriptase. Many cancer patients who are treated and become clinically “cancer-free” face the unnerving prospect of recurrent cancer. Because cancer cells may break free from a tumor mass and circulate to another location in a patient’s body and / or may resist treatment and persist in very small numbers after treatment is finished , recurrence may be local (recurring in the same Attorney Docket No.36024-790601 location on the patient’s body), regional (recurring in lymph nodes or other tissues proximate to the original site), or distant (found in tissue far from the original site). Recurrent cancer may emerge from cells that were able to resist treatment during a first round of cancer treatment. Further, distant recurrent cancers may often have a higher propensity to circulate and metastasize. Accordingly, recurrent cancer is often much more dangerous than an initial cancer. For clinicians and such patients, it may be highly advantageous to develop methods and systems that are affordable, minimally invasive, and provide early warning and / or early detection of recurrent cancer. There may be various clinical metrics for measuring the continuing or returning presence of cancers in patients whose cancer may be in remission. One such measure may be “minimal residual disease” (MRD). In some embodiments, MRD may be any cancer cells remaining in a patient’s system—e.g., following cancer treatment (for example, cells that cannot be detected by conventional clinical scans or tests). A patient may be clinically diagnosed as in “complete remission” because there may be no clinical evidence of cancer based on conventional scans or laboratory tests but may nevertheless harbor some MRD. MRD may be used to describe residual cancer cells in any type of cancer. Some existing methods of detecting MRD include flow-cytometric immunophenotyping, cell culture systems, fluorescent in situ hybridization, and polymerase chain reaction (PCR) amplification techniques. Each of these, however, may have significant drawbacks. For example, such methods may not be able to detect very small quantities of MRD, so they cannot provide clinicians and patients with the earliest possible warning of residual disease. Such methods may also be invasive, expensive, time-consuming, and require laboratory technicians. The methods may also not be able to detect MRD if phenotype and / or genetic mutations fail to meet certain necessary parameters for detection (for example if there is no suitable target for PCR analysis). In some embodiments, cell-free DNA (which may be obtained from patient’s blood) is utilized. In some embodiments, a circulating tumor DNA (ctDNA) is utilized. In some embodiments, a circulating tumor DNA (ctDNA) is detected and / or quantified. In some embodiments, the fraction of ctDNA (“ctDNA fraction”) in a cfDNA may be quantified. In some embodiments, the fraction of ctDNA may be the tumor fraction. In some embodiments, deriving accurate information and / or metrics from cfDNA might be challenging for several reasons. For example, detecting and accurately quantifying metrics related to cancer (e.g., metrics related to ctDNA) may be difficult for several reasons. For example, the fraction of ctDNA as compared to the patient’s normal somatic DNA may be extremely small. Additionally, ctDNA may comprise genetic mutations that may be unknown. Attorney Docket No.36024-790601 Adding to that, ctDNA may comprise only a very small number of single nucleotide (SNP) mutations. Additionally, extant genetic sequencing techniques may experience read errors that may be challenging to distinguish from true mutations. Detecting and accurately quantifying ctDNA fractions might also be highly challenging for several reasons. For example, the fraction of ctDNA versus the patient’s normal somatic DNA is likely to be extremely small; ctDNA often comprises genetic mutations that are unknown (e.g., to the clinician); ctDNA often comprises only a very small number of single-nucleotide (SNP) mutations; and extant genetic sequencing techniques experience read errors that are very difficult to distinguish from true mutations, especially when the mutations are often so subtle. In some embodiments, methods and / or systems are provided that overcome such challenges. In some embodiments, methods may comprise providing a sample obtained or derived from the subject, wherein the sample comprises a cell-free DNA (cfDNA). In some embodiments, cfDNA comprises ctDNA. In some embodiments, methods may comprise assaying a sample to generate genomic data (e.g., comprising sequence data, data related to biomarkers of cancers, etc.). In some embodiments, assaying a sample may comprise whole genome sequencing (WGS). In some embodiments, assaying a sample may comprise whole exome sequencing. In some embodiments, assaying a sample may comprise RNA sequencing. In some embodiments, assaying a sample may comprise metagenomic sequencing. In some embodiments, assaying a sample may comprise transcriptome sequencing. In some embodiments, assaying a sample may comprise ChIP sequencing. In some embodiments, assaying a sample may comprise ATAC sequencing. In some embodiments, assaying a sample may comprise shotgun sequencing. In some embodiments, assaying a sample may comprise amplicon sequencing. In some embodiments, assaying a sample may comprise targeted gene panel sequencing. In some embodiments, assaying a sample may comprise hybrid capture-based sequencing. In some embodiments, assaying a sample may comprise molecular inversion probes (MIPs). In some embodiments, assaying a sample may comprise tiling array sequencing. In some embodiments, assaying a sample may comprise CRISPR-Cas9 enrichment sequencing. In some embodiments, assaying a sample may comprise multiplex ligation-dependent probe amplification (MLPA). In some embodiments, assaying a sample may comprise anchored multiplex PCR. In some embodiments, more than one sample may be assayed. In some embodiments, assaying a sample or more than one sample may comprise centrifugation and isolation of portions of the sample (such as plasma, buffy coat and / or cell pellet). In some embodiments, the genomic data is processed into genomic segments. In some embodiments, an embedding may be produced based at least in part on the set of genomic segments. In some embodiments, an estimated frequency Attorney Docket No.36024-790601 that a genomic sequence from the genomic data arises from a source may be generated based at least in part on the embedding. In some embodiments, the fraction of ctDNA in the cfDNA is determined. In some embodiments, the genomic data may be processed in order to identify the presence or absence of cancer and / or ctDNA. In some embodiments, methods of determining a fraction of ctDNA in a cell-free DNA are provided. In some embodiments, assaying the sample may comprise assaying one or more types of analytes (e.g., in the sample) in order to generate an estimated frequency that a genomic sequence from the genomic data arises from a source, and / or to determine a presence or absence of cancer and / or ctDNA. In some embodiments, the one or more types of analytes may comprise DNA or RNA, for example cfDNA or cfRNA. In some embodiments, one or more types of analytes may be cfDNA, germline DNA, genomic DNA, and / or cfRNA. In some embodiments, by analyzing different analytes, the methods may allow for improved detection or determination (e.g., of a prognosis). In some embodiments, a genomic data may be produced (and / or ctDNA fraction calculated) by, at least in part, assaying a sample (e.g., diseased samples, tumor samples, normal samples, unknown samples, plasma samples, or a combination thereof, obtained from a subject). In some embodiments, the method(s) may proceed in various modes, including tumor-informed (such as, e.g., illustrated in FIG.8A and / or FIG.8B), tumor-naïve (such as, e.g., illustrated in FIG.8C and / or FIG.8D), and with or without “normal” somatic sequences (such as, e.g., illustrated in FIG.8B). Sample In some embodiments, a sample may be derived from a subject. In some embodiments, a sample may be cell-free. In some embodiments, a sample may be substantially cell-free. In some embodiments, a sample may be processed or fractionated to produce cell-free samples. For example, samples may include cell-free nucleic acids, cell-free ribonucleic acid (cfRNA), cell- free deoxyribonucleic acid (cfDNA), cell-free fetal DNA (cffDNA), plasma, serum, urine, saliva, and derivatives thereof. In some embodiments, a sample comprises a cfDNA. Cell-free biological samples may be obtained or derived from subjects using an ethylenediaminetetraacetic acid (EDTA) collection tube, a cell-free RNA collection tube (e.g., Streck), or a cell-free DNA collection tube (e.g., Streck). In some embodiments, samples may be derived from whole blood samples by fractionation. In some embodiments, samples or derivatives thereof may contain cells. For example, a sample may be a blood sample or a derivative thereof (e.g., blood collected by a collection tube or blood drops). In some embodiments, an estimated frequency that a genomic sequence from the genomic data arises from a source may be generated. In some embodiments, a molecular residual disease, Attorney Docket No.36024-790601 tumor mutational burden, and / or clonal hematopoiesis (CHiP) may be estimated. In some embodiments, the systems and methods provided herein may be used to generate an estimated frequency that a genomic sequence from the genomic data arises from a source. In some embodiments, the systems and methods provided herein may be used to estimate molecular residual disease, tumor mutational burden, and / or clonal hematopoiesis for one or more samples. In some embodiments, a sample is obtained (or derived) from a subject. The subject may have a cancer. The subject may have a cancer that comprises a hard tumor. The subject may be undergoing treatment for cancer. For example, the subject may be subjected to one or more cancer treatments (e.g., chemotherapy, immunotherapy). The subject may have completed treatment. The subject may be in remission from cancer. The subject may be suspected of having cancer. The subject may be suspected of having a residual disease. The subject may have residual disease. In some embodiments, a plurality of samples from the same subject are derived at different time points. For example, a first sample may be from a tumor biopsy and a second sample from plasma after treatment. In some embodiments, a sample may be a solid tumor sample. In some embodiments, a sample may be other than a solid tumor sample. For example, a sample that is other than a solid tumor sample may be a bodily fluid such as blood, urine, cerebrospinal fluid, saliva, and / or mucus. In some embodiments, a sample may be obtained or derived from a subject using a swab. In some embodiments, a sample may be obtained or derived from a subject using a saliva kit. In some embodiments, the sample comprises cell-free DNA (cfDNA). In some embodiments, the sample comprises genomic DNA (gDNA). In some embodiment, gDNA may be from intact cells. In some embodiments, gDNA may be DNA from intact cells found in the sample (such as cells in urine or blood). In some embodiments, gDNA may arise from healthy cells. In some embodiments, cells found in the sample may be healthy cells. In some embodiments, gDNA may arise from cells that are normal (not diseased / not-cancerous). For example, in some samples (such as blood) cells may exist that are normal cells (such as nucleated blood cells) which may be isolated (such as in a buffy coat after centrifugation) and assayed (e.g., their DNA may be sequenced). In some samples, cells found in the sample may be tumor cells. In some embodiments, some samples (such as blood or urine) may contain tumor cells that may be pelleted using centrifugation, the pellet may then be processed and sequenced to obtain tumor DNA without the use of a biopsy. In some embodiments, one or more sequences indicative of the source may be determined. In some embodiments, determining the one or more sequences indicative of the source (e.g., cancer) comprises comparing sequences present in the genomic DNA of the subject Attorney Docket No.36024-790601 and sequences present in cell-free nucleic acids of the subject. In some embodiments, gDNA (e.g., sequences present in gDNA) may be used in the production of an embedding. In some embodiments, gDNA (e.g., sequences present in gDNA) may be used in the production of a normal embedding. In some embodiments, cfDNA may be DNA that is not in a cell (such as DNA that is released into the blood when a cell undergoes apoptosis). In some embodiments, cfDNA may comprise DNA that originated from a single source. In some embodiments, cfDNA may comprise DNA that originated from a plurality of sources (such as different cells, different cell types, healthy cells, diseased cells, tumor cells, etc.). In some embodiments, a cfDNA molecule that originated from a tumor cell (i.e., the source of the cfDNA is from a tumor) may be circulating tumor DNA (ctDNA). In some embodiments, a single ctDNA molecule may originate from a tumor cell. In some embodiments, a population of ctDNA may originate from a tumor cell or a plurality of tumor cells. In some embodiments, a portion of cfDNA may be ctDNA. In some embodiments, cfDNA may comprise a percentage (e.g., a fraction) of ctDNA. In some embodiments, cfDNA may comprise a percentage (e.g., a fraction) of normal DNA (e.g., non- tumor). In some embodiments, methods and systems described herein may utilize a normal sample. In some embodiments, a normal sample may be a sample or portion of a sample (such as a buffy coat) that does not comprise ctDNA and / or tissue from a tumor biopsy. In some embodiments, a normal sample may be a sample or portion of a sample that is taken prior to a diagnosis of cancer. In some embodiments, a normal sample may be a sample or portion of a sample that is not diseased (e.g., a sample or portion that does not comprise molecules derived from a tumor cell). In some embodiments, methods and system described herein may comprise a diseased sample. In some embodiments, a diseased sample may be a tumor biopsy sample or a portion thereof. In some embodiments, a diseased sample may be a sample that is taken from a subject while the subject has a cancer (e.g., cancer confirmed by one or more of known methods). In some embodiments, a diseased sample may be a sample or portion of sample that comprises ctDNA. In some embodiments, a diseased sample may be a sample or portion of a sample that is cancerous (such as a cell pellet in a urine sample). In some embodiments, the methods and systems described herein may utilize a tumor sample. In some embodiments, a tumor sample may be a sample other than a solid tumor sample. In some embodiments, methods and systems described herein may utilize a sample. In some embodiments, the sample is with an unknown status (e.g., it might be unknown whether it Attorney Docket No.36024-790601 was derived from a subject who has a cancer). In some embodiments, the sample may comprise cfDNA. In some embodiments, it might be unknown (e.g., not yet determined) whether the sample or a portion thereof comprises ctDNA (e.g., an unknown sample). In some embodiments, it might be unknown whether the sample comprises any portion of a tumor. In some embodiments, it might be unknown whether the sample or any portion thereof comprises nucleotides which arise from a tumor or tumor cell. In some embodiments, the sample (e.g., with an unknown status) may be a blood sample (or a portion such as the plasma) derived from a patient for whom the status of cancer is unknown. In some embodiments, the sample may be a urine sample derived from a patient for whom the status of cancer is unknown. In some embodiments, a sample may comprise various types of samples (e.g., diseased sample, tumor sample, normal sample, unknown sample). In some embodiments, a sample may comprise portions of a different type (e.g., portions comprising or not ctDNA). In some embodiments, a sample may comprise separable or otherwise distinguishable (e.g., through computational methods) portions which may be or be treated as different types of samples. For example, a blood sample may comprise a plasma portion and a buffy coat portion. In some embodiments, the plasma portion may serve as an unknown sample. In some embodiments, the plasma may serve as a diseased sample. In some embodiments, the buffy coat portion may serve as a normal sample. As another example, a urine sample may comprise a cell pellet and a supernatant (e.g., a liquid) portion. In some embodiments, after centrifugation, the cell pellet may be used as a tumor or diseased sample. In some embodiments, the supernatant may be used as the unknown sample. In some embodiments, the supernatant may be used as the normal sample. In some embodiments, more than one sample of any type may be taken. In some embodiments, more than one sample of any type may be taken simultaneously. In some embodiments, more than one sample of any type may be taken at different time points. In some embodiments, more than one samples may comprise additional samples. Additional samples may be any type of sample. In some embodiments, additional samples of any type may be taken simultaneously. In some embodiments, additional samples of any type may be taken at different time points. Sources In some embodiments, the source of cfDNA or a fraction of cfDNA may be determined and / or estimated. In some embodiments, an estimated frequency that a genomic data (such as sequencing data (e.g., produced from a sample for which the status is unknown)) arises from the source may be generated. In some embodiments, the source may correspond to the origin of the cfDNA. In some embodiments, the source may refer to an initial cell or cell type (e.g., cancer Attorney Docket No.36024-790601 cell, skin cell, lung cell) that the cfDNA was generated from (e.g., synthesized in, or cleaved by enzymes or reactants in). In some embodiments, the source (e.g., of cfDNA) may be healthy cells, cancer cell, tumor cells, diseased cells, pre-cancerous cells, residual cancer cells, residual disease, metastatic cells, cells from tumors of unknown primaries (such as a metastasis where a primary tumor was not identified), or any combination thereof. In some embodiments, the tumor may be associated with cancer. In some embodiments, the tumor may be benign. In some embodiments, the cancer may be childhood lymphoblastic leukemia, leukemia, lymphoma, multiple myeloma, adrenocortical carcinoma, bladder cancer, bladder urothelial carcinoma, bone cancer, brain lower grade glioma, breast cancer, breast invasive carcinoma, cervical squamous cell carcinoma, endocervical adenocarcinoma, cholangiocarcinoma, chronic lymphocytic leukemia, chronic myeloid disorders, colon adenocarcinoma, colorectal cancer, early onset prostate cancer, esophageal adenocarcinoma, esophageal carcinoma, gallbladder cancer, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney cancer, kidney chromophobe, kidney renal clear cell carcinoma, kidney renal papillary cell carcinoma, liver cancer, liver hepatocellular carcinoma lower grade glioma, lung cancer, lung adenocarcinoma, lung squamous cell carcinoma, lymphoid neoplasm diffuse large b-cell lymphoma, malignant lymphoma, mesothelioma, neuroblastoma, oral cancer, ovarian serous cystadenocarcinoma, pancreatic cancer endocrine neoplasms, pancreatic adenocarcinoma, pediatric brain cancer, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectum adenocarcinoma, renal cancer, sarcoma, secretory cancer, skin cutaneous melanoma, stomach adenocarcinoma, testicular germ cell tumor, testicular germ cell tumors, thymoma, thyroid carcinoma, uterine carcinosarcoma, uterine corpus endometrial carcinoma, uveal melanoma, adrenocortical carcinoma. In some embodiments, a source may come from a location in the body. In some embodiments, identification of a source may include identifying a location in the body. In some embodiments, a location may be determined by assessing the tissue of origin. In some embodiments, the source may be a tissue type. In some embodiments, a tissue type may be a tissue of the body that is healthy (e.g., non-cancerous), pre-cancerous, or cancerous. In some embodiments, a sample comprising cfDNA may comprise nucleic acids from a plurality of sources. In some embodiments, an estimated frequency may be generated that a genomic sequence (e.g., a genomic sequence from a genomic data produced by assaying a sample comprising cfDNA) arises from a source. In some embodiments, a fraction of the cfDNA, such as a tumor fraction (e.g., the fraction of cfDNA from a tumor source) and / or normal fraction (e.g., the fraction of cfDNA from a normal / non-cancer source) may be estimated. In some embodiments, the source of cfDNA in a sample may be identified using methods and systems Attorney Docket No.36024-790601 disclosed herein. In some embodiments, the fraction of cfDNA belonging to a source may be calculated using the methods and systems described herein. In some embodiments, the method comprises assaying a sample (e.g., a sample comprising cfDNA). In some embodiments, the sample is obtained from a subject at one or more areas distal to a tumor. In some embodiments, a primary tumor has not been identified in relation to a metastatic cancer. In some embodiments, a primary tumor was removed prior to the development of the metastatic cancer. In some embodiments, the subject has previously had cancer. In some embodiments, the subject is in remission. In some embodiments, the subject is in treatment. In some embodiments, the subject has been cancer free for less than a year. In some embodiments, the subject has been in remission for less than a year. In some embodiments, the subject has been cancer free for more than a year. In some embodiments, the subject has been cancer free for less than 5 years. In some embodiments, the subject has been cancer free for more than 5 years. In some embodiments, the subject has been in remission for more than a year. In some embodiments, the subject has finished treatment. In some embodiments, the subject finished treatment less than a year ago. In some embodiments, the subject finished treatment less than 5 years ago, in some embodiments, the subject finished treatment more than 5 years ago. In some embodiments, the subject finished treatment more than a year ago. In some embodiments, a sample (such as a diseased sample) may be obtained from one or more areas distal to the tumor source. Sample Processing and Data Generation In some embodiments, a sample is assayed. In some embodiments, a sample is processed prior to assaying. In some embodiments, a sample comprises cfDNA. In some embodiments, the sample is processed to remove (e.g., at least partially) sample components other than the cell-free DNA. In some embodiments, the sample is processed to isolate a cell-free DNA. In some embodiments, the sample is obtained or derived from a blood sample. In some embodiments, the sample is obtained or derived from a urine sample. In some embodiments, the sample is obtained or derived from a saliva sample. In some embodiments, the sample is obtained or derived from a blood plasma sample. In some embodiments, the sample is obtained or derived from a stool sample. In some embodiments, the sample is obtained or derived from a cerebrospinal fluid sample. In some embodiments, a sample is from a tumor (e.g., a tumor sample). In some embodiments, the samples are collected using a swab, for example, to swab cells or bodily fluid from a subject. In some embodiments, a sample is from normal tissue (e.g., a normal sample). In some embodiments, a sample has an unknown status (e.g., an unknown sample). In some embodiments, the sample is a diseased sample. In some embodiments, for example, a diseased sample may be from an individual suffering from a condition, disease, or disorder. In some Attorney Docket No.36024-790601 embodiments, the diseased sample may comprise tissue associated with a disease (e.g., from a cancerous tissue (such as a tumor biopsy)). In some embodiments, the diseased sample may comprise nucleic acids (e.g., ctDNA, cfDNA) associated with a disease (e.g., from a cancerous tissue or cell). In some embodiments, a sample (e.g., an unknown sample) may be from a subject for whom the status of cancer is unknown, for example from a subject being screened for cancer where the presence of a tumor is currently unknown (e.g., the subject has not previously been screened for cancer), or a subject being monitored for residual disease after treatment (e.g., a subject that had cancer prior to a treatment (e.g., a successful treatment)). In some embodiments, the sample is a sample type other than a tumor sample (e.g., a non-tumor sample). In some embodiments, the sample is a diseased sample. In some embodiments, the sample does not comprise a solid tumor sample. In some embodiments, the sample does not comprise a tumor biopsy sample (e.g., a non-biopsy sample). In some embodiments, the methods and systems herein utilize an unknown sample, a tumor sample, a normal sample, a diseased sample, or any combination thereof. In some embodiments, the samples used in a method do not comprise a tumor sample (e.g., a non-tumor sample). In some embodiments, a diseased sample may be used. In some embodiments, the samples used in a method do not comprises a solid tumor sample. In some embodiments, the samples used in a method are other than a solid tumor sample (e.g., a non-solid tumor sample). In some embodiments, the samples used in a method do not comprises comprise a tumor biopsy sample (e.g., a non-biopsy sample). In some embodiments, samples comprising tumor cells collected in areas distal from the tumor may be used (such as cells found in a urine sample). In some embodiments, samples comprising ctDNA may be used (e.g., as opposed to a sample comprising tumor cells); in some embodiments such samples may be collected in areas distal from the tumor. In some embodiments, DNA from samples may be used to generate a proxy or model of an expected tumor sample. For example, by processing a plasma sample that contains ctDNA and comparing the sequences with the germline or genomic DNA of a subject, sequences that are specific to a tumor can be identified and can be used in place of a tumor sample. In some embodiments, samples comprising cell-free DNA are used. In some embodiments, cell-free DNA are fragments of DNA found circulating in a bloodstream. Fragments may differ in size; for example, they can be from 120 to 220 base pairs. In some embodiments, a portion of the cell-free DNA might be a ctDNA. In some embodiments, ctDNA may be a single- or double-stranded DNA released by the tumor cells into the blood. In some embodiments, ctDNA may comprise one or more mutations of an original tumor. Attorney Docket No.36024-790601 In some embodiments, a sample is obtained or derived from the subject. A variety of methods can be used to obtain or derive the sample. In some embodiments, the sample is obtained by drawing blood from the subject (e.g., from a vein). In some embodiments, the sample obtained from the subject is further processed (e.g., purified, extracted, fractioned, derivatized, and / or otherwise altered) before it undergoes further testing. In some embodiments, the sample comprises the cell-free DNA comprising the ctDNA. In some embodiments the DNA is extracted from the sample. In some embodiments, the DNA is cfDNA. In some embodiments, the DNA is genomic DNA (gDNA). In some embodiments, the DNA is tumor DNA (e.g., DNA from a tumor biopsy). In some embodiments, the sample is assayed. In some embodiments, the DNA is sequenced to produces genomic (e.g., sequencing) data. The genomic (e.g., sequencing) data may be a part of a sequence file. In some embodiments, the genomic data may be processed into a set of genomic segments. In some embodiments, an embedding is generated based at least in part on the set of genomic segments. In some embodiments, an algorithm is applied to the embedding to produce a frequency vector. In some embodiments, the embedding comprises a frequency vector. In some embodiments, the embedding comprises a major allelic frequency and / or a minor allelic frequency. In some embodiments, the embedding comprises a weighting. For example, FIG.1 shows an embedding, 3 columns of the embedding are shown where each column may be a genomic loci, the values in the columns may indicate a frequency of various SNPs / SNVs at those locations. As another example, FIG.1 shows a column with values of 0.8 and 0.2 present in the “A” and “Insertion” rows, respectively, in which this may be interpreted as a major and minor allelic frequency at a genomic loci where the higher value is the major allelic frequency and the lower being the minor allelic frequency. In some embodiments, an estimated frequency that the DNA arises from a source is generated. In some embodiments, an estimated frequency that the cfDNA arises from a source is generated. In some embodiments, generating the estimated frequency is based at least in part on a frequency vector. In some embodiments, the estimated frequency is generated that a genomic sequence arises from a source. In some embodiments, the estimated frequency may comprise, may be, or may be used to generate a tumor fraction. In some embodiments, the estimated frequency may comprise, may be, or may be used to generate a majority allelic fraction and / or a minority allelic fraction or some combination thereof. In some embodiments, the genomic sequence may comprise one or more mutations. In some embodiments, the genomic sequence may be associated with a genomic loci. In some embodiments, the genomic sequence may correspond to Attorney Docket No.36024-790601 one or more sequencing reads. In some embodiments, the genomic sequence may be inferred or determined via processing of one or more sequencing reads. Genomic Data In some embodiments, the method comprises assaying the sample to generate genomic data. In some embodiments, the genomic data may comprise a sequencing data. In some examples, the genomic data comprises sequence reads generated by assaying a sample. In some embodiments, for example, genomic (e.g., sequence) data may be generated based on sequencing of nucleic acids (e.g., gDNA, cfDNA) in a sample. In some embodiments, the genomic (e.g., sequence) data comprises DNA sequence reads. In some embodiments, the genomic (e.g., sequence) data comprises RNA sequence reads. In some embodiments, the genomic (e.g., sequence) data comprises sequence reads generated from normal and / or tumor DNA. In some embodiments, the genomic (e.g., sequence) data comprises sequence reads generated from cell- free DNA. In some embodiments, the genomic (e.g., sequence) data comprises sequence reads generated from ctDNA. In some embodiments, the genomic (e.g., sequence) data comprises information on biomarkers of cancers. In some embodiments, the genomic (e.g., sequence) comprises a file (e.g., a digital file with one or more of aforementioned information). In various embodiments, genomic (e.g., sequence) data is generated. In some embodiments, the nucleic acids are sequenced at a coverage of at least 10x, at least 20x, at least 30x, at least 40x, at least 50x, at least 60x, at least 70x, at least 80 x, at least 90x, at least 100x, at least 150x, at least 200x, or more. In some embodiments, the sequencing provides sequence data (e.g., a sequence of nucleotides). In some embodiments, the sequence data may comprise the sequence of one or more genomic loci. For example, the sequence data may comprise sequences corresponding to a SNV, insertion, deletion, or variants at one or more genomic loci. In some embodiments, a sample comprises gDNA. In some embodiments, gDNA may be present in and / or is extracted from cells in the sample (such as blood cells or other cells) that are (e.g., known to be) non-cancerous and / or not tumor cells, for example, nucleated blood cells. In some embodiments, gDNA is present in and / or is extracted from the sample (e.g., cells from the sample) suspected of being cancerous or known to be cancerous. For example, cells may be pelleted from a sample (e.g., urine). In some embodiments, gDNA is present in and / or is extracted from the sample (e.g., cells from the sample) that are suspected of being or known to be normal (non-cancerous). In some embodiments, gDNA may be present in and / or extracted from a buffy coat. In some embodiments, cfDNA is extracted from the sample. In some embodiments, cfDNA is extracted from the plasma of a blood sample. In some embodiments, the DNA is sequenced. Attorney Docket No.36024-790601 In some embodiments, when a diseased sample is assayed (e.g., sequenced) it may produce a diseased sequencing data. In some embodiments, when sequenced data is produced from an unknown sample it may be referred to as unknown sequencing data. When sequencing data is produced from a tumor sample it may be referred to as tumor sequencing data. When sequencing data is produced from a normal sample it may be referred to as normal sequencing data. In an aspect, the present disclosure provides methods of computationally embedding genomic information, the method comprising providing genomic sequence data in a standard digital computer-readable format, wherein the genomic sequence data includes multiple reads spanning the same homologous genomic locus; and creating a matrix for each base pair coordinate, wherein each matrix column represents a frequency of occurrence at that coordinate of A, T, C, G, an insertion, and a deletion. In some embodiments, the methods may encompass methods of extracting and / or embedding nucleotide frequencies (e.g., using a sequencing file). In some embodiments, the genomic data may comprise a sequencing file. In some embodiments, the sequence data may comprise a sequence file. In some embodiments, the genomic data (e.g., sequence data) may comprise a sequencing read. In some embodiments, the genomic data (e.g., sequence data) may comprise a plurality of sequencing reads. In some embodiments, the genomic data (e.g., sequence data) may comprise sets of sequencing reads. In some embodiments, the genomic data (e.g., sequence data) may comprise alignment data. In some embodiments, the genomic data (e.g., sequence data) may comprise genomic assemblies. In some embodiments, nucleotide frequencies may be calculated based at least in part on the genomic data (e.g., sequence data). In some embodiments, nucleotide frequencies for samples from different sources may be calculated based at least in part on a sample from each of the sources (such as a tumor source and a normal / healthy source). The genomic data may comprise a digital representation of a DNA sequence in the sample. The genomic data may comprise a digital representation of a cfDNA sequence in the sample. The digital representation of a DNA sequence may be a numerical representation of the DNA sequence. The sequencing file may be a Binary Alignment Map (BAM) file, which may be the sequencing data format. The sequencing file may be a Compressed Reference Alignment Map (CRAM) file. The sequencing file may be a Sequence alignment Map (SAM) file. The sequencing file may be an unmapped BAM (uBAM) file. The sequencing file may be a FASTQ file. The sequencing file may be a HDF5 file. The sequence file may be a BedGraph file. The sequencing file may be a BED file. In some embodiments, the sequencing file contains alignment information which may be used to assemble a genomic Attorney Docket No.36024-790601 sequence. In some embodiments, the genomic data comprises a representation of a genome locus (such as a single nucleotide, a region of the genome, a fragment, a k-mer, etc.) as a matrix with rows of the matrix representing a unit (such as nucleotides, k-mers, indels, insertions, deletions, etc.) and columns representing a genomic location (such as a location relative to the start of a read, a location in an alignment, a location in a reference genome, etc.) and the values being a measurements of the unit (rows) at a given genomic location (columns) , see e.g., FIG.1. In some embodiments, of the genomic data represents a genome locus as a matrix with relative base pair frequencies. In some embodiments, the genomic data can support indels. The relative base pair frequency may be calculated by normalizing the raw count of base pairs to another value (such as the length of the sequence or the number of reads at a given genomic location). In some embodiments, each column of the matrix records a frequency of occurrence of A, T, C, G, an insertion, or a deletion. See e.g., FIG.1. In some embodiments, the features are k-mers (i.e., multiple contiguous base pairs of length k). In some embodiments, the length of the k-mers is a fixed length. In some embodiments the length is at least 2 base pairs, 3 base pairs, 4 base pairs, 5 base pairs, 6 base pairs, 7 base pairs, 8 base pairs, 9 base pairs, 10 base pairs, 11 base pairs, 12 base pairs, 13 base pairs, 14 base pairs, 15 base pairs, 16 base pairs, 17 base pairs, 18 base pairs, 19 base pairs, 20 base pairs, 30 base pairs, 40 base pairs, 50 base pairs, 60 base pairs, 70 base pairs, 80 base pairs, 90 base pairs, 100 base pairs, or more. In some embodiments the length is at most 2 base pairs, 3 base pairs, 4 base pairs, 5 base pairs, 6 base pairs, 7 base pairs, 8 base pairs, 9 base pairs, 10 base pairs, 11 base pairs, 12 base pairs, 13 base pairs, 14 base pairs, 15 base pairs, 16 base pairs, 17 base pairs, 18 base pairs, 19 base pairs, 20 base pairs, 30 base pairs, 40 base pairs, 50 base pairs, 60 base pairs, 70 base pairs, 80 base pairs, 90 base pairs, 100 base pairs. In some embodiments the matrix is a one-hot matrix wherein a 0 indicates the absence of a feature and 1 indicates the presence of a feature. In some embodiments, the matrix is continuous, wherein the values of the columns and rows are a real-valued number. In some embodiments, the values of the matrix may be in a range. In some embodiments, the range may cover all numbers. In some embodiments, the range may be between -1 and 1 (inclusive or exclusive). In some embodiments, the range may be between 0 and 1 (inclusive or exclusive). In some embodiments, the range may be 10 and below. In some embodiments, the range may be -10 and above. In some embodiments, the values may be a probability. In some embodiments, the matrix may have features in columns and position on the sequence in rows. Genome Segmentation In some embodiments, the genomic data may be processed into genomic segments (e.g., a set of genomic segments). In some embodiments, the genomic data (e.g., sequence data) may be Attorney Docket No.36024-790601 segmented to produce a set of genomic segments. In some embodiments, a genomic segment in the set of genomic segments may have a length of at least 1 base pair, at least 2 base pairs, at least 3 base pairs, at least 4 base pairs, at least 5 base pairs, at least 6 base pairs, at least 7 base pairs, at least 8 base pairs, at least 9 base pairs, at least 10 base pairs, at least 20 base pairs, at least 30 base pairs, at least 40 base pairs, at least 50 base pairs, at least 100 base pairs, at least 200 base pairs, at least 300 base pairs, at least 400 base pairs, at least 500 base pairs, at least 600 base pairs, at least 700 base pairs, at least 800 base pairs, at least 900 base pairs, at least 1000 base pairs, or any combination thereof. In some embodiments, a genomic segment in the set of genomic segments may have a length of at most 10 base pairs, at most 20 base pairs, at most 30 base pairs, at most 40 base pairs, at most 50 base pairs, at most 100 base pairs, at most 200 base pairs, at most 300 base pairs, at most 400 base pairs, at most 500 base pairs, at most 600 base pairs, at most 700 base pairs, at most 800 base pairs, at most 900 base pairs, at most 1000 base pairs, or any combination thereof. In some embodiments, genomic segments may comprise one or more genetic variants, one or more structural variants, or any combination thereof. In some embodiments, a genomic segment in the set of genomic segments may comprise a SNV, a SNP, an indel, a deletion, an insertion, a duplication, an inversion, a translocation, a copy number variation, a tandem repeat, variable number tandem repeat, a mobile element insertion, a transposable element, a gene fusion, an epigenetic modification, or any combination thereof. In some embodiments, the set of genomic segments may comprise a SNV, a SNP, an indel, a deletion, an insertion, a duplication, an inversion, a translocation, a copy number variation, a tandem repeat, variable number tandem repeat, a mobile element insertion, a transposable element, a gene fusion, an epigenetic modification, or any combination thereof. In some embodiments, a genomic segment in the set of genomic segments may comprise one or more sequences indicative of a source. In some embodiments, a genomic segment in the set of genomic segments may comprise one or more sequences indicative of a disease or disorder (e.g., a cancer). In some embodiments, the sequences are specific to a subject’s cancer (e.g., the sequences are determined by sequencing nucleic acids from the subject). In some embodiments, the sequences are specific to a subject’s cancer (e.g., the sequences typical for the subject’s cancer may be used for such purpose). In some embodiments, the set of genomic segments is determined by a split criterion. In some embodiments, the split criterion is based at least in part on a sliding window. In some embodiments, the sliding window is a fixed size. In some embodiments, the sliding window is a variable size. In some embodiments, the sliding window moves along a sequence (e.g., a nucleotide sequence). The sliding window may move along the sequence in steps where at each Attorney Docket No.36024-790601 step the window starts and / or ends at a given position on the sequence. The positions of each step may be determined by a step size. For example, using a step size of j, if a window a sequence at position i in the sequence and the window would start at position i+j. In some embodiments, the sliding window calculates a value for the sequence it is positioned at. The calculation may comprise an entropy formula, a measure of homogeneity (such as Gini index) or a sum of squared error calculation for example. A genome may be split into segments. In some embodiments, segments may have the same major and minor copy number. The segmentation may be done on the one-dimensional signal of major allelic fraction. Secondary segmentation may be done according to a spectral- corrected count. For example, the signal may be modeled as: Where xi is the relative base pair frequency of the letter i which may be the genomic bases A,C,T, and G that are normalized to the total count which is equal to the Altcount+Refcount wherein the altcount is the count of alternate bases at that location and the refcount is the count of the bases occurring in the reference for that location. The reference for a given location may be from a reference sample. A reference sample may be sample from the subject, such as a normal sample and / or normal signature. A reference sample may be a reference genome or a portion thereof. A reference sample may be from a single nucleotide polymorphism database (e.g., dbSNP). A reference sample may be based on a population reference. A reference sample may be used for a given location or set of locations in the genome. An example of a split criterion may use moving windows and declare a split point if the following formula is constant. In this example, a step of a sliding window ends at position i of the sequence. The window size is N. The count xj denotes the count of the base j. The count of base j is summed over the two windows on either side of the step for the positions i-N and i+N. These summations are normalized to the length of the window. The difference of the summations of the two windows is then normalized to the standard deviations of the count of base j for both windows. Attorney Docket No.36024-790601 In some embodiments, the methods may permit generation of an estimated frequency that a genomic sequence (e.g., from the genomic data) arises from a source. In some embodiments, such generation may be based at least in part on the embedding. In some embodiments, the embedding and / or embedded data may comprise and / or be such as shown in FIG 1, FIG.4A, and / or FIG.4B. In some embodiments, the methods may permit calculation of ctDNA fraction using embedding or embedded data such as those shown in FIG.4A and FIG.4B. In some embodiments, the embedding may have a dimensionality of 2 or more. In some embodiments, the embedding may have a dimensionality that is equal to the number of features in the matrix. In some embodiments, the embedding may have a dimensionality of 4. The four dimensions may be the bases A, T, C, and G. The samples may be mapped to the embedding using a vector of the frequencies of features, such as the frequencies of the bases A, C, T, and G. The frequencies may be relative frequencies. In some embodiments, the embedding of a sample, such as a plasma samples, are compared against the embeddings of other samples such as diseased samples, tumor samples and / or normal such as the example shown in FIG.8A. In some embodiments, the embedding of a sample is compared against an embedding of a tumor sample or a diseased sample only, such as in the example shown in FIG.8B. In some embodiments, the embedding of the sample may be compared against an embedding of a normal sample only (e.g., tumor-naïve), such as shown in the example in FIG.8C. In some embodiments, the embedding of the sample may not be compared against an embedding (e.g., such as of a tumor), such as shown in the example in FIG.8D. In some embodiments, the embedding may be based on (e.g., at least in part) on a patient specific tumor signature. In some embodiments, at least a portion of the embedding may be used as the basis from a patient specific tumor signature. In some embodiments, the embedding may be generated through a machine learning model. In some embodiments, the machine learning model may comprise prior probabilities. In some embodiments, the embedding may comprise a weighting. In some embodiments, the weighting may be a value or set of values that may be applied (e.g., to the embedding). In some embodiments, a weighting may be a set of scalar values (e.g., weights) that modify the components of an embedding. In some embodiments, a weighting may be a set of scalar values (e.g., weights) that modify the components of a vector (such as components of a frequency vector or frequency vectors as a part of the embedding). In some embodiments, a weighting may modify elements of a vector at a loci (such as weighting a frequency vector and a minor / major allelic fraction). In some embodiments, the weighting may reflect the relative importance, contribution, or influence of components of the vector. In some embodiments, a weighting may be a component of a machine learning model such as a coefficient, an attention vector, and / or a token. Attorney Docket No.36024-790601 In some embodiments, the embedding may comprise a prior probability. In some embodiments, the embedding may comprise a coefficient. In some embodiments, the machine learning model may be a Bayesian model, a Markov model, a hidden Markov model, a logistic regression, a support vector machine a neural network, or any combination thereof. For example, “plasma DNA” may be a combination of the tumor fraction (e.g., ctDNA) and normal fraction, hence plasma DNA may be calculated as Plasma DNA = TF*(Tumor DNA) + (1-TF)*(Normal DNA). It may be possible to estimate the tumor fraction using the relative frequencies vector, using the following formula: ^^^^^^= ^^^^^^in location n, correspond to the T,N, and P (tumor DNA, normal DNA, and plasma DNA) BAM files, and CNF is the copy number factor (a measure of ploidy) per location, see Fig.2. This framework may integrate the various somatic features (e.g., SNV, CNV), as shown in the examples of FIGS.4A–4B. In various cases, the genomic data (e.g., sequence data) relating to tumor DNA or normal DNA might not be matched to a normal sample, tumor sample, or diseased sample. For example, a blood draw may be performed on a subject and plasma BAM files may be generated based on sequencing of the plasma. Data corresponding to normal DNA may be generated via sequencing of a subject’s genomic DNA from a blood draw (e.g., the same blood draw) and may be compared to a plasma DNA. Data relating to the tumor DNA may be inferred by comparing a gDNA against plasma DNA. Sequences (e.g., variants) that are present in the plasma DNA but not present in the gDNA may be inferred as derived from the tumor, and data corresponding to the tumor DNA may be determined. In some embodiments, a CNF may be expressed numerically. In some embodiments, a CNF may be expressed as a real number (such as any number). In some embodiments, a CNF may be bound between a range of 0 and 1 (exclusive or inclusive). In some embodiments, a CNF may be a fraction. In some embodiments, a CNF may measure the number of times a genomic segment occurs in a sample. In some embodiments, a CNF may measure the number of times a genomic segment occurs in a sample normalized by the coverage of the sample. In some embodiments, a CNF may comprise a summation or the major copy number and the minor copy number (where multiple sources are detected), such as described herein. Frequency Vectors In some embodiments, generating an estimated frequency (that a genomic sequence (e.g., from the genomic data) arises from a source) is based at least in part on a frequency vector. In some embodiments, data (e.g., sequence data) from a sample (e.g., diseased sample, tumor sample, a normal sample and / or an unknown sample) may be used to generate frequency vectors. In some embodiments, frequency vectors may be vectors of base pair frequencies for a location Attorney Docket No.36024-790601 in the genome (such as at a specific position or in a segment). In some embodiments, frequency vectors may be calculated for samples, such as tumor samples (tumor frequency), normal sample (normal frequency), and / or unknown samples (such as blood samples, saliva samples, urine samples where the level of ctDNA may be unknown). Frequency vectors may be referred to with reference to the specific sample type (such as plasma frequencies, saliva frequencies, urine frequencies)—e.g., to mean the frequency vector of the particular sample type. In some embodiments, a frequency vector may be an observed frequency where the frequency is observed in a sample (for example, plasma, urine, swab, saliva). In some embodiments, an observed frequency may be an observed cancer-related (e.g., tumor-related) frequency. In some embodiments, an observed frequency may be an observed relative frequency. In some embodiments, an observed frequency may be a major allelic frequency. In some embodiments, an observed frequency may be a minor allelic frequency. In some embodiments, an observed frequency may be a ploidy (e.g., a copy number factor). In some embodiments, frequency vectors may be generated for genomic segments. In some embodiments, an algorithm may be applied to produce a frequency vector. In some embodiments, frequency vectors may be generated per segment. In some embodiments, the frequency data may be relative frequency data. In some embodiments, relative frequency data may be the frequency data that is corrected for a measured quality. In some embodiments, the measured quality may be measured from the sample, such as segment length, sequencing depth / coverage and / or sequence length. In some embodiments, the measured quality may be from outside the sample such as a prior probability. In some embodiments, a frequency vector may comprise frequency information for features of the sequence. In some embodiments, a feature may be any feature discussed herein. In some embodiments, a feature may be a k-mer. In some embodiments, a feature may be a location in the sequence. In some embodiments, a feature may be two k-mers that are non-contiguous, the two k-mers may be separated by a known distance. In some embodiments, a feature may be a consensus sequence. In some embodiments, a feature may be calculated by a filter such as a convolutional filter / kernel trained to recognize a pattern in the sequence. In some embodiments, a filter may be a preprogrammed filter that is capable of being used to identify a pattern in the sequence. In some embodiments, a feature may be a tokenized pattern in the sequence such as those used in training and in inference of a neural network such as a diffusion model, a large language model, or a transformer. An example algorithm for carrying out embodiments of the methods of the present disclosure (e.g., generating frequency vectors) may comprise applying the following formula: ^^^^ ^^^^^^^^^^+^^^^^^ =^^∗ ^^^^+(1−^^^^)∗2∗ ^^^^^^ ^^ ^^^^^∗^^^^^^^^+(1−^^^^)∗2 , where ^^^^, ^^^^, ^^^^^in location n, correspond to a T, N, and P Attorney Docket No.36024-790601 (tumor DNA, normal DNA, and plasma DNA) sequence files (e.g., BAM files), and CNF is the copy number factor (e.g., ploidy) per location. In this example, the plasma relative frequencies vector f P may be dependent on the relative frequencies of f T, f N and the tumor Ploidy,according to the following formula (restating the above formula): =^^^^+^^^^^^^^^^^^ ∗ ^^^^+(1−^^^^)∗2∗ ^^^^ . In this example j a genomic location and given the purity and a measure of ploidy (e.g., copy number factor). f P may have a multinomial distribution with the parameter nj, which may be the coverage of genomic location j, and the categories may be the genomic bases A, C, T, and G. Xi may be the number of times the genomic bases represented by i appear in thesequencing data of genomic location j, ^^ ∈ {^^, ^^, ^^, ^^}, j = 1, . . . 3e9 (i.e., an index running overall 3 billion human base pairs). For example, pT-pC. The convergence of the multinomial distribution to the Poisson distribution may be used to model, for any genomic location j, the general equation for each component i (such asin the case where i = 1,2,3 in the relative frequencies vector f P): ^^(^^^^|^^^^) = , ^^ = 1,2,3 where xi may be the ith component if the observed relative vector (anobserved frequency), ^^^^^,^^^^^^^and given known ploidy and purity. Purity may be used as a quality control. Filtering germline variants In some embodiments, germline variation is a genetic variation which may be inherited, as such a germline variation may be present in every cell in the body. In some embodiments, germline variants provide a background for genetic variants that originate in a tumor cell. In some embodiments, filtering germline variants can reduce noise when generating a patient specific tumor signature and / or embedding. In some embodiments, germline variants may be any of the variants disclosed herein. In some embodiments, germline variants may be determined via analysis of a reference (e.g., a reference sample from the subject, a reference from a database such from population genetic data). In some embodiments, copy number variation (CNV) and single-nucleotide variation (SNV) somatic variation may be separated from germline variations. In some embodiments, CNV germline mutations may comprise maternally and / or paternally derived germline mutations. In some embodiments, CNV mutations may appear as heterozygous or homozygous for a given locus. For example, a CNV mutation on a heterozygous Attorney Docket No.36024-790601locus may have a normal frequency vector ^^^^ = (0.5,0.5,0). In some embodiments, the CNF issplit into the maternal-strand copies and the paternal-strand copies. In some embodiments, ca may be used to represent the number of paternal-strand copies, and cb may be used to represent the number of maternal-strand copies. In some embodiments, estimation of germline CNV may comprise the equation (0.5,0.5,0) may yield the equation ^^^^(^^^^) observed frequency fT not reveal which of the frequency components may be the maternal and which may be the paternal. For example, if (ca,cb) = (3,1), fT can have possible values of fTa = (0.75,0.25,0) or fTb = (0.25,0.75,0). The observed frequencies,^^^^^^^^^^, may be used in estimation of the prior probabilities for each state. The likelihood function may comprise being written as a conditional probability function using the low of total probability: where πa = P(#paternal_copies = ca) and πb = P(#maternal_copies = cb), and πa + πb = 1. πa can be estimated from the tumor sample by using the following equation where ^^(∙)is the Poisson probability mass function. In some embodiments, the likelihood function for CNV may be defined using thelow of total probability Lj(xi), = φ(^^^^|^^^^^^^^^^^^^^^^ ^^^^ ^^ℎ^^ ^^^^^^^^^^^^^^^^ ^^^^^^ ^^^^^^^^^^^^) ⋅ π(^^^^^^^^^^^^^^^^^^^^ ^^^^^^^^^^^^^^^^ ^^^^^^ ^^^^^^^^^^^^ ))+ ^^(^^^^|^^^^^^^^^^^^^^^^ ^^^^ ^^ℎ^^ ^^^^^^^^^^^^^^^^ ^^^^^^ ^^^^^^^^^^^^) ⋅ ^^(^^^^^^^^^^^^^^^^ ^^^^ ^^^^^^^^^^^^^^^^^^^^^^ ^^^^^^^^^^^^)) + ^^(^^^^|^^^^^^^^^^^^^^^^ ^^^^ ^^^^^^ℎ ^^^^^^^^^^^^^^^^ ^^^^^^ ^^^^^^^^^^^^^^^^ ^^^^^^ ^^^^^^^^^^^^) ⋅ ^^(^^^^^^^^^^^^^^^^^^^^ ^^^^^^ℎ ^^^^^^^^^^^^^^^^ ^^^^^^ ^^^^^^^^^^^^^^^^ ^^^^^^ ^^^^^^^^^^^^ )) + ^^(^^^^|^^^^ ^^^^^^ ^^^^^^^^^^^^^^^^, ^^, ^^^^) ⋅ ^^(^^^^ ^^^^^^^^^^^^^^^^^^^^^^)), where ^^(∙) is the Poisson probability mass function and where ^^(∙)may be prior probabilities estimated from ^^^^^^^^^^. In some embodiments, tumor SNV sites may be classified using pipelines (e.g., mutect + sterlka lab-specific excluded sites based on donor’s plasma). Computational input may include peripheral blood mononuclear cells (PBMCs), tumor sample, and an exclusion list. The output may include a VCV file or a VCF file with description of SNV sites. Normal sites classification may include detection of heterozygous sites. Input may include reference human genome, PBMCs sample, and the dbSNP database. Output may include a VCV file or a VCF file with description of heterozygous sites. Major Allelic Fraction Estimation. Attorney Docket No.36024-790601 The measured MAF may comprise a blend of measured DNA from various sources, such as normal tissue, diseased tissue, cancer tissue, and / or metastatic tissue. The combination of sources may be estimated using methods and systems described here. In some embodiments, the MAF may be determined per segment. In some embodiments, the measured MAF may comprise imbalanced contributions from among the various sources such that one source is the major contributor to the sample and another is a minor contributor to the sample. In some embodiments, the measured allelic fractions may comprise a major allelic fraction (MAF) and a minor allelic fraction based on the contributions of the various sources. In some embodiments, the methods and systems provided herein estimate the major allelic fraction (MAF). In some embodiments, the MAF may be estimated per segment. In some embodiments, the major allelic fraction or the minor allelic fraction is represented as a matrix of features. In some embodiments the features are base pairs, wherein each column of the matrix records a frequency of occurrence of A, T, C, G, an insertion, or a deletion. In some embodiments, the features are k-mers (i.e., multiple contiguous base pairs of length k). In some embodiments, the length of the k-mers is a fixed length. In some embodiments the length is at least 2 base pairs, 3 base pairs, 4 base pairs, 5 base pairs, 6 base pairs, 7 base pairs, 8 base pairs, 9 base pairs, 10 base pairs, 11 base pairs, 12 base pairs, 13 base pairs, 14 base pairs, 15 base pairs, 16 base pairs, 17 base pairs, 18 base pairs, 19 base pairs, 20 base pairs, 30 base pairs, 40 base pairs, 50 base pairs, 60 base pairs, 70 base pairs, 80 base pairs, 90 base pairs, 100 base pairs. In some embodiments the length is at most 2 base pairs, 3 base pairs, 4 base pairs, 5 base pairs, 6 base pairs, 7 base pairs, 8 base pairs, 9 base pairs, 10 base pairs, 11 base pairs, 12 base pairs, 13 base pairs, 14 base pairs, 15 base pairs, 16 base pairs, 17 base pairs, 18 base pairs, 19 base pairs, 20 base pairs, 30 base pairs, 40 base pairs, 50 base pairs, 60 base pairs, 70 base pairs, 80 base pairs, 90 base pairs, 100 base pairs. In some embodiments the matrix is a one-hot matrix wherein a 0 indicates the absence of a feature and 1 indicates the presence of a feature. In some embodiments, the matrix is continuous, wherein the values of the columns and rows are a real-valued number. In some embodiments, the values of the matrix may be in a range. In some embodiments, the range may cover all numbers. In some embodiments, the range may be between -1 and 1 (inclusive or exclusive). In some embodiments, the range may be between 0 and 1 (inclusive or exclusive). In some embodiments, the range may be 10 and below. In some embodiments, the range may be -10 and above. In some embodiments, the values may be a probability. In some embodiments, the value may be a likelihood. In some embodiments, the matrix may have features in columns and position on the sequence in rows. Attorney Docket No.36024-790601 In some embodiments, similar to the MAF, the copy number measured in the sample may be the result of contributions from various sources. In some embodiments, cancers may deviate from the normal copy number distribution. In some embodiments, as such, determining the major contributing source and minor contributing source per segment may prove useful to determining the tumor fraction of a sample. In some embodiments, the major copy number (major CN), minor copy number (minor CN), purity and / or a measure of ploidy (CNF) may vary between samples. In some embodiments, the major CN, minor CN, purity and / or a measure of ploidy (CNF) may be modeled. In some embodiments, the model may be optimized. In some embodiments, the optimization of the model may comprise a symmetric Poisson mixture on the samples. In some embodiments, the measured MAF may be derived from the measured purity, ploidy and tumor MAF. For example, tumor MAF and a measure of ploidy (such as CNF) may be estimated using: . As previously disclosed, ploidy may be comprised of different contributing sources such as a major CN and minor CN. The ploidy may also comprise other contributing sources. For example, In the above example, more contributing sources may be added to calculate the ploidy if needed or their contributions may be estimated and accounted for. In some embodiments, the tumor MAF may be estimated as the percent contribution of the major CN to the total ploidy, an example of such a calculation shown here. Minimal Residual Disease (MRD) In another aspect, the present disclosure provides a method comprising: assaying a diseased sample derived from a subject to produce a diseased sequencing data, wherein the Attorney Docket No.36024-790601 diseased sample is other than a solid tumor sample. In some embodiments, the method further comprises determining one or more sequences indicative of a source based at least in part on the diseased sequencing data. In some embodiments, the method further comprises assaying a sample comprising cell-free DNA (cfDNA) derived from the subject to produce a genomic data. In some embodiments, the method further comprises generating an estimated frequency that a genomic sequence from the genomic data arises from the source based at least in part on the one or more sequences indicative of the source. Generally, previous methods of determining the presence of a cancer in a subject may rely on obtaining a tissue sample from a tumor, for example via a biopsy. However, obtaining a tumor biopsy may rely on or require an invasive procedure on the subject. Invasive procedures may result in higher risk to the subject, and may cause additional injuries to the subject. As such, obtaining one or more samples without getting a solid tumor sample, or obtaining one or more samples from areas away from the tumor may be advantageous. In various aspects of the disclosure, samples comprising tumor cells or tumor DNA may be collected or obtained without performing a tumor biopsy. In another aspect, the present disclosure provides methods of determining the presence or absence of residual disease. Residual disease may be tumor cells that are present when a patient is in remission. Residual disease may be difficult to detect through imaging or using metabolic biomarkers. The use of cfDNA may facilitate the detection of residual disease. In some embodiments, methylation levels may be used to indicate residual disease. In some embodiments, genetic variants (such as single nucleotide polymorphisms (SNP’s), single nucleotide variants (SNV’s), copy number variations (CNV’s), indels, duplications, inversions, fusions, double strand breaks, and / or alterations in ploidy) may be used to indicate residual disease. In various embodiments, samples are assayed (e.g., to generate genomic data (e.g., sequence data)). In some embodiments, the assaying comprises a whole genome assay. In some embodiments, the assaying may comprise sequencing genomic DNA. In some embodiments, the assaying may comprise sequencing cell free nucleic acids (e.g., cfDNA). In some embodiments, the assaying may comprise sequencing genomic DNA and cell-free nucleic acids (e.g., cfDNA). In some embodiments, the whole genome assay may be a whole genome sequencing assay (WGS) such as shotgun sequencing, long-read sequencing, nanopore sequencing, Sanger sequencing, next generation sequencing, whole genome-bisulfate sequencing, and / or clone-by- clone sequencing. In some embodiments, the sequencing may be a targeted sequencing assay, and may comprise PCR, qPCR, amplicon sequencing, and or hybrid capture sequencing. Attorney Docket No.36024-790601 In some embodiments, a sample may be processed. In some embodiments, processing may comprise subjecting the samples to chemical, mechanical, or biological processes. For example, the samples may be centrifuged. In another example, the chemical or biological reagents may be added to samples to extract or separate out components in the samples. In some embodiments, the processing may generate a cell-free sample. In some embodiments, the processing of the sample(s) may generate a plasma portion and a buffy coat portion, wherein the plasma comprises cell-free DNA and the buffy coat portion comprises cells found in the. In some embodiments, the buffy coat may comprise normal cells. In some embodiments, normal cells are processed to provide the sequence for the genomic DNA (gDNA) of the subject. In some embodiments, the buffy coat may comprise non-tumor cells. In some embodiments, non-tumor cells are processed to provide the sequence for the genomic DNA (gDNA) of the subject. In some embodiments, the cell-free nucleic acids are processed to provide the sequence for a cell-free nucleic acid sequence. In some embodiments, the one or more sequences indicative of the cancer are determined. In some embodiments, sequencing data (e.g., a diseased sequencing data produced via assaying a diseased sample derived from a subject) may be processed to identify or generate a patient specific tumor signature. For example, the sequencing data may be processed to identify or generate an embedding. In some embodiments, the embedding may comprise a patient specific tumor signature. In some embodiments, determining the one or more sequences indicative of the cancer may comprise comparing sequences (e.g., sequence variations) present in the genomic DNA of the subject and sequences (e.g., sequence variations) present in the cell free nucleic acid of the subject. The patient specific tumor signature may comprise a set of genetic variants. In some embodiments, the patient specific tumor signature may comprise weights applied to the genetic variants in the set of genetic variants. In some embodiments, one or more biological samples may be used, wherein the one or more biological samples do not comprise a solid tumor sample. As discussed in this disclosure, the samples may be analyzed or processed such that a solid tumor sample is not needed. For example, a sample may be processed such to obtain ctDNA, specific sequences indicative of the cancer may be identified. In some embodiments, one or more biological samples may be obtained from a subject at one or more areas distal to a tumor. For example, the sample collection may comprise non-invasive techniques (e.g., do not require a tumor biopsy). In some embodiments, a sample may comprise a cell sample. In some embodiments, the cell sample may be derived from cells found in the sample(s). In some embodiments, the cells may comprise Attorney Docket No.36024-790601 normal cells. In some embodiments, the cells may comprise tumor cells. In some embodiments, the sample(s) may comprise a plasma sample. In various aspects, additional samples are collected at different time points. The additional samples may comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 additional samples or any combination thereof. In some embodiments, the additional samples may be the same type of sample as the (initial) diseased sample. In some embodiments, an additional sample may be obtained from one or more areas distal to the tumor source. In some embodiments, additional samples comprise cfDNA. In some embodiments, the additional sample is a sample comprising cell-free DNA (cfDNA). In some embodiments, all of the additional samples are from the same source. In some embodiments, all of the additional samples are from a different source. In some embodiments, some of the additional samples are from a same source, and others are from a different source. In some embodiments, the additional samples comprise a plasma sample. In some embodiments, the additional comprises a urine sample. In some embodiments, the additional samples comprise a blood sample. In some embodiment, the samples of the additional samples are from different sources (e.g., plasma, urine). In some embodiments, some of the samples of the additional samples are from the same source. In some embodiments, the time point at which the additional samples is taken may be at least 1 minute, 10 minutes, 20 minutes, 30 minutes, 1 hour, 2 hours, 3 hours, 4 hours, 5 hours, 6 hours, 7 hours, 8 hours, 9 hours, 10 hours, 20 hours, 1 day, 2 days 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 3 weeks, 1 month, 2 months 3 months, 4 months, 5 months, 6 months, 1 year, 2 years, 3 years, 4 years, 5 years, 6 years, 7 years, 8 years, 9 years, 10 years, 20 years, 30 years, 40 years and / or 50 years after an initial sample was taken – or any combination thereof. An initial sample may be any sample type. In some embodiments, the time point at which the additional samples is taken may be at most 1 minute, 10 minutes, 20 minutes, 30 minutes, 1 hour, 2 hours, 3 hours, 4 hours, 5 hours, 6 hours, 7 hours, 8 hours, 9 hours, 10 hours, 20 hours, 1 day, 2 days 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 3 weeks, 1 month, 2 months 3 months, 4 months, 5 months, 6 months, 1 year, 2 years, 3 years, 4 years, 5 years, 6 years, 7 years, 8 years, 9 years, 10 years, 20 years, 30 years, 40 years and / or 50 years after the first sample was taken – or any combination thereof. In some embodiments, the additional samples are from a subject that is suspected of having a metastasis. In some embodiments, the additional samples are from a subject that is suspected of having a recurrence. In some embodiments, the additional samples are from a subject that is suspected of having residual disease. In some embodiments, the additional samples are from a subject that is suspected of having a tumor of unknown origin. In some embodiments, the additional samples are from a subject that is suspected of having a second Attorney Docket No.36024-790601 primary cancer. In some embodiments, the additional samples are from a subject that has previously had cancer. In some embodiments, the additional samples are from a subject that is in remission. In some embodiments, the additional samples are from a subject that is in treatment. In some embodiments, the additional samples are from a subject that has been cancer free for less than a year. In some embodiments, the additional samples are from a subject that has been in remission for less than a year. In some embodiments, the additional samples are from a subject that has been cancer free for more than a year. In some embodiments, the additional samples are from a subject that has been cancer free for less than 5 years. In some embodiments, the additional samples are from a subject that has been cancer free for more than 5 years. In some embodiments, the additional samples are from a subject that has been in remission for more than a year. In some embodiments, the additional samples are from a subject that has finished treatment. In some embodiments, the additional samples are from a subject that finished treatment less than a year ago. In some embodiments, the additional samples are from a subject finished treatment less than 5 years ago, in some embodiments, the subject finished treatment more than 5 years ago. In some embodiments, the second biological sample(s) are from a subject that finished treatment more than a year ago. In some embodiments, the additional sample(s) may be assayed to produce an additional sets of sequence data where a set of sequence data is produced for each of the biological samples. In some embodiments the presence or absence of the cancer may be determined based at least in part on the patient specific tumor signature. In some embodiments, the presence or absence of residual disease may be determined based at least in part on the patient specific tumor signature (in some embodiments, it is further based on the additional sets of sequence data). In some embodiments, a tumor fraction may be determined (such as by using the methods disclosed herein). In some embodiments, a tumor fraction may be determined based at least in part on the patient specific tumor signature. In some embodiments, the patient specific tumor signature may be updated. In some embodiments, the patient specific tumor signature may be updated by identifying a set of sequences (e.g., genetic variants) in the additional sets of sequence data that may be altered from the patient specific tumor signature and altering the set of sequences (e.g., genetic variants) in the patient specific tumor signature, the weights of the sequences (e.g., genetic variants) in the patient specific tumor signature, or both. The altering may comprise increasing one or more weights. The altering may comprise decreasing one or more weights. The altering may comprise adding a sequence to the patient specific tumor signature. The altering may comprise removing a sequence from the patient specific tumor signature. Attorney Docket No.36024-790601 In some embodiments, weighting genetic variants may provide an advantage to the methods and systems disclosed herein as they may provide additional data about the prevalence of patient specific variants in when variants arise at the same genomic locus, such as when tumor heterogeneity causes variants to arise from different lineages in a tumor. Monitoring alterations in the prevalence of a set of variants at a locus, and the resulting alterations in weighting in the patient specific tumor signature may enable better understanding of what variants are present at various stages of monitoring or diagnosis. Weighting may also provide evidence of variants that may not be detected in a cfDNA sample, such as when a genomic locus is not present or has low abundance in a cfDNA sample. In such a situation the resulting sequencing data may not have high confidence and / or high signal at a given locus. The presence and prevalence of a different variant may be used as evidence for imputation of the level of the missing or low abundance genomic loci. In some embodiments, tumor heterogeneity may exist in a given tumor such that there is non-uniform genetic variation across a tumor and may arise from lineages of surviving tumor cells proliferating and accumulating additional genetic variations that do not necessarily exist in other lineages in the tumor. Such variation can lead to different areas of a tumor having different responses, or response levels, to various pressures such as circumvention of the immune systems, resistance to programmed cell death, cytotoxicity due to mutational burden, varying proliferative signaling, angiogenesis activation, evasion of growth suppressors, activation of invasion or metastasis, and response / sensitivity to treatment. The combination of variants present together in a lineage of cells and present overall in the tumor may play a part in protecting the individual cells and the tumor itself. When a suit of mutations is advantageous it may provide at least a portion of the tumor with the ability to persist. If the surviving cells are low enough in number, such as after treatment and / or in remission, the cells may be difficult to detect / monitor using imaging or conventional methods of monitoring such as through protein-based biomarkers (for which the level of secretion may be lower than the detection threshold). In such a case, methods, such as those described here, may be used to detect the residual disease and monitor its survival and evolution over time. In some embodiments, to account for the heterogeneity of the tumor and / or the residual disease variants in the patient specific tumor signature may be weighted. Weighting may reflect the relative levels at which variants at the same genomic location are present to one another. Weights may also reflect the levels at which variants are present relative to all variants at all locations. Weights may reflect known priors about variants such as their severity or contribution to qualities that allow cancer to proliferate (such as those listed in this example) or contribution to cancer severity or morbidity and / or mortality. Weights Attorney Docket No.36024-790601 may not be mutually exclusive, and multiple weights may be applied to the variants. Weights could be derived from a machine learning model such as a logistic regression model trained to predict cancer or cancer severity where the coefficients of the model could serve as feature importance weights in the patient specific tumor signature. In some embodiments, the weight can provide useful information that may allow for the imputation of missing values in the sequencing data. Missing values may arise from a sequence not being present, or not having sufficient levels for confident sequencing, in the plasma. Missing data may also arise from noisy sequencing disallowing for a high confidence read of the base at a given location. Weights may be used as a model of the correlation among bases present and may therefore be valuable for imputation of missing data in the sequencing data. A method such as a hidden Markov model may be used to ingest the know prior probabilities (e.g., the weights) and output the probability that a base or sequence exists at a given location, such as the location of the missing data. In some embodiments, data from a population may be processed to provide a non- tumor population signature. The non-tumor population signature may be compared against the first sequencing data. Somatic variants that are non-tumor related may be removed based at least in part on the non-tumor population signature and the first sequencing data. The non-tumor population signature may be compared against the second sequencing data. Somatic variants that are non-tumor related may be removed based at least in part on the non-tumor population signature and the second sequencing data. The methods and systems described herein may be used to update the patient specific tumor signature. One a patient specific tumor signature is generated it may be updated using subsequent sequencing data. To do so a second biological sample may be used to provide sequencing data for a second time point. Variants in the sample may be analyzed and compared against the patient specific tumor signature to assess any changes to somatic mutations or to the relative abundance of somatic mutations from the patient specific tumor signature. Changes to relative abundance or variants may be used to update the weights of the patient specific tumor signature, the variants in the patient specific tumor signature or both. Methods such as these may be used to monitor for changes in the tumor that and thereby provide insight into potential drug resistance for future therapies. In another aspect, the present disclosure provides methods of determining a likelihood of cancer recurrence in a subject, comprising: providing a first biological sample from a subject who has a cancer; sequencing the first biological sample to generate a first sequence data; identifying a set of sequences from the sequence data that is indicative of the cancer; Attorney Docket No.36024-790601 providing a second biological sample from the subject after the cancer has gone into remission; sequencing the first biological sample to generate a first sequence data; sequencing the second biological sample to generate a second sequence data; and comparing the second sequence data to the set of sequences from the sequence data that is indicative of the cancer to determine the presence or absence of recurrence of the cancer in the subject. In some embodiments, the set of mutations comprises a SNV, a CNV, a frequency vector, a CNF (ploidy), a measure of purity, or any combination thereof. In some embodiments, the second biological sample does not comprise data for some number of the mutations in the set of mutations. The missing data may be imputed by a machine learning model. In another aspect, the present disclosure provides methods of estimating the fraction of tumor DNA from a DNA sample from a subject, the method comprising: (a) sequencing a sample of DNA extracted from a subject; (b) recording the DNA sequence data in a standard digital computer-readable format; (c) providing to a computer algorithm the digital computer-readable formatted genomic data; (d) providing to the computer algorithm reference somatic genomic sequence data in a standard digital computer-readable format; (e) from previously sampled tumor material, providing to the computer algorithm reference tumor genomic sequence data in a standard digital computer-readable format; (f) creating a matrix for each base pair coordinate, wherein each matrix column records a frequency of occurrence at that coordinate of A, T, C, G, an insertion; and (g) estimating the DNA fraction circulating in the patient’s plasma according to the equation ^^^^= . In any embodiment the method, the sample of DNA may be drawn from: known tumor cells, known non-tumor cells, blood plasma, urine, saliva, or any combination thereof. In any embodiment of the method, the subject may be a human patient. Tumor Mutational Burden In addition to MRD, there may be other metrics for measuring and describing latent and / or residual risk of recurrent cancer. One such measure is tumor mutational burden (TMB), or more specifically in the case of blood cancers (e.g., leukemia), blood tumor mutational burden (bTMB). TMB measures the number of somatic coding alterations present in a particular cancer. The increased presence of such mutations can be favorable to treatment because the alterations contribute to immunogenicity through the generation of antigens targeted by T-cell response. A higher TMB may be associated with favorable immunotherapy response rate and survival across multiple types of cancer. Thus, assessing TMB in a subject or patient can help medical practitioners and patients make optimal treatment decisions. Attorney Docket No.36024-790601 Measuring TMB may require obtaining an accurate read of tumor DNA sequence. The methods of the present description can help alleviate a long felt need and improve estimates of a subject’s true TMB. For example, the CNV-TMB may be calculated using the following likelihood function: Thus, the TMB may be where c ≥ 0 is a given constant. In alternative embodiments TMB may A similar calculation may be applied for SNV-TMB (e.g., detected mutations per 3,000 million base pairs) component. Phase Machine Learning Machine learning (ML), broadly, is a class of computational methods that leverage existing data to build analytical models and improve performance in future computational tasks. In some embodiments, “training data” is introduced to the ML system from which the system learns about patterns in the data. In some embodiments, missing data may be imputed using a machine learning method. In some embodiments, the machine learning method may be a multi-phase approach. Multi-phase approaches to ML involve the use of algorithms to impute variables from a dataset, and another algorithm to predict values that can be treated successfully. Such ML approaches may be used to fill in a dataset where at least some of the data may be incorrect. Aggregated and / or Layered Data In some embodiments, the ctDNA fraction estimates achieved using the present methods may be further augmented using copy-number variant (CNV) data; tissue feature data; cancer type data; and the like. Clonal Hematopoiesis of Indeterminate Potential (ChIP) may refer to the presence of a hematologic malignancy-associated somatic mutation in blood cells of a subject. In subjects with ChIP, a clonally expanded hematopoietic stem cell may be caused by a leukemogenic mutation in the subject, which may not have evidence of a determinate origin such as a hematologic malignancy, dysplasia, or cytopenia. The presence of ChIP in a subject may be indicative of, or associated with, a disease, disorder, or condition (e.g., leukemia or cardiovascular disease). For example, ChIP may be associated with a 0.5% to 1% risk per year of leukemia. Using systems and methods of the present disclosure, ChIP events may be assessed in Attorney Docket No.36024-790601 a subject. For example, by applying systems and methods of the present disclosure to sequencing reads derived from PBMC samples of a subject, changes in the estimated MAF value for certain genomic regions can be indicative of ChIP events in the subject. ChIP assessment may comprise determining a presence or absence of mutations with a variant allele frequency of at least about 2%, wherein the mutations are located in genes that are affected in, or associated with, hematologic cancers. In some embodiments, the methods and systems disclosed herein may provide for a workflow, such as the one shown in FIG.7, comprising genomic segmentation, estimation of major allele fraction; optimization of major CN, minor CN, and purity; estimations of “pure” major allele fraction, and ploidy per segment; calculations of probabilities per site for tumor frequency. In some embodiments, an algorithm is used to generate an output (e.g., a frequency vector, an estimate of allelic fraction, an estimate of ploidy). In some embodiments, an algorithm may be a trained algorithm. In some embodiments, the trained algorithms may use one or more sequences (e.g., a variant, mutation) as an input and generate an output regarding the presence or absence of a cancer. In some embodiments, the trained algorithm may be trained on multiple samples. For example, the trained algorithm may be trained using at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 300, 400, 500 , 600 ,700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or more independent training samples. The trained algorithm may be trained using no more than 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 300, 400, 500 , 600 ,700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or less, independent training samples. In some embodiments, the training samples may be associated with a presence or an absence of the cancer. In some embodiments, the training samples may be associated with a relapse of cancer. In some embodiments, the training samples may be associated with cancer that is resistant to a particular drug or treatment. In some embodiments, an individual training sample may be positive for a particular cancer. In some embodiments, an individual training sample may be negative for a particular cancer. In some embodiments, the trained algorithm may be able to detect a cancer, determine a probability of recurrence or relapse of a cancer, or determine if a cancer comprises a set of biomarkers may be Attorney Docket No.36024-790601 resistant to a treatment. In some embodiments, the training sample may be associated with additional clinical health data of a subject. For example, additional clinical health data may comprise the gender, weight, height, or levels of metabolites or antibodies in a subject. In some embodiments, additional clinical health data may comprise indication of other diseases, disorders, or diseases conditions. The trained algorithms may be trained using multiple sets of training samples. The sets may comprise training samples as described elsewhere herein. For example, the training may be performed using a first set of independent training samples associated with a presence of the cancer and a second set of independent training samples associated with an absence of the cancer. Similarly, a first set may be associated with relapse and a second sample may be associated with the absence of relapse. The trained algorithm may also process additional clinical health data of the subject. For example, additional clinical health data may comprise the gender, weight, height, or levels of metabolites or antibodies in a subject. Additional clinical health data may comprise indication of other diseases, disorders, or diseases conditions that the subject may suffer from. By using the additional clinical health data, in conjunction with the biomarkers, the trained algorithm may output a presence or absences of cancer, probability of relapse, or resistance to drug treatment, which may be different from the output of an algorithm that does not process additional clinical health. In some embodiments, the trained algorithm may be trained, at least in part, using an unsupervised machine learning algorithm. For example, the unsupervised machine learning algorithm may utilize cluster analysis to identify attributes of interest. In some embodiments, the trained algorithm may be trained, at least in part, using a supervised machine learning algorithm. For example, the algorithm may be inputted with training data such to generate an expected or desired output. In some embodiments, the supervised learning algorithm may comprise a deep learning algorithm, a support vector machine (SVM), a neural network, or a Random Forest. In some embodiments, via the machine learning algorithm, the trained algorithm may be able to identify relationships of sample sequence data to a fraction of ctDNA. Without the trained algorithm, it may otherwise be difficult to identify such relationships. In various aspects, the systems and methods may comprise an accuracy, sensitivity, or specificity of detection of the cancer or a parameter of the cancer. For example, the methods or systems may comprise detecting the presence or the absence of cancer (or the presence of a parameter of the cancer, such as recurrence, relapse, or drug resistance) in the subject at an accuracy of at least about 60%, at least about 70%, at least about 75%, at least about Attorney Docket No.36024-790601 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. The methods or systems may comprise detecting the presence or the absence of cancer (or the presence of a parameter of the cancer, such as recurrence, relapse, or drug resistance) in the subject at a sensitivity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. The methods or systems may comprise detecting the presence or the absence of cancer (or the presence of a parameter of the cancer, such as recurrence, relapse, or drug resistance) in the subject at a specificity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.The methods or systems may comprise detecting the presence or the absence of cancer (or the presence of a parameter of the cancer, such as recurrence, relapse, or drug resistance) in the subject at a positive predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. The methods or systems may comprise detecting the presence or the absence of cancer (or the presence of a parameter of the cancer, such as recurrence, relapse, or drug resistance) in the subject at a negative predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. Networking In an additional aspect, the present disclosure provides a networked computation system, the system comprising: at least one computer terminal having a data input means for receiving digital genomic sequence data; and at least one server communicatively linked to the computer terminals, the server loaded with software configured to perform the methods of any of the paragraphs supra. In any embodiment, the computer terminals may be configured to receive the results of data of computational operations performed on the at least one server. Computer systems The present disclosure provides computer systems that are programmed to implement methods of the disclosure. A computer system may be programmed or otherwise configured to perform analysis or operations of the methods, for example determine a likelihood of the presence of a cancer based on a set of biomarkers of an individual or run an algorithm. The computer system can regulate various aspects of methods and systems of the present disclosure, such as, for example, perform an algorithm, input training data, analyze sets of biomarkers, or output a result for the user as to the presence or absence of cancer. The computer system can be Attorney Docket No.36024-790601 an electronic device of a user or a computer system that is remotely located with respect to the electronic device. The electronic device can be a mobile electronic device. The computer system includes a central processing unit (CPU, also “processor” and “computer processor” herein), which can be a single core or multi core processor, or a plurality of processors for parallel processing. The computer system also includes memory or memory location (e.g., random-access memory, read-only memory, flash memory), electronic storage unit (e.g., hard disk), communication interface (e.g., network adapter) for communicating with one or more other systems, and peripheral devices , such as cache, other memory, data storage and / or electronic display adapters. The memory, storage unit, interface and peripheral devices are in communication with the CPU through a communication bus (solid lines), such as a motherboard. The storage unit can be a data storage unit (or data repository) for storing data. The computer system can be operatively coupled to a computer network (“network”) with the aid of the communication interface. The network can be the Internet, an internet and / or extranet, or an intranet and / or extranet that is in communication with the Internet. The network in some cases is a telecommunication and / or data network. The network can include one or more computer servers, which can enable distributed computing, such as cloud computing. The network, in some cases with the aid of the computer system, can implement a peer-to-peer network, which may enable devices coupled to the computer system to behave as a client or a server. The CPU can execute a sequence of machine-readable instructions, which can be embodied in a program or software. The instructions may be stored in a memory location, such as the memory. The instructions can be directed to the CPU, which can subsequently program or otherwise configure the CPU to implement methods of the present disclosure. Examples of operations performed by the CPU can include fetch, decode, execute, and writeback. The CPU can be part of a circuit, such as an integrated circuit. One or more other components of the system can be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC). The storage unit can store files, such as drivers, libraries and saved programs. The storage unit can store user data, e.g., user preferences and user programs. The computer system in some cases can include one or more additional data storage units that are external to the computer system, such as located on a remote server that is in communication with the computer system through an intranet or the Internet. The computer system can communicate with one or more remote computer systems through the network. For instance, the computer system can communicate with a remote computer system of a user (e.g., a medical professional or patient). Examples of remote computer Attorney Docket No.36024-790601 systems include personal computers (e.g., portable PC), slate or tablet PC’s (e.g., Apple® iPad, Samsung® Galaxy Tab), telephones, Smart phones (e.g., Apple® iPhone, Android-enabled device, Blackberry®), or personal digital assistants. The user can access the computer system via the network. Methods as described herein can be implemented by way of machine (e.g., computer processor) executable code stored on an electronic storage location of the computer system, such as, for example, on the memory or electronic storage unit. The machine executable or machine-readable code can be provided in the form of software. During use, the code can be executed by the processor. In some cases, the code can be retrieved from the storage unit and stored on the memory for ready access by the processor. In some situations, the electronic storage unit can be precluded, and machine-executable instructions are stored on memory. The code can be pre-compiled and configured for use with a machine having a processer adapted to execute the code, or can be compiled during runtime. The code can be supplied in a programming language that can be selected to enable the code to execute in a pre- compiled or as-compiled fashion. Aspects of the systems and methods provided herein, such as the computer system, can be embodied in programming. Various aspects of the technology may be thought of as “products” or “articles of manufacture” typically in the form of machine (or processor) executable code and / or associated data that is carried on or embodied in a type of machine readable medium. Machine-executable code can be stored on an electronic storage unit, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. “Storage” type media can include any or all of the tangible memory of the computers, processors or the like, or associated modules thereof, such as various semiconductor memories, tape drives, disk drives and the like, which may provide non-transitory storage at any time for the software programming. All or portions of the software may at times be communicated through the Internet or various other telecommunication networks. Such communications, for example, may enable loading of the software from one computer or processor into another, for example, from a management server or host computer into the computer platform of an application server. Thus, another type of media that may bear the software elements includes optical, electrical and electromagnetic waves, such as used across physical interfaces between local devices, through wired and optical landline networks and over various air-links. The physical elements that carry such waves, such as wired or wireless links, optical links or the like, also may be considered as media bearing the software. As used herein, unless restricted to non-transitory, tangible “storage” Attorney Docket No.36024-790601 media, terms such as computer or machine “readable medium” refer to any medium that participates in providing instructions to a processor for execution. Hence, a machine readable medium, such as computer-executable code, may take many forms, including but not limited to, a tangible storage medium, a carrier wave medium or physical transmission medium. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices in any computer(s) or the like, such as may be used to implement the databases, etc. shown in the drawings. Volatile storage media include dynamic memory, such as main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media therefore include for example: a floppy disk, a flexible disk, hard disk, magnetic tape, any other magnetic medium, a CD-ROM, DVD or DVD-ROM, any other optical medium, punch cards paper tape, any other physical storage medium with patterns of holes, a RAM, a ROM, a PROM and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave transporting data or instructions, cables or links transporting such a carrier wave, or any other medium from which a computer may read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution. The computer system can include or be in communication with an electronic display that comprises a user interface (UI) for providing, for example, an input of biomarkers or sequencing data, or a visual output relating to a detection, diagnosis, or prognosis. Examples of UI’s include, without limitation, a graphical user interface (GUI) and web-based user interface. Methods and systems of the present disclosure can be implemented by way of one or more algorithms. An algorithm can be implemented by way of software upon execution by the central processing unit. The algorithm can, for example, determine a presence or absence of a cancer or cancer parameter based on a set of input sequencing data from a sample derived from a subject. Mobile application In some embodiments, a computer program includes a mobile application provided to a mobile computing device. In some embodiments, the mobile application is provided to a mobile computing device at the time it is manufactured. In other embodiments, the Attorney Docket No.36024-790601 mobile application is provided to a mobile computing device via the computer network described herein. In view of the disclosure provided herein, a mobile application is created by various techniques using hardware, languages, and development environments known to the art. Mobile applications may be implemented in various languages. Suitable programming languages include, by way of non-limiting examples, C, C++, C#, Objective-C, Java™, JavaScript, Pascal, Object Pascal, Python™, Ruby, Rails, VB.NET, WML, and XHTML / HTML with or without CSS, or combinations thereof. Suitable mobile application development environments may be used. Various development environments include, by way of non-limiting examples, AirplaySDK, alcheMo, Appcelerator®, Celsius, Bedrock, Flash Lite, .NET Compact Framework, Rhomobile, and WorkLight Mobile Platform. Other development environments are available including, by way of non-limiting examples, Lazarus, MobiFlex, MoSync, and Phonegap. Also, mobile device manufacturers distribute software developer kits including, by way of non-limiting examples, iPhone and iPad (iOS) SDK, Android™ SDK, BlackBerry® SDK, BREW SDK, Palm® OS SDK, Symbian SDK, webOS SDK, and Windows® Mobile SDK. Various commercial forums are available for distribution of mobile applications including, by way of non-limiting examples, Apple® App Store, Google® Play, Chrome WebStore, BlackBerry® App World, App Store for Palm devices, App Catalog for webOS, Windows® Marketplace for Mobile, Ovi Store for Nokia® devices, Samsung® Apps, and Nintendo® DSi Shop. Standalone application In some embodiments, a computer program includes a standalone application, which is a program that is run as an independent computer process, not an add-on to an existing process, e.g., not a plug-in. Standalone applications may be compiled. A compiler is a computer program(s) that transforms source code written in a programming language into binary object code such as assembly language or machine code. Suitable compiled programming languages include, by way of non-limiting examples, C, C++, Objective-C, COBOL, Delphi, Eiffel, Java™, Lisp, Python™, Visual Basic, and VB .NET, or combinations thereof. Compilation is often performed, at least in part, to create an executable program. In some embodiments, a computer program includes one or more executable complied applications. Web browser plug-in In some embodiments, the computer program includes a web browser plug-in (e.g., extension, etc.). In computing, a plug-in is one or more software components that add Attorney Docket No.36024-790601 specific functionality to a larger software application. Makers of software applications support plug-ins to enable third-party developers to create abilities which extend an application, to support easily adding new features, and to reduce the size of an application. When supported, plug-ins enable customizing the functionality of a software application. For example, plug-ins are commonly used in web browsers to play video, generate interactivity, scan for viruses, and display particular file types. Examples of web browser plug-ins include Adobe® Flash® Player, Microsoft® Silverlight®, and Apple® QuickTime®. In some embodiments, the toolbar comprises one or more web browser extensions, add-ins, or add-ons. In some embodiments, the toolbar comprises one or more explorer bars, tool bands, or desk bands. In view of the disclosure provided herein, various plug-in frameworks may be used that enable development of plug-ins in various programming languages, including, by way of non-limiting examples, C++, Delphi, Java™, PHP, Python™, and VB .NET, or combinations thereof. Web browsers (also called Internet browsers) are software applications, designed for use with network-connected computing devices, for retrieving, presenting, and traversing information resources on the World Wide Web. Suitable web browsers include, by way of non- limiting examples, Microsoft® Internet Explorer®, Mozilla® Firefox®, Google® Chrome, Apple® Safari®, Opera Software® Opera®, and KDE Konqueror. In some embodiments, the web browser is a mobile web browser. Mobile web browsers (also called microbrowsers, mini- browsers, and wireless browsers) are designed for use on mobile computing devices including, by way of non-limiting examples, handheld computers, tablet computers, netbook computers, subnotebook computers, smartphones, music players, personal digital assistants (PDAs), and handheld video game systems. Suitable mobile web browsers include, by way of non-limiting examples, Google® Android® browser, RIM BlackBerry® Browser, Apple® Safari®, Palm® Blazer, Palm® WebOS® Browser, Mozilla® Firefox® for mobile, Microsoft® Internet Explorer® Mobile, Amazon® Kindle® Basic Web, Nokia® Browser, Opera Software® Opera® Mobile, and Sony® PSP™ browser. Software modules In some embodiments, the platforms, systems, media, and methods disclosed herein include software, server, and / or database modules, or use of the same. In view of the disclosure provided herein, software modules are created by various techniques using machines, software, and languages. The software modules disclosed herein are implemented in a multitude of ways. In various embodiments, a software module comprises a file, a section of code, a programming object, a programming structure, a distributed computing resource, a cloud Attorney Docket No.36024-790601 computing resource, or combinations thereof. In further various embodiments, a software module comprises a plurality of files, a plurality of sections of code, a plurality of programming objects, a plurality of programming structures, a plurality of distributed computing resources, a plurality of cloud computing resources, or combinations thereof. In various embodiments, the one or more software modules comprise, by way of non-limiting examples, a web application, a mobile application, a standalone application, and a distributed or cloud computing application. In some embodiments, software modules are in one computer program or application. In other embodiments, software modules are in more than one computer program or application. In some embodiments, software modules are hosted on one machine. In other embodiments, software modules are hosted on more than one machine. In further embodiments, software modules are hosted on a distributed computing platform such as a cloud computing platform. In some embodiments, software modules are hosted on one or more machines in one location. In other embodiments, software modules are hosted on one or more machines in more than one location. Databases In some embodiments, the platforms, systems, media, and methods disclosed herein include one or more databases, or use of the same. In view of the disclosure provided herein, various databases are suitable for storage and retrieval of datasets described herein, or any combination thereof. In various embodiments, suitable databases include, by way of non-limiting examples, relational databases, non-relational databases, object-oriented databases, object databases, entity-relationship model databases, associative databases, XML databases, document oriented databases, and graph databases. Further non-limiting examples include SQL, PostgreSQL, MySQL, Oracle, DB2, Sybase, and MongoDB. In some embodiments, a database is Internet-based. In further embodiments, a database is web-based. In still further embodiments, a database is cloud computing based. In a particular embodiment, a database is a distributed database. In other embodiments, a database is based on one or more local computer storage devices. Data transmission The subject matter described herein, including methods and systems as described herein and may be configured to be performed in one or more facilities at one or more locations. Facility locations are not limited by country and include any country or territory. In some instances, one or more operations are performed in a different country than another operation of the method. In some embodiments, one or more method operations involving a computer system are performed in a different country than another operation of the methods provided herein. In some embodiments, data processing and storage are performed in a different country or location Attorney Docket No.36024-790601 than one or more operations of the methods described herein. In some embodiments, one or more products or data are transferred from one or more of the facilities to one or more different facilities for analysis or further analysis. Data includes, but is not limited to, information regarding the stratification of a subject, and any data produced by the methods disclosed herein. In some embodiments of the methods and systems described herein, the subject information is compiled, and a subsequent data transmission operation may transmit or store the subject information. In some embodiments, any operation of any method described herein is performed by a software program or module on a computer. In additional or further embodiments, data from any operation of any method described herein is transferred to and from facilities located within the same or different countries, including analysis performed in one facility in a particular location and the data shipped to another location or directly to an individual in the same or a different country. In additional or further embodiments, data from any operation of any method described herein is transferred to and / or received from a facility located within the same or different countries, including analysis of a data input, such as queries, objects, properties, types, filters, tables, or any combination thereof, performed in one facility in a particular location and corresponding data transmitted to another location. The methods described herein may utilize one or more computers. The computer may include a monitor or other user interface for displaying data, results, billing information, marketing information (e.g., demographics), customer information, or sample information. The computer may also include means for data or information input. The computer may include a processing unit and fixed or removable media or a combination thereof. The computer may be accessed by a user in physical proximity to the computer, for example via a keyboard and / or mouse, or by a user that does not necessarily have access to the physical computer through a communication medium such as a modem, an internet connection, a telephone connection, or a wired or wireless communication signal carrier wave. In some cases, the computer may be connected to a server or other communication device for relaying information from a user to the computer or from the computer to a user. In some cases, the user may store data or information obtained from the computer through a communication medium on media, such as removable media. It is envisioned that data relating to the methods can be transmitted over such networks or connections for reception and / or review by a party. The entity entering or reviewing information into a database for the purpose of one or more of the following: inventory tracking, order tracking, customer management, customer service, billing, and sales. Sample information may include, but is not limited to: Attorney Docket No.36024-790601 customer name, unique customer identification, or any information suitable for storage in a database. The database may be accessible by a user. Database access may take the form of electronic communication such as a computer or telephone. The database may be accessed through an intermediary such as a customer service representative, business representative, or consultant. The availability or degree of database access may change upon payment of a fee for products and services rendered or to be rendered. Embodiments The following are non-limiting examples of embodiments . Any of these example embodiments may be combined with any other of the embodiments described here or elsewhere in the specification and claims. A method comprising: assaying a cell-free DNA sample to produce genomic data; processing the genomic data into a set of genomic segments; producing an embedding based at least in part on the set of genomic segments; applying a machine learning method to the embedding to produce a plasma frequency; and generating, based at least in part on the plasma frequency, an estimated frequency that a genomic sequence arises from a source. The method of embodiment 1, further comprising: estimating, based at least in part on the set of genomic segments, a frequency of major alleles for at least one segment of the set of genomic segments; estimating, based at least in part on the set of genomic segments, a frequency of minor alleles for the at least one segment; and estimating a measure of ploidy for the at least one segment. The method of embodiment 2, further comprising correcting the estimated frequency of major alleles or the estimated frequency of minor alleles, based at least in part on the estimated measure of ploidy. The method of embodiment 2 or 3, wherein the estimated frequency of major alleles or the estimated frequency of minor alleles is based at least in part of a copy number factor. Attorney Docket No.36024-790601 The method of embodiment 2 or 3, wherein the embedding is based at least in part on the estimated frequency of major alleles or the estimated frequency of minor alleles for a genomic segment in the set of genomic segments that is associated with a diseased sample. The method of embodiment 2 or 3, wherein the embedding is based at least in part on the estimated frequency of major alleles or the estimated frequency of minor alleles for a genomic segment in the set of genomic segments that is associated with a reference sample. The method of embodiment 2, further comprising determining a measure of tumor mutational burden. The method of embodiment 7, wherein the measure of tumor mutational burden is based at least in part on one or more of a log likelihood function, the estimated frequency of minor alleles, the estimated frequency of major alleles, or an observed frequency. The method of embodiment 2, further comprising determining a measure of clonal hematopoiesis of indeterminate potential (ChIP). The method of embodiment 9, wherein the measure of ChIP is based at least in part on one or more of a log likelihood function, the estimated frequency of minor alleles, the estimated frequency of major alleles, or an observed frequency. The method of embodiment 1, wherein a source is a cancer. The method of embodiment 11, wherein the cancer is a metastatic cancer. The method of embodiment 11, wherein the cancer is selected from the group consisting of childhood lymphoblastic leukemia, leukemia, lymphoma, multiple myeloma, adrenocortical carcinoma, bladder cancer, bladder urothelial carcinoma, bone cancer, brain lower grade glioma, breast cancer, breast invasive carcinoma, cervical squamous cell carcinoma, endocervical adenocarcinoma, cholangiocarcinoma, chronic lymphocytic leukemia, chronic myeloid disorders, colon adenocarcinoma, colorectal cancer, early onset prostate cancer, esophageal adenocarcinoma, esophageal carcinoma, gallbladder cancer, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney cancer, kidney chromophobe, kidney renal clear cell carcinoma, kidney renal papillary cell carcinoma, liver cancer, liver hepatocellular carcinoma lower grade glioma, lung cancer, lung adenocarcinoma, lung squamous cell carcinoma, lymphoid neoplasm diffuse large b-cell lymphoma, malignant lymphoma, mesothelioma, neuroblastoma, oral cancer, ovarian serous cystadenocarcinoma, pancreatic cancer endocrine neoplasms, pancreatic adenocarcinoma, pediatric brain cancer, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectum adenocarcinoma, renal cancer, sarcoma, secretory cancer, skin cutaneous melanoma, stomach adenocarcinoma, testicular germ cell tumor, Attorney Docket No.36024-790601 testicular germ cell tumors, thymoma, thyroid carcinoma, uterine carcinosarcoma, uterine corpus endometrial carcinoma, uveal melanoma, adrenocortical carcinoma. The method of embodiment 13, wherein a primary tumor has not been identified in relation to the metastatic cancer. The method of embodiment 13, wherein a primary tumor was removed prior to development of the metastatic cancer. The method of embodiment 1, wherein the cell-free DNA sample is obtained or derived from a blood sample. The method of embodiment 1, wherein the cell-free DNA sample is obtained or derived from a urine sample. The method of embodiment 1, wherein the cell-free DNA sample is obtained or derived from a saliva sample. The method of embodiment 1, wherein the cell-free DNA sample is obtained or derived from a blood plasma sample. The method of embodiment 1, wherein the set of genomic segments is determined by a split criterion. The method of embodiment 20, wherein the split criterion is based at least in part on a sliding window. The method of embodiment 2, wherein the estimated frequency of major alleles or the estimated frequency of minor alleles is represented as a matrix of base pairs, wherein each column of the matrix records a frequency of occurrence of A, T, C, G, an insertion, or a deletion. The following examples are included for illustrative purposes only and are not intended to limit the scope of the inventive concepts. Examples Example 1. Estimation of Tumor Fraction Let CNF(Ploidyn) be the copy number factor per location. A hidden state for any location n may be calculated using where the vector copy number may be CNF*fT. Observed frequency: ^^^^^^^^^^,^^ = ^^(^^ ^^^^ , ^^), where ε may be a noise process.^^^^+^^^^^^ ∗ ^^^^+(1−^ ) ^^ ing the equation ^^^^=^^ ^^^^^ ∗2∗ ^^ Task: by us^^^^^^^^∗^^^^^^^^+(1−^^^^)∗2 , estimate ctDNAfraction (TF) given Attorney Docket No.36024-790601 Define the likelihood function: . Solve: This may be plugged in to the frameworks illustrated at Figs.5A–5B, 6, and 8A– 8D. FIG 11A and 11B examples demonstrate that this method is capable of accurate fraction estimation. cfDNA samples were simulated, in silico with known tumor fractions (for 2 examples, 5118 and 5234). The tumor fraction was then estimated accurately with estimations closely following the dashed line indicating the expected prediction. The method performed similarly in synthetic datasets produced from a combination of cell free plasma samples (FIG 11A) and cellular sample (FIG 11B). In vitro mixtures of samples, with known tumor fractions, were then produced for three cancer types (“blc” (bladder) “brca” (breast), and “crc” (colorectal) in FIGs 12A- 12 C respectively). The model had high performance over about 10-5 tumor fraction for all cancer types with the method correctly giving no tumor fraction to the ND samples which are controls containing no tumor DNA. This shows that the method is capable of detecting ctDNA at very low levels with high accuracy across cancer types. Tumor fraction can be accurately estimated using these methods. As shown FIG 13 where the true label indicates the known tumor fraction of sample, here samples with very low tumor fraction (100µ) and control samples with 0 tumor fraction. Samples with a tumor fraction are distinctly separated from control samples. FIG 14 shows that the error of the method (mean squared error, MSE) decreases for all priors as coverage increases. This shows that above about 20x coverage, the method can accurately identify tumor fraction in cfDNA regardless of prior. Example 2. Plasma tumor fraction estimation, given known fT and fN Assuming a SNV mutation in both DNA copies, without CNV mutation, Ploidy = 2, coverage = ni, Purity = 1 and the following frequencies vectors fT = (0,1,0), fN = (1,0,0). In this instance, there may be complete information on both tumor and normal samples, and theprobability function may be ^^ where x may be the observed tumor- related frequency component in fP, times nj. The maximum likelihood estimator of Poisson distribution may be the sample mean, and therefore the tumor fraction estimator will be ^^^^ = ^^⁄ ^^^^ .Example 3. Tumor Naïve Approach Attorney Docket No.36024-790601 In some embodiments, MRD (minimum residual disease or molecular residual disease) is a method for monitoring a cancer patient to determine if their cancer has recurred or grown post-treatment. Some cancer patients may have surviving tumor cells (such as treatment resistant tumor cells) post-treatment when they are in remission. These cells may be difficult to detect. As such, methods to identify MRD may provide evidence of residual disease earlier than other clinical methods (such as imaging) providing patients and doctors valuable lead times in monitoring, locating and treating potential relapses. In some embodiments, MRD may be detected in a tumor informed or tumor naïve paradigm. In some embodiments, the tumor informed methods utilize DNA from the tumor itself which is profiled. In these methods, the assay looks specifically for markers that were found in that patient’s tumor to detect tumor DNA at subsequent monitoring time points. In some embodiments, a biopsy of the tumor tissue itself may be required, limiting the ability of tumor informed methods to be used for cancers that are prohibitively difficult to biopsy or on patients for which no biopsy exists. In some patients or cancer types, it may be not possible to get DNA directly from the cancer tissue. In such cases tumor naïve methods, e.g., methods that are not dependent on DNA directly from tumor tissue, may be useful. Existing tumor naïve MRD methods typically look for biomarkers that are common in cancer across patients rather than applying patient specific models. In some embodiments, the lack of a patient specific model may limit the sensitivity of these assays. In some instances, certain tumor naïve MRD methods may rely on cfDNA methylation patterns rather than genomic variation (such as indels, SNV, CNV) or other genomic features. In some embodiments, a tumor-naïve scenario may include a normal sequence (e.g., a known non-cancer sequence) such as in FIG.8C, or no normal sequence such as in FIG. 8D. In the tumor-naïve scenario where no tumor-specific sequencing data is available for the patient (and, e.g., no tumor embedding is produced) (such as is shown in FIGs.8C-8D), the prior probabilities may be used for estimating tumor fraction—e.g., such as those derived from a representative set of tumor samples corresponding to the relevant cancer type. In some embodiments, these representative tumor samples can be used to establish a statistical baseline for mutational frequencies and ploidy adjustments, which can then be applied to the patient's plasma sample for estimating the likelihood of ctDNA presence. In some embodiments, this approach ensures that even in the absence of a matched tumor sample, an informed estimation of tumor fraction can still be made. In other embodiments, other approaches may also be used that do not utilize a tumor sample (e.g., that do not utilize a DNA from the tumor itself)—for example, using a Attorney Docket No.36024-790601 diseased sample other than a tumor sample (e.g., other than a solid tumor sample). In an example, Fig.15 shows the results of variants’ analysis in two replicate runs of the same tumor sample and a plasma sample. The diagram shows the number of shared variants for each overlap combination of tumor replicate 1 (T1), tumor replicate 2 (T2), and plasma (P).1651 variants (P unique) were found in the plasma only; without limiting to a particular theory, at least some of them can be attributed to non-cancer variants.16780 and 26753 (T1 and T2 Unique) variants are found in the individual tumor replicates; without limiting to a particular theory, at least some of them can be attributed to noise. As shown in the figure, there were 2941 overlaps between tumor replicate 1 and plasma which were not found to overlap between tumor replicate 2 and plasma (P in T1 & !P in T2); and there were 275 overlaps between tumor replicate 2 and plasma which were not found to overlap between tumor replicate 1 and plasma (P in T2 & !P in T1). The 5114 variants were found to be present in both tumor replicate runs (T1 in T2 & !T1 in P); without limiting to a particular theory, at least some of such variants might be attributed to variants only found in tumor samples. The 17587 variants were found to overlap between 2 tumor replicate runs and a plasma sample (T1 in T2 & T1 in P); without limiting to a particular theory, at least some of these can be attributed to tumor specific variants that are found in plasma. This demonstrates that there is a significant amount of overlap (e.g., in variants) between the tumors and the plasma, and that the presence of such overlapping sequences in a plasma sample can be used, e.g., to determine sequences indicative of a cancer without the use of a tumor sample (e.g., in the tumor naïve scenario). By identifying sequences in a sample that are indicative of cancer, a patient specific tumor signature can be created and used, e.g., in an MRD method(s), without the need for tumor DNA derived directly from the tumor. In some embodiments, this tumor naïve approach can be specific to each patient. In some embodiments, such approach may improve sensitivity over existing tumor-naïve MRD methods. For example, for a patient, a blood can be drawn prior to cancer remission. This initial blood draw, taken while the patient still has active cancer, may be used to generate a patient specific tumor signature of mutations, copy number variants and / or other genomic features of the tumor without the need for tumor sample. To generate the patient specific tumor signature gDNA may be extracted from blood cells (e.g., from buffy coat) and sequenced. Cell-free DNA (cfDNA) from plasma containing circulating tumor DNA (ctDNA) may be extracted from the patient’s blood (e.g., from the same blood draw). This cfDNA may be sequenced. By comparing the cfDNA to the gDNA, mutations that are found only in cfDNA can be identified and these can be inferred to be arising Attorney Docket No.36024-790601 from ctDNA originating from the tumor. These ctDNA signatures can then be used generate a patient specific tumor model containing mutations, insertions, deletions, fusions, copy-number variants, and other genomic features that are specific to that patient’s tumor. This can allow for calculation of tumor fraction without a tumor sample. Subsequent monitoring of the patient cfDNA at future time points can then be compared and analyzed based on the to their own patient specific tumor signature to identity tumor DNA (ctDNA). Example 4: VAF for detection of tumor DNA and residual disease In some examples, plasma samples show similar variant allelic frequencies (VAF) to tumor samples, as shown in FIGs 9A-9F. Paired plasma and tumor samples from 3 subjects with cancer were compared (FIGS 9A and 9B; FIGs 9C and 9D; FIGs 9E and 9F). Clusters of variations (indicated by numbered circles in FIGs 9A-9F) show similar distribution across the paired sample and tumor samples. The clusters shown here and in FIG 10, can be used to demonstrate how a method employing an embedding could be used to estimate the source or fractional sources of a sample. In some embodiments, this demonstrates that cfDNA from the tumor source (ctDNA) is detectable in a plasma sample. The distribution itself may be used as a proxy for the patient specific tumor sample demonstrating that the preservation of the distribution of clusters in plasma samples may be detected in cfDNA. Using these observations, techniques were developed to generate patient specific tumor signatures and to detect not only the presence of tumor DNA in plasma samples but the fraction that ctDNA makes up of the cfDNA. Further, ctDNA is detectable in cfDNA when residual disease is present, as demonstrated by FIGs 10A-10F where multiple samples at different time points were taken and assayed. A tumor sample was taken before treatment whose VAF is shown in FIG 10A, establishing a tumor signature. In some embodiments, a tumor sample is not required and instead a tumor signature may be derived via other means (e.g., in a tumor naïve scenario); however a tumor sample may also be used to establish a tumor signature, for example as shown in this example. Multiple samples at different time points were taken, with each of FIGs 10B-10F providing data for a sample, with each labeled as a plasma #1 – plasma #5 sample by the order in which they were taken. The first sample, FIG 10C, shows no discernable tumor signature, indicating that the subject has minimal to no recurrence or residual disease. As this is close to the end of treatment no signature would be expected. FIG 10E, 10B, 10D, 10F (in this order) show a progressively more discernable tumor signature over time. The reemergence of the tumor signature indicates residual disease is present and detectable in the cfDNA. Taken together, the indication of detectable tumor signature in cfDNA and the detection of the tumor signature over time where there is residual disease demonstrate, e.g., that Attorney Docket No.36024-790601 this method may be used to detect residual disease (and / or a tumor signature may be developed from ctDNA for use in such methods). Using these observations methods were developed where, e.g., one or more sequences indicative of a source (e.g., a caner) can be used to detect residual disease presence.
Claims
Attorney Docket No.36024-790601 CLAIMS WHAT IS CLAIMED IS:
1. A method comprising:a. assaying a sample comprising cell-free DNA (cfDNA) derived from a subject toproduce genomic data; b. processing the genomic data into a set of genomic segments;c. producing an embedding based at least in part on the set of genomic segments;and d. based at least in part on the embedding, generating an estimated frequency that agenomic sequence from the genomic data arises from a source.
2. The method of claim 1, wherein generating the estimated frequency is based at least inpart on a frequency vector.
3. The method of claim 2, wherein the frequency vector is a plasma frequency vector.
4. The method any one of claims 1 to 3, wherein the set of genomic segments comprises oneor more sequences indicative of the source.
5. The method of claim 4, wherein the one or more sequences indicative of the source aredetermined based at least in part on assaying a diseased sample derived from the subject.
6. The method of claim 5, wherein the diseased sample comprises a plasma sample.
7. The method of any one of claims 5 to 6, wherein the diseased sample comprises a tumorsample.
8. The method of any one of claims 5 to 6, wherein the diseased sample does not comprise atumor sample.
9. The method of claim 8, wherein the diseased sample derived from the subject is assayedto produce a diseased sequencing data and the one or more sequences indicative of the source are determined based at least in part on the diseased sequencing data.
10. The method of any one of claims 5 to 9, wherein the sample comprising cfDNA isobtained after the diseased sample is obtained from the subject.
11. The method of any one of claims 1 to 10, further comprising estimating, based at least inpart on the set of genomic segments, a major allelic fraction and / or a minor allelic fraction for at least one segment in the set of genomic segments.
12. The method of any one of claims 1 to 11, further comprising, estimating a measure ofploidy for at least one segment in the set of genomic segments.Attorney Docket No.36024-79060113. The method of any one of claims 1 to 12, further comprising, estimating a measure ofploidy for at least one segment, and correcting the major allelic fraction or the minor allelic fraction, based at least in part on the measure of ploidy.
14. The method of any one of claims 1 to 13, wherein the major allelic fraction or the minorallelic fraction is based at least in part on a copy number factor (CNF).
15. The method of any one of claims 1 to 14, wherein the embedding is based at least in parton the major allelic fraction or the minor allelic fraction for a genomic segment in the set of genomic segments that is associated with a diseased sample.
16. The method of any one of claims 1 to 15, wherein the embedding is based at least in parton filtering the major allelic fraction or the minor allelic fraction for a genomic segment in the set of genomic segments that is associated with a reference sample.
17. The method of any one of claims 1 to 16, further comprising determining a measure oftumor mutational burden.
18. The method of claim 17, wherein the measure of tumor mutational burden is based atleast in part on one or more of a log likelihood function, the minor allelic fraction, the major allelic fraction, or an observed frequency.
19. The method of any one of claims 1 to 18, further comprising determining a measure ofclonal hematopoiesis of indeterminate potential (ChIP).
20. The method of claim 19, wherein the measure of ChIP is based at least in part on one ormore of a log likelihood function, the minor allelic fraction, the major allelic fraction, or an observed frequency.
21. The method of any one of claims 1 to 20, further comprising determining a tumorfraction.
22. The method of any one of claims 1 to 21, wherein the generating an estimated frequencycomprises applying an algorithm to the embedding to produce a plasma frequency.
23. The method of claim 22, wherein the algorithm is a machine learning algorithm.
24. The method of any one of claims 1 to 23, wherein the source is a cancer cell.
25. The method of any one of claims 1 to 23, wherein the source is a tissue type.
26. The method of any one of claims 1 to 23, wherein the source is a tumor associated with acancer.
27. The method of any one of claims 24 to 26, wherein the cancer is a metastatic cancer.
28. The method of any one of claims 24 to 27, wherein the cancer is selected from the groupconsisting of childhood lymphoblastic leukemia, leukemia, lymphoma, multiple myeloma, adrenocortical carcinoma, bladder cancer, bladder urothelial carcinoma, boneAttorney Docket No.36024-790601 cancer, brain lower grade glioma, breast cancer, breast invasive carcinoma, cervical squamous cell carcinoma, endocervical adenocarcinoma, cholangiocarcinoma, chronic lymphocytic leukemia, chronic myeloid disorders, colon adenocarcinoma, colorectal cancer, early onset prostate cancer, esophageal adenocarcinoma, esophageal carcinoma, gallbladder cancer, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney cancer, kidney chromophobe, kidney renal clear cell carcinoma, kidney renal papillary cell carcinoma, liver cancer, liver hepatocellular carcinoma lower grade glioma, lung cancer, lung adenocarcinoma, lung squamous cell carcinoma, lymphoid neoplasm diffuse large b-cell lymphoma, malignant lymphoma, mesothelioma, neuroblastoma, oral cancer, ovarian serous cystadenocarcinoma, pancreatic cancer endocrine neoplasms, pancreatic adenocarcinoma, pediatric brain cancer, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectum adenocarcinoma, renal cancer, sarcoma, secretory cancer, skin cutaneous melanoma, stomach adenocarcinoma, testicular germ cell tumor, testicular germ cell tumors, thymoma, thyroid carcinoma, uterine carcinosarcoma, uterine corpus endometrial carcinoma, uveal melanoma, adrenocortical carcinoma.
29. The method of any one of claims 27 to 28, wherein a primary tumor has not beenidentified in relation to the metastatic cancer.
30. The method of any one of claims 27 to 29, wherein a primary tumor was removed prior todevelopment of the metastatic cancer.
31. The method of any one of claims 1 to 30, wherein the sample is obtained or derived froma blood sample, a urine sample, or a saliva sample.
32. The method of any one of claims 1 to 31, wherein the sample is obtained or derived froma blood plasma sample.
33. The method of any one of claims 1 to 32, wherein the set of genomic segments isdetermined by a split criterion.
34. The method of claim 33, wherein the split criterion is based at least in part on a slidingwindow.
35. The method of any one of claims 1 to 34, wherein the major allelic fraction or the minorallelic fraction is represented as a matrix of base pairs, comprising a frequency of occurrence of A, T, C, G, an insertion, or a deletion.
36. The method of any one of claims 1 to 35, wherein the genomic sequence comprises amutation.Attorney Docket No.36024-79060137. The method of any one of claims 1 to 36, wherein the genomic sequence comprises agenomic loci.
38. A method comprising:a. assaying a diseased sample derived from a subject to produce a diseasedsequencing data, wherein the diseased sample is other than a solid tumor sample; b. determining one or more sequences indicative of a source based at least in part onthe diseased sequencing data; and c. assaying a sample comprising cell-free DNA (cfDNA) derived from the subject toproduce genomic data; d. generating an estimated frequency that a genomic sequence from the genomicdata arises from the source based at least in part on the one or more sequences indicative of the source.
39. The method of any one of claims 38, wherein c comprises assaying at least two samplescomprising cfDNA obtained at different time points.
40. The method of any one of claims 38 to 39, wherein the source is a cancer cell.
41. The method of any one of claims 38 to 39, wherein the source is a tissue type.
42. The method of any one of claims 38 to 39, wherein the source is a tumor.
43. The method of any one of claims 5 to 10, and 38 to 42, wherein the diseased samplecomprises a bodily fluid.
44. The method of claim 43, wherein the bodily fluid comprises urine.
45. The method of claim 43, wherein the bodily fluid comprises blood.
46. The method of any one of claims 5 to 10, and 38 to 45, wherein the diseased sample isprocessed to generate a cell-free sample and / or comprises a cell sample.
47. The method of any one of claims 5 to 10, and 38 to 46, wherein the diseased samplecomprises a plasma sample.
48. The method of any one of claims 5 to 10, and 38 to 47, wherein assaying the diseasedsample comprises sequencing cell-free nucleic acids.
49. The method of any one of claims 5 to 10, and 38 to 48, wherein assaying the diseasedsample further comprises sequencing genomic DNA.
50. The method of any one of claims 5 to 10, and 38 to 49, wherein determining the one ormore sequences indicative of the cancer comprises comparing sequences present in the genomic DNA of the subject and sequences present in cell-free nucleic acids of the subject.Attorney Docket No.36024-79060151. The method of claim 50, wherein the diseased sample comprises a buffy coat portion anda plasma portion, and wherein the sequences present in the genomic DNA of the subject are determined by assaying the buffy coat portion and the sequences present in cell-free nucleic acids of the subject are determined by assaying the plasma portion.
52. The method of any one of claims 5 to 10, and 38 to 51, wherein the diseased samplecomprises ctDNA.
53. The method of any one of claims 5 to 10, and 38 to 52, wherein the diseased sample isderived from the subject before the sample comprising cfDNA is derived from the subject.
54. The method of any one of claims 5 to 10, and 38 to 53, wherein the one or moresequences indicative of a cancer make up at least a portion of a tumor signature.
55. The method of any one of claims 5 to 10, and 38 to 54, wherein the one or moresequences indicative of the cancer comprises a set of genetic variants .
56. The method of any one of the preceding claims, wherein the estimated frequencycomprises a tumor fraction.
57. The method of any one of the preceding claims, further comprising determining thepresence or absence of a cancer, a recurrence of a cancer, and / or a metastasis of a cancer.
58. The method of any one of the preceding claims, further comprising determining thepresence or absence of a residual disease from a cancer.
Citation Information
Patent Citations
Algorithms for disease diagnostics
US20190100809A1
Cancer Classification with Genomic Region Modeling
US20210313006A1