Methods for tumor fraction analysis using spectral analysis methods
Spectral analysis of sequence data from biological samples using machine learning algorithms addresses the challenges of detecting cancer recurrence by accurately quantifying ctDNA and MRD, enabling early and cost-effective detection.
Patent Information
- Application Number
- PCT/US2025/035939
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-01
- Filing Date
- 2025-06-30
- Publication Date
- 2026-01-08
AI Technical Summary
Current methods for detecting cancer recurrence after initial treatment are haphazard, invasive, and unreliable, often requiring frequent and expensive patient visits, and genomic detection of residual cancer cells is challenging due to their small fraction and genetic similarity to healthy cells.
A method involving spectral analysis of sequence data from biological samples, including spectral analysis of fragment counts and local maxima, to determine tumor fraction, using machine learning algorithms and noise suppression techniques, to accurately quantify circulating tumor DNA (ctDNA) and detect minimal residual disease (MRD).
Enables early and accurate detection of cancer recurrence through minimally invasive methods, reducing costs and anxiety by providing reliable quantification of ctDNA and MRD, allowing for timely intervention.
Smart Images

Figure US2025035939_08012026_PF_FP_ABST
Abstract
Description
METHODS FOR TUMOR FRACTION ANALYSIS USING SPECTRAL ANALYSISMETHODSCROSS-REFERENCE
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 666,371, filed July 1, 2024, which is incorporated herein by reference in its entirety.BACKGROUND
[0002] Cancer recurrence is of foremost concern for patients and medical practitioners following initial cancer treatment. After interventions such as chemotherapy, radiotherapy, targeted drugs, immunotherapies, surgical excision, etc., a patient’s cancer may become clinically undetectable and nominally “cured”, upon which a patient is shifted to post-treatment monitoring. Unfortunately, cancers frequently do recur. Recurrence of a treated cancer may be diagnosed weeks, months, or even many years after the cancer became clinically undetectable using conventional detection methods. Thus, there is a need for methods to assist in detecting cancer.SUMMARY
[0003] The continuing risk of cancer recurrence even many years after initial treatment may present many clinical challenges. For example, clinical monitoring for recurrence may be haphazard, imperfect, invasive and require frequent expensive and inconvenient patient visits. Detection may rely on clinical symptoms, imaging, cell count, and / or biochemical cues, which may only be detectable after a recurrence has matured. Even if recurrence never happens, the costs, inconveniences, and anxieties relating to recurrence risk are significant. An estimated 7 percent of post-treatment cancer patients experience debilitating fear of recurrence having a significant negative impact on life activities (as described by Butow P, et al. Fear of cancer recurrence: a practical guide for clinicians. Oncology. 2018; 32:32-38, which is incorporated by reference herein in its entirety).
[0004] Attempts have been made to meet these problems using genomic detection methods based. If a simple serum test may accurately detect and quantify minute quantities of genetic material and / or markers signaling lingering or recurring cancer, doctors may be able to detect recurrence early using low-cost minimally invasive techniques. However, (a) the residual cancer that may exist in a treated patient’s body represents an extremely small fraction of the overall genetic material present in that patient’s body; and (b) the geneticmaterial of cancer cells may differ very slightly from healthy cells. These conditions make accurate and reliable genomic detection and measurement a very challenging problem.
[0005] Consequently, there is a long-felt unmet need for methods and systems for accurately and reliably detecting, identifying, and quantifying cancer genetic material (if any) circulating in a patient.
[0006] In some aspects, the present disclosure provides systems and methods for determining whether a subject (e.g., patient) has certain genetic material circulating in the body.
[0007] In an aspect the present disclosure describes or provides a method of estimating a tumor fraction of a subject, comprising: (a) providing a biological sample obtained or derived from the subject; (b) assaying the biological sample to produce a sequence data; (c) applying a spectral analysis method to at least a portion of the sequence data to determine the tumor fraction of the biological sample.
[0008] In some embodiments, the sequence data comprises a measure of a fragment count for a genome position and / or a measure of an occurrence for the fragment count.
[0009] In some embodiments, the spectral analysis is applied to the sequence data comprising the measure of the occurrence for the fragment count versus the measure of the fragment count for the genome position.
[0010] In some embodiments, applying the spectral analysis method comprises detecting one or more local maxima in at least a portion of the sequence data.
[0011] In some embodiments, the detecting the one or more local maxima comprises detecting one or more local maxima in the at least a portion of the sequence data comprising the measure of the occurrence for the fragment count versus the measure of the fragment count for the genome position.
[0012] In some embodiments, applying the spectral analysis method comprises determining a frequency of the one or more local maxima.
[0013] In some embodiments, the tumor fraction is determined based at least in part on the one or more local maxima and / or the frequency of the one or more local maxima.
[0014] In some embodiments, the spectral analysis comprises determining a likelihood of a particular tumor fraction corresponding to the biological sample based at least in part on the frequency of the one or more local maxima.
[0015] In some embodiments, determining the tumor fraction comprises determining the particular tumor fraction with the maximum likelihood.
[0016] In some embodiments, the biological sample comprises a cell-free DNA (cfDNA), and wherein assaying the biological sample comprises assaying the cfDNA.
[0017] In some embodiments, the genome position is a bin.
[0018] The method of any one of the preceding claims, further comprising detecting and / or removing a set of outliers from the sequence data.
[0019] In some embodiments, wherein detecting and / or removing the set of outliers comprises applying a clustering algorithm.
[0020] In some embodiments, wherein the clustering algorithm comprises DBSCAN.
[0021] In some embodiments, the methods and systems provided comprise determining a pattern comprising one or more copy numbers corresponding to one or more of the genome positions.
[0022] The method of any one of the preceding claims, wherein the subject has or is suspected of having cancer.
[0023] In some embodiments, wherein the cancer comprises fibrosarcoma, myosarcoma, liposarcoma, chondrosarcoma, osteogenic sarcoma, chordoma, angiosarcoma, endotheliosarcoma, lymphangiosarcoma, lymphangio endotheliosarcoma, synovioma, mesothelioma, Ewing’s tumor, leiomyosarcoma, rhabdomyosarcoma, colon carcinoma, pancreatic cancer, breast cancer, ovarian cancer, prostate cancer, bladder cancer, lung cancer, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinomas, cystadenocarcinoma, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatoma, bile duct carcinoma, choriocarcinoma, seminoma, embryonal carcinoma, Wilms tumor, cervical cancer, uterine cancer, testicular cancer, lung carcinoma, small cell lung carcinoma, bladder carcinoma, kidney cancer, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pinealoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, schwannoma, meningioma, melanoma, neuroblastoma, retinoblastoma, colorectal cancer, endometrial cancer, endometroid cancer, clear cell carcinoma, or any combination thereof. In some embodiments, wherein the cancer comprises breast cancer. In some embodiments, wherein the cancer comprises lung cancer. In some embodiments, wherein the cancer comprises bladder cancer. In some embodiments, the cancer comprises kidney cancer.
[0024] In some embodiments, the subject has or is suspected of having cancer remission.
[0025] In some embodiments, the subject has or is suspected of having cancer recurrence.
[0026] In some embodiments, the biological sample is selected from the group consisting of: a blood sample, a plasma sample, a serum sample, a urine sample, a saliva sample, a cell sample, and a tissue sample. In some embodiments, the biological sample is the plasma sample.
[0027] In some embodiments, the assaying comprises a technique selected from the group consisting of: nucleic acid sequencing, nucleic acid amplification, nucleic acid enrichment, microarray, reverse transcription, and a combination thereof.
[0028] In some embodiments, the sequence data is normalized. In some embodiments, the sequence data is normalized by at least applying a neutral correction function. In some embodiments, the methods and systems provided comprise noise removal comprising using a sample with known characteristics.
[0029] In some embodiments, determining the tumor fraction comprises using a machine learning algorithm. In some embodiments, the machine learning algorithm comprises noise suppression, signal processing, or both. In some embodiments, the machine learning algorithm comprises at least one of a neutral learning model, and gradient boosted decision tree.
[0030] In some embodiments, the methods and systems provided comprise administering a therapy to the subject at least in part based on the tumor fraction determined in (c).
[0031] In some embodiments, the methods and systems provided comprise comprising determining a presence or an absence of minimal residual disease (MRD) in the subject, based at least in part on the tumor fraction. In some embodiments, the methods and systems provided comprise administering a therapy to the subject, thereby treating the MRD. In some embodiments, the methods and systems provided comprise determining a likelihood of recurrence of cancer based at least in part on the tumor fraction.
[0032] In some embodiments, the methods and systems provided comprise a therapy to the subject, based at least in part on the determined likelihood of recurrence of cancer. In some embodiments, the therapy comprises chemotherapy, radiotherapy, immunotherapy, targeted therapy, surgical resection, laser ablation, or any combination thereof.
[0033] In some embodiments, the methods and systems provided comprise providing a report comprising a copy number variation (CNV) data. In some embodiments, the sequence data comprises a copy number variation (CNV) data. In some embodiments, the sequence data further comprises single-nucleotide-variant (SNV) data, and wherein the SNV data is corrected based at least in part on the tumor fraction.INCORPORATION BY REFERENCE
[0034] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.BRIEF DESCRIPTION OF THE FIGURES
[0035] Certain novel features are set forth with particularity in the appended claims. A better understanding of the features and advantages will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles are utilized, and the accompanying drawings (also “Figure” and “FIG.” herein), of which:
[0036] Features, aspects, and advantages of some embodiments of the present disclosure may be better understood when read with reference to the accompanying nonlimiting drawings, wherein:
[0037] Fig. 1 shows an example schematic relating to tumor fraction (TF).
[0038] Figs. 2A-2B show example data of some embodiments of a method of the present disclosure.
[0039] Fig. 3 shows an example of some embodiments of a workflow (using, e.g., a tumor sample).
[0040] Fig. 4 shows an example of some embodiments of a workflow (using, e.g., a normal sample).
[0041] Fig. 5 shows an example of some embodiments of a workflow relating to a tumor fraction determination (using, e.g., a plasma sample).
[0042] Fig. 6 shows an example of some embodiments of a workflow relating to normalization.
[0043] Figs. 7A-7C shows an example of some embodiments relating to outlier removal.
[0044] Figs. 8A-8C show examples of some embodiments of data relating to and of a workflow relating to normalization.
[0045] Fig. 9 shows an example of some embodiments of a workflow relating to tumor fraction estimation.
[0046] Figs. 10A-10B show examples of some embodiments of a data generated in the process of a tumor fraction estimation workflow and of a block diagram of a workflow.
[0047] Figs. 11A-11B show an example of some embodiments of data generated in the process of tumor fraction estimation workflow and of a block diagram of a workflow.
[0048] Fig. 12 shows an example of some embodiments of a workflow relating to normalization.
[0049] Fig. 13 shows an example of some embodiments of workflow relating to tumor fraction estimation.
[0050] Fig. 14 shows an example of some embodiments of computer system that may be programmed or otherwise configured to implement methods provided herein.DETAILED DESCRIPTION
[0051] While various embodiments have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Variations, changes, and substitutions may occur to those skilled in the art. It should be understood that various alternatives to the embodiments described herein may be employed.
[0052] As used in the specification and claims, the singular form “a”, “an”, and “the” include plural references unless the context clearly dictates otherwise. For example, the term “a nucleic acid” includes a plurality of nucleic acids, including mixtures thereof.
[0053] As used herein, the term “subject,” generally refers to an entity or a medium that has testable or detectable genetic information. A subject can be a person, individual, or patient. A subject can be a vertebrate, such as, for example, a mammal. Non-limiting examples of mammals include humans, simians, farm animals, sport animals, rodents, and pets. The subject may have cancer or be suspected of having cancer. The subject may be displaying a symptom(s) indicative of cancer. As an alternative, the subject can be asymptomatic with respect to such cancer.
[0054] As used herein, the term “biological sample,” generally refers to a sample obtained from or derived from one or more subjects (e.g., with or without further processing steps prior to any analysis and / or obtaining any data from the sample). Biological samples may be cell-free biological samples or substantially cell-free biological samples, or may be processed or fractionated to produce cell-free biological samples. For example, cell-free biological samples may include cell-free ribonucleic acid (cfRNA), cell-free deoxyribonucleic acid (cfDNA), cell-free fetal DNA (cffDNA), plasma, serum, urine, saliva, and derivatives thereof. Cell-free biological samples may be obtained or derived fromsubjects using an ethylenediaminetetraacetic acid (EDTA) collection tube, a cell-free RNA collection tube (e.g., Streck), or a cell-free DNA collection tube (e.g., Streck). Cell-free biological samples may be derived from whole blood samples by fractionation. Biological samples or derivatives thereof may contain cells. For example, a biological sample may be a blood sample or a derivative thereof (e.g., blood collected by a collection tube or blood drops).
[0055] As used herein, the term “nucleic acid” generally refers to a polymeric form of nucleotides of any length, either deoxyribonucleotides (dNTPs) or ribonucleotides (rNTPs), or analogs thereof. Nucleic acids may have any three-dimensional structure, and may perform any function, known or unknown. Non-limiting examples of nucleic acids include deoxyribonucleic (DNA), ribonucleic acid (RNA), coding or non-coding regions of a gene or gene fragment, loci (locus) defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, short interfering RNA (siRNA), short-hairpin RNA (shRNA), micro-RNA (miRNA), ribozymes, cDNA, recombinant nucleic acids, branched nucleic acids, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. A nucleic acid may comprise one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure may be made before or after assembly of the nucleic acid. The sequence of nucleotides of a nucleic acid may be interrupted by non-nucleotide components. A nucleic acid may be further modified after polymerization, such as by conjugation or binding with a reporter agent.
[0056] As used herein, the term “target nucleic acid” generally refers to a nucleic acid molecule in a starting population of nucleic acid molecules having a nucleotide sequence whose presence, amount, and / or sequence, or changes in one or more of these, are desired to be determined. A target nucleic acid may be any type of nucleic acid, including DNA, RNA, and analogs thereof. As used herein, a “target ribonucleic acid (RNA)” generally refers to a target nucleic acid that is RNA. As used herein, a “target deoxyribonucleic acid (DNA)” generally refers to a target nucleic acid that is DNA.
[0057] As used herein, the terms “amplifying” and “amplification” generally refer to increasing the size or quantity of a nucleic acid molecule. The nucleic acid molecule may be single-stranded or double-stranded. Amplification may include generating one or more copies or “amplified product” of the nucleic acid molecule. Amplification may be performed, for example, by extension (e.g., primer extension) or ligation. Amplification may include performing a primer extension reaction to generate a strand complementary to a single-stranded nucleic acid molecule, and in some cases generate one or more copies of the strand and / or the single-stranded nucleic acid molecule. The term “DNA amplification” generally refers to generating one or more copies of a DNA molecule or “amplified DNA product.” The term “reverse transcription amplification” generally refers to the generation of deoxyribonucleic acid (DNA) from a ribonucleic acid (RNA) template via the action of a reverse transcriptase.(i) In various aspects, the present disclosure provides methods and systems of estimating a tumor fraction of a subject. The method may comprise: (a) providing a biological sample obtained or derived from the subject; (b) assaying the biological sample to determine a tumor signature; and (c) applying a spectral analysis method in conjunction with a machine learning algorithm to the tumor signature to determine the tumor fraction of the biological sample. In an aspect, the present disclosure provides a method of estimating a tumor fraction of a subject, comprising: (a) providing a biological sample obtained or derived from the subject; (b) assaying the biological sample to produce a sequence data; (c) applying a spectral analysis method to at least a portion of the sequence data to determine the tumor fraction of the biological sample.
[0058] In some embodiments, the subject has or is suspected of having cancer. In some embodiments, the cancer comprises fibrosarcoma, myosarcoma, liposarcoma, chondrosarcoma, osteogenic sarcoma, chordoma, angiosarcoma, endotheliosarcoma, lymphangiosarcoma, lymphangio endotheliosarcoma, synovioma, mesothelioma, Ewing’s tumor, leiomyosarcoma, rhabdomyosarcoma, colon carcinoma, pancreatic cancer, breast cancer, ovarian cancer, prostate cancer, bladder cancer, lung cancer, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinomas, cystadenocarcinoma, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatoma, bile duct carcinoma, choriocarcinoma, seminoma, embryonal carcinoma, Wilms tumor, cervical cancer, uterine cancer, testicular cancer, lung carcinoma, small cell lung carcinoma, bladder carcinoma, kidney cancer, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pinealoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, schwannoma, meningioma, melanoma, neuroblastoma, retinoblastoma, colorectal cancer, endometrial cancer, endometroid cancer, clear cell carcinoma, or any combination thereof. In some embodiments, the cancer comprises breast cancer. In some embodiments, the cancer comprises lung cancer. In some embodiments, the cancer comprises lung cancer. In some embodiments, the cancer comprises bladder cancer. In someembodiments, the cancer comprises bladder cancer. In some embodiments, the cancer comprises kidney cancer.
[0059] In some embodiments, the subject has or is suspected of having cancer remission.
[0060] In some embodiments, the subject has or is suspected of having cancer recurrence.
[0061] In some embodiments, the biological sample is selected from the group consisting of: a blood sample, a plasma sample, a serum sample, a urine sample, a saliva sample, a cell sample, and a tissue sample. In some embodiments, the biological sample is the plasma sample.
[0062] In some embodiments, the biological sample is a set of biological samples, wherein (b) comprises assaying the set of biological samples to determine a set of tumor signatures, and wherein (c) comprises determining the tumor fraction of the set of biological samples.
[0063] In some embodiments, (b) further comprises assaying nucleic acids obtained or derived from the biological sample to generate a set of nucleic acid sequences. In some embodiments, the assaying comprises a technique selected from the group consisting of: nucleic acid sequencing, nucleic acid amplification, nucleic acid enrichment, microarray, reverse transcription, and a combination thereof.
[0064] In some embodiments, the nucleic acids comprise cell-free deoxyribonucleic acid (cfDNA). In some embodiments, the biological sample comprises a cell-free DNA (cfDNA).
[0065] In some embodiments, the method further comprises determining a fraction of putative single-nucleotide-variant (SNV) reads from the set of nucleic acid sequences that are read errors, and / or a fraction of putative SNV reads from the set of nucleic acid sequences that are true mutations. In some embodiments, the method further comprises detecting and removing a set of outliers from the set of nucleic acid sequences.
[0066] In some embodiments, the method further comprises applying a normalization to the set of nucleic acid sequences, thereby generating normalized sequencing data. In some embodiments, the normalization comprises a neutral correction function. In some embodiments, the sequence data is normalized by at least applying a neutral correction function. In some embodiments, the neutral correction function is expressed by: Normalized count(i) = count(i) / fi, wherein f is a normalization factor and count(i) is a feature count.
[0067] In some embodiments, the method further comprises sorting the normalized sequencing data into a set of genomic bins, and determining quantitative measures of the normalized sequencing data in each of the set of genomic bins to produce bin-wise count data. In some embodiments, the method further comprises applying a transformation to the bin-wise count data to be within a given range. In some embodiments, the method further comprises producing the bin-wise count data according to the expression: count(i) / count_0(i) = (l+TF / 2) * CopyNumber(i), wherein “count(i)” is the normalized bin-wise count for the biological sample at a given position in the normalized sequencing data, wherein “count_0(i)” is a normalized bin-wise count for the biological sample at a position that is earlier than the given position in the normalized sequencing data, and wherein CopyNumber(i) is a function indicative of a measure of copy number variation for the biological sample. In some embodiments, the method further comprises determining a histogram of counts per genomic location in the normalized sequencing data. In some embodiments, the method further comprises determining a measure of peak frequency of the histogram. In some embodiments, the tumor fraction is determined based at least in part on the measure of peak frequency.
[0068] In some embodiments, the machine learning algorithm uses noise suppression, signal processing, or both.
[0069] In some embodiments, the method further comprises determining a presence or an absence of minimal residual disease (MRD) in the subject, based at least in part on the tumor fraction determined in (c). In some embodiments, the method further comprises administering a therapy to the subject, thereby treating the MRD. In some embodiments, the therapy comprises chemotherapy, radiotherapy, immunotherapy, targeted therapy, surgical resection, laser ablation, or any combination thereof. In some embodiments, the method further comprises administering the therapy to the subject based at least in part on a likelihood of recurrence of cancer.
[0070] In some embodiments, the method further comprises determining providing a copy -number variation (CNV) report.
[0071] In some embodiments, the method further comprises applying a clustering algorithm to produce clustering data. In some embodiments, the clustering algorithm comprises DBSCAN.
[0072] In some embodiments, the machine learning algorithm comprises at least one of a neutral learning model, gradient boosted decision tree, and xgboost.
[0073] In another aspect, the present disclosure provides a method of analyzing copy number variation (CNV) of a subject, comprising: (a) providing a biological sample obtained or derived from the subject; (b) assaying the biological sample to determine a tumor signature; and (c) applying a spectral analysis method in conjunction with a machine learning algorithm to the tumor signature to determine a measure of CNV of the biological sample.
[0074] In some embodiments, the subject has or is suspected of having cancer. In some embodiments, the cancer comprises fibrosarcoma, myosarcoma, liposarcoma, chondrosarcoma, osteogenic sarcoma, chordoma, angiosarcoma, endotheliosarcoma, lymphangiosarcoma, lymphangio endotheliosarcoma, synovioma, mesothelioma, Ewing’s tumor, leiomyosarcoma, rhabdomyosarcoma, colon carcinoma, pancreatic cancer, breast cancer, ovarian cancer, prostate cancer, bladder cancer, lung cancer, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinomas, cystadenocarcinoma, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatoma, bile duct carcinoma, choriocarcinoma, seminoma, embryonal carcinoma, Wilms tumor, cervical cancer, uterine cancer, testicular cancer, lung carcinoma, small cell lung carcinoma, bladder carcinoma, kidney cancer, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pinealoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, schwannoma, meningioma, melanoma, neuroblastoma, retinoblastoma, colorectal cancer, endometrial cancer, endometroid cancer, clear cell carcinoma, or any combination thereof. In some embodiments, the cancer comprises breast cancer. In some embodiments, the cancer comprises lung carcinoma. In some embodiments, the cancer comprises bladder carcinoma. In some embodiments, the cancer comprises kidney cancer.
[0075] In some embodiments, the subject has or is suspected of having cancer remission.
[0076] In some embodiments, the subject has or is suspected of having cancer recurrence.
[0077] In some embodiments, the biological sample is selected from the group consisting of: a blood sample, a plasma sample, a serum sample, a urine sample, a saliva sample, a cell sample, and a tissue sample. In some embodiments, the biological sample is the plasma sample.
[0078] In some embodiments, the biological sample is a set of biological samples, wherein (b) comprises assaying the set of biological samples to determine a set of tumorsignatures, and wherein (c) comprises determining the measure of CNV of the set of biological samples.
[0079] In some embodiments, (b) further comprises assaying nucleic acids obtained or derived from the biological sample to generate a set of nucleic acid sequences. In some embodiments, the assaying comprises a technique selected from the group consisting of: nucleic acid sequencing, nucleic acid amplification, nucleic acid enrichment, microarray, reverse transcription, and a combination thereof. In some embodiments, the nucleic acids comprise cell-free deoxyribonucleic acid (cfDNA). In some embodiments, the method further comprises determining a fraction of putative single-nucleotide-variant (SNV) reads from the set of nucleic acid sequences that are read errors, and / or a fraction of putative SNV reads from the set of nucleic acid sequences that are true mutations. In some embodiments, the method further comprises detecting and removing a set of outliers from the set of nucleic acid sequences. In some embodiments, the method further comprises applying a normalization to the set of nucleic acid sequences, thereby generating normalized sequencing data.
[0080] In some embodiments, the machine learning algorithm uses noise suppression, signal processing, or both.
[0081] In some embodiments, the method further comprises determining a presence or an absence of minimal residual disease (MRD) in the subject, based at least in part on the measure of CNV.
[0082] In some embodiments, the method further comprises administering a therapy to the subject based on a tumor fraction. In some embodiments, the method further comprises administering a therapy to the subject, thereby treating the MRD. In some embodiments, the therapy comprises chemotherapy, radiotherapy, immunotherapy, targeted therapy, surgical resection, laser ablation, or any combination thereof. In some embodiments, the method further comprises administering the therapy to the subject based at least in part on a likelihood of recurrence of cancer.
[0083] In some embodiments, the method further comprises applying a clustering algorithm to produce clustering data. In some embodiments, the clustering algorithm comprises DBSCAN.
[0084] In some embodiments, the machine learning algorithm comprises at least one of a neutral learning model, gradient boosted decision tree, and xgboost.
[0085] Another aspect of the present disclosure provides a non-transitory computer readable medium comprising machine executable code that, upon execution by one or more computer processors, implements any of the methods above or elsewhere herein.
[0086] Another aspect of the present disclosure provides a system comprising one or more computer processors and computer memory coupled thereto. The computer memory comprises machine executable code that, upon execution by the one or more computer processors, implements any of the methods above or elsewhere herein.
[0087] The systems and methods of the present disclosure may be used to determine whether a subject (e.g., patient) has certain genetic material, such as ctDNA, circulating in the body, and may aid in quantifying the genetic material.
[0088] In some embodiments, one important use of the methods and systems of this disclosure relate to detection and monitoring of minimal residual disease (MRD). Many cancer patients who are treated and become clinically “cancer-free” face the unnerving prospect of recurrent cancer. Because cancer cells may break free from a tumor mass and circulate to another location in a patient’s body, recurrence may be local (recurring in the same location on the patient’s body), regional (recurring in lymph nodes or other tissues proximate to the original site), or distant (found in tissue far from the original site). Recurrent cancer over emerges from cells that were selected to resist treatment during a first round of cancer treatment. Further, distant recurrent cancers are often selected for their propensity to circulate and metastasize. Accordingly recurrent cancer may be much more dangerous than an initial cancer.
[0089] For clinicians and such patients, there exists an urgent need for develop methods and systems for affordable, minimally invasive, early warning and / or early detection of recurrent cancer.
[0090] Various clinical metrics may be used for measuring the continuing or returning presence of cancers in patients whose cancer is in remission. One such measure is “minimal residual disease” (MRD). MRD may refer to a condition in which any cancer cells remain in a subject’s (e.g., patient’s) body following cancer treatment that cannot be detected by conventional clinical scans or tests. A patient may be clinically diagnosed as in “complete remission” because there is no clinical evidence of cancer based on conventional scans or laboratory tests, but may nevertheless harbor some MRD. MRD is frequently used in discussion of the treatment of blood cancers such as leukemias, lymphomas, and / or multiple myelomas, where the cancer cells circulate and do not necessarily comprise a defined tumor mass or carcinoma. However, MRD may be used to describe residual cancer cells in any type of cancer.
[0091] Some methods of detecting MRD include flow-cytometric immunophenotyping, cell culture systems, fluorescent in situ hybridization, and polymerasechain reaction (PCR) amplification techniques. Each of these, however, may have significant drawbacks. First, such methods are not able to detect very small quantities of MRD, meaning they cannot provide clinicians and patients with the earliest possible warning of residual disease. Second, such methods can be invasive, expensive, time-consuming, and require skilled laboratory technicians. Third, such extant methods may not be able to detect MRD if phenotype and / or genetic mutations fail to meet certain necessary parameters for detection (for example if there is no obvious target for PCR analysis).
[0092] The methods and systems of the present disclosure provide significant advantages over prior MRD detection methodologies. Any MRD in the patient’s body will leave a genetic signature in the form of ctDNA. The present disclosure provides systems and methods that would permit a clinician to draw a patient’s blood, perform relatively simply sequencing methods on blood plasma, and quantify through the detection of biomarkers (e.g., CNVs, SNVs) and application of algorithms disclosed herein the fraction of DNA circulating which is ctDNA (“ctDNA fraction”).
[0093] Detecting and accurately quantifying tiny ctDNA factions is highly challenging for several reasons, including at least: the fraction of ctDNA versus the patient’s normal somatic DNA is likely to be extremely small; ctDNA often comprises genetic mutations that are unknown to the clinician; ctDNA often comprises only a very small number of single-nucleotide (SNP) mutations; and extant genetic sequencing techniques experience read errors that are very difficult to distinguish from true mutations, especially when the mutations are often so subtle.
[0094] The present disclosure provides methods and systems that overcome such challenges. Provided herein are systems and methods for detection of the frequency of ctDNA in a subject. The systems and methods provided herein comprises assaying polynucleotides to identify biomarkers of cancers and circulating tumor DNA (ctDNA) from cancer (cfDNA) in a subject. The biomarkers may be processed in order to identify the presence or absence of cancer or cfDNA. The methods described herein may process multiple type of analytes in order to determine a presence or absence of cancer. The multiple types of analytes may comprise DNA or RNA, for example cfDNA or cfRNA. The multiple analytes may be cfDNA, germline DNA, and / or cfRNA. By analyzing a plurality of different analytes, the methods may allow for improved detection or determination of a prognosis as compared to methods performed on fewer analytes or only one of many different analytes.
[0095] FIG. 1 depicts a schematic conceptualizing of a copy -number variation for a non-cancerous (e.g., “normal”) sample (upper-left, having a constant copy-numberthroughout the sequence), a purely cancerous sample (upper-right, having varying copynumber along sequence loci), and a mixed normal + cancerous sample with a low fraction of cancerous DNA (bottom-center, showing some copy -number variation, with a muted variation signal because the fraction of cancerous DNA is low), in accordance with some embodiments. In some embodiments, in the case of minimal residual disease (MRD), the fraction of remaining tumor DNA circulating in the patient’s blood plasma may be expected to be very low. As depicted in FIG.l, in some embodiments, a low-tumor fraction sample (such as some plasma samples) may comprise circulating cell-free DNA (cfDNA) which arises from normal cells (e.g., as in a normal sample in FIG. 1) and cancer cells (e.g., as in a pure cancer sample). cfDNA samples (e.g., from plasma) may show CNV data that is representative of combination of the CNVs of normal cells and the cancer cells, with sample comprising a high tumor fraction appearing more similar to the data of cancer cells, and samples with low tumor fraction appearing more similar to data of the normal cells. As such, it may be possible to determine or estimate a tumor fraction of a sample by analyzing the CNV data in one of more genomic regions. In some embodiments, the present disclosure provides methods and systems to estimate tumor fraction of a sample in various scenarios, such as in a tumor naive and tumor informed scenario.
[0096] In some embodiments, the methods and systems provided herein may detect a ctDNA fraction (e.g., of a cfDNA) in a sample which may be derived from a subject. In some embodiments, the sample may be assayed (e.g., with or without further processing). In some embodiments, the processed sample may be sequenced to produce sequence data.
[0097] In some embodiments the methods and systems provided herein may estimate a tumor fraction of a subject. In some embodiments, the method and systems may comprise providing a biological sample obtained or derived from the subject; assaying the biological sample to produce a sequence data; and applying a spectral analysis method to at least a portion of the sequence data to determine the tumor fraction of the biological sample
[0098] In some embodiments, a sequence data may comprise data directly derived from a biological sample or data derived by further processing, modification, and / or alteration of the data directly derived from a biological sample. In some embodiments, sequence data may comprise data pertaining to a nucleic acid composition of a biological sample. In some embodiments, sequence data may comprise data generated from sequencing nucleic acid from a biological sample. For example, sequence data may comprise the nucleic acid sequences of the nucleic acids in the biological sample. In some embodiments, sequence data may comprise data pertaining to the nucleic acid molecules and / or data that can begenerated (e.g., with or without further processing) by sequencing the nucleic acids. For example, sequence data may comprise data pertaining to a number of sequencing reads or the number of nucleic acid fragments that map to a particular genomic region. In some embodiments, sequence data may comprise copy number variation (CNV) data. In some embodiments, sequence data may comprise single nucleotide variant (SNV) data. In some embodiments, sequence data may comprise data generated (e.g., with or without further processing) from sequencing DNA from a biological sample. In some embodiments, sequence data may comprise data from (e.g., with or without further processing) sequencing RNA from a biological sample. In some embodiments, sequence data may be data that is derived (e.g., with further processing) from data that arises from a biological sample. In some embodiments, sequence data may comprise data that has been processed (e.g., data that’s been generated upon processing of the data that directly arises from a biological sample). In some embodiments, sequence data may comprise data generated by one or more processes (e.g., such as separate datasets which underwent a different set of processing steps or were generated from different steps in a processing pipeline). In some embodiments, processes may include binning, normalization, log-transforms, producing a quantitative measure, converting data into a different format, determining dependency of one or more measures versus one or more of different measures, processing data using machine learning, or any combination thereof. In some embodiments, quantitative measures may include, but are not limited to, fragment counts, a measure of occurrence (e.g., frequency) for the fragment count, a measure of fragment count for a genome positions, a measure of occurrence of fragment count versus the measure of the fragment count for the genome position. In some embodiments, sequence data may comprise a measure of copy number variation. In some embodiments, sequence data may comprise quantitative data. In some embodiments, sequence data may comprise data that can be processed or manipulated based at least in part on information gathered from the sequence data (or other data obtained elsewhere). For example, the sequence data may comprise SNV data, and the SNV data may be corrected based at least in part on a tumor fraction (which may be determined via processing the sequence data (e.g., using a method described in this disclosure), or determined via an orthogonal or separate method).
[0099] In some embodiments, the sequence data may comprise a measure of a fragment count for a genome position and / or a measure of an occurrence for the fragment count. In some embodiments, a genome position may be a location or range of locationswithin sequence data. In some embodiments, a genome position may be a bin or set of bins. In some embodiments, a genome position may be a segment.
[0100] In various embodiments, sequence data (e.g., from plasma, donor, tumor, or normal samples) may be processed and / or comprise data that has been processed such that that it may be used for downstream processes (e.g., determining a tumor fraction). Processing may comprise binning or segmenting to generate bins or segments. Bins (e.g., segments) may be variable in length. Bins may be of the same length. Bins may be at least 1 bp in length, at least 10 bp in length, at least 20 bp in length, at least 30 bp in length, at least 40 bp in length, at least 50 bp in length, at least 60 bp in length, at least 70 bp in length, at least 80 bp in length, at least 90 bp in length, at least 100 bp in length, at least 110 bp in length, at least 120 bp in length, at least 140 bp in length, at least 150 bp in length, at least 160 bp in length, at least 170 bp in length, at least 180 bp in length, at least 190 bp in length, at least 200 bp in length, at least 300 bp in length, at least 400 bp in length, at least 500 bp in length, at least 1000 bp in length, at least 2000 bp in length, at least 3000 bp in length, at least 4000 bp in length, at least 5000 bp in length, at least 10000 bp in length, at least 40000 bp in length, at least 50000 bp in length, at least 60000 bp in length, at least 70000 bp in length, at least 80000 bp in length, at least 90000 bp in length, at least 100000 bp in length, or any combination thereof. Bins may be up to 200000 bp in length. Bins may be up to 300000 bp in length, bins may be up to 400000 bp in length, bins may be up to 500000 bp in length, bins may be up to 600000 bp in length, Bins may be up to 700000 bp in length, Bins may be up to 800000 bp in length, bins may be up to 900000 bp in length, bins may be up to 1000000 bp in length, bins may be up to 2000000 bp in length, bins may be up to 300000 bp in length. In some embodiments, the bins may comprise a copy number or a copy number variation (CNV). In some embodiments, bins may comprise one or more genetic variants, one or more structural variants, or any combination thereof. In some embodiments, a bin in the set of bins may comprise a SNV, a SNP, an indel, a deletion, an insertion, a duplication, an inversion, a translocation, a copy number variation, a tandem repeat, variable number tandem repeat, a mobile element insertion, a transposable element, a gene fusion, an epigenetic modification, or any combination thereof. In some embodiments, the set of genomic bins may comprise a SNV, a SNP, an indel, a deletion, an insertion, a duplication, an inversion, a translocation, a copy number variation, a tandem repeat, variable number tandem repeat, a mobile element insertion, a transposable element, a gene fusion, an epigenetic modification, or any combination thereof. In some embodiments, the reads in a given bin may be counted and the binned read counts may be outputted and / or used in downstream processes.
[0101] An example of an embodiment of a binning is depicted in FIG. 2A, wherein bins are generated of lOObp in length making 30000 bins (x-axis) from the human genome. The y- axis shows the count per bin. In some embodiments, a measure of the occurrence for the fragment count versus the measure of the fragment count for the genome position may be generated and / or used. In FIG. 2A, an example of a histogram of read count distribution is shown alongside the y-axis; FIG. 2B depicts an example of a histogram (the histogram alongside the y axis in FIG. 2A) split into sections based on peaks to indicate peak frequency.
[0102] In various embodiments, a histogram (e.g., such as in FIG. 10A) may be used to describe data (for example, the distribution of such data, a measure of a frequency or number of occurrences of an action or object, dependency of one or more measures versus one or more other measures). In some embodiments, a graphical representation of a histogram may be generated; in some embodiments, no graphical representation is generated. In various embodiments, a histogram may be referred to wherein an operation, function, or transformation may be applied, carried out on, or otherwise operated on and / or from. In various embodiments, a histogram may be an output (or part thereof; e.g., such as a report or a part thereof). In some embodiments, the output of a histogram may comprise outputting the numerical data corresponding to the histogram, a visualization of the histogram, or both. In various embodiments, the axis of a histogram may be referred to as pertaining to a metric or a parameter as to which the data relates and is not necessarily limited to a graphical representation. In some embodiments, the axis may be used in a graphical representation when appropriate.
[0103] In some embodiments, the sequence data (or a portion thereof) may be analyzed using computational methods to determine the tumor fraction (e.g., ctDNA fraction in a cfDNA) of the biological sample. The computational methods may be applied to the sequence data to carry out the analysis. In some embodiments, the computational method may output a report. In some embodiments, the computational methods may output a pattern. In some embodiments, the computational method may output a report and a pattern. In some embodiments, the sequencing process may introduce noise to the sequencing data of the sample. Various methods and processes may be used to reduce, eliminate, or otherwise account for noise. For example, in some embodiments, the computational method may comprise taking sequencing data from donor samples as input. In some embodiments, donor sequence data may be raw sequencing data, or processed sequencing data. In some embodiments, donor samples may comprise control samples or be used as a control sample.In some embodiments, donor samples may have known characteristics (such as tumor fraction, noise signatures, noise, CNV levels, SNV signatures, or any combination thereof).
[0104] In some embodiments, the method may comprise providing a sample derived from a tumor (such as a biopsy). In some embodiments, such methods may be referred to as “tumor informed”. In some embodiments, tumor derived samples may be more pure than other sample types that provide ctDNA (such as a plasma sample, a urine sample, or a fecal sample). This may provide a relatively noise free sample due to it having higher percentage of tumor derived nucleic acids. Additionally, tumor derived samples may provide more information about the heterogeneity present in a tumor. This may be, in part, due to the lower level of noise. If a mutation or CNV profile is present at very low levels in the tumor it may still be present in a cfDNA / ctDNA sample but be undetectable above the noise in the sample. In some embodiments, by collecting a tumor derived sample or using data derived from a tumor derived sample, a baseline (e.g., a pattern) of known mutations and / or CNVs may be established for a subject’s cancer (versus, for example, against a database). This baseline may be used to detect ctDNA fraction in higher noise settings / low ctDNA fraction settings (such as in MRD detection).
[0105] Some embodiments of a method of estimating a tumor fraction of a subject is depicted in FIG. 3. An example of some embodiments of a tumor informed method is also illustrated in FIG. 3. Circles denote possible outputs and / or inputs and boxes denote processing steps in the example method. As shown in FIG. 3, data for a tumor derived sample (e.g., which could be in a form of a Binary Alignment Map file format (the “BAM” input circle)) may comprise, e.g., sequencing reads. In some embodiments, the reads may be binned and counted and the binned read counts may be outputted (e.g., in the “wig” filetype output). In some embodiments, the binned read counts may then be normalized. An example of an embodiment of a normalized read counts is presented in FIGs. 8A-8B where an example of raw count data is depicted in FIG. 8A and an example of normalized count data is presented in FIG. 8B. In some embodiments, the normalized count may have less variability in the read count data, as normalization may remove noise inherent in the sample or sequencing method and / or normalize the data to a baseline. In some embodiments, downstream methods may thus benefit from the normalized data which may be less noisy and / or may have a more discernable signal. In some embodiments, the normalized data depicted in FIG. 3 may be used to estimate tumor fraction (TF) and a report may be generated. In some embodiments, the report may comprise the sequence data. In some embodiments, the report may comprise tumor fraction, a tumor score, a tumor fraction scoreor any combination thereof. In some embodiments, the report may comprise a bin count data, read data, histograms of counts per bin, patterns, statistics, copy number variant (CNV) data or combinations thereof. In some embodiments, a pattern (e.g., a copy number pattern (e.g., a tumor signature) may be extracted and outputted as depicted in the example of FIG. 3. In some embodiments, a pattern may be derived from a tumor derived sample and may be referred to as a tumor pattern, tumor profile, or tumor signature.
[0106] In various embodiments (e.g., as depicted in Fig. 4), the methods may use a sample derived from a normal tissue (e.g., non-diseased tissue). In some embodiments, normal (e.g., non-diseased) samples may be samples derived from patients without any known disease. In some embodiments, normal samples may be derived from tissues or samples known to be free of a disease (e.g., non-tumor samples, samples from patients without cancer). In some embodiments, normal samples may be samples from tissues that are not positive for a given disease (e.g., cancer) though may still have associations or mutations relevant to another disease (e.g., a non-cancer disease).
[0107] In some embodiments, normal samples may be derived from a subject (e.g., subject’s non-cancerous tissue). In some embodiments, normal samples may be derived from portions of a sample known to be disease free (such as a buffy coat portion of a blood sample). In some embodiments, normal samples may provide a pattern or baseline that is indicative of a subject’s genomic profile (e.g., somatic mutations, germline mutations, CNVs) outside of a tumor, or a normal portion of a cfDNA sample where tumor DNA (ctDNA) may be present. In some embodiments, a normal profile may be used to remove data associated with a normal sample (e.g., germline mutations, sequence data, normalized sequence data, or patterns (for example normal patterns), or normalized count data) from a cfDNA sample of a patient with or suspected of having a cancer.
[0108] In some embodiments, samples with known characteristics may be used to, for example, further correct sequence data. For example, samples derived other than from the subject (e.g., derived from donors) may be used. In some embodiments, data from these samples (e.g., donor samples) may be generated by the same sequencing equipment or a sequencing equipment of the same type as is used to generate the sample for which the method is being carried out. In some embodiments, donor data may be used to normalize data at various stages in some embodiments of a method (such as binned read count data, sequencing data, and / or count normalized data). In some embodiments, the normalization using donor data may be used to account for variability in the data introduced by the methods and equipment used to generate the data (such as a batch effect). In some embodiments,donor data is sequencing data (such as data comprising read sequences). In some embodiments, donor data is binned data (such as data comprising binned read counts). In some embodiments, donor data is used to produce a noise model corresponding to the donor(s). In some embodiments, a noise model may be a model of noise for a set of samples. In some embodiments, a noise model may be used to identify and / or remove noise from sequence data. In some embodiments, noise may be a distortion (or variation) in the data that may be introduced by methods, equipment, reagents, human error, machine errors, or any combination thereof. In some embodiments, data (such as sequence data) generated using the same or similar reagents, methods, equipment, or personnel may have similar or the same variations. In some embodiments, samples (such as donor samples) may have known characteristics (e.g., such as tumor fraction, noise signatures, noise, CNV levels, SNV signatures, or any combination thereof) and may be used to identify and remove variations from other data.
[0109] In some embodiments, a normal sample may be used. An example of the use of normal samples is depicted in FIG. 4 which depicts a method similar to that described in FIG. 3 but uses normal sample from the subject (and / or normal sample data) and / or donor sample (and / or donor data) to produce a pattern (e.g., a CNV pattern). In this figure, the donor data may be used during normalization to reduce noise in the count data (for example, noise attributable to the sequencer and / or methodology (such as batch effect)). In some embodiments, as such the donor data may come from similar equipment or the same equipment. Similarly, the donor data may be prepared using the methodology (such as laboratory steps or be performed by a same person) or similar methodology. This may be expected to produce noise in the donor data that is common to that methodology, machine(s), or lab. FIG. 4 further depicts some embodiments of a CHIP removal. As shown, FIG. 4 further depicts generating a normal pattern that can be used in conjunction with a tissue (e.g., tumor) derived pattern. For example, the pattern may be used for CHIP removal to generate a pattern without CHIP (“pattern. chip”). The example method depicted in FIG. 4, or portions thereof, may be applied to tumor derived sequence data.
[0110] In some embodiments, the method may be used to analyze samples that comprise cfDNA and / or ctDNA, which may be referred to as a cfDNA sample. In some embodiments, a portion of the cfDNA may be ctDNA which may be referred to as the ctDNA fraction and / or tumor fraction. In some embodiments, some cfDNA samples may comprise low tumor fractions (low TF). In some embodiments, in sequencing data derived from lower tumor fraction samples the detectable tumor signal / tumor signature / tumor pattern may bedifficult to discern from noise in the sequencing data. The noise may arise from a variety of sources. In some embodiments, one such source may be from the composition of the cfDNA arising from a heterogenous population of cells which may include tumor cells or any cells from the body. In some embodiments, in such cases the tumor fraction may be difficult to detect due to the large variety of other reads (which may comprise normal reads). In some embodiments, the tumor-derived reads may appear at such a low frequency that they are dismissed as sequencing noise. In some embodiments, sequencing noise may arise as a natural artifact of the sequencing and / or amplification process and errors propagated by the sequencing method are detected as valid reads. In some embodiments, when tumor fraction is very low it may fall around or below the sequencing error rate and be dismissed as noise. In some embodiments, a second type of error / noise may arise via a batch effect. In some embodiments, a batch effect noise which may be introduced due to variation in methodology such as small errors in settings (such as temperatures, volumes, flow rates etc.) in the machine or handling variation due to variations in equipment (such as pipettes) or personal habits (such as speed of pipetting). In some embodiments, batch effect may cause predictable variations that correlate with equipment types, batch numbers of reagents, personnel differences, environmental differences, etc. In some embodiments, batch effect may be present in addition to other forms of noise present in a sample. In some embodiments, batch effect and / or sequencing error may be dealt with using methods and systems provided herein. [OHl] In some embodiments, a tumor fraction (TF) may be low (e.g., expected and / or predicted to be low). In some embodiments, a method for determining a tumor fraction may be estimated as depicted in FIG. 5. In some embodiments, a method for estimating a tumor fraction may comprise a tumor informed method, where tumor derived data (e.g., a tumor pattern) may be used. In some embodiments, for example, such method reduces noise and / or improves ctDNA signal from a plasma sample. In some embodiments, as depicted in an example schematics, sequencing data derived from a plasma sample may be binned and then corrected using binned read counts of donor data (as denoted by the “wig” file format) and tumor pattern data (such as may be produced by methods similar to those depicted in FIG. 4 as applied to tumor data) is further used.
[0112] In some embodiments, as depicted in an example schematic provided in FIG.5, the donor data may be also applied (e.g., to the TF estimation as a model of noise). In some embodiments, the donor data may be used to correct the TF estimation using TF reports based on the donor data. In some embodiments, as depicted in an example, outputting may comprise outputting normalized count data, a partial report of tumor fraction, and a reportbased on the corrected tumor fraction, and / or a z-score. In some embodiments, a cfDNA may have high enough levels to be detected above noise. In some embodiments, donor data may be used to correct and / or normalize noise in the sample sequencing data. In some embodiments, donor data may be used to produce a bin noise model. In some embodiments, tumor fraction estimation may comprise a bin noise model. In some embodiments, donor data may be used to produce a report (such as a partial report) and / or a tumor fraction report. In some embodiments, the tumor fraction estimation may be corrected. In some embodiments, correction of the tumor fraction estimation may be based at least in part on a report or partial report produced based on donor data.
[0113] In some embodiments, the methods and systems provided herein may comprise normalization. In some embodiments, read counts may be normalized. In some embodiments, binned read counts may be normalized. In some embodiments, normalization may comprise the application of an outlier removal method. In some embodiments, the outlier removal method may comprise a clustering method such as DBScan, k-means clustering, isolation forest, local outlier factor (LOF), or any combination thereof. In some embodiments, outlier removal may comprise a machine learning method, such as log transformation, z-score normalization, decision tree, random forest, support vector machine (SVM), gradient-boosted trees (such as those implemented in the XGBOOST library), random forest, VAE, GAN, or any combination thereof. In some embodiments, normalization may comprise applying machine learning to clusters from a clustering method.
[0114] In some embodiments, the methods and systems provided herein may comprise neutral correction. In some embodiments, normalization may comprise neutral correction. Neutral correction may be used to account for or remove sources of variation and / or bias that are not related to biological differences between samples (such as the difference between a sample that comprises ctDNA and one that does not). In some embodiments, neutral correction may improve detectability of signals of interest in the sequencing data, allowing for more accurate performance of the methods and systems provided herein. In some embodiments, neutral correction may comprise correction for errors introduced due to sequencing depth. In some embodiments, neutral correction may comprise correction for errors introduced due to batch effect. In some embodiments, neutral correction may comprise correction for errors introduced due to GC content bias. In some embodiments, neutral correction may comprise correction for errors introduced due to fragmentation bias. In some embodiments, neutral correction may comprise correction for errors introduced due to road mapping bias. In some embodiments, neutral correction may comprise the use of donorsequencing data. In some embodiments, neutral correction may comprise the use of normal sequencing data. In some embodiments, neutral correction may comprise the use of tumor sequencing data. In some embodiments, neutral correction may comprise library size normalization. In some embodiments, neutral correction may comprise housekeeping gene normalization. In some embodiments, neutral correction may comprise normalization based on bias parameters (such as partitioning the genome into elements (e.g., bins or segments) based on bias parameters such as GC content and estimating bias within each element to correct for bias). In some embodiments, neutral correction may comprise mutual nearest neighbors (MNN) to correct for batch effect. In some embodiments, normalization may produce a normalized count histogram such as the example depicted in FIG. 10A. In some embodiments, the normalized count histogram may comprise the log-count data for various normalized counts. For example, the log count data depicted in FIG. 10A shows the count of binned reads with a copy number in the range of about 0.5 to 2.0. In some embodiments, normalization may comprise the use of donor binned count data. As depicted in FIG.12, normalization may be based on donor read count data (e.g. the wig file) where the donor data is used as additional input to the neutral model (such as a machine learning model), for example, during training to produce a model which has not learned biases (e.g., biases such as those present in the other training data such as GC bias, batch effect, etc.).
[0115] FIGs. 6 - 8 depict examples of some embodiments of normalization methods and workflows. FIG. 6 depicts an example workflow of a normalization method comprising inputting count data and removing outliers (e.g., using a clustering method or other machine learning methods (such as deep learning)), then processing the data using a neutral model method (such as a machine learning method). The output of the neutral model may be neutral corrected or have an additional neutral correction method applied resulting in a normalized count data.
[0116] FIG. 7 depicts example of some embodiments of the outlier removal (e.g., as could be performed in the workflow depicted in FIG. 6 and / or FIG. 7C (as indicated by the star). FIG. 7A depicts an example of a raw count data with the y-axis showing the read count and the x-axis showing position in the genome. FIG. 7B depicts an example of the same data clustered according to GC content and count. The outliers are shown as small dots; the main cluster is shown as the group of x’s, and subclusters are shown as the large dots. Outliers are identified on both x and y axis. FIG. 8A-8C depicts an example of some embodiments of data and the workflow (of, e.g., FIG. 6) at the neutral correction stage (such as is indicated by FIG. 8C by the star). FIG. 8A depicts example raw data and FIG. 8B depicts examplenormalized data after neutral correction showing the data points have far less variability and is around the natural line of 1 (CNV of 1, e.g., being the healthy or standard copy number).
[0117] Peak detection
[0118] In some embodiments, data (e.g., CNV data) in a cfDNA sample may show large amounts of variability even when normalized due to the heterogeneous mixture of DNA in the sample (DNA from multiple cell types, and potentially from cancer). This makes estimating CNV and / or tumor fraction difficult as even normalized data may show considerable variability. Therefore, in some embodiments, peak detection may be used to identify peaks (e.g., in CNV data to approximate CNV in the sample which may then be used to estimate the tumor fraction (and / or tumor fraction score) based on peak frequency.
[0119] In some embodiments, peaks of a histogram may be determined or identified. In some embodiments, determination of a tumor fraction of a biological sample may comprise detecting one or more local maxima (e.g., peaks) in a histogram and scoring the frequency of peaks (e.g., peaks frequency) such as shown in FIG. 9. In some embodiments, peak detection may be carried out on a histogram of normalized counts such as that depicted in FIG. 10A and indicated by a star in FIG. 10B. As shown in FIG. 10A, detected peaks are indicated by dots near under the peaks. This correlates to the normalized data in FIG. 8B which shows high concentrations of datapoints around, e.g., the 0.75, 1.0, 1.25, and 1.5 demarcations on the y-axis. As an example, the histogram of FIG. 10A is based on this data and the detected peaks are at these same values on the normalized count axis (x-axis). In some embodiments, detecting peaks transforms the read count data from continuous data to discrete data which allows the method to identify and approximate CNV for a sample where that sample may comprise multiple values for any given read. In some embodiments, peaks may be determined through thresholding. In some embodiments, thresholding may identify peaks by setting a fixed threshold and considering any data point exceeding that threshold as part of a peak. In some embodiments, peaks may be determined using a derivative. In some embodiments, a derivative may be a first derivative (e.g., the rate of change or slope). In some embodiments, peak detection may be based on identifying points where the first derivative changes sign (from positive to negative for a maximum). In some embodiments, peak detection may be based on a second derivative (e.g., amplification of the signal's curvature) which may be useful for detecting hidden or overlapping peaks. In some embodiments, a second derivative will be close to zero at peak locations.
[0120] In some embodiments, peak detection may be based on a local maximum / minimum, window search, residual after first derivative, Fourier self-deconvolution, quadratic fit, wavelet transform, dispersion by standard deviation, or any combination thereof.
[0121] In some embodiments, peaks may be identified using a moving average wherein the data (e.g., the histogram) may be smoothed using a moving average and peaks may then be identified using any of the methods described herein.
[0122] In some embodiments, applying spectral analysis may comprise determining peaks in a histogram.
[0123] Tumor Fraction estimation
[0124] In various embodiments, a tumor fraction may be estimated. In some embodiments, a tumor fraction may be determined based at least in part on one or more local maxima and / or the frequency of the one or more local maxima. In some embodiments, a tumor fraction may be estimated based on a measure of peak frequency. In some embodiments, peaks may be local maxima of data, such as sequence data. In some embodiments, sequence data may comprise one or more local maxima. In some embodiments, peak frequency (or a frequency of local maxima) may be a measure of an interval between peaks, such as the average interval between peaks. In some embodiments, peak frequency may be calculated along with a likelihood function of a fragment count for a genome position and / or a measure of an occurrence (e.g., frequency) for the fragment count (e.g., a count histogram) to arrive at the tumor fraction — for example, as indicated in the example workflow in FIG. 11B (such as indicated by the star) and discussed herein. In some embodiments, the calculation of the likelihood function may optimize for an estimation of tumor fraction corresponding to the distribution of the CNV in the read data (e.g., sequence data) when considering the mixture of normal and tumor DNA in a sample. In some embodiments, the methods and systems provided herein may comprise a determination of peak frequency. In some embodiments, peak frequency may be determined by the frequency of the peaks in the count histogram, an example of which is depicted in FIG. 2B. For example, in FIG. 10A the peaks (indicated by dots) occur at a frequency of approximately 0.2, FIG. 11A the peak frequency (indicated by the dot) is at approximately 0.2 (“2xl0_1”). In some embodiments, the likelihood of a tumor fraction may be indicated by the y-axis as, e.g., depicted on the plot in FIG. 11 A. FIG. 11A further depicts an example of the resulting tumor fraction (tf* or tumor score) above the plot which, in this example, is the maximum likelihood of the function at the peak frequency divided by 2 (for a diploid genome; in some embodiment this may be altered based on the estimated ploidy (e.g., of a sample, of a tumor,etc.)). In this example, it is 0.9 / 2. In some embodiments, a tumor fraction may be produced based on a bin noise model and / or a method noise model. For example, FIG. 13 depicts an example of some embodiments of a workflow of donor data being used to generate a bin noise model and a method noise model which are used in combination during tumor fraction estimation (TF estimation) before producing a report. In some embodiments, the tumor fraction may be compared to and / or corrected with the donor data (e.g., bias corrected). In some embodiments, a set of donors may produce a tumor fraction which may be used for such comparison and / or correction. In some embodiments, for example, a score may be generated (such as a z-score) comprising the tumor fraction estimation for subject minus the donors’ mean or median divided by a standard deviation of the donor data.
[0125] Reports and / or Partial reports
[0126] In some embodiments, the methods and systems provided herein may provide a report. In some embodiments, the methods and systems provided herein may provide a report comprising the copy number variation data. In some embodiments, the methods and systems herein may produce a partial report. In some embodiments, a report may comprise a tumor fraction estimation, a tumor score, a tumor fraction score, or any combination thereof. In some embodiments, a partial report may comprise at least one of a tumor fraction estimation, a tumor score, a tumor fraction score. In some embodiments, a partial report may be generated for donor data. In some embodiments, a report may be generated for donor data. In some embodiments, a partial report may be generated for tumor data In some embodiments, a report may be generated for tumor data. In some embodiments, a partial report may be generated for plasma data. In some embodiments, a report may be generated for plasma data. In some embodiments, a partial report may be generated for normal data. In some embodiments, a report may be generated for normal data.
[0127] Bin Noise Model
[0128] In some embodiments, the methods and systems provided herein may produce a bin noise model. In some embodiments, a bin noise model may be generated for donor data. In some embodiments, a bin noise model may be generated for tumor data. In some embodiments, a bin noise model may be generated for plasma data. In some embodiments, a bin noise model may be generated for normal data. In some embodiments, a bin noise model may be generated to model noise in a sample. In some embodiments, a bin noise model may be a machine learning model such as a linear regression model, a logistic regression model, a neural network, a VAE, a GAN, a diffusion model, a support vector machine (SVM), a decision tree, a gradient boosted tree, a random forest, a clustering method, or anycombination thereof. In some embodiments, a bin noise model may be a distribution. For example, a bin noise model may model the noise in a sample such as GC bias. A bin noise model may be used to correct regions of sequence data with high levels of repeats such as telomeres or centromeres.
[0129] Method noise model
[0130] In some embodiments, the methods and systems provided herein may produce a method noise model. In some embodiments, a method noise model may be generated for donor data. In some embodiments, a method noise model may be generated for tumor data. In some embodiments, a method noise model may be generated for plasma data. In some embodiments, a method noise model may be generated for normal data. In some embodiments, a method noise model may be generated to model noise in a method (such as a sequencing and / or amplification method). In some embodiments, a method noise model may be a machine learning model such as a linear regression model, a logistic regression model, a neural network, a VAE, a GAN, a diffusion model, a support vector machine (SVM), a decision tree, a gradient boosted tree, a random forest, a clustering method, or any combination thereof. For example, a method noise model may model the noise present due to batch effect.
[0131] Tumor Naive
[0132] In some embodiments, a tumor naive method may be implemented to detect tumor fraction without using tumor derived data. In some embodiments, in such cases, the mutations, variations or CNV of the tumor may not be known, or may be known or partially known from previous cfDNA samples. In some embodiments, tumor naive methods may utilize proxies or substitutes to perform the methods without a tumor sample. In some embodiments, for example, values derived from databases (e.g., publicly available sequence data of known cancer patients) may be used instead of data directly derived from a subject’s tumor. In some embodiments, if mutations, variations or CNV of the tumor are known or partially known the tumor informed method may be used with weighting on the tumor derived portions. In cases where there is no tumor data, methods similar to those already described herein may be used but without the tumor data.
[0133] In some embodiments, a sample is obtained or derived from the subject. A variety of methods can be used to obtain or derive the sample. In some embodiments, the sample is obtained by drawing blood from the subject (e.g., from a vein). In some embodiments, the sample obtained from the subject is further processed (e.g., purified, extracted, fractioned, derivatized, and / or otherwise altered) before it undergoes furthertesting. In some embodiments, the sample comprises the cell-free DNA comprising the ctDNA. In some embodiments, the sample is obtained or derived from a blood sample. In some embodiments, the sample is obtained or derived from a urine sample. In some embodiments, the sample is obtained or derived from a saliva sample. In some embodiments, the sample is obtained or derived from a blood plasma sample. In some embodiments, the sample is obtained or derived from a stool sample. In some embodiments, the sample is obtained or derived from a cerebrospinal fluid sample. In some embodiments, a sample is from a tumor (e.g., a tumor sample). In some embodiments, the samples are collected using a swab, for example, to swab cells or bodily fluid from a subject. In some embodiments, a sample is from normal tissue (e.g., a normal sample). In some embodiments, a sample has an unknown status (e.g., an unknown sample). In some embodiments, the sample is a diseased sample. In some embodiments, for example, a diseased sample may be from an individual suffering from a condition, disease, or disorder. In some embodiments, the diseased sample may comprise tissue associated with a disease (e.g., from a cancerous tissue (such as a tumor biopsy)). In some embodiments, the diseased sample may comprise nucleic acids (e.g., ctDNA, cfDNA) associated with a disease (e.g., from a cancerous tissue or cell). In some embodiments, a sample (e.g., an unknown sample) may be from a subject for whom the status of cancer is unknown, for example from a subject being screened for cancer where the presence of a tumor is currently unknown (e.g., the subject has not previously been screened for cancer), or a subject being monitored for residual disease after treatment (e.g., a subject that had cancer prior to a treatment (e.g., a successful treatment)). In some embodiments, the sample is a sample type other than a tumor sample (e.g., a non-tumor sample). In some embodiments, the sample is a diseased sample. In some embodiments, the sample does not comprise a solid tumor sample. In some embodiments, the sample does not comprise a tumor biopsy sample (e.g., a non-biopsy sample).
[0134] In some embodiments, samples comprising cell-free DNA are used. In some embodiments, cell-free DNA are fragments of DNA found circulating in a bloodstream. In some embodiments, fragments may differ in size; for example, they can be from 120 to 220 base pairs. In some embodiments, a portion of the cell-free DNA might be a ctDNA. In some embodiments, ctDNA may be a single- or double-stranded DNA released by the tumor cells into the blood. In some embodiments, ctDNA may comprise one or more mutations of an original tumor.
[0135] In some embodiments, a sample comprises gDNA. In some embodiments, gDNA may be present in and / or is extracted from cells in the sample (such as blood cells orother cells) that are (e.g., known to be) non-cancerous and / or not tumor cells, for example, nucleated blood cells. In some embodiments, gDNA is present in and / or is extracted from the sample (e.g., cells from the sample) suspected of being cancerous or known to be cancerous. For example, cells may be pelleted from a sample (e.g., urine). In some embodiments, gDNA is present in and / or is extracted from the sample (e.g., cells from the sample) that are suspected of being or known to be normal (non-cancerous). In some embodiments, gDNA may be present in and / or extracted from a buffy coat. In some embodiments, cfDNA is extracted from the sample. In some embodiments, cfDNA is extracted from the plasma of a blood sample. In some embodiments, the DNA is sequenced.
[0136] In various embodiments, a sample may be processed prior to being sequenced or assayed. For example, in some embodiments, a sample comprises cfDNA. The sample may be processed to remove (e.g., at least partially) sample components other than the cell- free DNA. In some embodiments, the sample is processed to isolate a cell-free DNA. In some embodiments the DNA is extracted from the sample. In some embodiments, the DNA is cfDNA. In some embodiments, the DNA is genomic DNA (gDNA). In some embodiments, the DNA is tumor DNA (e.g., DNA from a tumor biopsy).
[0137] In some embodiments, the methods and systems herein utilize an unknown sample, a tumor sample, a normal sample, a diseased sample, or any combination thereof. In some embodiments, the samples used in a method do not comprise a tumor sample (e.g., a non-tumor sample). In some embodiments, a diseased sample may be used. In some embodiments, the samples used in a method do not comprises a solid tumor sample. In some embodiments, the samples used in a method are other than a solid tumor sample (e.g., a nonsolid tumor sample). In some embodiments, the samples used in a method do not comprises comprise a tumor biopsy sample (e.g., a non-biopsy sample). In some embodiments, samples comprising tumor cells collected in areas distal from the tumor may be used (such as cells found in a urine sample). In some embodiments, samples comprising ctDNA may be used (e.g., as opposed to a sample comprising tumor cells); in some embodiments such samples may be collected in areas distal from the tumor. In some embodiments, DNA from samples may be used to generate a proxy or model of an expected tumor sample. For example, by processing a plasma sample that contains ctDNA and comparing the sequences with the germline or genomic DNA of a subject, sequences that are specific to a tumor can be identified and can be used in place of a tumor sample.In some embodiments, the sample is assayed via sequencing. In some embodiments, the method comprises assaying the sample to produce a sequence data. In some examples, thesequence data comprises sequence reads generated by assaying a sample. In some embodiments, the sequence data may be generated based on sequencing of nucleic acids (e.g., gDNA, cfDNA) in a sample. In some embodiments, the sequence data comprises DNA sequence reads. In some embodiments, the sequence data comprises RNA sequence reads. In some embodiments, the sequence data comprises sequence reads generated from normal and / or tumor DNA. In some embodiments, the sequence data comprises sequence reads generated from cell-free DNA. In some embodiments, the sequence data comprises sequence reads generated from ctDNA. In some embodiments, the sequence data comprises information on biomarkers of cancers. In some embodiments, the sequence data comprises a file (e.g., a digital file with one or more of aforementioned information).
[0138] In some embodiments, the nucleic acids are sequenced at a coverage of at least lOx, at least 20x, at least 30x, at least 40x, at least 50x, at least 60x, at least 70x, at least 80 x, at least 90x, at least lOOx, at least 150x, at least 200x, or more. In some embodiments, the sequencing provides sequence data. In some embodiments, the sequence data may comprise the sequence of one or more genomic loci. For example, the sequence data may comprise sequences corresponding to a CNV, SNV, insertion, deletion, or variants at one or more genomic loci.
[0139] In some embodiments, the sequence data may comprise a sequence file. In some embodiments, the sequence data may comprise a sequencing read. In some embodiments, the sequence data may comprise a plurality of sequencing reads. In some embodiments, the sequence data may comprise sets of sequencing reads. In some embodiments, the sequence data may comprise alignment data. In some embodiments, the sequence data may comprise genomic assemblies. In some embodiments, nucleotide frequencies may be calculated based at least in part on the sequence data. In some embodiments, nucleotide frequencies for samples from different sources may be calculated based at least in part on a sample from each of the sources (such as a tumor source and a normal / healthy source). In some embodiments, the sequence data may comprise a digital representation of a DNA sequence in the sample. In some embodiments, the sequence data may comprise a digital representation of a cfDNA sequence in the sample. In some embodiments, the digital representation of a DNA sequence may be a numerical representation of the DNA sequence. In some embodiments, the sequencing file may be a Binary Alignment Map (BAM) file, which may be the sequencing data format. In some embodiments, the sequencing file may be a Compressed Reference Alignment Map (CRAM) file. In some embodiments, the sequencing file may be a Sequence alignment Map (SAM)file. In some embodiments, the sequencing file may be an unmapped BAM (uBAM) file. In some embodiments, the sequencing file may be a FASTQ file. In some embodiments, the sequencing file may be a HDF5 file. In some embodiments, the sequence file may be a BedGraph file. In some embodiments, the sequencing file may be a BED file. In some embodiments, the sequencing file contains alignment information which may be used to assemble a genomic sequence. In some embodiments, the genomic data can support indels.
[0140] Spectral Methods
[0141] In some embodiments, a spectral analysis method is applied. In some embodiments, a spectral analysis method may refer to a class of approaches (e.g., a class of machine learning (ML) approaches) for extracting useful information from massive, noisy, and / or incomplete datasets. In some embodiments, a spectral method may refer to methods of optimizing for one or more variables in a machine learning setting or outside a machine learning setting. In some embodiments, spectral methods refer to a collection of algorithms built upon eigenvalues and eigenvectors of some properly designed matrices constructed from data. In some embodiments, a diverse array of applications may be found in machine learning, data science, and signal processing. In some embodiments, due to their simplicity and effectiveness, spectral methods may be used not only as a stand-alone estimator, but also to initialize other more sophisticated algorithms to improve performance.
[0142] In some embodiments, a spectral analysis method may be applied to at least a portion of the sequence data. In some embodiments, spectral analysis method may be applied to at least a portion of the sequence data to determine or estimate a tumor fraction (e.g., ctDNA fraction of cfDNA). In some embodiments, the spectral analysis method is applied to at least the portion of the sequence data comprising the measure of the occurrence for the fragment count versus the measure of the fragment count for the genome position. In some embodiments, a spectral method may be applied to measures corresponding to a number of sequencing reads or fragments at one or more genome positions. In some embodiments, the spectral analysis method may comprise detecting one or more local maxima (e.g., peaks). For example, in some embodiments, sequencing data may be generated such that the data comprises measures corresponding a coverage, copy number, or read count of one more genomic regions. In some embodiments, the spectral analysis method may be applied to the data corresponding to coverage, copy number variation, or read count, of one more genomic regions such that the coverages, copy numbers variation, or read counts that are more prevalent (e.g., those that are at one local maximums) can be determined. In some embodiments, for example the detection of these one or local maxima comprises usingspectral methods to solve for an optimization of a parameter. For example, the spectral analysis may comprise identifying a maximum value within a function (or range of a function), plot, or histogram, and may, based at least on the maximum value or optimization, determine the location of one or more local maxima (e.g., peaks). In some embodiments, a tumor fraction may be estimated based at least in part on a copy number variation (CNV) data. In some embodiments, spectral analysis may be applied to CNV data such as count data and / or normalized count data to determine or estimate tumor fraction. In some embodiments, the ctDNA fraction estimates achieved using the present methods may be further augmented using copy-number variant (CNV) data; tissue feature data; cancer type data; and the like. In some embodiments, the ctDNA fraction estimates may be further augmented using copynumber variant (CNV) data; tissue feature data; cancer type data; and the like.
[0143] In some embodiments, a spectral analysis method may be applied to at least a portion of the sequence data comprising a measure of the occurrence (e.g., frequency) for the fragment count versus a measure of the fragment count for the genome position (for example FIG. 2B and FIG. 10A). In some embodiments, the spectral analysis method comprises detecting one or more local maxima in at least a portion of the sequence data (for example, such as the dots in FIG.10A). In some embodiments, the one or more local maxima may also be referred to as peaks. In some embodiments, the detecting the one or more local maxima comprises detecting one or more local maxima in at least the potion of the sequence data comprising the measure of the occurrence for the fragment count versus the measure of the fragment count for the genome position. In some embodiments, the spectral analysis method comprises determining a frequency of the one or more local maxima (e.g., peaks frequency or the frequency of peaks). For example, in some embodiments, one or more local maxima are determined (e.g., using spectral analysis or non-spectral analysis) corresponding to coverages, copy number variations, or read counts, that are more prevalent in sample. The spectral analysis method may identify a frequency of the one or more local maxima. For example, in some embodiments, the spectral analysis may comprise a Fourier transform, or other algorithmic methodologies, for example, to transform, decompose, or deconvolve functions, plots, or histograms. For example, in some embodiments, transforming, decomposing, or deconvolving data may allow for the detecting of underlying frequencies of the data (e.g., a frequency of one or more local maxima). In some embodiments, the tumor fraction is determined based at least in part on the one or more local maxima and / or the frequency of the one or more local maxima. In some embodiments, the spectral analysis comprises determining a likelihood of a particular tumor fraction corresponding to the biological samplebased at least in part on the frequency of the one or more local maxima. In some embodiments, determining the tumor fraction comprises determining the particular tumor fraction with the highest (or maximum) likelihood (such as shown in FIG. 11 A). For example, the spectral analysis may comprise solving for an optimization of a parameter or determining a maximum value for a function, plot, or histogram. For example, a function, plot, or histogram that describes a likelihood of a particular peak frequency or tumor fraction may be analyzed via the spectral analysis. In some embodiments, a likelihood of a peak frequency or a tumor fraction may correspond to how likely a given peak frequency or tumor fraction would result in a particular set of CNV data. The optimization or determination of a maximum value may correspond to a specific tumor fraction that corresponds to a sample (or subject).
[0144] Machine Learning
[0145] In some embodiments, machine learning (ML) may refer to a class of computational methods that leverage data to build analytical models and improve performance in future computational tasks. For example, “training data” may be introduced to the ML system from which the system learns about patterns in the data.
[0146] In some embodiments, the rate of cancer recurrence varies widely according to many factors, especially cancer type and subtype — for example, glioblastomas seem to recur in about 100 percent of cases, whereas estimated risk of childhood acute lymphoblastic leukemia is 15-20 percent. Accordingly, pairing the ML methods with other clinical information may tend to improve overall predictions and treatment / prevention decisions.
[0147] In some embodiments, the frequency of ctDNA may be based upon biomarkers (e.g., CNVs, SNVs) . In some embodiments, biomarkers may comprise signals or features found in a sample or processed from signal or features from a sample. In some embodiments, machine learning may comprise noise suppression, signal processing (such as signal integration), or both. In some embodiments, a machine learning algorithm may comprise at least one of a neutral learning model, or gradient boosted decision tree.
[0148] In some embodiments, the sets of biomarkers may be processed using an algorithm. In some embodiments, the algorithm may be a trained algorithm. In some embodiments, the trained algorithms may use inputs (e.g., including the sets of biomarkers as an input) and generate an output regarding the presence or absence of a cancer. In some embodiments, the output may be specific to a type of cancer or subtype of cancer.
[0149] In some embodiments, the trained algorithm may be trained on multiple samples. For example, the trained algorithm may be trained using at least 2, 3, 4, 5, 6, 7, 8, 9,10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 300, 400, 500 , 600 ,700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or more independent training samples. In some embodiments, the trained algorithm may be trained using no more than 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 300, 400, 500 , 600 ,700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or less, independent training samples. In some embodiments, the training samples may be associated with a presence or an absence of the cancer. In some embodiments, the training samples may be associated with a relapse of cancer. In some embodiments, the training samples may be associated with cancer that is resistant to a particular drug or treatment. In some embodiments, an individual training sample may be positive for a particular cancer. In some embodiments, an individual training sample may be negative for a particular cancer. In some embodiments, by using training samples, the trained algorithm may be able to detect a cancer, determine a probability of recurrence or relapse of a cancer, or determine if a cancer comprises a set of biomarkers may be resistant to a treatment. In some embodiments, the training sample may be associated with additional clinical health data of a subject. For example, additional clinical health data may comprise the gender, weight, height, or levels of metabolites or antibodies in a subjects. In some embodiments, additional clinical health data may comprise indication of other diseases, disorders, or diseases conditions.
[0150] In some embodiments, the trained algorithms may be trained using multiple sets of training samples. In some embodiments, the sets may comprise training samples as described elsewhere herein. For example, the training may be performed using a first set of independent training samples associated with a presence of the cancer and a second set of independent training samples associated with an absence of the cancer. Similarly, a first set may be associated with relapse and a second sample may be associated with the absence of relapse.
[0151] In some embodiments, the trained algorithm may also process additional clinical health data of the subject. For example, additional clinical health data may comprise the gender, weight, height, or levels of metabolites or antibodies in a subjects. In some embodiments, additional clinical health data may comprise indication of other diseases, disorders, or diseases conditions that the subject may suffer from. In some embodiments, byusing the additional clinical health data, in conjunction with the biomarkers, the trained algorithm may output a presence or absences of cancer, probability of relapse, or resistance to drug treatment, that may be different from the output of an algorithm that does not process additional clinical health.
[0152] In some embodiments, the trained algorithm may be an unsupervised machine learning algorithm. For example, the unsupervised machine learning algorithm may utilize cluster analysis to identify attributes of interest. In some embodiments, the trained algorithm may be a supervised machine learning algorithm. For example, the algorithm may be inputted with training data such to generate an expected or desired output. In some embodiments, the supervised learning algorithm may comprise a deep learning algorithm, a support vector machine (SVM), a neural network, or a Random Forest. In some embodiments, via the machine learning algorithm, the trained algorithm may be able to identify relationships of biomarkers to a particular cancer prognosis or diagnosis. In some embodiments, without the trained algorithm, it may otherwise be difficult to identify relationships of the biomarkers to accurately identify the presence of a cancer or other parameters associated with the cancer.
[0153] In some embodiments, in various aspects, the systems and methods may comprise an accuracy, sensitivity, or specificity of detection of the cancer or a parameter of the cancer. For example, the methods or systems may comprise detecting the presence or the absence of cancer (or the presence of a parameter of the cancer, such as recurrence, relapse, or drug resistance) in the subject at an accuracy of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. In some embodiments, the methods or systems may comprise detecting the presence or the absence of cancer (or the presence of a parameter of the cancer, such as recurrence, relapse, or drug resistance) in the subject at a sensitivity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. In some embodiments, the methods or systems may comprise detecting the presence or the absence of cancer (or the presence of a parameter of the cancer, such as recurrence, relapse, or drug resistance) in the subject at a specificity of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. In some embodiments, the methods or systems may comprise detecting the presence or the absence of cancer (or the presence of a parameter of the cancer, such as recurrence, relapse, or drug resistance) in the subject at a positive predictive value of at least about 60%, at least about 70%, at least about 75%, at leastabout 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. In some embodiments, the methods or systems may comprise detecting the presence or the absence of cancer (or the presence of a parameter of the cancer, such as recurrence, relapse, or drug resistance) in the subject at a negative predictive value of at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.
[0154] Computer systems
[0155] The present disclosure provides computer systems that are programmed to implement methods of the disclosure. FIG. 14 shows a computer system 1401 that is programmed or otherwise configured to perform analysis or operations of the methods, for example determine a likelihood of the presence of a cancer based on a set of biomarkers of an individual or run an algorithm. The computer system 1401 can regulate various aspects of methods and systems of the present disclosure, such as, for example, perform an algorithm, input training data, analyze sets of biomarker, or output a result for the user as to the presence or absence of cancer. The computer system 1401 can be an electronic device of a user or a computer system that is remotely located with respect to the electronic device. The electronic device can be a mobile electronic device.
[0156] The computer system 1401 includes a central processing unit (CPU, also “processor” and “computer processor” herein) 1405, which can be a single core or multi core processor, or a plurality of processors for parallel processing. The computer system 1401 also includes memory or memory location 1410 (e.g., random-access memory, read-only memory, flash memory), electronic storage unit 1415 (e.g., hard disk), communication interface 1420 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 1425, such as cache, other memory, data storage and / or electronic display adapters. The memory 1410, storage unit 1415, interface 1420 and peripheral devices 1425 are in communication with the CPU 1405 through a communication bus (solid lines), such as a motherboard. The storage unit 1415 can be a data storage unit (or data repository) for storing data. The computer system 1401 can be operatively coupled to a computer network (“network”) 1430 with the aid of the communication interface 1420. The network 1430 can be the Internet, an internet and / or extranet, or an intranet and / or extranet that is in communication with the Internet. The network 1430 in some cases is a telecommunication and / or data network. The network 1430 can include one or more computer servers, which can enable distributed computing, such as cloud computing. The network 1430, in some caseswith the aid of the computer system 1401, can implement a peer-to-peer network, which may enable devices coupled to the computer system 1401 to behave as a client or a server.
[0157] The CPU 1405 can execute a sequence of machine-readable instructions, which can be embodied in a program or software. The instructions may be stored in a memory location, such as the memory 1410. The instructions can be directed to the CPU 1405, which can subsequently program or otherwise configure the CPU 1405 to implement methods of the present disclosure. Examples of operations performed by the CPU 1405 can include fetch, decode, execute, and writeback.
[0158] The CPU 1405 can be part of a circuit, such as an integrated circuit. One or more other components of the system 1401 can be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).
[0159] The storage unit 1415 can store files, such as drivers, libraries and saved programs. The storage unit 1415 can store user data, e.g., user preferences and user programs. The computer system 1401 in some cases can include one or more additional data storage units that are external to the computer system 1401, such as located on a remote server that is in communication with the computer system 1401 through an intranet or the Internet.
[0160] The computer system 1401 can communicate with one or more remote computer systems through the network 1430. For instance, the computer system 1401 can communicate with a remote computer system of a user (e.g., a medical professional or patient). Examples of remote computer systems include personal computers (e.g., portable PC), slate or tablet PC’s (e.g., Apple® iPad, Samsung® Galaxy Tab), telephones, Smart phones (e.g., Apple® iPhone, Android-enabled device, Blackberry®), or personal digital assistants. The user can access the computer system 1401 via the network 1430.
[0161] Methods as described herein can be implemented by way of machine (e.g., computer processor) executable code stored on an electronic storage location of the computer system 1401, such as, for example, on the memory 1410 or electronic storage unit 1415. The machine executable or machine readable code can be provided in the form of software.During use, the code can be executed by the processor 1405. In some cases, the code can be retrieved from the storage unit 1415 and stored on the memory 1410 for ready access by the processor 605. In some situations, the electronic storage unit 1415 can be precluded, and machine-executable instructions are stored on memory 1410.
[0162] The code can be pre-compiled and configured for use with a machine having a processer adapted to execute the code, or can be compiled during runtime. The code can besupplied in a programming language that can be selected to enable the code to execute in a pre-compiled or as-compiled fashion.
[0163] Aspects of the systems and methods provided herein, such as the computer system 1401, can be embodied in programming. Various aspects of the technology may be thought of as “products” or “articles of manufacture” typically in the form of machine (or processor) executable code and / or associated data that is carried on or embodied in a type of machine readable medium. Machine-executable code can be stored on an electronic storage unit, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. “Storage” type media can include any or all of the tangible memory of the computers, processors or the like, or associated modules thereof, such as various semiconductor memories, tape drives, disk drives and the like, which may provide non- transitory storage at any time for the software programming. All or portions of the software may at times be communicated through the Internet or various other telecommunication networks. Such communications, for example, may enable loading of the software from one computer or processor into another, for example, from a management server or host computer into the computer platform of an application server. Thus, another type of media that may bear the software elements includes optical, electrical and electromagnetic waves, such as used across physical interfaces between local devices, through wired and optical landline networks and over various air-links. The physical elements that carry such waves, such as wired or wireless links, optical links or the like, also may be considered as media bearing the software. As used herein, unless restricted to non-transitory, tangible “storage” media, terms such as computer or machine “readable medium” refer to any medium that participates in providing instructions to a processor for execution.
[0164] Hence, a machine readable medium, such as computer-executable code, may take many forms, including but not limited to, a tangible storage medium, a carrier wave medium or physical transmission medium. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices in any computer(s) or the like, such as may be used to implement the databases, etc. shown in the drawings. Volatile storage media include dynamic memory, such as main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media therefore include for example: a floppy disk, aflexible disk, hard disk, magnetic tape, any other magnetic medium, a CD-ROM, DVD or DVD-ROM, any other optical medium, punch cards paper tape, any other physical storage medium with patterns of holes, a RAM, a ROM, a PROM and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave transporting data or instructions, cables or links transporting such a carrier wave, or any other medium from which a computer may read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0165] The computer system 1401 can include or be in communication with an electronic display 1435 that comprises a user interface (UI) 1440 for providing, for example, an input of biomarkers or sequencing data, or an visual output relating to a detection, diagnosis, or prognosis. Examples of UI’s include, without limitation, a graphical user interface (GUI) and web-based user interface.
[0166] Methods and systems of the present disclosure can be implemented by way of one or more algorithms. An algorithm can be implemented by way of software upon execution by the central processing unit 1405. The algorithm can, for example, determine a presence or absence of a cancer or cancer parameter based on a set of input sequencing data from a sample derived from a subject.EmbodimentsEmbodiment 1. A method of estimating a tumor fraction of a subject, comprising:(a) providing a biological sample obtained or derived from the subject;(b) assaying the biological sample to determine a tumor signature; and(c) applying a spectral analysis method in conjunction with a machine learning algorithm to the tumor signature to determine the tumor fraction of the biological sample.Embodiment 2. The method of embodiment 1, wherein the subject has or is suspected of having cancer.Embodiment 3. The method of embodiment 2, wherein the cancer comprises fibrosarcoma, myosarcoma, liposarcoma, chondrosarcoma, osteogenic sarcoma, chordoma, angiosarcoma, endotheliosarcoma, lymphangiosarcoma, lymphangio endotheliosarcoma, synovioma, mesothelioma, Ewing’s tumor, leiomyosarcoma, rhabdomyosarcoma, colon carcinoma, pancreatic cancer, bladder cancer, lung cancer, breast cancer, ovarian cancer, prostate cancer, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinomas,cystadenocarcinoma, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatoma, bile duct carcinoma, choriocarcinoma, seminoma, embryonal carcinoma, Wilms tumor, cervical cancer, uterine cancer, testicular cancer, lung carcinoma, small cell lung carcinoma, bladder carcinoma, kidney cancer, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pinealoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, schwannoma, meningioma, melanoma, neuroblastoma, retinoblastoma, colorectal cancer, endometrial cancer, endometroid cancer, clear cell carcinoma, or any combination thereof.Embodiment 4. The method of embodiment 3, wherein the cancer comprises breast cancer.Embodiment 5. The method of embodiment 3, wherein the cancer comprises lung carcinoma.Embodiment 6. The method of embodiment 3, wherein the cancer comprises bladder carcinoma.Embodiment 7. The method of embodiment 3, wherein the cancer comprises kidney cancer.Embodiment 8. The method of embodiment 1, wherein the subject has or is suspected of having cancer remission.Embodiment 9. The method of embodiment 1, wherein the subject has or is suspected of having cancer recurrence.Embodiment 10. The method of embodiment 1, wherein the biological sample is selected from the group consisting of: a blood sample, a plasma sample, a serum sample, a urine sample, a saliva sample, a cell sample, and a tissue sample.Embodiment 11. The method of embodiment 10, wherein the biological sample is the plasma sample.Embodiment 12. The method of embodiment 1, wherein the biological sample is a set of biological samples, wherein (b) comprises assaying the set of biological samples to determine a set of tumor signatures, and wherein (c) comprises determining the tumor fraction of the set of biological samples.Embodiment 13. The method of embodiment 1, wherein (b) further comprises assaying nucleic acids obtained or derived from the biological sample to generate a set of nucleic acid sequences.Embodiment 14. The method of embodiment 13, wherein the assaying comprises a technique selected from the group consisting of: nucleic acid sequencing, nucleic acidamplification, nucleic acid enrichment, microarray, reverse transcription, and a combination thereof.Embodiment 15. The method of embodiment 13, wherein the nucleic acids comprise cell-free deoxyribonucleic acid (cfDNA).Embodiment 16. The method of embodiment 13, further comprising determining a fraction of putative single-nucleotide-variant (SNV) reads from the set of nucleic acid sequences that are read errors, and / or a fraction of putative SNV reads from the set of nucleic acid sequences that are true mutations.Embodiment 17. The method of embodiment 13, further comprising detecting and removing a set of outliers from the set of nucleic acid sequences.Embodiment 18. The method of embodiment 13, further comprising applying a normalization to the set of nucleic acid sequences, thereby generating normalized sequencing data.Embodiment 19. The method of embodiment 18, wherein the normalization comprises a neutral correction function.Embodiment 20. The method of embodiment 19, wherein the neutral correction function is expressed by:Normalized count(i) = count(i) / fi(a) wherein f is a normalization factor and count(i) is a feature count.Embodiment 21. The method of embodiment 18, further comprising sorting the normalized sequencing data into a set of genomic bins, and determining quantitative measures of the normalized sequencing data in each of the set of genomic bins to produce bin-wise count data.Embodiment 22. The method of embodiment 21, further comprising applying a transformation to the bin-wise count data to be within a given range.Embodiment 23. The method of embodiment 21, further comprising producing the binwise count data according to the expression: count(i) / count_0(i) = (l+TF / 2) * CopyNumber(i)(a) wherein “count(i)” is the normalized bin-wise count for the biological sample at a given position in the normalized sequencing data,(b) wherein “count_0(i)” is a normalized bin-wise count for the biological sample at a position that is earlier than the given position in the normalized sequencing data, and(c) wherein CopyNumber(i) is a function indicative of a measure of copy number variation for the biological sample.Embodiment 24. The method of embodiment 23, further comprising determining a histogram of counts per genomic location in the normalized sequencing data.Embodiment 25. The method of embodiment 24, further comprising determining a measure of peak frequency of the histogram.Embodiment 26. The method of embodiment 25, wherein the tumor fraction is determined based at least in part on the measure of peak frequency.Embodiment 27. The method of embodiment 1, wherein the machine learning algorithm uses noise suppression, signal integration, or both.Embodiment 28. The method of embodiment 1, further comprising determining a presence or an absence of minimal residual disease (MRD) in the subject, based at least in part on the tumor fraction determined in (c).Embodiment 29. The method of embodiment 28, further comprising administering a therapy to the subject, thereby treating the MRD.Embodiment 30. The method of embodiment 29, wherein the therapy comprises chemotherapy, radiotherapy, immunotherapy, targeted therapy, surgical resection, laser ablation, or any combination thereof.Embodiment 31. The method of embodiment 29, further comprising administering the therapy to the subject based at least in part on a likelihood of recurrence of cancer.Embodiment 32. The method of embodiment 1, further comprising determining providing a copy-number variant (CNV) report.Embodiment 33. The method of embodiment 1, further comprising applying a clustering algorithm to produce clustering data.Embodiment 34. The method of embodiment 33, wherein the clustering algorithm comprises DBSCAN.Embodiment 35. The method of embodiment 1, wherein the machine learning algorithm comprises at least one of a neutral learning model, gradient boosted decision tree, and xgboost.Embodiment 36. A method of analyzing copy number variation (CNV) of a subject, comprising:(a) providing a biological sample obtained or derived from the subject;(b) assaying the biological sample to determine a tumor signature; and(c) applying a spectral analysis method in conjunction with a machine learning algorithm to the tumor signature to determine a measure of CNV of the biological sample.Embodiment 37. The method of embodiment 36, wherein the subject has or is suspected of having cancer.Embodiment 38. The method of embodiment 37, wherein the cancer comprises fibrosarcoma, myosarcoma, liposarcoma, chondrosarcoma, osteogenic sarcoma, chordoma, angiosarcoma, endotheliosarcoma, lymphangiosarcoma, lymphangio endotheliosarcoma, synovioma, mesothelioma, Ewing’s tumor, leiomyosarcoma, rhabdomyosarcoma, colon carcinoma, pancreatic cancer, breast cancer, ovarian cancer, prostate cancer, bladder cancer, lung cancer, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinomas, cystadenocarcinoma, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatoma, bile duct carcinoma, choriocarcinoma, seminoma, embryonal carcinoma, Wilms tumor, cervical cancer, uterine cancer, testicular cancer, lung carcinoma, small cell lung carcinoma, bladder carcinoma, kidney cancer, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pinealoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, schwannoma, meningioma, melanoma, neuroblastoma, retinoblastoma, colorectal cancer, endometrial cancer, endometroid cancer, clear cell carcinoma, or any combination thereof.Embodiment 39. The method of embodiment 38, wherein the cancer comprises breast cancer.Embodiment 40. The method of embodiment 38, wherein the cancer comprises lung cancer.Embodiment 41. The method of embodiment 38, wherein the cancer comprises bladder cancer.Embodiment 42. The method of embodiment 38, wherein the cancer comprises kidney cancer.Embodiment 43. The method of embodiment 36, wherein the subject has or is suspected of having cancer remission.Embodiment 44. The method of embodiment 36, wherein the subject has or is suspected of having cancer recurrence.Embodiment 45. The method of embodiment 36, wherein the biological sample is selected from the group consisting of: a blood sample, a plasma sample, a serum sample, a urine sample, a saliva sample, a cell sample, and a tissue sample.Embodiment 46. The method of embodiment 45, wherein the biological sample is the plasma sample.Embodiment 47. The method of embodiment 36, wherein the biological sample is a set of biological samples, wherein (b) comprises assaying the set of biological samples to determine a set of tumor signatures, and wherein (c) comprises determining the measure of CNV of the set of biological samples.Embodiment 48. The method of embodiment 36, wherein (b) further comprises assaying nucleic acids obtained or derived from the biological sample to generate a set of nucleic acid sequences.Embodiment 49. The method of embodiment 48, wherein the assaying comprises a technique selected from the group consisting of: nucleic acid sequencing, nucleic acid amplification, nucleic acid enrichment, microarray, reverse transcription, and a combination thereof.Embodiment 50. The method of embodiment 48, wherein the nucleic acids comprise cell-free deoxyribonucleic acid (cfDNA).Embodiment 51. The method of embodiment 48, further comprising determining a fraction of putative single-nucleotide-variant (SNV) reads from the set of nucleic acid sequences that are read errors, and / or a fraction of putative SNV reads from the set of nucleic acid sequences that are true mutations.Embodiment 52. The method of embodiment 48, further comprising detecting and removing a set of outliers from the set of nucleic acid sequences.Embodiment 53. The method of embodiment 48, further comprising applying a normalization to the set of nucleic acid sequences, thereby generating normalized sequencing data.Embodiment 54. The method of embodiment 36, wherein the machine learning algorithm uses noise suppression, signal integration, or both.Embodiment 55. The method of embodiment 36, further comprising determining a presence or an absence of minimal residual disease (MRD) in the subject, based at least in part on the measure of CNV determined in (c).Embodiment 56. The method of embodiment 55, further comprising administering a therapy to the subject, thereby treating the MRD.Embodiment 57. The method of embodiment 56, wherein the therapy comprises chemotherapy, radiotherapy, immunotherapy, targeted therapy, surgical resection, laser ablation, or any combination thereof.Embodiment 58. The method of embodiment 56, further comprising administering the therapy to the subject based at least in part on a likelihood of recurrence of cancer.Embodiment 59. The method of embodiment 36, further comprising applying a clustering algorithm to produce clustering data.Embodiment 60. The method of embodiment 59, wherein the clustering algorithm comprises DBSCAN.Embodiment 61. The method of embodiment 36, wherein the machine learning algorithm comprises at least one of a neutral learning model, gradient boosted decision tree, and xgboost.EXAMPLES
[0167] Example 1: High tumor fraction (TF) estimation
[0168] Using methods and systems of some of the embodiments of the present disclosure, a spectral method is applied to the detection and quantification of TF by the following example computational operations:
[0169] Count normalization is performed.
[0170] Outliers removal is performed, including filtering for normalization only bins from one copy -number value (largest).
[0171] Modeling is performed, including outliers removal in the bin count and GC- content spaces.
[0172] Clustering is performed, including using DBSCAN and choosing a largest cluster.
[0173] Neutral model learning is performed, including normalizing for positiondependent neutral count change.
[0174] Modeling is performed using the following formula: count(i) = f(gc content(i), mappability (i), donor- 1 count(i), . . . donor-k count(i)).
[0175] Implementation is performed, including learning a function f using decisiontree gradient boosting (xgboost).
[0176] Neutral correction is performed, using the following formula: Normalized count(i) = count(i) / fi
[0177] Tumor fraction (TF) estimation is performed, including finding the tumor fraction that best fits the formula: count(i) / count_0(i) = (l+TF / 2) * CopyNumber(i).
[0178] Count histogram peaks are determined, and peaks frequency scoring is performed.
[0179] The following observation is made: (coverage(i) / coverage_0(i) ) % TF = K (constant).
[0180] An optimization problem is analyzed:
[0182] Where co represents the eigenvalue being optimized, x is the peak, i is the imaginary number ( ( — 1))
[0183] Histogram peaks are determined (e.g., as shown in Figs. 10A-10B).
[0184] Peak frequency is scored (e.g., as shown in Figs. 11 A-l IB).
[0185] Example 2: Low tumor fraction (TF) estimation
[0186] Using methods and systems of some of the embodiments of the present disclosure, a spectral method is applied to the detection and quantification of ctDNA by the following e computational operations:
[0187] A value of tumor fraction (TF) is determined that best fits the formula count(i) / count_0(i) = (l+TF / 2) * CopyNumber(i).
[0188] An assumption is made that most of the copy-number variant (CNV) pattern is preserved, and that copy -number in tissue equals the copy-number in plasma.
[0189] An optimization problem is analyzed:TF = where w, are chosen to balance the weight° of each chromosome, since in most of the samples copy-number segment is part of a chromosome or full number of chromosomes.
[0190] Panel of Normal (PON) bins correction methods are performed as follows:
[0191] Bias removal is performed, using the following formula: TFt= TFt— gPl0N
[0192] In the case of no tumor prior, the following formula is used: TFt=
Claims
CLAIMSWHAT IS CLAIMED IS:
1. A method of estimating a tumor fraction of a subject, comprising:(a) providing a biological sample obtained or derived from the subject;(b) assaying the biological sample to produce a sequence data; and(c) applying a spectral analysis method to at least a portion of the sequence data to determine the tumor fraction of the biological sample.
2. The method of claim 1, wherein the sequence data comprises a measure of a fragment count for a genome position and / or a measure of an occurrence for the fragment count.
3. The method of claim 2, wherein the spectral analysis method is applied to the at least the portion of the sequence data comprising the measure of the occurrence for the fragment count versus the measure of the fragment count for the genome position.
4. The method of any one of the preceding claims, wherein applying the spectral analysis method comprises detecting one or more local maxima in the at least the portion of the sequence data.
5. The method of claim 4, wherein the detecting the one or more local maxima comprises detecting one or more local maxima in the at least the portion of the sequence data comprising the measure of the occurrence for the fragment count versus the measure of the fragment count for the genome position.
6. The method of claim 4 or 5, wherein applying the spectral analysis method comprises determining a frequency of the one or more local maxima.
7. The method of any one of claims 4 to 6, wherein the tumor fraction is determined based at least in part on the one or more local maxima and / or the frequency of the one or more local maxima.
8. The method of claim 7, wherein the spectral analysis method comprises determining a likelihood of a particular tumor fraction corresponding to the biological sample based at least in part on the frequency of the one or more local maxima.
9. The method of claim 7 or 8, wherein determining the tumor fraction comprises determining the particular tumor fraction with the maximum likelihood.
10. The method of any one of the preceding claims, wherein the biological sample comprises a cell-free DNA (cfDNA), and wherein assaying the biological sample comprises assaying the cfDNA.
11. The method of any one of claims 2 to 10, wherein the genome position is a bin.
12. The method of any one of the preceding claims, further comprising detecting and / or removing a set of outliers from the sequence data.
13. The method of claim 12, wherein detecting and / or removing the set of outliers comprises applying a clustering algorithm.
14. The method of claim 13, wherein the clustering algorithm comprises DBSCAN.
15. The method of any one claims 2 to 14, further comprising determining a pattern comprising one or more copy numbers corresponding to one or more of the genome positions.
16. The method of any one of the preceding claims, wherein the subject has or is suspected of having cancer.
17. The method of claim 16, wherein the cancer comprises fibrosarcoma, myosarcoma, liposarcoma, chondrosarcoma, osteogenic sarcoma, chordoma, angiosarcoma, endotheliosarcoma, lymphangiosarcoma, lymphangio endotheliosarcoma, synovioma, mesothelioma, Ewing’s tumor, leiomyosarcoma, rhabdomyosarcoma, colon carcinoma, pancreatic cancer, breast cancer, ovarian cancer, prostate cancer, bladder cancer, lung cancer, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinomas, cystadenocarcinoma, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatoma, bile duct carcinoma, choriocarcinoma, seminoma, embryonal carcinoma, Wilms tumor, cervical cancer, uterine cancer, testicular cancer, lung carcinoma, small cell lung carcinoma, bladder carcinoma, kidney cancer, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pinealoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, schwannoma, meningioma, melanoma, neuroblastoma, retinoblastoma, colorectal cancer, endometrial cancer, endometroid cancer, clear cell carcinoma, or any combination thereof.
18. The method of claim 16 or 17, wherein the cancer comprises breast cancer.
19. The method of claim 16 or 17, wherein the cancer comprises lung cancer.
20. The method of claim 16 or 17, wherein the cancer comprises bladder cancer.
21. The method of claim 16 or 17, wherein the cancer comprises kidney cancer.
22. The method of any one of the preceding claims, wherein the subject has or is suspected of having cancer remission.
23. The method of any one of the preceding claims, wherein the subject has or is suspected of having cancer recurrence.
24. The method of any one of the preceding claims, wherein the biological sample is selected from the group consisting of: a blood sample, a plasma sample, a serum sample, a urine sample, a saliva sample, a cell sample, and a tissue sample.
25. The method of claim 24, wherein the biological sample is the plasma sample.
26. The method of any one of the preceding claims, wherein the assaying comprises a technique selected from the group consisting of: nucleic acid sequencing, nucleic acid amplification, nucleic acid enrichment, microarray, reverse transcription, and a combination thereof.
27. The method of any one of the preceding claims, wherein the sequence data is normalized.
28. The method of claim 27, wherein the sequence data is normalized by at least applying a neutral correction function.
29. The method of any one of the preceding claims, wherein the method further comprises noise removal comprising using a sample with known characteristics.
30. The method of any one of the preceding claims, wherein the tumor fraction is determined at least in part by using a machine learning algorithm.
31. The method of claim 30, wherein the machine learning algorithm comprises noise suppression, signal processing, or both.
32. The method of claim 30 or 31, wherein the machine learning algorithm comprises at least one of a neutral learning model, or gradient boosted decision tree.
33. The method of any one of the preceding claims, further comprising administering a therapy to the subject at least in part based on the tumor fraction determined in (c).
34. The method of any one of the preceding claims, further comprising determining a presence or an absence of a minimal residual disease (MRD) in the subject, based at least in part on the tumor fraction determined in (c).
35. The method of claim 34, further comprising administering a therapy to the subject, thereby treating the MRD.
36. The method of any one of the preceding claims, further comprising determining a likelihood of recurrence of cancer, based at least in part on the tumor fraction determined in (c).-SO-37. The method of claim 36, further comprising administering a therapy to the subject, based at least in part on the determined likelihood of recurrence of cancer.
38. The method of claims 33 to 37, wherein the therapy comprises chemotherapy, radiotherapy, immunotherapy, targeted therapy, surgical resection, laser ablation, or any combination thereof.
39. The method of any one of the preceding claims, wherein the sequence data comprises a copy number variation (CNV) data.
40. The method of any one of the preceding claims, further comprising providing a report comprising the copy number variation data.
41. The method of claims 39 or 40, wherein the sequence data further comprises single-nucleotide-variant (SNV) data, and wherein the SNV data is corrected based at least in part on the tumor fraction.
Citation Information
Patent Citations
Machine measurement metrology frame for a lithography system
WO2021257055A1