Validation of a bioinformatic model for classifying non-tumor variants in cell-free DNA liquid biopsy assays
A bioinformatics model improves the sensitivity and specificity of liquid biopsy tests by differentiating tumor and non-tumor nucleic acid variants in cfDNA using plasma-based datasets, addressing the challenges of clonal hematopoietic variants and reducing the need for WBC sequencing.
Patent Information
- Application Number
- JP2025528486
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-17
- Filing Date
- 2023-11-16
- Publication Date
- 2025-12-09
AI Technical Summary
Existing liquid biopsy tests face challenges in distinguishing circulating tumor DNA (ctDNA) from other cell-free DNA (cfDNA) due to the presence of clonal hematopoietic variants and biological noise, necessitating costly and complex workflows for genotyping white blood cells to remove non-tumor variants.
A bioinformatics model is developed to improve sensitivity and specificity in identifying non-tumor variants using cfDNA, employing a computer-based method to generate datasets and classifiers that differentiate tumor and non-tumor nucleic acid variants by analyzing ratios and prevalence of tumor-associated gene variants in plasma and white blood cells.
The model enhances the accuracy of cancer detection assays by effectively distinguishing tumor and non-tumor nucleic acid variants, reducing the need for costly WBC sequencing and improving biomarker evaluation.
Smart Images

Figure 2025539779000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 384,215, filed November 17, 2022, and relies on the filing date of that provisional patent application, the entire disclosure of which is incorporated herein by reference. [Background technology]
[0002] background Liquid biopsy tests can be used to profile circulating tumor nucleic acids in blood samples from patients, for example, to detect cancer at an early stage, select treatment, and monitor disease progression and / or minimal residual disease. Circulating plasma cell-free tumor DNA (ctDNA) is a small DNA fragment from apoptotic and necrotic tumor cells or circulating tumor cells (CTCs) introduced into the bloodstream. While ctDNA is only a fraction of cell-free DNA (cfDNA) specifically released from cancer cells, the majority of cfDNA in a given sample typically originates from normal, noncancerous cells, including normal white blood cells, hematopoietic stem cells (HSCs), or other early blood cell progenitors that undergo apoptosis or necrosis during the clonal hematopoietic process. One problem associated with many liquid biopsy tests is distinguishing ctDNA from other cfDNA in patient samples. Furthermore, the presence of clonal hematopoietic (CH) variants due to aging and treatment, as well as biological noise, can confound biomarker interpretation.
[0003] Currently, comprehensive methods for removing non-tumor variants require genotyping the white blood cell (WBC) fraction of paired plasma samples, which is a costly and complex workflow. Therefore, there remains a need for methods and related aspects for distinguishing between tumor and non-tumor origin nucleic acid variants detected in cell-free DNA (cfDNA) samples, with particular promise for achieving a plasma-only bioinformatic solution for identifying non-tumor variants for accurate biomarker evaluation in cell-free DNA (cfDNA).
[0004] We describe a bioinformatics model that demonstrates improved sensitivity for identifying non-tumor variants over WBC sequencing at low VAF (<0.6%). In a paired plasma and WBC late-stage cancer cohort, the majority of non-tumor variants were in known clonal hematopoietic genes and variants of unknown significance. The described analytical platform demonstrates high sensitivity and specificity for WBC sequencing to identify tumor and non-tumor variants using only cfDNA. Summary of the Invention
[0005] Summary of the Invention The present disclosure provides, inter alia, methods for distinguishing between tumor and non-tumor origin nucleic acid variants in cell-free nucleic acid (cfNA) samples, which improve the sensitivity and specificity of cancer detection assays and guide treatment strategies. Additional methods and related systems and computer-readable media are also provided.
[0006] In some embodiments, the present disclosure provides a method for distinguishing (e.g., differentiating) tumor- and non-tumor-origin nucleic acid variants in a cell-free nucleic acid (cfNA) sample obtained from a test subject, at least in part using a computer. The method includes generating or providing at least one tumor variant dataset by a computer, the tumor variant dataset including a population of reference tumor-associated gene variants. The tumor variant dataset includes observed frequency data for one or more tumor-associated gene variants in the population of reference tumor-associated gene variants among reference samples, the reference samples including a reference body fluid sample (e.g., a plasma sample, a serum sample, etc.) containing only plasma and / or a reference non-body fluid sample (e.g., a cell sample, a tissue sample, etc.) containing white blood cells. The reference samples may be obtained from a single reference subject and / or different reference subjects with the same cancer type. The method also includes determining by a computer one or more ratios of the observed frequency data for one or more tumor-associated gene variants among the reference samples in the population of reference tumor-associated gene variants to generate at least one MAF variance and / or relative prevalence dataset. Furthermore, the method also includes generating or providing at least one set of probabilities of non-tumor origin from the relative prevalence dataset by a computer, and using the set of probabilities of non-tumor origin to distinguish nucleic acid variants detected in a cfNA sample obtained from the test subject as being tumor-origin nucleic acid variants or non-tumor-origin nucleic acid variants. In some embodiments of the methods, systems, computer-readable media, and other aspects of the present disclosure, one or more other features are optionally utilized in addition to or instead of the ratios of observed frequency data. Some of these other features include, for example, uniformity of prevalence across cancer types, longitudinal mutation allele fraction (MAF) variation over time, rates in hematological cancers, and / or the like.
[0007] In another aspect, the present disclosure provides a method for distinguishing between tumor- and non-tumor-origin nucleic acid variants in a cell-free nucleic acid (cfNA) sample obtained from a test subject, using at least in part a computer. The method includes determining the relative prevalence of one or more tumor-associated gene variants observed only in one or more reference plasmas compared to one or more reference white blood cells to generate at least one relative prevalence dataset by a computer. Furthermore, the method also includes generating or providing at least one set of probabilities of non-tumor origin from the relative prevalence dataset by a computer, and using the set of probabilities of non-tumor origin to distinguish nucleic acid variants detected in the cfNA sample obtained from the test subject as tumor-origin nucleic acid variants or non-tumor-origin nucleic acid variants.
[0008] In some embodiments, the present disclosure provides a method for distinguishing between tumor- and non-tumor-origin nucleic acid variants in a cell-free nucleic acid (cfNA) sample obtained from a test subject, using at least a computer. The method includes determining, by a computer, the variance in mutant allele fraction (MAF) values and / or at least one statistic associated therewith (e.g., the mean, standard deviation, and / or chi-squared p-value of the variant MAF over time) for each of one or more tumor-associated and / or non-tumor-associated genetic variants observed only in one or more reference plasmas compared to one or more reference white blood cells for at least two different time points to generate at least one MAF variance and / or relative prevalence dataset. Furthermore, the method also includes generating or providing, by a computer, at least one set of probabilities of non-tumor origin from the MAF variance and / or relative prevalence dataset, and using the set of probabilities of non-tumor origin to distinguish nucleic acid variants detected in the cfNA sample obtained from the test subject as tumor-origin nucleic acid variants or non-tumor-origin nucleic acid variants.
[0009] In some embodiments, the present disclosure provides a method for distinguishing between tumor and non-tumor-origin nucleic acid variants in cell-free nucleic acid (cfNA) samples obtained from test subject, at least partially using a computer.The method includes: when the prevalence of the first nucleic acid variant detected in the cfNA sample is less than a threshold probability from a set of probabilities of non-tumor origin, by a computer, classifying at least a first nucleic acid variant detected in the cfNA sample obtained from test subject as tumor-origin nucleic acid variant; and when the prevalence of the second nucleic acid variant detected in the cfNA sample is greater than a threshold probability from a set of probabilities of non-tumor origin, by a computer, classifying at least a second nucleic acid variant detected in the cfNA sample obtained from test subject as non-tumor-origin nucleic acid variant, thereby distinguishing between tumor and non-tumor-origin nucleic acid variants in the cfNA sample obtained from test subject. The set of probabilities of non-tumor origin is generated by: generating or providing, by a computer, at least one tumor variant dataset comprising a population of reference tumor-associated gene variants, wherein the tumor variant dataset comprises observed frequency data between reference samples comprising reference plasma only and reference leukocytes for one or more tumor-associated gene variants in the population of reference tumor-associated gene variants, and the reference samples are obtained from a single reference subject and / or different reference subjects having the same cancer type; determining, by a computer, one or more ratios of the observed frequency data between the reference samples for one or more tumor-associated gene variants in the population of reference tumor-associated gene variants to generate at least one relative prevalence dataset; and generating, by a computer, a set of probabilities of non-tumor origin from the relative prevalence dataset.
[0010] In another aspect, the present disclosure provides a method for generating, at least in part, a classifier for distinguishing nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as being tumor-originated or non-tumor-originated, using at least one computer. The method includes generating or providing, by a computer, at least one tumor variant dataset comprising a population of reference tumor-associated gene variants, wherein the tumor variant dataset comprises observed frequency data among reference samples comprising reference plasma only and / or reference leukocytes for one or more tumor-associated gene variants in the population of reference tumor-associated gene variants, the reference samples being obtained from a single reference subject and / or different reference subjects having the same cancer type. The method also includes determining, by a computer, one or more ratios of the observed frequency data among the reference samples for one or more tumor-associated gene variants in the population of reference tumor-associated gene variants to generate at least one relative prevalence dataset. Additionally, the method also includes applying, by a computer, at least one machine learning model to the relative prevalence dataset to generate at least one set of probabilities of non-tumor origin, thereby generating a classifier that distinguishes nucleic acid variants detected in the cfNA sample as being of tumor origin or non-tumor origin.
[0011] In some embodiments, the present disclosure provides a method for distinguishing between tumor- and non-tumor-origin nucleic acid variants in a cell-free nucleic acid (cfNA) sample obtained from a test subject with a cancer type, using at least in part a computer. The method includes determining, by a computer, the prevalence of one or more genetic variants observed in the cfNA sample to generate a test subject prevalence dataset. The method includes comparing, by a computer, the prevalence of the one or more genetic variants in the test subject prevalence dataset with the prevalence of the genetic variant observed in a reference cfNA sample obtained from a reference subject with the cancer type. The method further includes classifying, by a computer, a given genetic variant in the test subject prevalence dataset as a non-tumor-origin nucleic acid variant if the prevalence of the given genetic variant in the test subject prevalence dataset is less than a predetermined threshold associated with the given genetic variant in the reference cfNA sample obtained from the reference subject with the cancer type.
[0012] In some aspects, the present disclosure provides a method for distinguishing between tumor- and non-tumor-origin nucleic acid variants in a cell-free nucleic acid (cfNA) sample obtained from a test subject, at least in part, using a computer. The method includes determining, by a computer, the prevalence of one or more genetic variants observed in the cfNA sample to generate a test subject prevalence dataset. The method also includes comparing, by a computer, the prevalence of the one or more genetic variants in the test subject prevalence dataset with the prevalence of the genetic variant observed in a reference cfNA sample obtained from a reference subject with leukemia, lymphoma, and / or hematological malignancy. Furthermore, the method also includes classifying, by a computer, a given genetic variant in the test subject prevalence dataset as a non-tumor-origin nucleic acid variant if the prevalence of the given genetic variant in the test subject prevalence dataset exceeds a predetermined threshold associated with the given genetic variant in the reference cfNA sample obtained from the reference subject with leukemia, lymphoma, and / or hematological malignancy.
[0013] In some embodiments, the methods disclosed herein include identifying genetic variants present in a cfNA sample from sequencing reads derived from cfNA molecules in the cfNA sample. In certain of these embodiments, the sequencing reads are obtained from targeted segments of cfNA molecules in the cfNA sample. In some embodiments, the reference tumor-associated gene variant population is obtained from a reference sample. In certain embodiments, the reference white blood cells include a reference tumor tissue sample and / or a reference white blood cell sample. In some embodiments, the methods disclosed herein include obtaining a cfNA sample from a test subject. In certain embodiments, the reference sample comprises at least about 25, at least about 50, at least about 100, at least about 200, at least about 300, at least about 400, at least about 500, at least about 600, at least about 700, at least about 800, at least about 900, at least about 1,000, at least about 5,000, at least about 10,000, at least about 15,000, at least about 20,000, at least about 25,000, at least about 30,000, or more bodily fluids and / or leukocytes. In some embodiments, the cfNA sample comprises cell-free deoxyribonucleic acid (cfDNA). In certain embodiments, the cfNA sample comprises cell-free ribonucleic acid (cfRNA). In some embodiments, the test subject is a mammalian subject. In certain embodiments, the test subject is a human subject. In some embodiments, the reference bodily fluid sample comprises a plasma sample. In certain embodiments, the reference bodily fluid sample comprises a serum sample. In some embodiments, the reference non-body fluid sample is a non-plasma sample. In some embodiments, the reference non-body fluid (e.g., non-plasma) sample comprises a cell sample. In certain embodiments, the reference non-body fluid (e.g., non-plasma) sample comprises a tissue sample.
[0014] In some embodiments, the methods disclosed herein include selecting one or more therapies for treating the cancer type if one or more tumor-originating nucleic acid variants associated with the cancer type are detected in a cfNA sample obtained from the test subject. In certain embodiments, the methods disclosed herein include administering to the test subject one or more therapies for treating the cancer type if one or more tumor-originating nucleic acid variants associated with the cancer type are detected in a cfNA sample obtained from the test subject.
[0015] In some embodiments, the cancer type is biliary tract cancer, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, glioma, astrocytoma, breast cancer, dysplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal cancer, colon cancer, hereditary nonpolyposis colorectal cancer, colorectal adenocarcinoma, gastrointestinal stromal tumor (GIST), endometrial cancer, endometrial stromal sarcoma, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, ocular melanoma, uveal melanoma, gallbladder cancer, gallbladder adenocarcinoma, renal cell carcinoma, clear cell renal cell carcinoma, transitional cell carcinoma, urothelial carcinoma, Wilms' tumor, leukemia, acute lymphocytic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic (CLL), chronic myelogenous (CML), chronic myelomonocytic (CMML), liver cancer, liver carcinoma carcinoma), hepatocellular carcinoma, hepatocellular carcinoma, cholangiocarcinoma, hepatoblastoma, lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B-cell lymphoma, non-Hodgkin's lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, T-cell lymphoma, non-Hodgkin's lymphoma, precursor T-lymphoblastic lymphoma / leukemia, peripheral T-cell lymphoma, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oropharyngeal carcinoma, oral squamous cell carcinoma, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic ductal adenocarcinoma, pseudopapillary neoplasm, acinar cell carcinoma, prostate cancer, prostate adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine cancer, stomach cancer, gastric carcinoma, gastrointestinal stromal tumor (GIST), uterine cancer, and uterine sarcoma. In certain embodiments, the reference tumor-associated genetic variant is selected from the group consisting of a single nucleotide variant (SNV), an insertion or deletion (indel), a copy number variant (CNV), a fusion, a transversion, a translocation, a frameshift, a duplication, a repeat expansion, and an epigenetic variant.
[0016] In some embodiments, the methods disclosed herein include randomly dividing a tumor variant dataset into a training dataset and a test dataset. In certain embodiments, the training dataset comprises about 80% of the tumor variant dataset, and the test dataset comprises about 20% of the tumor variant dataset. In some embodiments, the tumor variant dataset comprises observation frequency data among reference samples of a given cancer type for one or more tumor-associated gene variants in a population of reference tumor-associated gene variants. In some embodiments, the methods disclosed herein include training a machine learning model using at least a portion of the population of tumor-associated gene variants to generate a trained machine learning model, and tumor-originating and non-tumor-originating nucleic acid variants detected in a cfNA sample obtained from the test subject are distinguished from each other using the trained machine learning model. In some of these embodiments, the machine learning model is trained using one or more of logistic regression, probit regression, decision tree, random forest, gradient boosting, support vector machine, K-nearest neighbor, and neural network. In some embodiments, the methods disclosed herein include using a probability threshold of at least about the 30th percentile for a given genetic variant as a classification cutoff. In some embodiments, the methods disclosed herein comprise performing a logistic regression on at least one of the ratios to obtain a given probability of non-tumor origin.
[0017] In some embodiments, the tumor variant dataset comprises mutant allele fraction data observed among reference samples for one or more tumor-associated gene variants in a population of reference tumor-associated gene variants. In some embodiments, the methods disclosed herein comprise normalizing the tumor variant dataset using one or more data normalization techniques. In certain of these embodiments, the data normalization techniques comprise min-max normalization and / or z-score normalization. In certain embodiments, a ratio of the observed frequency data of a given genetic variant only in the reference plasma to the observed frequency data of the given genetic variant in the reference leukocytes greater than 1 (1.0) indicates a high probability that the given genetic variant is a nucleic acid variant of non-tumor origin. In certain embodiments, the reference leukocytes comprise a reference tumor tissue sample, the set of probabilities of non-tumor origin comprises at least one set of probabilities of clonal hematopoietic origin.
[0018] In some embodiments, the tumor variant dataset comprises mutant allele fraction data observed among reference samples for one or more tumor-associated gene variants in a population of reference tumor-associated gene variants. In some embodiments, the methods disclosed herein comprise normalizing the tumor variant dataset using one or more data normalization techniques. In certain of these embodiments, the data normalization techniques comprise min-max normalization and / or z-score normalization. In certain embodiments, a ratio of observed frequency data of a given genetic variant only in the reference plasma to observed frequency data of the given genetic variant in reference leukocytes that is less than 1 (1.0) indicates a high probability that the given genetic variant is a nucleic acid variant of non-tumor origin. In certain embodiments, where the reference leukocytes comprise a reference leukocyte sample, the set of probabilities of non-tumor origin comprises at least one set of probabilities of clonal hematopoietic origin.
[0019] In another aspect, the present disclosure provides a system comprising a controller that includes or can access a computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least the following: (a) generating or providing at least one tumor variant dataset comprising a population of reference tumor-associated gene variants, wherein the tumor variant dataset comprises observed frequency data between reference samples comprising reference plasma only and / or reference leukocytes for one or more tumor-associated gene variants in the population of reference tumor-associated gene variants, wherein the reference samples are obtained from a single reference subject and / or different reference subjects having the same cancer type; (b) determining one or more ratios of the observed frequency data between the reference samples for the one or more tumor-associated gene variants in the population of reference tumor-associated gene variants to generate at least one relative prevalence dataset; and (c) applying at least one machine learning model to the relative prevalence dataset to generate at least one set of probabilities of non-tumor origin to create a classifier that distinguishes nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as being tumor-originated nucleic acid variants or non-tumor-originated nucleic acid variants.
[0020] In another aspect, the present disclosure provides a system including a controller that includes or can access a computer-readable medium containing non-transitory computer-executable instructions that, when executed by at least one electronic processor, at least perform the following: (a) determine the relative prevalence of one or more tumor-associated gene variants observed only in one or more reference plasmas compared to one or more reference white blood cells to generate at least one relative prevalence dataset; and (b) generate at least one set of probabilities of non-tumor origin from the relative prevalence dataset to generate a classifier that distinguishes nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as being of tumor origin or non-tumor origin.
[0021] In another aspect, the present disclosure provides a system including a controller that includes or can access a computer-readable medium containing non-transitory computer-executable instructions that, when executed by at least one electronic processor, at least perform the following: (a) determine the variation in mutant allele fraction (MAF) values and / or at least one statistic associated therewith for at least two different time points for each of one or more tumor-associated gene variants observed only in one or more reference plasmas compared to one or more reference leukocytes to generate at least one MAF variance and / or relative prevalence dataset; and (b) generate at least one set of probabilities of non-tumor origin from the MAF variance and / or relative prevalence dataset to generate a classifier that distinguishes nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as being tumor-origin nucleic acid variants or non-tumor-origin nucleic acid variants.
[0022] In some embodiments, the systems disclosed herein include a nucleic acid sequencer operably connected to a controller, the nucleic acid sequencer configured to provide sequencing reads derived from cfNA molecules in the cfNA sample. In certain of these embodiments, the nucleic acid sequencer or another system component is configured to group sequence reads generated by the nucleic acid sequencer into families of sequence reads, each family containing sequence reads generated from a given cfNA molecule in the cfNA sample. In certain embodiments, the systems disclosed herein include a database operably connected to the controller, the database indexed to tumor-origin nucleic acid variants and containing one or more treatments. In some embodiments, the systems disclosed herein include a sample preparation component operably connected to the controller, the sample preparation component configured to prepare cfNA molecules in the cfNA sample to be sequenced by the nucleic acid sequencer. In certain embodiments, the systems disclosed herein include a nucleic acid amplification component operably connected to the controller, the nucleic acid amplification component configured to amplify at least targeted segments of cfNA molecules in the cfNA sample. In certain embodiments, the systems disclosed herein include a material transfer component operably connected to the controller, the material transfer component configured to transfer one or more materials between at least the nucleic acid sequencer and the sample preparation component.
[0023] In some aspects, the present disclosure provides a computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least: (a) generating or providing at least one tumor variant dataset comprising a population of reference tumor-associated gene variants, wherein the tumor variant dataset comprises observed frequency data among reference samples comprising reference plasma only and / or reference leukocytes for one or more tumor-associated gene variants in the population of reference tumor-associated gene variants, wherein the reference samples are obtained from a single reference subject and / or different reference subjects having the same cancer type; (b) determining one or more ratios of the observed frequency data among the reference samples for the one or more tumor-associated gene variants in the population of reference tumor-associated gene variants to generate at least one relative prevalence dataset; and (c) applying at least one machine learning model to the relative prevalence dataset to generate at least one set of probabilities of non-tumor origin to create a classifier that distinguishes nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as being tumor-originated nucleic acid variants or non-tumor-originated nucleic acid variants. In certain embodiments of the methods, systems, computer-readable media, and other aspects of the present disclosure, one or more other features are optionally utilized in addition to or in place of the observed frequency data ratios, including, for example, uniformity of prevalence across cancer types, longitudinal mutant allele fraction (MAF) variation over time, proportions in hematologic cancers, mutant gene name, location, cancer type, chromosomal location, and / or the like.
[0024] In another aspect, the present disclosure provides a computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least: (a) determine the relative prevalence of one or more tumor-associated gene variants observed only in one or more reference plasmas compared to one or more reference white blood cells to generate at least one relative prevalence dataset; and (b) generate at least one set of probabilities of non-tumor origin from the relative prevalence dataset to generate a classifier that distinguishes nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as being of tumor origin or non-tumor origin.
[0025] In another aspect, the present disclosure provides a computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, at least perform the following: (a) determine the variation in mutant allele fraction (MAF) values and / or at least one statistic associated therewith for at least two different time points for each of one or more tumor-associated gene variants observed only in one or more reference plasmas compared to one or more reference leukocytes to generate at least one MAF variance and / or relative prevalence dataset; and (b) generate at least one set of probabilities of non-tumor origin from the MAF variance and / or relative prevalence dataset to create a classifier that distinguishes nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as being of tumor origin or non-tumor origin.
[0026] In some embodiments of the system or computer-readable medium disclosed herein, the electronic processor further performs at least the step of dividing the tumor variant dataset into a training dataset and a test dataset (e.g., randomly or non-randomly). In certain embodiments of the system or computer-readable medium disclosed herein, the electronic processor further performs at least: training a machine learning model using at least a portion of the population of tumor-associated gene variants to generate a trained machine learning model; and using the trained machine learning model to distinguish nucleic acid variants detected in the cfNA sample as being tumor-origin nucleic acid variants or non-tumor-origin nucleic acid variants. In some embodiments of the system or computer-readable medium disclosed herein, the electronic processor further performs at least: performing a logistic regression on at least one of the ratios to obtain a given probability of non-tumor origin. In certain embodiments of the system or computer-readable medium disclosed herein, the electronic processor further performs at least: normalizing the tumor variant dataset using one or more data normalization techniques. In certain embodiments of the systems or computer-readable media disclosed herein, the electronic processor further performs, at least, selecting one or more therapies for treating the cancer type if one or more tumor-originating nucleic acid variants associated with the cancer type are detected in the cfNA sample.
[0027] In certain embodiments, the methods, systems, or computer-readable media disclosed herein distinguish between tumor and non-tumor origin nucleic acid variants in a cell-free nucleic acid (cfNA) sample based at least in part on (i) the uniformity of the prevalence of the nucleic acid variant across cancer types; (ii) the variation in the mutant allele fraction (MAF) of the nucleic acid variant over time; and / or (iii) the prevalence of the nucleic acid variant in hematological cancers, such as leukemia, lymphoma, and / or hematological malignancies.
[0028] In some embodiments, the results of the systems and methods disclosed herein are used as input to generate a report. The report may be in paper or electronic format. For example, the classification of nucleic acid variants detected in a cell-free nucleic acid sample as being of tumor or non-tumor origin, as determined by the methods and systems disclosed herein, can be directly displayed in such a report. In some embodiments, only nucleic acid variants classified as being of tumor origin are displayed in such a report.
[0029] The various steps of the methods disclosed herein, or steps performed by the systems disclosed herein, may be performed at the same or different times, in the same or different geographic locations, e.g., countries, and / or by the same or different people.
[0030] In other aspects, a subject may be administered a therapy based on a determination by the methods and systems disclosed herein that the variant is of tumor or non-tumor origin. In certain embodiments, administration of treatment to a subject may be discontinued based on a determination by the methods and systems disclosed herein that the variant is of tumor or non-tumor origin. BRIEF DESCRIPTION OF THE DRAWINGS [Brief explanation of the drawings]
[0031] [Figure 1] FIG. 1 is a flow chart outlining exemplary method steps for distinguishing between tumor and non-tumor origin nucleic acid variants according to some embodiments.
[0032] [Figure 2] FIG. 2 is a flow chart outlining exemplary method steps for distinguishing between tumor and non-tumor origin nucleic acid variants according to some embodiments.
[0033] [Figure 3]FIG. 3 is a flow chart outlining exemplary method steps for distinguishing between tumor and non-tumor origin nucleic acid variants according to some embodiments.
[0034] [Figure 4] FIG. 4 is an exemplary block diagram for creating a predictive model.
[0035] [Figure 5] FIG. 5 is a flow chart illustrating an exemplary training method.
[0036] [Figure 6] FIG. 6 is a diagram of an exemplary process flow for using a machine learning-based classifier.
[0037] [Figure 7] FIG. 7 is a schematic diagram of an exemplary system suitable for use in certain embodiments.
[0038] [Figure 8] Figure 8 utilizes an internal database of over 250,000 clinical patients. Model design included features engineered from internal and external public datasets and trained using 10-fold cross-validation with multiple models. Only the results of the logistic regression model are shown. Model validation was performed on an independent cohort of paired plasma and WBC late-stage samples sequenced with an epigenomic panel, as well as healthy donors sequenced with a genomic panel.
[0039] [Figure 9]Model performance demonstrated high ROC AUC and accuracy for predicted calls. Prediction of tumor and non-tumor status was compared to WBC confirmation for A) 713 somatic SNVs / indels from 72 paired plasma and WBC GuardantInfinity™ samples, and B) 243 somatic SNVs / indels from 76 paired plasma and healthy donors in GuardantOMNI™. The reduced confirmation rate in WBC sequencing observed for low VAF variants (<0.6%) may be due to the detection limits of WBC variant calls and / or possible non-WBC lineage origins.
[0040] [Figure 10] Assay-specific engineered features among the most important with feature importance, and examples include: A) Top 10 features ranked by relative importance in the validation dataset. Individual gene names were included with one-hot coding. B) Highly ranked engineered features include clonality (left), defined as VAF / tumor fraction measured by methylation or maximum somatic VAF (center), VAF variation across time points (right), and C) homogeneity of variant prevalence across solid tumor cancer types in the plasma database.
[0041] [Figure 11] Concordance in non-tumor prediction: correlation with gene prevalence and age. The number of variants within each gene predicted as non-tumor or tumor-derived that were confirmed by WBC or model prediction in the late validation cohort in Figure 2. For genes with the most commonly confirmed cfDNA variants in WBC samples, variant counts are shown along with counts in clinically actionable genes (BRCA1, BRAF, KRAS, ESR1, ATM, CHEK2). The most frequent WBC-confirmed genes are consistent with previous reports, including a high prevalence of clonal hematopoiesis in ATM and CHEK2 (*).
[0042] [Figure 12]Correlation between non-tumor calls and age. Percentage of variants detected in WBC or predicted as non-tumor in the later validation cohort by age range. As expected from the literature, variants predicted or confirmed as non-tumor are highly correlated with age. DETAILED DESCRIPTION OF THE INVENTION
[0043] definition In order to more readily understand this disclosure, certain terms are first defined below. Additional definitions of these terms and other terms may be found throughout the specification. In the event that a definition of a term set forth below conflicts with a definition in a patent application or issued patent incorporated by reference, the definition set forth in this application should be used to understand the meaning of the term.
[0044] As used herein and in the appended claims, the words "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to a "method" includes one or more methods and / or steps of the type described herein and / or that will become apparent to those of skill in the art upon reading this disclosure. It is also understood that there is an implicit "about" before temperatures, concentrations, times, numbers of bases or base pairs, coverage, etc. discussed in this disclosure, so that slight and insubstantial equivalents are within the scope of this disclosure. In this application, the use of the singular includes the plural unless otherwise stated. Additionally, the use of "comprise," "comprises," "comprising," "contain," "contains," "containing," "include," "includes," and "including" is not intended to be limiting.
[0045] It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting. Furthermore, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. In describing and claiming the methods, computer-readable media, and systems, the following terms, and grammatical variations thereof, will be used in accordance with the definitions set forth below.
[0046] About: As used herein, "about" or "approximately," as applied to one or more values or elements of interest, refers to a value or element similar to the stated reference value or element. In certain embodiments, the term "about" or "approximately" refers to a range of values or elements that falls within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1% or less in either direction (greater or less than) the stated reference value or element, unless otherwise stated or apparent from the context (except where such number exceeds 100% of possible values or elements).
[0047] Adapter: As used herein, "adapter" refers to a short nucleic acid (e.g., less than about 500 nucleotides, less than about 100 nucleotides, or less than about 50 nucleotides in length) that is typically at least partially double-stranded and used to ligate to either or both ends of a given sample nucleic acid molecule. The adapter may contain nucleic acid primer binding sites at both ends that allow amplification of the nucleic acid molecule flanked by the adapter, and / or sequencing primer binding sites, including primer binding sites for sequencing applications, such as various next-generation sequencing (NGS) applications. The adapter may also contain a binding site for a capture probe, such as an oligonucleotide attached to a flow cell support. The adapter may also contain a nucleic acid tag, as described herein. The nucleic acid tag is typically positioned relative to the amplification primer and sequencing primer binding sites so that the nucleic acid tag is included in the amplicon and sequencing reads of a given nucleic acid molecule. Adapters of the same or different sequences can be ligated to each end of a nucleic acid molecule. In certain embodiments, the same adapter is ligated to each end of a nucleic acid molecule, except that the sequences of the nucleic acid tags are different. In some embodiments, the adaptor is a Y-shaped adaptor having one end blunt or tailed with one or more complementary nucleotides as described herein for joining to a nucleic acid molecule. In yet another exemplary embodiment, the adaptor is a bell-shaped adaptor comprising a blunt or tailed end for joining to a nucleic acid molecule to be analyzed. Other exemplary adaptors include T-tail adaptors and C-tail adaptors.
[0048] Administer: As used herein, "administering" or "administering" a therapeutic agent (e.g., an immunological therapeutic agent) to a subject means giving, applying, or contacting the composition with the subject. Administration can be accomplished by any of several routes, including, for example, topical, oral, subcutaneous, intramuscular, intraperitoneal, intravenous, intrathecal, and intradermal.
[0049] Detailed Description Tumor-derived somatic variants in circulating nucleic acids, such as cell-free DNA (cfDNA), can be used for targeted therapy selection, long-term monitoring, and early cancer detection. Cell-free tumor DNA (ctDNA) is a small DNA fragment released into the bloodstream from necrotic / apoptotic tumor cells or circulating tumor cells (CTCs). The majority of cfDNA originates from normal cells, including normal white blood cells undergoing apoptosis or necrosis. Recent studies have demonstrated that a significant proportion of mutations detected in cfDNA can originate from non-tumor sources, particularly clonal hematopoiesis, resulting in the accumulation of somatic mutations in hematopoietic stem cells and contributing to cfDNA "noise." The presence of non-tumor variants in plasma / cfDNA can confound ctDNA interpretation. Therefore, methods and related aspects for distinguishing between them are highly sought after. In particular, clonal hematopoietic-derived mutations (also known as clonal hematopoietic origin) refer to the somatic acquisition of genomic mutations in hematopoietic stem cells and / or hematopoietic progenitor cells, resulting in clonal expansion. Additionally, clonal hematopoiesis of indeterminate potential ("CHIP") refers to hematopoiesis in an individual involving proliferation of hematopoietic stem cells that contain one or more somatic mutations (e.g., hematologic cancer-associated mutations and / or non-cancer-associated mutations) but otherwise lack diagnostic criteria for hematologic malignancy, such as definitive morphological evidence of dysplasia. CHIP is a common age-related phenomenon in which hematopoietic stem cells contribute to the formation of genetically distinct subpopulations of blood cells.
[0050] Current approaches to identifying nucleic acid variants derived from clonal hematopoiesis or otherwise derived from cancer tumors include sequencing white blood cells (WBCs) or peripheral blood mononuclear cells and removing these sequences from the nucleic acid variants in the plasma portion of a given blood sample, sequencing tissue and removing all nucleic acid variants except those in the plasma fraction, or a combination of both techniques (Id.). Attempted bioinformatic approaches include removing nucleic acid variants present in genes frequently mutated in hematologic malignancies, as they are likely derived from the hematologic fraction; comparing nucleic acid fragment sizes at single loci in wild-type and WBC cfDNA; and using absolute or relative variant minor allele frequency cutoffs for tumors. A challenge with these approaches lies in the requirement for matched WBCs and tissue, which are not always available and complicate sample processing. This disclosure presents novel bioinformatic methods and related embodiments for classifying nucleic acid variants or mutations detected in plasma or other bodily fluids as tumor or non-tumor in origin, regardless of the availability of matched WBCs or tumor tissue.
[0051] In the context of the methods and compositions described herein, cell-free nucleic acids or "cfNA" refers to nucleic acids that are not contained within or otherwise bound to cells. Cell-free nucleic acids can include, for example, all unencapsulated nucleic acids derived from bodily fluids (e.g., blood, plasma, serum, urine, cerebrospinal fluid (CSF), etc.) from a subject. Cell-free nucleic acids include DNA (cfDNA), RNA (cfRNA), and genomic DNA, mitochondrial DNA, circulating DNA, siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), and / or hybrids thereof containing fragments of any of these. Cell-free nucleic acids can be double-stranded, single-stranded, or hybrids thereof. Cell-free nucleic acids can be released into bodily fluids through secretion or cell death processes, such as cell necrosis, apoptosis, etc. Some cell-free nucleic acids are released into bodily fluids from cancer cells, such as circulating tumor DNA (ctDNA). Others are released from healthy cells. CtDNA can be unencapsulated tumor-derived fragmented DNA. Another example of cell-free nucleic acid is fetal DNA circulating freely in the maternal bloodstream, also referred to as cell-free fetal DNA (cffDNA). Cell-free nucleic acid can have one or more epigenetic modifications, for example, cell-free nucleic acid can be acetylated, 5-methylated, ubiquitinated, phosphorylated, sumoylated, ribosylated, and / or citrullinated. In some embodiments, for example, the term "cell-free nucleic acid" refers to nucleic acid that is not contained within or otherwise associated with cells when isolated from a given subject.
[0052] Furthermore, in the context of the methods and compositions described herein, the cellular origin of cell-free nucleic acid refers to the cell type from which a given cell-free nucleic acid molecule originates or is otherwise derived (e.g., via an apoptotic process, a necrotic process, etc.). In certain embodiments, for example, a given cell-free nucleic acid molecule may be derived from a tumor cell (e.g., a cancerous cell, etc.) or a non-tumor or normal cell (e.g., a non-cancerous cell, a hematopoietic stem cell, etc.).
[0053] Also, for methods and compositions described herein that involve bioinformatic processes, a classifier refers to algorithmic computer code that receives test data as input and produces as output a classification of the input data as belonging to one or another class (e.g., tumor DNA or non-tumor DNA).
[0054] Furthermore, in the context of the methods and compositions described herein, minor allele frequency relates to the frequency with which a minor allele (e.g., a less common allele) occurs in a given population of nucleic acids, such as a sample obtained from a subject. Genetic variants with low minor allele frequency are typically present relatively infrequently in a sample.
[0055] Furthermore, in the context of the methods and compositions described herein, the mutant allele fraction ("MAF") refers to the fraction of nucleic acid molecules that have an allelic change or mutation relative to a reference at a given genomic location in a given sample. MAF is generally expressed as a fraction or percentage. For example, the MAF is typically less than about 0.5, 0.1, 0.05, or 0.01 (i.e., less than about 50%, 10%, 5%, or 1%) of all somatic variants or alleles present at a given locus.
[0056] Furthermore, in the context of the methods and compositions described herein, tumor fraction refers to an estimate of the fraction of nucleic acid molecules derived from tumors in a given sample.For example, the tumor fraction of a sample can be the maximum mutant allele fraction (MAX MAF) of the sample or the range of the sample, or a measure derived from the length, epigenetic status, or other characteristics of the cfNA fragments in the sample, or any other selected feature of the sample.The term "MAX MAF" refers to the maximum or maximum MAF of all somatic variants present in a given sample.In some embodiments, the tumor fraction of a sample is equal to the MAX MAF of the sample.
[0057] Figure 1 is a flowchart outlining exemplary method steps for distinguishing between tumor- and non-tumor-origin nucleic acid variants in a cell-free nucleic acid (cfNA) sample obtained from a test subject, according to some embodiments. For example, the methods disclosed herein can be used to facilitate the removal or reduction of background noise generated by non-tumor-origin nucleic acid variants (e.g., cfDNA fragments derived from non-cancerous or normal cells) detected in a given sample from a test subject, thereby improving assay sensitivity. As shown, method 100 includes determining (e.g., by a computer) the relative prevalence of tumor-associated gene variants observed only in reference plasma compared to reference white blood cells (e.g., a cell sample, a tissue sample, etc.) to generate a relative prevalence dataset (step 102). Method 100 also includes generating (e.g., by a computer) a set of probabilities of non-tumor origin from the relative prevalence dataset (step 104). Additionally, method 100 further includes distinguishing nucleic acid variants detected in a cfNA sample obtained from the test subject as being tumor-origin nucleic acid variants or non-tumor-origin nucleic acid variants using the set of probabilities of non-tumor origin (step 106). Related systems and computer-readable media for implementing the methods disclosed herein are further described below.
[0058] To further illustrate, Figure 2 is a flowchart outlining exemplary method steps for distinguishing between tumor- and non-tumor-origin nucleic acid variants in a cell-free nucleic acid (cfNA) sample obtained from a test subject, according to some embodiments. As shown, method 200 includes generating (e.g., by a computer) a tumor variant dataset including a population of reference tumor-associated gene variants, where the tumor variant dataset includes observed frequency (prevalence) data for tumor-associated gene variants in the population of reference tumor-associated gene variants among reference samples including reference plasma-only fluid samples (e.g., plasma samples, serum samples, etc.) and / or white blood cell samples (e.g., cell samples, tissue samples, etc.) (step 202). The reference samples are typically obtained from a single reference subject and / or different reference subjects with the same cancer type. Method 200 also includes determining (e.g., by a computer) a ratio of observed frequency data for tumor-associated gene variants among the reference samples in the population of reference tumor-associated gene variants to generate at least one MAF variance and / or relative prevalence dataset (step 204). Method 200 further includes generating (e.g., by a computer) a set of probabilities of non-tumor origin from the MAF variance and / or relative prevalence dataset (step 206). Additionally, method 200 also includes using the set of probabilities of non-tumor origin to distinguish nucleic acid variants detected in a cfNA sample obtained from the test subject as being tumor-origin nucleic acid variants or non-tumor-origin nucleic acid variants (step 208).
[0059] 3 is a flowchart outlining exemplary method steps for distinguishing or classifying tumor- and non-tumor-origin nucleic acid variants in a cell-free nucleic acid (cfNA) sample obtained from a test subject, according to some embodiments. As shown, method 300 includes obtaining raw data (step 302), e.g., in the form of cancer and non-cancerous (i.e., normal or healthy) sample data and white blood cell and / or tissue sample data (e.g., from the COSMIC Cancer Database, The Cancer Genome Atlas (TCGA) data, Memorial Sloan Kettering Cancer Center (MSKCC) data, and / or another data source). In the feature engineering process, input features are generated, for example, by calculating mutant allele fraction (MAF) variation over time (step 303), calculating the raw number and prevalence of nucleic acid variants for all cancer types, and calculating the ratio between the prevalence of nucleic acid variants observed in plasma and / or other body fluid and tissue datasets for all cancer types (step 304), calculating the proportion of nucleic acid variants in hematological malignancies or other cancer types (step 305), and testing plasma and / or other body fluid sample prevalences for uniformity across cancer types (e.g., development uniformity score) (step 306). Bioinformatic data may include observed frequencies of genetic variants among samples of specific cancer types, including hematological malignancies; prevalence of variants in plasma and / or other body fluids, tumor tissue, leukocytes, mutant allele fractions of variants, etc. Additional or other data types may be used as needed for these feature engineering processes. Method 300 also includes transformation and cleanup processes (step 308), such as cleaning up for sample prevalence (e.g., adjusting for samples with low numbers of a given nucleic acid variant, small number of samples, etc.), performing a logarithmic transformation (e.g., Log(x+1) or Np.log1p), and performing normalization (e.g., Yeo-Johnson normalization, min-max normalization, z-score normalization, and / or the like).Method 300 also includes a machine learning step (step 310) of creating a machine learning model to provide the probability of a non-tumor nucleic acid variant being present in a given sample, using, for example, logistic regression or deep learning techniques. Exemplary models that can be used for training and further classification include, without limitation, logistic regression, probit regression, decision trees, random forests, gradient boosting, support vector machines, k-nearest neighbors, neural networks, or an ensemble of more than one of these methods. Ensemble methods are meta-algorithms that combine several machine learning techniques into a single predictive model to reduce variance (bagging), bias (boosting), or improve prediction (stacking). Most ensemble methods use a single base learning algorithm to generate homogeneous base learners, i.e., learners of the same type, resulting in a homogeneous ensemble. Some methods also use heterogeneous learners, i.e., learners of different types, resulting in a heterogeneous ensemble. For an ensemble method to be more accurate than any of its individual members, the base learners must be as accurate and diverse as possible.
[0060] The data set is divided into a training set and a test set using various methods as needed. In some embodiments, for example, the data set is randomly divided into a training set and a test set in an 80 / 20 ratio. In addition, method 300 also includes selecting a cutoff value to determine the threshold for classifying nucleic acid variants as being of tumor or non-tumor cell origin (step 312).
[0061] Body fluid:tissue ratio - binary classification
[0062] Some embodiments involve comparing the prevalence of variants observed in a body fluid sample (e.g., plasma sample) dataset compared to their occurrence in a tissue dataset of the same cancer origin. In certain of these embodiments, a logistic regression is performed on these ratios to obtain the probability of clonal hematopoietic origin.
[0063] In some embodiments, the values of the performance metric may include, for example, accuracy (i.e., the proportion of correct predictions), balanced_accuracy (defined as the average of the recall obtained for each class), precision_macro (which involves calculating a metric for each label and then finding their unweighted average, although this approach does not account for label imbalance), precision_micro (which involves calculating the metric globally by counting the total true positives, false negatives, and false positives), precision_weighted (which involves calculating a metric for each label and finding their average weighted by support (e.g., to determine the number of true instances of each label)), etc. In particular embodiments, the performance metric is estimated by 5-fold stratified cross-validation on a training set (e.g., folds are created by preserving the proportion of samples in each class).
[0064] The Box-Cox transformation is optionally used to transform non-normal distributions into normal distributions, but this technique does not work with negative numbers. In contrast, the Yeo-Johnson transformation allows for operation with negative numbers. For example, for both logistic regression and support vector machine (SVM) models, all features are optionally first transformed with the Yeo-Johnson transformation (a parametric monotonic transformation applied to make the data more Gaussian to stabilize variance and minimize skewness). In some embodiments, zero-mean, unit variance normalization is further applied to the transformed data.
[0065] In some embodiments, the basic inputs used to define the set of parameters are (1) the model type and (2) the set of hyperparameters. In particular embodiments, the resulting parameters are used for all future classifications. In some embodiments, the training set is used to perform a grid search with 5-fold stratified cross-validation over the following sets of hyperparameters (e.g., to define the cost of misclassification): kernel: linear, C: [0.001, 0.01, 0.1, 1, 10, 100, 1000], and kernel: radial basis function (rbf), C: [0.001, 0.01, 0.1, 1, 10, 100, 1000], gamma: [0.0001, 0.001, 0.01, 0.1, 0.2, 0.3, 0.4, 0.5, 1].
[0066] Direct training on the dataset Some embodiments use machine learning based on known datasets and features related to variant gene name, location, cancer type, chromosomal location, and other features to predict tumor / non-tumor origin. In these embodiments, the method typically involves training a machine learning model on clonal hematopoietic (CH) and tissue-specific training datasets to identify features specific to either origin and applying the model to historical variants observed in previous datasets to determine the probability that a given variant can be attributed to CH. In these embodiments, the method also typically involves determining which probability threshold is optimal for accurate classification of CH and applying this list of probabilities to a new dataset to classify its origin as tumor or clonal hematopoietic. In certain of these embodiments, the top 10 percentile of variants, while being a minority, have high predictive value for CH origin.
[0067] Prevalence in body fluids relative to tumor tissue Certain embodiments use the higher prevalence of a given variant in a body fluid (e.g., plasma or serum) database compared to its occurrence in tumor tissue, which may be less confounded by clonal hematopoiesis (CH) and thus may signal variants that are more likely to originate from CH. In some of these embodiments, the method involves determining the prevalence of specific variants present in the body fluid database and comparing them to the prevalence observed in a primary tissue database, such as the COSMIC database. Some of these embodiments involve determining the ratio of the prevalence of the variant observed only in plasma to the prevalence of the variant observed in tissue samples. Some of these embodiments involve calculating an odds ratio for the variant's prevalence and the probability that the value of the odds ratio is equal to, greater than, or less than 1. In these embodiments, the method also generally involves applying a machine learning model to these relative prevalence values of body fluid versus tissue prevalence to determine the probability that the variant is more likely to originate from CH or a non-tumorous condition. A probability threshold (e.g., about the 10th, 15th, 20th, 25th, 30th, 35th, 40th, or another percentile) is typically used as a cutoff for classification. In general, a minority of variants have a high predictive value for being non-tumor or CH.
[0068] Multiclass models using distribution or homogeneity tests across tumor types In some embodiments, tumor-specific variants have a specific distribution or selection depending on the biology of each cancer type. If a variant is not specific to a cancer type, it typically has a uniform distribution, which may indicate a passenger mutation or a non-tumor state. Therefore, certain methods determine the prevalence of variants across tumor types, or their relative proportions and representation, and train machine learning models to separate different tumor and non-tumor classes. Some of these methods involve using the coefficient of variation to determine the distribution and significant enrichment in specific tumor types. In certain of these embodiments, very few variants are predictive, tumor-type specific, and unlikely to be CH. Some variants have no demonstrable preference for a specific tumor type and have low prevalence across all tumor types, indicating a high likelihood of CH. In general, if a variant is substantially uniform across tumor types, it is likely to be non-tumor / CH in origin, while if a variant is very common in a particular tumor, it is likely to be under biological selection in the tumor. Current methods that strictly rely on patient age or absolute VAF, and that ignore expected relative prevalence in different tumor contexts, fail to consider these underlying disease-specific mechanisms (or lack thereof) that drive the observed VAF and key biological features indicative of variant origin.
[0069] Other input features to the machine learning model in these embodiments may include, for example, variant-based tumor classification (e.g., tumor type, or predicted tumor type), the presence or signature of methylation, other variants within a given sample (e.g., known CH variants present in the sample, increasing the likelihood that other variants in the sample are of non-tumor origin), the difference in family size of a given variant versus a reference allele, the nature of the observed nucleotide or other change in the variant, the absolute value of the MAF within the sample, the relative value of the MAF within the sample, how the value of the MAF within the sample changes over time relative to other variants, and / or the like.
[0070] Monitor the change in MAF value over time In certain cases, mutant clones of non-tumor variants are likely to remain more stable over time in a subject compared to tumor-derived variants. Thus, in some embodiments, the method involves calculating the coefficient of variation (CV, variance relative to the mean) of the rate of change over time for each patient with multiple time points (e.g., >3), and calculating the statistics and distribution of CV across all variants and patients. In these embodiments, known driver or tumor variants generally have a dynamic percentage across time points (due to tumor growth and shrinkage) and a larger CV compared to non-tumor variants. This can also be used as an input feature for the classifier. In contrast, non-tumor variant MAFs will typically be less dynamic, more stable over time, and have a lower CV over time compared to true tumor variants. The distribution of these CVs can be separated in a machine learning model to provide robust classification of tumor or non-tumor status. Other input features to the machine learning model in these embodiments may include, for example, variant clonality (VAF relative to tumor fraction) over time or across patients, fragmentation data points, fragment size, location, patient age (older patients are more likely to have CHIP), and / or the like. Current methods that can track VAF or VAF variance across time points in a single patient are less accurate than approaches that aggregate VAF across a large number of patients, especially if these patients are all serially tested on the same platform and bioinformatics pipeline, leading to consistent VAF and more robust measurements of variance. Furthermore, this classification method may use a static threshold that does not adjust the value of variance relative to absolute VAF; non-tumor variants with higher VAF may be confused with lower VAF variants with similar measures of variance over time. Machine learning models that take into account both absolute VAF and VAF variance across time points in a sufficiently large cohort of patients measured on the same platform have higher resolution for classification and are less likely to result in false positive or negative labeling of tumor / non-tumor status.
[0071] 4, an additional method for creating a predictive model (e.g., a classification model) is described. The described method may use machine learning (“ML”) techniques to train at least one ML module 430 configured to classify mutations detected in plasma as of tumor or non-tumor origin, which may result from clonal hematopoiesis or biological noise, based on analysis of one or more training datasets 410A-410N by a training module 420.
[0072] One or more training datasets 410A-410N may include cancer / non-cancerous (e.g., tumor / non-tumor) body fluid plasma-only sample data and cancer / non-cancerous (e.g., tumor / non-tumor) white blood cell and / or non-body fluid (e.g., tissue) sample data. One or more training datasets 410A-410N may include cancer / non-cancerous (e.g., tumor / non-tumor) body fluid plasma-only sample data and cancer / non-cancerous white blood cell and / or non-body fluid (e.g., tissue) sample data (e.g., from the COSMIC Cancer Database, The Cancer Genome Atlas (TCGA) data, and / or another data source). Subsets of the cancer / non-cancerous body fluid sample data and / or the cancer / non-cancerous non-body fluid sample data may be randomly assigned to the training dataset 410 or the test dataset. In some implementations, the assignment of data to the training dataset or the test dataset may not be completely random. In this case, one or more criteria may be used during the assignment. In general, any suitable method can be used to assign data to a training or testing data set while ensuring that the data distribution is somewhat similar in the training and testing data sets.
[0073] The training module 420 may train the ML module 430 by extracting feature sets from the cancer / non-cancer body fluid sample data and / or the cancer / non-cancer non-body fluid sample data in the training dataset 410 according to one or more feature selection techniques. The training module 420 may train the ML module 430 by extracting feature sets from the training dataset 410 that include statistically significant features.
[0074] The training module 420 may extract feature sets from the training dataset 410 in various ways. The training module 420 may perform feature extraction multiple times, each time using a different feature extraction technique. In one example, feature sets created using different techniques may each be used to create a different machine learning-based classification model 440. For example, the feature set with the highest quality metric may be selected for use in training. The training module 420 may use the feature set(s) to build one or more machine learning-based classification models 440A-440N configured to classify new variants (e.g., of unknown origin) as tumor or non-tumor in origin.
[0075] The training dataset 410 may be analyzed to determine any dependencies, associations, and / or correlations between features in the training dataset 410 and experimental parameters. The identified correlations may have the form of a list of features. As used herein, the term "feature" may refer to any characteristic of an item of data that can be used to determine whether the item of data falls within one or more particular categories. By way of example, features described herein may include one or more of the observed frequency of a genetic variant among samples of a particular cancer type, including a hematological malignancy; the prevalence of the variant in plasma, tumor tissue, or leukocytes; and / or the minor allele frequency of the variant.
[0076] The feature selection technique may include one or more feature selection rules. The one or more feature selection rules may include feature occurrence rules. The feature occurrence rules may include determining which features in the training data set 410 occur a threshold number of times and identifying those features that meet the threshold as features.
[0077] A single feature selection rule may be applied to select features, or multiple feature selection rules may be applied to select features. Feature selection rules may be applied in a cascade fashion, where the feature selection rules are applied in a specific order and build on the results of previous rules. For example, feature occurrence rules may be applied to the training dataset 410 to create a first list of features. The final list of features may be analyzed according to additional feature selection techniques to determine one or more feature groups (e.g., groups of features that can be used to classify variants as tumor or non-tumor in origin). Any feature selection technique, such as a filter, wrapper, and / or embedding method, may be used to identify feature groups using any suitable computational technique. One or more feature groups may be selected according to a filter method. Examples of filter methods include Pearson correlation, linear discriminant analysis, analysis of variance (ANOVA), chi-square, combinations thereof, and the like. Feature selection according to a filter method is independent of any machine learning algorithm. Instead, features may be selected based on their scores in various statistical tests for correlation with outcome variables.
[0078] As another example, one or more feature groups may be selected according to a wrapper method. The wrapper method may be configured to use a subset of features and train a machine learning model using the subset of features. Features may be added to and / or removed from the subset based on inferences drawn from previous models. Examples of wrapper methods include forward feature selection, backward feature reduction, recursive feature reduction, and combinations thereof. As an example, one or more feature groups may be identified using forward feature selection. Forward feature selection is an iterative method that starts with no features in the machine learning model. In each iteration, features that most improve the model are added until adding new variables does not improve the machine learning model's performance. As an example, one or more feature groups may be identified using backward reduction. Backward reduction is an iterative method that starts with all features in the machine learning model. In each iteration, the lowest-ranking features are removed until no improvement is observed with feature removal. Recursive feature reduction may be used to identify one or more feature groups. Recursive feature reduction is a greedy optimization algorithm whose goal is to find the best-performing feature subset. Recursive feature reduction iteratively builds a model, eliminating the best or worst-performing features at each iteration. Recursive feature reduction builds the next model in which features remain until all features are exhausted. Recursive feature reduction then ranks the features based on their order of reduction.
[0079] As a further example, one or more feature groups may be selected according to an embedding method. The embedding method combines the qualities of filter and wrapper methods. Examples of embedding methods include least absolute shrinkage and selection operator (LASSO) and ridge regression, which implement penalty functions to reduce overfitting. For example, LASSO regression uses L1 regularization, which adds a penalty equal to the absolute value of the coefficient magnitude, while ridge regression uses L2 regularization, which adds a penalty equal to the square of the coefficient magnitude.
[0080] After the training module 420 creates the feature set(s), the training module 420 may create a machine learning-based classification model 440 based on the feature set(s). A machine learning-based classification model may refer to a complex mathematical model for data classification created using machine learning techniques. In one example, the machine learning-based classification model 440 may include a map of support vectors representing boundary features. By way of example, the boundary features may be selected from the feature set and / or may represent the highest-ranking features therein.
[0081] The training module 420 may construct machine learning based classification models 440A-440N using the feature set determined or extracted from the training dataset 410. In some examples, the machine learning based classification models 440A-440N may be combined into a single machine learning based classification model 440. Similarly, the ML module 430 may represent a single classifier including single or multiple machine learning based classification models 440 and / or multiple classifiers including single or multiple machine learning based classification models 440.
[0082] The features may be combined in a classification model trained using machine learning techniques such as discriminant analysis; decision trees; nearest neighbor (NN) algorithms (e.g., k-NN models, replicator NN models, etc.); statistical algorithms (e.g., Bayesian networks, etc.); clustering algorithms (e.g., k-means, mean shift, etc.); neural networks (e.g., reservoir networks, artificial neural networks, etc.); support vector machines (SVMs); logistic regression algorithms; linear regression algorithms; Markov models or chains; principal component analysis (PCA) (e.g., for linear models); multilayer perceptron (MLP) ANNs (e.g., for nonlinear models); replicated reservoir networks (e.g., for nonlinear models, typically for time series); random forest classification; combinations thereof, and / or the like. The resulting ML module 430 may include a decision rule or mapping for each feature to determine the tumor / non-tumor origin of the variants.
[0083] In one embodiment, the training module 420 may train the machine learning-based classification model 440 as a convolutional neural network (CNN), which includes at least one convolutional feature layer and three fully connected layers leading to a final classification layer (softmax), which may finally be applied to combine the outputs of the fully connected layers using a softmax function, as known in the art.
[0084] The feature(s) and the ML module 430 can be used to predict the tumor / non-tumor origin of variants in a test dataset. In one example, the predicted result for each variant can include a confidence level corresponding to the likelihood or probability that the variant in the test dataset is associated with a tumor or non-tumor origin. The confidence level can be a value between 0 and 1. In one example, when there are two states (e.g., tumor origin and non-tumor origin), the confidence level can correspond to a value p, which refers to the likelihood that a particular variant belongs to the first state (e.g., tumor origin). In this case, the value 1-p can refer to the likelihood that a particular variant belongs to the second state (e.g., non-tumor origin). Generally, multiple confidence levels can be provided for each variant in the test dataset and for each feature when there are three or more states. Top-performing features can be determined by comparing the results obtained for each test variant with the known tumor / non-tumor origin of each test variant. Generally, top-performing features have results that closely match the known tumor / non-tumor origin states. The top performing feature(s) can be used to predict / classify the tumor / non-tumor origin status of a given variant.
[0085] 5 is a flowchart illustrating an example training method 500 for creating an ML module 430 using the training module 420. The training module 420 can implement supervised, unsupervised, and / or semi-supervised (e.g., reinforcement-based) machine learning-based classification models 440. The method 500 illustrated in FIG. 5 is an example of a method for supervised learning. Variations of this example training method are described below, although other training methods can be implemented to train unsupervised and / or semi-supervised machine learning models as well.
[0086] The training method 500 may determine (e.g., access, receive, retrieve, etc.) data in step 510. The data may include cancer / non-cancerous (e.g., tumor / non-tumor) body fluid sample data and cancer / non-cancerous (e.g., tumor / non-tumor) non-body fluid (e.g., tissue) sample data. The data may include one or more variants, each variant having an assigned tumor or non-tumor origin status.
[0087] The training method 500 may create a training data set and a test data set in step 520. The training data set and the test data set may be created by randomly assigning data to either the training data set or the test data set. In some implementations, the assignment of the computational parameters and associated experimental parameters as training or test data may not be completely random. As an example, a majority of the computational parameters and associated experimental parameters may be used to create the training data set. For example, 75% of the computational parameters and associated experimental parameters may be used to create the training data set, and 25% may be used to create the test data set. In another example, 80% of the computational parameters and associated experimental parameters may be used to create the training data set, and 20% may be used to create the test data set.
[0088] The training method 500 may, at step 530, determine (e.g., extract, select, etc.) one or more features that can be used by a classifier to distinguish between different classifications of, for example, tumorous versus non-tumorous conditions. As one example, the training method 500 may determine a set of features from the cancer / non-cancerous body fluid sample data and the cancer / non-cancerous non-body fluid sample data. In a further example, the set of features may be determined from data other than the cancer / non-cancerous body fluid sample data and the cancer / non-cancerous non-body fluid sample data in either the training dataset or the test dataset. Such other data may be used to determine an initial set of features, which may be further reduced using the training dataset.
[0089] The training method 500 may train one or more machine learning models using one or more features at step 540. In one example, the machine learning models may be trained using supervised learning. In another example, other machine learning techniques, including unsupervised learning and semi-supervised learning, may be used. The machine learning models trained at 540 may be selected based on different criteria depending on the problem to be solved and / or the data available in the training dataset. For example, machine learning classifiers may be subject to different degrees of bias. Thus, at step 550, multiple machine learning models may be trained, optimized, improved, and cross-validated at 540.
[0090] The training method 500 may select one or more machine learning models to build a predictive model at 560. The predictive model may be evaluated using a test dataset. The predictive model may analyze the test dataset and generate a predicted tumor / non-tumor origin status at step 570. The predicted tumor / non-tumor origin may be evaluated at step 580 to determine whether such value achieves a desired level of accuracy. The performance of a predictive model may be evaluated in several ways based on the true positive, false positive, true negative, and / or false negative classification of some of the data points represented by the predictive model.
[0091] For example, a false positive of a predictive model may refer to the number of times the predictive model incorrectly classified a variant as having a tumor origin when it actually had a non-tumor origin. Conversely, a false negative of a predictive model may refer to the number of times the machine learning model classified a variant as having a non-tumor origin when it actually had a tumor origin. True negatives and true positives may refer to the number of times the predictive model correctly classified one or more variants. Related to these measurements are the concepts of recall and precision. Generally, recall refers to the ratio of true positives to the sum of true positives and false negatives, which quantifies the sensitivity of a predictive model. Similarly, precision refers to the ratio of true positives, which are the sum of true positives and false positives. Once such a desired level of accuracy is reached, the training phase ends and a predictive model (e.g., ML module 430) may be output at step 590. However, if the desired level of accuracy is not reached, subsequent iterations of the training method 500 may be performed, beginning at step 510, using modifications, such as considering a larger set of data.
[0092] 6 is a diagram of an exemplary process flow for using a machine learning-based classifier to classify variants as tumor or non-tumor origin. As shown in FIG. 6, unclassified variants 610 may be provided as input to the ML module 430. The ML module 430 may process the unclassified variants 610 using the machine learning-based classifier to arrive at a prediction result 620. The prediction result 620 may identify one or more characteristics of the unclassified variants 610. For example, the classification result 620 may identify the origin state of the unclassified variants 610 (e.g., whether the variant is tumor or non-tumor origin). Thus, in one embodiment, a method is disclosed that is implemented using a network-based computer system having one or more processors, a network interface, and one or more memories, the method including: retrieving, by the computer system, genetic information and additional information for a plurality of tumor and non-tumor plasma-only and a plurality of tumor and non-tumor non-body fluid (e.g., tissue) samples from the one or more memories, where the additional information includes a state of tumor or non-tumor origin; and training, by the one or more processors, one or more machine learning models by fitting one or more models to the genetic information and additional information, where each of the one or more models is configured to receive an individual's genetic information as input and provide, as output, a prediction of the individual having or developing a tumor.
[0093] System and computer-readable medium The present disclosure also provides various systems, bioinformatics pipelines, and computer program products or machine-readable media. In some embodiments, for example, the methods described herein are optionally performed or facilitated, at least in part, using systems, distributed computing hardware and applications (e.g., cloud computing services), electronic communications networks, communications interfaces, computer program products, machine-readable media, electronic storage media, software (e.g., machine-executable code or logical instructions), and / or the like. For illustrative purposes, FIG. 7 provides a schematic diagram of an exemplary system suitable for implementing at least aspects of the methods disclosed herein. As shown, system 700 includes at least one controller or computer, such as a server 702 (e.g., a search engine server), including a processor 704 and memory, storage devices, or memory components 706, and one or more other communication devices 714 and 716 (e.g., client-side computer terminals, phones, tablets, laptops, other mobile devices, etc.) located remotely from and communicating with the remote server 702 via an electronic communications network 712, such as the Internet or other internetwork. The communication devices 714 and 716 typically include, for example, electronic displays (e.g., internet-enabled computers, etc.) in communication with the server 702 computer over the network 712, the electronic displays including user interfaces (e.g., graphical user interfaces (GUIs), web-based user interfaces, and / or the like) for displaying results from implementing the methods described herein. In certain embodiments, the communication network also encompasses the physical transfer of data from one location to another, for example, using hard drives, thumb drives, or other data storage mechanisms.System 700 also includes a program product 708 stored on a computer- or machine-readable medium, such as one or more various types of memory, such as memory 706 of server 702, readable by server 702 to facilitate a navigation search application or other executable file executable by one or more other communication devices, such as 714 (shown generally as a desktop or personal computer) and 716 (shown generally as a tablet computer). In some embodiments, system 700 also optionally includes at least one database server, such as server 710 associated with an online website that stores searchable data (e.g., nucleic acid variant lists, indexed treatments, etc.) directly or via search engine server 702. System 700 also optionally includes one or more other servers located remotely from server 702, each of which is associated with one or more database servers 710 located remotely from or locally relative to each of the other servers, as appropriate. The other servers can beneficially serve geographically dispersed users and enhance geographically distributed operations.
[0094] As will be appreciated by those skilled in the art, the memory 706 of the server 702 may include volatile and / or nonvolatile memory, including, for example, RAM, ROM, and magnetic or optical disks, among others. While illustrated as a single server, those skilled in the art will appreciate that the illustrated configuration of the server 702 is provided by way of example only, and that other types of servers or computers configured according to various other methodologies or architectures may also be used. The server 702 shown schematically in FIG. 7 represents a server or server cluster or server farm and is not limited to individual physical servers. A server site may be deployed as a server farm or server cluster managed by a server hosting provider. The number of servers and their architecture and configuration may be increased based on the use, demand, and capacity requirements of the system 700. Also, as will be appreciated by those skilled in the art, the other user communication devices 714 and 716 in these embodiments may be, for example, laptops, desktops, tablets, personal digital assistants (PDAs), mobile phones, servers, or other types of computers. As known and understood by those skilled in the art, network 712 may include portions of the Internet, an intranet, a telecommunications network, an extranet, or the World Wide Web, and / or a local or other area network of multiple computers / servers in communication with one or more other computers via a communications network.
[0095] As will be further understood by those skilled in the art, the exemplary program product or machine-readable medium 708 is in the form of microcode, programs, cloud computing formats, routines, and / or symbolic languages that provide one or more sets of ordered operations that control the functioning of and direct the operation of the hardware, as appropriate. According to exemplary embodiments, the program product 708 also need not reside entirely in volatile memory, but rather can be selectively loaded as needed according to various methods as will be known and understood by those skilled in the art.
[0096] As will be further understood by those skilled in the art, the terms “computer-readable medium” or “machine-readable medium” refer to any medium that participates in providing instructions to a processor for execution. By way of example, the terms “computer-readable medium” or “machine-readable medium” encompass distribution media, cloud computing formats, intermediate storage media, computer execution memory, and any other medium or device that can store, for example, for reading by a computer, a program product 708 that implements the functions or processes of various embodiments of the present disclosure. A “computer-readable medium” or “machine-readable medium” may take many forms, including, but not limited to, non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical or magnetic disks. Volatile media include dynamic memory, such as the main memory of a given system. Transmission media include coaxial cables, copper wire, and fiber optics, including the wires that comprise a bus. Transmission media can also take the form of acoustic or light waves, such as those created during radio wave and infrared data communications, among others. Exemplary forms of computer-readable media include a floppy disk, a flexible disk, a hard disk, a magnetic tape, a flash drive, or any other magnetic medium, a CD-ROM, any other optical medium, punch cards, paper tape, any other physical medium with a pattern of holes, RAM, PROM, and EPROM, FLASH-EPROM, any other memory chip or cartridge, a carrier wave, or any other medium from which a computer can read.
[0097] The program product 708 is copied from the computer-readable medium to a hard disk or similar intermediate storage medium, as needed. When the program product 708, or portions thereof, are executed, they are loaded, as needed, from their distribution medium, their intermediate storage medium, etc. into the execution memory of one or more computers, and configure the computer(s) to operate according to the functions or methods of the various embodiments. All such operations are well known to those skilled in the art of, for example, computer systems.
[0098] To further illustrate, in certain embodiments, the present application provides a system including one or more processors and one or more memory components in communication with the processor. The memory components typically include one or more instructions that, when executed, cause the processor to provide information (e.g., via communications devices 714, 716, etc.) that causes the processor to display at least one nucleic acid variant list, a variant classification call report or result, a selected treatment, and / or the like, and / or receive information (e.g., via communications devices 714, 716, etc.) from other system components and / or from a system user.
[0099] In some embodiments, the program product 708 includes non-transitory computer-executable instructions that, when executed by the electronic processor 704, perform at least the following: (i) creating a tumor variant dataset comprising a population of reference tumor-associated gene variants, wherein the tumor variant dataset comprises observed frequency data between reference samples comprising reference plasma only and / or reference leukocytes for tumor-associated gene variants in the population of reference tumor-associated gene variants, and the reference samples are obtained from a single reference subject and / or different reference subjects having the same cancer type; (ii) determining a ratio of observed frequency data between reference samples for tumor-associated gene variants in the population of reference tumor-associated gene variants to generate a relative prevalence dataset; (iii) creating a set of probabilities of non-tumor origin from the relative prevalence dataset; and (iv) using the set of probabilities of non-tumor origin to distinguish nucleic acid variants detected in a cfNA sample obtained from the test subject as being tumor-originated nucleic acid variants or non-tumor-originated nucleic acid variants.
[0100] System 700 also typically includes additional system components configured to implement various aspects of the methods described herein. In some of these embodiments, these one or more additional system components are located remotely from and communicate with the remote server 702 via an electronic communications network 712, while in other embodiments, these one or more additional system components are located locally and communicate with the server 702 (i.e., when the electronic communications network 712 is not present) or directly with, for example, a desktop computer 714.
[0101] In some embodiments, for example, additional system components include a sample preparation component 718 operably connected to the controller 702 (either directly or indirectly, e.g., via electronic communications network 712). The sample preparation component 718 is configured to prepare nucleic acids in a sample (e.g., to prepare a library of nucleic acids) to be amplified and / or sequenced by a nucleic acid amplification component (e.g., a thermal cycler, etc.) and / or a nucleic acid sequencer. In certain of these embodiments, the sample preparation component 718 is configured to isolate nucleic acids from other components in the sample, attach barcode-containing or adapters to nucleic acids as described herein, selectively enrich one or more regions from a genome or transcriptome prior to sequencing, and / or the like.
[0102] In certain embodiments, system 700 also includes a nucleic acid amplification component 720 (e.g., a thermal cycler, etc.) operably connected to controller 702 (either directly or indirectly (e.g., via electronic communication network 712)). Nucleic acid amplification component 720 is configured to amplify nucleic acids in a sample from a subject. For example, nucleic acid amplification component 720 is configured to amplify selectively enriched regions from a genome or transcriptome in a sample as described herein, as needed.
[0103] System 700 also typically includes at least one nucleic acid sequencer 722 operably connected to controller 702 (either directly or indirectly, e.g., via electronic communications network 712). Nucleic acid sequencer 722 is configured to provide sequence information from nucleic acids (e.g., amplified nucleic acids) in a sample from a subject. Essentially, any type of nucleic acid sequencer can be adapted for use in these systems. For example, nucleic acid sequencer 722 is optionally configured to perform pyrosequencing, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing by ligation, sequencing by hybridization, or other techniques on the nucleic acids to generate sequencing reads. Optionally, nucleic acid sequencer 722 is configured to group sequence reads into families of sequence reads, each family including sequence reads generated from nucleic acids in a given sample. In some embodiments, nucleic acid sequencer 722 generates sequencing reads using clonal single-molecule arrays derived from a sequencing library. In certain embodiments, the nucleic acid sequencer 722 includes at least one chip having an array of microwells for sequencing a sequencing library to generate sequencing reads.
[0104] To facilitate full or partial system automation, system 700 also typically includes a material transfer component 724 operably connected to controller 702 (either directly or indirectly (e.g., via electronic communications network 712)). Material transfer component 724 is configured to transfer one or more materials (e.g., nucleic acid samples, amplicons, reagents, and / or the like) to and / or from nucleic acid sequencer 722, sample preparation component 718, and nucleic acid amplification component 720.
[0105] Further details regarding computer systems and networks, databases, and computer program products are provided, for example, in Peterson, Computer Networks: A Systems Approach, Morgan Kaufmann, 5th Ed. (2011); Kurose, Computer Networking: A Top-Down Approach, Pearson, 7th Ed. (2016); Elmasri, Fundamentals of Database Systems, Addison Wesley, 6th Ed. (2010); Coronel, Database Systems: Design, Implementation, & Management, Cengage Learning, 11th Ed. (2014); Tucker, Programming Languages, McGraw-Hill Science / Engineering / Math, 2nd Ed. (2006); and Rhoton, Cloud Computing Architected: Solution Design Handbook, Recursive Press (2011), each of which is incorporated by reference in its entirety.
[0106] Sample collection and preparation The sample may be any biological sample isolated from a subject. The sample may include body fluids or body tissues (e.g., known or suspected solid tumors). The sample may include whole blood, platelets, serum, plasma, feces, red or white blood cells, endothelial cells, tissue biopsies, cerebrospinal fluid, synovial fluid, lymphatic fluid, ascites fluid, interstitial or extracellular fluid, fluid in the space between cells (including gingival crevicular fluid, bone marrow, pleural effusion, cerebrospinal fluid, saliva, mucous membranes, sputum, semen, sweat, and urine). The sample is preferably a body fluid, particularly blood and its fractions, and urine. Such samples contain nucleic acids released from tumors. The nucleic acids may include DNA and RNA, and may be in double-stranded and / or single-stranded form. Sample can be the form that is first isolated from subject, or can be further processed to remove or add components such as cells, to concentrate one component to another, or to convert one form of nucleic acid into another form, for example, RNA into DNA, or single-stranded nucleic acid into double-stranded.Therefore, for example, the body fluid for analysis is the plasma or serum that contains cell-free nucleic acid, for example, cell-free DNA (cfDNA).
[0107] In certain embodiments, polynucleotides can be enriched prior to sequencing. Enrichment can be performed for specific target regions ("target sequences") or non-specifically. In some embodiments, targeted regions of interest can be enriched using differential tiling and capture schemes with capture probes ("baits") selected for one or more bait set panels. Differential tiling and capture schemes use bait sets at different relative concentrations, subject to a set of constraints (e.g., sequencer constraints such as sequencing load, availability of each bait, etc.), to differentially tile (e.g., at different "resolutions") across genomic regions associated with the baits and capture them at a desired level for downstream sequencing. These target genomic regions of interest can include regions of the genome or transcriptome of interest. In some embodiments, biotin-labeled beads bearing probes for one or more regions of interest can be used to capture target sequences, followed by optional subsequent amplification of those regions to enrich for the regions of interest.
[0108] Sequence capture typically involves the use of oligonucleotide probes that hybridize to the target sequence. Probe set strategies can involve tiling probes across the region of interest. Such probes can be, for example, about 60-130 bases long. Sets can have a depth of about 2x, 3x, 4x, 5x, 6x, 8x, 9x, 10x, 15x, 30x, 50x, or more. The effectiveness of sequence capture depends, in part, on the length of the sequence in the target molecule that is complementary (or nearly complementary) to the sequence of the probe.
[0109] In some embodiments, the methods of the present disclosure involve selectively enriching regions from the genome or transcriptome of a subject prior to sequencing. In other embodiments, the methods of the present disclosure involve non-selectively enriching regions from the genome or transcriptome of a subject prior to sequencing.
[0110] In certain embodiments, a sample index sequence is introduced into the polynucleotide after enrichment. The sample index sequence may be introduced by PCR or, optionally, ligated to the polynucleotide as part of an adaptor.
[0111] The volume of the bodily fluid can depend on the desired read depth of the sequenced region. Exemplary volumes are 0.4-40 ml, 5-20 ml, and 10-20 ml. For example, the volume can be 0.5 ml, 1 ml, 5 ml, 10 ml, 20 ml, 30 ml, or 40 ml. The volume of the sampled bodily fluid can be 5-20 ml.
[0112] Sample can contain various amounts of nucleic acid, which contain genome equivalent.For example, the sample of about 30ng DNA can contain about 10,000 (104) haploid human genome equivalent, and in the case of cfDNA, about 200 billion (2 x 1011) individual polynucleotide molecules.Similarly, the sample of about 100ng DNA can contain about 30,000 haploid human genome equivalent, and in the case of cfDNA, about 600 billion individual molecules.
[0113] The sample can include nucleic acids from different sources, such as cells and acellular sources. The sample can include nucleic acids with mutations. For example, the sample can include DNA with germline mutations and / or somatic mutations. The sample can include DNA with cancer-related mutations (e.g., cancer-related somatic mutations).
[0114] Exemplary amounts of cell-free nucleic acid in a sample prior to amplification range from about 1 fg to about 1 μg, e.g., 1 pg to 200 ng, 1 ng to 100 ng, or 10 ng to 1000 ng. For example, the amount can be up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng of cell-free nucleic acid molecules. The amount can be at least 1 fg, at least 10 fg, at least 100 fg, at least 1 pg, at least 10 pg, at least 100 pg, at least 1 ng, at least 10 ng, at least 100 ng, at least 150 ng, or at least 200 ng of cell-free nucleic acid molecules. The amount can be up to 1 femtogram (fg), 10 fg, 100 fg, 1 picogram (pg), 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng, or 200 ng of cell-free nucleic acid molecules. The method can include obtaining between 1 femtogram (fg) and 200 ng.
[0115] Cell-free nucleic acids have an exemplary size distribution of about 100 to 500 nucleotides, with molecules of 110 to about 230 nucleotides accounting for about 90% of the molecules, with a mode of about 168 nucleotides in humans and a second minor peak in the range of 240 to 430 nucleotides. Cell-free nucleic acids can be about 160 to about 180 nucleotides, or about 320 to about 360 nucleotides, or about 430 to about 480 nucleotides.
[0116] Cell-free nucleic acids can be isolated from body fluids by a partitioning process, in which cell-free nucleic acids found in solution are separated from intact cells and other insoluble components of the body fluid. Partitioning can include techniques such as centrifugation or filtration. Alternatively, cells in the body fluid can be lysed, and the cell-free and cellular nucleic acids can be processed together. Generally, after the addition of buffer and washing steps, the cell-free nucleic acids can be precipitated with alcohol. Additional cleanup steps, such as silica-based columns, can be used to remove contaminants or salts. For example, nonspecific bulk carrier nucleic acids can be added to the entire reaction to optimize certain aspects of the procedure, such as yield.
[0117] After such processing, the sample may contain various forms of nucleic acids, including double-stranded DNA, single-stranded DNA, and single-stranded RNA. If necessary, single-stranded DNA and RNA can be converted to double-stranded form so that they can be included in subsequent processing and analysis steps.
[0118] amplification The sample nucleic acid flanked by the adaptor can be amplified by PCR and other amplification methods, typically primed by a primer that binds to the primer-binding site in the adaptor adjacent to the DNA molecule being amplified. The amplification method can involve cycles of extension, denaturation, and annealing resulting from thermal cycling, or can be isothermal, as in the case of transcription-mediated amplification. Other amplification methods include ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and self-sustaining sequence-based replication.
[0119] Barcodes can be introduced into nucleic acid molecules using conventional nucleic acid amplification methods by applying one or more amplifications. Amplification can be performed in one or more reaction mixtures. Molecular tags and sample indexes / tags can be introduced simultaneously or in any order. Molecular tags and sample indexes / tags can be introduced before and / or after sequence capture. In some cases, only molecular tags are introduced before probe capture, and sample indexes / tags are introduced after sequence capture. In some cases, both molecular tags and sample indexes / tags are introduced before probe capture. In some cases, sample indexes / tags are introduced after sequence capture. Sequence capture typically involves introducing a single-stranded nucleic acid molecule complementary to a target sequence, e.g., a coding sequence in a genomic region, where mutations in such regions are associated with cancer types. Typically, amplification generates multiple non-uniquely or uniquely tagged nucleic acid amplicons with molecular tags and sample indexes / tags ranging in size from 200 nt to 700 nt, 250 nt to 350 nt, or 320 nt to 550 nt. In some embodiments, the amplicons are approximately 300 nt in size. In some embodiments, the amplicon has a size of about 500 nt.
[0120] Barcode Barcodes can be incorporated into or otherwise joined to adapters by chemical synthesis, ligation, overlap extension PCR, among other methods. In general, assignment of unique or non-unique barcodes in reactions follows the methods and systems described by U.S. Patent Application Nos. 20010053519, 20110160078, and U.S. Patent Nos. 6,582,908, 7,537,898, and 9,598,731.
[0121] Tags can be linked to sample nucleic acids randomly or non-randomly. In some cases, they are introduced into microwells in a predicted ratio of identifiers (i.e., barcode combinations). A set of barcodes can be unique, for example, all barcodes have different nucleotide sequences. A set of barcodes can be non-unique, that is, some barcodes have the same nucleotide sequence and some barcodes have different nucleotide sequences. For example, identifiers can be loaded so that more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000, or 1,000,000,000 identifiers are loaded per genome sample. In some cases, identifiers may be loaded such that fewer than 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 or 1,000,000,000 identifiers are loaded per genomic sample. In some cases, the average number of identifiers loaded per sample genome is less than or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 or 1,000,000,000 identifiers per genome sample.
[0122] A preferred format uses 20-50 x 20-50 tags, i.e., 20-50 different tags ligated to both ends of a target molecule, generating 400-2500 tag combinations. Such a number of tags is sufficient so that different molecules with the same start and stop points have a high probability (e.g., at least 94%, 99.5%, 99.99%, 99.999%) of receiving different combinations of tags.
[0123] In some cases, the identifier may be a predetermined or random or semi-random sequence oligonucleotide. In other cases, multiple barcodes may be used, and the barcodes are not necessarily unique to each other among the multiple barcodes. In this example, the barcode may be attached (e.g., by ligation or PCR amplification) to individual molecules such that the combination of the barcode and the sequence to which it may be attached creates a unique sequence that can be individually tracked. As described herein, detection of a non-uniquely tagged barcode in combination with the beginning (start) and / or ending (stop) genomic coordinates of a given sequenced sample molecule (i.e., excluding sequence information derived from barcodes, adapters, etc.) can enable the assignment of a unique identity to a particular molecule. The length or number of base pairs of an individual sequenced sample molecule (i.e., excluding sequence information corresponding to barcodes, adapters, etc.) can also be used to assign a unique identity to such a molecule. As described herein, fragments from a single strand of nucleic acid that have been assigned a unique identity can thereby enable subsequent identification of fragments from the parental and / or complementary strands.
[0124] Sequencing Pipeline The adaptor-flanked sample nucleic acids, with or without prior amplification, can be subjected to sequencing, for example, by one or more sequencing devices 107. Sequencing methods include, for example, Sanger sequencing, high-throughput sequencing, pyrosequencing, sequencing-by-synthesis, single molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing-by-ligation, sequencing-by-hybridization, RNA-Seq (Illumina), digital gene expression (Helicos), next-generation sequencing, single molecule sequencing-by-synthesis (SMSS) (Helicos), massively parallel sequencing, clonal single molecule arrays (Solexa), shotgun sequencing, Ion Torrent, Oxford Nanopore, Roche Genia, Maxim-Gilbert sequencing, primer walking, sequencing using PacBio, SOLiD, Ion Torrent, or Nanopore platforms. The sequencing reactions can be performed in a variety of sample processing units, which can be multiple lanes, multiple channels, multiple wells, or other means of processing multiple sample sets substantially simultaneously. The sample processing units can also include multiple sample chambers to allow for the processing of multiple runs simultaneously.
[0125] The sequencing reaction can be performed on one or more fragment types known to contain markers for other cancer diseases. The sequencing reaction can also be performed on any nucleic acid fragment present in the sample. The sequence reaction can provide at least 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9% or 100% sequencing of a given genome. In other cases, the sequence reaction can provide less than 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9% or 100% sequencing of a given genome.
[0126] Simultaneous sequencing reaction can be carried out using multiplex sequencing.In some cases, cell-free polynucleotides can be sequenced with at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions.In other cases, cell-free polynucleotides can be sequenced with less than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions.Sequencing reactions can be carried out consecutively or simultaneously.Subsequent data analysis can be carried out on all or part of sequencing reactions. In some cases, data analysis may be performed on at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. In other cases, data analysis may be performed on fewer than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. An exemplary read depth is 1000 to 50,000 reads per locus (base).
[0127] Sequence analysis pipeline Nucleotide variations in sequenced nucleic acids can be determined by comparing the sequenced nucleic acids with a reference sequence. The reference sequence is often a known sequence, such as a known whole genome sequence or partial genome sequence from a subject, or a whole genome sequence from a human subject. The reference sequence can be hG19. The sequenced nucleic acid can represent a sequence determined directly for a nucleic acid in a sample, as described above, or a consensus sequence of an amplification product of such a nucleic acid. Comparison can be performed at one or more designated positions on the reference sequence. A subset of sequenced nucleic acids containing positions corresponding to designated positions in the reference sequence when the respective sequences are maximally aligned can be identified. Within such a subset, it can be determined which sequenced nucleic acids, if any, contain a nucleotide variation at the designated position, and optionally which nucleic acids, if any, contain a reference nucleotide (i.e., the same as the reference sequence). If the number of sequenced nucleic acids in the subset containing a nucleotide variant exceeds a threshold, the variant nucleotide can be called at the designated position. The threshold value can be a simple number, such as at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 sequenced nucleic acids in the subset containing the nucleotide variant, or, among other possibilities, a ratio of at least 0.5, 1, 2, 3, 4, 5, 10, 15, or 20 sequenced nucleic acids in the subset containing the nucleotide variant. The comparison can be repeated for any designated position of interest in the reference sequence. Sometimes, the comparison can be performed for designated positions occupying at least 20, 100, 200, or 300 consecutive positions on the reference sequence, for example, 20 to 500 or 50 to 300 consecutive positions.
[0128] The methods can also be used to diagnose the presence or absence of a condition, particularly cancer, in a subject, characterize the condition (e.g., staging the cancer or determining the heterogeneity of the cancer), monitor the response to treatment of the condition, and influence the prognostic risk of developing the condition or the subsequent course of the condition.
[0129] This method can be used to detect various cancers.Cancer cells, as most cells, can be characterized by the turnover rate at which old cells die and are replaced by newer cells.In general, dead cells in contact with the vasculature of a given subject can release DNA or DNA fragments into the bloodstream.This also applies to cancer cells in various stages of disease.Cancer cells can also be characterized by various genetic abnormalities, such as copy number variations and rare mutations, depending on the stage of the disease.This phenomenon can be used to detect the presence or absence of cancer in individuals using the methods and systems described herein.
[0130] The types and number of cancers that can be detected include blood cancer, brain cancer, lung cancer, skin cancer, nose cancer, throat cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, skin cancer, intestinal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, solid tumors, heterogeneous tumors, homogeneous tumors, etc.
[0131] Cancer can be detected from genetic variations including mutations, rare mutations, indels, copy number variations, transversions, translocations, inversions, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structural changes, gene fusions, chromosomal fusions, gene truncations, gene amplifications, gene duplications, chromosomal lesions, DNA lesions, abnormal changes in nucleic acid chemical modifications, and abnormal changes in epigenetic patterns.
[0132] Genetic data can also be used to characterize specific forms of cancer. Cancers are often heterogeneous in both composition and stage. Genetic profile data can enable characterization of specific subtypes of cancer, which can be important in diagnosing or treating that specific subtype. This information can also provide clues to the subject or practitioner regarding the prognosis of a particular type of cancer, allowing either the subject or practitioner to adapt treatment options as the disease progresses. Some cancers progress and become more aggressive and genetically unstable. Other cancers may remain benign, inactive, or dormant. The systems and methods of the present disclosure can be useful in determining disease progression.
[0133] This analysis is also useful for determining the effectiveness of a particular treatment option. Because more cancers may die and shed DNA, successful treatment may increase the amount of copy number variations or rare mutations detected in the subject's blood. In other cases, this may not occur. In another example, a particular treatment option may be correlated with the cancer's genetic profile over time. This correlation may be useful for selecting a treatment. Furthermore, if a cancer is observed to be in remission after treatment, this method can be used to monitor residual disease or disease recurrence.
[0134] The present method can also be used to detect genetic mutations for conditions other than cancer. Immune cells, such as B cells, can undergo rapid clonal expansion in the presence of certain diseases. Clonal expansion can be monitored using copy number variation detection to monitor specific immune states. In this example, copy number variation analysis can be performed over time to generate a profile of how a particular disease may progress. Copy number variation or even rare mutation detection can be used to determine how pathogen populations are changing over the course of infection. This can be particularly important during chronic infections, such as HIV / AIDS or hepatitis infections, where viruses can change life cycle states and / or mutate to more virulent forms over the course of infection. The methods of the present invention can be used to determine or profile the host's rejection activity, as immune cells attempt to destroy transplant tissue and monitor the status of the transplant, as well as alter the course of rejection treatment or prevention.
[0135] Furthermore, the disclosed methods can be used to characterize heterogeneity of abnormal conditions in a subject, the method comprising generating a genetic profile of extracellular polynucleotides in the subject, the genetic profile comprising multiple data obtained from copy number variation and rare mutation analysis. In some cases, including but not limited to cancer, diseases can be heterogeneous. Disease cells may not be identical. In the example of cancer, some tumors are known to contain different types of tumor cells, some cells at different stages of cancer. In other examples, heterogeneity can include multiple disease foci. Again, in the example of cancer, multiple tumor foci may be present, perhaps one or more foci being the result of metastasis spreading from the primary site.
[0136] The method can be used to generate or profile a fingerprint or data set that is the sum of genetic information from different cells in a heterogeneous disease, which data set can include copy number variation and rare mutation analysis, either alone or in combination.
[0137] The methods can be used to diagnose, prognose, monitor, or observe cancer or other diseases of fetal origin, i.e., these methodologies can be used to diagnose, prognose, monitor, or observe cancer or other diseases in pregnant subjects, fetal subjects, where DNA and other polynucleotides may co-circulate with maternal molecules.
[0138] Exemplary Precision Procedures and Applications The precise diagnosis provided by the computer system 700 can result in a precise treatment plan that can be identified by the computer system 700 (and / or managed by a medical professional). For example, in lung cancer and other diseases, the goal may be to ensure that no superior treatment options exist given the presence of a given mutation. For example, EGFR (L858R, exon 19 deletion), BRAF V600E, ALK, and ROS1 fusions can be treated with targeted therapies that may be more suitable than platinum therapy and chemotherapy. While these are examples of primary drivers, other targetable drivers exist, such as MET exon 14 skipping. In another example, in the case of colon cancer, the goal may be to avoid ineffective treatments. Chemotherapy with FOLFIRI or irinotecan regimens may be supplemented with cetuximab or panitumumab if KRAS or NRAS are wild-type. Therefore, confidence in whether KRAS and NRAS are wild-type will increase confidence that the addition of cetuximab or panitumumab is the correct treatment option and further testing may not be required. The biological explanation for this is that cetuximab or panitumumab targets EGFR and inhibits its activity. Because RAS (K / NRAS) is downstream of EGFR, if RAS is activated, inhibition of EGFR will have minimal or no effect, making cetuximab or panitumumab treatment inappropriate.
[0139] The mutants analyzed by the disclosed methods and systems may be loss-of-function mutants (such as ATM). For example, DNA damage repair (DDR) is a cellular process that functions to maintain genome integrity or stability. Defects or defects in a given DDR mechanism may lead to tumorigenesis or other diseases and may be used to identify test subjects or patients who may benefit from a given targeted therapy. Homologous recombination repair deficiency (HRD), as an example, is a cellular phenotype that may make a patient a candidate for the administration of a therapeutic agent such as a poly ADP-ribose polymerase (PARP) inhibitor. In certain embodiments, a treatment may be administered to a subject that includes at least one PARP inhibitor, and the mutant has been identified as being of tumor or non-tumor origin using the methods and systems described herein. In certain embodiments, the PARP inhibitor may include, among others, olaparib, talazoparib, rucaparib, or niraparib (trade name ZEJULA). In some embodiments, the treatment includes at least one base excision repair (BER) inhibitor. For example, olaparib may suppress BER. In certain embodiments, administration of therapy to a subject may be discontinued based on a determination that the subject has a variant of tumor or non-tumor origin using the methods and systems described herein.
[0140] Non-tumor variants can affect the determination of tumor mutational burden (TMB) scores, resulting in artificially high scores if not excluded or filtered from the TMB determination. TMB scores are typically used to predict whether a patient will respond to immunotherapy treatment. Thus, the methods and systems provided herein can be used to distinguish between variants of tumor or non-tumor origin as part of the TMB calculation, such as those described in PCT / US2019 / 042882, incorporated herein by reference. In another aspect, the present disclosure provides a method for classifying a subject as a candidate for immunotherapy by determining whether the subject has variants of tumor or non-tumor origin. In certain embodiments, the methods of the present disclosure include administering one or more immunotherapies to a subject based on determining whether variants are of tumor or non-tumor origin using the methods or systems disclosed herein, alone or in combination with a method for determining a TMB score. In some embodiments, the immunotherapy includes at least one checkpoint inhibitor antibody. In some embodiments, the immunotherapy comprises an antibody against PD-1, PD-2, PD-L1, PD-L2, CTLA-40, OX40, B7.1, B7He, LAG3, CD137, KIR, CCR5, CD27, or CD40. In some embodiments, the immunotherapy comprises administration of a pro-inflammatory cytokine to at least one tumor type. In some embodiments, the immunotherapy comprises administration of T cells to at least one tumor type. In some embodiments, the subject is administered a combination therapy (e.g., immunotherapy + PARPi + chemotherapy, etc.), among many other treatments further exemplified herein or known to one of skill in the art.
[0141] The methods and systems provided herein can be used to evaluate mutations for prognostic value regarding survival or response to treatment. For example, TP53 mutations can be evaluated for prognostic and predictive value of treatment with ALK inhibitors. Determining the tumor / non-tumor origin of the variants analyzed herein can also be used to register subjects for selected therapies (e.g., TP53 drugs). Another application of the methods and systems described herein can be for analyzing less studied mutations (e.g., FGFR2 mutations for FGFR inhibitors, or ERBB2 for ERBB2 inhibitors), where distinguishing between tumor and non-tumor origin of the variant can provide confidence that the variant is tumor-derived or not. In certain embodiments, the methods and systems described herein can be used to monitor molecular response by tracking tumor-only variants to determine the dynamics of the variant over time.
[0142] As additional treatments are developed for various diseases, the interpretation of negative predictions will become increasingly complex but important in the design of precision therapies. [Example]
[0143] Example Example 1: Variant calls were obtained from over 250,000 plasma samples, including healthy donors and patients with early- and late-stage cancers, sequenced on the Guardant360™, GuardantREVEAL™, GuardantOMNI™, and GuardantInfinity™ liquid biopsy panels, as well as public tissue datasets. Models were trained on paired plasma and WBC datasets and optimized with 10-fold cross-validation to generate non-tumor and tumor variant classifiers. To validate these calls, an independent cohort of 72 paired plasma and WBC advanced cancer samples was genotyped with the GuardantInfinity™ assay. A cohort of 76 healthy donor samples genotyped with the GuardantOMNI assay was also evaluated.
[0144] Example 2: An ensemble model was trained on a database of over 250,000 plasma samples, including healthy donors and patients with early- and late-stage cancers sequenced on the Guardant360™, GuardantREVEAL™, and GuardantOMNI™ liquid biopsy panels, as well as public tissue datasets. To generate non-tumor and tumor variant classifiers, the model was optimized using 5-fold cross-validation and hyperparameter tuning. To validate these calls, 116 paired plasma and WBC advanced cancer clinical samples were selected for their high prevalence of putative CH variants and sequenced and genotyped using an in-house bioinformatics pipeline. In the validation cohort, cfDNA variants were determined to be of non-tumor or CH origin if they had sufficient molecular support in WBCs, and cfDNA variants with no support in WBCs above 0.6% (the detection limit in gDNA) were determined to be of tumor origin.
[0145] Example 3: The validation cohort consisted of 2,150 somatic SNVs and indels, of which 956 were confirmed in WBC and 1,194 were confirmed in plasma only. Half of the confirmed CH variants (48%, 458 / 956) were present in known CH genes (e.g., DNMT3A, TET2, PPM1D), while the other half were present in genes such as TP53, ATM, NOTCH4, FAT1, and SRSF2. No clinically actionable variants were identified in WBC. 624 somatic variants were predicted as non-tumor or CH, and of these, 515 / 624 were correctly identified as CH for a positive predictive value (PPV) of 83%. Of all CH variants confirmed in WBC, 54% (553 / 956) had a CH or non-CH predictive value, and the CH predictive value had a positive percent agreement (PPA) of 91% (515 / 553) with WBC. The remaining variants without CH predictions (403 / 956) had low or no prevalence across the dataset and occurred primarily in LRP1B, TET2, TP53, and KMT2D. Nearly half (67%, n=109) of the CH predictions that did not occur in WBC occurred in CH genes. For non-CH gene variants, 16% of false-positive predictions occurred in six variants across four genes (ACVR2A, RNF43, B2M, and FLT3).
[0146] Example 4: We present a plasma-only method with high PPA and PPV by WBC genotyping to classify non-tumor CH variants in cfDNA. Further investigations are underway to improve the sensitivity of annotating rare CH variants. Accurate CH identification is important for treatment selection across targeted therapies, especially loss of function variants in DNA repair genes that may confer sensitivity to PARPi or ATRi therapy.
[0147] All patent applications, websites, other publications, accession numbers, etc. cited above or below are incorporated by reference in their entirety for all purposes to the same extent as if each individual item were specifically and individually indicated to be incorporated by reference. Where different versions of a sequence are associated with an accession number at different times, the version associated with the accession number as of the effective filing date of this application is meant. The effective filing date means the earlier of the actual filing date or the filing date of the priority application, if applicable, that references an accession number. Similarly, where different versions of a publication, website, etc. are published at different times, the version last published as of the effective filing date of the application is meant unless otherwise specified. Any feature, step, element, embodiment, or aspect of the present disclosure can be used in combination with any other, unless otherwise specified. While the present disclosure has been described in some detail by way of illustration and example for purposes of clarity and understanding, it will be apparent that certain changes and modifications can be practiced within the scope of the appended claims.
Claims
1. 1. A method for distinguishing between tumor and non-tumor origin nucleic acid variants in a cell-free nucleic acid (cfNA) sample obtained from a test subject using, at least in part, a computer, said method comprising: generating or providing by said computer at least one tumor variant dataset comprising a population of reference tumor-associated gene variants, said tumor variant dataset comprising observation frequency data among reference samples comprising a reference plasma-only sample and / or a reference leukocyte sample for one or more tumor-associated gene variants in said population of reference tumor-associated gene variants, said reference samples being obtained from a single reference subject and / or different reference subjects having the same cancer type; determining by the computer one or more ratios of the observed frequency data between the reference samples for one or more tumor-associated gene variants in the population of reference tumor-associated gene variants to generate at least one MAF distribution and / or relative prevalence data set; generating by said computer at least one set of probabilities of non-tumor origin from said MAF variance and / or relative prevalence dataset; using the set of probabilities of non-tumor origin to distinguish nucleic acid variants detected in the cfNA sample obtained from the test subject as being tumor-origin nucleic acid variants or non-tumor-origin nucleic acid variants; A method comprising:
2. 1. A method for distinguishing between tumor and non-tumor origin nucleic acid variants in a cell-free nucleic acid (cfNA) sample obtained from a test subject using, at least in part, a computer, said method comprising: determining by the computer a relative prevalence of one or more tumor-associated gene variants observed in one or more reference plasma-only samples compared to one or more reference white blood cell samples to generate at least one relative prevalence dataset; generating by the computer at least one set of probabilities of non-tumor origin from the relative prevalence dataset; using the set of probabilities of non-tumor origin to distinguish nucleic acid variants detected in the cfNA sample obtained from the test subject as being tumor-origin nucleic acid variants or non-tumor-origin nucleic acid variants; A method comprising:
3. 1. A method for distinguishing between tumor and non-tumor origin nucleic acid variants in a cell-free nucleic acid (cfNA) sample obtained from a test subject using, at least in part, a computer, said method comprising: determining by said computer the variation in mutant allele fraction (MAF) values and / or at least one statistic associated therewith for at least two different time points for each of one or more tumor-associated and / or non-tumor-associated genetic variants to generate at least one relative prevalence data set; generating by the computer at least one set of probabilities of non-tumor origin from the relative prevalence dataset; using the set of probabilities of non-tumor origin to distinguish nucleic acid variants detected in the cfNA sample obtained from the test subject as being tumor-origin nucleic acid variants or non-tumor-origin nucleic acid variants; A method comprising:
4. 1. A method for distinguishing between tumor and non-tumor origin nucleic acid variants in a cell-free nucleic acid (cfNA) sample obtained from a test subject using, at least in part, a computer, said method comprising: classifying, by the computer, at least the first nucleic acid variant detected in the cfNA sample obtained from the test subject as a tumor-origin nucleic acid variant if the prevalence of a first nucleic acid variant detected in the cfNA sample is less than a threshold probability from a set of probabilities of non-tumor origin; and classifying, by the computer, at least the second nucleic acid variant detected in the cfNA sample obtained from the test subject as a non-tumor-origin nucleic acid variant if the prevalence of a second nucleic acid variant detected in the cfNA sample is greater than a threshold probability from a set of probabilities of non-tumor origin, thereby distinguishing between the tumor and non-tumor-origin nucleic acid variants in the cfNA sample obtained from the test subject, wherein the set of probabilities of non-tumor origin is generating or providing by said computer at least one tumor variant dataset comprising a population of reference tumor-associated gene variants, said tumor variant dataset comprising observation frequency data among reference samples comprising a reference plasma-only fluid sample and / or a reference leukocyte sample for one or more tumor-associated gene variants in said population of reference tumor-associated gene variants, said reference samples being obtained from a single reference subject and / or different reference subjects having the same cancer type; determining by the computer one or more ratios of the observed frequency data between the reference samples for one or more tumor-associated gene variants in the population of reference tumor-associated gene variants to generate at least one relative prevalence data set; generating by the computer a set of probabilities of non-tumor origin from the relative prevalence dataset; Generated by the method.
5. 1. A method for generating, at least in part, a computational classifier for distinguishing nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as being tumor-originating nucleic acid variants or non-tumor-originating nucleic acid variants, the method comprising: generating or providing by said computer at least one tumor variant dataset comprising a population of reference tumor-associated gene variants, said tumor variant dataset comprising observation frequency data among reference samples comprising a reference plasma-only sample and / or a reference leukocyte sample for one or more tumor-associated gene variants in said population of reference tumor-associated gene variants, said reference samples being obtained from a single reference subject and / or different reference subjects having the same cancer type; determining by the computer one or more ratios of the observed frequency data between the reference samples for one or more tumor-associated gene variants in the population of reference tumor-associated gene variants to generate at least one relative prevalence data set; applying, by the computer, at least one machine learning model to the relative prevalence dataset to generate at least one set of probabilities of non-tumor origin, thereby generating the classifier that distinguishes the nucleic acid variants detected in the cfNA sample as being of tumor origin or non-tumor origin; A method comprising:
6. 1. A method for distinguishing between tumor and non-tumor origin nucleic acid variants in a cell-free nucleic acid (cfNA) sample obtained from a test subject having a cancer type, using at least in part a computer, said method comprising: determining by the computer a prevalence of one or more genetic variants observed in the cfNA sample to generate a test subject prevalence dataset; comparing, by the computer, the prevalence of one or more genetic variants in the test subject prevalence dataset to the prevalence of the genetic variants observed in a reference cfNA sample obtained from a reference subject having the cancer type; classifying by the computer a given genetic variant in the test subject prevalence dataset as a non-tumor-origin nucleic acid variant if the prevalence of the given genetic variant in the test subject prevalence dataset is below a predetermined threshold associated with the given genetic variant in the reference cfNA sample obtained from a reference subject having the cancer type, thereby distinguishing between the tumor and non-tumor-origin nucleic acid variants in the cfNA sample obtained from the test subject having the cancer type; A method comprising:
7. 1. A method for distinguishing between tumor and non-tumor origin nucleic acid variants in a cell-free nucleic acid (cfNA) sample obtained from a test subject using, at least in part, a computer, said method comprising: determining by the computer a prevalence of one or more genetic variants observed in the cfNA sample to generate a test subject prevalence dataset; comparing, by the computer, the prevalence of one or more genetic variants in the test subject prevalence dataset with the prevalence of the genetic variants observed in reference cfNA samples obtained from reference subjects with leukemia, lymphoma, and / or hematological malignancies; classifying by the computer the given genetic variant in the test subject prevalence dataset as a non-tumor-origin nucleic acid variant if the prevalence of the given genetic variant in the test subject prevalence dataset is above a predetermined threshold associated with the given genetic variant in the reference cfNA sample obtained from a reference subject having the leukemia, the lymphoma, and / or the hematological malignancy, thereby distinguishing between the tumor and non-tumor-origin nucleic acid variants in the cfNA sample obtained from the test subject having the leukemia, the lymphoma, and / or the hematological malignancy; A method comprising:
8. 10. The method of any one of the preceding claims, comprising identifying genetic variants present in the cfNA sample from sequencing reads derived from cfNA molecules in the cfNA sample.
9. 10. The method of any one of the preceding claims, wherein the sequencing reads are obtained from targeted segments of the cfNA molecules in the cfNA sample.
10. 10. The method of any one of the preceding claims, wherein said population of reference tumor-associated gene variants is obtained from said reference sample.
11. 10. The method of any one of the preceding claims, comprising randomly splitting the tumor variant dataset into a training dataset and a testing dataset.
12. 10. The method of any one of the preceding claims, wherein the training dataset comprises about 80% of the tumor variant dataset and the testing dataset comprises about 20% of the tumor variant dataset.
13. 10. The method of any one of the preceding claims, wherein said tumor variant dataset comprises observed frequency data among reference samples of a given cancer type for one or more tumor-associated gene variants in said population of reference tumor-associated gene variants.
14. 10. The method of any one of the preceding claims, comprising training a machine learning model using at least a portion of the population of tumor-associated gene variants to generate a trained machine learning model, wherein the tumor-origin nucleic acid variants and non-tumor-origin nucleic acid variants detected in the cfNA sample obtained from the test subject are distinguished from each other using the trained machine learning model.
15. 10. The method of any one of the preceding claims, wherein the machine learning model is trained using one or more of logistic regression, probit regression, decision tree, random forest, gradient boosting, support vector machine, K-nearest neighbors, and neural network.
16. 10. The method of any one of the preceding claims, comprising using a threshold of at least about the 30th percentile probability for a given genetic variant as a cutoff for classification.
17. 10. The method of any one of the preceding claims, comprising performing a logistic regression on at least one of said ratios to obtain a given probability of non-tumor origin.
18. 10. The method of any one of the preceding claims, wherein said tumor variant dataset comprises mutant allele fraction data observed among reference samples for one or more tumor-associated gene variants in said population of reference tumor-associated gene variants.
19. 10. The method of any one of the preceding claims, comprising normalizing the tumor variant dataset using one or more data normalization techniques.
20. 10. The method of any one of the preceding claims, wherein the data normalization technique comprises min-max normalization and / or z-score normalization.
21. 10. The method of any one of the preceding claims, wherein the reference non-body fluid sample comprises a reference tumor tissue sample and / or a reference leukocyte sample.
22. 10. The method of any one of the preceding claims, wherein a ratio of the observed frequency data of the given genetic variant in the reference plasma-only sample to the observed frequency data of the given genetic variant in the reference white blood cell sample that is greater than one (1.0) indicates that the given genetic variant is likely to be a nucleic acid variant of non-tumor origin.
23. 21. The method of any one of claims 1 to 20, wherein a ratio of observed frequency data for a given genetic variant in the plasma-only fluid sample to observed frequency data for the given genetic variant in a reference non-body fluid sample that is less than one (1.0) indicates that the given genetic variant is likely to be a nucleic acid variant of non-tumor origin.
24. 10. The method of any one of the preceding claims, wherein the set of probabilities of non-tumor origin comprises at least one set of probabilities of clonal hematopoietic origin.
25. 10. The method of any one of the preceding claims, comprising obtaining the cfNA sample from the test subject.
26. 10. The method of any one of the preceding claims, comprising selecting one or more therapies for treating said cancer type if one or more tumor-originating nucleic acid variants associated with said cancer type are detected in said cfNA sample obtained from said test subject.
27. 10. The method of any one of the preceding claims, comprising administering one or more therapies to the test subject to treat a cancer type if one or more tumor-origin nuclear variants associated with the cancer type are detected in the cfNA sample obtained from the test subject.
28. The cancer type is biliary tract cancer cancer), bladder cancer, transitional cell cancer, urothelial cancer, brain cancer, glioma, astrocytoma, breast cancer, dysplastic cancer, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal cancer, colon cancer, hereditary non-polyposis colorectal cancer, colorectal adenocarcinoma, gastrointestinal stromal tumor (GIST), endometrial cancer, endometrial stromal sarcoma, esophageal cancer, esophageal squamous cell carcinoma, food tract adenocarcinoma, ocular melanoma, uveal melanoma, gallbladder cancer, gallbladder adenocarcinoma, renal cell carcinoma, clear cell renal cell carcinoma, transitional cell carcinoma, urothelial carcinoma, Wilms tumor, leukemia, acute lymphocytic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic (CLL), chronic myeloid (CML), chronic myelomonocytic (CMML), liver cancer (liver) cancer), liver carcinoma carcinoma), hepatocellular carcinoma, hepatocellular carcinoma, cholangiocarcinoma, hepatoblastoma, lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B-cell lymphoma, non-Hodgkin's lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, T-cell lymphoma, non-Hodgkin's lymphoma, precursor T-lymphoblastic lymphoma / leukemia, peripheral T-cell lymphoma, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oropharyngeal carcinoma, oral squamous cell carcinoma, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic ductal adenocarcinoma, pseudopapillary neoplasm, acinar cell carcinoma, prostate cancer, prostate adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine cancer, stomach cancer, gastric carcinoma 10. The method of any one of the preceding claims, wherein the cancer is selected from the group consisting of gastrointestinal stromal tumor (GIST), uterine cancer, and uterine sarcoma.
29. 10. The method of any one of the preceding claims, wherein the reference tumor-associated genetic variants are selected from the group consisting of single nucleotide variants (SNVs), insertions or deletions (indels), copy number variants (CNVs), fusions, transversions, translocations, frameshifts, duplications, repeat expansions, and epigenetic variants.
30. 10. The method of any one of the preceding claims, wherein the reference sample comprises at least about 25, at least about 50, at least about 100, at least about 200, at least about 300, at least about 400, at least about 500, at least about 600, at least about 700, at least about 800, at least about 900, at least about 1,000, at least about 5,000, at least about 10,000, at least about 15,000, at least about 20,000, at least about 25,000, at least about 30,000, or more body fluid and / or non-body fluid samples.
31. 10. The method of any one of the preceding claims, wherein the cfNA sample comprises cell-free deoxyribonucleic acid (cfDNA).
32. 10. The method of any one of the preceding claims, wherein the cfRNA sample comprises cell-free ribonucleic acid (cfRNA).
33. 10. The method of any one of the preceding claims, wherein the test subject is a mammalian subject.
34. 10. The method of any one of the preceding claims, wherein the test subject is a human subject.
35. 10. The method of any one of the preceding claims, wherein the reference body fluid sample comprises a plasma sample.
36. 10. The method of any one of the preceding claims, wherein the reference body fluid sample comprises a serum sample.
37. 10. The method of any one of the preceding claims, wherein the reference non-body fluid sample comprises a cell sample.
38. 10. The method of any one of the preceding claims, wherein the reference non-body fluid sample comprises a tissue sample.
39. The method for distinguishing between tumor and non-tumor origin nucleic acid variants in a cell-free nucleic acid (cfNA) sample comprises, at least in part, (i) the uniformity of the prevalence of the nucleic acid variant across cancer types; (ii) the variation over time in the mutant allele fraction (MAF) of said nucleic acid variant; and / or (iii) the prevalence of said nucleic acid variants in hematological cancers, such as leukemia, lymphoma, and / or hematological malignancies.
10. The method of any one of the preceding claims, based on
40. A system including a controller that includes or has access to a computer-readable medium containing non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least (a) generating or providing at least one tumor variant dataset comprising a population of reference tumor-associated gene variants, said tumor variant dataset comprising observation frequency data among reference samples comprising a reference plasma-only sample and / or a reference leukocyte sample for one or more tumor-associated gene variants in said population of reference tumor-associated gene variants, said reference samples being obtained from a single reference subject and / or different reference subjects having the same cancer type; (b) determining one or more ratios of the observed frequency data between the reference samples for one or more tumor-associated gene variants in the population of reference tumor-associated gene variants to generate at least one relative prevalence dataset; (c) applying at least one machine learning model to said relative prevalence dataset to generate at least one set of probabilities of non-tumor origin to create a classifier that distinguishes said nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as being a tumor-origin nucleic acid variant or a non-tumor-origin nucleic acid variant; A system that implements the above.
41. A system including a controller that includes or has access to a computer-readable medium containing non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least (a) determining the relative prevalence of one or more tumor-associated gene variants observed in one or more reference plasma-only samples compared to one or more reference white blood cell samples to generate at least one relative prevalence dataset; (b) generating at least one set of probabilities of non-tumor origin from the relative prevalence dataset to generate a classifier that distinguishes nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as being tumor-origin nucleic acid variants or non-tumor-origin nucleic acid variants; A system that implements the above.
42. A system including a controller that includes or has access to a computer-readable medium containing non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least (a) determining by a computer the variance in mutant allele fraction (MAF) values and / or at least one statistic associated therewith for at least two different time points for each of one or more tumor-associated and / or non-tumor-associated genetic variants to generate at least one MAF variance and / or relative prevalence data set; (b) generating at least one set of probabilities of non-tumor origin from said MAF variance and / or relative prevalence dataset to generate a classifier that distinguishes nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as being tumor-origin nucleic acid variants or non-tumor-origin nucleic acid variants; A system that implements the above.
43. 10. The system of any one of the preceding claims, comprising a nucleic acid sequencer operably connected to the controller, the nucleic acid sequencer configured to provide sequencing reads derived from cfNA molecules in the cfNA sample.
44. 10. The system of any one of the preceding claims, wherein the nucleic acid sequencer or another system component is configured to group sequence reads generated by the nucleic acid sequencer into families of sequence reads, each family comprising sequence reads generated from a given cfNA molecule in the cfNA sample.
45. 10. The system of any one of the preceding claims, comprising a database operatively connected to said controller, said database comprising one or more treatments indexed against said tumor-originating nucleic acid variants.
46. 10. The system of any one of the preceding claims, comprising a sample preparation component operably connected to the controller, the sample preparation component configured to prepare the cfNA molecules in the cfNA sample to be sequenced by the nucleic acid sequencer.
47. 10. The system of any one of the preceding claims, comprising a nucleic acid amplification component operably connected to the controller, the nucleic acid amplification component configured to amplify at least targeted segments of the cfNA molecules in the cfNA sample.
48. 10. The system of any one of the preceding claims, comprising a material transfer component operably connected to the controller, the material transfer component configured to transfer one or more materials between at least the nucleic acid sequencer and the sample preparation component.
49. A computer-readable medium containing non-transitory computer-executable instructions that, when executed by at least one electronic processor, (a) generating or providing at least one tumor variant dataset comprising a population of reference tumor-associated gene variants, said tumor variant dataset comprising observation frequency data among reference samples comprising a reference plasma-only sample and / or a reference leukocyte sample for one or more tumor-associated gene variants in said population of reference tumor-associated gene variants, said reference samples being obtained from a single reference subject and / or different reference subjects having the same cancer type; (b) determining one or more ratios of the observed frequency data between the reference samples for one or more tumor-associated gene variants in the population of reference tumor-associated gene variants to generate at least one relative prevalence dataset; (c) applying at least one machine learning model to said relative prevalence dataset to generate at least one set of probabilities of non-tumor origin to create a classifier that distinguishes said nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as being a tumor-origin nucleic acid variant or a non-tumor-origin nucleic acid variant; 1. A computer-readable medium comprising non-transitory computer-executable instructions for implementing the method of claim 1.
50. A computer-readable medium containing non-transitory computer-executable instructions that, when executed by at least one electronic processor, (a) determining the relative prevalence of one or more tumor-associated gene variants observed in one or more reference body fluid samples compared to one or more reference non-body fluid samples to generate at least one MAF variance and / or relative prevalence dataset; (b) generating at least one set of probabilities of non-tumor origin from said MAF variance and / or relative prevalence dataset to generate a classifier that distinguishes nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as being tumor-origin nucleic acid variants or non-tumor-origin nucleic acid variants; 1. A computer-readable medium comprising non-transitory computer-executable instructions for implementing the method of claim 1.
51. A computer-readable medium containing non-transitory computer-executable instructions that, when executed by at least one electronic processor, (a) determining by a computer the variation in mutant allele fraction (MAF) values and / or at least one statistic associated therewith for at least two different time points for each of one or more tumor-associated and / or non-tumor-associated genetic variants to generate at least one relative prevalence data set; (b) generating at least one set of probabilities of non-tumor origin from the relative prevalence dataset to generate a classifier that distinguishes nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as being tumor-origin nucleic acid variants or non-tumor-origin nucleic acid variants; 1. A computer-readable medium comprising non-transitory computer-executable instructions for implementing the method of claim 1.
52. 10. The system or computer readable medium of any one of the preceding claims, wherein the electronic processor is further configured to at least split the tumor variant dataset into a training dataset and a test dataset.
53. 10. The system or computer-readable medium of any one of the preceding claims, wherein the electronic processor is further configured to: train a machine learning model using at least a portion of the population of tumor-associated gene variants to generate a trained machine learning model; and use the trained machine learning model to distinguish the nucleic acid variants detected in the cfNA sample as being tumor-origin nucleic acid variants or non-tumor-origin nucleic acid variants.
54. 10. The system or computer-readable medium of any one of the preceding claims, wherein the electronic processor is further configured to perform at least a logistic regression on at least one of the ratios to obtain a given probability of non-tumor origin.
55. 10. The system or computer readable medium of any one of the preceding claims, wherein the electronic processor is further configured to at least normalize the tumor variant dataset using one or more data normalization techniques.
56. The system or computer-readable medium of any one of the preceding claims, wherein the electronic processor further performs selecting one or more therapies for treating the cancer type at least if one or more tumor-originating nucleic acid variants associated with the cancer type are detected in the cfNA sample.