Verification of bioinformatics models for classifying non-tumor variations in cell-free DNA liquid biopsy assays

By computer analysis of tumor-related genetic variation data in reference samples, a collection of non-tumor sources of probability was generated, solving the problem of distinguishing tumor from non-tumor variants in cell-free nucleic acid samples, improving the sensitivity and specificity of detection, simplifying the process and reducing costs.

CN120226085APending Publication Date: 2025-06-27GUARDANT HEALTH INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202380080204.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-17
Filing Date
2023-11-16
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively distinguish between tumor-derived nucleic acid variants from non-tumor-derived nucleic acid variants in cell-free nucleic acid samples, and the existing methods require genotyping of white blood cells paired with plasma samples, which is complex and costly.

Method used

By computer-generating or providing a tumor variant data set including a reference tumor-related genetic variant population, the observation frequency data ratio of tumor-related genetic variants between reference samples is determined, and a set of non-tumor-derived probability sets are generated, and the nucleic acid variants detected in cfNA samples are then distinguished from tumor-derived or non-tumor-derived sources.

Benefits of technology

Improved sensitivity to distinguish low VAF (<0.6%) non-tumor variants, achieved accurate assessment of biomarkers in cell-free DNA, simplified the process and reduced costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120226085A_ABST
    Figure CN120226085A_ABST
Patent Text Reader

Abstract

Provided herein are methods of differentiating between tumor-derived nucleic acid variations and non-tumor-derived nucleic acid variations in a cell-free nucleic acid (cfNA) sample. Certain of the methods include generating a tumor variation dataset comprising a population of reference tumor-associated genetic variations, where the tumor variation dataset includes observed frequency data for tumor-associated genetic variations in the population of reference tumor-associated genetic variations in a reference sample, the reference sample comprises a reference plasma-only sample and a reference leukocyte sample; and determining a ratio of observed frequency data for tumor-associated genetic variations in the population of reference tumor-associated genetic variations between the reference samples to produce a relative incidence dataset. Additional methods and related systems and computer readable media are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 384,215, filed on November 17, 2022, and relies on the filing date of this U.S. Provisional Patent Application. The entire disclosure of this U.S. Provisional Patent Application is incorporated herein by reference.

[0003] Background

[0004] Liquid biopsy testing can be used to profile circulating tumor nucleic acids in blood samples from patients, for purposes such as early cancer detection, therapy selection, and monitoring of disease progression and / or minimal residual disease. Circulating plasma cell - free tumor DNA (ctDNA) is small DNA fragments from apoptotic and necrotic tumor cells or from circulating tumor cells (CTCs) that have been introduced into the bloodstream. CtDNA is only a fraction of the cell - free DNA (cfDNA) specifically released from cancer cells, while most cfDNA in a given sample typically originates from normal non - cancerous cells, including normal white blood cells, hematopoietic stem cells (HSCs), or other early blood cell progenitors that undergo apoptosis or necrosis during clonal hematopoiesis. One problem associated with many liquid biopsy tests is distinguishing ctDNA in a patient sample from other cfDNA. Additionally, the presence of clonal hematopoiesis (CH) variants and biological noise due to aging and therapies has the potential to confound biomarker interpretation.

[0005] Currently, comprehensive methods for filtering out non - tumor variants require genotyping of the white blood cell (WBC) fraction of paired plasma samples, which is an expensive and complex workflow. Thus, there remains a need for methods and related aspects for distinguishing tumor - derived nucleic acid variants and non - tumor - derived nucleic acid variants detected in cell - free nucleic acid (cfNA) samples, with particular focus on achieving plasma - only bioinformatics solutions for identifying non - tumor variants for accurate biomarker assessment in cell - free DNA (cfDNA).

[0006] Described herein are bioinformatics models that have improved sensitivity for identifying non - tumor variants with low variant allele frequencies (VAFs) (<0.6%) compared to WBC sequencing. In paired plasma and WBC advanced cancer cohorts, most non - tumor variants are in known clonal hematopoiesis genes and variants of unknown significance. The described analysis platform shows high sensitivity and specificity for distinguishing tumor and non - tumor using only cfDNA from WBCs. Summary of the Invention

[0008] The present disclosure provides methods for distinguishing tumor-derived nucleic acid variants from non-tumor-derived nucleic acid variants in cell-free nucleic acid (cfNA) samples, which improve the sensitivity and specificity of cancer detection assays and guide treatment strategies and other attributes. Also provided are additional methods as well as related systems and computer-readable media.

[0009] In some aspects, the present disclosure provides a method for distinguishing (e.g., discriminating) tumor-derived nucleic acid variants from non-tumor-derived nucleic acid variants in a cell-free nucleic acid (cfNA) sample obtained from a test subject, at least in part using a computer. The method includes generating or providing, by the computer, at least one tumor variant data set including a reference population of tumor-associated genetic variants. The tumor variant data set includes frequency of observance data of one or more tumor-associated genetic variants in the reference population of tumor-associated genetic variants in a reference sample, the reference sample including a reference body fluid sample (e.g., a plasma sample, a serum sample, or a similar sample), including plasma only, and / or a reference non-body fluid sample (e.g., a cell sample, a tissue sample, etc.), including white blood cells. The reference sample is obtained from a single reference subject and / or from different reference subjects having the same cancer type. The method further includes determining, by the computer, one or more ratios of the frequency of observance data of one or more tumor-associated genetic variants in the reference population of tumor-associated genetic variants between the reference samples to produce at least one MAF variance and / or relative prevalence data set; additionally, the method includes generating or providing, by the computer, at least one set of probabilities of non-tumor origin from the relative prevalence data set, and using the set of probabilities of non-tumor origin to classify nucleic acid variants detected in the cfNA sample obtained from the test subject as tumor-derived nucleic acid variants or non-tumor-derived nucleic acid variants. In some embodiments of the methods, systems, computer-readable media, and other aspects of the present disclosure, one or more other features are optionally used in combination with or in place of the ratios of the frequency of observance data. Some of these other features include, for example, consistency of prevalence across cancer types, changes in longitudinal mutant allele fraction (MAF) over time, proportion of blood cancers, and / or similar features.

[0010] In other aspects, the present disclosure provides a method for differentiating, at least in part using a computer, tumor-derived nucleic acid variants and non-tumor-derived nucleic acid variants in a cell-free nucleic acid (cfNA) sample obtained from a test subject. The method includes determining, by the computer, the relative incidence of one or more tumor-associated genetic variants observed in one or more reference plasma-only samples compared to one or more reference white blood cells to generate at least one relative incidence data set; additionally, the method includes generating or providing, by the computer, at least one non-tumor-derived probability set from the relative incidence data set, and using the non-tumor-derived probability set to classify nucleic acid variants detected in a cfNA sample obtained from the test subject as tumor-derived nucleic acid variants or non-tumor-derived nucleic acid variants.

[0011] In some aspects, the present disclosure provides a method for differentiating, at least in part using a computer, tumor-derived nucleic acid variants and non-tumor-derived nucleic acid variants in a cell-free nucleic acid (cfNA) sample obtained from a test subject. The method includes determining, by the computer, the change in the mutant allele fraction (MAF) value and / or at least one statistic associated therewith (e.g., mean, standard deviation, and / or chi-square p-value of the variant MAF over time) of each of one or more tumor-associated genetic variants and / or non-tumor-associated genetic variants observed in one or more reference plasma-only samples compared to one or more reference white blood cells at at least two different time points to generate at least one MAF variance and / or relative incidence data set. Additionally, the method includes generating or providing, by the computer, at least one non-tumor-derived probability set from the MAF variance and / or relative incidence data set, and using the non-tumor-derived probability set to classify nucleic acid variants detected in a cfNA sample obtained from the test subject as tumor-derived nucleic acid variants or non-tumor-derived nucleic acid variants.

[0012] In some aspects, the present disclosure provides a method for differentiating tumor-derived nucleic acid variants from non-tumor-derived nucleic acid variants in a cell-free nucleic acid (cfNA) sample obtained from a test subject, at least in part using a computer. The method includes classifying, by the computer, at least a first nucleic acid variant detected in a cfNA sample obtained from a test subject as a tumor-derived nucleic acid variant when an incidence rate of the first nucleic acid variant detected in the cfNA sample is less than a probability threshold from a non-tumor-derived probability set, and classifying, by the computer, at least a second nucleic acid variant detected in the cfNA sample obtained from the test subject as a non-tumor-derived nucleic acid variant when an incidence rate of the second nucleic acid variant detected in the cfNA sample is greater than the probability threshold from the non-tumor-derived probability set, thereby differentiating tumor-derived nucleic acid variants from non-tumor-derived nucleic acid variants in the cfNA sample obtained from the test subject. The non-tumor-derived probability set is generated by: generating or providing, by the computer, at least one tumor variant data set including a reference tumor-associated genetic variant population, wherein the tumor variant data set includes observed frequency data of one or more tumor-associated genetic variants in the reference tumor-associated genetic variant population in a reference sample, the reference sample including a reference plasma only and reference white blood cells, and wherein the reference sample is obtained from a single reference subject and / or from different reference subjects having the same cancer type; determining, by the computer, one or more ratios of the observed frequency data of one or more tumor-associated genetic variants in the reference tumor-associated genetic variant population between the reference samples to generate at least one relative incidence data set; and generating, by the computer, the non-tumor-derived probability set from the relative incidence data set.

[0013] In other aspects, the present disclosure provides a method of generating a classifier that at least partially uses a computer to classify nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as tumor-derived nucleic acid variants or non-tumor-derived nucleic acid variants. The method includes generating or providing, by a computer, at least one tumor variant data set including a reference population of tumor-associated genetic variants, wherein the tumor variant data set includes observed frequency data of one or more tumor-associated genetic variants in the reference population of tumor-associated genetic variants in a reference sample, the reference sample including reference plasma only and / or reference white blood cells, and wherein the reference sample is obtained from a single reference subject and / or from different reference subjects having the same cancer type. The method further includes determining, by a computer, one or more ratios of the observed frequency data of one or more tumor-associated genetic variants in the reference population of tumor-associated genetic variants between the reference samples to generate at least one relative incidence data set. Additionally, the method includes applying, by a computer, at least one machine learning model to the relative incidence data set to generate at least one non-tumor-derived probability set, thereby generating a classifier that classifies nucleic acid variants detected in a cfNA sample as tumor-derived nucleic acid variants or non-tumor-derived nucleic acid variants.

[0014] In some aspects, the present disclosure provides a method of at least partially using a computer to distinguish tumor-derived nucleic acid variants and non-tumor-derived nucleic acid variants in a cell-free nucleic acid (cfNA) sample obtained from a test subject having a cancer type. The method includes determining, by a computer, the incidence of one or more genetic variants observed in the cfNA sample to generate a test subject incidence data set. The method further includes comparing, by a computer, the incidence of one or more genetic variants in the test subject incidence data set with the incidence of genetic variants observed in a reference cfNA sample obtained from a reference test subject having the cancer type. Additionally, the method includes classifying, by a computer, a particular genetic variant in the test subject incidence data set as a non-tumor-derived nucleic acid variant when the incidence of the particular genetic variant in the test subject incidence data set is lower than a predetermined threshold associated with the particular genetic variant in the reference cfNA sample obtained from a reference test subject having the cancer type.

[0015] In some aspects, the present disclosure provides a method of using a computer at least in part to distinguish tumor-derived nucleic acid variants from non-tumor-derived nucleic acid variants in a cell-free nucleic acid (cfNA) sample obtained from a test subject. The method includes determining, by the computer, the incidence of one or more genetic variants observed in the cfNA sample to generate a test subject incidence data set. The method further includes comparing, by the computer, the incidence of one or more genetic variants in the test subject incidence data set to the incidence of genetic variants observed in a reference cfNA sample obtained from a reference subject with leukemia, lymphoma, and / or hematologic malignancies. Additionally, the method includes classifying, by the computer, a specific genetic variant in the test subject incidence data set as a non-tumor-derived nucleic acid variant when the incidence of the specific genetic variant in the test subject incidence data set is higher than a predetermined threshold associated with the specific genetic variant in the reference cfNA sample obtained from a reference subject with leukemia, lymphoma, and / or hematologic malignancies.

[0016] In some embodiments, the methods disclosed herein include identifying genetic variants present in a cfNA sample from sequencing reads derived from cfNA molecules in the cfNA sample. In certain of these embodiments, the sequencing reads are obtained from a targeted region of cfNA molecules in the cfNA sample. In some embodiments, a reference population of tumor-associated genetic variants is obtained from a reference sample. In certain embodiments, the reference leukocytes include a reference tumor tissue sample and / or a reference leukocyte sample. In some embodiments, the methods disclosed herein include obtaining a cfNA sample from a test subject. In certain embodiments, the reference sample comprises at least about 25, at least about 50, at least about 100, at least about 200, at least about 300, at least about 400, at least about 500, at least about 600, at least about 700, at least about 800, at least about 900, at least about 1,000, at least about 5,000, at least about 10,000, at least about 15,000, at least about 20,000, at least about 25,000, at least about 30,000, or more body fluid cells and / or leukocytes. In some embodiments, the cfNA sample comprises cell-free deoxyribonucleic acid (cfDNA). In certain embodiments, the cfNA sample comprises cell-free ribonucleic acid (cfRNA). In some embodiments, the test subject is a mammalian subject. In certain embodiments, the test subject is a human subject. In some embodiments, the reference body fluid sample comprises a plasma sample. In certain embodiments, the reference body fluid sample comprises a serum sample. In some embodiments, the reference non-body fluid sample is a non-plasma sample. In some embodiments, the reference non-body fluid (e.g., non-plasma) sample comprises a cell sample. In certain embodiments, the reference non-body fluid (e.g., non-plasma) sample comprises a tissue sample.

[0017] In some embodiments, the methods disclosed herein include selecting one or more therapies to treat a cancer type when one or more tumor-derived nucleic acid variants associated with the cancer type are detected in a cfNA sample obtained from a test subject. In certain embodiments, the methods disclosed herein include administering one or more therapies to a test subject to treat a cancer type when one or more tumor-derived nucleic acid variants associated with the cancer type are detected in a cfNA sample obtained from the test subject.

[0018] In some embodiments, the cancer type is selected from the group consisting of: biliary tract cancer, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, glioma, astrocytoma, breast cancer, metaplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal cancer, colon cancer, hereditary non-polyposis colorectal cancer, colorectal adenocarcinoma, gastrointestinal stromal tumor (GIST), endometrial cancer, endometrial stromal sarcoma, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, ocular melanoma, uveal melanoma, gallbladder cancer, gallbladder adenocarcinoma, renal cell carcinoma, clear cell renal cell carcinoma, transitional cell carcinoma, urothelial carcinoma, Wilms tumor, leukemia, acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), chronic myelomonocytic leukemia (CMML), liver cancer, hepatic epithelial carcinoma, hepatoma, hepatocellular carcinoma, cholangiocarcinoma, hepatoblastoma, lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B-cell lymphoma, non-Hodgkin lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, T-cell lymphoma, non-Hodgkin lymphoma, precursor T-lymphoblastic lymphoma / leukemia, peripheral T-cell lymphoma, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oral cancer, oral squamous cell carcinoma, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic ductal adenocarcinoma, solid pseudopapillary tumor, acinar cell carcinoma, prostate cancer, prostatic adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine cancer, gastric cancer, gastric epithelial carcinoma, gastrointestinal stromal tumor (GIST), uterine cancer, and uterine sarcoma. In certain embodiments, the reference tumor-associated genetic variants are selected from the group consisting of: single nucleotide variants (SNVs), insertions or deletions (insertions / deletions (indels)), copy number variants (CNVs), fusions, transversions, translocations, frameshifts, duplications, repeat expansions, and epigenetic variants.

[0019] In some embodiments, the methods disclosed herein include randomly dividing a tumor variant data set into a training data set and a test data set. In certain embodiments, the training data set comprises approximately 80% of the tumor variant data set, and the test data set comprises approximately 20% of the tumor variant data set. In some embodiments, the tumor variant data set comprises observed frequency data of one or more tumor-related genetic variants in a reference tumor-related genetic variant population in a reference sample for a particular cancer type. In some embodiments, the methods disclosed herein include using at least a portion of the tumor-related genetic variant population to train a machine learning model to produce a trained machine learning model, wherein tumor-derived nucleic acid variants and non-tumor-derived nucleic acid variants detected in a cfNA sample obtained from a test subject are distinguished from each other using the trained machine learning model. In some of these embodiments, the machine learning model is trained using one or more of: logistic regression, probit regression, decision tree, random forest, gradient boosting, support vector machine, K-nearest neighbor, and neural network. In some embodiments, the methods disclosed herein include using a probability threshold of at least approximately the 30th percentile for a particular genetic variant as a cut-off for classification. In some embodiments, the methods disclosed herein include performing logistic regression on at least one of the ratios to obtain a particular non-tumor-derived probability.

[0020] In some embodiments, the tumor variant data set comprises mutant allele fraction data observed in a reference sample for one or more tumor-related genetic variants in a reference tumor-related genetic variant population. In some embodiments, the methods disclosed herein include using one or more data normalization techniques to normalize the tumor variant data set. In certain of these embodiments, the data normalization techniques include min-max normalization and / or z-score normalization. In certain embodiments, a ratio of the observed frequency data of a particular genetic variant in reference plasma only relative to the observed frequency data of the particular genetic variant in reference white blood cells greater than one (1.0) indicates that the particular genetic variant may be a non-tumor-derived nucleic acid variant. In certain embodiments, wherein the reference white blood cells comprise a reference tumor tissue sample, the non-tumor-derived probability set comprises at least one clonal hematopoiesis-derived probability set.

[0021] In some embodiments, the tumor variant dataset includes mutant allele fraction data observed in a reference sample for one or more tumor-associated genetic variants in a reference population of tumor-associated genetic variants. In some embodiments, the methods disclosed herein include normalizing the tumor variant dataset using one or more data normalization techniques. In certain of these embodiments, the data normalization techniques include min-max normalization and / or z-score normalization. In certain embodiments, a ratio of the observed frequency data of a particular genetic variant in reference plasma only relative to the observed frequency data of the particular genetic variant in reference white blood cells less than one (1.0) indicates that the particular genetic variant may be a non-tumor-derived nucleic acid variant. In certain embodiments, wherein the reference white blood cells comprise a reference white blood cell sample, the non-tumor-derived probability set includes at least one clonal hematopoiesis-derived probability set.

[0022] In other aspects, the present disclosure provides a system that includes a controller that includes a computer-readable medium or is capable of accessing a computer-readable medium that includes non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least: (a) generating or providing at least one tumor variant dataset that includes a reference population of tumor-associated genetic variants, wherein the tumor variant dataset includes observed frequency data of one or more tumor-associated genetic variants in the reference population of tumor-associated genetic variants in a reference sample, the reference sample including reference plasma only and / or reference white blood cells, and wherein the reference sample is obtained from a single reference subject and / or from different reference subjects having the same cancer type; (b) determining one or more ratios of the observed frequency data of one or more tumor-associated genetic variants in the reference population of tumor-associated genetic variants between the reference samples to produce at least one relative incidence dataset; and (c) applying at least one machine learning model to the relative incidence dataset to produce at least one non-tumor-derived probability set to generate a classifier that classifies nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as tumor-derived nucleic acid variants or non-tumor-derived nucleic acid variants.

[0023] In other aspects, the present disclosure provides a system that includes a controller, the controller including a computer-readable medium or being able to access a computer-readable medium, the computer-readable medium including non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least: (a) determining a relative incidence of one or more tumor-associated genetic variations observed in one or more reference plasma-only samples compared to one or more reference white blood cells to generate at least one relative incidence data set; and (b) generating at least one non-tumor origin probability set from the relative incidence data set to generate a classifier that classifies nucleic acid variations detected in a cell-free nucleic acid (cfNA) sample as tumor-origin nucleic acid variations or non-tumor origin nucleic acid variations.

[0024] In other aspects, the present disclosure provides a system that includes a controller, the controller including a computer-readable medium or being able to access a computer-readable medium, the computer-readable medium including non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least: (a) determining a change in the mutant allele fraction (MAF) value of each of one or more tumor-associated genetic variations observed in one or more reference plasma-only samples compared to one or more reference white blood cells at at least two different time points and / or at least one statistic associated therewith; and (b) generating at least one non-tumor origin probability set from the MAF variance and / or relative incidence data set to generate a classifier that classifies nucleic acid variations detected in a cell-free nucleic acid (cfNA) sample as tumor-origin nucleic acid variations or non-tumor origin nucleic acid variations.

[0025] In some embodiments, the systems disclosed herein include a nucleic acid sequencer operably connected to a controller, the nucleic acid sequencer being configured to provide sequencing reads from cfNA molecules in a cfNA sample. In certain of these embodiments, the nucleic acid sequencer or another system component is configured to group sequence reads generated by the nucleic acid sequencer into sequence read families, each family containing sequence reads generated from a particular cfNA molecule in the cfNA sample. In certain embodiments, the systems disclosed herein include a database operably connected to the controller, the database including one or more therapies indexed to nucleic acid variants of tumor origin. In some embodiments, the systems disclosed herein include a sample preparation component operably connected to the controller, the sample preparation component being configured to prepare cfNA molecules in a cfNA sample to be sequenced by the nucleic acid sequencer. In certain embodiments, the systems disclosed herein include a nucleic acid amplification component operably connected to the controller, the nucleic acid amplification component being configured to amplify at least a targeted segment of cfNA molecules in the cfNA sample. In certain embodiments, the systems disclosed herein include a material transfer component operably connected to the controller, the material transfer component being configured to transfer one or more materials between at least the nucleic acid sequencer and the sample preparation component.

[0026] In some aspects, the present disclosure provides a computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least: (a) generating or providing at least one tumor variant data set comprising a reference tumor-associated genetic variant population, wherein the tumor variant data set comprises observed frequency data of one or more tumor-associated genetic variants in the reference tumor-associated genetic variant population in a reference sample, the reference sample comprising reference-only plasma and / or reference white blood cells, and wherein the reference sample is obtained from a single reference subject and / or from different reference subjects having the same cancer type; (b) determining one or more ratios of the observed frequency data of one or more tumor-associated genetic variants in the reference tumor-associated genetic variant population between the reference samples to produce at least one relative incidence data set; and (c) applying at least one machine learning model to the relative incidence data set to produce at least one non-tumor origin probability set to generate a classifier that classifies nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as tumor-origin nucleic acid variants or non-tumor origin nucleic acid variants. In certain embodiments of the methods, systems, computer-readable media, and other aspects of the present disclosure, one or more other features are optionally used in combination with or in place of the ratios of the observed frequency data. Some of these other features include, for example, consistency of incidence across cancer types, change in longitudinal mutant allele fraction (MAF) over time, proportions in blood cancers, variant gene names, locations, cancer types, chromosomal locations, and / or similar features.

[0027] In other aspects, the present disclosure provides a computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least: (a) determining the relative incidence of one or more tumor-associated genetic variants observed in one or more reference-only plasmas compared to one or more reference white blood cells to produce at least one relative incidence data set; and (b) generating at least one non-tumor origin probability set from the relative incidence data set to generate a classifier that classifies nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as tumor-origin nucleic acid variants or non-tumor origin nucleic acid variants.

[0028] In other aspects, the present disclosure provides a computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least: (a) determining a change in the mutant allele fraction (MAF) value of each of one or more tumor-associated genetic variants observed in one or more reference cell-free plasmas compared to one or more reference white blood cells at at least two different time points and / or at least one statistic associated therewith; and (b) generating at least one set of non-tumor origin probabilities from a MAF variance and / or relative incidence dataset to generate a classifier that classifies nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as tumor-origin nucleic acid variants or non-tumor origin nucleic acid variants.

[0029] In some embodiments of the systems or computer-readable media disclosed herein, the electronic processor also performs at least: dividing a tumor variant dataset (e.g., randomly or non-randomly) into a training dataset and a test dataset. In certain embodiments of the systems or computer-readable media disclosed herein, the electronic processor also performs at least: using at least a portion of a population of tumor-associated genetic variants to train a machine learning model to produce a trained machine learning model, and using the trained machine learning model to classify nucleic acid variants detected in a cfNA sample as tumor-origin nucleic acid variants or non-tumor origin nucleic acid variants. In some embodiments of the systems or computer-readable media disclosed herein, the electronic processor also performs at least: performing logistic regression on at least one of the ratios to obtain a specific non-tumor origin probability. In certain embodiments of the systems or computer-readable media disclosed herein, the electronic processor also performs at least: normalizing the tumor variant dataset using one or more data normalization techniques. In some embodiments of the systems or computer-readable media disclosed herein, the electronic processor also performs at least: selecting one or more therapies to treat a cancer type when one or more tumor-origin nucleic acid variants associated with the cancer type are detected in a cfNA sample.

[0030] In certain embodiments, the methods, systems or computer-readable media disclosed herein distinguish tumor-origin nucleic acid variants and non-tumor origin nucleic acid variants in a cell-free nucleic acid (cfNA) sample at least in part based on: (i) the consistency of the incidence of nucleic acid variants across cancer types; (ii) the change in the mutant allele fraction (MAF) of the nucleic acid variants over time; and / or (iii) the incidence of nucleic acid variants in blood cancers such as leukemia, lymphoma, and / or blood malignancies.

[0031] In some embodiments, the results of the systems and methods disclosed herein are used as inputs to generate a report. The report can be in paper or electronic format. For example, as determined by the methods and systems disclosed herein, the classification of nucleic acid variants detected in a cell-free nucleic acid sample as being of tumor origin or non-tumor origin can be directly displayed in such a report. In some embodiments, only nucleic acid variants classified as being of tumor origin are displayed in such a report.

[0032] The various steps of the methods disclosed herein, or the steps implemented by the systems disclosed herein, can be performed at the same or different times, at the same or different geographical locations (e.g., countries), and by the same or different individuals.

[0033] In other aspects, based on determining whether a variant is of tumor origin or non-tumor origin by the methods and systems disclosed herein, a therapy can be administered to a subject. In certain embodiments, administration of a treatment to a subject can be discontinued based on determining whether a variant is of tumor origin or non-tumor origin by the methods and systems disclosed herein. Brief Description of the Drawings

[0035] Figure 1 is a flow chart schematically depicting exemplary method steps for distinguishing nucleic acid variants of tumor origin and non-tumor origin according to some embodiments.

[0036] Figure 2 is a flow chart schematically depicting exemplary method steps for distinguishing nucleic acid variants of tumor origin and non-tumor origin according to some embodiments.

[0037] Figure 3 is a flow chart schematically depicting exemplary method steps for distinguishing nucleic acid variants of tumor origin and non-tumor origin according to some embodiments.

[0038] Figure 4 is an exemplary block diagram for generating a prediction model.

[0039] Figure 5 is a flow chart illustrating an exemplary training method

[0040] Figure 6 is an illustration of an exemplary process flow using a machine learning-based classifier.

[0041] Figure 7 is a schematic diagram of an exemplary system applicable to certain embodiments.

[0042] Figure 8An internal database of >250K clinical patients was utilized. Model design included features engineered from internal and external public datasets and was trained using multiple models with 10-fold cross-validation. Only results from the logistic regression model are shown. Model validation was performed on paired plasma and WBC late-stage samples sequenced by the epigenomic panel and an independent cohort of healthy donors sequenced by the genomic panel.

[0043] Figure 9 Model performance demonstrated high ROC AUC and accuracy for prediction calls. In A) 713 somatic SNV / insertions / deletions from 72 paired plasma and WBC GuardantInfinityTM samples and B) 243 somatic SNV / insertions / deletions from 76 paired plasma and healthy donors on GuardantOMNI TM, predictions of tumor and non-tumor status were compared to WBC confirmation. The lower confirmation rate of low VAF variants (<0.6%) observed in WBC sequencing may be attributed to the limit of detection of WBC variant calls and / or possible non-WBC lineage origin.

[0044] Figure 10 The most important assay-specific engineered features, where feature importance and examples include A) the top 10 features ranked by relative importance in the validation dataset. Each gene name is included in one-hot encoding. B) Engineered features ranked high include clonality, defined as VAF / tumor fraction (left panel) measured by methylation or maximum somatic VAF, VAF variation across time points (middle panel), average percentage (right panel); and C) consistency of variant incidence across solid tumor cancer types in the plasma database.

[0045] Figure 11 Consistency of non-tumor predictions: gene incidence and correlation with age. In the late-stage validation cohort from Figure 2 the number of variants predicted to be non-tumor or tumor origin within each gene as confirmed by WBC or model prediction. Variant counts for genes most commonly confirmed with cfDNA variants in WBC samples are shown, along with counts for clinically actionable genes (BRCA1, BRAF, KRAS, ESR1, ATM, CHEK2). The most common WBC-confirmed genes are consistent with previous reports, including a high incidence of clonal hematopoiesis in ATM and CHEK2 (*).

[0046] Figure 12Correlation between non-tumor determination and age. Proportion of variants detected or predicted as non-tumor in WBC in the late validation cohort divided by age range. As expected from the literature, variants predicted or confirmed as non-tumor are highly correlated with age.

[0047] Definition

[0048] To more readily understand the present disclosure, certain terms are first defined below. Additional definitions of the following terms and other terms may be set forth throughout the specification. If the definitions of the terms set forth below are inconsistent with the definitions in the patent applications or issued patents incorporated by reference, the definitions set forth in the present application are applied to understand the meaning of the term.

[0049] As used in this specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" include plural referents. Thus, for example, reference to "a method" includes one or more methods and / or steps of the type described herein and / or steps that will become apparent to those of ordinary skill in the art upon reading the present disclosure and the like. It should also be understood that there is an implicit "about" prior to the temperatures, concentrations, times, number of bases or base pairs, coverage, etc. discussed in the present disclosure, such that equivalents of minor and non-substantive differences are within the scope of the present disclosure. In this application, unless otherwise expressly stated, the use of the singular includes the plural. In addition, the use of "comprise", "comprises", "comprising", "contain", "contains", "containing", "include", "includes", and "including" is not intended to be limiting.

[0050] It should also be understood that the terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting. In addition, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. When describing and claiming these methods, computer-readable media, and systems, the following terms and their grammatical variations will be used according to the definitions set forth below.

[0051] About: As used herein, "about" or "approximately" applied to one or more values or elements of interest refers to a value or element that is similar to the reference value or element. In certain embodiments, the term "about" or "approximately" refers to a range of values or elements that fall within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1% or less of the reference value or element in either direction (greater than or less than) of the reference value or element, unless otherwise stated or otherwise apparent from the context (unless such numbers exceed 100% of the possible values or elements).

[0052] Adapter: As used herein, an "adapter" refers to a short nucleic acid (e.g., less than about 500 nucleotides in length, less than about 100 nucleotides in length, or less than about 50 nucleotides in length) that is typically at least partially double-stranded and is used to ligate either or both ends of a nucleic acid molecule of a given sample. An adapter can contain nucleic acid primer binding sites that permit amplification of a nucleic acid molecule flanked by adapters at both ends, and / or sequencing primer binding sites, including primer binding sites for sequencing applications such as various next-generation sequencing (NGS) applications. An adapter can also contain binding sites for capture probes, such as oligonucleotides attached to a flow cell support or the like. An adapter can also contain nucleic acid tags as described herein. The nucleic acid tags are typically positioned relative to the amplification primer and sequencing primer binding sites such that the nucleic acid tags are included in the amplicons and sequencing reads of a given nucleic acid molecule. Adapters of the same or different sequences can be ligated to the corresponding ends of a nucleic acid molecule. In certain embodiments, the same adapter is ligated to the corresponding ends of a nucleic acid molecule except that the sequences of the nucleic acid tags are different. In some embodiments, the adapter is a Y-shaped adapter, where one end is a blunt end or tailed as described herein for ligation to a nucleic acid molecule that is also blunt-ended or tailed with one or more complementary nucleotides. In yet other exemplary embodiments, the adapter is a bell-shaped adapter, comprising blunt ends or tailed ends for ligation to a nucleic acid molecule to be analyzed. Other exemplary adapters include T-tailed and C-tailed adapters.

[0053] Administer: As used herein, "administer" or "administering" a therapeutic agent (e.g., an immunotherapeutic agent) to a subject means to give, provide, or bring into contact a composition with the subject. Administration can be accomplished by any of a number of routes, including, for example, topical, oral, subcutaneous, intramuscular, intraperitoneal, intravenous, intrathecal, and intradermal.

[0054] Detailed description

[0055] Somatic variants of tumor origin in circulating nucleic acids, such as cell-free DNA (cfDNA), can be used for targeted therapy selection, longitudinal monitoring, and early detection of cancer. Cell-free tumor DNA (ctDNA) is small DNA fragments released from necrotic / apoptotic tumor cells or circulating tumor cells (CTCs) into the bloodstream. The vast majority of cfDNA is derived from normal cells, including normal white blood cells undergoing apoptosis or necrosis. Recent studies have shown that a large proportion of the mutations detected in cfDNA can be of non-tumor origin, particularly from clonal hematopoiesis that leads to the accumulation of somatic mutations in hematopoietic stem cells, resulting in cfDNA 'noise'. The presence of non-tumor variants in plasma / cfDNA can confound ctDNA interpretation; thus, methods and related aspects for distinguishing these have been extensively explored. In particular, clonal hematopoiesis-derived mutations (also known as clonal hematopoiesis origin) refer to the somatic acquisition of genomic mutations in hematopoietic stem cells and / or progenitor cells that lead to clonal expansion. Moreover, clonal hematopoiesis of indeterminate potential ("CHIP") refers to hematopoiesis in an individual involving the expansion of hematopoietic stem cells that contain one or more somatic mutations (e.g., blood cancer-related mutations and / or non-cancer-related mutations), but otherwise lack the diagnostic criteria for blood malignancies, such as clear morphological evidence of dysplasia. CHIP is a common age-related phenomenon in which hematopoietic stem cells give rise to genetically distinct subsets of blood cells.

[0056] Current methods for identifying nucleic acid variants derived from or otherwise originating from clonal hematopoiesis from cancer tumor nucleic acid variants include: sequencing white blood cells (WBCs) or peripheral blood mononuclear cells and removing these sequences from nucleic acid variants in the plasma portion of a specific blood sample, sequencing tissue and removing all nucleic acid variants outside of the tissue in the plasma fraction, or a combination of the two techniques (as above). Bioinformatics methods that have been attempted include removing nucleic acid variants that occur in genes frequently mutated in blood malignancies as they may be derived from the blood fraction, comparing the nucleic acid fragment sizes at individual loci in wild-type and WBC cfDNA, and using absolute or relative variant minor allele frequency cutoffs for the tumor. The challenge with these methods is the need for matched WBCs and tissue, which is not always available and complicates sample processing. The present disclosure presents novel bioinformatics methods and related aspects for classifying nucleic acid variants or mutations detected in plasma or other body fluids as being from tumor or non-tumor, without relying on the availability of matched WBCs or tumor tissue.

[0057] As related to the methods and compositions described herein, cell-free nucleic acid or “cfNA” refers to nucleic acids that are not contained within cells or otherwise associated with cells. Cell-free nucleic acids can include, for example, all unencapsulated nucleic acids derived from body fluids (e.g., blood, plasma, serum, urine, cerebrospinal fluid (CSF), etc.) from a subject. Cell-free nucleic acids include DNA (cfDNA), RNA (cfRNA), and hybrids thereof, including genomic DNA, mitochondrial DNA, circulating DNA, siRNA, miRNA, circular RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), and / or fragments of any of these. Cell-free nucleic acids can be double-stranded, single-stranded, or hybrids thereof. Cell-free nucleic acids can be released into body fluids through secretion or cell death programs, such as necrosis, apoptosis, or similar processes. Some cell-free nucleic acids are released from cancer cells into body fluids, such as circulating tumor DNA (ctDNA). Others are released from healthy cells. CtDNA can be unencapsulated, tumor-derived, fragmented DNA. Another example of cell-free nucleic acid is fetal DNA that freely circulates in the maternal bloodstream, also known as cell-free fetal DNA (cffDNA). Cell-free nucleic acids can have one or more epigenetic modifications, e.g., cell-free nucleic acids can be acetylated, 5-methylated, ubiquitinated, phosphorylated, sumoylated, ribosylated, and / or citrullinated. In some embodiments, for example, the term “cell-free nucleic acid” refers to nucleic acids that are not contained within cells or otherwise associated with cells when isolated from a particular subject.

[0058] Additionally, as related to the methods and compositions described herein, the cellular source of cell-free nucleic acid means the cell type from which a particular cell-free nucleic acid molecule is derived or otherwise originated (e.g., via an apoptotic process, a necrotic process, or a similar process). In certain embodiments, for example, a particular cell-free nucleic acid molecule can be derived from a tumor cell (e.g., a cancerous cell, etc.) or a non-tumor or normal cell (e.g., a non-cancerous cell, a hematopoietic stem cell, etc.).

[0059] Furthermore, for the methods and compositions described herein, including bioinformatics processes, a classifier involves algorithmic computer code that receives test data as input and produces a classification that classifies the input data as belonging to one or another category (e.g., tumor DNA or non-tumor DNA) as output.

[0060] In addition, with respect to the methods and compositions described herein, the minor allele frequency refers to the frequency at which a minor allele (e.g., not the most common allele) occurs in a particular nucleic acid population, such as a sample obtained from a subject. Genetic variants at low minor allele frequencies are generally present at relatively low frequencies in the sample.

[0061] In addition, with respect to the methods and compositions described herein, the mutant allele fraction (“MAF”) refers to the fraction of nucleic acid molecules having an allelic change or mutation relative to a reference at a particular genomic location in a particular sample. MAF is typically expressed as a fraction or percentage. For example, MAF is typically less than about 0.5, 0.1, 0.05, or 0.01 of all somatic variants or alleles present at a particular locus (i.e., less than about 50%, 10%, 5%, or 1%).

[0062] In addition, with respect to the methods and compositions described herein, the tumor fraction refers to an estimate of the fraction of nucleic acid molecules derived from a tumor in a particular sample. For example, the tumor fraction of a sample can be a measure derived from the maximum mutant allele fraction (MAX MAF) of the sample, or the coverage of the sample, or the length, epigenetic state, or other property of cfNA fragments in the sample, or any other selected feature of the sample. The term “MAX MAF” refers to the maximum or largest MAF of all somatic variants present in a particular sample. In some embodiments, the tumor fraction of a sample is equal to the MAX MAF of the sample.

[0063] Figure 1 is a flow chart schematically depicting exemplary method steps for distinguishing tumor-derived nucleic acid variants and non-tumor-derived nucleic acid variants in a cell-free nucleic acid (cfNA) sample obtained from a test subject. For example, the methods disclosed herein can be used to facilitate removal or reduction of background noise generated by non-tumor-derived nucleic acid variants (e.g., cfDNA fragments derived from non-cancerous or normal cells) detected in a particular sample from a test subject, thereby increasing assay sensitivity. As shown, method 100 includes determining (e.g., by a computer) the relative incidence of tumor-associated genetic variants observed in a reference-only plasma compared to a reference white blood cell (e.g., a cell sample, tissue sample, or similar sample) to generate a relative incidence data set (step 102). Method 100 also includes generating (e.g., by a computer) a set of non-tumor-derived probabilities from the relative incidence data set (step 104). In addition, method 100 further includes using the set of non-tumor-derived probabilities to classify nucleic acid variants detected in a cfNA sample obtained from a test subject as either tumor-derived nucleic acid variants or non-tumor-derived nucleic acid variants (step 106). Related systems and computer-readable media for implementing the methods disclosed herein are further described below.

[0064] For further illustration, Figure 2 is a flow chart schematically depicting exemplary method steps for distinguishing tumor-derived nucleic acid variants and non-tumor-derived nucleic acid variants in a cell-free nucleic acid (cfNA) sample obtained from a test subject according to some embodiments. As shown, method 200 includes generating (e.g., by computer) a tumor variant data set including a reference population of tumor-related genetic variants, wherein the tumor variant data set includes observed frequency (incidence) data of tumor-related genetic variants in the reference population of tumor-related genetic variants in a reference sample, the reference sample including a reference only plasma fluid sample (e.g., a plasma sample, a serum sample, or a similar sample) and / or a white blood cell sample (e.g., a cell sample, a tissue sample, or a similar sample) (step 202). The reference sample is typically obtained from a single reference subject and / or from different reference subjects having the same cancer type. Method 200 also includes determining (e.g., by computer) a ratio of the observed frequency data of tumor-related genetic variants in the reference population of tumor-related genetic variants between the reference samples to produce at least one MAF variance and / or relative incidence data set (step 204). Method 200 also includes generating (e.g., by computer) a non-tumor-derived probability set from the MAF variance and / or relative incidence data set (step 206). Additionally, method 200 includes using the non-tumor-derived probability set to classify nucleic acid variants detected in a cfNA sample obtained from a test subject as tumor-derived nucleic acid variants or non-tumor-derived nucleic acid variants (step 208).

[0065] As additional illustration, Figure 3FIG. 0 is a flow chart schematically depicting exemplary method steps for distinguishing or classifying tumor-derived nucleic acid variants and non-tumor-derived nucleic acid variants in a cell-free nucleic acid (cfNA) sample obtained from a test subject. As shown, method 300 includes obtaining raw data, such as raw data in the form of cancer and non-cancer (i.e., normal or healthy) sample data and white blood cell and / or tissue sample data (e.g., from the COSMIC cancer database, The Cancer Genome Atlas (TCGA) data, Memorial Sloan Kettering Cancer Center (MSKCC) data, and / or other data sources) (step 302). In the feature engineering step, input features are created by, for example, calculating the change in mutant allele fraction (MAF) over time (step 303), calculating the raw number and incidence of nucleic acid variants for all cancer types and calculating the ratio between the incidence of nucleic acid variants observed in plasma and / or other body fluids and the tissue datasets for all cancer types (step 304), calculating the proportion of nucleic acid variants in hematological malignancies or other cancer types (step 305), and testing the consistency of the incidence across plasma and / or other body fluid samples in different cancer types (e.g., developing a consistency score) (step 306). Bioinformatics data can include the observed frequency of genetic variants in samples of a particular cancer type (including hematological malignancies); the incidence of variants in plasma and / or other body fluids, tumor tissue, white blood cells, the mutant allele fraction of the variants, and others. Additional or other data types are optionally used for these feature engineering steps. Method 300 also includes transformation and cleaning processes, such as cleaning for sample incidence (e.g., adjustment for samples with low numbers of specific nucleic acid variants, low numbers of samples, etc.), performing a log transformation (e.g., Log(x+1) or Np.log1p), and performing normalization (e.g., Yeo-Johnson normalization, min-max normalization, z-score normalization, and / or similar normalization) (step 308). Method 300 also includes a machine learning step of generating a machine learning model using, for example, logistic regression or deep learning techniques to provide the probability of the presence of non-tumor nucleic acid variants in a particular sample (step 310). Exemplary models that can be used for training and further classification include, but are not limited to, logistic regression, probit regression, decision trees, random forests, gradient boosting, support vector machines, k-nearest neighbors, neural networks, or an ensemble of more than one of these methods. Ensemble methods are meta-algorithms that combine several machine learning techniques into one predictive model in order to reduce variance (bagging), bias (boosting), or improve prediction (stacking).Most ensemble methods use a single-base learning algorithm to produce homogeneous base learners, i.e., learners of the same type, resulting in a homogeneous ensemble. There are also some methods that use heterogeneous learners, i.e., learners of different types, resulting in a heterogeneous ensemble. To make the ensemble method more accurate than any of its individual members, the base learners must be as accurate as possible and as diverse as possible.

[0066] Optionally, various methods are used to divide the data set into a training set and a test set. In some embodiments, for example, the data set is randomly divided into a training data set and a test data set at an 80 / 20 ratio. Additionally, method 300 also includes selecting a cut-off value for determining a threshold for classifying a nucleic acid variant as being of tumor cell origin or non-tumor cell origin (step 312).

[0067] Body fluid: tissue ratio - binary classification

[0068] Some embodiments include comparing the incidence of variants observed in a body fluid sample (e.g., a plasma sample) data set relative to their occurrence in a tissue data set of the same cancer origin. In certain of these embodiments, logistic regression is performed on these ratios to obtain the probability of clonal hematopoiesis origin.

[0069] In some embodiments, the value of the performance metric can include, for example, accuracy (i.e., the fraction of correct predictions), balanced accuracy (defined as the average of the recall rates obtained for each class), precision macro (which involves calculating the metric for each label and then finding their unweighted average; however, this method does not account for label imbalance), precision micro (which involves globally calculating the metric by counting the total true positives, false negatives, and false positives), precision weighted (which involves calculating the metric for each label and finding their average weighted by the support (e.g., to determine the number of true instances for each label)), and similar performance metrics. In certain embodiments, the performance metric is estimated by stratified 5-fold cross-validation of the training set (e.g., where each fold is created by preserving the percentage of samples for each class).

[0070] The Box-Cox transformation is optionally used to transform non-normal distributions into normal distributions, but this method is not applicable to negative numbers. In contrast, the Yeo-Johnson transformation allows one to handle negative numbers. For example, for both logistic regression models and support vector machine (SVM) models, all features are optionally first transformed with the Yeo-Johnson transformation (a parametric, monotonic transformation that is applied to make the data more Gaussian-like in order to stabilize the variance and minimize skewness). In some embodiments, zero-mean, unit-variance normalization is further applied to the transformed data.

[0071] In some embodiments, the basic inputs for defining a set of parameters are: (1) the model type, and (2) a set of hyperparameters. In certain embodiments, the resulting parameters are used for all future classifications. In some embodiments, a training set is used to run a grid search with 5-fold stratified cross-validation on the following sets of hyperparameters (e.g., to define the cost of misclassification): kernel: linear, C: [0.001, 0.01, 0.1, 1, 10, 100, 1000], and kernel: radial basis function (rbf), C: [0.001, 0.01, 0.1, 1, 10, 100, 1000], γ: [0.0001, 0.001, 0.01, 0.1, 0.2, 0.3, 0.4, 0.5, 1].

[0072] Direct training on the dataset

[0073] Some embodiments use machine learning and features regarding mutated gene names, locations, cancer types, chromosomal locations, and other features based on known datasets to predict tumor / non-tumor origin. In these embodiments, the method generally includes: training a machine learning model on clonal hematopoiesis (CH) and tissue-specific training datasets to identify features specific to either origin, and applying the model to historical mutations observed in the previous dataset to determine the probability that a particular mutation is attributable to CH. In these embodiments, the method generally also includes determining which probability threshold is optimal for accurate classification of CH, and applying that list of probabilities to a new dataset in order to classify its origin as either tumor or clonal hematopoiesis. In some of these embodiments, the top 10th percentile of mutations will be a small number of mutations, but will have a high predictive value as a CH origin.

[0074] Incidence rate of body fluids relative to tumor tissue

[0075] Certain embodiments use the higher incidence of specific variants in a body fluid (e.g., plasma or serum) database relative to their occurrence in tumor tissue, where clonal hematopoiesis (CH) can be less confounding and can thus inform variants that may be from CH. In some of these embodiments, the method includes determining the incidence of specific variants that occur in the body fluid database and comparing it to the incidence observed in a primary tissue database such as the COSMIC database or a similar database. Some of these embodiments include determining the ratio of the incidence of variants observed in plasma only to the incidence of those variants observed in tissue samples. Some of these embodiments include calculating the odds ratio of the incidence of the variant and the probability that the value of the odds ratio is equal to, greater than, or less than 1. In these embodiments, the method generally also includes applying a machine learning model to these relative incidence values of body fluid relative to tissue incidence to determine the probability that the variant may be from CH or a non-tumor state. A probability threshold (e.g., approximately the 10th percentile, 15th percentile, 20th percentile, 25th percentile, 30th percentile, 35th percentile, 40th percentile, or other percentile) is typically used as a cutoff for classification. Generally, a small number of variants will have a high predictive value as non-tumor or CH.

[0076] Multi-class model using cross-tumor type distribution or concordance test

[0077] In some embodiments, tumor-specific variants will have a specific distribution or selection depending on the biology of each cancer type. If a variant is not specific to a cancer type, it will generally have a consistent distribution, which can indicate a passenger mutation or a non-tumor state. Thus, certain methods determine the incidence or their relative proportions and representation of variants across tumor types, and machine learning models can be trained to separate them into different tumor classes and non-tumor classes. Some of these methods include using the coefficient of variation to determine the distribution and any significant enrichment in a specific tumor type. In certain of these embodiments, a very small number of variants will be predictive and specific to the tumor type and are less likely to be CH. Some variants will have no apparent selectivity for a specific tumor type and have a low incidence across all tumor types, indicating that they may be CH. Generally, if a variant is generally consistent across tumor types, then it may be non-tumor / CH in origin, while if a variant is highly prevalent in certain tumors, then there is a greater likelihood that there is a biological selection for that variant in the tumor. Current methods that strictly rely on patient age or absolute VAF, and methods that ignore the expected relative incidences in different tumor contexts, will fail to account for these potential disease-specific mechanisms (or lack thereof) driving the observed VAF and the key biological features indicating the origin of the variant.

[0078] In these embodiments, other input features of the machine learning model include, for example, tumor classification based on mutations (e.g., tumor type or expected tumor type), methylation presence or signature, other mutations within a given sample (e.g., known CH mutations present in the sample, increasing the probability that other mutations in the sample are also non-tumor-derived), differences in family size of a particular mutation relative to a reference allele, the nature or other changes of nucleotides observed in the mutation, the absolute value of MAF within the sample, the relative value of MAF within the sample, how the value of MAF within the sample changes over time relative to other mutations, and / or similar features.

[0079] Monitoring the change of MAF value over time

[0080] In some cases, the clones of mutations that are non-tumor-derived will likely remain more stable over time in a subject compared to mutations that are tumor-derived. Thus, in some embodiments, the method includes, for each patient having multiple time points (e.g., >3), calculating the coefficient of variation (CV, deviation relative to the mean) of the percentage of mutations over time, and calculating the statistics and distribution of the CV across all mutations and patients. In these embodiments, known driver mutations or tumor mutations will generally have a dynamic percentage over time (due to tumor growth and shrinkage), with a large CV across time points compared to non-tumor mutations. This can also be used as an input feature for the classifier. In contrast, non-tumor mutation MAF is generally less dynamic and more stable over time compared to true tumor mutations, and will have a lower CV over time. The distribution of these CVs can be separated in the machine learning model and provide a robust classification of tumor or non-tumor status. In these embodiments, other input features of the machine learning model include, for example, mutation clonality (VAF relative to tumor fraction) over time or across patients, fragmentomics data points, fragment size, location, age of the patient (older patients have a higher probability of CHIP), and / or similar features. Current methods that track VAF or VAF dispersion across time points in an individual patient will be less accurate than methods that aggregate VAFs across patients, especially if these patients are all tested consecutively on the same platform and bioinformatics pipeline, resulting in consistent VAFs and more robust mutation measurements. Additionally, this classification method can use a static threshold that is not adjusted for the dispersion value relative to the absolute VAF, where non-tumor mutations with higher VAFs may be confused with lower VAF mutations having similar dispersion measurements over time. A machine learning model that considers both the absolute VAFs in a large enough cohort of patients measured on the same platform and the VAF dispersion across time points will have a higher classification resolution and is less likely to result in false positive or false negative labeling of tumor / non-tumor status.

[0081] Now turning to Figure 4, describes additional methods for generating a predictive model (e.g., a classification model). The described methods can use machine learning (“ML”) techniques to train at least one ML module 430 based on an analysis of one or more training datasets 410A - 410N by a training module 420, where the at least one ML module 430 is configured to classify mutations detected in plasma as being of tumor origin or non - tumor origin that can be from clonal hematopoiesis or biological noise.

[0082] One or more training datasets 410A - 410N can include cancer / non - cancer (e.g., tumor / non - tumor) body fluid only plasma sample data and cancer / non - cancer (e.g., tumor / non - tumor) white blood cell and / or non - body fluid (e.g., tissue) sample data. One or more training datasets 410A - 410N can include cancer / non - cancer (e.g., tumor / non - tumor) body fluid only plasma sample data and cancer / non - cancer white blood cell and / or non - body fluid (e.g., tissue) sample data (e.g., from the COSMIC cancer database, The Cancer Genome Atlas (TCGA) data, and / or other data sources). A subset of the cancer / non - cancer body fluid sample data and / or cancer / non - cancer non - body fluid sample data can be randomly assigned to the training dataset 410 or the test dataset. In some embodiments, the assignment of data to the training dataset or the test dataset may not be completely random. In such cases, one or more criteria can be used during the assignment. Generally, any suitable method can be used to assign data to the training dataset or the test dataset while ensuring that the data distribution is somewhat similar in the training dataset and the test dataset.

[0083] The training module 420 can train the ML module 430 by extracting a feature set from the cancer / non - cancer body fluid sample data and / or cancer / non - cancer non - body fluid sample data in the training dataset 410 according to one or more feature selection techniques. The training module 420 can train the ML module 430 by extracting a feature set from the training dataset 410 that includes statistically significant features.

[0084] The training module 420 can extract a feature set from the training dataset 410 in various ways. The training module 420 can perform feature extraction multiple times, each time using a different feature extraction technique. In an example, the feature sets generated using different techniques can each be used to generate different machine - learning - based classification models 440. For example, the feature set with the highest quality metric can be selected for training. The training module 420 can use the feature set to build one or more machine - learning - based classification models 440A - 440N that are configured to classify the origin of new variants (e.g., with unknown origin) as tumor or non - tumor.

[0085] The training dataset 410 can be analyzed to determine any dependencies, associations, and / or correlations between the features in the training dataset 410 and the experimental parameters. The identified correlations can be in the form of a list of features. As used herein, the term "feature" can refer to any characteristic of a data item that can be used to determine whether the data item falls within one or more specific categories. By way of example, the features described herein can include one or more of the following: the observed frequency of genetic variation in samples of a particular cancer type (including hematological malignancies); the incidence of variation in plasma, tumor tissue, or white blood cells; and / or the minor allele frequency of the variation.

[0086] Feature selection techniques can include one or more feature selection rules. The one or more feature selection rules can include a feature occurrence rule. The feature occurrence rule can include determining which features in the training dataset 410 occur more than a threshold number of times and identifying those features that meet the threshold as features.

[0087] A single feature selection rule can be applied to select features, or multiple feature selection rules can be applied to select features. The feature selection rules can be applied in a cascaded manner, where the feature selection rules are applied in a specific order and applied to the result of the previous rule. For example, the feature occurrence rule can be applied to the training dataset 410 to generate a first list of features. The final list of features can be analyzed according to additional feature selection techniques to determine one or more feature groups (e.g., a feature group that can be used to classify a variation as tumor-derived or non-tumor-derived). Any suitable computational technique can be used to identify the feature groups using any feature selection technique such as a filtering method, a wrapper method, and / or an embedded method. One or more feature groups can be selected according to a filtering method. Filtering methods include, for example, Pearson's correlation, linear discriminant analysis, analysis of variance (ANOVA), chi-square, combinations thereof, and similar methods. Feature selection according to a filtering method is independent of any machine learning algorithm. Instead, features can be selected based on a score of the correlation of the feature with the outcome variable in various statistical tests.

[0088] As another example, one or more feature groups can be selected based on a wrapper method. The wrapper method can be configured to use a subset of features and use the subset of features to train a machine learning model. Based on the inferences drawn from a previous model, features can be added to and / or removed from the subset. Wrapper methods include, for example, forward feature selection, backward feature elimination, recursive feature elimination, combinations thereof, and the like. As an example, forward feature selection can be used to identify one or more feature groups. Forward feature selection is an iterative method that starts with no features in a machine learning model. In each iteration, the feature that best improves the model is added until adding a new variable no longer improves the performance of the machine learning model. As an example, backward elimination can be used to identify one or more feature groups. Backward elimination is an iterative method that starts with all features in a machine learning model. In each iteration, the least important feature is removed until no improvement is observed when removing features. Recursive feature elimination can be used to identify one or more feature groups. Recursive feature elimination is a greedy optimization algorithm aimed at finding the subset of features with the best performance. Recursive feature elimination repeatedly creates models and retains the best or worst performing features at each iteration. Recursive feature elimination builds the next model with the remaining features until all features are exhausted. Then, recursive feature elimination ranks the features based on the order of feature elimination.

[0089] As a further example, one or more feature groups can be selected based on an embedded method. Embedded methods combine the characteristics of filter methods and wrapper methods. Embedded methods include, for example, Least Absolute Shrinkage and Selection Operator (LASSO) and Ridge regression, which implement penalty functions to reduce overfitting. For example, LASSO regression performs L1 regularization, which adds a penalty equal to the absolute value of the coefficient size, while Ridge regression performs L2 regularization, which adds a penalty equal to the square of the coefficient size.

[0090] After the training module 420 has generated the feature set, the training module 420 can generate a machine learning-based classification model 440 based on the feature set. A machine learning-based classification model can refer to a complex mathematical model for data classification generated using machine learning techniques. In one example, the machine learning-based classification model 440 can include a mapping of support vectors representing boundary features. By way of example, the boundary features can be selected from the feature set and / or represent the highest-ranked features in the feature set.

[0091] The training module 420 can use a feature set determined or extracted from the training data set 410 to construct machine learning-based classification models 440A - 440N. In some instances, the machine learning-based classification models 440A - 440N can be combined into a single machine learning-based classification model 440. Similarly, the ML module 430 can represent a single classifier that includes a single or more than one machine learning-based classification model 440 and / or multiple classifiers that include a single or more than one machine learning-based classification model 440.

[0092] Features can be combined in a classification model trained using machine learning methods such as discriminant analysis; decision trees; nearest neighbor (NN) algorithms (e.g., k-NN models, replicator NN models, etc.); statistical algorithms (e.g., Bayesian networks, etc.); clustering algorithms (e.g., k-means, mean shift, etc.); neural networks (e.g., Kohonen networks, artificial neural networks, etc.); support vector machines (SVM); logistic regression algorithms; linear regression algorithms; Markov models or chains; principal component analysis (PCA) (e.g., for linear models); multi-layer perceptron (MLP) ANN (e.g., for non-linear models); replicator Kohonen networks (e.g., for non-linear models, typically used for time series); random forest classification; combinations and / or similar methods thereof. The resulting ML module 430 can include decision rules or mappings for each feature to determine the origin of the variant tumor / non-tumor.

[0093] In an implementation, the training module 420 can train the machine learning-based classification model 440 as a convolutional neural network (CNN). The CNN includes at least one convolutional feature layer and three fully connected layers that lead to a final classification layer (softmax). The final classification layer can ultimately apply the softmax function, known in the art, to combine the outputs of the fully connected layers.

[0094] The features and the ML module 430 can be used to predict the tumor / non-tumor origin of the variations in the test dataset. In one instance, the prediction result for each variation can include a confidence level corresponding to the likelihood or probability that the variation in the test dataset is associated with a tumor origin or a non-tumor origin. The confidence level can be a value between 0 and 1. In one instance, when there are two states (e.g., tumor origin and non-tumor origin), the confidence level can correspond to a value p, which refers to the likelihood that a particular variation belongs to the first state (e.g., tumor origin). In this case, the value 1 - p can refer to the likelihood that a particular variation belongs to the second state (e.g., non-tumor origin). Generally, when there are more than two states, multiple confidence levels can be provided for each variation in the test dataset and for each feature. The best-performing features can be determined by comparing the results obtained for each test variation with the known tumor / non-tumor origin of each test variation. Generally, the best-performing features will have results that closely match the known tumor / non-tumor origin states. The best-performing features can be used to predict / classify the tumor / non-tumor origin state of a particular variation.

[0095] Figure 5 FIG. is a flow chart illustrating an exemplary training method 500 for generating the ML module 430 using the training module 420. The training module 420 can implement a supervised, unsupervised, and / or semi-supervised (e.g., reinforcement-based) machine learning-based classification model 440. Figure 5 The method 500 illustrated in FIG. is an example of a supervised learning method; variations of this example of the training method are discussed below, however, other training methods can be similarly implemented to train unsupervised and / or semi-supervised machine learning models.

[0096] The training method 500 can determine (e.g., access, receive, retrieve, etc.) data at step 510. The data can include cancer / non-cancer (e.g., tumor / non-tumor) body fluid sample data and cancer / non-cancer (e.g., tumor / non-tumor) non-body fluid (e.g., tissue) sample data. The data can include one or more variations, each having a specified tumor or non-tumor origin state.

[0097] The training method 500 can generate a training data set and a test data set at step 520. The training data set and the test data set can be generated by randomly allocating data to the training data set or the test data set. In some embodiments, the computational parameters and the associated experimental parameters can be allocated to training or test data not entirely randomly. As an example, most of the computational parameters and the associated experimental parameters can be used to generate the training data set. For example, 75% of the computational parameters and the associated experimental parameters can be used to generate the training data set, and 25% can be used to generate the test data set. In another example, 80% of the computational parameters and the associated experimental parameters can be used to generate the training data set, and 20% can be used to generate the test data set.

[0098] The training method 500 can determine (e.g., extract, select, etc.) one or more features at step 530, and the one or more features can be used by, for example, a classifier to distinguish different classifications of a tumor state from a non-tumor state. As an example, the training method 500 can determine a feature set from cancer / non-cancer body fluid sample data and cancer / non-cancer non-body fluid sample data. In additional examples, the feature set can be determined from data other than the cancer / non-cancer body fluid sample data and the cancer / non-cancer non-body fluid sample data in the training data set or the test data set. Such other data can be used to determine an initial feature set, which can be further reduced using the training data set.

[0099] The training method 500 can use one or more features to train one or more machine learning models at step 540. In one example, supervised learning can be used to train the machine learning models. In another example, other machine learning techniques can be employed, including unsupervised learning and semi-supervised learning. The machine learning models trained at 540 can be selected based on different criteria depending on the problem to be solved and / or the data available in the training data set. For example, machine learning classifiers may suffer from different degrees of bias. Therefore, more than one machine learning model can be trained at 540 and optimized, improved, and cross-validated at step 550.

[0100] The training method 500 can select one or more machine learning models to build a prediction model at 560. The test data set can be used to evaluate the prediction model. At step 570, the prediction model can analyze the test data set and generate a predicted tumor / non-tumor origin state. The predicted tumor / non-tumor origin can be evaluated at step 580 to determine whether such a value has reached the desired accuracy level. The performance of the prediction model can be evaluated in various ways based on multiple true positives, false positives, true negatives, and / or false negatives classifications of more than one data point indicated by the prediction model.

[0101] For example, false positives of a prediction model can refer to the number of times the prediction model incorrectly classifies a variant that is actually not of tumor origin as being of tumor origin. Conversely, false negatives of a prediction model can refer to the number of times the machine learning model classifies a variant as not being of tumor origin (whereas in fact the variant is of tumor origin). True negatives and true positives can refer to the number of times the prediction model correctly classifies one or more variants. Associated with these measurements are the concepts of recall and precision. Generally, recall is the ratio of true positives to the sum of true positives and false negatives, which quantifies the sensitivity of the prediction model. Similarly, precision is the ratio of true positives to the sum of true positives and false positives. When such an expected accuracy level is reached, the training phase ends, and the prediction model (e.g., ML module 430) can be output at step 590; however, when the expected accuracy level is not reached, subsequent iterations of the training method 500 can be performed starting from step 510, varying, for example, by considering a larger data set.

[0102] Figure 6 is a diagrammatic illustration of an exemplary process flow for classifying variants as being of tumor origin or not of tumor origin using a machine learning-based classifier. As Figure 6 illustrated, unclassified variants 610 can be provided as input to the ML module 430. The ML module 430 can process the unclassified variants 610 using a machine learning-based classifier to arrive at a prediction result 620. The prediction result 620 can identify one or more features of the unclassified variants 610. For example, the classification result 620 can identify the origin status of the unclassified variants 610 (e.g., whether the variant is of tumor origin or not of tumor origin). Thus, in an embodiment, a method implemented using a network-based computer system is disclosed, the computer system including one or more processors, a network interface, and one or more memories, the method including retrieving, by the computer system, genetic information and additional information from one or more memories of more than one tumor and non-tumor only plasma and more than one tumor and non-tumor non-bodily fluid (e.g., tissue) samples, where the additional information includes a tumor origin or non-tumor origin status; and training, by one or more processors, a machine learning model by fitting one or more models to the genetic information and the additional information, where each of the one or more models is configured to receive genetic information of an individual as input and provide a prediction of whether the individual has or will develop a tumor as output.

[0103] System and computer-readable medium

[0104] The present disclosure also provides various systems, bioinformatics pipelines, and computer program products or machine-readable media. For example, in some embodiments, the methods described herein are optionally performed or facilitated, at least in part, using systems, distributed computing hardware and applications (e.g., cloud computing services), electronic communication networks, communication interfaces, computer program products, machine-readable media, electronic storage media, software (e.g., machine-executable code or logical instructions), and the like. For illustration, Figure 7 FIG. shows a schematic diagram of an exemplary system suitable for implementing at least some aspects of the methods disclosed in the present application. As shown, system 700 includes at least one controller or computer, such as server 702 (e.g., a search engine server), which includes a processor 704 and a memory, storage device, or memory component 706, and one or more other communication devices 714 and 716 (e.g., client computer terminals, telephones, tablets, laptops, other mobile devices, etc.) located at a location remote from remote server 702 and communicating with remote server 702 via an electronic communication network 712 (such as the Internet or other interconnected network). Communication devices 714 and 716 typically include an electronic display (e.g., an Internet-enabled computer, etc.) that communicates with a computer such as server 702 via network 712, where the electronic display includes a user interface (e.g., a graphical user interface (GUI), a web-based user interface, etc.) for displaying results when implementing the methods described herein. In certain embodiments, the communication network also includes physically transferring data from one location to another, for example, using a hard disk drive, a thumb drive, or other data storage mechanisms. System 700 also includes a program product 708 stored on a computer or machine-readable medium, such as one or more various types of memories, such as memory 706 of server 702, which can be read by server 702 to facilitate, for example, a search application or other applications executable by one or more other communication devices such as 714 (schematically shown as a desktop or personal computer) and 716 (schematically shown as a tablet computer). In some embodiments, system 700 optionally further includes at least one database server, such as, for example, server 710 associated with an online website having data (e.g., a list of nucleic acid variants, indexed therapies, etc.) stored thereon that can be searched directly or via search engine server 702. System 700 optionally further includes one or more other servers located remote from server 702, each server optionally associated with one or more database servers 710 remote from or local to each other server. The other servers can beneficially serve geographically remote users and enhance geographically distributed operations.

[0105] As would be understood by one of ordinary skill in the art, the memory 706 of the server 702 optionally includes volatile and / or non-volatile memory, including, for example, RAM, ROM, and magnetic or optical disks, etc. One of ordinary skill in the art should also understand that although the server 702 is illustrated as a single server, the configuration of the illustrated server 702 is given by way of example only, and other types of servers or computers configured according to various other methods or architectures may also be used. Figure 7 The server 702 schematically illustrated in Figure 7 represents a server or a server cluster or a server farm, and is not limited to any single physical server. The server site may be deployed as a server farm or a server cluster managed by a server hosting provider. The number of servers and their architecture and configuration may be increased based on the usage, requirements, and capacity requirements of the system 700. As would also be understood by one of ordinary skill in the art, the other user communication devices 714 and 716 in these embodiments may be, for example, laptop computers, desktop computers, tablet computers, personal digital assistants (PDAs), mobile phones, servers, or other types of computers. As is known and understood by one of ordinary skill in the art, the network 712 may include the Internet, an intranet, a telecommunications network, an extranet, or the World Wide Web of more than one computer / server, which communicate with one or more other computers through a communication network, and / or a part of a local network or other local area network.

[0106] As would be further understood by one of ordinary skill in the art, the exemplary program product or machine-readable medium 708 is optionally in the form of microcode, programs, cloud computing formats, routines, and / or symbolic languages, which provide one or more sets of ordered operations that control the functions of the hardware and direct its operation. According to the exemplary embodiments, the program product 708 does not need to reside entirely in volatile memory, but may be selectively loaded as needed according to various methods known and understood by one of ordinary skill in the art.

[0107] As further understood by those of ordinary skill in the art, the term "computer-readable medium" or "machine-readable medium" refers to any medium that participates in providing instructions to a processor for execution. By way of illustration, the term "computer-readable medium" or "machine-readable medium" includes distribution media, cloud computing formats, intermediate storage media, the execution memory of a computer, and any other medium or device capable of storing a program product 708 that implements the functions or processes of the various embodiments of the present disclosure, for example, for reading by a computer. The "computer-readable medium" or "machine-readable medium" can take many forms, including but not limited to non-volatile media, volatile media, and transmission media. Non-volatile media includes, for example, optical discs or magnetic disks. Volatile media includes dynamic memory, such as the main memory of a given system. Transmission media includes coaxial cables, copper wire, and fiber optics, including the wires that make up a bus. Transmission media can also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications. Exemplary forms of computer-readable media include floppy disks, flexible disks, hard disks, magnetic tapes, flash memory disks, or any other magnetic medium, CD-ROMs, any other optical medium, punched cards, paper tapes, any other physical storage medium with hole patterns, RAM, PROM, and EPROM, FLASH-EPROM, any other memory chip or cartridge, carrier waves, or any other medium from which a computer can read.

[0108] The program product 708 is optionally copied from the computer-readable medium to a hard disk or similar intermediate storage medium. When the program product 708 or a portion thereof is to be run, it is optionally loaded from their distribution media, their intermediate storage media, etc. into the execution memory of one or more computers, configuring the computers to operate in accordance with the functions or methods of the various embodiments. All such operations are, for example, well known to those of ordinary skill in the computer system art.

[0109] For further illustration, in certain embodiments, the present application provides a system including one or more processors and one or more memory components in communication with the processors. The memory components generally include one or more instructions that, when executed, cause the processors to provide information that causes at least one nucleic acid variant list, variant classification determination report or result, selected therapy, etc. to be displayed (e.g., via communication devices 714, 716, etc.) and / or to receive information from other system components and / or from system users (e.g., via communication devices 714, 716, etc.).

[0110] In some embodiments, the program product 708 includes non-transitory computer-executable instructions that, when executed by the electronic processor 704, perform at least: (i) generating a tumor variant data set including a reference population of tumor-associated genetic variants, where the tumor variant data set includes observed frequency data of tumor-associated genetic variants in the reference population of tumor-associated genetic variants in a reference sample, the reference sample including a reference of only plasma and / or reference white blood cells, and where the reference sample is obtained from a single reference subject and / or from different reference subjects having the same cancer type, (ii) determining a ratio of the observed frequency data of tumor-associated genetic variants in the reference population of tumor-associated genetic variants between reference samples to produce a relative incidence data set, (iii) generating a non-tumor origin probability set from the relative incidence data set, and (iv) using the non-tumor origin probability set to classify nucleic acid variants detected in a cfNA sample obtained from a test subject as tumor-origin nucleic acid variants or non-tumor origin nucleic acid variants.

[0111] System 700 generally also includes additional system components configured to perform aspects of the methods described herein. In some of these embodiments, one or more of these additional system components are located at a location remote from the remote server 702 and communicate with the remote server 702 via an electronic communication network 712, while in other embodiments, one or more of these additional system components are located at a local location and communicate with the server 2102 (i.e., in the absence of an electronic communication network 712), or communicate directly with, for example, a desktop computer 714.

[0112] In some embodiments, for example, an additional system component including a sample preparation component 718 is operably connected (directly or indirectly (e.g., via the electronic communication network 712)) to the controller 702. The sample preparation component 718 is configured to prepare nucleic acids in a sample (e.g., prepare a nucleic acid library) for amplification and / or sequencing by a nucleic acid amplification component (e.g., a thermal cycler, etc.) and / or a nucleic acid sequencer. In certain of these embodiments, the sample preparation component 718 is configured to separate nucleic acids from other components in the sample, to ligate one or more adapters containing barcodes to the nucleic acids as described herein, to selectively enrich one or more regions from the genome or transcriptome prior to sequencing, and so on.

[0113] In certain embodiments, system 700 also includes a nucleic acid amplification component 720 (e.g., a thermal cycler, etc.) operably connected (directly or indirectly (e.g., via the electronic communication network 712)) to the controller 702. The nucleic acid amplification component 720 is configured to amplify nucleic acids in a sample from a subject. For example, the nucleic acid amplification component 720 is optionally configured to amplify regions selectively enriched from the genome or transcriptome in a sample as described herein.

[0114] System 700 generally also includes at least one nucleic acid sequencer 722, which is operably connected (either directly or indirectly (e.g., via electronic communication network 712)) to controller 702. The nucleic acid sequencer 722 is configured to provide sequence information from nucleic acids (e.g., amplified nucleic acids) in a sample from a subject. Basically any type of nucleic acid sequencer can be suitable for these systems. For example, the nucleic acid sequencer 722 is optionally configured to perform pyrosequencing, single molecule sequencing, nanopore sequencing, semiconductor sequencing, ligation sequencing, hybridization sequencing, or other techniques on nucleic acids to generate sequencing reads. Optionally, the nucleic acid sequencer 722 is configured to group the sequence reads into sequence read families, each family including sequence reads generated from nucleic acids in a given sample. In some embodiments, the nucleic acid sequencer 722 uses a clonal single molecule array derived from a sequencing library to generate sequencing reads. In certain embodiments, the nucleic acid sequencer 722 includes at least one chip having a micropore array for sequencing a sequencing library to generate sequencing reads.

[0115] To facilitate full or partial system automation, system 700 generally also includes a material transfer component 724 operably connected (either directly or indirectly (e.g., via electronic communication network 712)) to controller 702. The material transfer component 724 is configured to transfer one or more materials (e.g., nucleic acid samples, amplicons, reagents, etc.) to and / or from the nucleic acid sequencer 722, the sample preparation component 718, and the nucleic acid amplification component 720.

[0116] Additional details related to computer systems and networks, databases, and computer program products are also provided, for example, in: Peterson, Computer Networks: A Systems Approach, Morgan Kaufmann, 5th Edition (2011), Kurose, Computer Networking: A Top-Down Approach, Pearson, 7th Edition (2016), Elmasri, Fundamentals of Database Systems, Addison Wesley, 6th Edition (2010), Coronel, Database Systems: Design, Implementation, & Management, Cengage Learning, 11th Edition (2014), Tucker, Programming Languages, McGraw-Hill Science / Engineering / Math, 2nd Edition (2006), and Rhoton, Cloud Computing Architected: Solution Design Handbook, Recursive Press (2011), each of which is incorporated by reference in its entirety.

[0117] Sample collection and preparation

[0118] The sample can be any biological sample isolated from a subject. The sample can include a body fluid or body tissue (e.g., a known or suspected solid tumor). The sample can include whole blood, platelets, serum, plasma, feces, red blood cells, white blood cells (leukocytes), endothelial cells, tissue biopsies, cerebrospinal fluid, synovial fluid, lymphatic fluid, ascites, interstitial fluid or extracellular fluid, fluid in the space between cells (including gingival crevicular fluid), bone marrow, pleural effusion, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, urine. The sample is preferably a body fluid, particularly blood and its fractions, and urine. Such samples include nucleic acids shed from a tumor. The nucleic acids can include DNA and RNA and can be in double-stranded and / or single-stranded form. The sample can be in the form initially isolated from the subject or can have undergone additional processing to remove or add components, such as cells, enrich one component relative to another, or convert one form of nucleic acid to another, such as RNA to DNA, or single-stranded nucleic acid to double-stranded. Thus, for example, the body fluid for analysis is plasma or serum containing cell-free nucleic acids such as cell-free DNA (cfDNA).

[0119] In certain embodiments, the polynucleotide can be enriched prior to sequencing. The enrichment can be targeted to specific target regions (“target sequences”) or can be non-specific. In some embodiments, the targeted regions of interest can be enriched using capture probes (“baits”) selected against one or more bait set panels using differential tiling and capture schemes. The differential tiling and capture schemes use different relative concentrations of bait sets to differentially tile (e.g., at different “resolutions”) across genomic regions associated with the baits, subject to constraints (e.g., sequencer limitations such as sequencing throughput, utility of each bait, etc.), and capture them at levels desired for downstream sequencing. These targeted genomic regions of interest can include regions of the subject's genome or transcriptome. In some embodiments, biotinylated beads with probes for one or more regions of interest can be used to capture the target sequences, optionally followed by amplification of these regions to enrich the regions of interest.

[0120] Sequence capture generally involves using oligonucleotide probes that hybridize to the target sequences. Probe set strategies can involve tiling the probes across the region of interest. Such probes can be, for example, about 60 to 130 bases in length. The set can have a depth of about 2x, 3x, 4x, 5x, 6x, 8x, 9x, 10x, 15x, 30x, 50x or greater. The effectiveness of sequence capture depends in part on the length of the sequence in the target molecule that is complementary (or nearly complementary) to the sequence of the probe.

[0121] In some embodiments, the methods of the present disclosure include selectively enriching regions in the subject's genome or transcriptome prior to sequencing. In other embodiments, the methods of the present disclosure include non-selectively enriching regions in the subject's genome or transcriptome prior to sequencing.

[0122] In certain embodiments, a sample index sequence is introduced into the polynucleotide after enrichment. The sample index sequence can be introduced into the polynucleotide by PCR or ligated to the polynucleotide, optionally as part of an adaptor.

[0123] The volume of the body fluid can depend on the desired read depth for the region to be sequenced. Exemplary volumes are 0.4 - 40 ml, 5 - 20 ml, 10 - 20 ml. For example, the volume can be 0.5 ml, 1 ml, 5 ml, 10 ml, 20 ml, 30 ml or 40 ml. The volume of the body fluid sampled can be 5 ml to 20 ml.

[0124] The sample can contain various amounts of nucleic acids containing genomic equivalents. For example, a sample of about 30 ng of DNA can contain about 10,000 (10^4) haploid human genome equivalents, and in the case of cfDNA, can contain about 200 billion (2×10 11 ) individual polynucleotide molecules. Similarly, a sample of about 100 ng of DNA can contain about 30,000 haploid human genome equivalents, and in the case of cfDNA, can contain about 600 billion individual molecules.

[0125] The sample can contain nucleic acids from different sources, such as nucleic acids from cells and cell-free. The sample can contain nucleic acids carrying mutations. For example, the sample can contain DNA carrying germline mutations and / or somatic mutations. The sample can contain DNA carrying cancer-related mutations (e.g., cancer-related somatic mutations).

[0126] Exemplary amounts of cell-free nucleic acids in the sample prior to amplification range from about 1 fg to about 1 μg, such as 1 pg to 200 ng, 1 ng to 100 ng, 10 ng to 1000 ng. For example, the amount can be up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng or up to about 20 ng of cell-free nucleic acid molecules. The amount can be at least 1 fg, at least 10 fg, at least 100 fg, at least 1 pg, at least 10 pg, at least 100 pg, at least 1 ng, at least 10 ng, at least 100 ng, at least 150 ng or at least 200 ng of cell-free nucleic acid molecules. The amount can be up to 1 femtogram (fg), 10 fg, 100 fg, 1 picogram (pg), 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng or 200 ng of cell-free nucleic acid molecules. The method can include obtaining from 1 femtogram (fg) to 200 ng.

[0127] Cell-free nucleic acids have an exemplary size distribution of about 100 - 500 nucleotides, where molecules of 110 to about 230 nucleotides represent about 90% of the molecules, the mode in humans is about 168 nucleotides, and the second minor peak is in the range between 240 and 430 nucleotides. Cell-free nucleic acids can be about 160 to about 180 nucleotides, or about 320 to about 360 nucleotides, or about 430 to about 480 nucleotides.

[0128] Cell-free nucleic acids can be isolated from body fluids by a partitioning step in which cell-free nucleic acids present in solution are separated from intact cells and other insoluble components in the body fluid. Partitioning can include techniques such as centrifugation or filtration. Optionally, the cells in the body fluid can be lysed and the cell-free nucleic acids and cellular nucleic acids are processed together. Typically, after addition of buffer and a washing step, the cell-free nucleic acids can be precipitated with alcohol. Further cleaning steps such as silica-based columns can be used to remove contaminants or salts. For example, non-specific bulk carrier nucleic acids can be added throughout the reaction to optimize certain aspects of the procedure, such as yield.

[0129] After such processing, the sample can include various forms of nucleic acids, including double-stranded DNA, single-stranded DNA, and single-stranded RNA. Optionally, the single-stranded DNA and RNA can be converted to the double-stranded form so that they are included in subsequent processing and analysis steps.

[0130] Amplification

[0131] Sample nucleic acids flanked by adapters can be amplified by PCR and other amplification methods typically primed by primers from primer binding sites in the adapters flanking the DNA molecule to be amplified. The amplification methods can include cycles of extension, denaturation, and annealing generated by thermal cycling, or can be isothermal cycling such as in transcription-mediated amplification. Other amplification methods include ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and self-sustained sequence replication.

[0132] One or more amplifications can be applied to introduce barcodes into nucleic acid molecules using conventional nucleic acid amplification methods. The amplification can be carried out in one or more reaction mixtures. Molecular tags and sample indices / tags can be introduced simultaneously or in any order. Molecular tags and sample indices / tags can be introduced before and / or after sequence capture. In some cases, only molecular tags are introduced before probe capture, while sample indices / tags are introduced after sequence capture. In some cases, both molecular tags and sample indices / tags are introduced before probe capture. In some cases, sample indices / tags are introduced after sequence capture. Typically, sequence capture includes introducing single-stranded nucleic acid molecules complementary to the target sequence (e.g., the coding sequence of a genomic region and mutations in such regions are associated with cancer types). Typically, the amplification produces more than one non-uniquely or uniquely tagged nucleic acid amplicon, where the molecular tags and sample indices / tags range in size from 200 nt to 700 nt, 250 nt to 350 nt, or 320 nt to 550 nt. In some embodiments, the amplicon has a size of about 300 nt. In some embodiments, the amplicon has a size of about 500 nt.

[0133] Barcode

[0134] Barcodes can be incorporated into adapters by methods such as chemical synthesis, ligation, overlap extension PCR, or otherwise linked to the adapter. Generally, the assignment of unique or non-unique barcodes in the reaction follows the methods and systems described in U.S. Patent Application 20010053519, 20110160078, and U.S. Patent Nos. 6,582,908, 7,537,898, and US 9,598,731.

[0135] Labels can be linked to the sample nucleic acids randomly or non-randomly. In some cases, they are introduced at a ratio of the expected identifier (i.e., a combination of barcodes) to the microwells. The set of barcodes can be unique, e.g., all barcodes have different nucleotide sequences. The set of barcodes can be non-unique, i.e., some barcodes have the same nucleotide sequence and some barcodes have different nucleotide sequences. For example, identifiers can be loaded such that more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000, or 1,000,000,000 identifiers are loaded per genomic sample. In some cases, identifiers can be loaded such that fewer than 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000, or 1,000,000,000 identifiers are loaded per genomic sample. In some cases, the average number of identifiers loaded per genomic sample is less than or greater than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000, or 1,000,000,000 identifiers per genomic sample.

[0136] The preferred form uses 20 - 50 different tags attached to both ends of the target molecule, generating 20 - 50×20 - 50 tags, i.e., 400 - 2500 tag combinations. Such a number of tags is sufficient to give different molecules with the same start and end points a high probability (e.g., at least 94%, 99.5%, 99.99%, 99.999%) of receiving different combinations of tags.

[0137] In some cases, the identifier can be a predefined sequence oligonucleotide, or a random sequence oligonucleotide or a semi - random sequence oligonucleotide. In other cases, more than one barcode can be used such that the barcodes need not be unique relative to each other among the more than one barcode. In this instance, the barcodes can be attached (e.g., by ligation or PCR amplification) to individual molecules such that the combination of the barcode and the sequence to which it can be attached generates a unique sequence that can be traced individually. As described herein, the detection of non - uniquely tagged barcodes in combination with the start (initiation) and / or end (termination) genomic coordinates of a particular sequencing sample molecule (i.e., excluding sequence information obtained from barcodes, adaptors, etc.) can allow for the assignment of a unique identity to a particular molecule. The length or number of base pairs of an individual sequencing sample molecule (i.e., excluding sequence information corresponding to barcodes, adaptors, etc.) can also be used to assign a unique identity to such a molecule. As described herein, fragments from a single - stranded nucleic acid that has been assigned a unique identity can thereby allow for the subsequent identification of fragments from the parental strand and / or complementary strand.

[0138] Sequencing pipeline

[0139] It is possible to sequence a sample nucleic acid flanked by adapters that has or has not been pre-amplified, such as by one or more sequencing devices 107. Sequencing methods include, for example, Sanger sequencing, high-throughput sequencing, pyrosequencing, sequencing by synthesis, single molecule sequencing, nanopore sequencing, semiconductor sequencing, ligation sequencing, hybridization sequencing, RNA-Seq (Illumina), digital gene expression (Helicos), next-generation sequencing, single molecule synthesis sequencing (SMSS) (Helicos), massively parallel sequencing, clonal single molecule array (Solexa), shotgun sequencing, Ion Torrent, Oxford nanopore, Roche Genia, Maxim-Gilbert sequencing, primer walking, sequencing using PacBio, SOLiD, Ion Torrent, or nanopore platforms. The sequencing reaction can be carried out in a variety of sample processing units, which can be multi-lane, multi-channel, multi-well, or other devices that can process more than one group of samples substantially simultaneously. The sample processing unit can also include more than one sample chamber, enabling more than one run to be processed simultaneously.

[0140] A sequencing reaction can be performed on one or more fragment types known to contain markers for cancer or other diseases. The sequencing reaction can also be performed on any nucleic acid fragment present in the sample. The sequencing reaction can provide sequencing of at least 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9%, or 100% of a given genome. In other cases, the sequencing reaction can provide sequencing of less than 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9%, or 100% of a given genome.

[0141] Simultaneous sequencing reactions can be performed using multiplex sequencing. In some cases, cell-free polynucleotides can be sequenced with at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. In other cases, cell-free polynucleotides can be sequenced with fewer than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. The sequencing reactions can be performed sequentially or simultaneously. Subsequent data analysis can be performed on all or a portion of the sequencing reactions. In some cases, data analysis can be performed on at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. In other cases, data analysis is performed on fewer than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. An exemplary read depth is 1000 - 50000 reads / locus (base).

[0142] Sequence analysis pipeline

[0143] Nucleotide variations in the sequenced nucleic acid can be determined by comparing the sequenced nucleic acid to a reference sequence. The reference sequence is typically a known sequence, e.g., a known whole genome sequence or partial genome sequence from a subject, a whole genome sequence of a human subject. The reference sequence can be hG19. As described above, the sequenced nucleic acid can represent the directly determined sequence of the nucleic acid in the sample, or the consensus sequence of the amplification products of such nucleic acids. The comparison can be made at one or more specified positions on the reference sequence. When the corresponding sequences are maximally aligned, a subset of the sequenced nucleic acids can be identified, including the positions corresponding to the specified positions on the reference sequence. In such a subset, it can be determined which (if any) of the sequenced nucleic acids include nucleotide variations at the specified positions, and optionally which (if any) include the reference nucleotides (i.e., the same as in the reference sequence). If the number of sequenced nucleic acids in the subset that include nucleotide variations exceeds a threshold, the variant nucleotide can be determined at the specified position. The threshold can be a simple number, such as at least 1, 2, 3, 4, 5, 6, 7, 9, or 10 sequenced nucleic acids in the subset that include nucleotide variations, or the threshold can be the ratio of the sequenced nucleic acids in the subset that include nucleotide variations, such as at least 0.5, 1, 2, 3, 4, 5, 10, 15, or 20, and other possibilities. The comparison can be repeated for any specified position of interest in the reference sequence. Sometimes the comparison can be made for specified positions that occupy at least 20, 100, 200, or 300 consecutive positions on the reference sequence (e.g., 20 - 500 or 50 - 300 consecutive positions).

[0144] The methods of the invention can also be used to diagnose the presence or absence of a condition, particularly cancer, in a subject, to characterize a condition (e.g., stage a cancer or determine the heterogeneity of a cancer), to monitor the response to treatment of a condition, and to achieve a prognosis of the risk of developing a condition or the subsequent course of a condition.

[0145] Multiple cancers can be detected using the methods of the invention. Cancer cells, like most cells, can be characterized by a rate of turnover, where old cells die and are replaced by newer cells. Typically, dead cells in a particular subject that are in contact with the vascular system can release DNA or fragments of DNA into the bloodstream. This is also true for cancer cells during various stages of the disease. Cancer cells can also be characterized by various genetic aberrations such as copy number variations and rare mutations depending on the stage of the disease. This phenomenon can be used to detect the presence or absence of cancer in an individual using the methods and systems described herein.

[0146] The types and numbers of cancers that can be detected can include blood cancer, brain cancer, lung cancer, skin cancer, nasal cancer, laryngeal cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, skin cancer, intestinal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, gastric cancer, solid tumors, heterogeneous tumors, homogeneous tumors, etc.

[0147] Cancers can be detected based on genetic variations including the following: mutations, rare mutations, insertions / deletions (indels), copy number variations, transversions, translocations, inversions, deletions, aneuploidy, segmental aneuploidy, polyploidy, chromosomal instability, chromosomal structural alterations, gene fusions, chromosomal fusions, gene truncations, gene amplifications, gene duplications, chromosomal damage, DNA damage, abnormal alterations in nucleic acid chemical modifications, abnormal alterations in epigenetic patterns.

[0148] Genetic data can also be used to characterize specific forms of cancer. Cancers are generally heterogeneous both in composition and stage. Genetic profile data can allow for the characterization of specific subtypes of cancer, which may be important in the diagnosis or treatment of that specific subtype. This information can also provide clues to the prognosis of a specific type of cancer for a subject or practitioner, and allow the subject or practitioner to adjust treatment options based on the progression of the disease. Some cancers progress and become more aggressive and genetically unstable. Other cancers can remain benign, inactive, or dormant. The systems and methods of the present disclosure can be used to determine disease progression.

[0149] The analysis of the present invention can also be used to determine the efficacy of specific treatment options. If a treatment is successful, then a successful treatment option can increase the amount of copy number variations or rare mutations detected in a subject's blood as more cancer cells may die and shed DNA. In other instances, this may not occur. In another instance, perhaps certain treatment options may be associated with the genetic profile of the cancer over time. This correlation can be used to select therapies. Additionally, if a cancer is observed to be in remission after treatment, the methods of the present invention can be used to monitor for residual disease or recurrence of the disease.

[0150] The methods of the present invention can also be used to detect genetic variations in conditions other than cancer. After the occurrence of certain diseases, immune cells, such as B cells, can undergo rapid clonal expansion. Copy number variation detection can be used to monitor clonal expansion, and certain immune states can be monitored. In this instance, copy number variation analysis can be performed over time to generate a profile of how a particular disease might progress. Copy number variations or even rare mutation detection can be used to determine how a pathogen population changes during the course of an infection. This can be particularly important during chronic infections, such as HIV / AIDS or hepatitis infections, where the virus can change its life cycle state and / or mutate into a more virulent form during the course of the infection. When immune cells attempt to destroy transplanted tissue, the methods of the present invention can be used to determine or dissect the rejection activity of the host body to monitor the status of the transplanted tissue and alter the course of rejection treatment or prevention.

[0151] In addition, the methods of the present disclosure can be used to characterize the heterogeneity of an abnormal condition of a subject, the method comprising generating a genetic profile of extracellular polynucleotides of the subject, wherein the genetic profile comprises more than one data obtained from copy number variation and rare mutation analysis. In some cases, including but not limited to cancer, diseases can be heterogeneous. Lesion cells can be different. In the instance of cancer, some tumors are known to contain different types of tumor cells, and some cells are at different stages of cancer. In other instances, heterogeneity can include multiple foci of a disease. Again, in the instance of cancer, there can be multiple tumor foci, perhaps one or more of which are the result of metastases that have spread from the primary site.

[0152] The methods of the present invention can be used to generate or dissect a fingerprint or data set that is the sum of genetic information derived from different cells in a heterogeneous disease. The data set can include copy number variation and rare mutation analysis, either alone or in combination.

[0153] The methods of the present invention can be used to diagnose, prognose, monitor, or observe cancer or other diseases of fetal origin. That is, these methods can be used in pregnant subjects to diagnose, prognose, monitor, or observe cancer or other diseases of the unborn subject, the DNA and other polynucleotides of which can circulate co - with maternal molecules.

[0154] Exemplary precision therapies and applications

[0155] The accurate diagnosis provided by computer system 700 can result in an accurate treatment plan that is identified by computer system 700 (and / or selected by a healthcare professional). For example, in the case of lung cancer and other diseases, the goal can be to ensure that there are no better treatment options based on the presence of specific mutations. For example, EGFR (L858R, exon 19 deletion), BRAF V600E, ALK, and ROS1 fusions can be treated with targeted therapies that may be more appropriate than platinum therapy and chemotherapy. Although these are examples of primary drivers, there are other targetable drivers such as MET exon 14 skipping. In another example, for colon cancer, the goal can be to avoid ineffective treatments. If KRAS orNRAS is wild-type, chemotherapy using FOLFIRI or chemotherapy using an irinotecan regimen can be supplemented with cetuximab or panitumumab. Thus, the confidence that KRAS andNRAS are wild-type will increase the confidence that adding cetuximab or panitumumab is the correct treatment choice and may obviate the need for further testing. The biological explanation for this is that cetuximab or panitumumab targets EGFR and inhibits its activity. RAS (K / NRAS) is downstream of EGFR, so if RAS is activated, inhibiting EGFR will have minimal or no effect, and thus cetuximab or panitumumab treatment will be inappropriately administered.

[0156] The mutations analyzed by the methods and systems of the present disclosure can be loss-of-function mutations such as ATM. For example, DNA damage repair (DDR) is a cellular process for maintaining genomic integrity or stability. A defect or lack in a given DDR mechanism can lead to tumorigenesis or other diseases and can be used to identify test subjects or patients who may benefit from a given targeted therapy. As an example, homologous recombination repair deficiency (HRD) is a cellular phenotype that can make a patient a candidate for administration of a therapeutic agent such as a poly ADP ribose polymerase (PARP) inhibitor. In certain embodiments, a therapy comprising at least one PARP inhibitor can be administered to a subject, wherein it has been determined using the methods and systems described herein that the mutation is tumor-derived or non-tumor-derived. In certain embodiments, the PARP inhibitor can include olaparib, talazoparib, rucaparib, niraparib (trade name ZEJULA), etc. In some embodiments, the therapy comprises at least one base excision repair (BER) inhibitor. For example, olaparib can inhibit BER. In certain embodiments, based on determining using the methods and systems described herein that a subject has a tumor-derived or non-tumor-derived mutation, administration of the therapy to the subject can be discontinued.

[0157] Non-tumor variants may affect the determination of the tumor mutational burden (TMB) score, which, if not removed or filtered from the TMB determination, will result in a spurious high score. The TMB score is commonly used to predict whether a patient will respond to immunotherapy treatment. Accordingly, the methods and systems provided herein can be used to distinguish tumor-derived or non-tumor-derived variants as part of the TMB calculation, such as those described in PCT / US2019 / 042882, which is incorporated herein by reference. In another aspect, the present disclosure provides a method of classifying a subject as a candidate for immunotherapy by determining whether the subject has a tumor-derived or non-tumor-derived variant. In certain embodiments, the methods of the present disclosure include administering one or more immunotherapies to a subject based on determining whether a variant is tumor-derived or non-tumor-derived using the methods or systems disclosed herein alone or in combination with methods for determining the TMB score. In some embodiments, the immunotherapy includes at least one checkpoint inhibitor antibody. In some embodiments, the immunotherapy includes an antibody against: PD-1, PD-2, PD-L1, PD-L2, CTLA-40, OX40, B7.1, B7He, LAG3, CD137, KIR, CCR5, CD27, or CD40. In some embodiments, the immunotherapy includes administering a pro-inflammatory cytokine against at least one tumor type. In some embodiments, the immunotherapy includes administering T cells against at least one tumor type. In some embodiments, the subject is administered a combination therapy (e.g., immunotherapy + PARPi + chemotherapy, etc.), as well as many other therapies further exemplified herein or otherwise known to those of ordinary skill in the art.

[0158] The methods and systems provided herein can be used to evaluate the prognostic value of mutations with respect to survival or response to treatment. For example, the prognostic and predictive value of TP53 mutations for treatment with an ALK inhibitor can be evaluated. The determination of the tumor-derived / non-tumor-derived origin of the variants analyzed herein can also be used for subject recruitment for the selection of therapies (e.g., TP53 drugs). Other applications of the methods and systems herein can be used to analyze less studied mutations (e.g., FGFR2 mutations for FGFR inhibitors, or ERBB2 for ERBB2 inhibitors), where distinguishing the variant or tumor-derived or non-tumor-derived can provide confidence as to whether the variant originated from the tumor. In certain embodiments, the methods and systems described herein can be used to monitor molecular response by tracking only tumor variants to determine the dynamics of the variants over time.

[0159] As additional therapies for various diseases are developed, the interpretation of negative predictions will become increasingly complex and increasingly important in designing precision therapies. Examples

[0160] Example 1:Variant calls were obtained from >250,000 plasma samples, which included the Guardant360TM, GuardantREVEALTM, GuardantOMNITM, and GuardantInfinityTM liquid biopsy panels, as well as healthy donors, early cancer patients, and late-stage cancer patients sequenced on public tissue datasets. The model was trained on paired plasma and WBC datasets and optimized with 10-fold cross-validation to produce non-tumor and tumor variant classifiers. To validate these calls, an independent cohort of 72 paired plasma and WBC late-stage cancer samples was genotyped on the GuardantInfinityTM assay. A cohort of 76 healthy donor samples genotyped on the GuardantOMNI assay was also evaluated.

[0161] Example 2: The integrated model was trained on a database of >250,000 plasma samples, which included the Guardant360TM, GuardantREVEALTM, and GuardantOMNITM liquid biopsy panels, as well as healthy donors, early cancer patients, and late-stage cancer patients sequenced on public tissue datasets. The model was optimized with 5-fold cross-validation and hyperparameter tuning to produce non-tumor and tumor variant classifiers. To validate these calls, 116 paired plasma and WBC late-stage cancer clinical samples with a high incidence of putative CH variants were selected and sequenced and genotyped using an in-house bioinformatics pipeline. In the validation cohort, cfDNA variants were determined to be of non-tumor origin or CH origin if there was sufficient molecular support in the WBC; cfDNA variants above 0.6% (the limit of detection in gDNA) without support in the WBC were determined to be from the tumor.

[0162] Example 3:The validation cohort consisted of 2150 somatic SNVs and indels, of which 956 were confirmed in WBCs and 1194 were confirmed only in plasma. Half of the confirmed CH variants (48%, 458 / 956) occurred in known CH genes (such as DNMT3A, TET2, PPM1D), while the other half occurred in genes such as TP53, ATM, NOTCH4, FAT1, SRSF2. No clinically actionable variants were identified in WBCs. Non-tumor or CH prediction was performed on 624 somatic variants; among them, 515 / 624 were correctly identified as CH, with a positive predictive value (PPV) of 83%. Among all CH variants confirmed in WBCs, 54% (553 / 956) had CH or non-CH prediction; the CH prediction had a positive percent agreement (PPA) of 91% (515 / 553) with WBCs. The remaining variants without CH prediction (403 / 956) had low or no incidence in the dataset and mainly occurred in LRP1B, TET2, TP53, KMT2D. Nearly half (67%, n = 109) of the CH predictions not in WBCs occurred in CH genes. For non-CH gene variants, 16% of the false positive predictions occurred in 6 variants across 4 genes (ACVR2A, RNF43, B2M, FLT3).

[0163] Example 4: We present a plasma-only method for classifying non-tumor, CH variants in cfDNA that has high PPA and PPV with WBC genotyping. Further studies are underway to improve the sensitivity for annotating rare CH variants. Accurate CH identification is crucial for treatment selection across targeted therapies, especially for loss-of-function variants in DNA repair genes that may confer sensitivity to PARPi or ATRi therapies.

[0164] All patent applications, websites, other publications, accession numbers, etc., cited above or below are hereby incorporated by reference in their entirety for all purposes to the extent that each individual item is specifically and individually indicated to be so incorporated by reference. If different versions of a sequence are associated with an accession number at different times, the version associated with the accession number on the effective filing date of the present application is meant. The effective filing date means the earlier of the actual filing date of the application citing the accession number or the filing date of the priority application (if applicable). Similarly, if different versions of a publication, website, etc. are published at different times, the most recently published version at the effective filing date of the application is meant, unless otherwise indicated. Any feature, step, element, embodiment or aspect of the present disclosure may be used in combination with any other feature, step, element, embodiment or aspect, unless otherwise specifically indicated. Although the present disclosure has been described in considerable detail by way of illustration and example for purposes of clarity and understanding, it will be apparent that certain changes and modifications may be practiced within the scope of the appended claims.

Claims

1. A method of at least partially using a computer to distinguish tumor-derived nucleic acid variants from non-tumor-derived nucleic acid variants in a cell-free nucleic acid (cfNA) sample obtained from a test subject, the method comprising: generating or providing, by the computer, at least one tumor variant data set comprising a reference population of tumor-related genetic variants, wherein the tumor variant data set comprises observed frequency data of one or more tumor-related genetic variants in the reference population of tumor-related genetic variants in a reference sample, the reference sample comprising a reference plasma-only sample and / or a reference white blood cell sample, and wherein the reference sample is obtained from a single reference subject and / or from different reference subjects having the same cancer type; determining, by the computer, one or more ratios of the observed frequency data of one or more tumor-related genetic variants in the reference population of tumor-related genetic variants between the reference samples to produce at least one MAF variance and / or relative incidence data set; generating, by the computer, at least one non-tumor-derived probability set from the MAF variance and / or relative incidence data set; and using the non-tumor-derived probability set to classify nucleic acid variants detected in the cfNA sample obtained from the test subject as tumor-derived nucleic acid variants or non-tumor-derived nucleic acid variants.

2. A method of at least partially using a computer to distinguish tumor-derived nucleic acid variants from non-tumor-derived nucleic acid variants in a cell-free nucleic acid (cfNA) sample obtained from a test subject, the method comprising: determining, by the computer, the relative incidence of one or more tumor-related genetic variants observed in one or more reference plasma-only samples compared to one or more reference white blood cell samples to produce at least one relative incidence data set; generating, by the computer, at least one non-tumor-derived probability set from the relative incidence data set; and using the non-tumor-derived probability set to classify nucleic acid variants detected in the cfNA sample obtained from the test subject as tumor-derived nucleic acid variants or non-tumor-derived nucleic acid variants.

3. A method of at least partially using a computer to distinguish tumor-derived nucleic acid variants from non-tumor-derived nucleic acid variants in a cell-free nucleic acid (cfNA) sample obtained from a test subject, the method comprising: determining, by the computer, the change in the mutant allele fraction (MAF) value and / or at least one statistic associated therewith for each of one or more tumor-related genetic variants and / or non-tumor-related genetic variants at at least two different time points to produce at least one relative incidence data set; generating, by the computer, at least one non-tumor-derived probability set from the relative incidence data set; and using the non-tumor-derived probability set to classify nucleic acid variants detected in the cfNA sample obtained from the test subject as tumor-derived nucleic acid variants or non-tumor-derived nucleic acid variants.

4. A method for at least partially using a computer to distinguish tumor-derived nucleic acid variants and non-tumor-derived nucleic acid variants in a cell-free nucleic acid (cfNA) sample obtained from a test subject, the method comprising: When the incidence rate of a first nucleic acid variant detected in the cfNA sample is less than a probability threshold from a non-tumor-derived probability set, the computer classifies at least the first nucleic acid variant detected in the cfNA sample obtained from the test subject as a tumor-derived nucleic acid variant, and when the incidence rate of a second nucleic acid variant detected in the cfNA sample is greater than the probability threshold from the non-tumor-derived probability set, the computer classifies at least the second nucleic acid variant detected in the cfNA sample obtained from the test subject as a non-tumor-derived nucleic acid variant, thereby distinguishing the tumor-derived nucleic acid variants and non-tumor-derived nucleic acid variants in the cfNA sample obtained from the test subject, wherein the non-tumor-derived probability set is generated by: The computer generates or provides at least one tumor variant data set including a reference tumor-related genetic variant population, wherein the tumor variant data set includes observed frequency data of one or more tumor-related genetic variants in the reference tumor-related genetic variant population in a reference sample, the reference sample includes a reference plasma-only sample and / or a reference white blood cell sample, and wherein the reference sample is obtained from a single reference subject and / or from different reference subjects having the same cancer type; The computer determines one or more ratios of the observed frequency data of one or more tumor-related genetic variants in the reference tumor-related genetic variant population between the reference samples to generate at least one relative incidence data set; and The computer generates the non-tumor-derived probability set from the relative incidence data set.

5. A method for generating a classifier that at least partially uses a computer to classify nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as tumor-derived nucleic acid variants or non-tumor-derived nucleic acid variants, the method comprising: The computer generates or provides at least one tumor variant data set including a reference tumor-related genetic variant population, wherein the tumor variant data set includes observed frequency data of one or more tumor-related genetic variants in the reference tumor-related genetic variant population in a reference sample, the reference sample includes a reference plasma sample and / or a reference white blood cell sample, and wherein the reference sample is obtained from a single reference subject and / or from different reference subjects having the same cancer type; The computer determines one or more ratios of the observed frequency data of one or more tumor-related genetic variants in the reference tumor-related genetic variant population between the reference samples to generate at least one relative incidence data set; and The computer applies at least one machine learning model to the relative incidence dataset to generate at least one set of non-tumor origin probabilities, thereby generating a classifier that classifies the nucleic acid variants to be detected in the cfNA sample as tumor-origin nucleic acid variants or non-tumor origin nucleic acid variants.

6. A method for at least partially using a computer to distinguish tumor-origin nucleic acid variants and non-tumor origin nucleic acid variants in a cell-free nucleic acid (cfNA) sample obtained from a test subject having a cancer type, the method comprising: The computer determines the incidence of one or more genetic variants observed in the cfNA sample to generate a test subject incidence dataset; The computer compares the incidence of one or more genetic variants in the test subject incidence dataset with the incidence of genetic variants observed in a reference cfNA sample obtained from a reference subject having the cancer type; And When the incidence of a specific genetic variant in the test subject incidence dataset is lower than a predetermined threshold associated with the specific genetic variant in the reference cfNA sample obtained from a reference subject having the cancer type, the computer classifies the specific genetic variant in the test subject incidence dataset as a non-tumor origin nucleic acid variant, thereby distinguishing tumor-origin nucleic acid variants and non-tumor origin nucleic acid variants in the cfNA sample obtained from the test subject having the cancer type.

7. A method for at least partially using a computer to distinguish tumor-origin nucleic acid variants and non-tumor origin nucleic acid variants in a cell-free nucleic acid (cfNA) sample obtained from a test subject, the method comprising: The computer determines the incidence of one or more genetic variants observed in the cfNA sample to generate a test subject incidence dataset; The computer compares the incidence of one or more genetic variants in the test subject incidence dataset with the incidence of genetic variants observed in a reference cfNA sample obtained from a reference subject having leukemia, lymphoma, and / or hematological malignancy; and When the incidence of a specific genetic variant in the test subject incidence dataset is higher than a predetermined threshold associated with the specific genetic variant in the reference cfNA sample obtained from a reference subject having leukemia, lymphoma, and / or hematological malignancy, the computer classifies the specific genetic variant in the test subject incidence dataset as a non-tumor origin nucleic acid variant, thereby distinguishing tumor-origin nucleic acid variants and non-tumor origin nucleic acid variants in the cfNA sample obtained from the test subject having leukemia, lymphoma, and / or hematological malignancy.

8. The method according to any one of the preceding claims, comprising identifying the genetic variants present in the cfNA sample from sequencing reads derived from cfNA molecules in the cfNA sample.

9. The method according to any one of the preceding claims, wherein the sequencing reads are obtained from a targeted segment of cfNA molecules in the cfNA sample.

10. The method according to any one of the preceding claims, wherein the reference population of tumor-related genetic variations is obtained from a reference sample.

11. The method according to any one of the preceding claims, comprising randomly dividing the tumor variant data set into a training data set and a test data set.

12. The method according to any one of the preceding claims, wherein the training data set comprises approximately 80% of the tumor variant data set, and the test data set comprises approximately 20% of the tumor variant data set.

13. The method according to any one of the preceding claims, wherein the tumor variant data set comprises observed frequency data of one or more tumor-related genetic variations in a reference population of tumor-related genetic variations in a reference sample of a specific cancer type.

14. The method according to any one of the preceding claims, comprising using at least a portion of the reference population of tumor-related genetic variations to train a machine learning model to produce a trained machine learning model, wherein the tumor-derived nucleic acid variations and non-tumor-derived nucleic acid variations detected in the cfNA sample obtained from a test subject are distinguished from each other using the trained machine learning model.

15. The method according to any one of the preceding claims, wherein the machine learning model is trained using one or more of the following: logistic regression, probit regression, decision tree, random forest, gradient boosting, support vector machine, K-nearest neighbor, and neural network.

16. The method according to any one of the preceding claims, comprising using a probability threshold of at least approximately the 30th percentile for a specific genetic variation as a cut-off value for classification.

17. The method according to any one of the preceding claims, comprising performing logistic regression on at least one of the ratios to obtain a specific non-tumor-derived probability.

18. The method according to any one of the preceding claims, wherein the tumor variant data set comprises mutant allele fraction data observed in a reference sample for one or more tumor-related genetic variations in a reference population of tumor-related genetic variations.

19. The method according to any one of the preceding claims, comprising normalizing the tumor variant data set using one or more data normalization techniques.

20. The method according to any one of the preceding claims, wherein the data normalization technique comprises min-max normalization and / or z-score normalization.

21. The method according to any one of the preceding claims, wherein the reference non-body fluid sample comprises a reference tumor tissue sample and / or a reference white blood cell sample.

22. The method according to any one of the preceding claims, wherein a ratio of the observed frequency data of a specific genetic variation in a reference plasma sample only to the observed frequency data of the specific genetic variation in a reference white blood cell sample greater than one (1.0) indicates that the specific genetic variation may be a non-tumor-derived nucleic acid variation.

23. The method according to any one of claims 1-20, wherein a ratio of the observed frequency data of a specific genetic variation in a plasma fluid sample only to the observed frequency data of the specific genetic variation in a reference non-body fluid sample less than one (1.0) indicates that the specific genetic variation may be a non-tumor-derived nucleic acid variation.

24. The method according to any one of the preceding claims, wherein the non-tumor origin probability set comprises at least one clonal hematopoiesis origin probability set.

25. The method according to any one of the preceding claims, comprising obtaining the cfNA sample from a test subject.

26. The method according to any one of the preceding claims, comprising selecting one or more therapies to treat the cancer type when one or more tumor-origin nucleic acid variants associated with the cancer type are detected in the cfNA sample obtained from the test subject.

27. The method according to any one of the preceding claims, comprising administering one or more therapies to the test subject to treat the cancer type when one or more tumor-origin nucleic acid variants associated with the cancer type are detected in the cfNA sample obtained from the test subject.

28. The method according to any one of the preceding claims, wherein the cancer type is selected from the group consisting of: biliary tract cancer, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, glioma, astrocytoma, breast cancer, metaplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal cancer, colon cancer, hereditary non-polyposis colorectal cancer, colorectal adenocarcinoma, gastrointestinal stromal tumor (GIST), endometrial cancer, endometrial stromal sarcoma, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, ocular melanoma, uveal melanoma, gallbladder cancer, gallbladder adenocarcinoma, renal cell carcinoma, clear cell renal cell carcinoma, transitional cell carcinoma, urothelial carcinoma, Wilms tumor, leukemia, acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), chronic myelomonocytic leukemia (CMML), liver cancer, hepatic epithelial carcinoma, hepatocellular adenoma, hepatocellular carcinoma, cholangiocarcinoma, hepatoblastoma, lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B-cell lymphoma, non-Hodgkin lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, T-cell lymphoma, non-Hodgkin lymphoma, precursor T-lymphoblastic lymphoma / leukemia, peripheral T-cell lymphoma, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oral cancer, oral squamous cell carcinoma, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic ductal adenocarcinoma, solid pseudopapillary tumor, acinar cell carcinoma, prostate cancer, prostatic adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine cancer, gastric cancer, gastric epithelial carcinoma, gastrointestinal stromal tumor (GIST), uterine cancer, and uterine sarcoma.

29. The method according to any one of the preceding claims, wherein the reference tumor-related genetic variants are selected from the group consisting of: single nucleotide variants (SNVs), insertions or deletions (indels), copy number variants (CNVs), fusions, transversions, translocations, frameshifts, duplications, repeat expansions, and epigenetic variants.

30. The method according to any one of the preceding claims, wherein the reference sample comprises at least about 25, at least about 50, at least about 100, at least about 200, at least about 300, at least about 400, at least about 500, at least about 600, at least about 700, at least about 800, at least about 900, at least about 1,000, at least about 5,000, at least about 10,000, at least about 15,000, at least about 20,000, at least about 25,000, at least about 30,000 or more individual body fluid samples and / or non-body fluid samples.

31. The method according to any one of the preceding claims, wherein the cfNA sample comprises cell-free deoxyribonucleic acid (cfDNA).

32. The method according to any one of the preceding claims, wherein the cfNA sample comprises cell-free ribonucleic acid (cfRNA).

33. The method according to any one of the preceding claims, wherein the test subject is a mammalian subject.

34. The method according to any one of the preceding claims, wherein the test subject is a human subject.

35. The method according to any one of the preceding claims, wherein the reference body fluid sample comprises a plasma sample.

36. The method according to any one of the preceding claims, wherein the reference body fluid sample comprises a serum sample.

37. The method according to any one of the preceding claims, wherein the reference non-body fluid sample comprises a cell sample.

38. The method according to any one of the preceding claims, wherein the reference non-body fluid sample comprises a tissue sample.

39. The method according to any one of the preceding claims, wherein the method for distinguishing tumor-derived nucleic acid variants and non-tumor-derived nucleic acid variants in a cell-free nucleic acid (cfNA) sample is at least partially based on: (i) the consistency of the incidence of nucleic acid variants across cancer types; (ii) the change over time of the mutant allele fraction (MAF) of the nucleic acid variant; and / or (iii) the incidence of the nucleic acid variant in blood cancers, such as leukemia, lymphoma, and / or hematological malignancies.

40. A system, the system comprising a controller, the controller comprising a computer-readable medium or being able to access a computer-readable medium, the computer-readable medium comprising non-transitory computer-executable instructions, which when executed by at least one electronic processor, at least perform: (a) generating or providing at least one tumor variant data set comprising a reference tumor-associated genetic variant population, wherein the tumor variant data set comprises observed frequency data of one or more tumor-associated genetic variants in the reference tumor-associated genetic variant population in a reference sample, the reference sample comprising a reference only plasma sample and / or a reference white blood cell sample, and wherein the reference sample is obtained from a single reference subject and / or obtained from different reference subjects having the same cancer type; (b) Determine one or more ratios of the observed frequency data of one or more tumor-related genetic variations in the reference tumor-related genetic variation population between the reference samples to generate at least one relative incidence data set; and (c) Apply at least one machine learning model to the relative incidence data set to generate at least one non-tumor origin probability set to generate a classifier that classifies nucleic acid variations detected in a cell-free nucleic acid (cfNA) sample as tumor-origin nucleic acid variations or non-tumor origin nucleic acid variations.

41. A system, the system includes a controller, the controller includes a computer-readable medium or is capable of accessing a computer-readable medium, the computer-readable medium includes non-transitory computer-executable instructions, when the non-transitory computer-executable instructions are executed by at least one electronic processor, at least perform: (a) Determine the relative incidence of one or more tumor-related genetic variations observed in one or more reference plasma-only samples compared to one or more reference white blood cell samples to generate at least one relative incidence data set; and (b) Generate at least one non-tumor origin probability set from the relative incidence data set to generate a classifier that classifies nucleic acid variations detected in a cell-free nucleic acid (cfNA) sample as tumor-origin nucleic acid variations or non-tumor origin nucleic acid variations.

42. A system, the system includes a controller, the controller includes a computer-readable medium or is capable of accessing a computer-readable medium, the computer-readable medium includes non-transitory computer-executable instructions, when the non-transitory computer-executable instructions are executed by at least one electronic processor, at least perform: (a) The computer determines the change in the mutant allele fraction (MAF) value and / or at least one statistic associated therewith for each of one or more tumor-related genetic variations and / or non-tumor-related genetic variations at at least two different time points to generate at least one MAF variance and / or relative incidence data set; and (b) Generate at least one non-tumor origin probability set from the MAF variance and / or relative incidence data set to generate a classifier that classifies nucleic acid variations detected in a cell-free nucleic acid (cfNA) sample as tumor-origin nucleic acid variations or non-tumor origin nucleic acid variations.

43. The system according to any one of the preceding claims, including a nucleic acid sequencer operably connected to the controller, the nucleic acid sequencer being configured to provide sequencing reads derived from cfNA molecules in the cfNA sample.

44. The system according to any one of the preceding claims, wherein the nucleic acid sequencer or another system component is configured to group the sequence reads generated by the nucleic acid sequencer into sequence read families, each family containing sequence reads generated from a specific cfNA molecule in the cfNA sample.

45. The system according to any one of the preceding claims, comprising a database operably connected to the controller, the database including one or more therapies indexed to the tumor-derived nucleic acid variants.

46. The system according to any one of the preceding claims, comprising a sample preparation assembly operably connected to the controller, the sample preparation assembly being configured to prepare the cfNA molecules in the cfNA sample to be sequenced by a nucleic acid sequencer.

47. The system according to any one of the preceding claims, comprising a nucleic acid amplification assembly operably connected to the controller, the nucleic acid amplification assembly being configured to amplify at least a targeted segment of the cfNA molecules in the cfNA sample.

48. The system according to any one of the preceding claims, comprising a material transfer assembly operably connected to the controller, the material transfer assembly being configured to transfer one or more materials between at least a nucleic acid sequencer and a sample preparation assembly.

49. A computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least: (a) generating or providing at least one tumor variant data set comprising a reference tumor-associated genetic variant population, wherein the tumor variant data set comprises observed frequency data of one or more tumor-associated genetic variants in the reference tumor-associated genetic variant population in a reference sample, the reference sample comprising a reference plasma-only sample and / or a reference white blood cell sample, and wherein the reference sample is obtained from a single reference subject and / or from different reference subjects having the same cancer type; (b) determining one or more ratios of the observed frequency data of one or more tumor-associated genetic variants in the reference tumor-associated genetic variant population between the reference samples to produce at least one relative incidence data set; and (c) applying at least one machine learning model to the relative incidence data set to produce at least one non-tumor-derived probability set to generate a classifier that classifies nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as tumor-derived nucleic acid variants or non-tumor-derived nucleic acid variants.

50. A computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least: (a) determining the relative incidence of one or more tumor-associated genetic variants observed in one or more reference body fluid samples compared to one or more reference non-body fluid samples to produce at least one MAF variance and / or relative incidence data set; and (b) generating at least one non-tumor-derived probability set from the MAF variance and / or relative incidence data set to generate a classifier that classifies nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as tumor-derived nucleic acid variants or non-tumor-derived nucleic acid variants.

51. A computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least: (a) determining, by the computer, for each of one or more tumor-related genetic variations and / or non-tumor-related genetic variations, a change in the mutant allele fraction (MAF) value and / or at least one statistic associated therewith at at least two different time points to generate at least one relative incidence dataset; and (b) generating, from the relative incidence dataset, at least one non-tumor origin probability set to generate a classifier that classifies nucleic acid variations detected in a cell-free nucleic acid (cfNA) sample as tumor-origin nucleic acid variations or non-tumor-origin nucleic acid variations.

52. The system or computer-readable medium according to any one of the preceding claims, wherein the electronic processor further performs at least: dividing the tumor variant dataset into a training dataset and a test dataset.

53. The system or computer-readable medium according to any one of the preceding claims, wherein the electronic processor further performs at least: using at least a portion of a tumor-related genetic variation population to train a machine learning model to produce a trained machine learning model, and using the trained machine learning model to classify the nucleic acid variations detected in the cfNA sample as tumor-origin nucleic acid variations or non-tumor-origin nucleic acid variations.

54. The system or computer-readable medium according to any one of the preceding claims, wherein the electronic processor further performs at least: performing logistic regression on at least one of the ratios to obtain a specific non-tumor origin probability.

55. The system or computer-readable medium according to any one of the preceding claims, wherein the electronic processor further performs at least: normalizing the tumor variant dataset using one or more data normalization techniques.

56. The system or computer-readable medium according to any one of the preceding claims, wherein the electronic processor further performs at least: when one or more tumor-origin nucleic acid variations associated with a cancer type are detected in the cfNA sample, selecting one or more therapies to treat the cancer type.

Citation Information

Patent Citations

  • Oligonucleotides

    US20010053519A1

  • Digital Counting of Individual Molecules by Stochastic Attachment of Diverse Labels

    US20110160078A1

  • Oligonucleotides

    US6582908B2

  • Compositions and methods of selective nucleic acid isolation

    US7537898B2

  • Systems and methods to detect rare mutations and copy number variation

    US9598731B2