Method for classifying genetic mutations detected in cell-free nucleic acids as of tumor or non-tumor origin

A computational method using reference datasets and machine learning classifies tumor and non-tumor nucleic acid variants in cell-free samples, enhancing cancer detection accuracy and treatment guidance.

JP7813712B2Active Publication Date: 2026-02-13GUARDANT HEALTH INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2022553081
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-11
Filing Date
2021-03-11
Publication Date
2026-02-13
Estimated Expiration
2041-03-11

AI Technical Summary

Technical Problem

Existing liquid biopsy tests struggle to distinguish between nucleic acid variants of tumor and non-tumor origin in cell-free nucleic acid samples, leading to challenges in cancer detection sensitivity and specificity.

Method used

A method using a computer to generate datasets from reference samples, determine ratios and probabilities, and apply machine learning models to classify nucleic acid variants as tumor or non-tumor origin based on observed frequencies and mutant allele fractions.

Benefits of technology

Improves the sensitivity and specificity of cancer detection by accurately distinguishing between tumor and non-tumor nucleic acid variants, guiding effective treatment strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007813712000004
    Figure 0007813712000004
  • Figure 0007813712000005
    Figure 0007813712000005
  • Figure 0007813712000006
    Figure 0007813712000006
Patent Text Reader

Abstract

Provided herein are methods for distinguishing between nucleic acid variants of tumor origin and nucleic acid variants of non-tumor origin in cell-free nucleic acid (cfNA) samples.Certain of these methods include: generating a tumor variant dataset comprising a population of reference tumor-associated genetic variants, wherein the tumor variant dataset comprises the frequency of observation data between reference samples, comprising a reference body fluid (e.g., plasma) sample and a reference non-body fluid (e.g., non-plasma) sample, for tumor-associated genetic variants in the population of reference tumor-associated genetic variants; and determining the ratio of the frequency of observation data between reference samples for tumor-associated genetic variants in the population of reference tumor-associated genetic variants to generate a relative prevalence dataset.Further methods and related systems and computer-readable media are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of and relies on the filing date of U.S. Provisional Patent Application No. 62 / 988,306, filed March 11, 2020, the entire disclosure of which is hereby incorporated by reference herein. [Background technology]

[0002] background Liquid biopsy tests can be used to profile circulating tumor nucleic acids in patient-derived blood samples, for example, to detect cancer at an early stage, select therapy, and monitor disease progression and / or minimal residual disease. Circulating plasma cell-free tumor DNA (ctDNA) is a small DNA fragment derived from apoptotic and necrotic tumor cells or from circulating tumor cells (CTCs) introduced into the bloodstream. Although ctDNA is only a portion of the cell-free DNA (cfDNA) specifically released from cancer cells, the majority of cfDNA in a given sample typically originates from normal, non-cancerous cells, including normal white blood cells, hematopoietic stem cells (HSCs), or other early blood cell precursors that undergo apoptosis or necrosis during the clonal hematopoietic process. One problem associated with many liquid biopsy tests is distinguishing ctDNA from other cfDNA in patient samples. Summary of the Invention [Means for solving the problem]

[0003] Thus, there remains a need for methods and related embodiments for distinguishing between nucleic acid variants of tumor origin and nucleic acid variants of non-tumor origin detected in cell-free nucleic acid (cfNA) samples.

[0004] overview The present disclosure provides methods for distinguishing between nucleic acid variants of tumor and non-tumor origin in cell-free nucleic acid (cfNA) samples, which, among other properties, improve the sensitivity and specificity of cancer detection assays and guide treatment strategies. Additional methods and related systems and computer-readable media are also provided.

[0005] In some aspects, the present disclosure provides a method for distinguishing (e.g., distinguishing between) nucleic acid variants of tumor origin and nucleic acid variants of non-tumor origin in a cell-free nucleic acid (cfNA) sample obtained from a test subject, at least in part using a computer. The method includes generating or providing, by a computer, at least one tumor variant dataset comprising a population of reference tumor-associated genetic variants. The tumor variant dataset includes observed data frequencies between reference samples, including reference body fluid samples (e.g., plasma samples, serum samples, etc.) and / or reference non-body fluid samples (e.g., cell samples, tissue samples, etc.), for one or more tumor-associated genetic variants in the population of reference tumor-associated genetic variants. The reference samples are obtained from a single reference subject and / or from different reference subjects having the same cancer type. The method also includes determining, by a computer, one or more ratios of observed data frequencies between reference samples for one or more tumor-associated genetic variants in the population of reference tumor-associated genetic variants to generate at least one MAF variance and / or relative prevalence dataset. Furthermore, this method also includes generating or providing, by a computer, at least one set of probabilities of non-tumor origin from the relative prevalence data set, and using the set of probabilities of non-tumor origin to identify the nucleic acid variants detected in the cfNA sample obtained from the test subject as nucleic acid variants of tumor origin or nucleic acid variants of non-tumor origin.In some embodiments of the methods, systems, computer-readable media and other aspects of the present disclosure, one or more other features are optionally utilized in conjunction with or instead of the frequency ratio of the observed data.Some of these other features include, for example, the uniformity of prevalence across cancer types, the longitudinal variation of mutant allele fraction (MAF) over time, the proportion in hematological cancers, etc.

[0006] In another aspect, the present disclosure provides a method for distinguishing between nucleic acid variants of tumor origin and nucleic acid variants of non-tumor origin in a cell-free nucleic acid (cfNA) sample obtained from a test subject, at least in part using a computer. The method includes determining, by a computer, the relative prevalence of one or more tumor-associated genetic variants observed in one or more reference body fluid samples compared to one or more reference non-body fluid samples to generate at least one relative prevalence dataset. The method also includes generating or providing, by a computer, at least one set of probabilities of non-tumor origin from the relative prevalence dataset, and using the set of probabilities of non-tumor origin to identify nucleic acid variants detected in the cfNA sample obtained from the test subject as nucleic acid variants of tumor origin or nucleic acid variants of non-tumor origin.

[0007] In some aspects, the present disclosure provides a method for distinguishing between nucleic acid variants of tumor origin and nucleic acid variants of non-tumor origin in a cell-free nucleic acid (cfNA) sample obtained from a test subject, at least in part using a computer. The method includes determining, by a computer, the variation in mutant allele fraction (MAF) values ​​and / or at least one statistic associated therewith (e.g., the mean, standard deviation, and / or chi-squared p-value of the variant MAF over time) for each of one or more tumor-associated and / or non-tumor-associated genetic variants observed in one or more reference body fluid samples compared to one or more reference non-body fluid samples for at least two different time points to generate at least one MAF variance and / or relative prevalence dataset. Furthermore, the method also includes generating or providing, by a computer, at least one set of probabilities of non-tumor origin from the MAF variance and / or relative prevalence dataset, and using the set of probabilities of non-tumor origin to identify nucleic acid variants detected in the cfNA sample obtained from the test subject as nucleic acid variants of tumor origin or non-tumor origin.

[0008] In some aspects, the present disclosure provides a method for distinguishing between nucleic acid variants of tumor origin and nucleic acid variants of non-tumor origin in a cell-free nucleic acid (cfNA) sample obtained from a test subject, at least in part using a computer. The method includes: classifying, by a computer, at least a first nucleic acid variant detected in the cfNA sample obtained from the test subject as a nucleic acid variant of tumor origin if the prevalence of the first nucleic acid variant detected in the cfNA sample is lower than a threshold probability from a set of probabilities of non-tumor origin; and classifying, by a computer, at least a second nucleic acid variant detected in the cfNA sample obtained from the test subject as a nucleic acid variant of non-tumor origin if the prevalence of the second nucleic acid variant detected in the cfNA sample is higher than a threshold probability from a set of probabilities of non-tumor origin, thereby distinguishing between nucleic acid variants of tumor origin and nucleic acid variants of non-tumor origin in the cfNA sample obtained from the test subject. The set of probabilities of non-tumor origin is produced by the following steps: generating or providing, by a computer, at least one tumor variant dataset comprising a population of reference tumor-associated genetic variants, where the tumor variant dataset comprises frequencies of observed data between reference samples, including reference body fluid samples and reference non-body fluid samples, for one or more tumor-associated genetic variants in the population of reference tumor-associated genetic variants, where the reference samples are obtained from a single reference subject and / or from different reference subjects having the same cancer type; determining, by a computer, one or more ratios of the frequencies of observed data between the reference samples for the one or more tumor-associated genetic variants in the population of reference tumor-associated genetic variants to produce at least one relative prevalence dataset; and generating, by a computer, a set of probabilities of non-tumor origin from the relative prevalence dataset.

[0009] In another aspect, the present disclosure provides a method for producing a classifier that, at least in part, uses a computer to distinguish nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as nucleic acid variants of tumor origin or nucleic acid variants of non-tumor origin. The method includes generating or providing, by a computer, at least one tumor variant dataset comprising a population of reference tumor-associated genetic variants, where the tumor variant dataset comprises observed data frequencies between reference samples, including reference body fluid samples and / or reference non-body fluid samples, for one or more tumor-associated genetic variants in the population of reference tumor-associated genetic variants, where the reference samples are obtained from a single reference subject and / or from different reference subjects having the same cancer type. The method also includes determining, by a computer, one or more ratios of observed data frequencies between the reference samples for one or more tumor-associated genetic variants in the population of reference tumor-associated genetic variants to produce at least one relative prevalence dataset. Additionally, the method also includes applying, by a computer, at least one machine learning model to the relative prevalence dataset to generate at least one set of probabilities of non-tumor origin, thereby producing a classifier that identifies nucleic acid variants detected in the cfNA sample as nucleic acid variants of tumor origin or nucleic acid variants of non-tumor origin.

[0010] In some aspects, the present disclosure provides a method for distinguishing between nucleic acid variants of tumor origin and nucleic acid variants of non-tumor origin in a cell-free nucleic acid (cfNA) sample obtained from a test subject with a cancer type, at least in part using a computer. The method includes determining, by a computer, the prevalence of one or more genetic variants observed in the cfNA sample to generate a test subject prevalence dataset. The method also includes comparing, by a computer, the prevalence of one or more genetic variants in the test subject prevalence dataset with the prevalence of the genetic variants observed in a reference cfNA sample obtained from a reference subject with the cancer type. Furthermore, the method includes classifying, by a computer, a given genetic variant in the test subject prevalence dataset as a nucleic acid variant of non-tumor origin if the prevalence of the given genetic variant in the test subject prevalence dataset is below a predetermined threshold associated with the given genetic variant in the reference cfNA sample obtained from the reference subject with the cancer type.

[0011] In some aspects, the present disclosure provides a method for distinguishing between nucleic acid variants of tumor origin and nucleic acid variants of non-tumor origin in a cell-free nucleic acid (cfNA) sample obtained from a test subject, at least in part using a computer. The method includes determining, by a computer, the prevalence of one or more genetic variants observed in the cfNA sample to generate a test subject prevalence dataset. The method also includes comparing, by a computer, the prevalence of one or more genetic variants in the test subject prevalence dataset with the prevalence of the genetic variants observed in a reference cfNA sample obtained from a reference subject with leukemia, lymphoma, and / or hematological malignancy. Furthermore, the method also includes classifying, by a computer, a given genetic variant in the test subject prevalence dataset as a nucleic acid variant of non-tumor origin if the prevalence of the given genetic variant in the test subject prevalence dataset exceeds a predetermined threshold associated with the given genetic variant in the reference cfNA sample obtained from the reference subject with leukemia, lymphoma, and / or hematological malignancy.

[0012] In some embodiments, the methods disclosed herein include identifying genetic variants present in a cfNA sample from sequencing read data originating from cfNA molecules in the cfNA sample. In certain of these embodiments, the sequencing read data is obtained from targeted segments of cfNA molecules in the cfNA sample. In some embodiments, the reference tumor-associated genetic variant population is obtained from a reference sample. In certain embodiments, the reference non-body fluid sample includes a reference tumor tissue sample and / or a reference white blood cell sample. In some embodiments, the methods disclosed herein include obtaining a cfNA sample from a test subject. In certain embodiments, the reference sample comprises at least about 25, at least about 50, at least about 100, at least about 200, at least about 300, at least about 400, at least about 500, at least about 600, at least about 700, at least about 800, at least about 900, at least about 1,000, at least about 5,000, at least about 10,000, at least about 15,000, at least about 20,000, at least about 25,000, at least about 30,000, or more bodily fluid and / or non-bodily fluid samples. In some embodiments, the cfNA sample comprises cell-free deoxyribonucleic acid (cfDNA). In certain embodiments, the cfNA sample comprises cell-free ribonucleic acid (cfRNA). In some embodiments, the test subject is a mammalian subject. In certain embodiments, the test subject is a human subject. In some embodiments, the reference bodily fluid sample comprises a plasma sample. In certain embodiments, the reference bodily fluid sample comprises a serum sample. In some embodiments, the reference non-body fluid sample is a non-plasma sample. In some embodiments, the reference non-body fluid (e.g., non-plasma) sample comprises a cell sample. In certain embodiments, the reference non-body fluid (e.g., non-plasma) sample comprises a tissue sample.

[0013] In some embodiments, the methods disclosed herein include selecting one or more therapies to treat a cancer type if one or more tumor-originating nucleic acid variants associated with the cancer type are detected in a cfNA sample obtained from the test subject. In certain embodiments, the methods disclosed herein include administering one or more therapies to the test subject to treat the cancer type if one or more tumor-originating nucleic acid variants associated with the cancer type are detected in a cfNA sample obtained from the test subject.

[0014] In some embodiments, the cancer type is bile tract cancer, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, glioma, astrocytoma, breast cancer, metaplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal cancer, colon cancer, hereditary nonpolyposis colorectal cancer, colorectal adenocarcinoma, gastrointestinal stromal tumor (GIST), endometrial cancer, endometrial stromal sarcoma, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, ocular melanoma, uveal melanoma, gallbladder cancer, gallbladder adenocarcinoma, renal cell carcinoma, clear cell renal cell carcinoma, transitional cell carcinoma, urothelial carcinoma, Wilms' tumor, leukemia, acute lymphocytic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic (CLL), chronic myelogenous (CML), chronic myelomonocytic (CMML), liver cancer, liver cancer , hepatocellular carcinoma, hepatocellular carcinoma, cholangiocarcinoma, hepatoblastoma, lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B-cell lymphoma, non-Hodgkin's lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, T-cell lymphoma, non-Hodgkin's lymphoma, precursor T-lymphoblastic lymphoma / leukemia, peripheral T-cell lymphoma, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oropharyngeal cancer, oral squamous cell carcinoma, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic ductal adenocarcinoma, pseudopapillary neoplasm, acinar cell carcinoma, prostate cancer, prostate adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine cancer, gastric cancer, gastrointestinal stromal tumor (GIST), uterine cancer, and uterine sarcoma. In certain embodiments, the reference tumor-associated genetic variant is selected from the group consisting of a single nucleotide variant (SNV), an insertion or deletion (indel), a copy number variant (CNV), a fusion, a transversion, a translocation, a frameshift, a duplication, a repeat expansion, and an epigenetic variant.

[0015] In some embodiments, the methods disclosed herein include randomly dividing a tumor variant dataset into a training dataset and a test dataset. In certain embodiments, the training dataset comprises about 80% of the tumor variant dataset, and the test dataset comprises about 20% of the tumor variant dataset. In some embodiments, the tumor variant dataset comprises observed data on the frequency of one or more tumor-associated genetic variants among reference samples of a given cancer type in a population of reference tumor-associated genetic variants. In some embodiments, the methods disclosed herein include training a machine learning model using at least a portion of the population of tumor-associated genetic variants to produce a trained machine learning model, and nucleic acid variants of tumor origin and non-tumor origin detected in a cfNA sample obtained from the test subject are distinguished from each other using the trained machine learning model. In some of these embodiments, the machine learning model is trained using one or more of logistic regression, probit regression, decision tree, random forest, gradient boosting, support vector machine, K-nearest neighbor, and neural network. In some embodiments, the methods disclosed herein include using a threshold of at least about the 30th percentile probability for a given genetic variant as a cutoff for classification. In some embodiments, the methods disclosed herein include performing a logistic regression on at least one of the ratios to obtain a given probability of non-tumor origin.

[0016] In some embodiments, the tumor variant dataset comprises observed mutant allele fraction data among reference samples for one or more tumor-associated genetic variants in a population of reference tumor-associated genetic variants. In some embodiments, the methods disclosed herein comprise normalizing the tumor variant dataset using one or more data normalization techniques. In certain of these embodiments, the data normalization techniques comprise min-max normalization and / or z-score normalization. In certain embodiments, a ratio of the frequency of the observed data for a given genetic variant in the reference body fluid sample to the frequency of the observed data for the given genetic variant in the reference non-body fluid sample that is greater than one (1.0) indicates that the given genetic variant is likely to be a nucleic acid variant of non-tumor origin. In certain embodiments, where the reference non-body fluid sample comprises a reference tumor tissue sample, the set of probabilities of non-tumor origin comprises at least one set of probabilities of clonal hematopoietic origin.

[0017] In some embodiments, the tumor variant dataset includes observed mutant allele fraction data among reference samples for one or more tumor-associated genetic variants in a population of reference tumor-associated genetic variants. In some embodiments, the methods disclosed herein include normalizing the tumor variant dataset using one or more data normalization techniques. In certain of these embodiments, the data normalization techniques include min-max normalization and / or z-score normalization. In certain embodiments, a ratio of the frequency of the observed data for a given genetic variant in the reference body fluid sample to the frequency of the observed data for the given genetic variant in the reference non-body fluid sample that is less than one (1.0) indicates that the given genetic variant is likely to be a nucleic acid variant of non-tumor origin. In certain embodiments, the reference non-body fluid sample includes a reference leukocyte sample, the set of probabilities of non-tumor origin includes at least one set of probabilities of clonal hematopoietic origin.

[0018] In other aspects, the present disclosure provides a system that includes a controller that includes or is able to access a computer-readable medium that includes non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least the following: (a) generating or providing at least one tumor variant dataset comprising a population of reference tumor-associated genetic variants, the tumor variant dataset comprising frequency of observations among reference samples, including reference body fluid samples and / or reference non-body fluid samples, for one or more tumor-associated genetic variants in the population of reference tumor-associated genetic variants, the reference sample comprising a single (b) determining one or more ratios of frequencies of observations between the reference samples for one or more tumor-associated genetic variants in the population of reference tumor-associated genetic variants to produce at least one relative prevalence dataset; and (c) applying at least one machine learning model to the relative prevalence dataset to produce at least one set of probabilities of non-tumor origin to generate a classifier that identifies nucleic acid variants detected in the cell-free nucleic acid (cfNA) sample as nucleic acid variants of tumor origin or nucleic acid variants of non-tumor origin.

[0019] In other aspects, the present disclosure provides systems that include a controller that includes or is capable of accessing a computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least the following: (a) determine the relative prevalence of one or more tumor-associated genetic variants observed in one or more reference body fluid samples compared to one or more reference non-body fluid samples to produce at least one relative prevalence dataset; and (b) generate at least one set of probabilities of non-tumor origin from the relative prevalence dataset to generate a classifier that identifies nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as nucleic acid variants of tumor origin or nucleic acid variants of non-tumor origin.

[0020] In another aspect, the present disclosure provides a system including a controller that includes or can access a computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least the following: (a) determine the variation in mutant allele fraction (MAF) values, and / or at least one statistic associated therewith, for at least two different time points for each of one or more tumor-associated genetic variants observed in one or more reference body fluid samples compared to one or more reference non-body fluid samples to produce at least one MAF variance and / or relative prevalence dataset; and (b) generate at least one set of probabilities of non-tumor origin from the MAF variance and / or relative prevalence dataset to generate a classifier that identifies nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as being of tumor origin or non-tumor origin.

[0021] In some embodiments, the systems disclosed herein include a nucleic acid sequencer operably connected to a controller, the nucleic acid sequencer configured to provide sequencing read data originating from cfNA molecules in the cfNA sample. In certain of these embodiments, the nucleic acid sequencer or another system component is configured to group sequence read data generated by the nucleic acid sequencer into families of sequence read data, each family including sequence read data generated from a given cfNA molecule in the cfNA sample. In certain embodiments, the systems disclosed herein include a database operably connected to the controller, the database including one or more therapies indexed to tumor-origin nucleic acid variants. In some embodiments, the systems disclosed herein include a sample preparation component operably connected to the controller, the sample preparation component configured to prepare cfNA molecules in the cfNA sample to be sequenced by the nucleic acid sequencer. In certain embodiments, the systems disclosed herein include a nucleic acid amplification component operably connected to the controller, the nucleic acid amplification component configured to amplify at least targeted segments of cfNA molecules in the cfNA sample. In certain embodiments, the systems disclosed herein include a material transfer component operably connected to the controller, the material transfer component configured to transfer one or more materials between at least the nucleic acid sequencer and the sample preparation component.

[0022] In some aspects, the present disclosure provides a computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least the following: (a) generating or providing at least one tumor variant dataset comprising a population of reference tumor-associated genetic variants, wherein the tumor variant dataset comprises frequencies of observations among reference samples, including reference body fluid samples and / or reference non-body fluid samples, for one or more tumor-associated genetic variants in the population of reference tumor-associated genetic variants, wherein the reference samples are obtained from a single reference subject and / or from different reference subjects having the same cancer type; (b) determining one or more ratios of frequencies of observations among the reference samples for one or more tumor-associated genetic variants in the population of reference tumor-associated genetic variants to produce at least one relative prevalence dataset; and (c) applying at least one machine learning model to the relative prevalence dataset to produce at least one set of probabilities of non-tumor origin to generate a classifier that identifies nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as nucleic acid variants of tumor origin or non-tumor origin. In certain embodiments of the methods, systems, computer-readable media, and other aspects of the present disclosure, one or more other features are optionally utilized in conjunction with or instead of the frequency ratio of the observed data. Some of these other features include, for example, uniformity of prevalence across cancer types, longitudinal variant allele fraction (MAF) variation over time, proportion in hematological cancers, variant gene name, location, cancer type, chromosomal location, etc.

[0023] In other aspects, the present disclosure provides a computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least the following: (a) determine the relative prevalence of one or more tumor-associated genetic variants observed in one or more reference body fluid samples compared to one or more reference non-body fluid samples to produce at least one relative prevalence dataset; and (b) generate, from the relative prevalence dataset, at least one set of probabilities of non-tumor origin to generate a classifier that identifies nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as nucleic acid variants of tumor origin or nucleic acid variants of non-tumor origin.

[0024] In other aspects, the present disclosure provides a computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least the following: (a) determine the variation in mutant allele fraction (MAF) values, and / or at least one statistic associated therewith, for each of one or more tumor-associated genetic variants observed in one or more reference body fluid samples compared to one or more reference non-body fluid samples for at least two different time points to produce at least one MAF variance and / or relative prevalence dataset; and (b) generate at least one set of probabilities of non-tumor origin from the MAF variance and / or relative prevalence dataset to generate a classifier that identifies nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as being of tumor origin or non-tumor origin.

[0025] In some embodiments of the system or computer-readable medium disclosed herein, the electronic processor further performs at least the step of dividing the tumor variant dataset into a training dataset and a test dataset (e.g., randomly or non-randomly). In certain embodiments of the system or computer-readable medium disclosed herein, the electronic processor further performs at least the steps of training a machine learning model using at least a portion of the population of tumor-associated genetic variants to produce a trained machine learning model, and using the trained machine learning model to identify nucleic acid variants detected in the cfNA sample as nucleic acid variants of tumor origin or nucleic acid variants of non-tumor origin. In some embodiments of the system or computer-readable medium disclosed herein, the electronic processor further performs at least the step of performing logistic regression on at least one of the ratios to obtain a given probability of non-tumor origin. In certain embodiments of the system or computer-readable medium disclosed herein, the electronic processor further performs at least the step of normalizing the tumor variant dataset using one or more data normalization techniques. In some embodiments of the systems or computer-readable media disclosed herein, the electronic processor further performs the step of selecting one or more therapies for treating the cancer type if at least one or more tumor-origin nucleic acid variants associated with the cancer type are detected in the cfNA sample.

[0026] In certain embodiments, the methods, systems, or computer-readable media disclosed herein distinguish between nucleic acid variants of tumor origin and nucleic acid variants of non-tumor origin in a cell-free nucleic acid (cfNA) sample based at least in part on: (i) the uniformity of nucleic acid variant prevalence across cancer types; (ii) the variation in mutant allele fraction (MAF) of nucleic acid variants over time; and / or (iii) the prevalence of nucleic acid variants in hematological cancers, such as leukemia, lymphoma, and / or hematological malignancies.

[0027] In some embodiments, the results of the systems and methods disclosed herein are used as input to generate reports.The reports can be in paper or electronic format.For example, the classification of the nucleic acid variants detected in cell-free nucleic acid samples as being of tumor or non-tumor origin, as determined by the methods and systems disclosed herein, can be directly indicated in such reports.In some embodiments, only the nucleic acid variants classified as being of tumor origin are indicated in such reports.

[0028] Various steps of the methods disclosed herein, or steps performed by the systems disclosed herein, may be performed at the same or different times, in the same or different geographic locations, e.g., countries, and / or by the same or different people.

[0029] In other aspects, a subject may be administered a treatment based on a determination by the methods and systems disclosed herein that the variant is of tumor or non-tumor origin. In certain embodiments, administration of a treatment to a subject may be discontinued based on a determination by the methods and systems disclosed herein that the variant is of tumor or non-tumor origin. [Brief explanation of the drawings]

[0030] [Figure 1] FIG. 1 is a flow chart that schematically illustrates exemplary method steps for distinguishing between nucleic acid variants of tumor origin and nucleic acid variants of non-tumor origin, according to some embodiments.

[0031] [Figure 2] FIG. 2 is a flow chart that schematically illustrates exemplary method steps for distinguishing between nucleic acid variants of tumor origin and nucleic acid variants of non-tumor origin, according to some embodiments.

[0032] [Figure 3]FIG. 3 is a flow chart that schematically illustrates exemplary method steps for distinguishing between nucleic acid variants of tumor origin and nucleic acid variants of non-tumor origin, according to some embodiments.

[0033] [Figure 4] FIG. 4 is an example block diagram for generating a predictive model.

[0034] [Figure 5] FIG. 5 is a flow chart illustrating an exemplary training method.

[0035] [Figure 6] FIG. 6 is an illustration of an exemplary process flow for using a machine learning based classifier.

[0036] [Figure 7] FIG. 7 is a schematic diagram of an exemplary system suitable for use with certain embodiments.

[0037] [Figure 8] Figure 8, Panel A, is a plot showing the mean standard deviation (SD) separation of percentages across time points for non-tumor and tumor classes. Figure 8, Panel B, is a plot showing the area under the receiver operating characteristic (ROC) curve (AUC) of variability in longitudinal variant allele fraction (MAF) across time points as a single feature in a logistic regression model. Figure 8, Panel C, is a confusion matrix table showing the accuracy of the logistic regression model.

[0038] [Figure 9]Figure 9, panel A, is a plot of the mean prevalence ratio showing that non-oncogene variants have a higher prevalence in plasma compared to oncogene variants, regardless of the number of clinical samples observed. Known oncogene variants, such as KRAS G12D and KRAS G12V, have a lower prevalence ratio compared to the known clonal hematopoietic variant JAK2 V617F. Figure 9, panel B, is a volcano plot showing variants with fold change in magnitude (x-axis) and statistical significance (log10 of p-value, y-axis) in plasma across tissues. Known non-oncogene variants frequently observed in clonal hematopoiesis, such as JAK2 V617F and GNAS R201H (blue variants, upper right corner), show both a high fold change in magnitude and high statistical significance. Figure 9, panel C, is a plot and table showing the enrichment performance in plasma samples compared to tissue samples as a single feature. In particular, Figure 9, panel C, shows a ROC AUC plot and a confusion matrix table.

[0039] [Figure 10] Figure 10, Panel A is a plot showing that uniformly low prevalence is observed in non-tumor variants (upper panel (Panel A)) compared to known tumor variants (lower panel (Panel A)). Figure 10, Panel B shows the performance of variant prevalence across cancer types as input features to a logistic regression model. In particular, Figure 10, Panel B is a plot showing the ROC AUC, while Figure 10, Panel C is a confusion matrix table.

[0040] [Figure 11] Figure 11 (Panels A and B) are plots and tables showing the performance of the proportion of samples in hematological malignancies as a single feature. In particular, Figure 11, Panel A is a plot showing the ROC AUC, while Figure 11, Panel B is a confusion matrix table.

[0041] [Figure 12] FIG. 12 illustrates a schematic diagram of a machine learning modeling flow chart according to some embodiments.

[0042] [Figure 13] Figure 13 (Panels A and B) are plots and tables showing the performance of a random forest ensemble model on four input features. In particular, Figure 13, Panel A, is a ROC AUC and confusion matrix table showing the performance of a classifier trained with a max depth of 2 and 300 estimators. Figure 13, Panel B, is the performance of the classifier on a validation dataset with variants identified in plasma only (tumor) or in the white blood cell (WBC) fraction (non-tumor). DETAILED DESCRIPTION OF THE INVENTION

[0043] definition In order to more readily understand this disclosure, certain terms are first defined below. Additional definitions for these and other terms may be set forth throughout the specification. In the event that a definition of a term set forth below conflicts with a definition in a patent application or issued patent incorporated by reference, the definition set forth in this application should be used to understand the meaning of the term.

[0044] As used herein and in the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Thus, for example, reference to "a method" includes one or more methods and / or steps of the type described herein and / or that will become apparent to those of ordinary skill in the art upon reading this disclosure. It should also be understood that the implicit "about" precedes temperatures, concentrations, times, numbers of bases or base pairs, coverage, etc. discussed in this disclosure, so that even very small, infinitesimal equivalents are within the scope of this disclosure. In this application, the use of the singular includes the plural unless specifically stated otherwise. Also, the use of "comprise," "comprises," "comprising," "contain," "contains," "containing," "include," "includes," and "including" is not intended to be limiting.

[0045] It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting. Furthermore, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. In describing and claiming the methods, computer-readable media, and systems, the following terminology and grammatical variants thereof will be used in accordance with the definitions set forth below.

[0046] About: As used herein, "about" or "approximately," when applied to one or more values ​​or elements of interest, refers to a value or element similar to the stated reference value or element. In certain embodiments, the term "about" or "approximately" refers to a range of values ​​or elements that fall within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1% or less of the stated reference value or element in either direction (greater or less), unless otherwise stated or otherwise clear from the context (except where such number would exceed 100% of the possible values ​​or elements).

[0047] Adapter: As used herein, "adapter" refers to a short nucleic acid (e.g., less than about 500, less than about 100, or less than about 50 nucleotides in length) that is typically at least partially double-stranded and used to ligate to either or both ends of a given sample nucleic acid molecule. The adapter may contain nucleic acid primer binding sites at both ends that allow amplification of the nucleic acid molecule flanked by the adapter, and / or sequencing primer binding sites that include primer binding sites for sequencing applications, such as various next-generation sequencing (NGS) applications. The adapter may also contain a binding site for a capture probe, e.g., an oligonucleotide bound to a flow cell support. The adapter may also contain a nucleic acid tag as described herein. The nucleic acid tag is typically positioned relative to the amplification primer and sequencing primer binding sites so that the nucleic acid tag is included in the amplicon and sequencing read data of a given nucleic acid molecule. Adapters of the same or different sequences can be ligated to each end of a nucleic acid molecule. In certain embodiments, the same adapters are ligated to each end of a nucleic acid molecule, except that the sequences of the nucleic acid tags are different. In some embodiments, the adapter is a Y-shaped adapter, with one end blunt-ended or tailed as described herein for joining to a nucleic acid molecule that is also blunt-ended or tailed with one or more complementary nucleotides. In yet other exemplary embodiments, the adapter is a bell-shaped adapter containing a blunt or tailed end for joining to a nucleic acid molecule to be analyzed. Other exemplary adapters include T-tailed adapters and C-tailed adapters.

[0048] Administer: As used herein, "administering" or "administering" a therapeutic agent (e.g., an immunological therapeutic agent) to a subject means giving, applying, or contacting the composition with the subject. Administration can be accomplished by any of several routes, including, for example, topical, oral, subcutaneous, intramuscular, intraperitoneal, intravenous, intrathecal, and intradermal.

[0049] Allele: As used herein, "allele" or "allelic variant" refers to a specific genetic variant at a defined genomic location or locus. Allelic variants are typically present at a frequency of 50% (0.5) or 100%, depending on whether the allele is heterozygous or homozygous. For example, germline variants are inherited and typically have a frequency of 0.5 or 1. However, somatic variants are acquired variants and typically have a frequency of <0.5. The major and minor alleles of a locus refer to nucleic acids that carry the locus occupied by nucleotides of a reference sequence and by variant nucleotides that differ from the reference sequence, respectively. Measurements at a locus can take the form of allele fraction (AF), which measures the frequency with which an allele is observed in a sample.

[0050] Amplify: As used herein, "amplify" or "amplification," with respect to nucleic acids, refers to the production of multiple copies of a polynucleotide, or a portion of a polynucleotide, typically starting from a small amount of the polynucleotide (e.g., a single polynucleotide molecule), where an amplification product or amplicon is generally detectable. Polynucleotide amplification encompasses a variety of chemical and enzymatic processes.

[0051] Barcode: As used herein, "barcode," with respect to nucleic acids, refers to a nucleic acid molecule having a sequence that can function as a molecular identifier. For example, individual "barcode" sequences are typically added to each DNA fragment during next-generation sequencing (NGS) library preparation so that each read can be identified and sorted prior to final data analysis.

[0052] Cancer type: As used herein, "cancer," "cancer type," or "tumor type" refers to a type or subtype of cancer, for example, as defined by histopathology. Cancer types can be determined by any conventional criteria, for example, by their presence in a given tissue (e.g., blood cancer, central nervous system (CNS), brain cancer, lung cancer (small cell and non-small cell), skin cancer, nose cancer, throat cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, intestinal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, breast cancer, prostate cancer, ovarian cancer, lung cancer, intestinal cancer, soft tissue cancer, neuroendocrine cancer, gastroesophageal cancer, head and neck cancer, gynecological cancer, colorectal cancer, urothelial cancer, Cancers may be defined based on whether they are of the same cellular lineage (e.g., carcinoma, sarcoma, lymphoma, cholangiocarcinoma, leukemia, mesothelioma, melanoma, or glioblastoma), solid state cancer, heterogeneous cancer, homogeneous cancer), unknown primary origin, etc., and / or of the same cellular lineage (e.g., carcinoma, sarcoma, lymphoma, cholangiocarcinoma, leukemia, mesothelioma, melanoma, or glioblastoma), and / or exhibit cancer markers such as Her2, BRCA1, BRCA2, TP53, CA15-3, CA19-9, CA-125, CEA, AFP, PSA, HCG, KRAS, BRAF, NRAS, hormone receptors, and NMP-22. Cancers may also be classified by stage (e.g., stage 1, 2, 3, or 4) and whether they are of primary or secondary origin.

[0053] Cell-free nucleic acid: As used herein, "cell-free nucleic acid" or "cfNA" refers to nucleic acid that is not contained within or otherwise bound to a cell. Cell-free nucleic acid may include, for example, all unencapsulated nucleic acids provided by a subject's body fluids (e.g., blood, plasma, serum, urine, cerebrospinal fluid (CSF), etc.). Cell-free nucleic acid includes DNA (cfDNA), RNA (cfRNA), and hybrids thereof, including genomic DNA, mitochondrial DNA, circulating DNA, siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), and / or fragments of any of these. Cell-free nucleic acid may be double-stranded, single-stranded, or a hybrid thereof. Cell-free nucleic acid may be released into body fluids through secretion or cell death processes, such as cell necrosis, apoptosis, etc. Some cell-free nucleic acids, such as circulating tumor DNA (ctDNA), are released into body fluids from cancer cells. Others are released from healthy cells. ctDNA can be fragmented DNA from unencapsulated tumors. Another example of cell-free nucleic acid is fetal DNA, also called cell-free fetal DNA (cffDNA), that circulates freely in maternal bloodstream. Cell-free nucleic acid can have one or more epigenetic modifications, for example, cell-free nucleic acid can be acetylated, 5-methylated, ubiquitinated, phosphorylated, sumoylated, ribosylated and / or citrullinated. In some embodiments, for example, the term "cell-free nucleic acid" refers to nucleic acid that is not contained in cells or otherwise bound to cells at the time of isolation from a given subject.

[0054] Cellular origin: As used herein, "cellular origin" or "origin," with respect to cell-free nucleic acids, refers to the cell type from which or from which a given cell-free nucleic acid molecule is derived (e.g., via an apoptotic process, a necrotic process, etc.). In certain embodiments, for example, a given cell-free nucleic acid molecule may originate from a tumor cell (e.g., a cancerous cell, etc.) or a non-tumor or normal cell (e.g., a non-cancerous cell, a hematopoietic stem cell, etc.).

[0055] Classifier: As used herein, "classifier" generally refers to algorithmic computer code that receives test data as input and produces as output a classification of the input data as belonging to a class (e.g., tumor DNA or non-tumor DNA).

[0056] Clonal hematopoietic-derived mutation: As used herein, "clonal hematopoietic-derived mutation" or "clonal hematopoietic origin" refers to the somatic acquisition of a genomic mutation in hematopoietic stem and / or progenitor cells that results in clonal expansion.

[0057] Clonal hematopoiesis of undetermined potential: As used herein, "clonal hematopoiesis of undetermined potential" or "CHIP" refers to hematopoiesis in an individual involving the expansion of hematopoietic stem cells that contain one or more somatic mutations (e.g., hematological cancer-associated mutations and / or non-cancer-associated mutations) but otherwise lack diagnostic criteria for hematological malignancy, e.g., definitive morphological evidence of dysplasia. CHIP is a common age-associated phenomenon in which hematopoietic stem cells contribute to the formation of genetically distinct subpopulations of blood cells.

[0058] Copy number variant: As used herein, "copy number variant," "CNV," or "copy number variation" refers to the phenomenon in which a segment of the genome is repeated, resulting in variation in the number of repeats in the genome among individuals in the population under consideration.

[0059] Coverage: As used herein, "coverage" refers to the number of nucleic acid molecules that represent a particular base position.

[0060] Deoxyribonucleic acid or ribonucleic acid: As used herein, "deoxyribonucleic acid" or "DNA" refers to a natural or modified nucleotide having a hydrogen group at the 2'-position of the sugar moiety. DNA typically comprises a chain of nucleotides containing deoxyribonucleosides, each containing one of four types of nucleobases: adenine (A), thymine (T), cytosine (C), and guanine (G). As used herein, "ribonucleic acid" or "RNA" refers to a natural or modified nucleotide having a hydroxyl group at the 2'-position of the sugar moiety. RNA typically comprises a chain of nucleotides containing ribonucleosides, each containing one of four types of nucleobases: A, uracil (U), G, and C. As used herein, the term "nucleotide" refers to a natural or modified nucleotide. Certain pairs of nucleotides specifically bind to each other in a complementary manner (referred to as complementary base pairing). In DNA, adenine (A) pairs with thymine (T) and cytosine (C) pairs with guanine (G). In RNA, adenine (A) pairs with uracil (U) and cytosine (C) pairs with guanine (G). When a first nucleic acid strand binds to a second nucleic acid strand composed of nucleotides complementary to those in the first strand, the two strands combine to form a duplex. As used herein, "nucleic acid sequencing data," "nucleic acid sequencing information," "sequence information," "nucleic acid sequence," "nucleotide sequence," "genomic sequence," "gene sequence," or "fragment sequence" or "nucleic acid sequencing read data" refers to any information or data that indicates the order and identity of nucleotide bases (e.g., adenine, guanine, cytosine, and thymine or uracil) in a nucleic acid, e.g., DNA or RNA molecule (e.g., whole genome, whole transcriptome, exome, oligonucleotide, polynucleotide, or fragment).It should be understood that the present teachings contemplate sequence information obtained using all possible different techniques, platforms, or technologies, including, but not limited to, capillary electrophoresis, microarrays, ligation-based systems, polymerase-based systems, hybridization-based systems, direct or indirect nucleotide identification systems, pyrosequencing, ion- or pH-based detection systems, and electronic signature-based systems.

[0061] Detection: As used herein, "detect," "detecting," or "detection" refers to the act of determining the existence or presence of one or more target nucleic acids (e.g., nucleic acids having targeted mutations or other markers) in a sample.

[0062] Hematopoietic stem cell: As used herein, a "hematopoietic stem cell" or "HSC" is a stem cell that gives rise to other blood cells through the process of hematopoiesis.

[0063] Indel: As used herein, "indel" refers to a mutation involving the insertion or deletion of a nucleotide position in the genome of a subject.

[0064] Indexed: As used herein, "indexed" refers to a first element (e.g., clinical information) associated with a second element (e.g., a given sample, a recommended therapy, etc.).

[0065] Machine learning algorithm: As used herein, "machine learning algorithm" generally refers to a computer-implemented algorithm that automates analytical model building, for example, for clustering, classification, or pattern recognition. Machine learning algorithms can be supervised or unsupervised. Learning algorithms include, for example, artificial neural networks (e.g., backpropagation networks), discriminant analysis (e.g., Bayesian classifiers or Fisher analysis), support vector machines, decision trees (e.g., recursive partitioning processes, e.g., CART classification and regression trees, or random forests), linear classifiers (e.g., multiple linear regression (MLR), partial least squares (PLS) regression, and principal component regression), hierarchical clustering, and cluster analysis. The data set from which a machine learning algorithm learns can be referred to as "training data." A model produced using a machine learning algorithm is generally referred to herein as a "machine learning model."

[0066] Minor allele frequency: As used herein, "minor allele frequency" refers to the frequency at which a minor allele (e.g., not the most common allele) is present in a given population of nucleic acids, e.g., a sample obtained from a subject. Genetic variants with low minor allele frequency typically have a relatively low frequency of occurrence in samples.

[0067] Variant allele fraction: As used herein, "variant allele fraction" or "MAF" refers to the fraction of nucleic acid molecules in a given sample that carry an allele change or mutation relative to a reference at a given genomic location. MAF is generally expressed as a fraction or percentage. For example, MAF is typically less than about 0.5, 0.1, 0.05 or 0.01 (i.e., less than about 50%, 10%, 5% or 1%) of all somatic variants or alleles present at a given locus.

[0068] Mutation: As used herein, "mutation," "nucleic acid variant," "variant," or "genetic abnormality" refers to a variation from a known reference sequence, including mutations such as single nucleotide variants (SNVs), copy number variants or variations (CNVs) / abnormalities, insertions or deletions (indels), truncations, gene fusions, transversions, translocations, frameshifts, duplications, repeat expansions, and epigenetic variants. Mutations can be germline or somatic mutations. In some embodiments, the reference sequence for comparison is the wild-type genomic sequence of the species from which the test sample is provided, typically the human genome. In certain cases, the mutation or variant is a "tumor-associated genetic variant" that causes or at least contributes to carcinogenesis.

[0069] Next-generation sequencing: As used herein, "next-generation sequencing" or "NGS" refers to a sequencing technology that has increased throughput compared to traditional Sanger-based and capillary electrophoresis-based approaches, for example, the ability to generate hundreds of thousands of relatively small sequence reads at once. Some examples of next-generation sequencing techniques include, but are not limited to, sequencing by synthesis, sequencing by ligation, and sequencing by hybridization.

[0070] Nucleic acid tag: As used herein, "nucleic acid tag" refers to a short nucleic acid (e.g., less than about 500, about 100, about 50, or about 10 nucleotides in length) used to label nucleic acid molecules to distinguish nucleic acids from different samples of different types or that have undergone different processing (e.g., indicating a sample index), or to distinguish different nucleic acid molecules in the same sample (e.g., indicating a molecular tag). Nucleic acid tags can be single-stranded, double-stranded, or at least partially double-stranded. Nucleic acid tags have the same length or varying lengths, as appropriate. Nucleic acid tags can also include double-stranded molecules with one or more blunt ends, can include 5' or 3' single-stranded regions (e.g., overhangs), and / or can include one or more other single-stranded regions elsewhere within a given molecule. Nucleic acid tags can be attached to one or both ends of other nucleic acids (e.g., sample nucleic acids to be amplified and / or sequenced). Nucleic acid tags can be decoded to reveal information, such as the sample of origin, the form, or processing of a given nucleic acid. Nucleic acid tags can also be used to enable pooling and / or parallel processing of multiple samples containing nucleic acids carrying different nucleic acid tags and / or sample indexes, and these nucleic acids are subsequently deconvoluted by reading the nucleic acid tags. Nucleic acid tags can also be referred to as molecular identifiers or tags, sample identifiers, index tags, and / or barcodes. Additionally or alternatively, nucleic acid tags can be used to distinguish different molecules in the same sample. This includes, for example, uniquely tagging each different nucleic acid molecule in a given sample, or non-uniquely tagging such molecules. For non-unique tagging applications, tags with a limited number of different sequences can be used to tag each nucleic acid molecule so that different molecules can be distinguished in combination with at least one nucleic acid tag, for example, based on their start and / or stop positions mapped to a selected reference genome. Typically, a sufficient number of different nucleic acid tags are used so that the probability that any two molecules have the same start / stop position and also have the same nucleic acid tag is low (e.g., less than about 10%, less than about 5%, less than about 1%, or less than about 0.1% chance).Some nucleic acid tags include multiple molecular identifiers for labeling samples, forms of nucleic acid molecules within the samples, and nucleic acid molecules within the forms that have the same start and stop positions. Such nucleic acid tags may be referred to using the exemplary format "A1i," where the capital letter indicates the type of sample, the Arabic numeral indicates the form of the molecule within the sample, and the lowercase Roman numeral indicates the molecule within the form.

[0071] Polynucleotide: As used herein, "polynucleotide," "nucleic acid," "nucleic acid molecule," or "oligonucleotide" refers to a linear polymer of nucleosides (including deoxyribonucleosides, ribonucleosides, or their analogs) joined by internucleoside linkages. Typically, a polynucleotide contains at least three nucleosides. Oligonucleotides often range in size from a few monomeric units, e.g., 3-4, to several hundred monomeric units. When a polynucleotide is designated by a sequence of letters, e.g., "ATGCCTG," it is understood that the nucleotides are in 5'→3' order from left to right, and that, in the case of DNA, "A" represents deoxyadenosine, "C" represents deoxycytidine, "G" represents deoxyguanosine, and "T" represents deoxythymidine, unless otherwise indicated. The letters A, C, G, and T may be used to refer to the base itself, a nucleoside, or a nucleotide that comprises the base, as is standard in the art.

[0072] Prevalence: As used herein, "prevalence" or "frequency of observation," with respect to nucleic acid variants, refers to the degree, prevalence, or frequency with which a given nucleic acid variant is or has been observed in a given sample (e.g., a given bodily fluid sample, a given non-bodily fluid sample, etc.) or other population (e.g., a given population of bodily fluid samples, a given population of non-bodily fluid samples, etc.).

[0073] Reference Sample: As used herein, "reference sample" or "reference cfNA sample" refers to a sample of known composition and / or known to have or lack certain properties (e.g., known nucleic acid variant(s), known cellular origin, known tumor fraction, known coverage, etc.) that is analyzed along with or compared to a test sample to assess the accuracy of an analytical procedure, classify the test sample, etc. A reference sample dataset typically includes at least about 25 to at least about 30,000 or more reference samples. In some embodiments, the reference sample dataset comprises about 50, 75, 100, 150, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,500, 5,000, 7,500, 10,000, 15,000, 20,000, 25,000, 50,000, 100,000, 1,000,000 or more reference samples.

[0074] Reference sequence: As used herein, "reference sequence" or "reference genome" refers to a known sequence used for comparison with experimentally determined sequences. For example, the known sequence can be the entire genome, a chromosome, or any segment thereof. A reference sequence typically comprises at least about 20, at least about 50, at least about 100, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, at least about 500, at least about 1000, at least about 100,000, at least about 1,000,000, at least about 1,000,000, or more nucleotides. A reference sequence can be aligned with a single continuous sequence of a genome or chromosome, or can comprise non-contiguous segments that are aligned with different regions of a genome or chromosome. Exemplary reference sequences include, for example, human genomes, such as hG19 and hG38.

[0075] Sample: As used herein, "sample" refers to any biological sample that can be analyzed by the methods and / or systems disclosed herein. In certain embodiments of the present disclosure, the sample is a bodily fluid sample from which acellular (circulating, not contained within, or otherwise associated with, cells) nucleic acids are sourced, such as whole blood or a fraction thereof, lymph, urine, and / or cerebrospinal fluid, among other bodily fluid types. In certain implementations, the bodily fluid sample is a plasma sample, which is the fluid portion of whole blood minus cells such as red blood cells and white blood cells. In some implementations, the bodily fluid sample is a serum sample, i.e., plasma lacking fibrinogen. In some embodiments of the present disclosure, the sample is a "non-bodily fluid sample" or "non-plasma sample," i.e., a biological sample other than a "bodily fluid sample," e.g., a cell and / or tissue sample, from which nucleic acids other than cell-free nucleic acids are sourced.

[0076] Sensitivity: As used herein, "sensitivity" refers to the ability of an assay or method to detect and distinguish between targeted analytes (e.g., nucleic acid variants) and non-targeted analytes for a given assay or method.

[0077] Sequencing: As used herein, "sequencing" refers to any of several techniques used to determine the sequence (e.g., the identity and order of monomeric units) of a biomolecule, e.g., a nucleic acid, e.g., DNA or RNA. Exemplary sequencing methods include targeted sequencing, single molecule real-time sequencing, exon or exome sequencing, intron sequencing, electron microscope-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxytermination sequencing, whole genome sequencing, sequencing by hybridization, pyrosequencing, capillary electrophoresis, duplex sequencing, cycle sequencing, single base extension sequencing, solid-phase sequencing, high-throughput sequencing, massively parallel signature sequencing, emulsion PCR, co-amplification-PCR at low denaturation temperatures (COLD-PCR), multiplex PCR, reversible dye terminator sequencing, paired-end sequencing, near-term sequencing, exonuclease sequencing, sequencing by ligation, short-read sequencing, single molecule sequencing, sequencing by synthesis, real-time sequencing, reverse terminator sequencing, nanopore sequencing, 454 sequencing, Solexa Genome These include, but are not limited to, Analyzer sequencing, SOLiD™ sequencing, MS-PET sequencing, and combinations thereof. In some embodiments, sequencing can be performed by a genetic analyzer, such as a genetic analyzer commercially available from Illumina, Inc., Pacific Biosciences, Inc., or Applied Biosystems / Thermo Fisher Scientific, among many others.

[0078] Sequence information: As used herein, "sequence information," in reference to a nucleic acid polymer, means the order and identity of the monomer units (e.g., nucleotides) in the polymer.

[0079] Single nucleotide variant: As used herein, "single nucleotide variant" or "SNV" refers to a mutation or variation in a single nucleotide present at a particular position in the genome.

[0080] Somatic mutation: As used herein, "somatic mutation" refers to a mutation in the genome that occurs after conception. Somatic mutations can occur in any cell of the body except germ cells and therefore are not passed on to progeny.

[0081] Specificity: As used herein, "specificity," with respect to a diagnostic analysis or assay, refers to the degree to which the analysis or assay detects the intended target analyte, to the exclusion of other components of a given sample.

[0082] Subject: As used herein, "subject" or "test subject" refers to an animal, e.g., a mammalian species (e.g., a human) or an avian (e.g., an avian) species, or other organism, e.g., a plant. More specifically, the subject can be a vertebrate, e.g., a mammal, e.g., a mouse, a primate, a monkey, or a human. Animals include livestock (e.g., production cattle, dairy cattle, poultry, horses, pigs, etc.), sport animals, and companion animals (e.g., pets or support animals). A subject can be a healthy individual, an individual having or suspected of having a disease or predisposition to a disease, or an individual in need of therapy or suspected of needing therapy. The terms "individual" or "patient" are intended to be interchangeable with "subject." In some embodiments, the subject is a human having or suspected of having cancer. For example, the subject can be an individual diagnosed with cancer, an individual to receive cancer therapy, and / or an individual who has received at least one cancer therapy. The subject can be in remission of cancer. As another example, the subject can be an individual diagnosed with an autoimmune disease. As another example, the subject may be a female individual who is pregnant or planning to become pregnant, or who may have been diagnosed with or suspected of having a disease, e.g., cancer, autoimmune disease.

[0083] Threshold: As used herein, "threshold" refers to a discretely determined value used to characterize or classify experimentally determined values. In certain embodiments, for example, "threshold" refers to a selected value to which a quantitative value is compared to determine whether a given nucleic acid variant is a nucleic acid variant of tumor origin or a nucleic acid variant of non-tumor origin. In some of these embodiments, the selected value is a "probability threshold."

[0084] Tumor fraction: As used herein, "tumor fraction" refers to an estimate of the fraction of nucleic acid molecules derived from tumors in a given sample. For example, the tumor fraction of a sample can be a measure derived from the maximum mutant allele fraction (MAX MAF) of the sample or the coverage of the sample, or the length of the cfNA fragments in the sample, epigenetic status or other properties, or any other selected feature of the sample. The term "MAX MAF" refers to the maximum or highest MAF of all somatic variants present in a given sample. In some embodiments, the tumor fraction of a sample is equal to the MAX MAF of the sample.

[0085] Value: As used herein, a "value" generally refers to an entry in a dataset, which can be anything that characterizes the feature to which the value refers, including, but not limited to, a number, a word or phrase, a symbol (e.g., + or -), or a degree.

[0086] Detailed Description Tumor-derived somatic variants in circulating nucleic acids, such as cell-free DNA (cfDNA), can be used for targeted therapy selection, long-term longitudinal monitoring, and early detection of cancer. Cell-free tumor DNA (ctDNA) is a small DNA fragment released into the bloodstream from necrotic / apoptotic tumor cells or circulating tumor cells (CTCs). Most cfDNA originates from normal cells, including normal white blood cells undergoing apoptosis or necrosis. Recent studies have demonstrated that a significant proportion of mutations detected in cfDNA can originate from non-tumor sources, particularly clonal hematopoiesis, resulting in the accumulation of somatic mutations in hematopoietic stem cells, which contribute to cfDNA "noise" (Razavi et al., "High-intensity sequencing reveals the sources of plasma circulating cell-free DNA variants," Nature Medicine, 25:1928-1937 (2019)). The presence of non-tumor variants in plasma / cfDNA can confound the interpretation of ctDNA; therefore, methods and related aspects to distinguish between them are highly sought after.

[0087] Current approaches to identifying nucleic acid variants derived from clonal hematopoiesis or otherwise originating from cancer tumor nucleic acid variants include sequencing white blood cells (WBCs) or peripheral blood mononuclear cells to remove these sequences from nucleic acid variants in the plasma portion of a given blood sample, sequencing tissue to remove all nucleic acid variants except tissue in the plasma fraction, or a combination of both techniques (Id.). Attempted bioinformatic approaches include removing nucleic acid variants present in genes that are frequently mutated in hematologic malignancies (Coombs et al., “Therapy-related clonal hematopoiesis in patients with non-hematologic cancers is common and impacts clinical outcome,” Cell Stem Cell, 21(3):374-382 (2017)), comparing nucleic acid fragment sizes for single loci in wild-type and WBC cfDNA when they are likely to originate from the hematologic compartment (Hubbell et al., “Cell-free DNA (cfDNA) fragment length patterns of tumor- and blood-derived variants in participants with and without cancer,” 2019 AACR Meeting, Abstract 3372, March 29-April 3, 2019), and using absolute or relative variant minor allele frequency cutoffs for tumors (Li et al., “Ultra-deep next-generation sequencing of plasma cell-free DNA in patients with advanced lung cancers: results from the These approaches include the Actionable Genome Consortium,” Ann. Oncol., 30(4):597-603 (2019)). The challenge with these approaches is the requirement for matched WBCs and tissue, which are not always available and complicate sample processing.The present disclosure presents novel bioinformatics methods and related embodiments for classifying nucleic acid variants or mutations detected in plasma or other bodily fluids as being of tumor or non-tumor origin, independent of the availability of matched WBCs or tumor tissue.

[0088] 1 is a flow chart illustrating exemplary method steps for distinguishing between nucleic acid variants of tumor origin and nucleic acid variants of non-tumor origin in a cell-free nucleic acid (cfNA) sample obtained from a test subject, according to some embodiments. For example, the methods disclosed herein can be used to facilitate the removal or reduction of background noise created by nucleic acid variants of non-tumor origin (e.g., cfDNA fragments originating from non-cancerous or normal cells) detected in a given sample from a test subject, thereby improving assay sensitivity. As shown, method 100 includes a step (step 102) of determining (e.g., by a computer) the relative prevalence of tumor-associated genetic variants observed in a reference body fluid sample (e.g., a plasma sample, a serum sample, etc.) compared to a reference non-body fluid sample (e.g., a cell sample, a tissue sample, etc.) to generate a relative prevalence dataset. Method 100 also includes a step (step 104) of generating a set of probabilities of non-tumor origin from the relative prevalence dataset. Additionally, method 100 further includes identifying nucleic acid variants detected in a cfNA sample obtained from the test subject as being nucleic acid variants of tumor origin or nucleic acid variants of non-tumor origin using the set of probabilities of non-tumor origin (step 106). Related systems and computer-readable media for performing the methods disclosed herein are further described below.

[0089] To further illustrate, Figure 2 is a flow chart that schematically shows exemplary method steps for distinguishing between nucleic acid variants of tumor origin and nucleic acid variants of non-tumor origin in a cell-free nucleic acid (cfNA) sample obtained from a test subject, according to some embodiments. As shown, method 200 includes a step (step 202) of generating (e.g., by a computer) a tumor variant dataset comprising a population of reference tumor-associated genetic variants, where the tumor variant dataset comprises frequency of observation (prevalence) data among reference samples comprising reference body fluid samples (e.g., plasma samples, serum samples, etc.) and / or reference non-body fluid samples (e.g., cell samples, tissue samples, etc.) for tumor-associated genetic variants in the population of reference tumor-associated genetic variants. The reference samples are typically obtained from a single reference subject and / or from different reference subjects with the same cancer type. Method 200 also includes determining (e.g., by a computer) a ratio of the frequencies of observations between the reference samples for tumor-associated genetic variants in a population of reference tumor-associated genetic variants to generate at least one MAF variance and / or relative prevalence dataset (step 204). Method 200 further includes generating (e.g., by a computer) a set of probabilities of non-tumor origin from the MAF variance and / or relative prevalence dataset (step 206). Method 200 also includes identifying nucleic acid variants detected in a cfNA sample obtained from the test subject as being nucleic acid variants of tumor origin or nucleic acid variants of non-tumor origin (step 208) using the set of probabilities of non-tumor origin.

[0090] 3 is a flow chart that schematically illustrates exemplary method steps for distinguishing or classifying nucleic acid variants of tumor origin from nucleic acid variants of non-tumor origin in a cell-free nucleic acid (cfNA) sample obtained from a test subject, according to some embodiments. As shown, method 300 includes obtaining raw data (step 302), e.g., in the form of cancer and non-cancerous (i.e., normal or healthy) sample data and tissue sample data (e.g., from the COSMIC Cancer Database, The Cancer Genome Atlas (TCGA) data, Memorial Sloan Kettering Cancer Center (MSKCC) data, and / or another data source). 57 In the feature engineering step, input features are created, for example, by calculating the variation in variant allele fraction (MAF) over time (step 303), calculating the raw values ​​and prevalence of nucleic acid variants for all cancer types, calculating the ratio between the prevalence of nucleic acid variants observed in plasma and / or other body fluid and tissue datasets for all cancer types (step 304), calculating the proportion of nucleic acid variants in hematological malignancies or other cancer types (step 305), and testing the plasma and / or other body fluid sample prevalences for homogeneity across cancer types (e.g., generating a homogeneity score) (step 306). Bioinformatic data may include the frequency of genetic variant observations among samples of specific cancer types, including hematological malignancies; the prevalence of variants in plasma and / or other body fluids, tumor tissue, and leukocytes; the variant's variant allele fraction; etc. Additional or other data types are used as needed for these feature engineering steps. Method 300 also includes transformation and cleanup processes (step 308), such as cleaning up for sample prevalence (e.g., adjusting for samples with a low number of a given nucleic acid variant, etc.), performing a log transformation (e.g., Log(x+1) or Np.log1p), and performing normalization (e.g., Yeo-Johnson normalization, min-max normalization, z-score normalization, etc.). Method 300 also includes a machine learning step (step 310) that generates a machine learning model to provide the probability of a non-tumor nucleic acid variant being present in a given sample, for example, using logistic regression or deep learning techniques. Exemplary models that can be used for training and further classification include, but are not limited to, logistic regression, probit regression, decision trees, random forests, gradient boosting, support vector machines, K-nearest neighbors, neural networks, or an ensemble of more than one of these methods.Ensemble methods are meta-algorithms that combine several machine learning techniques into a single predictive model to reduce variance (bagging), reduce bias (boosting), or improve predictions (stacking). Most ensemble methods use a single base learning algorithm to produce uniform base learners, i.e., learners of the same type, resulting in a uniform ensemble. There are also some methods that use heterogeneous learners, i.e., learners of different types, resulting in a heterogeneous ensemble. For an ensemble method to be more accurate than any of its individual members, the base learners must be as accurate and as diverse as possible.

[0091] The data set is divided into a training set and a test set as needed using various approaches. In some embodiments, for example, the data set is randomly divided into a training data set and a test data set at an 80 / 20 ratio. In addition, method 300 also includes a step (step 312) of selecting a cutoff value for determining the threshold for classifying nucleic acid variants as being of tumor cell origin or non-tumor cell origin.

[0092] Body fluid:tissue ratio - binary classification

[0093] Some embodiments include comparing the prevalence of variants observed in a bodily fluid sample (e.g., plasma sample) dataset to their presence in a tissue dataset of the same cancer origin. In certain of these embodiments, logistic regression is performed on these ratios to obtain the probability of clonal hematopoietic origin.

[0094] In some embodiments, performance metric values ​​may include, for example, accuracy (i.e., fraction of correct predictions), balanced_accuracy (defined as the average of the recall rates obtained for each class), precision_macro (which involves calculating a metric for each label and then finding their unweighted average; however, this approach does not take label imbalance into account), precision_micro (which involves calculating a metric globally by counting total true positives, false negatives, and false positives), precision_weighted (which involves calculating a metric for each label and finding their average weighted by support (e.g., to determine the number of true instances for each label)), etc. In certain embodiments, performance metrics are estimated by stratified 5-fold cross-validation on the training set (e.g., where the split is made by preserving the percentage of samples for each class).

[0095] The Box-Cox transformation is used as needed to convert non-normal distributions to normal distributions, but this approach does not work with negative numbers. In contrast, the Yeo-Johnson transformation allows it to work with negative numbers. For example, for both logistic regression and support vector machine (SVM) models, all features are first transformed as needed by the Yeo-Johnson transformation (a parametric monotonic transformation applied to make the data more Gaussian-like in order to stabilize variance and minimize skewness). The Yeo-Johnson transformation is given by Equation 1:

number

[0096] In some embodiments, the basic inputs used to define the set of parameters are: (1) model type, and (2) a set of hyperparameters. In certain embodiments, the resulting parameters are used for all further classification. In some embodiments, the training set is used to perform a grid search with 5-fold stratified cross-validation on the following sets of hyperparameters (e.g., to define the cost of misclassification): kernel: linear, C: [0.001, 0.01, 0.1, 1, 10, 100, 1000], and kernel: radial basis function (rbf), C: [0.001, 0.01, 0.1, 1, 10, 100, 1000], gamma: [0.0001, 0.001, 0.01, 0.1, 0.2, 0.3, 0.4, 0.5, 1].

[0097] Direct training on the dataset

[0098] Some embodiments use machine learning and features of variant gene name, location, cancer type, chromosomal location, and other characteristics based on known datasets to predict tumor / non-tumor origin. In these embodiments, this method typically includes: training a machine learning model with clonal hematopoiesis (CH) and tissue-specific training datasets to identify features specific to either origin; and applying this model to historical variants observed in previous datasets to determine the probability that a given variant can be attributed to CH. In these embodiments, this method also typically includes determining which probability threshold is optimal for accurate classification of CH; and applying this list of probabilities to a new dataset to classify the origin as tumor or clonal hematopoiesis. In certain of these embodiments, the top 10 percentile of variants, although a minority of variants, have a high predictive value of CH origin.

[0099] Prevalence in body fluids compared with tumor tissue

[0100] Certain embodiments utilize the higher prevalence of a given variant in a body fluid (e.g., plasma or serum) database compared to its presence in tumor tissue, where clonal hematopoiesis (CH) may be less confounding and therefore may be informative about variants that are more likely to be CH-derived. In some of these embodiments, the method involves determining the prevalence of specific variants present in the body fluid database and comparing them to the prevalence observed in a primary tissue database, such as the COSMIC database. Some of these embodiments involve determining the ratio of the prevalence of the variant observed in body fluid samples to the prevalence of the variant observed in tissue samples. Some of these embodiments involve calculating an odds ratio for the variant's prevalence and the probability that the value of the odds ratio is equal to, greater than, or less than 1. In these embodiments, the method also generally involves applying a machine learning model to these relative prevalence values ​​of body fluid versus tissue prevalence to determine the probability that the variant is likely to originate from CH or a non-tumor status. A probability threshold (e.g., about the 10th, 15th, 20th, 25th, 30th, 35th, 40th, or another percentile) is typically used as a cutoff for classification. Generally, a small number of variants have a high predictive value for being non-tumor or CH.

[0101] Multiclass models using distribution or homogeneity tests across tumor types

[0102] In some embodiments, tumor-specific variants have a specific distribution or selection depending on the biology of each cancer type. If a variant is not specific to a cancer type, it typically has a uniform distribution, which may indicate a passenger mutation or non-tumor status. Therefore, certain methods determine the prevalence of variants across tumor types, or their relative proportions and representations, and machine learning models can be trained to separate them into distinct tumor and non-tumor classes. Some of these methods include using the coefficient of variation to determine the distribution and any significant enrichment in a specific tumor type. In certain of these embodiments, very few variants are predictive and tumor-type specific, and are unlikely to be CH. Some variants do not have demonstrable selectivity for a specific tumor type and have a low prevalence across all tumor types, indicating a high likelihood of being CH. In general, if a variant is substantially uniform across tumor types, it is likely to be non-tumor / CH in origin, but if a variant is highly prevalent in a specific tumor, it is more likely that there is biological selection for that variant in the tumor. Current methods that strictly rely on patient age or absolute VAF, and that ignore expected relative prevalence in different tumor settings, fail to consider these underlying disease-specific mechanisms (or their absence) that drive the observed VAF, as well as important biological features indicative of variant origin.

[0103] Other input features to the machine learning model, in these embodiments, include, for example, variant-based tumor classification (e.g., tumor type or expected tumor type), the presence or signature of methylation, other variants in a given sample (e.g., known CH variants present in a sample that increase the probability that other variants in the sample are also non-tumor in origin), differences in family size for a given variant versus a reference allele, the nature of the observed nucleotide or other change in the variant, the absolute value of the MAF in the sample, the relative value of the MAF in the sample, how the value of the MAF in the sample changes over time compared to other variants, etc.

[0104] Monitor MAF values ​​over time

[0105] In certain cases, variant clones of non-tumor variants are more likely to remain stable over time in a subject than variants originating from tumors. Thus, in some embodiments, this method includes calculating the coefficient of variation (CV, dispersion compared to the mean) of variant percentage over time for each patient at multiple time points (e.g., >3), and calculating the distribution of statistics and CV across all variants and patients. In these embodiments, known driving factors or tumor variants generally have dynamic percentages over time (due to tumor growth and shrinkage), and have larger CVs across time points compared to non-tumor variants. This can also be used as an input feature for classifiers. In contrast, non-tumor variant MAFs are typically less dynamic and more stable over time than true tumor variants, and have lower CVs over time. The distribution of these CVs can be separated in machine learning models to provide robust classification of tumor or non-tumor status. Other input features for the machine learning model in these embodiments include, for example, variant clonality (relative VAF to tumor fraction) over time or across patients, fragmentomics data points, fragment size, location, patient age (older patients have a higher probability of CHIP), etc. Current methods that can track VAF or VAF dispersion across time points in a single patient are less accurate than approaches that aggregate VAF across patients, especially when these patients are all tested consecutively on the same platform and bioinformatics pipeline, which results in consistent VAF and a more robust measure of variation. Furthermore, this method of classification may use a fixed threshold that does not adjust for dispersion values ​​compared to absolute VAF, and thus non-tumor variants with higher VAF may be confused by lower VAF variants with similar measures of dispersion over time.Machine learning models that take into account both absolute VAF as well as VAF scatter across time points in a sufficiently large cohort of patients measured on the same platform have higher resolution for classification and are less likely to produce false positive or negative labeling of tumor / non-tumor status.

[0106] 4, a further method for generating a predictive model (e.g., a classification model) is described. The described method may use machine learning (“ML”) techniques to train at least one ML module 430 configured to classify mutations detected in plasma as of tumor origin or non-tumor origin, which may be derived from clonal hematopoiesis or biological noise, based on analysis of one or more training datasets 410A-410N by a training module 420.

[0107] One or more training datasets 410A-410N may include cancer / non-cancerous (e.g., tumor / non-tumor) body fluid (e.g., blood, plasma, serum, cerebrospinal fluid, urine) sample data and cancer / non-cancerous (e.g., tumor / non-tumor) non-body fluid (e.g., tissue) sample data (e.g., from the COSMIC Cancer Database, The Cancer Genome Atlas (TCGA) data, and / or another data source). Subsets of the cancer / non-cancerous body fluid sample data and / or the cancer / non-cancerous non-body fluid sample data may be randomly assigned to the training dataset 410 or the test dataset. In some implementations, the assignment of data to the training dataset or the test dataset may not be completely random. In this case, one or more criteria may be used during the assignment. In general, any suitable method may be used to assign data to the training set or the test dataset while ensuring that the data distribution is somewhat similar in the training dataset and the test dataset.

[0108] The training module 420 may train the ML module 430 by extracting feature sets from the cancer / non-cancer body fluid sample data and / or the cancer / non-cancer non-body fluid sample data in the training dataset 410 according to one or more feature selection techniques. The training module 420 may train the ML module 430 by extracting feature sets from the training dataset 410 that include statistically significant features.

[0109] The training module 420 may extract feature sets from the training dataset 410 in various ways. The training module 420 may perform feature extraction multiple times, using a different feature extraction technique each time. In an example, feature sets generated using different techniques may each be used to generate a different machine learning-based classification model 440. For example, the feature set with the highest quality metric may be selected for use in training. The training module 420 may use the feature set(s) to build one or more machine learning-based classification models 440A-440N configured to classify new variants (e.g., with unknown origin) as tumor or non-tumor in origin.

[0110] The training dataset 410 may be analyzed to determine any dependencies, associations, and / or correlations between features in the training dataset 410 and experimental parameters. The identified correlations may have the form of a list of features. The term "feature," as used herein, may refer to any characteristic of an item of data that may be used to determine whether the item of data falls within one or more particular categories. By way of example, the features described herein may include one or more of the following: the frequency of observation of a genetic variant among samples of a particular cancer type, including hematological malignancies; the prevalence of the variant in plasma, tumor tissue, or leukocytes; and / or the minor allele frequency of the variant.

[0111] The feature selection technique may include one or more feature selection rules. The one or more feature selection rules may include feature presence rules. The feature presence rules may include determining which features in the training data set 410 are present more than a threshold number of times and identifying features that meet the threshold as features.

[0112] A single feature selection rule may be applied to select features, or multiple feature selection rules may be applied to select features. Feature selection rules may be applied in a cascade fashion, where feature selection rules are applied in a specific order and apply to the results of previous rules. For example, a feature presence rule may be applied to the training dataset 410 to generate a first list of features. The final list of features may be analyzed according to additional feature selection techniques to determine one or more feature groups (e.g., a group of features that can be used to classify variants as tumor or non-tumor origin). Any suitable computational technique may be used to identify feature groups using any feature selection technique, such as a filter, wrapper, and / or embedding method. One or more feature groups may be selected according to a filter method. Filter methods include, for example, Pearson's correlation, linear discriminant analysis, analysis of variance (ANOVA), chi-square, combinations thereof, etc. The selection of features according to a filter method is independent of any machine learning algorithm. Instead, features may be selected based on scores in various statistical tests for their correlation with outcome variables.

[0113] As another example, one or more feature sets may be selected according to a wrapper method. The wrapper method may be configured to use a subset of features and train a machine learning model using the subset of features. Features may be added to and / or removed from the subset based on inferences drawn from previous models. Wrapper methods include, for example, forward feature selection, backward feature elimination, recursive feature reduction, combinations thereof, and the like. As an example, forward feature selection may be used to identify one or more feature sets. Forward feature selection is an iterative method that starts with no features present in the machine learning model. In each iteration, the feature that best improves the model is added until adding a new variable no longer improves the performance of the machine learning model. As an example, backward elimination may be used to identify one or more feature sets. Backward elimination is an iterative method that starts with all features present in the machine learning model. In each iteration, the lowest-ranking feature is removed until no improvement is observed from removing the feature. Recursive feature reduction may be used to identify one or more feature sets. Recursive feature reduction is a greedy optimization algorithm that aims to find the best-performing feature subset. Recursive feature reduction iteratively creates models, setting aside the best- or worst-performing features in each iteration. Recursive feature reduction constructs the next model using the remaining features until all features are exhausted. Recursive feature reduction then ranks the features based on their order of reduction.

[0114] As a further example, one or more feature sets may be selected according to an embedding method. Embedding methods combine the properties of filter methods and wrapper methods. Embedding methods include, for example, Least Absolute Shrinkage and Selection Operator (LASSO) and ridge regression, which implements a penalty function to reduce overfitting. For example, LASSO regression implements L1 regularization, which adds a penalty equal to the absolute value of the magnitude of the coefficient, while ridge regression implements L2 regularization, which adds a penalty equal to the square of the magnitude of the coefficient.

[0115] After the training module 420 generates the feature set(s), the training module 420 may generate a machine learning-based classification model 440 based on the feature set(s). A machine learning-based classification model may refer to a complex mathematical model for data classification that is generated using machine learning techniques. In one example, the machine learning-based classification model 440 may include a map of support vectors that indicate boundary features. By way of example, the boundary features may be selected from and / or may indicate the highest ranked features in the feature set.

[0116] The training module 420 may use the feature set determined or extracted from the training dataset 410 to construct the machine learning based classification models 440A-440N. In some examples, the machine learning based classification models 440A-440N may be combined into a single machine learning based classification model 440. Similarly, the ML module 430 may represent a single classifier including single or multiple machine learning based classification models 440 and / or multiple classifiers including single or multiple machine learning based classification models 440.

[0117] Features can be combined in classification models trained using machine learning approaches, such as discriminant analysis; decision trees; nearest neighbor (NN) algorithms (e.g., k-NN models, replicator NN models, etc.); statistical algorithms (e.g., Bayesian networks, etc.); clustering algorithms (e.g., k-means, mean shift, etc.); neural networks (e.g., reservoir networks, artificial neural networks, etc.); support vector machines (SVMs); logistic regression algorithms; linear regression algorithms; Markov models or chains; principal component analysis (PCA) (e.g., for linear models); multilayer perceptron (MLP) ANNs (e.g., for nonlinear models); replicating reservoir networks (e.g., for nonlinear models, typically for time series); random forest classification; combinations thereof; and the like. The resulting ML module 430 can include a decision rule or mapping for each feature to determine tumor / non-tumor origin for variants.

[0118] In one embodiment, the training module 420 may train the machine learning-based classification model 440 as a convolutional neural network (CNN), which includes at least one convolutional feature layer and three fully connected layers leading to a final classification layer (softmax), which may finally be applied to combine the outputs of the fully connected layers using a softmax function known in the art.

[0119] The feature(s) and the ML module 430 can be used to predict the tumor / non-tumor origin of variants in the test dataset. In one example, the predicted result for each variant can include a confidence level corresponding to the likelihood or probability that the variant in the test dataset is associated with a tumor or non-tumor origin. The confidence level can be a value between zero and one. In one example, if two statuses (e.g., tumor origin and non-tumor origin) exist, the confidence level can correspond to a value p indicating the likelihood that a particular variant belongs to the first status (e.g., tumor origin). In this case, the value 1-p can indicate the likelihood that a particular variant belongs to the second status (e.g., non-tumor origin). Generally, multiple confidence levels can be provided for each variant in the test dataset, and for each feature if more than two statuses exist. Top-performing features can be determined by comparing the results obtained for each test variant with the known tumor / non-tumor origin for each test variant. Generally, top-performing features have results that closely match the known tumor / non-tumor origin status. The top performing feature(s) can be used to predict / classify the tumor / non-tumor origin status of a given variant.

[0120] 5 is a flowchart illustrating an example training method 500 for generating an ML model 430 using the training module 420. The training module 420 can perform supervised, unsupervised, and / or semi-supervised (e.g., reinforcement-based) machine learning-based classification models 440. The method 500 illustrated in FIG. 5 is one example of a supervised learning method; variations of this example training method are discussed below, although other training methods can be similarly performed to train unsupervised and / or semi-supervised machine learning models.

[0121] The training method 500 may determine (e.g., access, receive, retrieve, etc.) data at step 510. The data may include cancer / non-cancerous (e.g., tumor / non-tumor) body fluid sample data and cancer / non-cancerous (e.g., tumor / non-tumor) non-body fluid (e.g., tissue) sample data. The data may include one or more variants, each variant having an assigned tumor or non-tumor origin status.

[0122] The training method 500 may generate a training data set and a test data set at step 520. The training data set and the test data set may be generated by randomly assigning data to either the training data set or the test data set. In some implementations, the assignment of calculated parameters and associated experimental parameters as training data or test data may not be completely random. As an example, a majority of the calculated parameters and associated experimental parameters may be used to generate the training data set. For example, 75% of the calculated parameters and associated experimental parameters may be used to generate the training data set, and 25% may be used to generate the test data set. In another example, 80% of the calculated parameters and associated experimental parameters may be used to generate the training data set, and 20% may be used to generate the test data set.

[0123] The training method 500 may, at step 530, determine (e.g., extract, select, etc.) one or more features that can be used by a classifier to, for example, distinguish between different classifications of tumor vs. non-tumor status. By way of example, the training method 500 may determine a set of features from the cancer / non-cancer body fluid sample data and the cancer / non-cancer non-body fluid sample data. In a further example, the set of features may be determined from data other than the cancer / non-cancer body fluid sample data and the cancer / non-cancer non-body fluid sample data in either the training dataset or the test dataset. Such other data may be used to determine an initial set of features, which may be further reduced using the training dataset.

[0124] The training method 500 may train one or more machine learning models using one or more features at step 540. In one example, the machine learning models may be trained using supervised learning. In another example, other machine learning techniques, including unsupervised learning and semi-supervised learning, may be used. The machine learning models trained at 540 may be selected based on different criteria, depending on the problem to be solved and / or the available data in the training dataset. For example, machine learning classifiers may suffer from different degrees of bias. Thus, more than one machine learning model may be trained at 540 and optimized, refined, and cross-validated at step 550.

[0125] The training method 500 may select one or more machine learning models to build a predictive model at 560. The predictive model may be evaluated using a test dataset. The predictive model may analyze the test dataset at step 570 and generate a predicted tumor / non-tumor origin status. The predicted tumor / non-tumor origin may be evaluated at step 580 to determine whether such value reaches a desired level of accuracy. The performance of a predictive model may be evaluated in several ways based on true positive, false positive, true negative, and / or false negative classifications of some of the data points represented by the predictive model.

[0126] For example, a false positive of a predictive model may refer to the number of times the predictive model incorrectly classified a variant as having a tumor origin when it was actually not. Conversely, a false negative of a predictive model may refer to the number of times the machine learning model classified a variant as having a non-tumor origin when it was actually having a tumor origin. True negatives and true positives may refer to the number of times the predictive model correctly classified one or more variants. These measurements are related to the concepts of recall and precision. Generally, recall refers to the ratio of true positives to the sum of true positives and false negatives, and quantifies the sensitivity of the predictive model. Similarly, precision refers to the ratio of true positives to the sum of true positives and false positives. If such a desired accuracy level is reached, the training phase ends and a predictive model (e.g., ML module 430) may be output in step 590; however, if the desired accuracy level is not reached, subsequent iterations of the training method 500 may be performed, beginning at step 510, with variations, such as considering a larger collection of data.

[0127] 6 is an illustration of an exemplary process flow for using a machine learning-based classifier to classify a variant as being of tumor or non-tumor origin. As illustrated in FIG. 6, an unclassified variant 610 may be provided as input to an ML module 430. The ML module 430 may process the unclassified variant 610 using a machine learning-based classifier(s) to arrive at a prediction result 620. The prediction result 620 may identify one or more characteristics of the unclassified variant 610. For example, the classification result 620 may identify the origin status of the unclassified variant 610 (e.g., whether the variant is of tumor or non-tumor origin). Thus, in one embodiment, a method is disclosed that is performed using a network-based computer system including one or more processors, a network interface, and one or more memories, the method including: retrieving, by the computer system, genetic information and additional information for a plurality of tumor and non-tumor body fluid samples and a plurality of tumor and non-tumor non-body fluid (e.g., tissue) samples from the one or more memories, wherein the additional information includes tumor origin or non-tumor origin status; and training, by the one or more processors, machine learning models by fitting one or more models to the genetic information and additional information, wherein each of the one or more models is configured to receive as input an individual's genetic information and provide as output a prediction of the individual having or developing a tumor.

[0128] System and computer-readable medium

[0129] The present disclosure also provides various systems, bioinformatics pipelines, and computer program products or machine-readable media. In some embodiments, for example, the methods described herein are implemented or facilitated, at least in part, using systems, distributed computing hardware and applications (e.g., cloud computing services), electronic communications networks, communications interfaces, computer program products, machine-readable media, electronic storage media, software (e.g., machine-executable code or logical instructions), and the like, as appropriate. To illustrate, FIG. 7 provides a schematic diagram of an exemplary system suitable for use in performing at least aspects of the methods disclosed herein. As shown, system 700 includes at least one controller or computer, e.g., a server 702 (e.g., a search engine server) including a processor 704 and memory, storage devices, or memory components 706, and one or more other communication devices 714 and 716 (e.g., client-side computer terminals, phones, tablets, laptops, other mobile devices, etc.) located remote from and in communication with the remote server 702 via an electronic communications network 712, e.g., the Internet or other internetwork. Communication devices 714 and 716 typically include electronic displays (e.g., internet-enabled computers, etc.) in communication with, for example, server 702 computer via network 712, which include a user interface (e.g., a graphical user interface (GUI), a web-based user interface, etc.) for displaying the results of performing the methods described herein. In certain embodiments, the communication network also encompasses the physical transfer of data from one location to another, for example, using a hard drive, thumb drive, or other data storage mechanism.System 700 also includes a program product 708 stored on a computer or machine-readable medium, such as one or more various types of memory, such as memory 706 of server 702, that can be read by server 702, for example, to facilitate a guided search application executable by one or more other communication devices, such as 714 (shown diagrammatically as a desktop or personal computer) and 716 (shown diagrammatically as a tablet computer). In some embodiments, system 700 also optionally includes at least one database server, such as server 710, associated with an online website having stored data thereon (e.g., nucleic acid variant lists, indexed therapies, etc.) searchable either directly or via search engine server 702. System 700 also optionally includes one or more other servers located remotely from server 702, each associated with one or more database servers 710 that are located remotely or locally from each of the other servers, as appropriate. The other servers may beneficially provide services to geographically dispersed users and may enhance geographically distributed operations.

[0130] As will be appreciated by those skilled in the art, the memory 706 of the server 702 may include volatile and / or nonvolatile memory, including, for example, RAM, ROM, and magnetic or optical disks, among others, as appropriate. While illustrated as a single server, those skilled in the art will also appreciate that the illustrated configuration of the server 702 is provided by way of example only, and that other types of servers or computers configured according to various other methodologies or architectures may also be used. The server 702 shown schematically in FIG. 7 represents a server or server cluster or server farm and is not limited to any individual physical server. A server site may be deployed as a server farm or server cluster managed by a server hosting provider. The number of servers and their architecture and configuration may be increased based on the use, demand, and capacity requirements for the system 700. As will also be appreciated by those skilled in the art, the other user communication devices 714 and 716, in these embodiments, may be, for example, laptops, desktops, tablets, personal digital assistants (PDAs), mobile phones, servers, or other types of computers. As known and understood by those skilled in the art, network 712 may include the Internet, an intranet, a telecommunications network, an extranet, or the World Wide Web of multiple computers / servers in communication with one or more other computers via portions of a communications network and / or local or other area network.

[0131] As will be further understood by those skilled in the art, exemplary program product or machine-readable medium 708 is in the form of microcode, programs, cloud computing formats, routines, and / or symbolic languages ​​that provide one or more sets of ordered operations that control the functionality of and direct the operation of the hardware, as appropriate. Program product 708 according to exemplary embodiments also need not reside entirely in volatile memory, but may be selectively loaded, as appropriate, according to various methodologies known and understood by those skilled in the art.

[0132] As will be further understood by those skilled in the art, the terms “computer-readable medium” or “machine-readable medium” refer to any medium that participates in providing instructions to a processor for execution. To illustrate, the terms “computer-readable medium” or “machine-readable medium” encompass, for example, distribution media, cloud computing formats, intermediate storage media, computer execution memory, and any other media or devices capable of storing, for reading by a computer, a program product 708 that performs the functions or processes of various embodiments of the present disclosure. A “computer-readable medium” or “machine-readable medium” may take many forms, including, but not limited to, non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical or magnetic disks. Volatile media include dynamic memory, such as the main memory of a given system. Transmission media include coaxial cables, copper wire, and fiber optics, including the wires that comprise a bus. Transmission media may also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications, among others. Exemplary forms of computer-readable media include a floppy disk, flexible disk, hard disk, magnetic tape, flash drive, or any other magnetic medium, CD-ROM, any other optical medium, punch cards, paper tape, any other physical medium with a pattern of holes, RAM, PROM and EPROM, FLASH-EPROM, any other memory chip or cartridge, carrier wave, or any other medium from which a computer can read.

[0133] The program product 708 is copied from the computer-readable medium to a hard disk or similar intermediate storage medium, as needed. When the program product 708, or portions thereof, are executed, it is loaded, as needed, from its distribution medium, its intermediate storage medium, etc. into the execution memory of one or more computers that configure the computer(s) to operate according to the functions or methods of the various embodiments. All such operations are well known to those skilled in the art of, for example, computer systems.

[0134] To further illustrate, in certain embodiments, the present application provides a system including one or more processors and one or more memory components in communication with the processor. The memory component typically includes one or more instructions that, when executed, cause the processor to provide information that causes at least one nucleic acid variant list, variant classification call report or result, selected therapy, etc. to be displayed (e.g., via communication devices 714, 716, etc.) and / or receive information from other system components and / or from a system user (e.g., via communication devices 714, 716, etc.).

[0135] In some embodiments, the program product 708 includes non-transitory computer-executable instructions that, when executed by the electronic processor 704, perform at least the following: (i) generating a tumor variant dataset comprising a population of reference tumor-associated genetic variants, wherein the tumor variant dataset comprises frequencies of observations among reference samples, including reference body fluid samples and / or reference non-body fluid samples, for tumor-associated genetic variants in the population of reference tumor-associated genetic variants, wherein the reference samples are obtained from a single reference subject and / or from different reference subjects having the same cancer type; (ii) determining a ratio of the frequencies of observations among the reference samples for tumor-associated genetic variants in the population of reference tumor-associated genetic variants to produce a relative prevalence dataset; (iii) generating a set of probabilities of non-tumor origin from the relative prevalence dataset; and (iv) using the set of probabilities of non-tumor origin, identifying nucleic acid variants detected in a cfNA sample obtained from the test subject as nucleic acid variants of tumor origin or nucleic acid variants of non-tumor origin.

[0136] System 700 also typically includes additional system components configured to implement various aspects of the methods described herein. In some of these embodiments, one or more of these additional system components are located remote from and in communication with remote server 702 via electronic communications network 712, while in other embodiments, one or more of these additional system components are located locally, in communication with server 702 (i.e., in the absence of electronic communications network 712), or directly with desktop computer 714, for example.

[0137] In some embodiments, for example, additional system components include a sample preparation component 718 operably connected to the controller 702 (either directly or indirectly, e.g., via the electronic communications network 712). The sample preparation component 718 is configured to prepare nucleic acids in the sample to be amplified and / or sequenced by a nucleic acid amplification component (e.g., a thermal cycler, etc.) and / or a nucleic acid sequencer (e.g., prepare a library of nucleic acids). In certain of these embodiments, the sample preparation component 718 is configured to attach one or more adapters comprising barcodes to the nucleic acids as described herein to isolate the nucleic acids from other components in the sample, such as for selectively enriching one or more regions from a genome or transcriptome prior to sequencing.

[0138] In certain embodiments, system 700 also includes a nucleic acid amplification component 720 (e.g., a thermal cycler, etc.) operably connected to controller 702 (either directly or indirectly, e.g., via electronic communication network 712). Nucleic acid amplification component 720 is configured to amplify nucleic acids in a sample from a subject. For example, nucleic acid amplification component 720 is configured to amplify selectively enriched regions from a genome or transcriptome in the sample, as desired, as described herein.

[0139] System 700 typically also includes at least one nucleic acid sequencer 722 operably connected to controller 702 (either directly or indirectly, e.g., via electronic communications network 712). Nucleic acid sequencer 722 is configured to provide sequence information from nucleic acids (e.g., amplified nucleic acids) in a sample from a subject. Essentially any type of nucleic acid sequencer can be adapted for use in these systems. For example, nucleic acid sequencer 722 is configured to perform pyrosequencing, single molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing by ligation, sequencing by hybridization, or other techniques on nucleic acids, as needed, to generate sequencing read data. Optionally, nucleic acid sequencer 722 is configured to group sequence read data into families of sequence read data, each family comprising sequence read data generated from nucleic acids in a given sample. In some embodiments, nucleic acid sequencer 722 uses clonal single molecule arrays derived from a sequencing library to generate sequencing read data. In certain embodiments, the nucleic acid sequencer 722 includes at least one chip having an array of microwells for sequencing a sequencing library to generate sequencing read data.

[0140] To facilitate full or partial system automation, system 700 typically also includes a material transfer component 724 operably connected to controller 702 (either directly or indirectly (e.g., via electronic communications network 712)). Material transfer component 724 is configured to transfer one or more materials (e.g., nucleic acid samples, amplicons, reagents, etc.) to and / or from nucleic acid sequencer 722, sample preparation component 718, and nucleic acid amplification component 720.

[0141] For further details regarding computer systems and networks, databases, and computer program products, see, for example, Peterson, Computer Networks: A Systems Approach, Morgan Kaufmann, 5th Ed. (2011), Kurose, Computer Networking: A Top-Down Approach, Pearson, 7 th Ed. (2016), Elmasri, Fundamentals of Database Systems, Addison Wesley, 6th Ed. (2010), Coronel, Database Systems: Design, Implementation, & Management, Cengage Learning, 11 th Ed. (2014), Tucker, Programming Languages, McGraw-Hill Science / Engineering / Math, 2nd Ed. (2006), and Rhoton, Cloud Computing Architected: Solution Design Handbook, Recursive Press (2011), each of which is incorporated by reference in its entirety.

[0142] Sample collection and preparation

[0143] The sample can be any biological sample isolated from a subject. The sample can include body fluids or body tissues (e.g., known or suspected solid tumors). Samples can include whole blood, platelets, serum, plasma, feces, red blood cells, white blood cells or leucocytes, endothelial cells, tissue biopsies, cerebrospinal fluid, synovial fluid, lymphatic fluid, ascites, interstitial or extracellular fluid, fluid in the spaces between cells including gingival crevicular fluid, bone marrow, pleural fluid, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, and urine. The sample is preferably a body fluid, particularly blood and its fractions, and urine. Such samples contain nucleic acids shed from tumors. Nucleic acids can include DNA and RNA, and can be in double-stranded and / or single-stranded form. Sample can be the form that is originally separated from subject, or can be further processed, for example, to remove or add components, for example, cell, to enrich one component with respect to another component, or to convert one form of nucleic acid into another form, for example, RNA into DNA, or single-stranded nucleic acid into double-stranded.Therefore, for example, the body fluid for analysis is the plasma or serum that contains cell-free nucleic acid, for example, cell-free DNA (cfDNA).

[0144] In certain embodiments, polynucleotides can be enriched before sequencing. Enrichment can be performed for specific target regions ("target sequences") or non-specifically. In some embodiments, targeted regions of interest can be enriched using a differential tiling and capture scheme with capture probes ("baits") selected for one or more bait set panels. The differential tiling and capture scheme uses different relative concentrations of bait sets to differentially tile (e.g., at different "resolutions") across genomic regions related to the baits according to a set of constraints (e.g., sequencer constraints, e.g., sequencing load, availability of each bait, etc.) and capture them at a desired level for downstream sequencing. These targeted genomic regions of interest can include regions of the genome or transcriptome of interest. In some embodiments, biotin-labeled beads bearing probes for one or more regions of interest can be used to capture target sequences, and if necessary, the regions can then be amplified to enrich for the regions of interest.

[0145] Sequence capture typically involves the use of oligonucleotide probes that hybridize to the target sequence. Probe set strategies can involve tiling probes across the region of interest. Such probes can be, for example, about 60-130 bases long. Sets can have a depth of about 2x, 3x, 4x, 5x, 6x, 8x, 9x, 10x, 15x, 30x, 50x, or greater. The effectiveness of sequence capture depends in part on the length of the sequence in the target molecule that is complementary (or nearly complementary) to the sequence of the probe.

[0146] In some embodiments, the methods of the disclosure include selectively enriching regions from the genome or transcriptome of the subject prior to sequencing, hi other embodiments, the methods of the disclosure include non-selectively enriching regions from the genome or transcriptome of the subject prior to sequencing.

[0147] In certain embodiments, a sample index sequence is introduced into the polynucleotide after enrichment. The sample index sequence can be introduced into the polynucleotide via PCR, optionally as part of an adapter, or can be ligated to the polynucleotide.

[0148] The volume of the bodily fluid can depend on the desired read depth for the sequenced region. Exemplary volumes are 0.4-40 ml, 5-20 ml, and 10-20 ml. For example, the volume can be 0.5 ml, 1 ml, 5 ml, 10 ml, 20 ml, 30 ml, or 40 ml. The volume of the sampled bodily fluid can be 5-20 ml.

[0149] A sample can contain various amounts of nucleic acid, including a genome equivalent. For example, a sample of about 30 ng of DNA contains about 10,000 (10 4 ) haploid human genome equivalents, and for cfDNA, approximately 200 billion (2 × 10 11 ) individual polynucleotide molecules. Similarly, a sample of about 100 ng of DNA may contain about 30,000 haploid human genome equivalents, or in the case of cfDNA, about 600 billion individual molecules.

[0150] The sample may comprise nucleic acids from different sources, for example, from cells and acellular sources.The sample may comprise nucleic acids carrying mutations.For example, the sample may comprise DNA carrying germline mutations and / or somatic mutations.The sample may comprise DNA carrying cancer-related mutations (for example, cancer-related somatic mutations).

[0151] Exemplary amounts of cell-free nucleic acid in a sample prior to amplification range from about 1 fg to about 1 μg, e.g., 1 pg to 200 ng, 1 ng to 100 ng, or 10 ng to 1000 ng. For example, the amount can be up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng of cell-free nucleic acid molecules. The amount can be at least 1 fg, at least 10 fg, at least 100 fg, at least 1 pg, at least 10 pg, at least 100 pg, at least 1 ng, at least 10 ng, at least 100 ng, at least 150 ng, or at least 200 ng of cell-free nucleic acid molecules. The amount can be up to 1 femtogram (fg), 10 fg, 100 fg, 1 picogram (pg), 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng, or 200 ng of cell-free nucleic acid molecules. The method can include obtaining between 1 femtogram (fg) and 200 ng.

[0152] Cell-free nucleic acids have an exemplary size distribution of about 100 to 500 nucleotides, with molecules of 110 to about 230 nucleotides accounting for about 90% of the molecules, the mode in humans being about 168 nucleotides, and a second, smaller peak ranging between 240 and 430 nucleotides. Cell-free nucleic acids can be about 160 to about 180 nucleotides, or about 320 to about 360 nucleotides, or about 430 to about 480 nucleotides.

[0153] Cell-free nucleic acids can be isolated from body fluids through a partitioning step, in which the cell-free nucleic acids found in solution are separated from intact cells and other insoluble components of the body fluid. The partitioning step can include techniques such as centrifugation or filtration. Alternatively, cells in the body fluid can be lysed, and the cell-free and cellular nucleic acids are processed together. Generally, after adding a buffer and a washing step, the cell-free nucleic acids can be precipitated with alcohol. Additional cleanup steps, such as a silica-based column, can be used to remove contaminants or salts. For example, non-specific bulk carrier nucleic acids can be added throughout the reaction to optimize certain aspects of the procedure, such as yield.

[0154] After such processing, the sample may contain nucleic acids in various forms, including double-stranded DNA, single-stranded DNA, and single-stranded RNA. If necessary, single-stranded DNA and RNA may be converted to double-stranded form and then included in subsequent processing and analysis steps.

[0155] amplification Adapter-flanked sample nucleic acids can be amplified by PCR and other amplification methods, typically primed by primers that bind to primer-binding sites in the adapters adjacent to the DNA molecules to be amplified. Amplification methods can include cycles of extension, denaturation, and annealing resulting from thermocycling, or can be isothermal, such as in transcription-mediated amplification. Other amplification methods include ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and self-sustaining sequence-based replication.

[0156] One or more rounds of amplification can be applied to introduce barcodes into nucleic acid molecules using conventional nucleic acid amplification methods. Amplification can be performed in one or more reaction mixtures. Molecular tags and sample indexes / tags can be introduced simultaneously or in any sequential order. Molecular tags and sample indexes / tags can be introduced before and / or after sequence capture. In some cases, only molecular tags are introduced before probe capture, and sample indexes / tags are introduced after sequence capture. In some cases, both molecular tags and sample indexes / tags are introduced before probe capture. In some cases, sample indexes / tags are introduced after sequence capture. Sequence capture typically involves introducing a single-stranded nucleic acid molecule complementary to a targeted sequence, e.g., a coding sequence of a genomic region, where mutations in such regions are associated with cancer types. Typically, amplification generates multiple nucleic acid amplicons non-uniquely or uniquely tagged with molecular tags and sample indexes / tags, ranging in size from 200 nt to 700 nt, 250 nt to 350 nt, or 320 nt to 550 nt. In some embodiments, the amplicon has a size of about 300 nt. In some embodiments, the amplicon has a size of about 500 nt.

[0157] Barcode Barcodes can be incorporated into or otherwise attached to adapters by chemical synthesis, ligation, overlap-extension PCR, among other methods. In general, the assignment of unique or non-unique barcodes in a reaction follows the methods and systems described by U.S. Patent Application Publication Nos. 20010053519, 20110160078, and U.S. Patent Nos. 6,582,908, 7,537,898, and 9,598,731.

[0158] Tags can be linked to sample nucleic acids randomly or non-randomly. In some cases, they are introduced in an expected ratio of identifiers (i.e., barcode combinations) to microwells. A collection of barcodes can be unique, e.g., all barcodes have different nucleotide sequences. A collection of barcodes can be non-unique, i.e., some barcodes have the same nucleotide sequence and some barcodes have different nucleotide sequences. For example, identifiers can be loaded so that more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000, or 1,000,000,000 identifiers are loaded per genomic sample. In some cases, identifiers may be loaded such that less than 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 or 1,000,000,000 identifiers are loaded per genomic sample. In some cases, the average number of identifiers loaded per sample genome is less than or greater than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000, or 1,000,000,000 identifiers per genome sample.

[0159] A preferred format uses 20-50 different tags ligated to both ends of a target molecule, creating 20-50 x 20-50 tags, or 400-2500 tag combinations. This number of tags is sufficient so that different molecules with the same start and stop points have a high probability (e.g., at least 94%, 99.5%, 99.99%, 99.999%) of receiving different combinations of tags.

[0160] In some cases, the identifier may be an oligonucleotide of predetermined, random, or semi-random sequence. In other cases, multiple barcodes may be used, such that the barcodes are not necessarily unique to one another among the plurality. In this example, the barcode may be attached to individual molecules (e.g., by ligation or PCR amplification) so that the combination of the barcode and the sequence to which it may be attached creates a unique sequence that can be individually tracked. As described herein, detection of a non-uniquely tagged barcode in combination with the initial (start) and / or final (stop) genomic coordinates of a given sequenced sample molecule (i.e., excluding sequence information obtained from barcodes, adapters, etc.) may allow for the assignment of a unique identity to a particular molecule. The length or number of base pairs of an individual sequenced sample molecule (i.e., excluding sequence information corresponding to barcodes, adapters, etc.) may also be used to assign a unique identity to such a molecule. As described herein, fragments from a single strand of nucleic acid, whereby assigned a unique identity, may allow for subsequent identification of fragments from the parental and / or complementary strands.

[0161] Sequencing pipeline The adaptor-flanked sample nucleic acids, with or without prior amplification, may be subjected to sequencing, such as by one or more sequencing devices 107. Sequencing methods include, for example, Sanger sequencing, high-throughput sequencing, pyrosequencing, sequencing-by-synthesis, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing-by-ligation, sequencing-by-hybridization, RNA-Seq (Illumina), Digital Gene Expression (Helicos), next-generation sequencing, single-molecule sequencing-by-synthesis (SMSS) (Helicos), massively parallel sequencing, Clonal Single Molecule Array (Solexa), shotgun sequencing, Ion Torrent, Oxford Nanopore, Roche Genia, Maxim-Gilbert sequencing, primer walking, sequencing using PacBio, SOLiD, Ion Torrent, or Nanopore platforms. Sequencing reactions may be performed in a variety of sample processing units, which may be multiple lanes, multiple channels, multiple wells, or other means of processing multiple sample sets substantially simultaneously. The sample processing unit may also contain multiple sample chambers to allow for simultaneous processing of multiple runs.

[0162] The sequencing reaction can be performed on one or more fragment types known to contain markers for cancer or other diseases. The sequencing reaction can be performed on any nucleic acid fragment present in the sample. The sequencing reaction can provide sequencing of at least 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9% or 100% of a given genome. In other cases, the sequencing reaction can provide sequencing of less than 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9% or 100% of a given genome.

[0163] Simultaneous sequencing reactions can be carried out using multiplex sequencing.In some cases, cell-free polynucleotides can be sequenced at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions.In other cases, cell-free polynucleotides can be sequenced at less than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions.Sequencing reactions can be carried out sequentially or simultaneously.Subsequent data analysis can be carried out on all or part of sequencing reactions. In some cases, data analysis may be performed on at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. In other cases, data analysis may be performed on less than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. An exemplary read depth is 1000 to 50,000 reads per locus (base).

[0164] Sequence analysis pipeline

[0165] Nucleotide variations in sequenced nucleic acids can be determined by comparing the sequenced nucleic acids with a reference sequence.The reference sequence is often a known sequence, for example, a known whole genome sequence or partial genome sequence from a subject, or a whole genome sequence from a human subject.The reference sequence can be hG19.The sequenced nucleic acid can represent the sequence directly determined for the nucleic acid in the sample, or the consensus sequence of the amplification product of such nucleic acid, as described above.Comparison can be performed at one or more designated positions on the reference sequence.When each sequence is maximally aligned, a subset of sequenced nucleic acids can be identified that includes a position corresponding to the designated position of the reference sequence.Within this subset, it can be determined which sequenced nucleic acids, if any, contain nucleotide variations at designated positions, and, if necessary, which contain reference nucleotides (i.e., the same as in the reference sequence).If the number of sequenced nucleic acids in the subset that contain nucleotide variants exceeds a threshold, the variant nucleotide can be called at the designated position. The threshold value can be, among other possibilities, a simple number within the subset containing the nucleotide variant, for example, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 sequenced nucleic acids, or a ratio within the subset containing the nucleotide variant, for example, at least 0.5, 1, 2, 3, 4, 5, 10, 15, or 20 sequenced nucleic acids. Comparison can be repeated for any designated position of interest in the reference sequence. Sometimes, comparison can be performed for designated positions occupying at least 20, 100, 200, or 300 consecutive positions on the reference sequence, for example, 20 to 500 or 50 to 300 consecutive positions.

[0166] The methods of the invention may also be used to diagnose the presence or absence of a condition, particularly cancer, in a subject, to characterize the condition (e.g., to stage the cancer or determine the heterogeneity of the cancer), to monitor the response of the condition to treatment, or to provide a prognosis of the risk of developing the condition or the subsequent course of the condition.

[0167] Various cancers can be detected using the method of the present invention. Like most cells, cancer cells can be characterized by the rate of turnover, in which old cells die and are replaced by newer cells. Generally, dead cells in contact with the vasculature in a given subject can release DNA or DNA fragments into the bloodstream. This also applies to cancer cells during various stages of disease. Depending on the stage of disease, cancer cells can also be characterized by various genetic abnormalities, such as copy number variations and rare mutations. This phenomenon can be used to detect the presence or absence of cancer in individuals using the methods and systems described herein.

[0168] The types and number of cancers that may be detected may include blood cancer, brain cancer, lung cancer, skin cancer, nose cancer, throat cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, skin cancer, bowel cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, solid tumors, heterogeneous tumors, homogeneous tumors, and the like.

[0169] Cancer can be detected from genetic variations including mutations, rare mutations, indels, copy number variations, transversions, translocations, inversions, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structural alterations, gene fusions, chromosomal fusions, gene truncations, gene amplifications, gene duplications, chromosomal lesions, DNA lesions, abnormal changes in chemical modifications of nucleic acids, and abnormal changes in epigenetic patterns.

[0170] Genetic data can also be used to characterize specific forms of cancer. Cancers are often heterogeneous in both composition and stage classification. Genetic profile data can enable the characterization of specific subtypes of cancer, which can be important in the diagnosis or treatment of that specific subtype. This information can also provide subjects or practitioners with clues regarding the prognosis of specific types of cancer, allowing either subjects or practitioners to adapt treatment options as the disease progresses. As some cancers progress, they become more aggressive and genetically unstable. Other cancers may remain benign, inactive, or dormant. The systems and methods of the present disclosure can be useful in determining disease progression.

[0171] The analysis of the present invention is also useful in determining the effectiveness of a particular treatment option. Because more cancers die and shed DNA, if treatment is successful, successful treatment options may increase the amount of copy number variations or rare mutations detected in the subject's blood. In other cases, this may not occur. In another example, perhaps a particular treatment option may be correlated with the genetic profile of cancer over time. This correlation may be useful in selecting a therapy. Furthermore, if cancer is observed to be in remission after treatment, the method of the present invention can be used to monitor residual disease or disease recurrence.

[0172] The methods of the present invention can also be used to detect genetic variations in conditions other than cancer. Immune cells, such as B cells, can undergo rapid clonal expansion in the presence of certain diseases. Clonal expansion can be monitored using copy number variation detection to monitor certain immune conditions. In this example, copy number variation analysis can be performed over time to generate a profile of how a particular disease may be progressing. Detection of copy number variations or even rare mutations can be used to determine how pathogen populations are changing during the course of infection. This can be particularly important during chronic infections, such as HIV / AIDS or hepatitis infections, where viruses can change life cycle states and / or mutate into more virulent forms during the course of infection. Because immune cells attempt to destroy transplanted tissues, the methods of the present invention can be used to determine or profile the host body's rejection activity to monitor the status of transplanted tissues and to alter the course of rejection treatment or prevention.

[0173] Furthermore, the disclosed method can be used to characterize the heterogeneity of an abnormal condition in a subject, the method comprising generating a genetic profile of extracellular polynucleotides in the subject, the genetic profile comprising multiple data resulting from the analysis of copy number variations and rare mutations. In some cases, including but not limited to cancer, diseases can be heterogeneous. Disease cells may not be identical. In the example of cancer, some tumors are known to contain different types of tumor cells, and some cells are at different stages of cancer. In other examples, heterogeneity can include multiple foci of disease. Again, in the example of cancer, there may be multiple tumor foci, and in this case, perhaps one or more foci are the result of metastasis that has spread from the primary site.

[0174] The methods of the present invention can be used to generate or profile a fingerprint or set of data that is a summary of the genetic information from different cells in a heterogeneous disease. This set of data can include analysis of copy number variations and rare mutations, alone or in combination.

[0175] The methods of the present invention can be used to diagnose, prognose, monitor, or observe cancer or other diseases of fetal origin, i.e., these methodologies can be used in pregnant subjects to diagnose, prognose, monitor, or observe cancer or other diseases in prenatal subjects whose DNA and other polynucleotides may be co-circulating with maternal molecules.

[0176] Exemplary Precision Procedures and Applications

[0177] The precision diagnostics provided by the computer system 700 can result in precision treatment plans that can be identified by the computer system 700 (and / or curated by medical professionals). For example, in lung cancer and other diseases, the goal may be to ensure that no superior treatment options exist given the presence of a given variant. For example, EGFR (L858R, exon 19 deletion), BRAF V600E, ALK, and ROS1 fusions can be treated with targeted therapies that may be more appropriate than platinum therapy and chemotherapy. These are examples of major drivers, but other targetable drivers exist, such as MET exon 14 skipping. In another example, for colon cancer, the goal may be to avoid ineffective treatments. Chemotherapy with FOLFIRI or irinotecan regimens can be supplemented with cetuximab or panitumumab if KRAS or NRAS are wild-type. Therefore, the confidence that KRAS and NRAS are wild-type increases the confidence that adding cetuximab or panitumumab is the correct treatment option and no further testing is required.The biological explanation for this is that cetuximab or panitumumab targets EGFR and inhibits its activity.Since RAS (K / NRAS) is downstream of EGFR, when RAS is activated, inhibiting EGFR has minimal or no effect, and therefore cetuximab or panitumumab treatment is inappropriately administered.

[0178] The variants analyzed by the methods and systems of the present disclosure may be loss-of-function variants (e.g., ATM). For example, DNA damage repair (DDR) is a cellular process that functions to maintain genome integrity or stability. Defects or deficiencies in a given DDR mechanism may lead to tumorigenesis or other diseases and may be used to identify test subjects or patients who may benefit from a given targeted therapy. Homologous recombination repair deficiency (HRD), for example, is a cellular phenotype that may make a patient a candidate for the administration of a therapeutic agent, such as a poly ADP-ribose polymerase (PARP) inhibitor. In certain embodiments, a therapy comprising at least one PARP inhibitor may be administered to a subject, where the variant has been determined to be of tumor or non-tumor origin using the methods and systems described herein. In certain embodiments, the PARP inhibitor may include, among others, OLAPARIB, TALAZOPARIB, RUCAPARIB, NIRAPARIB (trade name ZEJULA). In some embodiments, the therapy comprises at least one base excision repair (BER) inhibitor. For example, OLAPARIB can inhibit BER. In certain embodiments, administration of therapy to a subject can be discontinued based on a determination using the methods and systems described herein that the subject has a variant of tumor or non-tumor origin.

[0179] Non-tumor variants can affect the determination of tumor mutation burden (TMB) scores, resulting in artificially high scores if not removed or filtered from the TMB determination. TMB scores are typically used to predict whether a patient will respond to immunotherapy treatment. Thus, the methods and systems provided herein can be used to distinguish between variants of tumor or non-tumor origin as part of a TMB calculation, such as that described in PCT / US2019 / 042882, incorporated herein by reference. In another aspect, the present disclosure provides a method for classifying a subject as a candidate for immunotherapy by determining whether the subject has a variant of tumor or non-tumor origin. In certain embodiments, the methods of the present disclosure include administering one or more immunotherapies to a subject based on a determination of whether a variant is of tumor or non-tumor origin using the methods or systems disclosed herein, alone or in combination with a method for determining a TMB score. In some embodiments, the immunotherapy includes at least one checkpoint inhibitor antibody. In some embodiments, the immunotherapy comprises an antibody against PD-1, PD-2, PD-L1, PD-L2, CTLA-40, OX40, B7.1, B7He, LAG3, CD137, KIR, CCR5, CD27, or CD40. In some embodiments, the immunotherapy comprises administration of a pro-inflammatory cytokine to at least one tumor type. In some embodiments, the immunotherapy comprises administration of T cells to at least one tumor type. In some embodiments, the subject is administered a combination therapy (e.g., immunotherapy + PARPi + chemotherapy), among many other therapies further exemplified herein or otherwise known to those of skill in the art.

[0180] The methods and systems provided herein can be used to evaluate mutations for their prognostic value regarding survival or response to treatment. For example, TP53 mutations can be evaluated for their prognostic and predictive value regarding treatment with ALK inhibitors. The tumor / non-tumor origin determination of variants analyzed herein can also be used to enroll subjects in selected therapies (e.g., TP53 drugs). Another application of the methods and systems described herein can be to analyze less well-studied mutations (e.g., FGFR2 mutations for FGFR inhibitors, or ERBB2 for ERBB2 inhibitors), where distinguishing between variants or between tumor and non-tumor origins can provide confidence that the variant originates from the tumor. In certain embodiments, the methods and systems described herein can be used to monitor molecular response by tracking variants only in tumors to determine variant dynamics over time.

[0181] As additional therapies are developed for various diseases, the interpretation of negative predictions becomes increasingly complex, yet is important for the design of precision therapies. [Example]

[0182] Example 1 Data Processing and Feature Engineering

[0183] A model was developed to predict the tumor or non-tumor origin of variants in an in-house database of over 180,000 plasma samples. The model was trained on multiple Guardant Health, Inc. and external public datasets with known tumor / non-tumor origin variants, and tested on a cohort of samples with matched WBC and plasma cfDNA to validate the results. The model was applied to over 150,000 variants using available data from both in-house and external datasets to obtain a list of variants with an associated probability of being non-tumor origin. At any point, additional data can be added to retrain and reclassify these variants.

[0184] Specifically, the development of this exemplary model implementation was achieved through several steps, which are outlined below.

[0185] 1) Curate a truth dataset for training tumor and non-tumor variants

[0186] To establish a training data set for classifying variants as tumor origin or non-tumor origin, a truth list consisting of well-established tumor variants with high confidence is curated based on authentic external sources of cancer variants, such as National Comprehensive Cancer Network guidelines for cancer treatment, MyCancerGenome, MSKCC OncoKB (classified by level of evidence), literature, and other sources of known variants related to therapy targetable or therapy resistance.The truth list of authentic variants related to clonal hematopoiesis is curated from literature and from the samples sequenced from normal healthy patients at Guardant Health, Inc.

[0187] 2) Sample selection for variant aggregation

[0188] The sample training set consisted of a combination of over 180,000 in-house clinical and research samples, in-house healthy normals, and healthy normal and cancer data from external sources: Jaiswal et al. Age-related clonal hematopoiesis associated with adverse outcomes. N Engl J Med. 2014 Dec 25;371(26):2488-98. Zehir et al. Mutational landscape of metastatic cancer revealed from prospective clinical sequencing of 10,000 patients. Nat Med. 2017 Jun;23(6):703-713. Tate et al. COSMIC: the Catalogue Of Somatic Mutations In Cancer. Nucleic Acids Res. 2019 Jan 8;47(D1):D941-D947. Variants were identified by either gene and mutated amino acid or mutated cDNA, or by variant position and mutated nucleotide. To identify variants originating from solid tissue tumors or originating in the blood due to clonal hematopoietic or hematological malignancies, all variants from healthy normals were pooled with samples from cancer types associated with hematological malignancies.

[0189] 3) Feature Engineering

[0190] A. Long-term longitudinal mutation allele ratio

[0191] The amplitude of the mutant allele fraction (MAF) over time in patients and the dynamic behavior of variants can both indicate the origin of variants, either as tumor or biological noise.Variants that at least two patients have samples from multiple time points (>1) were used in this analysis.For each variant and for each patient, the mean and variation (chi-square statistic, standard deviation) in MAF were calculated across multiple time points, and then this was collapsed for each variant to calculate the overall mean of the MAF mean over time, which indicates amplitude, and the mean of variation or variation, which indicates dynamic behavior.Together, the mean of the mean and the mean of the variation represent the characteristics for subsequent analysis.

[0192] B. Prevalence distribution of variants across cancer types in plasma as an indicator of tumor / non-tumor origin

[0193] The majority of true cancer variants have tumor-specific profiles driven by selection from the perspective of tumor biology. In contrast, variants arising from biological noise or clonal hematopoiesis should arise and thrive regardless of tissue cancer type. To determine the relative homogeneity of cancer type profiles for a given variant, which may indicate a source of biological noise, we calculated the prevalence of each variant in plasma as a percentage of total samples for each cancer category, as previously performed for relative enrichment in plasma across tissues. This prevalence was used as input to logistic regression / random forests to distinguish between background noise and non-random selection among cancer types.

[0194] C. Proportion of samples with hematological malignancies as predictors of tumor / non-tumor origin

[0195] Variants frequently observed in hematological malignancies, such as leukemia and lymphoma, are blood-specific and not solid tumor variants. The proportion of total samples belonging to hematological malignancies was calculated for each variant and used as the single input feature into the logistic regression.

[0196] D. Enrichment in plasma compared to tissue database

[0197] The strong enrichment of a variant in the plasma database compared to the relative prevalence of that variant for the same cancer type in the tissue database suggests that the variant may originate from clonal hematopoiesis rather than from the tumor. The relative ratio of prevalence in plasma compared to tissue and the p-value of the fold change (Fisher's exact test) were calculated and normalized for each variant.

[0198] Example 2 The variation in MAF across time points is related to tumor and non-tumor origin, with non-tumor variants having a lower spread in MAF across time points.

[0199] A dataset containing over 180,000 plasma samples processed in-house was filtered for patients with at least two samples corresponding to different time points, followed by filtering for variants seen in at least three patients. For each patient, the mean variant MAF over time, as well as the standard deviation of MAF over time and chi-squared p-value were calculated for each variant and summarized by the mean at the variant level across patients. The resulting mean MAF, standard deviation of MAF, and chi-squared p-value were used as input features into a logistic regression model that represents both the amplitude and uniformity of variant MAF across time points. Figure 8, Panel A, shows the mean standard deviation (SD) separation of percentages over time for non-tumor and tumor classes.

[0200] The dataset was filtered to a set of 2,509 variants with known tumor and non-tumor labels. Minor class upsampling was performed to ensure balanced class labeling. The resulting training and test sets were 459 and 197 variants, respectively. Five-fold cross-validation was used to measure model performance. Briefly, in K-fold cross-validation, the dataset is divided into k smaller sets, and the following procedure is used for each of the k "folds": a. Train the model using the k-1 folds as training data; b. The resulting model is validated on the remaining portion of the data (i.e., this is used as a test set to calculate a performance measure, eg, accuracy). As a result, the performance measure reported by k-fold cross-validation is the average of the values ​​calculated in the loop.

[0201] The ROC AUC was used to evaluate the performance of the classifier. Briefly, the ROC (Receiver Operating Characteristic) is a probability curve of the true positive rate versus the false positive rate at various threshold settings. The AUC (or "area under the ROC curve") measures the two-dimensional area under the ROC curve and provides a summary measure of performance across all possible classification thresholds, thereby describing the probability that the model will correctly classify a new variant. Because the AUC measures the quality of the model's prediction regardless of which classification threshold is selected, it is commonly used to evaluate binary classifiers because it is invariant with respect to scale (a measure of how well the predictions are ranked, rather than their absolute value) and classification threshold. The ROC AUC for a single input feature in a logistic regression model was 90%, indicating a high predictive value for this feature (Figure 8, Panel B). The accuracy of this model is: (TP + TN) / (TP + TN + FP + FN) = (322 + 356) / (322 + 356 + 56 + 62) = 85.2%, where TP is true positive, TN is true negative, FP is false positive, FN is false negative, and values ​​of 1 and 0 correspond to tumor and non-tumor, respectively (Figure 8, Panel C).

[0202] Example 3 Variant enrichment in plasma compared to tissue as an indicator of non-tumor origin

[0203] Relative enrichment of variants in plasma compared to tissue was performed as follows: samples annotated with cancer type were collapsed into a superset of cancer categories, and all samples without cancer type were removed. The prevalence of individual variants (the number of samples with the variant in question for that cancer type divided by all samples in that cancer type) was calculated separately for each cancer category in the plasma and tissue datasets. The total number of plasma samples was over 180,000, while the total number of tissue samples was 291,847. The odds ratio of prevalence in plasma compared to prevalence in tissue, and the corresponding p-value, were calculated as relative enrichment using Fisher's exact test and used as the single input feature in subsequent analyses. Figure 9, panels A and B, shows that known, well-established non-tumor variants (num_clin, panel A) have a higher ratio compared to tumor variants, regardless of the number of clinical samples observed. The odds ratios and p-values ​​were used as input features to a logistic regression model applying the Yeo-Johnson transformation given by:

number

[0204] The performance of this model is shown to be an ROC AUC of 81% (Figure 9, Panel C). The accuracy of this model is: (TP + TN) / (TP + TN + FP + FN) = (2333 + 3135) / (2333 + 3135 + 561 + 1295) = 5,468 / 7,324 (74.6%), where TP is a true positive, TN is a true negative, FP is a false positive, and FN is a false negative (Figure 9, Panel C).

[0205] Example 4 Non-tumor-specific variants show greater uniformity in prevalence across cancer types compared with known tumor drivers

[0206] The prevalence of each variant was calculated for each cancer type (the number of samples with that variant divided by the total number of samples in that cancer type) as described in Example 2. For variants of non-tumor origin, the prevalence in cancer types was uniformly low compared to tumor-specific variants (Figure 10, Panel A, upper panel), suggesting that variants may exhibit specific prevalence profiles driven by tumor selection and biology (Figure 10, Panel A, lower panel). The prevalence of each variant for each individual cancer type was used as input features into a logistic regression model. The ROC AUC in the test set was 83% (Figure 10, Panel B). The accuracy of this model was 75.5%, based on the following accuracy formula: (TP + TN) / (TP + TN + FP + FN) = (220 + 265) / (220 + 265 + 65 + 92), where TP is a true positive, TN is a true negative, FP is a false positive, and FN is a false negative (Figure 10, Panel C).

[0207] Example 5 High proportion of leukemia / lymphoma samples supporting variants highly indicative of non-tumor status

[0208] Using a tissue dataset from COSMIC containing 281,718 samples, we calculated the total number of samples for a given variant in leukemia / lymphoma / hematological malignancies or healthy individuals compared to all samples containing that variant across all cancer types. The proportion of samples with a variant in leukemia / lymphoma indicates the likelihood that the variant appears in the blood; the higher the probability of blood origin, the lower the probability that it is a tumor variant. Therefore, this proportion in "heme" (hematological malignancies) was used as the single input feature into a logistic regression model. The ROC AUC of this model's performance was 0.94% (Figure 11, Panel A), indicating a high predictive value for this feature. The accuracy of this model is: (TP + TN) / (TP + TN + FP + FN) = (531 + 546) / (531 + 546 + 88 + 72) = 87.1% (Figure 11, Panel B), where TP is a true positive, TN is a true negative, FP is a false positive, and FN is a false negative.

[0209] Example 6 Random Forest Machine Learning Algorithm

[0210] Using the above features, a random forest model was trained to classify somatic mutations detected in plasma according to their source of origin. Briefly, random forest is an ensemble learning method widely used for classification, regression, and other methods of supervised learning. Random forests used for classification are meta-estimators that fit several decision tree classifiers to various subsamples of a dataset during training, and use averaging to determine the predicted class in order to improve the model's predictive accuracy and control overfitting. Here, a random forest model was trained to classify mutations detected in plasma as originating from tumor or non-tumor origin, where the subsample size was always the same as the original input sample size, but samples were drawn with replacement. An implementation of the random forest classifier from the Scikit-learn machine learning library was used (searched from the internet).<URL:https: / / scikit-learn.org> [Retrieved July 25, 2019]). The training of the model (Model 1200) is described according to the machine learning modeling flowchart (see Figure 12).

[0211] Stratified sampling was used to split the original dataset into a training set (80% of the entire dataset) and a test set (20% of the entire dataset). This type of sampling and split was chosen to ensure that the training and test sets had approximately the same percentage of samples of each target class as the full set.

[0212] The best set of model hyperparameters was determined by a grid search technique. Hyperparameters are parameters that are not directly learned within the estimator (for random forests, these parameters are the maximum depth of the decision tree, the number of trees in the forest, etc.). The hyperparameter space was explored for the best cross-validation score by a grid search that exhaustively checks all possible combinations of parameter values, evaluates model performance, and retains the best combination of parameters. 10-fold cross-validation was used to measure the performance of the model given a set of hyperparameters. In K-fold cross-validation, the dataset was divided into k smaller sets. The following procedure was used for each of the k "folds": a. Train the model using the k-1 folds as training data; b. The resulting model is validated on the remaining portion of the data (i.e., this is used as a test set to calculate a performance measure, eg, accuracy). As a result, the performance measure reported by k-fold cross-validation is the average of the values ​​calculated in the loop.

[0213] The optimal set of parameters was determined to be: - Maximum tree depth: 2. - Number of estimators in the forest: 300.

[0214] Finally, the model was retrained using the optimal set of parameters identified during the previous step. The final performance of the model was evaluated using the test set.

[0215] The trained model was used to determine tumor or non-tumor predictions across all variants previously observed in-house, resulting in a list of high-confidence tumor or non-tumor variants.

[0216] Example 7 Random forest classifier (ensemble model) using four input features to predict tumor / non-tumor status

[0217] A random forest classifier was trained and its performance evaluated using a set of 2,509 variants with known tumor and non-tumor labels. The input labels were divided into training and testing, consisting of 1,997 and 500 variants, respectively. To build the model, data observed across more than 110,000 samples was preprocessed into the four features described in Examples 2, 3, 4, and 5, and into genes as variant inputs using one-hot encoding. A grid search, which is an iterative process of scanning through all hyperparameter combinations to find the optimal configuration for the model, was performed on the preprocessed dataset using a manually specified subset of hyperparameters estimated to be appropriate for this model: tree depths of 2, 3, and 4, and the number of estimators of 50, 100, 200, 300, and 500. The optimal hyperparameters for this dataset, as determined by grid search, were determined to be a max depth of 2 and 300 estimators, and a max depth of 3 and 500 estimators.

[0218] To evaluate the impact of each set of hyperparameters on the random forest model, 5-fold cross-validation was performed across input datasets with known tumor and non-tumor labels.

[0219] Figure 13, panel A, shows the performance of a classifier trained with a max depth of 2 and 300 estimators. The ROC AUC was 97% for the training dataset, demonstrating a high probability of correctly classifying new variants. To estimate the accuracy of the classifier, a confusion matrix was created for the predicted and known tumor labels for the test dataset. The accuracy of this model is based on the following accuracy formula: (TP + TN) / (TP + TN + FP + FN) = (148 + 568) / (148 + 568 + 15 + 28) = 94.0%, where TP is true positive, TN is true negative, FP is false positive, and FN is false negative. Thus, the model was able to correctly classify 94.0% of the variants in this test dataset. Figure 13, panel B, shows the performance of the classifier on a validation dataset with variants confirmed in plasma only (tumor) or in the white blood cell (WBC) fraction (non-tumor).

[0220] Performance of a classifier trained with a max depth of 3 and 500 estimators. The ROC AUC was also 91% on the training dataset, and the accuracy ((TP + TN) / (TP + TN + FP + FN)) was (115 + 309) / (115 + 309 + 45 + 31) = 354 / 500 = 70.8%, where TP is true positive, TN is true negative, FP is false positive, and FN is false negative. Given the lower accuracy and equivalent ROC AUC despite the larger depth and number of trees / estimators, we concluded that lower depth and fewer estimators (2 and 300, respectively) were better hyperparameters for this model.

[0221] Example 8 Logistic regression classifier (ensemble model) using four input features to predict tumor / non-tumor status

[0222] Training and test datasets with known tumor / non-tumor labels were created, and the input dataset with features was preprocessed as described in Example 7. For the logistic regression model, multiple feature scaling methods were applied to normalize the data, ensure Gaussian-like distribution, stabilize variance, and minimize skewness. Considering that the input features have different units, zero mean and unit variance were also used to further normalize the features to ensure that they were on the same scale, centered at 0 and with a standard deviation of 1, thus preventing any bias due to different scales or value ranges. Polynomial features were applied to generate a new feature matrix consisting of all polynomial combinations of features. Quadratic inputs were used to prevent overfitting.

[0223] Five-fold cross-validation was performed on the input data set with known tumor and non-tumor labels. ROC AUC was 92% in the test data. The accuracy of our model (TP+TN) / (TP+TN+FP+FN)=(103+334) / (103+334+20+43)=437 / 500(87.4%), which indicates that this model can correctly classify 87.4% of variants in this test data set (where TP is true positive, TN is true negative, FP is false positive, and FN is false negative).

[0224] Example 9 Performance of tumor / non-tumor classification on paired leukocyte and plasma cohorts

[0225] Subsequent validation was performed on a set of 38 paired plasma and white blood cell (WBC) fractions that had previously been sequenced and analyzed in-house and further discussed in Yen et al. Analysis of clonal hematopoiesis in cell-free DNA of advanced cancer patients (Poster #5396, AACR 2019), but were not used in training the model. Briefly, variants detected exclusively in the plasma fraction were labeled as tumor, while variants detected in both the plasma and WBC fractions were labeled as non-tumor. A subset of variants detected in the plasma-WBC set (98 / 648, 15%) was tested to predict tumor / non-tumor status, and its accuracy could be determined based on concordance with the plasma-WBC dataset. Compared with the 5-fold cross-validation performance of the ensemble random forest classifier (97% ROC AUC), 95 / 108 (88.0%) variants predicted as tumor-origin were confirmed as tumor-origin based on plasma-WBC data. Because this model was optimized for sensitivity of tumor variant accuracy, the sensitivity for calling tumor was slightly lower, with 95 / 112 (84.8%) confirmed tumor variants predicted as tumor. The specificity for non-tumor prediction was also lower (30 / 43, 70.0%), likely due to the smaller, more limited training dataset for non-tumor variants. Table 1. Agreement of tumor and non-tumor predictions with paired plasma and WBC samples [Table 1]

[0226] conclusion

[0227] A data-driven bioinformatics approach to classifying tumor and non-tumor variants in cfDNA is crucial for accurate reporting of tumor variants that inform therapy selection in clinical diagnostic settings and for the appropriate identification of biomarkers with predictive or prognostic value in clinical research. As sequenced cfDNA samples accumulate over time, the number of new tumor and non-tumor variants that can be classified using this model should improve along with classifier performance. Sensitive and specific tumor and non-tumor variant classifiers are difficult to obtain from patients, reducing confidence in accurate ctDNA detection in additional paired tissues or leukocyte fractions and complicating sample processing workflows.

[0228] All patent applications, websites, other publications, accession numbers, etc., listed above and below are incorporated by reference in their entirety for all purposes to the same extent as if each individual item were specifically and individually indicated to be incorporated by reference. Where different versions of a sequence are associated with an accession number at different times, the version associated with the accession number as of the effective filing date of this application is meant. The effective filing date means the earlier of the actual filing date or, if applicable, the filing date of the priority application that references that accession number. Similarly, where different versions of a publication, website, etc. are published at different times, the most recently published version as of the effective filing date of this application is meant unless otherwise specified. Any feature, step, element, embodiment, or aspect of the present disclosure may be used in combination with any other feature, step, element, embodiment, or aspect, unless otherwise specifically indicated. While the present disclosure has been described in some detail by way of illustration and example for purposes of clarity and understanding, it will be apparent that certain changes and modifications can be made within the scope of the appended claims. The present invention provides, for example, the following items. (Item 1) 1. A method for distinguishing between nucleic acid variants of tumor origin and nucleic acid variants of non-tumor origin in a cell-free nucleic acid (cfNA) sample obtained from a test subject, at least in part using a computer, comprising: generating or providing, by the computer, at least one tumor variant dataset comprising a population of reference tumor-associated genetic variants, wherein the tumor variant dataset comprises frequency of observations among reference samples, comprising reference body fluid samples and / or reference non-body fluid samples, for one or more tumor-associated genetic variants in the population of reference tumor-associated genetic variants, wherein the reference samples are obtained from a single reference subject and / or from different reference subjects having the same cancer type; determining, by the computer, one or more ratios of the frequencies of observed data between the reference samples for one or more tumor-associated genetic variants in the population of reference tumor-associated genetic variants to produce at least one MAF variance and / or relative prevalence dataset; generating, by the computer, at least one set of probabilities of non-tumor origin from the MAF variance and / or relative prevalence dataset; and using the set of probabilities of non-tumor origin to identify nucleic acid variants detected in the cfNA sample obtained from the test subject as being nucleic acid variants of tumor origin or nucleic acid variants of non-tumor origin. A method comprising: (Item 2) 1. A method for distinguishing between nucleic acid variants of tumor origin and nucleic acid variants of non-tumor origin in a cell-free nucleic acid (cfNA) sample obtained from a test subject, at least in part using a computer, comprising: determining, by the computer, a relative prevalence of one or more tumor-associated genetic variants observed in the one or more reference body fluid samples compared to one or more reference non-body fluid samples to produce at least one relative prevalence dataset; generating, by the computer, at least one set of probabilities of non-tumor origin from the relative prevalence dataset; and using the set of probabilities of non-tumor origin to identify nucleic acid variants detected in the cfNA sample obtained from the test subject as being nucleic acid variants of tumor origin or nucleic acid variants of non-tumor origin. A method comprising: (Item 3) 1. A method for distinguishing between nucleic acid variants of tumor origin and nucleic acid variants of non-tumor origin in a cell-free nucleic acid (cfNA) sample obtained from a test subject, at least in part using a computer, comprising: determining, by the computer, the variation in variant allele fraction (MAF) values, and / or at least one statistic associated therewith, for each of one or more tumor-associated and / or non-tumor-associated genetic variants for at least two different time points to generate at least one relative prevalence dataset; generating, by the computer, at least one set of probabilities of non-tumor origin from the relative prevalence dataset; and using the set of probabilities of non-tumor origin to identify nucleic acid variants detected in the cfNA sample obtained from the test subject as being nucleic acid variants of tumor origin or nucleic acid variants of non-tumor origin. A method comprising: (Item 4) 1. A method for distinguishing between nucleic acid variants of tumor origin and nucleic acid variants of non-tumor origin in a cell-free nucleic acid (cfNA) sample obtained from a test subject, at least in part using a computer, comprising: classifying, by the computer, at least the first nucleic acid variant detected in the cfNA sample obtained from the test subject as a nucleic acid variant of tumor origin if the prevalence of the first nucleic acid variant detected in the cfNA sample is lower than a threshold probability from a set of probabilities of non-tumor origin; and classifying, by the computer, at least the second nucleic acid variant detected in the cfNA sample obtained from the test subject as a nucleic acid variant of non-tumor origin if the prevalence of the second nucleic acid variant detected in the cfNA sample is higher than a threshold probability from the set of probabilities of non-tumor origin, thereby distinguishing between the nucleic acid variants of tumor origin and the nucleic acid variants of non-tumor origin in the cfNA sample obtained from the test subject, wherein the set of probabilities of non-tumor origin generating or providing, by the computer, at least one tumor variant dataset comprising a population of reference tumor-associated genetic variants, wherein the tumor variant dataset comprises frequency of observations among reference samples, comprising reference body fluid samples and / or reference non-body fluid samples, for one or more tumor-associated genetic variants in the population of reference tumor-associated genetic variants, wherein the reference samples are obtained from a single reference subject and / or from different reference subjects having the same cancer type; Determining, by the computer, one or more ratios of the frequencies of the observed data between the reference samples for one or more tumor-associated genetic variants in the population of reference tumor-associated genetic variants to generate at least one relative prevalence data set; and generating, by the computer, the set of probabilities of non-tumor origin from the relative prevalence dataset. A method produced by (Item 5) 1. A method for producing a classifier that uses, at least in part, a computer to distinguish nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as being nucleic acid variants of tumor origin or nucleic acid variants of non-tumor origin, comprising: generating or providing, by the computer, at least one tumor variant dataset comprising a population of reference tumor-associated genetic variants, wherein the tumor variant dataset comprises frequency of observations among reference samples, comprising reference body fluid samples and / or reference non-body fluid samples, for one or more tumor-associated genetic variants in the population of reference tumor-associated genetic variants, wherein the reference samples are obtained from a single reference subject and / or from different reference subjects having the same cancer type; Determining, by the computer, one or more ratios of the frequencies of the observed data between the reference samples for one or more tumor-associated genetic variants in the population of reference tumor-associated genetic variants to generate at least one relative prevalence data set; and applying, by the computer, at least one machine learning model to the relative prevalence dataset to generate at least one set of probabilities of non-tumor origin, thereby producing the classifier that identifies the nucleic acid variants detected in the cfNA sample as nucleic acid variants of tumor origin or nucleic acid variants of non-tumor origin. (Item 6) 1. A method for distinguishing between nucleic acid variants of tumor origin and nucleic acid variants of non-tumor origin in a cell-free nucleic acid (cfNA) sample obtained from a test subject having a type of cancer, at least in part using a computer, comprising: determining, by the computer, a prevalence of one or more genetic variants observed in the cfNA sample to generate a test subject prevalence dataset; comparing, by the computer, the prevalence of one or more genetic variants in the test subject prevalence dataset with the prevalence of the genetic variants observed in a reference cfNA sample obtained from a reference subject having the cancer type; and classifying, by the computer, the given genetic variant in the test subject prevalence dataset as a nucleic acid variant of non-tumor origin if the prevalence of the given genetic variant in the test subject prevalence dataset is below a predetermined threshold associated with the given genetic variant in the reference cfNA sample obtained from the reference subject having the cancer type, thereby distinguishing between the nucleic acid variant of tumor origin and the nucleic acid variant of non-tumor origin in the cfNA sample obtained from the test subject having the cancer type. (Item 7) 1. A method for distinguishing between nucleic acid variants of tumor origin and nucleic acid variants of non-tumor origin in a cell-free nucleic acid (cfNA) sample obtained from a test subject, at least in part using a computer, comprising: determining, by the computer, a prevalence of one or more genetic variants observed in the cfNA sample to generate a test subject prevalence dataset; comparing, by the computer, the prevalence of one or more genetic variants in the test subject prevalence dataset with the prevalence of the genetic variants observed in reference cfNA samples obtained from reference subjects with leukemia, lymphoma, and / or hematological malignancies; and classifying, by the computer, the given genetic variant in the test subject prevalence dataset as a nucleic acid variant of non-tumor origin if the prevalence of the given genetic variant in the test subject prevalence dataset is above a predetermined threshold associated with the given genetic variant in the reference cfNA sample obtained from the reference subject with the leukemia, the lymphoma, and / or the hematological malignancy, thereby distinguishing between the nucleic acid variant of tumor origin and the nucleic acid variant of non-tumor origin in the cfNA sample obtained from the test subject with the leukemia, the lymphoma, and / or the hematological malignancy. (Item 8) The method of any one of the preceding items, comprising identifying genetic variants present in the cfNA sample from sequencing read data originating from cfNA molecules in the cfNA sample. (Item 9) The method of any one of the preceding items, wherein the sequencing read data is obtained from a targeted segment of the cfNA molecule in the cfNA sample. (Item 10) 10. The method of any one of the preceding items, wherein the population of reference tumor-associated genetic variants is obtained from the reference sample. (Item 11) The method of any one of the preceding items, comprising randomly splitting the tumor variant dataset into a training dataset and a testing dataset. (Item 12) The method of any one of the preceding items, wherein the training dataset constitutes about 80% of the tumor variant dataset and the testing dataset constitutes about 20% of the tumor variant dataset. (Item 13) The method of any one of the preceding items, wherein the tumor variant dataset comprises frequencies of observations among reference samples of a given cancer type for one or more tumor-associated genetic variants in the population of reference tumor-associated genetic variants. (Item 14) The method of any one of the preceding items, comprising training a machine learning model using at least a portion of the population of tumor-associated genetic variants to produce a trained machine learning model, wherein the nucleic acid variants of tumor origin and the nucleic acid variants of non-tumor origin detected in the cfNA sample obtained from the test subject are distinguished from each other using the trained machine learning model. (Item 15) 10. The method of any one of the preceding items, wherein the machine learning model is trained using one or more of logistic regression, probit regression, decision tree, random forest, gradient boosting, support vector machine, k-nearest neighbors, and neural network. (Item 16) The method of any one of the preceding items, comprising using a threshold of at least about the 30th percentile probability for a given genetic variant as a cutoff for classification. (Item 17) The method of any one of the preceding items, comprising performing a logistic regression on at least one of the ratios to obtain a given probability of non-tumor origin. (Item 18) The method of any one of the preceding items, wherein the tumor variant dataset comprises mutant allele fraction data observed among reference samples for one or more tumor-associated genetic variants in the population of reference tumor-associated genetic variants. (Item 19) 10. The method of any one of the preceding items, comprising normalizing the tumor variant dataset using one or more data normalization techniques. (Item 20) 10. The method of any one of the preceding items, wherein the data normalization technique comprises min-max normalization and / or z-score normalization. (Item 21) 10. The method of any one of the preceding items, wherein the reference non-body fluid sample comprises a reference tumor tissue sample and / or a reference white blood cell sample. (Item 22) 10. The method of any one of the preceding items, wherein a ratio of the frequency of observations of the given genetic variant in the reference body fluid sample to the frequency of observations of the given genetic variant in the reference non-body fluid sample that is greater than one (1.0) indicates that the given genetic variant is likely to be a nucleic acid variant of non-tumor origin, and wherein the reference non-body fluid sample comprises a reference tumor tissue sample. (Item 23) 21. The method of any one of items 1 to 20, wherein a ratio of the frequency of the observed data of the given genetic variant in the reference body fluid sample to the frequency of the observed data of the given genetic variant in the reference non-body fluid sample that is less than one (1.0) indicates that the given genetic variant is likely to be a nucleic acid variant of non-tumor origin, and wherein the reference non-body fluid sample comprises a reference white blood cell sample. (Item 24) 10. The method of any one of the preceding items, wherein said set of probabilities of non-tumor origin includes at least one set of probabilities of clonal hematopoietic origin. (Item 25) The method of any one of the preceding items, comprising obtaining the cfNA sample from the test subject. (Item 26) The method of any one of the preceding items, comprising selecting one or more therapies for treating a cancer type if one or more tumor-origin nucleic acid variants associated with the cancer type are detected in the cfNA sample obtained from the test subject. (Item 27) The method of any one of the preceding items, comprising administering one or more therapies to the test subject to treat the cancer type if one or more tumor-origin nucleic acid variants associated with the cancer type are detected in the cfNA sample obtained from the test subject. (Item 28) The cancer types are cholangiocarcinoma, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, glioma, astrocytoma, breast cancer, metaplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal cancer, colon cancer, hereditary nonpolyposis colorectal cancer, colorectal adenocarcinoma, gastrointestinal stromal tumor (GIST), endometrial cancer, endometrial stromal sarcoma, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, ocular melanoma, uveal melanoma, gallbladder cancer, gallbladder adenocarcinoma, renal cell carcinoma, clear cell renal cell carcinoma, transitional cell carcinoma, urothelial carcinoma, Wilms' tumor, leukemia, acute lymphocytic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic (CLL), chronic myelogenous (CML), chronic myelomonocytic (CMML), liver cancer, hepatocellular carcinoma ... hi any one of the preceding items, the method is selected from the group consisting of: uterine cancer, cholangiocarcinoma, hepatoblastoma, lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B-cell lymphoma, non-Hodgkin's lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, T-cell lymphoma, non-Hodgkin's lymphoma, precursor T-lymphoblastic lymphoma / leukemia, peripheral T-cell lymphoma, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oropharyngeal cancer, oral squamous cell carcinoma, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic ductal adenocarcinoma, pseudopapillary neoplasm, acinar cell carcinoma, prostate cancer, prostate adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine cancer, gastric cancer, gastrointestinal stromal tumor (GIST), uterine cancer, and uterine sarcoma. (Item 29) The method of any one of the preceding items, wherein the reference tumor-associated genetic variant is selected from the group consisting of a single nucleotide variant (SNV), an insertion or deletion (indel), a copy number variant (CNV), a fusion, a transversion, a translocation, a frameshift, a duplication, a repeat expansion, and an epigenetic variant. (Item 30) 10. The method of any one of the preceding items, wherein the reference sample comprises at least about 25, at least about 50, at least about 100, at least about 200, at least about 300, at least about 400, at least about 500, at least about 600, at least about 700, at least about 800, at least about 900, at least about 1,000, at least about 5,000, at least about 10,000, at least about 15,000, at least about 20,000, at least about 25,000, at least about 30,000 or more body fluid and / or non-body fluid samples. (Item 31) The method of any one of the preceding items, wherein the cfNA sample comprises cell-free deoxyribonucleic acid (cfDNA). (Item 32) The method of any one of the preceding items, wherein the cfRNA sample comprises cell-free ribonucleic acid (cfRNA). (Item 33) The method of any one of the preceding items, wherein the test subject is a mammalian subject. (Item 34) The method of any one of the preceding items, wherein the test subject is a human subject. (Item 35) 10. The method of any one of the preceding items, wherein the reference body fluid sample comprises a plasma sample. (Item 36) 10. The method of any one of the preceding items, wherein the reference body fluid sample comprises a serum sample. (Item 37) 10. The method of any one of the preceding items, wherein the reference non-body fluid sample comprises a cell sample. (Item 38) 10. The method of any one of the preceding items, wherein the reference non-body fluid sample comprises a tissue sample. (Item 39) The method for distinguishing between nucleic acid variants of tumor origin and nucleic acid variants of non-tumor origin in a cell-free nucleic acid (cfNA) sample comprises: (i) the uniformity of the prevalence of the nucleic acid variant across cancer types; (ii) the variation in mutant allele fraction (MAF) of the nucleic acid variant over time; and / or (iii) the prevalence of the nucleic acid variant in hematological cancers, such as leukemia, lymphoma, and / or hematological malignancies. 10. The method of any one of the preceding items, based at least in part on (Item 40) A system including a controller that includes or has access to a computer-readable medium that includes non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least the following: (a) generating or providing at least one tumor variant dataset comprising a population of reference tumor-associated genetic variants, wherein said tumor variant dataset comprises frequency of observations among reference samples comprising reference body fluid samples and / or reference non-body fluid samples for one or more tumor-associated genetic variants in said population of reference tumor-associated genetic variants, wherein said reference samples are obtained from a single reference subject and / or from different reference subjects having the same cancer type; (b) determining one or more ratios of the frequencies of observed data between the reference samples for one or more tumor-associated genetic variants in the population of reference tumor-associated genetic variants to produce at least one relative prevalence dataset; and (c) applying at least one machine learning model to said relative prevalence dataset to generate at least one set of probabilities of non-tumor origin to generate a classifier that identifies said nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as being nucleic acid variants of tumor origin or nucleic acid variants of non-tumor origin. (Item 41) A system including a controller that includes or has access to a computer-readable medium that includes non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least the following: (a) determining the relative prevalence of one or more tumor-associated genetic variants observed in one or more reference body fluid samples compared to one or more reference non-body fluid samples to produce at least one relative prevalence dataset; and (b) generating at least one set of probabilities of non-tumor origin from the relative prevalence dataset to generate a classifier that identifies nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as nucleic acid variants of tumor origin or nucleic acid variants of non-tumor origin. (Item 42) A system including a controller that includes or has access to a computer-readable medium that includes non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least the following: (a) determining, by the computer, the variation in, and / or at least one statistic associated therewith, for each of one or more tumor-associated and / or non-tumor-associated genetic variants for at least two different time points to produce at least one MAF variance and / or relative prevalence data set; and (b) generating at least one set of probabilities of non-tumor origin from the MAF variance and / or relative prevalence dataset to generate a classifier that identifies nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as nucleic acid variants of tumor origin or nucleic acid variants of non-tumor origin. (Item 43) The system of any one of the preceding items, further comprising a nucleic acid sequencer operably connected to the controller, the nucleic acid sequencer configured to provide sequencing read data originating from cfNA molecules in the cfNA sample. (Item 44) The system of any one of the preceding items, wherein the nucleic acid sequencer or another system component is configured to group sequence reads generated by the nucleic acid sequencer into families of sequence reads, each family containing sequence reads generated from a given cfNA molecule in the cfNA sample. (Item 45) 10. The system of any one of the preceding items, comprising a database operably connected to the controller, the database comprising one or more therapies indexed to the nucleic acid variants of tumor origin. (Item 46) The system of any one of the preceding items, further comprising a sample preparation component operably connected to the controller, the sample preparation component configured to prepare the cfNA molecules in the cfNA sample to be sequenced by the nucleic acid sequencer. (Item 47) 10. The system of any one of the preceding items, comprising a nucleic acid amplification component operably connected to the controller, the nucleic acid amplification component configured to amplify at least a targeted segment of the cfNA molecule in the cfNA sample. (Item 48) 10. The system of any one of the preceding items, including a material transfer component operably connected to the controller, the material transfer component configured to transfer one or more materials between at least the nucleic acid sequencer and the sample preparation component. (Item 49) A computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least the following: (a) generating or providing at least one tumor variant dataset comprising a population of reference tumor-associated genetic variants, wherein said tumor variant dataset comprises frequency of observations among reference samples comprising reference body fluid samples and / or reference non-body fluid samples for one or more tumor-associated genetic variants in said population of reference tumor-associated genetic variants, wherein said reference samples are obtained from a single reference subject and / or from different reference subjects having the same cancer type; (b) determining one or more ratios of the frequencies of observed data between the reference samples for one or more tumor-associated genetic variants in the population of reference tumor-associated genetic variants to produce at least one relative prevalence dataset; and (c) applying at least one machine learning model to said relative prevalence dataset to generate at least one set of probabilities of non-tumor origin to generate a classifier that identifies nucleic acid variants detected in the cell-free nucleic acid (cfNA) sample as nucleic acid variants of tumor origin or nucleic acid variants of non-tumor origin. (Item 50) A computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least the following: (a) determining the relative prevalence of one or more tumor-associated genetic variants observed in one or more reference body fluid samples compared to one or more reference non-body fluid samples to produce at least one MAF variance and / or relative prevalence dataset; and (b) generating at least one set of probabilities of non-tumor origin from the MAF variance and / or relative prevalence dataset to generate a classifier that identifies nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as nucleic acid variants of tumor origin or nucleic acid variants of non-tumor origin. (Item 51) A computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least the following: (a) determining, by the computer, the variation in variant allele fraction (MAF) values, and / or at least one statistic associated therewith, for each of one or more tumor-associated and / or non-tumor-associated genetic variants for at least two different time points to produce at least one relative prevalence data set; and (b) generating at least one set of probabilities of non-tumor origin from the relative prevalence dataset to generate a classifier that identifies nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as nucleic acid variants of tumor origin or nucleic acid variants of non-tumor origin. (Item 52) 10. The system or computer-readable medium of any one of the preceding items, wherein the electronic processor further performs at least the step of splitting the tumor variant dataset into a training dataset and a testing dataset. (Item 53) The system or computer-readable medium of any one of the preceding items, wherein the electronic processor further performs at least the steps of training a machine learning model using at least a portion of the population of tumor-associated genetic variants to produce a trained machine learning model, and using the trained machine learning model to identify the nucleic acid variants detected in the cfNA sample as nucleic acid variants of tumor origin or nucleic acid variants of non-tumor origin. (Item 54) 10. The system or computer-readable medium of any one of the preceding items, wherein the electronic processor further performs at least the step of performing a logistic regression on at least one of the ratios to obtain a given probability of non-tumor origin. (Item 55) 10. The system or computer-readable medium of any one of the preceding items, wherein the electronic processor further performs at least the step of normalizing the tumor variant dataset using one or more data normalization techniques. (Item 56) The system or computer-readable medium of any one of the preceding items, wherein the electronic processor further performs the step of selecting one or more therapies for treating the cancer type if, at least, one or more tumor-origin nucleic acid variants associated with the cancer type are detected in the cfNA sample.

Claims

1. 1. A method for distinguishing between nucleic acid variants of tumor origin and nucleic acid variants of non-tumor origin in a cell-free nucleic acid (cfNA) sample obtained from a test subject using, at least in part, a computer, comprising: generating or providing, by the computer, at least one tumor variant dataset comprising a population of reference tumor-associated genetic variants, wherein the tumor variant dataset comprises frequency of observations among reference samples, comprising reference body fluid samples and / or reference non-body fluid samples, for one or more tumor-associated genetic variants in the population of reference tumor-associated genetic variants, wherein the reference samples are obtained from a single reference subject and / or from different reference subjects having the same cancer type; determining, by the computer, one or more ratios of the frequencies of observed data between the reference samples for one or more tumor-associated genetic variants in the population of reference tumor-associated genetic variants to generate at least one MAF variance and / or relative prevalence dataset; generating, by the computer, at least one set of probabilities of non-tumor origin from the MAF variance and / or relative prevalence dataset, comprising performing a logistic regression on at least one of the ratios to obtain the at least one set of probabilities of non-tumor origin from the MAF variance and / or relative prevalence dataset; and using said set of probabilities of non-tumor origin to distinguish nucleic acid variants detected in said cfNA sample obtained from said test subject as nucleic acid variants of tumor origin or nucleic acid variants of non-tumor origin. A method comprising:

2. 10. The method of claim 1, comprising identifying genetic variants present in the cfNA sample from sequencing read data originating from cfNA molecules in the cfNA sample.

3. 3. The method of claim 2, wherein the sequencing read data is obtained from a targeted segment of the cfNA molecule in the cfNA sample.

4. 4. The method of any one of claims 1, 2 and 3, wherein the population of reference tumor-associated genetic variants is obtained from the reference sample.

5. 5. The method of claim 1, comprising randomly splitting the tumor variant dataset into a training dataset and a test dataset.

6. 6. The method of claim 5, wherein the training dataset comprises 80% of the tumor variant dataset and the testing dataset comprises 20% of the tumor variant dataset.

7. 7. The method of any one of claims 1 to 6, wherein the tumor variant dataset comprises frequencies of observations among reference samples of a given cancer type for one or more tumor-associated genetic variants in the population of reference tumor-associated genetic variants.

8. 8. The method of any one of claims 1 to 7, comprising training a machine learning model using at least a portion of the population of tumor-associated genetic variants to produce a trained machine learning model, wherein the nucleic acid variants of tumor origin and the nucleic acid variants of non-tumor origin detected in the cfNA sample obtained from the test subject are distinguished from each other using the trained machine learning model.

9. 9. The method of claim 8, wherein the machine learning model is trained using one or more of logistic regression, probit regression, decision trees, random forests, gradient boosting, support vector machines, k-nearest neighbors, and neural networks.

10. 10. The method of any one of claims 1 to 9, comprising using a threshold of at least the 30±3 percentile probability for a given genetic variant as a cutoff for classification.

11. 11. The method of any one of claims 1 to 10, wherein the tumor variant dataset comprises mutant allele fraction data observed among reference samples for one or more tumor-associated genetic variants in the population of reference tumor-associated genetic variants.

12. 12. The method of any one of claims 1 to 11, comprising normalizing the tumor variant dataset using one or more data normalization techniques.

13. The method of claim 12 , wherein the data normalization technique comprises min-max normalization and / or z-score normalization.

14. The method of claim 1 , wherein the reference non-body fluid sample comprises a reference tumor tissue sample and / or a reference white blood cell sample.

15. 15. The method of any one of claims 1-14, wherein a ratio of the frequency of the observed data for the given genetic variant in the reference body fluid sample to the frequency of the observed data for the given genetic variant in the reference non-body fluid sample that is greater than one (1.0) indicates that the given genetic variant is likely to be a nucleic acid variant of non-tumor origin, and wherein the reference non-body fluid sample comprises a reference tumor tissue sample.

16. 16. The method of any one of claims 1-15, wherein a ratio of the frequency of the observed data for the given genetic variant in the reference body fluid sample to the frequency of the observed data for the given genetic variant in the reference non-body fluid sample that is less than one (1.0) indicates that the given genetic variant is likely to be a nucleic acid variant of non-tumor origin, and wherein the reference non-body fluid sample comprises a reference white blood cell sample.

17. 17. The method of any one of claims 1 to 16, wherein said set of probabilities of non-tumor origin comprises at least one set of probabilities of clonal hematopoietic origin.

18. 18. The method of any one of claims 1 to 17, wherein the cfNA sample is a sample obtained from the test subject.

19. 19. The method of any one of claims 1 to 18, comprising selecting one or more therapies for treating a cancer type if one or more tumor-origin nucleic acid variants associated with the cancer type are detected in the cfNA sample obtained from the test subject.

20. 20. The method of any one of claims 1 to 19, wherein if one or more tumor-origin nucleic acid variants associated with the cancer type are detected in the cfNA sample obtained from the test subject, it is indicated that one or more therapies be administered to the test subject to treat the cancer type.

21. The cancer types are cholangiocarcinoma, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, glioma, astrocytoma, breast cancer, metaplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal cancer, colon cancer, hereditary nonpolyposis colorectal cancer, colorectal adenocarcinoma, gastrointestinal stromal tumor (GIST), endometrial cancer, endometrial stromal sarcoma, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, ocular melanoma, uveal melanoma, gallbladder cancer, gallbladder adenocarcinoma, renal cell carcinoma, clear cell renal cell carcinoma, transitional cell carcinoma, urothelial carcinoma, Wilms' tumor, leukemia, acute lymphocytic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic (CLL), chronic myelogenous (CML), chronic myelomonocytic (CMML), liver cancer, hepatocellular carcinoma, hepatocellular carcinoma, hepatocellular carcinoma, 21. The method of claim 19 or claim 20, wherein the cancer is selected from the group consisting of: uterine cancer, cholangiocarcinoma, hepatoblastoma, lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B-cell lymphoma, non-Hodgkin's lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, T-cell lymphoma, non-Hodgkin's lymphoma, precursor T-lymphoblastic lymphoma / leukemia, peripheral T-cell lymphoma, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oropharyngeal cancer, oral squamous cell carcinoma, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic ductal adenocarcinoma, pseudopapillary neoplasm, acinar cell carcinoma, prostate cancer, prostate adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine cancer, gastric cancer, gastrointestinal stromal tumor (GIST), uterine cancer, and uterine sarcoma.

22. 22. The method of any one of claims 1 to 21, wherein the reference tumor-associated genetic variant is selected from the group consisting of a single nucleotide variant (SNV), an insertion or deletion (indel), a copy number variant (CNV), a fusion, a transversion, a translocation, a frameshift, a duplication, a repeat expansion, and an epigenetic variant.

23. 23. The method of any one of claims 1 to 22, wherein the reference sample comprises at least 25±2 body fluid and / or non-body fluid samples.

24. 24. The method of any one of claims 1 to 23, wherein the cfNA sample comprises cell-free deoxyribonucleic acid (cfDNA).

25. 25. The method of any one of claims 1 to 24, wherein the cfNA sample comprises cell-free ribonucleic acid (cfRNA).

26. 26. The method of any one of claims 1 to 25, wherein the test subject is a mammalian subject.

27. 27. The method of claim 26, wherein the test subject is a human subject.

28. 28. The method of any one of claims 1 to 27, wherein the reference body fluid sample comprises a plasma sample.

29. 29. The method of any one of claims 1 to 28, wherein the reference body fluid sample comprises a serum sample.

30. 30. The method of any one of claims 1 to 29, wherein the reference non-body fluid sample comprises a cell sample.

31. 31. The method of any one of claims 1 to 30, wherein the reference non-body fluid sample comprises a tissue sample.

32. The method for distinguishing between nucleic acid variants of tumor origin and nucleic acid variants of non-tumor origin in a cell-free nucleic acid (cfNA) sample comprises: (i) the uniformity of the prevalence of said nucleic acid variants across cancer types; (ii) the variation in mutant allele fraction (MAF) of the nucleic acid variant over time; and / or (iii) the prevalence of said nucleic acid variants in hematological cancers, such as leukemia, lymphoma, and / or hematological malignancies.

32. The method of any one of claims 1 to 31, based at least in part on

33. A system including a controller that includes or has access to a computer-readable medium that includes non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least the following: (a) generating or providing at least one tumor variant dataset comprising a population of reference tumor-associated genetic variants, said tumor variant dataset comprising frequency of observations among reference samples comprising reference body fluid samples and / or reference non-body fluid samples for one or more tumor-associated genetic variants in said population of reference tumor-associated genetic variants, said reference samples being obtained from a single reference subject and / or from different reference subjects having the same cancer type; (b) determining one or more ratios of the frequencies of observed data between the reference samples for one or more tumor-associated genetic variants in the population of reference tumor-associated genetic variants to produce at least one relative prevalence dataset; and (c) applying at least one machine learning model to said relative prevalence dataset to produce at least one set of probabilities of non-tumor origin to generate a classifier that identifies said nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as being nucleic acid variants of tumor origin or nucleic acid variants of non-tumor origin, comprising performing logistic regression on at least one of said ratios to obtain said at least one set of probabilities of non-tumor origin.

34. 34. The system of claim 33, comprising a nucleic acid sequencer operably connected to the controller, the nucleic acid sequencer configured to provide sequencing read data originating from cfNA molecules in the cfNA sample.

35. 35. The system of claim 34, wherein the nucleic acid sequencer or another system component is configured to group sequence read data generated by the nucleic acid sequencer into families of sequence read data, each family including sequence read data generated from a given cfNA molecule in the cfNA sample.

36. 36. The system of any one of claims 33 to 35, comprising a database operably connected to the controller, the database comprising one or more therapies indexed to nucleic acid variants of tumor origin.

37. 36. The system of claim 34 or claim 35, comprising a sample preparation component operably connected to the controller, the sample preparation component configured to prepare the cfNA molecules in the cfNA sample to be sequenced by the nucleic acid sequencer.

38. 38. The system of any one of claims 33 to 37, comprising a nucleic acid amplification component operably connected to the controller, the nucleic acid amplification component configured to amplify at least a targeted segment of the cfNA molecule in the cfNA sample.

39. 38. The system of claim 37, further comprising a material transfer component operably connected to the controller, the material transfer component configured to transfer one or more materials between at least the nucleic acid sequencer and the sample preparation component.

40. A computer-readable storage medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least the following: (a) generating or providing at least one tumor variant dataset comprising a population of reference tumor-associated genetic variants, said tumor variant dataset comprising frequency of observations among reference samples comprising reference body fluid samples and / or reference non-body fluid samples for one or more tumor-associated genetic variants in said population of reference tumor-associated genetic variants, said reference samples being obtained from a single reference subject and / or from different reference subjects having the same cancer type; (b) determining one or more ratios of the frequencies of observed data between the reference samples for one or more tumor-associated genetic variants in the population of reference tumor-associated genetic variants to produce at least one relative prevalence dataset; and (c) applying at least one machine learning model to said relative prevalence dataset to produce at least one set of probabilities of non-tumor origin to generate a classifier that identifies nucleic acid variants detected in a cell-free nucleic acid (cfNA) sample as being nucleic acid variants of tumor origin or nucleic acid variants of non-tumor origin, comprising performing logistic regression on at least one of said ratios to obtain said at least one set of probabilities of non-tumor origin.

41. 40. The system of any one of claims 33 to 39, wherein the electronic processor further performs at least the step of splitting the tumor variant dataset into a training dataset and a test dataset.

42. 42. The system of any one of claims 33-39 and 41, wherein the electronic processor further performs at least the steps of training a machine learning model using at least a portion of the population of tumor-associated genetic variants to produce a trained machine learning model, and using the trained machine learning model to identify the nucleic acid variants detected in the cfNA sample as nucleic acid variants of tumor origin or nucleic acid variants of non-tumor origin.

43. 43. The system of any one of claims 33 to 39, 41 and 42, wherein the electronic processor further performs at least the step of performing a logistic regression on at least one of the ratios to obtain a given probability of non-tumor origin.

44. 44. The system of any one of claims 33-39 and 41-43, wherein the electronic processor further performs at least the step of normalizing the tumor variant dataset using one or more data normalization techniques.

45. 45. The system of any one of claims 33-39 and 41-44, wherein the electronic processor further performs the step of selecting one or more therapies for treating the cancer type if at least one or more tumor-origin nucleic acid variants associated with the cancer type are detected in the cfNA sample.

Citation Information

Patent Citations

  • Ultra-sensitive detection of circulating tumor DNA through genome-wide integration

    WO2019169042A1

  • Methods and systems for determining the cellular origin of cell-free nucleic acids

    WO2019236478A1