Method and system for determining the cellular origin of cell-free nucleic acids

The method addresses the challenge of distinguishing hematopoietic stem cell and tumor cell nucleic acids by using subclonality scores for computational filtering, enhancing cancer detection and diagnosis.

JP2026012397APending Publication Date: 2026-01-23GUARDANT HEALTH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025185468
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-06-04
Filing Date
2025-11-04
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing methods struggle to accurately distinguish between nucleic acids originating from hematopoietic stem cells and tumor cells in cell-free nucleic acid samples, leading to challenges in early detection and characterization of diseases like cancer, particularly in cases of clonal hematopoiesis of undetermined potential (CHIP).

Method used

A method and system using computational analysis to identify and filter out nucleic acid variants from hematopoietic stem cells by assigning subclonality scores, allowing for the detection of tumor-derived nucleic acids through a target nucleic acid variant filter list, thereby improving the specificity and sensitivity of disease detection.

Benefits of technology

Enhances the ability to detect early-stage cancers by distinguishing between nucleic acids from target and non-target cells, facilitating accurate disease diagnosis and treatment, such as cancer, through improved computational filtering and classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026012397000008
    Figure 2026012397000008
  • Figure 2026012397000009
    Figure 2026012397000009
  • Figure 2026012397000010
    Figure 2026012397000010
Patent Text Reader

Abstract

To provide methods useful in determining the cellular origin of a cell-free nucleic acid (cfNA) fragment from a cfNA sample, such as a liquid biopsy sample.SOLUTION: The methods disclosed herein generally improve the specificity and / or sensitivity of an assay for detecting diseased cell nucleic acid (e.g., cancer cell DNA) in a cfNA sample by identifying variant alleles that are produced by non-target cells, such as hematopoietic stem cells in certain embodiments. Still other aspects include, inter alia, related systems and computer-readable media.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] background The detection and quantification of polynucleotides are important for molecular biology and medical applications such as diagnostic methods. Genetic testing is particularly useful for some diagnostic methods. For example, disorders caused by rare genetic alterations (e.g., sequence variants) or epigenetic marker changes (e.g., cancer and partial or complete aneuploidy) can be detected or more accurately characterized by DNA sequence information.

[0002] Early detection and monitoring of hereditary diseases such as cancer is often necessary for successful treatment or disease management.One approach can include monitoring the sample derived from cell-free nucleic acid, which is a polynucleotide group that can be found in various types of body fluids.In some cases, disease can be characterized or detected based on detecting genetic abnormalities (for example, the copy number variation and / or sequence variation of one or more nucleic acid sequences) or detecting the occurrence of other genetic changes.Cell-free DNA (cfDNA) can contain the genetic abnormalities associated with certain diseases. However, CfDNA present in blood can originate from several cellular sources, both cancerous and noncancerous. One source of potentially problematic cell-free DNA is hematopoietic stem cells, whose mutations can expand clonal populations of blood cells. The acquisition of such somatic mutations that drive clonal expansion without other signs of hematologic malignancy is referred to as cells derived from "clonal hematopoiesis of undetermined potential (CHIP)." See Steensma et al., Blood, 126:9-16 (2015). At least 10% of the elderly population over 70 years of age have CHIP, which results from oligoclonal expansion of mutated hematopoietic stem cells. See Jaiswal et al., N. Engl. J. Med., 371(26):2488-2498 (2014). Hematopoietic stem cells can contain genetic variants in genomic regions associated with cancer, even if the hematopoietic stem cells are noncancerous. It is therefore of interest to identify alleles that contribute to the sampled cfDNA population that are predominantly present in hematopoietic stem cells but absent in cancer cells. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] Steensma et al., Blood, 126:9-16 (2015) [Non-patent document 2] Jaiswal et al., N. Engl. J. Med., 371(26):2488-2498(2014) Summary of the Invention [Means for solving the problem]

[0004] Abstract The present disclosure provides methods, computer-readable media, and systems useful for determining the cellular origin of cfNA fragments from cell-free nucleic acid (cfNA) samples, such as liquid biopsy samples. These aspects, in certain embodiments, improve the specificity and / or sensitivity of assays typically used to detect diseased cell nucleic acids (e.g., cancer cell DNA) in cfNA samples by identifying variant alleles produced by non-target cells, such as hematopoietic stem cells. Furthermore, the methods disclosed herein facilitate the identification of the cellular origin of nucleic acids that are often present in very small amounts in cfNA samples, such as tumor-derived nucleic acids in early-stage cancers. Thus, the methods and related aspects disclosed herein facilitate early detection of disease, among many other uses.

[0005] In one embodiment, the present disclosure provides a method for detecting nucleic acid molecules originating from target cells in a subject, at least partially using a computer. The method includes: (a) receiving test sequence information by a computer, including sequence reads obtained from cell-free nucleic acid (cfNA) fragments from a test sample obtained from the subject. The method also includes (b) identifying that the test sequence information contains at least one allele variant that substantially matches at least one classification allele on the target nucleic acid variant filter list. The classification allele comprises a subclonality score below at least one selected cutoff threshold, thereby indicating that the classification allele is derived from a reference cfNA fragment originating from a target cell, thereby detecting nucleic acid molecules originating from the target cell in the subject. In some embodiments, for example, (b) comprises identifying at least one allelic variant in the test sequence information; mapping the allelic variant to at least one classification allele on the target nucleic acid variant filter list; identifying a subclonality score of the classification allele; and comparing the subclonality score to at least one selected cutoff threshold, wherein when the subclonality score is less than the selected cutoff threshold, it indicates that the classification allele is derived from a reference cfNA fragment originating from the target cell.

[0006] In one embodiment, the present disclosure provides a method for detecting nucleic acid molecules originating from tumor cells in a subject, at least partially using a computer. The method includes: (a) receiving, by a computer, test sequence information including sequence reads obtained from cell-free deoxyribonucleic acid (cfDNA) fragments in a test sample obtained from the subject; (b) removing (e.g., deleting, hiding, ignoring, etc.) one or more of the sequence reads (e.g., including at least a portion of the classification alleles) originating from the subject's hematopoietic stem cells from the test sequence information to generate filtered test sequence information; (c) identifying, by a computer, one or more of the sequence reads in the filtered test sequence information that substantially align with reference sequence information obtained from one or more reference subjects, where the reference sequence information originates from one or more tumor cells in the reference subject, thereby detecting nucleic acid molecules originating from tumor cells in the subject.

[0007] In one embodiment, the present disclosure provides a method for treating a disease in a subject. The method includes (a) receiving test sequence information including sequence reads obtained from cell-free nucleic acid (cfNA) fragments from a test sample obtained from the subject. The method also includes (b) identifying the presence of at least one allelic variant in the test sequence information that substantially matches at least one classification allele on the target nucleic acid variant filter list. The classification allele includes a subclonality score below at least one selected cutoff threshold, thereby indicating that the classification allele is derived from a reference cfNA fragment originating from a diseased cell, thereby diagnosing the disease in the subject. Furthermore, the method also includes (c) administering one or more therapies to the subject, thereby treating the disease in the subject.

[0008] In another aspect, the disclosure provides a method for generating a classifier, or at least a portion thereof, at least partially using a computer, the method comprising: (a) generating by a computer a subclonality score for each allele in a set of classification alleles from sequence information comprising sequence reads obtained from cell-free nucleic acid (cfNA) fragments from one or more reference samples, wherein each classification allele is potentially clinically significant and comprises a minor allele observed at a given locus in the reference samples. The method also includes (b) computationally comparing the subclonality score to at least one selected cutoff threshold, wherein classification alleles having a subclonality score above the selected cutoff threshold indicate that the classification alleles are derived from reference cfNA fragments originating from non-target cells and are added to a non-target nucleic acid variant filter list, and / or classification alleles having a subclonality score below the cutoff threshold indicate that the classification alleles are derived from reference cfNA fragments originating from target cells and are added to a target nucleic acid variant filter list, thereby generating a classifier.

[0009] In another aspect, the present disclosure provides a method for generating a classifier, or at least a portion thereof, at least partially using a computer. The method includes: (a) identifying a set of classification alleles from sequence information including sequence reads obtained from cell-free nucleic acid (cfNA) fragments from one or more reference samples, where each classification allele is potentially clinically significant and includes a minor allele observed at a given locus in the reference samples. The method also includes: (b) determining a minor allele frequency (MAF) value for each classification allele in each of the reference samples from the sequence information, and (c) determining a maximum minor allele frequency (maxMAF) value for each of the reference samples. The method also includes: (d) calculating, by a computer, the ratio of the MAF value to the maxMAF value for at least some of the reference samples for each classification allele observed in a given reference sample to generate a ratio value. The method also includes (e) calculating, by a computer, the ratio of the number of times a given classification allele in at least some of the reference samples had a ratio value less than at least one selected clonality cutoff value to the total number of times the given classification allele appeared in at least some of the reference samples, to generate a subclonality score for each of the classification alleles in at least some of the reference samples. The method further includes (f) comparing, by a computer, the subclonality score to at least one selected cutoff threshold, wherein classification alleles with a subclonality score above the selected cutoff threshold indicate that the classification alleles originated from reference cfNA fragments originating from non-target cells and are added to a non-target nucleic acid variant filter list, and / or classification alleles with a subclonality score below the cutoff threshold indicate that the classification alleles originated from reference cfNA fragments originating from target cells and are added to a target nucleic acid variant filter list, thereby generating a classifier.

[0010] In another aspect, the present disclosure provides a method for creating a database of subclonality scores for use in classifying the cellular origin of cell-free nucleic acid (cfNA) fragments in test samples obtained from a subject. The method includes: (a) identifying a set of classification alleles from sequence information including sequence reads obtained from cell-free nucleic acid (cfNA) fragments from one or more reference samples, where each classification allele is potentially clinically significant and includes a minor allele observed at a given locus in the reference samples. The method also includes: (b) determining a minor allele frequency (MAF) value for each classification allele in each of the reference samples from the sequence information; (c) determining a maximum minor allele frequency (maxMAF) value for each of the reference samples; and (d) calculating a ratio of the MAF value to the maxMAF value for at least some of the reference samples for each classification allele observed in a given reference sample to generate a ratio value. The method also includes (e) calculating by a computer the ratio of the number of times a given classification allele in at least a portion of the reference samples had a ratio value less than at least one selected clonality cutoff value to the total number of times that the given classification allele appeared in at least a portion of the reference samples to generate a subclonality score for each of the classification alleles in at least a portion of the reference samples. The method further includes (f) non-temporarily storing the subclonality scores indexed to the corresponding classification alleles in a database system, thereby creating a database of subclonality scores for use in classifying the cellular origin of cfNA fragments in a test sample obtained from a subject.

[0011] In some embodiments, the method disclosed herein comprises identifying the classification allele set, wherein the method comprises determining the MAF value for each somatic nucleic acid variant at each locus in a potentially clinically significant target genomic locus set from sequence information obtained from a reference sample, wherein the target genomic locus set is identical in each reference sample, and determining the maxMAF value for each reference sample to generate allele information.In certain embodiments, the MAF for each classification allele is less than about 2%.In some embodiments, the MAF for each classification allele is less than about 1%.

[0012] In certain embodiments, the methods disclosed herein include generating a classifier using clinical information indexed to a reference sample. In some embodiments, the methods disclosed herein include detecting nucleic acid molecules originating from target cells in a subject using clinical information indexed to a test sample. In certain embodiments, the clinical information is selected from the group consisting of age, sex, race, weight, body mass index (BMI), medical history, smoking, alcohol consumption, etc. In other exemplary embodiments, subclonal lists (e.g., target nucleic acid variant filter lists or non-target nucleic acid variant filter lists) are generated for various subsets of samples, e.g., based on minimum maxMAF, and maxMAFs are called based on known driver mutations, etc. In some embodiments, subclonal lists are generated based on specific indications (e.g., a given cancer type (e.g., lung, colorectal, etc.)). In certain embodiments, a machine learning classifier is trained based on one or more features, including mutant allele frequency, subclonal ratio, gene type, variants associated with hematological malignancies, patient age, observation of other CHIP variants, cancer type, etc.

[0013] In some embodiments, the methods disclosed herein include determining a subclonality score using the frequency of each MAF / maxMAF value for each of the classified alleles. In certain embodiments, the selected clonality boundary value is within the range of about 1% to about 99%. In some of these embodiments, for example, the selected clonality boundary value is about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90%. In some embodiments, the selected cutoff threshold is within the range of about 1% to about 99%. In some of these embodiments, for example, the selected cutoff threshold is about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90%.

[0014] In certain embodiments, the methods disclosed herein comprise comparing the subclonality score with a plurality of selected cutoff thresholds. In some of these embodiments, for example, the plurality of selected cutoff thresholds comprise a first cutoff threshold and a second cutoff threshold, the first cutoff threshold being greater than the second cutoff threshold, and classification alleles with subclonality scores greater than the first cutoff threshold indicate that the classification alleles are derived from reference cfNA fragments originating from non-target cells, and are added to a non-target nucleic acid variant filter list; and / or classification alleles with subclonality scores less than the second cutoff threshold indicate that the classification alleles are derived from reference cfNA fragments originating from target cells, and are added to a target nucleic acid variant filter list.

[0015] In some embodiments, the methods disclosed herein include classifying an allelic variant in the test sequence information that substantially matches at least one classification allele on the non-target nucleic acid variant filter list as originating from a target cell when the allelic variant comprises an MAF greater than about 1%. In certain embodiments, the methods disclosed herein include classifying an allelic variant in the test sequence information that substantially matches at least one classification allele on the non-target nucleic acid variant filter list as originating from a target cell when the allelic variant comprises a truncation, an indel, and / or a splice site variant.

[0016] In certain embodiments, the methods disclosed herein comprise determining the frequency of each ratio value for a given classification allele in at least a portion of a reference sample. In some embodiments, the methods disclosed herein comprise using the classifier to determine whether a test sample obtained from a subject contains cfNA fragments originating from the target cells. In certain embodiments, the methods disclosed herein comprise using the classifier to determine whether a test sample obtained from a subject contains cfNA fragments originating from non-target cells. In some embodiments, the database comprises a target nucleic acid variant filter list and / or a non-target nucleic acid variant filter list.

[0017] In certain embodiments, the non-target cells comprise non-diseased cells. In some embodiments, the non-target cells comprise hematopoietic stem cells. In certain embodiments, the non-target cells comprise non-tumor cells. In some embodiments, the non-target cells comprise maternal cells. In certain embodiments, the non-target cells comprise transplant recipient cells.

[0018] In certain embodiments, the target cells comprise diseased cells. In some embodiments, the target cells comprise tumor cells. In some embodiments, the target cells comprise fetal cells. In certain embodiments, the target cells comprise transplant donor cells.

[0019] In certain embodiments, the methods disclosed herein include treating a disease. In some of these embodiments, for example, the disease includes cancer and the therapy includes at least one immunotherapy. Typically, the subject is a mammalian subject (e.g., a human subject).

[0020] In some embodiments, the methods disclosed herein further comprise obtaining a test sample from a subject. The test sample is typically selected from the group consisting of blood, plasma, serum, sputum, urine, semen, vaginal fluid, stool, synovial fluid, cerebrospinal fluid, saliva, etc. In some embodiments, the methods disclosed herein further comprise generating test sequence information from cfNA fragments in the test sample. In some embodiments, the methods disclosed herein further comprise amplifying a segment of the cfNA fragment containing the target genomic locus to generate an amplified nucleic acid. In certain embodiments, the methods disclosed herein further comprise sequencing the cfNA fragments in the test sample to generate test sequence information. In some of these embodiments, the test sequence information is obtained from a targeted segment of the cfNA fragment in the test sample, the targeted segment being obtained by selectively enriching one or more regions from the cfNA fragment in the test sample prior to sequencing. In certain embodiments, the methods disclosed herein further comprise amplifying the obtained targeted segment prior to sequencing. In certain embodiments, the methods disclosed herein further comprise attaching one or more adapters comprising barcodes to the cfNA fragments and / or the amplified targeting segments prior to sequencing. In certain embodiments, the sequencing is selected from the group consisting of targeted sequencing, bisulfite sequencing, intron sequencing, exome sequencing, and whole genome sequencing.

[0021] In yet another aspect, the present disclosure provides a system comprising a computer-readable medium or a controller that can access the computer-readable medium, the system comprising computer-executable non-transitory instructions that, when executed by at least one electronic processor, at least: (a) receive test sequence information comprising sequence reads obtained from cell-free nucleic acid (cfNA) fragments from a test sample obtained from a subject; and (b) identify at least one allelic variant present in the test sequence information that substantially matches at least one classification allele on a target nucleic acid variant filter list, wherein the classification allele comprises a subclonality score below at least one selected cutoff threshold, thereby indicating that the classification allele is derived from a reference cfNA fragment originating from a target cell, thereby indicating that the allelic variant in the test sequence information originated from a target cell within the subject. In some embodiments, for example, (b) comprises identifying at least one allelic variant in the test sequence information; mapping the allelic variant to at least one classification allele on the target nucleic acid variant filter list; identifying a subclonality score of the classification allele; and comparing the subclonality score to at least one selected cutoff threshold, wherein when the subclonality score is less than the selected cutoff threshold, it suggests that the classification allele is derived from a reference cfNA fragment originating from the target cell.

[0022] In yet another aspect, the present disclosure provides a system comprising a computer-readable medium or a controller that can access the computer-readable medium, the system comprising computer-executable non-transitory instructions that, when executed by at least one electronic processor, at least: (a) receive test sequence information including sequence reads obtained from cell-free deoxyribonucleic acid (cfDNA) fragments in a test sample obtained from a subject; (b) remove (e.g., delete, hide, ignore, etc.) one or more sequence reads (e.g., including at least a portion of classification alleles) originating from the subject's hematopoietic stem cells from the test sequence information to generate filtered test sequence information; and (c) identify in the filtered test sequence information that the one or more sequence reads that substantially align with reference sequence information obtained from one or more reference subjects are present, where the reference sequence information originates from tumor cells in the reference subject, thereby indicating that the test sample contains one or more cfDNA fragments originating from tumor cells in the subject.

[0023] In yet another aspect, the present disclosure provides a system comprising a computer-readable medium or a controller that can access the computer-readable medium, the computer-executable non-transitory instructions, which when executed by at least one electronic processor, at least: (a) generate a subclonality score for each allele in a classification allele set from sequence information comprising sequence reads obtained from cell-free nucleic acid (cfNA) fragments from one or more reference samples, wherein each classification allele is potentially clinically significant and includes a minor allele observed at a given locus in the reference samples. and (b) comparing the subclonality score to at least one selected cutoff threshold, wherein classification alleles having a subclonality score above the selected cutoff threshold indicate that the classification alleles are derived from reference cfNA fragments originating from non-target cells, and the classification alleles are added to a non-target nucleic acid variant filter list, and / or classification alleles having a subclonality score below the cutoff threshold indicate that the classification alleles are derived from reference cfNA fragments originating from target cells, and the classification alleles are added to a target nucleic acid variant filter list.

[0024] In yet another aspect, the present disclosure provides a system comprising a computer-readable medium or a controller that can access the computer-readable medium, the system comprising computer-executable non-transitory instructions, which when executed by at least one electronic processor, at least: (a) identify a set of classification alleles from sequence information comprising sequence reads obtained from cell-free nucleic acid (cfNA) fragments from one or more reference samples, wherein each classification allele is potentially clinically significant and comprises a minor allele observed at a given locus in the reference samples; (b) determine from the sequence information a minor allele frequency (MAF) value for each classification allele in each of the reference samples; (c) determine, for each of the reference samples, a maximum minor allele frequency (maxMAF) value; (d) for each classification allele observed in a given reference sample, calculate a ratio of the MAF value to the maxMAF value for at least some of the reference samples to generate a ratio value; and (e) identify a set of classification alleles from sequence information comprising sequence reads obtained from cell-free nucleic acid (cfNA) fragments from one or more reference samples, wherein each classification allele is potentially clinically significant and comprises a minor allele observed at a given locus in the reference samples. (f) calculating for each of the classification alleles a ratio of the number of times that a given classification allele in at least a portion of the reference samples had a ratio value less than at least one selected clonality boundary value to the total number of times that the given classification allele appeared in at least a portion of the reference samples to generate a subclonality score for each of the classification alleles in at least a portion of the reference samples; and (f) comparing the subclonality score to at least one selected cutoff threshold, wherein classification alleles with subclonality scores above the selected cutoff threshold indicate that the classification alleles were derived from reference cfNA fragments originating from non-target cells, and the classification alleles are added to a non-target nucleic acid variant filter list, and / or classification alleles with subclonality scores below the cutoff threshold indicate that the classification alleles were derived from reference cfNA fragments originating from target cells, and the classification alleles are added to a target nucleic acid variant filter list.

[0025] In some embodiments, the systems disclosed herein include a nucleic acid sequencer operably connected to a controller, the nucleic acid sequencer configured to provide sequence information from cfNA fragments in a test sample and / or a reference sample. In certain of these embodiments, the nucleic acid sequencer is configured to perform pyrosequencing, bisulfite sequencing, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, ligation-based sequencing, or hybridization-based sequencing on the nucleic acids to generate sequencing reads.

[0026] In some embodiments, the system disclosed herein includes a sample preparation component operably connected to a controller, the sample preparation component configured to prepare cfNA fragments to be sequenced by a nucleic acid sequencing device. In some of these embodiments, the sample preparation component is configured to selectively enrich regions from the cfNA fragments in the test sample and / or the reference sample. In certain embodiments, the sample preparation component is configured to attach one or more adapters containing barcodes to the cfNA fragments.

[0027] In certain embodiments, the systems disclosed herein include a nucleic acid amplification component operably connected to a controller, the nucleic acid amplification component configured to amplify cfNA fragments in the test sample and / or the reference sample. In some of these embodiments, the nucleic acid amplification component is configured to amplify selectively enriched regions of cfNA fragments in the test sample and / or the reference sample. In some embodiments, the systems disclosed herein include a material transfer component operably connected to the controller, the material transfer component configured to transfer one or more materials between a nucleic acid sequencing device, a nucleic acid amplification component, and / or a sample preparation component. In certain embodiments, the systems disclosed herein include a database operably connected to the controller, the database including a non-target nucleic acid variant filter list and / or a target nucleic acid variant filter list.

[0028] In yet another aspect, the present disclosure provides a computer-readable medium comprising computer-executable non-transitory instructions that, when executed by at least one electronic processor, at least: (a) receive test sequence information comprising sequence reads obtained from cell-free nucleic acid (cfNA) fragments from a test sample obtained from a subject; and (b) identify at least one allelic variant present in the test sequence information that substantially matches at least one classification allele on a target nucleic acid variant filter list, wherein the classification allele comprises a subclonality score below at least one selected cutoff threshold, thereby indicating that the classification allele is derived from a reference cfNA fragment originating from a target cell, thereby indicating that the allelic variant in the test sequence information originated from a target cell within the subject. In some embodiments, for example, (b) comprises identifying at least one allelic variant in the test sequence information; mapping the allelic variant to at least one classification allele on the target nucleic acid variant filter list; identifying a subclonality score of the classification allele; and comparing the subclonality score to at least one selected cutoff threshold, wherein when the subclonality score is less than the selected cutoff threshold, it indicates that the classification allele is derived from a reference cfNA fragment originating from the target cell.

[0029] In another aspect, the present disclosure provides a computer-readable medium including computer-executable non-transitory instructions that, when executed by at least one electronic processor, at least: (a) receive test sequence information including sequence reads obtained from cell-free deoxyribonucleic acid (cfDNA) fragments in a test sample obtained from a subject; (b) remove (e.g., delete, hide, ignore, etc.) one or more of the sequence reads (e.g., including at least a portion of classification alleles) originating from the subject's hematopoietic stem cells from the test sequence information to generate filtered test sequence information; and (c) identify in the filtered test sequence information that the one or more sequence reads substantially align with reference sequence information obtained from one or more reference subjects, where the reference sequence information originated from tumor cells in the reference subject, thereby indicating that the test sample contains one or more cfDNA fragments originating from tumor cells in the subject.

[0030] In another aspect, the disclosure provides a computer-readable medium comprising computer-executable non-transitory instructions that, when executed by at least one electronic processor, at least: (a) generate a subclonality score for each allele in a set of classification alleles from sequence information comprising sequence reads obtained from cell-free nucleic acid (cfNA) fragments from one or more reference samples, wherein each classification allele is potentially clinically significant and comprises a minor allele observed at a given locus in the reference samples; and (b) compare the subclonality score to at least one selected cutoff threshold, wherein classification alleles with a subclonality score above the selected cutoff threshold indicate that the classification alleles are derived from reference cfNA fragments originating from non-target cells, and the classification alleles are added to a non-target nucleic acid variant filter list; and / or classification alleles with a subclonality score below the cutoff threshold indicate that the classification alleles are derived from reference cfNA fragments originating from target cells, and the classification alleles are added to a target nucleic acid variant filter list.

[0031] In another aspect, the present disclosure provides a computer-readable medium comprising computer-executable non-transitory instructions, which when executed by at least one electronic processor, at least: (a) identify a set of classification alleles from sequence information comprising sequence reads obtained from cell-free nucleic acid (cfNA) fragments from one or more reference samples, wherein each classification allele is potentially clinically significant and comprises a minor allele observed at a given locus in the reference samples; (b) determine from the sequence information a minor allele frequency (MAF) value for each classification allele in each of the reference samples; (c) determine for each of the reference samples a maximum minor allele frequency (maxMAF) value; (d) for each classification allele observed in a given reference sample, calculate a ratio of the MAF value to the maxMAF value for at least a portion of the reference samples to generate a ratio value; and (e) identify a set of classification alleles from sequence information comprising sequence reads obtained from cell-free nucleic acid (cfNA) fragments from one or more reference samples, wherein each classification allele is potentially clinically significant and comprises a minor allele observed at a given locus in the reference samples. (f) calculating for each of the classification alleles a ratio of the number of times the allele had a ratio value less than at least one selected clonality boundary value to the total number of times that given classification allele appeared in at least a portion of the reference samples to generate a subclonality score for each of the classification alleles in at least a portion of the reference samples; and (f) comparing the subclonality score to at least one selected cutoff threshold, wherein classification alleles with subclonality scores above the selected cutoff threshold indicate that the classification alleles are derived from reference cfNA fragments originating from non-target cells and are added to a non-target nucleic acid variant filter list, and / or classification alleles with subclonality scores below the cutoff threshold indicate that the classification alleles are derived from reference cfNA fragments originating from target cells and are added to a target nucleic acid variant filter list.

[0032] In some embodiments of the systems or computer-readable media disclosed herein, the computer-readable medium comprises computer-executable non-transitory instructions that, when executed by at least one electronic processor, determine, from sequence information obtained from reference samples, a value of a MAF for each potentially clinically significant somatic nucleic acid variant at each locus in a set of target genomic loci, where the set of target genomic loci are identical in each reference sample, and further at least determine a value of maxMAF for each of the reference samples to generate allele information.

[0033] In certain embodiments of the systems or computer-readable media disclosed herein, the computer-readable medium comprises computer-executable, non-transitory instructions that, when executed by at least one electronic processor, at least further perform the following: generate a non-target nucleic acid variant filter list and / or a target nucleic acid variant filter list using clinical information indexed to the reference sample. In certain embodiments of the systems or computer-readable media disclosed herein, the computer-executable, non-transitory instructions that, when executed by at least one electronic processor, at least further perform the following: detect cfNA fragments originating from target cells in the subject using clinical information indexed to the test sample. In certain embodiments of the systems or computer-readable media disclosed herein, the computer-executable, non-transitory instructions that, when executed by at least one electronic processor, at least further perform the following: determine a subclonality score using the frequency of each MAF / max-MAF value for each classification allele.

[0034] In certain embodiments of the systems or computer-readable media disclosed herein, the computer-readable medium comprises computer-executable non-transitory instructions which, when executed by at least one electronic processor, further perform at least the following: comparing the subclonality score to a plurality of selected cutoff thresholds, wherein the plurality of selected cutoff thresholds comprises a first cutoff threshold and a second cutoff threshold, the first cutoff threshold being greater than the second cutoff threshold; and wherein classification alleles having a subclonality score greater than the first cutoff threshold indicate that the classification alleles are derived from reference cfNA fragments originating from non-target cells, and the classification alleles are added to a non-target nucleic acid variant filter list; and / or wherein classification alleles having a subclonality score less than the second cutoff threshold indicate that the classification alleles are derived from reference cfNA fragments originating from target cells, and the classification alleles are added to a target nucleic acid variant filter list. In certain embodiments of the systems or computer-readable media disclosed herein, the computer-readable medium comprises computer-executable non-transitory instructions that, when executed by at least one electronic processor, at least further classify an allelic variant in the test sequence information that substantially matches at least one classification allele on the non-target nucleic acid variant filter list as originating from a target cell when the allelic variant in the test sequence information substantially matches at least one classification allele on the non-target nucleic acid variant filter list comprises a MAF greater than about 1%. In some embodiments of the systems or computer-readable media disclosed herein, the computer-readable medium comprises computer-executable non-transitory instructions that, when executed by at least one electronic processor, at least further classify an allelic variant in the test sequence information that substantially matches at least one classification allele on the non-target nucleic acid variant filter list as originating from a target cell when the allelic variant comprises a truncation, an indel, and / or a splice site variant.

[0035] In some embodiments of the systems or computer-readable media disclosed herein, the computer-readable medium comprises computer-executable non-transitory instructions that, when executed by at least one electronic processor, at least further determine the frequency of each ratio value for a given classification allele in at least a portion of the reference samples. In certain embodiments of the systems or computer-readable media disclosed herein, the computer-readable medium comprises computer-executable non-transitory instructions that, when executed by at least one electronic processor, at least further determine whether a test sample obtained from a subject contains cfNA fragments originating from target cells using the target nucleic acid variant filter list. In some embodiments of the systems or computer-readable media disclosed herein, the computer-readable medium comprises computer-executable non-transitory instructions that, when executed by at least one electronic processor, at least further determine whether a test sample obtained from a subject contains cfNA fragments originating from non-target cells using the non-target nucleic acid variant filter list. [Brief explanation of the drawings]

[0036] The accompanying drawings (also referred to herein as "Figure" and "FIG"), which are incorporated into and form a part of this specification, illustrate certain embodiments and, together with the specification, serve to explain certain principles of the methods, computer-readable media, and systems disclosed herein. The description provided herein is better understood when read in conjunction with the accompanying drawings, which are included by way of example and in no way by way of limitation. It will be understood that like reference numerals identify like components throughout the drawings, unless the context dictates otherwise. It will also be understood that some or all of the drawings may be schematic for illustrative purposes and do not necessarily indicate the actual relative size or position of the elements shown.

[0037] [Figure 1] 1A and 1B are histograms of two alleles (FIG. 1A shows classification allele 1, and FIG. 1B shows classification allele 2) distinguished by estimated subclonality scores based on a clonality boundary value set at a threshold of 50%. According to some embodiments of the present invention, allele 1 is negative (i.e., not indicative of a cancer cell origin (likely a hematopoietic stem cell)), whereas allele 2 is positive (i.e., indicative of a cancer cell origin) because it is present in more than 50% of the subjects in the reference sample database. In each of FIGS. 1A and 1B, the Y-axis shows the number of records, and the X-axis shows the distribution of the MAF / maxMAF ratio.

[0038] [Figure 2] FIG. 2 is a flow chart that schematically depicts exemplary method steps for detecting nucleic acid molecules originating from target cells in a subject, according to some embodiments of the present invention.

[0039] [Figure 3] FIG. 3 is a flow chart that schematically depicts exemplary method steps for detecting nucleic acid molecules originating from tumor cells in a subject, according to some embodiments of the present invention.

[0040] [Figure 4] FIG. 4 is a flow chart that schematically depicts exemplary method steps for treating a disease in a subject, according to some embodiments of the present invention.

[0041] [Figure 5] FIG. 5 is a flow chart that schematically illustrates exemplary method steps for generating a classifier, according to some embodiments of the present invention.

[0042] [Figure 6] FIG. 6 is a flow chart that schematically illustrates exemplary method steps for generating a classifier, according to some embodiments of the present invention.

[0043] [Figure 7] FIG. 7 is a schematic diagram of an exemplary system suitable for use with certain embodiments of the present invention.

[0044] [Figure 8] Figures 8A-C show Kaplan-Meier plots for patient data without filtering (Figure 8A), with tissue filtering (Figure 8B), and with classifier filtering (Figure 8C; i.e., using subclonality scores). In each plot shown in Figures 8A-C, the undetected curve is the top curve, and the detected curve is the bottom curve.

[0045] [Figure 9] Figure 9 shows a plot of the range of allele frequencies (x-axis) against the number of variants observed (y-axis) for the different filter scenarios shown in Figures 8A-C. DETAILED DESCRIPTION OF THE INVENTION

[0046] definition In order to more readily understand this disclosure, certain terms are first defined below. Additional definitions for these terms and other terms may be found throughout the specification. In the event that a definition of a term set forth below conflicts with a definition in a patent application or issued patent incorporated by reference, the definition set forth in this application should be used to understand the meaning of that term.

[0047] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a method" includes one or more methods and / or steps of the type described herein and / or that would become apparent to one of ordinary skill in the art upon reading this disclosure. It will also be recognized that there is an implicit "about" before temperatures, concentrations, times, numbers of bases or base pairs, coverage, etc. discussed in this disclosure, and that slight and insubstantial equivalents are within the scope of this disclosure. In this application, the use of the singular includes the plural unless specifically stated otherwise. Additionally, the use of "comprise," "comprises," "comprising," "contain," "contains," "containing," "include," "includes," and "including" is not intended to be limiting.

[0048] It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting. Furthermore, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. In describing and claiming the methods, computer-readable media, and systems, the following terminology and grammatical variations thereof will be used in accordance with the definitions set forth below.

[0049] About: As used herein, "about" or "approximately," when applied to one or more values ​​or elements of interest, refers to a value or element similar to a stated reference value or element. In certain embodiments, the term "about" or "approximately," unless otherwise stated or clear from the context, refers to a range of values ​​or elements that fall within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1% or less in either direction (greater or lower) of the stated reference value or element (except where such number exceeds 100% of the possible value or element).

[0050] Adapter: As used herein, "adapter" refers to a short nucleic acid (e.g., less than about 500 nucleotides in length, less than about 100 nucleotides in length, or less than about 50 nucleotides in length) that is typically at least partially double-stranded and used to ligate to one or both ends of a given sample nucleic acid molecule. The adapter may contain nucleic acid primer binding sites that enable amplification of the nucleic acid molecule adjacent to the adapter at both ends, and / or may contain sequencing primer binding sites, including primer binding sites for sequencing applications, such as various next-generation sequencing (NGS) applications. The adapter may also contain a binding site for a capture probe (e.g., an oligonucleotide attached to a flow cell support, etc.). The adapter may also contain a nucleic acid tag as described herein. The nucleic acid tag is typically positioned relative to the amplification primer binding site and the sequencing primer binding site so that the nucleic acid tag is included in the amplicon and sequencing read of a given nucleic acid molecule. The same or different adapters can be ligated to each end of a nucleic acid molecule. In certain embodiments, the same adapter is ligated to each end of a nucleic acid molecule, except for the nucleic acid tags. In some embodiments, the adapter is a Y-shaped adapter, one end of which is blunt-ended or tailed as described herein for binding to a nucleic acid molecule that is also blunt-ended or tailed with one or more complementary nucleotides. In yet other exemplary embodiments, the adapter is a bell-shaped adapter that includes a blunt end or a tailed end for binding to the nucleic acid molecule to be analyzed. Other exemplary adapters include T-tailed adapters and C-tailed adapters.

[0051] Administer: As used herein, "administer" or "administering" a therapeutic agent (e.g., an immunological therapeutic agent) to a subject means giving, applying, or contacting the composition with the subject. Administration can be accomplished by any of several routes, including, for example, topical, oral, subcutaneous, intramuscular, intraperitoneal, intravenous, intrathecal, and intradermal.

[0052] Allele: As used herein, "allele" or "allelic variant" refers to a specific genetic variant at a defined genomic location or locus. Allelic variants are usually expressed at a frequency of 50% (0.5) or 100%, depending on whether the allele is heterozygous or homozygous. For example, germline variants are inherited and usually have a frequency of 0.5 or 1. However, somatic variants are acquired variants and usually have a frequency of <0.5. The major and minor alleles of a locus refer to nucleic acids that carry the locus, where the locus is occupied by a nucleotide of a reference sequence and a variant nucleotide that differs from the reference sequence, respectively. Measurements at a locus can take the form of an allelic fraction (AF), which is the frequency at which an allele is observed in a sample.

[0053] Amplify: As used herein, "amplify" or "amplification" in the context of nucleic acids refers to the production of multiple copies of a polynucleotide or portion of a polynucleotide, usually starting from a small amount of the polynucleotide (e.g., a single polynucleotide molecule), and the amplification product or amplicon is usually detectable. Polynucleotide amplification includes a variety of chemical and enzymatic processes.

[0054] Barcode: As used herein, "barcode" in the context of nucleic acids refers to a nucleic acid molecule that contains a sequence that can serve as a molecular identifier. For example, individual "barcode" sequences are typically added to each DNA fragment during next-generation sequencing (NGS) library preparation so that each read can be identified and sorted before final data analysis.

[0055] Cancer type: As used herein, "cancer," "cancer type," or "tumor type" refers to the type or subtype of cancer, as defined, for example, by histopathology. The type of cancer can be determined by any conventional criteria, for example, by appearance in a given tissue (e.g., blood cancer, central nervous system (CNS), brain tumor, lung cancer (small cell and non-small cell), skin cancer, nose cancer, throat cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, bowel cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, breast cancer, prostate cancer, ovarian cancer, lung cancer, intestinal cancer, soft tissue cancer, neuroendocrine cancer, gastroesophageal cancer, head and neck cancer, gynecological cancer, colorectal cancer, etc.). Cancers may be defined based on whether they are of unknown primary origin, such as urothelial carcinoma, solid cancer, heterogeneous cancer, homogeneous cancer, or of unknown primary origin, and / or of the same cellular lineage (e.g., carcinoma, sarcoma, lymphoma, cholangiocarcinoma, leukemia, mesothelioma, melanoma, or glioblastoma), and / or exhibit cancer markers (e.g., Her2, CA15-3, CA19-9, CA-125, CEA, AFP, PSA, HCG, hormone receptors, and NMP-22). Cancers may also be classified by stage (e.g., stage 1, 2, 3, or 4) and by whether they are of primary or secondary origin.

[0056] Cell-free nucleic acids: As used herein, "cell-free nucleic acids" refers to nucleic acids that are not contained within or bound to cells, or in some embodiments, nucleic acids remaining in a sample after removal of intact cells. Cell-free nucleic acids can include, for example, any unencapsulated nucleic acids originating from a subject's body fluids (e.g., blood, plasma, serum, urine, cerebrospinal fluid (CSF), etc.). Cell-free nucleic acids include DNA (cfDNA), RNA (cfRNA), and hybrids thereof, including genomic DNA, mitochondrial DNA, circulating DNA, siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, small RNA (snoRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), and / or fragments of any of these. Cell-free nucleic acids can be double-stranded, single-stranded, or hybrids thereof. Cell-free nucleic acids can be released into body fluids by secretion or cell death processes, such as cell necrosis, apoptosis, etc. Cell-free nucleic acids can be found in efferosomes or exosomes. Some cell-free nucleic acids are released from cancer cells into body fluids (e.g., circulating tumor DNA (ctDNA)). Other cell-free nucleic acids are released from healthy cells. CtDNA can be unencapsulated fragmented DNA derived from tumors. Another example of cell-free nucleic acids is fetal DNA circulating freely in the maternal bloodstream, also known as cell-free fetal DNA (cffDNA). Cell-free nucleic acids can have one or more epigenetic modifications, for example, cell-free nucleic acids can be acetylated, 5-methylated, ubiquitinated, phosphorylated, sumoylated, ribosylated, and / or citrullinated.

[0057] Cellular origin: As used herein, "cellular origin" in the context of cell-free nucleic acids refers to the cell type from which a given cell-free nucleic acid molecule is derived or originates (e.g., via an apoptotic process, a necrotic process, etc.). In certain embodiments, for example, a given cell-free nucleic acid molecule may originate from a tumor cell (e.g., a cancerous lung cell, etc.) or a non-tumor or normal cell (e.g., a non-cancerous lung cell, a hematopoietic stem cell, etc.).

[0058] Classification allele: As used herein, a "classification allele" refers to an allelic variant whose presence in a given nucleic acid molecule identifies the origin (e.g., cellular origin) of the nucleic acid molecule. In certain embodiments, for example, the presence of a given classification allele in a nucleic acid molecule can identify the nucleic acid molecule as originating from a target cell (e.g., diseased cell, tumor cell, fetal cell, transplant donor cell, etc.) or a non-target cell (e.g., non-diseased cell, hematopoietic stem cell, maternal cell, transplant recipient cell, etc.) depending on the specific application. Typically, a given classification allele is associated with a subclonality score that can be used to assign the classification allele to a target nucleic acid variant filter list or a non-target nucleic acid variant filter list, depending on whether the subclonality score is below, above, or above the selected cutoff threshold used in a given application.

[0059] Classifier: As used herein, "classifier" generally refers to algorithmic computer code that takes test data as input and produces as output a classification of the input data as belonging to one class or another (e.g., tumor DNA or non-tumor DNA).

[0060] Clinical information: As used herein, "clinical information" refers to any information that can inform healthcare decisions for a subject. Examples of clinical information include, but are not limited to, genomic information, age, sex, race, weight, body mass index (BMI), medical history, drug use, smoking, and alcohol consumption, among others.

[0061] Clonal hematopoietic-derived mutations: As used herein, "clonal hematopoietic-derived mutations" refers to the somatic acquisition of genomic mutations in hematopoietic stem and / or progenitor cells that lead to clonal expansion.

[0062] Clonal hematopoiesis of undetermined potential: As used herein, "clonal hematopoiesis of undetermined potential" or "CHIP" refers to hematopoiesis in an individual, meaning the expansion of hematopoietic stem cells that contain one or more somatic mutations (e.g., mutations associated with hematopoietic malignancies and / or mutations that are not), but do not have diagnostic criteria for hematopoietic malignancies (e.g., definitive morphological evidence of dysplasia). CHIP is a common age-related phenomenon in which hematopoietic stem cells contribute to the formation of genetically distinct subpopulations of blood cells.

[0063] Clonality threshold: As used herein, "clonality threshold" refers to a selected value used in the calculation of a given subclonality score.

[0064] Control result: As used herein, "control result" or "reference result" refers to a result or set of results that can be compared to a given test sample or test result to identify one or more promising characteristics of that test sample or test result, and / or to identify one or more possible prognostic outcomes and / or one or more customized therapies for the subject from whom the test sample was taken or obtained. Control results are usually obtained from a set of reference samples (e.g., subjects with the same disease or cancer type as the test subject, and / or subjects receiving or having received the same therapy as the test subject).

[0065] Control Sample: As used herein, "control sample" or "control DNA sample" refers to a sample of known composition and / or known characteristics and / or parameters (e.g., known cellular origin, known tumor fraction, known coverage, etc.) that is analyzed along with or in comparison to a test sample to assess the accuracy of an analytical procedure. A data set of control samples typically includes at least about 25 to at least about 30,000 or more control samples. In some embodiments, the dataset of control samples includes about 50, 75, 100, 150, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,500, 5,000, 7,500, 10,000, 15,000, 20,000, 25,000, 50,000, 100,000, 1,000,000 or more control samples.

[0066] Coverage: As used herein, "coverage" refers to the number of nucleic acid molecules occupying a particular base position.

[0067] Cutoff threshold: As used herein, "cutoff threshold" refers to a selected value that is compared with a subclonality score to assign a classified allele with that subclonality score to a target nucleic acid variant filter list or a non-target nucleic acid variant filter list.

[0068] Deoxyribonucleic acid or ribonucleic acid: As used herein, "deoxyribonucleic acid" or "DNA" refers to a natural or modified nucleotide having a hydrogen group at the 2' position of the sugar moiety. DNA typically comprises a chain of nucleotides containing deoxyribonucleosides, each of which contains one of four types of nucleobases: adenine (A), thymine (T), cytosine (C), and guanine (G). As used herein, "ribonucleic acid" or "RNA" refers to a natural or modified nucleotide having a hydroxyl group at the 2' position of the sugar moiety. RNA typically comprises a chain of nucleotides containing ribonucleosides, each of which contains one of four types of nucleobases: A, uracil (U), G, and C. As used herein, the term "nucleotide" refers to a natural or modified nucleotide. Certain pairs of nucleotides bind specifically to each other in a complementary manner (called complementary base pairing). In DNA, adenine (A) pairs with thymine (T) and cytosine (C) pairs with guanine (G). In RNA, adenine (A) pairs with uracil (U) and cytosine (C) pairs with guanine (G). When a first nucleic acid strand binds to a second nucleic acid strand composed of nucleotides complementary to those in the first strand, the two strands combine to form a duplex. As used herein, "nucleic acid sequencing data," "nucleic acid sequencing information," "sequence information," "nucleic acid sequence," "nucleotide sequence," "genomic sequence," "gene sequence," or "fragment sequence" or "nucleic acid sequencing read" refers to any information or data that indicates the order and identity of nucleotide bases (e.g., adenine, guanine, cytosine, and thymine or uracil) in a molecule of nucleic acid such as DNA or RNA (e.g., a whole genome, a whole transcriptome, an exome, an oligonucleotide, a polynucleotide, or a fragment).It should be understood that the present teachings contemplate sequence information obtained using any available technique, platform, or technology, including, but not limited to, capillary electrophoresis, microarrays, ligation-based systems, polymerase-based systems, hybridization-based systems, direct or indirect nucleotide discrimination systems, pyrosequencing, ion- or pH-based detection systems, and electronic signature-based systems.

[0069] Fragment: As used herein, "fragment" in the context of cell-free nucleic acids refers to a nucleic acid molecule that is naturally present in a subject's body (or a sample obtained from a subject) and should not be construed as requiring that the fragmentation step be performed in vitro.

[0070] Hematopoietic stem cell: As used herein, a "hematopoietic stem cell" or "HSC" is a stem cell that gives rise to other blood cells through the process of hematopoiesis.

[0071] Immunotherapy: As used herein, "immunotherapy" refers to treatment with one or more agents that stimulate the immune system to kill cancer cells or at least inhibit their growth, preferably to suppress further growth, reduce the size of cancer, and / or eliminate cancer. Some of these agents bind to targets present on cancer cells; some bind to targets present on immune cells but not on cancer cells; and some bind to targets present on both cancer cells and immune cells. Such agents include, but are not limited to, checkpoint inhibitors and / or antibodies. Checkpoint inhibitors are inhibitors of immune system pathways that maintain self-tolerance and regulate the duration and amplitude of physiological immune responses in peripheral tissues to minimize collateral tissue damage (see, e.g., Pardoll, Nature Reviews Cancer 12, 252-264 (2012)). Exemplary agents include antibodies against PD-1, PD-2, PD-L1, PD-L2, CTLA-40, OX40, B7.1, B7He, LAG3, CD137, KIR, CCR5, CD27, or CD40. Other exemplary agents include pro-inflammatory cytokines, such as IL-1β, IL-6, and TNF-α. Other exemplary agents are T cells activated against tumors, for example, T cells activated by expressing a chimeric antigen that targets a tumor antigen recognized by the T cells.

[0072] Indel: As used herein, "indel" refers to a mutation that involves the insertion or deletion of a nucleotide position in a subject's genome.

[0073] Indexed: As used herein, "indexed" refers to the linking of a first element (e.g., clinical information) to a second element (e.g., a given sample).

[0074] Maximum minor allele frequency: As used herein, "maximum minor allele frequency," "maximum MAF," or "maxMAF" refers to the largest or highest MAF of all somatic variants present or observed in a given sample.

[0075] Minor allele frequency: As used herein, "minor allele frequency" or "MAF" refers to the frequency at which a minor allele (e.g., the least common allele) is present in a given nucleic acid population (e.g., a sample obtained from a subject). In other words, "minor allele frequency" refers to the frequency of alleles observed at a given locus in a given sample that are not the most prevalent allele observed at that locus in that sample. MAF is generally expressed as a proportion or percentage. For example, MAF is usually less than about 0.5, 0.1, 0.05, or 0.01 (i.e., less than about 50%, 10%, 5%, or 1%) of all somatic variants or all alleles present at a given locus.

[0076] Mutation: As used herein, "mutation" or "genetic abnormality" refers to a variation from a known reference sequence, including mutations such as single nucleotide variants (SNVs), copy number variants or copy number variations (CNVs) / copy number abnormalities, insertions or deletions (indels), truncations, gene fusions, transversions, translocations, frameshifts, duplications, repeat multiplications and epigenetic variants.Mutation can be germline mutations or somatic mutations.In some embodiments, the reference sequence for comparison is the wild-type genome sequence of the species of the subject that provides the test sample, usually the human genome.

[0077] Neoplasm: As used herein, the terms "neoplasm" and "tumor" are used interchangeably. They refer to an abnormal growth of cells in a subject. A neoplasm or tumor can be benign, potentially malignant, or malignant. A malignant tumor is referred to as a cancer or cancerous tumor.

[0078] Next-generation sequencing: as used herein, " next-generation sequencing " or " NGS " refers to the sequencing technology that has high throughput compared with the approach based on traditional Sanger method and the approach based on capillary electrophoresis, for example, can generate hundreds of thousands of relatively small sequence reads at once.Some examples of next-generation sequencing method include but are not limited to sequencing by synthesis, sequencing by ligation and sequencing by hybridization.

[0079] Nucleic acid tag: As used herein, "nucleic acid tag" refers to a short nucleic acid (e.g., less than about 500, about 100, about 50, or about 10 nucleotides in length) used to label nucleic acid molecules to distinguish nucleic acids from different samples (e.g., corresponding to a sample index) or to label nucleic acid molecules to distinguish different nucleic acid molecules in the same sample that have undergone different types or different treatments (e.g., corresponding to a molecular tag). Nucleic acid tags can be single-stranded, double-stranded, or at least partially double-stranded. Nucleic acid tags can have the same length or various lengths, as desired. Nucleic acid tags can also include double-stranded molecules with one or more blunt ends, can include 5' or 3' single-stranded regions (e.g., overhangs), and / or can include one or more other single-stranded regions at other locations within a given molecule. Nucleic acid tags can be attached to one or both ends of other nucleic acids (e.g., sample nucleic acids to be amplified and / or sequenced). Decoding nucleic acid tags can reveal information such as the sample origin, the form or treatment of a given nucleic acid. The use of nucleic acid tags can also enable the pooling and / or parallel processing of multiple samples containing nucleic acids with different nucleic acid tags and / or sample indexes, and these nucleic acids are then deconvoluted by reading the nucleic acid tags.Nucleic acid tags can also be referred to as molecular identifiers or molecular tags, sample identifiers, index tags, and / or barcodes.In addition or alternatively, nucleic acid tags can be used to distinguish different molecules in the same sample.This includes, for example, uniquely tagging each different nucleic acid molecule in a given sample, or non-uniquely tagging such molecules.For non-unique tagging applications, each nucleic acid molecule can be tagged with a limited number of tags, so that different molecules can be distinguished, for example, based on the start / end positions located in a selected reference genome along with at least one nucleic acid tag.A sufficient number of different nucleic acid tags are usually used so that the probability that any two molecules have the same start / end position and the same nucleic acid tag is low (for example, less than about 10%, less than about 5%, less than about 1%, or less than about 0.1%).Some nucleic acid tags contain multiple molecular identifiers that label samples, the types of nucleic acid molecules in the samples, and nucleic acid molecules in the types that have the same start and end positions.Such nucleic acid tags can be referred to using the exemplary format "A1i", where the capital letter indicates the sample type, the Arabic numerals indicate the types of molecules in the sample, and the lowercase Roman numerals indicate the molecules in a certain type.

[0080] Polynucleotide: As used herein, "polynucleotide," "nucleic acid," "nucleic acid molecule," or "oligonucleotide" refers to a linear polymer of nucleosides (including deoxyribonucleosides, ribonucleosides, or their analogs) joined by internucleoside linkages. Typically, a polynucleotide contains at least three nucleosides. Oligonucleotides often range in size from a few monomeric units, e.g., 3-4, to several hundred monomeric units. Whenever a polynucleotide is represented by a sequence of letters such as "ATGCCTG," it is understood that the nucleotides are in 5'→3' order from left to right, and that, in the case of DNA, "A" represents deoxyadenosine, "C" represents deoxycytidine, "G" represents deoxyguanosine, and "T" represents deoxythymidine, unless otherwise stated. The letters A, C, G, and T may be used, as is standard in the art, to refer to the base itself, the nucleoside, or the nucleotide that comprises the base.

[0081] Potential clinical significance: As used herein, "potential clinical significance" in the context of an allelic variant refers to the fact that the presence of an allele in a given nucleic acid molecule from a subject may affect healthcare decisions for that subject.

[0082] Reference sequence: As used herein, "reference sequence" or "reference genome" refers to a known sequence used for comparison with experimentally determined sequences. For example, the known sequence can be a whole genome, a chromosome, or any segment thereof. A reference sequence usually comprises at least about 20, at least about 50, at least about 100, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, at least about 500, at least about 1000 or more nucleotides. A reference sequence can be aligned with a single continuous sequence of a genome or chromosome, or can comprise discontinuous segments that align with various regions of a genome or chromosome. Exemplary reference sequences include, for example, the human genome, for example, hG19 and hG38.

[0083] Sample: As used herein, "sample" means anything that can be analyzed by the methods and / or systems disclosed herein.

[0084] Sensitivity: As used herein, "sensitivity" in the context of a given assay or method refers to the ability of the assay or method to detect and distinguish between target analytes (e.g., cfDNA fragments originating from tumor cells) and non-target analytes (e.g., cfDNA fragments originating from non-tumor cells).

[0085] Sequencing: As used herein, "sequencing" refers to any of several techniques used to determine the sequence (e.g., identity and order of monomer units) of biomolecules, for example, nucleic acids such as DNA or RNA.Exemplary sequencing methods include targeted sequencing, single-molecule real-time sequencing, exon or exome sequencing, intron sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxy termination sequencing, whole genome sequencing, hybridization sequencing, pyrosequencing, capillary electrophoresis, gel electrophoresis, double-strand sequencing, cycle sequencing, single-base extension sequencing, solid-phase sequencing, high-throughput sequencing, massively parallel signature sequencing, emulsion PCR, co-amplification at lower denaturation temperature PCR. temperature-PCR (COLD-PCR), multiplex PCR, reversible dye terminator sequencing, paired-end sequencing, near-term sequencing, exonuclease sequencing, ligation sequencing, short-read sequencing, single-molecule sequencing, sequencing-by-synthesis, real-time sequencing, reverse terminator sequencing, nanopore sequencing, 454 sequencing, Solexa Genome Analyzer sequencing, SOLiD TM In some embodiments, sequencing is performed using a genetic analyzer, such as a gene analyzer manufactured by Illumina, Inc., Pacific Biosciences, Inc., or Applied Pharma, Inc., among others. The assay may be performed by commercially available genetic analyzers from Biosystems / Thermo Fisher Scientific.

[0086] Sequence information: As used herein, "sequence information" in the context of a nucleic acid polymer means the order and identity of the monomer units (eg, nucleotides) in that polymer.

[0087] Somatic mutation: As used herein, "somatic mutation" refers to a mutation in the genome that occurs after conception. Somatic mutations can occur in any cell of the body except germ cells and are therefore not passed on to offspring.

[0088] Splice site variant: As used herein, "splice site variant" in the context of nucleic acid mutations refers to a genetic change in a given DNA sequence that occurs at the boundary between an exon and an intron (splice site). This change can interfere with RNA splicing, resulting in the loss of an exon or the inclusion of an intron and a change in the protein-coding sequence.

[0089] Specificity: As used herein, "specificity" in the context of a diagnostic analysis or assay refers to the degree to which the analysis or assay detects the intended target analyte to the exclusion of other components of a given sample.

[0090] Subclonality score: As used herein, "subclonality score" is the ratio of the number of times a given allele is observed to have a MAF / maxMAF ratio value below the clonality cutoff value to (i.e., divided by) the total number of times that the given allele is observed or occurs in a sample set.

[0091] Subject: As used herein, "subject" or "test subject" refers to an animal, e.g., a mammalian species (e.g., a human) or an avian (e.g., a bird) species, or other organisms such as plants. More particularly, the subject can be a vertebrate, e.g., a mammal, e.g., a mouse, a primate, a monkey, or a human. Animals include livestock (e.g., production cattle, dairy cattle, poultry, horses, pigs, etc.), sport animals, and companion animals (e.g., pets or service animals). A subject can be a healthy individual, an individual having or suspected of having a disease or predisposition to a disease, or an individual in need of therapy or suspected of needing therapy. The terms "individual" or "patient" are intended interchangeably with "subject." In some embodiments, the subject is a human who has or is suspected of having cancer. For example, the subject can be an individual who has been diagnosed with cancer, an individual who will receive cancer therapy, and / or an individual who has received at least one cancer therapy. The subject may be in remission from cancer. As another example, the subject may be an individual who has been diagnosed with an autoimmune disease. As another example, the subject may be a female individual who may have been diagnosed with or suspected of having a disease, such as cancer, an autoimmune disease, and who is pregnant or planning to become pregnant.

[0092] Substantial match: As used herein, "substantial match" means that at least one first value or element is at least approximately equal to at least one second value or element. In certain embodiments, for example, the cellular origin of a given allelic variant of a cfDNA sample is determined when there is at least one substantial or approximate match (e.g., sequence alignment and / or other clinical information or characteristics) between the allelic variant and a reference sample or a classification allele.

[0093] Substantially align: As used herein, the phrase "substantially align" in the context of aligning nucleic acid sequences means that a first nucleic acid sequence has at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity with at least one subsequence of a second nucleic acid sequence. In some embodiments, for example, a given sequence read is substantially aligned with a reference sequence when the given sequence read has 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with at least one subsequence or region or the entirety of the reference sequence.

[0094] Threshold: As used herein, "threshold" refers to a discrete, determined value used to characterize or classify experimentally determined values.

[0095] Truncation: As used herein, "truncation" in the context of nucleic acid mutations refers to sequence variations observed in a given DNA sequence that may truncate or shorten the polypeptide (e.g., protein) encoded by that DNA sequence when expressed.

[0096] Tumor fraction: As used herein, " tumor fraction " refers to the estimated proportion of nucleic acid molecules derived from tumor in a given sample.For example, the tumor fraction of a sample can be the maximum minor allele frequency (maxMAF) of that sample or the coverage of that sample, or a measure derived from the length, epigenetic status or other characteristics of the cfDNA fragments in that sample or any other selected feature of that sample.The term " maxMAF " refers to the maximum or highest MAF of all somatic variants present in a given sample.In some embodiments, the tumor fraction of a sample is equal to the maxMAF of that sample.

[0097] Value: As used herein, a "value" generally refers to an entry in a dataset that may characterize the feature to which the value refers, including, but not limited to, a number, a word or phrase, a sign (e.g., + or -), or a degree. Detailed Description

[0098] Introduction

[0099] Provided herein are methods, computer-readable media, and systems for improving the sensitivity and / or specificity of detecting cancer cell DNA or other target nucleic acids in cell-free nucleic acids (cfNA) present in samples obtained from patients. The subject methods, computer-readable media, and systems can be easily applied to tumor cfDNA analysis and other target cfNA (e.g., the techniques described in U.S. Patent No. 9,920,366 (B2), U.S. Patent No. 9,840,743 (B2), and PCT published patent application WO2017 / 181146 (A1) (each of which is incorporated by reference)). In some embodiments, provided herein are methods for identifying alleles that can be used to determine whether the alleles originate from cancer cells or hematopoietic stem cells. Once such informative alleles are identified, in certain exemplary embodiments, the alleles can be used to classify a sample as containing or not containing tumor cell DNA.

[0100] Methods for determining the cellular origin of CFNA and related embodiments

[0101] The present application discloses various methods related to determining whether a cell-free nucleic acid (cfNA) sample contains nucleic acid molecules or fragments originating from a given cell or tissue type. In some exemplary embodiments, these methods are used to determine whether a cfNA sample contains nucleic acid molecules (e.g., cell-free deoxyribonucleic acid (cfDNA) fragments and / or cell-free ribonucleic acid (cfRNA) fragments) originating from diseased cells (e.g., tumor cells), fetal cells, transplant donor cells, etc. Often, these types of nucleic acid molecules account for only a small fraction of the total nucleic acid molecules present in a given cfNA sample, which generally contains a large background of nucleic acid molecules originating from, for example, non-diseased cells, normal or healthy cells (e.g., hematopoietic stem cells or other non-tumor cells), maternal cells, transplant recipient cells, etc. Many existing analytical techniques are not sensitive enough to reliably detect and characterize nucleic acid molecules present in such small numbers in a cfNA sample. Information obtained from the methods disclosed herein is typically used to diagnose whether a subject from whom a cfNA sample was obtained has a given disease, disorder, or condition. In certain embodiments, the methods include administering a therapy to a subject or treating a diagnosed disease, disorder, or condition. The present application also discloses, for example, related methods for generating classifiers, as well as methods for creating databases of subclonality scores useful in classifying the cellular origin of cfNA fragments in a test sample.

[0102] In various embodiments of the subject method, multiple loci are sequenced to detect the allele variants of the loci and the allele frequency at each locus.The DNA can be derived from various cell sources (each producing cell-free DNA), thereby producing a mixture of cell-free DNA from different genomic sources for the same locus.The DNA source can be tumor cells (including several clonally different tumor cell variants present in the same subject) and non-tumor cells (particularly blood cells).In some embodiments, genomic regions are targeted for sequencing (as opposed to whole genome sequencing).To detect multiple alleles at the same locus and provide the allele frequency at that locus, multiple fragments of cfDNA from a sample can be sequenced simultaneously using a high-throughput DNA sequencing device.Clonal hematopoiesis with undefined potential (CHIP) is a common age-related phenomenon in which hematopoietic stem cells contribute to the formation of genetically distinct blood cell subpopulations. These hematopoietic stem cells can generate allelic information in cell-free DNA that can be confused with allelic variants generated in cancerous cells.

[0103] By using a database of allele information from reference subjects, alleles that can be used to classify cell-free DNA samples as containing or not containing tumor cell DNA can be discovered. These databases usually contain cell-free DNA sequence information from any subject suspected of having cancer. Generally, the larger the database, the more useful it is for identifying allele variants that can be used to discover allele variants that indicate the presence or absence of tumor cell DNA in cell-free DNA. In the database, multiple potentially clinically significant loci are sequenced for each patient, and for each sequenced locus, the frequency of each allele at that locus is determined. The minor allele frequency (MAF) for each locus is also determined. Due to genetic heterogeneity in a given cfDNA sample, each MAF can vary greatly depending on the locus. For example, a driver mutation at a locus is likely to have a higher MAF than a passenger mutation acquired in a later clone during tumor progression. For a given patient, the allele with the maximum MAF (maxMAF) among the analyzed allele set is identified, and the value of MAF relative to maxMAF is also identified. The database may also include other clinical information for each patient, and the other clinical information may be correlated with the genetic information of each patient. Examples of such clinical information include tumor detection, patient survival time, patient age, etc.

[0104] The allele information in the database can then be screened for alleles that can be used to classify cfDNA samples as containing or not containing clinically significant tumor DNA.For each potentially clinically significant given allele variant in the test sample, the ratio of minor allele frequency (MAF) to maxMAF is determined.Then, the calculated MAF / maxMAF ratio for the allele of interest is usually generated for many samples in the database.Then, the frequency of each MAF / maxMAF value for a given allele in the database (or a portion of the database) can be determined.For example, a histogram of MAF / maxMAF values ​​can be plotted.Then, a clonality boundary value can be set to calculate a subclonality score.The subclonality score is the ratio of the number of cases in which a given allele in the database has a MAF / maxMAF value that is less than the clonality boundary to the total number of cases in which a given allele is observed in the samples collected in the database.Then, a cutoff threshold can be set to determine whether a given allele indicates the presence of tumor DNA. Alleles with subclonality scores above the threshold can be used to distinguish alleles derived from non-tumor DNA, while alleles with subclonality scores below the threshold can be used to distinguish alleles derived from tumor DNA.

[0105] For example, a 50% clonality threshold can be set, for example, as shown in Figures 1A and 1B. As shown, allele 1 has an estimated high subclonality score based on the clonality threshold set at 50%, and allele 2 has a low subclonality score. Thus, in some embodiments, alleles can fall into one of two categories: those representing non-tumor DNA (negative, e.g., allele 1 in Figure 1A) or those representing tumor DNA (positive, e.g., allele 2 in Figure 1B). In the example provided in Figures 1A and 1B, allele 1 and allele 2 are located in different genes, i.e., are not variant alleles at the same locus. This analysis can be applied to multiple tested alleles in a database, and a set of positive and negative alleles is generated by placing a given allele in either the positive or negative category. If necessary, a more stringent selection threshold can be applied to exclude alleles from either category, so that the excluded alleles are not used to make a classification decision for a given sample. For example, alleles with a subclonality score below 25% can be positive, alleles with a subclonality score above 75% can be negative, and alleles within the excluded range (i.e., 25%-75%) are not used to classify samples as containing or not containing tumor DNA. Using the classified alleles, a list of alleles can be generated for classifying a given test sample obtained from a subject. Such a list is referred to as a "tumor variant filter list" or a "non-tumor variant filter list," depending on the context of the term's use. Figures 1A and 1B also provide an example in which the clonality threshold for the MAF / maxMAF distribution is set at 50% for samples collected in a database.

[0106] A low subclonality score usually indicates that the observed allele represents the presence of tumor DNA. For example, a score of zero may indicate that the MAF / maxMAF exceeds the clonality threshold in every sample collected in the database in which the allele was observed, which may indicate that the allele was the dominant minor allele in each sample in the database.

[0107] In addition to testing for the presence or absence of positive (i.e., tumor origin) and negative (i.e., non-tumor origin) alleles in a sample, other classification criteria may also be used as needed. To call clinically significant mutations, using the information in the positive and negative allele set discovered from the MAF / maxMAF ratio is generally most useful for alleles with low MAF. In some embodiments, for example, if a given variant allele is found to have a MAF greater than 1% and the allele is a negative allele, even if the allele is listed in the negative allele list, the sample will still be classified as containing tumor DNA. In another example, if a variant allele is found to have a MAF greater than 2%, in certain embodiments, the sample will be classified as containing tumor DNA, even if the allele is listed in the negative allele list. In other embodiments, a MAF threshold greater than 1% may be used.

[0108] Another exemplary classification criterion is the type of allelic variant observed. In some embodiments, for example, even if an allelic variant is classified as negative by subclonality score and MAF is lower than a selected value (for example, in some embodiments, less than 2%, in other embodiments, less than 1%), the variant is called to have clinical significance. Allelic variants such as truncation, indel, or splice site variants indicate cancer and are not usually present in hematopoietic stem cells.

[0109] In some embodiments, a cfDNA sample from a patient can be characterized as containing DNA from cancer cells if it meets any one of the following criteria: (1) it has an allelic variant that is a truncation, indel, or splice site variant, (2) it has a subclonality score-positive allele, or (3) it has a subclonality score-negative allele with a MAF of greater than 1%. In some embodiments, a cfDNA sample from a patient can be characterized as containing DNA from cancer cells if it meets any one of the following criteria: (1) it has an allelic variant that is a truncation, indel, or splice site variant, (2) it has a subclonality score-positive allele, or (3) it has a subclonality score-negative allele with a MAF of greater than 2%.

[0110] Because the frequency of CHIP mutations typically increases with patient age, in certain embodiments, patient age and / or other patient data can be utilized in the classification to determine whether a cell-free DNA sample contains tumor DNA. In other exemplary embodiments, subclonal lists (e.g., target nucleic acid variant filter lists or non-target nucleic acid variant filter lists) are generated across various sample subsets, e.g., based on minimum maxMAF, calling maxMAF, etc., based on known driver mutations. In some embodiments, subclonal lists are generated based on a specific indication (e.g., a given cancer type (e.g., lung, colorectal, etc.)). In certain embodiments, a machine learning classifier is trained based on one or more features, including mutant allele frequency, subclonal ratio, gene type, variants associated with hematological malignancies, patient age, observation of other CHIP variants, cancer type, etc.

[0111] To further illustrate aspects of the methods disclosed herein, Figure 2 provides a flow chart that schematically depicts exemplary method steps for detecting nucleic acid molecules originating from target cells (e.g., tumor cells) in a subject, at least in part, using a computer. As shown, method 200 includes, in step 202, receiving, by a computer, test sequence information that includes sequence reads obtained from cell-free nucleic acid (cfNA) fragments from a test sample obtained from the subject. Method 200 also includes, in step 204, identifying the presence of at least one allelic variant in the test sequence information that substantially matches at least one classification allele on the target nucleic acid variant filter list, where the classification allele comprises a subclonality score less than at least one selected cutoff threshold (e.g., about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or another value), thereby indicating that the classification allele is derived from a reference cfNA fragment originating from a target cell, thereby detecting a nucleic acid molecule originating from the target cell in the subject. In some embodiments, for example, step 204 comprises: identifying at least one allelic variant in the test sequence information; mapping the allelic variant to at least one classification allele on the target nucleic acid variant filter list; identifying a subclonality score for the classification allele; and comparing the subclonality score with at least one selected cutoff threshold, wherein a subclonality score below the selected cutoff threshold indicates that the classification allele is derived from a reference cfNA fragment originating from a target cell. Related systems, including computers and computer-readable media, are further described herein.

[0112] 3 provides a flowchart schematically illustrating exemplary method steps for detecting nucleic acid molecules originating from tumor cells in a subject using, at least in part, a computer, according to some embodiments. As shown, method 300 includes, in step 302, receiving, by a computer, test sequence information including sequence reads obtained from cell-free deoxyribonucleic acid (cfDNA) fragments in a test sample obtained from the subject. Method 300 also includes, in step 304, removing (e.g., deleting, hiding, ignoring, etc.) from the test sequence information one or more sequence reads originating from hematopoietic stem cells of the subject (e.g., including at least a portion of the classification alleles) to generate filtered test sequence information. Method 300 further includes, in step 306, identifying by a computer one or more sequence reads present in the filtered test sequence information that substantially align with reference sequence information obtained from one or more reference subjects, which reference sequence information originates from one or more tumor cells within the reference subjects, thereby detecting nucleic acid molecules originating from tumor cells within the subject.

[0113] Figure 4 provides a flowchart that schematically illustrates exemplary method steps for treating a disease in a subject. As shown, method 400 includes, in step 402, receiving test sequence information including sequence reads obtained from cell-free nucleic acid (cfNA) fragments from a test sample obtained from the subject. Method 400 further includes, in step 404, identifying the presence of at least one allelic variant in the test sequence information that substantially matches at least one classification allele on the target nucleic acid variant filter list, where the classification allele comprises a subclonality score below at least one selected cutoff threshold, thereby indicating that the classification allele is derived from a reference cfNA fragment originating from a diseased cell, thereby diagnosing the disease in the subject. Furthermore, method 400 also includes, in step 406, administering one or more therapies to the subject, thereby treating the disease in the subject. Exemplary therapies are further described herein.

[0114] 5 provides a flowchart that schematically depicts exemplary method steps for generating a classifier at least in part using a computer. As shown, method 500 includes, in step 502, computationally generating a subclonality score for each allele in a set of classification alleles from sequence information including sequence reads obtained from cell-free nucleic acid (cfNA) fragments from one or more reference samples, where each classification allele is potentially clinically significant and includes a minor allele observed at a given locus in the reference samples. Method 500 also includes, in step 504, computationally comparing the subclonality scores to at least one selected cutoff threshold (e.g., about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or another value), where classification alleles having a subclonality score above the selected cutoff threshold indicate that the classification allele is derived from a reference cfNA fragment originating from a non-target cell and those classification alleles are added to a non-target nucleic acid variant filter list, and / or classification alleles having a subclonality score below the cutoff threshold indicate that the classification allele is derived from a reference cfNA fragment originating from a target cell and those classification alleles are added to a target nucleic acid variant filter list, thereby generating a classifier.

[0115] 6 provides a flowchart that schematically illustrates exemplary method steps for generating a classifier at least in part using a computer. As shown, method 600 includes, in step 602, a step of computationally identifying a set of classification alleles from sequence information including sequence reads obtained from cell-free nucleic acid (cfNA) fragments from one or more reference samples, where each classification allele is potentially clinically significant and includes a minor allele observed at a given locus in the reference samples. Method 600 also includes, in step 604, a step of computationally determining a minor allele frequency (MAF) value for each classification allele in each reference sample from the sequence information, and in step 606, a step of computationally determining a maximum minor allele frequency (maxMAF) value for each reference sample. Method 600 also includes, in step 608, a step of computationally calculating, for each classification allele observed in a given reference sample, a ratio of the MAF value to the maxMAF value for at least some of the reference samples to generate a ratio value. Method 600 also includes, in step 610, calculating by a computer, for each classification allele, the ratio of the number of times that a given classification allele in at least a portion of the reference sample had a ratio value that is less than at least one selected clonality boundary value (e.g., about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or another value) to the total number of times that the given classification allele appeared in at least a portion of the reference sample, to generate a subclonality score for each classification allele in at least a portion of the reference sample.Further, method 600 includes, in step 612, computationally comparing the subclonality scores to at least one selected cutoff threshold (e.g., about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or another value), where classification alleles with subclonality scores above the selected cutoff threshold indicate that the classification alleles are derived from reference cfNA fragments originating from non-target cells and are added to a non-target nucleic acid variant filter list, and / or classification alleles with subclonality scores below the cutoff threshold indicate that the classification alleles are derived from reference cfNA fragments originating from target cells and are added to a target nucleic acid variant filter list, thereby generating a classifier.

[0116] In some embodiments, the method includes obtaining a cfDNA sample from a subject. Essentially any sample type may be used as needed. In certain embodiments, for example, the cfDNA sample is blood, plasma, serum, sputum, urine, semen, vaginal fluid, stool, synovial fluid, cerebrospinal fluid, saliva, or the like. Additional exemplary sample types that may be used as needed are further described herein. Typically, the subject is a mammalian subject (e.g., a human subject). Essentially any type of nucleic acid (e.g., DNA and / or RNA) may be evaluated according to the methods disclosed herein. Some examples include cell-free nucleic acid (e.g., cfDNA of tumor, fetal, maternal, etc. origin), cellular nucleic acid (including circulating tumor cells (e.g., obtained by lysing intact cells in a sample)), circulating tumor nucleic acid, and the like.

[0117] The methods disclosed herein generally involve obtaining sequence information from nucleic acids in a sample taken from a subject. In certain embodiments, the sequence information is obtained from a targeted segment of nucleic acid. Essentially any number of genomic regions may be targeted as desired. The targeting segment may comprise at least 10, at least 50, at least 100, at least 500, at least 1000, at least 2000, at least 5000, at least 10,000, at least 20,000, or at least 50,000 (e.g., 25, 50, 75, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 15,000, 25,000, 30,000, 35,000, 40,000, 45,000) different and / or overlapping genomic regions.

[0118] In these embodiments, the method also generally includes various sample preparation or library preparation steps to prepare nucleic acids for sequencing. Many different sample preparation methods are well known to those skilled in the art. Essentially any of these techniques can be used or adapted for use in carrying out the methods described herein. For example, in addition to various purification steps to isolate nucleic acids from other components in a given sample, typical steps for preparing nucleic acids for sequencing include tagging nucleic acids with molecular identifiers or barcodes, adding adapters (which may include, for example, barcodes), amplifying nucleic acids one or more times, enriching target segments of nucleic acids (for example, using various target capture strategies, etc.). Exemplary library preparation processes are further described herein. For further details regarding nucleic acid sample / library preparation, see, for example, van Dijk et al., Library preparation methods for next-generation sequencing: Tone down the bias, Experimental Cell Research, 322(1):12-20(2014), and Micic (Ed.), Sample Preparation Techniques for Soil, Plant, and Animal Samples (Springer Protocols Handbooks), 1 st Ed., Humana Press (2016) and Chiu, Next-Generation Sequencing and Sequence Data Analysis, Bentham Science Publishers (2018), each of which is incorporated by reference in its entirety.

[0119] The methods disclosed herein are typically used to diagnose the presence of a disease, disorder, or condition, particularly cancer, in a subject; characterize such a disease, disorder, or condition (e.g., to stage a given cancer, to characterize cancer heterogeneity, etc.); monitor response to treatment; assess the potential risk of developing a given disease, disorder, or condition; and / or evaluate the prognosis of the disease, disorder, or condition. The methods disclosed herein are also used, if desired, to characterize specific forms of cancer. Because cancers are often heterogeneous in both composition and staging, data generated using the methods disclosed herein may enable characterization of specific subtypes of cancer, thereby aiding in diagnosis and treatment selection. This information may also provide clues to the subject or medical practitioner regarding the prognosis of a particular type of cancer, allowing the subject and / or medical practitioner to adapt treatment options as the disease progresses. As some cancers progress, they become invasive and genetically unstable. Other tumors remain benign, inactive, or dormant.

[0120] sample

[0121] The sample can be any biological sample isolated from a subject. Samples can include body tissue, whole blood, platelets, serum, plasma, stool, red blood cells, white blood cells or leukocytes, endothelial cells, tissue biopsy material (e.g., biopsy material from a known or suspected solid tumor), cerebrospinal fluid, synovial fluid, lymphatic fluid, ascites, interstitial fluid or extracellular fluid (e.g., fluid derived from the intercellular spaces), gingival exudate, gingival crevicular fluid, bone marrow, pleural effusion, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, and urine. Samples are preferably body fluids, particularly blood and its fractions, and urine. Such samples contain nucleic acids shed from tumors. These nucleic acids can include DNA and RNA, and can be in double-stranded and single-stranded forms. Sample can be in the form that is isolated from subject, or can be subjected to further processing, such as removing or adding components such as cell, concentrating one component to another, or converting one form of nucleic acid into another form (for example, RNA into DNA, or single-stranded nucleic acid into double-stranded).Therefore, for example, the body fluid sample for analysis is the plasma or serum that contains cell-free nucleic acid, for example, cell-free DNA (cfDNA).

[0122] In some embodiments, the volume of the bodily fluid sample taken from the subject depends on the desired read depth for the region to be sequenced. Exemplary volumes are about 0.4 to 40 ml, about 5 to 20 ml, or about 10 to 20 ml. For example, the volume can be about 0.5 ml, about 1 ml, about 5 ml, about 10 ml, about 20 ml, about 30 ml, about 40 ml, or more. The volume of plasma sampled is typically about 5 ml to about 20 ml.

[0123] Samples can contain varying amounts of nucleic acid. Typically, the amount of nucleic acid in a given sample is equated to multiple genome equivalents. For example, a sample of about 30 ng of DNA contains about 10,000 (10 4 ) haploid human genome equivalent, for cfDNA, approximately 200 billion (2 × 10 11) individual polynucleotide molecules. Similarly, a sample of about 100 ng of DNA may contain about 30,000 haploid human genome equivalents, or in the case of cfDNA, about 600 billion individual molecules.

[0124] In some embodiments, the sample contains nucleic acids from various sources, such as cells and cell-free sources (e.g., blood samples, etc.). Typically, the sample contains nucleic acids having mutations. For example, the sample optionally contains DNA having germline mutations and / or somatic mutations. Typically, the sample contains DNA having mutations associated with cancer (e.g., somatic mutations associated with cancer).

[0125] Exemplary amounts of cell-free nucleic acid in a sample prior to amplification typically range from about 1 femtogram (fg) to about 1 microgram (μg), e.g., from about 1 picogram (pg) to about 200 nanograms (ng), from about 1 ng to about 100 ng, or from about 10 ng to about 1000 ng. In some embodiments, the sample contains up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng of cell-free nucleic acid molecules. Optionally, the amount is at least about 1 fg, at least about 10 fg, at least about 100 fg, at least about 1 pg, at least about 10 pg, at least about 100 pg, at least about 1 ng, at least about 10 ng, at least about 100 ng, at least about 150 ng, or at least about 200 ng of cell-free nucleic acid molecules. In certain embodiments, the amount is up to about 1 fg, about 10 fg, about 100 fg, about 1 pg, about 10 pg, about 100 pg, about 1 ng, about 10 ng, about 100 ng, about 150 ng, or about 200 ng of cell-free nucleic acid molecules. In some embodiments, the method includes obtaining from about 1 fg to about 200 ng of cell-free nucleic acid molecules from a sample.

[0126] Cell-free nucleic acids typically have a size distribution ranging from about 100 to about 500 nucleotides in length, with molecules between about 110 and about 230 nucleotides in length accounting for about 90% of the molecules in a sample, a mode of about 168 nucleotides in length, and a second minor peak in the range of about 240 to about 440 nucleotides in length. In certain embodiments, the cell-free nucleic acids are about 160 to about 180 nucleotides in length, about 320 to about 360 nucleotides in length, or about 440 to about 480 nucleotides in length.

[0127] In some embodiments, cell-free nucleic acids are isolated from bodily fluids by a partitioning process that separates the cell-free nucleic acids found in solution from intact cells and other insoluble components of the bodily fluid. In some of these embodiments, partitioning involves techniques such as centrifugation or filtration. Alternatively, cells in the bodily fluid are lysed, and the cell-free and cellular nucleic acids are processed together. Generally, after the addition of a buffer and a washing step, the cell-free nucleic acids are precipitated, for example, with alcohol. In certain embodiments, an additional cleanup step, such as a silica-based column, is used to remove contaminants or salts. Certain aspects of the exemplary procedure (e.g., yield) are optimized throughout the reaction, for example, by adding nonspecific bulk carrier nucleic acid as needed. After such processing, the sample typically contains various forms of nucleic acids, including double-stranded DNA, single-stranded DNA, and / or single-stranded RNA. If necessary, the single-stranded DNA and / or single-stranded RNA is converted to a double-stranded form, which is then included in subsequent processing and analysis steps.

[0128] Nucleic Acid Tags

[0129] In certain embodiments, tags that provide molecular identifiers or barcodes are incorporated into or otherwise attached to the adapters by, inter alia, chemical synthesis, ligation, or overlap extension PCR. In some embodiments, the assignment of unique or non-unique identifiers or molecular barcodes in the reaction is according to the methods and utilizes the systems described therein, for example, in U.S. Patent Application Nos. 20010053519, 20030152490, 20110160078, and U.S. Patent Nos. 6,582,908, 7,537,898, and 9,598,731 (each of which is incorporated by reference).

[0130] Tags are randomly or non-randomly linked to sample nucleic acids. In some embodiments, tags are introduced with expected identifier (for example, unique barcode and / or non-unique barcode combination) and microwell ratio. For example, identifiers can be loaded so that more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 or 1,000,000,000 identifiers are loaded per genome sample. In some embodiments, the identifiers are loaded such that less than about 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 or 1,000,000,000 identifiers are loaded per genomic sample. In certain embodiments, the average number of identifiers loaded per genomic sample is about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 0 or less than 1,000,000,000, or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000, or 1,000,000,000 identifiers. The identifiers are generally unique and / or non-unique.

[0131] One exemplary format uses about 2 to about 1,000,000 different tags, or about 5 to about 150 different tags, or about 20 to about 50 different tags ligated to both ends of a target nucleic acid molecule. 20-50 x 20-50 tags yields a total of 400-2500 tags. Such a number of tags is typically sufficient to ensure that different molecules with the same start and end points have a high probability (e.g., at least 94%, 99.5%, 99.99%, 99.999%) that they will receive different tag combinations.

[0132] In some embodiments, the identifier is a predetermined, random, or semi-random sequence oligonucleotide. In other embodiments, multiple barcodes may be used such that the multiple barcodes are not necessarily unique to each other within the plurality. In these embodiments, the barcode is generally attached to each molecule (e.g., by ligation or PCR amplification) so that the combination of the barcode and the sequence to which it is attached generates a unique sequence that can be individually tracked. As described herein, detecting a non-uniquely tagged barcode along with sequence data at the beginning (start) and end (end) of the sequence read typically allows a unique identity to be assigned to a particular molecule. The length of each sequence read, i.e., the number of base pairs, can also be used as needed to assign a unique identity to a given molecule. As described herein, fragments derived from a single-stranded nucleic acid that have been assigned a unique identity can subsequently identify fragments derived from the parent strand and / or the complementary strand.

[0133] Nucleic Acid Amplification

[0134] The sample nucleic acid adjacent to the adaptor is usually amplified by PCR and other amplification methods using nucleic acid primers that bind to the primer binding sites in the adaptor adjacent to the DNA molecule to be amplified.In some embodiments, the amplification method includes cycles of extension, denaturation and annealing, which are brought about by thermocycling, or can be isothermal, such as in transcription-mediated amplification.Other exemplary amplification methods that can be used as needed include, among others, ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification and self-sustaining sequence-based replication.

[0135] To introduce molecular tags and / or sample indexes / tags into nucleic acid molecules using conventional nucleic acid amplification methods, one or more amplification cycles are typically applied. Such amplification is typically performed in one or more reaction mixtures. Molecular tags and sample indexes / tags are optionally introduced simultaneously or in any sequential order. In some embodiments, molecular tags and sample indexes / tags are introduced before and / or after a sequence capture step is performed. In some embodiments, only molecular tags are introduced before probe capture, and sample indexes / tags are introduced after the sequence capture step. In certain embodiments, both molecular tags and sample indexes / tags are introduced before a probe-based capture step is performed. In some embodiments, sample indexes / tags are introduced after the sequence capture step. Typically, a sequence capture protocol includes introducing a targeted nucleic acid sequence, for example, a coding sequence in a genomic region and a single-stranded nucleic acid molecule complementary to a mutation in such region associated with a cancer type. Typically, the amplification reaction generates a plurality of non-uniquely or uniquely tagged nucleic acid amplicons having molecular tags and sample indexes / tags ranging in size from about 200 nucleotides (nt) to about 700 nt, 250 nt to about 350 nt, or about 320 nt to about 550 nt. In some embodiments, the amplicons have a size of about 300 nt. In some embodiments, the amplicons have a size of about 500 nt.

[0136] Nucleic acid enrichment

[0137] In some embodiments, sequences are enriched before sequencing nucleic acids. Enrichment can be performed for specific target regions or non-specifically ("target sequences"), as needed. In some embodiments, targeted regions of interest can be enriched using nucleic acid capture probes ("baits") selected for one or more bait set panels using differential tiling and capture schemes. Differential tiling and capture schemes generally use different relative concentrations of bait sets to differentially tile (e.g., at different "resolutions") across genomic sections associated with those baits, subject to a set of constraints (e.g., sequencing load, availability of each bait, etc.) to capture targeted nucleic acids at a desired level for downstream sequencing. These targeted genomic sections of interest optionally include natural or synthetic nucleotide sequences of nucleic acid constructs. In some embodiments, biotin-labeled beads bearing probes for one or more sections of interest can be used to capture target sequences, and then optionally amplify the sections to enrich for the regions of interest.

[0138] Sequence capture typically requires the use of oligonucleotide probes that hybridize to target nucleic acid sequences. In certain embodiments, a probe set strategy involves tiling probes across a section of interest. Such probes can be, for example, about 60 to about 120 nucleotides in length. The set can have a depth of about 2x, 3x, 4x, 5x, 6x, 8x, 9x, 10x, 15x, 20x, 50x, or more. The effectiveness of sequence capture generally depends in part on the length of the sequence in the target molecule that is complementary (or nearly complementary) to the sequence of the probe.

[0139] Nucleic Acid Sequencing

[0140] The sample nucleic acid, optionally flanked by adapters, is typically subjected to sequencing, with or without pre-amplification. Optionally, the sequencing method or commercially available format may include, for example, Sanger sequencing, high-throughput sequencing, bisulfite sequencing, pyrosequencing, sequencing by synthesis, single-molecule sequencing, nanopore-based sequencing, semiconductor sequencing, sequencing by ligation, hybridization sequencing, RNA-Seq (Illumina), Digital Gene Expression (Helicos), next-generation sequencing (NGS), Single Molecule Sequencing by Synthesis (SMSS) (Helicos), massively parallel sequencing, Clonal Single Molecule Array (Solexa), shotgun sequencing, Ion Torrent, Oxford Nanopore, Roche Genia, Maxam-Gilbert sequencing, primer walking, sequencing using PacBio, SOLiD, Ion Torrent, or nanopore platforms. Sequencing reactions can be performed in a variety of sample processing devices that may include multiple lanes, multiple channels, multiple wells, or other means for processing multiple sample sets substantially simultaneously. Sample processing devices can also include multiple sample chambers, allowing for the processing of multiple runs simultaneously.

[0141] Sequencing reactions can be performed on one or more nucleic acid fragment types or sections known to contain cancer or other disease markers. Sequencing reactions can also be performed on any nucleic acid fragment present in a sample. Sequencing reactions can provide a genome sequence coverage of at least about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9% or 100% of the genome. In other cases, the genome sequence coverage can be less than about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9% or 100% of the genome.

[0142] Simultaneous sequencing reaction can be carried out using multiplex sequencing method.In some embodiments, cell-free polynucleotide is sequenced by at least about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000 or 100,000 sequencing reaction.In other embodiments, cell-free polynucleotide is sequenced by less than about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000 or 100,000 sequencing reaction.Sequencing reaction is usually carried out consecutively or simultaneously.Subsequent data analysis is generally carried out on all or part of sequencing reaction. In some embodiments, data analysis is performed on at least about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. In other embodiments, data analysis may be performed on less than about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. An exemplary read depth is about 1000 to about 50,000 reads per locus (base position).

[0143] In some embodiments, a nucleic acid population for sequencing is prepared by enzymatically forming blunt ends on double-stranded nucleic acids with single-stranded overhangs at one or both ends. In these embodiments, the population is typically treated with an enzyme having 5'-3' DNA polymerase activity and 3'-5' exonuclease activity in the presence of nucleotides (e.g., A, C, G, and T or U). Exemplary enzymes or their catalytic fragments, if necessary, include Klenow Large Fragment and T4 polymerase. For 5' overhangs, the enzyme typically extends the recessed 3' end of the opposite strand until it aligns with the 5' end, generating a blunt end. For 3' overhangs, the enzyme typically digests from its 3' end to the 5' end of the opposite strand, sometimes beyond the 5' end of the opposite strand. If this digestion proceeds beyond the 5' end of the opposite strand, the gap can be filled by an enzyme with the same polymerase activity used for the 5' overhang. The creation of blunt ends to double-stranded nucleic acids facilitates, for example, the attachment of adapters and subsequent amplification.

[0144] In some embodiments, the population of nucleic acids is subjected to further processing (e.g., conversion of single-stranded nucleic acids to double-stranded nucleic acids and / or conversion of RNA to DNA). These forms of nucleic acids are also optionally ligated to adapters and amplified.

[0145] The nucleic acid that has been subjected to the above-described blunt-end forming process, with or without prior amplification, and optionally other nucleic acids in sample, can be sequenced to produce sequenced nucleic acid.Sequenced nucleic acid can refer to either the sequence of nucleic acid (i.e., sequence information) or the nucleic acid whose sequence has been determined.Sequencing can be carried out to provide the sequence data of each nucleic acid molecule in sample directly or indirectly from the consensus sequence of the amplification product of each nucleic acid molecule in sample.

[0146] In some embodiments, after blunt-end formation, double-stranded nucleic acids in the sample with single-stranded overhangs are ligated at both ends to adapters containing barcodes, and the nucleic acid sequence and the in-line barcode introduced by the adapter are determined by sequencing. These blunt-ended DNA molecules are optionally ligated to the blunt end of an at least partially double-stranded adapter (e.g., a Y-shaped or bell-shaped adapter). Alternatively, the blunt ends of the sample nucleic acid and the adapter can be tailed with complementary nucleotides to facilitate ligation (e.g., sticky-end ligation).

[0147] A nucleic acid sample is typically contacted with a sufficient number of adapters so that the probability that any two copies of the same nucleic acid will receive the same combination of adapter barcodes from the adapters ligated at both ends is low (e.g., <1 or 0.1%). The use of adapters in this manner allows for the identification of families of nucleic acid sequences that have the same start and end points on a reference nucleic acid and are ligated to the same combination of barcodes. Such families correspond to the sequences of the amplification products of nucleic acids in a sample before amplification. The sequences of these family members can be compiled to derive consensus nucleotides or complete consensus sequences for nucleic acid molecules in the original sample that have been modified by blunt-end formation and adapter attachment. In other words, a nucleotide occupying a particular position in a nucleic acid in the sample is determined to be the consensus of the nucleotides occupying its corresponding position in the family member sequences. A family can include sequences of one or both strands of a double-stranded nucleic acid. When a family member comprises the sequence of both strands of double-stranded nucleic acid, the sequence of one strand is converted to its complementary strand in order to compile all sequences to derive consensus nucleotide or consensus sequence.Some families only comprise one member sequence.In this case, this sequence can be considered as the sequence of the nucleic acid in the sample before amplification.Alternatively, families that only comprise one member sequence can be excluded from subsequent analysis.

[0148] Nucleotide variations in sequenced nucleic acids can be revealed by comparing the sequenced nucleic acids with a reference sequence.The reference sequence is often a known sequence, for example, a known whole genome sequence or partial genome sequence of a subject (for example, a whole genome sequence of a human subject).The reference sequence can be, for example, hG19 or hG38.The sequenced nucleic acid can correspond to the consensus of the sequence determined directly for the nucleic acid in the sample or the sequence of the amplification product of such nucleic acid, as described above.Comparison can be performed at one or more designated positions on the reference sequence.A subset of sequenced nucleic acids that includes a position corresponding to the designated position of the reference sequence can be identified when the respective sequences are maximally aligned. Within such a subset, it can determine which sequenced nucleic acids (if any) contain nucleotide variations at a specified position, the length of a given cfDNA fragment based on where its endpoints (i.e., the 5'- and 3'-terminal nucleotides) map to the reference sequence, the offset of the midpoint of the cfDNA fragment from the midpoint of the genomic region in the given cfDNA fragment, and, if necessary, which nucleic acids (if any) contain the reference nucleotide (i.e., the same nucleotide as in the reference sequence).If the number of sequenced nucleic acids in the subset containing a nucleotide variant exceeds a selected threshold, a variant nucleotide can be called at the specified position.The threshold can be a simple number (e.g., at least 1, 2, 3, 4, 5, 6, 7, 9, or 10 sequenced nucleic acids in the subset containing the nucleotide variant), or a ratio (e.g., at least 0.5, 1, 2, 3, 4, 5, 10, 15, or 20, among others) relative to the sequenced nucleic acids in the subset containing the nucleotide variant. Comparison can be repeated for any desired designated position in the reference sequence. Sometimes, comparison can be made for designated positions occupying at least about 20, 100, 200, or 300 contiguous positions on the reference sequence, for example, about 20-500 or about 50-300 contiguous positions.

[0149] Further details regarding nucleic acid sequencing, including the formats and uses described herein, can be found in, for example, Levy et al., Annual Review of Genomics and Human Genetics, 17:95-115 (2016); Liu et al., J. of Biomedicine and Biotechnology, Volume 2012, Article ID 251364:1-11 (2012); Voelkerding et al., Clinical Chem., 55:641-658 (2009); MacLean et al., Nature Rev. Microbiol., 7:287-296 (2009); Astier et al., J Am Chem. Soc.,128(5):1705-10(2006), U.S. Patent No. 6,210,891, U.S. Patent No. 6,258,568, U.S. Patent No. 6,833,246, U.S. Patent No. 7,115,400, U.S. Patent No. 6,969,488, U.S. Patent No. 5,912,148, U.S. Patent No. 6,130,073, U.S. Patent No. 7,169,560, U.S. Patent No. 7,282,337, U.S. Patent No. 7,482,122 Nos. 7,170,050, 7,302,146, 7,313,308, and 7,476,503 (each of which is incorporated by reference in its entirety).

[0150] Data analysis

[0151] In some embodiments, raw sequencing data can include sequence read sets, which can be provided in various file formats (e.g., FASTQ, VCF, CRAM or BAM).The file containing raw sequencing data can include sequence data for one strand or both strands, as in paired-end reads.In one example, raw sequencing data is provided as a FASTQ file for both strands, i.e., the sense strand and antisense strand generated from paired-end sequencing procedure.These files can include additional codes that provide information about the quality of the read, and can also provide quality scores.The raw sequencing data of each polynucleotide molecule can be stored on a local drive, cloud or server.

[0152] In some cases, the sequence reads generated by sequencing reaction can be aligned with or mapped to reference sequence for bioinformatics analysis.Reference sequence is often a known sequence, for example, a known whole genome sequence or partial genome sequence of a subject, or a whole genome sequence of a human subject.Reference sequence can be hG19.As described above, the sequenced nucleic acid can correspond to the sequence directly determined for the nucleic acid in sample or the consensus sequence of the amplification product of such nucleic acid.Comparison can be performed at one or more designated positions on the reference sequence.

[0153] Sequence reads can be aligned with a reference sequence using a mapping tool, non-limiting examples of which include Burrow's Wheeler Transform (BWA), Novoalign, and Bowtie. These mapping tools generate an alignment file describing the alignment parameters used, the position (e.g., coordinates) of the sequence read relative to the reference sequence, and the mapping quality score. These alignment parameters (e.g., the number of differences allowed between the sequencing read and the reference sequence, the number of gaps allowed, and gap opening penalties, the number of gap extensions, etc.) can be defined by the user. In some cases, reads are aligned with a human reference genome such as hg19 using the BWA mapping tool with default alignment parameters. The BWA tool provides an output file, a BAM file, containing alignment statistics. The alignment statistics may include the coordinates of the reference sequence to which the processed reads are aligned. The alignment statistics may also provide a MapQ score, which indicates the uniqueness of the read when mapped to the reference sequence. The processed reads can then be sorted using the molecular barcodes and coordinates on a reference sequence.

[0154] A subset of sequenced nucleic acids containing positions corresponding to designated positions in the reference sequence can be identified when the respective sequences are maximally aligned. Within such a subset, it can be determined which sequenced nucleic acids contain nucleotide variations (if any) at the designated positions, and optionally which contain the reference nucleotide (i.e., the same nucleotide as in the reference sequence). Comparison can be repeated for any desired designated position in the reference sequence. Sometimes, comparison can be performed for designated positions occupying at least 20, 100, 200, or 300 consecutive positions on the reference sequence, for example, 20 to 500 or 50 to 300 consecutive positions.

[0155] A sample can be contacted with a sufficient number of different molecular barcodes so that the probability that any two copies of the same nucleic acid will receive the same combination of adapters containing molecular barcodes from the adapters ligated at one or both ends is low (e.g., <1 or 0.1%). The use of adapters in this manner can enable grouping of sequence reads with the same start and end points aligned (or mapped) to a reference nucleic acid and ligated to the same combination of molecular barcodes into families of reads generated from the same original molecule. Such families can correspond to the sequences of amplification products of nucleic acids in the sample before amplification.

[0156] The sequences of family members can be compiled to derive a consensus nucleotide or a complete consensus sequence for the nucleic acid molecules in the original sample, modified by blunting and adapter attachment. In other words, a nucleotide occupying a particular position in a nucleic acid in the sample can be determined to be the consensus of the nucleotide occupying its corresponding position in the family member sequences. The consensus nucleotide can be determined by methods such as voting or confidence scores, to name two non-limiting exemplary methods. A family can include sequences of one or both strands of a double-stranded nucleic acid. When a family member includes sequences of both strands of a double-stranded nucleic acid, the sequence of one strand is converted to its complementary strand in order to compile all sequences to derive a consensus nucleotide or consensus sequence. Some families may contain only one member sequence. In this case, this sequence can be considered the sequence of the nucleic acid in the sample before amplification. Alternatively, a family containing only one member sequence can be excluded from subsequent analysis.

[0157] In some embodiments, the results of the systems and methods disclosed herein are used as input to generate a report. The report may be in paper format. For example, the report may indicate the presence or absence of a therapeutic nucleic acid construct in a biological sample. In some embodiments, the report may include an indication of the level of the therapeutic nucleic acid construct in the biological sample.

[0158] The various steps of the methods disclosed herein, or steps performed by the systems disclosed herein, may be performed at the same time or at different times, in the same or different geographic locations, e.g., the same country or different countries, and / or by the same or different people.

[0159] Sequencing Panel

[0160] To improve the possibility of detecting tumors that show mutations, the region of DNA that is sequenced can include a panel of genes or genomic regions.By selecting a limited region (e.g., a limited panel) for sequencing, the total amount of sequencing required (e.g., the total amount of nucleotides that are sequenced) can be reduced.By targeting multiple different genes or regions, sequencing panels can detect a single cancer, a series of cancers, or all cancers.Alternatively, DNA can be sequenced by whole genome sequencing (WGS) or other unbiased sequencing methods that do not use sequencing panels.

[0161] In some embodiments, a panel that targets multiple different genes or genomic regions is selected so that a predetermined ratio of subjects with cancer shows genetic variants or tumor markers in one or more different genes in the panel.The panel can be selected to limit the region for sequencing to a fixed number of base pairs.The panel can be selected to sequence a desired amount of DNA.The panel can further be selected to achieve a desired sequence read depth.The panel can be selected to achieve a desired sequence read depth or sequence read coverage for a certain amount of sequenced base pairs.The panel can be selected to achieve a theoretical sensitivity, theoretical specificity and / or theoretical accuracy for detecting one or more genetic variants in a sample.

[0162] The probe for detecting the panel of regions can include the probe for detecting the genomic region of interest (hot spot region) and the nucleosome recognition probe (for example, KRAS codon 12 and 13), and these probes can be designed to optimize capture based on the analysis of cfDNA coverage and fragment size variation that is affected by nucleosome binding pattern and GC sequence composition.Region as used herein can also include non-hot spot regions that are optimized based on nucleosome position and GC model. The panel may include multiple subpanels, including a subpanel for identifying tissue of origin (e.g., using previously published literature to define 50-100 baits representing genes (not necessarily promoters) with the most diverse transcriptional profiles across tissues), a subpanel for identifying whole genome scaffolds (e.g., for sparse tiling with just a handful of probes across chromosomes to identify ultraconserved genomic content and for the purpose of lining copy number bases), and a subpanel for identifying transcription start sites (TSSs) / CpG islands (e.g., to capture differentially methylated regions (e.g., differentially methylated regions (DMRs)) in promoters of tumor suppressor genes (e.g., SEPT9 / VIM in colorectal cancer). In some embodiments, the markers for tissue of origin are tissue-specific epigenetic markers.

[0163] Some example lists of genomic locations of interest can be found in Table 1 and Table 2. In some embodiments, the genomic locations used in the methods of the disclosure include at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, or 97 genes from Table 1. In some embodiments, the genomic locations used in the methods of the disclosure include at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 SNVs from Table 1. In some embodiments, a genomic location used in the methods of the disclosure comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 CNVs from Table 1. In some embodiments, a genomic location used in the methods of the disclosure comprises at least 1, at least 2, at least 3, at least 4, at least 5, or 6 fusions from Table 1. In some embodiments, a genomic location used in the methods of the disclosure comprises at least a portion of at least 1, at least 2, or 3 indels from Table 1. In some embodiments, the genomic locations used in the methods of the disclosure comprise at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 105, at least 110, or 115 of the genes in Table 2.In some embodiments, a genomic location used in the methods of the disclosure comprises at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 SNVs from Table 2. In some embodiments, a genomic location used in the methods of the disclosure comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 CNVs from Table 2. In some embodiments, a genomic location used in the methods of the disclosure comprises at least 1, at least 2, at least 3, at least 4, at least 5, or 6 fusions from Table 2. In some embodiments, the genomic locations used in the methods of the disclosure comprise at least a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 indels from Table 2. Each of these genomic locations of interest can be identified as a scaffold region or a hotspot region for a given bait set panel. An example list of hotspot genomic locations of interest can be found in Table 3. In some embodiments, the genomic locations used in the methods of the disclosure comprise at least a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 genes from Table 3.Each hotspot genomic location is listed along with several characteristics, including the associated gene, the chromosome on which it resides, the genomic start and end positions corresponding to the gene's locus, the length of the gene's locus in base pairs, the exons covered by the gene, and the important features that a given genomic location of interest may seek to capture (e.g., type of mutation). [Table 1-1] [Table 1-2] [Table 2] [Table 3-1] [Table 3-2] [Table 3-3]

[0164] In some embodiments, one or more regions in the panel include one or more loci of one or more genes for detecting residual cancer after surgery. This detection can be earlier than existing cancer detection methods. In some embodiments, one or more genomic locations in the panel include one or more loci of one or more genes for detecting cancer in high-risk patient populations. For example, smokers have a significantly higher rate of lung cancer than the general population. In addition, smokers may develop other lung symptoms (e.g., the development of irregular nodules in the lungs) that make cancer detection difficult. In some embodiments, the methods described herein detect cancer in high-risk patients earlier than existing cancer detection methods.

[0165] Based on the number of subjects with cancer that have tumor markers in a gene or region, a genomic location can be selected for inclusion in sequencing panel.Based on the prevalence of subjects with cancer and the tumor markers present in that gene, a genomic location can be selected for inclusion in sequencing panel.The presence of tumor markers in a region can indicate that a subject has cancer.

[0166] In some cases, the panel can be selected using information from one or more databases. Information about cancer can be obtained from cancer tumor biopsies or cfDNA assays. The database can include information describing the population of sequenced tumor samples. The database can include information about mRNA expression in tumor samples. The database can include information about regulatory elements or genomic regions in tumor samples. The information related to sequenced tumor samples can include the frequency of various genetic variants and can describe the genes or regions in which these genetic variants exist. These genetic variants can be tumor markers. A non-limiting example of such a database is COSMIC. COSMIC is a catalog of somatic mutations found in various cancers. For a specific cancer, COSMIC ranks genes based on the frequency of mutations. A gene can be selected for inclusion in a panel by having a high frequency of mutations in a given gene. For example, COSMIC suggests that 33% of sequenced breast cancer samples have mutations in TP53, and 22% of sampled breast cancer samples have mutations in KRAS. Other ranked genes, including APC, have mutations that are found in only about 4% of sequenced breast cancer samples. TP53 and KRAS may be included in a sequencing panel based on their relatively high frequency in sampled breast cancers (e.g., compared to APC, which is present at a frequency of about 4%). COSMIC is provided as a non-limiting example; however, any database or information set that associates cancer with tumor markers located in a gene or gene region may be used. In another example, COSMIC found that of 1,156 cholangiocarcinoma samples, 380 samples (33%) had mutations in TP53. Several other genes (e.g., APC) have mutations in 4-8% of all samples. Therefore, TP53 may be selected for inclusion in the panel based on its relatively high frequency in a population of cholangiocarcinoma samples.

[0167] A gene or genomic section can be selected for a panel if the frequency of the tumor marker in sampled tumor tissue or circulating tumor DNA is significantly higher than the frequency found in a given background population. A combination of genomic locations can be selected for inclusion in a panel so that at least a majority of subjects with cancer have a tumor marker or genomic region present in at least one of the genomic locations or genes in the panel. The combination of genomic locations can be selected based on data showing that for a particular cancer or set of cancers, a majority of subjects have one or more tumor markers in one or more of the selected regions. For example, to detect cancer 1, a panel including regions A, B, C, and / or D can be selected based on data showing that 90% of subjects with cancer 1 have tumor markers in regions A, B, C, and / or D of the panel. Alternatively, tumor markers can be shown to be independently present in two or more regions in subjects with cancer, and the tumor markers in those two or more regions, when combined, are present in a majority of the population of subjects with cancer. For example, to detect cancer 2, a panel including regions X, Y, and Z may be selected based on data showing that 90% of subjects have tumor markers in one or more regions, and that in 30% of such subjects, the tumor marker is detected only in region X, while in the remaining subjects in which the tumor marker is detected, the tumor marker is detected only in regions Y and / or Z. Tumor markers present at one or more genomic locations previously shown to be associated with one or more cancers may indicate or predict that a subject has cancer if a tumor marker is detected in 50% or more of those regions. Computational approaches (e.g., models using the conditional probability of detecting cancer given the cancer frequencies for a set of tumor markers in one or more regions) can be used to predict which regions, alone or in combination, may predict cancer.Other approaches to panel selection include the use of databases that describe information from studies using comprehensive genomic profiling of tumors using large panels and / or whole-genome sequencing (WGS, RNA-seq, Chip-seq, bisulfate sequencing, ATAC-seq, etc.). Information gleaned from the literature can also describe pathways that are commonly affected and mutated in a particular cancer. Panel selection can be further characterized by using ontologies that describe genetic information.

[0168] The genes included in the panel for sequencing can include the fully transcribed region, promoter region, enhancer region, regulatory element and / or downstream sequence.To further increase the possibility of detecting tumors that show mutations, only exons can be included in the panel.The panel can include all exons of a selected gene, or only one or more exons of the exons of a selected gene.The panel can include exons of multiple different genes.The panel can include at least one exon of multiple different genes.

[0169] In some embodiments, the panel of exons for each of a plurality of different genes is selected such that a predetermined proportion of subjects with cancer exhibit a genetic variant in at least one exon within the panel of exons.

[0170] At least one full-length exon of each different gene in the panel of genes can be sequenced. The sequenced panel can include exons from multiple genes. The panel can include exons from 2 to 100 different genes, 2 to 70 genes, 2 to 50 genes, 2 to 30 genes, 2 to 15 genes, or 2 to 10 genes.

[0171] The selected panel may include a varying number of exons. The panel may include 2-3,000 exons. The panel may include 2-1,000 exons. The panel may include 2-500 exons. The panel may include 2-100 exons. The panel may include 2-50 exons. The panel may include 2-50 exons. The panel may include 300 or fewer exons. The panel may include 200 or fewer exons. The panel may include 100 or fewer exons. The panel may include 50 or fewer exons. The panel may include 40 or fewer exons. The panel may include 30 or fewer exons. The panel may include 25 or fewer exons. The panel may include 20 or fewer exons. The panel may include 15 or fewer exons. The panel may include 10 or fewer exons. The panel may include 9 or fewer exons. The panel may include 8 or fewer exons. The panel may include up to seven exons.

[0172] The panel may include one or more exons of a plurality of different genes. The panel may include one or more exons of each of a proportion of the plurality of different genes. The panel may include at least two exons of at least 25%, 50%, 75%, or 90% of the different genes. The panel may include at least three exons of at least 25%, 50%, 75%, or 90% of the different genes. The panel may include at least four exons of at least 25%, 50%, 75%, or 90% of the different genes.

[0173] The size of a sequencing panel can vary. A sequencing panel can be large or small (in terms of nucleotide size), depending on several factors (e.g., the total amount of nucleotides sequenced or the number of unique molecules sequenced for a particular region within the panel). A sequencing panel can be 5 kb to 50 kb in size. A sequencing panel can be 10 kb to 30 kb in size. A sequencing panel can be 12 kb to 20 kb in size. A sequencing panel can be 12 kb to 60 kb in size. Sequencing panels can be at least 10kb, 12kb, 15kb, 20kb, 25kb, 30kb, 35kb, 40kb, 45kb, 50kb, 60kb, 70kb, 80kb, 90kb, 100kb, 110kb, 120kb, 130kb, 140kb, or 150kb in size. Sequencing panels can be less than 100kb, 90kb, 80kb, 70kb, 60kb, or 50kb in size.

[0174] The panel selected for sequencing can include at least 1, 5, 10, 15, 20, 25, 30, 40, 50, 60, 80 or 100 genomic locations (e.g., each containing a genomic region of interest).In some cases, the genomic locations in the panel are selected so that the size of these locations is relatively small.In some cases, the region in the panel has a size of about 10kb or less, about 8kb or less, about 6kb or less, about 5kb or less, about 4kb or less, about 3kb or less, about 2.5kb or less, about 2kb or less, about 1.5kb or less, or about 1kb or less. In some cases, genomic locations within a panel have a size of about 0.5 kb to about 10 kb, about 0.5 kb to about 6 kb, about 1 kb to about 11 kb, about 1 kb to about 15 kb, about 1 kb to about 20 kb, about 0.1 kb to about 10 kb, or about 0.2 kb to about 1 kb. For example, a region within a panel can have a size of about 0.1 kb to about 5 kb.

[0175] The panel selected herein may enable deep sequencing sufficient to detect low-frequency genetic variants (e.g., in cell-free nucleic acid molecules obtained from a sample). The amount of genetic variants in a sample may be referred to in terms of the minor allele frequency for a given genetic variant. Minor allele frequency may refer to the frequency at which a minor allele (e.g., the least common allele) is present in a given nucleic acid population (e.g., a sample). A genetic variant with a low minor allele frequency may have a relatively low occurrence frequency in a sample. In some cases, the panel may enable detection of genetic variants with a minor allele frequency of at least 0.0001%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, or 0.5%. The panel may enable detection of genetic variants with a minor allele frequency of 0.001% or more. The panel may enable detection of genetic variants with a minor allele frequency of 0.01% or more. The panel may allow for the detection of genetic variants present in a sample at frequencies as low as 0.0001%, 0.001%, 0.005%, 0.01%, 0.025%, 0.05%, 0.075%, 0.1%, 0.25%, 0.5%, 0.75%, or 1.0%. The panel may allow for the detection of tumor markers present in a sample at frequencies of at least 0.0001%, 0.001%, 0.005%, 0.01%, 0.025%, 0.05%, 0.075%, 0.1%, 0.25%, 0.5%, 0.75%, or 1.0%. The panel may allow for the detection of tumor markers present in a sample at frequencies as low as 1.0%. The panel may allow for the detection of tumor markers present in a sample at frequencies as low as 0.75%. The panel may allow for the detection of tumor markers present in a sample at frequencies as low as 0.5%. The panel may allow for the detection of tumor markers present in a sample at frequencies as low as 0.25%. The panel may allow for the detection of tumor markers present in a sample at frequencies as low as 0.1%. The panel may allow for the detection of tumor markers present in a sample at frequencies as low as 0.075%.The panel may enable detection of tumor markers present in a sample at frequencies as low as 0.05%. The panel may enable detection of tumor markers present in a sample at frequencies as low as 0.025%. The panel may enable detection of tumor markers present in a sample at frequencies as low as 0.01%. The panel may enable detection of tumor markers present in a sample at frequencies as low as 0.005%. The panel may enable detection of tumor markers present in a sample at frequencies as low as 0.001%. The panel may enable detection of tumor markers present in a sample at frequencies as low as 0.0001%. The panel may enable detection of tumor markers in sequenced cfDNA present in a sample at frequencies between 1.0% and 0.0001%. The panel may enable detection of tumor markers in sequenced cfDNA present in a sample at frequencies between 0.01% and 0.0001%.

[0176] Genetic variants can be presented as a percentage of the population of subjects with disease (for example, cancer).In some cases, at least 1%, 2%, 3%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95% or 99% of the population with cancer show one or more genetic variants in at least one of the regions in the panel.For example, at least 80% of the population with cancer can show one or more genetic variants in at least one of the genome positions in the panel.

[0177] The panel may include one or more locations containing genomic regions of interest from one or more individual genes. In some cases, the panel may include one or more locations containing genomic regions of interest from at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, or 80 individual genes. In some cases, the panel may include one or more locations containing genomic regions of interest from at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, or 80 individual genes. In some cases, the panel may include one or more locations containing genomic regions of interest from about 1 to about 80, 1 to about 50, about 3 to about 40, 5 to about 30, or 10 to about 20 different genes.

[0178] The regions in the panel can be selected so that they comprise the sequence that is differentially transcribed across one or more tissues.In some cases, the location that comprises genomic region can comprise the sequence that is transcribed in certain tissues at a higher level than in other tissues.For example, the location that comprises genomic region can comprise the sequence that is transcribed in certain tissues but not in other tissues.

[0179] The genome location in the panel may include coding and / or non-coding sequences. For example, the genome location in the panel may include one or more sequences in exons, introns, promoters, 3' untranslated regions, 5' untranslated regions, regulatory elements, transcription start sites and / or splice sites. In some cases, the region in the panel may include other non-coding sequences, including pseudogenes, repeat sequences, transposons, viral elements and telomeres. In some cases, the genome location in the panel may include sequences in non-coding RNAs, such as ribosomal RNAs, transfer RNAs, Piwi-interacting RNAs and microRNAs.

[0180] The genomic locations within the panel can be selected to detect (diagnose) cancer with a desired level of sensitivity (e.g., by detecting one or more genetic variants). For example, regions within the panel can be selected to detect cancer (e.g., by detecting one or more genetic variants) with at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% sensitivity. Genomic locations within the panel can be selected to detect cancer with 100% sensitivity.

[0181] The genomic locations within the panel can be selected to detect (diagnose) cancer with a desired level of specificity (e.g., by detecting one or more genetic variants). For example, the genomic locations within the panel can be selected to detect cancer (e.g., by detecting one or more genetic variants) with at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% specificity. The genomic locations within the panel can be selected to detect one or more genetic variants with 100% specificity.

[0182] The genomic locations within the panel can be selected to detect (diagnose) cancer with a desired positive predictive value. Positive predictive value can be increased by increasing sensitivity (e.g., the probability of detecting an actual positive) and / or specificity (e.g., the probability of not mistaking an actual negative for a positive). As a non-limiting example, the genomic locations within the panel can be selected to detect one or more genetic variants with a positive predictive value of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. Regions within the panel can be selected to detect one or more genetic variants with a positive predictive value of 100%.

[0183] The genome position in the panel can be selected to detect (diagnose) cancer with desired accuracy.As used herein, the term "accuracy" can refer to the ability of a test to distinguish between disease symptoms (such as cancer) and healthy symptoms.Accuracy can be quantified using measures such as sensitivity and specificity, predictive value, likelihood ratio, area under ROC curve, Youden index and / or diagnostic odds ratio.

[0184] Accuracy can be presented as a percentage, referring to the ratio of the number of tests that gave correct results to the total number of tests performed. The regions within the panel can be selected to detect cancer with at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% accuracy. The genomic locations within the panel can be selected to detect cancer with 100% accuracy.

[0185] The panel can be selected to be highly sensitive and detect low-frequency genetic variants. For example, the panel can be selected so that genetic variants or tumor markers present in samples at frequencies as low as 0.01%, 0.05%, or 0.001% can be detected with at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% sensitivity. The genomic locations within the panel can be selected to detect tumor markers present in samples at frequencies of 1% or less with a sensitivity of 70% or greater. A panel may be selected to detect tumor markers present at frequencies as low as 0.1% in a sample with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5% or 99.9%. A panel may be selected to detect tumor markers present at frequencies as low as 0.01% in a sample with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5% or 99.9%. The panel may be selected to detect tumor markers present in samples at frequencies as low as 0.001% with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5% or 99.9%.

[0186] The panel can be selected to be highly specific and detect low frequency genetic variants.For example, the panel can be selected so that genetic variants or tumor markers present in samples at frequencies as low as 0.01%, 0.05% or 0.001% can be detected with at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5% or 99.9% specificity.The genomic locations within the panel can be selected to detect tumor markers present in samples at frequencies of 1% or less with 70% or more specificity. A panel may be selected to detect tumor markers present at frequencies as low as 0.1% in a sample with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% specificity. A panel may be selected to detect tumor markers present at frequencies as low as 0.01% in a sample with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% specificity. A panel may be selected to detect tumor markers present at frequencies as low as 0.001% in a sample with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% specificity.

[0187] A panel can be selected to have high accuracy and detect low-frequency genetic variants. A panel can be selected so that genetic variants or tumor markers present in a sample at frequencies as low as 0.01%, 0.05%, or 0.001% can be detected with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% accuracy. Genomic locations within the panel can be selected to detect tumor markers present in a sample at frequencies of 1% or less with 70% or greater accuracy. A panel can be selected to detect tumor markers present in a sample at frequencies as low as 0.1% with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% accuracy. A panel may be selected to detect tumor markers present in a sample at frequencies as low as 0.01% with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5% or 99.9% accuracy. A panel may be selected to detect tumor markers present in a sample at frequencies as low as 0.001% with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5% or 99.9% accuracy.

[0188] Panels can be selected to be highly predictive and to detect low frequency genetic variants. Panels can be selected such that a genetic variant or tumor marker present in a sample at a frequency as low as 0.01%, 0.05%, or 0.001% can have a positive predictive value of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%.

[0189] The concentration of the probe or bait used in the panel can be increased (2-6 ng / μL) to capture more nucleic acid molecules in the sample. The concentration of the probe or bait used in the panel can be at least 2 ng / μL, 3 ng / μL, 4 ng / μL, 5 ng / μL, 6 ng / μL, or more. The concentration of the probe can be about 2 ng / μL to about 3 ng / μL, about 2 ng / μL to about 4 ng / μL, about 2 ng / μL to about 5 ng / μL, or about 2 ng / μL to about 6 ng / μL. The concentration of the probe or bait used in the panel can be 2 ng / μL or more to 6 ng / μL or less. In some cases, this can allow more molecules in a biologic to be analyzed, thereby enabling detection of low-frequency alleles.

[0190] Cancer and other diseases

[0191] In certain embodiments, the methods and aspects disclosed herein are used to diagnose a given disease, disorder, or condition in a patient. Typically, the disease under consideration is some type of cancer. Non-limiting examples of such cancers include cholangiocarcinoma, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain tumor, glioma, astrocytoma, breast carcinoma, metaplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal carcinoma, colon cancer, hereditary nonpolyposis colorectal cancer, colorectal adenocarcinoma, gastrointestinal stromal tumor (GIST), endometrial carcinoma, endometrial stromal sarcoma, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, ocular melanoma, uveal melanoma, gallbladder carcinoma, gallbladder adenocarcinoma, renal cell carcinoma, renal clear cell carcinoma, transitional cell carcinoma, urothelial carcinoma, Wilms' tumor, leukemia, acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), chronic myelogenous leukemia (CML), chronic myelomonocytic leukemia (CML), and chronic myelomonocytic leukemia (CML). ML), liver cancer, liver carcinoma, hepatoma, hepatocellular carcinoma, cholangiocarcinoma, hepatoblastoma, lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B-cell lymphoma, non-Hodgkin's lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, T-cell lymphoma, non-Hodgkin's lymphoma, precursor T-lymphoblastic lymphoma / leukemia, peripheral T-cell lymphoma, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oropharyngeal carcinoma, oral squamous cell carcinoma, osteosarcoma, ovarian carcinoma, pancreatic cancer, pancreatic ductal adenocarcinoma, pseudopapillary tumor, acinar cell carcinoma, prostate cancer, prostate adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine carcinoma, gastric cancer, gastrointestinal stromal tumor (GIST), uterine cancer or uterine sarcoma.

[0192] Non-limiting examples of other genetically based diseases, disorders, or conditions that may optionally be evaluated using the methods and systems disclosed herein include achondroplasia, alpha-1 antitrypsin deficiency, antiphospholipid syndrome, autism, autosomal dominant polycystic kidney disease, Charcot-Marie-Tooth (CMT), cri du chat syndrome, Crohn's disease, cystic fibrosis, Dercum's disease, Down's syndrome, Duane's syndrome, Duchenne muscular dystrophy, Factor V Leiden thrombophilia, familial hypercholesterolemia, familial Mediterranean fever, fragile X syndrome, Gaucher's disease, hemochromatosis, hemophilia, holoprosencephaly, Huntington's disease, Klinefelter's syndrome, Marfan's syndrome, myotonic dystrophy, neurofibromatosis, Noonan's syndrome, osteogenesis imperfecta, Parkinson's disease, phenylketonuria, Poland malformation, These include: porphyria, progeria, retinitis pigmentosa, severe combined immunodeficiency (SCID), sickle cell disease, spinal muscular atrophy, Tay-Sachs disease, thalassemia, trimethylaminuria, Turner syndrome, palatocardiofacial syndrome, WAGR syndrome, and Wilson's disease.

[0193] Customized Therapy and Related Administration

[0194] In some embodiments, the methods disclosed herein relate to the identification and administration of a therapy to a patient with a given disease, disorder, or condition. Essentially any cancer therapy (e.g., surgical therapy, radiation therapy, chemotherapy, etc.) can be included as part of these methods. Typically, the therapy includes at least one immunotherapy (or immunotherapeutic agent). Immunotherapy generally refers to a method of enhancing the immune response to a given type of cancer. In certain embodiments, immunotherapy refers to a method of enhancing T-cell response to tumors or cancer.

[0195] In some embodiments, immunotherapy or immunotherapeutic agents target immune checkpoint molecules. Certain tumors can evade the immune system by exploiting immune checkpoint pathways. Thus, targeting immune checkpoints has emerged as an effective approach to counter the ability of tumors to evade the immune system and to activate anti-tumor immunity against certain cancers. Pardoll, Nature Reviews Cancer,2012,12:252-264.

[0196] In certain embodiments, the immune checkpoint molecule is an inhibitory molecule that attenuates signals involved in T cell responses to antigens. For example, CTLA4 is expressed on T cells and plays a role in downregulating T cell activation by binding to CD80 (also known as B7.1) or CD86 (also known as B7.2) on antigen-presenting cells. PD-1 is another inhibitory checkpoint molecule expressed on T cells. PD-1 limits T cell activity in peripheral tissues during inflammatory responses. Furthermore, the ligand for PD-1 (PD-L1 or PD-L2) is typically upregulated on the surface of many different tumors, resulting in downregulation of anti-tumor immune responses in the tumor microenvironment. In certain embodiments, the inhibitory immune checkpoint molecule is CTLA4 or PD-1. In other embodiments, the inhibitory immune checkpoint molecule is a ligand for PD-1, such as PD-L1 or PD-L2. In other embodiments, the inhibitory immune checkpoint molecule is a ligand for CTLA4, such as CD80 or CD86. In other embodiments, the inhibitory immune checkpoint molecule is lymphocyte activation gene 3 (LAG3), killer cell immunoglobulin-like receptor (KIR), T-cell membrane protein 3 (TIM3), galectin 9 (GAL9), or adenosine A2a receptor (A2aR).

[0197] Antagonists that target these immune checkpoint molecules can be used to enhance antigen-specific T cell responses against certain cancers. Thus, in certain embodiments, the immunotherapy or immunotherapeutic agent is an antagonist of an inhibitory immune checkpoint molecule. In certain embodiments, the inhibitory immune checkpoint molecule is PD-1. In certain embodiments, the inhibitory immune checkpoint molecule is PD-L1. In certain embodiments, the antagonist of an inhibitory immune checkpoint molecule is an antibody (e.g., a monoclonal antibody). In certain embodiments, the antibody or monoclonal antibody is an anti-CTLA4, anti-PD-1, anti-PD-L1, or anti-PD-L2 antibody. In certain embodiments, the antibody is a monoclonal anti-PD-1 antibody. In some embodiments, the antibody is a monoclonal anti-PD-L1 antibody. In certain embodiments, the monoclonal antibody is a combination of an anti-CTLA4 antibody and an anti-PD-1 antibody, a combination of an anti-CTLA4 antibody and an anti-PD-L1 antibody, or a combination of an anti-PD-L1 antibody and an anti-PD-1 antibody. In certain embodiments, the anti-PD-1 antibody is one or more of pembrolizumab (Keytruda®) or nivolumab (Opdivo®). In certain embodiments, the anti-CTLA4 antibody is ipilimumab (Yervoy®). In certain embodiments, the anti-PD-L1 antibody is one or more of atezolizumab (Tecentriq®), avelumab (Bavencio®), or durvalumab (Imfinzi®).

[0198] In certain embodiments, the immunotherapy or immunotherapeutic agent is an antagonist (e.g., an antibody) against CD80, CD86, LAG3, KIR, TIM3, GAL9, or A2aR. In other embodiments, the antagonist is a soluble version of an inhibitory immune checkpoint molecule (e.g., a soluble fusion protein comprising the extracellular domain of an inhibitory immune checkpoint molecule and the Fc domain of an antibody). In certain embodiments, the soluble fusion protein comprises the extracellular domain of CTLA4, PD-1, PD-L1, or PD-L2. In some embodiments, the soluble fusion protein comprises the extracellular domain of CD80, CD86, LAG3, KIR, TIM3, GAL9, or A2aR. In one embodiment, the soluble fusion protein comprises the extracellular domain of PD-L2 or LAG3.

[0199] In certain embodiments, the immune checkpoint molecule is a costimulatory molecule that amplifies signals involved in T cell responses to an antigen. For example, CD28 is a costimulatory receptor expressed on T cells. When a T cell binds to an antigen via its T cell receptor, CD28 binds to CD80 (also known as B7.1) or CD86 (also known as B7.2) on an antigen-presenting cell, amplifying T cell receptor signaling and promoting T cell activation. Because CD28 binds to the same ligands (CD80 and CD86) as CTLA4, CTLA4 can counteract or regulate costimulatory signaling mediated by CD28. In certain embodiments, the immune checkpoint molecule is a costimulatory molecule selected from CD28, inducible T cell costimulator (ICOS), CD137, OX40, or CD27. In other embodiments, the immune checkpoint molecule is a ligand for a costimulatory molecule, such as CD80, CD86, B7RP1, B7-H3, B7-H4, CD137L, OX40L, or CD70.

[0200] By using agonists that target these costimulatory checkpoint molecules, antigen-specific T cell responses to certain cancers can be enhanced. Thus, in certain embodiments, the immunotherapy or immunotherapeutic agent is an agonist of a costimulatory checkpoint molecule. In certain embodiments, the agonist of the costimulatory checkpoint molecule is an agonist antibody, preferably a monoclonal antibody. In certain embodiments, the agonist antibody or monoclonal antibody is an anti-CD28 antibody. In other embodiments, the agonist antibody or monoclonal antibody is an anti-ICOS, anti-CD137, anti-OX40, or anti-CD27 antibody. In other embodiments, the agonist antibody or monoclonal antibody is an anti-CD80, anti-CD86, anti-B7RP1, anti-B7-H3, anti-B7-H4, anti-CD137L, anti-OX40L, or anti-CD70 antibody.

[0201] Therapeutic options for treating particular genetically based diseases, disorders or conditions other than cancer are widely known to those of skill in the art and will be apparent upon consideration of the particular disease, disorder or condition under consideration.

[0202] In certain embodiments, the customized therapies described herein are typically administered parenterally (e.g., intravenously or subcutaneously). Pharmaceutical compositions containing immunotherapeutic agents are typically administered intravenously. Certain therapeutic agents are administered orally. However, customized therapies (e.g., immunotherapeutic agents, etc.) can also be administered by any method known in the art, including, for example, buccal, sublingual, rectal, vaginal, intraurethral, ​​topical, intraocular, intranasal, and / or intraauricular, including tablets, capsules, granules, aqueous suspensions, gels, sprays, suppositories, salves, ointments, etc.

[0203] System and computer-readable medium

[0204] The present disclosure also provides various systems and computer program products or machine-readable media. In some embodiments, for example, the methods described herein are performed or facilitated, as appropriate, at least in part, using systems, distributed computing hardware and applications (e.g., cloud computing services), electronic communications networks, communications interfaces, computer program products, machine-readable media, electronic storage media, software (e.g., machine-executable code or logical instructions), and the like. For example, FIG. 7 provides a schematic diagram of an exemplary system suitable for at least performing and using aspects of the methods disclosed herein. As shown, system 700 includes at least one controller or computer, e.g., a server 702 (e.g., a search engine server) including a processor 704 and memory, storage devices, or memory components 706, as well as one or more other communication devices 714 and 716 (e.g., client-side computer terminals, phones, tablets, laptops, other mobile devices, etc.) located remotely from the remote server 702 and in communication with the remote server 702 via an electronic communications network 712 (e.g., the Internet or other internetwork). The communication devices 714 and 716 are typically For example, the system 700 includes an electronic display (e.g., an Internet-enabled computer, etc.) in communication with the server 702 computer via a network 712, where the electronic display includes a user interface (e.g., a graphical user interface (GUI), a web-based user interface, etc.) for displaying results from performing the methods described herein. In certain embodiments, the communication network also encompasses the physical movement of data from one location to another, for example, using a hard drive, thumb drive, or other data storage mechanism. The system 700 also includes a program product 708 stored on a computer-readable or machine-readable medium (e.g., one or more various types of memory readable by the server 702, e.g., memory 706 of the server 702) for use in, for example, a guided search application or other application executable by one or more other communication devices (e.g., 714 (shown schematically as a desktop or personal computer) and 716 (shown schematically as a tablet computer)). In some embodiments, system 700 also optionally includes at least one database server (e.g., a server 710 associated with an online website storing searchable data (e.g., classifier scores, control sample or comparator result data, indexed customized therapies, etc.) either directly or via search engine server 702). System 700 also optionally includes one or more other servers located remotely from server 702, each associated with one or more database servers 710 located remotely or proximately to each of the other servers, as appropriate. The other servers can beneficially serve geographically dispersed users and can enhance geographically distributed operations.

[0205] As one skilled in the art would understand, the memory 706 of the server 702 may include volatile and / or nonvolatile memory, including, for example, RAM, ROM, and magnetic or optical disks, among others, as appropriate. While depicted as a single server, one skilled in the art would understand that the depicted server 702 configuration is provided merely as an example, and that other types of servers or computers configured according to various other methods or architectures may also be used. The server 702 depicted schematically in FIG. 7 may be a server, server cluster, or server farm, and is not limited to any individual physical server. Server locations may be deployed as a server farm or server cluster managed by a server hosting provider. The number of servers and their configuration and arrangement may increase based on usage, demand, and capacity requirements for the system 700. As one skilled in the art would also understand, the other user communication devices 714 and 716 in these embodiments may be, for example, laptops, desktops, tablets, personal digital assistants (PDAs), mobile phones, servers, or other types of computers. As those skilled in the art will know and understand, network 712 may include the Internet, an intranet, a telecommunications network, an extranet, or the World Wide Web, of multiple computers / servers in communication with one or more other computers via portions of a communications network and / or local area network or other area network.

[0206] As those skilled in the art will further appreciate, exemplary program product or machine-readable medium 708 is in the form of microcode, programs, cloud computing formats, routines, and / or symbolic languages ​​that provide one or more sequenced operations to control and direct the operation of the hardware, as appropriate. Program product 708, according to exemplary embodiments, also need not reside entirely in volatile memory, but may be selectively loaded as needed according to various methods, as those skilled in the art will know and appreciate.

[0207] As will be further understood by those skilled in the art, the term “computer-readable medium” or “machine-readable medium” refers to any medium that participates in providing instructions to a processor for execution. For example, the term “computer-readable medium” or “machine-readable medium” encompasses distribution media, cloud computing formats, intermediate storage media, computer execution memory, and any other medium or device that can store a program product 708 that performs the functions or processes of various embodiments of the present disclosure, for example, for reading by a computer. A “computer-readable medium” or “machine-readable medium” can take many forms, including, but not limited to, non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical or magnetic disks. Volatile media include dynamic memory, such as the main memory of a given system. Transmission media include coaxial cables, copper wire, and fiber optics, including the wires that comprise a bus. Transmission media can also take the form of acoustic or light waves (e.g., those generated during radio wave and infrared communications, among others). Exemplary forms of computer-readable media include a floppy disk, a flexible disk, a hard disk, magnetic tape, a flash drive or any other magnetic medium, a CD-ROM, any other optical medium, punch cards, paper tape, any other physical medium with a pattern of holes, RAM, PROM and EPROM, FLASH-EPROM, any other memory chip or cartridge, a carrier wave, or any other medium from which a computer can read.

[0208] The program product 708 is copied from the computer-readable medium to a hard disk or similar intermediate storage medium, as needed. When the program product 708, or portions thereof, are executed, they are loaded, as needed, from their distribution medium, intermediate storage medium, etc. into the execution memory of one or more computers, and those computers are configured to operate according to the functions or methods of the various embodiments. All such operations are well known to those skilled in the art of, for example, computer systems.

[0209] By way of further illustration, in certain embodiments, the present application provides systems that include one or more processors and one or more memory components in communication with the processors. The memory components typically contain one or more instructions that, when executed, cause the processor to provide information for display (e.g., via communication devices 714, 716, etc.), such as sequence information, subclonality scores, classifier scores, test results, control or comparator results, customized therapies, etc., and / or receive information (e.g., via communication devices 714, 716, etc.) from other system components and / or system users.

[0210] In some embodiments, the program product 708 comprises computer-executable non-transitory instructions that, when executed by the electronic processor 704, at least: (a) generate a subclonality score for each allele in the set of classification alleles from sequence information comprising sequence reads obtained from cell-free nucleic acid (cfNA) fragments from one or more reference samples, wherein each classification allele is potentially clinically significant and comprises a minor allele observed at a given locus in the reference samples; and b) compare the subclonality score to at least one selected cutoff threshold, wherein classification alleles with a subclonality score above the selected cutoff threshold indicate that the classification alleles are derived from reference cfNA fragments originating from non-target cells, and the classification alleles are added to a non-target nucleic acid variant filter list, and / or classification alleles with a subclonality score below the cutoff threshold indicate that the classification alleles are derived from reference cfNA fragments originating from target cells, and the classification alleles are added to a target nucleic acid variant filter list. Further computer-readable medium embodiments are described herein.

[0211] System 700 also typically includes additional system components configured to perform various aspects of the methods described herein. In some of these embodiments, one or more of these additional system components are located remotely from and in communication with remote server 702 via electronic communications network 712, while in other embodiments, one or more of these additional system components are located locally and in communication with server 702 (i.e., in the absence of electronic communications network 712) or directly with, for example, desktop computer 714.

[0212] In some embodiments, for example, additional system components include a sample preparation component 718 operably connected to the controller 702 (either directly or indirectly, e.g., via the electronic communication network 712). The sample preparation component 718 is configured to prepare nucleic acids in the sample (e.g., prepare a library of nucleic acids) to be amplified and / or sequenced by a nucleic acid amplification component (e.g., a thermal cycler, etc.) and / or a nucleic acid sequencing device. In certain of these embodiments, the sample preparation component 718 is configured to isolate nucleic acids from other components in the sample, to attach one or more adapters comprising barcodes to nucleic acids as described herein, to selectively enrich one or more regions from a genome or transcriptome prior to sequencing, etc.

[0213] In certain embodiments, system 700 also includes a nucleic acid amplification component 720 (e.g., a thermal cycler, etc.) operably connected to controller 702 (either directly or indirectly, e.g., via electronic communication network 712). Nucleic acid amplification component 720 is configured to amplify nucleic acids in a sample from a subject. For example, nucleic acid amplification component 720 is configured to amplify selectively enriched regions from a genome or transcriptome in the sample, as desired, as described herein.

[0214] The system 700 also typically includes at least one nucleic acid sequencer 722 operably connected to the controller 702 (either directly or indirectly, e.g., via the electronic communications network 712). The nucleic acid sequencer 722 is configured to provide sequence information from nucleic acids (e.g., amplified nucleic acids) in a sample from a subject. Essentially any type of nucleic acid sequencer can be adapted for use in these systems. For example, the nucleic acid sequencer 722 is optionally configured to perform bisulfite sequencing, pyrosequencing, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing by ligation, sequencing by hybridization, or other techniques on the nucleic acids to generate sequencing reads. Optionally, the nucleic acid sequencer 722 is configured to group sequence reads into families of sequence reads, each family including sequence reads generated from nucleic acids in a given sample. In some embodiments, the nucleic acid sequencer 722 generates sequencing reads using clonal single-molecule arrays obtained from a sequencing library. In certain embodiments, the nucleic acid sequencing device 722 comprises at least one chip having an array of microwells for sequencing a sequencing library to generate sequencing reads.

[0215] To facilitate full or partial automation of the system, system 700 also typically includes a material transfer component 724 operably connected (directly or indirectly, e.g., via electronic communications network 712) to controller 702. Material transfer component 724 is configured to transfer one or more materials (e.g., nucleic acid samples, amplicons, reagents, etc.) to and / or from nucleic acid sequencing device 722, sample preparation component 718, and nucleic acid amplification component 720.

[0216] For further details regarding computer systems and networks, databases, and computer program products, see, for example, Peterson, Computer Networks: A Systems Approach, Morgan Kaufmann, 5th Ed. (2011), Kurose, Computer Networking: A Top-Down Approach, Pearson, 7 th Ed. (2016), Elmasri, Fundamentals of Database Systems, Addison Wesley, 6th Ed. (2010), Coronel, Database Systems: Design, Implementation, & Management, Cengage Learning, 11 th Ed. (2014), Tucker, Programming Languages, McGraw-Hill Science / Engineering / Math, 2nd Ed. (2006), and Rhoton, Cloud Computing Architected: Solution Design Handbook, Recursive Press (2011), each of which is incorporated by reference in its entirety.

[0217] [Example]

[0218] Example 1: Circulating tumor cell-free DNA

[0219] Circulating tumor cell-free DNA (CtDNA) in patients with colorectal cancer (CRC) after surgery correlates with molecular residual disease and may be useful for determining prognosis and guiding adjuvant therapy decision-making.

[0220] Postoperative ctDNA is strongly associated with disease recurrence in patients with metastatic CRC who underwent curative surgery (p=0.004). See Overman et al. (2017). Using highly sensitive panels, detection of circulating tumor DNA (ctDNA) can detect minimal residual disease and predict disease recurrence after hepatic resection. JCO35(suppl). Early studies used clinically unreasonable assays indexed to individual patient-specific tumor tissue-derived mutations or were confounded by non-tumor-related somatic alterations, including variants related to clonal hematopoiesis. See Tie J. et al. (2016). Sci Transl Med 8(346).

[0221] By using a highly sensitive CRC next-generation sequencing (NGS) panel, we provide data demonstrating that the detection of postoperative ctDNA does not require the presence of known somatic alterations. With the goal of increasing the specificity of ctDNA detection in postoperative CRC patients, we used a variant classifier to further distinguish between tumor- and non-tumor-derived alterations (see classifier filters in Figure 8A-C).

[0222] Patients with CRC scheduled for liver metastasectomy were prospectively enrolled in an IRB-approved clinical trial. Pre- and post-operative plasma was sequenced to high depth using a 38-gene NGS panel with a theoretical sensitivity of 96% for CRC. Fifty-one metastatic colorectal cancer patients with both pre- and post-operative ctDNA results were recruited at a single institution (Table 4: Cohort demographics). Tumor tissue was sequenced using this panel or a local test. ctDNA profiles from 17,700 CRC patients (Guardant Health, Redwood City, CA) were used to train a variant classifier to exclude non-tumor-derived alterations. The classifier was designed to identify cfDNA mutations of tumor origin. [Table 4]

[0223] Predicting recurrence using only somatic variant detection after surgery is associated with a high clinical false-positive rate. Many mutations of non-tumor origin occur at low allele frequencies. However, a simple threshold for allele frequency may exclude many clinically relevant mutations. Filtering using tumor tissue is effective but may be clinically unreasonable due to the increased complexity and cost. Filtering using a novel variant classifier without prior knowledge of tumor genotype eliminated false positives while maintaining clinically acceptable sensitivity. A priori variant classification may enable clinically feasible ctDNA diagnostics for adjuvant treatment decision-making in early-stage disease.

[0224] Although the foregoing disclosure has been described in some detail by way of illustration and example for purposes of clarity and understanding, it will be apparent to those skilled in the art upon reading this disclosure that various changes in form and detail may be made without departing from the true scope of the disclosure and within the purview of the appended claims. For example, all methods, systems, computer-readable media and / or component features, steps, elements or other aspects thereof may be used in various combinations.

[0225] All patents, patent applications, websites, other publications or documents, accession numbers, etc. cited herein are incorporated by reference in their entirety for all purposes to the same extent as if each individual item were specifically and individually indicated to be incorporated by reference. Where different versions of a sequence are associated with accession numbers at various times, the version associated with that accession number as of the effective filing date of this application is meant. The effective filing date means the actual filing date or, if applicable, the filing date of the priority application referencing the accession number, whichever is earlier. Similarly, where different versions of publications, websites, etc. are published at different times, the version published closest to the effective filing date of this application is meant unless otherwise indicated. The present invention provides, for example, the following items. (Item 1) 1. A method for detecting nucleic acid molecules originating from target cells in a subject using, at least in part, a computer, the method comprising: (a) receiving, by the computer, test sequence information comprising sequence reads obtained from cell-free nucleic acid (cfNA) fragments from a test sample obtained from the subject; and (b) identifying at least one allelic variant in the test sequence information; (c) mapping the allelic variant to at least one classification allele on a target nucleic acid variant filter list; (d) identifying a subclonality score for the classified allele; and (e) comparing the subclonality score to at least one selected cutoff threshold, wherein when the subclonality score is less than the selected cutoff threshold, it indicates that the classification allele is derived from a reference cfNA fragment originating from the target cell, thereby detecting a nucleic acid molecule originating from the target cell in the subject. A method comprising: (Item 2) 1. A method for detecting nucleic acid molecules originating from tumor cells in a subject using, at least in part, a computer, the method comprising: (a) receiving, by the computer, test sequence information comprising sequence reads obtained from cell-free deoxyribonucleic acid (cfDNA) fragments in a test sample obtained from the subject; (b) removing, by the computer, one or more of the sequence reads or classification alleles originating from hematopoietic stem cells of the subject from the test sequence information to generate filtered test sequence information; and (c) identifying, by the computer, one or more sequence reads present in the filtered test sequence information that substantially align with reference sequence information obtained from one or more reference subjects, which reference sequence information originates from one or more tumor cells within the reference subjects, thereby detecting the nucleic acid molecules originating from the tumor cells within the subject. A method comprising: (Item 3) 1. A method of treating a disease in a subject, the method comprising: (a) receiving test sequence information comprising sequence reads obtained from cell-free nucleic acid (cfNA) fragments from a test sample obtained from the subject; (b) identifying at least one allelic variant present in the test sequence information that substantially matches at least one classification allele on the target nucleic acid variant filter list, wherein the classification allele comprises a subclonality score below at least one selected cutoff threshold, thereby indicating that the classification allele is derived from a reference cfNA fragment originating from a diseased cell, thereby diagnosing the disease in the subject; and (c) administering one or more therapies to said subject, thereby treating the disease in said subject. A method comprising: (Item 4) 1. A method of generating a classifier using, at least in part, a computer, the method comprising: (a) generating by the computer a subclonality score for each allele in a classification allele set from sequence information comprising sequence reads obtained from cell-free nucleic acid (cfNA) fragments from one or more reference samples, wherein each classification allele is potentially clinically significant and comprises a minor allele observed at a given locus in the reference samples; and (b) comparing, by the computer, the subclonality score to at least one selected cutoff threshold, wherein classification alleles having a subclonality score above the selected cutoff threshold indicate that the classification alleles are derived from reference cfNA fragments originating from non-target cells, and the classification alleles are added to a non-target nucleic acid variant filter list; and / or Classifying alleles having a subclonality score below the cutoff threshold indicate that the classifying alleles are derived from a reference cfNA fragment originating from the target cell, and the classifying alleles are added to a target nucleic acid variant filter list; thereby generating the classifier. A method comprising: (Item 5) 1. A method of generating a classifier using, at least in part, a computer, the method comprising: (a) identifying by said computer a set of classification alleles from sequence information comprising sequence reads obtained from cell-free nucleic acid (cfNA) fragments from one or more reference samples, wherein each classification allele is potentially clinically significant and comprises a minor allele observed at a given locus in said reference samples; (b) determining, by the computer, from the sequence information, a minor allele frequency (MAF) value for each classification allele in each of the reference samples; (c) determining by said computer a maximum minor allele frequency (maxMAF) value for each of said reference samples; (d) calculating, by the computer, for each classification allele observed in a given reference sample, a ratio of the value of the MAF to the value of the maxMAF for at least a portion of the reference sample to generate a ratio value; (e) calculating by the computer, for each of the classification alleles, a ratio of the number of times that the given classification allele in at least some of the reference samples had a ratio value that was less than at least one selected clonality boundary value to the total number of times that the given classification allele appeared in at least some of the reference samples, to generate a subclonality score for each of the classification alleles in at least some of the reference samples; and (f) comparing, by the computer, the subclonality score to at least one selected cutoff threshold, wherein classification alleles having a subclonality score above the selected cutoff threshold indicate that the classification alleles are derived from reference cfNA fragments originating from non-target cells, and the classification alleles are added to a non-target nucleic acid variant filter list; and / or Classifying alleles having a subclonality score below the cutoff threshold indicate that the classifying alleles are derived from a reference cfNA fragment originating from the target cell, and the classifying alleles are added to a target nucleic acid variant filter list; thereby generating the classifier. A method comprising: (Item 6) 1. A method for generating a database of subclonality scores for use in classifying the cellular origin of cell-free nucleic acid (cfNA) fragments in a test sample obtained from a subject, the method comprising: (a) computationally identifying a set of classification alleles from sequence information comprising sequence reads obtained from cell-free nucleic acid (cfNA) fragments from one or more reference samples, wherein each classification allele is potentially clinically significant and comprises a minor allele observed at a given locus in said reference samples; (b) determining, by the computer, from the sequence information, a minor allele frequency (MAF) value for each classification allele in each of the reference samples; (c) determining by said computer a maximum minor allele frequency (maxMAF) value for each of said reference samples; (d) for each classification allele observed in a given reference sample, calculating by the computer a ratio of the value of the MAF to the value of the maxMAF for at least a portion of the reference sample to generate a ratio value; (e) calculating by the computer, for each of the classification alleles, a ratio of the number of times that the given classification allele in at least some of the reference samples had a ratio value that was less than at least one selected clonality boundary value to the total number of times that the given classification allele appeared in at least some of the reference samples, to generate a subclonality score for each of the classification alleles in at least some of the reference samples; and (f) non-temporarily storing the subclonality scores indexed to the corresponding classification alleles in a database system, thereby creating a database of the subclonality scores for use in classifying the cellular origin of cfNA fragments in a test sample obtained from a subject. A method comprising: (Item 7) 10. The method of claim 1, wherein identifying the classification allele set comprises determining, from the sequence information obtained from the reference samples, a MAF value for each somatic nucleic acid variant at each locus within a set of potentially clinically significant target genomic loci, wherein the set of target genomic loci is identical in each reference sample; and determining a maxMAF value for each of the reference samples to generate allele information. (Item 8) 10. The method of any one of the preceding items, wherein the MAF for each classification allele is less than about 2%. (Item 9) 10. The method of any one of the preceding items, wherein the MAF for each classification allele is less than about 1%. (Item 10) 10. The method of any one of the preceding items, comprising using clinical information indexed to the reference samples to generate the classifier. (Item 11) 10. The method of any one of the preceding items, comprising detecting the nucleic acid molecule originating from the target cell in the subject using clinical information indexed to the test sample. (Item 12) The method of any one of the preceding items, wherein the clinical information is selected from the group consisting of age, sex, race, weight, body mass index (BMI), medical history, smoking, and alcohol consumption. (Item 13) 10. The method of any one of the preceding items, comprising determining a subclonality score using the frequency of each MAF / max-MAF value for each of the classification alleles. (Item 14) 10. The method of any one of the preceding items, wherein the selected clonality boundary value is within the range of about 1% to about 99%. (Item 15) 10. The method of any one of the preceding items, wherein the selected clonality boundary value is about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90%. (Item 16) 10. The method of any one of the preceding items, wherein the selected cutoff threshold is within the range of about 1% to about 99%. (Item 17) The method of any one of the preceding items, wherein the selected cutoff threshold is about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90%. (Item 18) The method of any one of the preceding items, comprising comparing the subclonality score to a plurality of selected cutoff thresholds. (Item 19) 10. The method of claim 1, wherein the plurality of selected cutoff thresholds comprises a first cutoff threshold and a second cutoff threshold, the first cutoff threshold being greater than the second cutoff threshold, and wherein classification alleles having a subclonality score greater than the first cutoff threshold indicate that the classification alleles are derived from reference cfNA fragments originating from non-target cells, and the classification alleles are added to the non-target nucleic acid variant filter list; and / or classification alleles having a subclonality score less than the second cutoff threshold indicate that the classification alleles are derived from reference cfNA fragments originating from target cells, and the classification alleles are added to the target nucleic acid variant filter list. (Item 20) 10. The method of any one of the preceding items, comprising classifying an allelic variant in the test sequence information as originating from a target cell when the allelic variant substantially matches at least one classification allele on a non-target nucleic acid variant filter list and comprises a MAF greater than about 1%. (Item 21) 10. The method of any one of the preceding items, comprising classifying an allelic variant in the test sequence information as originating from a target cell when the allelic variant substantially matches at least one classified allele on a non-target nucleic acid variant filter list and comprises a truncation, an indel, and / or a splice site variant. (Item 22) 10. The method of any one of the preceding items, comprising determining the frequency of each ratio value for a given classification allele in at least some of the reference samples. (Item 23) The method of any one of the preceding items, comprising using the classifier to determine whether a test sample obtained from a subject contains cfNA fragments originating from the target cells. (Item 24) The method of any one of the preceding items, comprising using the classifier to determine whether a test sample obtained from a subject contains cfNA fragments originating from the non-target cells. (Item 25) A database comprising the target nucleic acid variant filter list and / or the non-target nucleic acid variant filter list of any one of the preceding items. (Item 26) Item 11. The method of any one of the preceding items, wherein the non-target cells comprise non-diseased cells. (Item 27) Item 10. The method of any one of the preceding items, wherein the non-target cells comprise hematopoietic stem cells. (Item 28) Item 10. The method of any one of the preceding items, wherein the non-target cells comprise non-tumor cells. (Item 29) 10. The method of any one of the preceding items, wherein the non-target cells comprise maternal cells. (Item 30) 10. The method of any one of the preceding items, wherein the non-target cells comprise transplant recipient cells. (Item 31) 10. The method of any one of the preceding items, wherein the target cells comprise diseased cells. (Item 32) Item 11. The method of any one of the preceding items, wherein the target cells comprise tumor cells. (Item 33) 10. The method of any one of the preceding items, wherein the target cells comprise fetal cells. (Item 34) 10. The method of any one of the preceding items, wherein the target cells comprise transplant donor cells. (Item 35) The method of any one of the preceding items, wherein the disease comprises cancer and the therapy comprises at least one immunotherapy. (Item 36) The method of any one of the preceding items, wherein the subject is a mammalian subject. (Item 37) 38. The method of any one of the preceding claims, wherein the mammalian subject is a human subject. The method of any one of the preceding items, further comprising obtaining the test sample from the subject. (Item 39) 8. The method of any one of the preceding items, wherein the test sample is selected from the group consisting of blood, plasma, serum, sputum, urine, semen, vaginal fluid, stool, synovial fluid, cerebrospinal fluid, and saliva. (Item 40) The method of any one of the preceding items, further comprising generating the test sequence information from the cfNA fragments in the test sample. (Item 41) 2. The method of any one of the preceding items, further comprising amplifying a segment of the cfNA fragment that includes a target genomic locus to produce an amplified nucleic acid. (Item 42) The method of any one of the preceding items, further comprising sequencing the cfNA fragments in the test sample to generate the test sequence information. (Item 43) The method of any one of the preceding items, wherein the test sequence information is obtained from a targeted segment of the cfNA fragment in the test sample, and the targeted segment is obtained by selectively enriching one or more regions from the cfNA fragment in the test sample prior to sequencing. (Item 44) 10. The method of any one of the preceding items, further comprising amplifying the obtained targeted segment prior to sequencing. (Item 45) The method of any one of the preceding items, further comprising attaching one or more adapters comprising a barcode to the cfNA fragments and / or the amplified targeting segments prior to sequencing. (Item 46) 10. The method of any one of the preceding items, wherein the sequencing is selected from the group consisting of targeted sequencing, bisulfite sequencing, intron sequencing, exome sequencing, and whole genome sequencing. (Item 47) 1. A system comprising a computer-readable medium or a controller that can access said computer-readable medium, said instructions including computer-executable non-transitory instructions, which when executed by at least one electronic processor, perform at least: (a) receiving test sequence information including sequence reads obtained from cell-free nucleic acid (cfNA) fragments from a test sample obtained from a subject; and (b) identifying the presence of at least one allelic variant in the test sequence information that substantially matches at least one classification allele on the target nucleic acid variant filter list, wherein the classification allele comprises a subclonality score below at least one selected cutoff threshold, thereby indicating that the classification allele is derived from a reference cfNA fragment originating from a target cell, thereby indicating that the allelic variant in the test sequence information originated from the target cell within the subject. The system. (Item 48) 1. A system comprising a computer-readable medium or a controller that can access said computer-readable medium, said instructions including computer-executable non-transitory instructions, which when executed by at least one electronic processor, perform at least: (a) receiving test sequence information comprising sequence reads obtained from cell-free deoxyribonucleic acid (cfDNA) fragments in a test sample obtained from the subject; (b) removing one or more of the sequence reads or classification alleles originating from hematopoietic stem cells of the subject from the test sequence information to generate filtered test sequence information; and (c) identifying one or more sequence reads present in the filtered test sequence information that substantially align with reference sequence information obtained from one or more reference subjects, wherein the reference sequence information originates from tumor cells within the reference subjects, thereby indicating that the test sample contains one or more cfDNA fragments that originated from the tumor cells within the subject. The system. (Item 49) 1. A system comprising a computer-readable medium or a controller that can access said computer-readable medium, said instructions including computer-executable non-transitory instructions, which when executed by at least one electronic processor, perform at least: (a) generating a subclonality score for each allele in a classification allele set from sequence information comprising sequence reads obtained from cell-free nucleic acid (cfNA) fragments from one or more reference samples, wherein each classification allele is potentially clinically significant and comprises a minor allele observed at a given locus in the reference samples; and (b) comparing the subclonality score with at least one selected cutoff threshold. and wherein classification alleles having a subclonality score above the selected cutoff threshold indicate that the classification alleles are derived from reference cfNA fragments originating from non-target cells, and the classification alleles are added to a non-target nucleic acid variant filter list; and / or Classifying alleles having a subclonality score below the cutoff threshold indicate that the classifying alleles are derived from the reference cfNA fragment originating from the target cell, and the classifying alleles are added to a target nucleic acid variant filter list. system. (Item 50) 1. A system comprising a computer-readable medium or a controller that can access said computer-readable medium, said instructions including computer-executable non-transitory instructions, which when executed by at least one electronic processor, perform at least: (a) identifying a set of classification alleles from sequence information comprising sequence reads obtained from cell-free nucleic acid (cfNA) fragments from one or more reference samples, wherein each classification allele is potentially clinically significant and comprises a minor allele observed at a given locus in said reference samples; (b) determining a minor allele frequency (MAF) value for each classification allele in each of the reference samples from the sequence information; (c) determining a maximum minor allele frequency (maxMAF) value for each of said reference samples; (d) for each classification allele observed in a given reference sample, calculating a ratio of the value of the MAF to the value of the maxMAF for at least a portion of the reference sample to generate a ratio value; (e) calculating, for each of the classification alleles, a ratio of the number of times that the given classification allele in at least some of the reference samples had a ratio value that was less than at least one selected clonality boundary value to the total number of times that the given classification allele appeared in at least some of the reference samples to generate a subclonality score for each of the classification alleles in at least some of the reference samples; and (f) comparing said subclonality score with at least one selected cutoff threshold. and wherein classification alleles having a subclonality score above the selected cutoff threshold indicate that the classification alleles are derived from reference cfNA fragments originating from non-target cells, and the classification alleles are added to a non-target nucleic acid variant filter list; and / or Classifying alleles having a subclonality score below the cutoff threshold indicate that the classifying alleles are derived from the reference cfNA fragment originating from the target cell, and the classifying alleles are added to a target nucleic acid variant filter list. system. (Item 51) The system of any one of the preceding items, further comprising a nucleic acid sequencing device operably connected to the controller, the nucleic acid sequencing device configured to provide the sequence information from the cfNA fragments in the test sample and / or the reference sample. (Item 52) 10. The system of any one of the preceding items, wherein the nucleic acid sequencer is configured to perform pyrosequencing, bisulfite sequencing, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, ligation-based sequencing, or hybridization-based sequencing on the nucleic acids to generate sequencing reads. (Item 53) The system of any one of the preceding items, further comprising a sample preparation component operably connected to the controller, the sample preparation component configured to prepare the cfNA fragments to be sequenced by a nucleic acid sequencer. (Item 54) 10. The system of any one of the preceding items, wherein the sample preparation component is configured to selectively enrich regions from the cfNA fragments in the test sample and / or the reference sample. (Item 55) 10. The system of any one of the preceding items, wherein the sample preparation component is configured to attach one or more adapters comprising a barcode to the cfNA fragments. (Item 56) 10. The system of any one of the preceding items, comprising a nucleic acid amplification component operably connected to the controller, the nucleic acid amplification component configured to amplify the cfNA fragments in the test sample and / or the reference sample. (Item 57) 10. The system of any one of the preceding items, wherein the nucleic acid amplification component is configured to amplify selectively enriched regions from the cfNA fragments in the test sample and / or the reference sample. (Item 58) 10. The system of any one of the preceding items, further comprising a material transfer component operably connected to the controller, the material transfer component configured to transfer one or more materials between a nucleic acid sequencing device, a nucleic acid amplification component, and / or a sample preparation component. (Item 59) 10. The system of any one of the preceding items, further comprising a database operably connected to the controller, the database including the non-target nucleic acid variant filter list and / or the target nucleic acid variant filter list. (Item 60) A computer-readable medium containing computer-executable non-transitory instructions, the instructions, when executed by at least one electronic processor, at least: (a) receiving test sequence information including sequence reads obtained from cell-free nucleic acid (cfNA) fragments from a test sample obtained from a subject; and (b) identifying the presence of at least one allelic variant in the test sequence information that substantially matches at least one classification allele on the target nucleic acid variant filter list, wherein the classification allele comprises a subclonality score below at least one selected cutoff threshold, thereby indicating that the classification allele is derived from a reference cfNA fragment originating from a target cell, thereby indicating that the allelic variant in the test sequence information originated from the target cell within the subject. A computer-readable medium for performing the above steps. (Item 61) A computer-readable medium containing computer-executable non-transitory instructions, the instructions, when executed by at least one electronic processor, at least: (a) receiving test sequence information comprising sequence reads obtained from cell-free deoxyribonucleic acid (cfDNA) fragments in a test sample obtained from the subject; (b) removing one or more of the sequence reads or classification alleles originating from hematopoietic stem cells of the subject from the test sequence information to generate filtered test sequence information; and (c) identifying one or more sequence reads present in the filtered test sequence information that substantially align with reference sequence information obtained from one or more reference subjects, which reference sequence information originated from tumor cells within the reference subjects, thereby indicating that the test sample contains one or more cfDNA fragments originating from the tumor cells within the subject. A computer-readable medium for performing the above steps. (Item 62) A computer-readable medium containing computer-executable non-transitory instructions, the instructions, when executed by at least one electronic processor, at least: (a) generating a subclonality score for each allele in a classification allele set from sequence information comprising sequence reads obtained from cell-free nucleic acid (cfNA) fragments from one or more reference samples, wherein each classification allele is potentially clinically significant and comprises a minor allele observed at a given locus in the reference samples; and (b) comparing the subclonality score with at least one selected cutoff threshold. and wherein classification alleles having a subclonality score above the selected cutoff threshold indicate that the classification alleles are derived from reference cfNA fragments originating from non-target cells, and the classification alleles are added to a non-target nucleic acid variant filter list; and / or Classifying alleles having a subclonality score below the cutoff threshold indicate that the classifying alleles are derived from the reference cfNA fragment originating from the target cell, and the classifying alleles are added to a target nucleic acid variant filter list. Computer-readable medium. (Item 63) A computer-readable medium containing computer-executable non-transitory instructions, the instructions, when executed by at least one electronic processor, at least: (a) identifying a set of classification alleles from sequence information comprising sequence reads obtained from cell-free nucleic acid (cfNA) fragments from one or more reference samples, wherein each classification allele is potentially clinically significant and comprises a minor allele observed at a given locus in said reference samples; (b) determining a minor allele frequency (MAF) value for each classification allele in each of the reference samples from the sequence information; (c) determining a maximum minor allele frequency (maxMAF) value for each of said reference samples; (d) for each classification allele observed in a given reference sample, calculating a ratio of the value of the MAF to the value of the maxMAF for at least a portion of the reference sample to generate a ratio value; (e) calculating, for each of the classification alleles, a ratio of the number of times that the given classification allele in at least some of the reference samples had a ratio value that was less than at least one selected clonality boundary value to the total number of times that the given classification allele appeared in at least some of the reference samples to generate a subclonality score for each of the classification alleles in at least some of the reference samples; and (f) comparing said subclonality score with at least one selected cutoff threshold. and wherein classification alleles with a subclonality score above the selected cutoff threshold indicate that the classification alleles are derived from reference cfNA fragments originating from non-target cells, and the classification alleles are added to a non-target nucleic acid variant filter list; and / or classification alleles with a subclonality score below the cutoff threshold indicate that the classification alleles are derived from reference cfNA fragments originating from target cells, and the classification alleles are added to a target nucleic acid variant filter list. Computer-readable medium. (Item 64) 10. The system or computer-readable medium of any one of the preceding items, wherein the computer-readable medium comprises computer-executable, non-transitory instructions that, when executed by the at least one electronic processor, further perform at least the following: determine, from the sequence information obtained from the reference samples, a value of a MAF for each somatic nucleic acid variant at each locus in a set of target genomic loci that are potentially clinically significant, wherein the set of target genomic loci are identical in each reference sample; and determine a value of maxMAF for each of the reference samples to generate allele information. (Item 65) 10. The system or computer-readable medium of any one of the preceding items, wherein the computer-readable medium comprises computer-executable, non-transitory instructions that, when executed by the at least one electronic processor, at least further generate the non-target nucleic acid variant filter list and / or the target nucleic acid variant filter list using clinical information indexed to the reference sample. (Item 66) 67. The system or computer-readable medium of any one of the preceding items, wherein the computer-readable medium comprises computer-executable, non-transitory instructions that, when executed by the at least one electronic processor, at least further perform the following: detect cfNA fragments originating from the target cells in the subject using clinical information indexed to the test sample. The system or computer-readable medium of any one of the preceding items, wherein the computer-readable medium comprises computer-executable non-transitory instructions that, when executed by the at least one electronic processor, at least further determine a subclonality score using the frequency of each MAF / max-MAF value for each of the classification alleles. (Item 68) 2. The system or computer-readable medium of any one of the preceding items, wherein the computer-readable medium comprises computer-executable non-transitory instructions that, when executed by the at least one electronic processor, further perform at least the following: comparing the subclonality score to a plurality of selected cutoff thresholds, wherein the plurality of selected cutoff thresholds comprises a first cutoff threshold and a second cutoff threshold, the first cutoff threshold being greater than the second cutoff threshold; and wherein classification alleles having a subclonality score greater than the first cutoff threshold indicate that the classification alleles are derived from reference cfNA fragments originating from non-target cells and the classification alleles are added to the non-target nucleic acid variant filter list; and / or wherein classification alleles having a subclonality score less than the second cutoff threshold indicate that the classification alleles are derived from reference cfNA fragments originating from target cells and the classification alleles are added to the target nucleic acid variant filter list. (Item 69) 10. The system or computer-readable medium of any one of the preceding items, wherein the computer-readable medium comprises computer-executable, non-transitory instructions that, when executed by the at least one electronic processor, further perform at least the following: classifying an allelic variant in the test sequence information that substantially matches at least one classification allele on a non-target nucleic acid variant filter list as originating from a target cell when the allelic variant comprises a MAF greater than about 1%. (Item 70) 10. The system or computer-readable medium of any one of the preceding items, wherein the computer-readable medium comprises computer-executable, non-transitory instructions that, when executed by the at least one electronic processor, further perform at least the following: classifying an allelic variant in the test sequence information as originating from a target cell when the allelic variant substantially matches at least one classification allele on a non-target nucleic acid variant filter list and the allelic variant comprises a truncation, an indel, and / or a splice site variant. (Item 71) 10. The system or computer-readable medium of any one of the preceding items, wherein the computer-readable medium comprises computer-executable, non-transitory instructions that, when executed by the at least one electronic processor, at least further determine the frequency of each ratio value for a given classification allele in at least a portion of the reference samples. (Item 72) The system or computer-readable medium of any one of the preceding items, wherein the computer-readable medium comprises computer-executable, non-transitory instructions that, when executed by the at least one electronic processor, at least further determine whether a test sample obtained from a subject contains cfNA fragments originating from the target cell using the target nucleic acid variant filter list. (Item 73) The system or computer-readable medium of any one of the preceding items, wherein the computer-readable medium comprises computer-executable, non-transitory instructions that, when executed by the at least one electronic processor, at least further perform the following: use the non-target nucleic acid variant filter list to determine whether a test sample obtained from a subject contains cfNA fragments originating from the non-target cells.

Claims

[Claim 1] The invention as set forth in the drawings.