Significance modeling of clonal absence of targeted variants
Significance modeling using tumor fraction estimates and likelihood values improves the detection of genetic variants in cfDNA samples, addressing sensitivity limitations and enabling precise cancer treatment decisions without tissue biopsies.
Patent Information
- Application Number
- JP2022545998
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-01-31
- Filing Date
- 2021-01-29
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2041-01-29
AI Technical Summary
Existing methods for detecting genetic variants in cell-free nucleic acid samples, such as ctDNA, have limited sensitivity due to low shedding capacity, impacting the reliability of determining wild-type status for genes like KRAS, NRAS, and BRAF, which is crucial for guiding anti-EGFR therapy in colorectal cancer treatment.
A method for significance modeling using computer-based tumor fraction estimates and likelihood values to determine the clonal absence of genetic variants, incorporating allele frequency and covariate information to generate a quantitative value for negative predictive value, enabling accurate detection of variants like KRAS, NRAS, and BRAF mutations in cfDNA samples.
Enhances the sensitivity and accuracy of detecting the absence of genetic variants, reducing the need for invasive tissue biopsies and guiding precise treatment decisions in cancer therapy.
Smart Images

Figure 0007763764000022 
Figure 0007763764000023 
Figure 0007763764000024
Abstract
Description
[Technical Field]
[0001] cross reference This application claims the benefit of the priority date of U.S. Provisional Patent Application No. 62 / 968,507, filed January 31, 2020, which is incorporated herein by reference in its entirety for all purposes. [Background technology]
[0002] background In advanced colorectal cancer (CRC), guidelines have recommended the use of anti-EGFR therapy only in patients whose tumors are wild-type for KRAS, NRAS, and BRAF. To date, cell-free circulating tumor DNA (ctDNA) testing has been used as a powerful test for positively detecting tumor-derived genomic alterations and microsatellite instability (MSI), with high concordance with tissue sequencing (Gupta et al., Oncologist, 24:1-9 (2019), Parikh et al., Nat Med., 25(9):1415-1421 (2019)). However, the ability to exclude such mutations is limited by the low shedding capacity of ctDNA, which impacts detection sensitivity. Using ctDNA or other nucleic acids to reliably determine the wild-type status of specific genes in tumors would facilitate timely treatment decisions and avoid the need for tissue biopsies to confirm wild-type status. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Gupta et al., Oncologist, 24:1-9 (2019), Parikh et at., Nat Med., 25(9):1415-1421 (2019) Summary of the Invention [Means for solving the problem]
[0004] Therefore, there is a need to identify genetic variants or their absence in order to diagnose and / or guide treatment of diseases that are detectable through genetic analysis, particularly from cell-free nucleic acid (cfNA) samples. Abstract The present disclosure relates to technologies that generate precision diagnoses based on determining the various states of nucleic acids, such as DNA or RNA, from genomes, chromosomes, or other genetic segments sequenced from a sample. Detection of target variants can help guide treatment plans.
[0005] When genetic variants are not detected, it may be equally important to determine whether genetic variants are not detected because the variants are actually absent at the clonal level in the sample (true negative result), or whether genetic variants are actually present at the clonal level but not detected (false negative result).Herein, we describe improvements related to significance modeling of negative prediction, such as whether genetic variants are not detected or whether they are actually absent in the sample.In a specific example, significance modeling can generate and use computer-based tumor fraction (TF) estimates of tumor variants or mutations based on the nucleic acid sequence read data (sequence reads) generated from the sample.
[0006] Alternatively or additionally, significance modeling may determine and use the occurrence and / or diversity of other variants detected or not detected in the sample. For example, significance modeling may use the detection of co-occurring variants with the target variant, or mutually exclusive variants that do not usually co-occur with the target variant. A negative predictive value ("NPV") may be generated based on the TF estimates and / or diversity of variants detected or not detected in the sample. The results may be used to provide a confidence level of a negative diagnosis (e.g., the absence of a given variant at the locus of interest) and / or further guide a treatment plan based on a negative diagnosis. For example, in the context of cancer diagnosis, co-occurring variants may include driver variants that tend to promote tumorigenesis, and mutually exclusive variants may include tumor suppressor variants that tend to suppress tumorigenesis.
[0007] In one aspect, the present disclosure provides a method for determining the probability that a first variant of interest is absent at a first locus at a clonal level in a nucleic acid sample obtained from a subject.The method includes: accessing a plurality of sequence read data of the nucleic acid in the sample; and determining that the first variant has not been detected at the first locus in the sample based on the plurality of sequence read data.The method also includes: generating a first likelihood value based on the probability that the first variant is absent at a clonal level, and a second likelihood value based on the probability that the first variant is present at a clonal level; determining a quantitative value based on the first likelihood value and the second likelihood value; comparing the quantitative value with a threshold; and determining that the first variant of interest is absent at a clonal level based on the comparison.
[0008] In one aspect, the present disclosure provides a method for determining that a first variant of interest is absent (and a negative prediction) at a first locus in a cell-free nucleic acid (cfNA) sample from a human subject at the clonal level. The method includes accessing a plurality of sequence reads of the cfNA sample; and determining, based on the plurality of sequence reads, that the first variant has not been detected at the first locus in the sample. The method also includes generating a first likelihood value based on the probability that the first variant is absent at the clonal level and / or a second likelihood value based on the probability that the first variant is present at the clonal level; and classifying, based on the comparison, that the first variant of interest is absent at the first locus at the clonal level.
[0009] In one aspect, the present disclosure provides a method for determining that a first variant of interest is absent at a first locus in a cell-free deoxyribonucleic acid (cfDNA) sample of a human subject at a clonal level (and a negative prediction).The method includes: accessing a plurality of sequence read data of the cfDNA sample; and determining that the first variant has not been detected at the first locus in the sample based on the plurality of sequence read data.The method also includes: generating a first likelihood value based on the probability that the first variant is absent at a clonal level, and / or a second likelihood value based on the probability that the first variant is present at a clonal level; optionally, determining a quantitative value based on the first likelihood value and / or the second likelihood value; comparing the quantitative value and / or the first likelihood value and / or the second likelihood value with a threshold; and determining (for example, classifying or calling in this context) that the first variant of interest is absent at a clonal level based on the comparison.
[0010] In some embodiments, generating the first and second likelihood values comprises determining a tumor fraction estimate for the sample, wherein the first and second likelihood values are based on the tumor fraction estimate. In certain embodiments, determining the tumor fraction estimate comprises determining a maximum mutant allele frequency (MAX MAF) of a tumor mutation in the sample. In some of these embodiments, determining the MAX MAF comprises determining a molecular count associated with the tumor mutation based on multiple sequence read data. In certain embodiments, generating the first and second likelihood values comprises determining an allele frequency of at least a second variant, wherein the first and second likelihood values are further based on the allele frequency and the MAX MAF. In certain of these embodiments, the method further comprises comparing the allele frequency to a second threshold value based on the MAX MAF, wherein determining that the first variant of interest is absent at the first locus at the clonal level is further based on comparing the MAF to the second threshold. In certain of these embodiments, determining the allele frequency comprises determining a first molecular count associated with the first variant based on the plurality of sequence read data. In some embodiments, determining the quantitative value comprises accessing covariate information indicating the historical prevalence of one or more variants that co-occur and / or are mutually exclusive with the first variant, wherein the quantitative value is based on the covariate information. In some of these embodiments, the method further comprises determining the occurrence of at least a second variant in the cfDNA sample, wherein the quantitative value is further based on the covariate information.
[0011] In certain embodiments, determining the quantitative value includes accessing covariate information indicating the historical prevalence of one or more variants that co-occur and / or mutually exclusive with the first variant, wherein the quantitative value is based on the covariate information. In some of these embodiments, the method further includes determining the occurrence of at least a second variant in the cfDNA sample, wherein the quantitative value is further based on the occurrence of the second variant. In certain embodiments, the quantitative value is based on the ratio of the first likelihood value to the second likelihood value. In certain embodiments, the method further includes determining a confidence level that the first variant is absent at the clonal level in the cfDNA sample based on the quantitative value. In some embodiments, the method further includes determining the generation of a treatment plan for treating a disease in a human subject. In some of these embodiments, the disease is cancer. In certain embodiments, the method further includes determining the occurrence of at least a second variant in the cfDNA sample and adjusting the quantitative value based on the occurrence of at least a second variant in the cfDNA sample.
[0012] In another aspect, the present disclosure provides a method for determining, at least in part by computer, that a first target nucleic acid variant is absent from a first locus in a cell-free nucleic acid (cfNA) sample obtained from a subject with a given type of cancer. The method includes the steps of: determining, by a computer, that the first target nucleic acid variant is not detected at the first locus in the cfNA sample; determining, by a computer, a coverage of the first locus from sequence information generated from the cfNA sample; and determining, by a computer, a tumor fraction from the sequence information generated from the cfNA sample. The method also includes the steps of determining, by a computer, a probability that the first target nucleic acid variant is present at the first locus in the cfNA sample from the coverage and tumor fraction to generate a quantitative value; and determining (e.g., classifying or calling in this context) that the first target nucleic acid variant is absent from the first locus in the cfNA sample if the quantitative value is different from a threshold value.
[0013] In another aspect, the present disclosure provides a method for determining, at least in part, using a computer, that a first target nucleic acid variant is absent from a first locus in a cell-free nucleic acid (cfNA) sample obtained from a subject. The method includes: determining that the first target nucleic acid variant is not detected in the cfNA sample obtained from the subject to generate a first test result; determining that at least a second target nucleic acid variant is detected in the cfNA sample obtained from the subject to generate a second test result; and determining, by the computer, a first probability that the first target nucleic acid variant is absent from the cfNA sample given the second test result and / or a second probability that the first target nucleic acid is present in the cfNA sample given the second test result. The method also includes generating, by the computer, a quantitative value using the first probability, the second probability, and / or a ratio thereof; and determining (e.g., classifying or calling in this context) that the first target nucleic acid variant is absent from the first locus in the cfNA sample if the quantitative value differs from a threshold value.
[0014] In another aspect, the present disclosure provides a method for determining, at least in part, using a computer, that a first target nucleic acid variant is absent from a first locus in a cell-free nucleic acid (cfNA) sample obtained from a subject with a given type of cancer. The method includes the steps of determining that the first target nucleic acid variant is not detected in the cfNA sample obtained from the subject; generating at least one tumor proportion-based value by a computer; generating at least one mutual exclusivity value by a computer; and determining (e.g., classifying or calling in this context) that the first target nucleic acid variant is absent from the first locus in the cfNA sample using the tumor proportion-based value and / or the mutual exclusivity value.
[0015] In some embodiments, the quantitative value is less than a threshold value, while in other embodiments, the quantitative value is greater than a threshold value. In certain embodiments, the quantitative value comprises a log-likelihood ratio (LLR) threshold value. Typically, the first and second test results are dependent on each other. In certain embodiments, the method disclosed herein comprises determining that a plurality of other selected target nucleic acid variants are absent from one or more other loci (e.g., a panel of selected or target loci).
[0016] In certain embodiments, the method includes determining that a first target nucleic acid variant is absent from a first locus in multiple reference cfNA samples to generate a threshold. In some of these embodiments, the threshold includes a clonality or subclonality threshold. In some embodiments of the methods disclosed herein, the first target nucleic acid variant includes a driver mutation. In certain embodiments, the method further includes administering one or more therapies to the subject based on determining that the first target nucleic acid variant is absent from the first locus in the cfNA sample. In some embodiments, the method includes estimating the probability of detecting the first target nucleic acid variant at the first locus in the cfNA sample using tumor fraction and a binomial model. In certain of these embodiments, the binomial model includes information about a predetermined cancer type and / or a second target nucleic acid variant. Other models may also be used as needed.
[0017] In some embodiments of the method disclosed herein, determining that the first target nucleic acid variant does not exist at the first locus in cfNA sample indicates that the first locus is wild-type.In certain embodiments, the predetermined cancer type is colorectal cancer, the first locus is KRAS, BRAF or NRAS, and determining that the first target nucleic acid variant does not exist at the first locus in cfNA sample indicates that the first locus is wild-type KRAS, BRAF or NRAS.In certain of these embodiments, the method further comprises administering cetuximab and / or panitumumab to the subject.In some embodiments, cfNA comprises cfDNA and / or cfRNA.
[0018] In certain embodiments, the methods disclosed herein further include repeating the method one or more times to monitor whether the first target nucleic acid variant is absent at the first locus in different cfNA samples obtained from the subject at different time points. In certain embodiments, the method further includes performing one or more additional tests to confirm or refute a determination that the first target nucleic acid variant is absent at the first locus in the cfNA sample. In some embodiments, the method includes determining a maximum mutant allele frequency (MAX MAF) for the cfNA sample and using the MAX MAF as a tumor fraction estimate. In certain embodiments, the method includes determining that the first target nucleic acid variant is not detected at the first locus in the cfNA sample based on multiple sequencing read data obtained from the cfNA sample. In some embodiments, the method includes determining that the first target nucleic acid variant is absent at the clonal level in the cfNA sample. In certain embodiments, the method includes generating a first likelihood value based on a first probability and a second likelihood value based on a second probability. In certain embodiments, the method includes determining a quantitative value based on the first likelihood value and the second likelihood value.
[0019] In some embodiments of the methods disclosed herein, generating the first and second likelihood values includes determining a tumor fraction estimate for the cfNA sample, where the first and second likelihood values are based on the tumor fraction estimate. In certain embodiments, the method includes determining the tumor fraction estimate, where the determining step includes determining a maximum mutant allele frequency (MAX MAF) of a tumor mutation in the cfNA sample. In certain embodiments, the method includes determining the MAX MAF, where the determining step includes determining a molecular count associated with the tumor mutation based on multiple sequence read data. In some embodiments, the method includes generating the first and second likelihood values, where the generating step includes determining an allele frequency of at least a second variant, where the first and second likelihood values are further based on the allele frequency and the MAX MAF. In some of these embodiments, the method further comprises comparing the allele frequency to a second threshold value based on the MAX MAF, wherein determining that the first target nucleic acid variant of interest is absent at the first locus at the clonal level is further based on comparing the MAF to the second threshold value.
[0020] In some embodiments, determining the first allele frequency comprises determining a first molecular count associated with the first target nucleic acid variant based on the plurality of sequence read data. In certain embodiments, determining the quantitative value comprises accessing covariate information indicating the historical prevalence of one or more variants that co-occur and / or mutually exclusive with the first variant, wherein the quantitative value is based on the covariate information. In some embodiments, the method further comprises determining the occurrence of at least a second target nucleic acid variant in the cfDNA sample, wherein the quantitative value is further based on the covariate information. In certain embodiments, the method further comprises determining the occurrence of at least a second target nucleic acid variant in the cfNA sample, wherein the quantitative value is further based on the occurrence of the second target nucleic acid variant. In some of these embodiments, the quantitative value is based on the ratio of the first likelihood value to the second likelihood value. In certain of these embodiments, the method further comprises determining a confidence level that the first target nucleic acid variant is absent at a clonal level in the cfNA sample based on the quantification value. In some of these embodiments, the method further comprises determining the occurrence of at least a second target nucleic acid variant in the cfNA sample, and adjusting the quantification value based on the occurrence of at least the second target nucleic acid variant in the cfNA sample.
[0021] In some embodiments of the methods disclosed herein, the ratio comprises a log-posterior probability ratio (LPPR), which is equal to the sum of the log-likelihood tumor fraction value, the log-likelihood mutual exclusivity value, and the log-prior value. In certain embodiments, the first locus or the second locus comprises a second target nucleic acid variant. In certain embodiments, the quantitative value comprises a negative predictive value (NPV) score. In some embodiments, the predetermined cancer type comprises lung cancer, and the first target nucleic acid variant is a mutation in a gene selected from the group consisting of EGFR, BRAF (e.g., V600E), ALK (e.g., fusion), ROS1 (e.g., fusion), and MET. In some embodiments, the predetermined cancer type comprises colorectal cancer, and the first target nucleic acid variant is a mutation in a gene selected from the group consisting of KRAS (e.g., G12X, G13X, Q61X, K117N, A146P / 146T / 146V), BRAF, and NRAS.
[0022] In another aspect, the present disclosure provides a system comprising a controller comprising or capable of accessing a computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least the following steps: accessing a plurality of sequence reads of a cfDNA sample; determining, based on the plurality of sequence reads, that a first variant was not detected at a first locus in the sample; generating a first likelihood value based on the probability that the first variant is absent at a clonal level and a second likelihood value based on the probability that the first variant is present at a clonal level; determining a quantification value based on the first likelihood value and the second likelihood value; comparing the quantification value to a threshold value; and determining, based on the comparison, that a first variant of interest is absent at a clonal level (e.g., classifying or calling, in this context).
[0023] In another aspect, the present disclosure provides a system including a controller that includes or is capable of accessing a computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least the steps of: accessing sequence information generated from a cell-free nucleic acid (cfNA) sample obtained from a subject having a predetermined type of cancer; determining from the sequence information that a first target nucleic acid variant is not detected at a first locus in the cfNA sample; determining from the sequence information a coverage of the first locus; determining a tumor fraction from the sequence information; determining from the coverage and the tumor fraction a probability that the first target nucleic acid variant is present at the first locus in the cfNA sample to generate a quantitative value; and, if the quantitative value differs from a threshold value, determining (e.g., classifying or calling in this context) that the first target nucleic acid variant is not present at the first locus in the cfNA sample.
[0024] In another aspect, the present disclosure provides a system comprising a controller comprising or capable of accessing a computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least the following steps: accessing sequence information generated from a cell-free nucleic acid (cfNA) sample obtained from the subject; determining from the sequence information that a first target nucleic acid variant is not detected in the cfNA sample to generate a first test result; determining from the sequence information that at least a second target nucleic acid variant is detected in the cfNA sample to generate a second test result; determining a first probability that the first target nucleic acid variant is not present in the cfNA sample given the second test result and / or a second probability that the first target nucleic acid is present in the cfNA sample given the second test result; generating a quantitative value from the first probability, the second probability, and / or a ratio thereof; and determining (e.g., classifying or calling in this context) that the first target nucleic acid variant is not present at the first locus in the cfNA sample if the quantitative value differs from a threshold value.
[0025] In another aspect, the present disclosure provides a system comprising a controller comprising or capable of accessing a computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least the steps of: accessing sequence information generated from a cell-free nucleic acid (cfNA) sample obtained from a subject; determining from the sequence information that a first target nucleic acid variant is not detected in the cfNA sample; generating at least one tumor fraction-based value; generating at least one mutual exclusivity value; and determining (e.g., classifying or calling in this context) that a first target nucleic acid variant is not present at a first locus in the cfNA sample using the tumor fraction-based value and / or the mutual exclusivity value.
[0026] In another aspect, the disclosure provides a computer-readable medium comprising non-transitory computer-executable instructions that, when executed by an electronic processor, perform at least the steps of accessing a plurality of sequence reads of a cfDNA sample; determining, based on the plurality of sequence reads, that a first variant was not detected at a first locus in the sample; generating a first likelihood value based on the probability that the first variant is absent at a clonal level and a second likelihood value based on the probability that the first variant is present at a clonal level; determining a quantification value based on the first likelihood value and the second likelihood value; comparing the quantification value to a threshold value; and determining (e.g., classifying or calling in this context) that a first variant of interest is absent at a clonal level at the first locus based on the comparison.
[0027] In another aspect, the disclosure provides a computer-readable medium comprising non-transitory computer-executable instructions that, when executed by an electronic processor, perform at least the steps of: accessing sequence information generated from a cell-free nucleic acid (cfNA) sample obtained from a subject having a predetermined type of cancer; determining from the sequence information that a first target nucleic acid variant is not detected at a first locus in the cfNA sample; determining from the sequence information a coverage of the first locus; determining from the sequence information a tumor fraction; determining from the coverage and the tumor fraction a probability that the first target nucleic acid variant is present at the first locus in the cfNA sample to generate a quantitative value; and determining (e.g., classifying or calling in this context) that the first target nucleic acid variant is not present at the first locus in the cfNA sample if the quantitative value differs from a threshold value.
[0028] In another aspect, the disclosure provides a computer-readable medium comprising non-transitory computer-executable instructions that, when executed by an electronic processor, perform at least the following steps: accessing sequence information generated from a cell-free nucleic acid (cfNA) sample obtained from the subject; determining from the sequence information that a first target nucleic acid variant is not detected in the cfNA sample to generate a first test result; determining from the sequence information that at least a second target nucleic acid variant is detected in the cfNA sample to generate a second test result; determining a first probability that the first target nucleic acid variant is not present in the cfNA sample given the second test result and / or a second probability that the first target nucleic acid is present in the cfNA sample given the second test result; generating a quantitative value using the first probability, the second probability, and / or a ratio thereof; and determining (e.g., classifying or calling in this context) that the first target nucleic acid variant is not present at the first locus in the cfNA sample if the quantitative value differs from a threshold value.
[0029] In another aspect, the present disclosure provides a computer-readable medium comprising non-transitory computer-executable instructions that, when executed by an electronic processor, perform at least the steps of: accessing sequence information generated from a cell-free nucleic acid (cfNA) sample obtained from the subject; determining from the sequence information that a first target nucleic acid variant is not detected in the cfNA sample; generating at least one tumor fraction-based value; generating at least one mutual exclusivity value; and determining (e.g., classifying or calling in this context) that the first target nucleic acid variant is not present at a first locus in the cfNA sample using the tumor fraction-based value and / or the mutual exclusivity value.
[0030] In some embodiments of the systems or computer-readable media disclosed herein, the quantification value is less than the threshold value, while in other exemplary embodiments, the quantification value is greater than the threshold value. In some of these embodiments, the first and second test results are dependent on each other. In certain of these embodiments, the non-transitory computer-executable instructions include determining that a plurality of other selected target nucleic acid variants are absent from one or more other loci. In some of these embodiments, the quantification value comprises a log-likelihood ratio (LLR) threshold. In certain of these embodiments, the non-transitory computer-executable instructions include determining that a first target nucleic acid variant is absent from the first locus in a plurality of reference cfNA samples to generate a threshold. In some of these embodiments, the threshold comprises a clonality or subclonality threshold. In some of these embodiments, the first target nucleic acid variant comprises a driver mutation. In some of these embodiments, the instructions further perform at least the step of outputting one or more therapy recommendations for the subject based on determining that the first target nucleic acid variant is absent from the first locus in the cfNA sample.
[0031] In some embodiments of the system or computer-readable medium disclosed herein, the instructions further perform at least the steps of: estimating the probability of detecting a first target nucleic acid variant at a first locus in a cfNA sample using a tumor fraction and a binomial model. In some of these embodiments, the instructions further perform at least the steps of: determining a maximum mutant allele frequency (MAX MAF) for the cfNA sample and using the MAX MAF as a tumor fraction estimate. In some of these embodiments, the instructions further perform at least the steps of: determining that the first target nucleic acid variant is absent at the clonal level in the cfNA sample. In certain of these embodiments, the instructions further perform at least the steps of: generating a first likelihood value based on a first probability and a second likelihood value based on a second probability. In certain of these embodiments, the instructions further perform at least the steps of: determining a quantitative value based on the first likelihood value and the second likelihood value.
[0032] In some embodiments of the system or computer-readable medium disclosed herein, the instructions further perform at least the steps of: generating a first likelihood value and a second likelihood value by determining a tumor fraction estimate for the cfNA sample, wherein the first likelihood value and the second likelihood value are based on the tumor fraction estimate. In certain of these embodiments, the instructions further perform at least the steps of determining the tumor fraction estimate by determining a maximum mutant allele frequency (MAX MAF) of a tumor mutation in the cfNA sample. In certain of these embodiments, the instructions further perform at least the steps of determining the MAX MAF by determining a molecular count associated with the tumor mutation based on the plurality of sequence read data. In certain of these embodiments, the instructions further perform at least the steps of generating the first likelihood value and the second likelihood value by determining the allele frequency of at least a second variant, wherein the first likelihood value and the second likelihood value are further based on the allele frequency and the MAX MAF. In some of these embodiments, the instructions further perform at least the steps of: comparing the allele frequency to a second threshold based on the MAX MAF, and determining that the first target nucleic acid variant of interest is absent at the first locus at a clonal level further based on comparing the MAF to the second threshold. In some of these embodiments, the instructions further perform at least the steps of determining the allele frequency by determining a first molecule count associated with the first target nucleic acid variant based on the plurality of sequence read data.
[0033] In some embodiments of the system or computer-readable medium disclosed herein, the instructions further perform at least the following steps: determining a quantitative value by accessing covariate information indicating the historical prevalence of one or more variants that co-occur and / or mutually exclusive with the first variant, wherein the quantitative value is based on the covariate information. In some of these embodiments, the instructions further perform at least the following steps: determining the occurrence of at least a second target nucleic acid variant in the cfDNA sample, wherein the quantitative value is further based on the covariate information. In some of these embodiments, the instructions further perform at least the following steps: determining a quantitative value by accessing covariate information indicating the historical prevalence of one or more variants that co-occur and / or mutually exclusive with the first target nucleic acid variant, wherein the quantitative value is further based on the covariate information. In certain of these embodiments, the instructions further perform at least the following steps: determining the occurrence of at least a second target nucleic acid variant in the cfNA sample, wherein the quantitative value is further based on the occurrence of the second target nucleic acid variant. In certain of these embodiments, the instructions further perform at least the steps of: determining a confidence level that the first target nucleic acid variant is absent at a clonal level in the cfNA sample based on the quantitative value; and adjusting the quantitative value based on the occurrence of at least the second target nucleic acid variant in the cfNA sample. In certain of these embodiments, the ratio comprises a log-posterior probability ratio (LPPR), which is equal to the sum of the log-likelihood tumor fraction value, the log-likelihood mutual exclusivity value, and the log-prior value.
[0034] In some embodiments, the results of the system and method disclosed herein are used as input for generating a report.The report can be in paper or electronic format.For example, when obtained by the method and system disclosed herein, such report can directly display the classification that the first variant of interest is not present at the first locus at the clonal level.Alternatively or additionally, the report can include diagnostic information or therapeutic recommendations based on the probability that the first variant of interest is not present at the first locus at the clonal level.
[0035] If the determination is based on a quantitative value different from the threshold, the quantitative value used in the determination may be less than the threshold or greater than the threshold, depending on the nature of the threshold. Thus, the quantitative value may or may not meet the threshold.
[0036] In certain aspects, the present disclosure provides methods of treating a disease in a subject, the method comprising: accessing a plurality of sequence reads of a cell-free deoxyribonucleic acid (cfDNA) sample obtained from the subject; determining, based on the plurality of sequence reads, that a first variant of interest was not detected at a first locus in the cfDNA sample; generating a first likelihood value based on a probability that the first variant is not present at a clonal level and / or a second likelihood value based on the probability that the first variant is present at a clonal level; determining a quantification value based on the first likelihood value and / or the second likelihood value; comparing the quantification value and / or the first likelihood value and / or the second likelihood value to a threshold; determining, based on the comparison, that the first variant of interest is not present at the first locus at a clonal level; and administering one or more therapies to the subject based at least in part on a determination that the first variant of interest is not present at the first locus at a clonal level, thereby treating the disease in the subject. In certain embodiments, one or more therapies are discontinued from being administered to the subject based at least in part on the determination that the first variant of interest is not present at the first locus at the clonal level, thereby treating the disease in the subject.In certain embodiments, the methods described herein are carried out on a plurality of subjects.In certain embodiments, a subset of subjects are administered one or more therapies based at least in part on the determination that the first variant of interest is not present at the first locus at the clonal level, and another subset of subjects are discontinued from the one or more therapies previously administered to these subjects.In certain embodiments, the subject is administered a therapy that is different from the therapy previously administered to the subject based at least in part on the determination that the first variant of interest is not present at the first locus at the clonal level.
[0037] In certain aspects, the present disclosure provides methods of treating disease in a subject, the method comprising administering or discontinuing administration of one or more therapies to the subject based at least in part on a determination that a first variant of interest is not clonally present at a first locus in a cell-free deoxyribonucleic acid (cfDNA) sample obtained from the subject, wherein the determination is obtained by: accessing a plurality of sequence reads of the cfDNA sample; determining, based on the plurality of sequence reads, that the first variant was not detected at the first locus in the sample; generating a first likelihood value based on a probability that the first variant is not clonally present and / or a second likelihood value based on a probability that the first variant is present at a clonally present level; determining a quantification value based on the first likelihood value and / or the second likelihood value; comparing the quantification value and / or the first likelihood value and / or the second likelihood value to a threshold; and determining, based on the comparison, that the first variant of interest is not clonally present at the first locus.
[0038] In certain aspects, the disclosure provides methods of treating cancer in a subject, the method comprising: determining that a first target nucleic acid variant is not detectable at a first locus in a cell-free nucleic acid (cfNA) sample obtained from a subject with cancer; determining coverage of the first locus from sequence information generated from the cfNA sample; determining a tumor fraction from the sequence information generated from the cfNA sample; determining a probability that the first target nucleic acid variant is present at the first locus in the cfNA sample from the coverage and tumor fraction to generate a quantitative value; determining, if the quantitative value differs from a threshold, that the first target nucleic acid variant is not present at the first locus in the cfNA sample; and administering or discontinuing administration of one or more therapies to the subject based at least in part on a determination that the first target nucleic acid variant is not present at the first locus in the cfNA sample, thereby treating the cancer in the subject.
[0039] In certain aspects, the present disclosure provides methods of treating cancer in a subject, the method comprising administering or discontinuing administration of one or more therapies to the subject based at least in part on a determination that a first target nucleic acid variant is not present at a first locus in a cell-free deoxyribonucleic acid (cfDNA) sample obtained from the subject with cancer, where the determination is obtained by the following steps: determining that the first target nucleic acid variant is not detected at the first locus in the cfNA sample; determining a coverage of the first locus from sequence information generated from the cfNA sample; determining a tumor fraction from the sequence information generated from the cfNA sample; determining a probability that the first target nucleic acid variant is present at the first locus in the cfNA sample from the coverage and the tumor fraction to generate a quantitative value; and determining that the first target nucleic acid variant is not present at the first locus in the cfNA sample if the quantitative value is different from a threshold value.
[0040] In certain aspects, the present disclosure provides methods of treating a disease in a subject, the method comprising: determining that a first target nucleic acid variant is not detected in a cell-free nucleic acid (cfNA) sample obtained from the subject to generate a first test result; determining that at least a second target nucleic acid variant is detected in the cfNA sample obtained from the subject to generate a second test result; determining a first probability that the first target nucleic acid variant is not present in the cfNA sample given the second test result, and / or a second probability that the first target nucleic acid is present in the cfNA sample given the second test result; generating a quantitative value using the first probability, the second probability, and / or a ratio thereof; determining, if the quantitative value differs from a threshold, that the first target nucleic acid variant is not present at the first locus in the cfNA sample; and administering or discontinuing administration of one or more therapies to the subject based at least in part on the determination that the first target nucleic acid variant is not present at the first locus, thereby treating the disease in the subject.
[0041] In certain aspects, the present disclosure provides methods of treating a disease in a subject, the method comprising administering or discontinuing administration of one or more therapies to the subject based at least in part on a determination that a first target nucleic acid variant is not present at a first locus in a cell-free nucleic acid (cfNA) sample obtained from the subject, the determination being obtained by: determining that the first target nucleic acid variant is not detected in the cfNA sample obtained from the subject to generate a first test result; determining that at least a second target nucleic acid variant is detected in the cfNA sample obtained from the subject to generate a second test result; determining a first probability that the first target nucleic acid variant is not present in the cfNA sample given the second test result and / or a second probability that the first target nucleic acid is present in the cfNA sample given the second test result; generating a quantitative value using the first probability, the second probability, and / or a ratio thereof; and determining that the first target nucleic acid variant is not present at the first locus in the cfNA sample if the quantitative value differs from a threshold value.
[0042] In certain aspects, the present disclosure provides methods of treating cancer in a subject, the method comprising: determining that a first target nucleic acid variant is absent from a first locus in a cell-free nucleic acid (cfNA) sample obtained from a subject having a predetermined type of cancer; generating at least one tumor fraction-based value; generating at least one mutual exclusivity value; determining using the tumor fraction value and / or the mutual exclusivity value that the first target nucleic acid variant is absent from the first locus in the cfNA sample; and administering or discontinuing administration of one or more therapies to the subject based at least in part on a determination that the first target nucleic acid variant is absent from the first locus in the cfNA sample, thereby treating the cancer in the subject.
[0043] In certain aspects, the present disclosure provides methods of treating cancer in a subject, comprising administering or discontinuing administration of one or more therapies to the subject based at least in part on a determination that a first target nucleic acid variant is not present at a first locus in a cell-free nucleic acid (cfNA) sample obtained from a subject having a given type of cancer, wherein the determination is obtained by determining that the first target nucleic acid variant is not detectable in the cfNA sample obtained from the subject; generating at least one tumor fraction-based value; generating at least one mutual exclusivity value; and determining that the first target nucleic acid variant is not present at the first locus in the cfNA sample using the tumor fraction-based value and / or the mutual exclusivity value.
[0044] The various steps of the disclosed methods, or steps performed by the systems disclosed herein, may be performed at the same or different times, in the same or different geographic locations, e.g., countries, and / or by the same or different people. [Brief explanation of the drawings]
[0045] [Figure 1] FIG. 1 illustrates an example of a system for generating a negative prediction of a target variant in a subject's sample according to an embodiment of the present disclosure.
[0046] [Figure 2] FIG. 2 illustrates a schematic diagram of the inputs and outputs of a negative prediction analyzer according to one embodiment.
[0047] [Figure 3] FIG. 3 illustrates an example of a method for generating a negative prediction of a target variant in a subject's sample according to one embodiment of the present disclosure.
[0048] [Figure 4]FIG. 4A illustrates a graph of a test hypothesis that the target variant (the target variant) is absent in the sample (or present in a subclonal MAF), according to one embodiment.
[0049] FIG. 4B illustrates a graph of the null hypothesis that the target variant is present in the sample according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0050] definition In order to more readily understand this disclosure, certain terms are first defined below. Additional definitions for these and other terms may be found throughout this specification. In the event that the definitions of terms set forth below are inconsistent with definitions in patent applications or issued patents incorporated herein by reference, the definitions set forth in this application should be used to understand the meaning of the terms.
[0051] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include the plural unless the context clearly dictates otherwise. Thus, for example, reference to "a method" includes one or more methods and / or steps of the type described herein and / or that will become apparent to those of ordinary skill in the art upon reading this disclosure. There is an implicit "about" before temperatures, concentrations, times, numbers of bases or base pairs, coverage, etc. discussed in this disclosure, and it is also recognized that minor and insubstantial equivalents are within the scope of this disclosure. In this application, the use of the singular includes the plural unless specifically stated otherwise. Similarly, the use of "comprises," "comprising," "containing," "contains," "containing," "include," "includes," and "including" is intended to be non-limiting.
[0052] It is also understood that the terminology used herein is intended for the purpose of describing particular embodiments only and is not intended to be limiting. Furthermore, unless otherwise defined, all technical chemical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. When describing and claiming the methods, computer-readable media, and systems, the following terms and grammatical variations thereof will be used in accordance with the definitions set forth below.
[0053] About: As used herein, "about" or "approximately," when applied to one or more values or elements of interest, refers to a value or element that is similar to a stated reference value or element. In certain embodiments, the term "about" or "approximately" refers to a range of values or elements that are within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less in either direction (greater or less) of the stated reference value or element, unless otherwise stated or apparent from the context (except where such number exceeds 100% of the possible values or elements).
[0054] Adapter: As used herein, "adapter" refers to a short nucleic acid (e.g., less than about 500 nucleotides, less than about 100 nucleotides, or less than about 50 nucleotides in length) that is typically at least partially double-stranded and used to ligate either or both ends of a given sample nucleic acid molecule. An adapter can include nucleic acid primer binding sites at both ends to enable amplification of the nucleic acid molecule flanked by the adapter, and / or sequencing primer binding sites, including primer binding sites for sequencing applications, such as various next-generation sequencing (NGS) applications. An adapter can also include a binding site for a capture probe, such as an oligonucleotide bound to a flow cell support or other device. An adapter can also include a nucleic acid tag, as described herein. The nucleic acid tag is typically positioned relative to the amplification primer and sequencing primer binding sites so that the nucleic acid tag is included in the amplicon and sequencing read data of a given nucleic acid molecule. Adapters of the same or different sequences can be ligated to each end of a nucleic acid molecule. In certain embodiments, adapters of the same sequence are ligated to each end of a nucleic acid molecule, except that the sequences of the nucleic acid tags differ. In some embodiments, the adapter is a Y-shaped adapter, which has one end blunt or tailed as described herein for binding to nucleic acid molecules that are also blunt or tailed with one or more complementary nucleotides. In yet other exemplary embodiments, the adapter is a bell-shaped adapter that includes a blunt or tailed end for binding to the nucleic acid molecule to be analyzed. Other exemplary adapters include T-tail and C-tail adapters.
[0055] Administer: As used herein, "administering" or "administering" a therapeutic agent (e.g., an immunotherapeutic agent) to a subject means giving, applying, or contacting the composition with the subject. Administration can be accomplished by any of several routes, including, for example, topical, oral, subcutaneous, intramuscular, intraperitoneal, intravenous, intrathecal, and intradermal.
[0056] Allele: As used herein, "allele" or "allelic variant" refers to a specific genetic variant at a defined genomic location or locus. Allelic variants typically occur at a frequency of 50% (0.5) or 100%, depending on whether the allele is heterozygous or homozygous. For example, germline variants are inherited and typically have a frequency of 0.5 or 1. However, somatic variants are acquired variants and typically have a frequency of <0.5. The major and minor alleles of a locus refer to nucleic acids in which the locus is occupied by a nucleotide of a reference sequence and a variant nucleotide that differs from the reference sequence, respectively. Measurements at a locus can take the form of an allelic fraction (AF), which measures the frequency with which an allele is observed in a sample.
[0057] Amplify: As used herein, "amplify" or "amplification" in the context of nucleic acids refers to the production of multiple copies of a polynucleotide or portion of a polynucleotide, typically starting from a small amount of the polynucleotide (e.g., a single polynucleotide molecule), where the amplification product or amplicon is generally detectable. Polynucleotide amplification encompasses a variety of chemical and enzymatic processes.
[0058] Barcode: As used herein, "barcode" in the context of nucleic acids refers to a nucleic acid molecule having a sequence that can act as a molecular identifier. For example, individual "barcode" sequences are typically added to each DNA fragment during next-generation sequencing (NGS) library preparation so that each read can be identified and sorted before final data analysis.
[0059] Cancer type: As used herein, "cancer," "cancer type," or "tumor type" refers to a type or subtype of cancer, as defined, for example, by histopathology. Cancer type can be determined by any conventional criteria, for example, by occurrence in a given tissue (e.g., blood cancer, central nervous system (CNS), brain cancer, lung cancer (small cell and non-small cell), skin cancer, nose cancer, throat cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, intestinal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, breast cancer, prostate cancer, ovarian cancer, lung cancer, intestinal cancer, soft tissue cancer, neuroendocrine cancer, gastroesophageal cancer, head and neck cancer, gynecological cancer, colorectal cancer, urinary cancer, etc.). Cancers can be defined based on cancer type (e.g., urothelial carcinoma, solid carcinoma, heterogeneous carcinoma, homogeneous carcinoma), unknown primary origin, and other, and / or cancers of the same cell lineage (e.g., carcinoma, sarcoma, lymphoma, cholangiocarcinoma, leukemia, mesothelioma, melanoma, or glioblastoma), and / or cancers that display cancer markers such as Her2, CA15-3, CA19-9, CA-125, CEA, AFP, PSA, HCG, KRAS, BRAF, NRAS, hormone receptors, and NMP-22. Cancers can also be classified by stage (e.g., stage 1, 2, 3, or 4) and whether they are primary or secondary in origin.
[0060] Cell-free nucleic acid: As used herein, "cell-free nucleic acid" refers to nucleic acid that is not contained within or otherwise associated with a cell. Cell-free nucleic acid can include all unencapsulated nucleic acids, for example, originating from bodily fluids (e.g., blood, plasma, serum, urine, cerebrospinal fluid (CSF), etc.) from a subject. Cell-free nucleic acid includes DNA (cfDNA), RNA (cfRNA), and hybrids thereof, including, for example, genomic DNA, mitochondrial DNA, circulating DNA, siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, small nuclear RNA (snoRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), and / or any fragment thereof. Cell-free nucleic acid can be double-stranded, single-stranded, or a hybrid thereof. Cell-free nucleic acid can be released into bodily fluids through secretion or cell death processes, such as cell necrosis, apoptosis, or other processes. Some cell-free nucleic acids, such as circulating tumor DNA (ctDNA), are released into bodily fluids from cancer cells. Others are released from healthy cells. ctDNA can be non-encapsulated tumor-derived fragmented DNA.Another example of cell-free nucleic acid is fetal DNA that circulates freely in the mother's bloodstream, also called cell-free fetal DNA (cffDNA).Cell-free nucleic acid can have one or more epigenetic modifications, for example, cell-free nucleic acid can be acetylated, 5-methylated, ubiquitinated, phosphorylated, sumoylated, ribosylated, and / or citrullinated.
[0061] Clonality: As used herein, "clonal" refers to a population of nucleic acids that contain nucleotide sequences that are substantially or completely identical to each other (e.g., target variants) at least at a given locus of interest.
[0062] Confidence interval: As used herein, "confidence interval" or "confidence level" means a range of values defined such that there is a specified probability that the value of a given parameter falls within that range of values.
[0063] Copy number variant: As used herein, "copy number variant," "CNV," or "copy number variation" refers to the phenomenon in which segments of the genome are repeated and the number of repeated sequences in the genome varies among individuals in the population under consideration.
[0064] Coverage: As used herein, "coverage" refers to the number of nucleic acid molecules that represent a particular base position.
[0065] Deoxyribonucleic acid or ribonucleic acid: As used herein, "deoxyribonucleic acid" or "DNA" refers to natural or modified nucleotides having a hydrogen group at the 2' position of the sugar moiety. DNA typically comprises a chain of nucleotides containing deoxyribonucleosides, each containing one of four types of nucleobases: adenine (A), thymine (T), cytosine (C), and guanine (G). As used herein, "ribonucleic acid" or "RNA" refers to natural or modified nucleotides having a hydroxyl group at the 2' position of the sugar moiety. RNA typically comprises a chain of nucleotides containing ribonucleosides, each containing one of four types of nucleobases: A, uracil (U), G, and C. As used herein, the term "nucleotide" refers to natural or modified nucleotides. Certain pairs of nucleotides specifically bind to each other in a complementary manner (referred to as complementary base pairing). In DNA, adenine (A) pairs with thymine (T) and cytosine (C) pairs with guanine (G). In RNA, adenine (A) pairs with uracil (U) and cytosine (C) pairs with guanine (G). When a first nucleic acid strand binds to a second nucleic acid strand composed of nucleotides that are complementary to the nucleotides in the first strand, the two strands combine to form a duplex. As used herein, "nucleic acid sequencing data," "nucleic acid sequencing information," "sequence information," "nucleic acid sequence," "nucleotide sequence," "genomic sequence," "gene sequence," or "fragment sequence," or "nucleic acid sequencing read data" refers to any information or data that indicates the order and identity of nucleotide bases (e.g., adenine, guanine, cytosine, and thymine, or uracil) in a molecule of nucleic acid, such as DNA or RNA (e.g., an entire genome, entire transcriptome, exome, oligonucleotide, polynucleotide, or fragment).It should be understood that the present teachings contemplate sequence information obtained using all available techniques, platforms, or technologies, including, but not limited to, capillary electrophoresis, microarrays, ligation-based systems, polymerase-based systems, hybridization-based systems, direct or indirect nucleotide identification systems, pyrosequencing, ion- or pH-based detection systems, and electronic signature-based systems.
[0066] Detect: As used herein, "detect," "detecting," or "detection" refers to the act of determining the existence or presence of one or more target nucleic acids (e.g., nucleic acids having targeted mutations or other markers) in a sample.
[0067] Driver mutation: As used herein, "driver mutation" means a mutation that promotes cancer progression.
[0068] Historical prevalence: As used herein, "historical prevalence" refers to sequence information obtained from or data derived from one or more reference samples (e.g., from reference subjects with a given cancer type) and / or a given subject.
[0069] Immunotherapy: As used herein, "immunotherapy" refers to treatment with one or more agents that stimulate the immune system to kill cancer cells or at least inhibit their growth, preferably to reduce further cancer growth, reduce the size of cancer, and / or eliminate cancer. Some such agents bind to targets present on cancer cells, some bind to targets present on immune cells but not on cancer cells, and some bind to targets present on both cancer cells and immune cells. Such agents include, but are not limited to, checkpoint inhibitors and / or antibodies. Checkpoint inhibitors are inhibitors of immune system pathways that maintain self-tolerance and modulate the duration and magnitude of physiological immune responses in peripheral tissues to minimize collateral tissue damage (see, e.g., Pardoll, Nature Reviews Cancer 12, 252-264 (2012)). Exemplary agents include antibodies against PD-1, PD-2, PD-L1, PD-L2, CTLA-4, OX40, B7.1, B7He, LAG3, CD137, KIR, CCR5, CD27, CD40, or CD47. Other exemplary agents include pro-inflammatory cytokines, such as IL-1β, IL-6, and TNF-α. Other exemplary agents are T cells activated against tumors, such as T cells activated by expressing a chimeric antigen that targets a tumor antigen recognized by the T cell.
[0070] Indel: As used herein, "indel" refers to a mutation involving the insertion or deletion of a nucleotide position in the genome of a subject.
[0071] Log prior data: As used herein, "log prior data" refers to the logarithm of the ratio of nucleic acid variant(s) or mutation(s) (e.g., target nucleic acid variant(s) or mutation(s)) to the wild-type variant in a sample population.
[0072] Maximum mutant allele frequency: As used herein, "maximum mutant allele frequency," "maximum MAF," or "MAX MAF" refers to the largest or greatest MAF of all somatic variants present or observed in a given sample.
[0073] Mutant allele frequency: As used herein, "mutant allele frequency," or "MAF," refers to the frequency with which a mutant allele occurs in a given population of nucleic acids, such as a sample obtained from a subject. MAF is typically expressed as a proportion or percentage.
[0074] Mutation: As used herein, "mutation," "variant," or "genetic abnormality" refers to a variation from a known reference sequence, including, for example, single nucleotide variants (SNVs), copy number variants or polymorphisms (CNVs) / abnormalities, insertions or deletions (indels), truncations, gene fusions, transversions, translocations, frameshifts, duplications, repeat expansions, and epigenetic variants. Mutations can be germline or somatic mutations. In some embodiments, the reference sequence for comparison purposes is the wild-type genomic sequence of the species from which the test sample is provided, typically the human genome.
[0075] Next-generation sequencing: As used herein, "next-generation sequencing" or "NGS" refers to a sequencing technology that has increased throughput compared to traditional Sanger and capillary electrophoresis-based approaches, for example, the ability to generate hundreds of thousands of relatively small sequence reads at once. Some examples of next-generation sequencing technologies include, but are not limited to, sequencing-by-synthesis, sequencing-by-ligation, and sequencing-by-hybridization.
[0076] Nucleic acid tag: As used herein, "nucleic acid tag" refers to a short nucleic acid (e.g., less than about 500 nucleotides in length, less than about 100 nucleotides in length, less than about 50 nucleotides in length, or less than about 10 nucleotides in length) used to label nucleic acid molecules to distinguish between nucleic acids from different samples (e.g., representing a sample index) or between different nucleic acid molecules in the same sample that have undergone different types or treatments (e.g., representing a molecular tag). Nucleic acid tags can be single-stranded, double-stranded, or at least partially double-stranded. Nucleic acid tags can have the same or different lengths as desired. Nucleic acid tags can also include double-stranded molecules with one or more blunt ends, can include 5' or 3' single-stranded regions (e.g., overhangs), and / or can include one or more other single-stranded regions at other locations within a given molecule. Nucleic acid tags can be attached to one or both ends of other nucleic acids (e.g., sample nucleic acids to be amplified and / or sequenced). Nucleic acid tags can be decoded to reveal information such as the sample of origin, form, or treatment of a given nucleic acid. Nucleic acid tags can also be used to enable pooling and / or parallel processing of multiple samples containing nucleic acids with different nucleic acid tags and / or nucleic acids with sample indices in which the nucleic acids are subsequently deconvoluted by reading the nucleic acid tags. Nucleic acid tags can also be referred to as molecular identifiers or tags, sample identifiers, index tags, and / or barcodes. Additionally or alternatively, nucleic acid tags can be used to distinguish different molecules within the same sample. This includes, for example, uniquely tagging each different nucleic acid molecule in a given sample, or non-uniquely tagging such molecules. For non-unique tagging applications, each nucleic acid molecule can be tagged using tags with a limited number of different sequences, thereby allowing different molecules to be distinguished, for example, based on the start and / or stop codon positions they map to a selected reference genome in combination with at least one nucleic acid tag.Typically, a sufficient number of different nucleic acid tags are used so that the probability that any two molecules have the same start / stop position and also have the same nucleic acid tag is low (e.g., less than about 10%, less than about 5%, less than about 1%, or less than about 0.1%). Some nucleic acid tags include multiple molecular identifiers to label samples, forms of nucleic acid molecules within the sample, and nucleic acid molecules within the form that have the same start and stop positions. Such nucleic acid tags can be referenced using the exemplary format "A1i", where the capital letter indicates the type of sample, the Arabic numerals indicate the form of molecules within the sample, and the lowercase Roman numerals indicate the molecules within the form.
[0077] Polynucleotide: As used herein, "polynucleotide," "nucleic acid," "nucleic acid molecule," or "oligonucleotide" refers to a linear polymer of nucleosides (including deoxyribonucleosides, ribonucleosides, or analogs thereof) linked by internucleoside linkages. Typically, a polynucleotide contains at least three nucleosides. Oligonucleotides often range in size from a small number of monomeric units, e.g., 3-4, to several hundred monomeric units. Whenever a polynucleotide is represented by a string of characters such as "ATGCCTG," the nucleotides are understood to be in 5'→3' order from left to right, and in the case of DNA, "A" denotes deoxyadenosine, "C" denotes deoxycytidine, "G" denotes deoxyguanosine, and "T" denotes deoxythymidine, unless otherwise indicated. The letters A, C, G, and T may be used to refer to the base itself, a nucleoside, or a nucleotide containing the base, as is standard in the art.
[0078] Reference Sample: As used herein, "reference sample" or "reference cfNA sample" refers to a sample of known composition and / or known to have or lack certain characteristics (e.g., known nucleic acid variant(s), known cellular origin, known tumor fraction, known coverage, and / or other) that is analyzed together with or compared to a test sample to assess the accuracy of an analytical procedure. A reference sample dataset typically includes at least about 25 to at least about 30,000 or more reference samples. In some embodiments, the reference sample dataset comprises about 50, 75, 100, 150, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,500, 5,000, 7,500, 10,000, 15,000, 20,000, 25,000, 50,000, 100,000, 1,000,000, or more reference samples.
[0079] Reference sequence: As used herein, "reference sequence" or "reference genome" refers to a known sequence used for comparison purposes with empirically determined sequences. For example, the known sequence can be a whole genome, a chromosome, or any segment thereof. A reference sequence typically comprises at least about 20, at least about 50, at least about 100, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, at least about 500, at least about 1000, or more nucleotides. A reference sequence can be aligned with a single continuous sequence of a genome or chromosome, or can comprise discontinuous segments that align with different regions of a genome or chromosome. Exemplary reference sequences include, for example, the human genome, for example, hG19 and hG38.
[0080] Sample: As used herein, "sample" means anything that can be analyzed by the methods and / or systems disclosed herein.
[0081] Sensitivity: As used herein, "sensitivity" in the context of a given assay or method refers to the ability of the assay or method to detect and distinguish between target analytes (e.g., nucleic acid variants) and non-target analytes.
[0082] Sequencing: As used herein, "sequencing" refers to any of several technologies used to determine the sequence (e.g., the identity and order of monomeric units) of a biomolecule, e.g., a nucleic acid such as DNA or RNA. Exemplary sequencing methods include targeted sequencing, single molecule real-time sequencing, exon or exome sequencing, intron sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxytermination sequencing, whole genome sequencing, sequencing by hybridization, pyrosequencing, capillary electrophoresis, double-stranded sequencing, cycle sequencing, single-base extension sequencing, solid-phase sequencing, high-throughput sequencing, massively parallel signature sequencing, emulsion PCR, co-amplification at low denaturation temperature-PCR (COLD-PCR), multiplex PCR, reversible dye terminator sequencing, paired-end sequencing, near-term sequencing, exonuclease sequencing, sequencing by ligation, short-read sequencing, single molecule sequencing, sequencing by synthesis, real-time sequencing, reverse-terminator sequencing, nanopore sequencing, 454 sequencing, Solexa Genome These include, but are not limited to, Genetic Analyzer sequencing, SOLiD™ sequencing, MS-PET sequencing, and combinations thereof. In some embodiments, sequencing can be performed by a genetic analyzer, such as a commercially available genetic analyzer from Illumina, Inc., Pacific Biosciences, Inc., or Applied Biosystems / Thermo Fisher Scientific, among many others.
[0083] Sequence information: As used herein, "sequence information" in the context of a nucleic acid polymer means the order and identity of the monomeric units (e.g., nucleotides, etc.) in that polymer.
[0084] Single nucleotide variant: As used herein, "single nucleotide variant" or "SNV" refers to a mutation or polymorphism of a single nucleotide that occurs at a specific position in the genome.
[0085] Somatic mutation: As used herein, "somatic mutation" refers to a mutation in the genome that occurs after conception. Somatic mutations can occur in any cell of the body except germ cells and are therefore not passed on to offspring.
[0086] Specificity: As used herein, "specificity" in the context of a diagnostic analysis or assay refers to the degree to which the analysis or assay detects the intended target analyte to the exclusion of other components of a given sample.
[0087] Subclonal: As used herein, "subclonal" in the context of nucleic acids refers to subpopulations of nucleic acids that contain nucleotide sequences that are substantially or completely identical to each other at least at a given locus of interest (e.g., target variant).
[0088] Subject: As used herein, "subject," or "test subject," refers to an animal, such as a mammalian species (e.g., a human) or an avian (e.g., a bird) species, or other organism, such as a plant. More specifically, the subject can be a vertebrate, such as a mammal, such as a mouse, a primate, a monkey, or a human. Animals include farm animals (e.g., beef cattle, dairy cattle, poultry, horses, pigs, and others), sport animals, and companion animals (e.g., pets or service animals). A subject can be a healthy individual, an individual with or suspected of having a disease or a predisposition to a disease, or an individual in need of therapy or suspected of needing therapy. The terms "individual," or "patient," are intended interchangeably with "subject." In some embodiments, the subject is a human with or suspected of having cancer. For example, the subject can be an individual who has been diagnosed with cancer, will receive cancer therapy, and / or has received at least one cancer therapy. The subject can be in remission from cancer. As another example, the subject may be an individual diagnosed with an autoimmune disease. As another example, the subject may be a female individual who is pregnant or planning to become pregnant, who may have been diagnosed with or suspected of having a disease, e.g., cancer, an autoimmune disease.
[0089] Threshold: As used herein, "threshold" refers to an individually determined value used to characterize or classify experimentally determined values. In certain embodiments, for example, "threshold" refers to a selected value to which a quantitative value is compared to determine that a given target nucleic acid variant is absent at a given locus.
[0090] Tumor fraction: As used herein, "tumor fraction" refers to an estimate of the proportion of nucleic acid molecules derived from tumors in a given sample. For example, the tumor fraction of a sample can be a measurement derived from the maximum mutant allele frequency (MAX MAF) of the sample or the coverage of the sample, or the length, epigenetic status, or other characteristics of the cfNA fragments in the sample, or any other selected feature of the sample. The term "MAX MAF" refers to the maximum or highest MAF of all somatic variants present in a given sample. In some embodiments, the tumor fraction of a sample is equal to the MAX MAF of the sample.
[0091] Value: As used herein, a "value" refers to an entry in a dataset, which can generally be anything that characterizes the trait to which the value refers, including, but not limited to, a number, a word or phrase, a symbol (e.g., + or -), or a degree. Detailed Description
[0092] FIG. 1 illustrates an example of a system 100 for generating a negative prediction of a target variant in a sample of a subject 111 according to an embodiment of the present disclosure. The system 100 may process one or more samples 101 from the subject 111 to generate sequence read data for variant detection and negative prediction. The system 100 may include a laboratory system 102, a computer system 110, and / or other components. It should be noted that the laboratory system 102 and the computer system 110 may be remote from each other or may be connected to each other by a computer network (not illustrated). The laboratory system 102 may include a sample collection and preparation pipeline 103, a sequencing pipeline 105, a sequence read data data store 109, and / or other components. The sequencing pipeline 105 may include one or more sequencing devices 107 (illustrated in FIG. 1 as sequencing devices 107a...n).
[0093] The computer system 110 may include a sequence analysis pipeline 112, a processor 120, a storage device 122, a variant detection pipeline 130, and / or other components.
[0094] The sequence analysis pipeline 112 may include a sequence quality control (QC) component 113 that can trim or discard sequence read data from the laboratory system 102, other analysis components 115 that may perform preliminary alignments with a reference genome, and an analysis QC component 116 that may perform quality control on the output of the analysis component 115. Output, for example, sequence read data for sample 101 of subject 111 from the sequence analysis pipeline 112, may be stored in an analysis data store 117.
[0095] Generally speaking, processor 120 may implement (be programmed by) various components of variant detection pipeline 130, such as variant detector 132, negative prediction analyzer 134, and / or other components. It should be noted that each of these components of variant detection pipeline 130 may alternatively comprise hardware modules. While illustrated separately for convenience, one or more of the various components or instructions, such as variant detector 132 and negative prediction analyzer 134, may be integrated with one another. In any event, variant detection pipeline 130 enables computer system 110 to identify variants, variant-derived diseases (diagnoses), negative predictions, and / or treatment regimens. The diagnoses and treatment regimens may be stored in a repository, such as clinical results store 160 or diagnostic results store 150.
[0096] Variant detector 132 may determine that a target variant was not detected based on analysis of sequence read data from laboratory system 102. It should be noted that while at least one sequence read and / or at least one molecule sequenced may confirm a target variant, this is not sufficient for variant detector 132 to detect the target variant. By way of example, in some embodiments, variant detector 132 may detect a target variant only if the number of sequence reads (and / or the number of molecules sequenced) confirming the target variant is greater than a threshold. Additionally or alternatively, variant detector 132 may detect a target variant only if the target variant confirmed by the sequence reads and / or the molecules sequenced meets a quality threshold. Target variants confirmed by at least one sequence read and / or at least one molecule sequenced, but that do not meet the threshold, may thus be ignored as false positives in some embodiments and may not be detected by variant detector 132. Other methods of determining that a target variant has not been detected based on analysis of sequence read data may be used, but further details for making this determination are omitted for clarity.
[0097] The negative prediction analyzer 134 may access the output of the variant detector 132 to confirm the negative prediction as an add-on to the variant detector. Alternatively or additionally, the negative prediction analyzer 134 may be incorporated with the variant detector 132.
[0098] 2 illustrates a schematic diagram of exemplary inputs and outputs of negative prediction analyzer 134 according to one embodiment. Negative prediction analyzer 134 may use covariate information 202, coverage information at target sites 204, disease type 206, and / or other input information for significance modeling. Negative prediction analyzer 134 may generate a quantitative value output 210, which may represent the likelihood that the negative prediction is correct or not, and a negative prediction assessment 212, which may include a confidence level or precision diagnosis based on the quantitative value output 210.
[0099] For example, sequence read data from the laboratory system 102 may be aligned with a reference genome, particularly various loci in the reference genome, to determine covariate information 202. The covariate information 202 may include covariate variant information, which may include historical mutual exclusivity and / or co-occurrence data of variants. Covariate variants may refer to two or more variants that have a negative correlation (mutual exclusivity) or a positive correlation (co-occurrence) with each other based on historical observations of sequence data from the laboratory system 102 and / or other data sources. For example, mutually exclusive variants may include variants that tend not to be observed with each other. Co-occurrence variants may be observed to occur when another variant is observed, such as a driver variant mutation and its co-occurring variant.
[0100] In certain examples, significance modeling may generate and use a computational estimate of the tumor fraction (TF) of a target variant based on nucleic acid sequence read data generated from the sample. Alternatively or additionally, significance modeling may determine and use the diversity of other variants detected or not detected in the sample. For example, significance modeling may use the detection of covariant variants that typically (based on past covariant variant information) co-occur with the target variant, or mutually exclusive variants that typically (based on past covariant variant information) do not co-occur with the target variant. A negative predictive value ("NPV") may be generated based on the TF estimates and / or diversity of variants detected or not detected in the sample. The results may be used to provide a confidence level in a negative diagnosis and / or further guide a treatment plan based on a negative diagnosis. For example, in the context of cancer diagnosis, covariant variants may include driver variants that tend to promote tumorigenesis, and mutually exclusive variants may include tumor suppressor variants that tend to suppress tumorigenesis.
[0101] Negative prediction
[0102] FIG. 3 illustrates an example of a method 300 for generating a negative prediction of a target variant in a subject's sample according to one embodiment of the present disclosure.
[0103] The methods of the present invention can be used to determine the absence of a target variant (e.g., absence at the clonal level) as a true negative result. Thus, with reference to FIG. 3, at 302, method 300 can include accessing a plurality of sequence read data of a cfDNA sample. At 304, method 300 can include determining, based on the plurality of sequence read data, that a target variant (the target variant) was not detected at a first locus in a sample (e.g., a cfNA sample). In some examples, the target variant (and / or other variants described herein) can include a somatic variant. In some examples, the target variant (and / or other variants described herein) can exclude a germline variant.
[0104] Negative predictive evaluation
[0105] At 306, method 300 may include generating a first likelihood value based on the probability that the target variant is absent at the clonal level and a second likelihood value based on the probability that the target variant is present at the clonal level. At 308, method 300 may include determining a quantification value based on the first likelihood value and the second likelihood value. At 310, method 300 may include comparing the quantification value to a threshold value. At 312, method 300 may include determining, based on the comparison, that the target variant is absent at the clonal level at the first locus. For example, method 300 may include determining that the allele frequency of the target variant does not exceed a threshold value (e.g., the subclonality threshold described with reference to Figures 4A and 4B).
[0106] Evaluation of negative prediction based on tumor fraction estimates
[0107] In some examples, method 300 and / or negative prediction analyzer 134 (by implementing method 300) may model the probability that the target variant is absent at the clonal level (or present at the subclonal level of the tumor variant) as a test or alternative hypothesis (H1) to generate a first likelihood value. For example, FIG. 4A illustrates a graph 400A of a test hypothesis that the target variant (the target variant) is absent from the sample (or present at the subclonal level of the tumor variant) according to one embodiment. Correspondingly, negative prediction analyzer 134 may model the probability that the target variant is present at the clonal level as a null hypothesis ((H0)) to generate a second likelihood value. For example, FIG. 4B illustrates a graph 400B of a null hypothesis that the target variant is present in the sample (and correlates with the allele frequency of the tumor variant) according to one embodiment. In both graphs 400A and 400B, "C" reflects the minor allele at the target locus. The value "0.3" reflects the weight applied to α1 (a TF estimate based on the mutant allele frequency of the tumor variant) such that the product of 0.3 x α1 acts as a threshold for subclonality. An allele frequency (α2) of the target variant in sample 101 of subject 111 that exceeds the threshold for subclonality may indicate that the target variant is correlated with the tumor variant.
[0108] In these examples, the negative prediction analyzer 134 may generate the first likelihood value and the second likelihood value by determining a tumor fraction (TF) estimate (e.g., α1 in the formulas described herein) for the sample. The TF estimate may indicate the proportion of tumor DNA detected in the sample. In some examples, the TF estimate may be determined by determining the allele frequency (referred to as the MAX MAF) of the tumor variant in the sample. The MAX MAF may be determined by determining a molecular count associated with the tumor variant based on multiple sequence read data. The first likelihood value, based on the probability that the target variant is not present at the clonal level (e.g., L1 in the formulas described herein), and the second likelihood value, based on the probability that the target variant is present at the clonal level or at the subclonal level (e.g., L0 in the formulas described herein), may be based on the TF estimate.
[0109] In some embodiments, the negative prediction analyzer 134 may use the TF estimates to generate a quantitative value that evaluates the quality of the negative prediction (e.g., by indicating the probability that the negative prediction is correct or false). For example, the negative prediction analyzer 134 may determine a first allele frequency of a target variant. The negative prediction analyzer 134 may determine the first allele frequency by determining a first molecular count associated with the target variant based on multiple sequence read data. The negative prediction analyzer 134 may use the first allele frequency together with the MAX MAF to determine first and second likelihood values further based on the first allele frequency and the MAX MAF.
[0110] With reference to Figure 4A, the probability that a target variant is absent at the clonal level (or present at the subclonal level) is determined by the subclonal threshold (0.3 *The first likelihood value and the second likelihood value may be based on a subclone weight (illustrated as α1). This may be a subclone weight (illustrated as 0.3) multiplied by a tumor fraction estimate (illustrated as an allele frequency, such as the MAX MAF, of a tumor variant). The subclone threshold may be determined based on a specific gene, cancer type, or other predictive value. These values may be within the range of 0.01 to 0.99, including, but not limited to, 0.01, 0.10, 0.20, 0.30, 0.40, 0.50, 0.60, 0.70, 0.80, 0.90, and 0.99. Equations 1-3 below relate to generating the first and second likelihood values and the resulting quantitative values in certain embodiments.
[0111]
number
[0112]
number
[0113]
number
[0114] Referring to Equations 1 to 3, L1 refers to the likelihood value of the test hypothesis that the variant is not present at the clonal level. The null hypothesis was generated using the same formula as for L1, but with alpha2 having a different range of values (e.g., 0.3 to 1). α1 refers to the allele frequency of the tumor variant, which can be used as a TF estimate. α2 refers to the allele frequency of the target variant (the target variant). M v refers to the molecular count that evidences the tumor variant at the tumor variant locus. M r refers to molecular counts corroborating the reference wild type at the tumor variant locus. M v' refers to the molecular count that evidences the target variant at the target variant locus. Mr' refers to the molecular count verifying the reference wild type at the locus of the targeted variant. ε refers to the error rate of the TF estimate. ε′ refers to the error rate of the target variant. Error rates are typically derived from sequence information (eg, z-scores or otherwise) obtained from samples obtained from healthy or normal subjects.
[0115] α2=t*α1 (Equation 4). This equation is for simplicity purposes (same as Equation 1), but is easier to calculate than the integral in Equation 1.
[0116] The ε and ε' error rates (maxmaf) in tumor proportion and target variants correspond to
[0117]
number
[0118]
number
[0119] Epsilon (ε) is obtained from the calculation of a z-score derived from sequence information obtained from samples obtained from healthy or normal subjects.
[0120] In the following formula: T - indicates that the target variant is absent at the clonal level. T + indicates that the target variant is present at the clonal level. V i + indicates the presence of variants (other than the target) (i = 1, ..., n, all other called variants). L i refers to the likelihood value (base hypothesis, i=0; test hypothesis, i=1).
[0121] Adjusting quantitative values based on the occurrence of other variants
[0122] In some examples, the negative prediction analyzer 134 may adjust the quantitative value determined from the TF estimate based on the presence of one or more variants other than the target variant in the sample 101 of the subject 111. For example, the negative prediction analyzer 134 may determine the occurrence of at least a second variant in the cfDNA sample 101 and adjust the quantitative value based on the occurrence of the at least a second variant.
[0123] For example, the occurrence data may be determined according to Equations 7 and 8:
[0124]
number
[0125]
number
[0126] The likelihood value (L1) that the test hypothesis is correct is adjusted based on Equation 9, and the adjusted likelihood value (L 1a ) may be generated, and the likelihood ratio (LR a ) may be generated according to Equation 10:
[0127]
number
[0128]
number
[0129] Equation 10 is the likelihood ratio using the conditional dependent property.
[0130] Evaluation of LLR-based negative prediction
[0131] In some examples, the quantitative value may be based on the LLR between the first likelihood value and the second likelihood value. Thus, the quantitative value may be based on the ratio between the first likelihood value (e.g., L1 in Equation 14) and the second likelihood value (e.g., L0 in Equation 15). In some examples, the negative prediction analyzer 134 may calculate the LLR (e.g., LLR illustrated in Equation 16) based on the TF. tf The negative prediction analyzer 134 may generate a quantitative value (e.g., LLR) based on Equation 11:
[0132] LLR=LLR tf +LLR me (Equation 11) (tumor ratio (LLR tf ) and mutual exclusivity (LLR me ) log-likelihood ratio (LLR).
[0133] Evaluating Negative Prediction Using LLR Based on Covariance (Mutually Exclusive) Data
[0134] In some examples, the quantitative value may be based on the LLR of the covariance data. For example, the negative prediction analyzer 134 may generate an LLR that reflects the covariance data, as illustrated in Equation 18: me (conditional probability of the number of times the variants are observed together).
[0135]
number
[0136]
number
[0137] Evaluating negative prediction using combinations of LLRs
[0138] In some embodiments, the quantitative value may be expressed as a log-posteriori probability ratio (LPPR) based on a combination of a TF-based log-likelihood of whether the null or test hypothesis is correct, a covariance-based (e.g., mutual exclusivity) log-likelihood of whether the null or test hypothesis is correct, and log-posteriori probability ratio (LPPR) based on prior data, as represented in Equations 19 and 21 below. In some examples, the quantitative value (e.g., LLR in Equation 11) may be further based on log-prior data based on past observed data that is not necessarily limited to the sample 101 of the subject 111. Such log-prior data may be based on covariate information indicating the past prevalence of one or more variants that co-occur with and / or mutually exclusive with the target variant. For example, the log-prior data may be
number
[0139]
number
[0140]
number
[0141]
number
[0142]
number
[0143]
number
[0144]
number
[0145]
number
[0146]
number
[0147]
number
[0148] It should be understood that in the previous examples, the negative prediction analyzer 134 is described as implementing the method 300 and performing the additional operations described above. It should be further understood that the additional operations described above may be part of, and extend, the method 300.
[0149] The various processes, operations, and / or methods shown in the figures may be achieved using some or all of the system components described in detail herein, and in some implementations, various operations may be performed in a different order or various operations may be omitted. Additional operations may be performed along with some or all of the operations shown in the depicted flow charts. One or more operations may be performed simultaneously. Accordingly, the operations illustrated (and described in more detail herein) are provided by way of example and should not be considered limiting as such.
[0150] Computer Implementation
[0151] The methods of the present invention may be computer-implemented, such that any or all of the operations described herein or in the appended claims, other than the wet chemistry steps, can be performed on a suitable programmed computer. The computer may be a mainframe, personal computer, tablet, smartphone, cloud, online data storage, remote data storage, or other. The computer may be operated in one or more locations.
[0152] Various operations of the methods of the present invention may utilize information and / or programs to generate results, which may be stored on computer-readable media (e.g., hard drives, secondary memory, expansion memory, servers, databases, portable memory devices (e.g., CD-Rs, DVDs, ZIP disks, flash memory cards), and the like).
[0153] The present disclosure also includes articles of manufacture for analyzing nucleic acid populations that include machine-readable media containing one or more programs that, when executed, implement the steps of the methods of the invention.
[0154] The present disclosure can be implemented in hardware and / or software. For example, different aspects of the present disclosure can be implemented in either client-side logic or server-side logic. The present disclosure, or components thereof, can be embodied in fixed-medium program components containing logic instructions and / or data that, when loaded into a suitably configured computing device, cause the device to perform in accordance with the present disclosure. The fixed medium containing the logic instructions can be sent to the viewer on the fixed medium for physical loading into the viewer's computer, or the fixed medium containing the logic instructions may reside on a remote server that the viewer accesses through a transmission medium to download the program components.
[0155] The present disclosure provides a computer control system programmed to implement the disclosed methods. The processor 120 may include a single-core or multi-core processor, or multiple processors for parallel processing. The storage device 122 may include random access memory, read-only memory, flash memory, a hard disk, and / or other types of storage. The computer system 110 may include a communication interface (e.g., a network adapter) for communicating with one or more other systems, as well as peripheral devices such as cache, other memory, data storage, and / or an electronic display adapter. Components of the computer system 110 may communicate with each other through an internal communication bus, such as a motherboard. The storage device 122 may be a data storage unit (or data repository) for storing data. The computer system 110 may be operably coupled to a computer network ("network") with the aid of a communication interface. The network may be the Internet, an internet and / or an extranet, or an intranet and / or extranet communicating with the Internet. In some cases, the network is a telecommunications and / or data network. The network may include a local area network. The network may include one or more computer servers, which may enable distributed computing such as cloud computing. The network may, in some examples, implement a peer-to-peer network, which may enable devices coupled to computer system 120 to act as clients or servers, with the help of computer system 110.
[0156] Processor 120 may execute sequences of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location, such as storage device 122. The instructions may be directed to processor 120, which may then program or otherwise configure processor 120 to implement the methods of the present disclosure. Examples of operations performed by processor 120 may include fetch, decode, execute, and write-back.
[0157] The processor 120 may be part of a circuit, such as an integrated circuit. One or more other components of the system 100 may also be included in the circuit. In some examples, the circuit may include an application-specific integrated circuit (ASIC).
[0158] Storage device 122 may store files such as drivers, libraries, and saved programs. Storage device 122 can store user data, such as user preferences and user programs. Computer system 110 may, in some examples, include one or more additional data storage units located external to computer system 110, for example, on remote servers that communicate with computer system 110 through an intranet or the Internet.
[0159] Computer system 110 can communicate with one or more remote computer systems through a network. By way of example, computer system 110 can communicate with a user's remote computer system. Examples of remote computer systems include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., an Apple® iPad®, a Samsung® Galaxy Tab), a telephone, a smartphone (e.g., an Apple® iPhone®, an Android®-type device, a Blackberry®), or a personal digital assistant. A user can access computer system 110 through a network.
[0160] The methods described herein may be implemented by machine (e.g., computer processor-) executable code stored in an electronic storage location of computer system 110, such as storage device 122. The machine-executable or machine-readable code may be provided in the form of software (e.g., a computer-readable medium). In use, the code may be executed by processor 120. In some examples, the code may be retrieved from storage device 122 and stored on storage device 122 for easy access by processor 120.
[0161] The code may be precompiled and configured for use with a machine having a processor adapted to execute the code, or may be compiled during the routine. The code may be supplied in a programming language that can be selected to enable the code to be executed as precompiled or compiled.
[0162] Aspects of the systems and methods provided herein, such as the computer system 110, can be embodied in programming. Various aspects of the technology may be considered to be "products" or "articles of manufacture," typically in the form of machine- (or processor-) executable code and / or associated data embodied in or embodied in a type of machine-readable medium. The machine-executable code can be stored in memory (e.g., read-only memory, random-access memory, flash memory) or an electronic storage unit, such as a hard disk.
[0163] A "storage" type medium may include any or all of a computer's tangible memory, processor, or other, or associated modules, such as various semiconductor memories, tape drives, disk drives, and the like, which may provide non-transitory storage at any time for software programming. All or a portion of the software may sometimes be communicated over the Internet or various other telecommunications networks. Such communication may enable, for example, loading of the software from one computer or processor to another, such as from a management server or host computer to an application server computer platform. Thus, other types of media that may bear software elements include light waves, radio waves, and electromagnetic waves used across physical interfaces between local devices, for example, through wired and optical telephone networks and by various air links. Physical elements bearing such waves, such as wired or wireless links, optical links, or the like, may also be considered media bearing software. As used herein, when not limited to non-transitory tangible storage media, "media" may include other types of (intangible) media.
[0164] Terms such as "storage" medium, computer or machine "readable medium" refer to any tangible (e.g., physical), non-transitory medium that participates in providing instructions to a processor for execution.
[0165] Thus, machine-readable media, e.g., computer-executable code, may take many forms, including, but not limited to, tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media include, for example, optical or magnetic disks, such as any storage device in any computer(s) or otherwise, such as those that may be used to implement a database shown in the drawings. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables, copper wire, or fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punch cards, paper tape, any other physical storage media with hole patterns, RAM, ROM, PROMs, and EPROMs, FLASH-EPROMs, any other memory chips or cartridges, carrier waves transporting data or instructions, cables or links transporting such carrier waves, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer-readable media may be involved in transporting one or more sequences of one or more instructions to a processor for execution.
[0166] The computer system 110 may include or be in communication with an electronic display 935 that includes a user interface (UI), for example to provide reports. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0167] The methods and systems of the present disclosure may be implemented by one or more algorithms, which may be implemented by software when executed by a processor.
[0168] Sample collection and analysis pipeline
[0169] The sample 101 can be any biological sample isolated from a subject. Samples can include body tissues, such as known or suspected solid tumors, whole blood, platelets, serum, plasma, stool, red blood cells, white blood cells or leucocytes, endothelial cells, tissue biopsies, cerebrospinal fluid, synovial fluid, lymphatic fluid, ascites, interstitial or extracellular fluid, fluids in the spaces between cells, including gingival crevicular fluid, bone marrow, pleural effusion, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, and urine. Samples are preferably body fluids, particularly blood and its fractions, and urine. Such samples contain nucleic acids excreted from tumors. Nucleic acids can include DNA and RNA, and can be in double-stranded and / or single-stranded forms. Sample can be the form that is originally separated from subject, or can be further processed, such as removing or adding components such as cell, enriching one component compared with other components, or converting one form of nucleic acid into another form, for example, converting RNA into DNA, or converting single-stranded nucleic acid into double-stranded nucleic acid.Thus, for example, the body fluid that is to be analyzed is the plasma or serum that contains cell-free nucleic acid, for example, cell-free DNA (cfDNA).
[0170] In certain embodiments, polynucleotides can be enriched prior to sequencing. Enrichment can be performed for specific target regions ("target sequences") or non-specifically. In some embodiments, targeted regions of interest can be enriched by capture probes ("baits") selected for one or more bait set panels using a differential tiling and capture scheme. Differential tiling and capture schemes use bait sets at different relative concentrations to differentially tile (e.g., at different "resolutions") across genomic regions related to the baits, subject to a set of constraints (e.g., sequencer constraints, e.g., sequencing load, availability of each bait, etc.), and capture them at a desired level for downstream sequencing. These targeted genomic regions of interest can include regions of the genome or transcriptome of interest. In some embodiments, biotin-labeled beads bearing probes for one or more regions of interest can be used to capture target sequences, which can then be optionally amplified and enriched for the regions of interest.
[0171] Sequence capture typically involves the use of oligonucleotide probes that hybridize to target sequences. Probe set strategies can involve tiling probes across a region of interest. Such probes can be, for example, about 60-130 bases long. Sets can have depths of about 2x, 3x, 4x, 5x, 6x, 8x, 9x, 10x, 15x, 30x, 50x, or more. The effectiveness of sequence capture depends in part on the length of the sequence in the target molecule that is complementary (or nearly complementary) to the sequence of the probe.
[0172] In some embodiments, the methods of the present disclosure involve selectively enriching regions from a subject's genome or transcriptome prior to sequencing. In other embodiments, the methods of the present disclosure involve non-selectively enriching regions from a subject's genome or transcriptome prior to sequencing.
[0173] In certain embodiments, a sample index sequence is introduced into the polynucleotide after enrichment. The sample index sequence may be introduced through PCR or, optionally, ligated to the polynucleotide as part of an adaptor.
[0174] The volume of plasma can depend on the desired read depth for the sequenced region. Exemplary volumes are 0.4-40 ml, 5-20 ml, and 10-20 ml. For example, the volume can be 0.5 ml, 1 ml, 5 ml, 10 ml, 20 ml, 30 ml, or 40 ml. The volume of plasma from which the sample was taken can be 5-20 ml.
[0175] A sample can contain various amounts of nucleic acid, including genome equivalents. For example, a sample of about 30 ng of DNA contains about 10,000 (10 4 ) haploid human genome equivalents, and in the case of cfDNA, approximately 200 billion (2 × 10 11 ) individual polynucleotide molecules. Similarly, about 100 ng of DNA sample can contain about 30,000 haploid human genome equivalents, or about 600 billion individual molecules in the case of cfDNA.
[0176] The sample may contain nucleic acids from different sources, such as cells and acellular nucleic acids. The sample may contain nucleic acids with mutations. For example, the sample may contain DNA with germline mutations and / or somatic mutations. The sample may contain DNA with cancer-associated mutations (e.g., cancer-associated somatic mutations).
[0177] Exemplary amounts of cell-free nucleic acid in a sample prior to amplification range from about 1 fg to about 1 μg, e.g., 1 pg to 200 ng, 1 ng to 100 ng, or 10 ng to 1000 ng. For example, the amount can be up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng of cell-free nucleic acid molecules. The amount can be at least 1 fg, at least 10 fg, at least 100 fg, at least 1 pg, at least 10 pg, at least 100 pg, at least 1 ng, at least 10 ng, at least 100 ng, at least 150 ng, or at least 200 ng of cell-free nucleic acid molecules. The amount can be up to 1 femtomole (fg), 10 fg, 100 fg, 1 picogram (pg), 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng, or 200 ng of cell-free nucleic acid molecules. The method can include obtaining 200 ng from 1 femtogram (fg).
[0178] Cell-free nucleic acids have an exemplary size distribution of about 100 to 500 nucleotides, with molecules of 110 to about 230 nucleotides representing about 90% of the molecules, with the mode in humans being about 168 nucleotides and a second, smaller peak in the range of 240 to 430 nucleotides. Cell-free nucleic acids can be about 160 to about 180 nucleotides, or about 320 to about 360 nucleotides, or about 430 to about 480 nucleotides.
[0179] Cell-free nucleic acids can be isolated from body fluids through a partitioning step, which separates the cell-free nucleic acids found in solution from intact cells and other insoluble components of the body fluid. Partitioning can include techniques such as centrifugation or filtration. Alternatively, cells in the body fluid can be lysed, and both the cell-free and cellular nucleic acids can be processed. Generally, after adding a buffer and a washing step, the cell-free nucleic acids can be precipitated with alcohol. Additional washing steps, such as using a silica-based column, can be used to remove contaminants or salts. For example, non-specific bulk carrier nucleic acids can be added to the entire reaction to optimize certain aspects of the procedure, such as yield.
[0180] After such processing, the sample may contain various forms of nucleic acids, including double-stranded DNA, single-stranded DNA, and single-stranded RNA. If necessary, the single-stranded DNA and RNA can be converted to double-stranded form so that they can be included in subsequent processing and analysis steps. amplification
[0181] The sample nucleic acid flanked by adapters can be amplified by PCR, and other amplification methods can typically be primed by primers that bind to the primer binding sites on the adapters adjacent to the DNA molecules to be amplified.The amplification method can involve cycles of extension, denaturation, and annealing due to thermocycling, or can be isothermal, such as transcription-mediated amplification.Other amplification methods include ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and self-sustaining sequence-based amplification.
[0182] One or more amplifications can be applied to introduce barcodes into nucleic acid molecules using conventional nucleic acid amplification methods. Amplification can be performed in one or more reaction mixtures. Molecular tags and sample indexes / tags can be introduced simultaneously or in any sequential order. Molecular tags and sample indexes / tags can be introduced before and / or after sequence capture. In some cases, only molecular tags are introduced before probe capture, and sample indexes / tags are introduced after sequence capture. In some examples, both molecular tags and sample indexes / tags are introduced before probe capture. In some examples, sample indexes / tags are introduced after sequence capture. Sequence capture typically involves introducing a single-stranded nucleic acid molecule complementary to a target sequence, e.g., a coding sequence of a genomic region, where mutations in such regions are associated with a type of cancer. Typically, amplification generates multiple non-uniquely or uniquely tagged nucleic acid amplicons with molecular tags and sample indexes / tags ranging in size from 200 nt to 700 nt, 250 nt to 350 nt, or 320 nt to 550 nt. In some embodiments, the amplicon has a size of about 300 nt. In some embodiments, the amplicon has a size of about 500 nt. Barcode
[0183] Barcodes can be incorporated into or otherwise attached to adapters by chemical synthesis, ligation, overlap extension PCR, among other methods. Generally, assignment of unique or non-unique barcodes during a reaction follows the methods and systems described in U.S. Patent Application Publication Nos. 20010053519, 20110160078, and U.S. Patent Nos. 6,582,908, 7,537,898, and 9,598,731.
[0184] The tags can be linked to the sample nucleic acids randomly or non-randomly. In some cases, they are introduced in a predicted ratio to the microwells of the identifiers (i.e., barcode combinations). The collection of barcodes can be unique, e.g., every barcode has a different nucleotide sequence. The collection of barcodes can be non-unique, i.e., some barcodes have the same nucleotide sequence and some barcodes have different nucleotide sequences. For example, identifiers may be loaded such that more than 1, more than 2, more than 3, more than 4, more than 5, more than 6, more than 7, more than 8, more than 9, more than 10, more than 20, more than 50, more than 100, more than 500, more than 1000, more than 5000, more than 10000, more than 50,000, more than 100,000, more than 500,000, more than 1,000,000, more than 10,000,000, more than 50,000,000, or more than 1,000,000,000 identifiers are loaded per genomic sample. In some examples, identifiers may be loaded such that fewer than 2, fewer than 3, fewer than 4, fewer than 5, fewer than 6, fewer than 7, fewer than 8, fewer than 9, fewer than 10, fewer than 20, fewer than 50, fewer than 100, fewer than 500, fewer than 1000, fewer than 5000, fewer than 10000, fewer than 50,000, fewer than 100,000, fewer than 500,000, fewer than 1,000,000, fewer than 10,000,000, fewer than 50,000,000, or fewer than 1,000,000,000 identifiers are loaded per genomic sample.In some examples, the average number of identifiers loaded per sample genome is less than about 1 or more, less than about 2 or more, less than about 3 or more, less than about 4 or more, less than about 5 or more, less than about 6 or more, less than about 7 or more, less than about 8 or more, less than about 9 or more, less than about 10 or more, less than about 20 or more, less than about 50 or more, less than about 100 or more, Less than or more than about 500, less than or more than about 1000, less than or more than about 5000, less than or more than about 10000, less than or more than about 50,000, less than or more than about 100,000, less than or more than about 500,000, less than or more than about 1,000,000, less than or more than about 10,000,000, less than or more than about 50,000,000, or less than or more than about 1,000,000,000 identifiers.
[0185] A preferred format uses 20-50 different tags ligated to both ends of a target molecule, creating 20-50 x 20-50 tags, or 400-2500 tag combinations. Such a large number of tags is sufficient to ensure that different molecules with the same start and stop points have a high probability (e.g., at least 94%, 99.5%, 99.99%, or 99.999%) of receiving different tag combinations.
[0186] In some examples, the identifier may be a predetermined, or a random or semi-random sequence oligonucleotide. In other examples, multiple barcodes may be used, such that the barcodes are not necessarily unique to each other among the plurality. In this example, the barcode may be attached to individual molecules (e.g., by ligation or PCR amplification) such that the combination of the barcode and the sequence to which it may be attached creates a unique sequence that can be individually tracked. As described herein, detection of a non-uniquely tagged barcode in combination with the genomic coordinates of the beginning (start) and / or end (stop) of a given sequenced sample molecule (i.e., excluding sequence information obtained from barcodes, adapters, and so forth) may allow for the assignment of a unique identity to a particular molecule. The length or number of base pairs of an individual sequenced sample molecule (i.e., excluding sequence information corresponding to barcodes, adapters, and so forth) may also be used to assign a unique identity to such a molecule. As described herein, fragments from a single strand of nucleic acid that have been assigned a unique identity may thereby allow for the subsequent identification of fragments from the parental strand and / or complementary strand. Sequencing pipeline
[0187] The adaptor-flanked sample nucleic acids, with or without prior amplification, can be subjected to sequencing, for example, by one or more sequencing devices 107. Sequencing methods include, for example, Sanger sequencing, high-throughput sequencing, pyrosequencing, sequencing-by-synthesis, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing-by-ligation, sequencing-by-hybridization, RNA-Seq (Illumina), Digital Gene Expression (Helicos), next-generation sequencing, single-molecule sequencing-by-synthesis (SMSS) (Helicos), massively parallel sequencing, Clonal Single Molecule Array (Solexa), shotgun sequencing, sequencing using Ion Torrent, Oxford Nanopore, Roche Genia, Maxam-Gilbert sequencing, primer walking, PacBio, SOLiD, Ion Torren, or Nanopore platforms. Sequencing reactions can be performed in a variety of sample processing units, which can be multiple lanes, multiple channels, multiple wells, or other means of processing multiple sample sets substantially simultaneously. The sample processing unit can also include multiple sample chambers capable of processing multiple runs simultaneously.
[0188] Sequencing reactions can be performed on one or more fragment types known to contain markers for other cancer diseases.Sequencing reactions can also be performed on any nucleic acid fragments present in a sample.Sequencing reactions can provide for sequencing at least 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9%, or 100% of a given genome.In other examples, sequencing reactions can provide for sequencing less than 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9%, or 100% of a given genome.
[0189] Multiplex sequencing can be used to carry out simultaneous sequencing reactions.In some examples, cell-free polynucleotides can be sequenced by at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions.In other examples, cell-free polynucleotides can be sequenced by less than 1000, less than 2000, less than 3000, less than 4000, less than 5000, less than 6000, less than 7000, less than 8000, less than 9000, less than 10000, less than 50000, less than 100,000 sequencing reactions.Sequencing reactions can be carried out sequentially or simultaneously.Subsequent data analysis can be carried out for all or part of sequencing reactions. In some examples, data analysis can be performed on at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. In other examples, data analysis can be performed on less than 1000, less than 2000, less than 3000, less than 4000, less than 5000, less than 6000, less than 7000, less than 8000, less than 9000, less than 10000, less than 50,000, or less than 100,000 sequencing reactions. An exemplary read data depth is 1000 to 50,000 read data per locus (base).
[0190] Sequence analysis pipeline
[0191] The methods of the invention can be used to diagnose the presence or absence of a condition, particularly cancer, in a subject, characterize the condition (e.g., determine the stage of the cancer or determine the heterogeneity of the cancer), monitor the response to treatment of the condition, or determine the prognostic risk of developing the condition or the subsequent course of the condition.
[0192] Various cancers can be detected using the method of the present invention. Like most cells, cancer cells can be characterized by the metabolic rate at which old cells die and new cells take their place. Generally, dead cells can release DNA or DNA fragments into the bloodstream when they come into contact with blood vessels in a given subject. This also applies to cancer cells during various stages of disease. Depending on the stage of disease, cancer cells can also be characterized by various genetic abnormalities, such as copy number variations and rare mutations. This phenomenon can be used to detect the presence or absence of cancer in individuals using the methods and systems described herein.
[0193] The types and number of cancers that may be detected may include blood cancer, brain cancer, lung cancer, skin cancer, nose cancer, throat cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, skin cancer, intestinal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, solid tumors, heterogeneous tumors, homogeneous tumors, and others.
[0194] Cancer can be detected from genetic variations including mutations, rare mutations, indels, copy number variations, transversions, translocations, inversions, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structural abnormalities, gene fusions, chromosomal fusions, gene truncations, gene amplifications, gene duplications, chromosomal lesions, DNA lesions, abnormal changes in nucleic acid chemical modifications, and abnormal changes in epigenetic patterns.
[0195] Genetic data can also be used to characterize specific types of cancer. Cancers are often heterogeneous in both composition and staging. Genetic profile data can enable characterization of specific subtypes of cancer, which may be important in diagnosing or treating that specific subtype. This information can also provide a subject or physician with clues regarding the prognosis of a specific type of cancer, allowing the subject or physician to take treatment options depending on the progression of the disease. Some cancers progress, become more aggressive, and become genetically unstable. Other cancers may remain benign, inactive, or dormant. The systems and methods of the present disclosure can be useful for determining disease progression.
[0196] The analysis of the present invention is also useful for determining the effectiveness of a particular treatment option. The success of a treatment option may increase the amount of copy number variations or rare mutations detected in the subject's blood, since more cancer cells die and release DNA if the treatment is successful. In other cases, this may not occur. In another example, a particular treatment option may possibly be correlated with the genetic profile of cancer over time. This correlation may be useful in selecting a therapy. In addition, if cancer is observed to be in remission after treatment, the method of the present invention can be used to monitor residual disease or disease recurrence.
[0197] The methods of the present invention can also be used to detect genetic mutations in conditions other than cancer. Immune cells, such as B cells, can undergo rapid clonal expansion based on the presence of a particular disease. Clonal expansion can be monitored using the detection of copy number variation, and a particular immune state can be monitored. In this example, copy number variation analysis can be performed over time to generate a profile of how a particular disease may progress. Detection of copy number variation or even rarer mutations can be used to determine how pathogen populations change over the course of infection. This can be particularly important during chronic infections, such as HIV / AIDS or hepatitis infections, where the virus can change life cycle states and / or mutate to a more virulent form over the course of infection. Because immune cells attempt to destroy transplanted tissue, the methods of the present invention can be used to determine or profile the host's rejection activity to alter the state of the transplanted tissue and the course of treatment or to monitor the prevention of rejection.
[0198] Furthermore, the methods of the present disclosure may be used to characterize heterogeneity of abnormal conditions in a subject, the method including generating a genetic profile of extracellular polynucleotides in the subject, the genetic profile including multiple data resulting from copy number variation and rare mutation analysis. In some cases, including but not limited to cancer, diseases may be heterogeneous. Disease cells may not be identical. In the example of cancer, it is known that some tumors contain different types of tumor cells, some cells at different stages of cancer. In other cases, heterogeneity may include multiple disease foci. Again, in the example of cancer, multiple tumor foci may be present, perhaps one or more foci being the result of metastasis that has spread from the primary site.
[0199] The methods of the present invention can be used to generate or profile a fingerprint, or set of data, that is a summary of genetic information from different cells in a heterogeneous disease. This data set can include copy number variation and rare mutation analysis, alone or in combination.
[0200] The methods of the present invention can be used to diagnose, prognose, monitor, or observe cancer or other diseases of fetal origin, i.e., these methodologies may be used on pregnant subjects to diagnose, prognose, monitor, or observe cancer or other diseases in unborn subjects whose DNA and other polynucleotides may co-circulate with maternal molecules.
[0201] Exemplary Precision Procedures
[0202] The precision diagnosis provided by the improved computer system 110 can result in a precision treatment plan that can be identified by the computer system 110 (and / or curated by a health professional). For example, in lung cancer and other diseases, the goal can be to ensure that no superior treatment options exist given the presence of a given variant. For example, EGFR (L858R, exon 19 deletion), BRAF V600E, ALK, and ROS1 fusions can be treated with targeted therapies that may be more suitable than platinum and chemotherapy. These are examples of primary drivers, but other targetable drivers exist, such as MET exon 14 skipping. In another example, in the case of colon cancer, the goal can be to avoid ineffective treatments. If KRAS or NRAS are wild-type, chemotherapy with FOLFIRI or irinotecan regimens can be supplemented with cetuximab or panitumumab. Thus, confirmation that KRAS and NRAS are wild-type increases confidence that adding cetuximab or panitumumab is the correct treatment option and no further testing is necessary. The biological explanation is that cetuximab or panitumumab targets EGFR and inhibits its activity. RAS (K / NRAS) is downstream of EGFR, so when RAS is activated, inhibiting EGFR has little or no effect, making cetuximab or panitumumab treatment inappropriate.
[0203] As additional therapies are developed for various diseases, the interpretation of negative predictive values becomes increasingly complex, yet is crucial in the design of precision therapies.
[0204] Another goal may be to guide whether downstream diagnostic procedures are performed. For example, determining the absence of a variant may allow for the avoidance (or recommendation of avoiding) expensive or invasive diagnostic tests, such as imaging procedures, scans (e.g., CT, MRI, or PET scans), endoscopic procedures, and / or solid tissue biopsies (e.g., needle biopsies). Similarly, it may also be possible to avoid (or recommend steps to avoid) separate liquid biopsy tests (e.g., blood, plasma, urine, cerebrospinal fluid), or stool tests. Blood assay-based results may be used in this way to guide reflexive tissue testing, avoiding the need for solid tissue biopsies to confirm the wild-type status of any potential variants of interest. The above-mentioned negative predictions may be used to assess the probability of the absence of clinically significant mutations in the liquid biopsy, providing confidence that the liquid biopsy is sufficient to detect the potential presence of a variant of interest and that downstream diagnostic procedures are not necessary. This may also facilitate timely treatment decisions.
[0205] Nucleotide polymorphisms in sequenced nucleic acids can be determined by comparing the sequenced nucleic acids with a reference sequence. The reference sequence is often a known sequence, such as a known entire or partial genome sequence obtained from a subject, or the entire genome sequence of a human subject. The reference sequence can be hG19. The sequenced nucleic acid can represent a sequence determined directly for the nucleic acid in the sample, or a consensus sequence of the amplification product of such nucleic acid. Comparison can be performed at one or more designated positions of the reference sequence. A subset of sequenced nucleic acids containing positions corresponding to the designated positions of the reference sequence when the reference sequence is maximally aligned can be identified. Within such a subset, it can be determined whether the sequenced nucleic acids, if any, contain nucleotide polymorphisms at the designated positions, and, if necessary, whether they contain the reference nucleotide, if any (i.e., the same as the reference sequence). If the number of sequenced nucleic acids in the subset containing a nucleotide variant exceeds a threshold, the variant nucleotide can be called at the designated position. The threshold value can be a base number, for example, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 sequenced nucleic acids in the subset that contain nucleotide variants, or, among other possibilities, the ratio of sequenced nucleic acids in the subset that contain nucleotide variants, for example, at least 0.5, 1, 2, 3, 4, 5, 10, 15, or 20. Comparison can be repeated at any designated position of interest in the reference sequence. Sometimes, comparison can be performed for designated positions that occupy at least 20, 100, 200, or 300 consecutive positions in the reference sequence, for example, 20 to 500, or 50 to 300 consecutive positions. [Example]
[0206] Example 1 Liquid biopsy wild-type prediction of negative predictors for anti-EGFR therapy in advanced colorectal cancer (CRC)
[0207] method
[0208] An analytical method was developed for the Guardant360® ctDNA test (Guardant Health, Redwood City, CA), which jointly analyzes estimated tumor fraction and the presence of mutually exclusive mutations to provide a yes / no / inevaluable wild-type status for clonal activating RAS / RAF mutations.
[0209] result
[0210] To validate the reliability of this method and model in determining clonal wild-type status, we used a subset of patient samples with CRC and known positive RAS / RAF mutation status by tissue sequencing who underwent clinical Guardant360 testing (n=98). Seventy-nine samples had concordant detection of RAS / RAF, while 19 had no RAS / RAF detected by Guardant360, which could be used to confirm model predictions. The model correctly identified all 19 samples as unevaluable for wild-type status and did not provide a wild-type call with high confidence regarding the presence of a known RAS / RAF mutation. To assess overall performance, we applied this method to a cohort of over 8,500 patient samples with CRC, allowing us to determine with high confidence whether the patient had a RAS / RAF mutant (40.7%) or clonal wild-type status (21.3%), significantly expanding the cohort of patients for whom a final determination of RAS / RAF status can be readily achieved through ctDNA testing.
[0211] conclusion
[0212] The Guardant360 ctDNA test can reliably determine the wild-type status of RAS / RAF genes in the majority of patients with advanced CRC and can reliably guide anti-EGFR therapy decisions.
[0213] Example 2 Mutual exclusivity and mutational co-occurrence observed in liquid biopsies of advanced cancer
[0214] Introduction
[0215] Somatic mutations in patients with untreated solid tumors tend to be clonal and often show histology-specific stereotypes of mutation occurrence. For example, in untreated non-small cell lung cancer (NSCLC) patients, EGFR exon 19 deletions have not been observed to co-occur with other driver mutations, such as MET exon 14 skipping deletions or EML4-ALK fusions (TCGA, 2017). In contrast, tumors from patients with previously treated disease are subjected to different biological and drug environments that affect their tumor biology and mutation patterns. Using the Guardant360 cell-free circulating tumor DNA (ctDNA) plasma assay, we characterized mutation patterns in a very large cohort of advanced NSCLC and colorectal (CRC) cancers.
[0216] method
[0217] De-identified results from patients with advanced NSCLC (n = 59,589) and CRC (n = 13,116) who underwent the clinical Guardant360 trial (Guardant Health, Redwood City, CA) were used for analyses of mutual exclusivity and variant co-occurrence. Patients included both untreated and previously treated patients. Variants included in the analysis required at least 200 observations, each with a variant allele proportion greater than 0.01. Variants that met the criteria were assessed using Fisher's exact test and corrected for multiple testing by the Bonferroni method.
[0218] result
[0219] Previously reported tissue analysis findings of mutual exclusivity between known NSCLC drivers, such as EGFR exon 19 deletions, and MET exon 14 skipping alterations were confirmed in over 59,000 ctDNA samples from patients with advanced NSCLC. An additional 70 pairs of mutually exclusive mutations were discovered, including novel pairs with mutations in STK11, TERT, and BRAF (class 3), which were observed to be mutually exclusive with known NSCLC driver mutations. Similarly, exclusive co-occurrence of EGFR resistance mutations T790M and C797S with EGFR drivers was also observed, recapitulating the co-occurrence observed in TCGA. In CRC, a cancer type not classically known to have mutually exclusive driver mutations, analysis of over 13,000 cases identified previously undescribed mutual exclusivity between the variants BRAF V600E and APC R876. * , p<0.005. Additional pairs of specific mutually exclusive mutations were found in KRAS, BRAF, APC, and TP53.
[0220] conclusion
[0221] Using a very large cohort of advanced NSCLC and CRC patients tested by ctDNA plasma-based comprehensive genomic profiling, we confirmed previously reported patterns of mutually exclusive driver mutations and discovered novel patterns of co-occurrence and exclusivity. These results highlight the utility of ctDNA for identifying clinically relevant mutations and novel biological mutation patterns.
[0222] All patent applications, websites, other publications, accession numbers, and the like, cited above or below, are hereby incorporated by reference in their entirety for all purposes to the same extent as if each individual item were specifically and individually indicated to be incorporated by reference herein. Where different versions of a sequence are associated with an accession number at different times, the version associated with the accession number as of the effective filing date of this application is meant. The effective filing date means the earlier of the actual filing date or, if applicable, the filing date of a priority application referencing the accession number. Similarly, where different versions of a publication, website, or other publication are published at different times, the most recently published version as of the effective publication date of that application is meant unless otherwise indicated. Any feature, step, element, embodiment, or aspect of the present disclosure can be used in combination with any other unless specifically indicated otherwise. While the present disclosure has been described in some detail by way of illustration and example for purposes of clarity and understanding, it will be apparent that certain changes and modifications can be practiced within the scope of the appended claims. The present invention provides, for example, the following items. (Item 1) 1. A method for determining that a first variant of interest is clonally absent from a first locus in a cell-free deoxyribonucleic acid (cfDNA) sample of a human subject, comprising: accessing a plurality of sequence reads of the cfDNA sample; determining, based on the plurality of sequence reads, that the first variant was not detected at the first locus in the sample; generating a first likelihood value based on the probability that the first variant is absent at a clonal level and / or a second likelihood value based on the probability that the first variant is present at a clonal level; optionally determining a quantitative value based on said first likelihood value and / or said second likelihood value; comparing the quantitative value and / or the first likelihood value and / or the second likelihood value with a threshold; and determining, based on said comparison, that said first variant of interest is clonally absent from said first locus. A method comprising: (Item 2) generating the first likelihood value and the second likelihood value comprises: determining a tumor fraction estimate for the sample, wherein the first likelihood value and the second likelihood value are based on the tumor fraction estimate; The method according to item 1, comprising: (Item 3) determining said tumor fraction estimate; determining the maximum mutant allele frequency (MAX MAF) of tumor mutations in said sample; Item 3. The method according to item 2, comprising: (Item 4) 4. The method of claim 3, wherein determining the MAX MAF comprises determining a molecular count associated with the tumor mutation based on the plurality of sequence read data. (Item 5) generating the first likelihood value and the second likelihood value comprises: determining an allele frequency of at least a second variant, wherein said first likelihood value and said second likelihood value are further based on said allele frequency and said MAX MAF. Item 3. The method according to item 3, comprising: (Item 6) 6. The method of claim 5, further comprising comparing the allele frequency to a second threshold based on the MAX MAF, wherein determining that the first variant of interest is absent from the first locus at a clonal level is further based on comparing the MAF to the second threshold. (Item 7) determining the allele frequencies 6. The method of claim 5, further comprising determining a first molecule count associated with the first variant based on the plurality of sequence reads. (Item 8) The step of determining the quantitative value comprises: 6. The method of claim 5, comprising accessing covariate information indicative of historical prevalence of one or more variants that exhibit co-occurrence and / or mutual exclusivity with the first variant, wherein the quantitative value is based on the covariate information. (Item 9) 9. The method of claim 8, further comprising determining an occurrence of at least a second variant in the cfDNA sample, wherein the quantification value is further based on the covariate information. (Item 10) The step of determining the quantitative value comprises: 2. The method of claim 1, comprising accessing covariate information indicative of historical prevalence of one or more variants that exhibit co-occurrence and / or mutual exclusivity with the first variant, wherein the quantitative value is based on the covariate information. (Item 11) 11. The method of claim 10, further comprising determining an occurrence of at least a second variant in the cfDNA sample, wherein the quantification value is further based on the occurrence of the second variant. (Item 12) Item 10. The method of item 1, wherein the quantitative value is based on a ratio of the first likelihood value to the second likelihood value. (Item 13) 2. The method of claim 1, further comprising determining a confidence level that the first variant is not present at a clonal level in the cfDNA sample based on the quantification value. (Item 14) 2. The method of claim 1, further comprising determining a treatment plan for treating the disease in the human subject. (Item 15) Item 15. The method of item 14, wherein the disease is cancer. (Item 16) determining the occurrence of at least a second variant in the cfDNA sample; and adjusting the quantitative value based on the occurrence of the at least second variant in the cfDNA sample. Item 1, the method of claim 1 further comprising: (Item 17) 1. A method for determining, at least in part using a computer, that a first target nucleic acid variant is absent from a first locus in a cell-free nucleic acid (cfNA) sample obtained from a subject having a predetermined type of cancer, the method comprising: determining that the first target nucleic acid variant is not detected at the first locus in the cfNA sample; determining, by the computer, coverage of the first locus from sequence information generated from the cfNA sample; determining, by the computer, tumor fraction from sequence information generated from the cfNA sample; determining, by the computer, the probability that the first target nucleic acid variant is present at the first locus in the cfNA sample from the coverage and the tumor fraction to generate a quantitative value; and determining that the first target nucleic acid variant is not present at the first locus in the cfNA sample if the quantification value differs from a threshold value; A method comprising: (Item 18) 1. A method for determining, at least in part, using a computer, that a first target nucleic acid variant is absent from a first locus in a cell-free nucleic acid (cfNA) sample obtained from a subject, comprising: determining that the first target nucleic acid variant is not detectable in the cfNA sample obtained from the subject to generate a first test result; determining that at least a second target nucleic acid variant is detected in the cfNA sample obtained from the subject to generate a second test result; determining, by the computer, a first probability that the first target nucleic acid variant is not present in the cfNA sample given the second test results and / or a second probability that the first target nucleic acid variant is present in the cfNA sample given the second test results; generating, by the computer, a quantitative value using the first probability, the second probability, and / or a ratio thereof; and determining that the first target nucleic acid variant is not present at the first locus in the cfNA sample if the quantification value differs from a threshold value; A method comprising: (Item 19) 1. A method for determining, at least in part using a computer, that a first target nucleic acid variant is absent from a first locus in a cell-free nucleic acid (cfNA) sample obtained from a subject having a predetermined cancer type, comprising: determining that the first target nucleic acid variant is not detectable in the cfNA sample obtained from the subject; generating, by said computer, at least one tumor fraction-based value; generating, by said computer, at least one mutual exclusivity value; and using the tumor fraction-based value and / or the mutual exclusivity value to determine that the first target nucleic acid variant is absent from the first locus in the cfNA sample. A method comprising: (Item 20) 10. The method of any one of the preceding items, wherein the quantitative value is less than the threshold value. (Item 21) 10. The method of any one of the preceding items, wherein the quantitative value is greater than the threshold value. (Item 22) 10. The method of any one of the preceding items, wherein the first and second test results are dependent on each other. (Item 23) The method of any one of the preceding items, comprising determining that a plurality of other selected target nucleic acid variants are absent at one or more other loci. (Item 24) 10. The method of any one of the preceding items, wherein the quantitative value comprises a log-likelihood ratio (LLR) threshold. (Item 25) 10. The method of any one of the preceding items, comprising determining that the first target nucleic acid variant is absent at the first locus in a plurality of reference cfNA samples to generate a threshold value. (Item 26) 26. The method of item 25, wherein the threshold comprises a clonality or subclonality threshold. (Item 27) 10. The method of any one of the preceding items, wherein the first target nucleic acid variant comprises a driver mutation. (Item 28) The method of any one of the preceding items, further comprising administering one or more therapies to the subject based on a determination that the first target nucleic acid variant is not present at the first locus in the cfNA sample. (Item 29) The method of any one of the preceding items, comprising using the tumor fraction and a binomial model to estimate the probability of detecting the first target nucleic acid variant at a first locus in the cfNA sample. (Item 30) 30. The method of claim 29, wherein the binomial model includes information about a predetermined cancer type and / or a second target nucleic acid variant. (Item 31) 10. The method of any one of the preceding items, wherein a determination that the first target nucleic acid variant is not present at the first locus in the cfNA sample indicates that the first locus is wild type. (Item 32) 10. The method of any one of the preceding items, wherein the predetermined cancer type is colorectal cancer, the first locus is KRAS, BRAF, or NRAS, and a determination that the first target nucleic acid variant is not present at the first locus in the cfNA sample indicates that the first locus is wild-type KRAS, BRAF, or NRAS. (Item 33) 33. The method of item 32, further comprising administering cetuximab and / or panitumumab to the subject. (Item 34) The method of any one of the preceding items, wherein the cfNA comprises cfDNA. (Item 35) The method of any one of the preceding items, wherein the cfNA comprises cfRNA. (Item 36) The method of any one of the preceding items, further comprising repeating the method one or more times to monitor whether the first target nucleic acid variant is present at the first locus in different cfNA samples obtained from the subject at different time points. (Item 37) The method of any one of the preceding items, further comprising performing one or more additional tests to confirm or refute a determination that the first target nucleic acid variant is not present at the first locus in the cfNA sample. (Item 38) 10. The method of any one of the preceding items, comprising determining a maximum mutant allele frequency (MAX MAF) for said cfNA sample and using said MAX MAF as a tumor fraction estimate. (Item 39) 10. The method of any one of the preceding items, comprising determining, based on a plurality of sequence reads obtained from the cfNA sample, that the first target nucleic acid variant is not detected at the first locus in the cfNA sample. (Item 40) 10. The method of any one of the preceding items, comprising determining that the first target nucleic acid variant is absent at the clonal level in the cfNA sample. (Item 41) 10. The method of any one of the preceding items, further comprising generating a first likelihood value based on the first probability and a second likelihood value based on the second probability. (Item 42) 10. The method of any one of the preceding items, comprising determining a quantitative value based on the first likelihood value and the second likelihood value. (Item 43) 10. The method of claim 1, wherein generating the first likelihood value and the second likelihood value comprises determining a tumor fraction estimate for the cfNA sample, and wherein the first likelihood value and the second likelihood value are based on the tumor fraction estimate. (Item 44) 44. The method of claim 43, wherein determining the tumor fraction estimate comprises determining the maximum mutant allele frequency (MAX MAF) of tumor mutations in the cfNA sample. (Item 45) 45. The method of claim 44, wherein determining the MAX MAF comprises determining a molecular count associated with a tumor mutation based on the plurality of sequence read data. (Item 46) 46. The method of claim 45, wherein generating the first likelihood value and the second likelihood value comprises determining an allele frequency of at least a second variant, wherein the first likelihood value and the second likelihood value are further based on the allele frequency and the MAX MAF. (Item 47) 47. The method of claim 46, further comprising comparing the allele frequency to a second threshold based on the MAX MAF, wherein determining that the first target nucleic acid variant of interest is absent at the first locus at a clonal level is further based on comparing the MAF to the second threshold. (Item 48) 47. The method of claim 46, wherein determining the allele frequency comprises determining a first molecule count associated with the first target nucleic acid variant based on the plurality of sequence reads. (Item 49) 47. The method of claim 46, wherein determining the quantitative value comprises accessing covariate information indicative of historical prevalence of one or more variants that exhibit co-occurrence and / or mutual exclusivity with the first variant, and wherein the quantitative value is based on the covariate information. (Item 50) 50. The method of claim 49, further comprising determining an occurrence of at least the second target nucleic acid variant in the cfDNA sample, wherein the quantification value is further based on the covariate information. (Item 51) 43. The method of claim 42, wherein determining the quantitative value comprises accessing covariate information indicative of historical prevalence of one or more variants that exhibit co-occurrence and / or mutual exclusivity with the first target nucleic acid variant, and wherein the quantitative value is based on the covariate information. (Item 52) 52. The method of claim 51, further comprising determining the occurrence of at least the second target nucleic acid variant in the cfNA sample, wherein the quantification value is further based on the occurrence of the second target nucleic acid variant. (Item 53) Item 43. The method of item 42, wherein the quantitative value is based on a ratio of the first likelihood value to the second likelihood value. (Item 54) 43. The method of claim 42, further comprising determining a confidence level that the first target nucleic acid variant is not present at a clonal level in the cfNA sample based on the quantification value. (Item 55) 43. The method of claim 42, further comprising determining the occurrence of at least the second target nucleic acid variant in the cfNA sample and adjusting the quantitative value based on the occurrence of at least the second target nucleic acid variant in the cfNA sample. (Item 56) 10. The method of any one of the preceding items, wherein the ratio comprises a log-posterior probability ratio (LPPR) equal to the sum of a log-likelihood tumor fraction value, a log-likelihood mutual exclusivity value, and a log-prior value. (Item 57) 10. The method of any one of the preceding items, wherein the first locus or the second locus comprises the second target nucleic acid variant. (Item 58) The method of any one of the preceding items, wherein the quantitative value comprises a negative predictive value (NPV) score. (Item 59) 10. The method of any one of the preceding items, wherein the predetermined cancer type comprises lung cancer and the first target nucleic acid variant is a mutation in a gene selected from the group consisting of EGFR, BRAF, ALK, ROS1, and MET. (Item 60) 10. The method of any one of the preceding items, wherein the predetermined type of cancer comprises colorectal cancer and the first target nucleic acid variant is a mutation in a gene selected from the group consisting of KRAS, BRAF, and NRAS. (Item 61) When executed by at least one electronic processor: accessing a plurality of sequence reads of the cfDNA sample; determining, based on the plurality of sequence reads, that a first variant was not detected at a first locus in the sample; generating a first likelihood value based on the probability that the first variant is absent at a clonal level and a second likelihood value based on the probability that the first variant is present at a clonal level; determining a quantitative value based on the first likelihood value and the second likelihood value; comparing the quantified value to a threshold value; and determining, based on said comparison, that said first variant of interest is clonally absent from said first locus. 1. A system including a controller that includes or has access to a computer-readable medium that includes non-transitory computer-executable instructions that at least implement the above. (Item 62) When executed by at least one electronic processor: accessing sequence information generated from a cell-free nucleic acid (cfNA) sample obtained from a subject having a predetermined type of cancer; determining from the sequence information that a first target nucleic acid variant is not detected at a first locus in the cfNA sample; determining coverage of the first locus from the sequence information; determining the tumor proportion from the sequence information; determining the probability that the first target nucleic acid variant is present at the first locus in the cfNA sample from the coverage and the tumor fraction to generate a quantitative value; and determining that the first target nucleic acid variant is not present at the first locus in the cfNA sample if the quantification value differs from a threshold value. 1. A system including a controller that includes or has access to a computer-readable medium that includes non-transitory computer-executable instructions that at least implement the above. (Item 63) When executed by at least one electronic processor: accessing sequence information generated from a cell-free nucleic acid (cfNA) sample obtained from the subject; determining from the sequence information that a first target nucleic acid variant is not detected in the cfNA sample to generate a first test result; determining from said sequence information that at least a second target nucleic acid variant is detected in said cfNA sample to generate a second test result; determining a first probability that the first target nucleic acid variant is not present in the cfNA sample given the second test result and / or a second probability that the first target nucleic acid variant is present in the cfNA sample given the second test result; generating a quantitative value using the first probability, the second probability, and / or a ratio thereof; and determining that the first target nucleic acid variant is not present at the first locus in the cfNA sample if the quantification value differs from a threshold value; 1. A system including a controller that includes or has access to a computer-readable medium that includes non-transitory computer-executable instructions that at least implement the above. (Item 64) When executed by at least one electronic processor: accessing sequence information generated from a cell-free nucleic acid (cfNA) sample obtained from the subject; determining from the sequence information that a first target nucleic acid variant is not detected in the cfNA sample; generating at least one tumor fraction-based value; generating at least one mutual exclusivity value; and using the tumor fraction-based value and / or the mutual exclusivity value to determine that the first target nucleic acid variant is absent from the first locus in the cfNA sample. 1. A system including a controller that includes or has access to a computer-readable medium that includes non-transitory computer-executable instructions that at least implement the above. (Item 65) When executed by at least an electronic processor: accessing a plurality of sequence reads of the cfDNA sample; determining, based on the plurality of sequence reads, that a first variant was not detected at a first locus in the sample; generating a first likelihood value based on the probability that the first variant is absent at a clonal level and a second likelihood value based on the probability that the first variant is present at a clonal level; determining a quantitative value based on the first likelihood value and the second likelihood value; comparing the quantified value to a threshold value; and determining, based on said comparison, that said first variant of interest is clonally absent from said first locus. 1. A computer-readable medium comprising non-transitory computer-executable instructions that at least implement the method of claim 1. (Item 66) When executed by at least an electronic processor: accessing sequence information generated from a cell-free nucleic acid (cfNA) sample obtained from a subject having a predetermined type of cancer; determining from the sequence information that a first target nucleic acid variant is not detected at a first locus in the cfNA sample; determining coverage of the first locus from the sequence information; determining the tumor proportion from the sequence information; determining the probability that the first target nucleic acid variant is present at the first locus in the cfNA sample from the coverage and the tumor fraction to generate a quantitative value; and determining that the first target nucleic acid variant is not present at the first locus in the cfNA sample if the quantification value differs from a threshold value; 1. A computer-readable medium comprising non-transitory computer-executable instructions that at least implement the method of claim 1. (Item 67) When executed by at least an electronic processor: accessing sequence information generated from a cell-free nucleic acid (cfNA) sample obtained from the subject; determining from the sequence information that a first target nucleic acid variant is not detected in the cfNA sample to generate a first test result; determining from said sequence information that at least a second target nucleic acid variant is detected in said cfNA sample to generate a second test result; determining a first probability that the first target nucleic acid variant is not present in the cfNA sample given the second test result, and / or a second probability that the first target nucleic acid variant is present in the cfNA sample given the second test result; generating a quantitative value using the first probability, the second probability, and / or a ratio thereof; and determining that the first target nucleic acid variant is not present at the first locus in the cfNA sample if the quantification value differs from a threshold value; 1. A computer-readable medium comprising non-transitory computer-executable instructions that at least implement the method of claim 1. (Item 68) When executed by at least an electronic processor: accessing sequence information generated from a cell-free nucleic acid (cfNA) sample obtained from the subject; determining from the sequence information that a first target nucleic acid variant is not detected in the cfNA sample; generating at least one tumor fraction-based value; generating at least one mutual exclusivity value; and using the tumor fraction-based value and / or the mutual exclusivity value to determine that the first target nucleic acid variant is absent from the first locus in the cfNA sample. 1. A computer-readable medium comprising non-transitory computer-executable instructions that at least implement the method of claim 1. (Item 69) 10. The system or computer-readable medium of any one of the preceding items, wherein the quantitative value is less than the threshold value. (Item 70) 10. The system or computer-readable medium of any one of the preceding items, wherein the quantitative value is greater than the threshold value. (Item 71) 10. The system or computer-readable medium of any one of the preceding items, wherein the first and second test results are dependent on each other. (Item 72) 10. The system or computer-readable medium of any one of the preceding items, comprising determining that a plurality of other selected target nucleic acid variants are absent from one or more other loci. (Item 73) 10. The system or computer-readable medium of any one of the preceding items, wherein the quantitative value comprises a log-likelihood ratio (LLR) threshold. (Item 74) 10. The system or computer-readable medium of any one of the preceding items, comprising determining that the first target nucleic acid variant is absent at the first locus in a plurality of reference cfNA samples to generate the threshold value. (Item 75) 75. The system or computer-readable medium of item 74, wherein the threshold comprises a clonality or subclonality threshold. (Item 76) 10. The system or computer-readable medium of any one of the preceding items, wherein the first target nucleic acid variant comprises a driver mutation. (Item 77) 10. The system or computer-readable medium of any one of the preceding items, wherein the instructions further perform at least the step of: outputting one or more therapy recommendations for the subject based on a determination that the first target nucleic acid variant is not present at the first locus in the cfNA sample. (Item 78) 10. The system or computer-readable medium of any one of the preceding items, wherein the instructions further perform at least the step of estimating the probability of detecting the first target nucleic acid variant at the first locus in the cfNA sample using the tumor fraction and a binomial model. (Item 79) The instructions further comprise: determining a maximum mutant allele frequency (MAX) for the cfNA sample; 10. The system or computer-readable medium of any one of the preceding items, further comprising at least the steps of determining a MAX MAF (maximum average fraction of tumors) and using the MAX MAF as the tumor fraction estimate. (Item 80) 10. The system or computer-readable medium of any one of the preceding items, wherein the instructions further perform at least the step of determining that the first target nucleic acid variant is absent at a clonal level in the cfNA sample. (Item 81) 10. The system or computer-readable medium of claim 1, wherein the instructions further perform at least the step of generating a first likelihood value based on the first probability and a second likelihood value based on the second probability. (Item 82) 10. The system or computer-readable medium of any one of the preceding items, wherein the instructions further perform at least the step of determining the quantitative value based on the first likelihood value and the second likelihood value. (Item 83) 10. The system or computer-readable medium of any one of the preceding items, wherein the instructions further perform at least the step of generating the first likelihood value and the second likelihood value by determining the tumor fraction estimate for the cfNA sample, wherein the first likelihood value and the second likelihood value are based on the tumor fraction estimate. (Item 84) 84. The system or computer-readable medium of claim 83, wherein the instructions further perform at least the step of determining the tumor fraction estimate by determining a maximum mutant allele frequency (MAX MAF) of tumor mutations in the cfNA sample. (Item 85) 85. The system or computer-readable medium of claim 84, wherein the instructions further perform at least the step of determining the MAX MAF by determining a molecular count associated with the tumor mutation based on the plurality of sequence reads. (Item 86) 85. The system or computer-readable medium of claim 84, wherein the instructions further perform at least the step of generating the first likelihood value and the second likelihood value by determining an allele frequency of at least a second variant, wherein the first likelihood value and the second likelihood value are further based on the allele frequency and the MAX MAF. (Item 87) 87. The system or computer-readable medium of Item 86, wherein the instructions further perform at least the steps of: comparing the allele frequency to the second threshold based on the MAX MAF; and determining that the first target nucleic acid variant of interest is absent at the first locus at a clonal level, further based on a comparison of the MAF to the second threshold. (Item 88) 87. The system or computer-readable medium of claim 86, wherein the instructions further perform at least the step of determining the allele frequency by determining a first molecule count associated with the first target nucleic acid variant based on the plurality of sequence reads. (Item 89) 87. The system or computer-readable medium of claim 86, wherein the instructions further perform at least the step of determining the quantitative value by accessing covariate information indicative of historical prevalence of one or more variants that exhibit co-occurrence and / or mutual exclusivity with the first variant, wherein the quantitative value is based on the covariate information. (Item 90) 90. The system or computer-readable medium of claim 89, wherein the instructions further perform at least the step of determining an occurrence of at least the second target nucleic acid variant in the cfDNA sample, wherein the quantification value is further based on the covariate information. (Item 91) 84. The system or computer-readable medium of claim 83, wherein the instructions further perform at least the step of determining the quantitative value by accessing covariate information indicative of a historical prevalence of one or more variants that exhibit co-occurrence and / or mutual exclusivity with the first target nucleic acid variant, wherein the quantitative value is based on the covariate information. (Item 92) 92. The system or computer-readable medium of claim 91, wherein the instructions further perform at least the step of determining an occurrence of at least the second target nucleic acid variant in the cfNA sample, wherein the quantification value is further based on the occurrence of the second target nucleic acid variant. (Item 93) 84. The system or computer-readable medium of claim 83, wherein the instructions further perform at least the step of determining a confidence level that the first target nucleic acid variant is absent at a clonal level in the cfNA sample based on the quantification value. (Item 94) 84. The system or computer-readable medium of claim 83, wherein the instructions further perform at least the steps of determining an occurrence of at least the second target nucleic acid variant in the cfNA sample and adjusting the quantitative value based on the occurrence of at least the second target nucleic acid variant in the cfNA sample. (Item 95) 10. The system or computer-readable medium of any one of the preceding items, wherein the ratio comprises a log-posteriori probability ratio (LPPR) equal to the sum of a log-likelihood tumor fraction value, a log-likelihood mutual exclusivity value, and a log-priori value. (Item 96) 10. The method or system of any one of the preceding items, further comprising generating a report optionally comprising information regarding the absence of said first target nucleic acid variant at said first locus in said sample and / or information derived therefrom. (Item 97) 97. The method or system of claim 96, further comprising the step of communicating the report to the subject from whom the sample was derived or to a third party, such as a medical professional.
Claims
1. 1. A computer-implemented method for determining that a first variant of interest is clonally absent from a first locus in a cell-free deoxyribonucleic acid (cfDNA) sample of a human subject, the method comprising: accessing a plurality of sequence reads of the cfDNA sample; determining, based on the plurality of sequence reads, that the first variant of interest was not detected at the first locus in the cfDNA sample; generating a first likelihood value based on a first probability, the first probability being the probability that the first variant of interest is absent at a clonal level, and / or a second likelihood value based on a second probability, the second probability being the probability that the first variant of interest is present at a clonal level; determining a quantitative value, including a negative predictive value (NPV) score, based on the first likelihood value and / or the second likelihood value; comparing the quantitative value and / or the first likelihood value and / or the second likelihood value with a threshold; and determining, based on said comparison, that said first variant of interest is clonally absent from said first locus. wherein the first likelihood value and the second likelihood value are based on a tumor fraction estimate for the cfDNA sample.
2. generating the first likelihood value and the second likelihood value comprises: determining the tumor fraction estimate for the cfDNA sample, wherein the tumor fraction estimate is determined by determining the maximum mutant allele frequency (MAX MAF) of tumor mutations in the cfDNA sample; The method of claim 1 , comprising:
3. 3. The method of claim 2, wherein determining the MAX MAF comprises determining a molecular count associated with the tumor mutation based on the plurality of sequence reads.
4. generating the first likelihood value and the second likelihood value comprises: determining an allele frequency of at least a second variant, wherein the first likelihood value and the second likelihood value are further based on the allele frequency and the MAX MAF, and determining that the first variant of interest is absent at the first locus at a clonal level comprises determining the allele frequency of the second variant and the MAX MAF. Further based on a comparison with a second threshold based on the MAF. The method of claim 3, comprising:
5. determining the allele frequencies determining a first molecule count associated with the first variant of interest based on the plurality of sequence reads; The method of claim 4.
6. determining the quantitative value comprises:
6. The method of any one of claims 1 to 5, comprising accessing covariate information indicative of the historical prevalence of one or more variants that exhibit co-occurrence and / or mutual exclusivity with the first variant of interest, wherein said quantitative value is based on said covariate information.
7. (i) determining the occurrence of at least a second variant in the cfDNA sample, wherein the quantification value is further based on the occurrence of the second variant; or (ii) determining the occurrence of at least a second variant in the cfDNA sample, wherein the quantitative value is further based on the covariate information. The method of claim 6.
8. The method of claim 1 , wherein the quantitative value is based on a ratio of the first likelihood value to the second likelihood value.
9. 2. The method of claim 1, further comprising determining a confidence level that the first variant of interest is absent at a clonal level in the cfDNA sample based on the quantification value.
10. 10. The method of claim 1, further comprising determining a treatment plan for treating the disease in the human subject.
11. determining the occurrence of at least a second variant in the cfDNA sample; and adjusting the quantitative value based on the occurrence of the at least second variant in the cfDNA sample. The method of claim 1 further comprising:
12. 1. A system comprising: When executed by at least one electronic processor: accessing a plurality of sequence reads of the cfDNA sample; determining, based on the plurality of sequence reads, that a first variant of interest was not detected at a first locus in the cfDNA sample; generating a first likelihood value based on a first probability, the first probability being the probability that the first variant of interest is absent at a clonal level, and a second likelihood value based on a second probability, the second probability being the probability that the first variant of interest is present at a clonal level; determining a quantitative value comprising a negative predictive value (NPV) score based on the first likelihood value and the second likelihood value; comparing the quantitative value to a threshold value; and determining, based on said comparison, that said first variant of interest is clonally absent from said first locus. wherein the first likelihood value and the second likelihood value are based on a tumor fraction estimate for the cfDNA sample.
13. 1. A computer-readable medium, comprising: When executed by at least an electronic processor: accessing a plurality of sequence reads of the cfDNA sample; determining, based on the plurality of sequence reads, that a first variant of interest was not detected at a first locus in the cfDNA sample; generating a first likelihood value based on a first probability, the first probability being the probability that the first variant of interest is absent at a clonal level, and a second likelihood value based on a second probability, the second probability being the probability that the first variant of interest is present at a clonal level; determining a quantitative value comprising a negative predictive value (NPV) score based on the first likelihood value and the second likelihood value; comparing the quantitative value to a threshold value; and determining, based on said comparison, that said first variant of interest is clonally absent from said first locus. wherein the first likelihood value and the second likelihood value are based on a tumor fraction estimate for the cfDNA sample.
14. (a) the quantitative value is less than the threshold value, or the quantitative value is greater than the threshold value, and / or (b) the quantitative value comprises a log-likelihood ratio (LLR) threshold; The method according to any one of claims 1 to 11.
15. (a) generating a first likelihood value based on the first probability and a second likelihood value based on the second probability; and / or (b) determining the quantitative value based on the first likelihood value and the second likelihood value; The method of any one of claims 1 to 11 and 14, further comprising:
16. 9. The method of claim 8, wherein the ratio comprises a log-posterior probability ratio (LPPR) equal to the sum of a log-likelihood tumor fraction value, a log-likelihood mutual exclusivity value, and a log-prior value.
17. 17. The method of any one of claims 1 to 11 and 14 to 16, further comprising generating a report optionally comprising information regarding the absence of said first variant of interest at said first locus in said cfDNA sample and / or information derived therefrom, and optionally said method further comprising communicating said report to a third party. (a) the quantitative value is less than the threshold value, or the quantitative value is greater than the threshold value, and / or (b) the quantitative value comprises a log-likelihood ratio (LLR) threshold; The system of claim 12.
19. The method of claim 19, wherein the instructions further comprise generating a first likelihood value based on the first probability and a second likelihood value based on the second probability; and / or (b) the instructions further determine the quantitative value based on the first likelihood value and the second likelihood value.
20. The system according to claim 12 or 18, which performs at least the following:
20. A system as described in any one of claims 12, 18, and 19, further comprising the step of generating a report optionally including information regarding the absence of the first variant of interest at the first locus in the cfDNA sample and / or information derived therefrom, and optionally the system further comprising the step of communicating the report to a third party. (a) the quantitative value is less than the threshold value, or the quantitative value is greater than the threshold value, and / or (b) the quantitative value comprises a log-likelihood ratio (LLR) threshold; The computer-readable medium of claim 13.
22. (a) the instructions further include generating a first likelihood value based on the first probability and a second likelihood value based on the second probability; and / or (b) the instructions further determine the quantitative value based on the first likelihood value and the second likelihood value.
22. The computer-readable medium of claim 13 or 21, which performs at least the steps of:
Citation Information
Patent Citations
Improvements in variant detection
WO2019170773A1