Methods of tumor-informed detection of ctdna using tumor-only somatic calling
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-02-09
- Publication Date
- 2026-08-13
Smart Images

Figure US2026014575_13082026_PF_FP_ABST
Abstract
Description
Atty. Dkt. No.: 131588-1668METHODS OF TUMOR-INFORMED DETECTION OF CTDNA USING TUMOR-ONLY SOMATIC CALLINGCROSS-REFERENCE TO RELATED PATENT APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Application Serial No.63 / 756,701, filed February 10, 2025. The entire disclosure of the prior application is incorporated by reference herein.BACKGROUND[00021 The following description of the background of the present technology is provided simply as an aid in understanding the present technology and is not admitted to describe or constitute prior art to the present technology.
[0003] The discovery of cell free deoxyribonucleic acid has promoted the non-invasive detection of alterations in genomic sequences that occur in various disease states. However, in some instances, e.g., cancer, the ability to determine the presence of disease by detecting disease-associated mutations has been hindered by the extremely low levels of cell free tumor DNA, operational expense to perform whole genome sequencing on tumor and matched nontumor samples, risk of mis-match of tumor and non-tumor samples. Methods that allow for the accurate detection of disease-associated mutations remain needed.SUMMARY
[0004] The present disclosure provides methods of tumor-informed minimal residual disease (MRD) assay using only a tumor sample (i.e., without a germline sample). The disclosed methods can utilize variant calling models coupled with post-sequencing error correction to accurately select tumor-specific and patient-specific somatic mutations that can be used in a panel for MRD assays. Among other advantages, the disclosed methods reduce the amount of input data needed to prepare a personalized MRD signature panel by using only one sample and allow for the selection of tumor-specific somatic variants that are best suited for patients tracking (i.e., least likely to cause false positives or false negatives resulting from sequencing errors, amplification errors, or biological artifacts).Atty. Dkt. No.: 131588-1668[00051 In one aspect, the present disclosure provides methods of detecting circulating tumor DNA (ctDNA) in a sample, comprising: (a) sequencing DNA from a tumor sample obtained from a patient, thereby obtaining a first set of sequencing reads of the DNA from the tumor sample; (b) determining a set of putative tumor-specific somatic variants comprising a plurality of putative somatic variant sites; (c) sequencing cell-free DNA (cfDNA) from a sample of blood, plasma, or serum from the cancer patient, thereby obtaining a second set of sequencing reads; and (d) detecting ctDNA sequences in the second set of sequencing reads based on a sequencing read comprising one or more putative somatic variant sites from the set of putative tumor-specific somatic variants.[0006J In some embodiments, determining a set of putative tumor-specific somatic variants is based on allele balance, copy number alteration (CNA), or a combination thereof for each putative somatic variant site in the first set of sequencing reads. In some embodiments, determining a set of putative tumor-specific somatic variants comprises training a machine learning model to select the plurality of putative somatic variant sites. In some embodiments, determining a set of putative tumor-specific somatic variants comprises calling, via a computer processor, the plurality of putative somatic variant sites based on one or more assemblies of haplotypes.
[0007] In some embodiments, the method does not comprise sequencing DNA from a nontumor sample from the patient for determining the set of tumor-specific somatic variants.
[0008] In some embodiments, the cfDNA from the sample of blood, plasma, or serum from the cancer patient is not enriched prior to sequencing.
[0009] In some embodiments, determining the set of tumor-specific somatic variants further comprises determining tumor purity of the tumor sample.[0010| In some embodiments, sequencing DNA from the tumor sample comprises whole genome sequencing, whole exome sequencing, targeted sequencing, or subtractive hybridization.
[0011] In some embodiments, detecting the presence or absence of ctDNA sequences in the second set of sequencing reads comprises correcting for, using a computer processor,Atty. Dkt. No.: 131588-1668inclusion of latent variant classes in the second set of sequencing reads. In some embodiments, correcting for inclusion of latent variant classes in the second set of sequencing reads comprises comparing the second set of sequencing reads to a reference dataset.
[0012] In some embodiments, sequencing the cfDNA comprises whole genome sequencing , whole exome sequencing, targeted sequencing, or subtractive hybridization.
[0013] In some embodiments, detecting the presence or absence of ctDNA sequences in the second set of sequencing reads comprises applying a classification model to the putative somatic variant sites to classify each as either tumor-specific, non-tumor-specific, or germline variant.
[0014] In some embodiments, the methods may further comprise determining a probability that each putative somatic variant site belongs to the class based on relative likelihoods of different binomial models, gaussian models, or Poisson models. In some embodiments, the non-tumor somatic variant and / or germline variant classes use a fixed parameter for the probability of observing a variant count.
[0015] In some embodiments, each putative somatic variant site in the set of putative tumorspecific somatic variants belongs to a class selected from (i) a tumor-specific somatic variant, (ii) a non-tumor-specific somatic variant or mutation due to technical or biological noise, and (iii) a germline variant incorrectly identified as somatic.[0016J In some embodiments, a probability of observing an alternate allele count for each class is modeled as a statistical distribution. In some embodiments, the statistical distribution is a binomial distribution, a negative binomial distribution, gaussian distribution, or Poisson distribution. In some embodiments, the statistical distribution includes a probability parameter determined by analysis of one or more reference sets of non-tumor somatic variants and / or germline variants.
[0017] In some embodiments, the methods may further comprise calculating a total likelihood of observing variant and non-variant counts for each putative somatic variant site.Atty. Dkt. No.: 131588-1668[0O18| In some embodiments, the methods may further comprise calculating the fraction of ctDNA present in the sample of cfDNA. In some embodiments, calculating the fraction of ctDNA comprises fitting a mixture model of variant counts and total counts across the entire panel of variants assayed, where the mixture components are the tumor, non-tumor, or germline variants and the weighting of each class is determined by the probability that a variant belongs to that class. In some embodiments, the mixture model is a binomial mixture model, a negative binomial mixture model, gaussian mixture model, or Poisson mixture model.
[0019] In some embodiments, detecting the presence or absence of ctDNA sequences in the second set of sequencing reads comprises: (i) calculating, using a computer processor, an allele balance of the somatic variants that is optionally adjusted based on purity of the tumor sample obtained from the subject, (ii) calculating, using a computer processor, a copy number for each somatic variant, (iii) correcting for, using a computer processor, an error rate for each somatic variant, or (iv) any combination thereof. In some embodiments, calculating the copy number for each somatic variant comprises a copy number for each somatic variant in the tumor sequencing data, the liquid sample sequencing data, or both the tumor sequencing data and the liquid sample sequencing data. In some embodiments, correcting for the error rate comprises an error rate for each somatic variant based on the liquid sample sequencing data, training data, or both the liquid sample sequencing data and training data.
[0020] In some embodiments, detecting the presence or absence of ctDNA sequences in the second set of sequencing reads comprises calculating, using a computer processor, a dropout rate of sites that do not contribute to a ctDNA signal, wherein the sites are not non-tumor, noise, or germline.1 021] In another aspect, the present disclosure provides methods of detecting circulating tumor DNA (ctDNA) in a sample, comprising: sequencing DNA from a tumor sample obtained from a patient, thereby obtaining a first set of sequencing reads of the DNA from the tumor sample; determining a set of putative tumor-specific somatic variants comprising a plurality of putative somatic variant sites; and sequencing cell-free DNA (cfDNA) from a sample of blood, plasma, or serum from the cancer patient, thereby obtaining a second set ofAtty. Dkt. No.: 131588-1668sequencing reads; and detecting ctDNA sequences in the second set of sequencing reads based on a sequencing read comprising one or more putative somatic variant sites from the set of putative tumor-specific somatic variants, wherein detecting comprises classifying each putative somatic variant site as (i) a tumor-specific somatic variant, (ii) a non-tumor-specific somatic variant or mutation due to technical or biological noise, and (iii) a germline variant incorrectly identified as somatic.
[0022] In some embodiments, classifying each putative somatic variant site as (i) a tumorspecific somatic variant further comprises adjusting a likelihood of the classification, using a computer processor, based on a dropout rate of sites that do not contribute to a ctDNA signal, wherein the sites are not non-tumor, noise, or germline.
[0023] In some embodiments, determining a set of putative tumor-specific somatic variants is based on allele balance, copy number alteration (CNA), or a combination thereof for each putative somatic variant site in the first set of sequencing reads.
[0024] In some embodiments, determining a set of putative tumor-specific somatic variants comprises training a machine learning model to select the plurality of putative somatic variant sites. In some embodiments, determining a set of putative tumor-specific somatic variants comprises calling, via a computer processor, the plurality of putative somatic variant sites based on one or more assemblies of haplotypes.
[0025] In some embodiments, the method does not comprise sequencing DNA from a nontumor sample from the patient for determining the set of tumor-specific somatic variants.
[0026] In some embodiments, the cfDNA from the sample of blood, plasma, or serum from the cancer patient is not enriched prior to sequencing.
[0027] In some embodiments, detecting ctDNA sequences in the second set of sequencing reads further comprises: (i) calculating, using a computer processor, an allele balance of the somatic variants that is optionally adjusted based on purity of the tumor sample obtained from the subject, (ii) calculating, using a computer processor, a copy number for each somatic variants, (iii) correcting for, using a computer processor, an error rate for each somatic variant, or (iv) any combination thereof. In some embodiments, calculating the copyAtty. Dkt. No.: 131588-1668number for each somatic variant comprises a copy number for each somatic variant in the tumor sequencing data, the liquid sample sequencing data, or both the tumor sequencing data and the liquid sample sequencing data.
[0028] In some embodiments, correcting for the error rate comprises an error rate for each somatic variant based on the liquid sample sequencing data, training data, or both the liquid sample sequencing data and training data.[0029| In some embodiments, detecting the presence or absence of ctDNA sequences in the second set of sequencing reads comprises correcting for, using a computer processor, inclusion of latent variant classes in the second set of sequencing reads. In some embodiments, correcting for inclusion of latent variant classes in the second set of sequencing reads comprises comparing the second set of sequencing reads to a reference dataset.
[0030] In another aspect, the present disclosure provides methods of detecting circulating tumor DNA (ctDNA) in a sample, comprising: sequencing cell-free DNA (cfDNA) from a sample of blood, plasma, or serum from a cancer patient, thereby obtaining a set of sequencing reads, wherein the cfDNA from the sample of blood, plasma, or serum from the cancer patient is not enriched prior to sequencing; and detecting ctDNA sequences in the set of sequencing reads based on a sequencing read comprising one or more putative somatic variant sites from a set of putative tumor-specific somatic variants, wherein detecting comprises (a) classifying each putative somatic variant site as (i) a tumor-specific somatic variant, (ii) a non-tumor-specific somatic variant, and (iii) a germline variant incorrectly identified as somatic, and (b) adjusting variant calls, using a computer processor, based on a dropout rate of sites that do not contribute to a ctDNA signal, wherein the sites are not nontumor, noise, or germline.
[0031] In some embodiments, the set of putative tumor-specific somatic variants is tumor-informed based on based on allele balance, copy number alteration (CNA), or a combination thereof of sequence reads of DNA from a tumor sample of the cancer patient without sequencing DNA from a non-tumor sample of the cancer patient.Atty. Dkt. No.: 131588-1668[00321 In some embodiments, detecting ctDNA sequences in the second set of sequencing reads further comprises: (i) calculating, using a computer processor, an allele balance of the somatic variants that is optionally adjusted based on purity of the tumor sample obtained from the subject, (ii) calculating, using a computer processor, a copy number for each somatic variants, (iii) correcting for, using a computer processor, an error rate for each somatic variant, or (iv) any combination thereof. In some embodiments, calculating the copy number for each somatic variant comprises a copy number for each somatic variant in the tumor sequencing data, the liquid sample sequencing data, or both the tumor sequencing data and the liquid sample sequencing data.
[0033] In some embodiments, correcting for the error rate comprises an error rate for each somatic variant based on the liquid sample sequencing data, training data, or both the liquid sample sequencing data and training data.
[0034] In some embodiments, detecting the presence or absence of ctDNA sequences in the second set of sequencing reads comprises correcting for, using a computer processor, inclusion of latent variant classes in the second set of sequencing reads. In some embodiments, correcting for inclusion of latent variant classes in the second set of sequencing reads comprises comparing the second set of sequencing reads to a dataset.
[0035] In some embodiments, classifying each putative somatic variant site comprises modeling a probability of observing an alternate allele count for each class as a statistical distribution. In some embodiments, the statistical distribution is a binomial distribution, a negative binomial distribution, gaussian distribution, or Poisson distribution. In some embodiments, the statistical distribution includes a probability parameter determined by analysis of one or more reference sets of non-tumor somatic variants and / or germline variants.
[0036] In some embodiments, the methods further comprise determining a probability that each putative somatic variant site belongs to the class based on relative likelihoods of different binomial models, gaussian models, or Poisson models. In some embodiments, the non-tumor somatic variant and / or germline variant classes use a fixed parameter for the probability of observing a variant count.Atty. Dkt. No.: 131588-1668[0O37| In some embodiments, the methods further comprise calculating a total likelihood of observing variant and non-variant counts for each putative somatic variant site.
[0038] In some embodiments, the methods further comprise calculating the fraction of ctDNA present in the sample of cfDNA.
[0039] In some embodiments of any of the foregoing aspects or embodiments, the patient has completed at least one cancer treatment prior to obtaining the tumor sample. In some embodiments, the cancer treatment is selected from chemotherapy, radiotherapy, surgery, immunotherapy, cell therapy, or biologic therapy.
[0040] In some embodiments of any of the foregoing aspects or embodiments, the methods further comprise repeating sequencing cfDNA from a sample of blood, plasma, or serum from the cancer patient at 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more successive time points. In some embodiments, sequencing cfDNA from a sample of blood, plasma, or serum from the cancer patient is repeated one or more times while the patient is in remission. In some embodiments, sequencing cfDNA from a sample of blood, plasma, or serum from the cancer patient is repeated one or more times while the patient is undergoing treatment for the cancer. In some embodiments, sequencing cfDNA from a sample of blood, plasma, or serum from the cancer patient is repeated one or more times coinciding with or prior to surgery; following, during, or prior to administration of chemotherapy; following, during, or prior to radiation therapy; following, during, or prior to administration of an immunotherapy; following, during, or prior to administration of a cell therapy; or following, during, or prior to administration of a biologic therapy.
[0041] In some embodiments of any of the foregoing aspects or embodiments, the subject has, had, or is suspected of having a tumor from a cancer selected from bladder cancer, breast cancer, colon / colorectal cancer, gynecologic cancers, head and neck cancers, hematological cancers, liver cancer, lung cancer, and skin cancer.BRIEF DESCRIPTION OF THE FIGURES
[0042] FIG. 1 : shows a block diagram illustrating an example computer environment for implementing methods and processes described herein, according to an embodiment.Atty. Dkt. No.: 131588-1668[0043| FIG. 2 : shows a graph illustrating the number of targets detected across the samples scaled based on the tumor fraction expected from a contrived mixture proportion.
[0044] FIG. 3 : shows a graph that illustrates that the claimed methods provide a consistent, linear series of estimates for tumor fraction and are highly concordant with the expected tumor fraction based on the contrived mixture proportions.DETAILED DESCRIPTION
[0045] In a tumor-informed MRD assay, tumor-specific somatic variants are used to identify the presence and / or quantity of circulating tumor DNA (ctDNA) in a sample of cell-free DNA (cfDNA). The most common way to identify tumor-specific somatic variants is by comparing the sequence data from matched tumor and normal samples, where tumor-specific somatic variants are called at sites primarily based on whether an alternate allele is observed in the tumor, but not the normal sample. The present disclosure, in contrast, provides methods of determining tumor-specific somatic variants from a tumor sample alone by applying particular models, which are described in further detail below. These methods improve existing technology by allowing for the use a set of low quality somatic variant calls that can be classified, via the disclosed methods, as tumor, noise or germline. In this way, the disclosed methods marginalize out (discard) any signal coming from non-tumor variants. Because there will be a higher proportion of false targets called due to noise in the tumor sample, the dropout rate can be adjusted to account for this. The difference between tumor / normal and tumor-only somatic variant calling can be determined by evaluating training data using both calling methods.
[0046] It is to be appreciated that certain aspects, modes, embodiments, variations and features of the present methods are described below in various levels of detail in order to provide a substantial understanding of the present technology.
[0047] The present disclosure is not to be limited in terms of the particular embodiments described in this application, which are intended as single illustrations of individual aspects of the disclosure. All the various embodiments of the present disclosure will not be described herein. Many modifications and variations of the disclosure can be made without departingAtty. Dkt. No.: 131588-1668from its spirit and scope, as will be apparent to those skilled in the art. Functionally equivalent methods and apparatuses within the scope of the disclosure, in addition to those enumerated herein, will be apparent to those skilled in the art from the foregoing descriptions. Such modifications and variations are intended to fall within the scope of the appended claims. The present disclosure is to be limited only by the terms of the appended claims, along with the full scope of equivalents to which such claims are entitled.
[0048] It is to be understood that the present disclosure is not limited to particular uses, methods, reagents, compounds, compositions or biological systems, which can, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting.I. Definitions
[0049] As used herein, the singular form “a,” “an,” and “the” include singular and plural references unless the context clearly dictates otherwise. For example, the term “a cell” includes a single cell as well as a plurality of cells, including mixtures thereof.
[0050] As used herein, the term “approximately” or “about” means plus or minus 10% as well as the specified number. For example, “about 10” should be understood as both “10” and “9-11”.
[0051] Unless otherwise indicated, nucleic acids are written left to right in 5' to 3' orientation; amino acid sequences are written left to right in amino to carboxy orientation, respectively.
[0052] As used herein, the term “comparable” refers to two (or more) sets of conditions, circumstances, individuals, or populations that are sufficiently similar to one another to permit comparison of results obtained or phenomena observed. In some embodiments, comparable sets of conditions, circumstances, individuals, or populations are characterized by a plurality of substantially identical features and one or a small number of varied features. Those of ordinary skill in the art will appreciate that sets of circumstances, individuals, or populations are comparable to one another when characterized by a sufficient number and type of substantially identical features to warrant a reasonable conclusion that differences in results obtained or phenomena observed under or with different sets of circumstances,Atty. Dkt. No.: 131588-1668individuals, or populations are caused by or indicative of the variation in those features that are varied. Those skilled in the art will appreciate that relative language used herein (e.g., enhanced, activated, reduced, inhibited, etc.) will typically refer to comparisons made under comparable conditions.
[0053] As used herein, “comprising” is to be interpreted as specifying the presence of the stated features, integers, steps, or components as referred to, but does not preclude the presence or addition of one or more features, integers, steps, or components, or groups thereof. Moreover, each of the terms “by”, “comprising,” “comprises”, “comprised of,” “including,” “includes,” “included,” “involving,” “involves,” “involved,” and “such as” are used in their open, non-limiting sense and may be used interchangeably. Further, the term “comprising” is intended to include examples and aspects encompassed by the terms “consisting essentially of’ and “consisting of.” Similarly, the term “consisting essentially of’ is intended to include examples encompassed by the term “consisting of.”
[0054] As used herein, the term “derived from” encompasses the terms “originated from,” “obtained from,” “obtainable from,” “isolated from,” and “created from,” and generally indicates that one specified material (e.g., a biological sample) finds its origin in another specified material or individual or has features that can be described with reference to the another specified material.
[0055] As used herein, “genomic DNA” refers to DNA of a cellular genome. In some embodiments, the genomic DNA can be cellular, i.e., contained within a cell. In some embodiments, the genomic DNA can be cell-free.
[0056] As used herein, the term “genotyping” refers to a process of determining the alleles an individual at particular genetic loci by examining an individual’s DNA. Genotyping differs from sequencing in which all of the nucleotides comprising a specific length of DNA are assessed.
[0057] As used herein, “mutation” refers to a change introduced into a reference sequence, including, but not limited to, substitutions, insertions, deletions (including truncations) relative to the reference sequence. Mutations can involve large sections of DNA (e.g., copyAtty. Dkt. No.: 131588-1668number variation). Mutations can involve whole chromosomes (e.g., aneuploidy). Mutations can involve small sections of DNA. Examples of mutations involving small sections of DNA include, e.g., point mutations or single nucleotide polymorphisms (SNPs), single nucleotide variant (SNV), multiple nucleotide polymorphisms, insertions (e.g., insertion of one or more nucleotides at a locus but less than the entire locus), multiple nucleotide changes, deletions (e.g., deletion of one or more nucleotides at a locus), inversions (e.g., reversal of a sequence of one or more nucleotides), an genomic rearrangements (e.g., deletions, duplications, inversions, and translocations). In some embodiments, the reference sequence is a parental sequence. In some embodiments, the reference sequence is a reference genome. Non-limiting examples of reference genomes include hgl8, hgl9, GRCh38.pl4, or GRCh37.pl3. In some embodiments, the mutation is inherited. In some embodiments, the mutation is spontaneous or de nova. In some embodiments, the mutation is a somatic mutation or somatic variant.
[0058] As used herein, “somatic variant” or “somatic mutation” refers to a variant arising after conception, in non-germline DNA of an individual. Somatic variants may include single-nucleotide variants (SNVs), multi -nucleotide variants, insertions and deletions (e.g., indel variants), and genomic rearrangements for example. “Somatic variant” and “somatic mutation” are used interchangeably herein. In some embodiments, the terms “somatic variant” or “somatic mutation” refers to a collection of somatic variants that are specific to a patient.
[0059] As used herein, “tumor-specific somatic variant” or “tumor-specific somatic mutation” refers to variant(s) (e.g., nucleic acid changes in DNA) that occurs in a somatic cell before or during tumor development and not present within germline cells. In some embodiments, the “tumor-specific somatic variant” can be specific to the subject. In some embodiments, the “tumor-specific somatic variant” can be common to a specific tumor or cancer.
[0060] As used herein, “patient-specific panel” or “putative tumor-specific somatic variants” refers to a collection of sequences comprising somatic variants that are specific to a patient, or markers that distinguish between two or more individuals.Atty. Dkt. No.: 131588-1668[00611 As used herein, “reference dataset” or “reference set” refers to a population-level collection of sequences prepared from non-tumor or germline samples. In some embodiments, the samples used to prepare the reference dataset are from a healthy subjects. Such a reference dataset can be used to correct for inclusion of latent variant classes including but not limited to, germline variants, sites with systematic noise, low frequency clonal hematopoiesis of indeterminate potential (CHIP) variants or other anomalous sites. Examples of reference datasets include, but are not limited to, those provided by the Genome Aggregation Database (gnomAD; gnomad.broadinstitute.org) and the Single Nucleotide Polymorphism database (dbSNP; ncbi.nlm.nih.gov / snp / ).
[0062] As used herein, “tumor fraction” or “fraction of circulating cell-free tumor DNA” or fractions of ctDNA” refers to the proportion of circulating cell-free tumor DNA (ctDNA) relative to the total amount of cell-free DNA (cfDNA). In some embodiments, the tumor fraction or fraction of ctDNA may be indicative of the size of the tumor.
[0063] As used herein, “sample” or “biological sample,” refers to a biological sample obtained or derived from a source of interest. In some embodiments, the sample can be any biological sample presumed to contain nucleic acids (e.g., RNA, DNA, such as, genomic DNA). In some embodiments, the biological sample can be obtained from a subject or patient. In some embodiments, a biological sample is or comprises biological tissue or fluid. In some embodiments, a biological sample may be or comprise bone marrow; blood (or a fraction thereof); blood cells; ascites; tissue or fine needle biopsy samples; cell-containing body fluids; free floating nucleic acids (e.g., cell free DNA); sputum; saliva; urine; cerebrospinal fluid, peritoneal fluid; pleural fluid; lymph; gynecological fluids; skin swabs; vaginal swabs; oral swabs; nasal swabs; washings or lavages such as a ductal lavages or broncheoalveolar lavages; aspirates; scrapings; bone marrow specimens; tissue biopsy specimens; surgical specimens; feces, other body fluids, secretions, and / or excretions; and / or cells therefrom. In some embodiments, a biological sample is or comprises cells obtained from an individual. In some embodiments, obtained cells are or include cells from an individual from whom the sample is obtained. In some embodiments, a sample is a “primary sample” obtained directly from a source of interest by any appropriate means. For example, in some embodiments, a primary biological sample is obtained by methods selected from theAtty. Dkt. No.: 131588-1668group consisting of a swab, biopsy (e.g., fine needle aspiration or tissue biopsy), surgery, collection of body fluid (e.g., blood, lymph, feces etc.), etc. In some embodiments, a sample may be a preparation that is obtained by processing (e.g., by removing one or more components of and / or by adding one or more agents to) a primary sample. For example, filtering using a semi-permeable membrane. In some embodiments, a processed sample may comprise nucleic acids or proteins extracted from a sample or obtained by subjecting a primary sample to techniques such as amplification or reverse transcription of mRNA, isolation and / or purification of some components, etc. In some embodiments, may include, but is not limited to, tissue, blood, plasma, saliva, urine, semen, amniotic fluid, oocytes, skin, hair, feces, cheek swabs, or pap smear lysate from an individual. In some embodiments, the sample is blood, plasma, or serum.
[0064] As used herein, “amplification,” with respect to nucleic acid sequences, refers to methods that increase the representation of a population of nucleic acid sequences in a sample. Copies of a particular target nucleic acid sequence generated in vitro in an amplification reaction are called “amplicons” or “amplification products”. Amplification may be exponential or linear. A target nucleic acid may be DNA (such as, for example, genomic DNA, cfDNA, ctDNA, and cDNA) or RNA. Amplification can be achieved using polymerase chain reaction (PCR) as well as numerous other methods such as isothermal methods, rolling circle methods, etc.
[0065] As used herein, “anneal,” “hybridize,” or “bind,” refer to two polynucleotide sequences, segments or strands, and can be used interchangeably, that forming hydrogen bonds with complementary bases to produce a double-stranded polynucleotide or a doublestranded region of a polynucleotide. Two complementary sequences (e.g., DNA and / or RNA) can anneal or hybridize.
[0066] As used herein, “treat,” “treatment,” and “treating” refer to the reduction or amelioration of the progression, severity, and / or duration of a proliferative disorder e.g., cancer, or the amelioration of a proliferative disorder resulting from the administration of one or more therapies.Atty. Dkt. No.: 131588-1668[0O67| As used herein, “small nucleotide polymorphism” or “SNP” refers to a singlenucleotide variant (SNV), a multi -nucleotide variant (MNV), or an indel variant about 100 base pairs or less.
[0068] As used herein, “multi-nucleotide variant” or “MNV” refers to a variant having 2 or more adjacent nucleotide changes.
[0069] As used herein, “copy number variant” or “copy number” refers to the number of each somatic variant in the set of somatic variants.
[0070] As used herein, “copy number variation” or “CNV” refers to the number of copies of a DNA segment. As used herein, copy number variation originates from changes in copy number in germline cells.
[0071] As used herein, “copy number alteration” or “CNA” refers to the number of copies of a DNA segment. As used herein, copy number alterations are changes in copy number in somatic cells.
[0072] As used herein, “allele balance” refers to a ratio of a variant allele to a reference allele. In some embodiments, the variant allele is a variant allele from each somatic variant in the set of somatic variants.
[0073] As used herein, “heterozygous” refers to a loci or gene having different nucleotides or nucleotide sequences on each allele. For example, a heterozygous site may comprise one single nucleotide polymorphism (SNP) on one allele, and one wild-type nucleotide at the corresponding site on the other allele.
[0074] As used herein, “homozygous” refers to a loci or gene having the same nucleotides or nucleotide sequences on each allele.
[0075] As used herein, the term “homozygous alternate” refers to a loci comprising two identical, non-reference (“variant”) nucleotides at the loci (e.g., one variant nucleotide on the sense strand and one variant nucleotide on the anti-sense strand) or a gene comprising two identical, non-reference (“variant”) alleles.Atty. Dkt. No.: 131588-1668
[0076] As used herein, “Next Generation Sequencing” or “NGS” refers to sequencing methods that allow for massively parallel sequencing of clonally amplified and of single nucleic acid molecules during which a plurality, e.g., millions, of nucleic acid fragments from a single sample or from multiple different samples are sequenced in unison. Non-limiting examples of NGS include sequencing-by-synthesis, sequencing-by-ligation, real-time sequencing, and nanopore sequencing.
[0077] As used herein, “sequence read” or “read” refers to sequence information of a nucleic acid fragment obtained through a sequencing assay, such as a next generation sequencing (NGS) assay. In some embodiments, a sequence read refers to data representing a sequence of nucleotide bases that were measured using a clonal sequencing method. Clonal sequencing may produce sequence data representing single, or clones, or clusters of one original DNA molecule. A sequence read may also have associated quality score at each base position of the sequence indicating the probability that nucleotide has been called correctly.
[0078] As used herein, a “set” of reads refers to all sequencing reads with a common parent nucleic acid strand, which may or may not have had errors introduced during sequencing or amplification of the parent nucleic acid strand.
[0079] As used herein, “subject,” “patient,” or “individual” refers to any organism upon which embodiments of the present disclosure may be used or administered, e.g., for experimental, screening, diagnostic, prophylactic, and / or therapeutic purposes. Typical subjects include animals (e.g., mammals such as mice, rats, rabbits, non-human primates, and humans; insects; worms; etc.).
[0080] As used herein, the term “target sequence” refers to a selected target polynucleotide, e.g., a sequence present in a cfDNA molecule, whose presence, amount, and / or nucleotide sequence, or changes in these, are desired to be determined. Target sequences can be interrogated for the presence or absence of a somatic and / or germline variant. The target polynucleotide can be a region of gene associated with a disease. In some embodiments, the region is an exon. The disease can be cancer.Atty. Dkt. No.: 131588-1668[00811 As used herein in the context of molecules, e.g., nucleic acids, proteins, or small molecules, a “variant” refers to a molecule that shows significant structural identity with a reference molecule but differs structurally from the reference molecule, e.g., in the presence or absence or in the level of one or more chemical moieties as compared to the reference entity. In some embodiments, a variant also differs functionally from its reference molecule. In some embodiments, a variant, is a distinct molecule that shares one or more such characteristic structural elements but differs in at least one aspect from the reference molecule. In some embodiments, a variant can be a polypeptide comprised of a plurality of amino acids having designated positions relative to one another in linear or three-dimensional space and / or contributing to a particular structural motif and / or biological function. In some embodiments, a variant can be a nucleic acid comprised of a plurality of nucleotide residues having designated positions relative to one another in linear or three-dimensional space. In some embodiments, a variant polypeptide or nucleic acid may differ from a reference polypeptide or nucleic acid as a result of one or more differences in amino acid or nucleotide sequence and / or one or more differences in chemical moieties (e.g., carbohydrates, lipids, phosphate groups) that are covalently components of the polypeptide or nucleic acid (e.g., that are attached to the polypeptide or nucleic acid backbone). In some embodiments, a variant polypeptide or nucleic acid shows an overall sequence identity with a reference polypeptide or nucleic acid that is at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, or 99%. In some embodiments, a variant polypeptide or nucleic acid does not share at least one characteristic sequence element with a reference polypeptide or nucleic acid. In some embodiments, a reference polypeptide or nucleic acid has one or more biological activities. In some embodiments, a variant polypeptide or nucleic acid shares one or more of the biological activities of the reference polypeptide or nucleic acid. In some embodiments, a variant polypeptide or nucleic acid lacks one or more of the biological activities of the reference polypeptide or nucleic acid. In some embodiments, a variant polypeptide or nucleic acid shows a reduced level of one or more biological activities as compared to the reference polypeptide or nucleic acid. In some embodiments, a polypeptide or nucleic acid of interest is considered to be a “variant” of a reference polypeptide or nucleic acid if it has an amino acid or nucleotide sequence that is identical to that of the reference but for a small number of sequence alterations at particular positions. Typically, fewer than aboutAtty. Dkt. No.: 131588-166820%, about 15%, about 10%, about 9%, about 8%, about 7%, about 6%, about 5%, about 4%, about 3%, or about 2% of the residues in a variant are substituted, inserted, or deleted, as compared to the reference. In some embodiments, a variant polypeptide or nucleic acid comprises about 10, about 9, about 8, about 7, about 6, about 5, about 4, about 3, about 2, or about 1 substituted residues as compared to a reference. Often, a variant polypeptide or nucleic acid comprises a very small number (e.g., fewer than about 5, about 4, about 3, about 2, or about 1) number of substituted, inserted, or deleted, functional residues (i.e., residues that participate in a particular biological activity) relative to the reference. In some embodiments, a variant polypeptide or nucleic acid comprises not more than about 5, about 4, about 3, about 2, or about 1 addition or deletion, and, in some embodiments, comprises no additions or deletions, as compared to the reference. In some embodiments, a variant polypeptide or nucleic acid comprises fewer than about 25, about 20, about 19, about 18, about 17, about 16, about 15, about 14, about 13, about 10, about 9, about 8, about 7, about 6, and commonly fewer than about 5, about 4, about 3, or about 2 additions or deletions as compared to the reference. In some embodiments, a reference polypeptide or nucleic acid is one found in nature. In some embodiments, a reference polypeptide or nucleic acid is a human polypeptide or nucleic acid.II. Minimal Residual Disease Detection
[0082] The goal of a minimum residual disease (MRD) assay is to detect and / or quantify circulating tumor DNA (ctDNA) so researchers and clinicians can detect recurrence early and monitor the progress of the disease through treatment. In general, an MRD assay will rely on a patient-specific and tumor-specific panel (i.e., “a set of putative tumor-specific somatic variants”) for assessing the presence of ctDNA in a patient sample. The set of putative tumorspecific somatic variants can be prepared with the general steps of (1) profiling a tumor or cancer sample from a patient, and (2) identifying a set of putative somatic mutations to target, and, at one or more later time points, (3) taking a subsequent sample from the patient, (4) enriching cell-free DNA (cfDNA) for the target somatic mutation sites, and (5) determining or estimating the ctDNA content of the cfDNA given the tumor profile and sequencing data. As disclosed herein, the methods of the present technology relate to detecting circulating tumor DNA (ctDNA) in a sample, comprising: (a) sequencing DNA from a tumor sampleAtty. Dkt. No.: 131588-1668obtained from a patient, thereby obtaining a first set of sequencing reads of the DNA from the tumor sample; (b) determining a set of putative tumor-specific somatic variants comprising a plurality of putative somatic variant sites; (c) sequencing cell-free DNA (cfDNA) from a sample of blood, plasma, or serum from the cancer patient, thereby obtaining a second set of sequencing reads; and (d) detecting ctDNA sequences in the second set of sequencing reads based on a sequencing read comprising one or more putative somatic variant sites from the set of putative tumor-specific somatic variants
[0083] The tumor sample may be a solid tumor sample, such as a biopsy or other tissue sample, or a liquid sample or a fluid sample, such as blood (in the case of a hematological cancer) or specific fractions of blood. In some embodiment, the tumor sample comprises a tumor biopsy or fluid sample. In some embodiments, the fluid sample is selected from blood, blood plasma, blood serum, urine, saliva, and cerebral spinal fluid (CSF). In some embodiments, the subject has, had, or is suspected of having a cancer. In some embodiments, the cancer is selected from bladder cancer, breast cancer, colon / colorectal cancer, gynecologic cancers, head and neck cancers, hematological cancers, liver cancer, lung cancer, and skin cancer.
[0084] Once a patient-specific and tumor-specific panel (i.e., a “set of putative tumor-specific somatic variants”) has been established, the set of tumor-specific somatic variants can be used to detect ctDNA (i.e., fragments that include a tumor-specific somatic mutation or variant) in subsequent samples taken from the cancer patient. The subsequent samples may be taken from a patient at various time points during the course of treatment or during a period of remission. For example, after a surgical removal of a tumor, the tumor may be profiled as described herein to determine tumor-specific somatic mutations, and at one or more subsequent time points a subsequent sample may be taken from the subject to search for the presence of any ctDNA comprising any one of the identified tumor-specific somatic mutations. The detection or presence of ctDNA comprising a tumor-specific somatic mutation may be indicative of cancer recurrence. Additionally or alternatively, similar assessment can be performed throughout the course of a patient’s treatment (e.g., with chemotherapy, radiation, immunotherapy, cell therapy, etc.) to detect or quantify ctDNA and determine whether the amount of ctDNA is increasing or decreasing, as this may beAtty. Dkt. No.: 131588-1668indicative of responsiveness to the therapy. Accordingly, assessment of a subsequent sample may be repeated 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more times throughout the course of a patient’s remission or treatment. The assessment of a subsequent sample may be repeated monthly, every other month, once every three months, once every four months, once every five months, once every six months, once every seven months, once every eight months, once every nine months, once every ten months, once every eleven months, or annually.
[0085] The type of sample used for the one or more subsequent samples is generally a blood sample, a plasma sample, a serum sample, or urine sample, but any biological sample that contains cfDNA and potentially contains ctDNA would be acceptable. In some embodiments, the one or more subsequent samples are cell-free samples.
[0086] A signature panel, a set of somatic variants, or set of putative tumor-specific somatic variants may comprise, for example, 10-10,000 tumor-specific somatic variants or putative tumor-specific somatic variants. For example, a set of tumor-specific somatic variants or putative tumor-specific somatic variants may comprise 10-20, 10-30, 10-40, 10-50, 10-60, 10-70, 10-80, 10-90, 10-100, 10-200, 10-250, 10-300, 10-350, 10-400, 10-4000, 10-3000, 10-2500, 10-2000, 10-1500, 10-1000, 10-950, 10-900, 10-850, 10-800, 10-750, 10-700, 10-650, 10-600, 10-550, 10-500, 50-100, 50-200, 50-250, 50-300, 50-400, 50-5000, 50-4000, 50-3000, 50-2500, 50-2000, 50-1500, 50-1000, 50-950, 50-900, 50-850, 50-800, 50-750, 50-700, 50-650, 50-600, 50-550, 50-500, 100-5000, 100-4000, 100-3000, 100-2500, 100-2000, 100-1500, 100-1000, 100-950, 100-900, 100-850, 100-800, 100-750, 100-700, 100-650, 100-600, 100-550, 100-500, 200-5000, 200-4000, 200-3000, 200-2500, 200-2000, 200-1500, 200-1000, 200-950, 200-900, 200-850, 200-800, 200-750, 200-700, 200-650, 200-600, 200-550, 200-500, 300-5000, 300-4000, 300-3000, 300-2500, 300-2000, 300-1500, 300-1000, SOO-OSO, 300-900, 300-850, 300-800, 300-750, 300-700, 300-650, 300-600, 300-550, 300-500, 400-5000, 400-4000, 400-3000, 400-2500, 400-2000, 400-1500, 400-1000, 400-950, 400-900, 400-850, 400-800, 400-750, 400-700, 400-650, 400-600, 400-550, 400-500, 500-5000, 500-4000, 500-3000, 500-2500, 500-2000, 500-1500, 500-1000, 500-950, 500-900, 500-850, 500-800, 500-750, 500-700, 500-650, 500-600, or 500-550 tumor-specific somatic variants or putative tumor-specific somatic variants. In some embodiments, a signature panel, a set of somatic variants or a subset of somatic variants may comprise or consist of about 10, aboutAtty. Dkt. No.: 131588-166820, about 30, about 40, about 50, about 75, about 100, about 150, about 200, about 250, about 300, about 350, about 400, about 450, about 500, about 550, about 600, about 650, about 700, about 750, about 800, about 850, about 900, about 950, about 1000, about 1100, about 1150, about 1200, about 1250, about 1300, about 1350, about 1400, about 1450, about 1500, about 1550, about 1600, about 1650, about 1700, about 1750, about 1800, about 1850, about 1900, about 1950, or about 2000 or more tumor-specific somatic mutations. In some embodiments, a signature panel , a set of somatic variants, or a set of putative tumor-specific somatic variants may comprise at least 10, at least 20, at least 30, at least 40, at least 50, at least 75, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1150, at least 1200, at least 1250, at least 1300, at least 1350, at least 1400, at least 1450, at least 1500, at least 1550, at least 1600, at least 1650, at least 1700, at least 1750, at least 1800, at least 1850, at least 1900, at least 1950, or at least 2000 tumor-specific somatic variants or putative tumorspecific somatic variants. The set of tumor-specific somatic variants or putative tumorspecific somatic variants may be in introns, exons, or a combination thereof. In some embodiments, the set of tumor-specific somatic variants or putative tumor-specific somatic variants may comprise or consist of one or more somatic mutations selected from SNVs, insertions, deletions and translocations. In some embodiments, the set of tumor-specific somatic variants or putative tumor-specific somatic variants may comprise or consist of one or more multi -nucleotide variants (MNVs).
[0087] After the set of tumor-specific somatic variants are selected, cell-free DNA (cfDNA) is sequenced. In some embodiments, the cfDNA is from a sample of blood, plasma, serum, or urine from the patient, (i.e., cancer patient). This sequencing may be performed by, for example Next Generation Sequencing (NGS). Deep sequencing may allow for more sensitive detection, and so the depth of the sequencing may be at least 50X, at least 100X, at least 150X, at least 200X, at least 250X, at least 300X, at least 350X, at least 400X, at least 450X, at least 500X, at least 550X, at least 600X, at least 650X, at least 700X, at least 750X, at least 800X, at least 850X, at least 900X, at least 950X, or at least 1000X. In other words, the depth of the sequencing may be about 50X, about 100X, about 150X, about 200X, about 250X, about 300X, about 350X, about 400X, about 450X, about 500X, about 550X, about 600X,Atty. Dkt. No.: 131588-1668about 650X, about 700X, about 750X, about 800X, about 850X, about 900X, about 950X, or about 1000X.
[0088] The ctDNA sequences are detected in the sequenced cfDNA reads based on the cfDNA sequences reads comprising one or more of the selected somatic variant sites from the set of tumor-specific somatic variants.
[0089] Prior to sequencing the cfDNA, the cfDNA sample may optionally be enriched for fragments comprising one or more of the selected somatic variant sites from the set of tumorspecific somatic variants. Such enrichment can be performed via hybrid capture enrichment, PCR-based enrichment, or on-sequencer enrichment. Briefly, enrichment may comprise extracting cfDNA from a subsequent sample taken from the cancer patient and contacting the extracted cfDNA with a plurality of oligonucleotides (i.e., oligonucleotide probes or primers), wherein each oligonucleotide in the plurality of oligonucleotides comprises a nucleic acid sequence that is capable of hybridizing to a cfDNA fragment comprising one of the tumorspecific somatic mutation sequences. Thus, enrichment may utilize a set of oligonucleotide probes / primers to selectively enrich ctDNA that may be in a subsequent sample by binding to and / or amplifying DNA fragments comprising previously identified tumor-specific somatic mutation sequences. In some embodiments, the cfDNA sample is not enriched prior to sequencing.
[0090] Although MRD methods generally provide significant clinical utility in tracking treatment, recurrence, and prognosis of cancer patients, certain aspects of prior MRD processes can be improved with the disclosed methods. Specifically, MRD generally requires tumor / cancer profiling compared to non-tumor and / or non-cancer samples to identify somatic variants present in a tumor of interest. More specifically, preparing the patient-specific and tumor-specific panel (i.e., a “signature panel” or “a set of somatic variants”) generally comprises, for example, (a) obtaining a tumor sample and a non-tumor sample from a cancer patient; (b) sequencing DNA from the tumor sample (e.g., genomic DNA) and sequencing DNA from the non-tumor sample (e.g., cfDNA), thereby obtaining sequences DNA or sequence reads from the tumor sample and the non-tumor sample; and (c) comparing the sequences of the tumor sample and the non-tumor sample to determine any tumor-specificAtty. Dkt. No.: 131588-1668somatic mutations that are present in the sequences of DNA from the tumor sample but not present in the sequences of DNA from the non-tumor sample. Additionally, most current MRD assays generally requires enrichment of cfDNA to be tested.
[0091] The present disclosure uses a tumor sample only to establish a patient-specific and tumor-specific panel (i.e., “a set of tumor-specific somatic variants”), and the disclosed methods do not require enrichment steps (though enrichment prior to sequencing can optionally be performed). In some embodiments, the methods as disclosure herein, do not comprise sequencing DNA from a non-tumor sample from the patient for determining the set of tumor-specific somatic variants and do not require comparing the sequences of the tumor sample to a non-tumor sample. In some embodiments, the methods, as disclosed herein, comprise somatic variant calling approaches that do not require a non-tumor sample from the patient. In some embodiments, the cfDNA from the sample from the cancer patient is not enriched prior to sequencing. In some embodiments, the cfDNA is from the sample of blood, plasma, or serum from the cancer patient.
[0092] The current methods can be achieved by prioritizing selection of sites to generate a set of tumor-specific somatic variants based on a plurality of features such as a relationship between allele balance and copy number, though other approaches such as machine learning, and assemblies of haplotypes can additionally or alternatively be used to aid in selection of tumor-specific variants useful for MRD analysis. The disclosed methods of selecting tumorspecific variants without comparison to a non-tumor sample may also incorporate accounting for variant background error and subsequence correcting therefrom.[0093 [ Thus, the disclosed methods allow for selection of tumor-only-informed signature panels. These methods can utilize machine learning- or artificial intelligence-based models to determine the subset of variants selected for such panels. ML and / or Al models utilized in accordance with technologies of the present disclosure can determine, by at least one processor, (a) allele balance, (b) dropout rate, (c) latent variant classes, and / or (d) sample or panel-specific error rates. ML and / or Al models utilized in accordance with technologies of the present disclosure can also determine haplotype or prepare haplotype assemblies. ML-and Al-based analysis of these features generally improves the overall predictiveness andAtty. Dkt. No.: 131588-1668selection of a patient-specific panel, such that the use of a germline comparison sample is no longer strictly required. The disclosed analysis of allele balance, dropout rate, latent variant classes, sample or panel-specific error rates, and / or haploidy to select an optimized signature panel would not be feasible than with human analysis alone, and the disclosed methodology represents a tangible improvement over current MRD methods and cancer treatment / management, as these methods require fewer subject samples and greater accuracy for tracking patient recurrence and / or response to treatment.
[0094] Specific aspects of MRD processes are discussed in more detail below.III. Phase I - Signature PanelA. DNA library preparation
[0095] In some embodiments of the methods disclosed herein, a DNA library is obtained or prepared from DNA or cfDNA obtained from a patient, e.g., a cancer patient. In some embodiments, a DNA library is obtained or prepared from the genome of the patient. In some embodiments, the DNA has been previously sequenced, and mutations or variants identified.[0096| When producing a DNA library from genomic DNA, the genomic DNA can be fragmented, for example by using a hydrodynamic shear or other mechanical force, or fragmented by chemical or enzymatic digestion, such as restriction digesting. This fragmentation process allows the DNA molecules present in the genome to be sufficiently short for analysis, such as sequencing or digital PCR. cfDNA, however, is generally sufficiently short such that no fragmentation is necessary. In some embodiment, cfDNA is fragmented. In some embodiments, cfDNA is not fragmented. cfDNA originates from genomic DNA. A portion of the cfDNA obtained from a sample of a cancer patient may originate from cancer cells (i.e., ctDNA) and a portion of the cfDNA may originate from noncancer cells.B. Panel of mutations / marker
[0097] In some embodiments, DNA from a tumor sample obtained from a patient is sequenced, thereby obtaining a set of sequence reads. In some embodiments, the set ofAtty. Dkt. No.: 131588-1668sequence reads is a first set of sequence reads of the DNA from the tumor sample. In some embodiments, the sequencing of the nucleic acid from the sample is performed using whole genome sequencing (WGS). In some embodiments, targeted sequencing (e.g., subtractive hybridization) is performed and may be either DNA or RNA sequencing. The targeted sequencing may be to a subset of the whole genome. In some embodiments, the targeted sequencing is to introns, exons, non-coding sequences, or a combination thereof. In some embodiments, targeted whole exome sequencing (WES) of the DNA from the sample is performed. In some embodiments, sequencing DNA from the tumor sample comprises whole genome sequencing (WGS), whole exome sequencing (WES), targeted sequencing, or subtractive hybridization. In some embodiments, whole genome sequencing (WGS) of the DNA (e.g., ctDNA and cfDNA) is performed. In some embodiments, Whole Exome Sequencing (WES) of the tumor DNA is performed. WES comprises selecting DNA sequences that encode proteins, and sequencing that DNA using any high throughput DNA sequencing technology. Methods that can be used to target exome DNA include the use of polymerase chain reaction (PCR), molecular inversion probes (MIP), hybrid capture, and insolution capture. The utility of targeted genome approaches is well established, and commercially available methods for WES include the Roche NimbleGen Capture Array (Roche NimbleGen Inc., Madison, WI), Agilent SureSelect (Agilent Technologies, Santa Clara, CA), and RainDance Technologies emulsion PCR (RainDance Technologies, Lexington, MA), IDT xGen® Exome Research Panel and others.
[0098] In some embodiments, the DNA is sequenced using a next generation sequencing platform (NGS), such as, massively parallel sequencing. NGS technologies provide high throughput sequence information, and provide digital quantitative information, in that each sequence read that aligns to the sequence of interest is countable. In some embodiments, clonally amplified DNA templates or single DNA molecules are sequenced in a massively parallel fashion within a flow cell. In addition to high-throughput sequence information, NGS provides quantitative information, in that each sequence read is countable and represents an individual clonal DNA template or a single DNA molecule. The sequencing technologies of NGS include pyrosequencing, sequencing-by-synthesis with reversible dye terminators, sequencing by oligonucleotide probe ligation and ion semiconductor sequencing. DNA from individual samples can be sequenced individually (i.e., singleplex sequencing) or DNA fromAtty. Dkt. No.: 131588-1668multiple samples can be pooled and sequenced as indexed genomic molecules (i.e., multiplex sequencing) on a single sequencing run, to generate up to several hundred million reads of DNA sequences. Commercially available platforms include, e.g., platforms for sequencing-by-synthesis, ion semiconductor sequencing, pyrosequencing, reversible dye terminator sequencing, sequencing by ligation, single-molecule sequencing, sequencing by hybridization, and nanopore sequencing. Platforms for sequencing by synthesis are available from, e.g., Illumina, 454 Life Sciences, Helicos Biosciences, and Qiagen. Illumina platforms can include, e.g., Illumina’s Solexa platform, Illumina’s Genome Analyzer. Life Science platforms include, e.g., the GS Flex and GS Junior, and are described in U.S. Pat. No.7,323,305. Platforms from Helicos Biosciences include the True Single Molecule Sequencing platform. Ion Torrent, an alternative NGS system, is available from ThermoScientific and is a semiconductor-based technology that detects hydrogen ions that are released during polymerization of nucleic acids. Any detection method that allows for the detection of segregatable markers may be used with the assay provided for herein.[0099| In some embodiments, the tumor-specific somatic mutations identified will be analyzed and filtered to generate a subset panel of markers. For example, the subset panel of markers may comprise one or more types of somatic mutation, including but not limited to single-nucleotide variants (SNVs), multi -nucleotide variants (MNVs), insertions and deletions (e.g., indel variants), and genomic rearrangements. In some embodiments, the subset panel of somatic mutations can include greater than 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50, and up to 100, up to 200, up to 300, up to 400, up to 500, up to 600, up to 700, up to 800, up to 900, up to 1,000, up to 1,500, up to 2,000, up to 2,500, up to 3,000, up to 4,000, up to 5,000, up to 6,000, up to 7,000, up to 8,000, up to 9,000, up to 10,000, up to 11, 000, up to 12,000, up to 13,000, up to 14,000, up to 15,000, or more mutations. In other embodiments, the subset panel includes between 50 and 15,000 mutations, between 100 and 15,000 mutations, between 500 and 13,000 mutations, between 1,000 and 10,000 mutations, between 2,000 and 8,000 mutations, or between 4,000 and 6,000 mutations.
[0100] The methods as disclosed herein may select a subset of tumor-specific somatic variants from a broader group of putative tumor-specific somatic variants based on tumorAtty. Dkt. No.: 131588-1668allele balance and copy number using tumor purity, machine learning, and / or one or more assemblies of haplotypes.C. Correction based on allele balance, copy number, and tumor purity
[0101] In some embodiments, the methods as disclosed herein, prioritize putative somatic variant sites for determining a set of putative tumor-specific somatic variants based on allele balance, copy number (e.g., copy number alteration (CNA)) or a combination thereof for each putative somatic variant site in the DNA sequence reads. In some embodiments, the sequence reads are the first set of sequences reads as disclosed herein. In some embodiments, determining a set of putative tumor-specific somatic variants comprises determining the tumor purity of the tumor sample from the patient.
[0102] As disclosed herein, a correction for tumor allele balance, is a ratio of variant allele to reference allele in sequencing data, using tumor purity. In some embodiments, the sequence data is the first set of sequence reads. Tumor samples can be contaminated with normal tissue, either due to errors in dissection or immune cell infiltration. This can reduce the observed allele balance for somatic variant sites, resulting in an artificially inflated inferred tumor fraction overall. This approach is modeled on the relationship between the probability of observing alternative allele counts, / ?, at a site given allele balance in the tumor ABT) copy number of the site in both tumor CNT) and normal (CAv) and the fraction of tumor DNA in the sample, the tumor purity (TP) and the error rate estimated from control site data (err '.
[0103] Each variant will be ranked by an expected p of 0.005% TF using allele balance and copy number in the tumor sample, assuming copy number in the non-tumor sample is 2. This approach accounts for the sub-clonal variants on detection sensitivity.
[0104] Other approaches utilizing allele balance and / or copy number are also feasible and would rely on the same underlying concepts. Namely, the disclosed methods can leverage the fact that most tumor samples are contaminated with normal DNA. Based on copy number alterations and allele balances for SNPs in the sequence data, an algorithm can be used to determine the purity of the tumor sample. Based on that purity, one can infer whetherAtty. Dkt. No.: 131588-1668mutations observed in the sequence data are patient-specific SNPs (germline) or tumorspecific SNVs (somatic) based on expected allele balances and copy numbers. For example, if a tumor sample was 50% pure and a heterozygous site was observed at 0.5 allele balance (i.e., half of all reads) in a region that is copy number 2, this indicates that the site is germline because any germline site is expected to have an allele balance of 0.5 or 1.0. Alternatively, if a tumor sample was 50% pure and a heterozygous site was observed at 0.25 allele balance, this indicates that the site is somatic since 0.25 is equivalent to a heterozygous somatic site at 50% purity (0.5 x 0.5). In other words, by taking into account sample purity, copy number, and expected allele balances, somatic variants can be identified based on the observed allele balance of any sequenced loci.[0105| In some embodiments, the allele balance and copy number of each of the somatic variants in the set of somatic variants is calculated using a computer processor.D. Haplotype Analysis
[0106] In some embodiments, the methods as disclosed herein, prioritize putative somatic variant sites for determining a set of putative tumor-specific somatic variants based assemblies of haplotypes for each putative somatic variant site in the DNA sequence reads. Such methods can aid in such somatic calling or refining the results from a set of putative tumor-specific somatic mutations. In some embodiments, the sequence reads used for a haplotype assembly are the first set of sequences reads as disclosed herein. In some embodiments, determining a set of putative tumor-specific somatic variants comprises calling, via a computer processor, the plurality of putative somatic variant sites based on one or more assemblies of haplotypes. Similar to the methods described above, once the haplotypes are assembled, variations in allele balance can be used to determine whether a given SNP or SNV is germline or a somatic variant.E. Machine Learning
[0107] In some embodiments, the methods as disclosed herein, utilize machine learning models trained with genomic data (e.g., DNA sequencing data, whole genome sequencing data, whole exome sequencing data, targeted genomic sequencing data, cfDNA sequencingAtty. Dkt. No.: 131588-1668data, chromatin immunoprecipitation sequencing data, reference genome data, transcriptomics data, epigenomics data, or proteomics data) are also utilized to select the subset of somatic variants. In some embodiments, determining a set of putative tumorspecific somatic variant comprises training a machine learning model to select the plurality of putative somatic variant sites.[0108| In some embodiments, data associated with the tumor sample from a cancer patient is retrieved by a processor. In some embodiments, the data is sequenced DNA (e.g., cfDNA or genomic DNA). In some embodiments, the set of putative tumor-specific somatic variants is generated by a computer device. In some embodiments, a training data set is generated by a processor comprising sequenced DNA and a set of putative tumor-specific somatic variants as disclosed herein. The training dataset can be used to train one or more machine learning models, such that when trained the machine learning model can ingest a set of mutations for a new patient and select a subset of the mutations that are suitable for tracking a patient’s progress over the course of treatment or cancer recurrence.
[0109] The tumor somatic variants are filtered to generate a set of putative tumor-specific somatic variants based on the plurality of features as disclosed herein and / or based on a plurality of features selected by a machine learning model trained with genomic data. In some embodiments, the genomic data comprises DNA sequencing data, whole genome sequencing data, whole exome sequencing data, targeted genomic sequencing data, cfDNA sequencing data, chromatin immunoprecipitation sequencing data, reference genome data, transcriptomics data, epigenomics data, or proteomics data.[0110| The subset of somatic variants serves as a signature panel for the patient that can be sequenced at various stages of the disease, i.e., the signature panel can be screened to determine the presence of cancer at surgery following diagnosis; during cancer treatment, e.g., at intervals during chemotherapy or radiation therapy, to monitor the efficacy of the treatment; at intervals during remission to confirm continued absence of disease; and / or to detect recurrence of the disease. The composition of the selected somatic variants for the subset is a key determinant for the sensitivity and specificity of the methods described herein.Atty. Dkt. No.: 131588-1668[01111 Various actions discussed herein, such as methods and processes discussed herein may be performed (at least partially) by one or more processors implementing / utilizing one or more machine learning models.
[0112] Various actions discussed herein, such as methods and processes discussed herein may be performed (at least partially) by one or more processors implementing / utilizing one or more machine learning models.[01131 Figure 1 illustrates an example computer environment 100 that can be used to provide a network-based implementation of the methods and processes described herein.Specifically, The Figure illustrates components of a system 100 for a mutation analysis system, according to an embodiment. The system 100 may include an analytics server 110a, system database 110b, a machine learning model 111, electronic data sources 120a-d (collectively electronic data sources 120), end-user devices 140a-c (collectively end-user devices 140), and an administrator computing device 150.
[0114] The system 100 is not confined to the components described herein and may include additional or other components not shown for brevity, which are to be considered within the scope of the embodiments described herein.
[0115] The above-mentioned components may be connected to each other through a network 130. Examples of the network 130 may include, but are not limited to, private or public local-area-networks (LAN), wireless LAN (WLAN) networks, metropolitan area networks (MAN), wide-area networks (WAN), and the Internet. The network 130 may include wired and / or wireless communications according to one or more standards and / or via one or more transport media. Communication over the network 130 may be performed in accordance with various communication protocols such as Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), and Institute of Electrical and Electronics Engineers (IEEE) communication protocols. In one example, the network 130 may include wireless communications according to Bluetooth specification sets or another standard or proprietary wireless communication protocol. In another example, the network 130 may also include communications over a cellular network, including, e.g., GSM (Global System forAtty. Dkt. No.: 131588-1668Mobile Communications), CDMA (Code Division Multiple Access), or EDGE (Enhanced Data for Global Evolution) networks.
[0116] The analytics server 110a may generate and display an electronic platform configured to receive information and output results of execution of the machine learning model 111. The electronic platform may include a graphical user interface (GUI) displayed on the electronic data sources 120, the end-user devices 140, and / or the administrator computing device 150. An example of the electronic platform generated and hosted by the analytics server 110a may be a web-based application or a website configured to be displayed on various electronic devices, such as mobile devices, tablets, personal computers, and the like. Simply put, the analytics server 110a may implement the platform to receive data and instructions from end users (using the end-user devices 140); The analytics server 110a may then execute the machine learning model 111 accordingly and display the results (e.g., a list of “suitable” or “best” mutations for tracking a patient’s progress over the course of treatment or cancer recurrence) on the platform.
[0117] The analytics server 110a may be any computing device comprising a processor and non-transitory, machine-readable storage capable of executing the various tasks and processes described herein. The analytics server 110a may employ various processors such as a central processing unit (CPU) and graphics processing unit (GPU), among others. Nonlimiting examples of such computing devices may include workstation computers, laptop computers, server computers, and the like. While the system 100 includes a single analytics server 110a, the analytics server 110a may include any number of computing devices operating in a distributed computing environment, such as a cloud environment.
[0118] The electronic data sources 120 may represent various sources that contain, retrieve, and / or access data needed to train the machine learning model 111. For instance, the analytics server 110a may use a laboratory computer 120a, medical professional device 120b, server 120c (associated with a laboratory), and / or database 120d (associated with a research lab, a clinic, and / or any third party providing data) to retrieve and receive data. As used herein, the electronic data sources may include any electronic source containing data (e.g., WGS data) that can be used to generate a training dataset in order to (ultimately) train theAtty. Dkt. No.: 131588-1668machine learning model 111. Even though referred to herein as “laboratory” devices, these devices may not always be operated in laboratories. Therefore, no limitation is intended by this term.10119] When generating the training data, the analytics server 110a may execute various algorithms to translate raw data received or retrieved from the electronic data sources 120 into machine-readable objects that can be stored and processed by other analytical processes as described herein.
[0120] End-user devices 140 may be any computing device comprising a processor and a non-transitory, machine-readable storage medium capable of performing the various tasks and processes described herein. Non-limiting examples of an end-user device 140 may be a workstation computer, laptop computer, tablet computer, or server computer. During operation, various users may use end-user devices 140 to access the GUI operationally managed by the analytics server 110a. Specifically, the end-user devices 140 may include laboratory computer 140a, laboratory server 140b, and a user device 140c. Even though referred to herein as “end-user” or “laboratory” devices, these devices may not always be operated by end-users or in laboratories. Therefore, no limitation is intended by these terms.[01211 The administrator computing device 150 may represent a computing device operated by a system administrator. The administrator computing device 150 may be configured to display various attributes and predictions generated by the analytics server 110a (e.g., various analytic metrics determined during training of one or more machine learning models and / or systems); monitor various models 111 utilized by the analytics server 110a, electronic data sources 120, and / or end-user devices 140; review feedback; and / or facilitate training or retraining (calibration) of the machine learning model 111 that are maintained by the analytics server 110a.|0122] The machine learning model 111 may be stored in the system database 110b. The machine learning model 111 may be trained using data received or retrieved from the electronic data sources 120 and may be executed using data received from the end-user devices 140. In some embodiments, the machine learning model 111 may reside within a data repository that is local or specific to a laboratory or an end user (e.g., client). In otherAtty. Dkt. No.: 131588-1668embodiments, the machine learning model may be stored centrally where access to its predictions are controlled per client basis.
[0123] It should be understood that any alternative and / or additional machine learning model(s) may be used to implement similar learning engines. As described herein, the analytics server 110a may store the machine learning model 111 (e.g., neural networks, random forest, support vector machines, regression models, recurrent models, etc.) in an accessible data repository. The analytics server 110a may retrieve the machine learning model 111 and train it to predict a suitable subset of mutations for tracking a patient’s progress over the course of treatment or cancer recurrence. Various machine learning techniques can be used to train the machine learning model 111, such as supervised learning techniques, unsupervised learning techniques, or semi-supervised learning techniques, among others.
[0124] In operation, the analytics server may train the machine learning model 111, such that (at the inference phase), the machine learning model 111 may determine one or more mutations from a set of available mutations. Using a higher number of target sites may increase the number of opportunities to detect ctDNA. Thus, if the machine learning model 111 chooses a fixed number of target sites for the assay, there may be a difference in the chance of observing ctDNA (e.g., if the accuracy in identifying target sites is 100% versus 90%). Using the methods and systems discussed herein, the analytics server can use one or more processors discussed herein to improve the accuracy of target site identification to increase the average number of valid target sites included in the MRD panels. The analytics server 110a may use the machine learning model 111 and / or an algorithmic approach in order to achieve this.
[0125] Using the training data (as generated using the methods and protocols discussed herein), the machine learning model 111 can be trained. In some embodiments, the machine learning model 111 may refer to (or utilize) a random forest model with features and variant classifications. In some embodiments, the training data can be segmented into training and validation / testing folds. For instance, the data can be folded into a 60 / 40 split (60% training / 40% test). During training, various hyperparameter tuning protocols can beAtty. Dkt. No.: 131588-1668implemented to select the optimal number of trees and the number of variables in each tree (of the random forest). These hyperparameters can then be used to train the machine learning model 111.
[0126] The analytics server may use various methodologies to evaluate the performance of the machine learning model 111. For instance, in some embodiments, out-of-bag error to assess the performance rather than k-fold cross-validation; however, other implementations and embodiments may include other methodologies.
[0127] In some embodiments, determining a set of putative tumor-specific somatic variants comprises training a machine learning model to select the plurality of putative somatic variant sites.IV. Detection and Monitoring of TumorsA. cfDNA
[0128] Cell-free nucleic acids, including cfDNA, can be obtained by various methods from biological samples including but not limited to plasma, serum, and urine. Other biological fluid samples include, but are not limited to blood, sweat, tears, sputum, ear flow, lymph, saliva, cerebrospinal fluid, ravages, bone marrow suspension, vaginal flow, transcervical lavage, brain fluid, ascites, milk, secretions of the respiratory, intestinal and genitourinary tracts, amniotic fluid, milk, and leukophoresis samples. In some embodiments, the sample is a sample that is easily obtainable by non-invasive procedures, e.g., blood, plasma, serum, sweat, tears, sputum, urine, ear flow, saliva or feces. In some embodiments, the sample is a peripheral blood sample, or the plasma and / or serum fractions of a peripheral blood sample. In some embodiments, the biological sample is a swab or smear, a biopsy specimen, or a cell culture. In some embodiments, the sample is a mixture of two or more biological samples, e.g., a biological sample can comprise two or more of a biological fluid sample, a tissue sample, and a cell culture sample.[0129[ cfDNA is present as fragments averaging about 170 bp. Accordingly, further fragmentation of cfDNA is not needed. In some embodiments, sufficient cfDNA is obtained from a 10 ml blood sample to confidently determine the presence or absence of cancer in aAtty. Dkt. No.: 131588-1668patient. The blood samples used in the method provided can be of about 5 ml, about 10 ml, about 15 ml, about 20 ml, about 25 ml or more than 25 ml. Typically, 20 ml of blood plasma contains between 5,000 and 10,000 genome equivalents, and provides more than sufficient cfDNA for determining tumor fraction according to the method provided. In some embodiments, sufficient cfDNA is obtained from 10 ml to 20 ml of blood to determine tumor fraction.
[0130] To separate cfDNA from cells in a sample, various methods including, but not limited to fractionation, centrifugation (e.g., density gradient centrifugation), DNA-specific precipitation, or high-throughput cell sorting and / or other separation methods can be used. Commercially available kits for manual and automated separation of cfDNA are available (Roche Diagnostics, Indianapolis, Ind., Qiagen, Germantown, MD).[01311 cfDNA can be end-repaired, and optionally dA tailed, and double-stranded adaptors comprising sequences complementary to amplification and sequencing primers are ligated to the ends of the cfDNA molecules to enable NGS sequencing, e.g., using an Illumina platform. Additionally, each of the double-stranded adaptors further comprises a non-random barcode sequence, which serves to differentiate individual cfDNA molecules. In some embodiments, the barcode sequences are random sequences. In other embodiments, the barcode sequences are non-random barcode sequences. Non-random barcode sequences provide a significant advantage over random barcode sequences because non-random barcode sequences enable unambiguous identification of the sequencing reads described below. The non-random barcode sequences are designed specifically to be base-balance both within and across all barcodes. Additionally, in some embodiments, the non-random barcodes can comprise a T nucleotide at the 3' end, which is complementary to the A nucleotide of dA-tailed cfDNA molecules. In embodiments utilizing a T nucleotide overhang at the 3' end of the barcode, barcodes of three different lengths can be designed to avoid a single base flashing across the entire flowcell of the sequencer. Non-random barcode sequences can be present in adaptors as sequences of 13, 14, and 15 bp; 10, 11, and 12 bp; 11, 12, and 13 bp; 13, 14, and 15 bp; 14, 15, and 16 bp; 15, 16, and 17 bp, and the like. In some embodiments, the shortest barcode sequence can be 8 bp and the longest barcode sequence can be 100 bp.Atty. Dkt. No.: 131588-1668B. Sequencing and Analysis[01321 The disclosed methods generally comprise sequencing one or more samples.Sequencing methods include, but are not limited to, Maxam- Gilbert sequencing-based techniques, chain-termination-based techniques, shotgun sequencing, bridge PCR sequencing, single-molecule real-time sequencing, ion semiconductor sequencing (Ion Torrent sequencing), nanopore sequencing, pyrosequencing (454), sequencing by synthesis, sequencing by ligation (SOLiD sequencing), sequencing by electron microscopy, dideoxy sequencing reactions (Sanger method), massively parallel sequencing, polony sequencing, duplex sequencing, and DNA nanoball sequencing. In some embodiments, sequencing involves hybridizing a primer to the template to form a template / primer duplex, contacting the duplex with a polymerase enzyme in the presence of a detectably labeled nucleotides under conditions that permit the polymerase to add nucleotides to the primer in a templatedependent manner, detecting a signal from the incorporated labeled nucleotide, and sequentially repeating the contacting and detecting steps at least once, wherein sequential detection of incorporated labeled nucleotide determines the sequence of the nucleic acid. In some embodiments, the sequencing comprises obtaining paired end reads. The accuracy or average accuracy of the sequence information may be greater than 80%, 90%, 95%, 99% or 99.98%. In some embodiments, the sequence information obtained is more than 50 bp, 100 bp or 200 bp. The sequence information may be obtained in less than 1 month, 2 weeks, 1 week 1 day, 3 hours, 1 hour, 30 minutes, 10 minutes, or 5 minutes. The sequence accuracy or average accuracy may be greater than 95% or 99%. Examples of detectable labels include radiolabels, florescent labels, enzymatic labels, etc. In some embodiments, the detectable label may be an optically detectable label, such as a fluorescent label. Examples of fluorescent labels include cyanine, rhodamine, fluorescien, coumarin, BODIPY, alexa, or conjugated multi-dyes. In some embodiments, the nucleotide is flagged if one or more of its sequence segments are substantially similar to one or more sequence segments of another nucleotide within the same partition. In some embodiments, sequences can be selectively and preferentially sequenced from the region of interest.
[0133] Captured sequences can be analyzed using the sequencing-by— synthesis technology of Illumina, which uses fluorescent reversible terminator deoxyribonucleotides. The readsAtty. Dkt. No.: 131588-1668generated by the sequencing process are aligned to a reference sequence and associated with a sequence of the somatic sequence panel specific for the patient. Mapping of the sequence reads can be achieved by comparing the sequence of the reads with the sequence of the reference genome to determine the specific genetic information, and optionally the chromosomal origin of the sequenced nucleic acid (e.g., cfDNA) molecule. A number of computer algorithms are available for aligning sequences, including without limitation BLAST (Altschul et al., 1990), BLITZ (MPsrch) (Sturrock & Collins, 1993), FASTA (Person & Lipman, 1988), BOWTIE (Langmead et al, Genome Biology 10:R25.1-R25.10
[2009] ), or ELAND (Illumina, Inc., San Diego, Calif., USA). In one embodiment, the sequencing data is processed by bioinformatic alignment analysis for the Illumina Genome Analyzer, which uses the Efficient Large-Scale Alignment of Nucleotide Databases (ELAND) software. Additional software includes SAMtools (SAMtools, Bioinformatics, 2009, 25(16):2078-9), and the Burroughs-Wheeler block sorting compression procedure which involves block sorting or preprocessing to make compression more efficient.[0134| The barcoded cfDNA fragments isolated form the patient's fluid sample, e.g., blood sample, can be amplified, e.g., by PCR, and captured using the hybrid probes. Capturing of the barcoded fragments comprises obtaining single strands of barcoded cfDNA, and hybridizing the barcoded cfDNA with different hybrid probes. Each of the different hybrid probes hybridizes to a single-stranded barcoded cfDNA target sequence to form a targethybrid probe duplex. The duplex is isolated from unhybridized cfDNA by binding the purification binding moiety comprised in the hybrid probe to the corresponding purification moiety binding partner. As described elsewhere herein, the corresponding purification moiety binding partner can be immobilized on a solid surface, e.g., a magnetic bead, which facilitates the separation of the capture duplex from unhybridized cfDNA molecules in solution. The barcoded cfDNA of the duplex is released, and is subjected to sequencing using an NGS instrument.
[0135] In some embodiments, sequencing of cfDNA comprises whole genome sequencing (WGS), whole exome sequencing (WES), targeted sequencing, or subtractive hybridization. In some embodiments, the sequencing is in cfDNA from a sample of blood, plasma, or serum from the cancer patient, thereby obtaining a second set of sequencing reads.Atty. Dkt. No.: 131588-1668[0136[ In some embodiments, the ctDNA sequences in the second set of sequencing reads based on a sequencing read comprising one or more putative somatic variant sites from the set of putative tumor-specific somatic variants can be detected.
[0137] In some embodiments, the tumor fraction can then be calculated based on the presents of ctDNA in the sample of cfDNA.C. Correction based on tumor allele balance, dropout rate, latent variant classes and sample or panel-specific error rates.
[0138] The methods as disclosed herein, utilizes correction based on tumor allele balance and correction of sites that do not contribute to ctDNA signal in plasma (e.g., dropout rate), and / or inclusion of latent variant classes to account for somatic calling errors. Additionally, correction based on defining and accounting for sample and / or panel-specific error rates.
[0139] The methods as disclosed herein, may utilize various types of modeling in the cfDNA analysis and ctDNA detection (e.g., tumor fraction detection) in which a set of putative tumor-specific somatic variants identified in the first set of sequence reads (e.g., sequenced DNA from a tumor sample) is modeled as a mixture of variants coming from the tumor, nontumor somatic variants (e.g., CHIP mutations), error-prone sites, and germline variants. For each variant “class” a likelihood function was developed describing the probability of observing the ALT-bearing reads given the total set of reads at a site if the variant is in a given class. The total likelihood of the data for each variant is the average of the per-class likelihoods weighted by the proportion of variants in each class, and the total likelihood of the data for all variants is the product of per-variant likelihoods. Only the tumor class includes tumor fraction (e.g., detection of ctDNA) as a parameter of the likelihood function, so fitting the model allows quantification of tumor fraction while marginalizing out the signal from non-tumor variants, which reduces the chance of a false positive when the target set includes non-tumor variants.
[0140] For the purposes of the methods described herein, all class likelihoods were modeled as different binomial models on the count of variant-bearing molecules given the totalAtty. Dkt. No.: 131588-1668molecular depth at a site. However, in some embodiments, the different models could be binomial models, gaussian models, or Poisson models.
[0141] The “probability” parameter of the binomial model can be set to 0.5 for the germline class, 0.01 - 0.1 for the CHIP+error-prone class (i.e., the “noise” class), and for the tumor class the probability parameter is a freely varying value in (0,1) calculated as a function of the copy number and genotype in the tumor at each target site. Future iterations could expand on the per-class likelihood functions by introducing information about insert sizes or molecule start sites, both of which can help differentiate tumor from germline or noise sites. As explained further below, the model is fit with expectation-maximization.i. Correction based on allele balance.
[0142] Allele balance can also be utilized for correction of detecting the presence or absence of ctDNA sequences in the cfDNA (e.g, second set of sequence reads) which is a ratio of variant allele to reference allele in sequencing data, using tumor purity. This approach is modeled on the relationship between the probability of observing alternative allele counts, / ?, at a site given allele balance in the tumor ABT) copy number of the site in both tumor (CNT) and normal (CNN) and the fraction of tumor DNA in the sample, the tumor purity TP) and the error rate estimated from control site data (err '.TF • CNT
[10143] 1p = - — — - ■ AB / TP + errTF-CN + (1 - TF) -CNN
[0144] Each variant will be ranked by an expected p of 0.005% TF using allele balance and copy number in the tumor sample, assuming copy number in the non-tumor sample is 2. This approach accounts for the sub-clonal variants on detection sensitivity.ii. Correction for dropout
[0145] Correction of detecting the presence or absence of ctDNA sequences in the cfDNA (e.g, second set of sequence reads) can be performed by correcting for sites that do not contribute to ctDNA signal. Such sites include sites due to biological reasons (e.g., lack of shedding into the blood stream) or technical reasons (e.g., false positive somatic variant call). These factors are collectively referred to as a dropout rate (DR) which is the false discoveryAtty. Dkt. No.: 131588-1668rate (FDR) + (1 - FDR) multiplied by biological false negative rate (alternatively plasma dropout rate) (PDR).DR = FDR + (1 - FDR) ■ PDRThe dropout rate (DR) is used to weight a binomial mixture model of the base model + an error- only model:where:L(p\n, k) = Pr(n,k\p) and k = ALT countsWhereby a higher likelihood ration indicates a greater confidence in tumor recurrence:iii. Correction for variant class
[0146] Correction of detecting the presence or absence of ctDNA sequences in the cfDNA (e.g., second set of sequence reads) can be performed by correcting for sites that do not contribute to ctDNA signal, such as, applying a classification model. In some embodiments, detecting the presence or absence of ctDNA sequences in the second set of sequencing reads comprises applying a classification model to the putative somatic variant sites to classify each as either tumor-specific, non-tumor-specific, or germline variant. Occasionally, germline variants, sites with systematic noise, low frequency clonal hematopoiesis of indeterminate potential (CHIP) variants or other anomalous sites may be selected. Detection of such sites can negatively influence the fit of the base model and result in a false positive MRD status.
[0147] To avoid an incorrect fit of the base + dropout mixture model, the mixture model can be expanded to include correction for latent variant classes. In some embodiments, detecting the presence or absence of ctDNA sequences in a second set of sequencing reads comprises correcting for, using a computer processor, inclusion of latent variant classes in the second set of sequencing reads. In some embodiments, the second set of sequencing reads is from the sequenced cfDNA. In some embodiments, correcting for inclusion of latent variant classes in the second set of sequencing reads comprises comparing the second set of sequencing reads to a reference dataset.Atty. Dkt. No.: 131588-1668
[0148] . In some embodiments, the probability that each putative somatic variant site belongs to a class is determined based on relative likelihoods of different binomial models, gaussian models, or Poisson models. In some embodiments, the non-tumor somatic variant and / or germline variant classes use a fixed parameter for the probability of observing a variant count. In some embodiments, each putative somatic variant site in the set of putative tumorspecific somatic variants belongs to a class selected from (i) a tumor-specific somatic variant, (ii) a non-tumor-specific somatic variant or mutation due to technical or biological noise, and (iii) a germline variant incorrectly identified as somatic. In some embodiments, the probability of observing an alternate allele count for each class is modeled as a statistical distribution (e.g., a binomial distribution, a negative binomial distribution, gaussian distribution, or Poisson distribution). In some embodiments, the statistical distribution includes a probability parameter determined by analysis of one or more reference sets of nontumor somatic variants and / or germline variants. In some embodiments, comprising a total likelihood of observing variant and non-variant counts for each putative somatic variant site is calculated.iv. Correction for sample or panel-specific error rates
[0149] Sequencing DNA from the sample from the subject may optionally further comprises error correction with duplex unique molecule identifiers (UMI). Duplex UMIs are short sequences that tag both strands of a DNA molecule. This technique is used to identify and correct errors in sequencing, and to detect mutations in DNA. This is done by identifying reads from each strand of the original fragment by looking for reads with the same alignment position, complementary orientations, and swapped UMIs. In general, sequences with the same barcode are collapsed to form a consensus sequence that minimizes random sequencing and late cycle PCR errors.
[0150] While the use of unique molecular identifiers (UMIs), tagging individual DNA molecules with a unique barcode that are associated with all sequences deriving from that molecule, is one approach to error-reduction in next generation sequencing (NGS), early cycle PCR errors are less likely to be corrected by this UMI approach. The presence of this residual error interferes with the interpretation of somatic variant detection from samples withAtty. Dkt. No.: 131588-1668minute fractions of ctDNA. Thus, for each site, a control site or a set of control sites can be selected.
[0151] A control site can be in any sequence reads, or are local to the somatic variant site, within about, for example, 160 nucleotides, 120 nucleotides, or 20 nucleotides and have the same reference base as the somatic variant site. In some embodiments, the control site is located within about 1 nucleotide to about 3 nucleotides upstream or downstream. At the control sites, counts of the reference and variant bases are collected. For each mutation type, (e.g., A > G) a statistical model (e.g., a binomial distribution, a negative binomial distribution, gaussian distribution, or Poisson distribution) for error rate can fit by maximum likelihood estimation. The mutation type-specific error rates are weighted by the proportion of that mutation type included in the MRD panel to derive a sample / panel-specific error rate. That rate can be used as the background probability of detecting a variant allele in a tumor fraction inference model. The advantage of defining a sample / panel-specific error rate is that it controls for differences in how samples were collected and processed (e.g., if PCR conditions vary slightly between NGS library preparations) as well as differences in composition of panels between patients (e.g., if a panel include more of an error-prone mutation type).
[0152] In some embodiments, the methods as disclosed herein, comprises calculating the fraction of ctDNA present in the sample of cfDNA. Such calculations can be performed by fitting a mixture model of variant counts and total counts across the entire panel of variants assayed, where the mixture components are the tumor, non-tumor, or germline variants and the weighting of each class is determined by the probability that a variant belongs to that class. In some embodiments, the mixture model is a binomial mixture model, a negative binomial mixture model, gaussian mixture model, or Poisson mixture model.
[0153] In some embodiments, detecting the presence or absence of ctDNA sequences in the sequence reads of the sequenced cfDNA e.g., second set of sequencing reads) comprises: (i) calculating, using a computer processor, an allele balance of the somatic variants that is optionally adjusted based on purity of the tumor sample obtained from the subject, (ii) calculating, using a computer processor, a copy number for each somatic variant, (iii)Atty. Dkt. No.: 131588-1668correcting for, using a computer processor, an error rate for each somatic variant, or (iv) any combination thereof. In some embodiments, calculating the copy number for each somatic variant comprises a copy number for each somatic variant in the tumor sequencing data (e.g., first set of sequence reads) the liquid sample sequencing data (e.g., second set of sequence reads) or both the tumor sequencing data and the liquid sample sequencing data. In some embodiments, correcting for the error rate comprises an error rate for each somatic variant based on the liquid sample sequencing data, training data, or both the liquid sample sequencing data and training data.
[0154] In some embodiments, detecting ctDNA sequences in the second set of sequencing reads (e.g., sequenced cfDNA) comprises calculating, using a computer processor, dropout rate of sites that do not contribute to a ctDNA signal, wherein the sites are not non-tumor, noise, or germline. In some embodiments, detecting ctDNA sequences in the second set of sequencing reads (e.g., sequenced cfDNA) is based on a sequencing read comprising one or more putative somatic variant sites from the set of putative tumor-specific somatic variants, by classifying each putative somatic variant site as (i) a tumor-specific somatic variant, (ii) a non-tumor-specific somatic variant or mutation due to technical or biological noise, and (iii) a germline variant incorrectly identified as somatic. In some embodiments, classifying each putative somatic variant site as a tumor-specific somatic variant comprises adjusting a likelihood of the classification, using a computer processor, based on a dropout rate of sites that do not contribute to a ctDNA signal, wherein the sites are not non-tumor, noise, or germline.
[0155] The disclosed methods are applicable to MRD testing, wherein the patient has previously been treated for a cancer, and may be considered in remission, however a small number of cancer cells remain in the body. The number of remaining cells may be so small that they do not cause any physical signs or symptoms and often cannot even be detected through traditional methods, such as viewing cells under a microscope and / or by tracking abnormal serum proteins in the blood. An MRD positive test results means that residual (remaining) disease was detected. A negative result means that residual disease was not detected. MRD testing may be used to measure the effectiveness of treatment and to predict if a patient is at risk of relapse. When a patient tests positive for MRD, it means that there areAtty. Dkt. No.: 131588-1668still residual cancer cells in the body after treatment. When MRD is detected, this is known as “MRD positivity.” When a patient tests negative, no residual cancer cells were found. When no MRD is detected, this is known as “MRD negativity.”
[0156] In some embodiments, the patient has completed at least one cancer treatment (e.g., chemotherapy, radiotherapy, surgery, immunotherapy, cell therapy, or biologic therapy) prior to obtaining the tumor sample for determination of the set of putative tumor-specific sites.[0157| The set of putative tumor-specific sites serves as a signature panel for the patient and cfDNA from the patient can be sequenced at various stages of the disease, i.e., the signature panel can be screened to determine the presence of cancer at surgery following diagnosis; during cancer treatment, e.g, at intervals during chemotherapy or radiation therapy, to monitor the efficacy of the treatment; at intervals during remission to confirm continued absence of disease; and / or to detect recurrence of the disease. The composition of the selected somatic variants for the subset is a key determinant for the sensitivity and specificity of the methods described herein.
[0158] In some embodiments, repeating sequencing cfDNA from the cancer patient can be at 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more successive time points. In some embodiments, the cfDNA is from a sample of blood, plasma, or serum from the cancer patient obtained at the successive time. In some embodiments, sequencing cfDNA from a sample of blood, plasma, or serum from the cancer patient is repeated one or more times while the patient is in remission; is repeated one or more times while the patient is undergoing treatment for the cancer; is repeated one or more times coinciding with or prior to surgery; is repeated one or more times coinciding with following, during, or prior to administration of chemotherapy; is repeated one or more times coinciding with following, during, or prior to radiation therapy; is repeated one or more times coinciding with following, during, or prior to administration of an immunotherapy; is repeated one or more times coinciding with following, during, or prior to administration of a cell therapy; or is repeated one or more times coinciding with or following, during, or prior to administration of a biologic therapy.[0159[ The foregoing corrections can be performed via one or more machine learning of AI-based models. Thus, models utilized in accordance with technologies of the presentAtty. Dkt. No.: 131588-1668disclosure can thus determine, by at least one processor, (a) allele balance, (b) dropout rate, (c) latent variant classes, and / or (d) sample or panel-specific error rates. In some embodiments, the model is selected from a convolutional neural network, a support vector machine, a boosted regression tree, a random forest, a fully-connected neural network, a transformer neural network, and a recurrent neural network. In some embodiments, the model comprises a convolutional neural network. In some embodiments, the model comprises a support vector machine. In some embodiments, the model comprises a boosted regression tree. In some embodiments, the model comprises a random forest. In some embodiments, the model comprises a fully-connected neural network. In some embodiments, the model comprises a transformer neural network. In some embodiments, the model comprises a recurrent neural network. Such an ML or Al model can be trained with tumor DNA samples and / or germline DNA camples with known clinical outcomes for the patients.
[0160] In some embodiments, the subject has, had, or is suspected of having a cancer or a tumor. The type of cancer or tumor can is not particularly limited and can include, but is not limited to, bladder cancer, breast cancer, colon / colorectal cancer, gynecologic cancers, head and neck cancers, hematological cancers, liver cancer, lung cancer, and skin cancer.EXAMPLES
[0161] The present invention is described in further detail in the following examples which are not in any way intended to limit the scope of the invention as claimed. All references cited are herein specifically incorporated by reference for all that is described therein. The following examples are offered to illustrate, but not to limit the claimed invention.Example 1: Improved sensitivity and estimation ofMRD - correcting for allele balance of the somatic variants based on purity of the tumor sample, a dropout rate of sites, and / or inclusion of latent variant classes
[0162] Preparing somatic variant panel: DNA is extracted from a normal and tumor sample from a patient and a sequencing library is prepared for each sample. The samples are sequenced by whole genome sequencing and somatic variants identified. A panel of the somatic variants is then selected by selecting sites having an expected variant alleleAtty. Dkt. No.: 131588-1668frequency based on copy number and allele balance reference cut off and / or a plurality of features selected by a machine learning model trained with whole genome sequencing data. Hybrid capture probes are then generated for the somatic variants of the subpanel.
[0163] Correction of ctDNA sequences: cfDNA is extracted from the patient and selectively enriched using hybrid capture probes for the somatic variants panel. The enriched library is sequenced to generate sequencing reads for each of the somatic variant panel. A computer processor is used to calculate the allele balance of the somatic variants based on purity of the tumor sample, a dropout rate of sites that do not contribute to a ctDNA signal, and / or inclusion of latent variant classes in the sequence library by comparing the sequence library to a reference dataset).
[0164] Diagnosing MRD: MRD is diagnosed based on the presence of the corrected patient specific somatic variants in the cfDNA sample.Example 2: Improved sensitivity and estimation of MRD - correcting sample or panel specific error rates
[0165] Preparing somatic variant panel: DNA is extracted from a normal and tumor sample from a patient and a sequencing library is prepared for each sample. The samples are sequenced by whole genome sequencing and somatic variants identified. A panel of the somatic variants is then selected by selecting sites having an expected variant allele frequency based on copy number and allele balance reference cut off and / or a plurality of features selected by a machine learning model trained with whole genome sequencing data. Hybrid capture probes are then generated for the somatic variants of the subpanel.
[0166] Correction using a control site: cfDNA is extracted from the patient and selectively enriched using hybrid capture probes for the somatic variant and a corresponding control site, where each corresponding control site for each somatic variant is located within 20 nucleotide bases of the somatic variant on the DNA fragment and each corresponding control site comprises a reference based that is the same as the base of the somatic variant. The enriched library is then sequenced to generate sequencing reads for each of the somatic variant panel.Atty. Dkt. No.: 131588-1668[0167| Diagnosing MRD: MRD is diagnosed based on the presence of the corrected patient specific somatic variants in the cfDNA sample.Example 3: Methods for tumor fraction / ctDNA detectionVariables
[0168] The data considered in this model includes: X: the count of ALT alleles at a somatic target site z; and m: the total sequencing depth at a somatic target site i.
[0169] Additional data is passed from the “tumor profile” step: ct the copy number of the tumor at site z; at . the allele balance of the target allele in the tumor FFPE sample at site z (note this is what is observed in the sample before accounting for tumor purity); tp. the tumor purity, or genome-equivalents proportion of DNA from the FFPE tumor sample coming from tumor vs patient normal tissue; cn / the copy number of the patient normal at site z. For convenience it is assumed that this is 2 on all autosomes and either 1 or 2 on allosomes depending on patient sex; an / the allele balance of the target allele in the patient normal (buffy coat) sample at site z. Additional variables include: tf. tumor fraction. The genomeequivalent proportion of cfDNA coming from the tumor, at / ', the expected allele balance if variant z if it came from mixture class k.The Model
[0170] Target sites can come from three classes: (1) somatic variants in the tumor, (2) nontumor somatic variants (e.g., a variant that arises from clonal hematopoiesis of indeterminate potential, a “CHIP mutation”) present in normal tissue or immune cells infiltrating the tumor FFPE sample, or (3) germline heterozygous sites incorrectly identified as somatic due to a masking failure in WGS somatic calling. The variants identified in (2) and (3) are examples of latent variants. Ideally all the variants would be from the tumor, however in practice other variants are not always dropped, and high-frequency variants significantly inflate tumor fraction estimates and detection likelihood ratios. Hard allele balance or insert size cutoffs have also been attempted to just mask non-tumor variants, but these were difficult to set in a principled way and were not successful in experiments, in which there were a number ofAtty. Dkt. No.: 131588-1668samples with significant numbers of CHIP-like detections, accordingly, the variants should be dealt with flexibly in the model.
[0171] The likelihood of the data (X, n) as a mixture of the three classes listed above can be modeled, indexed by k. The likelihood of the data is the standard mixture model formula (equation 1):
[0172] 7tk variable is the “mixture weights”, or the proportion of variants coming from each k class.
[0173] L(Xiik\ ... ) is the likelihood of observing X ALT reads at site i if it is in class k. The binomial probability mass function is used for all the classes, which has parameters n (the total (deduplicated) sequencing depth) and p (the probability of sampling an ALT-bearing read from the pool of molecules covering a target site).Detection Probabilities
[0174] Each of the k mixture classes is modeled as a binomial probability B X\n. p). The p values are the probability of success, i.e. the probability of drawing a target mutation from the pool of molecules sequenced at a site.Tumor DNA[0175| If k = 1 and the site is from the tumor, then the probability of observing an ALT-bearing read is the expected allele balance of the target mutation in the cell-free DNA mixture. This depends both on the proportion of tumor molecules in the cfDNA (the tumor fraction), and the proportion of target mutations in the pure tumor (itself a function of the copy number, genotype, and subclonality of the target variant, plus the tumor purity which is used to correct the observed allele balance in the FFPE sample for what it would look like if it was 100% pure tumor). The full form of this is (equation 2):Atty. Dkt. No.: 131588-1668where tf is the tumor fraction, and e is the error rate (discussed more below). The two terms in the numerator give the number of ALT alleles expected to come from the tumor and the patient, and the four terms on the bottom give the total number of ALT or REF alleles coming from tumor or patient normal, with the goal being to get to the standard allele balance formula: a = ALT / (ALT+REF).CHIP mutations
[0176] CHIP mutations are somatic variants found in white blood cells but not (necessarily) associated with the tumor. Because buffy is used as the “normal” for tumor-normal somatic calling, most high-frequency CHIP mutations should be dropped at that stage. But low-frequency CHIP mutations may get through because the typical WGS depth (~30x) is not high enough to regularly detect somatic variants present at just 1-2%.
[0177] Currently, there is no principled way of modeling the expected frequency of CHIP variants. By definition, it can be said that they don't come from tumor, so t_f= 0. But their copy number or allele balance in patient normal is not known. Empirically, looking at previously generated data, the average allele balance of target mutations detected in the buffy coat is 1.7%.Germline Variants
[0178] The somatic calling and site prioritization algorithm is designed to mask out any position variable in the patient normal, but some errors inevitably get through. This can happen, for example, in cases where there is relatively low depth on the patient normal sample which fails to sample any reads bearing one of the two alleles at a heterozygous site or homozygous site. If the variant is observed in the tumor sample sequence reads, it appears as a high-frequency somatic mutation and will be upweighted by the selection algorithm. It is possible that these germline sites are interpreted as coming from the tumor, because they also have very high frequency in the cfDNA, and will therefore significantly throw off theAtty. Dkt. No.: 131588-1668maximum likelihood fit for tumor fraction. The germline sites can be either germline heterozygous sites, germline homozygous sites (i.e., homozygous alternative “HOMALT” targets), or a combination thereof.|0179] A germline heterozygous site by definition has at= 0 and an= 0.5. It is also assumed that it is CN2 in tumor and normal (ct = 2, cn= 2). Furthermore, because it is coming from the patient normal only, it can be modeled using equation 2 with the above values filled in, giving the expected value of 0.5.Error Rates
[0180] On a per-read level target mutations can be observed through genuine non-biological error. The deduplication and error correction pipeline is designed to minimize these, but they still happen (currently at a rate of about 1 error per million bases). In particular, PCR errors occurring in the first few cycles are likely to get through the error correction pipeline because they will be propagated to up to half the molecules eventually sequenced and therefore a majority of error-containing molecules may end up being sampled.[01811 The separate “error model” described here estimates this rate. Essentially, for each target mutation, all the other positions covered by the same probe matching the target reference base are found, and the number of times a read matching the target somatic mutation is observed are counted (for example, if the target is T->A at chrl: 100, all the T positions from chrl: 50-150 are found and it is determined how often a read bearing A at those positions is observed). Alternatively, all sites with the same target reference based are found, not limited to any base size range. For the purposes of the tumor fraction model an error rate value is obtained that can be plugged in to e above.Fitting the Model
[0182] In this example, it is assumed that e is known. The other free parameters are / / (the tumor fraction), and the vector of mixture weights it.
[0183] Expectation-maximization (EM) was used to fit the model. Plausible estimations are used for the parameters of interest, e.g., tf= 0.001 and it = [0.95, 0.04, 0.01] (listing it in theAtty. Dkt. No.: 131588-1668order of [tumor, CHIP, germline]. Under these parameter values an initial likelihood matrix are filled out with one row per mixture class (tumor / CHIP / germline) and one column per site. Each entry is the binomial log-likelihood (i.e., the probability mass) of observing X ALT reads given total sequencing depth n and a probability of detection p determined by the equations above (equation 2 for tumor variants, 0.01725 for CHIPs, and 0.5 for germline variants).
[0184] The EM loop is then entered. In the “E” step the mixture weights were updated. The most likely mixture class for each variant is the class with the maximum likelihood value in each column of the likelihood matrix, which is achieved by running an argmax function down the columns. The proportion of variants in each mixture class is then calculated and used as a new set of mixture weights.[0185| In the “M” step, the new mixture weights are used to find a new optimum tumor fraction estimate. Finally, the total likelihood of the data is calculated using equation 1 (but in log space and using the logSumExp function to approximate log(B(X|n, p)).10186] This procedure can then be repeated an arbitrary number of times. Each iteration should improve the total likelihood, and repeats are stopped when the change in likelihood falls below a threshold.Example 4: Proof of Concept Study
[0187] DNA was extracted from a tumor cell line and was whole genome sequenced (WGS).Somatic variants were called by Mutect2 in tumor-only mode using a panel of normal sequences generated from WGS of buffy coat (white blood cell) DNA from clinical samples. Somatic variants were filtered to enrich for bonafide somatic variants. Filters included depth in the tumor WGS (>50x), the variant allele fraction (>0.3), presence of the variant allele in population sequencing (dbSNP database) and overlap with a pre-selected region of interest generated by a process that excluded poorly mapping regions of the genome, excessively high or low GC content, and repetitive sequence. After filtering, 465 targets remained.
[0188] Counts of alleles at all remaining target sites and adjacent control sites (for error rate estimation) were gathered from deduplicated and error-corrected targeted sequencing dataAtty. Dkt. No.: 131588-1668from 4 contrived samples. Contrived samples were generated by mixing fragmented and size selected DNA from the tumor cell line characterized by WGS and a matched normal cell line from the same patient at pre-defined proportions spanning 4 orders of magnitude.
[0189] Target site counts were fit to a multi-level binomial mixture model with 4 latent target classifications: tumor, noise, germline HET, and germline HOMALT. The tumor class included components that accounted for error rate and the probability that a somatic target is a false positive (dropout rate).
[0190] The number of targets detected across the samples scaled based on the tumor fraction expected from the contrived mixture proportion, as shown in FIG. 2. Notably, the model consistently classifies approximately 14% of targets as originating from germline variants.10191] Despite the high prevalence of germline targets included in the target set and the high proportion of germline targets detected in the low tumor fraction samples, the model provided a consistent, linear series of estimates for tumor fraction across 4 orders of magnitude, as shown in FIG. 3. These estimates were slightly biased in this sample set, but highly concordant with the expected tumor fraction based on the contrived mixture proportions.EQUIVALENTS
[0192] The present technology is not to be limited in terms of the particular embodiments described in this application, which are intended as single illustrations of individual aspects of the present technology. Many modifications and variations of this present technology can be made without departing from its spirit and scope, as will be apparent to those skilled in the art. Functionally equivalent methods and apparatuses within the scope of the present technology, in addition to those enumerated herein, will be apparent to those skilled in the art from the foregoing descriptions. Such modifications and variations are intended to fall within the scope of the present technology. It is to be understood that this present technology is not limited to particular methods, reagents, compounds, compositions or biological systems, which can, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting.Atty. Dkt. No.: 131588-1668[0193| All patents, patent applications, provisional applications, and publications referred to or cited herein are incorporated by reference in their entirety, including all figures and tables, to the extent that are not inconsistent with the explicit teachings of this specification.
Claims
Atty. Dkt. No.: 131588-1668WHAT IS CLAIMED IS:
1. A method of detecting circulating tumor DNA (ctDNA) in a sample, comprising:(a) sequencing DNA from a tumor sample obtained from a patient, thereby obtaining a first set of sequencing reads of the DNA from the tumor sample;(b) determining a set of putative tumor-specific somatic variants comprising a plurality of putative somatic variant sites;(c) sequencing cell-free DNA (cfDNA) from a sample of blood, plasma, or serum from the cancer patient, thereby obtaining a second set of sequencing reads; and (d) detecting ctDNA sequences in the second set of sequencing reads based on a sequencing read comprising one or more putative somatic variant sites from the set of putative tumor-specific somatic variants.
2. The method of claim 1, wherein determining a set of putative tumor-specific somatic variants is based on allele balance, copy number alteration (CNA), or a combination thereof for each putative somatic variant site in the first set of sequencing reads,3. The method of claim 1, wherein determining a set of putative tumor-specific somatic variants comprises training a machine learning model to select the plurality of putative somatic variant sites,4. The method of claim 1, wherein determining a set of putative tumor-specific somatic variants comprises calling, via a computer processor, the plurality of putative somatic variant sites based on one or more assemblies of haplotypes.
5. The method of any one of claims 1-4, wherein the method does not comprise sequencing DNA from a non-tumor sample from the patient for determining the set of tumor-specific somatic variants.
6. The method of any one of claims 1-5, wherein the cfDNA from the sample of blood, plasma, or serum from the cancer patient is not enriched prior to sequencing.
7. The method of any one of claims 1-6, wherein determining the set of tumor-specific somatic variants further comprises determining tumor purity of the tumor sample.Atty. Dkt. No.: 131588-16688. The method of any one of claims 1-7, wherein sequencing DNA from the tumor sample comprises whole genome sequencing, whole exome sequencing, targeted sequencing, or subtractive hybridization.
9. The method of any one of claims 1-8, wherein detecting the presence or absence of ctDNA sequences in the second set of sequencing reads comprises correcting for, using a computer processor, inclusion of latent variant classes in the second set of sequencing reads.
10. The method of claim 9, wherein correcting for inclusion of latent variant classes in the second set of sequencing reads comprises comparing the second set of sequencing reads to a reference dataset.
11. The method of any one of claims 1-10, wherein sequencing the cfDNA comprises whole genome sequencing , whole exome sequencing, targeted sequencing, or subtractive hybridization.
12. The method of any one of claims 1-11, wherein detecting the presence or absence of ctDNA sequences in the second set of sequencing reads comprises applying a classification model to the putative somatic variant sites to classify each as either tumor-specific, non-tumor-specific, or germline variant.
13. The method of any one of claims 1-12, further comprising determining a probability that each putative somatic variant site belongs to the class based on relative likelihoods of different binomial models, gaussian models, or Poisson models.
14. The method of claim 13, wherein the non-tumor somatic variant and / or germline variant classes use a fixed parameter for the probability of observing a variant count.
15. The method of any one of claims 1-14, wherein each putative somatic variant site in the set of putative tumor-specific somatic variants belongs to a class selected from (i) a tumor-specific somatic variant, (ii) a non-tumor-specific somatic variant or mutation due to technical or biological noise, and (iii) a germline variant incorrectly identified as somatic.Atty. Dkt. No.: 131588-166816. The method of any one of claims 1-15, wherein a probability of observing an alternate allele count for each class is modeled as a statistical distribution.
17. The method of claim 16, wherein the statistical distribution is a binomial distribution, a negative binomial distribution, gaussian distribution, or Poisson distribution.
18. The method of claim 16 or 17, wherein the statistical distribution includes a probability parameter determined by analysis of one or more reference sets of nontumor somatic variants and / or germline variants.
19. The method of any one of claims 1-18, further comprising calculating a total likelihood of observing variant and non-variant counts for each putative somatic variant site.
20. The method of any one of claims 1-19, further comprising calculating the fraction of ctDNA present in the sample of cfDNA.
21. The method of claim 20, wherein calculating the fraction of ctDNA comprises fitting a mixture model of variant counts and total counts across the entire panel of variants assayed, where the mixture components are the tumor, non-tumor, or germline variants and the weighting of each class is determined by the probability that a variant belongs to that class.
22. The method of claim 21, wherein the mixture model is a binomial mixture model, a negative binomial mixture model, gaussian mixture model, or Poisson mixture model.
23. The method of any one of claims 1-22, wherein detecting the presence or absence of ctDNA sequences in the second set of sequencing reads comprises:(i) calculating, using a computer processor, an allele balance of the somatic variants that is optionally adjusted based on purity of the tumor sample obtained from the subject,(ii) calculating, using a computer processor, a copy number for each somatic variant, (iii) correcting for, using a computer processor, an error rate for each somatic variant,Atty. Dkt. No.: 131588-1668or(iv) any combination thereof.
24. The method of claim 23, wherein calculating the copy number for each somatic variant comprises a copy number for each somatic variant in the tumor sequencing data, the liquid sample sequencing data, or both the tumor sequencing data and the liquid sample sequencing data.
25. The method of claim 23 or 24, wherein correcting for the error rate comprises an error rate for each somatic variant based on the liquid sample sequencing data, training data, or both the liquid sample sequencing data and training data.
26. The method of any one of claims 1-25, wherein detecting the presence or absence of ctDNA sequences in the second set of sequencing reads comprises calculating, using a computer processor, a dropout rate of sites that do not contribute to a ctDNA signal, wherein the sites are not non-tumor, noise, or germline.
27. A method of detecting circulating tumor DNA (ctDNA) in a sample, comprising: sequencing DNA from a tumor sample obtained from a patient, thereby obtaining a first set of sequencing reads of the DNA from the tumor sample;determining a set of putative tumor-specific somatic variants comprising a plurality of putative somatic variant sites; andsequencing cell-free DNA (cfDNA) from a sample of blood, plasma, or serum from the cancer patient, thereby obtaining a second set of sequencing reads; and detecting ctDNA sequences in the second set of sequencing reads based on a sequencing read comprising one or more putative somatic variant sites from the set of putative tumor-specific somatic variants, wherein detecting comprises classifying each putative somatic variant site as (i) a tumor-specific somatic variant, (ii) a non- tumor-specific somatic variant or mutation due to technical or biological noise, and (iii) a germline variant incorrectly identified as somatic.
28. The method of claim 27, wherein classifying each putative somatic variant site as (i) a tumor-specific somatic variant further comprises adjusting a likelihood of theAtty. Dkt. No.: 131588-1668classification, using a computer processor, based on a dropout rate of sites that do not contribute to a ctDNA signal, wherein the sites are not non-tumor, noise, or germline.
29. The method of claim 27 or 28, wherein determining a set of putative tumor-specific somatic variants is based on allele balance, copy number alteration (CNA), or a combination thereof for each putative somatic variant site in the first set of sequencing reads.
30. The method of claim 27 or 28, wherein determining a set of putative tumor-specific somatic variants comprises training a machine learning model to select the plurality of putative somatic variant sites.
31. The method of claim 27 or 28, wherein determining a set of putative tumor-specific somatic variants comprises calling, via a computer processor, the plurality of putative somatic variant sites based on one or more assemblies of haplotypes.
32. The method of any one of claims 27-31, wherein the method does not comprise sequencing DNA from a non-tumor sample from the patient for determining the set of tumor-specific somatic variants.
33. The method of any one of claims 27-32, wherein the ctDNA from the sample of blood, plasma, or serum from the cancer patient is not enriched prior to sequencing.
34. The method of any one of claims 27-33, wherein detecting ctDNA sequences in the second set of sequencing reads further comprises:(i) calculating, using a computer processor, an allele balance of the somatic variants that is optionally adjusted based on purity of the tumor sample obtained from the subject,(ii) calculating, using a computer processor, a copy number for each somatic variants, (iii) correcting for, using a computer processor, an error rate for each somatic variant, or(iv) any combination thereof.Atty. Dkt. No.: 131588-166835. The method of claim 34, wherein calculating the copy number for each somatic variant comprises a copy number for each somatic variant in the tumor sequencing data, the liquid sample sequencing data, or both the tumor sequencing data and the liquid sample sequencing data.
36. The method of claim 34 or 35, wherein correcting for the error rate comprises an error rate for each somatic variant based on the liquid sample sequencing data, training data, or both the liquid sample sequencing data and training data.
37. The method of any one of claims 27-36, wherein detecting the presence or absence of ctDNA sequences in the second set of sequencing reads comprises correcting for, using a computer processor, inclusion of latent variant classes in the second set of sequencing reads.
38. The method of claim 37, wherein correcting for inclusion of latent variant classes in the second set of sequencing reads comprises comparing the second set of sequencing reads to a reference dataset.
39. A method of detecting circulating tumor DNA (ctDNA) in a sample, comprising: sequencing cell-free DNA (cfDNA) from a sample of blood, plasma, or serum from a cancer patient, thereby obtaining a set of sequencing reads, wherein the cfDNA from the sample of blood, plasma, or serum from the cancer patient is not enriched prior to sequencing; anddetecting ctDNA sequences in the set of sequencing reads based on a sequencing read comprising one or more putative somatic variant sites from a set of putative tumorspecific somatic variants, wherein detecting comprises(a) classifying each putative somatic variant site as (i) a tumor-specific somatic variant, (ii) a non-tumor-specific somatic variant, and (iii) a germline variant incorrectly identified as somatic, and(b) adjusting variant calls, using a computer processor, based on a dropout rate of sites that do not contribute to a ctDNA signal, wherein the sites are not non-tumor, noise, or germline.Atty. Dkt. No.: 131588-166840. The method of claim 39, wherein the set of putative tumor-specific somatic variants is tumor-informed based on based on allele balance, copy number alteration (CNA), or a combination thereof of sequence reads of DNA from a tumor sample of the cancer patient without sequencing DNA from a non-tumor sample of the cancer patient.
41. The method of claim 39 or 40, wherein detecting ctDNA sequences in the second set of sequencing reads further comprises:(i) calculating, using a computer processor, an allele balance of the somatic variants that is optionally adjusted based on purity of the tumor sample obtained from the subject,(ii) calculating, using a computer processor, a copy number for each somatic variants, (iii) correcting for, using a computer processor, an error rate for each somatic variant, or(iv) any combination thereof.
42. The method of claim 41, wherein calculating the copy number for each somatic variant comprises a copy number for each somatic variant in the tumor sequencing data, the liquid sample sequencing data, or both the tumor sequencing data and the liquid sample sequencing data.
43. The method of claim 41 or 42, wherein correcting for the error rate comprises an error rate for each somatic variant based on the liquid sample sequencing data, training data, or both the liquid sample sequencing data and training data.
44. The method of any one of claims 39-43, wherein detecting the presence or absence of ctDNA sequences in the second set of sequencing reads comprises correcting for, using a computer processor, inclusion of latent variant classes in the second set of sequencing reads.
45. The method of claim 44, wherein correcting for inclusion of latent variant classes in the second set of sequencing reads comprises comparing the second set of sequencing reads to a dataset.Atty. Dkt. No.: 131588-166846. The method of any one of claims 39-45, wherein classifying each putative somatic variant site comprises modeling a probability of observing an alternate allele count for each class as a statistical distribution.
47. The method of claim 46, wherein the statistical distribution is a binomial distribution, a negative binomial distribution, gaussian distribution, or Poisson distribution.
48. The method of claim 46 or 47, wherein the statistical distribution includes a probability parameter determined by analysis of one or more reference sets of nontumor somatic variants and / or germline variants.
49. The method of any one of claims 39-48, further comprising determining a probability that each putative somatic variant site belongs to the class based on relative likelihoods of different binomial models.
50. The method of claim 49, wherein the non-tumor somatic variant and / or germline variant classes use a fixed parameter for the probability of observing a variant count.
51. The method of any one of claims 39-50, further comprising calculating a total likelihood of observing variant and non-variant counts for each putative somatic variant site.
52. The method of any one of claims 39-51, further comprising calculating the fraction of ctDNA present in the sample of cfDNA.
53. The method of any one of claims 1-52, wherein the patient has completed at least one cancer treatment prior to obtaining the tumor sample.
54. The method of claim 53, wherein the cancer treatment is selected from chemotherapy, radiotherapy, surgery, immunotherapy, cell therapy, or biologic therapy.
55. The method of any one of claims 1-54, further comprising repeating sequencing cfDNA from a sample of blood, plasma, or serum from the cancer patient at 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more successive time points.Atty. Dkt. No.: 131588-166856. The method of claim 55, wherein sequencing cfDNA from a sample of blood, plasma, or serum from the cancer patient is repeated one or more times while the patient is in remission.
57. The method of claim 55, wherein sequencing cfDNA from a sample of blood, plasma, or serum from the cancer patient is repeated one or more times while the patient is undergoing treatment for the cancer.
58. The method of claim 55, wherein sequencing cfDNA from a sample of blood, plasma, or serum from the cancer patient is repeated one or more times coinciding with or prior to surgery; following, during, or prior to administration of chemotherapy; following, during, or prior to radiation therapy; following, during, or prior to administration of an immunotherapy; following, during, or prior to administration of a cell therapy; or following, during, or prior to administration of a biologic therapy.
59. The method of any one of claims 1-58, wherein the tumor is from a cancer selected from bladder cancer, breast cancer, colon / colorectal cancer, gynecologic cancers, head and neck cancers, hematological cancers, liver cancer, lung cancer, and skin cancer.