Methods for detecting contamination of DNA samples using germline variant allele fractions

WO2026170134A1PCT designated stage Publication Date: 2026-08-13MYRIAD WOMENS HEALTH INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-02-09
Publication Date
2026-08-13
Patent Text Reader

Abstract

The present disclosure provides, among other things, methods of detecting contamination in a sample, such as a DNA sample. The disclosed methods can be used to detect contamination using germline variants, and avoid the need for additional molecular workflow or molecular spike-in prior to sequencing.
Need to check novelty before this filing date? Find Prior Art

Description

Atty. Dkt. No.: 131588-1665METHODS FOR DETECTING CONTAMINATION OF DNA SAMPLES USING GERMLINE VARIANT ALLELE FRACTIONS CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of and priority to U.S. Provisional Application No.63 / 756,686, filed on February 10, 2025, the entire disclosure of which is incorporated by reference herein.BACKGROUND[00021 The following description of the background of the present technology is provided simply as an aid in understanding the present technology and is not admitted to describe or constitute prior art to the present technology.

[0003] Cell-free DNA assays often require high sensitivity for detecting specific mutations in a sample. For example, molecular residual disease (MRD) assays may need to detect somatic variants from a patient’s tumor at about 1 part per million sensitivity. Contamination of patient samples by external DNA sources could cause false negatives if a sample is diluted below an assay’s limit of detection, or false positives if the contaminant bears a variant expected to be present in the primary sample. Existing cross-contamination assays generally use a molecular spike-in approach, but this cannot detect contamination prior to spike-in addition. Alternative methods of detecting contamination generally involve comparing a sample’s DNA sequence to an external source (for example, sequence data from a previous sequencing run of the same individual). Accordingly, there is a need to simplify, streamline, and improve contamination detection in nucleic acid samples.SUMMARY OF THE INVENTION

[0004] The present disclosure provides methods for detecting contamination in a nucleic acid sample that avoid the need for additional molecular workflow updates and allow contamination detection without prior sequencing of an individual. The present disclosure provides, among other things, methods of detecting contamination using germline variant allele fractions. Thus, the methods of the present disclosure improve current methods of -1- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665detecting contamination in a nucleic acid sample, which are useful in, for example, biomedical and forensic applications.

[0005] In one aspect, the present disclosure provides methods of detecting contamination in a sample, comprising: sequencing DNA from a sample from the subject; and detecting contamination in the sample based on (i) the presence of a sequencing read comprising an alt-allele at one or more of a plurality of germline variants from a panel of germline variants and (ii) an allele fraction of the alt-allele.

[0006] In some embodiments, the plurality of germline variants comprise at least one common germline variant. In some embodiments, the plurality of germline variants consists of common germline variants.

[0007] In some embodiments, the sample is a sample of blood, plasma, serum, or urine.

[0008] In some embodiments, the subject was previously diagnosed with cancer.

[0009] In some embodiments, the methods further comprise detecting in the sample sequence reads comprising one or more tumor-specific somatic mutations that are determined for the subject prior to sequencing the sample.

[0010] In some embodiments, the methods further comprise enriching the DNA from the sample for DNA comprising one or more germline variants of the panel of germline variants prior to sequencing. In some embodiments, enriching the DNA from the sample comprises hybrid capture enrichment or PCR-based enrichment.

[0011] In some embodiments, detecting contamination does not comprise a spike-in.

[0012] In some embodiments, sequencing DNA from the sample from the subject further comprises error correction, optionally with duplex unique molecule identifiers (UMI).

[0013] In another aspect, the present disclosure provides methods of detecting contamination in a sample, comprising: sequencing DNA from a germline sample from a subject, thereby obtaining a set of germline sequencing reads, and sequencing DNA from a tumor sample obtained from the subject, thereby obtaining a set of tumor sequencing reads; determining a -2- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665panel of tumor-specific somatic mutations based on the set of germline sequencing reads and the set of tumor sequencing reads; sequencing cell-free DNA (cfDNA) from a sample from the subject, thereby obtaining a set of cfDNA sequencing reads; and detecting contamination in the sample based on (i) the presence of a sequencing read in the set of cfDNA sequencing reads comprising an alt-allele at one or more of a plurality of germline variants of a panel of germline variants and (ii) an allele fraction of the alt-allele.

[0014] In some embodiments, the plurality of germline variants comprise at least one common germline variant. In some embodiments, the plurality of germline variants consists of common germline variants.

[0015] In some embodiments, the sample is a sample of blood, plasma, serum, or urine.

[0016] In some embodiments, the methods further comprise enriching the DNA from the sample for DNA comprising one or more of the panel of common germline variants prior to sequencing. In some embodiments, enriching the DNA from the sample comprises hybrid capture enrichment or PCR-based enrichment.

[0017] In some embodiments, detecting contamination in the sample does not comprise a spike-in.

[0018] In some embodiments, sequencing cfDNA from the sample from the subject further comprises error correction with duplex unique molecule identifiers (UMI).

[0019] In some embodiments, in some embodiments, the methods further comprise detecting the presence of circulating tumor DNA (ctDNA) in the sample based on the present of sequencing read in the set of cfDNA sequencing reads comprising one or more of the panel of tumor-specific somatic mutations.

[0020] In some embodiments of any of the foregoing aspects or embodiments, the germline variants in the panel of germline variants have a minor allele frequency of about 50%.

[0021] In some embodiments of any of the foregoing aspects or embodiments, detecting contamination comprises quantifying the sequence reads comprising an alt-allele at one or more of the germline variants and quantifying the total number of sequencing reads of each -3- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665germline variant in the panel of germline variants, thereby obtaining an allele count comprising an alt-allele count and a total allele count. In some embodiments, the methods may further comprise modeling the allele count as a statistical distribution, optionally a binomial distribution, a negative binomial distribution, gaussian distribution, or Poisson distribution. In some embodiments, the statistical distribution includes a probability parameter determined by analysis of one or more reference sets of germline variants.

[0022] In some embodiments of any of the foregoing aspects or embodiments, the methods further comprise determining a probability that each alt-allele belongs to the subject or originated from contamination based on relative likelihoods or expected allele fraction. In some embodiments, the expected allele fraction uses a fixed parameter for the probability of observing contamination.[00231 In some embodiments of any of the foregoing aspects or embodiments, the methods further comprising calculating a total likelihood for contamination for each germline variant in the set of germline variants.10024] In some embodiments of any of the foregoing aspects or embodiments, the methods further comprise calculating an allele fraction of genotypes of each germline variant in the set of germline variants. In some embodiments, calculating the allele fraction comprises fitting a model of alt allele counts and total allele counts across the panel of germline variants. In some embodiments, the model is a binomial model, a negative binomial model, gaussian model, or Poisson model10025] In some embodiments of any of the foregoing aspects or embodiments, the panel of germline variants comprises at least one germline variant that is not expected to be homozygous relative to its reference allele.

[0026] In some embodiments of any of the foregoing aspects or embodiments, sequencing DNA from the sample comprises whole genome sequencing, whole exome sequencing, targeted sequencing, or subtractive hybridization.

[0027] In some embodiments, sequencing DNA from the germline sample and / or sequencing DNA from the tumor sample comprises whole genome sequencing, whole exome sequencing,-4- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665targeted sequencing, or subtractive hybridization. In some embodiments, the subject has completed at least one cancer treatment prior to obtaining the tumor sample. In some embodiments, the cancer treatment is selected from chemotherapy, radiotherapy, surgery, immunotherapy, cell therapy, or biologic therapy.

[0028] In some embodiments of any of the foregoing aspects or embodiments, the methods further comprise sequencing DNA from another sample from the subject at 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more successive time points. In some embodiments, sequencing DNA from another sample from the subject is repeated one or more times while the patient is in remission. In some embodiments, sequencing DNA from another sample from the subject is repeated one or more times while the patient is undergoing treatment for the cancer. In some embodiments, sequencing DNA from another sample from the subject is repeated one or more times coinciding with or prior to surgery; following, during, or prior to administration of chemotherapy; following, during, or prior to radiation therapy; following, during, or prior to administration of an immunotherapy; following, during, or prior to administration of a cell therapy; or following, during, or prior to administration of a biologic therapy.

[0029] In some embodiments of any of the foregoing aspects or embodiments, the subject has, had, or is suspected of having a cancer. In some embodiments, the cancer is selected from bladder cancer, breast cancer, colon / colorectal cancer, gynecologic cancers, head and neck cancers, hematological cancers, liver cancer, lung cancer, and skin cancer.

[0030] In another aspect, the present disclosure provides methods of preparing an enriched DNA fraction, comprising: obtaining a sample comprising cfDNA from a subject; and enriching from the cfDNA a DNA fraction comprising one or more of a plurality of germline variants from a panel of germline variants by contacting the cfDNA with a plurality of primer pairs or probes, wherein each primer pair or probe in the plurality is specific for a DNA fragments comprising a germline variant of the panel of germline variants.

[0031] In some embodiments, the sample is a sample of blood, plasma, serum, or urine. In some embodiments, the subject was previously diagnosed with cancer. In some embodiments, the methods further comprise enriching from the cfDNA a DNA fraction comprising one or-5- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665more tumor-specific somatic mutations that are determined for the subject prior to enrichment.

[0032] In some embodiments, enriching the DNA fraction comprises hybrid capture enrichment or PCR-based enrichment.

[0033] In some embodiments, the plurality of germline variants comprise at least one common germline variant. In some embodiments, the plurality of germline variants consists of common germline variants.

[0034] In some embodiments, the methods further comprise sequencing the DNA fraction. In some embodiments, the methods further comprise detecting contamination in the sample based on (i) the presence of a sequencing read comprising an alt-allele at one or more of a plurality of germline variants from a panel of germline variants and (ii) an allele fraction of the alt-allele.DETAILED DESCRIPTION

[0035] The present disclosure provides, among other things, methods of detecting contamination in a sample, comprising: sequencing DNA from a sample from the subject; and detecting contamination in the sample based on (i) counting the number of reads including non-reference alleles at common germline variant positions (ii) identifying variant positions with atypical allele fractions, and (iii) fitting a statistical model to estimate the most likely contamination rate of the sample. The present disclosure also provides methods of detecting contamination in a sample, comprising: sequencing DNA from a germline sample from a subject, thereby obtaining a set of sequencing reads, and sequencing DNA from a tumor sample obtained from the subject, thereby obtaining a set of tumor sequencing reads; determining a panel of tumor-specific somatic mutations based on the set of germline sequencing reads and the set of tumor sequencing reads; sequencing cell-free DNA (cfDNA) from a sample from the subject, thereby obtaining a set of cfDNA sequencing reads; and detecting contamination in the sample based on (i) the presence of a sequencing read in the set of cfDNA sequencing reads comprising an alt-allele at one or more of a plurality of germline variants of a panel of germline variants and (ii) an allele fraction of the alt-allele.-6- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665

[0036] Unlike prior methodologies, the disclosed methods allow for detecting contamination of patient samples (e.g., DNA samples) without prior knowledge of an individual’s genome sequence and without a spike-in. In general, a set of germline single nucleotide variants are identified and added to a sequencing panel (the sequencing panel could be, for example, a patient-specific MRD panel targeting somatic variants in the patient’s tumor, a multiplex PCR of several regions of the genome, or an in-silico subset of a whole genome sequencing run). The germline variant positions are selected such that they have minor allele frequency that are preferably near 50% in the intended patient population. DNA from regions around these SNV positions can then be captured and sequenced along with the primary targets of an assay’s analysis (e.g., the patient-specific and tumor-specific somatic mutations of an MRD panel). Contamination from external sources of DNA can be detected by fitting a statistical model to read counts covering the germline SNV positions based on the allele fraction of the germline SNVs.

[0037] The present disclosure also provides, among other things, methods of preparing an enriched DNA fraction, comprising: obtaining a sample comprising cfDNA from a subject; and enriching from the cfDNA a DNA fraction comprising one or more of a plurality of germline variants from a panel of germline variants by contacting the cfDNA with a plurality of primer pairs or probes, wherein each primer pair or probe in the plurality is specific for a DNA fragments comprising a germline variant of the panel of germline variants.

[0038] It is to be appreciated that certain aspects, modes, embodiments, variations and features of the present methods are described below in various levels of detail in order to provide a substantial understanding of the present technology.

[0039] The present disclosure is not to be limited in terms of the particular embodiments described in this application, which are intended as single illustrations of individual aspects of the disclosure. All the various embodiments of the present disclosure will not be described herein. Many modifications and variations of the disclosure can be made without departing from its spirit and scope, as will be apparent to those skilled in the art. Functionally equivalent methods and apparatuses within the scope of the disclosure, in addition to those enumerated herein, will be apparent to those skilled in the art from the foregoing descriptions.-7- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665Such modifications and variations are intended to fall within the scope of the appended claims. The present disclosure is to be limited only by the terms of the appended claims, along with the full scope of equivalents to which such claims are entitled.

[0040] It is to be understood that the present disclosure is not limited to particular uses, methods, reagents, compounds, compositions or biological systems, which can, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting.Definitions[0041 [ Unless defined otherwise, all technical and scientific terms used herein have the meaning commonly understood by a person skilled in the art to which this disclosure belongs. The following references provide one of skill with a general definition of many of the terms used in the present disclosure. Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed.1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, the following terms have the meanings ascribed to them below, unless specified otherwise. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure.[0042 [ As used herein, the term “alt-allele fraction” refers to the proportion (or “fraction”) of sequencing reads covering a variant position that do not match the reference (e.g., wild-type) allele.

[0043] As used herein, the term “amplification,” with respect to nucleic acid sequences, refers to methods that increase the representation of a population of nucleic acid sequences in a sample. Copies of a particular target nucleic acid sequence generated in vitro in an amplification reaction are called “amplicons” or “amplification products”. Amplification may be exponential or linear. A target nucleic acid may be DNA (such as, for example, genomic DNA, cfDNA, ctDNA, and cDNA) or RNA. Amplification can be achieved using polymerase-8- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665chain reaction (PCR) as well as numerous other methods such as isothermal methods, rolling circle methods, etc.

[0044] As used herein, the term “approximately” or “about” means plus or minus 10% as well as the specified number. For example, “about 10” should be understood as both “10” and “9-11”.

[0045] As used herein, the term “biopsy” refers to a tissue sample excised from a subject (e.g., a subject with cancer). Tissue samples may be obtained using any suitable method, including, but not limited to, needle biopsies, aspiration, scraping, excision using surgical equipment, etc.

[0046] As used herein, the term “comparable” refers to two (or more) sets of conditions, circumstances, individuals, or populations that are sufficiently similar to one another to permit comparison of results obtained or phenomena observed. In some embodiments, comparable sets of conditions, circumstances, individuals, or populations are characterized by a plurality of substantially identical features and one or a small number of varied features. Those of ordinary skill in the art will appreciate that sets of circumstances, individuals, or populations are comparable to one another when characterized by a sufficient number and type of substantially identical features to warrant a reasonable conclusion that differences in results obtained or phenomena observed under or with different sets of circumstances, individuals, or populations are caused by or indicative of the variation in those features that are varied. Those skilled in the art will appreciate that relative language used herein (e.g., enhanced, activated, reduced, inhibited, etc.) will typically refer to comparisons made under comparable conditions.[0047J As used herein, the term “derived from” encompasses the terms “originated from,” “obtained from,” “obtainable from,” “isolated from,” and “created from,” and generally indicates that one specified material (e.g., a biological sample) finds its origin in another specified material or individual or has features that can be described with reference to another specified material.-9- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665

[0048] As used herein, term “genomic DNA” refers to DNA of a cellular genome. The genomic DNA can be cellular, i.e., contained within a cell, or it can be cell-free.

[0049] As used herein, the term “common” when used in the context of “common germline variant(s)” refers to variant or SNPs for which the minor allele frequency is approximately 50%. For example, to be considered a “common germline variant” the minor allele frequency is inclusive of 40% to 60%.[0050| As used herein, the term “minor allele frequency” refers to the proportion (“frequency”) of a less prevalent (e.g., second most prevalent) allele in a population. For example, if a particular loci on an allele contains an adenosine in 60% of a population and a thymine 40% of the population, the “minor allele frequency of a thymine at that loci is 40%.

[0051] As used herein, the term “heterozygous” refers to a loci or gene having different nucleotides or nucleotide sequences on each allele. For example, a heterozygous gene may comprise one Single Nucleotide Polymorphism (SNP) on one allele, and one wild-type nucleotide at the corresponding site on the other allele.

[0052] As used herein, the term “homozygous alternate” refers to a loci comprising two identical, non-wild-type (“variant”) nucleotides at the loci (e.g., one variant nucleotide on the sense strand and one variant nucleotide on the anti-sense strand) or a gene comprising two identical, non-wild-type (“variant”) alleles.[0053J As used herein, the term “homozygous reference” refers to a loci comprising two identical, wild-type (“non-varianf ’) nucleotides at the loci (e.g., one non-variant nucleotide on the sense strand and one non-variant nucleotide on the anti-sense strand) or a gene comprising two identical, wild-type (“non-variant”) alleles.

[0054] As used herein, the term “genotyping” refers to a process of determining the alleles an individual at particular genetic loci by examining an individual’s DNA. Genotyping differs from sequencing in which all of the nucleotides comprising a specific length of DNA are assessed.-10- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665

[0055] As used herein, the terms “improved”, “increased”, or “reduced”, or grammatically comparable comparative terms, indicate values that are relative to a comparable reference measurement. For example, in some embodiments, an assessed value achieved with an agent of interest may be “improved” relative to that obtained with a comparable reference agent. Alternatively or additionally, in some embodiments, an assessed value achieved in a subject or system of interest may be “improved” relative to that obtained in the same subject or system under different conditions, or in a different, comparable subject (e.g., in a comparable subject or system that differs from the subject or system of interest in presence of one or more indicators of a particular disease, disorder or condition of interest, or in prior exposure to a condition or agent, etc.). In some embodiments, comparative terms refer to statistically relevant differences (e.g., that are of a prevalence and / or magnitude sufficient to achieve statistical relevance). Those skilled in the art will be aware, or will readily be able to determine, in a given context, a degree and / or prevalence of difference that is required or sufficient to achieve such statistical significance.[0056| As used herein, the term “multi-nucleotide variant” or “MNV” refers to a variant having 2 or more adjacent nucleotide changes.

[0057] As used herein, the term “Next Generation Sequencing” or “NGS” refers to sequencing methods that allow for massively parallel sequencing of clonally amplified and of single nucleic acid molecules during which a plurality, e.g., millions, of nucleic acid fragments from a single sample or from multiple different samples are sequenced in unison. Non-limiting examples of NGS include sequencing-by-synthesis, sequencing-by-ligation, real-time sequencing, and nanopore sequencing.

[0058] As used herein, the term “patient-specific panel” or “patient-specific somatic variants” refers to a collection of sequences comprising somatic mutations that are specific to a patient, or markers that distinguish between two or more individuals. A signature panel may distinguish one sample from another.

[0059] As used herein, the term “sample” or “biological sample,” refers to a biological sample obtained or derived from a source of interest, as described herein. In certain embodiments, a source of interest comprises an organism, such as a microbe, a plant, an -11- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665animal or a human. In certain embodiments, a biological sample is or comprises biological tissue or fluid. In certain embodiments, a biological sample may be or comprise bone marrow; blood (or a fraction thereof, such as plasma or serum); blood cells; ascites; tissue or fine needle biopsy samples; cell-containing body fluids; free floating nucleic acids (e.g., cell free DNA); sputum; saliva; urine; cerebrospinal fluid, peritoneal fluid; pleural fluid; lymph; gynecological fluids; skin swabs; vaginal swabs; oral swabs; nasal swabs; washings or lavages such as a ductal lavages or broncheoalveolar lavages; aspirates; scrapings; bone marrow specimens; tissue biopsy specimens; surgical specimens; feces, other body fluids, secretions, and / or excretions; and / or cells therefrom, etc. In certain embodiments, a biological sample is or comprises cells obtained from an individual. In certain embodiments, obtained cells are or include cells from an individual from whom the sample is obtained. In certain embodiments, a sample is a “primary sample” obtained directly from a source of interest by any appropriate means. For example, in certain embodiments, a primary biological sample is obtained by methods selected from the group consisting of a swab, biopsy (e.g., fine needle aspiration or tissue biopsy), surgery, collection of body fluid (e.g, blood, lymph, feces etc.), etc. In certain embodiments, as will be clear from context, the term “sample” refers to a preparation that is obtained by processing (e.g, by removing one or more components of and / or by adding one or more agents to) a primary sample. For example, filtering using a semi-permeable membrane. Such a processed “sample” may comprise, for example nucleic acids or proteins extracted from a sample or obtained by subjecting a primary sample to techniques such as amplification or reverse transcription of mRNA, isolation and / or purification of certain components, etc.

[0060] As used herein, the term “sequence read,” or simply “read,” refers to sequence information of a nucleic acid fragment obtained through a sequencing assay, such as a next generation sequencing (NGS) assay. In some embodiments, a sequence read refers to data representing a sequence of nucleotide bases that were measured using a clonal sequencing method. Clonal sequencing may produce sequence data representing a single molecule, or clones or clusters of one original DNA molecule. A sequence read may also have associated quality score at each base position of the sequence indicating the probability that nucleotide has been called correctly.-12- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665[00611 As used herein, a “set” of reads refers to all sequencing reads with a common parent nucleic acid strand.

[0062] As used herein, the term “Single Nucleotide Polymorphism” or “SNP” refers to a single base pair variation in a nucleic acid sequence. SNPs can also be referred to as Single-Nucleotide Variants (SNVs).

[0063] As used herein, the term “somatic variant” or “somatic mutation” refers to a variant arising after conception, in non-germline DNA of an individual. Somatic variants may include single-nucleotide variants (SNVs), multi -nucleotide variants, insertions and deletions (e.g., indel variants), and genomic rearrangements for example. The terms “somatic variant” and “somatic mutation” are used interchangeably herein.

[0064] As used herein, the term “spike-in” or “molecular spike-in” refers to a process in which a nucleic acid molecule (which may be synthetic) of known sequence and quantity is added to a biological sample prior to sequencing as a control.

[0065] As used herein, the term “subject” or “patient” or “individual” refers to any organism upon which embodiments of the present disclosure may be used or administered, e.g., for experimental, screening, diagnostic, prophylactic, and / or therapeutic purposes. Typical subjects include animals (e.g., mammals such as mice, rats, rabbits, non-human primates, and humans; insects; worms; etc.).

[0066] As used herein, the term “target sequence” refers to a selected target polynucleotide, e.g., a sequence present in a cfDNA molecule, whose presence, amount, and / or nucleotide sequence, or changes in these, are desired to be determined. Target sequences can be interrogated for the presence or absence of a somatic and / or germline variant. The target polynucleotide can be a region of gene associated with a disease. In some embodiments, the region is an exon. The disease can be cancer.

[0067] As used herein, the term “tumor fraction” refers to the proportion of circulating cell-free tumor DNA (ctDNA) relative to the total amount of cell-free DNA (cfDNA). Tumor fraction may be indicative of the size of the tumor.-13- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665[0O68| As used herein, the term “tumor-specific somatic mutations” refer to nucleic acid e.g., DNA) changes (“variants”) that occur in a somatic cell before or during tumor development. This type of variant is not present within the germline.

[0069] As used herein in the context of molecules, e.g., nucleic acids, proteins, or small molecules, the term “variant” refers to a molecule that shows significant structural identity with a reference molecule but differs structurally from the reference molecule, e.g., in the presence or absence or in the level of one or more chemical moieties as compared to the reference entity. In some embodiments, a variant also differs functionally from its reference molecule. In general, whether a particular molecule is properly considered to be a “variant” of a reference molecule is based on its degree of structural identity with the reference molecule. As will be appreciated by those skilled in the art, any biological or chemical reference molecule has certain characteristic structural elements. A variant, by definition, is a distinct molecule that shares one or more such characteristic structural elements but differs in at least one aspect from the reference molecule. To give but a few examples, a polypeptide may have a characteristic sequence element comprised of a plurality of amino acids having designated positions relative to one another in linear or three-dimensional space and / or contributing to a particular structural motif and / or biological function; a nucleic acid may have a characteristic sequence element comprised of a plurality of nucleotide residues having designated positions relative to one another in linear or three-dimensional space. In some embodiments, a variant polypeptide or nucleic acid may differ from a reference polypeptide or nucleic acid as a result of one or more differences in amino acid or nucleotide sequence and / or one or more differences in chemical moieties (e.g., carbohydrates, lipids, phosphate groups) that are covalently components of the polypeptide or nucleic acid (e.g., that are attached to the polypeptide or nucleic acid backbone). In some embodiments, a variant polypeptide or nucleic acid shows an overall sequence identity with a reference polypeptide or nucleic acid that is at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, or 99%. In some embodiments, a variant polypeptide or nucleic acid does not share at least one characteristic sequence element with a reference polypeptide or nucleic acid. In some embodiments, a reference polypeptide or nucleic acid has one or more biological activities. In some embodiments, a variant polypeptide or nucleic acid shares one or more of the biological activities of the reference polypeptide or nucleic acid. In some -14- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665embodiments, a variant polypeptide or nucleic acid lacks one or more of the biological activities of the reference polypeptide or nucleic acid. In some embodiments, a variant polypeptide or nucleic acid shows a reduced level of one or more biological activities as compared to the reference polypeptide or nucleic acid. In some embodiments, a polypeptide or nucleic acid of interest is considered to be a “variant” of a reference polypeptide or nucleic acid if it has an amino acid or nucleotide sequence that is identical to that of the reference but for a small number of sequence alterations at particular positions. Typically, fewer than about 20%, about 15%, about 10%, about 9%, about 8%, about 7%, about 6%, about 5%, about 4%, about 3%, or about 2% of the residues in a variant are substituted, inserted, or deleted, as compared to the reference. In some embodiments, a variant polypeptide or nucleic acid comprises about 10, about 9, about 8, about 7, about 6, about 5, about 4, about 3, about 2, or about 1 substituted residues as compared to a reference. Often, a variant polypeptide or nucleic acid comprises a very small number (e.g., fewer than about 5, about 4, about 3, about 2, or about 1) number of substituted, inserted, or deleted, functional residues (i.e., residues that participate in a particular biological activity) relative to the reference. In some embodiments, a variant polypeptide or nucleic acid comprises not more than about 5, about 4, about 3, about 2, or about 1 addition or deletion, and, in some embodiments, comprises no additions or deletions, as compared to the reference. In some embodiments, a variant polypeptide or nucleic acid comprises fewer than about 25, about 20, about 19, about 18, about 17, about 16, about 15, about 14, about 13, about 10, about 9, about 8, about 7, about 6, and commonly fewer than about 5, about 4, about 3, or about 2 additions or deletions as compared to the reference. In some embodiments, a reference polypeptide or nucleic acid is one found in nature. In some embodiments, a reference polypeptide or nucleic acid is a human polypeptide or nucleic acid.Methods for Detecting Contamination in a Sample[0070| This disclosure describes improved methods for detecting contamination in nucleic acid samples (e.g., DNA-containing samples) without prior knowledge of the genome sequence of the individual from which the sample was derived and without a spike-in. These methods decrease the need for additional molecular workflow and allow for detection of types of contamination that would be missed by conventional spike-in methods. In general,-15- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665the disclosed methodologies include selecting a set of germline variants, which may be single nucleotide variants (SNV) and may be “common” variants, as described in more detail herein. This set or panel of germline variants can be added to whatever sequencing panel for which the sample was intended. DNA from regions around these germline variants is then captured and sequenced along with the primary targets of an assay’s analysis (i.e., the sequencing panel for which the sample was intended). Contamination from external sources of DNA can then be detected based on the allele frequency of these variants by fitting a statistical model to read counts of sequences of the common the panel of germline variants.

[0071] Those skilled in the art will recognize that there are numerous genetic sequencing assays that are used for biomedical, diagnostic, prognostic, and forensic applications (among others), and the disclosed methods can be incorporated into any such application. For example, the sequencing panel for which a given sample is intended could be a patientspecific MRD panel targeting tumor-somatic variants of the patient, a multiplex PCR of several regions of the genome, or an in-silico subset of a whole genome sequencing run. Regardless of the application of the target sequencing panel, the disclosed methods and the panels of germline variants described here may be used to detect contamination in the sample being sequenced.

[0072] The disclosed methods of detecting contamination in a sample (e.g., a DNA sample) can comprise: sequencing DNA from a sample from the subject; and detecting contamination in the sample based on (i) counting the number of reads including non-reference alleles at common germline variant positions (ii) identifying variant positions with atypical allele fractions, and (iii) fitting a statistical model to estimate the most likely contamination rate of the sample. The disclosed methods of detecting contamination in a sample (e.g., a DNA sample) can comprise: sequencing DNA from a sample from the subject; and detecting contamination in the sample based on (i) the presence of a sequencing read comprising an alt-allele at one or more of a plurality of germline variants from a panel of germline variants and (ii) an allele fraction of the alt-allele. Additionally or alternatively, the disclosed methods of detecting contamination in a sample (e.g., a DNA sample) can comprise: sequencing DNA from a germline sample from a subject, thereby obtaining a set of germline sequencing reads, and sequencing DNA from a tumor sample obtained from the subject, thereby obtaining a set-16- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665of tumor sequencing reads; determining a panel of tumor-specific somatic mutations based on the set of germline sequencing reads and the set of tumor sequencing reads; sequencing cell-free DNA (cfDNA) from a sample from the subject, thereby obtaining a set of cfDNA sequencing reads; and detecting contamination in the sample based on (i) the presence of a sequencing read in the set of cfDNA sequencing reads comprising an alt-allele at one or more of a plurality of germline variants of a panel of germline variants and (ii) an allele fraction of the alt-allele. Germline samples are sample that contain germline DNA and can include, but are not limited to blood, plasma, and buffy coat samples as well as fibroblasts.

[0073] For the purposes of the present disclosure, the panel of germline variants can comprise at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, or more common germline variants. In some embodiments, the panel of germline variants comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, or more common germline variants. In some embodiments, the panel of germline variants comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more common germline variants. In some embodiments, the panel of germline variants consists of a plurality of common germline variants.

[0074] As used herein, a “common germline variant” may have a minor allele frequency of 40% to 60%, 45% to 55%, or about 50%. For example, a “common germline variant” may have a minor allele frequency of 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, or any value in between. Additionally, it should be understood that while common germline variants are useful for practicing the disclosed methods, they are not required. Rather, the disclosed methods can utilize any panel of germline variants, regardless of minor allele frequency, but as the frequency of minor allele decreases, more sites may be required to obtain the same level of power that would have been achieved with relatively fewer common germline variants. In some embodiments, the germline variants in the panel of germline variants have a minor allele frequency of about 50%.-17- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665[0075[ In some embodiments, the sample is a sample of blood, plasma, serum, or urine. The sample may comprise cell free DNA (cfDNA), circulating tumor DNA (ctDNA), or both cfDNA and ctDNA. The sample may comprise DNA, RNA, or both.

[0076] The subject from which the sample was derived is generally a human or another mammal. The subject can have cancer, be suspected of having cancer, or was previously diagnosed with cancer. Alternatively, the suspect may have or be suspected of having another disease, disorder, and or condition for which genetic analysis / sequencing is needed.Additionally or alternatively, the subject may not be a pregnant female and / or the sample may not comprise fetal DNA.

[0077] When the disclosed methods of detecting contamination are incorporated into an MRD analysis, the methods can further comprise detecting in the sample sequence reads comprising one or more tumor-specific somatic mutations that are determined for the subject prior to sequencing the sample. Such embodiments generally comprise detecting the presence of circulating tumor DNA (ctDNA) in the sample based on the present of sequencing read in the set of cfDNA sequencing reads comprising one or more of the panel of tumor-specific somatic mutations.

[0078] Regardless of whether the sample is intended for MRD analysis or another form of genetic analysis / sequencing, the methods can further comprise enriching the DNA from the sample for DNA comprising one or more of the panel of germline variants prior to sequencing. Enriching the DNA from the sample may comprise, for example, hybrid capture enrichment, PCR-based enrichment, on-sequencer enrichment, or other forms of enrichment (e.g., size-based enrichment).

[0079] As noted throughout, this disclosure, use of the disclosed methods avoids the need for a molecular spike-in, but the two approaches to detecting contamination are not mutually exclusive. Accordingly, and in general, the disclosed method of detecting contamination do not comprise a spike-in. However, in some embodiments, the methods may optionally comprise a spike-in.-18- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665[00801 Further, sequencing DNA from the sample from the subject may optionally further comprise error correction, such as error correction with duplex unique molecule identifiers (UMI). Duplex UMIs are short sequences that tag both strands of a DNA molecule. This technique is used to identify and correct errors in sequencing, and to detect mutations in DNA. This is done by identifying reads from each strand of the original fragment by looking for reads with the same alignment position, complementary orientations, and swapped UMIs.[0081) Detecting contamination based on (i) the presence of a sequencing read in the set of cfDNA sequencing reads comprising an alt-allele at one or more of a plurality of germline variants of a panel of germline variants and (ii) an allele fraction of the alt-allele can comprise modeling the allele count or allele fraction of any detected alt-alleles. The allele counts or allele fraction may be modeled as, for example, a binomial distribution.Alternatively, the allele counts or fraction could be modeled as a negative binomial distribution, gaussian distribution, or Poisson distribution. For example, in some embodiments, such a model can be prepared as follows:The data (X) is the count of ALT-bearing read pairs and total read pairs at each common germline variant position. For an unmixed sample there should be 3 genotypes present (0 / 0, 0 / 1, 1 / 1). In a mixed sample of 2 individuals there are 9 genotype combinations: 0 / 0:0 / 0, 0 / 0:0 / l, 0 / 0: 1 / 1, 0 / 1 :0 / 0, 0 / 1 :0 / l, 0 / 1: 1 / 1, 1 / 1 :0 / 0, 1 / 1 :0 / l, 1 / 1: 1 / 1, wherein the genotypes of sample 1 and sample 2 are written in VCF- style notation with a colon separating each sample’s genotype.

[0082] The expected allele fraction of a variant in each genotype class is the average of each genotype's frequency weighted by the contamination rate c:E(af) = (1-c) * af_l + c * af_2af_l and af_2 above are the expected allele fractions in each individual given the genotype.

[0083] For example, a 0 / 0 genotype has allele fraction 0 and an 0 / 1 genotype has expected allele fraction 0.5. Expected allele fractions in the mixture can be detected from 0 and 1 by an error rate which can be specific to each mutation (for example, stratifying mutations by-19- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665trinucleotide context), or an average across all variants. In the context of an MRD assay, one can use error rates stratified by complementary dinucleotide pairs which are estimated by analyzing reads at “control” sites in other regions of the genome.

[0084] In some embodiments, allele counts can be modeled as binomially distributed with n set to the total count of read pairs covering a variant position and p set to the expected allele frequency (as shown above). In some embodiments, allele counts can be modeled as a negative binomial distribution, gaussian distribution, or Poisson distribution. The total likelihood of the data is:L(X) = prod_i sum_g p(gl & g2) * p(alts_i|depth_i,E(af_g))where i indexes variants and g indexes genotype combinations, and alts i is the count of read pairs bearing an ALT allele at variant position i.

[0085] The value p(gl & g2) is the prior probability of drawing each genotype mixture from the Population given a variant’s minor allele frequency (i.e., the product of the expected Hardy-Weinberg proportions of each individual genotype in the population).

[0086] The likelihood can be maximized with respect to contamination rate using numerical optimization and a likelihood ratio test to compare the likelihood when c=0 to the maximum likelihood estimate (c_ml). A likelihood ratio test statistic can then be calculated as:LRT = -2 * ( log( L(X|c=0) ) - log( L(X|c_ml) )and approximate the probability of observing the data in an uncontaminated sample with a chi-squared distribution with one degree of freedom. Samples with a sufficiently low p value can be classified as contaminated and the maximum likelihood estimate of c can be interpreted as the fraction of reads in the sample originating from a contaminant.

[0087] The foregoing model and other models disclosed herein can be utilized for methods of detecting contamination in a sample, comprising: sequencing DNA from a sample from the subject; and detecting contamination in the sample based on (i) the presence of a sequencing read comprising an alt-allele at one or more of a plurality of germline variants from a panel of germline variants and (ii) an allele fraction of the alt-allele.-20- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665[0088| The subject may be a healthy subject, a subject with a disease, disorder, or condition, or a subject that is suspected of having a disease, disorder, or condition and is therefore seeking a diagnosis or prognosis. In some embodiments, the subject was previously diagnosed with cancer.

[0089] The disclosed methods of detecting contamination can be combined with any sequencing-based method or assays that rely on nucleic acid samples (e.g., DNA or RNA samples). For example, as discussed in further detail herein, the disclosed methods can be utilized in the context of an MRD assay to detect contamination. Accordingly, in some embodiments, the disclosed methods may further comprise detecting sequence reads comprising one or more tumor-specific somatic mutations in the sample. Such tumor-specific somatic mutations can be determined for the subject prior to sequencing the sample (e.g., as part of a patient-specific signature panel).

[0090] When the disclosed methods of detection are utilized in an MRD assay, the methods of detecting contamination in a sample may comprise: sequencing DNA from a germline sample from a subject, thereby obtaining a set of germline sequencing reads, and sequencing DNA from a tumor sample obtained from the subject, thereby obtaining a set of tumor sequencing reads; determining a panel of tumor-specific somatic mutations based on the set of germline sequencing reads and the set of tumor sequencing reads; sequencing cell-free DNA (cfDNA) from a sample from the subject, thereby obtaining a set of cfDNA sequencing reads; and detecting contamination in the sample based on (i) the presence of a sequencing read in the set of cfDNA sequencing reads comprising an alt-allele at one or more of a plurality of germline variants of a panel of germline variants and (ii) an allele fraction of the alt-allele. Some embodiments may further comprise detecting the presence of circulating tumor DNA (ctDNA) in the sample based on the presence of at least one sequencing read in the set of cfDNA sequencing reads comprising one or more of the panel of tumor-specific somatic mutations. Some embodiments may further comprise sequencing DNA from another sample from the subject at 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more successive time points. In some embodiments, sequencing DNA from another sample from the subject is repeated one or more times while the patient is in remission. In some embodiments, sequencing DNA from another sample from the subject is repeated one or more times while the patient is undergoing-21- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665treatment for the cancer. In some embodiments, sequencing DNA from another sample from the subject is repeated one or more times coinciding with or prior to surgery; following, during, or prior to administration of chemotherapy; following, during, or prior to radiation therapy; following, during, or prior to administration of an immunotherapy; following, during, or prior to administration of a cell therapy; or following, during, or prior to administration of a biologic therapy.

[0091] The plurality of germline variants can comprise at least one common germline variant (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more common germline variants). In some embodiments, the plurality of germline variants consists of common germline variants.

[0092] In general, the sample may be a biological sample, such as a tissue sample or a biological fluid, such as blood, plasma, serum, or urine.

[0093] Some embodiments of the disclosed methods may comprise enriching the DNA from the sample for DNA comprising one or more of the panel of germline variants prior to sequencing. In embodiments involving an MRD assay, the methods may also or alternatively comprise enriching the DNA for tumor-specific somatic mutations, such as a pre-determined panel (i.e., a “signature panel”) of tumor-specific somatic mutations that are unique to the subject. Enriching the DNA from the sample may comprise, for example, hybrid capture enrichment or PCR-based enrichment.

[0094] In some embodiments, the germline variants in the panel of germline variants have a minor allele frequency of about 50%.

[0095] Because the disclosed methods are able to detect contamination based on the allele fraction of specific germline variants, the disclosed methods of detecting contamination in the sample do not require a spike-in. Thus, embodiments of the disclosed methods do not comprise a molecular spike-in, but some embodiments can optionally include a molecular spike-in if desired.

[0096] Further, the disclosed methods can comprise error correction, such as methods of error correction utilizing duplex unique molecule identifiers (UMI), which are discussed in more detail herein.-22- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665

[0097] The disclosed methods may comprise detecting contamination by quantifying the sequence reads comprising an alt-allele at one or more of the germline variants and quantifying the total number of sequencing reads of each germline variant in the panel of germline variants, thereby obtaining an allele count comprising an alt-allele count and a total allele count. Such methods can further comprise modeling the allele count as a statistical distribution, optionally a binomial distribution, a negative binomial distribution, gaussian distribution, or Poisson distribution. In some embodiments, the statistical distribution can include a probability parameter determined by analysis of one or more reference sets of germline variants.

[0098] The disclosed methods may further comprise determining a probability that each alt-allele belongs to the subject or originated from contamination based on relative likelihoods or expected allele fraction. In such embodiments, the expected allele fraction may use a fixed parameter for the probability of observing contamination. For such embodiments, the probability may be determined by calculating a probability of sequencing read counts for a given alt-allele of the given germline variant times a probability of the given alt-allele from a pool with known alt-allele fraction, marginalized over all alleles of the given germline variant.

[0099] The disclosed methods may further comprise calculating a total likelihood for contamination for each germline variant in the set of germline variants.

[0100] The disclosed methods may further comprise calculating an allele fraction of genotypes of each germline variant in the set of germline variants. In some embodiments, calculating the allele fraction comprises fitting a model of alt allele counts and total allele counts across the panel of germline variants. In such embodiments, the model can be, for example, a binomial model, a negative binomial model, gaussian model, or Poisson model.

[0101] Other methods of detecting contamination based on the present or allele fraction of germline variants generally require that the germline variants being assessed are homozygous relative to their respective reference alleles. However, the presently disclosed methods are not so constrained and can utilize germline variants regardless of whether the variant is homozygous relative to its reference allele. Accordingly, in some embodiments the panel of -23- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665germline variants comprises at least one germline variant (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10, or more) that is not expected to be homozygous relative to its reference allele.

[0102] In some embodiments, sequencing DNA from the sample comprises whole genome sequencing, whole exome sequencing, targeted sequencing, or subtractive hybridization.

[0103] In some embodiments, wherein the subject has, had, or is suspected of having a cancer or tumor. The type of cancer or tumor can is not particularly limited and can include, but is not limited to, bladder cancer, breast cancer, colon / colorectal cancer, gynecologic cancers, head and neck cancers, hematological cancers, liver cancer, lung cancer, and skin cancer.

[0104] As noted throughout the disclosure, the disclosed methods of detecting contamination can optionally comprise a step of enriching a DNA sample for DNA fragments that comprise the germline variants from the pre-determined panel of germline variants. Accordingly, the present disclosure provides methods of preparing an enriched DNA fraction, comprising: obtaining a sample comprising cfDNA from a subject; and enriching from the cfDNA a DNA fraction comprising one or more of a plurality of germline variants from a panel of germline variants by contacting the cfDNA with a plurality of primer pairs or probes, wherein each primer pair or probe in the plurality is specific for a DNA fragments comprising a germline variant of the panel of germline variants. In some embodiments, the sample is a sample of blood, plasma, serum, or urine. In some embodiments, the subject was previously diagnosed with cancer. Some embodiments may further comprise enriching from the cfDNA a DNA fraction comprising one or more tumor-specific somatic mutations that are determined for the subject prior to enrichment. Enriching the DNA fraction may comprise hybrid capture enrichment or PCR-based enrichment. In some embodiments, the plurality of germline variants comprise at least one common germline variant. In some embodiments, the plurality of germline variants consists of common germline variants. Some embodiments may further comprise sequencing the DNA fraction, and optionally further comprise detecting contamination in the sample based on (i) the presence of a sequencing read comprising an alt-allele at one or more of a plurality of germline variants from a panel of germline variants and (ii) an allele fraction of the alt-allele.-24- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665[01O5| Methods of the present disclosure can also comprise extracting DNA from a sample. DNA can be extracted from a sample by (1) harvesting cells or tissue (e.g., a tumor biopsy); (2) lysing the cells; (3) inactivating DNAses; (4) capturing DNA; and (5) separating the DNA from at least some of the components with which it was associated when initially produced (e.g., RNA, proteins, other cellular components), whether in nature and / or in an experimental setting; and (6) resuspending or eluting the DNA.

[0106] Harvesting tissues or cells may be completed by, for example, blood draw, needle biopsy, aspiration, scraping, or excision using surgical equipment. Subsequently, the cells may be lysed using lysis buffer (e.g., comprising a chaotropic agent) and / or by mechanical disruption. In step (3), DNAses may be inactivated by, for example, heat treatment (e.g., 5 minutes at 75°C) and / or by the use of DNAse inhibitors. Following lysis and DNAse inactivation, DNA can be captured by binding to a surface, such as a silica surface.Separating DNA from at least some of the components with which it was associated when initially produced can be completed by, for example, degrading RNA with RNAse and degrading protein with proteinase K. Alternative or additional methods include dissolving the sample in buffers containing certain salts (e.g., guanidinium salts) to remove proteins and / or washing away components with which the DNA was associated when initially produced while the DNA is bound to a solid-support (e.g., silica beads). Finally, the extracted DNA can be eluted or resuspended in water or a buffer (e.g., a buffer suitable for use in downstream applications and / or analyses).

[0107] In some embodiments, the extracted DNA is quantified and / or the quality of the DNA is assessed prior to sequencing (e.g., low-coverage sequencing) DNA (e.g., DNA from a first, second, and / or subsequent sample). Spectroscopic and / or electrophoretic methods can be used to quantify DNA or to assess its quality.

[0108] Methods of the present disclosure can comprise amplifying the DNA, for example, amplifying the DNA of a panel of germline variants. DNA can be amplified by, for example, polymerase chain reaction (PCR), such as real-time PCR (e.g., TaqMan) and quantitative PCR (qPCR).-25- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665[01O9| DNA for use in accordance with technologies described herein can be sequenced from a sample (e.g., a first sample, a second sample, a subsequent (“further”) sample) by whole genome sequencing, whole exome sequencing, subtractive hybridization, or targeted sequencing. In some embodiments, methods of the present disclosure further comprise detecting in the second sample sequence reads comprising one or more tumor-specific somatic mutations that are determined for the subject prior to sequencing the second sample. In some embodiments, the methods further comprise enriching the DNA from the second sample for DNA comprising one or more of the tumor-specific somatic mutations prior to sequencing. Enriching the DNA comprising one or more of the tumor-specific somatic mutations can comprise the use of, for example, hybrid capture enrichment or PCR-based enrichment.

[0110] Methods of the present disclosure can further comprise enriching the DNA from a sample (e.g., a blood, plasma, or serum sample) for DNA comprising one or more of the panel of germline variants prior to sequencing. DNA enrichment can comprise the use of hybrid capture enrichment or PCR-based enrichment. In some embodiments, methods of the present disclosure further comprise enriching the DNA from a further (“subsequent”) sample for DNA comprising one or more of the panel of germline variants prior to sequencing. DNA enrichment can comprise the use of hybrid capture enrichment or PCR-based enrichment. Thus, in some embodiments, enriching the DNA from a sample of DNA comprising one or more of the panel of germline variants comprises hybrid capture enrichment or PCR-based enrichment.Subjects and Samples

[0111] The present disclosure provides, among other things, methods for detecting contamination in a sample. The sample for use in accordance with the methods described herein can be a biological sample, such as a nucleic acid sample or, more specifically, a DNA sample or an RNA sample. In some embodiments, the sample is a sample of blood (e.g., whole blood, a blood fraction), plasma, serum, urine, or tumor. In some embodiments, the sample comprises cell free DNA (cfDNA), circulating tumor DNA (ctDNA), or both cfDNA and ctDNA.-26- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665[01121 In some embodiments, the sample is obtained from a subject that has been previously diagnosed with cancer. In such embodiments, the disclosed methods of detecting contamination may be implemented in diagnostic or prognostic assays to detect or track the cancer, such as an MRD assay. The cancer can be selected from, for example, bladder cancer, breast cancer, colon / colorectal cancer, gynecologic cancers, head and neck cancers, hematological cancers, liver cancer, lung cancer, and skin cancer.

[0113] In some embodiments, the subject has completed at least one cancer treatment. The cancer treatment can be selected from, for example, chemotherapy, radiotherapy, surgery, immunotherapy, cell therapy, or biologic therapy. In some embodiments, the subject has completed at least one cancer treatment prior to obtaining the sample.

[0014] In some embodiments, methods of the present disclosure further comprise repeating sequencing DNA from a further (“subsequent”) sample from the subject at 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more successive time points. In some embodiments, sequencing DNA from a further sample from the subject is repeated one or more times while the subject is in remission (e.g., from cancer).

[0115] In some embodiments, sequencing DNA from a further (“subsequent”) sample from the subject is repeated one or more times while the subject is undergoing treatment for cancer. In some embodiments, sequencing DNA from a further sample from the subject is repeated one or more times coinciding with or prior to surgery; following, during, or prior to administration of chemotherapy; following, during, or prior to radiation therapy; following, during, or prior to administration of an immunotherapy; following, during, or prior to administration of a cell therapy; or following, during, or prior to administration of a biologic therapy.

[0116] In some embodiments, the subject is a healthy subject (i.e., a subject not previously diagnosed with cancer or another disease / condition). In some embodiments, the subject is a subject having or suspected of having a disease, disorder, and / or condition. In some embodiments, the subject is not a pregnant female.-27- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665Panels of Germline Variants[0117| A common germline variant in accordance with the technologies of the present disclosure is a variant in a germline e.g., reproductive) cell that can be passed on to offspring and has a minor allele frequency of about 40-60%. As disclosed herein, the methods provide may utilize a panel of germline variants comprising or consisting of a plurality of common germline variants. However, as noted above, other germline variants with less common allele frequencies can also be utilized. In some embodiments, the panel of germline variants may comprise a germline variant with a minor allele frequency of about 25%. In some embodiments, the panel of germline variants may comprise a germline variant with a minor allele frequency of about 30%. In some embodiments, the panel of germline variants may comprise a germline variant with a minor allele frequency of about 35%. In some embodiments, the panel of germline variants may comprise a germline variant with a minor allele frequency of about 40%. In some embodiments, the panel of germline variants may comprise a germline variant with a minor allele frequency of about 45%. In some embodiments, the panel of germline variants may comprise a germline variant with a minor allele frequency of about 50%. In some embodiments, the panel of germline variants may comprise a germline variant with a minor allele frequency of about 55%. In some embodiments, the panel of germline variants may comprise a germline variant with a minor allele frequency of about 60%. In some embodiments, the panel of germline variants may comprise a germline variant with a minor allele frequency of about 65%. In some embodiments, the panel of germline variants may comprise a germline variant with a minor allele frequency of about 70%. In some embodiments, the panel of germline variants may comprise a germline variant with a minor allele frequency of about 75%.

[0118] Panels of germline variants comprise a plurality of germline variants (e.g., common germline variants). While evaluation of a single common germline variant may allow for discrimination of whether two or more samples are from the same or different subjects, addition of each evaluated common germline variant or germline variant to a panel of germline variants increases the likelihood that the combination of particular germline variants genotypes from the panel of germline variants are a unique combination associated with a particular subject. Accordingly, such panels of germline variants can be utilized in methods-28- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665of the present disclosure to determine the source of a sample, e.g., to determine if (i) a first sample and a second sample are from the same subject; or (ii) a first sample and a second sample are from different subjects.10119] In some embodiments, panels of germline variants for use in accordance with the technologies described herein are not sample-informed or subject-specific (e.g., the selection of variants for the panel is not dependent on or unique to a particular sample or subject). Rather, the disclosed methods may utilize the same panel of germline variants for all sample evaluation, which is possible, at least in part, due to their minor allele frequency across a plurality of subjects.Use in Minimal Residual Disease Detection

[0120] Technologies of the present disclosure can be particularly useful in determining the source of samples obtained over a period of time e.g., as from the same or different subjects, e.g., for disease monitoring), including in detecting Minimal Residual Disease (MRD). The goal of a MRD assay is to detect and / or quantify circulating tumor DNA (ctDNA) so researchers and clinicians can detect recurrence early and monitor the progress of the disease (e.g., cancer) through treatment. In general, a MRD assay will rely on a patient-specific and tumor-specific panel (i.e., a “signature panel” or a “panel of patient-specific somatic variants” or a “panel of tumor-specific somatic variants”) for assessing the presence of ctDNA in a patient (“subject”) sample. The signature panel can be prepared with the general steps of (1) profiling a tumor or cancer sample from a patient, and (2) identifying a subset of somatic mutations to target, and, at one or more later time points, (3) taking a subsequent sample from the patient, (4) enriching cell-free DNA (cfDNA) for the target somatic mutation sites, and (5) detecting, determining, or estimating the ctDNA content of cell free DNA (cfDNA) given the tumor profile and sequencing data.

[0121] More specifically, preparing the patient-specific and tumor-specific panel (i.e., a “signature panel”) may comprise, for example, (a) obtaining a tumor sample and, optionally, a non-tumor sample from a cancer patient; (b) sequencing DNA (e.g., genomic DNA) from the tumor sample and, optionally, sequencing DNA (e.g., cell free DNA or “cfDNA”) from the non-tumor sample; and (c) determining tumor-specific somatic mutations to prepare a -29- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665patient-specific “signature panel” (e.g., by comparing the sequences of the tumor sample and the non-tumor sample to determine any tumor-specific somatic mutations that are present in the sequences of DNA from the tumor sample but not present in the sequences of DNA from the non-tumor sample). Sequencing of the DNA from the tumor sample and non-tumor sample may comprise whole genome sequencing or various types of targeted sequencing, such as whole exome sequencing or subtractive hybridization.

[0122] A comparison of the tumor and non-tumor sequences can be performed by, for example, aligning the sequences of DNA (e.g., genomic DNA) from the tumor sample to a reference human genome that is not from the patient and aligning the sequences of DNA (e.g., cfDNA) from the non-tumor sample to the reference genome that is not from the patient. The reference genome can be, for example, a publicly available human genome assembly, such as hgl8, hgl9, GRCh38.pl4, GRCh37.pl3, or other assemblies from the Genome Reference Consortium. Alternatively, the comparison of the tumor and non-tumor sequences can be performed by, for example, aligning the sequences of DNA (e.g., genomic DNA) from the tumor sample to sequences of DNA (e.g., cfDNA) from the non-tumor sample. With either approach, the skilled artisan is able to detect and identify tumor-specific somatic mutations that are present in the tumor sample but not in the non-tumor sample. In some embodiments, determining tumor-specific somatic mutations to prepare a patientspecific “signature panel” does not comprise genotyping or preparing a genotype for the tumor sample and / or the non-tumor sample.[0123 { In some embodiments, preparing the tumor-specific panel (i.e., a “signature panel”) may comprise, for example (a) obtaining a tumor sample from a cancer patient; (b) sequencing DNA (e.g., genomic DNA) from the tumor sample, thereby obtaining sequences of DNA or sequence reads from the tumor sample; and (c) comparing the sequences of the tumor sample to one or more reference genomes or non-tumor samples from the subject to determine any tumor-specific somatic mutations that are present in the sequences of DNA from the tumor sample but not present in the sequences of DNA from the one or more reference genomes or non-tumor samples. This comparison may be performed by, for example, aligning the sequences of DNA (e.g., genomic DNA) from the tumor sample to the sequences of the one or more reference genomes. Again, the reference genome can be, for-30- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665example, a publicly available human genome assembly, such as hgl8, hgl9, GRCh38.pl 4, GRCh37.pl3, or other assemblies from the Genome Reference Consortium. Additionally or alternatively, genomic sequences from a non-tumor sample from the same patient may also be used as a reference genome or in conjunction with another reference genome (e.g., hgl8, hgl9, GRCh38.pl4, GRCh37.pl3, etc.) to determine tumor-specific somatic variants. Such alignment, again, allows the skilled artisan to detect and identify tumor-specific somatic mutations that are present in the tumor sample. Sequencing of the DNA from the tumor sample may comprise whole genome sequencing or various types of targeted sequencing, such as whole exome sequencing or subtractive hybridization.

[0124] In some embodiments, rather than compare the sequences of the tumor sample to one or more reference genomes, mathematical algorithms and / or artificial intelligence are utilized to determine the likelihood of any identified potential somatic mutations as being a tumorspecific somatic mutation that is present in the sequences of DNA from the tumor sample but not present in the sequences of DNA from one or more reference genomes and / or non-tumor samples. In some embodiments, such mathematical algorithms and / or machine learning are utilized in combination with comparing the sequences of the tumor sample to one or more reference genomes and / or non-tumor samples.

[0125] The tumor sample may be a solid tumor sample, such as a biopsy or other tissue sample, or a liquid sample, such as blood (in the case of a hematological cancer) or specific fractions of blood. The non-tumor sample may be tissue-matched with the tumor sample or it may be from a different tissue. For example, the non-tumor sample may be selected from a healthy (i.e., non-cancerous or non-tumor) tissue sample, blood or specific fractions of blood such as buffy coat, leukocytes, fibroblast, or any other biological sample comprising cfDNA or genomic DNA.

[0126] Once a patient-specific and tumor-specific panel (i.e., a “signature panel”) has been established, such a signature panel can be used to enrich ctDNA (e.g., fragments that include a target sequence corresponding to a tumor-specific somatic mutation or variant) in subsequent samples taken from the cancer patient. The subsequent (or “further”) samples may be taken from a patient at various time points during the course of treatment or during a-31- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665period of remission. For example, after a surgical removal of a tumor, the tumor may be profiled as described herein to determine tumor-specific somatic mutations, and at one or more subsequent time points a subsequent sample may be taken from the subject to search for the presence of any ctDNA comprising any one of the identified tumor-specific somatic mutations. The detection or presence of ctDNA comprising a tumor-specific somatic mutation may be indicative of cancer recurrence. Additionally or alternatively, similar assessment can be performed throughout the course of a patient’s treatment (e.g., with chemotherapy, radiation, immunotherapy, cell therapy, etc.) to detect or quantify ctDNA and determine whether the amount of ctDNA is increasing or decreasing, as this may be indicative of responsiveness to the therapy. Accordingly, assessment of a subsequent (or “further”) sample may be repeated 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more times throughout the course of a patient’s remission or treatment. The assessment of a subsequent sample may be repeated monthly, every other month, once every three months, once every four months, once every five months, once every six months, once every seven months, once every eight months, once every nine months, once every ten months, once every eleven months, or annually. Methods of determining the source of a sample, as described herein, can be utilized to determine if a first sample and a second and / or subsequent sample assessed in a MRD assay are from the same subject or from different subjects.

[0127] The type of sample used for the one or more subsequence samples is generally a blood sample, a plasma sample, or a serum sample, but any biological sample that contains cfDNA and potential contains ctDNA would be acceptable. In some embodiments, the one or more subsequent samples are cell-free samples.10128] Enrichment of ctDNA (e.g., fragments that include a target sequence corresponding to a tumor-specific somatic mutation or variant) in the one or more subsequent samples can be performed by methods including, but not limited to, hybrid capture-based enrichment, PCR-target enrichment, or on-sequencer enrichment. Briefly, enrichment may comprise extracting cfDNA from a subsequent sample taken from the cancer patient and contacting the extracted cfDNA with a plurality of oligonucleotides (i.e., oligonucleotide probes), wherein each oligonucleotide in the plurality of oligonucleotides comprises a nucleic acid sequence that is capable of hybridizing to a cfDNA fragment comprising one of the tumor-specific somatic-32- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665mutation sequences identified by comparing the sequences of the patients tumor DNA and non-tumor DNA. Thus, enrichment may utilize a set of oligonucleotide probes to selectively enrich ctDNA that may be in the subsequent sample by binding to previously identified tumor-specific somatic mutation sequences.

[0129] A signature panel may comprise 10-5000 tumor-specific somatic mutations. For example, a signature panel may comprise 10-4000, 10-3000, 10-2500, 10-2000, 10-1500, 10-1000, 10-950, 10-900, 10-850, 10-800, 10-750, 10-700, 10-650, 10-600, 10-550, 10-500, 50-5000, 50-4000, 50-3000, 50-2500, 50-2000, 50-1500, 50-1000, 50-950, 50-900, 50-850, 50-800, 50-750, 50-700, 50-650, 50-600, 50-550, 50-500, 100-5000, 100-4000, 100-3000, 100-2500, 100-2000, 100-1500, 100-1000, 100-950, 100-900, 100-850, 100-800, 100-750, 100-700, 100-650, 100-600, 100-550, 100-500, 200-5000, 200-4000, 200-3000, 200-2500, 200-2000, 200-1500, 200-1000, 200-950, 200-900, 200-850, 200-800, 200-750, 200-700, 200-650, 200-600, 200-550, 200-500, 300-5000, 300-4000, 300-3000, 300-2500, 300-2000, 300-1500, 300-1000, 300-950, 300-900, 300-850, 300-800, 300-750, 300-700, 300-650, 300-600, 300-550, 300-500, 400-5000, 400-4000, 400-3000, 400-2500, 400-2000, 400-1500, 400-1000, 400-950, 400-900, 400-850, 400-800, 400-750, 400-700, 400-650, 400-600, 400-550, 400-500, 500-5000, 500-4000, 500-3000, 500-2500, 500-2000, 500-1500, 500-1000, SOO-OSO, 500-900, 500-850, 500-800, 500-750, 500-700, 500-650, 500-600, or 500-550 tumorspecific somatic mutations. In some embodiments, a signature panel may comprise or consist of about 10, about 20, about 30, about 40, about 50, about 75, about 100, about 150, about 200, about 250, about 300, about 350, about 400, about 450, about 500, about 550, about 600, about 650, about 700, about 750, about 800, about 850, about 900, about 950, about 1000, about 1100, about 1150, about 1200, about 1250, about 1300, about 1350, about 1400, about 1450, about 1500, about 1550, about 1600, about 1650, about 1700, about 1750, about 1800, about 1850, about 1900, about 1950, or about 2000 or more tumor-specific somatic mutations. In some embodiments, a signature panel may comprise at least 10, at least 20, at least 30, at least 40, at least 50, at least 75, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1150, at least 1200, at least 1250, at least 1300, at least 1350, at least 1400, at least 1450, at least 1500, at least 1550, at least 1600, at least 1650, at least -33- 4923-7836-0082.2Atty. Dkt. No.: 131588-16651700, at least 1750, at least 1800, at least 1850, at least 1900, at least 1950, or at least 2000 tumor-specific somatic mutations. The tumor-specific somatic mutations may be in introns, exons, intergenic regions, or a combination thereof. In some embodiments, the tumor-specific somatic mutations may be one or more somatic mutations selected from Single Nucleotide Variants, insertions, deletions, and translocations.[0130| After enrichment or concurrently with enrichment of ctDNA (e.g., fragments that include a target sequence corresponding to a tumor-specific somatic mutation or variant), the enriched DNA is sequenced. This sequencing may be performed by, for example Next Generation Sequencing (NGS). Deep sequencing may allow for more sensitive detection, and so the depth of the sequencing may be at least 50X, at least 100X, at least 150X, at least 200X, at least 250X, at least 300X, at least 350X, at least 400X, at least 450X, at least 500X, at least 550X, at least 600X, at least 650X, at least 700X, at least 750X, at least 800X, at least 850X, at least 900X, at least 950X, or at least 1000X. In other words, the depth of the sequencing may be about 50X, about 100X, about 150X, about 200X, about 250X, about 300X, about 350X, about 400X, about 450X, about 500X, about 550X, about 600X, about 650X, about 700X, about 750X, about 800X, about 850X, about 900X, about 950X, or about 1000X. The detection sensitivity of the disclosed MRD methods may be about 20 to about 50 ctDNA fragments comprising one or more of the set of somatic mutations in the fluid sample per a total background of about 500,000 cfDNA fragments.

[0131] MRD methods may be used for tracking and assessing recurrence in a cancer patient. For example, the cancer patient may have a cancer selected from, but not limited to, bladder cancer, breast cancer, colon / colorectal cancer, gynecologic cancers, head and neck cancers, hematological cancers, liver cancer, lung cancer, and skin cancer.10132] In MRD assays, the obtaining and testing of subsequent samples from a cancer patient, may be repeated one or more times following completion of a cancer treatment; one or more times while the cancer patient is in remission; one or more times coinciding with or prior to surgery; following, during, or prior to administration of chemotherapy; following, during, or prior to radiation therapy; following, during, or prior to immunotherapy; or following, during, or prior to cell therapy. MRD assays may also be repeated at times prior to,-34- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665coinciding with, and / or following an imaging test, such as a PET scan, a PET / CT scan, an MRI, or an X-ray. Subsequent (or “further”) samples may be evaluated relative to a first and / or second sample using methods of determining the source of a sample described herein, e.g., to determine if a first, second, and / or subsequent samples are from the same subject or are from different subjects.[0133| MRD methods also allow for detecting ctDNA or determining the tumor fraction from a biological sample from a patient that has, previously had, or is suspected of having cancer.EXAMPLESExample 1 : Method for Detecting Contamination in a DNA Sample[0134| The following workflow is used to detect contamination in a DNA sample. The allele count or allele fraction of germline variant is a sample is modeled as a binomial distribution as follows:The data (X) is the count of ALT-bearing read pairs and total read pairs at each common germline variant position. For an unmixed sample there should be 3 genotypes present (0 / 0, 0 / 1, 1 / 1). In a mixed sample of 2 individuals there are 9 genotype combinations: 0 / 0:0 / 0, 0 / 0:0 / l, 0 / 0: 1 / 1, 0 / E0 / 0, 0 / E0 / 1, 0 / 1: 1 / 1, 1 / E0 / 0, 1 / E0 / 1, 1 / 1: 1 / 1, wherein the genotypes of sample 1 and sample 2 are written in VCF- style notation with a colon separating each sample’s genotype.

[0135] The expected allele fraction of a variant in each genotype class is the average of each genotype’s frequency weighted by the contamination rate c:E(af) = (1-c) * af_l + c * af_2af_l and af_2 above are the expected allele fractions in each individual given the genotype.

[0136] A 0 / 0 genotype has allele fraction 0 and an 0 / 1 genotype has expected allele fraction 0.5. Expected allele fractions in the mixture can be detected from 0 and 1 by an error rate which can be specific to each mutation (for example, stratifying mutations by trinucleotide context), or an average across all variants. Optionally, error rates can be stratified by-35- 4923-7836-0082.2Atty. Dkt. No.: 131588-1665complementary dinucleotide pairs which are estimated by analyzing reads at “control” sites in other regions of the genome.

[0137] Allele counts can be modeled as binomially distributed with n set to the total count of read pairs covering a variant position and p set to the expected allele frequency (as shown above). The total likelihood of the data is:L(X) = prod_i sum_g p(gl & g2) * p(alts_i|depth_i,E(af_g))where i indexes variants and g indexes genotype combinations, and alts i is the count of read pairs bearing an ALT allele at variant position i.

[0138] The value p(gl & g2) is the prior probability of drawing each genotype mixture from the Population given a variant’s minor allele frequency (i.e., the product of the expected Hardy-Weinberg proportions of each individual genotype in the population).

[0139] The likelihood is maximized with respect to contamination rate using numerical optimization and a likelihood ratio test to compare the likelihood when c=0 to the maximum likelihood estimate (c_ml). A likelihood ratio test statistic is then calculated as:LRT = -2 * ( log( L(X|c=0) ) - log( L(X|c_ml) )and approximates the probability of observing the data in an uncontaminated sample with a chi-squared distribution with one degree of freedom. Samples with a sufficiently low p value are classified as contaminated and the maximum likelihood estimate of c is interpreted as the fraction of reads in the sample originating from a contaminant.

[0140] The foregoing workflow is used to detect contamination based on (i) the presence of a sequencing read in the set of cfDNA sequencing reads comprising an alt-allele at one or more of a plurality of germline variants of a panel of germline variants and (ii) an allele fraction of the alt-allele. It should be recognized that this workflow is non-limiting and other models or workflows my also be used to detect contamination as disclosed herein.-36- 4923-7836-0082.2

Claims

Atty. Dkt. No.: 131588-1665WHAT IS CLAIMED IS:

1. A method of detecting contamination in a sample, comprising:sequencing DNAfrom a sample from the subject; anddetecting contamination in the sample based on (i) the presence of a sequencing read comprising an alt-allele at one or more of a plurality of germline variants from a panel of germline variants and (ii) an allele fraction of the alt-allele.

2. The method of claim 1, wherein the plurality of germline variants comprise at least one common germline variant.

3. The method of claim 1 or 2, wherein the plurality of germline variants consists of common germline variants.

4. The method of any one of claims 1-3, wherein the sample is a sample of blood, plasma, serum, or urine.

5. The method of any one of claims 1-4, wherein the subject was previously diagnosed with cancer.

6. The method of any one of claims 1-5, further comprising detecting in the sample sequence reads comprising one or more tumor-specific somatic mutations that are determined for the subject prior to sequencing the sample.

7. The method of any one of claims 1-6, further comprising enriching the DNA from the sample for DNA comprising one or more germline variants of the panel of germline variants prior to sequencing.

8. The method of claim 7, wherein enriching the DNA from the sample comprises hybrid capture enrichment or PCR-based enrichment.

9. The method of any one of claims 1-8, wherein detecting contamination does not comprise a spike-in.-37- 4923-7836-0082.2Atty. Dkt. No.: 131588-166510. The method of any one of claims 1-9, wherein sequencing DNA from the sample from the subject further comprises error correction, optionally with duplex unique molecule identifiers (UMI).

11. A method of detecting contamination in a sample, comprising:sequencing DNA from a germline sample from a subject, thereby obtaining a set of germline sequencing reads, and sequencing DNA from a tumor sample obtained from the subject, thereby obtaining a set of tumor sequencing reads;determining a panel of tumor-specific somatic mutations based on the set of germline sequencing reads and the set of tumor sequencing reads;sequencing cell-free DNA (cfDNA) from a sample from the subject, thereby obtaining a set of cfDNA sequencing reads; anddetecting contamination in the sample based on (i) the presence of a sequencing read in the set of cfDNA sequencing reads comprising an alt-allele at one or more of a plurality of germline variants of a panel of germline variants and (ii) an allele fraction of the alt-allele.

12. The method of claim 11, wherein the plurality of germline variants comprise at least one common germline variant.

13. The method of claim 11 or 12, wherein the plurality of germline variants consists of common germline variants.

14. The method of any one of claims 11-13, wherein the sample is a sample of blood, plasma, serum, or urine.

15. The method of any one of claims 11-14, further comprising enriching the DNAfrom the sample for DNA comprising one or more of the panel of common germline variants prior to sequencing.

16. The method of claim 15, wherein enriching the DNA from the sample comprises hybrid capture enrichment or PCR-based enrichment.-38- 4923-7836-0082.2Atty. Dkt. No.: 131588-166517. The method of any one of claims 11-16, wherein detecting contamination in the sample does not comprise a spike-in.

18. The method of any one of claims 11-17, wherein sequencing cfDNAfrom the sample from the subject further comprises error correction with duplex unique molecule identifiers (UMI).

19. The method of any one of claims 11-18, further comprising detecting the presence of circulating tumor DNA (ctDNA) in the sample based on the present of sequencing read in the set of cfDNA sequencing reads comprising one or more of the panel of tumor-specific somatic mutations.

20. The method of any one of claims 1-19, wherein the germline variants in the panel of germline variants have a minor allele frequency of about 50%.

21. The method of any one of claims 1-20, wherein detecting contamination comprises quantifying the sequence reads comprising an alt-allele at one or more of the germline variants and quantifying the total number of sequencing reads of each germline variant in the panel of germline variants, thereby obtaining an allele count comprising an alt-allele count and a total allele count.

22. The method of claim 21, further comprising modeling the allele count as a statistical distribution, optionally a binomial distribution, a negative binomial distribution, gaussian distribution, or Poisson distribution.

23. The method of claim 22, wherein the statistical distribution includes a probability parameter determined by analysis of one or more reference sets of germline variants.

24. The method of any one of claims 1-23, further comprising determining a probability that each alt-allele belongs to the subject or originated from contamination based on relative likelihoods or expected allele fraction.

25. The method of claim 24, wherein the expected allele fraction uses a fixed parameter for the probability of observing contamination.-39- 4923-7836-0082.2Atty. Dkt. No.: 131588-166526. The method of any one of claims 1-25, further comprising calculating a total likelihood for contamination for each germline variant in the set of germline variants.

27. The method of any one of claims 1-26, further comprising calculating an allele fraction of genotypes of each germline variant in the set of germline variants.

28. The method of claim 27, wherein calculating the allele fraction comprises fitting a model of alt allele counts and total allele counts across the panel of germline variants.

29. The method of claim 28, wherein the model is a binomial model, a negative binomial model, gaussian model, or Poisson model30. The method of any one of claims 1-29, wherein the panel of germline variants comprises at least one germline variant that is not expected to be homozygous relative to its reference allele.

31. The method of any one of claims 1-30, wherein sequencing DNA from the sample comprises whole genome sequencing, whole exome sequencing, targeted sequencing, or subtractive hybridization.

32. The method of any one of claims 11-31, wherein sequencing DNA from the germline sample and / or sequencing DNA from the tumor sample comprises whole genome sequencing, whole exome sequencing, targeted sequencing, or subtractive hybridization.

33. The method of any one of claims 11-32, wherein the subject has completed at least one cancer treatment prior to obtaining the tumor sample.

34. The method of claim 33, wherein the cancer treatment is selected from chemotherapy, radiotherapy, surgery, immunotherapy, cell therapy, or biologic therapy.

35. The method of any one of claims 1-34, further comprising sequencing DNA from another sample from the subject at 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more successive time points.-40- 4923-7836-0082.2Atty. Dkt. No.: 131588-166536. The method of claim 35, wherein sequencing DNAfrom another sample from the subject is repeated one or more times while the patient is in remission.

37. The method of claim 35, wherein sequencing DNAfrom another sample from the subject is repeated one or more times while the patient is undergoing treatment for the cancer.

38. The method of claim 35, wherein sequencing DNAfrom another sample from the subject is repeated one or more times coinciding with or prior to surgery; following, during, or prior to administration of chemotherapy; following, during, or prior to radiation therapy; following, during, or prior to administration of an immunotherapy; following, during, or prior to administration of a cell therapy; or following, during, or prior to administration of a biologic therapy.

39. The method of any one of claims 1-38, wherein the subject has, had, or is suspected of having a cancer selected from bladder cancer, breast cancer, colon / colorectal cancer, gynecologic cancers, head and neck cancers, hematological cancers, liver cancer, lung cancer, and skin cancer.

40. A method of preparing an enriched DNA fraction, comprising:obtaining a sample comprising cfDNAfrom a subject; andenriching from the cfDNA a DNA fraction comprising one or more of a plurality of germline variants from a panel of germline variants by contacting the cfDNA with a plurality of primer pairs or probes, wherein each primer pair or probe in the plurality is specific for a DNA fragments comprising a germline variant of the panel of germline variants.

41. The method of claim 40, wherein the sample is a sample of blood, plasma, serum, or urine.

42. The method of claim 40 or 41, wherein the subject was previously diagnosed with cancer.-41- 4923-7836-0082.2Atty. Dkt. No.: 131588-166543. The method of claim 42, further comprising enriching from the cfDNA a DNA fraction comprising one or more tumor-specific somatic mutations that are determined for the subject prior to enrichment.

44. The method of any one of claims 40-43, wherein enriching the DNA fraction comprises hybrid capture enrichment or PCR-based enrichment.

45. The method of any one of claims 40-44, wherein the plurality of germline variants comprise at least one common germline variant.

46. The method of any one of claims 40-45, wherein the plurality of germline variants consists of common germline variants.

47. The method of any one of claims 40-46, further comprising sequencing the DNA fraction.

48. The method of claim 47, further comprising detecting contamination in the sample based on (i) the presence of a sequencing read comprising an alt-allele at one or more of a plurality of germline variants from a panel of germline variants and (ii) an allele fraction of the alt-allele.-42- 4923-7836-0082.2