Method for identifying DNA samples from the same individual based on low-coverage sequencing data
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-02-09
- Publication Date
- 2026-08-13
Abstract
Description
Atty. Dkt. No.: 131588-1666METHOD FOR IDENTIFYING DNA SAMPLES FROM THE SAME INDIVIDUAL BASED ON LOW-COVERAGE SEQUENCING DATA CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of and priority to U.S. Provisional Application No.63 / 756,695, filed on February 10, 2025, the entire disclosure of which is incorporated by reference herein.BACKGROUND[00021 The following description of the background of the present technology is provided simply as an aid in understanding the present technology and is not admitted to describe or constitute prior art to the present technology.
[0003] Methods of determining the source of a sample, such as a nucleic acid sample, are useful in, for example, biomedical and forensic applications. In some cases, such methods determine two or more samples (e.g., nucleic acid (e.g., DNA) samples) as being from the same subject by comparing nucleic acid sequencing information, also referred to as “genetic relatedness” or “genetic identity.” These methods generally require the use of genotype calls at corresponding sites for comparison, such as Single Nucleotide Polymorphisms (SNPs) or Short Tandem Repeats (STRs). The data for genotype calls are typically produced by high-coverage sequencing methods. However, there are still challenges in identifying whether two or more samples as being from the same or different subjects remain, particularly if the samples do not include enough nucleic acid and / or the nucleic acid is of insufficient quality.SUMMARY OF THE INVENTION
[0004] The present disclosure provides, among other things, methods of determining the source of a sample. In particular, the disclosed methods facilitate determining (i) a first sample and a second sample are from the same subject; or (i) a first sample and a second sample are from different subjects. The methods described herein are particularly useful for determining the source of the sample, wherein both a first sample and a second sample (and / or a “further” / “ another” sample) are sequenced at low-coverage. The ability to utilize -1- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666low-coverage sequencing data can (i) reduce sequencing costs; (ii) reduce sequencing time; and / or (iii) allow for samples with a lower amount of nucleic acid and / or nucleic acid of insufficient quality for high-coverage sequencing to be utilized. Thus, the methods of the present disclosure improve current methods of determining the source of a sample, which require high-coverage sequencing data from at least one of the evaluated samples.
[0005] In one aspect, the present disclosure provides a method of determining the source of a sample, comprising: sequencing at low coverage DNA from a first sample from a subject, thereby obtaining a first set of sequencing reads; counting heterozygous and homozygous-reference-bearing sequencing reads at germline variants in the first set of sequencing reads that corresponds to a panel of germline variants comprising a plurality of common germline variants; sequencing at low coverage DNA from a second sample, thereby obtaining a second set of sequencing reads; counting heterozygous and homozygous-reference-bearing sequencing reads at germline variants from the panel of common germline variants sequenced in the second set of sequencing reads; determining (i) the first sample and the second sample are from the same subject or (ii) the first sample and the second sample are from different subjects based on a likelihood of observing the counts of sequencing reads for the panel of common germline variants for each sample.
[0006] In some embodiments, the likelihood of observing the counts of sequencing reads for the panel of germline variants is determined by calculating a genotype likelihood selected from (i) homozygous-reference, (ii) heterozygous, and (iii) homozygous alternate for each of the germline variants in the first set of sequencing reads and the second set of sequencing reads.
[0007] In some embodiments, calculating the genotype likelihood comprises modeling a statistical distribution, optionally a binomial distribution.
[0008] In some embodiments, a genotype likelihood of a given germline variant in the first sample or the second sample, independently, is a probability of counts of sequencing reads for a given genotype of the given germline variant times a probability of the given genotype from a pool with known alt-allele fraction, marginalized over all genotypes of the given germline variant.-2- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666[0OO9| In some embodiments, calculating the genotype likelihood comprises calculating a mean of a likelihood of a given germline variant in the first sample and a likelihood of the given germline variant in the second sample.
[0010] In some embodiments, the first sample and the second sample are from the same subject when the likelihood of observing the counts of sequencing reads for the panel of germline variants for each sample is above a predetermined threshold. In some embodiments, the first sample and the second sample are from different subjects when the likelihood of observing the counts of sequencing reads for the panel of germline variants for each sample is below a predetermined threshold.
[0011] In some embodiments, the first sample is a sample of blood, plasma, serum, or urine. In some embodiments, the first sample is a tumor sample. In some embodiments, the second sample is a sample of blood, plasma, serum, or urine.[0012 [ In some embodiments, the subject was previously diagnosed with cancer.
[0013] In some embodiments, methods of the present disclosure further comprise detecting in the second sample sequence reads comprising one or more tumor-specific somatic mutations that are determined for the subject prior to sequencing the second sample. In some embodiments, such methods further comprise enriching the DNA from the second sample for DNA comprising one or more of the tumor-specific somatic mutations prior to sequencing. In some embodiments, enriching the DNA comprising one or more of the tumor-specific somatic mutations comprises hybrid capture enrichment or PCR-based enrichment.
[0014] In some embodiments, methods of the present disclosure further comprise enriching the DNA from the second sample for DNA comprising one or more of the panel of germline variants prior to sequencing. In some embodiments, enriching the DNA from the second sample for DNA comprising one or more of the panel of germline variants comprises hybrid capture enrichment or PCR-based enrichment.[0015[ In some embodiments, sequencing DNA from the first sample comprises whole genome sequencing, whole exome sequencing, or targeted sequencing. In some-3- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666embodiments, sequencing DNA from the second sample comprises whole genome sequencing, whole exome sequencing, or targeted sequencing.
[0016] In some embodiments, the subject has completed at least one cancer treatment prior to obtaining the tumor sample. In some embodiments, the cancer treatment is selected from chemotherapy, radiotherapy, surgery, immunotherapy, cell therapy, or biologic therapy.
[0017] In some embodiments, methods of the present disclosure further comprise sequencing DNA from another sample from the subject at 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more successive time points. In some embodiments, sequencing DNA from another sample from the subject is repeated one or more times while the patient is in remission. In some embodiments, sequencing DNA from another sample from the subject is repeated one or more times while the patient is undergoing treatment for the cancer. In some embodiments, sequencing DNA from another sample from the subject is repeated one or more times coinciding with or prior to surgery; following, during, or prior to administration of chemotherapy; following, during, or prior to radiation therapy; following, during, or prior to administration of an immunotherapy; following, during, or prior to administration of a cell therapy; or following, during, or prior to administration of a biologic therapy.
[0018] In some embodiments, the subject has, had, or is suspected of having a cancer. In some embodiments, the cancer is selected from bladder cancer, breast cancer, colon / colorectal cancer, gynecologic cancers, head and neck cancers, hematological cancers, liver cancer, lung cancer, and skin cancer.DETAILED DESCRIPTION
[0019] It is to be appreciated that certain aspects, modes, embodiments, variations and features of the present methods are described below in various levels of detail in order to provide a substantial understanding of the present technology.
[0020] The present disclosure is not to be limited in terms of the particular embodiments described in this application, which are intended as single illustrations of individual aspects of the disclosure. All the various embodiments of the present disclosure will not be described herein. Many modifications and variations of the disclosure can be made without departing -4- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666from its spirit and scope, as will be apparent to those skilled in the art. Functionally equivalent methods and apparatuses within the scope of the disclosure, in addition to those enumerated herein, will be apparent to those skilled in the art from the foregoing descriptions. Such modifications and variations are intended to fall within the scope of the appended claims. The present disclosure is to be limited only by the terms of the appended claims, along with the full scope of equivalents to which such claims are entitled.
[0021] It is to be understood that the present disclosure is not limited to particular uses, methods, reagents, compounds, compositions or biological systems, which can, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting.
[0022] Methods of determining the source of a sample, such as a nucleic acid sample, are useful in, for example, biomedical and forensic applications. In some cases, such methods determine two or more samples (e.g., nucleic acid (e.g., DNA) samples) as being from the same or different subjects by comparing nucleic acid sequencing information, also referred to as “genetic relatedness” or “genetic identity.” These methods generally require the use of genotype calls at corresponding sites for comparison, such as Single Nucleotide Polymorphisms (SNPs) or Short Tandem Repeats (STRs). The data for genotype calls are typically produced by high-coverage sequencing methods. However, challenges in identifying two or more samples as being from the same or different subjects remain. For example, some samples do not include enough nucleic acid and / or nucleic acid of sufficient quality which can result in low-coverage sequencing data that can preclude the ability to generate accurate genotype calls and / or complete enough genotype calls to compare at corresponding sites between samples. Accordingly, there remains a need for improved methods of determining the source of a biological sample, including identifying different samples as being from the same or different subjects, by comparing nucleic acid sequencing information from low-coverage sequencing data.
[0023] The present disclosure provides, among other things, methods of determining the source of a sample. In particular, the disclosed methods facilitate determining (i) a first sample and a second sample are from the same subject; or (i) a first sample and a second-5- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666sample are from different subjects. The methods described herein are particularly useful for determining the source of the sample, wherein both a first and a second sample (and / or “further” / " another” sample) are sequenced at low-coverage. The ability to utilize low-coverage sequencing data can (i) reduce sequencing costs; (ii) reduce sequencing time; and / or (iii) allow for samples with a lower amount of nucleic acid and / or nucleic acid of insufficient quality for high-coverage sequencing to be utilized. Thus, the methods of the present disclosure improve current methods of determining the source of a sample, which require high-coverage sequencing information from at least one of the evaluated samples.Definitions
[0024] Unless defined otherwise, all technical and scientific terms used herein have the meaning commonly understood by a person skilled in the art to which this disclosure belongs. The following references provide one of skill with a general definition of many of the terms used in the present disclosure. Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed.1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, the following terms have the meanings ascribed to them below, unless specified otherwise. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure.
[0025] As used herein, the term “alt-allele fraction” refers to the proportion (or “fraction”) of sequencing reads covering a variant position that does not match the reference (e.g., wildtype) allele.
[0026] As used herein, the term “amplification,” with respect to nucleic acid sequences, refers to methods that increase the representation of a population of nucleic acid sequences in a sample. Copies of a particular target nucleic acid sequence generated in vitro in an amplification reaction are called “amplicons” or “amplification products”. Amplification may be exponential or linear. A target nucleic acid may be DNA (such as, for example, genomic DNA, cfDNA, ctDNA, and cDNA) or RNA. Amplification can be achieved using polymerase-6- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666chain reaction (PCR) as well as numerous other methods such as isothermal methods, rolling circle methods, etc.
[0027] As used herein, the term “approximately” or “about” means plus or minus 10% as well as the specified number. For example, “about 10” should be understood as both “10” and “9-11”.
[0028] As used herein, the term “biopsy” refers to a tissue sample excised from a subject (e.g., a subject with cancer). Tissue samples may be obtained using any suitable method, including, but not limited to, needle biopsies, aspiration, scraping, excision using surgical equipment, etc.
[0029] As used herein, the term “comparable” refers to two (or more) sets of conditions, circumstances, individuals, or populations that are sufficiently similar to one another to permit comparison of results obtained or phenomena observed. In some embodiments, comparable sets of conditions, circumstances, individuals, or populations are characterized by a plurality of substantially identical features and one or a small number of varied features. Those of ordinary skill in the art will appreciate that sets of circumstances, individuals, or populations are comparable to one another when characterized by a sufficient number and type of substantially identical features to warrant a reasonable conclusion that differences in results obtained or phenomena observed under or with different sets of circumstances, individuals, or populations are caused by or indicative of the variation in those features that are varied. Those skilled in the art will appreciate that relative language used herein (e.g., enhanced, activated, reduced, inhibited, etc.) will typically refer to comparisons made under comparable conditions.[0030J As used herein, the term “derived from” encompasses the terms “originated from,” “obtained from,” “obtainable from,” “isolated from,” and “created from,” and generally indicates that one specified material (e.g., a biological sample) finds its origin in another specified material or individual or has features that can be described with reference to another specified material.-7- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666[00311 As used herein, term “genomic DNA” refers to DNA of a cellular genome. The genomic DNA can be cellular, i.e., contained within a cell, or it can be cell-free.
[0032] As used herein, the term “common” when used in the context of “common germline variant(s)” refers to variant or SNPs for which the minor allele frequency is approximately 50%. For example, to be considered a “common germline variant” the minor allele frequency is inclusive of 40% to 60%.[00331 As used herein, the term “high-coverage” refers to sequencing data wherein a relatively larger number of reads are sequenced per region of nucleic acid such that heterozygous genotypes can be called with very high confidence . The number of reads sequenced per region of nucleic acid is also referred to herein as “depth” or “sequencing depth.” “High-coverage” sequencing data refers to data wherein the depth of the sequencing is above about 15X. In some embodiments, “high-coverage” sequencing data refers to data wherein the depth of the sequencing is about 20X, 30X, 40X, 50X, 60X, 70X, 80X, 90X, 100X, 150X, 200X, 250X, 300X, 350X, 400X, 450X, or 500X. “High-coverage” sequencing depth may depend, in part, on the type of sequencing technologies utilized.
[0034] As used herein, the term “low-coverage” refers to sequencing data wherein a smaller number of reads are sequenced per region of nucleic acid such that heterozygous genotype calls cannot be made with high confidence. The number of reads sequenced per region of nucleic acid is also referred to herein as “depth” or “sequencing depth.” “Low-coverage” sequencing data refers to data wherein the depth of the sequencing is below about 15X. In some embodiments, “low-coverage” sequencing data refers to data wherein the depth of the sequencing is about 14X, 13X, 12X, 11X, 10X, 9X, 8X, 7X, 6X, 5X, 4X, 3X, 2X, or IX. “Low-coverage” sequencing depth may depend, in part, on the type of sequencing technologies utilized.
[0035] As used herein, the term “minor allele frequency” refers to the proportion (“frequency”) of a less prevalent (e.g., second most prevalent) allele in a population. For example, if a particular loci on an allele contains an adenosine in 60% of a population and a thymine 40% of the population, the “minor allele frequency of a thymine at that loci is 40%.-8- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666
[0036] As used herein, the term “heterozygous” refers to a loci or gene having different nucleotides or nucleotide sequences on each allele. For example, a heterozygous gene may comprise one Single Nucleotide Polymorphism (SNP) on one allele, and one wild-type nucleotide at the corresponding site on the other allele.
[0037] As used herein, the term “homozygous alternate” refers to a loci comprising two identical, non-wild-type (“variant”) nucleotides at the loci (e.g., one variant nucleotide on the sense strand and one variant nucleotide on the anti-sense strand) or a gene comprising two identical, non-wild-type (“variant”) alleles.
[0038] As used herein, the term “homozygous-reference” refers to a loci comprising two identical, wild-type (“non-varianf ’) nucleotides at the loci (e.g., one non-variant nucleotide on the sense strand and one non-variant nucleotide on the anti-sense strand) or a gene comprising two identical, wild-type (“non-variant”) alleles.[00391 As used herein, the term “genotyping” refers to a process of determining the alleles an individual at particular genetic loci by examining an individual’s DNA. Genotyping differs from sequencing in which all of the nucleotides comprising a specific length of DNA are assessed.
[0040] As used herein, the terms “improved”, “increased”, or “reduced”, or grammatically comparable comparative terms, indicate values that are relative to a comparable reference measurement. For example, in some embodiments, an assessed value achieved with an agent of interest may be “improved” relative to that obtained with a comparable reference agent. Alternatively or additionally, in some embodiments, an assessed value achieved in a subject or system of interest may be “improved” relative to that obtained in the same subject or system under different conditions, or in a different, comparable subject (e.g., in a comparable subject or system that differs from the subject or system of interest in presence of one or more indicators of a particular disease, disorder or condition of interest, or in prior exposure to a condition or agent, etc.). In some embodiments, comparative terms refer to statistically relevant differences (e.g, that are of a prevalence and / or magnitude sufficient to achieve statistical relevance). Those skilled in the art will be aware, or will readily be able to-9- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666determine, in a given context, a degree and / or prevalence of difference that is required or sufficient to achieve such statistical significance.
[0041] As used herein, the term “multi-nucleotide variant” or “MNV” refers to a variant having 2 or more adjacent nucleotide changes.
[0042] As used herein, the term “Next Generation Sequencing” or “NGS” refers to sequencing methods that allow for massively parallel sequencing of clonally amplified and of single nucleic acid molecules during which a plurality, e.g., millions, of nucleic acid fragments from a single sample or from multiple different samples are sequenced in unison. Non-limiting examples of NGS include sequencing-by-synthesis, sequencing-by-ligation, real-time sequencing, and nanopore sequencing.
[0043] As used herein, the term “patient-specific panel” or “patient-specific somatic variants” refers to a collection of sequences comprising somatic mutations that are specific to a patient, or markers that distinguish between two or more individuals. A signature panel may distinguish one sample from another.
[0044] As used herein, the term “sample” or “biological sample,” refers to a biological sample obtained or derived from a source of interest, as described herein. In certain embodiments, a source of interest comprises an organism, such as a microbe, a plant, an animal or a human. In certain embodiments, a biological sample is or comprises biological tissue or fluid. In certain embodiments, a biological sample may be or comprise bone marrow; blood (or a fraction thereof); blood cells; ascites; tissue or fine needle biopsy samples; cell-containing body fluids; free floating nucleic acids (e.g., cell free DNA); sputum; saliva; urine; cerebrospinal fluid, peritoneal fluid; pleural fluid; lymph; gynecological fluids; skin swabs; vaginal swabs; oral swabs; nasal swabs; washings or lavages such as a ductal lavages or broncheoalveolar lavages; aspirates; scrapings; bone marrow specimens; tissue biopsy specimens; surgical specimens; feces, other body fluids, secretions, and / or excretions; and / or cells therefrom, etc. In certain embodiments, a biological sample is or comprises cells obtained from an individual. In certain embodiments, obtained cells are or include cells from an individual from whom the sample is obtained. In certain embodiments, a sample is a “primary sample” obtained directly from a source of interest by -10- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666any appropriate means. For example, in certain embodiments, a primary biological sample is obtained by methods selected from the group consisting of a swab, biopsy (e.g., fine needle aspiration or tissue biopsy), surgery, collection of body fluid (e.g., blood, lymph, feces etc.), etc. In certain embodiments, as will be clear from context, the term “sample” refers to a preparation that is obtained by processing (e.g., by removing one or more components of and / or by adding one or more agents to) a primary sample. For example, filtering using a semi-permeable membrane. Such a processed “sample” may comprise, for example nucleic acids or proteins extracted from a sample or obtained by subjecting a primary sample to techniques such as amplification or reverse transcription of mRNA, isolation and / or purification of certain components, etc.[0045 [ As used herein, the term “sequence read” or simply “read” refers to sequence information of a nucleic acid fragment obtained through a sequencing assay, such as a next generation sequencing (NGS) assay. In some embodiments, a sequence read refers to data representing a sequence of nucleotide bases that were measured using a clonal sequencing method. Clonal sequencing may produce sequence data representing single, or clones, or clusters of one original DNA molecule. A sequence read may also have associated quality score at each base position of the sequence indicating the probability that nucleotide has been called correctly.
[0046] As used herein, a “set” of reads refers to all sequencing reads with a common parent nucleic acid strand, which may or may not have had errors introduced during sequencing or amplification of the parent nucleic acid strand.[0047| As used herein, the term “Single Nucleotide Polymorphism” or “SNP” refers to a single base pair variation in a nucleic acid sequence. SNPs can also be referred to as Single-Nucleotide Variants (SNVs).
[0048] As used herein, the term “somatic variant” or “somatic mutation” refers to a variant arising after conception, in non-germline DNA of an individual. Somatic variants may include single-nucleotide variants (SNVs), multi -nucleotide variants, insertions and deletions (e.g., indel variants), and genomic rearrangements for example. The terms “somatic variant” and “somatic mutation” are used interchangeably herein.-11- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666[0O49| As used herein, the term “subject” or “patient” or “individual” refers to any organism upon which embodiments of the present disclosure may be used or administered, e.g., for experimental, screening, diagnostic, prophylactic, and / or therapeutic purposes. Typical subjects include animals (e.g., mammals such as mice, rats, rabbits, non-human primates, and humans; insects; worms; etc.).
[0050] As used herein, the term “target sequence” refers to a selected target polynucleotide, e.g., a sequence present in a cfDNA molecule, whose presence, amount, and / or nucleotide sequence, or changes in these, are desired to be determined. Target sequences can be interrogated for the presence or absence of a somatic and / or germline variant. The target polynucleotide can be a region of gene associated with a disease. In some embodiments, the region is an exon. The disease can be cancer.[00511 As used herein, term “tumor fraction” refers to the proportion of circulating cell-free tumor DNA (ctDNA) relative to the total amount of cell-free DNA (cfDNA). Tumor fraction may be indicative of the size of the tumor.10052] As used herein, the term “tumor-specific somatic mutations” refer to nucleic acid e.g., DNA) changes (“variants”) that occur in a somatic cell before or during tumor development. This type of variant is not present within the germline.
[0053] As used herein in the context of molecules, e.g., nucleic acids, proteins, or small molecules, the term “variant” refers to a molecule that shows significant structural identity with a reference molecule but differs structurally from the reference molecule, e.g., in the presence or absence or in the level of one or more chemical moieties as compared to the reference entity. In some embodiments, a variant also differs functionally from its reference molecule. In general, whether a particular molecule is properly considered to be a “variant” of a reference molecule is based on its degree of structural identity with the reference molecule. As will be appreciated by those skilled in the art, any biological or chemical reference molecule has certain characteristic structural elements. A variant, by definition, is a distinct molecule that shares one or more such characteristic structural elements but differs in at least one aspect from the reference molecule. To give but a few examples, a polypeptide may have a characteristic sequence element comprised of a plurality of amino acids having -12- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666designated positions relative to one another in linear or three-dimensional space and / or contributing to a particular structural motif and / or biological function; a nucleic acid may have a characteristic sequence element comprised of a plurality of nucleotide residues having designated positions relative to one another in linear or three-dimensional space. In some embodiments, a variant polypeptide or nucleic acid may differ from a reference polypeptide or nucleic acid as a result of one or more differences in amino acid or nucleotide sequence and / or one or more differences in chemical moieties (e.g., carbohydrates, lipids, phosphate groups) that are covalently components of the polypeptide or nucleic acid (e.g., that are attached to the polypeptide or nucleic acid backbone). In some embodiments, a variant polypeptide or nucleic acid shows an overall sequence identity with a reference polypeptide or nucleic acid that is at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, or 99%. In some embodiments, a variant polypeptide or nucleic acid does not share at least one characteristic sequence element with a reference polypeptide or nucleic acid. In some embodiments, a reference polypeptide or nucleic acid has one or more biological activities. In some embodiments, a variant polypeptide or nucleic acid shares one or more of the biological activities of the reference polypeptide or nucleic acid. In some embodiments, a variant polypeptide or nucleic acid lacks one or more of the biological activities of the reference polypeptide or nucleic acid. In some embodiments, a variant polypeptide or nucleic acid shows a reduced level of one or more biological activities as compared to the reference polypeptide or nucleic acid. In some embodiments, a polypeptide or nucleic acid of interest is considered to be a “variant” of a reference polypeptide or nucleic acid if it has an amino acid or nucleotide sequence that is identical to that of the reference but for a small number of sequence alterations at particular positions. Typically, fewer than about 20%, about 15%, about 10%, about 9%, about 8%, about 7%, about 6%, about 5%, about 4%, about 3%, or about 2% of the residues in a variant are substituted, inserted, or deleted, as compared to the reference. In some embodiments, a variant polypeptide or nucleic acid comprises about 10, about 9, about 8, about 7, about 6, about 5, about 4, about 3, about 2, or about 1 substituted residues as compared to a reference. Often, a variant polypeptide or nucleic acid comprises a very small number (e.g., fewer than about 5, about 4, about 3, about 2, or about 1) number of substituted, inserted, or deleted, functional residues (i.e., residues that participate in a particular biological activity) relative to the reference. In some-13- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666embodiments, a variant polypeptide or nucleic acid comprises not more than about 5, about 4, about 3, about 2, or about 1 addition or deletion, and, in some embodiments, comprises no additions or deletions, as compared to the reference. In some embodiments, a variant polypeptide or nucleic acid comprises fewer than about 25, about 20, about 19, about 18, about 17, about 16, about 15, about 14, about 13, about 10, about 9, about 8, about 7, about 6, and commonly fewer than about 5, about 4, about 3, or about 2 additions or deletions as compared to the reference. In some embodiments, a reference polypeptide or nucleic acid is one found in nature. In some embodiments, a reference polypeptide or nucleic acid is a human polypeptide or nucleic acid.Methods for Determining the Source of a Sample
[0054] The present disclosure provides, among other things, methods of determining the source of a sample. Such methods can comprise sequencing at low coverage DNA from a first sample from a subject, thereby obtaining a first set of sequencing reads; preparing a panel of germline variants comprising a plurality of germline variants (e.g., common germline variants); genotyping each germline variant from the panel of germline variants sequenced in the first set of sequencing reads; sequencing at low coverage DNA from a second sample, thereby obtaining a second set of sequencing reads; counting heterozygous and homozygous reference-bearing sequencing reads at each germline variant from the panel of germline variants sequenced in the second set of sequencing reads; determining (i) the first sample and the second sample are from the same subject or (ii) the first sample and the second sample are from different subjects based on a likelihood of observing the counts of sequencing reads for the panel of germline variants for each sample. In some embodiments, the panel of germline variants comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, or more common germline variants. In some embodiments, the panel of germline variants comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, or more common germline variants. In some embodiments, the panel of germline variants comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more common germline variants. In some-14- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666embodiments, the panel of germline variants consists of a plurality of common germline variants.
[0055] For the purposes of the present disclosure, a “common germline variant” may have a minor allele frequency of 40% to 60%, 45% to 55%, or about 50%. For example, a “common germline variant” may have a minor allele frequency of 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, or any value in between. Additionally, it should be understood that while common germline variants are useful for practicing the disclosed methods, they are not required. Rather, the disclosed methods can utilize any panel of germline variants, regardless of minor allele frequency, but as the frequency of minor allele decreases, more sites may be required to obtain the same level of power that would have been achieved with relatively fewer common germline variants.
[0056] No current methods for identifying the source of a sample (e.g., determining (i) a first sample and a second sample are from the same subject or (ii) a first sample and a second sample are from different subjects) can utilize low-coverage DNA sequencing information from both the first and the second (and / or further / another) samples. However, the methods of the present disclosure are useful and effective in determining the source of a sample even when DNA from both the first and second (and / or further / another) samples are sequenced at low-coverage (e.g., do not include enough nucleic acid and / or nucleic acid of sufficient quality to be sequenced at high-coverage to more reliably produce genotype calls).Accordingly, the ordered combination of steps of technologies of the present disclosure are not well-known, routine, or conventional and constitute an improvement in the field of determining the source of a sample. The ability to utilize low-coverage sequencing data to determine if a first sample and a second (and / or further / another) sample are from the same or different subjects can (i) reduce sequencing costs; (ii) reduce sequencing time; and / or (iii) allow for samples with a lower amount of nucleic acid and / or nucleic acid of insufficient quality for high-coverage sequencing to be utilized. Thus, the methods of the present disclosure improve current methods of determining the source of a sample, which required high-coverage sequencing information from at least one of the evaluated samples.-15- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666[0057| The low-coverage sequencing data used in accordance with the technologies of the present disclosure can be prepared by listing a variant’s coordinates, the homozygous reference and homozygous alternate alleles, and minor allele frequencies (e.g., in a file, such as a CSV file). Binary Alignment Maps (BAMs) from each sample can be analyzed to extract the count of alternative-bearing and total read pairs covering variant positions. This class can use data from paired-end sequencing and count read pairs as single molecules. Read counts can be stored as a set of Polars DataFrames and used for downstream analyses.
[0058] To determine if (i) the first sample and the second (and / or further / another) sample are from the same subject or (ii) the first sample and the second (and / or further / another) sample are from different subjects based on a likelihood of observing the counts of sequencing reads (also referred to herein as “genotype likelihood”) for the panel of germline variants (e.g., a panel of common germline variants) for each sample, it is determined whether the first and second (and / or further) samples are more likely to be from the same subject or different subjects (e.g., two unrelated subjects). Genotype likelihoods, rather than genotype calls, are utilized to, for example, account for the low-coverage DNA sequencing which can result in artifactual genotype call mismatches at a given germline variant. In some embodiments, the likelihood of observing the counts of sequencing reads for the panel of germline variants is determined by calculating a genotype likelihood selected from (i) homozygous-reference, (ii) heterozygous, and (iii) homozygous alternate for each of the germline variants in the first set of sequencing reads and the second set of sequencing reads.[00591 Calculating the genotype likelihood can comprise modeling a statistical distribution, such as a binomial distribution. Other suitable statistical distributions include, but are not limited to, negative binomial, Poisson, and gamma distributions. For example, in some embodiments, a binominal model determines the genotype likelihoods from homozygous-reference, heterozygous, and homozygous-alternate genotypes using the formulas:p(genotype=0 / 0|read_counts)=dbinom(x=n_alt_reads,n=n_total_reads,p=0+error)p(genotype=0 / l|read_counts)=dbinom(x=n_alt_reads,n=n_total_reads,p=0.5)p(genotype=l / l|read_counts)=dbinom(x=n_alt_reads,n=n_total_reads,p=l-error)-16- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666where p is a probability, dbinom is a function returning the probability mass of a binomial distribution given a set of parameters, “n alt reads” is the number of sequence reads comprising an alt allele, and “n total reads” is the total number of sequence reads, and error is the error rate of the sequencing and library preparation process.
[0060] In some embodiments, genotype likelihood of a given germline variant in a first sample or a second (or further) sample, independently, is a probability of sequencing read counts for a given genotype of the given germline variant times a probability of the given genotype from a pool with known alt-allele fraction, marginalized over all genotypes of the given germline variant. Such a genotype likelihood of a given variant can be determined using the formula:L(site, one sample) = \sum_k p(read_counts|genotype_k) * p(genotype_k|af), wherein p(genotype_k|af) is the mendelian expectation.[00611 The mendelian expectation is also the binomial probability mass (e.g., with allele frequencies p=0.5 and q=0.5, the probability of a subject being heterozygous at a given loci is binom.pmf(x=l,n=2,p=0.5) = 0.5 = 2pq).
[0062] In some embodiments, calculating the genotype likelihood comprises calculating a mean of a likelihood of a given germline variant in the first sample and a likelihood of the given germline variant in the second (and / or further) sample. The likelihood of two samples’ genotype call at one site can be determined as the product of the per-sample likelihoods:L(site, two samples) = L(site,samplel) * L(site,sample2)[00631 To determine if (i) the first sample and the second sample are from the same subject or (ii) the first sample and the second sample are from different subjects, the likelihoods of observing the counts of sequencing reads for both the first and second samples if they were drawn from a subject with the same genotype can be compared to the likelihood of observing the counts of sequencing reads for the first and second samples if they were drawn independently from two random individuals.-17- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666[0O64| If the genotype of the focal individual (e.g., a subject of interest) was known, the likelihood that the two samples came from the same individual could be calculated with: L(match) = \prod_sites L(site|gt), where gt is the alt-allele fraction in the focal individual’s genotypes. For technologies of the present disclosure, these genotypes are typically unknown. Further, in view of the low-coverage sequencing used in accordance with methods described herein, it is expected that the genotypes of each sample cannot be called without error. To allow calculation of L(match) in these cases, we can complete a reciprocal swap. A reciprocal swap comprises first determining the genotype likelihood of a germline variant from the first sample and then determining the probability of observing the counts of sequencing reads from the second sample if they were obtained from a subject with the genotype likelihood from the first sample. Subsequently, the genotype likelihood of a germline variant from the second sample is determined and the probability of observing counts of sequencing reads from the first sample if they were obtained from a subject with the genotype likelihood from the second sample is determined. The likelihood of a “match” (e.g., the first sample and the second sample being from the same subject) can be determined as the mean of those two likelihoods:L(match)=(L(match|E(gt|samplel_read_counts))+ L(match|E(gt|sample2_read_counts))) / 2
[0065] The likelihood of the data if the samples are from different individuals is calculated with the formula: L(no match) = \prod_sites L(site|maf), where maf is the minor allele frequency of the corresponding allele in the population.]0066| The likelihood of observing the genotypes for the panel of germline variants for each sample can then be determined as a likelihood ratio for “match” (samples from the same subject) versus “non-match” (samples from different subjects):LR=L(match) / L(random)
[0067] A higher likelihood ratio (LR) indicates relatively higher confidence that the first sample and the second sample are from the same subject; whereas a lower LR indicates that the first sample and the second sample are likely from different subjects.-18- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666[0O68| Thus, in some embodiments, the first sample and the second sample are from the same subject when the likelihood of observing the counts of sequencing reads for the panel of germline variants for each sample is above a predetermined threshold. In some embodiments, the first sample and the second sample are from different subjects when the likelihood of observing the counts of sequencing reads for the panel of germline variants for each sample is below a predetermined threshold.
[0069] The predetermined threshold can be set by analyzing distributions of likelihood ratios calculated for a large number of pairs of known matched and unmatched samples. The cutoff can be set at a value that yields a pre-specified level of sensitivity and / or specificity on the known matched and unmatched data. For example, the cutoff may be about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, or about 15 or more. In some embodiments, the cutoff is 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 or more. In some embodiments, the cutoff is 10, yielding > 99% sensitivity to unmatched sample pairs.
[0070] Methods of the present disclosure can also comprise extracting DNA from a sample. DNA can be extracted from a sample by (1) harvesting cells or tissue (e.g., a tumor biopsy); (2) lysing the cells; (3) inactivating DNAses; (4) capturing DNA; and (5) separating the DNA from at least some of the components with which it was associated when initially produced (e.g., RNA, proteins, other cellular components), whether in nature and / or in an experimental setting; and (6) resuspending or eluting the DNA.
[0071] Harvesting tissues or cells may be completed by, for example, blood draw, needle biopsy, aspiration, scraping, or excision using surgical equipment. Subsequently, the cells may be lysed using lysis buffer (e.g., comprising a chaotropic agent) and / or by mechanical disruption. In step (3), DNAses may be inactivated by, for example, heat treatment (e.g., 5 minutes at 75°C) and / or by the use of DNAse inhibitors. Following lysis and DNAse inactivation, DNA can be captured by binding to a surface, such as a silica surface.Separating DNA from at least some of the components with which it was associated when initially produced can be completed by, for example, degrading RNA with RNAse and degrading protein with proteinase K. Alternative or additional methods include dissolving the-19- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666sample in buffers containing certain salts (e.g., guanidinium salts) to remove proteins and / or washing away components with which the DNA was associated when initially produced while the DNA is bound to a solid-support (e.g., silica beads). Finally, the extracted DNA can be eluted or resuspended in water or a buffer e.g., a buffer suitable for use in downstream applications and / or analyses).[00721 In some embodiments, the extracted DNA is quantified and / or the quality of the DNA is assessed prior to sequencing (e.g., low-coverage sequencing) DNA (e.g., DNA from a first, second, and / or subsequent sample). Spectroscopic and / or electrophoretic methods can be used to quantify DNA or to assess its quality.
[0073] Methods of the present disclosure can comprise amplifying the DNA, for example, amplifying the DNA of a panel of germline variants. DNA can be amplified by, for example, polymerase chain reaction (PCR), such as real-time PCR (e.g., TaqMan) and quantitative PCR (qPCR).
[0074] DNA for use in accordance with technologies described herein can be low-coverage sequenced from a sample (e.g., a first sample, a second sample, a subsequent (“further”) sample) by whole genome sequencing, whole exome sequencing, or targeted sequencing. In some embodiments, sequencing DNA from the first sample comprises whole genome sequencing, whole exome sequencing, or targeted sequencing. In some embodiments, sequencing DNA from the first sample comprises whole genome sequencing. In some embodiments, sequencing DNA from the first sample comprises whole exome sequencing. In some embodiments, sequencing DNA from the first sample comprises targeted sequencing. In some embodiments, sequencing DNA from the second sample comprises whole genome sequencing, whole exome sequencing, or targeted sequencing. In some embodiments, sequencing DNA from the second sample comprises whole genome sequencing. In some embodiments, sequencing DNA from the second sample comprises whole exome sequencing. In some embodiments, sequencing DNA from the second sample comprises targeted sequencing. In some embodiments, sequencing DNA from a further sample comprises whole genome sequencing, whole exome sequencing, or targeted sequencing. In some embodiments, sequencing DNA from a further sample comprises whole genome sequencing.-20- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666In some embodiments, sequencing DNA from a further sample comprises whole exome sequencing. In some embodiments, sequencing DNA from a further sample comprises targeted sequencing.
[0075] In some embodiments, methods of the present disclosure further comprise detecting in the second sample sequence reads comprising one or more tumor-specific somatic mutations that are determined for the subject prior to sequencing the second sample. In some embodiments, the methods further comprise enriching the DNA from the second sample for DNA comprising one or more of the tumor-specific somatic mutations prior to sequencing. Enriching the DNA comprising one or more of the tumor-specific somatic mutations can comprise the use of, for example, hybrid capture enrichment or PCR-based enrichment.
[0076] Methods of the present disclosure can further comprise enriching the DNA from a sample (e.g., a first sample, second sample, subsequent (further / another) sample) for DNA comprising one or more of the panel of germline variants prior to sequencing. In some embodiments, methods of the present disclosure further comprise enriching the DNA from the first sample for DNA comprising one or more of the panel of germline variants prior to sequencing. In some embodiments, methods of the present disclosure further comprise enriching the DNA from the second sample for DNA comprising one or more of the panel of germline variants prior to sequencing. DNA enrichment can comprise the use of hybrid capture enrichment or PCR-based enrichment. In some embodiments, methods of the present disclosure further comprise enriching the DNA from a further (“subsequent’ ’ / “another”) sample for DNA comprising one or more of the panel of germline variants prior to sequencing. DNA enrichment can comprise the use of hybrid capture enrichment or PCR-based enrichment. Thus, in some embodiments, enriching the DNA from the first sample for DNA comprising one or more of the panel of germline variants comprises hybrid capture enrichment or PCR-based enrichment. In some embodiments, enriching the DNA from the second sample for DNA comprising one or more of the panel of germline variants comprises hybrid capture enrichment or PCR-based enrichment. In some embodiments, enriching the DNA from a further sample for DNA comprising one or more of the panel of germline variants comprises hybrid capture enrichment or PCR-based enrichment.-21- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666Subjects and Samples
[0077] The present disclosure provides, among other things, methods for determining the source of the sample, wherein a first sample from a subject and a second sample can be determined as (i) from the same subject; or (ii) from different subjects. Thus, the subject of such methods can be any subject and a sample (e.g, a first sample, a second sample) of such methods can be any sample (e.g, a biological sample) obtained or derived from a source of interest (e.g., a subject).
[0078] A sample (e.g., a first sample, or a second sample) for use in accordance with the methods described herein can be a biological sample. In some embodiments, a sample e.g., a first sample, a second sample) is a sample of blood (e.g., whole blood, a blood fraction), plasma, serum, urine, or tumor. In some embodiments, a first sample for use in accordance with the technologies described herein in a sample of blood, plasma, serum, or urine. In some embodiments, a first sample is a tumor sample. In some embodiments, a second sample for use in accordance with the technologies described herein is a sample of blood, plasma, serum, or urine. In some embodiments, a second sample is a tumor sample.
[0079] In some embodiments, a first sample is obtained at a first time point and a second sample is obtained a second (e.g., later) time point. In some embodiments, a further (or “subsequent” or “another”) sample is obtained (e.g., from a subject) at a plurality of successive time points. For example, in some embodiments, a further sample is obtained at 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more successive time points. Obtaining a second and / or further samples can be useful in, for example, tracking a subject’s condition, disease, and / or disorder over time.
[0080] In some embodiments, the subject is a healthy subject. In some embodiments, the subject is a subject with a disease, disorder, and / or condition. In some embodiments, the subject was previously diagnosed with cancer. In some embodiments, the subject has cancer. In some embodiments, the subject is in remission from cancer. In some embodiments, wherein the subject has, had, or is suspected of having a cancer or tumor. The type of cancer or tumor can is not particularly limited and can include, but is not limited to, bladder cancer,-22- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666breast cancer, colon / colorectal cancer, gynecologic cancers, head and neck cancers, hematological cancers, liver cancer, lung cancer, and skin cancer.
[0081] In some embodiments, the subject has completed at least one cancer treatment. The cancer treatment can be selected from, for example, chemotherapy, radiotherapy, surgery, immunotherapy, cell therapy, or biologic therapy. In some embodiments, the first and / or second sample is a tumor sample, and the subject has completed at least one cancer treatment prior to obtaining the tumor sample.
[0082] In some embodiments, methods of the present disclosure further comprise repeating sequencing DNA from another (“further”, “subsequent”) sample from the subject at 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more successive time points. In some embodiments, sequencing DNA from a further sample from the subject is repeated one or more times while the subject is in remission e.g., from cancer).[00831 In some embodiments, sequencing DNA from a further (“subsequent”, “another”) sample from the subject is repeated one or more times while the subject is undergoing treatment for cancer. In some embodiments, sequencing DNA from a further sample from the subject is repeated one or more times coinciding with, or prior to, surgery; following, during, or prior to administration of chemotherapy; following, during, or prior to radiation therapy; following, during, or prior to administration of an immunotherapy; following, during, or prior to administration of a cell therapy; or following, during, or prior to administration of a biologic therapy.Panels of Germline Variants
[0084] A common germline variant in accordance with the technologies of the present disclosure is a variant in a germline (e.g., reproductive) cell that can be passed on to offspring and has a minor allele frequency of about 40-60%. As disclosed herein, the methods provided may utilize a panel of germline variants comprising or consisting of a plurality of common germline variants. However, as noted above, other germline variants with less common allele frequencies can also be utilized. In some embodiments, the panel of germline variants may comprise a germline variant with a minor allele frequency of about 25%. In some-23- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666embodiments, the panel of germline variants may comprise a germline variant with a minor allele frequency of about 30%. In some embodiments, the panel of germline variants may comprise a germline variant with a minor allele frequency of about 35%. In some embodiments, the panel of germline variants may comprise a germline variant with a minor allele frequency of about 40%. In some embodiments, the panel of germline variants may comprise a germline variant with a minor allele frequency of about 45%. In some embodiments, the panel of germline variants may comprise a germline variant with a minor allele frequency of about 50%. In some embodiments, the panel of germline variants may comprise a germline variant with a minor allele frequency of about 55%. In some embodiments, the panel of germline variants may comprise a germline variant with a minor allele frequency of about 60%.
[0085] Panels of germline variants comprise a plurality of germline variants (e.g., common germline variants). While evaluation of a single common germline variant may allow for discrimination of whether two or more samples are from the same or different subjects, addition of each evaluated common germline variant or germline variant to a panel of germline variants increases the likelihood that the combination of particular germline variants genotypes from the panel of germline variants are a unique combination associated with a particular subject. Accordingly, such panels of germline variants can be utilized in methods of the present disclosure to determine the source of a sample, e.g., to determine if (i) a first sample and a second sample are from the same subject; or (ii) a first sample and a second sample are from different subjects.
[0086] In some embodiments, panels of germline variants for use in accordance with the technologies described herein are not sample-informed or subject-specific (e.g., the selection of variants for the panel is not dependent on or unique to a particular sample or subject). Rather, the disclosed methods may utilize the same panel of germline variants for all sample evaluation, which is possible, at least in part, due to their minor allele frequency across a plurality of subjects.Use in Minimal Residual Disease Detection-24- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666[0087J Technologies of the present disclosure can be particularly useful in determining the source of samples obtained over a period of time (e.g., as from the same or different subjects, e.g., for disease monitoring), including in detecting Minimal Residual Disease (MRD). The goal of a MRD assay is to detect and / or quantify circulating tumor DNA (ctDNA) so researchers and clinicians can detect recurrence early and monitor the progress of the disease (e.g., cancer) through treatment. In general, a MRD assay will rely on a patient-specific and tumor-specific panel (i.e., a “signature panel” or a “panel of patient-specific somatic variants”) for assessing the presence of ctDNA in a patient (“subject”) sample. The signature panel can be prepared with the general steps of (1) profiling a tumor or cancer sample from a patient, and (2) identifying a subset of somatic mutations to target, and, at one or more later time points, (3) taking a subsequent sample from the patient, (4) enriching cell-free DNA (cfDNA) for the target somatic mutation sites, and (5) determining or estimating the ctDNA content of cell free DNA (cfDNA) given the tumor profile and sequencing data.
[0088] More specifically, preparing the patient-specific and tumor-specific panel (i.e., a “signature panel”) may comprise, for example, (a) obtaining a tumor sample and a non-tumor sample from a cancer patient; (b) sequencing DNA (e.g., genomic DNA) from the tumor sample and sequencing DNA (e.g., cell free DNA or “cfDNA”) from the non-tumor sample, thereby obtaining sequences of DNA or sequence reads from the tumor sample and the non-tumor sample; and (c) comparing the sequences of the tumor sample and the non-tumor sample to determine any tumor-specific somatic mutations that are present in the sequences of DNA from the tumor sample but not present in the sequences of DNA from the non-tumor sample. Sequencing of the DNA from the tumor sample and non-tumor sample may comprise whole genome sequencing or various types of targeted sequencing, such as whole exome sequencing.
[0089] This comparison of the tumor and non-tumor sequences can be performed by, for example, aligning the sequences of DNA (e.g., genomic DNA) from the tumor sample to a reference human genome that is not from the patient and aligning the sequences of DNA (e.g., cfDNA) from the non-tumor sample to the reference genome that is not from the patient. The reference genome can be, for example, a publicly available human genome assembly, such as hgl8, hgl9, GRCh38.pl4, GRCh37.pl3, or other assemblies from the-25- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666Genome Reference Consortium. Alternatively, the comparison of the tumor and non-tumor sequences can be performed by, for example, aligning the sequences of DNA (e.g., genomic DNA) from the tumor sample to sequences of DNA (e.g., cfDNA) from the non-tumor sample. With either approach, the skilled artisan is able to detect and identify tumor-specific somatic mutations that are present in the tumor sample but not in the non-tumor sample.[0090| In some embodiments, preparing the tumor-specific panel (i.e., a “signature panel”) may comprise, for example (a) obtaining a tumor sample from a cancer patient; (b) sequencing DNA e.g., genomic DNA) from the tumor sample, thereby obtaining sequences of DNA or sequence reads from the tumor sample; and (c) comparing the sequences of the tumor sample to one or more reference genomes or non-tumor samples from the subject to determine any tumor-specific somatic mutations that are present in the sequences of DNA from the tumor sample but not present in the sequences of DNA from the one or more reference genomes or non-tumor samples. This comparison may be performed by, for example, aligning the sequences of DNA (e.g., genomic DNA) from the tumor sample to the sequences of the one or more reference genomes. Again, the reference genome can be, for example, a publicly available human genome assembly, such as hgl8, hgl9, GRCh38.pl 4, GRCh37.pl3, or other assemblies from the Genome Reference Consortium. Additionally or alternatively, genomic sequences from a non-tumor sample from the same patient may also be used as a reference genome or in conjunction with another reference genome (e.g., hgl8, hgl9, GRCh38.pl4, GRCh37.pl3, etc.) to determine tumor-specific somatic variants. Such alignment, again, allows the skilled artisan to detect and identify tumor-specific somatic mutations that are present in the tumor sample. Sequencing of the DNA from the tumor sample may comprise whole genome sequencing or various types of targeted sequencing, such as whole exome sequencing or subtractive hybridization.[00911 In some embodiments, rather than compare the sequences of the tumor sample to one or more reference genomes, mathematical algorithms and / or artificial intelligence are utilized to determine the likelihood of any identified potential somatic mutations as being a tumorspecific somatic mutation that is present in the sequences of DNA from the tumor sample but not present in the sequences of DNA from one or more reference genomes and / or non-tumor samples. In some embodiments, such mathematical algorithms and / or machine learning are-26- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666utilized in combination with comparing the sequences of the tumor sample to one or more reference genomes and / or non-tumor samples.
[0092] The tumor sample may be a solid tumor sample, such as a biopsy or other tissue sample, or a liquid sample, such as blood (in the case of a hematological cancer) or specific fractions of blood. The non-tumor sample may be tissue-matched with the tumor sample or it may be from a different tissue. For example, the non-tumor sample may be selected from a healthy (i.e., non-cancerous or non-tumor) tissue sample, blood or specific fractions of blood such as buffy coat, leukocytes, fibroblast, or any other biological sample comprising cfDNA or genomic DNA.
[0093] Once a patient-specific and tumor-specific panel (i.e., a “signature panel”) has been established, such a signature panel can be used to enrich ctDNA (e.g., fragments that include a target sequence corresponding to a tumor-specific somatic mutation or variant) in subsequent samples taken from the cancer patient. The subsequent (or “further”, “another”) samples may be taken from a patient at various time points during the course of treatment or during a period of remission. For example, after a surgical removal of a tumor, the tumor may be profiled as described herein to determine tumor-specific somatic mutations, and at one or more subsequent time points a subsequent sample may be taken from the subject to search for the presence of any ctDNA comprising any one of the identified tumor-specific somatic mutations. The detection or presence of ctDNA comprising a tumor-specific somatic mutation may be indicative of cancer recurrence. Additionally or alternatively, similar assessment can be performed throughout the course of a patient’s treatment (e.g., with chemotherapy, radiation, immunotherapy, cell therapy, etc.) to detect or quantify ctDNA and determine whether the amount of ctDNA is increasing or decreasing, as this may be indicative of responsiveness to the therapy. Accordingly, assessment of a subsequent (or “further”, “another”) sample may be repeated 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more times throughout the course of a patient’s remission or treatment. The assessment of a subsequent sample may be repeated monthly, every other month, once every three months, once every four months, once every five months, once every six months, once every seven months, once every eight months, once every nine months, once every ten months, once every eleven months, or annually. Methods of determining the source of a sample, as described herein, can-27- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666be utilized to determine if a first sample and a second and / or subsequent sample assessed in a MRD assay are from the same subject or from different subjects.
[0094] The type of sample used for the one or more subsequence samples is generally a blood sample, a plasma sample, or a serum sample, but any biological sample that contains cfDNA and potential contains ctDNA would be acceptable. In some embodiments, the one or more subsequent samples are cell-free samples.[00951 Enrichment of ctDNA (e.g., fragments that include a target sequence corresponding to a tumor-specific somatic mutation or variant) in the one or more subsequent samples can be performed by methods including, but not limited to, hybrid capture-based enrichment, PCR-target enrichment, or on-sequencer enrichment. Briefly, enrichment may comprise extracting cfDNA from a subsequent sample taken from the cancer patient and contacting the extracted cfDNA with a plurality of oligonucleotides (i.e., oligonucleotide probes), wherein each oligonucleotide in the plurality of oligonucleotides comprises a nucleic acid sequence that is capable of hybridizing to a cfDNA fragment comprising one of the tumor-specific somatic mutation sequences identified by comparing the sequences of the patients tumor DNA and non-tumor DNA. Thus, enrichment may utilize a set of oligonucleotide probes to selectively enrich ctDNA that may be in the subsequent sample by binding to previously identified tumor-specific somatic mutation sequences.
[0096] A signature panel may comprise 10-5000 tumor-specific somatic mutations. For example, a signature panel may comprise 10-4000, 10-3000, 10-2500, 10-2000, 10-1500, 10-1000, 10-950, 10-900, 10-850, 10-800, 10-750, 10-700, 10-650, 10-600, 10-550, 10-500, 50-5000, 50-4000, 50-3000, 50-2500, 50-2000, 50-1500, 50-1000, 50-950, 50-900, 50-850, 50-800, 50-750, 50-700, 50-650, 50-600, 50-550, 50-500, 100-5000, 100-4000, 100-3000, 100-2500, 100-2000, 100-1500, 100-1000, 100-950, 100-900, 100-850, 100-800, 100-750, 100-700, 100-650, 100-600, 100-550, 100-500, 200-5000, 200-4000, 200-3000, 200-2500, 200-2000, 200-1500, 200-1000, 200-950, 200-900, 200-850, 200-800, 200-750, 200-700, 200-650, 200-600, 200-550, 200-500, 300-5000, 300-4000, 300-3000, 300-2500, 300-2000, 300-1500, 300-1000, 300-950, 300-900, 300-850, 300-800, 300-750, 300-700, 300-650, 300-600, 300-550, 300-500, 400-5000, 400-4000, 400-3000, 400-2500, 400-2000, 400-1500, 400--28- 4920-4080-9613.1Atty. Dkt. No.: 131588-16661000, 400-950, 400-900, 400-850, 400-800, 400-750, 400-700, 400-650, 400-600, 400-550, 400-500, 500-5000, 500-4000, 500-3000, 500-2500, 500-2000, 500-1500, 500-1000, SOO-OSO, 500-900, 500-850, 500-800, 500-750, 500-700, 500-650, 500-600, or 500-550 tumorspecific somatic mutations. In some embodiments, a signature panel may comprise or consist of about 10, about 20, about 30, about 40, about 50, about 75, about 100, about 150, about 200, about 250, about 300, about 350, about 400, about 450, about 500, about 550, about 600, about 650, about 700, about 750, about 800, about 850, about 900, about 950, about 1000, about 1100, about 1150, about 1200, about 1250, about 1300, about 1350, about 1400, about 1450, about 1500, about 1550, about 1600, about 1650, about 1700, about 1750, about 1800, about 1850, about 1900, about 1950, or about 2000 or more tumor-specific somatic mutations. In some embodiments, a signature panel may comprise at least 10, at least 20, at least 30, at least 40, at least 50, at least 75, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1000, at least 1100, at least 1150, at least 1200, at least 1250, at least 1300, at least 1350, at least 1400, at least 1450, at least 1500, at least 1550, at least 1600, at least 1650, at least 1700, at least 1750, at least 1800, at least 1850, at least 1900, at least 1950, or at least 2000 tumor-specific somatic mutations. The tumor-specific somatic mutations may be in introns, exons, intergenic regions, or a combination thereof. In some embodiments, the tumor-specific somatic mutations may be one or more somatic mutations selected from Single Nucleotide Variants, insertions, deletions, and translocations.
[0097] After enrichment or concurrently with enrichment of ctDNA (e.g., fragments that include a target sequence corresponding to a tumor-specific somatic mutation or variant), the enriched DNA is sequenced. This sequencing may be performed by, for example Next Generation Sequencing (NGS). Deep sequencing may allow for more sensitive detection, and so the depth of the sequencing may be at least 50X, at least 100X, at least 150X, at least 200X, at least 250X, at least 300X, at least 350X, at least 400X, at least 450X, at least 500X, at least 550X, at least 600X, at least 650X, at least 700X, at least 750X, at least 800X, at least 850X, at least 900X, at least 950X, or at least 1000X. In other words, the depth of the sequencing may be about 50X, about 100X, about 150X, about 200X, about 250X, about 300X, about 350X, about 400X, about 450X, about 500X, about 550X, about 600X, about -29- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666650X, about 700X, about 750X, about 800X, about 850X, about 900X, about 950X, or about 1000X. The detection sensitivity of the disclosed MRD methods may be about 20 to about 50 ctDNA fragments comprising one or more of the set of somatic mutations in the fluid sample per a total background of about 500,000 cfDNA fragments.
[0098] MRD methods may be used for tracking and assessing recurrence in a cancer patient. For example, in some embodiments, the subject has, had, or is suspected of having a cancer including, but not limited to, bladder cancer, breast cancer, colon / colorectal cancer, gynecologic cancers, head and neck cancers, hematological cancers, liver cancer, lung cancer, and skin cancer.
[0099] In MRD assays, the obtaining and testing of subsequent samples from a cancer patient, may be repeated one or more times following completion of a cancer treatment; one or more times while the cancer patient is in remission; one or more times coinciding with or prior to surgery; following, during, or prior to administration of chemotherapy; following, during, or prior to radiation therapy; following, during, or prior to immunotherapy; or following, during, or prior to cell therapy. MRD assays may also be repeated at times prior to, coinciding with, and / or following an imaging test, such as a PET scan, a PET / CT scan, an MRI, or an X-ray. Subsequent (or “further”, “another”) samples may be evaluated relative to a first and / or second sample using methods of determining the source of a sample described herein, e.g., to determine if a first, second, and / or subsequent samples are from the same subject or are from different subjects.
[0100] MRD methods also allow for detecting ctDNA or determining the tumor fraction from a biological sample from a patient that has, previously had, or is suspected of having cancer.EXAMPLESExample 1 : Method for Identifying DNA Samples From the Same Subject Based on Low-Coverage Sequencing Data
[0101] To identify DNA samples from the same subject based on low-coverage sequencing data, multiple DNA samples are first sequenced at a matching set of common germline variant positions and then fit to a statistical model.-30- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666[01021 To prepare the low-coverage sequencing data for evaluation, data covering a set of common germline variants are provided in a file listing the variants’ coordinates, the reference, the alternate alleles, and the minor allele frequency. Binary Alignment Maps (BAMs) from each sample are analyzed to extract the count of alternative-bearing and total read pairs covering variant positions. This class requires data from paired-end sequencing and count read pairs as single molecules. Read counts can be stored as a set of Polars DataFrames and used for downstream analyses.
[0103] To complete sample matching, the function asks if two samples are more likely to be drawn from the same subject or two unrelated subjects. First, the genotype likelihood of homozygous-reference, heterozygous, and homozygous alternate is calculated using a binominal model:p(genotype=0 / 0|read_counts)=dbinom(x=n_alt_reads,n=n_total_reads,p=0+error)p(genotype=0 / l|read_counts)=dbinom(x=n_alt_reads,n=n_total_reads,p=0.5)p(genotype=l / l|read_counts)=dbinom(x=n_alt_reads,n=n_total_reads,p=l-error)[01041 Genotype likelihood is utilized rather than genotype calls to account for low-coverage sequencing which could create artifactual identifying mismatches, such as identifying SNP mismatches, in called genotypes.
[0105] For a single sample and a single variant site (“site”), the likelihood of the data is the probability of the read counts given the genotype likelihood times the probably of drawing that genotype from a pool with known alternate fraction marginalized over all genotypes (indexed by k):L(site, one sample) = \sum_k p(read_counts|genotype_k) * p(genotype_k|af)[0106| p(genotype|af) is the mendelian expectation, which is also the binominal probability mass. For example, with allele frequences p=0.5 and q=0.5, the probability of drawing one heterozygote is binom.pmf(x=l,n=2,p=0.5) = 0.5 = 2pq).-31- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666[01O7| The likelihood of two samples’ data at one variant site is the product of the persample likelihoods:L(site, two samples) = L(site,samplel) * L(site,sample2)
[0108] To ask if two samples come from the same subject, the likelihoods of observing the data from both samples if they were drawn from an individual with the same genotype versus the likelihood of observing the data for both samples if they were drawn independently from the global pool with alternate fraction of each site to the global minor allele fraction in, for example, the 1000 genomes data is calculated:L(match) = \prod_sites L(site|gt)L(random) = \prod_sites L(site|maf)
[0109] Gt corresponds to the alternate fraction in the focal individual at each identifying SNP site. While this genotype is not actually known, particularly when comparing two samples sequenced with low-coverage, both samples may have genotyping errors in the identifying SNP. To address this challenge, a reciprocal swap can be completed. To do so, genotypes are called on sample 1 (i.e., assuming the maximum likelihood genotype is true) and then the probability of observing the read counts of sample 2 is calculated if they were from an individual with the sample 1 genotypes. Subsequently, the genotypes can be called on sample 1 and the likelihood of the sample 1 genotype calls given the sample 2 genotype calls can be calculated. The likelihood of a mismatch is the mean of these two likelihoods:L(match)=(L(match|E(gt|samplel_read_counts))+L(match|E(gt|sample2_read_counts))) / 2
[0110] Then, the test statistic is a likelihood ratio for match versus random:LR=L(match) / L(random)-32- 4920-4080-9613.1Atty. Dkt. No.: 131588-1666[01111 A higher value is indicative of increased confidence in a match between two samples. Such methods can facilitate, given a large group of subjects, the ability to check all pairwise combinations of samples, or to identify the closest match to a given sample.4920-4080-9613.1
Claims
Atty. Dkt. No.: 131588-1666WHAT IS CLAIMED IS:
1. A method of determining the source of a sample, comprising:sequencing at low coverage DNA from a first sample from a subject, thereby obtaining a first set of sequencing reads;counting heterozygous and homozygous-reference-bearing sequencing reads at germline variants in the first set of sequencing reads that corresponds to a panel of germline variants comprising a plurality of common germline variants; sequencing at low coverage DNA from a second sample, thereby obtaining a second set of sequencing reads;counting heterozygous and homozygous-reference-bearing sequencing reads at germline variants from the panel of common germline variants sequenced in the second set of sequencing reads;determining (i) the first sample and the second sample are from the same subject or (ii) the first sample and the second sample are from different subjects based on a likelihood of observing the counts of sequencing reads for the panel of common germline variants for each sample.
2. The method of claim 1, wherein the likelihood of observing the counts of sequencing reads for the panel of germline variants is determined by calculating a genotype likelihood selected from (i) homozygous-reference, (ii) heterozygous, and (iii) homozygous alternate for each of the germline variants in the first set of sequencing rads and the second set of sequencing reads.
3. The method of claim 2, wherein calculating the genotype likelihood comprises modeling a statistical distribution, optionally a binomial distribution.
4. The method of claim 2 or 3, wherein a genotype likelihood of a given germline variant in the first sample or the second sample, independently, is a probability of counts of sequencing reads for a given genotype of the given germline variant times a probability of the given genotype from a pool with known alt-allele fraction, marginalized over all genotypes of the given germline variant.-34- 4920-4080-9613.1Atty. Dkt. No.: 131588-16665. The method of any one of claims 1-4, wherein calculating the genotype likelihood comprises calculating a mean of a likelihood of a given germline variant in the first sample and a likelihood of the given germline variant in the second sample.
6. The method of any one of claims 1-5, wherein the first sample and the second sample are from the same subject when the likelihood of observing the counts of sequencing reads for the panel of germline variants for each sample is above a predetermined threshold.
7. The method of any one of claims 1-6, wherein the first sample and the second sample are from different subjects when the likelihood of observing the counts of sequencing reads for the panel of germline variants for each sample is below a predetermined threshold.
8. The method of any one of claims 1-7, wherein the first sample is a sample of blood, plasma, serum, or urine.
9. The method of any one of claims 1-7, wherein the first sample is a tumor sample.
10. The method of any one of claims 1-9, wherein the second sample is a sample of blood, plasma, serum, or urine.
11. The method of any one of claims 1-10, wherein the subject was previously diagnosed with cancer.
12. The method of any one of claims 1-11, further comprising detecting in the second sample sequence reads comprising one or more tumor-specific somatic mutations that are determined for the subject prior to sequencing the second sample.-35- 4920-4080-9613.1Atty. Dkt. No.: 131588-166613. The method of claim 12, further comprising enriching the DNA from the second sample for DNA comprising one or more of the tumor-specific somatic mutations prior to sequencing.
14. The method of claim 13, wherein enriching the DNA comprising one or more of the tumor-specific somatic mutations comprises hybrid capture enrichment or PCR-based enrichment.
15. The method of any one of claims 1-14, further comprising enriching the DNA from the second sample for DNA comprising one or more of the panel of germline variants prior to sequencing.
16. The method of claim 15, wherein enriching the DNA from the second sample for DNA comprising one or more of the panel of germline variants comprises hybrid capture enrichment or PCR-based enrichment.
17. The method of any one of claims 1-16, wherein sequencing DNA from the first sample comprises whole genome sequencing, whole exome sequencing, or targeted sequencing.
18. The method of any one of claims 1-17, wherein sequencing DNA from the second sample comprises whole genome sequencing, whole exome sequencing, or targeted sequencing.
19. The method of any one of claims 1-18, wherein the subject has completed at least one cancer treatment prior to obtaining the tumor sample.
20. The method of claim 19, wherein the cancer treatment is selected from chemotherapy, radiotherapy, surgery, immunotherapy, cell therapy, or biologic therapy.-36- 4920-4080-9613.1Atty. Dkt. No.: 131588-166621. The method of any one of claims 1-20, further comprising sequencing DNA from another sample from the subject at 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more successive time points.
22. The method of claim 21, wherein sequencing DNA from another sample from the subject is repeated one or more times while the patient is in remission.
23. The method of claim 21, wherein sequencing DNA from another sample from the subject is repeated one or more times while the patient is undergoing treatment for the cancer.
24. The method of claim 21, wherein sequencing DNA from another sample from the subject is repeated one or more times coinciding with or prior to surgery; following, during, or prior to administration of chemotherapy; following, during, or prior to radiation therapy; following, during, or prior to administration of an immunotherapy; following, during, or prior to administration of a cell therapy; or following, during, or prior to administration of a biologic therapy.
25. The method of any one of claims 1-24, wherein the subject has, had or is suspected of having a cancer selected from bladder cancer, breast cancer, colon / colorectal cancer, gynecologic cancers, head and neck cancers, hematological cancers, liver cancer, lung cancer, and skin cancer.-37- 4920-4080-9613.1