Methods for analysis of circulating cells

The method addresses the challenge of analyzing small numbers of circulating cells by using multiplex amplification and next-generation sequencing to accurately identify and characterize circulating cells, enhancing non-invasive diagnostics for fetal health and cancer monitoring.

JP2025148486APending Publication Date: 2025-10-07NATERA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025117302
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-01-28
Filing Date
2025-07-11
Publication Date
2025-10-07

AI Technical Summary

Technical Problem

Existing liquid biopsy methods face challenges in accurately analyzing and characterizing small numbers of circulating cells, such as circulating fetal cells and circulating tumor cells, due to the difficulty in isolating and analyzing very small cell samples, which limits the effectiveness of non-invasive diagnostic and monitoring techniques.

Method used

A method involving multiplex amplification of single nucleotide polymorphism (SNP) loci followed by next-generation sequencing is used to generate amplicons from both cellular DNA and cell-free DNA, allowing for the determination of the origin and purity of circulating cells, even in samples containing a single cell, through genotyping and calculating allelic ratios.

Benefits of technology

Enables accurate identification and genomic analysis of single circulating cells, providing reliable non-invasive diagnostic tools for fetal health and cancer monitoring by determining cell type, DNA purity, and measuring sample quality, even in samples with very few cells.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025148486000001_ABST
    Figure 2025148486000001_ABST
Patent Text Reader

Abstract

To provide methods for characterizing and analyzing circulating cells, in particular, to provide methods for confirming the identity of an individual cell and confirming that the obtained sample is derived from a single cell of a defined identity.SOLUTION: Additional methods are provided for analyzing a single cell sample to determine copy number variation and aneuploidy in the context of circulating fetal cells, microdeletions, single nucleic acid variations associated with cancer, or early relapse of cancer and metastasis.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 62 / 780,724, filed December 17, 2018, and U.S. Provisional Patent Application No. 62 / 797,518, filed January 28, 2019, which applications are incorporated herein by reference in their entireties.

[0002] There is currently considerable interest in using liquid biopsy as an alternative to more invasive tissue biopsies for prenatal diagnosis, diagnosis of diseases and conditions, and testing and monitoring of disease onset and treatment response. For example, liquid biopsy of pregnant mothers offers a noninvasive method for fetal diagnosis and testing, avoiding the risks associated with invasive fetal or amniotic fluid biopsies. For cancer onset and treatment response diagnosis and monitoring, liquid biopsy can also avoid the need for invasive tumor tissue biopsies, which carry the risk of metastasis or surgical complications. Liquid biopsy-based methods may also be more sensitive than imaging-based methods, which are not sensitive enough to detect early-stage recurrence or metastasis.

[0003] Liquid biopsy is based on collecting blood or urine samples from patients and isolating cell-free DNA or circulating cells shed from fetuses, tumors, or other target organs. Analysis of cell-free DNA and circulating cells has provided a powerful, non-invasive diagnostic method. However, because cell-free DNA is fragmented and does not encompass the entire genome, analysis of cell-free DNA does not provide the same amount of information as analyzing whole genomic DNA directly from cells and / or tissues. Meanwhile, analysis of circulating cells is challenged by the difficulty of characterizing and analyzing very small cell samples. Therefore, to maximize the benefits of liquid biopsy-based methods, methods are needed to isolate and analyze samples containing only a few cells or a single cell. Summary of the Invention

[0004] Disclosed herein are methods for analyzing and characterizing circulating cell samples that may contain only a small number of circulating cells or a single circulating cell. According to aspects exemplified herein, in one embodiment, a method for analyzing and characterizing circulating cells suspected of fetal origin derived from a blood sample or cell line of a mother carrying a fetus includes the steps of: a) generating a first set of amplicons of cellular DNA isolated from one or more circulating cells suspected of fetal origin obtained from the blood sample or cell line, and a second set of amplicons of cell-free DNA obtained from the plasma fraction of the blood sample or a mixture of DNA from the child's cell line and the mother's cell line, wherein the first set of amplicons and the second set of amplicons are obtained by performing a multiplex amplification reaction of a plurality of single nucleotide polymorphism (SNP) loci; b) genotyping the first set of amplicons and the second set of amplicons by next-generation sequencing; and c) determining, based on the genotypes of the first set of amplicons and the second set of amplicons, that the one or more circulating cells are one or more maternal cells, one or more fetal cells, or a mixture of maternal and fetal cells.

[0005] In one embodiment, the method disclosed herein further includes generating a third set of maternal cell DNA isolated from one or more maternal cells obtained from the buffy coat fraction of the blood sample, and genotyping the third set of amplicons by next-generation sequencing to determine the maternal genotype.

[0006] In one embodiment, the method disclosed herein further comprises obtaining from the blood sample circulating cells suspected to be of fetal origin, a plasma fraction containing cell-free DNA, and a buffy coat fraction containing maternal cells.

[0007] In one embodiment, determining that the one or more circulating cells are one or more maternal cells, one or more fetal cells, or a mixture of maternal and fetal cells includes calculating maternal and fetal matches between cell-free DNA obtained from the plasma fraction (or cell-free DNA known to be a mixture of child and maternal DNA) and cellular DNA obtained from circulating cells suspected of fetal origin (or a cell line from the child), and calculating the fetal mixture fraction (or child fraction).

[0008] In one embodiment, the first set of amplicons and the second set of amplicons are obtained by performing a multiplex amplification reaction of at least 1,000 SNP loci.

[0009] In one embodiment, the methods disclosed herein further comprise estimating the fetal fraction of a mixture of maternal and fetal cells (or a mixture of maternal and fetal cell lines) by measuring the abundance of SNP loci.

[0010] In one embodiment, the methods disclosed herein further comprise determining the allele ratio at the SNP.

[0011] In one embodiment, the estimation of fetal fraction uses allelic ratios.

[0012] In one embodiment, the method disclosed herein includes dividing a plurality of circulating cells into individual reaction volumes, isolating cellular DNA in each reaction volume, and attaching a sample barcode to the isolated cellular DNA in each reaction volume.

[0013] In one embodiment, cellular DNA is isolated from a single circulating cell suspected to be a fetal cell.

[0014] In one embodiment, the method further comprises performing a non-invasive prenatal test using the genotype determined to be one or more fetal cells.

[0015] In one embodiment, the method further comprises detecting copy number variations or aneuploidy of a target chromosome or target chromosome segment of interest in one or more circulating cells determined to be one or more fetal cells.

[0016] In one embodiment, the target chromosome or target chromosome segment of interest is chromosome 13, chromosome 18, chromosome 21, a sex chromosome, and / or a chromosome segment thereof.

[0017] In one embodiment, the methods disclosed herein further comprise detecting a microdeletion in one or more circulating cells determined to be one or more pure fetal cells.

[0018] In one embodiment, the microdeletion is a 22q11.2 deletion associated with DiGeorge syndrome, a microdeletion associated with Prader-Willi syndrome, a microdeletion associated with Angelman syndrome, a 1p36 deletion, and / or a microdeletion associated with Criss-Cat syndrome.

[0019] In another exemplary embodiment, the present disclosure provides a method for monitoring and detecting early recurrence or metastasis of a tumor in a cancer patient by obtaining one or more circulating cells determined to be derived from the tumor, the method comprising: a) selecting one or more patient-specific mutations based on mutations identified in tumor samples of patients diagnosed with cancer; b) longitudinally collecting one or more blood samples from the patient after the patient has been treated with surgery, first-line chemotherapy, and / or adjuvant therapy; c) generating a first set of amplicons from cellular DNA isolated from one or more circulating cells suspected of being circulating tumor cells and a second set of amplicons from cell-free DNA obtained from the blood sample obtained from the cancer patient, wherein the tumor cells and normal cells are obtained from each of the cancer patient's blood sample or fractions thereof, and the first set of amplicons and the second set of amplicons are obtained by performing a multiplex amplification reaction of patient-specific mutations associated with cancer; d) sequencing the first set of amplicons and the second set of amplicons by next generation sequencing; e) determining the origin of one or more circulating cells based on the sequences of the first amplicon set, wherein the sequences of the second amplicon set are used as a reference, and detection of one or more patient-specific mutations from the first amplicon set generated from circulating cells determined to be of tumor origin is indicative of early recurrence or metastasis of the cancer.

[0020] In another aspect, the disclosure herein provides a method for treating a cancer patient, the method comprising: a) treating a cancer patient with surgery, first-line chemotherapy, and / or adjuvant therapy; b) longitudinally collecting one or more blood samples from the patient after the patient has been treated with surgery, first-line chemotherapy, and / or adjuvant therapy; c) generating a first set of amplicons from cellular DNA isolated from one or more circulating cells suspected of being circulating tumor cells and a second set of amplicons from cell-free DNA obtained from the blood sample obtained from the cancer patient, wherein the tumor cells and normal cells are obtained from each of the cancer patient's blood sample or fractions thereof, and the first set of amplicons and the second set of amplicons are obtained by performing a multiplex amplification reaction of patient-specific mutations associated with cancer; d) sequencing the first set of amplicons and the second set of amplicons by next generation sequencing; e) determining the origin of one or more circulating cells based on the sequences of the first amplicon set, wherein the sequences of the second amplicon set are used as a reference, and detection of one or more patient-specific mutations from the first amplicon set generated from circulating cells determined to be of tumor origin is indicative of early recurrence or metastasis of the cancer; f) administering to said patient a compound, said compound being found to be effective in treating a cancer having said one or more patient-specific mutations detected in a blood sample.

[0021] The improved methods provided herein allow for accurate analysis and characterization of samples containing very small numbers of CTCs. The number of CTCs in a sample analyzed by the methods of the present invention can be less than 100 cells, less than 75 cells, less than 50 cells, less than 25 cells, less than 20 cells, less than 15 cells, less than 10 cells, less than 5 cells, or even one cell. Using the methods provided herein, accurate genomic information can be obtained from one circulating tumor cell, two CTCs, three CTCs, four CTCs, five CTCs, six CTCs, seven CTCs, eight CTCs, nine CTCs, or ten CTCs.

[0022] In some embodiments, the method includes monitoring and detecting early tumor recurrence or metastasis in a cancer patient by determining the identity of circulating cells suspected to be circulating tumor cells based on predetermined patient-specific mutations. In some embodiments, the patient-specific mutations include cancer-associated single nucleotide variants (SNVs), copy number variants (CNVs), indels, deletions, or gene fusions. In some embodiments, the presence of at least two cancer-associated patient-specific mutations indicates that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least eight cancer-associated patient-specific mutations indicates that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least 16 cancer-associated patient-specific mutations indicates that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least 50 cancer-associated patient-specific mutations indicates that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least 100 cancer-associated patient-specific mutations indicates that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least 100 to about 1000 patient-specific mutations associated with cancer indicates that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least 100 to about 1000 patient-specific mutations associated with cancer indicates that the one or more circulating cells are tumor cells.

[0023] In one aspect, multiplex amplification is performed to obtain amplicons containing patient-specific mutations associated with cancer. In some embodiments, the first set of amplicons and the second set of amplicons are obtained by performing a multiplex amplification reaction of 50 to 1000 cancer-associated patient-specific mutations.

[0024] In one aspect, a method for monitoring and detecting early tumor recurrence and metastasis in cancer patients includes separating a plurality of circulating cells into individual reaction volumes, isolating cellular DNA in each reaction volume, and attaching a sample barcode to the isolated cellular DNA in each reaction volume. In some embodiments, the cellular DNA is isolated from a single circulating cell suspected to be a tumor cell.

[0025] In some embodiments, the method further comprises performing a non-invasive cancer test using the genotype in which the one or more circulating cells are determined to be one or more tumor cells.

[0026] In another illustrative aspect, in one embodiment, a method for analyzing and characterizing circulating cells suspected of being donor cells derived from a blood sample of a transplant recipient, the method comprising: a) generating a first set of amplicons of cellular DNA isolated from one or more circulating cells suspected of being donor cells obtained from the blood sample and a second set of amplicons of cell-free DNA obtained from the plasma fraction of the blood sample, the first set of amplicons and the second set of amplicons being obtained by performing a multiplex amplification reaction of a plurality of single nucleotide polymorphism (SNP) loci; b) genotyping the first set of amplicons and the second set of amplicons by next-generation sequencing; and c) determining, based on the genotyping of the first set of amplicons and the second set of amplicons, that the one or more circulating cells are one or more donor cells, one or more recipient cells, or a mixture of donor and recipient cells. [Brief explanation of the drawings]

[0027] [Figure 1] Figure 1 shows the workflow of an exemplary method for acquiring and analyzing circulating cells.

[0028] [Figure 2]FIG. 2 shows the workflow of an exemplary method for confirming the identity of individual circulating cells.

[0029] [Figure 3] Figure 3 shows hetrate plots of the genetic data obtained in Example 2. Figure 3A is a hetrate plot for sample 1F, which contains one fetal cell. Figure 3B is a hetrate plot for sample 2F, which contains two fetal cells. Figure 3C is a hetrate plot for sample 4F, which contains four fetal cells. Figure 3D is a hetrate plot for sample 10F, which contains ten fetal cells. Figure 3E is a hetrate plot for sample 20F, which contains twenty fetal cells.

[0030] [Figure 4-1] Figure 4 shows hetrate plots of the genetic data obtained in Example 3. Figure 4A is a hetrate plot for sample 10F, which contains 10 fetal cells. Figure 4B is a hetrate plot for sample 9F1M, which contains 9 fetal cells and 1 maternal cell. Figure 4C is a hetrate plot for sample 7F3M, which contains 7 fetal cells and 3 maternal cells. Figure 4D is a hetrate plot for sample 5F5M, which contains 5 fetal cells and 5 maternal cells. Figure 4E is a hetrate plot for sample 3F7M, which contains 3 fetal cells and 7 maternal cells. Figure 4F is a hetrate plot for sample 1F4M, which contains 1 fetal cell and 4 maternal cells. Figure 4G is a hetrate plot for sample 1F9M, which contains 1 fetal cell and 9 maternal cells. Figure 4H is a hetrate plot for sample 1F19M, which contains 1 fetal cell and 19 maternal cells. Figure 4I is a hetrate plot for Sample 10M containing 10 maternal cells. Figure 4J is a hetrate plot for Sample 10% CF mixture containing 10% fetal DNA and 90% maternal DNA. [Figure 4-2] Same as above.

[0031] [Figure 5]FIG. 5 shows an exemplary workflow and conditions for a multiplex PCR amplification reaction according to Example 1.

[0032] The presently disclosed embodiments will be further described with reference to the accompanying drawings, in which like structure is represented by like numerals throughout the several views. The drawings shown are not necessarily to scale, emphasis instead being placed on illustrating the principles of the presently disclosed embodiments.

[0033] The above-identified figures are provided by way of illustration and not by way of limitation. DETAILED DESCRIPTION OF THE INVENTION

[0034] 1. Overview The present disclosure provides methods and compositions for the characterization and genomic analysis of circulating cells. In particular, the present disclosure provides methods for confirming the identity of an isolated single circulating cell and analyzing genomic data from a single cell, which are unprecedented in the literature.

[0035] The isolation and genomic analysis of single cells has attracted considerable interest from clinical and industrial researchers. Single-cell genomics provides genetic information at the cellular level. Genotyping profiles and genomic information obtained from isolated single cells have many potential applications in clinical diagnostics, risk prediction, drug development, and disease treatment. However, the development of methods for obtaining genomic information from single cells has been hampered by the significant challenge of analyzing data from samples containing only a small number of cells, let alone single-cell samples. The disclosure herein provides both biochemical and analytical methods for accurately obtaining and analyzing genomic information from single-cell samples.

[0036] A method for accurately obtaining and analyzing genomic information, especially from single cells, would have a significant impact on liquid biopsy-based diagnostics. Liquid biopsy refers to the collection of a non-solid biological tissue sample, most commonly blood, but also saliva, urine, cerebrospinal fluid, and other bodily fluids. For example, a liquid biopsy from a pregnant mother could be used to obtain a sample containing both cell-free DNA (cfDNA) and fetal circulating cells, enabling non-invasive fetal diagnosis. Intact circulating fetal cells (CFCs) can enter the maternal circulation. Therefore, analysis of fetal genetic material could provide non-invasive prenatal genetic diagnosis (NIPGD or NPD). However, CFCs are highly diluted among the billions of maternal blood cells. Another example is the use of liquid biopsies in cancer patients to diagnose or monitor tumors without biopsying the tumor itself. Liquid biopsies of cancer patients allow the isolation of circulating tumor cells (CTCs) and circulating tumor DNA (ctDNA) for tumor diagnosis and monitoring, as well as for examining tumor response to treatment. Although CTCs are shed from primary and metastatic lesions into the bloodstream of cancer patients, CTCs are relatively rare among the patient's blood, even in patients with advanced metastatic disease. Accurate tumor diagnosis based on liquid biopsies therefore relies on the accurate isolation and analysis of cell samples containing only a few cells or a single cell. Liquid biopsies may also be used to isolate and analyze circulating cells originating from transplanted organs for transplant organ monitoring.

[0037] Therefore, the full potential of liquid biopsy-based diagnostics depends on the ability of diagnostic methods to accurately isolate and analyze CRCs, such as circulating fetal cells (CFCs) and circulating tumor cells (CTCs). For example, when CFCs are used to diagnose a pregnant mother's fetus, analysis of a single circulating cell from the pregnant mother's blood sample relies on the isolated single circulating cell being identified as fetal in origin and providing useful information. Similarly, analysis of a single circulating cell from a cancer patient relies on the isolated single circulating cell being identified as tumor in origin and providing useful information. Additionally, diagnostics based on circulating rare cells (CRCs) also rely on the ability to analyze samples containing only a small number of cells, or even only a single cell.

[0038] The present disclosure fulfills this need by providing a method for accurately analyzing isolated single circulating cells and characterizing them qualitatively and quantitatively. The present disclosure provides a SNP-based next-generation sequencing (NGS) method for analyzing isolated single circulating cell samples, which can identify cell type, estimate DNA purity, and measure sample quality. For example, the disclosed method may be used to confirm that isolated cells suspected to be fetal cells are pure single fetal cells, pure maternal cells, or a mixture of fetal and maternal cells. In another embodiment, the disclosed method may be used to confirm that an isolated cell sample suspected to be circulating tumor cells is either an isolated single circulating tumor cell or normal cells derived from a patient's blood.

[0039] Figure 1 is an overview of the workflow of an exemplary embodiment of the present disclosure for the analysis of circulating fetal cells isolated from a maternal blood sample. The workflow outlined in Figure 1 may therefore be adapted to achieve the analysis of circulating tumor cells or any other circulating rare cells isolated from a subject. In one exemplary embodiment, the method of the present disclosure outlined in Figure 1 and described in further detail below comprises: 1. collecting a blood sample from a mother carrying a fetus; 2. Isolating a plasma fraction and / or a buffy coat fraction from the blood sample; 3. Isolating fetal and maternal cells from the blood sample; 4. (a) extracting cell-free DNA (cfDNA) from the plasma fraction, and optionally (b) extracting genomic DNA from the buffy coat fraction; 5. Lysing the isolated (a) fetal cells and (b) maternal cells; 6. Preparing a cfDNA library for sequencing; 7. Performing massively multiplexed PCR (mmPCR) (a) directly on the cell lysate, (b) on the maternal DNA, and (c) on the cfDNA library; 8. Adding DNA barcodes to the amplified DNA samples to label each sample; 9. Preparing a DNA sample pool for sequencing; 10. Running the sample through a next generation sequencing (NGS) protocol; 11. (a) preprocessing the NGS data; and (b) analyzing the NGS data; and 12. Interpreting the results of the NGS data processing and analysis.

[0040] The results of the above-described method can provide information about the match between cfDNA, fetal DNA (cell lysate), and maternal DNA, allowing for the determination of the purity of a cell sample. As shown in Figure 2 and the Examples herein, the disclosed method can be used to confirm that an analyzed DNA sample originates from a single circulating cell. The Examples provided herein demonstrate that, following the workflow outlined in Figures 1 and 2 and using cfDNA as a reference, the method can determine the origin of cellular DNA from an unknown sample and the composition of the sample. For example, in the case of fetal / maternal cells isolated from the blood of a pregnant woman, the method can determine whether the sample is pure fetal cells or pure maternal cells (Figure 2 and Examples).

[0041] Additionally, the data can reveal chromosomal abnormalities in the circulating cells. If the sample is determined to be a mixture of maternal and fetal cells, the fetal fraction can be determined. In another aspect, the methods provided herein can be used to determine the purity of a circulating tumor cell sample, or to analyze the genomic information of the circulating tumor cells. In some embodiments, the methods provided herein allow for the determination that circulating cells suspected to be derived from a transplanted organ are pure, isolated cells originating from the transplanted organ.

[0042] 2. Analyzing a single cell to confirm its identity The present disclosure provides data analysis methods that can analyze genomic information from samples of very few cells, even single cells. In particular, the present disclosure provides methods that can determine the identity of a single circulating cell isolated from a blood sample and analyze the genomic information of the single circulating cell. For example, as further described below and in the Examples herein, the present disclosure shows how to confirm the identity of circulating fetal cells derived from a blood sample of a pregnant mother, and then analyze the genomic information of the confirmed pure fetal cell sample. The methods described below can be adapted to confirm the identity of any cell sample by using corresponding cell-free DNA as a reference.

[0043] According to aspects exemplified herein, in one embodiment, a method for analyzing and characterizing circulating cells suspected of fetal origin from a blood sample of a mother carrying a fetus includes: a) generating a first set of amplicons of cellular DNA isolated from one or more cell lines or one or more circulating cells suspected of fetal origin obtained from the blood sample, and a second set of amplicons of cell-free DNA obtained from the plasma fraction of the blood sample, or a DNA mixture obtained from the child's DNA and the mother's DNA, wherein the first set of amplicons and the second set of amplicons are obtained by performing a multiplex amplification reaction of a plurality of single nucleotide polymorphism (SNP) loci; b) genotyping the first set of amplicons and the second set of amplicons by next-generation sequencing; and c) determining that the one or more circulating cells are one or more maternal cells, one or more fetal cells, or a mixture of maternal and fetal cells based on the genotyping of the first set of amplicons and the second set of amplicons. In one embodiment, the method disclosed herein further comprises generating a third set of maternal cell DNA isolated from one or more maternal cells obtained from the buffy coat fraction of the blood sample or isolated from a maternal cell line, and genotyping the third set of amplicons by next-generation sequencing to determine the maternal genotype.

[0044] In one embodiment, determining that the one or more circulating cells are one or more maternal cells, one or more fetal cells, or a mixture of maternal and fetal cells includes calculating maternal and fetal matches between cell-free DNA obtained from the plasma fraction or a mixture of DNA from the child and the mother and cellular DNA obtained from circulating cells suspected of fetal origin, and calculating the fetal mix fraction.

[0045] The first step to calculate all the following indices is to analyze the allele ratio (AR) of the cfDNA sample and classify them as maternal genotype / fetal genotype = RR / RR, RR / RM, RM / *, MM / MM, or MM / MR. The allele ratio is calculated by the following method: reference counts / total counts for a given SNP. In some embodiments, the basic genotyper for cfDNA classifies SNPs by AR. The calculation algorithm can be adjusted for various applications. In an exemplary but non-limiting embodiment, the basic genotyper for cfDNA classifies SNPs by AR in the following manner:

number

[0046] Alternatively, in some embodiments, the genotype of the cfDNA may be classified by using a PANORAMA® genotyping device.

[0047] As a second step, the SNPs in the unknown sample are classified as homozygous-R (HoR), homozygous-M (HoM), or heterozygous (Het) based on the allelic ratio (AR). The algorithm for this calculation can be adjusted for various applications. In an exemplary, but non-limiting embodiment, the SNPs in the unknown sample are classified as homozygous-R (HoR), homozygous-M (HoM), or heterozygous (Het) based on the allelic ratio (AR) in the following manner:

number

[0048] As a third step, by using the genotyping / classification of these SNPs, the following indices were calculated that can be used to classify unknown samples:

[0049] Maternal concordance = the proportion of homozygous SNPs (RR / RR or MM / MM) in the cfDNA sample that have allele ratios (HoR and HoM, respectively) that are within the expected range in the unknown sample.

number

[0050] If the maternal match is high (greater than approximately 85%), the unknown sample is likely to have come from the same mother as the cfDNA sample.

[0051] Fetal concordance = the proportion of fetal band SNPs (RR / RM or MM / MR) in the cfDNA sample that have an allele ratio (Het) within the expected range in the unknown sample.

number

[0052] If the fetal match is high (greater than about 85%), then the unknown sample likely contains at least some fetal DNA.

[0053] Fractional Fetal Admixture = Estimate of the contribution of fetal DNA in an unknown sample. It is 1 minus the RR / RM and MM / MR separation of the fetal bands in the unknown sample. If the sample is pure maternal cells, the separation will be close to 1 and the fractional fetal admixture will be 0. If the sample is pure fetal cells, the separation will be close to 0 and the fractional fetal admixture will be close to 1. Admixture will be a number between these two values.

number

[0054] In one exemplary embodiment, the unknown sample is classified according to Table 1 below by using the calculated maternal concordance, fetal concordance, and fetal admixture percentage. According to the exemplary embodiment shown in Table 1, a pure fetal sample is defined as having a fetal admixture percentage greater than 95% and fetal and maternal concordance greater than 85%. In some embodiments, a sample is a pure fetal sample if the fetal admixture percentage is greater than 75%, 80%, 85%, 90%, 95%, or 98% and the maternal and fetal concordance are greater than 75%, 80%, 85%, 90%, or 95%. In one specific embodiment, a sample is a pure fetal sample if the fetal admixture percentage is greater than 95% and the fetal concordance is greater than 85%. A sample may be a pure maternal sample if the maternal concordance is greater than 75%, 80%, 85%, 90%, or 95% and the fetal admixture percentage is less than 15%, 10%, 5%, or 2%. In one particular embodiment, a sample is pure maternal when the maternal concordance is greater than 85% and the fetal admixture is less than 5%. Note that certain combinations (e.g., 100% fetal concordance and 0% fetal admixture) are impossible or highly unlikely (e.g., high fetal concordance and low maternal concordance) and are therefore excluded from the classification table.

[0055] [Table 1]

[0056] This method has been validated in the Examples herein for the Pano® V2 target set for chromosomes 13, 18, and 21, and is expected to be useful for any SNP target set (provided it is of sufficient size and minor allele frequency). For example, the SPECTRUM® target set may be a useful genotyping device for the disclosed methods.

[0057] 3. Isolation of Circulating Cells Circulating fetal cells (CFCs), circulating tumor cells (CTCs), and circulating cells from transplanted organs are relatively rare in liquid biopsy samples. For example, the concentration of fetal cells in maternal blood depends on the stage of pregnancy and the condition of the fetus, but estimates range from 1 to 40 fetal cells per milliliter of maternal blood, or less than 1 fetal cell per 100,000 maternal nucleated cells. While current technology can isolate small amounts of fetal cells from maternal blood, it is difficult to enrich them to an arbitrary level of purity.

[0058] A number of methods have been developed to enrich for circulating rare cells, which are based on properties that distinguish circulating tumor cells from surrounding normal cells, or circulating fetal cells from maternal cells.

[0059] Such distinguishing characteristics may be physical properties such as size, density, charge, and deformability, allowing enrichment based on the use of methods such as filtration through specialized filters, microscopy combined with micropipetting, microfluidic systems, centrifugation, electrophoresis, or flow cytometry.

[0060] Alternatively, the distinguishing characteristic may be a biological property, such as cell surface protein expression or viability, that can label and distinguish circulating cells from normal or maternal cells. For example, antibodies may be used to detect specific cell surface markers expressed by circulating cells but not by surrounding cells, and the circulating cells may then be isolated by flow cytometry, magnetic beads, or placement on a solid support coated with an agent that binds the antibody, thereby isolating cells with the bound antibody. For example, circulating tumor cells have been isolated using antibodies against epithelial cell adhesion molecule (EpCAM). See Alix-Panabieres and Pantel, Clin. Chem., 59(1):110-118 (2013). Alternatively, circulating cells may be enriched by negative selection using antibodies that recognize antigens present on normal cells but not on circulating cells, followed by removal of antibody-labeled cells. Another negative selection method is selective lysis of red blood cells.

[0061] 4. Patient population As used herein, "subject," "patient," or "individual" refers to any subject, patient, or individual, and the terms are used interchangeably herein. In this regard, "subject," "patient," and "individual" include mammals, and particularly humans. When used in conjunction with "in need thereof," the term "subject," "patient," or "individual" contemplates any subject, patient, or individual who has or is at risk for a particular condition or disorder.

[0062] The methods and compositions of the present disclosure can be used to diagnose, monitor, or treat a variety of subjects, including humans and non-human animals, including mammals, and immature and mature animals, including human children and adults, suffering from a variety of conditions or disorders. The human subject being diagnosed or treated can be an embryo, fetus, infant, child, school-age child, teenager, adolescent, adult, or geriatric patient.

[0063] In some embodiments, the subject being diagnosed or monitored by isolating and analyzing circulating fetal cells obtained from a maternal blood sample is a fetus. In some embodiments, the methods disclosed herein may be used to determine the paternity of a fetus. In some embodiments, the methods disclosed herein may be used to determine copy number variations in a chromosome or chromosome segment of interest in a fetus of a pregnant mother. In some embodiments, the methods disclosed herein may be used to determine copy number variations in chromosomes 13, 18, 21, sex chromosomes, and / or chromosome segments thereof. In some embodiments, the methods disclosed herein may be used to detect microdeletions in fetal chromosomes. In some embodiments, the microdeletions detected in the fetus are the 22q11.2 deletion associated with DiGeorge syndrome, the microdeletion associated with Prader-Willi syndrome, the microdeletion associated with Angelman syndrome, the 1p36 deletion, and the microdeletion associated with Criss-O-Cat syndrome. In some embodiments, the methods disclosed herein may be used to detect single nucleotide variations in a fetus of a pregnant mother.

[0064] In some embodiments, the methods disclosed herein may be used to diagnose or monitor cancer patients. In some embodiments, the methods disclosed herein may be used to detect copy number variations in tumor tissue by isolating and detecting circulating tumor cells from a patient's blood sample. In some embodiments, the methods disclosed herein may be used to detect one or more cancer-associated patient-specific mutations in a cancer patient. In some embodiments, the methods disclosed herein may be used to detect at least two cancer-associated patient-specific mutations in a cancer patient's tumor. In some embodiments, the methods disclosed herein may be used to detect mutations associated with early tumor recurrence or metastasis in a cancer patient. In some embodiments, the methods disclosed herein may be used to monitor a cancer patient's response to treatment.

[0065] In some embodiments, the methods disclosed herein may be used to monitor transplanted organ recipients by isolating and analyzing circulating cells originating from the transplanted organ. In some embodiments, the methods disclosed herein may be used to diagnose graft-versus-host disease in transplanted organ recipients. In some embodiments, the methods disclosed herein may be used to monitor the response of organ transplant patients to immunosuppressive therapy.

[0066] 5. Obtaining Genotype Data 5.1 Sources of genetic data Any relevant individual genetic data can be obtained from: bulk diploid tissue of an individual, one or more diploid cells from an individual, one or more haploid cells from an individual, one or more blastomeres from a target individual, extracellular genetic material present in an individual, extracellular genetic material from an individual present in the mother's blood or the blood of a cancer patient, cells from an individual present in the mother's blood, one or more embryos generated from the gametes of a related individual, one or more blastomeres removed from the embryo, extracellular genetic material present in a related individual, genetic material known to originate from a related individual, and combinations thereof. In exemplary embodiments, the methods provided herein are used to analyze free DNA originating from the genome of a target sample from target cells, such as fetal cells or tumor cells. In one embodiment, the methods disclosed herein further comprise obtaining, from a blood sample, circulating cells suspected of fetal origin, a plasma fraction containing cell-free DNA, and a buffy coat fraction containing maternal cells. In one embodiment, the methods disclosed herein further comprise obtaining, from a blood sample, circulating cells suspected of being tumor cells, a plasma fraction containing cell-free DNA, and a buffy coat fraction containing normal cells.

[0067] As used herein, the term "cell-free DNA" refers to DNA available for analysis without the need for cell lysis. Cell-free DNA can be found in blood or other bodily fluids. Cell-free DNA can be obtained from various tissues, including tissues in liquid form, such as blood, lymph, ascites, and cerebrospinal fluid. Cell-free DNA can be derived from various cell sources. In some instances, cell-free DNA is composed of DNA derived from fetal cells. Cell-free DNA can be a mixture of DNA derived from target and non-target cells. For DNA analysis of fetal aneuploidy, in some embodiments, cell-free DNA can be obtained from the blood of a pregnant woman, in which case the cell-free DNA includes a mixture of maternal and fetal cell-free DNA. In other embodiments, cell-free DNA can be derived from cancerous tumor cells. Cell-free DNA can include a mixture of tumor cell-derived DNA and non-tumor cell-free DNA from other parts of the body.

[0068] In one illustrative embodiment, the sample analyzed by the method of the present invention is a blood sample or a fraction thereof. In one embodiment, the method provided herein is particularly suitable for amplifying DNA fragments, particularly tumor DNA fragments found in circulating tumor DNA (ctDNA). Such fragments are typically about 160 nucleotides in length.

[0069] It is known in the art that cell-free nucleic acids (cfNA), such as cfNA, can be released into the circulation through various forms of cell death, including apoptosis, necrosis, autophagy, and necroptosis. cfDNA is fragmented, with fragment size distributions ranging from 150-350 bp to over 10,000 bp (see Kalnina et al. World J Gastroenterol. 2015 Nov 7;21(41):11636-11653). For example, the size distribution of plasma DNA fragments from hepatocellular carcinoma (HCC) patients ranges from 100-220 bp in length with a peak count frequency of approximately 166 bp, and tumor DNA concentrations are highest in fragments 150-180 bp in length (see Jiang et al. Proc Natl Acad Sci USA 112:E1317-E1325).

[0070] In an illustrative embodiment, circulating tumor DNA (ctDNA) was isolated from blood using EDTA-2Na tubes after centrifugation to remove cellular debris and platelets. Plasma samples can be stored at -80°C until DNA extraction using, for example, a QIAamp DNA Mini Kit (Qiagen, Hilden, Germany) (see, e.g., Hamakawa et al., Br J Cancer. 2015;112:352-356). reported that the median concentration of extracted cell-free DNA across all samples was 43.1 ng / ml of plasma (ranging from 9.5 to 1338 ng / ml), with mutant fragments ranging from 0.001 to 77.8%, with a median of 0.90%.

[0071] Genetic data, such as DNA sequence data, can be obtained from a mixture of DNA containing DNA derived from one or more target cells and DNA derived from one or more non-target cells. The target cells and non-target cells differ from each other at the genomic level based on other criteria. The term "derived" is used to indicate the cell is the ultimate source of the DNA. Thus, for example, cell-free DNA obtained from the maternal blood of a pregnant woman is derived from cells derived from the fetal placenta, which are typically genetically identical to the fetus's own and the mother's cells. The method employs a combination of patients.

[0072] 5.2 Amplification (e.g., PCR) reaction mixture: In some embodiments, obtaining genotype data involves sequencing (e.g., shotgun sequencing, single molecule sequencing, or other next-generation sequencing techniques), SNP arrays for detecting polymorphic loci, or multiplex PCR. In some embodiments, genotyping involves the use of SNP arrays to detect polymorphic loci, such as at least 100, 200, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 20,000, 25,000, 30,000, 40,000, 50,000, 75,000, or 100,000 different polymorphic loci. In some embodiments, genotyping involves the use of multiplex PCR. In some embodiments, the method involves contacting a sample in a fraction with a library of primers that simultaneously hybridize to at least 100, 200, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 13,000, 15,000, 20,000, 25,000, 30,000, 40,000, 50,000, 75,000, or 100,000 different polymorphic loci (e.g., SNPs) to generate a reaction mixture, subjecting the reaction mixture to primer extension reaction conditions to generate amplification products, which are measured using a high-throughput sequencer to generate sequence data. In some embodiments, RNA (e.g., mRNA) is sequenced. Because mRNA contains only exons, sequencing mRNA allows alleles to be determined for polymorphic loci (e.g., SNPs) over long distances in the genome, e.g., several megabases.

[0073] In one embodiment, multiplexed target PCR is performed directly on lysed cell samples. In some embodiments, cells are lysed by Protease K digestion. In some embodiments, cells are lysed by Protease K digestion in saline containing potassium chloride (KCl), magnesium chloride (MgCl), and Tris-hydrochloride. In some embodiments, to prepare cells for PCR amplification, cells are lysed by Protease K digestion in saline without potassium chloride (KCl), magnesium chloride (MgCl), and Tris-hydrochloride. See the Examples section for further details.

[0074] In some embodiments, the sample barcode is added to the DNA during the amplicon generation step.

[0075] In some embodiments, the methods of the invention include forming an amplification reaction mixture. The reaction mixture is typically formed by combining a polymerase, nucleotide triphosphates, nucleic acid fragments from a nucleic acid library generated from a sample, a set of target-specific forward outer primers, and a first-strand reverse outer universal primer. In another illustrative embodiment, the reaction mixture includes a target-specific forward inner primer in place of the target-specific forward outer primer, and amplicons from a first PCR reaction using the outer primer in place of a nucleic acid fragment from the nucleic acid library. The reaction mixtures provided herein, in illustrative embodiments, are each formed in an independent aspect of the invention. In an illustrative embodiment, the reaction mixture is a PCR reaction mixture. The PCR reaction mixture typically includes magnesium.

[0076] In some embodiments, the reaction mixture includes ethylenediaminetetraacetic acid (EDTA), magnesium, tetramethylammonium chloride (TMAC), or a combination thereof. In some embodiments, the concentration of TMAC is between 20 and 70 mM, inclusive. Without intending to be bound by any particular theory, it has been found that TMAC binds to DNA, stabilizes the duplex, increases primer specificity, and / or equalizes the melting temperatures of various primers. In some embodiments, TMAC increases the uniformity of the amount of amplification product for various targets. In some embodiments, the concentration of magnesium (e.g., from magnesium chloride) is between 1 and 8 mM.

[0077] A large number of primers for multiplex PCR of a large number of targets can chelate a large amount of magnesium (two phosphates in the primer chelate one magnesium). For example, if enough primers are used so that the phosphate concentration from the primers is about 9 mM, the primers can reduce the effective magnesium concentration by about 4.5 mM. In some embodiments, EDTA is used to reduce the amount of magnesium available as a cofactor for polymerase, because high concentrations of magnesium can cause PCR errors, such as amplification of non-target loci. In some embodiments, the concentration of EDTA reduces the amount of available magnesium to between 1 and 5 mM (e.g., between 3 and 5 mM).

[0078] In some embodiments, the pH is between 7.5 and 8.5, such as between 7.5 and 8, 8 and 8.3, or 8.3 and 8.5 (inclusive). In some embodiments, a Tris buffer is used at a concentration between 10 and 100 mM, such as between 10 and 25 mM, 25 and 50 mM, 50 and 75 mM, or 25 and 75 mM (inclusive). In some embodiments, all of these concentrations of Tris buffer are used at a pH between 7.5 and 8.5. In some embodiments, a combination of KCl and (NH)SO is used, such as between 50 and 150 mM KCl and between 10 and 90 mM (NH)SO, inclusive. In some embodiments, the KCl concentration is between 0 and 30 mM, between 50 and 100 mM, or between 100 and 150 mM (inclusive). In some embodiments, the concentration of (NH4)2SO4 is between 10 and 50 mM, 50 and 90 mM, 10 and 20 mM, 20 and 40 mM, 40 and 60 mM, or 60 and 80 mM (NH4)2SO4, inclusive. In some embodiments, the concentration of ammonium [NH4+] is between 0 and 160 mM, such as between 0 and 50, 50 and 100, or 100 and 160 mM, inclusive. In some embodiments, the concentration of the sum of potassium and ammonium ([K+]+[NH4+]) is between 0 and 160 mM, such as between 0 and 25, 25 and 50, 50 and 150, 50 and 75, 75 and 100, 100 and 125, or 125 and 160 mM, inclusive. An exemplary buffer with 120 mM [K+]+[NH4+] contains 20 mM KCl and 50 mM (NH4)2SO4. In some embodiments, the buffer comprises 25 to 75 mM Tris buffer having a pH of 7.2 to 8, 0 to 50 mM KCl, 10 to 80 mM ammonium sulfate, and 3 to 6 mM magnesium, inclusive.In some embodiments, the buffer comprises 25 to 75 mM Tris buffer, pH 7 to 8.5, 3 to 6 mM MgCl, 10 to 50 mM KCl, and 20 to 80 mM (NH)SO, inclusive. In some embodiments, 100 to 200 Units / mL of polymerase is used. In some embodiments, 100 mM KCl, 50 mM (NH)SO, 3 mM MgCl, 7.5 nM of each primer in the library, 50 mM TMAC, and 7 μl of DNA template are used in a final volume of 20 μl at pH 8.1.

[0079] In some embodiments, a crowding agent such as polyethylene glycol (PEG, e.g., PEG 8,000) or glycerol is used. In some embodiments, the amount of PEG (e.g., PEG 8,000) is between 0.1 and 20%, such as between 0.5 and 15%, 1 and 10%, 2 and 8%, or 4 and 8%, inclusive. In some embodiments, the amount of glycerol is between 0.1 and 20%, such as between 0.5 and 15%, 1 and 10%, 2 and 8%, or 4 and 8%, inclusive. In some embodiments, the crowding agent allows for lower polymerase concentrations and / or shorter annealing times to be used. In some embodiments, the crowding agent improves DOR uniformity and / or reduces dropout (undetected alleles). Some embodiments use polymerases with proofreading activity, polymerases with no or negligible proofreading activity, or mixtures of polymerases with and without (or negligible) proofreading activity. Some embodiments use hot start polymerases, non-hot start polymerases, or mixtures of hot start and non-hot start polymerases. Some embodiments use HotStarTaq DNA polymerase (see, e.g., Qiagen catalog no. 203203). Some embodiments use AmpliTaq Gold® DNA polymerase. Some embodiments use PrimeSTAR GXL DNA polymerase, a high-fidelity polymerase that provides efficient PCR amplification when excess template is present in the reaction mixture and when amplifying long products (Takara Clontech, Mountain View, CA).In some embodiments, KAPA Taq DNA polymerase or KAPA Taq HotStart DNA polymerase is used, which is a single-subunit, wild-type Taq DNA polymerase derived from the thermophilic bacterium Thermus aquaticus. KAPA Taq and KAPA Taq HotStart DNA polymerases possess 5'-3' polymerase and 5'-3' exonuclease activity, but lack 3'-to-5' exonuclease (proofreading) activity (see, e.g., KAPA BIOSYSTEMS catalog no. BK1000). In some embodiments, Pfu DNA polymerase is used, a highly thermostable DNA polymerase from the hyperthermophilic archaeon Pyrococcus furiosus. The enzyme catalyzes the template-dependent polymerization of nucleotides into double-stranded DNA in the 5'-to-3' direction. Pfu DNA polymerase also exhibits 3' to 5' exonuclease (proofreading) activity, which allows the polymerase to correct nucleotide incorporation errors. The polymerase lacks 5' to 3' exonuclease activity (see, e.g., Thermo Scientific Catalog No. EP0501). In some embodiments, Klentaq1 is used, which is a Klenow-fragment analog of Taq DNA polymerase and lacks exonuclease or endonuclease activity (see, e.g., DNA POLYMERASE TECHNOLOGY, Inc., St. Louis, MO, Catalog No. 100). In some embodiments, the polymerase is a PHUSION DNA polymerase, such as PHUSION High-Fidelity DNA Polymerase (M0530S, New England BioLabs, Inc.) or PHUSION Hot Start Flex DNA Polymerase (M0535S, New England BioLabs, Inc.).In some embodiments, the polymerase is a Q5® DNA polymerase, such as Q5® High-Fidelity DNA Polymerase (M0491S, New England BioLabs, Inc.) or Q5® Hot Start High-Fidelity DNA Polymerase (M0493S, New England BioLabs, Inc.). In some embodiments, the polymerase is T4 DNA polymerase (M0203S, New England BioLabs, Inc.).

[0080] In some embodiments, a polymerase having between 5 and 600 Units / mL (units per mL of reaction volume) is used, such as between 5 and 100, 100 and 200, 200 and 300, 300 and 400, 400 and 500, or 500 and 600 Units / mL, inclusive.

[0081] 5.3 PCR method In some embodiments, hot-start PCR is used to reduce or prevent polymerization before PCR thermocycling. Exemplary hot-start PCR methods involve inhibiting the initial DNA polymerase until the reaction mixture reaches a high temperature, or physically separating the reaction components. Some embodiments utilize the gradual release of magnesium. Because DNA polymerase requires magnesium ions for activity, magnesium is chemically separated from the reaction by binding it to a chemical, releasing it into solution only at high temperatures. Some embodiments utilize non-covalent binding of an inhibitor. In this method, a peptide, antibody, or aptamer non-covalently binds to the enzyme at low temperatures, inhibiting its activity. After incubation at elevated temperatures, the inhibitor is released, initiating the reaction. Some embodiments utilize a cold-sensitive Taq polymerase, such as a modified DNA polymerase that has little activity at low temperatures. Some embodiments utilize chemical modifications. In this method, a molecule is covalently attached to the side chain of an amino acid in the active site of the DNA polymerase. This molecule is released from the enzyme by incubating the reaction mixture at elevated temperature. Once released, the enzyme becomes active.

[0082] In some embodiments, the amount of template nucleic acid (e.g., RNA or DNA sample) is between 20 and 5,000 ng, such as between 20 and 200, 200 and 400, 400 and 600, 600 and 1,000, 1,000 and 1,500, or 2,000 and 3,000 ng, inclusive.

[0083] In some embodiments, a Qiagen Multiplex PCR kit is utilized (Qiagen Catalog No. 206143). For 100 x 50 μL multiplex PCR reactions, the kit contains 2 x QIAGEN Multiplex PCR Master Mix (3 x 0.85 ml, giving a final concentration of 3 mM MgCl), 5 x Q-Solution (1 x 2.0 ml), and RNase-free water (2 x 1.7 ml). QIAGEN Multiplex PCR Master Mix (MM) contains a combination of KCl and (NH4)2SO4, as well as a PCR additive, Factor MP, which increases the local concentration of primers on the template. Factor MP specifically stabilizes bound primers, allowing efficient primer extension by HotStarTaq DNA polymerase. HotStarTaq DNA polymerase is a modified version of Taq DNA polymerase that lacks polymerase activity at room temperature. In some embodiments, HotStarTaq DNA polymerase is activated by a 15 minute incubation at 95°C, which can be incorporated into any existing thermocycling program.

[0084] In some embodiments, a final concentration of 1×QIAGEN MM (recommended), 7.5 nM of each primer in the library, 50 mM TMAC, and 7 μl of DNA template in a final volume of 20 μl are used. In some embodiments, PCR thermocycling conditions include 95° C. for 10 minutes (hot start), 20 cycles of 96° C. for 30 minutes, 65° C. for 15 minutes and 72° C. for 30 seconds, then 72° C. for 2 minutes (final extension), followed by a 4° C. hold.

[0085] In some embodiments, a final concentration of 2x QIAGEN MM (twice the recommended concentration), 2 nM of each primer in the library, 70 mM TMAC, and 7 μl of DNA template are used in a total volume of 20 μl. Some embodiments also include up to 4 mM EDTA. In some embodiments, PCR thermocycling conditions include 95°C for 10 minutes (hot start), 25 cycles of 96°C for 30 seconds, 65°C for 20, 25, 30, 45, 60, 120, or 180 minutes, and optionally 72°C for 30 seconds, followed by 72°C for 2 minutes (final extension), followed by a 4°C hold.

[0086] Another exemplary set of conditions involves a semi-nested PCR approach: The first PCR reaction uses a 20 μl reaction volume containing 2× QIAGEN MM final concentration, 1.875 nM of each primer in the library (forward and reverse outer primers), and DNA template.

[0087] Thermocycling parameters included 95°C for 10 minutes, 25 cycles of 96°C for 30 seconds, 65°C for 1 minute, 58°C for 6 minutes, 60°C for 8 minutes, 65°C for 4 minutes, and 72°C for 30 seconds, followed by a 72°C for 2 minutes, followed by a 4°C hold. Two microliters of the resulting product, diluted 1:200, was then used as input for a second PCR reaction. This reaction used a 10-microliter reaction volume containing a final concentration of 1x QIAGEN MM, 20 nM each of the forward inner primers, and 1 μM of the reverse primer tag. Thermocycling parameters included 95°C for 10 minutes, 15 cycles of 95°C for 30 seconds, 65°C for 1 minute, 60°C for 5 minutes, 65°C for 5 minutes, and 72°C for 30 seconds, followed by a 72°C for 2 minutes, followed by a 4°C hold. The annealing temperature can optionally be higher than the melting temperature of some or all of the primers, as described herein (see U.S. Patent Application No. 14 / 918,544, filed October 20, 2015, which is incorporated by reference in its entirety).

[0088] The melting temperature (Tm) is the temperature at which half (50%) of an oligonucleotide (e.g., a primer) DNA duplex dissociates into single-stranded DNA. The annealing temperature (TA) is the temperature at which a PCR protocol is performed. Traditionally, it is typically 5°C lower than the lowest Tm of the primer used, ensuring nearly all possible duplexes are formed (so that essentially all primer molecules bind to the template nucleic acid). While this is highly efficient, lower temperatures inevitably result in more nonspecific reactions. One consequence of a too-low TA is that primers may anneal to sequences other than their true targets because internal single-base mismatches or partial annealing may be tolerated. In some embodiments of the present invention, the TA is higher than the Tm, such that only a small fraction of the target anneals to the primer at any given time (e.g., only about 1-5%). As they are extended, they are taken out of the equilibrium state of primer-target annealing and dissociation (because extension quickly raises the Tm above 70°C), and approximately 1-5% of the new targets have primers. This allows for long annealing times in the reaction, resulting in approximately 100% copied targets per cycle.

[0089] In various embodiments, the annealing temperature is between 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13°C, to 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or 15°C at the upper end of the range, and is at least 25, 50, 60, 70, 75, 80, 90, 95 or 100% higher than the melting temperature (e.g., empirically measured or calculated Tm) of the mismatched primer. In various embodiments, the annealing temperature is at least 25, 50, 75, 100, 300, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 15,000, 19,000, 20,000, 25,000, 27,000, 28,000, 30,000, 40,000, 50,000, 75, 000, 100,000, or all melting points (e.g., experimentally measured or calculated Tm) by 1 to 15°C (e.g., 1°C to 10°C, 1°C to 5°C, 1°C to 3°C, 3°C to 5°C, 5°C to 10°C, 5°C to 8°C, 8°C to 10°C, 10°C to 12°C, or 12°C to 15°C). In various embodiments, the annealing temperature is between 1 and 15°C (e.g., between 1 and 10, 1 and 5, 1 and 3, 3 and 5, 3 and 8, 5 and 10, 5 and 8, 8 and 10, 10 and 12, or 12 and 15°C, inclusive), and is higher than the melting temperatures (e.g., empirically measured or calculated Tms) of at least 25%, 50%, 60%, 70%, 75%, 80%, 90%, 95% of the mismatched primers or all of the mismatched primers, and the length of the annealing step (per PCR cycle) is between 5 and 180 minutes, such as between 15 and 120 minutes, 15 and 60 minutes, 15 and 45 minutes, or 20 and 60 minutes, inclusive.

[0090] Exemplary Multiplex PCR Methods

[0091] In various embodiments, long annealing times (as described herein and exemplified in Example 12) and / or low primer concentrations are used. Indeed, in some embodiments, limited primer concentrations and / or conditions are used. In various embodiments, the length of the annealing step is between 15, 20, 25, 30, 35, 40, 45, or 60 minutes at the lower end of the range and 20, 25, 30, 35, 40, 45, 60, 120, or 180 minutes at the higher end of the range. In various embodiments, the length of the annealing step (per PCR cycle) is between 30 and 180 minutes. For example, the annealing step can be between 30 and 60 minutes, and the concentration of each primer can be less than 20, 15, 10, or 5 nM. In other embodiments, primer concentrations are from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20 or 25 nM at the lower end of the range to 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25 and 50 at the upper end of the range.

[0092] In high-concentration multiplexing, the solution may become viscous due to the large amount of primers in the solution.If the solution is too viscous, the primer concentration can be reduced to a level that is still sufficient for the primers to bind to template DNA.In various embodiments, 1,000 to 100,000 different primers are used, and the concentration of each primer is less than 20 nM, for example, less than 10 nM, or between 1 and 10 nM (inclusive).

[0093] 5.4 Exemplary Multiplex PCR Methods In one aspect, the invention features a method for amplifying target loci in a nucleic acid sample, including: (i) contacting the nucleic acid sample with a library of primers that simultaneously hybridize to at least 1,000; 2,000; 5,000; 7,500; 10,000; 15,000; 19,000; 20,000; 25,000; 27,000; 28,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different target loci to generate a reaction mixture; and (ii) subjecting the reaction mixture to primer extension reaction conditions (e.g., PCR conditions) to generate amplification products that include target amplicons. In some embodiments, the method also includes determining the presence or absence of at least one target amplicon (e.g., at least 50, 60, 70, 80, 90, 95, 96, 97, 98, 99, or 99.5% of the target amplicons). In some embodiments, the method also includes determining the sequence of at least one target amplicon (e.g., at least 50, 60, 70, 80, 90, 95, 96, 97, 98, 99, or 99.5% of the target amplicons). In some embodiments, at least 50, 60, 70, 80, 90, 95, 96, 97, 98, 99, or 99.5% of the target loci are amplified. In some embodiments, at least 25; 50; 75; 100; 300; 500; 750; 1,000; 2,000; 5,000; 7,500; 10,000; 15,000; 19,000; 20,000; 25,000; 27,000; 28,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different target loci are amplified by at least 5, 10, 20, 40, 50, 60, 80, 100, 120, 150, 200, 300, or 400 times. In some embodiments, at least 50, 60, 70, 80, 90, 95, 96, 97, 98, 99, 99.5, or 100% of the target loci are amplified by at least 5, 10, 20, 40, 50, 60, 80, 100, 120, 150, 200, 300, or 400 fold. In various embodiments, less than 60, 50, 40, 30, 20, 10, 5, 4, 3, 2, 1, 0.5, 0.25, 0.1, or 0.05% of the amplification products are primer dimers.In some embodiments, the methods involve multiplex PCR and sequencing (e.g., high-throughput sequencing).

[0094] In various embodiments, long annealing times and / or low primer concentrations are used. In various embodiments, the length of the annealing step is greater than 3 minutes, 5 minutes, 8 minutes, 10 minutes, 15 minutes, 20 minutes, 30 minutes, 45 minutes, 60 minutes, 75 minutes, 90 minutes, 120 minutes, 150 minutes, or 180 minutes. In various embodiments, the length of the annealing step (per PCR cycle) is greater than 5 minutes and less than 180 minutes, e.g., greater than 5 minutes and less than 60 minutes, greater than 10 minutes and less than 60 minutes, greater than 5 minutes and less than 30 minutes, or greater than 10 minutes and less than 30 minutes. In various embodiments, the length of the annealing step is greater than 5 minutes (e.g., greater than 10 or 15 minutes), and the concentration of each primer is less than 20 nM. In various embodiments, the length of the annealing step is greater than 5 minutes (e.g., greater than 10 or 15 minutes), and the concentration of each primer is greater than 1 nM to 20 nM, or greater than 1 nM to 10 nM. In various embodiments, the length of the annealing step is greater than 20 minutes (e.g., greater than 30, 45, 60, or 90 minutes), and the concentration of each primer is less than 1 nM.

[0095] In high-concentration multiplexing, the solution may become viscous due to the large amount of primers in the solution. If the solution is too viscous, the primer concentration can be reduced to a level that is still sufficient for the primers to bind to the template DNA. In various embodiments, fewer than 60,000 different primers are used, and the concentration of each primer is less than 20 nM, such as less than 10 nM, or from 1 nM to 10 nM. In various embodiments, more than 60,000 different primers (e.g., 60,000 to 120,000 different primers) are used, and the concentration of each primer is less than 10 nM, such as less than 5 nM, or from 1 nM to 10 nM.

[0096] It has been found that the annealing temperature can optionally be higher than the melting temperature of some or all of the primers (in contrast to other methods that use annealing temperatures lower than the melting temperatures of the primers). m The annealing temperature (T) is the temperature at which half (50%) of the DNA duplex of an oligonucleotide (e.g., a primer) dissociates from its perfect complementary strand to form single-stranded DNA. A ) is the temperature at which the PCR protocol is carried out. In traditional methods, the lowest T m The temperature is usually 5°C lower than that of the standard PCR product, so that nearly all possible duplexes are formed (because essentially all primer molecules bind to the template nucleic acid). This is highly efficient, but the lower temperature inevitably leads to a more nonspecific reaction. A One consequence of having T that is too low is that internal single base mismatches or partial annealing may be tolerated, allowing primers to anneal to sequences other than the true target. A (T m ), and only a small fraction of the targets anneal to the primers at any given time (e.g., only about 1-5%). As these are extended, they are taken out of the equilibrium state of primer-target annealing and dissociation (T m At temperatures above 70°C (because the temperature quickly rises above 70°C), approximately 1–5% of the target copies contain primers. This allows for approximately 100% of the target copies per cycle by annealing the reaction for a long time. Therefore, the most stable molecular pairs (those with a perfect DNA pair between the primer and template DNA) are preferentially extended, generating the correct target amplicon. For example, the same experiment was performed at annealing temperatures of 57°C and 63°C using primers with melting temperatures below 63°C. When the annealing temperature was 57°C, the percentage of mapped reads in the amplified PCR products was as low as 50% (approximately 50% of the amplified products were primer dimers). When the annealing temperature was 63°C, the percentage of primer dimers in the amplified products dropped to approximately 2%.

[0097] In various embodiments, the annealing temperature is greater than or equal to the melting temperature (e.g., experimentally measured or calculated T m In some embodiments, the annealing temperature is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 15° C. higher than the melting temperatures (e.g., experimentally measured or calculated T ) of at least 25, 50, 75, 100, 300, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 15,000, 19,000, 20,000, 25,000, 27,000, 28,000, 30,000, 40,000, 50,000, 75,000, 100,000, or all of the non-identical primers. m ), and the length of the annealing step (per PCR cycle) is greater than 1, 3, 5, 8, 10, 15, 20, 30, 45, 60, 75, 90, 120, 150, or 180 minutes.

[0098] In various embodiments, the annealing temperature is greater than or equal to the melting temperature (e.g., experimentally measured or calculated T m) by 1 to 15°C (for example, 1°C to 10°C, 1°C to 5°C, 1°C to 3°C, 3°C to 5°C, 5°C to 10°C, 5°C to 8°C, 8°C to 10°C, 10°C to 12°C, or 12°C to 15°C). In various embodiments, the annealing temperature is greater than or equal to the melting temperature (e.g., experimentally measured or calculated T m ) is 1 to 15°C (for example, 1°C or more and 10°C or less, 1°C or more and 5°C or less, 1°C or more and 3°C or more and 3°C or more and 5°C or less, 5°C or more and 10°C or less, 5°C or more and 8°C or more and 8°C or more and 10°C or less, 10°C or more and 12°C or more and 12°C or more and 15°C or less), and the length of the annealing step (per PCR cycle) is, for example, 5 minutes to 180 minutes, such as 5 minutes to 60 minutes, 10 minutes to 60 minutes, 5 minutes to 30 minutes, or 10 minutes to 30 minutes.

[0099] In some embodiments, the annealing temperature is determined based on the highest melting temperature of the primers (e.g., experimentally measured or calculated T m In some embodiments, the annealing temperature is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 15° C. higher than the highest melting temperature of the primers (e.g., experimentally measured or calculated T m ), and the length of the annealing step (per PCR cycle) is greater than 1, 3, 5, 8, 10, 15, 20, 30, 45, 60, 75, 90, 120, 150, or 180 minutes.

[0100] In some embodiments, the annealing temperature is determined based on the highest melting temperature of the primers (e.g., experimentally measured or calculated Tm ) is 1 to 15°C (e.g., 1 to 10°C, 1 to 5°C, 1 to 3°C, 3 to 5°C, 5 to 10°C, 5 to 8°C, 8 to 10°C, 10 to 12°C, or 12 to 15°C). In some embodiments, the annealing temperature is 1 to 15°C higher than the highest melting point (e.g., experimentally measured or calculated T m ) is higher by 1 to 15°C (for example, 1 to 10°C, 1 to 5°C, 1 to 3°C, 3 to 5°C, 5 to 10°C, 5 to 8°C, 8 to 10°C, 10 to 12°C, or 12 to 15°C), and the length of the annealing step (per PCR cycle) is, for example, 5 to 180 minutes, such as 5 to 60 minutes, 10 to 60 minutes, 5 to 30 minutes, or 10 to 30 minutes.

[0101] In some embodiments, the annealing temperature is determined based on the average melting temperature (e.g., experimentally measured or calculated T ) of at least 25, 50, 75, 100, 300, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 15,000, 19,000, 20,000, 25,000, 27,000, 28,000, 30,000, 40,000, 50,000, 75,000, 100,000, or all of the non-identical primers. m In some embodiments, the annealing temperature is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 15° C. higher than the average melting temperature (e.g., experimentally measured or calculated T ) of at least 25, 50, 75, 100, 300, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 15,000, 19,000, 20,000, 25,000, 27,000, 28,000, 30,000, 40,000, 50,000, 75,000, 100,000, or all of the non-identical primers. m), and the length of the annealing step (per PCR cycle) is greater than 1, 3, 5, 8, 10, 15, 20, 30, 45, 60, 75, 90, 120, 150, or 180 minutes.

[0102] In some embodiments, the annealing temperature is determined based on the average melting temperature (e.g., experimentally measured or calculated T ) of at least 25, 50, 75, 100, 300, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 15,000, 19,000, 20,000, 25,000, 27,000, 28,000, 30,000, 40,000, 50,000, 75,000, 100,000, or all of the non-identical primers. m ) by 1 to 15°C (for example, 1°C to 10°C, 1°C to 5°C, 1°C to 3°C, 3°C to 5°C, 5°C to 10°C, 5°C to 8°C, 8°C to 10°C, 10°C to 12°C, or 12°C to 15°C). In some embodiments, the annealing temperature is adjusted to the average melting temperature (e.g., experimentally measured or calculated T ) of at least 25, 50, 75, 100, 300, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 15,000, 19,000, 20,000, 25,000, 27,000, 28,000, 30,000, 40,000, 50,000, 75,000, 100,000, or all of the non-identical primers. m ) is 1 to 15°C (for example, 1°C or more and 10°C or less, 1°C or more and 5°C or less, 1°C or more and 3°C or more and 3°C or more and 5°C or less, 5°C or more and 10°C or less, 5°C or more and 8°C or more and 8°C or more and 10°C or less, 10°C or more and 12°C or more and 12°C or more and 15°C or less), and the length of the annealing step (per PCR cycle) is, for example, 5 minutes to 180 minutes, such as 5 minutes to 60 minutes, 10 minutes to 60 minutes, 5 minutes to 30 minutes, or 10 minutes to 30 minutes.

[0103] In some embodiments, the annealing temperature is from 50° C. to 70° C., e.g., from 55° C. to 60° C., from 60° C. to 65° C., or from 65° C. to 70° C. In some embodiments, the annealing temperature is from 50° C. to 70° C., e.g., from 55° C. to 60° C., from 60° C. to 65° C., or from 65° C. to 70° C., and (i) the length of the annealing step (per PCR cycle) is longer than 3, 5, 8, 10, 15, 20, 30, 45, 60, 75, 90, 120, 150, or 180 minutes, or (ii) the length of the annealing step (per PCR cycle) is from 5 minutes to 180 minutes, e.g., from 5 minutes to 60 minutes, from 10 minutes to 60 minutes, from 5 minutes to 30 minutes, or from 10 minutes to 30 minutes.

[0104] In some embodiments, one or more of the following conditions is true: T m used in experimental measurements of or assumed to be T m is calculated using the following conditions: temperature is 60.0°C, primer concentration is 100 nM, and / or salt concentration is 100 mM. In some embodiments, other conditions are used, such as those used for multiplex PCR with libraries. In some embodiments, 100 mM KCl, 50 mM (NH4)2SO4, 3 mM MgCl2, 7.5 nM each primer, and 50 mM TMAC, pH 8.1 are used. In some embodiments, T mis calculated using the Primer3 program (libprimer3, release 2.2.3) using the built-in SantaLucia parameters (available on the World Wide Web at primer3.sourceforge.net, incorporated herein by reference in its entirety). In some embodiments, the calculated primer melting temperature is the temperature at which half of the primer molecules are expected to anneal. As noted above, even at temperatures higher than the calculated melting temperature, a certain percentage of the primers will anneal and thus allow PCR extension. In some embodiments, the experimentally measured Tm (actual Tm) is determined using a thermostatically controlled cell in a UV spectrophotometer. In some embodiments, plotting temperature versus absorbance generates an S-shaped curve with two plateaus. The absorbance value midway between the plateaus corresponds to the Tm.

[0105] In some embodiments, absorbance at 260 nm is measured as a function of temperature on an Ultrospec 2100 pr UV / Visible Spectrophotometer (Amersham Biosciences) (see, e.g., Takiya et al., "An empirical approach for thermal stability (Tm) prediction of PNA / DNA duplexes," Nucleic Acids Symp Ser(Oxf);(48):131-2, 2004, which is incorporated herein by reference in its entirety). In some embodiments, absorbance at 260 nm is measured by decreasing the temperature from 95°C to 20°C at a rate of 2°C per minute. In some embodiments, a primer and its perfect complement (e.g., 2 μM of each pair of oligomers) are mixed, and then annealed by heating the sample to 95°C over the course of 30 minutes, maintaining that temperature for 5 minutes, and then cooling to room temperature and maintaining the sample at 95°C for at least 60 minutes. In some embodiments, the melting temperature is determined by analyzing the data using SWIFT Tm software. In some embodiments of any of the methods of the invention, the method comprises experimentally measuring or calculating (by computer) the melting temperature of at least 50, 80, 90, 92, 94, 96, 98, 99, or 100% of the primers in the library either before or after the primers are used for PCR amplification of the target locus.

[0106] In some embodiments, the library comprises a microarray. In some embodiments, the library does not comprise a microarray.

[0107] In some embodiments, most or all of the primers are extended to form amplification products. When all primers are consumed in a PCR reaction, the same or nearly the same number of primer molecules are converted into target amplicons for each target locus, thereby increasing the uniformity of amplification of various target loci. In some embodiments, at least 80, 90, 92, 94, 96, 98, 99, or 100% of the primer molecules are extended to form amplification products. In some embodiments, for at least 80, 90, 92, 94, 96, 98, 99, or 100% of the target loci, at least 80, 90, 92, 94, 96, 98, 99, or 100% of the primer molecules for that target locus are extended to form amplification products. In some embodiments, multiple cycles are performed until that percentage of primers is consumed. In some embodiments, multiple cycles are performed until all or substantially all of the primers are consumed. If desired, a higher percentage of primers can be consumed by reducing the initial primer concentration and / or increasing the number of PCR cycles performed.

[0108] In some embodiments, PCR may be performed in microliter reaction volumes, which may make it difficult to achieve specific PCR amplification (due to the low local concentration of template nucleic acid) compared to the nanoliter or picoliter reaction volumes used in microfluidic applications. In some embodiments, the reaction volume is between 1 L and 60 L, e.g., between 5 L and 50 L, between 10 L and 50 L, between 10 L and 20 L, between 20 L and 30 L, between 30 L and 40 L, or between 40 L and 50 L.

[0109] In one embodiment, the method disclosed herein uses highly efficient, highly multiplexed targeted PCR to amplify DNA, followed by high-throughput sequencing to determine the allele frequency at each target locus.The ability to multiplex more than about 50 or 100 PCR primers in one reaction volume so that the majority of the sequence reads obtained are mapped to target loci is novel and non-trivial.One of the techniques that can perform highly multiplexed targeted PCR in a highly efficient manner is to design primers that are unlikely to hybridize with each other. PCR probes, typically referred to as primers, are selected by generating thermodynamic models of potentially harmful interactions between at least 300; at least 500; at least 750; at least 1,000; at least 2,000; at least 5,000; at least 7,500; at least 10,000; at least 20,000; at least 25,000; at least 30,000; at least 40,000; at least 50,000; at least 75,000; or at least 100,000 possible primer pairs, or unintended interactions between primers and sample DNA, and then using the models to eliminate designs that are incompatible with other designs in the pool. Another method that can perform highly multiplexed targeted PCR in a highly efficient manner is to use a partial or complete nesting approach to targeted PCR. Using one or a combination of these methods, it is possible to multiplex at least 300, at least 800, at least 1,200, at least 4,000, or at least 10,000 primers in a single pool, resulting in amplified DNA in which the majority of DNA molecules map to the target locus when sequenced. Using one or a combination of these methods, it is possible to multiplex a large number of primers in a single pool, resulting in amplified DNA in which more than 50%, more than 60%, more than 67%, more than 80%, more than 90%, more than 95%, more than 96%, more than 97%, more than 98%, more than 99%, or more than 99.5% of the DNA molecules map to the target locus.

[0110] In some embodiments, detection of target genetic material may be performed in a multiplexed manner. The number of target gene sequences that can be run in parallel can range from 1 to 10, 10 to 100, 100 to 1,000, 1,000 to 10,000, 10,000 to 100,000, 100,000 to 1 million, or 1 million to 10 million. Attempts to multiplex more than 100 primers per pool have resulted in significant problems, such as undesirable side reactions, such as primer-dimer formation.

[0111] 5.5 Targeted PCR In some embodiments, PCR can be used to target specific locations in the genome. In plasma samples, the original DNA is highly fragmented (typically less than 500 bp, with an average length of less than 200 bp). In PCR, both the forward and reverse primers can anneal to and amplify the same fragment. Therefore, if the fragment is short, the PCR assay should amplify even a relatively short region. As with MIPS, if the polymorphism location is too close to the polymerase binding site, bias in amplification from various alleles may occur. Currently, PCR primers targeting polymorphic regions, such as those containing SNPs, are typically designed so that the 3' end of the primer hybridizes to the base immediately adjacent to the polymorphic base. In an embodiment of the present disclosure, the 3' ends of both the forward and reverse PCR primers are designed to hybridize to bases one or several positions away from the variant position (polymorphic site) of the target allele. The number of bases between the polymorphic site (SNP or another polymorphism) and the base at which the 3' end of the primer is designed to hybridize may be 1, 2, 3, 4, 5, 6, 7 to 10, 11 to 15, or 16 to 20. The forward and reverse primers may be designed to hybridize at various numbers of bases away from the polymorphic site.

[0112] Although PCR assays can be performed in large numbers, interactions between different PCR assays make it difficult to multiplex more than about 100 assays. While various complex molecular approaches can be used to increase the level of multiplexing, the number of assays per reaction is still limited to less than 100, perhaps 200, or even 500. Samples containing large amounts of DNA are divided into multiple subreactions and then remixed before sequencing. For samples in which either the entire sample or a subset of DNA molecules is limited, dividing the sample introduces statistical noise. In one embodiment, a small or limited amount of DNA can refer to amounts less than 10 pg, 10-100 pg, 100 pg-1 ng, 1-10 ng, or 10-100 ng. This method is particularly useful for small amounts of DNA, where other methods involve dividing into multiple pools, which can pose significant problems related to the resulting stochastic noise. Furthermore, this method also offers the advantage of minimizing bias when performed on samples with any DNA amount. Under these circumstances, a general pre-amplification step may be used to increase the overall sample volume, which ideally should not significantly alter the allele distribution.

[0113] In one embodiment, the disclosed method generates PCR products specific to a large number of target loci, specifically 1,000-5,000 loci, 5,000-10,000 loci, or more than 10,000 loci, allowing for genotyping by sequencing or some other genotyping method from a limited sample, such as DNA from a single cell or body fluid. Currently, performing multiplex PCR reactions for more than 5-10 targets is a significant challenge and is often hindered by primer by-products, such as primer dimers and other artifacts. When detecting target sequences using microarrays with hybridization probes, primer dimers and other artifacts are undetectable and can be ignored. However, when using sequencing as a detection method, the majority of sequence reads will sequence these artifacts rather than the desired target sequences in the sample. Prior art methods used to multiplex more than 50 or 100 reactions in a single reaction volume followed by sequencing often result in more than 20%, even more than 50%, often more than 80%, and in some cases more than 90% of the sequence reads being off-target.

[0114] Typically, to perform targeted sequencing of multiple (n) targets (>50, >100, >500, or >1,000) in a sample, the sample can be split into multiple parallel reactions, amplifying each target individually. This method has been performed in PCR multiwell plates or can be performed on commercially available platforms, such as the FLUIDIGM ACCESS ARRAY (in a microfluidic chip, 48 reactions per sample) or Rain Dance Technology's Droplet PCR (hundreds to thousands of targets). Unfortunately, these splitting and pooling methods are problematic for samples with limited DNA, where the genome copy number is often insufficient to ensure one copy of each genomic region is present in each well. This is particularly problematic when polymorphic loci are targeted and the relative proportions of alleles at polymorphic loci are desired, as stochastic noise introduced by splitting and pooling can result in highly inaccurate measurements of the allele proportions present in the original DNA sample. Described herein is a method for effectively and efficiently amplifying multiple PCR reactions that is applicable when only limited amounts of DNA are available. In one embodiment, the method may be applied to the analysis of DNA mixtures, such as single cells, bodily fluids such as plasma, free-floating DNA present in biopsies, environmental and / or forensic samples.

[0115] In one embodiment, targeted sequencing can involve one, several, or all of the following steps: a) generating and amplifying a library with adapter sequences at both ends of the DNA fragments; b) dividing the library into multiple reactions after library amplification; c) generating and optionally amplifying a library with adapter sequences at both ends of the DNA fragments; d) performing 1000-10,000-multiplexed amplifications of selected targets using one target-specific "forward" primer and one tag-specific primer per target; e) performing a second amplification of the products using a "reverse" target-specific primer and one (or more) primers specific to the universal tag introduced as part of the target-specific forward primer in the first round; f) performing a 1000-multiplexed pre-amplification of selected targets for a limited number of cycles; g) dividing the products into multiple aliquots and amplifying subpools of targets in individual reactions (e.g., 50-500-multiplexed, although single-plexing can also be used); h) pooling the products of the parallel subpool reactions. i) During these amplifications, the primers carry sequence matching tags (partial or full length) so that the products can be sequenced.

[0116] 5.6 Highly multiplexed PCR Disclosed herein are methods that enable targeted amplification of hundreds to tens of thousands of target sequences (e.g., SNP loci) from a nucleic acid sample, such as genomic DNA obtained from plasma. The amplified sample may be relatively free of primer-dimer products, resulting in low allele bias at the target loci. Analysis of these products can be performed by sequencing if sequence-matched adapters are added to the products during or after amplification.

[0117] Even when performing highly multiplexed PCR amplification using methods known in the art, unsuitable primer dimer products for sequencing are generated in excess of the desired amplification products. These products can be experimentally reduced by removing the primers that form these products or by performing in silico primer selection. However, the larger the number of assays, the more difficult this problem becomes.

[0118] One solution is to split the 5000-multiplex reaction into several lower-multiplex amplifications, e.g., 100 50-multiplex reactions or 50 100-multiplex reactions, or to split the sample into individual PCR reactions. However, when sample DNA is limited, such as in noninvasive prenatal testing from pregnancy plasma, splitting the sample into multiple reactions creates a bottleneck and should be avoided.

[0119] Described herein is a method for initially broadly amplifying plasma DNA from a sample and then splitting the sample into multiple multiplexed target enrichment reactions with a more moderate number of target sequences per reaction. In one embodiment, the disclosed method can be used to preferentially enrich a DNA mixture at multiple loci. The method includes one or more of the following steps: generating and amplifying a library from a DNA mixture, where molecules in the library have adapter sequences ligated to both ends of the DNA fragments; splitting the amplified library into multiple reactions; and performing a first round of multiplexed amplification of selected targets using one or more target-specific "forward" primers and one or more adapter-specific universal "reverse" primers per target. In one embodiment, the disclosed method further includes performing a second round of amplification using "reverse" target-specific primers and one or more primers specific to the universal tags introduced as part of the target-specific forward primers in the first round. In one embodiment, the method may involve fully nested, hemi-nested, semi-nested, one-sided fully nested, one-sided hemi-nested, or one-sided semi-nested PCR. In one embodiment, the disclosed method is used to preferentially enrich a DNA mixture at multiple loci, comprising multiplexed preamplification of selected targets for a limited number of cycles, dividing the products into multiple aliquots, and amplifying subpools of targets in individual reactions, and pooling the products of the parallel subpool reactions. Note that this method can be used to perform targeted amplification for 50-500 loci, 500-5,000 loci, 5,000-50,000 loci, or even 50,000-500,000 loci with low levels of allelic bias. In one embodiment, the primers carry partial or full-length sequence match tags.

[0120] A workflow may involve (1) extracting DNA, e.g., plasma DNA, (2) preparing a fragment library with universal adapters at both ends of the fragments, (3) amplifying the library using universal primers specific to the adapters, (4) dividing the amplified sample "library" into multiple aliquots, (5) multiplexing the aliquots (e.g., about 100-, 1,000-, or 10,000-multiplexing, with one target-specific primer per target and a tag-specific primer), (6) pooling the aliquots of one sample, (7) barcoding the samples, (8) mixing the samples and adjusting the concentration, and (9) sequencing the samples. A workflow may include multiple substeps, each of which may include one of the listed steps (e.g., library preparation step (2) may involve three enzymatic steps (blunt-end, dA tailing, and adapter ligation) and three purification steps). Workflow steps may be combined, split, or performed in a different order (e.g., barcoding and pooling samples).

[0121] Note that library amplification can be performed in a biased manner to favor more efficient amplification of short fragments. This allows for preferential amplification of short sequences, such as mononucleosomal DNA fragments, such as cell-free fetal DNA (placental origin) present in the circulation of pregnant women. Note that PCR assays can also contain tags, such as sequence tags (usually truncated, 15-25 base pairs). After multiplexing, the PCR multiplexes of samples are pooled, and tagging is then completed (including barcoding) by tag-specific PCR (which can also be performed by ligation). Additionally, full-length sequence tags can be added in the same reaction as multiplexing. Targets can be amplified using target-specific primers in the first cycle, followed by SQ adapter sequencing following the tag-specific primers. PCR primers do not have to carry tags. Sequence tags can also be added to amplification products by ligation.

[0122] In one embodiment, highly multiplexed PCR followed by clonal sequencing can be used to evaluate the amplified products for various applications, such as detecting fetal aneuploidy. Conventional multiplex PCR simultaneously evaluates up to 50 loci. However, the methods described herein can simultaneously evaluate more than 50 loci, more than 100 loci, more than 500 loci, more than 1,000 loci, more than 5,000 loci, more than 10,000 loci, more than 50,000 loci, and more than 100,000 loci. Experiments have shown that up to 10,000 and even more than 10,000 distinct loci can be simultaneously evaluated in a single reaction, with sufficient efficiency and specificity for highly accurate, non-invasive prenatal aneuploidy diagnosis and / or copy number determination. Assays may be combined in a single reaction using the entire sample, e.g., a cfDNA sample isolated from plasma, a fraction thereof, or a further processed cfDNA derivative. The sample (e.g., cfDNA or derivative) may be split into multiple parallel multiplexed reactions. Optimal sample splitting and multiplexing are determined by trading off various performance specifications. Due to limited material, splitting a sample into multiple fractions can introduce sampling noise, require additional processing time, and increase the potential for error. Conversely, the higher the multiplexing, the greater the amount of spurious amplification and the greater the amplification imbalance, both of which can reduce test performance.

[0123] In the application of the methods described herein, two important related concerns exist: the limited volume of the original sample (e.g., plasma) and the limited number of original molecules in the material from which allele frequency or other measurements are obtained. Below a certain level, random sampling noise becomes significant and can affect test accuracy. Typically, when measurements are performed on samples containing 500–1000 original molecules per target locus, sufficient quality data can be obtained for noninvasive prenatal aneuploidy diagnosis. There are also many ways to increase the number of separate measurements, for example, by increasing the sample volume. Each manipulation performed on the sample can also result in material loss. It is essential to characterize and avoid losses caused by various manipulations, or, if necessary, improve the yield of specific manipulations to avoid losses that could reduce test performance.

[0124] In one embodiment, all or a fraction of the original sample (e.g., a cfDNA sample) can be amplified to reduce potential losses in subsequent steps. Various methods are available for amplifying all of the genetic material in a sample, increasing the amount available for downstream procedures. In one embodiment, DNA fragments in ligation-mediated PCR (LM-PCR) are amplified by PCR after ligation of one, two, or many separate adapters. In one embodiment, all DNA is amplified isothermally using phi-29 polymerase in multiple displacement amplification (MDA). DOP-PCR and its variants use random priming to amplify the original DNA material. Each method has its own characteristics, such as uniformity of amplification across all expressed regions of the genome, efficiency of capture and amplification of original DNA, and amplification performance as a function of fragment length.

[0125] In one embodiment, LM-PCR may be used with a single heteroduplex adapter containing a 3-prime tyrosine. The heteroduplex adapter allows for the use of a single adapter molecule, which can be converted into two different sequences on the 5-prime and 3-prime ends of the original DNA fragments during the first round of PCR. In one embodiment, the amplified library can be divided by size separation or by products such as AMPURE, TASS, or other similar methods. Prior to ligation, the sample DNA may be blunt-ended, and then a single adenosine base is added to the 3-prime end. Prior to ligation, the DNA may be cleaved using a restriction enzyme or some other cleavage method. During ligation, the 3-prime adenosine of the sample fragment and the complementary 3-prime tyrosine overhang of the adapter enhance ligation efficiency. The extension step of PCR amplification may be time-limited to reduce amplification from fragments longer than about 200 bp, about 300 bp, about 400 bp, about 500 bp, or about 1,000 bp. Multiple reactions were performed using the conditions specified by the commercially available kit. As a result, less than 10% of the sample DNA molecules were successfully ligated. However, through a series of optimizations of the reaction conditions, ligation was improved to approximately 70%.

[0126] In one embodiment, the method described herein performs highly efficient and multiplexed targeted PCR for at least 1000 target loci, and then performs next-generation sequencing to obtain genotype data.In some embodiments, multiplexed targeted PCR is performed on at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 11000, 12000, 13000, 14000, 15000, 20000, 25000, 30000, 35000, 40000, 45000, or 50000 target loci.In one embodiment, multiplexed targeted PCR is performed on 13000 target loci.In some embodiments, the target loci are single nucleotide polymorphisms (SNPs).

[0127] 5.7 Mini-PCR The following mini-PCR method is desirable for samples containing short, digested, or fragmented nucleic acids, such as cfDNA. While traditional PCR assay designs result in significant loss of distinct fetal molecules, this loss can be significantly reduced by designing a very short PCR assay called a mini-PCR assay. Fetal cfDNA in maternal serum is highly fragmented, with fragment sizes distributed in an approximately Gaussian pattern, with an average of 160 bp, a standard deviation of 15 bp, a minimum size of approximately 100 bp, and a maximum size of approximately 220 bp. Regarding targeted polymorphisms, the distribution of fragment start and end positions is not necessarily random, but varies significantly between individual targets and among all targets collectively, and the polymorphic site of a specific target locus can occupy any position between the start and end positions among the various fragments derived from that locus. It should be noted that the term mini-PCR can equally refer to conventional PCR without additional restrictions or limitations.

[0128] During PCR, amplification occurs only from template DNA fragments containing both the forward and reverse primer sites. Because fetal cfDNA fragments are short, the likelihood of both primer sites being present is the likelihood of a fetal fragment of length L containing both the forward and reverse primer sites, which is the ratio of the amplicon length to the fragment length. Under ideal conditions, assays with amplicons of 45, 50, 55, 60, 65, or 70 bp will successfully amplify 72%, 69%, 66%, 63%, 59%, or 56% of the available template fragment molecules, respectively. The amplicon length is the distance between the 5-prime ends of the forward and reverse priming sites. Amplicon lengths shorter than those typically used by those skilled in the art can result in more efficient measurement of the desired polymorphic locus by requiring only short sequence reads. In one embodiment, a significant proportion of the amplicons are less than 100 bp, less than 90 bp, less than 80 bp, less than 70 bp, less than 65 bp, less than 60 bp, less than 55 bp, less than 50 bp, or less than 45 bp.

[0129] In prior art methods, short assays such as those described herein are typically avoided because they are not necessary and because they impose significant constraints on primer design due to limited primer length, annealing properties, and distance between forward and reverse primers.

[0130] It should also be noted that biased amplification may occur if the 3-prime end of either primer is within approximately 1-6 bases of the polymorphic site. This single-base difference in the initial polymerase binding site may result in preferential amplification of one allele, thereby altering the observed allele frequency and potentially reducing performance. All of these constraints make it extremely difficult to identify primers that successfully amplify specific loci, as well as to design large primer sets that are compatible in the same multiplex reaction. In one embodiment, the 3' ends of the inner forward and reverse primers are designed to hybridize to DNA regions upstream from the polymorphic site and are separated from the polymorphic site by a small number of bases. Ideally, the number of bases may be 6-10, but may also be 4-15, 3-20, 2-30, or 1-60 bases, achieving substantially the same results.

[0131] Multiplex PCR can involve a single round of PCR in which all targets are amplified, or it can involve a single round of PCR followed by one or more rounds of nested PCR or some variation of nested PCR. Nested PCR consists of subsequent rounds of PCR amplification using one or more new primers that bind at least one base pair internal to the primer used in the previous round. Nested PCR reduces the number of spurious amplification targets by amplifying only those amplification products from the previous amplification that have the correct internal sequence in the subsequent reaction. Reducing spurious amplification targets improves the number of useful measurements that can be obtained, especially in sequencing. Nested PCR typically involves designing primers completely internal to the previous primer binding site, which necessarily increases the minimum DNA segment size required for amplification. For samples with highly fragmented DNA, such as plasma cfDNA, as the assay size increases, the number of individual cfDNA molecules from which measurements can be obtained decreases. In one embodiment, to counteract this effect, a partially nested approach may be used in which one or both of the second round primers overlap with the first binding site and are extended inward by a few bases, providing more specificity while minimizing the increase in total assay size.

[0132] In one embodiment, a multiplexed pool of PCR assays is designed to amplify potential heterozygous SNPs or other polymorphic or non-polymorphic loci on one or more chromosomes, and these assays are used in a single reaction to amplify DNA. The number of PCR assays can be 50-200 PCR assays, 200-1,000 PCR assays, 1,000-5,000 PCR assays, or 5,000-20,000 PCR assays (50-200-multiplexed, 200-1,000-multiplexed, 1,000-5,000-multiplexed, 5,000-20,000-multiplexed, and over 20,000-multiplexed, respectively). In one embodiment, a multiplexed pool of approximately 10,000 PCR assays (10,000-multiplexed) is designed to amplify potential heterozygous SNP loci on chromosomes X, Y, 13, 18, and 21, and 1 or 2. These assays are used in a single reaction to amplify cfDNA obtained from maternal plasma, chorionic villus samples, amniocentesis samples, single or small numbers of cells, other body fluids or tissues, cancer, or other genetic material. The SNP frequency of each locus may be determined by clonal or other methods of amplicon sequencing. Statistical analysis of allele frequency distributions or ratios across all assays may be used to determine whether the sample contains one or more trisomies for the chromosomes included in the test. In another embodiment, the original cfDNA sample is split into two samples, and parallel 5,000-multiplexed assays are performed. In another embodiment, the original cfDNA sample is divided into n samples, and parallel (approximately 10,000 / n) multiplexed assays are performed. In this case, n is 2 to 12, or 12 to 24, or 24 to 48, or 48 to 96. Data are collected and analyzed in a manner similar to that previously described. Note that this method is equally applicable to the detection of translocations, deletions, duplications, and other chromosomal abnormalities.

[0133] In one embodiment, tails with no homology to the target genome may be added to the 3-prime or 5-prime ends of any of the primers. These tails facilitate subsequent manipulations, procedures, or measurements. In one embodiment, the tail sequence may be the same as the forward or reverse target-specific primer. In one embodiment, different tail sequences may be used for the forward or reverse target-specific primer. In one embodiment, multiple different tails may be used for different loci or sets of loci. A particular tail may be common to all loci or to a subset of loci. For example, using forward and reverse tails corresponding to the forward and reverse primers required by any current sequencing platform allows for direct sequencing after amplification. In one embodiment, the tails can be used as priming sites common to all amplified targets or to add other useful sequences. In some embodiments, inner primers may contain regions designed to hybridize either upstream or downstream of the target locus (e.g., polymorphic locus). In some embodiments, the primers may contain molecular barcodes. In some embodiments, the primers may contain universal priming sequences designed to allow for PCR amplification.

[0134] In one embodiment, a 10,000-multiplexed PCR assay pool is created in which the forward and reverse primers have tails corresponding to the requisite forward and reverse sequences required for high-throughput sequencing instruments (often referred to as massively parallel sequencing instruments), such as the HISEQ, GAIIX, or MYSEQ available from ILLUMINA. Additionally, the 5-prime contained in the sequencing tails is an additional sequence that can be used as a priming site in subsequent PCR to add nucleotide barcode sequences to the amplicons, allowing for multiplexed sequencing of multiple samples in a single lane of the high-throughput sequencing instrument.

[0135] In one embodiment, a 10,000-multiplexed PCR assay pool is created in which the reverse primer has a tail corresponding to the required reverse sequence required for high-throughput sequencing equipment. After amplification in the first 10,000-multiplexed assay, subsequent PCR amplifications can be performed using another 10,000-multiplexed pool with partially nested forward primers (e.g., nested by 6 bases) for all targets and reverse primers corresponding to the reverse sequence tails included in the first round. Performing subsequent rounds of partially nested amplification using only one target-specific primer and a universal primer limits the required assay size and reduces sample noise, but also significantly reduces the number of spurious amplicons. Sequence tags can be added to attached ligation adapters and / or as part of the PCR probe, so that the tags become part of the final amplicons.

[0136] The tumor fraction influences test performance. Numerous methods exist for increasing the tumor fraction of DNA present in patient plasma. This can be achieved by the previously discussed LM-PCR method, as well as by targeted removal of long fragments. In one embodiment, prior to multiplex PCR amplification of the target loci, an additional multiplex PCR reaction can be performed to selectively remove long maternal fragments corresponding to the loci targeted in the subsequent multiplex PCR. Additional primers are designed to anneal at sites at a greater distance from the polymorphisms expected to be present in cell-free fetal DNA fragments. These primers can be used in a single-cycle multiplex PCR reaction, followed by multiplex PCR of the target polymorphic loci. These distal primers are tagged with a molecule or moiety that allows selective recognition of the tagged fragment of DNA. In one embodiment, these DNA molecules can be covalently modified with biotin molecules after the single-cycle PCR, allowing removal of newly formed double-stranded DNA containing these primers. The double-stranded DNA formed during the first round is likely to be maternal in origin. Removal of hybrid material may be achieved by using magnetic streptavidin beads. Other tagging methods exist that function similarly. In one embodiment, a size selection method may be used to enrich the sample for shorter DNA strands, for example, strands less than about 800 bp, less than about 500 bp, or less than about 300 bp. Amplification of short fragments can then proceed as usual.

[0137] The mini-PCR method described in this disclosure enables highly multiplexed amplification and analysis of hundreds, thousands, or even millions of loci from a single sample in a single reaction. Simultaneously, detection of the amplified DNA can be multiplexed. Dozens to hundreds of samples can be multiplexed in a single sequencing lane using barcoded PCR. This multiplexed detection has been successfully tested up to 49-multiplexing, with even greater multiplexing possible. Effectively, this allows hundreds of samples to be genotyped with thousands of SNPs in a single sequencing run. For these samples, the method allows for determination of genotype and heterozygosity rate, as well as simultaneous copy number determination, both of which may be used for aneuploidy detection. It may also be used as part of a mutation burden method. The method may be used with any amount of DNA or RNA, and the target region may be a SNP, other polymorphic region, non-polymorphic region, or a combination thereof.

[0138] In some embodiments, ligation-mediated universal PCR amplification of fragmented DNA can be used.Ligation-mediated universal PCR amplification can be used to amplify plasma DNA, and then it can be divided into multiple parallel reactions.Furthermore, this can be used to preferentially amplify short fragments, thereby increasing tumor fraction.In some embodiments, adding tags to fragments by ligation allows for the detection of shorter fragments, the use of shorter target sequence-specific primer parts, and / or annealing at high temperatures, which reduces non-specific reactions.

[0139] The methods described herein may be used for a number of purposes involving a set of target DNA mixed with a certain amount of contaminating DNA. In some embodiments, the target DNA and contaminating DNA may be from genetically related individuals. For example, fetal (target) genetic abnormalities may be detected from maternal plasma, which also contains fetal (target) DNA and maternal (contaminating) DNA. Such abnormalities include whole chromosomal abnormalities (e.g., aneuploidies), partial chromosomal abnormalities (e.g., deletions, duplications, inversions, translocations), polynucleotide polymorphisms (e.g., STRs), single nucleotide polymorphisms, and / or other genetic abnormalities or variations. In some embodiments, the target DNA and contaminating DNA may be from the same individual, but the target DNA and contaminating DNA differ by one or more mutations, such as in the case of cancer. (See, e.g., H. Mamon et al., Preferential Amplification of Apoptotic DNA from Plasma: Potential for Enhancing Detection of Minor DNA Alterations in Circulating DNA. Clinical Chemistry 54:9 (2008)). In some embodiments, DNA can be found in cell culture (apoptotic) supernatant. In some embodiments, apoptosis can be induced in biological samples (e.g., blood), followed by library preparation, amplification, and / or sequencing. Numerous workflows and protocols for achieving this goal are presented elsewhere in this disclosure.

[0140] In some embodiments, the target DNA may be from a single cell, from a DNA sample consisting of less than one copy of the target genome, from small amounts of DNA, from DNA from mixed sources (e.g., plasma and tumor of a cancer patient, mixed healthy and cancerous DNA, transplants, etc.), from other bodily fluids, from cell cultures, from culture supernatants, from forensic DNA samples, from ancient DNA samples (e.g., insects trapped in amber), from other DNA samples, and combinations thereof.

[0141] In some embodiments, short amplicon sizes may be used, which are particularly suitable for fragmented DNA (see, e.g., A. Sikora, et al. Detection of increased amounts of cell-free fetal DNA with short PCR amplicons. Clin Chem. 2010 Jan;56(1):136-8).

[0142] Using a short amplicon size can provide several significant benefits. A short amplicon size can result in optimized amplification efficiency. A short amplicon size typically results in shorter products. Therefore, nonspecific priming is less likely to occur. Because the clusters are smaller, short products can be clustered more densely on the sequencing flow cell. Note that the methods described herein can work equally well with longer PCR amplicons. The amplicon length may be increased as needed, for example, when sequencing longer sequences. Experiments using 146-multiplexed target amplification with assays of 100 bp to 200 bp length as the first step of a nested PCR protocol were performed on single cells and on genomic DNA, with positive results.

[0143] In some embodiments, the methods described herein may be used to amplify and / or detect SNPs, copy number, nucleotide methylation, mRNA levels, other types of RNA expression levels, other genetic and / or epigenetic characteristics. The mini-PCR methods described herein may be used in conjunction with next-generation sequencing, or other downstream methods such as microarrays, digital PCR counting, real-time PCR, mass spectrometry, etc.

[0144] In some embodiments, the mini-PCR amplification method described herein may be used as part of a method for accurate quantification of minority populations. It may be used for absolute quantification using spike calibrators. It may be used for quantification of mutations / minor alleles via very deep sequencing and may be performed in a highly multiplexed manner. It may be used for standard paternity testing and for kinship or ancestry identity testing in humans, animals, plants, or other organisms. It may be used for forensic testing. It may be used for rapid genotyping and copy number analysis (CN) of any type of material, such as amniotic fluid and CVS, sperm, and products of conception (POC). It may be used for single-cell analysis, such as genotyping of samples biopsied from embryos. It may be used for rapid embryo analysis (within one, one, or two days of biopsy) by targeted sequencing using mini-PCR.

[0145] In some embodiments, mini-PCR amplification can be used for tumor analysis. Tumor biopsies are often a mixture of healthy and tumor cells. Targeted PCR allows deep sequencing of SNPs and loci with little to no background sequence. It can be used for copy number analysis and loss of heterozygosity analysis of tumor DNA. The tumor DNA can be present in many different body fluids or tissues of tumor patients. It can be used to detect tumor recurrence and / or tumor screening. It can be used for seed quality control testing. It can be used for breeding or fishing purposes. Note that any of these methods can equally well be used to target non-polymorphic loci for ploidy classification purposes.

[0146] Some references that describe some of the fundamental methods underlying the methods disclosed herein include the following: (1) Wang HY, Luo M, Tereshchenko IV, Frikker DM, Cui X, Li JY, Hu G, Chu Y, Azaro MA, Lin Y, Shen L, Yang Q, Kambouris ME, Gao R, Shih W, Li H. Genome Res. 2005 Feb;15(2):276-83. Department of Molecular Genetics, Microbiology and Immunology / The Cancer Institute of New Jersey, Robert Wood Johnson Medical School, New Brunswick, New Jersey 08903, USA. (2) High-throughput genotyping of single nucleotide polymorphisms with high sensitivity. Li H, Wang HY, Cui X, Luo M, Hu G, Greenawalt DM, Tereshchenko IV, Li JY, Chu Y, Gao R. Methods Mol Biol. 2007;396 - PubMed PMID: 18025699. (3) A method comprising multiplexing of an average of 9 assays for sequencing is described in: Nested Patch PCR enables highly multiplexed mutation discovery in candidate genes. Varley KE, Mitra RD. Genome Res. 2008 Nov;18(11):1844-50. Epub 2008 Oct 10. Note that the method disclosed herein allows for a higher degree of multiplexing than the above references.

[0147] 6. Methods for analyzing genomic information from circulating fetal cells for non-invasive prenatal testing. In one embodiment, the method further comprises performing a non-invasive prenatal test using the genotype determined to be one or more fetal cells.

[0148] In one embodiment, the method further comprises detecting copy number variations or aneuploidy of a target chromosome or target chromosome segment of interest in one or more circulating cells determined to be one or more fetal cells.

[0149] Ploidy classification, also known as "chromosome copy number classification" or "copy number calling" (CNC), refers to the determination of the amount and identity of one or more chromosomes present in a cell. Aneuploidy refers to the presence of the wrong number of chromosomes in a cell. In human somatic cells, it refers to the absence of 22 pairs of autosomes and one pair of sex chromosomes. In human gametes, it refers to the absence of one of the 23 chromosomes. In the case of a single chromosome, it refers to the presence of more or less than two homologous but non-identical chromosomes, where each of the two chromosomes originates from a different parent. Ploidy state refers to the amount and identity of one or more chromosomes in a cell.

[0150] CNVs are often assigned to one of two major categories based on the length of the affected sequence. The first category includes copy number polymorphisms (CNPs), which are common in the general population and occur at an overall frequency greater than 1%. CNPs are typically small (most less than 10 kilobases in length) and are often enriched for genes encoding proteins important for drug detoxification and immunity. A subset of these CNPs is highly variable in copy number. As a result, different human chromosomes can have a wide range of copy numbers (e.g., 2, 3, 4, 5, etc.) for a particular set of genes. Recently, CNPs associated with immune response genes have been associated with susceptibility to complex genetic disorders, including psoriasis, Crohn's disease, and glomerulonephritis.

[0151] The second type of CNV includes relatively rare variants, which are much longer than CNPs, ranging in length from hundreds of thousands of base pairs to more than one million base pairs.In some cases, these CNVs may have occurred during the sperm or egg production that gave birth to a particular individual, or may have only been passed down within a few generations within a family.These large and rare structural variants are disproportionately observed in subjects with mental retardation, developmental delay, schizophrenia, and autism.Their occurrence in such subjects has led to speculation that large and rare CNVs may be more important in neurocognitive diseases than other inherited variants, including single-base substitutions.

[0152] Gene copy number can be altered in cancer cells. For example, Chrlp duplication is common in breast cancer, and EGFR copy number can be higher than normal in non-small cell lung cancer. Additional target genes for cancer diagnosis are provided in Section 9 of this specification. Cancer is one of the leading causes of death. Therefore, early diagnosis and treatment of cancer is important and can improve patient outcomes (e.g., increasing the probability of remission and the duration of remission). Early diagnosis also provides patients with fewer drastic treatment options. Many current treatments that destroy cancerous cells also affect normal cells, resulting in a variety of potential side effects, such as nausea, vomiting, low blood counts, increased risk of infection, hair loss, and mucosal ulcers. Therefore, early detection of cancer is desirable, thereby reducing the amount and / or number of treatments (e.g., chemotherapy or radiation) required to eliminate the cancer.

[0153] Copy number variation is also associated with severe mental and physical disabilities and idiopathic learning disabilities.Non-invasive prenatal testing (NIPT) using cell-free DNA (cfDNA) can be used to detect abnormalities such as fetal trisomy 13, 18 and 21, triploidy and sex chromosome aneuploidy.Therefore, in one embodiment, the target chromosome or target chromosome segment of interest is chromosome 13, chromosome 18, chromosome 21, sex chromosome, and / or their chromosome segments.

[0154] In addition, subchromosomal microdeletions, which may cause severe mental and physical disabilities, are more difficult to detect due to their small size. Eight of the microdeletion syndromes have a total incidence rate of more than 1 in 1000, which is approximately the same frequency as fetal autosomal trisomy. Thus, in one embodiment, the method disclosed herein further comprises detecting a microdeletion in one or more circulating cells determined to be one or more pure fetal cells. In one embodiment, the microdeletion is a 22q11.2 deletion associated with DiGeorge syndrome, a microdeletion associated with Prader-Willi syndrome, a microdeletion associated with Angelman syndrome, a 1p36 deletion, and / or a microdeletion associated with Criss-Cat syndrome.

[0155] Furthermore, high copy numbers of CCL3L1 are associated with reduced susceptibility to HIV infection, and low copy numbers of FCGR3B (CD16 cell surface immunoglobulin receptor) may increase susceptibility to systemic lupus erythematosus and similar inflammatory autoimmune disorders.

[0156] The method for measuring chromosome copy number in fetal cells based on counting the number of reads from the DNA sequence mapped to a given chromosome or chromosome segment is conveniently referred to as "counting method" or "quantification method" for analyzing chromosome copy number or chromosome segment copy number. Examples of such methods can be found in, among others, published patent application US2013 / 0172211A1, U.S. Patent No. 8,008,018, U.S. Patent No. 8,467,976B2, and U.S. published patent application US2012 / 0003637A1. In many cases, such methods involve generating a reference value (cutoff value) for the number of reads from the DNA sequence mapped to a specific chromosome, and if the number of reads exceeds the value, it indicates a specific genetic abnormality.

[0157] Therefore, the method for analyzing and characterizing a cell sample suspected to be of fetal origin may be combined with a method and system for determining copy number or a method and system for detecting aneuploidy of a chromosome or chromosome segment of interest in a target cell confirmed to be a fetal cell. These methods for determining copy number or detecting aneuploidy are performed using a chromosome or chromosome segment of interest to set a bias model. That is, to set test parameters, a sample identified as a diploid sample with high confidence is used in the same parallel analysis to analyze the aneuploidy of the same chromosome or chromosome segment of interest in other samples in a test sample set. Confidence refers to the statistical likelihood that the classified SNP, allele, set of alleles, ploidy classification, or determined copy number of a chromosome segment correctly represents the actual genetic status of an individual. Such a method for determining copy number variation and aneuploidy from fetal DNA is described in U.S. Patent Publication No. 2018 / 0173846, the entire contents of which are incorporated herein.

[0158] The abundance of each locus detected by sequencing DNA preparations obtained from target and non-target cells may vary for reasons other than the starting abundance of the locus in the initial sample material prior to preparation for sequencing, e.g., prior to amplification steps such as PCR. Variables such as PCR primer binding efficiency, amplicon length, and GC content can contribute to variations in the representation of individual loci during sequencing or sequencing. Such factors can introduce locus-specific bias, resulting in over- or under-representation of individual loci. Further bias can arise from sample-specific discrepancies. For example, one sample may have more DNA than another sample in a sample set due to pipetting or other measurement errors during physical processing of the sample. In exemplary embodiments, these sample-specific biases are accounted for by sample-specific parameters. These specific sample-specific parameter embodiments are disclosed in U.S. Patent Publication No. 2018 / 0173846, which is incorporated herein in its entirety.

[0159] Regardless of the specific method used to generate genetic information, the amount of genetic sequence information from each locus depends on the relative amount of copy numbers of the locus in the original sample. Loci that are considered to be on the same chromosomal segment, or in some embodiments, loci that are considered to be on the same chromosome, are assumed to have the same starting amount. Thus, for example, multiple loci present on chromosome 21 of the genome of a target cell (or the genome of a non-target cell) are assumed to be present in approximately equal amounts in genomic DNA. Therefore, differences in the observed amount of genetic information between loci on the same chromosome are the result of locus-specific bias. For example, if SNP1 and SNP2 are assumed to be located on the same chromosome and have the same copy number on the same chromosome, and SNP1 is found to have a read depth of 0.1% and SNP2 is found to have a read depth of 0.4%, this can be explained by a quantitative locus-specific bias that favors the production of DNA sequences from SNP2 over SNP1. This bias can be further normalized by considering the distribution of possible sampling results for two different SNPs. Thus, the methods of the present invention analyze bias and provide bias models, as discussed in more detail in U.S. Patent Publication 2018 / 0173846, which is incorporated herein in its entirety.

[0160] In certain embodiments of the present invention, one or two maximum likelihood methods are used. As described above, a first maximum likelihood method can be used to identify diploid samples in a sample set and determine a first probability that other samples in the sample set are aneuploid. Thus, in certain embodiments, one or more or all of the chromosomes or chromosome segments of interest are determined to be disomic using the first maximum likelihood method. This method is further described and illustrated in U.S. Patent Publication No. 2018 / 0173846, which is incorporated herein in its entirety.

[0161] In another embodiment, the copy number of the target chromosome or chromosome segment is determined by generating a plurality of second ploidy hypotheses or a set of second ploidy hypotheses, where each second ploidy hypothesis is associated with a specific copy number of the target chromosome or chromosome segment in the target cell.The model is then used to test how well the genetic data from each patient fits each of the second hypotheses.The degree of fit for each second hypothesis is determined.A second probability value is calculated for each second hypothesis, where the second probability value indicates the likelihood that the genome of the target cell has the number of chromosomes or chromosome segments specified by the second hypothesis.Therefore, the copy number of the chromosome or chromosome segment in the genome of the target cell can be determined by selecting the second hypothesis with the greatest likelihood.These first and second hypotheses can be considered in combination to increase the reliability of aneuploidy determination, as discussed in detail in U.S. Patent Publication No. 2018 / 0173846, the entire contents of which are incorporated herein.

[0162] In one embodiment of the present disclosure, the method for determining the ploidy status of a fetus further includes taking into account the percentage of fetal DNA in the sample. In one embodiment of the present disclosure, the method involves calculating the percentage of DNA in the sample that is of fetal or placental origin. In one embodiment of the present disclosure, the aneuploidy classification threshold is adaptively normalized based on the calculated fetal DNA percentage. In some embodiments, a method for estimating the percentage of DNA that is of fetal origin in a DNA mixture includes obtaining a mixed sample containing maternal genetic material and fetal genetic material, obtaining a genetic sample from the father of the fetus, measuring DNA in the mixed sample, measuring DNA in the father's sample, and calculating the percentage of DNA that is of fetal origin in the mixed sample using the DNA measurement results of the mixed sample and the father's sample.

[0163] In one embodiment, the methods described herein allow for confirmation that a cell sample contains both true fetal and true maternal cells, and the fetal fraction can be estimated by using cell-free DNA (cfDNA) as a reference. As described in U.S. Patent Application Publication No. 2018 / 0298439, which is incorporated herein in its entirety, the genotype data of the maternal DNA, the mixture of fetal and maternal DNA, and the estimated fetal DNA fraction can be used to determine the paternity of the fetus. The genotype data obtained from the mixture of the mother's and child's cfDNA and the mixture of the fetal and maternal DNA can be used to determine the fetal fraction of the sample, enabling paternity testing.

[0164] Some embodiments may be used in combination with the PARENTAL SUPPORT™ (PS) method, which is described in U.S. Patent Application No. 11 / 603,406 (U.S. Publication No. 20070184467), U.S. Patent Application No. 12 / 076,348 (U.S. Publication No. 20080243398), U.S. Patent Application No. 13 / 110,685, PCT Application No. PCT / US09 / 52730 (PCT Publication No. WO / 2010 / 017214), PCT Application No. PCT / US10 / 050824 (PCT Publication No. WO / 2011 / 041485), PCT Application No. PCT / US2011 / 037018 (PCT Publication No. WO / 2011 / 146632), and PCT Application No. PCT / US2011 / 61506, which applications are incorporated herein by reference in their entireties. PARENTAL SUPPORT™ is an informatics-based method that can be used to analyze genetic data. In some embodiments, the methods disclosed herein may be considered part of the PARENTAL SUPPORT™ method. In some embodiments, the PARENTAL SUPPORT™ method is a collection of methods that can be used to determine with high accuracy the genetic data of a target individual, a single cell or a small number of cells from that individual, or a DNA mixture consisting of the target individual's DNA, or the DNA of one or more other individuals, particularly to determine disease-associated alleles, determine alleles of interest, determine the ploidy state of one or more chromosomes in the target individual, and / or determine the degree of relatedness of the target individual to another individual. PARENTAL SUPPORT™ may refer to any of these methods. PARENTAL SUPPORT™ is an example of an informatics-based method.

[0165] In some embodiments of the present invention, improved reliability in determining aneuploidy can be achieved by determining the aneuploidy of a sample using a quantitative non-allelic threshold or cutoff method, and then determining the aneuploidy of the same sample using a likelihood determination method. If a sample is identified as having aneuploidy in a chromosome or chromosome segment of interest using a threshold method, and the sample is identified as having aneuploidy with high confidence using a likelihood determination for a set of hypotheses, the sample is identified as having aneuploidy in the chromosome or chromosome segment of interest for one or more target cells in the subject from which the sample was obtained. This method is further described and illustrated in U.S. Patent Publication No. 2018 / 0173846, which is incorporated herein in its entirety.

[0166] The method of performing non-allele threshold analysis, especially NIPT, is known in the art.For example, United States Patent No. 7,888,017 and United States Patent No. 8,318,430, which are incorporated herein by reference in their entirety, provide a method for determining fetal aneuploidy by counting the number of reads that are mapped to estimated chromosome, and comparing this with the number of reads that are mapped to reference chromosome, and using the assumption that if there is excess read on estimated chromosome, this corresponds to the triploidy of fetus at this chromosome.The teachings in these documents, including the depth and reference value of sequence read, can be useful for carrying out the embodiments of the present invention.

[0167] 7. Methods for analyzing genomic information from circulating tumor cells. The methods and compositions provided herein improve the detection, diagnosis, staging, screening, treatment, and management of cancer (e.g., breast, bladder, or colorectal cancer) by using liquid biopsy to obtain a sample containing circulating tumor cells. In illustrative embodiments, the methods provided herein analyze cancer-associated mutations in circulating fluids, particularly circulating tumor cells (CTCs). The methods offer the advantage of identifying most mutations found in tumors, not only clonal mutations but also subclonal mutations, in a single test utilizing a tumor sample, all of which are valid, rather than the multiple tests previously required.

[0168] Thus, in one aspect, the present disclosure provides a method for monitoring and detecting early recurrence or metastasis of a tumor in a cancer patient by obtaining one or more circulating cells determined to be derived from the tumor, the method comprising: a) selecting one or more patient-specific mutations based on mutations identified in tumor samples of patients diagnosed with cancer; b) longitudinally collecting one or more blood samples from the patient after the patient has been treated with surgery, first-line chemotherapy, and / or adjuvant therapy; c) generating a first set of amplicons from cellular DNA isolated from one or more circulating cells suspected of being circulating tumor cells and a second set of amplicons from cell-free DNA obtained from the blood sample obtained from the cancer patient, wherein the tumor cells and normal cells are obtained from each of the cancer patient's blood sample or fractions thereof, and the first set of amplicons and the second set of amplicons are obtained by performing a multiplex amplification reaction of patient-specific mutations associated with cancer; d) sequencing the first set of amplicons and the second set of amplicons by next generation sequencing; e) determining the origin of one or more circulating cells based on the sequences of the first amplicon set, wherein the sequences of the second amplicon set are used as a reference, and detection of one or more patient-specific mutations from the first amplicon set generated from circulating cells determined to be of tumor origin is indicative of early recurrence or metastasis of the cancer.

[0169] In another aspect, the disclosure herein provides a method for treating a cancer patient, the method comprising: a) treating a cancer patient with surgery, first-line chemotherapy, and / or adjuvant therapy; b) longitudinally collecting one or more blood samples from the patient after the patient has been treated with surgery, first-line chemotherapy, and / or adjuvant therapy; c) generating a first set of amplicons from cellular DNA isolated from one or more circulating cells suspected of being circulating tumor cells and a second set of amplicons from cell-free DNA obtained from the blood sample obtained from the cancer patient, wherein the tumor cells and normal cells are obtained from each of the cancer patient's blood sample or fractions thereof, and the first set of amplicons and the second set of amplicons are obtained by performing a multiplex amplification reaction of patient-specific mutations associated with cancer; d) sequencing the first set of amplicons and the second set of amplicons by next generation sequencing; e) determining the origin of one or more circulating cells based on the sequences of the first amplicon set, wherein the sequences of the second amplicon set are used as a reference, and detection of one or more patient-specific mutations from the first amplicon set generated from circulating cells determined to be of tumor origin is indicative of early recurrence or metastasis of the cancer; f) administering to said patient a compound, said compound being found to be effective in treating a cancer having said one or more patient-specific mutations detected in a blood sample.

[0170] The improved methods provided herein allow for accurate analysis and characterization of samples containing very small numbers of CTCs. The number of CTCs in a sample analyzed by the methods of the present invention can be less than 100 cells, less than 75 cells, less than 50 cells, less than 25 cells, less than 20 cells, less than 15 cells, less than 10 cells, less than 5 cells, or even one cell. Using the methods provided herein, accurate genomic information can be obtained from one circulating tumor cell, two CTCs, three CTCs, four CTCs, five CTCs, six CTCs, seven CTCs, eight CTCs, nine CTCs, or ten CTCs.

[0171] In some embodiments, the method includes monitoring and detecting early tumor recurrence or metastasis in a cancer patient by determining the identity of circulating cells suspected to be circulating tumor cells based on predetermined patient-specific mutations. In some embodiments, the patient-specific mutations include cancer-associated single nucleotide variants (SNVs), copy number variants (CNVs), indels, deletions, or gene fusions. In some embodiments, the presence of at least two cancer-associated patient-specific mutations indicates that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least eight cancer-associated patient-specific mutations indicates that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least 16 cancer-associated patient-specific mutations indicates that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least 50 cancer-associated patient-specific mutations indicates that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least 100 cancer-associated patient-specific mutations indicates that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least 1000 cancer-associated patient-specific mutations indicates that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least 5,000 patient-specific mutations associated with cancer indicates that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least 10,000 patient-specific mutations associated with cancer indicates that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least 15,000 patient-specific mutations associated with cancer indicates that the one or more circulating cells are tumor cells.

[0172] In some embodiments, the presence of at least 100 to about 1000 cancer-associated patient-specific mutations indicates that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least 100 to about 1000 cancer-associated patient-specific mutations indicates that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least 1000 to about 5000 cancer-associated patient-specific mutations indicates that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least 5000 to about 10,000 cancer-associated patient-specific mutations indicates that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least 10,000 to about 15,000 cancer-associated patient-specific mutations indicates that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least 15,000 to about 20,000 cancer-associated patient-specific mutations indicates that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least 20,000 to about 25,000 cancer-associated patient-specific mutations indicates that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least 50,000 cancer-associated patient-specific mutations indicates that the one or more circulating cells are tumor cells.

[0173] In some embodiments, the presence of at least 5,000 patient-specific mutations associated with cancer indicates that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least 10,000 patient-specific mutations associated with cancer indicates that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least 15,000 patient-specific mutations associated with cancer indicates that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least 25,000 patient-specific mutations associated with cancer indicates that the one or more circulating cells are tumor cells.

[0174] In one aspect, multiplex amplification is performed to obtain amplicons containing patient-specific mutations associated with cancer. In some embodiments, the first set of amplicons and the second set of amplicons are obtained by performing a multiplex amplification reaction of 50 to 1000 cancer-associated patient-specific mutations.

[0175] In one aspect, a method for monitoring and detecting early tumor recurrence and metastasis in cancer patients includes separating a plurality of circulating cells into individual reaction volumes, isolating cellular DNA in each reaction volume, and attaching a sample barcode to the isolated cellular DNA in each reaction volume. In some embodiments, the cellular DNA is isolated from a single circulating cell suspected to be a tumor cell.

[0176] In some embodiments, the method further comprises performing a non-invasive cancer test using the genotype in which the one or more circulating cells are determined to be one or more tumor cells.

[0177] The methods and compositions may be useful by themselves or may be useful when used in conjunction with other methods for the detection, diagnosis, staging, screening, treatment and management of cancer (e.g., breast, bladder or colorectal cancer), e.g., to support the results of these other methods and provide more reliable and / or conclusive results.

[0178] The terms "cancer" and "cancerous" refer to or describe the physiological condition in animals that is typically characterized by uncontrolled cell growth. A "tumor" contains one or more cancerous cells. There are several main types of cancer. Carcinoma is cancer that begins in the skin or in the tissues that line or cover internal organs. Sarcoma is cancer that begins in bone, cartilage, fat, muscle, blood vessels, or other connective or supportive tissue. Leukemia is cancer that begins in blood-forming tissues, such as the bone marrow, and causes large numbers of abnormal blood cells to be produced and enter the bloodstream. Lymphoma and multiple myeloma are cancers that begin in cells of the immune system. Central nervous system cancer is cancer that begins in the tissues of the brain and spinal cord.

[0179] In some embodiments, the cancer is acute lymphoblastic leukemia, acute myeloid leukemia, adrenocortical carcinoma, AIDS-related cancer, AIDS-related lymphoma, anal cancer, appendix cancer, astrocytoma, atypical teratoid / rhabdoid tumor, basal cell carcinoma, bladder cancer, brain stem glioma, brain tumor (brain stem glioma, atypical teratoid / rhabdoid tumor of the central nervous system, embryonal tumor of the central nervous system, astrocytoma, craniopharyngioma, ependymoblastoma, ependymoma, medulloblastoma, medulloepithelioma, pineal gland tumor, Intermediately differentiated parenchymal tumors, supratentorial primitive neuroectodermal tumors, and pineoblastoma), breast cancer, bronchial tumors, Burkitt's lymphoma, carcinoma of unknown primary site, carcinoid tumor, carcinoma of unknown primary site, atypical teratoid / rhabdoid tumor of the central nervous system, embryonal tumor of the central nervous system, cervical cancer, childhood cancer, chordoma, chronic lymphocytic leukemia, chronic myeloid leukemia, chronic myeloproliferative disorder, colon cancer, colorectal cancer, craniopharyngioma, cutaneous T-cell lymphoma, endocrine pancreatic Islet cell tumor, endometrial cancer, ependymoblastoma, ependymoma, esophageal cancer, nasal neuroblastoma, Ewing's sarcoma, extracranial germ cell tumor, extragonadal germ cell tumor, extrahepatic bile duct cancer, gallbladder cancer, gastric cancer, gastrointestinal carcinoid tumor, gastrointestinal stromal cell tumor, gastrointestinal stromal tumor (GIST), gestational trophoblastic tumor, glioma, hairy cell leukemia, head and neck cancer, cardiac cancer, Hodgkin's lymphoma, hypopharyngeal cancer, intraocular melanoma, islet cell tumor, Kaposi's sarcoma, renal cancer, Langerhans cell histiocytosis Cancer of the nasal cavity, nasopharyngeal cancer, liver cancer, malignant fibrous histiocytoma, bone cancer, medulloblastoma, medulloepithelioma, melanoma, Merkel cell carcinoma, Merkel cell skin cancer, mesothelioma, metastatic squamous cell carcinoma of the neck of unknown primary origin, oral cancer, multiple endocrine neoplasia syndrome, multiple melanoma, multiple melanoma / plasmacytoma, mycosis fungoides, myelodysplastic syndrome, myeloproliferative neoplasm, nasal cancer, nasopharyngeal cancer, neuroblastoma, non-Hodgkin's lymphoma, non-melanoma skin cancer, non-small cell lung cancer, oral cancer, oral cavity cancercancer), oropharyngeal cancer, osteosarcoma, other brain and spinal cord tumors, ovarian cancer, epithelial ovarian cancer, ovarian germ cell tumor, ovarian low malignant potential tumor, pancreatic cancer, papillomatosis, sinonasal cancer, parathyroid cancer, pelvic cancer, penile cancer, pharyngeal cancer, intermediately differentiated pineal parenchymal tumor, pineoblastoma, pituitary tumor, plasmacytoma / multiple myeloma, pleuropulmonary blastoma, primary central nervous system (CNS) lymphoma, primary hepatocellular carcinoma, prostate cancer, rectal cancer, renal cancer, renal cell (kidney) cancer, renal cell carcinoma, respiratory tract cancer, retinoblastoma, rhabdomyosarcoma, salivary gland cancer, Sezary syndrome, small cell lung cancer, small intestine cancer, soft tissue sarcoma, squamous cell carcinoma, squamous cell cervical carcinoma, gastric cancer, supratentorial primitive neuroectodermal tumor, T-cell lymphoma, testicular cancer, throat cancer, thymic carcinoma, thymoma, thyroid cancer, transitional cell carcinoma, transitional cell carcinoma of the renal pelvis and ureter, trophoblastic tumor, ureteral cancer, urethral cancer, uterine cancer, uterine sarcoma, vaginal cancer, vulvar cancer, Waldenstrom's macroglobulinemia, or Wilms' tumor.

[0180] In other embodiments, provided herein are methods for detecting cancer (e.g., breast cancer, bladder cancer, or colorectal cancer) in a blood sample or fraction thereof from an individual, e.g., an individual suspected of having cancer, the method comprising determining single base variants present in a sample containing circulating tumor cells and a ctDNA sample using the ctDNA SNV amplification / sequencing workflow provided herein. The presence of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 SNVs at the lower end of the range, or 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 40 or 50 SNVs at the upper end of the range in a sample at multiple single base loci indicates the presence of cancer (e.g., breast cancer, bladder cancer, or colorectal cancer).

[0181] In certain examples of such embodiments, the cancer is stage 1a, 1b, or 2a breast cancer, bladder cancer, or colorectal cancer. In certain examples of such embodiments, the cancer is stage 1a or 1b breast cancer, bladder cancer, or colorectal cancer. In certain examples of such embodiments, the individual has not undergone surgery. In certain examples of such embodiments, the individual has not undergone a biopsy. In certain examples of such embodiments, prior to performing targeted amplification on CTCs from an individual, data is provided regarding SNVs present in tumors derived from the individual. Thus, in these embodiments, an SNV amplification / sequencing reaction is performed on one or more tumor samples from the individual. In this method, the CTC SNV amplification / sequencing reaction provided herein is also advantageous because it provides a liquid biopsy of clonal and subclonal mutations. Additionally, as provided herein, in an individual with cancer (e.g., breast cancer, bladder cancer, or colorectal cancer), a clonal mutation can be more clearly identified if a high percentage of VAF, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10% VAF or greater, is determined for a given SNV in a ctDNA sample from the individual.

[0182] In an illustrative embodiment, the set of single base variant loci of any of the methods herein includes all of the single base variant loci identified in the TCGA and COSMIC datasets for cancer (e.g., breast cancer, bladder cancer, or colorectal cancer).

[0183] In certain embodiments of any of the methods herein, the set of single nucleotide variant loci includes, at the lower end of the range, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 75, 100, 250, 500, 1000, 2500, 5000, or 10,000 single nucleotide variant loci known to be associated with cancer (e.g., breast cancer, bladder cancer, or colorectal cancer), and at the higher end of the range, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 75, 100, 250, 500, 1000, 2500, 5000, 10,000, 20,000, and 25,000 single nucleotide variant loci known to be associated with cancer (e.g., breast cancer, bladder cancer, or colorectal cancer).

[0184] In another embodiment, provided herein is a method for supporting a diagnosis of cancer (e.g., breast cancer, bladder cancer, or colorectal cancer) in an individual, such as an individual suspected of having cancer (e.g., breast cancer, bladder cancer, or colorectal cancer), from a sample of blood or a fraction thereof from the individual, the method comprising performing a CTC SNV amplification / sequencing workflow provided herein to determine whether one or more single nucleotide variants are present in a plurality of single nucleotide variant loci. In such embodiments, the following factors, opinions, guidelines, or rules apply: the absence of a single nucleotide variant supports a diagnosis of stage 1a, 1b, or 2a adenocarcinoma; the presence of a single nucleotide variant supports a diagnosis of squamous cell carcinoma or stage 2b or 3a adenocarcinoma; and / or the presence of 10 or more single nucleotide variants supports a diagnosis of squamous cell carcinoma or stage 2b or 3 adenocarcinoma.

[0185] In some embodiments, the method further includes determining a treatment regimen, a therapy, and / or administering to the individual a compound that targets one or more clonal single nucleotide variants. In some examples, subclonal SNVs and / or other clonal SNVs are not targeted by the therapy. Some treatments and associated mutations are provided elsewhere herein and are known in the art. Thus, in some examples, the method further includes administering to the individual a compound that is known to be particularly effective in treating cancers (e.g., breast, bladder, or colorectal cancer) that have one or more of the determined single nucleotide variants.

[0186] In some embodiments, the methods for detecting SNVs described herein can be used to guide treatment regimens. Treatments targeting specific mutations associated with ADC and SCC are available and under development ( Nature Review Cancer. 14:535-551 (2014)). For example, detection of EGFR mutations at L858R or T790M can be beneficial for treatment selection. Erlotinib, gefitinib, afatinib, AZK9291, CO-1686, and HM61713 are currently approved in the United States or in clinical trials and target specific EGFR mutations. In another example, a G12D, G12C, or G12V mutation in KRAS can be used to treat an individual with a combination of selumetinib and docetaxel. In another example, a V600E mutation in BRAF can be used to treat a subject with vemurafenib, dabrafenib, and trametinib.

[0187] In exemplary embodiments, the target genes of the present invention are cancer-associated genes, and in many illustrative embodiments, they are cancer-associated genes. A cancer-associated gene (e.g., a cancer-associated gene or a bladder cancer-associated gene or a colorectal cancer-associated gene) refers to a gene associated with an altered risk of cancer (e.g., breast cancer, bladder cancer, or colorectal cancer) or a gene associated with an altered prognosis of cancer. Exemplary cancer-associated genes that promote cancer include oncogenes, genes that enhance cell proliferation, invasion, or metastasis, genes that suppress apoptosis, and pro-angiogenesis genes. Cancer-associated genes that suppress cancer include, but are not limited to, tumor suppressor genes, genes that suppress cell proliferation, invasion, or metastasis, genes that promote apoptosis, and anti-angiogenesis genes.

[0188] One embodiment of the mutation detection method begins with selecting a region of a gene to target, using the region containing a known mutation to amplify primers for mPCR-NGS to amplify and detect the mutation.

[0189] The methods provided herein can be used to detect virtually any type of mutation, particularly mutations known to be associated with cancer, and in particular the methods provided herein are directed to mutations, particularly SNVs, associated with cancer, particularly breast, bladder or colorectal cancer. Exemplary SNVs may occur in one or more of the following genes: EGFR, FGFR1, FGFR2, ALK, MET, ROS1, NTRK1, RET, HER2, DDR2, PDGFRA, KRAS, NF1, BRAF, PIK3CA, MEK1, NOTCH1, MLL2, EZH2, TET2, DNMT3A, SOX2, MYC, KEAP1, CDKN2A, NRG1, TP53, LKB1, and PTEN, which have been identified as mutated, gaining copy number, or fused with other genes, and combinations thereof, in various lung cancer samples (Non-small-cell lung cancers: a heterogeneous set of diseases. Chen et al. Nat. Rev. Cancer. 2014 Aug 14(8):535-551). In other examples, the list of genes is as listed above, and the SNVs are as reported in the cited Chen et al. references, etc.

[0190] 8. Exemplary Embodiments of Analysis Methods 8.1 Exemplary Embodiments for Detecting Cancer In certain embodiments, the analyzing step in the method for determining whether circulating tumor nucleic acid is present comprises analyzing a set of chromosomal segments known to exhibit aneuploidy in cancer. In certain embodiments, the analyzing step in the method for determining whether circulating tumor nucleic acid is present comprises analyzing 1,000 to 50,000 or 100 to 1000 polymorphic loci related to ploidy. In certain embodiments, the analyzing step in the method for determining whether circulating tumor nucleic acid is present comprises analyzing 100 to 1000 single-base variant sites. For example, in these embodiments, the analyzing step comprises performing multiplex PCR to amplify amplicons across 1000 to 50,000 polymer loci and 100 to 1000 single-base variant sites. This multiplex reaction can be configured as a single reaction or as a pool of different subsets of multiplex reactions. The multiplex reactions provided herein, such as the massively multiplex PCR disclosed herein, provide exemplary processes for performing amplification reactions that improve multiplexing and, therefore, help achieve improved levels of sensitivity.

[0191] In certain embodiments, multiplex PCR reactions are performed under limiting primer conditions for at least 10%, 20%, 25%, 50%, 75%, 90%, 95%, 98%, 99%, or 100% of the reactions. Improved conditions for performing large-scale multiplex reactions provided herein can be used.

[0192] In certain aspects, the above-described methods for determining whether circulating tumor nucleic acid is present in a sample in an individual, and all embodiments thereof, can be implemented using a system. The present disclosure provides teachings regarding specific functional and structural characteristics for implementing the methods. As non-limiting examples, systems include:

[0193] an input processor configured to analyze data from the sample to determine ploidy at a set of polymorphic loci on a chromosomal segment in the individual; and

[0194] A modeler configured to determine the level of allelic imbalance present at a polymorphic locus based on a ploidy determination, where an allelic imbalance of 0.5% or greater indicates the presence of cycling.

[0195] 8.2 Exemplary Embodiments for Detecting Single-Base Variants In certain aspects, provided herein are methods for detecting single-base variants. The improved methods provided herein can achieve a detection limit of 0.015, 0.017, 0.02, 0.05, 0.1, 0.2, 0.3, 0.4, or 0.5 percent SNVs in a sample. All embodiments for detecting SNVs can be implemented using a system. The present disclosure provides teachings regarding specific functional and structural characteristics for implementing the methods. Further provided herein are embodiments that include a non-transitory computer-readable medium containing computer-readable code that, when executed by a processing device, causes the processing device to perform the SNV detection method provided herein.

[0196] Thus, in one embodiment, there is provided herein a method for determining whether a single base variant is present at a set of genomic locations in a sample from an individual, the method comprising: generating, for each genomic location, an estimate of the efficiency and error rate per cycle of the amplicon spanning the genomic location; using a training dataset; receiving observed nucleotide identity information for each genomic location in the sample; determining a set of probabilities for the percentage of single base variants arising from one or more true mutations at each genomic location using the observed nucleotide identity information at each genomic location and a model of the percentage of various variants, independently estimated amplification efficiencies and error rates per cycle for each genomic location; and determining a percentage and confidence of the most likely true variant from the set of probabilities for each genomic location.

[0197] In an exemplary embodiment of the method for determining whether single base variants exist, efficiency and per-cycle error rate estimates are generated for the amplicon set spanning genome position.For example, 2, 3, 4, 5, 10, 15, 20, 25, 50, 100 or more amplicons spanning genome position can be included.

[0198] In an exemplary embodiment of the method for determining whether a single base variant is present, the observed nucleotide identity information includes the total number of reads observed for each genomic location and the number of reads of the variant allele observed for each genomic location.

[0199] In an exemplary embodiment of the method for determining whether a single base variant is present, the sample is a plasma sample and the single base variant is present in circulating tumor DNA of the sample. In an exemplary embodiment of the method for determining whether a single base variant is present, the sample is a plasma sample and the single base variant is present in circulating fetal DNA of the sample.

[0200] In another embodiment, provided herein is a method for estimating the percentage of single base variants present in a sample from an individual, comprising: generating, at a set of genomic locations, estimates of efficiencies and error rates per cycle for one or more amplicons spanning the genomic locations; using a training dataset; receiving measured nucleotide identity information for each genomic location in the sample; using the amplification efficiencies and error rates per cycle of the amplicons to generate estimated means and variances for the total number of molecules, background error molecules, and true mutant molecules for a search space containing an initial percentage of true mutant molecules; and determining the percentage of single base variants present in the sample that result from true mutations by fitting the distribution using the estimated means and variances to the measured nucleotide identity information in the sample to determine the most likely percentage of true single base variants.

[0201] The training dataset for this embodiment of the present invention preferably includes samples from a healthy population. In certain exemplary embodiments, the training dataset is analyzed on the same day or even in the same run as one or more test samples. For example, a training dataset can be generated using samples from 2, 3, 4, 5, 10, 15, 20, 25, 30, 36, 48, 96, 100, 192, 200, 250, 500, 1000, or more healthy individuals. If data from more healthy individuals, e.g., 96 or more, are available, confidence in estimating amplification efficiency increases, even if runs are performed before the method is performed on the test samples. Because the PCR error rate is an error rate per amplicon, nucleic acid sequence information generated not only for the position of the SNV base but also for the entire amplified region surrounding the SNV can be used. For example, using samples from 50 individuals and sequencing 20-base-pair amplicons surrounding the SNV, error frequency data from 1000-base reads can be used to determine the error frequency rate.

[0202] Typically, amplification efficiency is estimated by estimating the mean and standard deviation of the amplification efficiency for the amplified segments, and then fitting it to a distribution model, such as a binomial or beta-binomial distribution. The error rate is determined for a PCR reaction with a known number of cycles, and then the error rate per cycle is estimated.

[0203] In certain exemplary embodiments, estimating starting molecules for the test dataset further comprises updating the efficiency estimate for the test dataset using the number of starting molecules estimated in step (b) if the observed number of reads differs significantly from the estimated number of reads. The estimate can then be updated for the new efficiency and / or starting molecules.

[0204] The search space used to estimate the total number of molecules, background error molecules, and true mutant molecules can include a search space from 0.1%, 0.2%, 0.25%, 0.5%, 1%, 2.5%, 5%, 10%, 15%, 20%, or 25% for the lower limit of the copy number of the base at the SNV position, which is the SNV base, and 1%, 2%, 2.5%, 5%, 10%, 12.5%, 15%, 20%, 25%, 50%, 75%, 90%, or 95% for the upper limit. A low range of 0.1%, 0.2%, 0.25%, 0.5%, or 1% for the lower limit and a high range of 1%, 2%, 2.5%, 5%, 10%, 12.5%, or 15% for the upper limit can be used in an exemplary embodiment for plasma samples in which the method detects circulating tumor DNA. A high range is used for tumor samples.

[0205] A distribution is fitted to the total number of error molecules (background errors plus true mutations) among the total molecules, and the likelihood or probability for each possible true mutation in the search space is calculated. This distribution can be a binomial or beta-binomial distribution.

[0206] The most likely true mutation is determined by determining the percentage of the most likely true mutation and calculating the confidence level using the data from the distribution fit.As an illustrative example, and not limiting the clinical interpretation of the method provided herein, when the average mutation rate is high, the confidence level required to make a positive SNV determination is low.For example, if the average mutation rate of SNVs in samples using the most likely hypothesis is 5% and the confidence level is 99%, a positive SNV classification is made.On the other hand, for this illustrative example, if the average mutation rate of SNVs in samples using the most likely hypothesis is 1% and the confidence level is 50%, in certain circumstances, a positive SNV classification is not made.It should be understood that the clinical interpretation of data is a function of sensitivity, specificity, prevalence, and the availability of alternatives.

[0207] In another embodiment, provided herein is a method for detecting one or more single base variants in a test sample from an individual, the method according to this embodiment comprising the steps of:

[0208] determining a median variant allele frequency of a plurality of control samples from each of a plurality of normal individuals for each positional base variant position in the set of single base variant positions based on results generated in the sequencing run, identifying selected single base variants having a median allele variant frequency in the normal samples below a threshold, and determining a background error for each of the single base variant positions after removing outlier samples for each of the single base variant positions; determining an observed read depth-weighted mean and variance for the selected single base variant positions for the test sample based on data generated in the sequencing run of the test sample; and using a computer to identify one or more single base variant positions with a read depth-weighted mean that is statistically significant compared to the background error for that position, thereby detecting one or more single base variants.

[0209] In certain embodiments of the subject methods for detecting one or more SNVs, the plurality of control samples comprises at least 25 samples. In certain exemplary embodiments, the plurality of control samples is at least 5, 10, 15, 20, 25, 50, 75, 100, 200, or 250 samples for the lower limit and 10, 15, 20, 25, 50, 75, 100, 200, 250, 500, and 1000 samples for the upper limit.

[0210] In certain embodiments of the method for detecting one or more SNVs, outliers are removed from data generated by high-throughput sequencing, and an average weighted by observed read depth and observed variance are determined. In certain embodiments of the method for detecting one or more SNVs, the read depth for each single-base variant position in the test sample is at least 100 reads.

[0211] In certain embodiments of the present methods for detecting one or more SNVs, the sequencing run includes a multiplexed amplification reaction performed under defined primer reaction conditions. These embodiments are performed in exemplary embodiments using improved methods for performing multiplexed amplification reactions provided herein.

[0212] Without being limited by theory, the method of the present invention utilizes a background error model using normal plasma samples, sequenced in the same sequencing run as the test sample, to account for run-specific artifacts. Noisy positions with median normal variant allele frequencies above thresholds, such as >0.1%, 0.2%, 0.25%, 0.5%, 0.75%, and 1.0%, are removed.

[0213] Outlier samples are repeatedly removed from the model that accounts for noise and contamination.For each base substitution of all genome loci, the read depth-weighted average and error standard deviation are calculated.In certain exemplary embodiments, the sample that comprises at least a threshold number of reads, for example, at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 250, 500 or 1000 variant reads, and in certain embodiments, the a1 Z score is higher than 2.5, 5, 7.5 or 10 against background error model, is counted as candidate mutation, for example, circulating fetal cells, circulating tumor cells or cell-free plasma samples.

[0214] In certain embodiments, the read depth for the lower limit of range is greater than 100, 250, 500, 1,000, 2000, 2500, 5000, 10,000, 20,000, 25,0000, 50,000 or 100,000, and the read depth for the upper limit of range is greater than 2000, 2500, 5,000, 7,500, 10,000, 25,000, 50,000, 100,000, 250,000 or 500,000, in the sequence run for each single base variant position in the single base variant position set.Typically, the sequence run is a high-throughput sequence run.In exemplary embodiments, the average or median value generated for test samples is weighted by read depth. Thus, in a sample with one variant allele detected in 1,000 reads, the likelihood that the variant allele call is true is weighted higher than in a sample with one variant allele detected in 10,000 reads. Because variant allele (i.e., mutation) calls are not made with 100% confidence, identified single-nucleotide variants can be considered candidate variants or mutations.

[0215] 8.3 Example Test Statistics for the Analysis of Phased Data Exemplary test statistics for analyzing the phased data of a sample that is known or suspected to be a mixed sample containing DNA or RNA originating from two or more genetically non-identical cells are described below. f indicates the proportion of DNA or RNA of interest, such as the proportion of DNA or RNA with a CNV of interest, or the proportion of DNA or RNA from target cells, such as cancer cells. In some embodiments of cancer testing, f indicates the proportion of DNA or RNA derived from cancer cells in a mixture of cancer cells and normal cells, or f indicates the proportion of cancer cells in a mixture of cancer cells and normal cells. Note that this refers to the proportion of DNA derived from target cells, assuming that two copies of DNA are provided from each target cell. This is different from the proportion of DNA derived from target cells in deleted or duplicated segments.

[0216] The possible alleles of each SNP are designated A and B. AA, AB, BA, and BB are used to designate all possible ordered allele pairs. In some embodiments, SNPs with alleles ordered AB or BA are analyzed. i denotes the number of sequence reads for the i-th SNP, and A i and B i denotes the number of reads of the i-th SNP, representing alleles A and B, respectively. The following is assumed:

[0217] N i =A i +B i .

[0218] Allele ratio R i is provided:

[0219]

number

[0220] T indicates the number of targeted SNPs.

[0221] Without loss of generality, some embodiments focus on one chromosome segment. For clarity, in this specification, the phrase "a first homologous chromosome segment compared with a second homologous chromosome segment" refers to the first homolog of the chromosome segment and the second homolog of the chromosome segment. In some embodiments, all of the target SNPs are contained in the target chromosome segment. In other embodiments, multiple chromosome segments are analyzed for possible copy number variations.

[0222] MAP estimation

[0223] This method exploits knowledge of ordered allele-mediated phasing to detect deletions or duplications of target segments. For each SNP i, define

[0224]

number

[0225] Then, define

[0226]

number

[0227] X under various copy number hypotheses (e.g., disomy, deletion of the first or second homolog, or duplication of the first or second homolog) i and the distribution of S is given below.

[0228] Disomy hypothesis

[0229] Under the hypothesis that the target segment is not deleted or duplicated,

[0230]

[0231]

number

[0232] During the ceremony,

[0233]

number

[0234] Assuming that the depth of the lead N is constant, we are given a binomial distribution S with the following parameters:

[0235]

number

[0236] Deletion hypothesis

[0237] Under the hypothesis that the first homolog is deleted (i.e., the AB SNP becomes B and the BA SNP becomes A), R i is the parameter for AB SNPs

number

number

[0238]

number

[0239] Assuming that the depth of the lead N is constant, we are given a binomial distribution S with the following parameters:

[0240]

number

[0241] Under the hypothesis that the second homolog is deleted (i.e., the AB SNP becomes A and the BA SNP becomes B), R i is the parameter for AB SNPs

number

number

[0242]

number

[0243] Assuming that the depth of the lead N is constant, we are given a binomial distribution S with the following parameters:

[0244]

number

[0245] overlap hypothesis

[0246] Under the assumption that the first homolog is duplicated (i.e., the AB SNP becomes AAB and the BA SNP becomes BBA), R i is the parameter for AB SNPs

number

number

[0247]

number

[0248] Assuming that the depth of the lead N is constant, we are given a binomial distribution S with the following parameters:

[0249]

number

[0250] Under the assumption that the second homolog is duplicated (i.e., the AB SNP becomes ABB and the BA SNP becomes BAA), R i is the parameter for AB SNPs

number

number

[0251]

number

[0252] Assuming that the depth of the lead N is constant, we are given a binomial distribution S with the following parameters:

[0253]

number

[0254] classification

[0255] As shown in the section above, X i is a binary random variable.

[0256]

number

[0257] This allows the probability of the test statistic S under each hypothesis to be calculated. The probability of each hypothesis given the measured data can be calculated. In some embodiments, the hypothesis with the highest probability is selected. If desired, the distribution of S can be calculated for each N i can be simplified by approximating it with a constant read depth N or truncating the read depth to a constant N. This simplification gives

[0258]

number

[0259] The value of f can be estimated by selecting the most likely value of f given the measured data, e.g., the value of f that produces the best data fit using an algorithm (e.g., a search algorithm), such as maximum likelihood estimation, maximum a posteriori probability estimation, or Bayesian estimation. In some embodiments, multiple chromosomal segments are analyzed, and a value of f is estimated based on the data for each segment. If all target cells have these duplications or deletions, the estimates of f based on the data for these various segments will be similar. In some embodiments, f is measured experimentally, e.g., by determining the proportion of DNA or RNA from cancer cells based on the methylation difference (hypomethylated or hypermethylated) between cancerous and non-cancerous DNA or RNA.

[0260] Rejection of a single hypothesis

[0261] The distribution of S for the disomy hypothesis does not depend on f. Therefore, the probability of measured data can be calculated for the disomy hypothesis without calculating f. A single-hypothesis rejection test can be used for the null hypothesis of disomy. In some embodiments, the probability of S under the disomy hypothesis is calculated, and if the probability is below a given threshold (e.g., less than 1 in 1,000), the disomy hypothesis is rejected. This indicates the presence of a duplication or deletion of a chromosomal segment. If desired, the false positive rate can be changed by adjusting the threshold.

[0262] 8.4 Exemplary Methods for Analysis of Phased Data Exemplary methods for analyzing data from known or suspected mixed samples containing DNA or RNA originating from two or more genetically non-identical cells are described below. In some embodiments, phased data is used. In some embodiments, the method involves, for each calculated allele ratio, determining whether the calculated allele ratio is above or below the expected allele ratio and the degree of difference for a particular locus. In some embodiments, for a particular hypothesis, a likelihood distribution of allele ratios at the locus is determined, and the closer the calculated allele ratio is to the center of the likelihood distribution, the more likely the hypothesis is correct. In some embodiments, the method involves, for each locus, determining the likelihood that the hypothesis is correct. In some embodiments, the method involves, for each locus, determining the likelihood that the hypothesis is correct and combining the probabilities of the hypotheses for each locus, and selecting the hypothesis with the highest combined probability. In some embodiments, the method involves, for each locus, determining the likelihood that the hypothesis is correct and determining each possible ratio of DNA or RNA from one or more target cells to total DNA or RNA in the sample. In some embodiments, the combined probability for each hypothesis is determined by combining the probability of that hypothesis for each locus with each possible ratio, and the hypothesis with the highest combined probability is selected.

[0263] In one embodiment, the following hypothesis is considered: 11 (All cells are normal), H 10 (presence of cells with only homolog 1 and lacking homolog 2), H 01 (presence of cells with only homolog 2 and lacking homolog 1), H 21 (presence of cells with homolog 1 duplication), H 12 (Presence of cells with a duplication of homolog 2). For a proportion f of target cells (or proportion of DNA or RNA from target cells), e.g., cancer cells or mosaic cells, the expected allele fraction of a heterozygous (AB or BA) SNP can be found as follows:

[0264] Formula (1):

[0265]

number

[0266] Correction of bias, contamination and sequencing errors:

[0267] Measured D at SNP S is the number of original mapped reads in which each allele is present, n A 0 and n B 0 Then, the predicted bias in A and B allele amplification is used to calculate the corrected read n A and n B can be found.

[0268] c a indicates atmospheric contamination (e.g., contamination from airborne or environmental DNA), and r(c a ) indicates the allele ratio for atmospheric contamination (initially set to 0.5). Furthermore, c g indicates the contamination rate (e.g., contamination from another sample) during genotyping, and r(c g ) is the allele ratio for the contaminant. s e (A,B) and s e (B,A) indicates a sequencing error that results in classification of one allele as a different allele (e.g., incorrect detection of the A allele when the B allele is present).

[0269] By correcting for atmospheric contamination, genotyping contamination, and sequencing errors, the observed allele ratio q(r, c) is calculated for a given expected ratio r. a , r(c a ), c g , r(c g ), s e (A,B), s e (B,A)) can be found.

[0270] Since the contaminating genotype is unknown, population frequencies are used to calculate P(r(c g ) can be found. More specifically, let p be the population frequency of one of the alleles (sometimes called the reference allele). Then, P(r(c g )=0)=(1-p) 2 , P(r(c g )=0)=2p(1-p) 、 and P(r(cg)=0)=p 2 Let r(c g ) to obtain E[q(r, c a , r(c a ), c g , r(c g ), s e (A,B), s e (B, A))] can be determined. Note that atmospheric contamination and genotyping contamination are determined using homozygous SNPs and are therefore not affected by the presence or absence of deletions or duplications. Furthermore, if desired, atmospheric contamination and genotyping contamination can be measured using a reference chromosome.

[0271] Likelihood of each SNP:

[0272] The following equation is the ratio of n to r, where r is the allele ratio. A and n B is the formula that gives the probability of observing:

[0273] Formula (2):

[0274]

number

[0275] D s indicates the data of SNP s. h ε{H 11 , H 01 , H 10 , H 21 , H 12}, in equation (1), r = r(AB, h) or r = r(BA, h), and r(cg ) and calculate the conditional expectation of the observed allele ratio E[q(r, c a , r(c a ), c g , r(c g Then, in equation (2), r = E[q(r, c a , r(c a ), c g , r(c g ), s e (A,B), s e (B,A))], P(D s |h,f) can be determined.

[0276] Search algorithm:

[0277] In some embodiments, SNPs with allele ratios that appear to be outliers are ignored (e.g., SNPs with allele ratios at least 2 or 3 standard deviations above or below the mean are ignored or excluded). Note that an identified advantage of this method is that when the mosaicism rate is high, the variability in allele ratios may also be high, and therefore SNPs are not excluded due to mosaicism.

[0278] F={f1,….,f N} indicates the search space for mosaic fraction (e.g., tumor or fetal fraction). P(D s |h,f) can be determined and the derivations of all SNPs can be combined.

[0279] The algorithm searches for each f in each hypothesis. Using a search method, it concludes that mosaicing exists if there exists a range F* of f in which the confidence in the deletion or duplication hypotheses is higher than the confidence in the no-deletion or no-duplication hypotheses. In some embodiments, P(D s A maximum likelihood estimate for |h,f) is determined. If desired, a conditional expectation for f ε F* may be determined. If desired, the confidence of each hypothesis can be determined.

[0280] In some embodiments, a beta-binomial distribution is used instead of a binomial distribution. In some embodiments, a reference chromosome or a reference chromosome segment is used to determine sample-specific parameters of the beta-binomial distribution.

[0281] 8.5 Exemplary Methods for Detecting Deletions and Duplications Without Phased Data In some embodiments, unphased genetic data is used to determine whether there is an overrepresentation of a copy number of a first homologous chromosomal segment compared to a second homologous chromosomal segment in the individual's genome (e.g., the genome in one or more cells, or cfDNA or cfRNA). In some embodiments, phased genetic data is used, but the phase is ignored. In some embodiments, the DNA or RNA sample is a mixed cfDNA or cfRNA sample from an individual, containing cfRNA or cfRNA from two or more genetically different cells. In some embodiments, the method utilizes the magnitude of the difference between the calculated allele ratio for each locus and the expected allele ratio.

[0282] In some embodiments, the methods involve obtaining genetic data at a set of polymorphic loci on chromosomes or chromosome segments in a DNA or RNA sample derived from one or more cells from an individual by measuring the abundance of each allele at each locus. In some embodiments, allele ratios are calculated for loci that are heterozygous in at least one cell from which the sample was derived. In some embodiments, the calculated allele ratio for a particular locus is the measured abundance of one of the alleles divided by the total amount calculated for all alleles at that locus. In some embodiments, the calculated allele ratio for a particular locus is the measured abundance of one of the alleles (e.g., an allele on a first homologous chromosome segment) divided by the measured abundance of one or more other alleles at that locus (e.g., an allele on a second homologous chromosome segment). Calculated allele ratios and predicted allele ratios may be calculated using any of the methods described herein or using any standard method (e.g., any mathematical transformation of the calculated or predicted allele ratios described herein).

[0283] In some embodiments, a test statistic is calculated based on the magnitude of the difference between the calculated allele ratio and the expected allele ratio for each locus. In some embodiments, the test statistic Δ is calculated using the following formula:

number

[0284] In the formula, δ i is the magnitude of the difference between the calculated allele ratio and the predicted allele ratio for the i-th locus;

[0285] In the formula, μ i is δ i is the average value of; and

[0286] In the formula, σ i 2 is δ i· is the standard deviation of

[0287] For example, if the predicted allele ratio is 0.5, then δ is calculated as follows: i The following can be stipulated.

[0288]

number

[0289] μ i and σ i The value of R i is a binomial random variable. In some embodiments, the standard deviation is assumed to be the same for all loci. In some embodiments, the mean or weighted mean of the standard deviation, or an estimate of the standard deviation, is σ i 2 In some embodiments, the test statistic is assumed to have a normal distribution. For example, the central limit theorem implies that as the number of loci (e.g., the number of SNPs T) increases, the distribution of Δ converges to a standard normal.

[0290] In some embodiments, a set of one or more hypotheses specifying the copy number of a chromosome or chromosome segment in the genome of one or more cells is enumerated. In some embodiments, the most likely hypothesis is selected based on a test statistic, thereby determining the copy number of a chromosome or chromosome segment in the genome of one or more cells. In some embodiments, a hypothesis is selected if the probability that the test statistic belongs to the distribution of test statistics for that hypothesis is above an upper threshold. One or more of the hypotheses is rejected if the probability that the test statistic belongs to the distribution of test statistics for that hypothesis is below a lower threshold. Alternatively, a hypothesis is not selected or rejected if the probability that the test statistic belongs to the distribution of test statistics for that hypothesis is between the lower and upper thresholds, or if the probability cannot be determined with a sufficiently high degree of confidence. In some embodiments, the upper and / or lower thresholds are determined from an empirical distribution, such as a distribution from training data (e.g., samples with known copy numbers, such as diploid samples, or samples known to have a particular deletion or duplication). Such an empirical distribution can be used to select a threshold for rejecting a single hypothesis. Note that the test statistic Δ is independent of S, so both can be used independently if desired.

[0291] 8.6 Exemplary Methods for Detecting Deletions and Duplications Using Allele Distributions or Patterns This section includes methods for determining whether a first homologous chromosomal segment is overrepresented in copy number relative to a second homologous chromosomal segment. In some embodiments, the methods involve enumerating (i) a plurality of hypotheses specifying the copy number of a chromosome or chromosomal segment present in the genome of one or more cells (e.g., cancer cells) of an individual, or (ii) a plurality of hypotheses specifying the degree of copy number overrepresentation of a first homologous chromosomal segment relative to a second homologous chromosomal segment in the genome of one or more cells of the individual. In some embodiments, the methods involve obtaining genetic data from an individual for a plurality of polymorphic loci (e.g., SNP loci) on a chromosome or chromosomal segment. In some embodiments, for each hypothesis, a probability distribution of a predicted individual genotype is generated. In some embodiments, a data fit between the obtained individual genetic data and the probability distribution of a predicted individual genotype is calculated. In some embodiments, one or more hypotheses are ranked according to the data fit, and the highest-ranked hypothesis is selected. In some embodiments, a technique or algorithm, such as a search algorithm, is used in one or more of the following steps: calculating the data fit, ranking the hypotheses, or selecting the highest-ranked hypothesis. In some embodiments, the data fitting is a fit to a beta-binomial distribution or a fit to a binomial distribution. In some embodiments, the technique or algorithm is selected from the group consisting of maximum likelihood estimation, maximum a posteriori estimation, Bayesian estimation, dynamic estimation (e.g., dynamic Bayesian estimation), and expectation-maximization estimation. In some embodiments, the method includes applying the technique or algorithm to the obtained genetic data or the predicted genetic data.

[0292] In some embodiments, the method involves enumerating (i) a plurality of hypotheses specifying the copy number of a chromosome or chromosomal segment present in the genome of one or more cells (e.g., cancer cells) of the individual, or (ii) a plurality of hypotheses specifying the degree of copy number overrepresentation of a first homologous chromosomal segment relative to a second homologous chromosomal segment in the genome of one or more cells of the individual. In some embodiments, the method involves obtaining genetic data from the individual for a plurality of polymorphic loci (e.g., SNP loci) on a chromosome or chromosomal segment. In some embodiments, the genetic data includes allele counts for the plurality of polymorphic loci. In some embodiments, for each hypothesis, a joint distribution model for expected allele counts at the plurality of polymorphic loci on the chromosome or chromosomal segment is generated. In some embodiments, the relative probabilities for one or more hypotheses are determined using the joint distribution model and the allele counts measured in the sample, and the hypothesis with the highest probability is selected.

[0293] In some embodiments, the distribution or pattern of alleles (e.g., the pattern of calculated allele ratios) is used to determine the presence or absence of a CNV, e.g., a deletion or duplication. If desired, the parental origin of the CNV can be determined based on the pattern.

[0294] 8.7 Exemplary Counting / Quantitative Methods In some embodiments, one or more counting methods (also referred to as quantification methods) are used to detect one or more CNS defects, such as deletions or duplications of chromosome segments or entire chromosomes. In some embodiments, one or more counting methods are used to determine whether an overrepresentation of the copy number of a first homologous chromosome segment is due to duplication of the first homologous chromosome segment or deletion of a second homologous chromosome segment. In some embodiments, one or more counting methods are used to determine the number of extra copies of a duplicated chromosome segment or chromosome (e.g., whether there are 1, 2, 3, 4, or more extra copies). In some embodiments, one or more counting methods are used to distinguish between samples with a large number of duplications and a small tumor fraction and samples with a small number of duplications and a large tumor fraction. For example, one or more counting methods are used to distinguish between samples with four extra chromosome copies and a 10% tumor fraction and samples with two extra chromosome copies and a 20% tumor fraction. Exemplary methods are disclosed, for example, in U.S. Patent Publications 2007 / 0184467, 2013 / 0172211, and 2012 / 0003637, U.S. Patent Nos. 8,467,976, 7,888,017, 8,008,018, 8,296,076, and 8,195,415, U.S. Patent Application No. 62 / 008,235 filed June 5, 2014, and U.S. Patent Application No. 62 / 032,785 filed August 4, 2014, each of which is incorporated by reference herein in its entirety.

[0295] In some embodiments, an f value (e.g., tumor fraction or fetal fraction) is used in CNV determination, e.g., to compare the observed difference between the dosage of two chromosomes or chromosome segments to the difference expected for a particular type of CNV given the f value (see, e.g., U.S. Patent Publication No. 2012 / 0190020, U.S. Patent Publication No. 2012 / 0190021, U.S. Patent Publication No. 2012 / 0190557, U.S. Patent Publication No. 2012 / 0191358, each of which is incorporated by reference in its entirety). For example, as tumor fraction increases, the difference in dosage of a chromosome segment duplicated in the tumor increases compared to a reference chromosome segment that is disomic. In some embodiments, the method includes comparing the relative frequency of a chromosome or chromosome segment of interest to a reference chromosome or reference chromosome segment (e.g., a chromosome or chromosome segment predicted or known to be disomic) with respect to the f value to determine the likelihood of the CNV. For example, for various possible CNVs (e.g., one or two extra copies of the chromosomal segment of interest), the difference in dosage between a first chromosome or chromosomal segment and a reference chromosome or reference chromosomal segment can be compared to the difference expected given the f value.

[0296] 8.8 Exemplary enumeration / quantitative methods using reference samples Exemplary quantitative methods using one or more reference samples are described in U.S. Patent Application No. 62 / 008,235, filed June 5, 2014, and U.S. Patent Application No. 62 / 032,785, filed August 4, 2014, which are incorporated by reference herein in their entireties. In some embodiments, one or more reference samples most likely to not have any CNVs in one or more chromosomes or chromosomes of interest are identified by selecting the sample with the highest tumor DNA fraction, selecting the sample with a z-score closest to zero, selecting a sample whose data fit a hypothesis that does not correspond to a CNV with the highest confidence or likelihood, selecting a sample that is known to be normal, selecting a sample from an individual with the lowest likelihood of having cancer (e.g., young age, male for breast cancer screening, no family history, etc.), selecting a sample with the highest DNA input, selecting a sample with the highest signal-to-noise ratio, selecting samples based on other criteria believed to correlate with the likelihood of having cancer, or selecting samples using a combination of criteria. Once the reference set is selected, these cases are assumed to be disomy, and then the bias per SNP, i.e., experiment-specific amplification and other processing biases for each locus, can be estimated. This experiment-specific bias estimate is then used to correct for bias in the measurement of the chromosome of interest, such as the locus on chromosome 21, and optionally other chromosome loci, for samples that are not part of the subset and are assumed to be disomy with respect to chromosome 21. Once bias is corrected for these samples with unknown ploidy, the data for these samples can then be analyzed twice using the same or a different method to determine whether an individual has trisomy 21. For example, a quantitative method can be used for the remaining samples with unknown ploidy, and z-scores can be calculated using the corrected measured genetic data for chromosome 21. Alternatively, the tumor fraction of samples from individuals suspected of having cancer can be calculated as part of a preliminary estimate of the ploidy status of chromosome 21.For cases with a given tumor fraction, the predicted corrected read fraction for disomy (disomy hypothesis) and the predicted corrected read fraction for trisomy (trisomy hypothesis) can be calculated. Alternatively, if tumor fractions have not been measured previously, a set of disomy and trisomy hypotheses can be generated for different tumor fractions. For each example, a predicted distribution of corrected read fractions can be calculated, taking into account expected statistical variation in the selection and measurement of various DNA loci. The observed corrected read fraction can be compared to the distribution of predicted corrected read fractions, and likelihood ratios can be calculated for the disomy and trisomy hypotheses for each sample with unknown ploidy. The ploidy state associated with the hypothesis with the highest calculated likelihood can be selected as the correct ploidy state.

[0297] 8.9 Exemplary Reference Chromosomes or Reference Chromosome Segments In some embodiments, any of the methods described herein are performed on one or more reference chromosomes or reference chromosome segments, and the results are compared to one or more chromosomes or chromosome segments of interest.

[0298] In some embodiments, a reference chromosome or reference chromosome segment is used as a control for those predicted to be free of CNVs. In some embodiments, the reference is the same chromosome or chromosome segment as one or more different samples known or predicted to have no deletions or duplications in that chromosome or chromosome segment. In some embodiments, the reference is a different chromosome or chromosome segment from the test sample predicted to be disomic. In some embodiments, the reference is a different segment from one of the chromosomes of interest in the same test sample. For example, the reference may be one or more segments outside the potential deletion or duplication region. Having a reference on the same chromosome as the test chromosome avoids variations between different chromosomes, such as differences between chromosomes in metabolism, apoptosis, histones, inactivation, and / or amplification. Analysis of a segment on the same chromosome as the test chromosome that does not have a CNV can be performed to determine differences between homologs in metabolism, apoptosis, histones, inactivation, and / or amplification, thereby determining the level of variation between homologs in the absence of CNV and comparing it to results from potential CNVs. In some embodiments, for a potential CNV, the degree of difference between the calculated allele ratio and the expected allele ratio is greater than the corresponding degree for the reference, thereby confirming the presence of the CNV.

[0299] In some embodiments, a reference chromosome or reference chromosome segment is used as a control for a CNV, such as a deletion or duplication of particular interest, that is predicted to exist. In some embodiments, the reference is the same chromosome or chromosome segment as one or more different samples known or predicted to have a deletion or duplication in the chromosome or chromosome segment. In some embodiments, the reference is a different chromosome or chromosome segment from the test sample known or predicted to have a CNV. In some embodiments, for a potential CNV, the degree of difference between the calculated allele ratio and the predicted allele ratio is less (e.g., not significantly different) than the corresponding degree for the reference CNV, thereby confirming the presence of the CNV. In some embodiments, for a potential CNV, the degree of difference between the calculated allele ratio and the predicted allele ratio is less (e.g., significantly less) than the corresponding degree for the reference CNV, thereby confirming the absence of the CNV. In some embodiments, one or more loci at which the genotype of cancer cells (or DNA or RNA derived from cancer cells, such as cfDNA or cfRNA) differs from the genotype of non-cancerous cells (or DNA or RNA derived from non-cancerous cells, such as cfDNA or cfRNA) are used to determine tumor fraction. The tumor fraction can be used to determine whether the overrepresentation of a copy number of a first homologous chromosomal segment is due to a duplication of the first homologous chromosomal segment or a deletion of a second homologous chromosomal segment. The tumor fraction can also be used to determine the excess copy number of the duplicated chromosomal segment or chromosome (e.g., whether one, two, three, four, or more excess copies are present), e.g., to distinguish between a sample with four excess chromosomal copies and a 10% tumor fraction and a sample with two excess chromosomal copies and a 20% tumor fraction. The tumor fraction can also be used to determine how well observed data matches predicted data for potential CNVs. In some embodiments, the degree of overrepresentation of CNVs is used to select a particular treatment or treatment regimen for an individual. For example, some therapeutic agents are only effective against chromosomal segments with at least four, six, or more copies.

[0300] In some embodiments, one or more loci used to determine tumor proportion are located on a reference chromosome or reference chromosome segment, such as a chromosome or chromosome segment that is known or predicted to be disomic, a chromosome or chromosome segment that is rarely duplicated or deleted in cancer cells in general, or in a specific type of cancer that an individual is known to have or is at increased risk of having, or a chromosome or chromosome segment that is unlikely to be aneuploid (a segment that is predicted to cause cell death if deleted or duplicated).In some embodiments, using any of the methods of the present invention, the reference chromosome or chromosome segment is confirmed to be disomic in both cancer cells and non-cancerous cells.In some embodiments, one or more chromosomes or chromosome segments that have high confidence in disomic classification are used.

[0301] Examples of loci that can be used to determine tumor incidence include polymorphisms or mutations (e.g., SNPs) in cancer cells (or DNA or RNA, such as cfDNA or cfRNA, derived from cancer cells) that are not present in non-cancerous cells (or DNA or RNA derived from non-cancerous cells) of an individual. In some embodiments, tumor incidence is determined by identifying polymorphic loci in which cancer cells (or DNA or RNA derived from cancer cells) have alleles that are not present in non-cancerous cells (or DNA or RNA derived from non-cancerous cells) in a sample (e.g., a plasma sample or tumor biopsy) from an individual, and determining tumor incidence in the sample using the amount of alleles unique to cancer cells at one or more of the identified polymorphic loci. In some embodiments, the non-cancerous cells are homozygous for a first allele at the polymorphic locus, and the cancer cells are (i) heterozygous for the first allele and a second allele at the polymorphic locus, or (ii) homozygous for the second allele. In some embodiments, non-cancerous cells are heterozygous for a first allele and a second allele at a polymorphic locus, and cancer cells have (i) one or two copies of a third allele at the polymorphic locus. In some embodiments, cancer cells are assumed or known to have only one copy of an allele not present in non-cancerous cells. For example, if the genotype of non-cancerous cells is AA and the genotype of cancer cells is AB, and 5% of the signals at the locus in the sample are from the B allele and 95% are from the A allele, the tumor fraction of the sample is 10%. In some embodiments, cancer cells are assumed or known to have two copies of an allele not present in non-cancerous cells. For example, if the genotype of non-cancerous cells is AA and the genotype of cancer cells is BB, and 5% of the signals at the locus in the sample are from the B allele and 95% are from the A allele, the tumor fraction of the sample is 5%. In some embodiments, multiple loci at which cancer cells have alleles that are absent in non-cancerous cells are analyzed to determine which loci in the cancer cells are heterozygous and which are homozygous.For example, for loci where non-cancerous cells are AA, if the signal from the B allele is about 5% at some loci and about 10% at some loci, then the cancer cells are assumed to be heterozygous at loci with about 5% of the B alleles and homozygous at loci with about 10% of the B alleles (indicating a tumor fraction of about 10%).

[0302] Examples of loci that can be used to determine tumor ratio include loci where cancer cells and non-cancerous cells share one allele (for example, where cancer cells are AB and non-cancerous cells are BB, or where cancer cells are BB and non-cancerous cells are AB). The amount of A signal, the amount of B signal, or the ratio of A signal to B signal in a mixed sample (containing DNA or RNA from cancer cells and non-cancerous cells) is compared with the corresponding values ​​of (i) a sample containing DNA or RNA from cancer cells only, or (ii) a sample containing DNA or RNA from non-cancerous cells only. The difference in values ​​is used to determine the tumor ratio of the mixed sample.

[0303] In some embodiments, loci that can be used to determine tumor fraction are selected based on the genotype of (i) a sample containing DNA or RNA derived only from cancer cells, and / or (ii) a sample containing DNA or RNA derived only from non-cancerous cells. In some embodiments, loci are selected based on the analysis of mixed samples, for example, loci where the absolute or relative amount of each allele is different from the amount expected if both cancer cells and non-cancerous cells have the same genotype at a particular locus. For example, if cancer cells and non-cancerous cells have the same genotype, a locus is expected to produce 0% B signal if all cells are AA, 50% B signal if all cells are AB, or 100% B signal if all cells are BB. If the B signal is any other value, it indicates that the genotypes of cancer cells and non-cancerous cells are different at that locus, and the locus can be used to determine tumor fraction.

[0304] In some embodiments, the tumor incidence calculated based on alleles at one or more loci is compared to the tumor incidence calculated using one or more of the counting methods disclosed herein.

[0305] 8.10 Examples of Phenotypic Detection Methods or Multiplex Mutation Analysis Methods In some embodiments, the method involves analyzing a sample for a set of mutations associated with a disease or disorder (e.g., cancer) or an increased risk of a disease or disorder. There is a strong correlation between events within a class (e.g., M or C cancer class) that can be used to improve the signal-to-noise ratio of the method and classify tumors into distinct clinical subsets. For example, an equivocal result for several mutations (e.g., several CNVs) on one or more chromosomes or chromosomal segments considered together can be a very strong signal. In some embodiments, determining the presence or absence of multiple polymorphisms or mutations of interest (e.g., 2, 3, 4, 5, 8, 10, 12, 15, or more) increases the sensitivity and / or specificity of determining the presence or absence of a disease or disorder, such as cancer, or the presence or absence of an increased risk of having a disease or disorder, such as cancer. In some embodiments, correlations between events across multiple chromosomes are used to consider signals more strongly than considering each signal individually. The design of the method itself can be optimized to best categorize tumors. This can be very useful for early detection and screening of recurrence, where sensitivity to one particular mutation / CNV may be most important. In some embodiments, the events are not necessarily correlated, but have correlated probabilities. In some embodiments, a matrix estimation formula is used with a noise covariance matrix that has off-diagonal terms.

[0306] In some embodiments, the invention features methods for detecting an individual's phenotype (e.g., a cancer phenotype), where the phenotype is defined by the presence of at least one mutation in a set of mutations. In some embodiments, the methods include obtaining DNA or RNA measurements from one or more cells from the individual, where one or more of the cells are suspected of having the phenotype, and analyzing the DNA or RNA measurements to determine, for each mutation in the set of mutations, the likelihood that at least one of the cells has the mutation. In some embodiments, the methods include determining that the individual has the phenotype if either (i) for at least one of the mutations, the likelihood that at least one of the cells contains the mutation is greater than a threshold, or (ii) for at least one of the mutations, the likelihood that at least one of the cells has the mutation is less than a threshold, and for a plurality of the mutations, the combined likelihood that at least one of the cells has at least one of the mutations is greater than a threshold. In some embodiments, one or more cells have a subset or all of the mutations in the set of mutations. In some embodiments, the subset of mutations is associated with cancer or an increased risk of cancer. In some embodiments, the set of mutations includes a subset or all of the mutations in the M class of cancer mutations (Ciriello, Nat Genet. 45(10):1127-1133, 2013, doi:10.1038 / ng.2762, which is incorporated herein by reference in its entirety). In some embodiments, the set of mutations includes a subset or all of the mutations in the C class of cancer mutations (Ciriello, supra). In some embodiments, the sample includes cell-free DNA or RNA. In some embodiments, the DNA or RNA measurements include measurements at a set of polymorphic loci (e.g., the dosage of each allele at each locus) on one or more chromosomes or chromosomal segments of interest.

[0307] 8.11 Example of Combining Methods To improve the accuracy of the results, two or more methods for detecting the presence or absence of CNV (e.g., any of the methods of the present invention or any known method) are performed. In some embodiments, one or more methods for analyzing factors indicating the presence or absence of a disease or disorder, or an increased risk of a disease or disorder (e.g., a method described herein or any known method) are performed.

[0308] In some embodiments, standard mathematical techniques are used to calculate covariances and / or correlations between two or more methods. Standard mathematical methods may also be used to determine the combined probability of a particular hypothesis based on two or more tests. Exemplary techniques include meta-analysis, analysis, Fisher's combined probability test for independent testing, Brown's method for combining known covariances and dependent p-values, and Kost's method for combining unknown covariances and dependent p-values. When likelihoods are determined by a first method that is in some sense orthogonal or unrelated to those determined for a second method, combining the likelihoods is straightforward and can be done by multiplication or normalization, or using, for example, the following formula:

[0309] R comb =R1R2 / [R1R2+(1-R1)(1-R2)]

[0310] R comb is the combined likelihood, and R1 and R2 are the individual likelihoods. For example, if the likelihood of trisomy from method 1 is 90% and the likelihood of trisomy from method 2 is 95%, then by combining the results from the two methods, the clinician will obtain (0.90)(0.95) / [(0.90)(0.95)+(1-0.90)(1-0.95)] = 99.42% likelihood of concluding that the fetus is trisomic. If the first and second methods are not orthogonal, i.e., there is a correlation between the two methods, the likelihoods can still be combined.

[0311] Examples of methods for analyzing multiple factors or variables are disclosed in U.S. Patent No. 8,024,128, issued September 20, 2011, U.S. Publication No. 2007 / 0027636, filed July 31, 2006, and U.S. Publication No. 2007 / 0178501, filed December 6, 2006, which are incorporated by reference in their entireties.

[0312] In various embodiments, the combined probability of a particular hypothesis or diagnosis is greater than 80, 85, 90, 92, 94, 96, 98, 99, or 99.9%, or greater than some other threshold.

[0313] 8.12 Exemplary Methods for Phasing Genetic Data In some embodiments, the genetic data is phased using the methods described herein or any known method for phasing genetic data (e.g., PCT Publication WO2009 / 105531 filed February 9, 2009; PCT Publication WO2010 / 017214 filed August 4, 2009; U.S. Patent Publication 2013 / 0123120 filed November 21, 2012; U.S. Patent Publication 2011 / 0033862 filed October 7, 2010; U.S. Patent Publication 2011 / 0033862 filed August 19, 2010; See U.S. Patent Publication No. 2011 / 0033862, filed February 3, 2011; U.S. Patent Publication No. 2011 / 0178719, filed February 3, 2011; U.S. Patent No. 8,515,679, filed March 17, 2008; U.S. Patent Publication No. 2007 / 0184467, filed November 22, 2006; U.S. Patent Publication No. 2008 / 0243398, filed March 17, 2008; and U.S. Application No. 61 / 994,791, filed May 16, 2014, each of which is incorporated herein by reference in its entirety.

[0314] In one embodiment, an individual's genetic data is phased using a computer program that uses population-based haplotype frequencies to estimate the most likely phase, such as HapMap-based phasing. For example, haploid datasets can be directly inferred from diploid data using statistical methods, utilizing known haplotype blocks in the general population (e.g., populations generated for the publicly available HapMap Project and the Perlegen Human Haplotype Project). Haplotype blocks are essentially sets of correlated alleles that occur repeatedly in various populations. Because these haplotype blocks are often common since ancient times, they may be used to predict haplotypes from diploid genotypes. Publicly available algorithms that accomplish this task include incomplete descent methods, Bayesian methods based on conjugate priors, and Bayesian methods based on priors from population genetics. Some of these algorithms use hidden Markov models.

[0315] In one embodiment, an individual's genetic data is phased using an algorithm that infers haplotypes from genotype data, such as localized haplotype clustering (see, e.g., Browning and Browning, "Rapid and Accurate Haplotype Phasing and Missing-Data Inference for Whole-Genome Association Studies By Use of Localized Haplotype Clustering" Am J Hum Genet. Nov 2007;81(5):1084-1097, which is incorporated herein by reference in its entirety). An exemplary program is Beagle version 3.3.2 or version 4 (available on the World Wide Web at hfaculty.washington.edu / browning / beagle / beagle.html, which is incorporated herein by reference in its entirety).

[0316] In one embodiment, the individual's genetic data is processed using an algorithm that infers haplotypes from genotypic data, e.g., an algorithm that uses linkage disequilibrium decay with distance, ordering and spacing of genotypic markers, imputation of missing data, recombination rate estimation, or a combination thereof (see, e.g., Stephens and Scheet, "Accounting for Decay of Linkage Disequilibrium in Haplotype Inference and Missing-Data Imputation," Am. J. Hum. Genet. 76:449-462, 2005, which is incorporated herein by reference in its entirety). An exemplary program is PHASE v.2.1 or v2.1.1 (available on the World Wide Web at stephenslab.uchicago.edu / software.html, which is incorporated herein by reference in its entirety).

[0317] In one embodiment, an individual's genetic data is phased using an algorithm that infers haplotypes from population genotype data, e.g., an algorithm that sequentially varies cluster membership along chromosomes according to a hidden Markov model. This method is flexible and allows for both "block-like" patterns of linkage disequilibrium and gradual reduction of linkage disequilibrium over distance (see, e.g., Scheet and Stephens, "A fast and flexible statistical model for large-scale population genotype data: applications to inferring missing genotypes and haplotypic phase," Am J Hum Genet, 78:629-644, 2006, which is incorporated herein by reference in its entirety). An exemplary program is fastPHASE (available on the World Wide Web at stephenslab.uchicago.edu / software.html, which is incorporated herein by reference in its entirety).

[0318] In one embodiment, an individual's genetic data is phased using a genotype imputation method, such as a method using one or more of the following reference datasets: the HapMap dataset, a control dataset genotyped on multiple SNP chips, and high-density typed samples from the 1,000 Genomes Project. An exemplary method is a flexible modeling framework that increases accuracy and combines information across multiple reference panels (see, e.g., Howie, Donnelly, and Marchini (2009) "A flexible and accurate genotype imputation method for the next generation of genome-wide association studies," PLoS Genetics 5(6):e1000529, 2009, which is incorporated herein by reference in its entirety). An exemplary program is IMPUTE or IMPUTE version 2 (also known as IMPUTE2) (available on the World Wide Web at mathgen.stats.ox.ac.uk / impute / impute_v2.html, which is incorporated herein by reference in its entirety).

[0319] In one embodiment, an individual's genetic data is phased using a haplotype inference algorithm, such as an algorithm that infers haplotypes under a coalescent genetic model with recombination, e.g., the algorithm developed by Stephens in PHASE v2.1. A key algorithmic improvement relies on the use of binary trees to represent the set of candidate haplosets for each individual. These binary tree representations (1) speed up the calculation of haplotype posterior probabilities by avoiding the redundant operations created in PHASE v2.1, and (2) overcome the exponential aspects of the haplotype inference problem by performing smart searches for plausible paths (i.e., haplotypes) in the binary tree (see, e.g., Delaneau, Coulonges, and Zagury, "Shape-IT: a new rapid and accurate algorithm for haplotype inference," BMC Bioinformatics 9:540, 2008 doi:10.1186 / 1471-2105-9-540, which is incorporated herein by reference in its entirety). An exemplary program is SHAPEIT (available on the World Wide Web at mathgen.stats.ox.ac.uk / genetics_software / shapeit / shapeit.html, which is incorporated herein by reference in its entirety).

[0320] In one embodiment, an individual's genetic data is phased using an algorithm that infers haplotypes from population genotype data, e.g., an algorithm that uses haplotype-fragment frequencies to obtain empirical probabilities for longer haplotypes. In some embodiments, the algorithm reconstructs haplotypes so that they have maximum local consistency (see, e.g., Eronen, Geerts, and Toivonen, "HaploRec: Efficient and accurate large-scale reconstruction of haplotypes," BMC Bioinformatics 7:542, 2006, which is incorporated herein by reference in its entirety). An exemplary program is HaploRec, e.g., HaploRec version 2.3 (available on the World Wide Web at cs.helsinki.fi / group / genetics / haplotyping.html, which is incorporated herein by reference in its entirety).

[0321] In one embodiment, the individual's genetic data is phased using algorithms that infer haplotypes from population genotype data, such as algorithms that use partition-ligation strategies and expectation-maximization-based algorithms (see, e.g., Qin, Niu, and Liu, "Partition-Ligation-Expectation-Maximization Algorithm for Haplotype Inference with Single-Nucleotide Polymorphisms," Am J Hum Genet. 71(5):1242-1247, 2002, which is incorporated herein by reference in its entirety). An exemplary program is PL-EM (available on the World Wide Web at people.fas.harvard.edu / ~junliu / plem / click.html, which is incorporated herein by reference in its entirety).

[0322] In one embodiment, the individual's genetic data is phased using an algorithm that infers haplotypes from population genotype data, such as an algorithm that simultaneously phases genotypes into haplotypes and block partitions. In some embodiments, an expectation-maximization algorithm is used (see, e.g., Kimmel and Shamir, "GERBIL: Genotype Resolution and Block Identification Using Likelihood," Proceedings of the National Academy of Sciences of the United States of America (PNAS) 102:158-162, 2005, which is incorporated herein by reference in its entirety). An exemplary program is GERBIL, available as part of the GEVALT version 2 program (available on the World Wide Web at acgt.cs.tau.ac.il / gevalt / , which is incorporated herein by reference in its entirety).

[0323] In one embodiment, the individual's genetic data is phased using an algorithm that infers haplotypes from population genotype data, such as an algorithm that uses the EM algorithm to calculate ML estimates of haplotype frequencies given unphased genotype measurements. The algorithm also misses some genotype measurements (e.g., due to PCR failures). The algorithm also allows for multiple imputation of individual haplotypes (see, e.g., Clayton, D. (2002), "SNPHAP: A Program for Estimating Frequencies of Large Haplotypes of SNPs," which is incorporated herein by reference in its entirety). An exemplary program is SNPHAP.

[0324] In one embodiment, the individual's genetic data is phased using an algorithm that infers haplotypes from population genotype data, e.g., an algorithm for haplotype inference based on genotype statistics collected for SNP pairing. This software can be used to perform relatively accurate phasing of large numbers of long genomic sequences obtained, for example, from DNA arrays. An exemplary program takes a genotype matrix as input and outputs a corresponding haplotype matrix (see, e.g., Brinza and Zelikovsky, "2SNP: scalable phasing based on 2-SNP haplotypes," Bioinformatics. 22(3):371-3, 2006, which is incorporated herein by reference in its entirety). An exemplary program is 2SNP (available on the World Wide Web at alla.cs.gsu.edu / ~software / 2SNP, which is incorporated herein by reference in its entirety).

[0325] In various embodiments, an individual's genetic data is phased using data on the probability of chromosomal crossover at different locations on a chromosome or chromosomal segment (e.g., using recombination data that can be found in the HapMap database, which generates a recombination risk score for any interval), and the dependence between polymorphic alleles on the chromosome or chromosomal segment is modeled. In some embodiments, the number of alleles at the polymorphic locus is calculated computationally based on sequence data or SNP array data. In some embodiments, multiple hypotheses are generated (e.g., computationally generated), each associated with a different possible state of the chromosome or chromosomal segment (e.g., overrepresentation of a first homologous chromosomal segment compared to a second homologous chromosomal segment in the genome of one or more cells from the individual, duplication of the first homologous chromosomal segment, duplication of the second homologous chromosomal segment, or equivalent distribution of the first and second homologous chromosomal segments). A model (e.g., a joint distribution model) for the expected number of alleles at the polymorphic locus on the chromosome is constructed (e.g., computationally constructed) for each hypothesis. The relative probability of each hypothesis is determined (e.g., determined on a computer) using the joint distribution model and the number of alleles. The hypothesis with the highest probability is then selected. In some embodiments, the steps of constructing a joint distribution model for the number of alleles and determining the relative probability of each hypothesis are performed using a method that does not require the use of a reference chromosome.

[0326] In some embodiments, a sample from an individual (e.g., a tumor biopsy, blood sample, plasma sample, serum sample, or another sample likely to contain the predominant or only cells, DNA, or RNA harboring the CNV of interest) is analyzed to determine the phase of one or more regions known or suspected to contain the CNV of interest (e.g., a deletion or duplication). In some embodiments, the sample has a high tumor fraction (e.g., 30, 40, 50, 60, 70, 80, 90, 95, 98, 99, or 100%, etc.). In some embodiments, the sample is a blood sample from a mother carrying a fetus. In some embodiments, the blood sample from a mother carrying a fetus contains circulating fetal cells and / or fetal DNA.

[0327] In some embodiments, the sample has a haplotype imbalance or any aneuploidy. In some embodiments, the sample contains any mixture of two types of DNA, where the two types of DNA have different ratios of two haplotypes and share at least one haplotype. In some embodiments, at least 10, 100, 500, 1,000, 2,000, 3,000, 5,000, 8,000, or 10,000 polymorphic loci are analyzed to determine the phase of some or all alleles at the loci. In some embodiments, the sample is derived from cells or tissues that have been treated to become aneuploid, for example, aneuploidy induced by long-term cell culture.

[0328] In some embodiments, most or all of the DNA or RNA in the sample contains the CNV of interest. In some embodiments, the ratio of DNA or RNA from one or more target cells containing the CNV of interest to the total DNA or RNA in the sample is at least 80, 85, 90, 95, or 100%. For samples with deletions, only one haplotype is present for the cells (or DNA or RNA) with the deletion. This first haplotype can be determined using standard methods to determine the identity of the alleles present in the deleted region. In samples containing only cells (or DNA or RNA) containing the deletion, only the signal from the first haplotype present in those cells is present. In samples that also contain a small number of cells (or DNA or RNA) without the deletion (such as a small number of non-cancerous cells), the weak signal from the second haplotype in these cells (or DNA or RNA) can be ignored. The second haplotype present in other cells, DNA, or RNA from individuals without the deletion can be determined by inference. For example, if the genotype of a cell from an individual without a deletion is (AB, AB) and the phasing data for that individual indicates that the first haplotype is (A, A), then the other haplotype can be inferred to be (B, B).

[0329] For samples in which both cells (or DNA or RNA) with and without deletions are present, the phase can be further determined. For example, a plot can be created in which the x-axis represents the linear position of each locus along a chromosome and the y-axis represents the number of A allele reads as a percentage of the total (A+B) allele reads. In some embodiments involving deletions, the pattern includes two central bands representing SNPs for which the individual is heterozygous (the upper band represents AB from cells without deletions and A from cells with deletions, and the lower band represents AB from cells without deletions and B from cells with deletions). In some embodiments, as the percentage of cells, DNA, or RNA with deletions increases, the separation of these two bands also increases. Thus, the uniqueness of the A allele can be used to determine a first haplotype, and the uniqueness of the B allele can be used to determine a second haplotype.

[0330] For samples with duplications, extra copies of the haplotype are present in the cells (or DNA or RNA) that contain the duplication. The haplotype in the overlapped region can be determined using standard methods for determining the identity of alleles present in increased abundance in the overlapped region, or the haplotype in the non-overlapping region can be determined using standard methods for determining the identity of alleles present in decreased abundance. Once one haplotype is determined, the other haplotype can be determined by inference.

[0331] For samples containing both cells (or DNA or RNA) with duplications and cells (or DNA or RNA) without duplications, phase can be further determined using methods similar to those described above for deletions. For example, a plot can be created in which the x-axis represents the linear position of each locus along a chromosome and the y-axis represents the number of A allele reads as a percentage of the total (A+B) allele reads. In some embodiments for deletions, the pattern includes two central bands representing SNPs for which the individual is heterozygous (the upper band represents AB from cells without duplications and AAB from cells with duplications, and the lower band represents AB from cells without duplications and ABB from cells with duplications). In some embodiments, the separation of these two bands increases as the percentage of cells, DNA, or RNA with duplications increases. Thus, the uniqueness of the A allele can be used to determine a first haplotype, and the uniqueness of the B allele can be used to determine a second haplotype. In some embodiments, a sample (e.g., a tumor biopsy or plasma sample) from an individual known to have cancer is phased for one or more CNV regions (e.g., the phase of at least 50, 60, 70, 80, 90, 95, or 100% of the polymorphic loci in the measured regions), and subsequent samples from the same individual are analyzed to monitor the progression of the cancer (e.g., to monitor remission or recurrence of the cancer). In some embodiments, a tumor-rich sample (e.g., a tumor biopsy or plasma sample from an individual with a high tumor burden) is used to obtain phase determination data, which is then used to analyze subsequent samples with a low tumor burden (e.g., a plasma sample from an individual who has been treated for cancer or is in remission).

[0332] In some embodiments, genetic data for an individual is phased using two or more of the methods described herein. In some embodiments, bioinformatics methods (e.g., methods that use population-based haplotype frequencies to estimate the most likely phase) and molecular biology methods (e.g., any of the molecular phasing methods described herein to obtain actual phasing data rather than bioinformatics-based estimated phasing data) are used. In some embodiments, phasing data from other subjects (e.g., past subjects) is used to improve the population data. For example, phasing data from other subjects can be added to the population data to calculate a prior probability distribution of possible haplotypes for another subject. In some embodiments, phasing data from other subjects (e.g., past subjects) is used to calculate a prior probability distribution of possible haplotypes for another subject.

[0333] In some embodiments, probabilistic data may be used. For example, due to the probabilistic nature of the occurrence of DNA molecules in a sample and various amplification and measurement biases, the relative number of DNA molecules measured from two different loci or from different alleles at a given locus does not always represent the relative number of molecules in a mixture or the relative number of molecules in an individual. When attempting to determine the genotype of a normal diploid individual at a given locus on an autosome by sequencing DNA from the individual's plasma, it is expected that only one allele will be observed (homozygosity) or that two alleles will be observed in approximately equal numbers (heterozygosity). If 10 A alleles and 2 B alleles are observed at that allele, it is unclear whether the individual is homozygous at that locus and the two B alleles are due to noise or contamination, or whether the individual is heterozygous and the low number of B allele molecules is due to random statistical fluctuations in the number of DNA molecules in plasma, amplification bias, contamination, or various other causes. In this case, the probability that an individual is homozygous and the corresponding probability that an individual is heterozygous can be calculated, and these probabilistic genotypes can be used for further calculations.

[0334] It should be noted that for a given allele ratio, the likelihood that the ratio closely represents the proportion of DNA molecules in an individual increases with the number of molecules observed. For example, if 100 molecules of A and 100 molecules of B are measured, the likelihood that the actual ratio was 50% is much higher than if 10 molecules of A and 10 molecules of B are measured. In one embodiment, Bayesian theory is combined with a detailed model of the data to determine the likelihood that a particular hypothesis is correct given empirical measurements. For example, when considering two hypotheses, one corresponding to a trisomic individual and the other corresponding to a disomic individual, the probability that the disomic hypothesis is correct is much higher if 100 molecules each with two alleles are observed compared to if 10 molecules each with two alleles are observed. As the data become noisier due to bias, contamination, or other noise sources, or as the number of empirical measurements at a given locus decreases, the probability that the maximum likelihood hypothesis is true given the empirical data decreases. In practice, it is possible to aggregate probabilities at many loci to increase the confidence with which the maximum likelihood hypothesis can be determined to be correct. In some embodiments, the probabilities are simply aggregated without considering recombination, hi some embodiments, the calculation takes crossover into account.

[0335] In one embodiment, probabilistically phased data is used to determine copy number variations. In some embodiments, the probabilistically phased data is, for example, population-based haplotype block frequency data from a data source such as the HapMap database. In some embodiments, the probabilistically phased data is haplotype data obtained by molecular methods, for example, phased by dilution, where individual segments of a chromosome are diluted to one molecule per reaction, but due to stochastic noise, the identity of the haplotype is not fully known. In some embodiments, the probabilistically phased data is haplotype data obtained by molecular methods, where the identity of the haplotype may be known with a high degree of certainty.

[0336] Consider the hypothetical case in which a physician wishes to determine whether an individual has any cells in their body that have a deletion in a particular chromosomal segment by measuring plasma DNA from that individual. The physician may use the knowledge that if all of the cells from which the plasma DNA was derived are diploid and of the same genotype, then for a heterozygous locus, the relative number of DNA molecules observed for each of the two alleles will fall into a single distribution that encompasses 50% A alleles and 50% B alleles. However, if a fraction of the cells from which the plasma DNA was derived contain a deletion in a particular chromosomal segment, then the relative number of DNA molecules observed for each of the two alleles for that locus will be expected to fall into two distributions: one that encompasses more than 50% A alleles for loci with deletions in chromosomal segments containing the B allele, and one that encompasses less than 50% for loci with deletions in chromosomal segments containing the A allele. As the proportion of cells from which the plasma DNA was derived increases, these two distributions will move further away from 50%.

[0337] In this hypothetical example, consider a physician who wants to determine whether an individual has a deletion of a chromosomal region in a portion of the individual's cells. The physician may collect blood from the individual into a vacutainer or other type of blood tube and centrifuge the blood to separate the plasma layer. The physician isolates DNA from the plasma and, optionally, enriches the DNA at the targeted locus using targeted amplification or other amplification, locus capture technology, size enrichment, or other enrichment techniques. The physician may perform analysis, for example, by measuring the number of alleles at a set of SNPs, i.e., generating allele frequency data, by generating enriched and / or amplified DNA using an assay such as qPCR, sequencing, microarray, or other technique that measures the amount of DNA in a sample. We consider data analysis for the case in which a physician uses targeted amplification technology to amplify cell-free plasma DNA and then sequences the amplified DNA to obtain the following possible exemplary data for six SNPs present on a chromosomal segment that is indicative of cancer. In this case, the individual was heterozygous at the SNPs:

[0338] SNP1: 460 reads A allele, 540 reads B allele (46% A)

[0339] SNP2: 530 reads A allele, 470 reads B allele (53% A)

[0340] SNP3: 40 reads A allele, 60 reads B allele (40% A)

[0341] SNP4: A allele in 46 reads, B allele in 54 reads (46% A)

[0342] SNP5: 520 reads A allele, 480 reads B allele (52% A)

[0343] SNP6: 200 reads A allele, 200 reads B allele (50% A)

[0344] From this set of data, it can be difficult to distinguish between cases where the individual is normal and all cells are disomy, and cases where the individual may have cancer, where some of the cells have deletions or duplications in the chromosome, resulting in cell-free DNA present in plasma. For example, two maximum likelihood hypotheses are that the individual has a deletion in this chromosome segment, with a 6% tumor incidence, and that the deleted segment of the chromosome has the genotype (A, B, A, A, B, B) or (A, B, A, A, B, A) for the six SNPs. In representing the individual's genotype for the SNP set, the first letter in parentheses corresponds to the genotype of the haplotype for SNP1, and the second letter corresponds to SNP2.

[0345] Using a method to determine an individual's haplotype at a chromosomal segment, if one finds that the haplotype for one of two chromosomes is (A, B, A, A, B, B), the maximum likelihood hypothesis is met, and the calculated likelihood that the individual has a deletion at that segment, and therefore potentially has cancerous or precancerous cells, increases significantly. On the other hand, if the individual is found to have the haplotype (A, A, A, A, A, A, A), the likelihood that the individual has a deletion at that chromosomal segment decreases significantly, and the likelihood of the no-deletion hypothesis likely increases (the actual likelihood value will depend on other parameters, e.g., measurement noise, among other things, in the system).

[0346] There are many methods for determining an individual's haplotype, many of which are described elsewhere herein. A partial list is provided here and is not intended to be exhaustive. One method is biological, in which individual DNA molecules are diluted until there is approximately one molecule from each chromosomal region in any given reaction volume, and then the genotype is measured using a method such as sequencing. Another method is informatics-based, which can probabilistically use population data on various haplotypes linked to frequencies. Another method is to measure the diploid data of an individual with one or more related individuals, who are predicted to share a haplotype block with the individual and suggest the haplotype block. Another method is to collect tissue samples containing a high concentration of deleted or duplicated segments and measure haplotypes based on allelic imbalance. For example, genotype measurements from tumor tissue samples with deletions can be used to determine phase determination data for the region, which can then be used to determine whether the cancer has regrown after resection.

[0347] In practice, typically more than 20 SNPs, more than 50 SNPs, more than 100 SNPs, more than 500 SNPs, more than 1,000 SNPs, or more than 5,000 SNPs are measured on a given chromosomal segment.

[0348] 9. Examples of Mutations Examples of mutations associated with a disease or disorder, such as cancer, or an increased risk (e.g., above-normal risk) of a disease or disorder, such as cancer, include single nucleotide variants (SNVs), multiple nucleotide mutations, deletions (e.g., deletions of 2-30 million base pair regions), duplications, or tandem repeats. In some embodiments, the mutations are in DNA, such as cfDNA, cell-free mitochondrial DNA (cfmDNA), nuclear DNA (cfnDNA), cellular DNA, or cell-free DNA originating from mitochondrial DNA. In some embodiments, the mutations are in RNA, such as cfRNA, cellular RNA, cytoplasmic RNA, coding cytoplasmic RNA, non-coding cytoplasmic RNA, mRNA, miRNA, mitochondrial RNA, rRNA, or tRNA. In some embodiments, the mutations are more frequently present in subjects with a disease or disorder (e.g., cancer) than in subjects without the disease or disorder (e.g., cancer). In some embodiments, the mutations are indicative of cancer, e.g., causative mutations. In some embodiments, the mutations are driver mutations that play a causative role in the disease or disorder. In some embodiments, the mutations are not causative mutations. For example, some cancers accumulate multiple mutations, but some of them are not causative mutations.Even non-causative mutations (such as mutations that are more frequently present in subjects with disease or disorder than in subjects without disease or disorder) can be useful for diagnosing disease or disorder.In some embodiments, mutation is loss of heterozygosity (LOH) at one or more microsatellites.

[0349] In some embodiments, subjects are screened for one or more polymorphisms or mutations that the subject is known to have (e.g., to test for their presence, changes in the amount of cells, DNA, or RNA that have those polymorphisms or mutations, or for remission or recurrence of cancer). In some embodiments, subjects are screened for one or more polymorphisms or mutations for which the subject is known to be at risk (e.g., subjects who have relatives that have the polymorphism or mutation). In some embodiments, subjects are screened for a panel of polymorphisms or mutations (e.g., at least 5, 10, 50, 100, 200, 300, 500, 750, 1,000, 1,500, 2,000, or 5,000 polymorphisms or mutations) that are associated with a disease or disorder, such as cancer.

[0350] Many coding variants associated with cancer are described in Abaan et al., "The Exomes of the NCI-60 Panel: A Genomic Resource for Cancer Biology and Systems Pharmacology," Cancer Research, July 15, 2013, and are available on the World Wide Web at dtp.nci.nih.gov / branches / btb / characterizationNCI60.html, which is incorporated herein by reference in its entirety. The NCI-60 human cancer cell line panel consists of 60 diverse cell lines representing cancers of the lung, colon, brain, ovary, breast, prostate, and kidney, as well as leukemia and melanoma. The genetic variations identified in these cell lines are of two types: Type I variants present in the normal population, and cancer-specific Type II variants.

[0351] Exemplary polymorphisms or mutations (e.g., deletions or duplications) are in one or more of the following genes: TP53, PTEN, PIK3CA, APC, EGFR, NRAS, NF2, FBXW7, ERBBs, ATAD5, KRAS, BRAF, VEGF, EGFR, HER2, ALK, p53, BRCA, BRCA1, BRCA2, SETD2, LRP1B, PBRM, SPTA1, DNMT3A, ARID1A, GRIN2A, TRRAP, STAG2, EPHA3 / 5 / 7, POLE, SYNE1, C20orf80, CSMD1, CTNNB1, ERBBs 2.FBXW7, KIT, MUC4, ATM, CDH1, DDX11, DDX12, DSPP, EPPK1, FAM186A, GNAS, HRNR, KRTAP4-11, MAP2K4, MLL3, NRAS, RB1, SMAD4, TTN, ABCC9, ACVR1B, ADAM 29, ADAMTS19, AGAP10, AKT1, AMBN, AMPD2, ANKRD30A, ANKRD40, APOBR, AR, BIRC6, BMP2, BRAT1, BTNL8, C12orf4, C1QTNF7, C20orf186, CAPRIN2, CBWD1, C CDC30, CCDC93, CD5L, CDC27, CDC42BPA, CDH9, CDKN2A, CHD8, CHEK2, CHRNA9, CIZ1, CLSPN, CNTN6, COL14A1, CREBBP, CROCC, CTSF, CYP1A2, DCLK1, DHDDS, DHX32, DKK2, DLEC1, DNAH14, DNAH5, DNAH9, DNASE1L3, DUSP16, DYNC2H1, ECT2, EFHB, RRN3P2, TRIM49B, TUBB8P5, EPHA7, ERBB3, ERCC6, FAM21A, FAM21C, FCGBP, FGFR2, FLG2, FLT1, FOLR2, FRYL, FSCB, GAB1, GABRA4, GABRP, GH2, GOLGA6L1, GPHB5, GPR32, GPX5, GTF3C3, HECW1, HIST1H3B, HLA-A, HRAS, HS3ST1 , HS6ST1, HSPD1, IDH1, JAK2, KDM5B, KIAA0528, KRT15, ​​KRT38, KRTAP21-1, KRTAP4-5, KRTAP4-7, KRTAP5-4, KRTAP5-5, LAMA4, LATS1, LMF1, LPAR4, LPPR4,<h2 style=";text-align:left;direction:ltr">LRRFIP1、LUM、LYST、MAP2K1、MARCH1、MARCO、MB21D2、MEGF10、MMP16、MORC1 、MRE11A、MTMR3、MUC12、MUC17、MUC2、MUC20、NBPF10、NBPF20、NEK1、NFE2L2 、NLRP4、NOTCH2、NRK、NUP93、OBSCN、OR11H1、OR2B11、OR2M4、OR4Q3、OR5D13 、OR8I2、OXSM、PIK3R1、PPP2R5C、PRAME、PRF1、PRG4、PRPF19、PTH2、PTPRC、PT PRJ、RAC1、RAD50、RBM12、RGPD3、RGS22、ROR1、RP11-671M22.1、RP13-996F3 4、RP1L1、RSBN1L、RYR3、SAMD3、SCN3A、SEC31A、SF1、SF3B1、SLC25A2、SLC44 A1、SLC4A11、SMAD2、SPTA1、ST6GAL2、STK11、SZT2、TAF1L、TAX1BP1、TBP、TG FBI、TIF1、TMEM14B、TMEM74、TPTE、TRAPPC8、TRPS1、TXNDC6、USP32、UTP20、V ASN、VPS72、WASH3P、WWTR1、XPO1、ZFHX4、ZMIZ1、ZNF167、ZNF436、ZNF492、ZNF598、ZRSR2、ABL1、AKT2、AKT3、ARAF、ARFRP1、ARID2、ASXL1、ATR、ATRX、AU RKA、AURKB、AXL、BAP1、BARD1、BCL2、BCL2L2、BCL6、BCOR、BCORL1、BLM、BRIP 1、BTK、CARD11、CBFB、CBL、CCND1、CCND2、CCND3、CCNE1、CD79A、CD79B、CDC73 、CDK12、CDK4、CDK6、CDK8、CDKN1B、CDKN2B、CDKN2C、CEBPA、CHEK1、CIC、CRK L、CRLF2、CSF1R、CTCF、CTNNA1、DAXX、DDR2、DOT1L、EMSY(C11orf30)、EP300、 EPHA3、EPHA5、EPHB1、ERBB4、ERG、ESR1、EZH2、FAM123B(WTX)、FAM46C、FANC A、FANCC、FANCD2、FANCE、FANCF、FANCG、FANCL、FGF10、FGF14、FGF19、FGF23、FGF3, FGF4, FGF6, FGFR1, FGFR2, FGFR3, FGFR4, FLT3, FLT4, FOXL2, GATA1, GATA2, GATA3, GID4(C17orf39), G NA11, GNA13, GNAQ, GNAS, GPR124, GSK3B, HGF, IDH1, IDH2, IGF1R, IKBKE, IKZF1, IL7R, INHBA, IRF4, IRS2, JA K1, JAK3, JUN, KAT6A(MYST3), KDM5A, KDM5C, KDM6A, KDR, KEAP1, KLHL6, MAP2K2, MAP2K4, MAP3K1, MCL1, MDM2 , MDM4, MED12, MEF2B, MEN1, MET, MITF, MLH1, MLL, MLL2, MPL, MSH2, MSH6, MTOR, MUTYH, MYC, MYCL1, MYCN, MYD8 8, NF1, NFKBIA, NKX2-1, NOTCH1, NPM1, NRAS, NTRK1, NTRK2, NTRK3, PAK3, PALB2, PAX5, PBRM1, PDGFRA, PDGFR B, PDK1, PIK3CG, PIK3R2, PPP2R1A, PRDM1, PRKAR1A, PRKDC, PTCH1, PTPN11, RAD51, RAF1, RARA, RET, RICTOR, R NF43, RPTOR, RUNX1, SMARCA4, SMARCB1, SMO, SOCS1, SOX10, SOX2, SPEN, SPOP, SRC, STAT4, SUFU, TET2, TGFBR2, TNFAIP3, TNFRSF14, TOP1, TP53, TSC1, TSC2, TSHR, VHL, WISP3, WT1, ZNF217, ZNF703, and combinations thereof (Su et al., J Mol Diagn 2011, 13:74-84; DOI:10.1016 / j.jmoldx.2010.11.010; and Abaan et al., "The Exomes of the NCI-60 Panel: A Genomic Resource for Cancer Biology and Systems Pharmacology," Cancer Research, July 15, 2013, which are incorporated by reference in their entireties). In some embodiments, the duplication is a chromosome 1p (Chr1p) duplication associated with breast cancer. In some embodiments, the one or more polymorphisms or mutations are:In some embodiments, the polymorphism or mutation is in BRAF, e.g., a V600E mutation. In some embodiments, the one or more polymorphisms or mutations are in K-ras. In some embodiments, a combination of one or more polymorphisms or mutations in K-ras and APC is present. In some embodiments, a combination of one or more polymorphisms or mutations in K-ras and p53 is present. In some embodiments, a combination of one or more polymorphisms or mutations in APC and p53 is present. In some embodiments, a combination of one or more polymorphisms or mutations in K-ras, APC and p53 is present. In some embodiments, a combination of one or more polymorphisms or mutations in K-ras and EGFR is present. Exemplary polymorphisms or mutations are in one or more of the following microRNAs: miR-15a, miR-16-1, miR-23a, miR-23b, miR-24-1, miR-24-2, miR-27a, miR-27b, miR-29b-2, miR-29c, miR-146, miR-155, miR-221, miR-222, and miR-223 (Calin et al., "A microRNA signature associated with prognosis and progression in chronic lymphocytic leukemia," N Engl J Med 353:1793-801, 2005, which is incorporated herein by reference in its entirety).

[0352] In some embodiments, the deletion is at least 0.01 kb, 0.1 kb, 1 kb, 10 kb, 100 kb, 1 mb, 2 mb, 3 mb, 5 mb, 10 mb, 15 mb, 20 mb, 30 mb, or 40 mb. In some embodiments, the deletion is 1 kb to 40 mb, e.g., 1 kb to 100 kb, 100 kb to 1 mb, 1 to 5 mb, 5 to 10 mb, 10 to 15 mb, 15 to 20 mb, 20 to 25 mb, 25 to 30 mb, or 30 to 40 mb.

[0353] In some embodiments, the overlap is at least 0.01 kb, 0.1 kb, 1 kb, 10 kb, 100 kb, 1 mb, 2 mb, 3 mb, 5 mb, 10 mb, 15 mb, 20 mb, 30 mb, or 40 mb. In some embodiments, the overlap is at least 1 kb to 40 mb, e.g., at least 1 kb to 100 kb, at least 100 kb to 1 mb, at least 1 to 5 mb, at least 5 to 10 mb, at least 10 to 15 mb, at least 15 to 20 mb, at least 20 to 25 mb, at least 25 to 30 mb, or at least 30 to 40 mb.

[0354] In some embodiments, the tandem repeat is a repeat of 2 to 60 nucleotides, e.g., 2 to 6, 7 to 10, 10 to 20, 20 to 30, 30 to 40, 40 to 50, or 50 to 60 nucleotides. In some embodiments, the tandem repeat is a repeat of two nucleotides (dinucleotide repeat). In some embodiments, the tandem repeat is a repeat of three nucleotides (trinucleotide repeat).

[0355] In some embodiments, the polymorphism or mutation is predictive. Exemplary predictive mutations include, for example, K-ras mutations, such as K-ras mutations that are indicative of postoperative disease recurrence in colorectal cancer (Ryan et al., "A prospective study of circulating mutant KRAS2 in the serum of patients with colorectal neoplasia: a strong prognostic indicator in postoperative follow-up," Gut 52:101-108, 2003; and Lecomte T et al., Detection of free-circulating tumor-associated DNA in plasma of colorectal cancer patients and its association with prognosis," Int J Cancer 100:542-548, 2002, each of which is incorporated herein by reference in its entirety).

[0356] In some embodiments, the polymorphism or mutation is associated with an altered response to a particular treatment (e.g., increased or decreased efficacy or side effects). An example is K-ras mutations, which are associated with decreased responsiveness to EGFR-based treatments in non-small cell lung cancer (Wang et al., "Potential clinical significance of a plasma-based KRAS mutation analysis in patients with advanced non-small cell lung cancer," Clin Canc Res 16:1324-1330, 2010, which is incorporated herein by reference in its entirety).

[0357] K-ras is an oncogene and is activated in many cancers. Exemplary K-ras mutations are mutations at codons 12, 13, and 61. K-ras cfDNA mutations have been identified in pancreatic cancer, lung cancer, colorectal cancer, bladder cancer, and gastric cancer (Fleischhacker & Schmidt "Circulating nucleic acids (CNAs) and caner - a survey," Biochim Biophys Acta 1775:181-232, 2007, which is incorporated herein by reference in its entirety).

[0358] p53 is a tumor suppressor that is mutated in many cancers, contributing to tumor progression (Levine & Oren, "The first 30 years of p53: growing ever more complex. Nature Rev Cancer," 9:749-758, 2009, which is incorporated herein by reference in its entirety). Many different codons can be mutated, including Ser249. p53 cfDNA mutations have been identified in breast, lung, ovarian, bladder, gastric, pancreatic, colorectal, colon, and hepatocellular carcinoma (Fleischhacker & Schmidt, "Circulating nucleic acids (CNAs) and caner—a survey," Biochim Biophys Acta 1775:181-232, 2007, which is incorporated herein by reference in its entirety).

[0359] BRAF is an oncogene downstream of Ras. BRAF mutations have been identified in glial neoplasms, melanoma, thyroid cancer, and lung cancer (Dias-Santagata et al., BRAF V600E mutations are common in pleomorphic xanthoastrocytoma: diagnostic and therapeutic implications. PLOS ONE 2011;6:e17948, 2011; Shinozaki et al., Utility of circulating B-RAF DNA mutations in serum for monitoring melanoma patients receiving biochemotherapy. Clin Canc Res 13:2068-2074, 2007; and Board et al., Detection of BRAF mutations in the tumor and serum of patients enrolled in the AZD6244 (ARRY-142886) advanced melanoma phase II study. Brit J Canc 2009;101:1724-1730, each of which is incorporated herein by reference in its entirety). The BRAF V600E mutation occurs, for example, in melanoma and is more common in advanced stages. The V600E mutation has been detected in cfDNA.

[0360] EGFR contributes to cell proliferation and is misregulated in many cancers (Downward J. Targeting RAS signaling pathways in cancer therapy. Nature Rev Cancer 3:11-22, 2003; and Levine & Oren, "The first 30 years of p53: growing ever more complex. Nature Rev Cancer," 9:749-758, 2009, which are incorporated by reference in their entireties). Exemplary EGFR mutations include those in exons 18-21, which have been identified in lung cancer patients. EGFR cfDNA mutations have been identified in lung cancer patients (Jia et al., "Prediction of epidermal growth factor receptor mutations in the plasma / pleural effusion to efficacy of gefitinib treatment in advanced non-small cell lung cancer," J Canc Res Clin Oncol 2010;136:1341-1347, 2010, which are incorporated by reference in their entireties).

[0361] Exemplary polymorphisms or mutations associated with breast cancer include LOH at microsatellites (Kohler et al., "Levels of plasma circulating cell-free nuclear and mitochondrial DNA as potential biomarkers for breast tumors," Mol Cancer 8:doi:10.1186 / 1476-4598-8-105, 2009, which is incorporated herein by reference in its entirety), p53 mutations (e.g., mutations in exons 5-8) (Garcia et al., "Extracellular tumor DNA in plasma and overall survival in breast cancer patients," Genes, Chromosomes & Cancer 45:692-701, 2006, which is incorporated herein by reference in its entirety), and HER2 (Sorensen et al., "Circulating HER2 DNA after trastuzumab treatment predicts survival and response in breast cancer," Anticancer Res 30:2463-2468, 2010, which is incorporated herein by reference in its entirety), and polymorphisms or mutations in PIK3CA, MED1, and GAS6 (Murtaza et al., "Non-invasive analysis of acquired resistance to cancer therapy by sequencing of plasma DNA," Nature 2013; doi:10.1038 / nature12065, 2013, which is incorporated herein by reference in its entirety).

[0362] Increased cfDNA levels and LOH are associated with decreased overall survival and disease-free survival. p53 mutations (exons 5-8) are associated with decreased overall survival. Decreased circulating HER2 cfDNA levels are associated with better response to HER2-targeted therapy in HER2-positive breast cancer patients. Activating mutations in PIK3CA, truncations in MED1, and splicing mutations in GAS6 result in treatment resistance.

[0363] Examples of polymorphisms or mutations associated with colorectal cancer include mutations in p53, APC, K-ras, and thymidylate synthase, as well as methylation of the p16 gene (Wang et al., "Molecular detection of APC, K-ras, and p53 mutations in the serum of colorectal cancer patients as circulating biomarkers," World J Surg 28:721-726, 2004; Ryan et al., "A prospective study of circulating mutant KRAS2 in the serum of patients with colorectal neoplasia: a strong prognostic indicator in postoperative follow-up," Gut 52:101-108, 2003; Lecomte et al., "Detection of free-circulating tumor-associated DNA in plasma of colorectal cancer patients and its association with prognosis," Int J Cancer 100:542-548, 2002; Schwarzenbach et al., "Molecular analysis of the polymorphisms of "Thymidylate synthase on cell-free circulating DNA in blood of patients with advanced colorectal carcinoma," Int J Cancer 127:881-888, 2009, each of which is incorporated by reference in its entirety.) Postoperative detection of serum K-ras mutations is a strong predictor of disease recurrence. Detection of K-ras mutations and p16 gene methylation is associated with decreased survival and increased disease recurrence. Detection of mutations in K-ras, APC, and / or p53 is associated with recurrence and / or metastasis.Using cfDNA, polymorphisms (including loss of heterozygosity, single nucleotide polymorphisms, variable number of tandem repeats, and deletions) in the thymidylate synthase (a target of fluoropyrimidine chemotherapy) gene may be associated with treatment response.

[0364] Examples of polymorphisms or mutations associated with lung cancer (e.g., non-small cell lung cancer) include K-ras (e.g., codon 12 mutations) and EGFR mutations. Exemplary predictive mutations include EGFR mutations (exon 19 deletions or exon 21 mutations), which are associated with increased overall survival and progression-free survival, and K-ras mutations (codon 12 and 13 mutations) are associated with decreased progression-free survival (Jian et al., "Prediction of epidermal growth factor receptor mutations in the plasma / pleural effusion to the efficacy of gefitinib treatment in advanced non-small cell lung cancer," J Canc Res Clin Oncol 136:1341-1347, 2010; Wang et al., "Potential clinical significance of a plasma-based KRAS mutation analysis in patients with advanced non-small cell lung cancer," Clin Canc Res 16:1324-1330, 2010, each of which is incorporated herein by reference in its entirety). Examples of polymorphisms or mutations that are indicative of response to treatment include EGFR mutations (exon 19 deletion or exon 21 mutation) that improve response to treatment, and K-ras mutations (codons 12 and 13) that decrease response to treatment. Resistance-contributing mutations have been identified in EFGR (Murtaza et al., "Non-invasive analysis of acquired resistance to cancer therapy by sequencing of plasma DNA," Nature doi:10.1038 / nature12065, 2013, which is incorporated herein by reference in its entirety).

[0365] Examples of polymorphisms or mutations associated with melanoma (e.g., uveal melanoma) include polymorphisms or mutations in GNAQ, GNA11, BRAF, and p53. Exemplary GNAQ and GNA11 mutations include R183 and Q209 mutations. Q209 mutations in GNAQ or GNA11 are associated with bone metastasis. BRAF V600E mutations can be detected in patients with metastatic / advanced melanoma. BRAF V600E is an indicator of invasive melanoma. The presence of BRAF V600E mutations after chemotherapy is associated with non-responsiveness to treatment.

[0366] Examples of polymorphisms or mutations associated with pancreatic cancer include K-ras and p53 (e.g., p53 Ser249) polymorphisms or mutations. p53 Ser249 is also associated with hepatitis B infection and hepatocellular carcinoma, as well as ovarian cancer and non-Hodgkin's lymphoma.

[0367] Even polymorphisms or mutations that exist at low frequencies in samples can be detected using the method of the present invention.For example, polymorphisms or mutations that exist at a frequency of 1 in 1 million can be measured 10 times by performing 10 million sequence reads.If desired, the number of sequence reads can be changed according to the desired sensitivity level.In some embodiments, the sample is reanalyzed or another sample from the subject is analyzed using a larger number of sequence reads to improve sensitivity.For example, if no polymorphisms or mutations are detected that are associated with cancer or associated with increased risk of cancer, or only a small number (for example, 1, 2, 3, 4, or 5), the sample is reanalyzed or another sample is tested.

[0368] In some embodiments, multiple polymorphisms or mutations are required for cancer or metastatic cancer. In such cases, screening for multiple polymorphisms or mutations improves the ability to accurately diagnose cancer or metastatic cancer. In some embodiments, a subject has a subset of multiple polymorphisms or mutations required for cancer or metastatic cancer, and the subject can be re-screened later to determine whether the subject has acquired additional mutations.

[0369] In some embodiments where multiple polymorphisms or mutations are required for cancer or metastatic cancer, the frequencies of each polymorphism or mutation can be compared to determine whether they occur at similar frequencies. For example, if two mutations (listed as "A" and "B") are required for cancer, some cells will have none, some cells will have A, some cells will have B, and some cells will have A and B. If A and B are observed at similar frequencies, the subject is more likely to have some cells that have both A and B. If A and B are observed at different frequencies, the subject is more likely to have different cell populations.

[0370] In some embodiments where multiple polymorphisms or mutations are required for cancer or metastatic cancer, the number or uniqueness of such polymorphisms or mutations present in a subject can be used to predict the likelihood or when the subject will have the disease or disorder. In some embodiments where polymorphisms or mutations tend to occur in a particular order, subjects may be tested periodically to determine if they have acquired other polymorphisms or mutations.

[0371] In some embodiments, determining the presence or absence of multiple polymorphisms or mutations (e.g., 2, 3, 4, 5, 8, 10, 12, 15, or more) increases the sensitivity and / or specificity of determining the presence or absence of a disease or disorder, e.g., cancer, or the presence or absence of an increased risk of having a disease or disorder, e.g., cancer.

[0372] In some embodiments, a polymorphism or mutation is detected directly, hi some embodiments, a polymorphism or mutation is detected indirectly by detection of one or more sequences (e.g., polymorphic loci such as SNPs) linked to the polymorphism or mutation.

[0373] 9.1 Exemplary Nucleic Acid Modifications In some embodiments, there are changes in the integrity of RNA or DNA (e.g., size changes in fragmented cfRNA or cfDNA, or changes in nucleosome composition) associated with diseases or disorders, such as cancer, or increased risk of diseases or disorders, such as cancer. In some embodiments, there are changes in the methylation pattern of RNA or DNA (e.g., hypermethylation of tumor suppressor genes) associated with diseases or disorders, such as cancer, or increased risk of diseases or disorders, such as cancer. For example, it has been suggested that methylation of CpG islands in the promoter regions of tumor suppressor genes induces local gene silencing. Abnormal methylation of p16 tumor suppressor genes occurs in subjects with liver cancer, lung cancer, and breast cancer. Other frequently methylated tumor suppressor genes, including APC, Ras association domain family protein 1A (RASSF1A), glutathione S-transferase P1 (GSTP1), and DAPK, have been detected in various types of cancer, such as nasopharyngeal carcinoma, colorectal cancer, lung cancer, esophageal cancer, prostate cancer, bladder cancer, melanoma, and acute leukemia. Methylation of certain tumor suppressor genes, such as p16, has been reported as an early event in cancer formation and is useful for early cancer screening.

[0374] In some embodiments, methylation patterns are determined using bisulfite conversion or non-bisulfite methods using methylation-sensitive restriction enzyme digestion (Hung et al., J Clin Pathol 62:308-313, 2009, which is incorporated herein by reference in its entirety). In the bisulfite conversion method, methylated cytosines remain as cytosines and unmethylated cytosines are converted to uracil. Methylation-sensitive restriction enzymes (e.g., BstUI) cleave unmethylated DNA sequences at specific recognition sites (e.g., 5'-CG v CG-3 for BstUI), while methylated sequences remain intact. In some embodiments, intact methylated sequences are detected. In some embodiments, stem-loop primers are used to selectively amplify unmethylated restriction enzyme-digested fragments without co-amplifying undigested methylated DNA.

[0375] 9.2 Examples of Alterations in mRNA Splicing In some embodiments, altered mRNA splicing is associated with a disease or disorder, such as cancer, or an increased risk of a disease or disorder, such as cancer. In some embodiments, the altered mRNA splicing is present in one or more of the following nucleic acids associated with cancer or an increased risk of cancer: DNMT3B, BRCA1, KLF6, Ron, or Gemin5. In some embodiments, the detected mRNA splice variant is associated with a disease or disorder, such as cancer. In some embodiments, multiple mRNA splice variants are produced by healthy cells (e.g., non-cancerous cells), and changes in the relative amounts of mRNA splice variants are associated with a disease or disorder, such as cancer. In some embodiments, altered mRNA splicing is due to changes in mRNA sequence (e.g., mutations in splice sites), changes in the levels of splicing factors, changes in the amount of available splicing factors (e.g., reduced amounts of available splicing factors due to splicing factor binding to repeats), changes in splicing regulation, or the tumor microenvironment.

[0376] The splicing reaction is carried out by a multiprotein / mRNA complex called the spliceosome (Fackenthal and Godley, Disease Models & Mechanisms 1:37-42, 2008, doi:10.1242 / dmm.000331, incorporated herein by reference in its entirety). The spliceosome recognizes intron-exon boundaries and removes the intervening intron through two transesterification reactions, resulting in the ligation of two adjacent exons. The fidelity of this reaction must be utmost, as erroneous ligation can reduce the likelihood of encoding a normal protein. For example, if exon skipping preserves the triplet codon reading frame that specifies the identity and order of amino acids during translation, alternatively spliced ​​mRNAs may specify proteins lacking critical amino acid residues. More commonly, exon skipping disrupts the translation reading frame and results in premature stop codons. These mRNAs are often degraded by at least 90% through a process known as nonsense-mediated mRNA decay, which reduces the likelihood that such defective messages will accumulate and generate truncated protein products. If mis-spliced ​​mRNAs escape this pathway, truncated, mutant, or unstable proteins will be produced.

[0377] Alternative splicing, a means of expressing multiple or many different transcripts from the same genomic DNA, results from the integration of a subset of available exons for a particular protein. By excluding one or more exons, specific protein domains can be lost from the encoded protein, potentially resulting in loss-of-function or gain-of-function of the protein. Several types of alternative splicing have been reported: exon skipping, alternative 5' or 3' splice sites, mutually exclusive exons, and, very rarely, intron retention. Using bioinformatics methods, some researchers have compared the amount of alternative splicing in cancer and normal cells and determined that cancer exhibits lower levels of alternative splicing than normal cells. Furthermore, the distribution of types of alternative splicing events differed between cancer and normal cells. Cancer cells showed less exon skipping than normal cells but more alternative 5' and 3' splice site selection and intron retention. When exonization (the use of sequences primarily used as introns in other tissues as exons) was examined, genes associated with exonization in cancer cells were preferentially associated with mRNA processing. This suggests a direct link between cancer cells and the generation of aberrant mRNA splice forms.

[0378] 9.3 Examples of changes in DNA or RNA levels In some embodiments, there is a change in the total amount or concentration of one or more types of DNA (e.g., cfDNA, cfmDNA, cfnDNA, cellular DNA, or mitochondrial DNA) or RNA (cfRNA, cellular RNA, cytoplasmic RNA, coding cytoplasmic RNA, non-coding cytoplasmic RNA, mRNA, miRNA, mitochondrial RNA, rRNA, or tRNA). In some embodiments, there is a change in the amount or concentration of one or more specific DNA (e.g., cfDNA, cfmDNA, cfnDNA, cellular DNA, or mitochondrial DNA) molecules or RNA (cfRNA, cellular RNA, cytoplasmic RNA, coding cytoplasmic RNA, non-coding cytoplasmic RNA, mRNA, miRNA, mitochondrial RNA, rRNA, or tRNA) molecules. In some embodiments, one allele is more highly expressed than another allele at the locus of interest. Exemplary miRNAs are short 20-22 nucleotide RNA molecules that regulate gene expression. In some embodiments, there is a change in the transcriptome, such as a change in the identity or abundance of one or more RNA molecules.

[0379] In some embodiments, an increase in the amount or concentration of cfDNA or cfRNA is associated with a disease or disorder, such as cancer, or an increased risk of a disease or disorder, such as cancer. In some embodiments, the total concentration of a certain type of DNA (e.g., cfDNA, cfmDNA, cfnDNA, cellular DNA, or mitochondrial DNA) or RNA (e.g., cfRNA, cellular RNA, cytoplasmic RNA, coding cytoplasmic RNA, non-coding cytoplasmic RNA, mRNA, miRNA, mitochondrial RNA, rRNA, or tRNA) is increased by at least 2, 3, 4, 5, 6, 7, 8, 9, 10, or more times compared to the total concentration of that type of DNA or RNA in a healthy (e.g., non-cancerous) subject. In some embodiments, a total concentration of cfDNA of 75 to 100 ng / mL, 100 to 150 ng / mL, 150 to 200 ng / mL, 200 to 300 ng / mL, 300 to 400 ng / mL, 400 to 600 ng / mL, 600 to 800 ng / mL, or 800 to 1,000 ng / mL, or a total concentration of cfDNA greater than 100 ng, e.g., greater than 200, 300, 400, 500, 600, 700, 800, 900, or 1,000 ng / mL, is indicative of cancer, an increased risk of cancer, an increased risk of tumors that are more malignant than benign, a decrease in cancer possibly going into remission, or a worsening prognosis for the cancer. In some embodiments, the amount of a type of DNA (e.g., cfDNA, cf mDNA, cf nDNA, cellular DNA, or mitochondrial DNA), or RNA (cfRNA, cellular RNA, cytoplasmic RNA, coding cytoplasmic RNA, non-coding cytoplasmic RNA, mRNA, miRNA, mitochondrial RNA, rRNA, or tRNA) that has one or more polymorphisms / mutations (e.g., deletions or duplications) associated with a disease or disorder, e.g., cancer, or an increased risk of a disease or disorder, e.g., cancer, is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 14, 16, 18, 20, or 25% of the total amount of DNA or RNA of that type.In some embodiments, at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 14, 16, 18, 20, or 25% of the total amount of a type of DNA (e.g., cfDNA, cf mDNA, cf nDNA, cellular DNA, or mitochondrial DNA), or RNA (cfRNA, cellular RNA, cytoplasmic RNA, coding cytoplasmic RNA, non-coding cytoplasmic RNA, mRNA, miRNA, mitochondrial RNA, rRNA, or tRNA) has a particular polymorphism or mutation (e.g., a deletion or duplication) associated with a disease or disorder, e.g., cancer, or an increased risk of a disease or disorder, e.g., cancer.

[0380] In some embodiments, the cfDNA is encapsulated. In some embodiments, the cfDNA is unencapsulated.

[0381] In some embodiments, the percentage of tumor DNA in total DNA (e.g., the percentage of tumor cfDNA in total cfDNA, or the percentage of tumor cfDNA in total cfDNA with a specific mutation) is determined. In some embodiments, the percentage of tumor DNA may be determined for multiple mutations, where the mutations may be single base variants, copy number variants, differential methylation, or a combination thereof. In some embodiments, the average tumor fraction calculated for one or a set of mutations with the highest calculated tumor fraction is considered to be the actual tumor fraction in the sample. In some embodiments, the average tumor fraction calculated for all of the mutations is considered to be the actual tumor fraction in the sample. In some embodiments, this tumor fraction is used to stage the cancer (as a higher tumor fraction may indicate a more advanced stage of cancer). In some embodiments, the tumor fraction is used to size the cancer, as tumor DNA fraction in plasma and tumor size may correlate. In some embodiments, the tumor fraction is used to size the tumor fraction with one or more mutations, as there may be a correlation between the observed tumor fraction in a plasma sample and the size of tissue with a given mutation genotype. For example, the amount of tissue carrying a given mutant genotype may correlate with the proportion of tumor DNA, which can be calculated by focusing on that particular mutation.

[0382] 10. Exemplary Biochemical Methods and Compositions 10.1 Sample Examples In some embodiments of any of the aspects of the invention, the sample contains cellular and / or extracellular genetic material from cells suspected of having a deletion or duplication, such as cells suspected of being cancerous or circulating cells suspected of fetal origin. In some embodiments, the sample includes any tissue or bodily fluid suspected of containing cells, DNA, or RNA with a deletion or duplication. The genetic measurements used as part of these methods may be performed on any sample containing DNA or RNA, including, but not limited to, tissue, blood, serum, plasma, urine, hair, tears, saliva, skin, nails, stool, bile, lymph, cervical mucus, semen, tumor, fetus, or other cells or materials containing nucleic acids. The sample may include or be used with any cell type, or DNA or RNA derived from any cell type (e.g., cells from any organ or tissue suspected of being cancerous or neuronal). In some embodiments, the sample contains nuclear DNA and / or mitochondrial DNA. In some embodiments, the sample is from any of the target individuals disclosed herein. In some embodiments, the target individual is a cancer patient.

[0383] Exemplary samples include samples containing cfDNA or cfRNA. In some embodiments, cfDNA can be used for analysis without the need to dissolve cells. Cell-free DNA can be obtained from various tissues, such as liquid tissues, such as blood, plasma, lymph, ascites, or cerebrospinal fluid. In some cases, cfDNA is composed of DNA derived from fetal cells. In some cases, cfDNA is isolated from plasma, which is separated from whole blood that has been centrifuged to remove cellular material. cfDNA can also be a mixture of DNA derived from target cells (e.g., cancer cells) and non-target cells (e.g., non-cancerous cells).

[0384] In some embodiments, sample contains or is suspected to contain a mixture of DNA (or RNA), such as a mixture of DNA (or RNA) originating from cancer cells and DNA (or RNA) originating from non-cancerous (i.e., normal) cells.In some embodiments, at least 0.5, 1, 3, 5, 7, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 92, 94, 95, 96, 98, 99 or 100% of the cells in sample are cancer cells.In some embodiments, at least 0.5, 1, 3, 5, 7, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 92, 94, 95, 96, 98, 99 or 100% of the DNA (e.g., cfDNA) or RNA (e.g., cfRNA) in sample are derived from cancer cells. In various embodiments, the percent of cells in a sample that are cancerous is between 0.5% and 99%, e.g., between 1% and 95%, between 5% and 95%, between 10% and 90%, between 5% and 70%, between 10% and 70%, between 20% and 90%, or between 20% and 70%. In some embodiments, the sample is enriched for cancer cells or enriched for DNA or RNA from cancer cells. In some embodiments in which the sample is enriched for cancer cells, at least 0.5, 1, 2, 3, 4, 5, 6, 7, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 92, 94, 95, 96, 98, 99, or 100% of the cells in the enriched sample are cancer cells. In some embodiments where the sample is enriched for DNA or RNA from cancer cells, at least 0.5, 1, 2, 3, 4, 5, 6, 7, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 92, 94, 95, 96, 98, 99, or 100% of the DNA or RNA in the enriched sample is from cancer cells.In some embodiments, cell sorting (such as, for example, fluorescence-activated cell sorting (FACS)) is used to enrich for cancer cells (see Barteneva et al., Biochim Biophys Acta., 1836(1):105-22, Aug 2013. doi:10.1016 / j.bbcan.2013.02.004. Epub 2013 Feb 24, and Ibrahim et al., Adv Biochem Eng Biotechnol. 106:19-39, 2007, each of which is incorporated herein by reference in its entirety).

[0385] In some embodiments, the sample is enriched for fetal cells. In some embodiments, where the sample is enriched for fetal cells, at least 0.5, 1, 2, 3, 4, 5, 6, 7% or more of the cells in the enriched sample are fetal cells. In some embodiments, the percentage of cells in the sample that are fetal cells is between 0.5% and 100%, e.g., between 1% and 99%, between 5% and 95%, between 10% and 95%, between 10% and 95%, between 20% and 90%, or between 30% and 70%. In some embodiments, the sample is enriched for fetal DNA. In some embodiments, where the sample is enriched for fetal DNA, at least 0.5, 1, 2, 3, 4, 5, 6, 7% or more of the DNA in the enriched sample is fetal DNA. In some embodiments, the percent of DNA in the sample that is fetal DNA is between 0.5% and 100%, e.g., between 1% and 99%, between 5% and 95%, between 10% and 95%, between 10% and 95%, between 20% and 90%, or between 30% and 70%.

[0386] In some embodiments, a sample contains a single cell or DNA and / or RNA derived from a single cell. In some embodiments, multiple individual cells (e.g., at least 5, 10, 20, 30, 40, or 50 cells from the same or different subjects) are analyzed in parallel. In some embodiments, cells from multiple samples from the same individual are combined, thereby reducing the amount of work compared to analyzing the samples separately. Combining multiple samples also allows multiple tissues to be tested for cancer simultaneously (which can be used to provide more thorough screening for cancer or to determine whether cancer has metastasized to other tissues).

[0387] In some embodiments, the sample contains a single cell or a small number of cells, e.g., 2, 3, 5, 6, 7, 8, 9, or 10 cells. In some embodiments, the sample has 1 to 100, 100 to 500, or 500 to 1,000 cells. In some embodiments, the sample contains between 1 picogram and 10 picograms, between 10 picograms and 100 picograms, between 100 picograms and 1 nanogram, between 1 nanogram and 10 nanograms, between 10 nanograms and 100 nanograms, or between 100 nanograms and 1 microgram of RNA and / or DNA.

[0388] In some embodiments, the sample is paraffin-embedded. In some embodiments, the sample is preserved using a preservative, such as formaldehyde, and optionally paraffin-embedded, which may cross-link DNA and make less DNA available for PCR. In some embodiments, the sample is a formaldehyde-fixed, paraffin-embedded (FFPE) sample. In some embodiments, the sample is a fresh sample (e.g., a sample obtained within one or two days of analysis). In some embodiments, the sample is frozen before analysis. In some embodiments, the sample is a clinical history sample.

[0389] These samples can be used in any of the methods of the present invention.

[0390] 10.2 Sample Preparation Method Examples In some embodiments, the method involves isolating or purifying DNA and / or RNA. There are many standard methods known in the art to achieve this goal. In some embodiments, the sample may be centrifuged and the various layers may be separated. In some embodiments, DNA or RNA may be isolated using filtration. In some embodiments, DNA or RNA preparation may include amplification, separation, chromatographic purification, liquid separation, isolation, preferential enrichment, preferential amplification, targeted amplification, or any of many other techniques known in the art or described herein. In some embodiments involving DNA isolation, RNase is used to degrade RNA. In some embodiments involving RNA isolation, DNase (e.g., DNase from Invitrogen, Carlsbad, CA, USA) is used to degrade DNA. In some embodiments, RNA is isolated using the RNeasy mini kit (Qiagen) according to the manufacturer's protocol. In some embodiments, small RNA molecules are isolated using the mirVana PARIS kit (Ambion, Austin, Texas, USA) according to the manufacturer's protocol (Gu et al., J. Neurochem. 122:641-649, 2012, which is incorporated herein by reference in its entirety). RNA concentration and purity may optionally be determined using Nanovue (GE Healthcare, Piscataway, New Jersey, USA), and RNA integrity may optionally be measured using the 2100 Bioanalyzer (Agilent Technologies, Santa Clara, California, USA) (Gu et al., J. Neurochem. 122:641-649, 2012, which is incorporated herein by reference in its entirety). In some embodiments, TRIZOL or RNAlater (Ambion) are used to stabilize RNA during storage.

[0391] In some embodiments, universally tagged adapters are added to create a library. Prior to ligation, sample DNA may be blunt-ended, and then a single adenosine base is added to the 3-prime end. Prior to ligation, the DNA may be cleaved using a restriction enzyme or some other cleavage method. During ligation, the 3-prime adenosine of the sample fragment and the complementary 3-prime tyrosine overhang of the adapter enhance ligation efficiency. In some embodiments, adapter ligation is performed using a ligation kit such as in the AGILENT SURESELECT kit. In some embodiments, the library is amplified using universal primers. In one embodiment, the amplified library is fractionated by size separation or by using a product such as AGENCOURT AMPURE beads or other similar methods. In some embodiments, target loci are amplified using PCR amplification. In some embodiments, the amplified DNA is sequenced (e.g., using an ILLUMINA IIGAX or HiSeq sequencer). In some embodiments, the amplified DNA is sequenced from each end of the amplified DNA to reduce sequencing errors. If there is a sequence error at a particular base when sequencing from one end of the amplified DNA, there is less likely to be a sequence error at the complementary base when sequencing from the other end of the amplified DNA (compared to sequencing from the same end of the amplified DNA multiple times).

[0392] In some embodiments, a whole genome application (WGA) is used to amplify nucleic acid samples. Numerous methods are available for WGA: ligation-mediated PCR (LM-PCR), degenerate oligonucleotide primer PCR (DOP-PCR), and multiple displacement amplification (MDA). In LM-PCR, short DNA sequences called adapters are ligated to blunt ends of DNA. These adapters contain universal amplification sequences that are used for DNA amplification by PCR. In DOP-PCR, random primers that also contain universal amplification sequences are used in the first round of annealing and PCR. The sequences are then further amplified with universal primer sequences in a second round of PCR. MDA uses phi-29 polymerase, a highly processive, nonspecific enzyme that replicates DNA and has been used in single-cell analysis. In some embodiments, WGA is not performed.

[0393] In some embodiments, selective amplification or enrichment is used to amplify or enrich target loci. In some embodiments, amplification and / or selective enrichment techniques can involve PCR, such as ligation-mediated PCR, hybridization fragment capture, molecular inversion probes, or other circularizing probes. In some embodiments, real-time quantitative PCR (RT-qPCR), digital PCR, or emulsion PCR, single-allele base extension reactions followed by mass spectrometry are used (Hung et al., J Clin Pathol 62:308-313, 2009, incorporated herein by reference in its entirety). In some embodiments, hybridization capture using hybrid capture probes is used to preferentially enrich DNA. In some embodiments, amplification or selective enrichment methods involve the use of probes where, when correctly hybridized to a target sequence, the 3-prime or 5-prime end of the nucleotide probe is separated from the polymorphic site of the polymorphic allele by a small number of nucleotides. This separation reduces the preferential amplification of one allele, known as allele bias. This is an improvement over methods involving the use of probes in which the 3-prime or 5-prime end of a correctly hybridized probe is directly adjacent to or very close to the polymorphic site of the allele. In one embodiment, probes whose hybridizing region likely or definitely encompasses the polymorphic site are excluded. Polymorphic sites at the hybridization site can cause uneven hybridization, and for some alleles, hybridization can be completely inhibited, resulting in preferential amplification of specific alleles. These embodiments are an improvement over other methods involving targeted amplification and / or selective enrichment in that they better preserve the original allele frequency of the sample at each polymorphic locus, regardless of whether the sample is a pure genomic sample from a single individual or a mixture of individuals.

[0394] In some embodiments, PCR (referred to as mini-PCR) is used to generate very short amplicons (see U.S. Patent Application No. 13 / 683,604, filed November 21, 2012; U.S. Patent Application Publication No. 2013 / 0123120; U.S. Patent Application No. 13 / 300,235, filed November 18, 2011; U.S. Patent Application Publication No. 2012 / 0270212, filed November 18, 2011; and U.S. Patent Application No. 61 / 994,791, filed May 16, 2014, which are incorporated herein by reference in their entirety). cfDNA (e.g., necrotically or apoptotically released cancerous cfDNA) is highly fragmented. For fetal cfDNA, fragment sizes are distributed approximately Gaussianly, with a mean of 160 bp, a standard deviation of 15 bp, a minimum size of approximately 100 bp, and a maximum size of approximately 220 bp. A polymorphic site at a particular target locus can occupy any position from the start to the end of various fragments originating from that locus. Because cfDNA fragments are short, the likelihood of both primer sites being present is the likelihood of a fragment of length L containing both the forward and reverse primer sites, which is the ratio of the amplicon length to the fragment length. Under ideal conditions, assays with amplicons of 45, 50, 55, 60, 65, or 70 bp will successfully amplify 72%, 69%, 66%, 63%, 59%, or 56% of the available template fragment molecules, respectively. In certain embodiments related to cfDNA from samples of individuals suspected of having cancer, most preferred for cfDNA, the cfDNA is amplified using primers that produce a maximum amplicon length of 85, 80, 75, or 70 bp, with a melting temperature of 50-65°C, and in particularly preferred embodiments, 54-60.5°C, resulting in a maximum amplicon length of 85, 80, 75, or 70 bp. The amplicon length is the distance between the 5-prime ends of the forward and reverse priming sites. Amplicon lengths shorter than those typically used by those skilled in the art can result in more efficient measurement of the desired polymorphic locus by requiring only short sequence reads.In one embodiment, a substantial proportion of the amplicons are less than 100 bp, less than 90 bp, less than 80 bp, less than 70 bp, less than 65 bp, less than 60 bp, less than 55 bp, less than 50 bp, or less than 45 bp.

[0395] In some embodiments, the amplification is carried out using direct multiplex PCR, sequential PCR, nested PCR, double nested PCR, one-and-a-half sided nested PCR, fully nested PCR, one-sided fully nested PCR, one-sided nested PCR, hemi-nested PCR, hemi-nested PCR, triplex hemi-nested PCR, semi-nested PCR, one-sided semi-nested PCR, reverse semi-nested PCR, or one-sided PCR. These are described in U.S. Patent Application No. 13 / 683,604, filed November 21, 2012, U.S. Patent Application Publication No. 2013 / 0123120, U.S. Patent Application No. 13 / 300,235, filed November 18, 2011, U.S. Patent Application Publication No. 2012 / 0270212, and U.S. Patent Application No. 61 / 994,791, filed May 16, 2014, which are incorporated by reference in their entirety. Any of these methods can be used for mini-PCR, if desired.

[0396] If desired, the extension step of PCR amplification can be limited in time to reduce amplification of fragments longer than 200, 300, 400, 500, or 1,000 nucleotides, which can result in enrichment of fragmented or short DNA (e.g., fetal DNA or DNA from cancer cells that have undergone apoptosis or necrosis) and improve test performance.

[0397] In some embodiments, multiplex PCR is used. In one embodiment, a method for amplifying target loci in a nucleic acid sample includes: (i) contacting the nucleic acid sample with a primer library that simultaneously hybridizes to at least 100, 200, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 20,000, 25,000, 30,000, 40,000, 50,000, 75,000, or 100,000 different target loci to generate a reaction mixture; and (ii) subjecting the reaction mixture to primer extension reaction conditions (e.g., PCR conditions) to generate amplification products containing target amplicons. In some embodiments, at least 50, 60, 70, 80, 90, 95, 96, 97, 98, 99, or 99.5% of the target loci are amplified. In various embodiments, less than 60, 50, 40, 30, 20, 10, 5, 4, 3, 2, 1, 0.5, 0.25, 0.1, or 0.05% of the amplification products are primer dimers. In some embodiments, the primers are in solution (e.g., dissolved in a liquid phase rather than a solid phase). In some embodiments, the primers are in solution and not immobilized on a solid support. In some embodiments, the primers are not part of a microarray. In some embodiments, the primers do not comprise a molecular inversion probe (MIP).

[0398] In some embodiments, two or more (e.g., three or four) target amplicons (e.g., amplicons from the mini-PCR methods disclosed herein) are ligated together, and the ligation product is then sequenced. Combining multiple amplicons into a single ligation product increases the efficiency of subsequent sequencing steps. In some embodiments, the target amplicons are shorter than 150, 100, 90, 75, or 50 base pairs before ligation. Selective enrichment and / or amplification can involve tagging each individual molecule with a different tag, molecular barcode, amplification tag, and / or sequencing tag. In some embodiments, the amplified products are analyzed by sequencing (e.g., high-throughput sequencing) or by hybridization to an array, such as, for example, a SNP array, an ILLUMINA INFINIUM array, or an AFFYMETRIX gene chip. In some embodiments, nanopore sequencing is used, such as the nanopore sequencing technology developed by Genia (see, for example, the World Wide Web at geniachip.com / technology, the contents of which are incorporated herein by reference in their entirety). In some embodiments, double-stranded sequencing is used (Schmitt et al., "Detection of ultra-rare mutations by next-generation sequencing," Proc Natl Acad Sci U S A. 109(36):14508-14513, 2012, the contents of which are incorporated herein by reference in their entirety). This method significantly reduces errors by individually tagging and sequencing each strand of a DNA duplex. Because the two strands are complementary, true mutations are found at the same position in both strands. In contrast, PCR or sequencing errors result in mutations in only one strand and can be deducted as technical errors. In some embodiments, the methods involve tagging both strands of double-stranded DNA with random but complementary double-stranded nucleotide sequences called Duplex Tags.Double-stranded tag sequences are incorporated into standard sequencing adapters by first introducing a single-stranded randomized nucleotide sequence into one adapter strand and then using DNA polymerase to extend the opposite strand, generating a complementary double-stranded tag. After ligating the tagged adapter to sheared DNA, the individually labeled strands are PCR amplified from asymmetric primer sites on the adapter tails and paired-end sequenced. In some embodiments, a sample (e.g., a DNA or RNA sample) is divided into multiple fractions, such as different wells (e.g., wells of a WaferGen SmartChip). Because some wells contain a higher proportion of molecules with mutations than the total sample, dividing the sample into different fractions (e.g., at least 5, 10, 20, 50, 75, 100, 150, 200, or 300 fractions) can increase analytical sensitivity. In some embodiments, each fraction has less than 500, 400, 200, 100, 50, 20, 10, 5, 2, or 1 DNA or RNA molecules. In some embodiments, the molecules in each fraction are sequenced separately. In some embodiments, all molecules in the same fraction are labeled with the same barcode (e.g., a random sequence or a non-human sequence) (e.g., by amplification using barcode-containing primers or barcode ligation), and different barcodes are labeled for molecules in different fractions. The barcoded molecules can be pooled and sequenced together. In some embodiments, the molecules are amplified, then pooled and sequenced, for example, by nested PCR. In some embodiments, one forward primer and two reverse primers, or two forward primers and one reverse primer, are used.

[0399] In some embodiments, the mutation (for example, SNV or CNV) present in less than 10, 5, 2, 1, 0.5, 0.1, 0.05, 0.01 or 0.005% of DNA molecules or RNA molecules in sample (for example, cfDNA or cfRNA sample) is detected (or can be detected).In some embodiments, the mutation (for example, SNV or CNV) present in less than 1,000, 500, 100, 50, 20, 10, 5, 4, 3 or 2 original DNA molecules or RNA molecules (before amplification) in sample (for example, cfDNA or cfRNA sample derived from blood sample, etc.) is detected (or can be detected).In some embodiments, the mutation (for example, SNV or CNV) present in only 1 original DNA molecule or RNA molecule (before amplification) in sample (for example, cfDNA or cfRNA sample derived from blood sample, etc.) is detected (or can be detected).

[0400] For example, if the detection limit for a mutation (e.g., a single-nucleotide variant (SNV)) is 0.1%, a mutation present at 0.01% can be detected by dividing the fraction into multiple fractions, e.g., 100 wells. Most of the wells will have no copies of the mutation. For the few wells that have a mutation, the mutation will result in a very high read percentage. In one example, there are 20,000 initial DNA copies from the target locus, and two of those copies contain the SNV of interest. If the sample is divided into 100 wells, 98 wells will have an SNV, and two wells will have an SNV at 0.5%. The DNA in each well is barcoded, amplified, pooled with DNA from other wells, and sequenced. The wells without SNVs can be used to measure the background amplification / sequencing error rate and determine whether the signal from outlier wells exceeds the background noise level.

[0401] In some embodiments, the amplification products are detected using an array, such as, for example, a microarray containing probes for one or more chromosomes of interest (e.g., chromosomes 13, 18, 21, X, Y, or any combination thereof). It is understood that commercially available SNP detection microarrays may be used, such as, for example, genotyping assays from Illumina (San Diego, CA), GoldenGate, DASL, Infinium, or CytoSNP-12, or SNP detection microarray products from Affymetrix, such as, for example, OncoScan microarrays.

[0402] In some embodiments involving sequencing, read depth is the number of sequence reads mapped to a given locus. Read depth may be normalized to the total number of reads. In some embodiments relating to the read depth of a sample, read depth is the average depth of reads to the target locus. In some embodiments relating to the read depth of a locus, read depth is the number of reads measured by a sequencer mapping to that locus. Generally, the greater the read depth of a locus, the closer the allele ratio at that locus tends to be to the allele ratio in the original DNA sample. Read depth can be expressed in a variety of different ways, including but not limited to, a percentage or ratio. Thus, for example, a highly parallel DNA sequencer such as an Illumina HISEQ may sequence, for example, 1 million clones, and sequence a locus 3,000 times, resulting in a read depth of 3,000 at that locus. The percentage of reads at that locus is 3,000 divided by 1 million total reads, or 0.3% of the total reads.

[0403] In some embodiments, allele data is obtained, where the allele data includes a quantitative measurement indicating the copy number of a specific allele at a polymorphic locus. In some embodiments, the allele data includes a quantitative measurement indicating the copy number of each observed allele at a polymorphic locus. Typically, quantitative measurements are obtained for all possible alleles at a polymorphic locus of interest. A quantitative measurement of the copy number of a specific allele at a polymorphic locus can be generated using any of the methods discussed in the previous paragraph for determining alleles at SNP or SNV loci, such as microarrays, qPCR, or DNA sequencing, including high-throughput DNA sequencing. This quantitative measurement is referred to herein as allele frequency data or observed genetic allele data. Methods that use allele data are sometimes referred to as quantitative allele methods. This method is in contrast to quantitative methods that use only quantitative data from non-polymorphic or polymorphic loci, regardless of allele specificity. When allele data is measured using high-throughput sequencing, the allele data typically includes the number of reads for each allele mapped to the locus of interest.

[0404] In some embodiments, non-allelic data are obtained, in which case the non-allelic data include quantitative measurements indicating the copy number of a particular locus. A locus can be polymorphic or non-polymorphic. In some embodiments where a locus is non-polymorphic, the non-allelic data does not include information about the relative or absolute amounts of individual alleles that may be present at the locus. Methods that use only non-allelic data (i.e., quantitative data from non-polymorphic loci, or quantitative data from polymorphic loci regardless of the allele specificity of each fragment) are called quantitative methods. Typically, quantitative measurements are obtained for all possible alleles of a polymorphic locus of interest, and a single value is associated with the measured amount of all alleles at the locus in total. Non-allelic data for a polymorphic locus can be obtained by summing the quantitative alleles for each allele of the locus. When allele data is measured using high-throughput sequencing, the non-allelic data typically include the number of reads mapped to the locus of interest. The sequence measurements may indicate the relative and / or absolute number of each allele present at a locus, and the non-allelic data includes the sum of reads mapped to that locus, regardless of allele specificity. In some embodiments, the same set of sequence measurements can be used to generate both allelic and non-allelic data. In some embodiments, the allelic data is used as part of a method to determine copy number at a chromosome of interest, and the generated non-allelic data can be used as part of a different method to determine copy number at a chromosome of interest. In some embodiments, the two methods are statistically orthogonal and are combined to make a more accurate determination of copy number at a chromosome of interest.

[0405] In some embodiments, obtaining genetic data includes (i) obtaining DNA sequence information by laboratory techniques, such as by using an automated high-throughput DNA sequencer, or (ii) obtaining information previously obtained by laboratory techniques, where the information is transmitted or obtained electronically, for example, by a computer over the internet or by electronic transfer from a sequencing device.

[0406] Additional exemplary sample preparation, amplification, and quantification methods are described in U.S. Patent Application No. 13 / 683,604, filed November 21, 2012 (U.S. Patent Application Publication No. 2013 / 0123120, and U.S. Patent Application No. 61 / 994,791, filed May 16, 2014), which are incorporated herein by reference in their entireties. These methods can be used to analyze any of the samples disclosed herein.

[0407] 10.3 Example of cell-free DNA quantification method If desired, the amount or concentration of cfDNA or cfRNA can be measured using standard methods.In some embodiments, the amount or concentration of cell-free mitochondrial DNA (cf mDNA) is determined.In some embodiments, the amount or concentration of cell-free DNA originating from nuclear DNA (cf nDNA) is determined.In some embodiments, the amount or concentration of cf mDNA and cf nDNA is determined simultaneously.

[0408] In some embodiments, qPCR is used to measure cf nDNA and / or cfm DNA (Kohler et al., "Levels of plasma circulating cell free nuclear and mitochondrial DNA as potential biomarkers for breast tumors." Mol Cancer 8:105, 2009, 8:doi:10.1186 / 1476-4598-8-105, which is incorporated by reference in its entirety). For example, one or more loci from cf nDNA (e.g., glyceraldehyde-3-phosphadehydrogenase, GAPDH) and one or more loci from cf mDNA (ATPase 8, MTATP 8) can be measured using multiplex qPCR. In some embodiments, cf nDNA and / or cf mDNA are measured using fluorescently labeled PCR (Schwarzenbach et al., "Evaluation of cell-free tumor DNA and RNA in patients with breast cancer and benign breast disease." Mol Biosys 7:2848-2854, 2011, which is incorporated by reference in its entirety). If desired, normal distribution of the data can be d...

Claims

1. 1. A method for determining the origin of circulating cells suspected of fetal origin, comprising: a) generating a first set of amplicons from cellular DNA isolated from a pediatric cell line or one or more circulating cells suspected to be of fetal origin obtained from a blood sample of a mother carrying a fetus, and a second set of amplicons from a DNA mixture of the child and the mother or cell-free DNA obtained from the plasma fraction of said blood sample, wherein said first set of amplicons and said second set of amplicons are obtained by performing a multiplexed amplification reaction of a plurality of single nucleotide polymorphism (SNP) loci; b) sequencing the first set of amplicons and the second set of amplicons by next generation sequencing; and c) determining the origin of the one or more circulating cells based on the sequences of the first amplicon set, wherein the sequences of the second amplicon set are used as a reference.

2. 10. The method of claim 1, further comprising generating a third set of amplicons from maternal cell DNA isolated from a maternal cell line or from one or more maternal cells obtained from a buffy coat fraction of the blood sample, and genotyping the third set of amplicons by next-generation sequencing to determine the maternal genotype.

3. 3. The method of claim 2, further comprising obtaining from the blood sample the circulating cells suspected to be of fetal origin, the plasma fraction containing the cell-free DNA, and the buffy coat fraction containing the maternal cells.

4. 2. The method of claim 1, wherein step (c) comprises calculating maternal and fetal matches between the mixture of child and mother's DNA or the cell-free DNA obtained from the plasma fraction and the cellular DNA obtained from the circulating cells suspected to be of fetal origin, and calculating a fetal admixture percentage.

5. 2. The method of claim 1, wherein the first set of amplicons and the second set of amplicons are obtained by performing a multiplex amplification reaction of at least 1,000 SNP loci.

6. 2. The method of claim 1, further comprising estimating the fetal fraction of the mixture of maternal and fetal cells by measuring the amount of the SNP locus.

7. 7. The method of claim 6, wherein the method further comprises determining an allele ratio at the SNP.

8. The method of claim 7, wherein the estimation of the fetal proportion uses the allele ratio.

9. 10. The method of claim 1, wherein the method comprises dividing a plurality of circulating cells or cell lines into individual reaction volumes, isolating cellular DNA in each reaction volume, and attaching a sample barcode to the isolated cellular DNA in each reaction volume.

10. 10. The method of claim 1, wherein the cellular DNA is isolated from a single circulating cell suspected to be a fetal cell.

11. 10. The method of claim 1, wherein the method further comprises performing a non-invasive prenatal test using a genotype determined to be one or more fetal cells.

12. 12. The method of claim 11, wherein the method further comprises detecting copy number variations or aneuploidy of a target chromosome or target chromosome segment of interest in the one or more circulating cells determined to be one or more fetal cells.

13. 13. The method of claim 12, wherein the target chromosome or target chromosome segment of interest is chromosome 13, chromosome 18, chromosome 21, a sex chromosome, and / or a chromosome segment thereof.

14. 12. The method of claim 11, wherein the method further comprises detecting a microdeletion in the one or more circulating cells determined to be one or more pure fetal cells.

15. 15. The method of claim 14, wherein the microdeletion is a 22q11.2 deletion associated with DiGeorge syndrome, a microdeletion associated with Prader-Willi syndrome, a microdeletion associated with Angelman syndrome, a 1p36 deletion, and / or a microdeletion associated with Cri-Cat syndrome.

16. 1. A method for monitoring and detecting early recurrence or metastasis of tumors in cancer patients, said method comprising: a) selecting one or more patient-specific mutations based on somatic mutations identified in tumor samples of patients diagnosed with cancer; b) longitudinally collecting one or more blood samples from said patient after said patient has been treated with surgery, first-line chemotherapy, and / or adjuvant therapy; c) generating an amplicon set from cellular DNA isolated from one or more circulating cells suspected to be circulating tumor cells and obtained from a blood sample of said patient, said amplicon set being generated by multiplexed amplification of genomic loci encompassing said patient-specific mutations associated with cancer; d) sequencing the amplicon set by next-generation sequencing and determining the origin of the one or more circulating cells based on the presence of the one or more patient-specific mutations in the amplicon set, wherein detection of one or more circulating tumor cells containing the one or more patient-specific mutations indicates early recurrence or metastasis of the cancer.

17. 17. The method of claim 16, wherein the patient-specific mutation comprises a cancer-associated single nucleotide variant (SNV), copy number variant (CNV), indel, and / or gene fusion.

18. 17. The method of claim 16, wherein the amplicon set is generated by multiplexed amplification of genomic loci encompassing at least eight patient-specific mutations associated with cancer.

19. 17. The method of claim 16, wherein the amplicon set is generated by multiplexed amplification of genomic loci encompassing at least 16 patient-specific mutations associated with cancer.

20. 17. The method of claim 16, wherein the presence of at least two patient-specific mutations associated with cancer indicates that the one or more circulating cells are tumor cells.

21. 17. The method of claim 16, wherein the presence of at least five patient-specific mutations associated with cancer indicates that the one or more circulating cells are tumor cells.

22. 17. The method of claim 16, wherein detection of at least two circulating tumor cells containing the one or more patient-specific mutations indicates early recurrence or metastasis of the cancer.

23. 17. The method of claim 16, wherein the method comprises dividing a plurality of circulating cells into individual reaction volumes, isolating cellular DNA in each reaction volume, and attaching a sample barcode to the isolated cellular DNA in each reaction volume.

24. 17. The method of claim 16, wherein the cellular DNA is isolated from a single circulating cell suspected to be a tumor cell.

25. 17. The method of claim 16, further comprising identifying additional patient-specific mutations associated with cancer from the genotype of one or more circulating cells determined to be tumor cells.

26. 17. The method of claim 16, wherein the patient is suffering from lung cancer.

27. 17. The method of claim 16, wherein the patient is suffering from breast cancer.

28. 17. The method of claim 16, wherein the patient is suffering from bladder cancer.

29. 17. The method of claim 16, wherein the patient is suffering from colorectal cancer.

30. 1. A method for determining the origin of circulating cells suspected to be donor cells, comprising: a) generating a first set of amplicons from cellular DNA isolated from one or more circulating cells suspected to be donor cells obtained from a blood sample of a transplant recipient, and a second set of amplicons from cell-free DNA obtained from the plasma fraction of said blood sample, wherein said first set of amplicons and said second set of amplicons are obtained by performing a multiplexed amplification reaction of a plurality of single nucleotide polymorphism (SNP) loci; b) sequencing the first set of amplicons and the second set of amplicons by next generation sequencing; and c) determining the origin of the one or more circulating cells based on the sequences of the first amplicon set, wherein the sequences of the second amplicon set are used as a reference.

31. 1. A method for monitoring and detecting early recurrence or metastasis of tumors in cancer patients, said method comprising: a) selecting one or more patient-specific mutations based on somatic mutations identified in tumor samples of patients diagnosed with cancer; b) longitudinally collecting one or more blood samples from said patient after said patient has been treated with surgery, first-line chemotherapy, and / or adjuvant therapy; c) generating a first set of amplicons from cell-free DNA isolated from a blood sample of said patient, said first set of amplicons being generated by multiplexed amplification of genomic loci encompassing said patient-specific mutations associated with cancer; d) sequencing the first amplicon set by next generation sequencing and detecting the presence of one or more patient-specific mutations in the amplicon set, wherein the presence of the one or more patient-specific mutations indicates early recurrence or metastasis of the cancer; e) generating a second set of amplicons from cellular DNA isolated from one or more circulating cells obtained from said patient's blood sample, said second set of amplicons being generated by multiplexed amplification of genomic loci encompassing said patient-specific mutations associated with cancer; f) sequencing the second amplicon set by next generation sequencing and determining the origin of the one or more circulating cells based on the presence of one or more patient-specific mutations in the amplicon set, wherein the presence of the one or more patient-specific mutations indicates that the one or more circulating cells are tumor cells; and g) identifying one or more additional mutations associated with cancer from the sequence of said one or more circulating cells determined to be tumor cells.

32. 32. The method of claim 31 , wherein the patient-specific mutation comprises a cancer-associated single nucleotide variant (SNV), copy number variant (CNV), indel, and / or gene fusion.

33. 32. The method of claim 31 , wherein the first and / or second amplicon sets are generated by multiplexed amplification of genomic loci encompassing at least eight patient-specific mutations associated with cancer.

34. 32. The method of claim 31 , wherein the first and / or second amplicon sets are generated by multiplexed amplification of genomic loci encompassing at least 16 patient-specific mutations associated with cancer.

35. 32. The method of claim 31 , wherein in step (d), the presence of at least two or at least five patient-specific mutations associated with cancer indicates early recurrence or metastasis of the cancer.

36. 32. The method of claim 31 , wherein in step (f), the presence of at least two or at least five patient-specific mutations associated with cancer indicates that the one or more circulating cells are tumor cells.

37. 32. The method of claim 31 , wherein the method comprises dividing a plurality of circulating cells into individual reaction volumes, isolating cellular DNA in each reaction volume, and attaching a sample barcode to the isolated cellular DNA in each reaction volume.

38. 32. The method of claim 31, wherein the cellular DNA is isolated from a single circulating cell suspected to be a tumor cell.

39. 32. The method of claim 31, wherein the patient is suffering from lung cancer, breast cancer, bladder cancer, or colorectal cancer.

40. 32. The method of claim 31, further comprising treating the patient based on the one or more additional mutations associated with cancer.