Analysis method for circulating cells

The method addresses the challenge of analyzing small numbers of circulating cells by using multiplex amplification and sequencing, enabling accurate identification and characterization for non-invasive diagnostics and monitoring.

JP7713054B2Active Publication Date: 2025-07-24NATERA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024050560
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-01-28
Filing Date
2024-03-26
Publication Date
2025-07-24
Estimated Expiration
2039-12-16

AI Technical Summary

Technical Problem

Current liquid biopsy methods face challenges in accurately analyzing and characterizing small numbers of circulating cells, such as circulating fetal cells and circulating tumor cells, due to the difficulty in isolating and analyzing these cells, which are highly diluted in maternal or patient blood samples, limiting the effectiveness of non-invasive diagnostic methods.

Method used

A method involving multiplex amplification of single nucleotide polymorphism (SNP) loci for circulating cells, followed by next-generation sequencing, to determine the origin and purity of isolated cells, enabling accurate genomic analysis even from a single cell.

Benefits of technology

Enables accurate identification and characterization of circulating cells, allowing for non-invasive prenatal diagnosis, cancer monitoring, and transplanted organ monitoring by confirming the origin and purity of isolated cells, even when only a few cells are present.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007713054000054
    Figure 0007713054000054
  • Figure 0007713054000055
    Figure 0007713054000055
  • Figure 0007713054000056
    Figure 0007713054000056
Patent Text Reader

Abstract

To provide methods for characterizing and analyzing circulating cells, in particular, to provide methods for confirming the identity of an individual cell and confirming that the obtained sample is derived from a single cell of a defined identity.SOLUTION: Additional methods are provided for analyzing a single cell sample to determine copy number variation and aneuploidy in the context of circulating fetal cells, microdeletions, single nucleic acid variations associated with cancer, or early relapse of cancer and metastasis.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 62 / 780,724, filed on Dec. 17, 2018, and U.S. Provisional Patent Application No. 62 / 797,518, filed on Jan. 28, 2019, which are hereby incorporated by reference in their entirety.

[0002] Currently, there is great interest in using liquid biopsy as an alternative to invasive tissue biopsy methods for prenatal diagnosis, diagnosis of diseases and conditions, and examination and monitoring of disease development and treatment response. For example, liquid biopsy of a pregnant mother provides a non - invasive method for fetal diagnosis and examination, avoiding the risks associated with performing invasive biopsies on the fetus or amniotic fluid. Even when diagnosing and monitoring cancer development and treatment response, liquid biopsy can avoid performing invasive tumor tissue biopsies that carry risks that may cause metastasis or surgical complications. Also, methods based on liquid biopsy may be more sensitive than imaging - based detection methods that are insufficiently sensitive to detect early - stage recurrence or metastasis.

[0003] Liquid biopsy is based on collecting a blood or urine sample from a patient and isolating cell - free DNA and circulating cells shed from a fetus, tumor, or other target organ. By analyzing cell - free DNA and circulating cells, a powerful and non - invasive diagnostic method has been provided. However, since cell - free DNA is fragmented and does not cover the entire genome, less information can be obtained from the analysis of cell - free DNA compared to directly analyzing whole - genome DNA from cells and / or tissues. On the other hand, the analysis of circulating cells has problems due to the difficulty of characterizing and analyzing very small amounts of cell samples. Therefore, in order to maximize the utilization of methods based on liquid biopsy, methods for isolating and analyzing samples containing only a small number of cells or only one cell are required.

Summary of the Invention

[0004] Disclosed herein is a method for analyzing and characterizing a circulating cell sample that may contain only a small number of circulating cells or only one circulating cell. According to an embodiment illustrated herein, in one embodiment, a method for analyzing and characterizing circulating cells that are presumed to be of fetal origin, derived from a blood sample of a mother carrying a fetus or from a cell line, the method comprising: a) a first amplicon set of cell DNA isolated from one or more circulating cells presumed to be of fetal origin obtained from the blood sample or the cell line, and a second amplicon set of cell-free DNA obtained from the plasma fraction of the blood sample or mixed DNA of a child cell line and a mother cell line, wherein the first amplicon set and the second amplicon set are obtained by performing a multiplex amplification reaction of a plurality of single nucleotide polymorphism (SNP) loci; b) genotyping the first amplicon set and the second amplicon set by next-generation sequencing; and c) determining, based on the genotypes of the first amplicon set and the second amplicon set, whether the one or more circulating cells are one or more maternal cells, one or more fetal cells, or a mixture of maternal and fetal cells.

[0005] In one embodiment, the method disclosed herein further comprises generating a third amplicon set of maternal cell DNA isolated from one or more maternal cells obtained from the buffy coat fraction of the blood sample, and genotyping the third amplicon set by next-generation sequencing to determine the genotype of the mother.

[0006] In one embodiment, the method disclosed herein further comprises obtaining from the blood sample a circulating cell presumed to be of fetal origin, a plasma fraction containing cell-free DNA, and a buffy coat fraction containing maternal cells.

[0007] In one embodiment, determining that the one or more circulating cells are one or more maternal cells, one or more fetal cells, or a mixture of maternal and fetal cells involves calculating maternal and fetal concordances between cell-free DNA obtained from the plasma fraction (or cell-free DNA known to be a mixture of the child's and mother's DNA) and cellular DNA obtained from circulating cells (or the child's cell line) that are suspected to be of fetal origin, and calculating the fetal mixture ratio (or child ratio).

[0008] In one embodiment, the first amplicon set and the second amplicon set are obtained by performing a multiplex amplification reaction of at least 1,000 SNP loci.

[0009] In one embodiment, the method disclosed herein further includes estimating the fetal fraction in a mixture of maternal and fetal cells (or a mixture of the mother's and child's cell lines) by measuring the amount of SNP loci.

[0010] In one embodiment, the method disclosed herein further includes determining the allelic ratio at the SNP.

[0011] In one embodiment, estimating the fetal fraction uses the allelic ratio.

[0012] In one embodiment, the method disclosed herein includes dividing a plurality of circulating cells into individual reaction volumes, isolating the cellular DNA in each reaction volume, and attaching a sample barcode to the cellular DNA isolated in each reaction volume.

[0013] In one embodiment, the cellular DNA is isolated from one circulating cell suspected to be a fetal cell.

[0014] In one embodiment, the method further includes performing a non-invasive prenatal test using the genotype determined that one or more circulating cells are one or more fetal cells.

[0015] In one embodiment, the method further includes detecting a copy number variation or aneuploidy of a target chromosome or a target chromosomal segment in one or more circulating cells determined to be one or more fetal cells.

[0016] In one embodiment, the target chromosome or the target chromosomal segment is chromosome 13, chromosome 18, chromosome 21, a sex chromosome, and / or a chromosomal segment thereof.

[0017] In one embodiment, the method disclosed herein further includes detecting a microdeletion in one or more circulating cells determined to be one or more pure fetal cells.

[0018] In one embodiment, the microdeletion is a 22q11.2 deletion associated with DiGeorge syndrome, a microdeletion associated with Prader-Willi syndrome, a microdeletion associated with Angelman syndrome, a 1p36 deletion, and / or a microdeletion associated with cri du chat syndrome.

[0019] In another exemplary aspect, in one aspect, the present disclosure provides a method for monitoring and detecting early recurrence or metastasis of a tumor in a cancer patient by obtaining one or more circulating cells determined to be derived from the tumor, the method comprising: a) selecting one or more patient-specific mutations based on mutations identified in a tumor sample of a patient diagnosed with cancer; b) longitudinally collecting one or more blood samples from the patient after the patient has been treated with surgery, first-line chemotherapy, and / or adjuvant therapy; c) A step of generating a first amplicon set from cell DNA isolated from one or more circulating cells presumed to be circulating tumor cells obtained from the blood sample obtained from the cancer patient, and a second amplicon set from cell-free DNA, wherein the tumor cells and normal cells are obtained from each of the blood sample of the cancer patient or a fraction thereof, and the first amplicon set and the second amplicon set are obtained by performing a multiplex amplification reaction of patient-specific mutations associated with cancer. d) A step of sequencing the first amplicon set and the second amplicon set by next-generation sequencing. e) A step of determining the origin of one or more circulating cells based on the sequence of the first amplicon set, wherein the sequence of the second amplicon set is used as a reference, and the detection of one or more patient-specific mutations from the first amplicon set generated from the circulating cells determined to be of tumor origin suggests early recurrence or metastasis of cancer. The method includes the above steps.

[0020] In another aspect, the disclosure herein provides a method for treating a cancer patient, the method comprising: a) Treating the cancer patient with surgery, first-line chemotherapy, and / or adjuvant therapy. b) After the patient is treated with surgery, first-line chemotherapy, and / or adjuvant therapy, long-term collecting one or more blood samples from the patient. c) A step of generating a first amplicon set from cell DNA isolated from one or more circulating cells presumed to be circulating tumor cells obtained from the blood sample obtained from the cancer patient, and a second amplicon set from cell-free DNA, wherein the tumor cells and normal cells are obtained from each of the blood sample of the cancer patient or a fraction thereof, and the first amplicon set and the second amplicon set are obtained by performing a multiplex amplification reaction of patient-specific mutations associated with cancer. d) sequencing the first amplicon set and the second amplicon set by next-generation sequencing; e) determining the origin of one or more circulating cells based on the sequence of the first amplicon set, wherein the sequence of the second amplicon set is used as a reference, and detecting one or more patient-specific mutations from the first amplicon set generated from the circulating cells determined to be of tumor origin, which suggests early recurrence or metastasis of cancer; f) administering a compound to the patient, wherein the compound has been found to be effective in treating cancer having the one or more patient-specific mutations detected from the blood sample. A method comprising the steps.

[0021] The improved method provided herein enables accurate analysis and characterization of samples containing very few CTCs. The number of CTCs in a sample analyzed by the method of the present invention can be less than 100 cells, less than 75 cells, less than 50 cells, less than 25 cells, less than 20 cells, less than 15 cells, less than 10 cells, less than 5 cells, or one cell. By using the method provided herein, accurate genomic information can be obtained from one circulating tumor cell, 2 CTCs, 3 CTCs, 4 CTCs, 5 CTCs, 6 CTCs, 7 CTCs, 8 CTCs, 9 CTCs, or 10 CTCs.

[0022] In some embodiments, the method includes monitoring and detecting early recurrence or metastasis of a tumor in a cancer patient by determining the identity of circulating cells that are presumed to be circulating tumor cells based on pre-determined patient-specific mutations. In some embodiments, the patient-specific mutations include single nucleotide variants (SNVs) associated with cancer, copy number variants (CNVs), indels, deletions, or gene fusions. In some embodiments, the presence of at least two patient-specific mutations associated with cancer suggests that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least eight patient-specific mutations associated with cancer suggests that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least sixteen patient-specific mutations associated with cancer suggests that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least fifty patient-specific mutations associated with cancer suggests that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least one hundred patient-specific mutations associated with cancer suggests that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least one hundred to about one thousand patient-specific mutations associated with cancer suggests that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least one hundred to about one thousand patient-specific mutations associated with cancer suggests that the one or more circulating cells are tumor cells.

[0023] In one aspect, multiplex amplification is performed to obtain amplicons that include patient-specific mutations associated with cancer. In some embodiments, a first amplicon set and a second amplicon set are obtained by performing a multiplex amplification reaction of 50 to 1000 cancer-associated patient-specific mutations.

[0024] In one aspect, a method for monitoring and detecting early recurrence and metastasis of tumors in cancer patients includes dividing a plurality of circulating cells into individual reaction volumes, isolating cellular DNA in each reaction volume, and attaching a sample barcode to the cellular DNA isolated in each reaction volume. In some embodiments, the cellular DNA is isolated from a single circulating cell suspected of being a tumor cell.

[0025] In some embodiments, the method further includes performing a non-invasive cancer test using a genotype determined that one or more circulating cells are one or more tumor cells.

[0026] In another illustrative aspect, in one embodiment, a method for analyzing and characterizing circulating cells suspected of being donor cells, derived from a blood sample of a transplant recipient, the method comprising: a) generating a first amplicon set of cellular DNA isolated from one or more circulating cells suspected of being donor cells obtained from the blood sample and a second amplicon set of cell-free DNA obtained from the plasma fraction of the blood sample, wherein the first amplicon set and the second amplicon set are obtained by performing a multiplex amplification reaction of a plurality of single nucleotide polymorphism (SNP) loci; b) genotyping the first amplicon set and the second amplicon set by next-generation sequencing; c) determining, based on the genotyping of the first amplicon set and the second amplicon set, that the one or more circulating cells are one or more donor cells, one or more recipient cells, or a mixture of donor and recipient cells. BRIEF DESCRIPTION OF THE DRAWINGS

[0027]

Figure 1

[0028]

Figure 2

[0029]

Figure 3

[0030]

Figure 4-1

Figure 4-2

[0031]

Figure 5

[0032] The embodiments disclosed herein will be further described with reference to the accompanying drawings, and the same structures are represented by the same reference numerals throughout the plurality of drawings. The drawings shown are not necessarily to scale and are shown emphasized instead of in a normal arrangement for the explanation of the principles of the embodiments disclosed herein.

[0033] The figures identified above are provided for presentation and not for limitation.

Mode for Carrying Out the Invention

[0034] 1. Overview The present disclosure provides methods and compositions for the characterization and genomic analysis of circulating cells. In particular, the present disclosure provides a method for confirming the individuality of a single isolated circulating cell and a method for analyzing genomic data from one cell, which are methods without precedent in the literature.

[0035] The isolation and genomic analysis of one cell have attracted the interest of many clinical and corporate researchers. Single-cell genomics provides genetic information at the cell level. Genotyping profiles and genomic information obtained from a single isolated cell hold the potential for many applications in clinical diagnosis, risk prediction, drug development, and disease treatment. However, the development of methods for obtaining genomic information from one cell has been hindered by the great challenge of analyzing data from a sample of one cell as well as a sample containing a small number of cells. The disclosure herein provides both biochemical and analytical methods for accurately obtaining and analyzing genomic information from a single-cell sample.

[0036] A method of accurately obtaining and analyzing genomic information from a single cell, in particular, will have a great impact on diagnosis based on liquid biopsy. Liquid biopsy refers to the collection of samples of non-solid biological tissues, most commonly blood, but also saliva, urine, cerebrospinal fluid, and other body fluids. For example, by using liquid biopsy of a pregnant mother to obtain a sample containing both cell-free DNA (cfDNA) of fetal origin and fetal circulating cells, non-invasive diagnosis of the fetus becomes possible. Intact circulating fetal cells (CFCs) can enter the maternal blood circulation. Therefore, by analyzing the genetic material of the fetus, it is possible to provide non-invasive prenatal genetic diagnosis (NIPGD or NPD). However, CFCs are highly diluted among billions of maternal blood cells. Another example is the use of liquid biopsy in cancer patients to diagnose or monitor tumors without performing a biopsy of the tumor itself. Liquid biopsy of cancer patients allows the isolation of circulating tumor cells (CTCs) and circulating tumor DNA (ctDNA), enabling the diagnosis and monitoring of tumors and the examination of the tumor's responsiveness to treatment. CTCs exfoliate from primary and metastatic lesions into the bloodstream of cancer patients, but even in patients with progressive metastatic disease, CTCs are relatively rare among the patient's blood cells. Therefore, accurate diagnosis of tumors based on liquid biopsy depends on the accurate isolation and analysis of cell samples containing only a small number of cells or a single cell. Additionally, liquid biopsy can be used to isolate and analyze circulating cells of transplanted organ origin to monitor the transplanted organ.

[0037] Therefore, the maximum utilization of liquid biopsy-based diagnosis depends on the ability of a diagnostic method to accurately isolate and analyze CRCs such as circulating fetal cells (CFCs) and circulating tumor cells (CTCs). For example, when diagnosing a fetus in a pregnant mother using CFCs, the analysis of a single circulating cell derived from the pregnant mother's blood sample depends on confirming that the isolated single circulating cell is of fetal origin and providing useful information. Similarly, the analysis of a single circulating cell from a cancer patient depends on confirming that the isolated single circulating cell is of tumor origin and providing useful information. In addition, diagnosis based on circulating rare cells (CRCs) also depends on the ability to analyze samples containing only a small number of cells or only a single cell.

[0038] The present disclosure meets this need by providing a method that can accurately analyze isolated single circulating cells and characterize them qualitatively and quantitatively. The disclosure herein provides a next-generation sequencing (NGS) method based on SNPs for analyzing an isolated single circulating cell sample, which method can identify the cell type, estimate the purity of its DNA, and measure the quality of the sample. For example, using the method of the present disclosure, it may be confirmed that an isolated cell presumed to be a fetal cell is a pure single isolated fetal cell, a pure maternal cell, or a mixture of fetal and maternal cells. In another embodiment, using the method of the present disclosure, it may be confirmed whether an isolated cell sample presumed to be a circulating tumor cell is an isolated single circulating tumor cell or a normal cell derived from the patient's blood.

[0039] FIG. 1 is an overview of the workflow of an exemplary embodiment of the present disclosure for the analysis of circulating fetal cells isolated from a maternal blood sample. Accordingly, the workflow outlined in FIG. 1 may be adapted to enable the analysis of circulating tumor cells or any other circulating rare cells isolated from a subject. In one exemplary embodiment, the method of the present disclosure outlined in FIG. 1 and described in further detail below 1. A step of collecting a blood sample from a mother who is pregnant with a fetus, 2. A step of isolating a plasma fraction and / or a buffy coat fraction from the blood sample, 3. A step of isolating fetal cells and maternal cells from the blood sample, 4. A step of extracting cell-free DNA (cfDNA) from the plasma fraction, and optionally, (b) genomic DNA from the buffy coat fraction, 5. A step of lysing the isolated (a) fetal cells and (b) maternal cells, 6. A step of preparing a cfDNA library for sequencing, 7. A step of performing massively-multiplexed PCR (mmPCR) on (a) the cell lysate directly, (b) the maternal DNA, and (c) the cfDNA library, 8. A step of adding a DNA barcode to the amplified DNA sample and labeling each sample, 9. A step of preparing a DNA sample pool for sequencing, 10. A step of running the sample via a next-generation sequencing (NGS) protocol, 11. A step of performing (a) preprocessing of the NGS data and (b) analysis of the NGS data, and 12. A step of interpreting the results of the processing and analysis of the NGS data.

[0040] As a result of the above method, information regarding the matching between cfDNA, fetal DNA (cell lysate), and maternal DNA can be provided, and the purity determination of the cell sample may be made possible. As shown in FIGS. 2 and the Examples of the present specification, using the disclosed method, it can be confirmed that the analyzed DNA sample originated from one circulating cell. The Examples provided herein show that, following the workflows outlined in FIGS. 1 and 2 and using cfDNA as a reference, the method can determine the origin of cell DNA from an unknown sample and the composition of the sample. For example, in the case of fetal / maternal cells isolated from the blood of a pregnant woman, the method can determine whether the sample is a pure fetal cell or a pure maternal cell (FIGS. 2 and the Examples).

[0041] Furthermore, chromosomal abnormalities in circulating cells can be revealed from the data. When the sample is determined to be a mixture of maternal and fetal cells, the fetal proportion can be determined. In another aspect, the method provided herein may be used to determine the purity of a circulating tumor cell sample and analyze the genomic information of circulating tumor cells. In some embodiments, the method provided herein enables the determination that circulating cells presumed to be derived from a transplanted organ are pure isolated cells originating from the transplanted organ.

[0042] 2. A method for analyzing one cell to confirm its identity The disclosure of this specification provides a data analysis method capable of analyzing genomic information even from a sample of a very small number of cells, or even a single cell. In particular, the present disclosure provides a method capable of determining the identity of a single circulating cell isolated from a blood sample and analyzing the genomic information of the single circulating cell. For example, as further described below and in the examples of this specification, the present disclosure shows how to confirm the identity of circulating fetal cells derived from a blood sample of a pregnant mother and then how to analyze the genomic information of the confirmed pure fetal cell sample. The methods described below can be adapted to confirm the identity of any cell sample by using the corresponding cell-free DNA as a reference.

[0043] According to the aspects illustrated in this specification, in one embodiment, a method for analyzing and characterizing circulating cells presumed to be of fetal origin from a blood sample of a mother carrying a fetus includes: a) generating a first amplicon set of cell DNA isolated from one or more cell lines or one or more circulating cells presumed to be of fetal origin obtained from the blood sample, and a second amplicon set of cell-free DNA obtained from the plasma fraction of the blood sample or a DNA mixture obtained from the DNA of the child and the mother, wherein the first amplicon set and the second amplicon set are obtained by performing a multiplex amplification reaction on a plurality of single nucleotide polymorphism (SNP) loci; b) genotyping the first amplicon set and the second amplicon set by next-generation sequencing; and c) determining, based on the genotyping of the first amplicon set and the second amplicon set, whether the one or more circulating cells are one or more maternal cells, one or more fetal cells, or a mixture of maternal and fetal cells. In one embodiment, the method disclosed herein further includes generating a third amplicon set of maternal cell DNA isolated from one or more maternal cells obtained from the buffy coat fraction of the blood sample or isolated from a maternal cell line, and genotyping the third amplicon set by next-generation sequencing to determine the genotype of the mother.

[0044] In one embodiment, determining whether the one or more circulating cells are one or more maternal cells, one or more fetal cells, or a mixture of maternal and fetal cells includes calculating maternal concordance and fetal concordance between the cell-free DNA obtained from the plasma fraction or the DNA mixture of the child and the mother and the cell DNA obtained from the circulating cells presumed to be of fetal origin, and calculating the fetal mixture ratio.

[0045] The first step for calculating all of the following metrics is to analyze the allele ratio (AR) of the cfDNA sample and classify them as maternal genotype / fetal genotype = RR / RR, RR / RM, RM / *, MM / MM, or MM / MR. The allele ratio is calculated by the following method: reference count for a given SNP / total count. In some embodiments, the basic genotyper for cfDNA classifies SNPs by AR. The algorithm for this calculation is adjustable for various applications. In an exemplary but non-limiting embodiment, the basic genotyper for cfDNA classifies SNPs by AR in the following manner:

Number

[0046] Alternatively, in some embodiments, the genotype of cfDNA can be classified by using the PANORAMA® genotyper.

[0047] As a second step, SNPs in the unknown sample are classified as homozygous-R (HoR), homozygous-M (HoM), or heterozygous (Het) based on the allele ratio (AR). The algorithm for this calculation is adjustable for various applications. In an exemplary but non-limiting embodiment, SNPs in the unknown sample are classified as homozygous-R (HoR), homozygous-M (HoM), or heterozygous (Het) based on the allele ratio (AR) in the following manner:

Number

[0048] As a third step, by using the genotyping / classification of these SNPs, the following metrics can be calculated and used to classify the unknown sample:

[0049] Maternal concordance = the proportion of homozygous SNPs (RR / RR or MM / MM) in a cfDNA sample that has allele ratios (HoR and HoM respectively) within the predicted range in an unknown sample.

Number

[0050] When maternal concordance is high (above about 85%), the unknown sample is likely to be from the same mother as the cfDNA sample.

[0051] Fetal concordance = the proportion of fetal band SNPs (RR / RM or MM / MR) in a cfDNA sample that has an allele ratio (Het) within the predicted range in an unknown sample.

Number

[0052] When fetal concordance is high (above about 85%), the unknown sample is likely to contain at least some fetal DNA.

[0053] Fetal mixture ratio = an estimate of the contribution by fetal DNA in an unknown sample. This is obtained by subtracting from 1 the separation of RR / RM and MM / MR of fetal bands in the unknown sample. When the sample is pure maternal cells, the separation is close to 1 and the fetal mixture ratio is 0. When the sample is pure fetal cells, the separation is close to 0 and the fetal ratio is close to 1. The mixture is a numerical value between these two values.

Number

[0054] In one illustrative embodiment, an unknown sample is classified according to Table 1 below by using the calculated maternal match, fetal match, and fetal mixture ratio. According to the illustrative embodiment shown in Table 1, a pure fetal sample is defined as having a fetal mixture ratio greater than 95%, as well as a fetal match and maternal match greater than 85%. In some embodiments, a sample is a pure fetal sample when the fetal mixture ratio exceeds 75%, 80%, 85%, 90%, 95%, or 98%, and the maternal match and fetal match exceed 75%, 80%, 85%, 90%, or 95%. In one particular embodiment, a sample is a pure fetal sample when the fetal mixture ratio exceeds 95% and the fetal match exceeds 85%. A sample can be a pure maternal sample when the maternal match exceeds 75%, 80%, 85%, 90%, or 95% and the fetal mixture ratio is less than 15%, 10%, 5%, or 2%. In one particular embodiment, when the maternal match exceeds 85% and the fetal mixture ratio is less than 5%, the sample is a pure maternal cell. Note that certain combinations (e.g., 100% fetal match and 0% fetal mixture ratio) are impossible or highly unlikely (e.g., high fetal match and low maternal match) and are therefore excluded from the classification table.

[0055]

Table 1

[0056] This method has been verified in the examples herein for the Pano® V2 target set with respect to chromosomes 13, 18, and 21 and is expected to be useful for any SNP target set (provided it is of sufficient size and minor allele frequency). For example, the SPECTRUM® target set can be a genotyping device useful for the methods of the present disclosure.

[0057] 3. Isolation of Circulating Cells Circulating fetal cells (CFCs), circulating tumor cells (CTCs), and circulating cells from transplanted organs are relatively rare in liquid biopsy samples. For example, the concentration of fetal cells in maternal blood depends on the stage of pregnancy and the condition of the fetus, but as an estimated value, it ranges from 1 to 40 fetal cells per milliliter of maternal blood, or less than 1 fetal cell per 100,000 maternal nucleated cells. Even with current technology, it is possible to isolate a small amount of fetal cells from the mother's blood, but it is difficult to concentrate the fetal cells to purity in any amount.

[0058] Numerous methods have been developed for concentrating rare circulating cells, and these methods are based on the property of identifying circulating tumor cells from surrounding normal cells for circulating tumor cells, or the property of identifying circulating fetal cells from maternal cells for circulating fetal cells.

[0059] Such identifying properties may be physical properties such as size, density, charge, and deformability, and concentration based on the use of methods such as filtration through a special filter, microscopy combined with micropipetting, microfluidic systems, centrifugation, electrophoresis, or flow cytometry is possible.

[0060] Alternatively, the identifying characteristic may be a biological characteristic such as, for example, the expression of a cell surface protein or viability, and it is possible to label circulating cells and distinguish them from normal cells or progenitor cells. For example, an antibody may be used to detect a specific cell surface marker expressed by circulating cells rather than surrounding cells, and then the circulating cells may be isolated by placing them on a solid support coated with a flow cytometry, magnetic beads, or an agent that binds an antibody and isolating a cell together with the bound antibody. For example, circulating tumor cells have been isolated using an antibody against epithelial cell adhesion molecule (EpCAM). See Alix-Panabieres and Pantel, Clin. Chem., 59(1):110-118(2013). Alternatively, circulating cells may be enriched by negative selection using an antibody that recognizes an antigen that is on normal cells but not on circulating cells, and then removing the antibody-labeled cells. Another negative selection method is the selective lysis of red blood cells.

[0061] 4. Patient Population As used herein, the terms "subject", "patient", or "individual" refer to any subject, patient, or individual, and the terms are used interchangeably herein. In this regard, "subject", "patient", and "individual" include mammals, and particularly humans. When used in conjunction with "in need thereof", the terms "subject", "patient" or "individual" are intended to refer to any subject, patient, or individual having or at risk of having a particular condition or disorder.

[0062] Using the methods and compositions of the present disclosure, various subjects can be diagnosed, monitored, or treated, including humans and non-human animals, including mammals, as well as immature and mature animals, including human children and adults suffering from various conditions or disorders. The human subject to be diagnosed or treated can be an embryo, fetus, infant, child, school-aged child, adolescent, young adult, adult, or elderly patient.

[0063] In some embodiments, the subject to be diagnosed or monitored by isolating and analyzing circulating fetal cells obtained from a maternal blood sample is a fetus. In some embodiments, the methods disclosed herein may be used for paternity testing of a fetus. In some embodiments, the methods disclosed herein may be used to determine copy number variations in a target chromosome or chromosomal segment in a fetus of a pregnant mother. In some embodiments, the methods disclosed herein may be used to determine copy number variations in chromosome 13, chromosome 18, chromosome 21, sex chromosomes, and / or segments of those chromosomes. In some embodiments, the methods disclosed herein may be used to detect microdeletions in a fetus's chromosomes. In some embodiments, the microdeletions detected in a fetus are 22q11.2 deletions associated with DiGeorge syndrome, microdeletions associated with Prader-Willi syndrome, microdeletions associated with Angelman syndrome, 1p36 deletions, and microdeletions associated with cri du chat syndrome. In some embodiments, the methods disclosed herein may be used to detect single nucleotide variations in a fetus of a pregnant mother.

[0064] In some embodiments, the methods disclosed herein may be used to diagnose or monitor cancer patients. In some embodiments, the methods disclosed herein may be used to detect copy number variations in tumor tissue by isolating and detecting circulating tumor cells from a patient's blood sample. In some embodiments, the methods herein may be used to detect one or more patient-specific mutations associated with cancer in a cancer patient. In some embodiments, the methods herein may be used to detect at least two patient-specific mutations associated with cancer in a tumor of a cancer patient. In some embodiments, the methods herein may be used to detect mutations associated with early recurrence or metastasis of a tumor in a cancer patient. In some embodiments, the methods disclosed herein may be used to monitor the response of a cancer patient to treatment.

[0065] In some embodiments, the methods disclosed herein may be used to monitor recipients of transplanted organs by isolating and analyzing circulating cells originating from the transplanted organ. In some embodiments, the methods disclosed herein may be used to diagnose graft-versus-host disease in recipients of transplanted organs. In some embodiments, the methods disclosed herein may be used to monitor the response of organ transplant patients to immunosuppressive therapy.

[0066] 5. Acquisition of Genotype Data 5.1 Sources of Gene Data Any relevant individual genetic data can be obtained from: the bulk diploid tissue of an individual, one or more diploid cells derived from an individual, one or more haploid cells derived from an individual, one or more blastomeres derived from a target individual, extracellular genetic material present in an individual, extracellular genetic material of an individual present in the mother's blood or the blood of a cancer patient, cells of an individual present in the mother's blood, one or more embryos generated from the gametes of a related individual, one or more blastomeres collected from the embryo, extracellular genetic material present in a related individual, genetic material known to originate from a related individual, and combinations thereof. In an exemplary embodiment, the methods provided herein are used to analyze cell-free DNA originating from the genome of a target sample from target cells such as, for example, fetal cells or tumor cells. In one embodiment, the methods disclosed herein further comprise obtaining from a blood sample circulating cells presumed to be of fetal origin, a plasma fraction containing cell-free DNA, and a buffy coat fraction containing maternal cells. In one embodiment, the methods disclosed herein further comprise obtaining from a blood sample circulating cells presumed to be tumor cells, a plasma fraction containing cell-free DNA, and a buffy coat fraction containing normal cells.

[0067] As used herein, the term "cell-free DNA" refers to DNA that is available for analysis without the need for a cell lysis step. Cell-free DNA can be found in blood or other body fluids. Cell-free DNA may be obtained from various tissues. Such tissues may be, for example, tissues in liquid form such as blood, lymph, ascites, cerebrospinal fluid, etc. Cell-free DNA may be derived from various cell sources. In some examples, cell-free DNA is composed of DNA derived from fetal cells. Cell-free DNA may be a mixture of DNA derived from target cells and non-target cells. In the case of analyzing DNA for fetal aneuploidy, in some embodiments, cell-free DNA may be obtained from the blood of a pregnant woman, in which case the cell-free DNA includes a mixture of cell-free DNA derived from the mother and cell-free DNA derived from the fetus. In other embodiments, cell-free DNA may be derived from cancerous tumor cells. Cell-free DNA may include a mixture of cell-free DNA derived from tumor cells and cell-free DNA derived from cells that are not tumors in other parts of the body.

[0068] In certain illustrative embodiments, the sample analyzed by the methods of the present invention is a blood sample or a fraction thereof. In certain embodiments, the methods provided herein are particularly suitable for the amplification of DNA fragments, particularly tumor DNA fragments found in circulating tumor DNA (ctDNA). Such fragments typically have a length of about 160 nucleotides.

[0069] For example, it is known in this technical field that cell-free nucleic acids (cfNA), such as cfDNA, can be released into the circulatory system through various forms of cell death, such as apoptosis, necrosis, autolysis, and necroptosis. cfDNA is fragmented, and the size distribution of the fragments varies from 150 to 350 bp to over 10,000 bp (see Kalnina et al. World J Gastroenterol. 2015 Nov 7;21(41):11636-11653). For example, the size distribution of plasma DNA fragments in patients with hepatocellular carcinoma (HCC) ranges from 100 to 220 bp with a peak in the count frequency at approximately 166 bp, and the tumor DNA concentration was highest in fragments of 150 to 180 bp (see Jiang et al. Proc Natl Acad Sci USA 112:E1317-E1325).

[0070] In an illustrative embodiment, circulating tumor DNA (ctDNA) was isolated from blood using an EDTA-2Na tube after removing cell debris and platelets by centrifugation. Plasma samples can be stored at -80°C until DNA extraction, for example, using the QIAamp DNA Mini Kit (Qiagen, Hilden, Germany) (see, for example, Hamakawa et al., Br J Cancer. 2015;112:352-356). Hamakava et al. reported that the median concentration of the extracted cell-free DNA in the total samples was 43.1 ng per ml of plasma (ranging from 9.5 to 1338 ng / ml), the mutant fragments ranged from 0.001 to 77.8%, and the median was 0.90%.

[0071] For example, genetic data such as DNA sequence data can be obtained from a mixture of DNA containing DNA derived from one or more target cells and DNA derived from one or more non-target cells. The target cells and non-target cells are genomically different from each other based on other criteria. The term "derived" is used to indicate that the cell is the ultimate source of the DNA. Thus, for example, cell-free DNA obtained from the maternal blood of a pregnant woman is derived from cells of fetal placental origin, which are typically genetically identical to the cells of the fetus itself and the mother. The method employs a combination of patients.

[0072] 5.2 Amplification (e.g., PCR) reaction mixture: In some embodiments, the acquisition of genotype data includes sequencing (e.g., shotgun sequencing, single molecule sequencing, or other next generation sequencing technologies, etc.), SNP arrays for detecting polymorphic loci, or multiplex PCR. In some embodiments, genotype analysis involves the use of SNP arrays that detect polymorphic loci such as, for example, at least 100, 200, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 20,000, 25,000, 30,000, 40,000, 50,000, 75,000, or 100,000 different polymorphic loci. In some embodiments, genotype analysis involves the use of multiplex PCR. In some embodiments, the method involves contacting a sample in a fraction with a library of primers that simultaneously hybridize to at least 100, 200, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 13,000, 15,000, 20,000, 25,000, 30,000, 40,000, 50,000, 75,000, or 100,000 different polymorphic loci (e.g., SNPs) to generate a reaction mixture, placing the reaction mixture under primer extension reaction conditions to generate an amplification product, and measuring with a high throughput sequencer to generate sequence data. In some embodiments, RNA (e.g., mRNA) is sequenced. Since mRNA contains only exons, sequencing mRNA can determine alleles for polymorphic loci (e.g., SNPs, etc.) over long distances in the genome, such as, for example, several megabases.

[0073] In one embodiment, multiplex target PCR is performed directly on the lysed cell sample. In some embodiments, the cells are lysed by protease K digestion. In some embodiments, the cells are lysed by protease K digestion in a saline solution containing potassium chloride (KCl), magnesium chloride (MgCl2), and tris - hydrochloride. In some embodiments, in order to prepare cells for PCR amplification, the cells are lysed by protease K digestion in a saline solution that does not contain potassium chloride (KCl), magnesium chloride (MgCl2), and tris - hydrochloride. For further details, refer to the Examples section.

[0074] In some embodiments, the sample barcode is added to the DNA during the amplicon generation step.

[0075] In certain embodiments, the method of the present invention includes forming an amplification reaction mixture. This reaction mixture is typically formed by combining a polymerase, nucleotide triphosphates, nucleic acid fragments from a nucleic acid library generated from a sample, a series of target - specific forward outer primers, and a single - strand (first strand) reverse outer universal primer. In other illustrative embodiments, the reaction mixture includes a target - specific forward inner primer, which is an alternative to the target - specific forward outer primer, and an amplicon from a first PCR reaction using an outer primer, which is an alternative to the nucleic acid fragment from the nucleic acid library. This reaction mixture provided herein is, in illustrative embodiments, each formed as an independent aspect of this invention. In illustrative embodiments, this reaction mixture is a PCR reaction mixture. The PCR reaction mixture typically contains magnesium.

[0076] In some embodiments, the reaction mixture includes ethylenediaminetetraacetic acid (EDTA), magnesium, tetramethylammonium chloride (TMAC), or combinations thereof. In some embodiments, the concentration of TMAC is between 20 and 70 mM (including both end values). Without intending to be bound by any particular theory, it has been found that TMAC binds to DNA, stabilizes double strands, increases primer specificity, and / or equalizes the melting points of various primers. In some embodiments, TMAC increases the uniformity of the amount of amplification products for various targets. In some embodiments, the concentration of magnesium (such as magnesium from magnesium chloride) is between 1 and 8 mM.

[0077] A large number of primers for multiplex PCR of a large number of targets can chelate a large amount of magnesium (two phosphates in the primer chelate one magnesium). For example, if sufficient primers are used such that the phosphate concentration from the primers is about 9 mM, the primers can reduce the effective magnesium concentration by about 4.5 mM. In some embodiments, EDTA is used to reduce the amount of magnesium available as a cofactor for polymerase because high concentrations of magnesium can sometimes cause PCR errors such as amplification of non-target loci. In some embodiments, the concentration of EDTA reduces the amount of available magnesium to between 1 and 5 mM (such as between 3 and 5 mM).

[0078] In some embodiments, the pH is between 7.5 and 8.5, such as between 7.5 and 8, between 8 and 8.3, or between 8.3 and 8.5 (including the values at both ends). In some embodiments, Tris buffer is used at a concentration, for example, between 10 and 100 mM, such as between 10 and 25 mM, between 25 and 50 mM, between 50 and 75 mM, or between 25 and 75 mM (including the values at both ends). In some embodiments, Tris buffers at these concentrations are all used at a pH between 7.5 and 8.5. In some embodiments, a combination of KCl and (NH4)2SO4 is used, such as KCl between 50 and 150 mM and (NH4)2SO4 between 10 and 90 mM (including the values at both ends). In some embodiments, the concentration of KCl is between 0 and 30 mM, between 50 and 100 mM, or between 100 and 150 mM (including the values at both ends). In some embodiments, the concentration of (NH4)2SO4 is between 10 and 50 mM, between 50 and 90 mM, between 10 and 20 mM, between 20 and 40 mM, between 40 and 60 mM, or between 60 and 80 mM (including the values at both ends). In some embodiments, the concentration of ammonium [NH4+] is between 0 and 160 mM, such as between 0 and 50, between 50 and 100, or between 100 and 160 mM (including the values at both ends). In some embodiments, the total concentration of potassium and ammonium ([K+]+[NH4+]) is between 0 and 160 mM, such as between 0 and 25, between 25 and 50, between 50 and 150, between 50 and 75, between 75 and 100, between 100 and 125, or between 125 and 160 mM (including the values at both ends). An exemplary buffer with [K+]+[NH4+] being 120 mM has 20 mM of KCl and 50 mM of (NH4)2SO4. In some embodiments, the buffer contains 25 to 75 mM of Tris buffer with a pH of 7.2 to 8, 0 to 50 mM of KCl, 10 to 80 mM of ammonium sulfate, and 3 to 6 mM of magnesium (including the values at both ends).In some embodiments, the buffer contains 25 to 75 mM of Tris buffer with a pH of 7 to 8.5, 3 to 6 mM of MgCl2, 10 to 50 mM of KCl and 20 to 80 mM of (NH4)2SO4 (including the values at both ends). In some embodiments, 100 to 200 Units / mL of polymerase is used. In some embodiments, in a final volume of 20 μl at pH 8.1, 100 mM of KCl, 50 mM of (NH4)2SO4, 3 mM of MgCl2, 7.5 nM of each primer in the library, 50 mM of TMAC, and 7 μl of DNA template are used.

[0079] In some embodiments, a crowding agent such as polyethylene glycol (PEG, e.g., PEG8,000) or glycerol is used. In some embodiments, the amount of PEG (e.g., PEG8,000) is from 0.1 to 20%, such as between 0.5 and 15%, between 1 and 10%, between 2 and 8% or between 4 and 8% (including the values at both ends). In some embodiments, the amount of glycerol is between 0.1 and 20%, such as between 0.5 and 15%, between 1 and 10%, between 2 and 8% or between 4 and 8% (including the values at both ends). In some embodiments, the crowding agent enables a reduction in the concentration of the polymerase used and / or a shortening of the annealing time. In some embodiments, the crowding agent improves the uniformity of the DOR and / or reduces dropout (undetected alleles). In some embodiments, a polymerase having proofreading activity, a polymerase having no (or very little) proofreading activity, or a mixture of a polymerase having proofreading activity and a polymerase having no (or very little) proofreading activity is used. In some embodiments, a hot start polymerase, a non-hot start polymerase, or a mixture of a hot start polymerase and a non-hot start polymerase is used. In some embodiments, HotStarTaq DNA polymerase is used (see, e.g., Catalog No. 203203 of Qiagen). In some embodiments, AmpliTaq Gold® DNA polymerase is used. In some embodiments, PrimeSTAR GXL DNA polymerase, a high-fidelity polymerase that provides effective PCR amplification when an excess of template is present in the reaction mixture and when amplifying long products, is used (Mountain View, California, Takara Clontech).In some embodiments, KAPA Taq DNA polymerase or KAPA Taq HotStart DNA polymerase is used, which are single-subunit, wild-type Taq DNA polymerase-based thermophilic bacteria, Thermus aquaticus. KAPA Taq and KAPA Taq HotStart DNA polymerases have 5'-3' polymerase and 5'-3' exonuclease activities, but do not have 3' to 5' exonuclease (proofreading) activity (see, for example, KAPA BIOSYSTEMS catalog No. BK1000). In some embodiments, Pfu DNA polymerase is used, which is a highly thermostable DNA polymerase from the hyperthermophilic archaeon Pyrococcus furiosus. The enzyme catalyzes the template-dependent polymerization of nucleotides into double-stranded DNA in the 5'→3' direction. Pfu DNA polymerase also exhibits 3'→5' exonuclease (proofreading) activity that allows it to correct nucleotide incorporation errors for the polymerase. The polymerase does not have 5'→3' exonuclease activity (see, for example, Thermo Scientific catalog No. EP0501). In some embodiments, Klentaq1 is used, which is a Klenow-fragment analog of Taq DNA polymerase and does not have exonuclease or endonuclease activity (see, for example, DNA POLYMERASE TECHNOLOGY catalog No. 100, St. Louis, Missouri). In some embodiments, the polymerase is PHUSION DNA polymerase, such as PHUSION High-Fidelity DNA polymerase (M0530S, New England BioLabs, Inc.) or PHUSION Hot Start Flex DNA polymerase (M0535S, New England BioLabs, Inc.).In some embodiments, the polymerase is Q5® DNA polymerase, such as Q5® High-Fidelity DNA polymerase (M0491S, New England BioLabs, Inc.) or Q5® Hot Start High-Fidelity DNA polymerase (M0493S, New England BioLabs, Inc.). In some embodiments, the polymerase is T4 DNA polymerase (M0203S, New England Biolabs).

[0080] In some embodiments, a polymerase between 5 and 600 Units / mL (units per mL of reaction volume) is used, such as between 5 and 100, 100 and 200, 200 and 300, 300 and 400, 400 and 500, or 500 and 600 Units / mL (including the end values).

[0081] 5.3 PCR method In some embodiments, hot start PCR is used to reduce or prevent polymerization before performing the PCR thermocycle. Exemplary hot start PCR methods include suppressing the initial DNA polymerase until the reaction mixture reaches a high temperature, or physically separating the reactions of the reaction components. In some embodiments, the slow release of magnesium is utilized. Since DNA polymerase requires magnesium ions for activity, magnesium is chemically separated from the reaction by binding it to a chemical substance and released into the solution only at high temperatures. In some embodiments, non-covalent binding of an inhibitor is utilized. In this method, a peptide, antibody or aptamer binds non-covalently to the enzyme at low temperature and suppresses its activity. After incubation at an elevated temperature, the inhibitor is released and the reaction begins. In some embodiments, a cold-sensitive Taq polymerase, such as a modified DNA polymerase that has little activity at low temperature, is utilized. In some embodiments, chemical modification is utilized. In this method, a molecule is covalently bound to the side chain of an amino acid within the active site of the DNA polymerase. This molecule is released from the enzyme by incubating the reaction mixture at an elevated temperature. When the molecule is released, the enzyme is activated.

[0082] In some embodiments, the amount of template nucleic acid (such as an RNA or DNA sample) is between 20 and 5,000 ng, such as between 20 and 200, 200 and 400, 400 and 600, 600 and 1,000, 1,000 and 1,500 or 2,000 and 3,000 ng (including both end values).

[0083] In some embodiments, the multiplex PCR kit from Qiagen is utilized (Qiagen catalog No. 206143). For a 100×50 μL multiplex PCR reaction, the kit contains 2× QIAGEN Multiplex PCR Master Mix (providing a final concentration of 3 mM MgCl2, 3×0.85 ml), 5× Q-Solution (1×2.0 ml), and RNase-free water (2×1.7 ml). QIAGEN Multiplex PCR Master Mix (MM) contains not only a combination of KCl and (NH4)2SO4, but also PCR additives and Factor MP that increases the local concentration of primers in the template. Factor MP specifically stabilizes the bound primers to enable effective primer extension by HotStarTaq DNA polymerase. HotStarTaq DNA polymerase is a modified form of Taq DNA polymerase and has no polymerase activity at room temperature. In some embodiments, HotStarTaq DNA polymerase is activated by incubation at 95 °C for 15 minutes, which can be incorporated into any existing thermocycle program.

[0084] In some embodiments, for a 1× QIAGEN MM final concentration (recommended concentration), each primer in the library is 7.5 nM, TMAC is 50 mM, and 7 μl of DNA template is used in a final volume of 20 μl. In some embodiments, the PCR thermocycle conditions include a 10-minute incubation at 95 °C (hot start), 20 cycles of 30 seconds at 96 °C, 15 minutes at 65 °C, and 30 seconds at 72 °C, followed by a 2-minute incubation at 72 °C (final extension), and then holding at 4 °C.

[0085] In some embodiments, in a total volume of 20 μl, 2× QIAGEN MM final concentration (2 times the recommended concentration), each primer in the library at 2 nM, TMAC at 70 mM, and 7 μl of DNA template are used. In some embodiments, EDTA up to 4 mM is also included. In some embodiments, the PCR thermocycle conditions include 10 minutes at 95 °C (hot start), 25 cycles of 30 seconds at 96 °C, 20, 25, 30, 45, 60, 120, or 180 minutes at 65 °C and optionally 30 seconds at 72 °C, then 2 minutes at 72 °C (final extension), and then holding at 4 °C.

[0086] Another exemplary set of conditions includes a semi-nested PCR approach. For the first PCR reaction, a 20 μl reaction volume is used that includes 2× QIAGEN MM final concentration, each primer in the library at 1.875 nM (forward and reverse outer primers), and a DNA template.

[0087] The thermocycling parameters include 10 minutes at 95 °C, 30 seconds at 96 °C, 1 minute at 65 °C, 6 minutes at 58 °C, 8 minutes at 60 °C, 4 minutes at 65 °C, and 30 seconds at 72 °C for 25 cycles, then 2 minutes at 72 °C, and then holding at 4 °C. Next, 2 μl of the resulting product diluted 1:200 is used as the input for a second PCR reaction. For this reaction, a 10 μl reaction volume is used that includes 1× QIAGEN MM final concentration, each forward inner primer at 20 nM, and a reverse primer tag at 1 μM. The thermocycling parameters include 10 minutes at 95 °C, 30 seconds at 95 °C, 1 minute at 65 °C, 5 minutes at 60 °C, 5 minutes at 65 °C, and 30 seconds at 72 °C for 15 cycles, then 2 minutes at 72 °C, and then holding at 4 °C. The annealing temperature can optionally be higher than the melting point of some or all of the primers, as described herein (see U.S. Patent Application No. 14 / 918,544, filed October 20, 2015, which is incorporated herein by reference in its entirety).

[0088] The melting point (Tm) is the temperature at which half (50%) of the DNA double strand of an oligonucleotide (e.g., a primer) and its complete complementary strand dissociate to become single-stranded DNA. The annealing temperature (TA) is the temperature at which the PCR protocol is carried out. In conventional methods, it is usually 5°C lower than the lowest Tm of the primers used, such that all possible double strands are formed almost (essentially all primer molecules bind to the template nucleic acid). This is highly efficient, but at low temperatures, it is inevitable that more non-specific reactions will occur. As a result of the TA being too low, internal single-base mismatches or partial annealing may be tolerated, so that the primer may anneal to sequences other than the true target. In some embodiments of the present invention, the TA is higher than the Tm, and at a certain moment only a very small part of the target (e.g., only about 1 - 5%) anneals the primer. If these extend, they break away from the equilibrium state of primer and target annealing and dissociation (because the Tm immediately rises above 70°C upon extension), and about 1 - 5% of the new target has the primer. Thus, by making the reaction a long annealing, it is possible to obtain about 100% of the copied target per cycle.

[0089] In various embodiments, the annealing temperature is between 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 °C and 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or 15 °C at the upper limit of the range, and is higher than the melting point of at least 25, 50, 60, 70, 75, 80, 90, 95 or 100% of the mismatched primers (such as Tm measured or calculated empirically). In various embodiments, the annealing temperature is 1 to 15 °C (such as 1 °C or more to 10 °C or less, 1 °C or more to 5 °C or less, 1 °C or more to 3 °C or less, 3 °C or more to 5 °C or less, 5 °C or more to 10 °C or less, 5 °C or more to 8 °C or less, 8 °C or more to 10 °C or less, 10 °C or more to 12 °C or less, or 12 °C or more to 15 °C or less) higher than the melting point of at least 25, 50, 75, 100, 300, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 15,000, 19,000, 20,000, 25,000, 27,000, 28,000, 30,000, 40,000, 50,000, 75,000, 100,000 of the non-identical primers, or all of the melting points (e.g., Tm measured or calculated experimentally). In various embodiments, the annealing temperature is between 1 and 15 °C (such as between 1 and 10, 1 and 5, 1 and 3, 3 and 5, 3 and 8, 5 and 10, 5 and 8, 8 and 10, 10 and 12 or 12 and 15 °C (including both end values)), and is higher than the melting point of at least 25%, 50%, 60%, 70%, 75%, 80%, 90%, 95% of the mismatched primers (such as Tm measured or calculated empirically) or all of the melting points of the mismatched primers (such as Tm measured or calculated empirically), and the length of the annealing step (per 1 PCR cycle) is between 5 and 180 minutes, such as between 15 and 120 minutes, 15 and 60 minutes, 15 and 45 minutes or 20 and 60 minutes (including both end values).

[0090] Exemplary multiplex PCR method

[0091] In various embodiments, long annealing times (as described herein and exemplified in Example 12) and / or low primer concentrations are used. In fact, in certain embodiments, restricted primer concentrations and / or conditions are used. In various embodiments, the length of the annealing step ranges from a lower limit of 15, 20, 25, 30, 35, 40, 45 or 60 minutes to an upper limit of 20, 25, 30, 35, 40, 45, 60, 120 or 180 minutes. In various embodiments, the length of the annealing step (per PCR cycle) is between 30 and 180 minutes. For example, the annealing step can be between 30 and 60 minutes, and the concentration of each primer can be less than 20, 15, 10 or 5 nM. In other embodiments, the primer concentration ranges from a lower limit of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20 or 25 nM to an upper limit of 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25 and 50.

[0092] In high-concentration multiplexing, due to the large amount of primers in the solution, the solution can become viscous. If the solution is too viscous, it is possible to reduce the primer concentration to an amount still sufficient for the primer to bind to the template DNA. In various embodiments, various primers from 1,000 to 100,000 are used, with the concentration of each primer being less than 20 nM, for example less than 10 nM, or between 1 and 10 nM (including both end values).

[0093] 5.4 Exemplary Multiplex PCR Method In one aspect, the present invention features a method of amplifying target loci in a nucleic acid sample, comprising: (i) contacting the nucleic acid sample with a primer library that simultaneously hybridizes to at least 1,000; 2,000; 5,000; 7,500; 10,000; 15,000; 19,000; 20,000; 25,000; 27,000; 28,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different target loci to generate a reaction mixture; and (ii) subjecting the reaction mixture to primer extension reaction conditions (such as PCR conditions) to generate an amplification product comprising a target amplicon. In some embodiments, the method also includes determining the presence or absence of at least one target amplicon (e.g., at least 50, 60, 70, 80, 90, 95, 96, 97, 98, 99, or 99.5% of the target amplicon). In some embodiments, the method also includes determining the sequence of at least one target amplicon (e.g., at least 50, 60, 70, 80, 90, 95, 96, 97, 98, 99, or 99.5% of the target amplicon). In some embodiments, at least 50, 60, 70, 80, 90, 95, 96, 97, 98, 99, or 99.5% of the target loci are amplified. In some embodiments, at least 25; 50; 75; 100; 300; 500; 750; 1,000; 2,000; 5,000; 7,500; 10,000; 15,000; 19,000; 20,000; 25,000; 27,000; 28,000; 30,000; 40,000; 50,000; 75,000; or 100,000 different target loci are amplified at least 5, 10, 20, 40, 50, 60, 80, 100, 120, 150, 200, 300, or 400-fold. In some embodiments, at least 50, 60, 70, 80, 90, 95, 96, 97, 98, 99, 99.5, or 100% of the target loci are amplified at least 5, 10, 20, 40, 50, 60, 80, 100, 120, 150, 200, 300, or 400-fold. In various embodiments, less than 60, 50, 40, 30, 20, 10, 5, 4, 3, 2, 1, 0.5, 0.25, 0.1, or 0.05% of the amplification product is primer dimer.In some embodiments, the method involves multiplex PCR and sequencing (e.g., high-throughput sequencing).

[0094] In various embodiments, long annealing times and / or low concentrations of primers are used. In various embodiments, the length of the annealing step is longer than 3 minutes, 5 minutes, 8 minutes, 10 minutes, 15 minutes, 20 minutes, 30 minutes, 45 minutes, 60 minutes, 75 minutes, 90 minutes, 120 minutes, 150 minutes, or 180 minutes. In various embodiments, the length of the annealing step (per PCR cycle) is, for example, 5 minutes or more to 60 minutes or less, 10 minutes or more to 60 minutes or less, 5 minutes or more to 30 minutes or less, or 10 minutes or more to 30 minutes or less, etc., 5 minutes or more to 180 minutes or less. In various embodiments, the length of the annealing step is longer than 5 minutes (e.g., longer than 10 minutes or 15 minutes), and the concentration of each primer is less than 20 nM. In various embodiments, the length of the annealing step is longer than 5 minutes (e.g., longer than 10 minutes or 15 minutes), and the concentration of each primer is 1 nM or more to 20 nM or less, or 1 nM or more to 10 nM or less. In various embodiments, the length of the annealing step is longer than 20 minutes (e.g., longer than 30 minutes, 45 minutes, 60 minutes or 90 minutes), and the concentration of each primer is less than 1 nM.

[0095] In high-concentration multiplexing, due to a large amount of primers in the solution, the solution can become viscous. If the solution is too viscous, it is possible to reduce the primer concentration to an amount still sufficient for the primers to bind to the template DNA. In various embodiments, fewer than 60,000 various primers are used, and the concentration of each primer is less than 20 nM, such as, for example, less than 10 nM, or 1 nM or more to 10 nM or less. In various embodiments, more than 60,000 various primers (e.g., 60,000 to 120,000 various primers) are used, and the concentration of each primer is less than 10 nM, such as, for example, less than 5 nM, or 1 nM or more to 10 nM or less.

[0096] The annealing temperature has been found to optionally be higher than the melting point of some or all of the primers (in contrast, other methods use an annealing temperature lower than the melting point of the primers). The melting point (T m ) is the temperature at which half (50%) of the DNA double strand of an oligonucleotide (e.g., a primer) dissociates from its complete complementary strand to become single-stranded DNA. The annealing temperature (T A ) is the temperature at which the PCR protocol is carried out. Conventionally, it is usually 5°C lower than the lowest T m of the primers used, so that almost all possible double strands were formed (since essentially all primer molecules bind to the template nucleic acid). This is highly efficient, but at low temperatures, it is inevitable that more non-specific reactions will occur. As a result of T A being too low, internal single-base mismatches or partial annealing may be tolerated, so that the primer may anneal to sequences other than the true target. In some embodiments of the present invention, T A is higher than (T m ), and at a certain moment only a very small part of the target anneals the primer (e.g., only about 1-5%). If these are extended, they will break away from the equilibrium state of primer and target annealing and dissociation (since T m immediately rises above 70°C during extension), and a new about 1-5% of the target has primers. Thus, by making the reaction a long annealing, it is possible to obtain about 100% of the copied target per cycle. Therefore, the most stable molecular pairs (molecular pairs with complete DNA pairs of primers and template DNA) are preferentially extended, and the correct target amplicons are generated. For example, the same experiment was performed using a primer with a melting point of less than 63°C, with an annealing temperature of 57°C and an annealing temperature of 63°C. When the annealing temperature was 57°C, the proportion of mapped reads of the amplified PCR product was as low as 50% (about 50% of the amplified product was primer dimer). When the annealing temperature was 63°C, the proportion of primer dimer in the amplified product decreased to about 2%.

[0097] In various embodiments, the annealing temperature is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 15 °C higher than the melting point of at least 25, 50, 75, 100, 300, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 15,000, 19,000, 20,000, 25,000, 27,000, 28,000, 30,000, 40,000, 50,000, 75,000, 100,000 of the non-identical primers, or all of the melting points (e.g., experimentally measured or calculated T m ) Also, in some embodiments, the annealing temperature is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 15 °C higher than the melting point of at least 25, 50, 75, 100, 300, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 15,000, 19,000, 20,000, 25,000, 27,000, 28,000, 30,000, 40,000, 50,000, 75,000, 100,000 of the non-identical primers, or all of the melting points (e.g., experimentally measured or calculated T m ) and the length of the annealing step (per PCR cycle) is longer than 1, 3, 5, 8, 10, 15, 20, 30, 45, 60, 75, 90, 120, 150, or 180 minutes.

[0098] In various embodiments, the annealing temperature is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 15 °C higher than the melting point of at least 25, 50, 75, 100, 300, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 15,000, 19,000, 20,000, 25,000, 27,000, 28,000, 30,000, 40,000, 50,000, 75,000, 100,000 of the non-identical primers, or all of the melting points (e.g., experimentally measured or calculated T m) is 1 to 15 °C higher (for example, 1 °C or more to 10 °C or less, 1 °C or more to 5 °C or less, 1 °C or more to 3 °C or less, 3 °C or more to 5 °C or less, 5 °C or more to 10 °C or less, 5 °C or more to 8 °C or less, 8 °C or more to 10 °C or less, 10 °C or more to 12 °C or less, or 12 °C or more to 15 °C or less). In various embodiments, the annealing temperature is at least 25, 50, 75, 100, 300, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 15,000, 19,000, 20,000, 25,000, 27,000, 28,000, 30,000, 40,000, 50,000, 75,000, 100,000, or all melting points (for example, experimentally measured or calculated T m ) is 1 to 15 °C higher (for example, 1 °C or more to 10 °C or less, 1 °C or more to 5 °C or less, 1 °C or more to 3 °C or less, 3 °C or more to 5 °C or less, 5 °C or more to 10 °C or less, 5 °C or more to 8 °C or less, 8 °C or more to 10 °C or less, 10 °C or more to 12 °C or less, or 12 °C or more to 15 °C or less), and the length of the annealing step (per PCR cycle) is, for example, 5 minutes or more to 180 minutes or less, such as 5 minutes or more to 60 minutes or less, 10 minutes or more to 60 minutes or less, 5 minutes or more to 30 minutes or less, or 10 minutes or more to 30 minutes or less.

[0099] In some embodiments, the annealing temperature is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 15 °C higher than the highest melting point of the primer (for example, experimentally measured or calculated T m ) is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 15 °C higher than the highest melting point of the primer (for example, experimentally measured or calculated T m ) is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 15 °C higher than the highest melting point of the primer, and the length of the annealing step (per PCR cycle) is longer than 1, 3, 5, 8, 10, 15, 20, 30, 45, 60, 75, 90, 120, 150, or 180 minutes.

[0100] In some embodiments, the annealing temperature is the highest melting point of the primer (for example, experimentally measured or calculated Tm ) is 1 to 15 °C higher (for example, within 1 °C or more to 10 °C, within 1 °C or more to 5 °C, within 1 °C or more to 3 °C, within 3 °C or more to 5 °C, within 5 °C or more to 10 °C, within 5 °C or more to 8 °C, within 8 °C or more to 10 °C, within 10 °C or more to 12 °C, or within 12 °C or more to 15 °C) than. In some embodiments, the annealing temperature is the highest melting point of the primer (for example, experimentally measured or calculated T m ) is 1 to 15 °C higher (for example, within 1 °C or more to 10 °C, within 1 °C or more to 5 °C, within 1 °C or more to 3 °C, within 3 °C or more to 5 °C, within 5 °C or more to 10 °C, within 5 °C or more to 8 °C, within 8 °C or more to 10 °C, within 10 °C or more to 12 °C, or within 12 °C or more to 15 °C), and the length of the annealing step (per PCR cycle) is, for example, within 5 minutes or more to 60 minutes, within 10 minutes or more to 60 minutes, within 5 minutes or more to 30 minutes, or within 10 minutes or more to 30 minutes, etc., within 5 minutes or more to 180 minutes.

[0101] In some embodiments, the annealing temperature is at least 25, 50, 75, 100, 300, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 15,000, 19,000, 20,000, 25,000, 27,000, 28,000, 30,000, 40,000, 50,000, 75,000, 100,000 of the primers that are not the same, or the average melting point of all (for example, experimentally measured or calculated T m ) is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 15 °C higher. In some embodiments, the annealing temperature is at least 25, 50, 75, 100, 300, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 15,000, 19,000, 20,000, 25,000, 27,000, 28,000, 30,000, 40,000, 50,000, 75,000, 100,000 of the primers that are not the same, or the average melting point of all (for example, experimentally measured or calculated T mAt least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or 15 °C higher than , and the length of the annealing step (per PCR cycle) is longer than 1, 3, 5, 8, 10, 15, 20, 30, 45, 60, 75, 90, 120, 150, or 180 minutes.

[0102] In some embodiments, the annealing temperature is 1 - 15 °C (e.g., 1 °C or more to 10 °C or less, 1 °C or more to 5 °C or less, 1 °C or more to 3 °C or less, 3 °C or more to 5 °C or less, 5 °C or more to 10 °C or less, 5 °C or more to 8 °C or less, 8 °C or more to 10 °C or less, 10 °C or more to 12 °C or less, or 12 °C or more to 15 °C or less) higher than the average melting point of at least 25, 50, 75, 100, 300, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 15,000, 19,000, 20,000, 25,000, 27,000, 28,000, 30,000, 40,000, 50,000, 75,000, 100,000 of the non-identical primers, or all of the average melting points (e.g., measured experimentally or calculated T m ) 1 - 15 °C (e.g., 1 °C or more to 10 °C or less, 1 °C or more to 5 °C or less, 1 °C or more to 3 °C or less, 3 °C or more to 5 °C or less, 5 °C or more to 10 °C or less, 5 °C or more to 8 °C or less, 8 °C or more to 10 °C or less, 10 °C or more to 12 °C or less, or 12 °C or more to 15 °C or less) higher than . In some embodiments, the annealing temperature is 1 - 15 °C (e.g., 1 °C or more to 10 °C or less, 1 °C or more to 5 °C or less, 1 °C or more to 3 °C or less, 3 °C or more to 5 °C or less, 5 °C or more to 10 °C or less, 5 °C or more to 8 °C or less, 8 °C or more to 10 °C or less, 10 °C or more to 12 °C or less, or 12 °C or more to 15 °C or less) higher than the average melting point of at least 25, 50, 75, 100, 300, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 15,000, 19,000, 20,000, 25,000, 27,000, 28,000, 30,000, 40,000, 50,000, 75,000, 100,000 of the non-identical primers, or all of the average melting points (e.g., measured experimentally or calculated T m ) 1 - 15 °C (e.g., 1 °C or more to 10 °C or less, 1 °C or more to 5 °C or less, 1 °C or more to 3 °C or less, 3 °C or more to 5 °C or less, 5 °C or more to 10 °C or less, 5 °C or more to 8 °C or less, 8 °C or more to 10 °C or less, 10 °C or more to 12 °C or less, or 12 °C or more to 15 °C or less) higher than , and the length of the annealing step (per PCR cycle) is, for example, 5 minutes or more to 180 minutes or less, such as 5 minutes or more to 60 minutes or less, 10 minutes or more to 60 minutes or less, 5 minutes or more to 30 minutes or less, or 10 minutes or more to 30 minutes or less.

[0103] In some embodiments, the annealing temperature is from 50°C to 70°C, such as from 55°C to 60°C, from 60°C to 65°C, or from 65°C to 70°C. In some embodiments, the annealing temperature is from 50°C to 70°C, such as from 55°C to 60°C, from 60°C to 65°C, or from 65°C to 70°C, and (i) the length of the annealing step (per PCR1 cycle) is longer than 3, 5, 8, 10, 15, 20, 30, 45, 60, 75, 90, 120, 150 or 180 minutes, or (ii) the length of the annealing step (per PCR1 cycle) is, for example, from 5 minutes to 180 minutes, such as from 5 minutes to 60 minutes, from 10 minutes to 60 minutes, from 5 minutes to 30 minutes, or from 10 minutes to 30 minutes.

[0104] In some embodiments, one or more of the following conditions are used for experimental measurement of T m or are assumed and T m is calculated: the temperature is 60.0°C, the primer concentration is 100 nM, and / or the salt concentration is 100 mM. In some embodiments, other conditions are used, such as those used for multiplex PCR using a library. In some embodiments, 100 mM KCl, 50 mM (NH4)2SO4, 3 mM MgCl2, the primer is 7.5 nM, and 50 mM TMAC, pH 8.1 are used. In some embodiments, T mis calculated using the Primer3 program (libprimer3, release 2.2.3) with the internal SantaLucia parameters (world wide web of primer3.sourceforge.net, incorporated herein by reference in its entirety). In some embodiments, the calculated primer melting point is the temperature at which half of the primer molecules are expected to anneal. As described above, even at a temperature higher than the calculated melting point, a certain percentage of the primers will anneal and accordingly PCR extension is also possible. In some embodiments, the experimentally measured Tm (actual Tm) is determined by using a thermostatically controlled cell in a UV spectrophotometer. In some embodiments, by plotting temperature against absorbance, an S-shaped curve with two plateaus is generated. The absorbance value in the middle between the plateaus corresponds to the Tm.

[0105] In some embodiments, the absorbance at 260 nm with an ultraspec 2100 pr UV / visible spectrophotometer (Amersham Biosciences) is measured as a function of temperature (see, e.g., Takiya et al., “An empirical approach for thermal stability (Tm) prediction of PNA / DNA duplexes,” Nucleic Acids Symp Ser (Oxf); (48):131-2, 2004, which is hereby incorporated by reference in its entirety). In some embodiments, the absorbance at 260 nm is measured by decreasing the temperature from 95° C. to 20° C. at a pace of 2° C. per minute. In some embodiments, a primer and its complete complementary strand (e.g., 2 μM of each pair of oligomers) are mixed, then the sample is heated to 95° C. for 30 minutes and maintained at that temperature for 5 minutes, and then cooled to room temperature, and annealing is performed by maintaining the sample at 95° C. for at least 60 minutes. In some embodiments, the melting point is determined by analyzing the data using SWIFT Tm software. In some embodiments of any part of the method of the present invention, the method experimentally measures or calculates (computer calculation) the melting point of at least 50, 80, 90, 92, 94, 96, 98, 99, or 100% of the primers in the library either before or after the primer is used for PCR amplification of the target locus.

[0106] In some embodiments, the library includes a microarray. In some embodiments, the library does not include a microarray.

[0107] In some embodiments, most or all of the primers are extended to form amplification products. When all primers are consumed in the PCR reaction, the same or approximately the same number of primer molecules are converted into target amplicons for each target locus, thus increasing the uniformity of amplification of various target loci. In some embodiments, at least 80, 90, 92, 94, 96, 98, 99, or 100% of the primer molecules are extended to form amplification products. In some embodiments, for at least 80, 90, 92, 94, 96, 98, 99, or 100% of the target loci, at least 80, 90, 92, 94, 96, 98, 99, or 100% of the primer molecules for that target locus are extended to form amplification products. In some embodiments, multiple cycles are performed until that percentage of the primers is consumed. In some embodiments, multiple cycles are performed until all or substantially all of the primers are consumed. If desired, a high percentage of the primers can be consumed by reducing the initial primer concentration and / or increasing the number of PCR cycles performed.

[0108] In some embodiments, the PCR method may be performed in a reaction volume in the microliter range, but compared to the nanoliter or picoliter reaction volumes used in microfluidic applications, it may be difficult to perform specific PCR amplification (due to the low local concentration of the template nucleic acid). In some embodiments, the reaction volume is from 1 L or more to 60 L or less, for example, from 5 L or more to 50 L or less, from 10 L or more to 50 L or less, from 10 L or more to 20 L or less, from 20 L or more to 30 L or less, from 30 L or more to 40 L or less, or from 40 L or more to 50 L or less.

[0109] In one embodiment, the method disclosed herein amplifies DNA using highly efficient and highly multiplexed targeted PCR, followed by high-throughput sequencing to determine allele frequencies at each target locus. The ability to multiplex about 50 or more than 100 PCR primers in a single reaction volume such that most of the resulting sequence reads map to the target locus is novel and non-obvious. One technique that enables highly multiplexed targeted PCR to be performed in a very efficient manner is to design primers that are unlikely to hybridize to each other. PCR probes, typically referred to as primers, generate a thermodynamic model for potentially harmful interactions between at least 300; at least 500; at least 750; at least 1,000; at least 2,000; at least 5,000; at least 7,500; at least 10,000; at least 20,000; at least 25,000; at least 30,000; at least 40,000; at least 50,000; at least 75,000; or at least 100,000 possible primer pairs, or unintended interactions between primers and sample DNA, and then the model is used to select by excluding designs that are not compatible with other designs in the pool. Another method that enables highly multiplexed targeted PCR to be performed in a very efficient manner is to use a partial or complete nested approach to targeted PCR. By using one or a combination of these methods, multiplexing of at least 300, at least 800, at least 1,200, at least 4,000, or at least 10,000 primers in a single pool is possible, and the resulting amplified DNA consists mostly of DNA molecules that map to the target locus when sequenced. By using one or a combination of these methods, multiplexing of the majority of primers in a single pool is possible, and the resulting amplified DNA consists of DNA molecules that map to the target locus in more than 50%, more than 60%, more than 67%, more than 80%, more than 90%, more than 95%, more than 96%, more than 97%, more than 98%, more than 99%, or more than 99.5%.

[0110] In some embodiments, detection of the target genetic material may be performed in a multiplexed manner. The number of target gene sequences that can be performed in parallel can range from 1 to 10, 10 to 100, 100 to 1000, 1000 to 10,000, 10,000 to 100,000, 100,000 to 1,000,000, or 1,000,000 to 10,000,000. Attempts to multiplex more than 100 primers per pool have resulted in significant problems such as unwanted side reactions such as the formation of primer dimers.

[0111] 5.5 Targeted PCR In some embodiments, PCR can be used to target specific locations in the genome. In plasma samples, the original DNA is highly fragmented (typically less than 500 bp, with an average length of less than 200 bp). In PCR, both the forward primer and the reverse primer can anneal to the same fragment and amplify. Thus, when the fragment is short, the PCR assay should amplify even in relatively short regions. When polymorphic positions are too close to the polymerase binding site, as in MIPS, bias can occur in amplification from various alleles. Currently, PCR primers targeting polymorphic regions, such as those containing SNPs, are typically designed such that the 3' end of the primer hybridizes to a base directly adjacent to the polymorphic base. In embodiments of the present disclosure, the 3' ends of both the forward PCR primer and the reverse PCR primer are designed to hybridize to bases at a position one or several bases away from the variant position (polymorphic site) of the target allele. The number of bases between the polymorphic site (SNP or another polymorphism) and the base to which the 3' end of the primer is designed to hybridize can be 1 base, 2 bases, 3 bases, 4 bases, 5 bases, 6 bases, 7 to 10 bases, 11 to 15 bases, or 16 to 20 bases. The forward primer and the reverse primer may be designed to hybridize at various numbers of bases away from the polymorphic site.

[0112] Multiple PCR assays can be performed, but due to interactions between different PCR assays, it is difficult to multiplex more than about 100 assays. Various complex molecular approaches can be used to increase the level of multiplexing, but still, it is limited to 100, perhaps 200, or in some cases less than 500 assays per reaction. Samples containing large amounts of DNA are divided into multiple sub-reactions and then re-mixed before sequencing. For samples where either the entire sample or a sub-population of a portion of the DNA molecules is limited, dividing the sample will introduce statistical noise. In one embodiment, a small or limited amount of DNA may refer to an amount less than 10 pg, 10 - 100 pg, 100 pg - 1 ng, 1 - 10 ng, or 10 - 100 ng. Other methods involve splitting into multiple pools, and this method is particularly useful in cases of small amounts of DNA where significant problems may arise related to the resulting stochastic noise. It should also be noted that this method also offers the advantage of minimizing bias when performed on samples of any DNA amount. Under these circumstances, a general pre-amplification step may be used to increase the overall sample amount. Ideally, this pre-amplification step should not significantly change the allele distribution.

[0113] In one embodiment, the disclosed method generates PCR products specific to a large number of target loci, specifically 1,000 to 5,000 loci, 5,000 to 10,000 loci, or more than 10,000 loci, and can perform genotyping by sequencing or some other genotyping method from a limited sample such as DNA from a single cell or body fluid. Currently, performing multiplex PCR reactions with more than 5 to 10 targets is a major challenge and is often hampered by primer by-products such as primer dimers and other artifacts. When using a microarray with hybridization probes to detect target sequences, primer dimers and other artifacts are not detected and can be ignored. However, when using sequencing as the detection method, most of the sequencing reads are not the desired target sequences in the sample, but rather sequence these artifacts. Prior art methods used to multiplex 50 or more reactions in a single reaction volume and then perform sequencing often result in off-target sequencing reads of more than 20%, more often more than 50%, in many cases more than 80%, and in some cases more than 90%.

[0114] Generally, to perform targeted sequencing of multiple (n) samples of more than 50, more than 100, more than 500, or more than 1,000 targets, the sample can be divided into a number of parallel reactions and each individual target amplified. This method is carried out in a PCR multi-well plate or can be performed on commercial platforms such as, for example, FLUIDIGM ACCESS ARRAY (48 reactions per sample in a microfluidic chip) or DROPLET PCR by RAIN DANCE TECHNOLOGY (hundreds to thousands of targets). Unfortunately, these partitioning and pooling methods are problematic with samples where the amount of DNA is limited, and often the genomic copy number is not sufficient to ensure that each genomic region is present at one copy per well. The stochastic noise generated by partitioning and pooling results in very inaccurate measurement results with respect to the allelic ratios present in the original DNA sample, which is a particularly serious problem when polymorphic loci are targeted and when the relative ratios of alleles at polymorphic loci are required. Described herein are methods for effectively and efficiently amplifying a number of PCR reactions applicable when only limited amounts of DNA are available. In one embodiment, the method may be applied to the analysis of DNA mixtures such as free floating DNA present in a single cell, body fluid such as plasma, biopsy, environmental and / or forensic samples.

[0115] In one embodiment, the targeted sequencing may involve one, multiple, or all of the following steps: a) generating and amplifying a library having adapter sequences at both ends of DNA fragments; b) splitting the library into multiple reactions after library amplification; c) generating a library having adapter sequences at both ends of DNA fragments and optionally amplifying it; d) performing 1000- to 10,000-multiplex amplification of selected targets using one target-specific "forward" primer and one tag-specific primer per target; e) performing a second amplification of the product using a "reverse" target-specific primer and one (or more) primer specific to the universal tag introduced as part of the target-specific forward primer in the first round; f) performing pre-amplification of 1000-multiplex of selected targets in a limited number of cycles; g) dividing the product into multiple aliquots and amplifying sub-pools of targets in individual reactions (e.g., 50- to 500-multiplex, but singleplex can also be used); h) pooling the products of parallel sub-pool reactions; i) during these amplifications, the primers carry sequence-compatible tags (partial or full-length), whereby the products can be sequenced.

[0116] 5.6 Highly Multiplexed PCR Disclosed herein are methods that enable targeted amplification of hundreds to tens of thousands of target sequences (e.g., SNP loci) from nucleic acid samples such as genomic DNA obtained from, for example, plasma. The amplified samples may relatively contain few primer dimer products and have low allelic bias at the target loci. If sequence-compatible adapters are added to the products during or after amplification, the analysis of these products can be performed by sequencing.

[0117] Even when highly multiplexed PCR amplification is performed using methods known in the art, primer dimer products that are not suitable for the sequence and exceed the desired amplification product are generated. The product can be experimentally reduced by removing the primers forming the product or by performing in-silico selection of the primers. However, the more assays there are, the more difficult this problem becomes.

[0118] One solution is to divide the 5000-multiplex reaction into several low-multiplex amplifications, such as 100 50-multiplex reactions or 50 100-multiplex reactions, or to divide the sample into individual PCR reactions. However, when the sample DNA is limited, such as in non-invasive prenatal diagnosis from, for example, pregnancy plasma, dividing the sample into multiple reactions causes a bottleneck and should be avoided.

[0119] In this specification, a method is described in which plasma DNA of a sample is first amplified extensively and then the sample is divided into a plurality of multiplexed target enrichment reactions having a moderately large number of target sequences per reaction. In one embodiment, the disclosed method can be used to preferentially enrich a DNA mixture at a plurality of loci. The method includes one or more of the following steps: generating and amplifying a library from the DNA mixture, wherein the molecules in the library have adapter sequences ligated to both ends of the DNA fragment; dividing the amplified library into a plurality of reactions; performing a first round of multiplexed amplification of selected targets using one or more target-specific "forward" primers per target and one or more adapter-specific universal "reverse" primers. In one embodiment, the method of the present disclosure further includes performing a second amplification using a "reverse" target-specific primer and one or more primers specific to a universal tag introduced as part of the target-specific forward primer in the first round. In one embodiment, the method may involve a nested, hemi-nested, semi-nested, one-sided nested, one-sided hemi-nested, or one-sided semi-nested PCR method. In one embodiment, the method of the present disclosure is used to preferentially enrich a DNA mixture at a plurality of loci, and the method includes amplifying the pre-multiplexed amplification of selected targets in a limited number of cycles, dividing the product into a plurality of aliquots, and amplifying sub-pools of the targets in individual reactions, and pooling the products of the parallel sub-pool reactions. It should be noted that targeted amplification can be performed using this method with low allelic bias for 50 to 500 loci, 500 to 5,000 loci, 5,000 to 50,000 loci, or even 50,000 to 500,000 loci. In one embodiment, the primers carry partial or full-length sequence-compatible tags.

[0120] The workflow may involve: (1) extracting DNA, such as plasma DNA for example; (2) preparing a fragment library having universal adapters at both ends of the fragments; (3) amplifying the library using universal primers specific to the adapters; (4) dividing the amplified sample “library” into a plurality of aliquots; (5) performing multiplexing (e.g., about 100-multiplexing, 1,000- or 10,000-multiplexing with one target-specific primer and one tag-specific primer per target) on the aliquots; (6) pooling aliquots of one sample; (7) barcoding the samples; (8) mixing the samples and adjusting the concentration; and (9) sequencing the samples. The workflow may include a plurality of sub-steps, and the sub-steps may include one of the steps listed (e.g., the library preparation step (2) may involve three enzyme steps (blunt end, dA tailing, and adapter ligation) and three purification steps). The steps of the workflow may be grouped, divided, or performed in a different order (e.g., barcoding and pooling of samples).

[0121] Note that library amplification can be performed in a biased manner such that shorter fragments are amplified more efficiently. In this way, it is possible to preferentially amplify short sequences such as mononucleosomal DNA fragments, such as cell-free fetal DNA (of placental origin) present in the circulation of pregnant women. Note that PCR assays can have tags such as sequence tags (usually truncated forms of 15-25 bases). After multiplexing, the PCR multiplexes of the samples are pooled, and then tagging is completed by tag-specific PCR (which can also be done by ligation) (including barcoding). Additionally, full-length sequence tags can be added in the same reaction as multiplexing. The target may be amplified using target-specific primers in the first cycle, and then the SQ adapter sequence is completed by continuing with tag-specific primers. The PCR primers do not have to carry tags. The sequence tags may be added to the amplification products by ligation.

[0122] In one embodiment, various applications such as the detection of fetal aneuploidy may be performed by performing clonal sequencing following high - throughput multiplex PCR to evaluate the amplification products. Conventional multiplex PCR evaluates a maximum of 50 loci simultaneously. On the other hand, by using the method described herein, it becomes possible to simultaneously evaluate more than 50 loci, more than 100 loci, more than 500 loci, more than 1,000 loci, more than 5,000 loci, more than 10,000 loci, more than 50,000 loci, and more than 100,000 loci simultaneously. Experiments have shown that it is possible to simultaneously evaluate up to 10,000 and more distinct loci in one reaction, and that it can have sufficient efficiency and specificity for non - invasive prenatal aneuploidy diagnosis and / or copy number determination with high accuracy. The assay may be combined in one reaction using the entire sample, such as a cfDNA sample isolated from plasma, a fraction thereof, or even a derivative of the processed cfDNA. The sample (e.g., cfDNA or derivative) may be divided into multiple parallel multiplexed reactions. Optimal sample splitting and multiplexing are determined by trading off various performance specifications. Since the amount of material is limited, splitting the sample into multiple fractions may introduce sampling noise, require time for processing, and increase the possibility of errors. Conversely, the higher the degree of multiplexing, the greater the amount of false amplification and the greater the amplification imbalance. Both can potentially reduce the test performance.

[0123] In the application of the methods described herein, there are two important related concerns: the limited amount of the original sample (e.g., plasma), and the limited number of original molecules in the substance for obtaining allele frequencies or other measurements. When the number of original molecules falls below a certain level, random sampling noise becomes prominent, which may affect the accuracy of the test. Typically, measurements are made on samples containing original molecules corresponding to 500 - 1000 per target locus, and sufficient quality data can be obtained for non-invasive prenatal aneuploidy diagnosis. For example, there are also many ways to increase the number of separate measurements, such as increasing the sample amount. Each operation performed on the sample may also result in loss of the substance. It is essential to characterize and avoid the losses caused by various operations, or improve the yield of specific operations as needed to avoid losses that may reduce the test performance.

[0124] In one embodiment, it is possible to reduce the potential for loss in subsequent steps by amplifying all or a fraction of the original sample (e.g., cfDNA sample). Various methods are available to amplify all of the genetic material in the sample and increase the amount available for downstream procedures. In one embodiment, DNA fragments of ligation-mediated PCR (LM-PCR) are amplified by PCR after ligation of either one separate adapter, two separate adapters, or many separate adapters. In one embodiment, phi-29 polymerase of multiple displacement amplification (MDA) is used to amplify all DNA isothermally. In DOP-PCR and its variants, random priming is used to amplify the original material DNA. Each method has characteristics such as the uniformity of amplification across all expressed regions of the genome, the efficiency of capture and amplification of the original DNA, and the amplification performance as a function of fragment length.

[0125] In one embodiment, LM-PCR may be used with a single heteroduplex adapter having a 3-prime tyrosine. The heteroduplex adapter allows for the use of a single adapter molecule that can be converted into two different sequences on the 5-prime and 3-prime ends of the original DNA fragment during the first round of PCR. In one embodiment, the amplified library can be separated by size, or by products such as AMPURE, TASS, or by other similar methods. Prior to ligation, the sample DNA may be blunt-ended and then one adenosine base is added to the 3-prime end. Prior to ligation, the DNA may be cleaved using a restriction enzyme or some other cleavage method. During ligation, the ligation efficiency is enhanced by the 3-prime adenosine of the sample fragment and the complementary 3-prime tyrosine overhang of the adapter. The elongation step of PCR amplification may be limited to reduce amplification from fragments longer than about 200bp, about 300bp, about 400bp, about 500bp, or about 1,000bp from a time perspective. A number of reactions were performed using the conditions specified by a commercially available kit. As a result, less than 10% of the sample DNA molecules were successfully ligated. Through a series of optimizations of the reaction conditions for this, the ligation was improved to approximately 70%.

[0126] In one embodiment, the methods described herein perform highly efficient and highly multiplexed targeted PCR of at least 1000 target loci and then perform next-generation sequencing to obtain genotype data. In some embodiments, the multiplexed targeted PCR is performed on at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 11000, 12000, 13000, 14000, 15000, 20000, 25000, 30000, 35000, 40000, 45000, or 50000 target loci. In one embodiment, the multiplexed targeted PCR is performed on 13000 target loci. In some embodiments, the target loci are single nucleotide polymorphisms (SNPs).

[0127] 5.7 MiniPCR The following MiniPCR method is desirable for samples containing short nucleic acids, digested nucleic acids, or fragmented nucleic acids such as, for example, cfDNA. The design of conventional PCR assays results in a significant loss of distinct fetal molecules, but this loss can be greatly reduced by designing a very short PCR assay called a MiniPCR assay. Fetal cfDNA in maternal serum is highly fragmented, with a fragment size that is on average 160 bp, standard deviation 15 bp, minimum size of approximately 100 bp, and maximum size of approximately 220 bp, and is approximately Gaussian distributed. With respect to the targeted polymorphisms, the distribution of the start and end positions of the fragments is not necessarily random, but varies greatly between individual targets and collectively between all targets, and the polymorphic site of a particular target locus can occupy any position from the start position to the end position between the various fragments derived from that locus. Note that the term MiniPCR may, without additional limitation or qualification, also refer equally to normal PCR.

[0128] During PCR, amplification occurs only from template DNA fragments that contain both the forward primer site and the reverse primer site. Because fetal cfDNA fragments are short, the likelihood that both primer sites are present is the likelihood of a fetal fragment of length L that contains both the forward primer site and the reverse primer site, which is the ratio of the amplicon length to the fragment length. Under ideal conditions, assays with amplicons that are 45, 50, 55, 60, 65, or 70 bp are successful in amplifying from 72%, 69%, 66%, 63%, 59%, or 56%, respectively, of the available template fragment molecules. The amplicon length is the distance between the 5-prime ends of the forward and reverse priming sites. Shorter amplicon lengths than are typically used by those of ordinary skill in the art can result in more efficient measurement of desired polymorphic loci by requiring only short sequence reads. In one embodiment, a substantial proportion of the amplicons are less than 100 bp, less than 90 bp, less than 80 bp, less than 70 bp, less than 65 bp, less than 60 bp, less than 55 bp, less than 50 bp, or less than 45 bp.

[0129] In prior art methods, short assays such as those described herein are typically avoided. This is because they are not necessary and because significant constraints are placed on primer design due to the limited length of the primers, annealing characteristics, and distance between the forward and reverse primers.

[0130] Also, it should be noted that when the 3-prime end of any of the primers is within approximately 1 to 6 bases of the polymorphic site, there is a possibility of biased amplification. This single-base difference at the initial polymerase binding site can cause preferential amplification of one allele, resulting in a change in the measured allele frequency and a possible decrease in performance. All of these constraints make it very difficult to identify primers that successfully amplify a particular locus, and even more difficult to design a large primer set that is compatible in the same multiplex reaction. In one embodiment, the 3' ends of the inner forward and reverse primers are designed to hybridize to a DNA region upstream from the polymorphic site and are separated from the polymorphic site by a small number of bases. Ideally, the number of bases may be 6 to 10 bases, but may also be 4 to 15 bases, 3 to 20 bases, 2 to 30 bases, or 1 to 60 bases, and substantially the same result can be achieved.

[0131] Multiplex PCR may involve one round of PCR in which all targets are amplified, or it may involve one round of PCR followed by one or more rounds of nested PCR or some modified version of nested PCR. Nested PCR consists of subsequent rounds of PCR amplification using one or more new primers that bind internally to the primers used in the previous round by at least one base pair. Nested PCR reduces the number of false amplification targets by amplifying only the amplification products from previous amplifications that have the correct internal sequence in subsequent reactions. The reduction of false amplification targets improves the number of useful measurements obtainable, especially in sequences. Nested PCR typically involves designing primers that are completely internal to the previous primer binding sites, which necessarily increases the minimum DNA segment size required for amplification. For samples such as plasma cfDNA where the DNA is highly fragmented, as the assay size increases, the number of individual cfDNA molecules from which measurements can be obtained decreases. In one embodiment, to counteract this effect, a partial nested method may be used in which one or both of the second round primers overlap the first binding site by several bases and are extended internally to be more specific while the overall size of the assay increases minimally.

[0132] In one embodiment, the multiplexing pool of PCR assays is designed to amplify potential heterozygous SNPs or other polymorphic or non-polymorphic loci on one or more chromosomes, and these assays are used in one reaction to amplify DNA. The number of PCR assays can be 50-200 PCR assays, 200-1,000 PCR assays, 1,000-5,000 PCR assays, or 5,000-20,000 PCR assays (50-200-multiplexing, 200-1,000-multiplexing, 1,000-5,000-multiplexing, 5,000-20,000-multiplexing, over 20,000-multiplexing), respectively. In one embodiment, a multiplexing pool of about 10,000 PCR assays (10,000-multiplexing) is designed to amplify potential heterozygous SNP loci on the X chromosome, Y chromosome, chromosome 13, chromosome 18, and chromosome 21, and chromosome 1 or 2, and these assays are used in one reaction to amplify cfDNA obtained from maternal plasma, chorionic villus samples, amniocentesis samples, one or a few cells, other body fluids or tissues, cancer, or other genetic materials. The SNP frequency of each locus may be determined by a clonal method or any other method of amplicon sequencing. Statistical analysis of allele frequency distribution, or the ratio of all assays, may be used to determine whether the sample contains trisomy of one or more of the chromosomes included in the test. In another embodiment, the original cfDNA sample is split into two samples and a parallel 5,000-multiplexing assay is performed. In another embodiment, the original cfDNA sample is split into n samples and a parallel (about 10,000 / n) multiplexing assay is performed. In this case, n is 2-12, or 12-24, or 24-48, or 48-96. The data is collected and analyzed in a manner similar to the methods already described. Note that this method is equally well applicable to the detection of translocations, deletions, duplications, and other chromosomal abnormalities.

[0133] In one embodiment, a tail with no homology to the target genome may be added to either the 3-prime or 5-prime end of any of the primers. These tails facilitate subsequent operations, procedures, or measurements. In one embodiment, the tail sequence may be the same as the forward or reverse target-specific primer. In one embodiment, different tail sequences may be used for the forward or reverse target-specific primers. In one embodiment, multiple different tails may be used for different loci, or sets of loci. A particular tail may be common to all loci, or common to a subset of loci. For example, by using forward and reverse tails corresponding to the forward and reverse primers required by any current sequencing platform, direct sequencing after amplification becomes possible. In one embodiment, the tail can be used as a priming site common to all amplified targets and can be used to add other useful sequences. In some embodiments, the inner primer may contain a region designed to hybridize either upstream or downstream of the target locus (e.g., a polymorphic locus). In some embodiments, the primer may contain a molecular barcode. In some embodiments, the primer may contain a universal priming sequence designed to enable PCR amplification.

[0134] In one embodiment, a 10,000-multiplex PCR assay pool is created such that the forward and reverse primers have tails corresponding to the essential forward and reverse sequences required by a high-throughput sequencing device (often also referred to as a massively parallel sequencing device), such as HISEQ, GAIIX, or MYSEQ, available from ILLUMINA, for example. Further, the 5-prime contained for the sequence tail is an additional sequence that can be used as a priming site in subsequent PCR to add a nucleotide barcode sequence to the amplicon, enabling multiplex sequencing of multiple samples in one lane of a high-throughput sequencing device.

[0135] In one embodiment, the 10,000-multiplex PCR assay pool is generated such that the reverse primer has a tail corresponding to the essential reverse sequence required for the high-throughput sequencing device. After amplification in the first 10,000-multiplex assay, subsequent PCR amplification may be performed using another 10,000-multiplex pool having forward primers that are partially nested (e.g., 6-base nested) for all targets, and reverse primers corresponding to the reverse sequence tails included in the first round. By performing subsequent rounds of partially nested amplification using only one target-specific primer and a universal primer, the required assay size is limited, sample noise is reduced, and the number of false amplicons is also significantly reduced. The sequence tag may be added to the attached ligation adapter and / or as part of the PCR probe, such that the tag becomes part of the final amplicon.

[0136] The tumor fraction affects the test performance. There are numerous methods for increasing the tumor fraction of DNA present in patient plasma. The tumor fraction can be increased by the aforementioned LM-PCR method that has already been studied, as well as by targeted removal of long fragments. In one embodiment, prior to performing multiplex PCR amplification of target loci, an additional multiplex PCR reaction may be carried out to selectively remove long parental fragments corresponding to the loci targeted in the subsequent multiplex PCR. The additional primers are designed to anneal at sites that are at a distance longer than the distance from polymorphisms expected to be present in cell-free fetal DNA fragments. These primers are used in a single cycle of multiplex PCR reaction, and then multiplex PCR of the target polymorphic loci may be performed. These distal primers are tagged with a molecule or moiety that enables selective recognition of the tagged fragments of DNA. In one embodiment, these DNA molecules may be covalently modified with biotin molecules after one cycle of PCR, thereby enabling removal of the newly formed double-stranded DNA containing these primers. The double-stranded DNA formed during the first round is likely of maternal origin. Removal of the hybrid material may be achieved by using magnetic streptavidin beads. There are other tagging methods that function similarly. In one embodiment, a size selection method may be used to enrich samples of shorter DNA strands, such as strands less than about 800 bp, less than about 500 bp, or less than about 300 bp. Amplification of the short fragments can then proceed as normal.

[0137] The miniPCR method described in this disclosure enables highly multiplexed amplification and analysis of hundreds to thousands, and even millions of loci, in one reaction, from one sample. At the same time, the detection of amplified DNA can be multiplexed. Dozens to hundreds of samples can be multiplexed in one sequencing lane using barcoded PCR. This multiplexed detection has been successfully tested up to 49-multiplex, and higher multiplexing is also possible. In fact, this enables genotyping of hundreds of samples at thousands of SNPs in one sequencing run. For these samples, the method can determine genotype and heterozygosity rate, and at the same time determine copy number, and both can be used for the purpose of aneuploidy detection. It may also be used as part of a method for determining the amount of mutation. The method may be used for any amount of DNA or RNA, and the target regions may be SNPs, other polymorphic regions, non-polymorphic regions, and combinations thereof.

[0138] In some embodiments, ligation-mediated universal PCR amplification of fragmented DNA may be used. Plasma DNA can be amplified using ligation-mediated universal PCR amplification and then divided into multiple parallel reactions. Furthermore, this can be used to preferentially amplify short fragments, resulting in an increased tumor fraction. In some embodiments, adding tags to the fragments by ligation enables detection of shorter fragments, use of shorter target sequence-specific primer portions, and / or annealing at high temperatures to reduce non-specific reactions.

[0139] The methods described herein may be used for a number of purposes when there is a set of target DNA mixed with an amount of contaminating DNA. In some embodiments, the target DNA and the contaminating DNA may be from genetically related individuals. For example, fetal (target) genetic abnormalities may be detected from maternal plasma that also contains fetal (target) DNA and maternal (contaminating) DNA. Such abnormalities include whole chromosome abnormalities (e.g., aneuploidy), partial chromosome abnormalities (e.g., deletions, duplications, inversions, translocations), polynucleotide polymorphisms (e.g., STRs), single nucleotide polymorphisms, and / or other genetic abnormalities or differences. In some embodiments, the target DNA and the contaminating DNA may be from the same individual, but the target DNA and the contaminating DNA may differ by one or more mutations, such as in the case of cancer. (See, e.g., H. Mamon et al. Preferential Amplification of Apoptotic DNA from Plasma: Potential for Enhancing Detection of Minor DNA Alterations in Circulating DNA. Clinical Chemistry 54:9 (2008)). In some embodiments, the DNA may be found in cell culture (apoptosis) supernatants. In some embodiments, apoptosis can be induced in a biological sample (e.g., blood), followed by library preparation, amplification, and / or sequencing. A number of workflows and protocols for achieving this purpose are presented elsewhere in this disclosure.

[0140] In some embodiments, the target DNA may be derived from a single cell, may be from a DNA sample consisting of less than 1 copy of the target genome, may be from a small amount of DNA, may be from DNA of mixed origin (e.g., plasma and tumor of a cancer patient, mixture of healthy DNA and cancerous DNA, transplantation, etc.), may be from other body fluids, may be from cell cultures, may be from culture supernatants, may be from forensic DNA samples, may be from ancient DNA samples (e.g., insects trapped in amber), may be from other DNA samples, and may be combinations thereof.

[0141] In some embodiments, short amplicon sizes may be used. Short amplicon sizes are particularly suitable for fragmented DNA (see, e.g., A. Sikora, et al. Detection of increased amounts of cell-free fetal DNA with short PCR amplicons. Clin Chem. 2010 Jan;56(1):136-8).

[0142] Using short amplicon sizes can provide several significant benefits. Short amplicon sizes can result in optimized amplification efficiency. Short amplicon sizes typically produce short products. Therefore, the possibility of non-specific priming is low. Due to the smaller clusters, short products can cluster at higher density on the sequencing flow cell. Note that the methods described herein can function as well as longer PCR amplicons. The length of the amplicon may be increased as needed, for example, when sequencing longer sequences. Experiments using 146-multiplexed target amplification with an assay of length from 100 bp to 200 bp as the first step of an nested PCR protocol were performed on a single cell and on genomic DNA, and positive results were obtained.

[0143] In some embodiments, the methods described herein may be used to amplify and / or detect SNPs, copy numbers, nucleotide methylation, mRNA levels, other types of RNA expression levels, other genetic characteristics, and / or epigenetic characteristics. The miniPCR methods described herein may be used with next-generation sequencing and may also be used with other downstream methods such as, for example, microarrays, counting by digital PCR, real-time PCR, mass spectrometry, and the like.

[0144] In some embodiments, the miniPCR amplification methods described herein may be used as part of a method for accurate quantification of a minority population. It may be used for absolute quantification using spike calibrators. It may be used for quantification of mutations / minor alleles via very deep sequencing and may be performed in a highly multiplexed manner. It may be used for standard paternity testing and may also be used for identity testing of relatives or ancestors in humans, animals, plants, or other organisms. It may be used for forensic testing. For example, it may be used for rapid genotyping and copy number analysis (CN) of any type of material such as amniotic fluid and CVS, sperm, products of conception (POC), and the like. It may be used for single-cell analysis such as, for example, genotyping of samples biopsied from embryos. It may be used for rapid embryo analysis (within 1 day, 1 day, or 2 days from biopsy) by targeted sequencing using miniPCR.

[0145] In some embodiments, the miniPCR amplification method can be used for tumor analysis. Tumor biopsies are often a mixture of healthy and tumor cells. Targeted PCR enables deep sequencing of SNPs and loci with little to no background sequence. It may be used for copy number analysis of tumor DNA and analysis of loss of heterozygosity. Such tumor DNA can be present in many different body fluids or tissues of tumor patients. It may be used for detection of tumor recurrence and / or tumor screening. It may be used for quality control testing of seeds. It may be used for breeding purposes or fishery purposes. Note that any of these methods can also be suitably used to target non-polymorphic loci for the purpose of ploidy classification.

[0146] Some of the documents that describe some of the basic methods underlying the methods disclosed herein include the following: (1) Wang HY, Luo M, Tereshchenko IV, Frikker DM, Cui X, Li JY, Hu G, Chu Y, Azaro MA, Lin Y, Shen L, Yang Q, Kambouris ME, Gao R, Shih W, Li H. Genome Res. 2005 Feb;15(2):276-83. Department of Molecular Genetics, Microbiology and Immunology / The Cancer Institute of New Jersey, Robert Wood Johnson Medical School, New Brunswick, New Jersey 08903, USA. (2) High-throughput genotyping of single nucleotide polymorphisms with high sensitivity. Li H, Wang HY, Cui X, Luo M, Hu G, Greenawalt DM, Tereshchenko IV, Li JY, Chu Y, Gao R. Methods Mol Biol. 2007;396 - PubMed PMID:18025699. (3) A method comprising multiplexing of an average of 9 assays for sequencing is described in: Nested Patch PCR enables highly multiplexed mutation discovery in candidate genes. Varley KE, Mitra RD. Genome Res. 2008 Nov;18(11):1844-50. Epub 2008 Oct 10. Note that the methods disclosed herein enable a higher degree of multiplexing than the above references.

[0147] 6. A method for analyzing genomic information from circulating fetal cells for non-invasive prenatal testing. In one aspect, the method further includes performing a non-invasive prenatal test using a genotype determined to be one or more fetal cells in one or more circulating cells.

[0148] In one embodiment, the method further includes detecting a copy number variation or aneuploidy of a target chromosome or a target chromosomal segment in one or more circulating cells determined to be one or more fetal cells.

[0149] Ploidy classification, also referred to as "chromosome copy number classification" or "copy number calling" (CNC), refers to the determination of the amount and identity of one or more chromosomes present in a cell. Aneuploidy refers to a state in which an incorrect number of chromosomes are present in a cell. In the case of human somatic cells, it refers to the case where a cell does not contain 22 pairs of autosomes and one pair of sex chromosomes. In the case of human gametes, it refers to the case where a cell does not contain one of each of the 23 chromosomes. In the case of one chromosome, it refers to the case where there are more than two or less than two homologous but not identical chromosomes, and each of the two chromosomes originates from a different parent. The ploidy state refers to the amount and identity of one or more chromosomes in a cell.

[0150] CNVs are often assigned to one of two major categories based on the length of the affected sequence. The first category includes copy number polymorphisms (CNPs), which are common to the general population and occur with an overall frequency higher than 1%. Typically, CNPs are small (most are less than 10 kilobases in length) and are often enriched for genes encoding proteins important for drug detoxification and immunity. A subset of these CNPs is highly variable with respect to copy number. As a result, different human chromosomes can have a wide range of copy numbers (e.g., 2, 3, 4, 5, etc.) for a particular set of genes. More recently, CNPs associated with immune response genes have been associated with susceptibility to complex genetic diseases including psoriasis, Crohn's disease, and glomerulonephritis.

[0151] The second type of CNV includes relatively rare variants that are much longer than CNP, ranging in size from hundreds of thousands to over a million base pairs. In some cases, these CNVs may have occurred during the production of the sperm or egg that gave rise to a particular individual, or may have been inherited only within several generations of a family. These large and rare structural variants have been observed disproportionately in subjects with mental retardation, developmental delay, schizophrenia, and autism. The occurrence in such subjects has led to the speculation that large and rare CNVs may be more important in neurocognitive diseases than other types of genetic variants including single nucleotide substitutions.

[0152] Gene copy number can vary in cancer cells. For example, duplication of Chrlp is common in breast cancer, and the copy number of EGFR can be higher than normal in non-small cell lung cancer. Further target genes for cancer diagnosis are provided in Section 9 of this specification. Cancer is one of the major causes of death. Therefore, early diagnosis and treatment of cancer are important and can improve the patient's outcome (e.g., increase the probability of remission and the duration of remission). Also, with early diagnosis, patients can receive fewer radical treatment options. Many of the current treatments that destroy cancerous cells also affect normal cells, resulting in various possible side effects such as nausea, vomiting, decreased blood cell count, increased risk of infection, hair loss, and mucosal ulcers. Therefore, early detection of cancer is desirable, which can reduce the amount and / or frequency of treatment (e.g., chemotherapy agents or radiation) necessary to remove the cancer.

[0153] Copy number variations are also associated with severe mental and physical disorders, as well as idiopathic learning disabilities. Non-invasive prenatal testing (NIPT) using cell-free DNA (cfDNA) can be used to detect abnormalities such as fetal trisomies 13, 18, and 21, triploidy, and sex chromosome aneuploidy. Thus, in one embodiment, the target chromosome or target chromosome segment of interest is chromosome 13, chromosome 18, chromosome 21, the sex chromosomes, and / or segments of those chromosomes.

[0154] In addition, subchromosomal microdeletions that can cause severe mental and physical disabilities are more difficult to detect due to their small size. Eight of the microdeletion syndromes have an overall incidence higher than 1 in 1000, which is approximately the same frequency as fetal autosomal trisomy. Thus, in one embodiment, the method disclosed herein further includes detecting microdeletions in one or more circulating cells determined to be one or more pure fetal cells. In one embodiment, the microdeletions are 22q11.2 deletions associated with DiGeorge syndrome, microdeletions associated with Prader-Willi syndrome, microdeletions associated with Angelman syndrome, 1p36 deletions, and / or microdeletions associated with cri-du-chat syndrome.

[0155] Furthermore, a higher copy number of CCL3L1 is associated with lower susceptibility to HIV infection, and a lower copy number of FCGR3B (CD16 cell surface immunoglobulin receptor) may increase susceptibility to systemic lupus erythematosus and similar inflammatory autoimmune disorders.

[0156] A method of measuring the chromosomal copy number in fetal cells based on counting the number of reads based on a DNA sequence mapped to a given chromosome or chromosomal segment is conveniently referred to as a "counting method" or "quantitative method" for analyzing chromosomal copy number or chromosomal segment copy number. Examples of such methods can be found, in particular, in published patent application US2013 / 0172211A1, U.S. Patent No. 8,008,018, U.S. Patent No. 8,467,976B2, and published U.S. patent application US2012 / 0003637A1. In many cases, such methods involve generating a reference value (cutoff value) for the number of reads of a DNA sequence mapped to a specific chromosome, and a number of reads exceeding the value indicates a specific genetic abnormality.

[0157] Accordingly, a method for analyzing and characterizing a cell sample suspected of being of fetal origin may be combined with a method and system for determining a copy number, or a method and system for detecting aneuploidy of a target chromosome or chromosomal segment in a cell that has been confirmed to be a fetal cell. These methods for determining a copy number or detecting aneuploidy are performed using the target chromosome or chromosomal segment and establish a bias model. That is, to set the test parameters, in the same parallel analysis, a sample identified as a diploid sample with high reliability is used in the analysis to analyze the aneuploidy of the same target chromosome or chromosomal segment in other samples in the set of test samples. Reliability refers to the statistical likelihood that the classified SNPs, alleles, sets of alleles, ploidy classification, or determined copy number of a chromosomal segment correctly represents the actual genetic state of an individual. Such methods for determining copy number variations and aneuploidy from fetal DNA are described in U.S. Patent Publication 2018 / 0173846, which is incorporated herein by reference in its entirety.

[0158] The amount of each locus detected by sequencing DNA preparations obtained from target and non-target cells can vary from locus to locus due to reasons other than the starting amount of the locus in the initial sample material prior to preparation for sequencing, such as before performing an amplification step such as PCR. For example, variables such as PCR primer binding efficiency, amplicon length, and GC content can cause variation in the appearance of individual loci during preparation for sequencing or during sequencing. Such factors can introduce locus-specific biases and cause over- or under-representation of individual loci. Additionally, biases can arise from sample-specific discrepancies. For example, due to pipetting errors or other measurement errors during physical processing of the sample, one sample may have more DNA than another sample in the sample set. In an exemplary embodiment, these sample-specific biases are accounted for by sample-specific parameters. These specific embodiments regarding sample-specific parameters are disclosed in U.S. Patent Publication 2018 / 0173846, which is hereby incorporated by reference in its entirety.

[0159] Regardless of the particular method used to generate genetic information, the amount of genetic sequence information from each locus depends on the relative amount of the copy number of the locus in the original sample. Loci that are thought to be on the same chromosomal segment, or in some embodiments, loci that are thought to be on the same chromosome, are presumed to have the same starting amount. Thus, for example, multiple loci present on chromosome 21 of the genome of a target cell (or the genome of a non-target cell) are presumed to be present in approximately equal amounts in genomic DNA. Thus, differences in the amount of genetic information measured between loci on the same chromosome are the result of locus-specific bias. For example, if SNP1 and SNP2 are located on the same chromosome and are assumed to have the same copy number on the same chromosome, and it is found that SNP1 has a read depth of 0.1% and SNP2 has a read depth of 0.4%, this can be explained by a quantitative locus-specific bias that is more favorable for the production of DNA sequences from SNP2 than from SNP1. This bias can be further normalized by considering the distribution of possible sampling results for two different SNPs. Thus, the method of the present invention analyzes bias and provides a bias model, as discussed in more detail in U.S. Patent Publication 2018 / 0173846, which is incorporated herein by reference in its entirety.

[0160] In certain embodiments of the present invention, one or two maximum likelihood methods are used. As described above, the first maximum likelihood method can be used to identify diploid samples in a sample set and to determine a first probability that the other samples in the sample set are aneuploid. Thus, in certain embodiments, one or more or all of the chromosomes or chromosomal segments of interest are determined to be diploid using the first maximum likelihood method. The method is further described and illustrated in U.S. Patent Publication 2018 / 0173846, which is incorporated herein by reference in its entirety.

[0161] In another embodiment, the copy number of a target chromosome or chromosomal segment is determined by generating a plurality of second multiple hypotheses or a set of second multiple hypotheses, where each of the second multiple hypotheses is associated with a particular copy number of the target chromosome or chromosomal segment in the target cell. A model is then used to test how well the genetic data from each patient fits each of the second hypotheses. A goodness of fit for each of the second hypotheses is determined. A second probability value is calculated for each of the second hypotheses, where the second probability value indicates the likelihood that the genome of the target cell has the number of chromosomes or chromosomal segments specified by the second hypothesis. Thus, by selecting the second hypothesis with the maximum likelihood, the copy number of the chromosome or chromosomal segment in the genome of the target cell can be determined. Such first and second hypotheses can be considered in combination to increase the confidence in aneuploidy determination, as detailed in U.S. Patent Publication 2018 / 0173846, which is incorporated herein by reference in its entirety.

[0162] In one embodiment of the disclosure regarding methods used to determine the ploidy state of a fetus, the method further includes taking into account the proportion of fetal DNA in the sample. In one embodiment of the disclosure, the method involves calculating the percentage of DNA in a sample that is of fetal or placental origin. In one embodiment of the disclosure, the threshold for aneuploidy classification is adaptively normalized based on the calculated fetal DNA percentage. In some embodiments, methods for estimating the percentage of DNA of fetal origin in a DNA mixture include obtaining a mixed sample containing maternal and fetal genetic material, obtaining a genetic sample from the father of the fetus, measuring the DNA in the mixed sample, measuring the DNA in the father's sample, and using the DNA measurement results of the mixed sample and the father's sample to calculate the percentage of DNA of fetal origin in the mixed sample.

[0163] In one aspect, the method described herein enables confirmation that a sample is a mixture of fetal and maternal cells when the cell sample contains both true fetal cells and true maternal cells, and the fetal fraction can be estimated by using cell-free DNA (cfDNA) as a reference. As described in U.S. Patent Application No. 2018 / 0298439, which is incorporated herein by reference in its entirety, determination of fetal paternity is possible based on the genotype data of maternal DNA, the mixture of fetal and maternal DNA, and the estimated fraction of fetal DNA. Using the genotype data obtained from the mixture of maternal and fetal cfDNA and the mixture of fetal and maternal DNA, the fetal fraction of the sample can be determined, enabling paternity testing.

[0164] Some embodiments may be used in combination with the PARENTAL SUPPORT (trademark) (PS) method, which is described in U.S. Patent Application 11 / 603,406 (U.S. Publication 20070184467), U.S. Patent Application 12 / 076,348 (U.S. Publication 20080243398), U.S. Patent Application 13 / 110,685, PCT Application PCT / US09 / 52730 (PCT Publication: WO / 2010 / 017214), PCT Application PCT / US10 / 050824 (PCT Publication: WO / 2011 / 041485), PCT Application PCT / US2011 / 037018 (PCT Publication: WO / 2011 / 146632), and PCT Application PCT / US2011 / 61506, which are hereby incorporated by reference in their entirety. PARENTAL SUPPORT (trademark) is an information science-based method that can be used to analyze genetic data. In some embodiments, the methods disclosed herein may be considered as part of the PARENTAL SUPPORT (trademark) method. In some embodiments, the PARENTAL SUPPORT (trademark) method is for determining genetic data of a target individual, one or a few cells derived from the individual, or a DNA mixture consisting of the DNA of the target individual, or the DNA of one or more other individuals, particularly for determining disease-related alleles, determining alleles of interest, determining the ploidy state of one or more chromosomes in the target individual, and / or determining the degree of relatedness between another individual and the target individual. PARENTAL SUPPORT (trademark) may refer to any of these methods. PARENTAL SUPPORT (trademark) is an example of an information science-based method.

[0165] In some embodiments of the present invention, an improvement in the reliability regarding aneuploidy determination can be obtained by determining the aneuploidy of a sample using a quantitative non-allelic threshold or cutoff method, and can also be obtained by determining the aneuploidy of the same sample using a method for determining a likelihood. If a sample is identified as having aneuploidy in a target chromosome or chromosomal segment by a threshold method, and the sample is identified as having aneuploidy with high confidence using likelihood determination for a hypothesis set, the sample is identified as a sample having aneuploidy in the target chromosome or chromosomal segment with respect to one or more target cells in a subject that is the source of the sample. The method is further described and illustrated in U.S. Patent Publication 2018 / 0173846, which is hereby incorporated by reference in its entirety.

[0166] Methods for performing non-allelic threshold analysis, such as NIPT in particular, are known in the art. For example, U.S. Pat. Nos. 7,888,017 and 8,318,430, which are hereby incorporated by reference in their entireties, count the number of reads mapped to a putative chromosome and compare this to the number of reads mapped to a reference chromosome, and if there are excess reads on the putative chromosome, provide a method for determining fetal aneuploidy using the assumption that this corresponds to triploidy of the fetus in that chromosome. The teachings in this document, including the depth of sequence reads and reference values, may be useful in the practice of embodiments of the present invention.

[0167] 7. A method for analyzing genomic information from circulating tumor cells. The methods and compositions provided herein improve the detection, diagnosis, staging, screening, treatment, and management of cancer (e.g., breast cancer, bladder cancer, or colorectal cancer) by obtaining a sample containing circulating tumor cells using a liquid biopsy. The methods provided herein, in illustrative embodiments, analyze cancer-associated mutations in a circulating fluid, particularly in circulating tumor cells (CTCs). The method provides the advantage of identifying most of the variants found in a tumor, not only clonal mutations but also subclonal mutations, in a single test that utilizes a tumor sample, if available, rather than multiple tests that may be required.

[0168] Accordingly, in one aspect, the present disclosure provides a method for monitoring and detecting early recurrence or metastasis of a tumor in a cancer patient by obtaining one or more circulating cells determined to be of tumor origin, the method comprising: a) selecting one or more patient-specific mutations based on mutations identified in a tumor sample of a patient diagnosed with cancer; b) longitudinally collecting one or more blood samples from the patient after the patient has been treated with surgery, first-line chemotherapy, and / or adjuvant therapy; c) generating a first amplicon set from cell DNA isolated from one or more circulating cells presumed to be circulating tumor cells obtained from the blood sample obtained from the cancer patient, and a second amplicon set from cell-free DNA, wherein the tumor cells and normal cells are obtained from each of the blood sample or a fraction thereof of the cancer patient, and the first amplicon set and the second amplicon set are obtained by performing a multiplex amplification reaction of patient-specific mutations associated with cancer; d) sequencing the first amplicon set and the second amplicon set by next-generation sequencing; e) a step of determining the origin of one or more circulating cells based on the sequence of the first amplicon set, wherein the sequence of the second amplicon set is used as a reference, and the detection of one or more patient-specific mutations from the first amplicon set generated from the circulating cells determined to be of tumor origin suggests early recurrence or metastasis of cancer.

[0169] In another aspect, the disclosure herein provides a method for treating a cancer patient, the method comprising: a) treating the cancer patient with surgery, first-line chemotherapy, and / or adjuvant therapy; b) longitudinally collecting one or more blood samples from the patient after the patient has been treated with surgery, first-line chemotherapy, and / or adjuvant therapy; c) generating a first amplicon set from cell DNA isolated from one or more circulating cells presumed to be circulating tumor cells obtained from the blood sample obtained from the cancer patient, and a second amplicon set from cell-free DNA, wherein the tumor cells and normal cells are obtained from each of the blood sample or a fraction thereof of the cancer patient, and the first amplicon set and the second amplicon set are obtained by performing a multiplex amplification reaction of patient-specific mutations associated with cancer; d) sequencing the first amplicon set and the second amplicon set by next-generation sequencing; e) a step of determining the origin of one or more circulating cells based on the sequence of the first amplicon set, wherein the sequence of the second amplicon set is used as a reference, and the detection of one or more patient-specific mutations from the first amplicon set generated from the circulating cells determined to be of tumor origin suggests early recurrence or metastasis of cancer; f) administering a compound to the patient, wherein the compound has been found to be effective in treating cancer having the one or more patient-specific mutations detected from the blood sample.

[0170] The improved method provided herein enables accurate analysis and characterization of samples containing a very small number of CTCs. The number of CTCs in a sample analyzed by the method of the present invention can be less than 100 cells, less than 75 cells, less than 50 cells, less than 25 cells, less than 20 cells, less than 15 cells, less than 10 cells, less than 5 cells, or one cell. By using the method provided herein, accurate genomic information can be obtained from one circulating tumor cell, 2 CTCs, 3 CTCs, 4 CTCs, 5 CTCs, 6 CTCs, 7 CTCs, 8 CTCs, 9 CTCs, or 10 CTCs.

[0171] In some embodiments, the method includes monitoring and detecting early recurrence or metastasis of a tumor in a cancer patient by determining the identity of circulating cells that are suspected to be circulating tumor cells based on pre-determined patient-specific mutations. In some embodiments, the patient-specific mutations include single nucleotide variants (SNVs), copy number variants (CNVs), indels, deletions, or gene fusions associated with cancer. In some embodiments, the presence of at least two patient-specific mutations associated with cancer suggests that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least eight patient-specific mutations associated with cancer suggests that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least sixteen patient-specific mutations associated with cancer suggests that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least fifty patient-specific mutations associated with cancer suggests that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least one hundred patient-specific mutations associated with cancer suggests that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least one thousand patient-specific mutations associated with cancer suggests that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least five thousand patient-specific mutations associated with cancer suggests that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least ten thousand patient-specific mutations associated with cancer suggests that the one or more circulating cells are tumor cells. In some embodiments, the presence of at least fifteen thousand patient-specific mutations associated with cancer suggests that the one or more circulating cells are tumor cells.

[0172] In some embodiments, when there are at least 100 to about 1,000 patient-specific mutations associated with cancer, it is suggested that the one or more circulating cells are tumor cells. In some embodiments, when there are at least 100 to about 1,000 patient-specific mutations associated with cancer, it is suggested that the one or more circulating cells are tumor cells. In some embodiments, when there are at least 1,000 to about 5,000 patient-specific mutations associated with cancer, it is suggested that the one or more circulating cells are tumor cells. In some embodiments, when there are at least 5,000 to about 10,000 patient-specific mutations associated with cancer, it is suggested that the one or more circulating cells are tumor cells. In some embodiments, when there are at least 10,000 to about 15,000 patient-specific mutations associated with cancer, it is suggested that the one or more circulating cells are tumor cells. In some embodiments, when there are at least 15,000 to about 20,000 patient-specific mutations associated with cancer, it is suggested that the one or more circulating cells are tumor cells. In some embodiments, when there are at least 20,000 to about 25,000 patient-specific mutations associated with cancer, it is suggested that the one or more circulating cells are tumor cells. In some embodiments, when there are at least 50,000 patient-specific mutations associated with cancer, it is suggested that the one or more circulating cells are tumor cells.

[0173] In some embodiments, when there are at least 5,000 patient-specific mutations associated with cancer, it is suggested that the one or more circulating cells are tumor cells. In some embodiments, when there are at least 10,000 patient-specific mutations associated with cancer, it is suggested that the one or more circulating cells are tumor cells. In some embodiments, when there are at least 15,000 patient-specific mutations associated with cancer, it is suggested that the one or more circulating cells are tumor cells. In some embodiments, when there are at least 25,000 patient-specific mutations associated with cancer, it is suggested that the one or more circulating cells are tumor cells.

[0174] In one aspect, multiplex amplification is performed to obtain amplicons that include patient-specific mutations associated with cancer. In some embodiments, the first amplicon set and the second amplicon set are obtained by performing a multiplex amplification reaction of 50 to 1000 cancer-associated patient-specific mutations.

[0175] In one aspect, a method for monitoring and detecting early recurrence and metastasis of tumors in cancer patients includes separating a plurality of circulating cells into individual reaction volumes, isolating cellular DNA in each reaction volume, and attaching a sample barcode to the cellular DNA isolated in each reaction volume. In some embodiments, the cellular DNA is isolated from one circulating cell that is presumed to be a tumor cell.

[0176] In some embodiments, the method further includes performing a non-invasive cancer test using a genotype determined that one or more circulating cells are one or more tumor cells.

[0177] The method and composition may be useful by themselves or when used with other methods for the detection, diagnosis, staging, screening, treatment, and management of cancer (e.g., breast cancer, bladder cancer, or colorectal cancer), for example, to assist the results of these other methods to provide reliable and / or final results.

[0178] The terms “cancer” and “cancerous” refer to, or describe, a physiological state in an animal that is typically characterized by unregulated cell growth. A “tumor” includes one or more cancerous cells. There are several main types of cancer. Carcinomas are cancers that begin in the skin or in tissues that line or cover internal organs. Sarcomas are cancers that begin in bone, cartilage, fat, muscle, blood vessels, or other connective or supportive tissue. Leukemias are cancers that begin in blood-forming tissues such as the bone marrow and produce large numbers of abnormal blood cells that enter the bloodstream. Lymphomas and multiple myelomas are cancers that begin in cells of the immune system. Cancers of the central nervous system are cancers that begin in the tissues of the brain and spinal cord.

[0179] In some embodiments, the cancer is acute lymphoblastic leukemia, acute myeloid leukemia, adrenocortical carcinoma, AIDS-related cancer, AIDS-related lymphoma, anal cancer, appendiceal cancer, astrocytoma, atypical teratoid / rhabdoid tumor, basal cell carcinoma, bladder cancer, brainstem glioma, brain tumor (brainstem glioma, atypical teratoid / rhabdoid tumor of the central nervous system, germinoma of the central nervous system, astrocytoma, craniopharyngioma, ependymoblastoma, ependymoma, medulloblastoma, medulloepithelioma, pineal parenchymal tumor of intermediate differentiation, supratentorial primitive neuroectodermal tumor, and pineoblastoma), breast cancer, bronchial tumor, Burkitt lymphoma, cancer of unknown primary site, carcinoid tumor, carcinoma of unknown primary site, atypical teratoid / rhabdoid tumor of the central nervous system, germinoma of the central nervous system, cervical cancer, childhood cancer, chordoma, chronic lymphocytic leukemia, chronic myeloid leukemia, chronic myeloproliferative disorder, colon cancer, colorectal cancer, craniopharyngioma, cutaneous T-cell lymphoma, endocrine pancreatic islet cell tumor, endometrial cancer, ependymoblastoma, ependymoma, esophageal cancer, nasal neuroblastoma, Ewing sarcoma, extracranial germ cell tumor, extragonadal germ cell tumor, extrahepatic bile duct cancer, gallbladder cancer, gastric cancer, gastrointestinal carcinoid tumor, gastrointestinal stromal cell tumor, gastrointestinal stromal tumor (GIST), gestational trophoblastic tumor, glioma, hairy cell leukemia, head and neck cancer, heart cancer, Hodgkin lymphoma, hypopharyngeal cancer, intraocular melanoma, islet cell tumor, Kaposi sarcoma, kidney cancer, Langerhans cell histiocytosis, laryngeal cancer, lip cancer, liver cancer, malignant fibrous histiocytoma, bone cancer, medulloblastoma, medulloepithelioma, melanoma, Merkel cell carcinoma, Merkel cell skin cancer, mesothelioma, metastatic squamous neck cancer of unknown primary, oral cancer, multiple endocrine neoplasia syndrome, multiple melanoma, multiple melanoma / plasmacytoma, mycosis fungoides, myelodysplastic syndrome, myeloproliferative neoplasm, nasal cancer, nasopharyngeal cancer, neuroblastoma, non-Hodgkin lymphoma, non-melanoma skin cancer, non-small cell lung cancer, oral cancer, oral cavitycancer, oral pharyngeal cancer, osteosarcoma, other brain tumors and spinal cord tumors, ovarian cancer, epithelial ovarian cancer, ovarian germ cell tumors, low-grade ovarian tumors, pancreatic cancer, papillomatosis, paranasal sinus cancer, parathyroid cancer, pelvic cancer, penile cancer, pharyngeal cancer, intermediate-differentiated pineal parenchymal tumors, pineoblastoma, pituitary tumors, plasmacytoma / multiple myeloma, pleuropulmonary blastoma, primary central nervous system (CNS) lymphoma, primary hepatocellular carcinoma, prostate cancer, rectal cancer, renal cancer, renal cell (kidney) cancer, renal cell carcinoma, airway cancer, retinoblastoma, rhabdomyosarcoma, salivary gland cancer, Sézary syndrome, small cell lung cancer, small intestine cancer, soft tissue sarcoma, squamous cell carcinoma, squamous cell carcinoma of the neck, gastric cancer, supratentorial primitive neuroectodermal tumors, T cell lymphoma, testicular cancer, laryngeal cancer, thymic cancer, thymoma, thyroid cancer, transitional cell carcinoma, transitional cell carcinoma of the renal pelvis and ureter, gestational trophoblastic tumors, ureteral cancer, urethral cancer, uterine cancer, uterine sarcoma, vaginal cancer, vulvar cancer, Waldenström macroglobulinemia, or Wilms tumor.

[0180] In other embodiments, provided herein is a method for detecting cancer (e.g., breast cancer, bladder cancer, or colorectal cancer) in a blood sample or a fraction thereof from an individual, such as an individual suspected of having cancer. The method includes determining single nucleotide variants present in the sample by using the ctDNA SNV amplification / sequencing workflow provided herein to determine single nucleotide variants present in a sample containing circulating tumor cells and a ctDNA sample. The presence of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 SNVs at a plurality of single nucleotide loci in the sample, at the lower limit of the range, and 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 40, or 50 SNVs at the upper limit of the range indicates the presence of cancer (e.g., breast cancer, bladder cancer, or colorectal cancer).

[0181] In certain examples of the embodiment, the cancer is breast cancer, bladder cancer, or colorectal cancer at stage 1a, 1b, or 2a. In certain examples of the embodiment, the cancer is breast cancer, bladder cancer, or colorectal cancer at stage 1a or 1b. In certain examples of the embodiment, the individual has not undergone surgery. In certain examples of the embodiment, the individual has not undergone a biopsy. In certain embodiments of the methods provided herein, prior to performing target amplification on CTCs from an individual, data is provided regarding the SNVs present in the tumor from which the individual is derived. Thus, in these embodiments, the SNV amplification / sequencing reaction is performed on one or more tumor samples from the individual. In this method, the CTC SNV amplification / sequencing reaction provided herein is also advantageous because it provides a liquid biopsy of clonal and subclonal mutations. Also, as provided herein, in an individual having cancer (e.g., breast cancer, bladder cancer, or colorectal cancer), for a certain SNV, if a high percentage VAF of, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10% VAF or more is determined in the ctDNA sample from the individual, it can be more clearly identified as a clonal mutation.

[0182] In illustrative embodiments, a set of single nucleotide variant loci of any of the methods herein includes all of the single nucleotide variant loci identified in the TCGA and COSMIC datasets for cancer (e.g., breast cancer, bladder cancer, or colorectal cancer).

[0183] In certain embodiments of any of the methods herein, a set of single nucleotide variant loci includes from a lower limit of 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 75, 100, 250, 500, 1000, 2500, 5000, or 10,000 to an upper limit of 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 75, 100, 250, 500, 1000, 2500, 5000, 10,000, 20,000, and 25,000 single nucleotide variant loci that have been found to be associated with cancer (e.g., breast cancer, bladder cancer, or colorectal cancer).

[0184] In other embodiments, provided herein is a method for corroborating a diagnosis of cancer (e.g., breast cancer, bladder cancer, or colorectal cancer) in an individual, such as an individual suspected of having cancer (e.g., breast cancer, bladder cancer, or colorectal cancer), from a sample of blood or a fraction thereof from the individual, the method comprising performing the CTC SNV amplification / sequencing workflow provided herein to determine whether one or more single nucleotide variants are present in a plurality of single nucleotide variant loci. In such embodiments, the following factors, observations, guidelines, or rules apply: the absence of a single nucleotide variant corroborates a diagnosis of adenocarcinoma at stage 1a, 1b, or 2a; the presence of a single nucleotide variant corroborates a diagnosis of squamous cell carcinoma or adenocarcinoma at stage 2b or 3a; and / or the presence of 10 or more single nucleotide variants corroborates a diagnosis of squamous cell carcinoma or adenocarcinoma at stage 2b or 3.

[0185] In certain embodiments, the method further includes determining a treatment plan, a treatment method, and / or administering to the individual a compound that targets one or more clonal single nucleotide variants. In certain examples, subclonal SNVs and / or other clonal SNVs are not targeted by the treatment method. Certain treatment methods and related mutations are provided in other parts of this specification and are known in the art. Thus, in certain examples, the method further includes administering to the individual a compound that is known to be particularly effective in treating cancer (such as breast cancer, bladder cancer, or colorectal cancer) having one or more of the determined single nucleotide variants.

[0186] In certain embodiments, it is possible to perform a treatment regimen using the method for detecting SNVs herein. Treatment methods targeting specific mutations associated with ADC and SCC are available and in development (Nature Review Cancer. 14:535-551 (2014)). For example, the detection of EGFR mutations at L858R or T790M may be beneficial for treatment selection. Erlotinib, gefitinib, afatinib, AZK9291, CO-1686, and HM61713 are currently approved treatments in the United States or in clinical trials, and these target specific EGFR mutations. In other examples, mutations at G12D, G12C, or G12V in KRAS can be utilized to treat an individual with a combination of selumetinib and docetaxel. As another example, mutations at V600E in BRAF can be utilized to treat a subject with vemurafenib, dabrafenib, and trametinib.

[0187] The target gene of the present invention is a cancer-related gene in exemplary embodiments and is a cancer-related gene in many illustrative embodiments. A cancer-related gene (e.g., a cancer-related gene or a bladder cancer-related gene or a colorectal cancer-related gene) refers to a gene related to a change in the risk of cancer (e.g., breast cancer, bladder cancer or colorectal cancer), or a gene related to a change in the prognosis of cancer. Exemplary cancer-related genes that promote cancer include oncogenes, genes that enhance cell proliferation, wetting or metastasis, genes that suppress apoptosis, and pro-angiogenesis genes. Cancer-related genes that suppress cancer include, but are not limited to, tumor suppressor genes, genes that suppress cell proliferation, wetting or metastasis, genes that promote apoptosis, and anti-angiogenesis genes.

[0188] In one embodiment of the mutation detection method, it starts with selecting a region of the target gene. Using a region containing known mutations, primers for mPCR-NGS are grown, amplified, and the mutations are detected.

[0189] Using the methods provided herein, virtually any type of mutation can be detected, particularly those mutations that have been found to be associated with cancer. In particular, the methods provided herein are directed to mutations such as SNVs that are associated with cancer, particularly breast cancer, bladder cancer, or colorectal cancer. Exemplary SNVs can be present in one or more of the following genes: EGFR, FGFR1, FGFR2, ALK, MET, ROS1, NTRK1, RET, HER2, DDR2, PDGFRA, KRAS, NF1, BRAF, PIK3CA, MEK1, NOTCH1, MLL2, EZH2, TET2, DNMT3A, SOX2, MYC, KEAP1, CDKN2A, NRG1, TP53, LKB1, and PTEN, as identified as being mutated, having increased copy number, or being fused to other genes, and combinations thereof, in various lung cancer samples (Non-small-cell lung cancers: a heterogeneous set of diseases. Chen et al. Nat. Rev. Cancer. 2014 Aug 14(8):535-551). In other embodiments, the list of genes is as listed above, and the SNVs are reported in the cited reference of Chen et al. and the like.

[0190] 8. Exemplary embodiments of the analysis method 8.1 Exemplary embodiments for detecting cancer In certain embodiments, the analysis step in the method for determining whether circulating tumor nucleic acids are present comprises analyzing a set of chromosomal segments known to exhibit aneuploidy in cancer. In certain embodiments, the analysis step in the method for determining whether circulating tumor nucleic acids are present comprises analyzing polymorphic loci for ploidy in the range of 1,000 to 50,000, or multiples of 100 to 1,000. In certain embodiments, the analysis step in the method for determining whether circulating tumor nucleic acids are present comprises analyzing 100 to 1,000 single nucleotide variant sites. For example, in these embodiments, the analysis step comprises performing multiplex PCR to amplify amplicons across 1,000 to 50,000 polymerase loci and 100 to 1,000 single nucleotide variant sites. This multiplex reaction can be set up as a single reaction or as a pool of different subsets of multiplex reactions. Multiplex reactions provided herein, such as the large-scale multiplex PCR disclosed herein, provide exemplary processes for performing amplification reactions that improve multiplexing and thereby assist in achieving improved sensitivity levels.

[0191] In certain embodiments, the multiplex PCR reaction is performed under limiting primer conditions for at least 10%, 20%, 25%, 50%, 75%, 90%, 95%, 98%, 99%, or 100% of the reaction. Improved conditions for performing the large-scale multiplex reactions provided herein can be used.

[0192] In certain aspects, the above-described methods for determining whether circulating tumor nucleic acids are present in a sample in an individual, and all embodiments thereof, can be implemented using a system. The present disclosure provides teachings regarding specific functional and structural characteristics for implementing the methods. By way of non-limiting example, the system includes:

[0193] An input processor configured to analyze data from a sample to determine ploidy at a set of polymorphic loci on chromosomal segments in an individual; and

[0194] A model configured to determine the level of allelic imbalance present at polymorphic loci based on ploidy determination, wherein an allelic imbalance of 0.5% or more in this case indicates the presence of circulation.

[0195] 8.2 Exemplary embodiments for detecting single nucleotide variants In certain aspects, provided herein is a method for detecting single nucleotide variants. The improved methods provided herein can achieve detection limits of 0.015, 0.017, 0.02, 0.05, 0.1, 0.2, 0.3, 0.4, or 0.5 percent of SNVs in a sample. All embodiments for detecting SNVs can be implemented using a system. The present disclosure provides teachings regarding specific functional and structural characteristics for implementing the methods. Further provided herein are embodiments including a non-transitory computer-readable medium including computer-readable code that, when executed by a processing device, causes the processing device to execute the SNV detection methods provided herein.

[0196] Accordingly, in one embodiment, provided herein is a method for determining whether a single nucleotide variant is present in a set of genomic positions in a sample from an individual, the method comprising, for each genomic position, generating an estimate of the efficiency of an amplicon spanning the genomic position and the error rate per cycle, using a training dataset, receiving nucleotide-specific information measured for each genomic position in the sample, determining a set of probabilities of the percentage of single nucleotide variants arising from one or more true variants at each genomic position using the nucleotide-specific information measured at each genomic position and a model of the percentages of various variants using the amplification efficiency and error rate per cycle estimated independently for each genomic position, and determining the percentage and confidence level of the most likely true variant from the set of probabilities for each genomic position.

[0197] In an exemplary embodiment of a method for determining whether a single nucleotide variant is present, an estimate of efficiency and error rate per cycle is generated for an amplicon set spanning genomic positions. For example, an amplicon set may contain 2, 3, 4, 5, 10, 15, 20, 25, 50, 100, or more amplicons spanning genomic positions.

[0198] In an exemplary embodiment of a method for determining whether a single nucleotide variant is present, the measured nucleotide-specific information includes the total number of reads measured for each genomic position and the number of reads of variant alleles measured for each genomic position.

[0199] In an exemplary embodiment of a method for determining whether a single nucleotide variant is present, the sample is a plasma sample and the single nucleotide variant is present in the circulating tumor DNA of the sample. In an exemplary embodiment of a method for determining whether a single nucleotide variant is present, the sample is a plasma sample and the single nucleotide variant is present in the circulating fetal DNA of the sample.

[0200] In another embodiment, provided herein is a method for estimating the percentage of single nucleotide variants present in a sample derived from an individual. The method includes the following steps: generating an estimate of efficiency and error rate per cycle for one or more amplicons spanning a genomic position in a genomic position set; using a training data set; receiving measured nucleotide-specific information for each genomic position in the sample; using the amplification efficiency and error rate per cycle of the amplicon to generate the total number of molecules, background error molecules, and estimated mean and variance of true variant molecules for a search space including an initial percentage of true variant molecules; and determining the percentage of the most likely true single nucleotide variants present in the sample resulting from true mutations by fitting a distribution using the estimated mean and variance to the measured nucleotide-specific information in the sample.

[0201] The training data set for this embodiment of the present invention preferably includes samples from a population of healthy individuals. In certain exemplary embodiments, the training data set is analyzed on the same day as, and even in the same run as, one or more samples during the test. For example, a training data set can be generated using samples from 2, 3, 4, 5, 10, 15, 20, 25, 30, 36, 48, 96, 100, 192, 200, 250, 500, 1000 or more healthy individuals. If more healthy individuals, e.g., 96 or more data, are available, the confidence in the estimation of amplification efficiency increases even if a run was performed before applying the method to the samples during the test. Since the PCR error rate is the error rate per amplicon, nucleic acid sequence information can be used not only for the position of the SNV base but also for the overall amplification region around the SNV. For example, by using samples from 50 individuals and sequencing a 20 base pair amplicon around the SNV, the error frequency rate can be determined using error frequency data from 1000 base reads.

[0202] Typically, the amplification efficiency is estimated by estimating the mean and standard deviation of the amplification efficiency for the amplified segments and then fitting it to a distribution model such as, for example, a binomial distribution or a beta-binomial distribution. The error rate is determined for a PCR reaction for which the number of cycles is known, and then the error rate per cycle is estimated.

[0203] In certain exemplary embodiments, the estimation of the starting molecules of the test data set further includes updating the estimated value of the efficiency for the test data set using the number of starting molecules estimated in step (b) when the measured number of reads is significantly different from the estimated number of reads. Thereafter, the estimation can be updated for the new efficiency and / or starting molecules.

[0204] The search space used to estimate the total number of molecules, background error molecules, and true variant molecules can include search spaces from 0.1%, 0.2%, 0.25%, 0.5%, 1%, 2.5%, 5%, 10%, 15%, 20%, or 25% for the lower limit of the copy number of the base at the SNV position, which is the SNV base, and from 1%, 2%, 2.5%, 5%, 10%, 12.5%, 15%, 20%, 25%, 50%, 75%, 90%, or 95% for the upper limit. 0.1%, 0.2%, 0.25%, 0.5%, or 1% for the lower limit in the low range and 1%, 2%, 2.5%, 5%, 10%, 12.5%, or 15% for the upper limit in the high range can be used in exemplary embodiments of plasma samples where the method is detecting circulating tumor DNA. The high range is used for tumor samples.

[0205] The distribution is fitted to the total number of error molecules (background error and true variants) in the total molecules, and the likelihood or probability for each possible true variant in the search space is calculated. This distribution can be a binomial distribution or a beta-binomial distribution.

[0206] The most likely true variant is determined by determining the percentage of the most likely true variant and calculating the confidence using data from the distribution fit. As an exemplary example, and not a limitation on the clinical interpretation of the methods provided herein, when the average mutation rate is high, the confidence required to make a positive determination for an SNV is low. For example, if the average mutation rate of an SNV in a sample using the most likely hypothesis is 5% and the confidence is 99%, a positive SNV classification is made. On the other hand, for this exemplary example, if the average mutation rate of an SNV in a sample using the most likely hypothesis is 1% and the confidence is 50%, in certain situations, a positive SNV classification is not made. It should be understood that the clinical interpretation of the data is a function of sensitivity, specificity, prevalence, and the availability of alternatives.

[0207] In another embodiment, provided herein is a method for detecting one or more single nucleotide variants in a test sample from an individual. The method according to this embodiment includes the following steps:

[0208] Determine the median of the variant allele frequencies of a plurality of control samples from each of a plurality of normal individuals for each position of each position base variant in a set of positions of single nucleotide variants based on the results generated in a sequence run, identify selected single nucleotide variants having a median of allele variant frequencies in normal samples below a threshold value, and after removing outlier samples for each of the single nucleotide variant positions, determine the background error for each of the single nucleotide variant positions; determine, for a test sample, the mean and variance weighted by the measured read depth for the selected single nucleotide variant positions based on the data generated in the sequence run of the test sample; and using a computer, identify one or more single nucleotide variant positions with a statistically significant read depth weighted mean compared to the background error at that position, thereby detecting one or more single nucleotide variants.

[0209] In certain embodiments of the method for detecting one or more SNVs, the plurality of control samples include at least 25 samples. In certain exemplary embodiments, the plurality of control samples are at least 5, 10, 15, 20, 25, 50, 75, 100, 200, or 250 samples for the lower limit, and 10, 15, 20, 25, 50, 75, 100, 200, 250, 500, and 1000 samples for the upper limit.

[0210] In certain embodiments of the method for detecting one or more SNVs, outliers are removed from the data generated by high-throughput sequencing, and the mean weighted by the measured read depth and the measured variance are determined. In certain embodiments of the method for detecting one or more SNVs, the read depth for each of the single nucleotide variant positions of the test sample is at least 100 reads.

[0211] In certain embodiments of the method for detecting one or more SNVs, the sequence run includes a multiplex amplification reaction performed under limited primer reaction conditions. These embodiments in the exemplary embodiments are implemented using the improved method for performing the multiplex amplification reaction provided herein.

[0212] Without being bound by theory, the method of the present invention utilizes a background error model using a normal plasma sample, which is sequenced in the same sequence run as the test sample and accounts for run-specific artifacts. For example, noise positions with median normal variant allele frequencies exceeding thresholds such as >0.1%, 0.2%, 0.25%, 0.5%, 0.75%, and 1.0% are removed.

[0213] Outlier samples are repeatedly removed from the model explaining noise and contamination. For each base substitution at all genomic loci, the mean weighted by read depth and the standard deviation of the error are calculated. In certain exemplary embodiments, samples such as, for example, circulating fetal cells, circulating tumor cells or cell-free plasma samples, including at least a threshold number of reads such as at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 250, 500, or 1000 variant reads, and in certain embodiments having a single nucleotide variant position with an a1 Z-score higher than 2.5, 5, 7.5, or 10 relative to the background error model, are counted as candidate mutations.

[0214] In certain embodiments, for the lower limit of the range, a read depth higher than 100, 250, 500, 1,000, 2,000, 2,500, 5,000, 10,000, 20,000, 25,000, 50,000, or 100,000, and for the upper limit of the range, a read depth higher than 2,000, 2,500, 5,000, 7,500, 10,000, 25,000, 50,000, 100,000, 250,000, or 500,000 is obtained in the sequence run for each single nucleotide variant position in the set of single nucleotide variant positions. Typically, the sequence run is a run of high-throughput sequencing. In an exemplary embodiment, the mean or median value generated for the test sample is weighted by the read depth. Thus, in a sample having one variant allele detected at 1,000 reads, the likelihood that the determination of the variant allele is true is weighted higher than in a sample having one variant allele detected at 10,000 reads. Since the determination of the variant allele (i.e., the mutation) is not made with 100% confidence, the identified single nucleotide variant can be considered a candidate variant or candidate mutation.

[0215] 8.3 Exemplary test statistics for the analysis of phased data Exemplary test statistics for the analysis of phased data of a sample that has been determined or is suspected to be a mixed sample containing DNA or RNA originating from two or more genetically non-identical cells are described below. f represents the proportion of the DNA or RNA of interest, such as the proportion of DNA or RNA with the CNV of interest, or the proportion of DNA or RNA from the cells of interest such as cancer cells. In some embodiments of cancer testing, f represents the proportion of DNA or RNA derived from cancer cells in a mixture of cancer cells and normal cells, or f represents the proportion of cancer cells in a mixture of cancer cells and normal cells. Note that this refers to the proportion of DNA derived from the cells of interest, assuming that 2 copies of DNA are provided from each of the cells of interest. This is different from the proportion of DNA derived from the cells of interest in segments that are deleted or duplicated.

[0216] The possible alleles of each SNP are designated as A and B. All possible allele pairs in all possible orders are shown using AA, AB, BA, and BB. In some embodiments, SNPs having alleles in the order of AB or BA are analyzed. N i represents the number of sequence reads of the i-th SNP, and A i and B i represent the number of reads of the i-th SNP that indicate allele A and allele B, respectively. The following are assumed:

[0217] N i = A i + B i .

[0218] The allele ratio R i is defined as:

[0219]

Equation

[0220] T represents the number of SNPs targeted.

[0221] Without loss of generality, some embodiments focus on one chromosomal segment. As a further clarity issue, the phrase "a first homologous chromosomal segment compared to a second homologous chromosomal segment" in the specification means the first homolog of the chromosomal segment and the second homolog of the chromosomal segment. In some embodiments, all of the target SNPs are contained in the chromosomal segment of interest. In other embodiments, multiple chromosomal segments are analyzed for the possibility of copy number variation.

[0222] MAP Estimation

[0223] This method utilizes the knowledge of phase determination through ordered alleles to detect deletions or duplications in the target segment. For each SNP i, define

[0224]

Number

[0225] Next, define

[0226]

Number

[0227] The distributions of X and S under various copy number hypotheses (e.g., hypotheses regarding deletion or duplication of two chromosomes, the first or second homolog) are described below. i and S are described below.

[0228] Two - chromosome hypothesis

[0229] Under the hypothesis that the target segment is not deleted or duplicated,

[0230]

[0231]

Number

[0232] where

[0233]

Number

[0234] Assuming that the depth of read N is constant, a binomial distribution S with the following parameters is given

[0235]

Number

[0236] Deletion hypothesis

[0237] Under the hypothesis that the first homolog is deleted (i.e., the AB SNP becomes B and the BA SNP becomes A), R i has parameters

Number

Number

[0238]

Number

[0239] assuming that the depth of read N is constant, a binomial distribution S with the following parameters is given

[0240]

Number

[0241] Under the hypothesis that the second homolog is deleted (i.e., the AB SNP becomes A and the BA SNP becomes B), R i has parameters

Number

Number

[0242]

Number

[0243] assuming that the depth of read N is constant, a binomial distribution S with the following parameters is given

[0244] [Numerical]

[0245] Duplication hypothesis

[0246] Under the hypothesis that the first homolog is duplicated (i.e., the AB SNP becomes AAB and the BA SNP becomes BBA), R i has parameters for the AB SNP [Numerical] as well as parameters for the BA SNP [Numerical] and has a binomial distribution. Thus,

[0247] [Numerical]

[0248] Assuming that the depth of read N is constant, a binomial distribution S with the following parameters is given

[0249] [Numerical]

[0250] Under the hypothesis that the second homolog is duplicated (i.e., the AB SNP becomes ABB and the BA SNP becomes BAA), R i has parameters for the AB SNP [Numerical] as well as parameters for the BA SNP [Numerical] and has a binomial distribution. Thus,

[0251]

Number

[0252] Assuming that the depth of lead N is constant, a binomial distribution S with the following parameters is given

[0253]

Number

[0254] Classification

[0255] As shown in the above section, X i is a binary random variable.

[0256]

Number

[0257] Thus, the probability of the test statistic S under each hypothesis can be calculated. The probability of each hypothesis given the measurement data can be calculated. In some embodiments, the hypothesis with the highest probability is selected. If desired, the distribution of S can be simplified by approximating each N i at a constant lead depth N or by truncating the lead depth to a constant N. With this simplification, it is given.

[0258]

Number

[0259] The value of f can be estimated by selecting the most likely value of f that gives the measurement data, e.g., the value of f that results in the best data fit using an algorithm (e.g., a search algorithm) such as maximum likelihood estimation, maximum a posteriori probability estimation, or Bayesian estimation. In some embodiments, multiple chromosomal segments are analyzed and the value of f is estimated based on the data for each segment. If all of the target cells have these duplications or deletions, the estimated values of f based on the data for these various segments will be similar. In some embodiments, f is measured experimentally by determining the proportion of DNA or RNA from cancer cells, e.g., based on differences in methylation (hypomethylation or hypermethylation) between cancerous and non-cancerous DNA or RNA.

[0260] Rejection of a single hypothesis

[0261] The distribution of S for the two-chromosome hypothesis is independent of f. Thus, the probability of the measured data can be calculated for the two-chromosome hypothesis without calculating f. A test for single hypothesis rejection can be used for the two-chromosome null hypothesis. In some embodiments, the probability of S under the two-chromosome hypothesis is calculated and if the probability is below a given threshold (e.g., less than 1 in 1,000), the two-chromosome hypothesis is rejected. This indicates the presence of duplications or deletions of chromosomal segments. If desired, the false positive rate can be varied by adjusting the threshold.

[0262] 8.4 Exemplary methods for analysis of phase-determined data An exemplary method for analyzing data from a sample that has been determined or is suspected to be a mixed sample containing DNA or RNA originating from two or more genetically non-identical cells is described below. In some embodiments, phased data is used. In some embodiments, the method involves determining, for each calculated allele ratio, whether the calculated allele ratio is above or below a predicted allele ratio and the degree of difference for a particular locus. In some embodiments, for a particular hypothesis, a likelihood distribution of allele ratios at a locus is determined, and the closer the calculated allele ratio is to the center of the likelihood distribution, the more likely the hypothesis is correct. In some embodiments, the method involves determining, for each locus, the likelihood that the hypothesis is correct. In some embodiments, the method includes determining, for each locus, the likelihood that the hypothesis is correct and combining the probabilities of the hypothesis for each locus, and the hypothesis with the highest combined probability is selected. In some embodiments, the method involves determining, for each locus, the likelihood that the hypothesis is correct and determining each possible ratio of the DNA or RNA of one or more target cells to the total DNA or total RNA in the sample. In some embodiments, the combined probability for each hypothesis is determined by combining the probability of the hypothesis for each locus with each possible ratio, and the hypothesis with the highest combined probability is selected.

[0263] In one embodiment, the following hypotheses are considered: H 11 (All cells are normal), H 10 (Presence of cells having only homolog 1 and lacking homolog 2), H 01 (Presence of cells having only homolog 2 and lacking homolog 1), H 21 (Presence of cells having a duplication of homolog 1), H 12 (Presence of cells having a duplication of homolog 2). For the proportion f of target cells (or the proportion of DNA or RNA from target cells) such as cancer cells or mosaic cells, the predicted allele ratios for heterozygous (AB or BA) SNPs can be found as follows:

[0264] Equation (1):

[0265]

Number

[0266] Correction of bias, contamination, and sequence error:

[0267] Observed D at SNP S is the number of original mapped reads in which each allele is present, n A 0 and n B 0 consists of. Next, using the predicted bias in the amplification of alleles A and B, the corrected reads n A and n B can be found.

[0268] c a indicates atmospheric contamination (e.g., contamination from DNA in the air or environment), and r(c a ) indicates the allele ratio for atmospheric contamination (initially set to 0.5). Further, c g indicates the contamination rate during genotyping (e.g., contamination from another sample, etc.), and r(c g ) is the allele ratio for the contaminant. s e (A,B) and s e (B,A) indicate sequence errors that classify one allele as a different allele (e.g., misdetecting an A allele when a B allele is present, etc.).

[0269] By correcting for atmospheric contamination, contamination during genotyping, and sequence errors, the observed allele ratio q(r, c a , r(c a ), c g , r(c g ), s e (A,B), s e (B,A)) for a given predicted ratio r can be found.

[0270] Since the genotype of the admixture is unknown, the population frequency can be used to find P(r(c g ))). More specifically, let p be the population frequency of one of the alleles (which may also be referred to as the reference allele). Next, P(r(c g ) = 0) = (1 - p) 2 , P(r(c g ) = 0) = 2p(1 - p) 、 and P(r(cg) = 0) = p 2 be assumed. Using the conditional expectation value for r(c g ), E[q(r, c a , r(c a ), c g , r(c g ), s e (A,B), s e (B,A))] can be determined. It should be noted that since the atmospheric admixture and the admixture during genotyping are determined using homozygous SNPs, they are not affected by the presence or absence of deletions or duplications. Additionally, if desired, it is also possible to measure the atmospheric admixture and the admixture during genotyping using the reference chromosome.

[0271] Likelihood of each SNP:

[0272] The following equation gives the probability of observing n A and n B when the allele ratio is r:

[0273] Equation (2):

[0274]

Number

[0275] D s represents the data of SNP s. For each hypothesis h ε {H 11 , H 01 , H 10 , H 21 , H 12}, in Equation (1), r = r(AB, h) or r = r(BA, h) can be assumed, and r(cg )Find the conditional expectation for and determine E[q(r, c a , r(c a ), c g , r(c g ))] of the measured allele ratio. Then, in Equation (2), set r = E[q(r, c a , r(c a ), c g , r(c g ), s e (A,B), s e (B,A))] to determine P(D s |h,f).

[0276] Search algorithm:

[0277] In some embodiments, SNPs having allele ratios that appear to be outliers are ignored (e.g., SNPs having allele ratios that are at least 2 or 3 standard deviations above or below the mean are ignored or excluded). It should be noted that the advantage identified in this way is that when the mosaicism rate is high, the variability of the allele ratio may also be high, and thus SNPs are not excluded due to mosaicism.

[0278] F = {f1, …., f N} represents the search space for the mosaicism rate (e.g., the rate of tumor or fetus). P(D s |h,f) can be determined for each SNP s and f ε F, and the inductions of all SNPs can be combined.

[0279] The algorithm explores each f for each hypothesis. Using the search method, if there exists a range F* of f for which the confidence of the deletion hypothesis or the duplication hypothesis is higher than the confidence of the non - deletion hypothesis or the non - duplication hypothesis, it is concluded that mosaicism exists. In some embodiments, the maximum likelihood estimate for P(D s |h,f) in F* is determined. Optionally, the conditional expectation for f ε F* may be determined. Optionally, the confidence of each hypothesis can be determined.

[0280] In some embodiments, a beta-binomial distribution is used instead of a binomial distribution. In some embodiments, a sample-specific parameter of the beta-binomial distribution is determined using a reference chromosome or a reference chromosomal segment.

[0281] 8.5 Exemplary methods for detecting deletions and duplications without using phased data In some embodiments, unphased genetic data is used to determine whether there is an overrepresentation of the copy number of a first homologous chromosomal segment compared to a second homologous chromosomal segment in an individual's genome (e.g., the genome in one or more cells, or cfDNA or cfRNA). In some embodiments, phased genetic data is used, but the phase is ignored. In some embodiments, the DNA or RNA sample is a mixed sample of cfDNA or cfRNA from an individual, including cfRNA or cfRNA derived from two or more genetically different cells. In some embodiments, the method utilizes the magnitude of the difference between the calculated allele ratio and the predicted allele ratio for each locus.

[0282] In some embodiments, the method involves obtaining genetic data at a set of polymorphic loci on a chromosome or chromosomal segment in a sample of DNA or RNA derived from one or more cells from an individual by measuring the amount of each allele at each locus. In some embodiments, the allele ratio is calculated for loci that are heterozygous in at least one cell from which the sample was derived. In some embodiments, the allele ratio calculated for a particular locus is the measured amount of one of the alleles divided by the total amount calculated for all alleles at that locus. In some embodiments, the calculated allele ratio for a particular locus is the measured amount of one of the alleles (e.g., the allele on the first homologous chromosomal segment) divided by the measured amount of one or more other alleles (e.g., the allele on the second homologous chromosomal segment) at that locus. The calculated allele ratios and predicted allele ratios may be calculated using any of the methods described herein or using any standard method (e.g., any mathematical transformation of calculating allele ratios or predicted allele ratios described herein).

[0283] In some embodiments, the test statistic is calculated based on the magnitude of the difference between the calculated allele ratio and the predicted allele ratio for each locus. In some embodiments, the test statistic Δ is given by the following formula

Number

[0284] where δ i is the magnitude of the difference between the calculated allele ratio and the predicted allele ratio for the i-th locus;

[0285] where μ i is the mean of δ i ; and

[0286] where σ i 2 is the standard deviation of δ i· .

[0287] For example, when the predicted allergy ratio is 0.5, δ i can be defined as follows.

[0288] [Number]

[0289] μ i and σ i values can be computed using a computer using the fact that R i is a binomial random variable. In some embodiments, the standard deviation is assumed to be the same for all seats. In some embodiments, the mean or weighted mean of the standard deviation, or an estimate of the standard deviation, is used for the value of σ i 2 In some embodiments, the test statistic is assumed to have a normal distribution. For example, the central limit theorem means that as the number of seats (e.g., the number of SNPs T, etc.) increases, the distribution of Δ converges to a standard normal.

[0290] In some embodiments, a set of one or more hypotheses that specify the copy number of a chromosome or chromosomal segment in the genome of one or more cells is enumerated. In some embodiments, the hypothesis that is most likely based on the test statistic is selected, and accordingly the copy number of a chromosome or chromosomal segment in the genome of one or more cells is determined. In some embodiments, the hypothesis is selected if the probability that the test statistic belongs to the distribution of test statistics for the hypothesis exceeds an upper threshold. One or more of the hypotheses are rejected if the probability that the test statistic belongs to the distribution of test statistics for the hypothesis is below a lower threshold. Or if the probability that the test statistic belongs to the test statistic for the hypothesis is between the lower and upper thresholds, or if the probability is not determined with a sufficiently high confidence level, the hypothesis is neither selected nor rejected. In some embodiments, the upper and / or lower thresholds are determined from an empirical distribution, such as a distribution from training data (e.g., a sample with a known copy number such as a diploid sample, or a sample known to have a particular deletion or duplication). Such an empirical distribution can be used to select a threshold for a single hypothesis rejection test. Note that since the test statistic Δ is independent of S, both can be used independently if desired.

[0291] 8.6 Exemplary methods for detecting deletions and duplications using the distribution or pattern of alleles This section includes a method for determining whether the copy number of a first homologous chromosomal segment is present in excess compared to a second homologous chromosomal segment. In some embodiments, the method involves (i) a plurality of hypotheses specifying the copy number of a chromosome or chromosomal segment present in the genome of one or more cells (e.g., cancer cells) of an individual, or (ii) enumerating a plurality of hypotheses specifying the degree of overrepresentation of the copy number of the first homologous chromosomal segment when compared to a second homologous chromosomal segment in the genome of one or more cells of the individual. In some embodiments, the method involves obtaining genetic data from an individual for a plurality of polymorphic loci (e.g., SNP loci) on a chromosome or chromosomal segment. In some embodiments, for each hypothesis, a probability distribution of the predicted genotype of the individual is generated. In some embodiments, the data fit between the obtained genetic data of the individual and the probability distribution of the predicted genotype of the individual is calculated. In some embodiments, one or more hypotheses are ranked according to the data fit, and the hypothesis ranked highest is selected. In some embodiments, a technique or algorithm, such as a search algorithm, is used in one or more of the following steps: calculating the data fit, ranking the hypotheses, or selecting the hypothesis ranked highest. In some embodiments, the data fit is a fit to a beta-binomial distribution or a binomial distribution. In some embodiments, a technique or algorithm is selected from the group consisting of maximum likelihood estimation, maximum a posteriori probability estimation, Bayesian estimation, dynamic estimation (e.g., dynamic Bayesian estimation), and expectation maximization estimation. In some embodiments, the method includes applying the technique or algorithm to the obtained genetic data or the predicted genetic data.

[0292] In some embodiments, the method involves enumerating a plurality of hypotheses that specify the copy number of a chromosome or chromosomal segment present in the genome of one or more cells (e.g., cancer cells) of an individual, or (ii) the degree of overrepresentation of the copy number of a first homologous chromosomal segment when compared to a second homologous chromosomal segment in the genome of one or more cells of the individual. In some embodiments, the method involves obtaining genetic data from an individual for a plurality of polymorphic loci (e.g., SNP loci) on a chromosome or chromosomal segment. In some embodiments, the genetic data includes allele counts for the plurality of polymorphic loci. In some embodiments, for each hypothesis, a simultaneous distribution model for the allele counts predicted at the plurality of polymorphic loci on the chromosome or chromosomal segment is generated. In some embodiments, the relative probabilities for one or more hypotheses are determined using the simultaneous distribution model and the allele counts measured in the sample, and the hypothesis with the highest probability is selected.

[0293] In some embodiments, the presence or absence of a CNV, such as a deletion or duplication, is determined using the distribution or pattern of alleles (e.g., the pattern of calculated allele ratios). Optionally, the parental origin of the CNV can be determined based on the pattern.

[0294] 8.7 Exemplary Counting / Quantitative Methods In some embodiments, one or more counting methods (also referred to as quantification methods) are used to detect one or more copy number alterations (CNAs), such as deletions or duplications of chromosomal segments or entire chromosomes. In some embodiments, one or more counting methods are used to determine whether an overrepresentation of the copy number of a first homologous chromosomal segment is due to a duplication of the first homologous chromosomal segment or a deletion of a second homologous chromosomal segment. In some embodiments, one or more counting methods are used to determine the number of duplicated chromosomal segments or extra copies of a chromosome (e.g., whether there are 1, 2, 3, 4 or more extra copies). In some embodiments, one or more counting methods are used to distinguish between samples having a large number of duplications and a low tumor fraction and samples having a small number of duplications and a high tumor fraction. For example, one or more counting methods are used to distinguish between a sample having four extra chromosomal copies and a 10% tumor fraction and a sample having two extra chromosomal copies and a 20% tumor fraction. Exemplary methods are disclosed in, for example, U.S. Patent Publications 2007 / 0184467, 2013 / 0172211, and 2012 / 0003637, U.S. Patents 8,467,976, 7,888,017, 8,008,018, 8,296,076, and 8,195,415, U.S. Patent Application 62 / 008,235 filed Jun. 5, 2014, and U.S. Application 62 / 032,785 filed Aug. 4, 2014, each of which is incorporated herein by reference in its entirety.

[0295] In some embodiments, an f value (e.g., tumor fraction or fetal fraction) is used in CNV determination, such that, for example, the measured difference between the amounts of two chromosomes or chromosomal segments is compared to the difference predicted for a particular type of CNV given the value of f (see, e.g., U.S. Patent Publication 2012 / 0190020, U.S. Patent Publication 2012 / 0190021, U.S. Patent Publication 2012 / 0190557, U.S. Patent Publication 2012 / 0191358, each of which is incorporated herein by reference in its entirety). For example, as the tumor fraction increases, the difference in the amount of chromosomal segments duplicated in the tumor also increases compared to a diploid reference chromosomal segment. In some embodiments, the method involves comparing the relative frequency of a subject chromosome or chromosomal segment to a reference chromosome or chromosomal segment (e.g., a chromosome or chromosomal segment predicted or determined to be diploid) for an f value to determine the likelihood of a CNV. For example, for various possible CNVs (e.g., one or two extra copies of a subject chromosomal segment), the difference in amount between a first chromosome or chromosomal segment and a reference chromosome or chromosomal segment can be compared to the difference predicted taking into account the f value.

[0296] 8.8 Exemplary Counting / Quantitative Methods Using Reference Samples Exemplary quantitative methods that use one or more reference samples are described in U.S. Patent Application 62 / 008,235, filed on June 5, 2014, and U.S. Patent Application 62 / 032,785, filed on August 4, 2014, which are hereby incorporated by reference in their entirety. In some embodiments, for one or more chromosomes or the chromosome of interest, one or more reference samples that are most likely to have no CNVs are selected by selecting the sample with the highest tumor DNA fraction, selecting the sample with the z-score closest to zero, selecting the sample for which the data fits a hypothesis not corresponding to the CNV with the highest confidence or likelihood, selecting the sample that has been determined to be normal, selecting the sample from an individual with the lowest likelihood of having cancer (e.g., young age, male in the case of breast cancer screening, no family history, etc.), selecting the sample with the highest DNA input amount, selecting the sample with the highest signal-to-noise ratio, selecting the sample based on other criteria thought to correlate with the likelihood of having cancer, or selecting the sample using a combination of several criteria. When the reference set is selected, assume that these cases are diploid, and then the bias per SNP, i.e., the experiment-specific amplification for each locus and other processing biases, can be estimated. Then, this estimated experiment-specific bias is used to correct the bias in the measurements of the chromosome of interest, such as the loci of chromosome 21, and optionally the loci of other chromosomes, for samples that are not part of the subset and are assumed to be diploid with respect to chromosome 21. In these samples with unknown ploidy, once the bias is corrected, the data of these samples can then be analyzed twice using the same or another method to determine whether the individual has trisomy 21. For example, the quantitative method can be used for the remaining samples with unknown ploidy, and the z-score can be calculated using the corrected measured gene data of chromosome 21. Alternatively, as part of a preliminary estimate of the ploidy state of chromosome 21, the tumor fraction of the sample from an individual suspected of having cancer can be calculated.For cases with that tumor fraction, the fraction of corrected reads predicted in the case of disomy (the disomy hypothesis), and the fraction of corrected reads predicted in the case of trisomy (the trisomy hypothesis) can be calculated. Alternatively, if the tumor fraction has not been measured in the past, a set of disomy and trisomy hypotheses can be generated for different tumor fractions. For each case, considering the statistical variations predicted in the selection and measurement of various DNA loci, a predicted distribution regarding the fraction of corrected reads can be calculated. The measured fraction of corrected reads can be compared with the distribution of predicted corrected read fractions, and for each sample with unknown ploidy, a likelihood ratio can be calculated for the disomy and trisomy hypotheses. The ploidy state associated with the hypothesis having the highest calculated likelihood can be selected as the correct ploidy state.

[0297] 8.9 Exemplary reference chromosomes or reference chromosome segments In some embodiments, any of the methods described herein are performed on one or more reference chromosomes or reference chromosome segments, and the results are compared to the one or more chromosomes or chromosome segments of interest.

[0298] In some embodiments, the reference chromosome or reference chromosomal segment is used as a control for a prediction of the absence of a CNV. In some embodiments, the reference is the same chromosome or chromosomal segment as one or more various samples that have been found or predicted to have no deletions or duplications in that chromosome or chromosomal segment. In some embodiments, the reference is a different chromosome or chromosomal segment than the test sample predicted to be diploid. In some embodiments, the reference is a different segment derived from one of the target chromosomes in the same sample being tested. For example, the reference may be one or more segments outside of the potential deletion or duplication region. Having the reference on the same chromosome as the one being tested avoids variations between different chromosomes, such as differences between chromosomes in metabolism, apoptosis, histones, inactivation, and / or amplification. Analysis of segments having no CNV on the same chromosome as the chromosome being tested can determine differences between homologs in metabolism, apoptosis, histones, inactivation, and / or amplification, and as a result, the level of variation between homologs in the absence of a CNV can be determined and compared to the results from potential CNVs. In some embodiments, with respect to a potential CNV, the degree of difference between the calculated allele ratio and the predicted allele ratio is greater than the corresponding degree with respect to the reference, thereby confirming the presence of the CNV.

[0299] In some embodiments, the reference chromosome or reference chromosomal segment is used as a control for those predicted to have CNVs, such as deletions or duplications of particular interest. In some embodiments, the reference is the same chromosome or chromosomal segment as one or more various samples known or predicted to have deletions or duplications in that chromosome or chromosomal segment. In some embodiments, the reference is a different chromosome or chromosomal segment than the test sample known or predicted to have a CNV. In some embodiments, with respect to a potential CNV, the degree of difference between the calculated allele ratio and the predicted allele ratio is less than (e.g., not significantly different from) a corresponding degree with respect to the CNV reference, thereby confirming the presence of the CNV. In some embodiments, with respect to a potential CNV, the degree of difference between the calculated allele ratio and the predicted allele ratio is smaller than (e.g., significantly smaller than) a corresponding degree with respect to the CNV reference, thereby confirming the absence of the CNV. In some embodiments, genotypes of cancer cells (or DNA or RNA derived from cancer cells such as cfDNA or cfRNA) are used at one or more loci that are different from genotypes of non-cancerous cells (or DNA or RNA derived from non-cancerous cells such as cfDNA or cfRNA) to determine the tumor fraction. The tumor fraction is used to determine whether an overrepresentation of the copy number of a first homologous chromosomal segment is due to a duplication of the first homologous chromosomal segment or a deletion of a second homologous chromosomal segment. Further, the tumor fraction is used to determine the copy number of the duplicated chromosomal segment or chromosomal excess (e.g., to determine whether there are 1, 2, 3, 4 or more excess copies), for example, to distinguish a sample with four excess chromosomal copies and a 10% tumor fraction from a sample with two excess chromosomal copies and a 20% tumor fraction. Further, the tumor fraction can be used to determine how well the measured data fits the predicted data for possible CNVs. In some embodiments, the degree of overrepresentation of the CNV is used to select a particular treatment or treatment regimen for an individual. For example, some therapeutic agents are only effective against chromosomal segments with at least four, six, or more copy numbers.

[0300] In some embodiments, one or more loci used to determine the tumor fraction are on a reference chromosome or reference chromosome segment, such as a chromosome or chromosome segment that has been found or is predicted to be diploid, typically in cancer cells or in a particular type of cancer that an individual has been found to have or has an increased risk of having, a chromosome or chromosome segment that rarely has duplications or deletions, or a chromosome or chromosome segment that is unlikely to be aneuploid (a segment that is predicted to cause cell death if duplicated or deleted). In some embodiments, using any of the methods of the invention, it is confirmed that the reference chromosome or chromosome segment is diploid in both cancer cells and non-cancerous cells. In some embodiments, one or more chromosomes or chromosome segments with a high confidence of diploid classification are used.

[0301] Examples of loci that can be used to determine tumor fraction include polymorphisms or mutations (e.g., SNPs) in cancer cells (or DNA or RNA such as cfDNA or cfRNA derived from cancer cells) that are not present in the non-cancerous cells of an individual (or DNA or RNA derived from non-cancerous cells). In some embodiments, the tumor fraction is determined by identifying polymorphic loci in which cancer cells (or DNA or RNA derived from cancer cells) have alleles that are not present in non-cancerous cells (or DNA or RNA derived from non-cancerous cells) in a sample from an individual (e.g., a plasma sample or a tumor biopsy), and using the amount of the allele specific to cancer cells at one or more of the identified polymorphic loci to determine the tumor fraction in the sample. In some embodiments, the non-cancerous cells are homozygous with respect to a first allele at the polymorphic locus, and the cancer cells are (i) heterozygous with respect to the first and a second allele or (ii) homozygous with respect to the second allele at the polymorphic locus. In some embodiments, the non-cancerous cells are heterozygous with respect to a first and a second allele at the polymorphic locus, and the cancer cells have (i) one or two copies of a third allele at the polymorphic locus. In some embodiments, it is assumed or determined that cancer cells have only one copy of an allele that is not present in non-cancerous cells. For example, if the genotype of non-cancerous cells is AA, the genotype of cancer cells is AB, and 5% of the signal at that locus in the sample is derived from the B allele and 95% is derived from the A allele, the tumor fraction of the sample is 10%. In some embodiments, it is assumed or determined that cancer cells have two copies of an allele that is not present in non-cancerous cells. For example, if the genotype of non-cancerous cells is AA, the genotype of cancer cells is BB, and 5% of the signal at that locus in the sample is derived from the B allele and 95% is derived from the A allele, the tumor fraction of the sample is 5%. In some embodiments, multiple loci in which cancer cells have alleles not present in non-cancerous cells are analyzed to determine which of the loci in cancer cells are heterozygous and which are homozygous.For example, for a locus where the non-cancerous cells are AA, if the signal from the B allele is about 5% at some loci and about 10% at some loci, the cancer cells are assumed to be heterozygous at the loci with about 5% B allele and homozygous at the loci with about 10% B allele (indicating a tumor fraction of about 10%).

[0302] Examples of loci that can be used to determine the tumor fraction include loci where cancer cells and non-cancerous cells can commonly have one allele in common (e.g., a locus where the cancer cells are AB and the non-cancerous cells are BB, or the cancer cells are BB and the non-cancerous cells are AB). The amount of the A signal, the amount of the B signal, or the ratio of the A signal to the B signal in a mixed sample (containing DNA or RNA derived from cancer cells and non-cancerous cells) is compared to the corresponding value in (i) a sample containing DNA or RNA derived from cancer cells only, or (ii) a sample containing DNA or RNA derived from non-cancerous cells only. The difference in the values is used to determine the tumor fraction of the mixed sample.

[0303] In some embodiments, the loci that can be used to determine the tumor fraction are selected based on the genotype of (i) a sample containing DNA or RNA derived from cancer cells only, and / or (ii) a sample containing DNA or RNA derived from non-cancerous cells only. In some embodiments, the loci are selected based on the analysis of the mixed sample, e.g., loci where the absolute or relative amount of each allele is different from the amount predicted when both cancer cells and non-cancerous cells have the same genotype at a particular locus. For example, when cancer cells and non-cancerous cells have the same genotype, the locus is predicted to produce 0% B signal when all cells are AA, 50% B signal when all cells are AB, or 100% B signal when all cells are BB. If the B signal is a different value, it is shown that the genotypes of the cancer cells and non-cancerous cells are different at that locus and the tumor fraction can be determined using that locus.

[0304] In some embodiments, the tumor fraction calculated based on alleles at one or more loci is compared to the tumor fraction calculated using one or more of the counting methods disclosed herein.

[0305] 8.10 Examples of methods for detecting phenotypes or examples of multiplex mutation analysis methods In some embodiments, the method includes analyzing a sample with respect to a set of mutations associated with a disease or disorder (e.g., cancer) or associated with an increased risk of a disease or disorder. There is a strong correlation between events within a class (e.g., an M or C cancer class) that can be used to improve the signal-to-noise ratio of the method and to classify tumors into distinct clinical subsets. For example, ambiguous results for some mutations (e.g., some CNVs) on one or more chromosomes or chromosomal segments considered together can be very strong signals. In some embodiments, by determining the presence or absence of a plurality of polymorphisms or mutations (e.g., 2, 3, 4, 5, 8, 10, 12, 15, or more) of interest, the sensitivity and / or specificity of determining the presence or absence of a disease or disorder such as cancer, or the presence or absence of an increased risk of having a disease or disorder such as cancer, is increased. In some embodiments, correlations between events across multiple chromosomes are used so that signals are considered more powerfully than considering each signal individually. The design of the method itself can be optimized to best categorize tumors. This can be very useful for early detection and screening of recurrence where sensitivity to one particular mutation / CNV may be most important. In some embodiments, the events are not necessarily correlated but have a probability of being correlated. In some embodiments, a matrix estimation formula with a noise covariance matrix having off-diagonal conditions is used.

[0306] In some embodiments, the present invention features a method for detecting a phenotype of an individual (e.g., a cancer phenotype), the phenotype being defined by the presence of at least one of a set of mutations. In some embodiments, the method comprises obtaining a measurement of DNA or RNA from a sample of DNA or RNA derived from one or more cells from the individual, wherein one or more of the cells are presumed to have the phenotype, and analyzing the measurement of the DNA or the RNA to determine, for each mutation in the set of mutations, the likelihood that at least one of the cells has the mutation. In some embodiments, the method comprises determining that the individual has the phenotype if (i) for at least one of the mutations, the likelihood that at least one of the cells contains the mutation is higher than a threshold value, or (ii) for at least one of the mutations, the likelihood that at least one of the cells has the mutation is lower than a threshold value, and for a plurality of the mutations, the combined likelihood that at least one of the cells has at least one of the mutations is higher than a threshold value. In some embodiments, one or more of the cells have a subset or all of the mutations in the set of mutations. In some embodiments, the subset of mutations is associated with cancer or an increased risk of cancer. In some embodiments, the set of mutations includes a subset or all of the mutations in the M class of cancer mutations (Ciriello, Nat Genet. 45(10):1127-1133, 2013, doi:10.1038 / ng.2762, which is incorporated herein by reference in its entirety). In some embodiments, the set of mutations includes a subset or all of the mutations in the C class of cancer mutations (Ciriello, supra). In some embodiments, the sample includes cell-free DNA or RNA. In some embodiments, the measurement of DNA or RNA includes measurements (such as the amount of each allele at each locus) at a set of polymorphic loci on one or more chromosomes or chromosomal segments of interest.

[0307] 8.11 Examples of Method Combinations To increase the accuracy of the results, two or more methods for detecting the presence or absence of CNV (e.g., any of the methods of the present invention or any known method) are implemented. In some embodiments, one or more methods for analyzing the presence or absence of a disease or disorder, or a factor indicating an increased risk of a disease or disorder (e.g., the methods described herein or any known method) are implemented.

[0308] In some embodiments, standard mathematical techniques are used to calculate the covariance and / or correlation between two or more methods. Additionally, standard mathematical techniques may be used to determine the combined probability of a particular hypothesis based on two or more tests. Exemplary techniques include meta-analysis, Narisys, Fisher's combined probability test for independent tests, Brown's method of combining known covariance and dependent p-values, and Kost's method of combining unknown covariance and dependent p-values. When the likelihood is determined for a second method, and the likelihood is determined by a first method that is in some sense orthogonal or unrelated, combining the likelihoods is straightforward and can be done by multiplication or normalization, or for example using the following equation:

[0309] R comb =R1R2 / [R1R2+(1 - R1)(1 - R2)]

[0310] R comb is the combined likelihood, and R1 and R2 are the individual likelihoods. For example, if the likelihood of trisomy from method 1 is 90% and the likelihood of trisomy from method 2 is 95%, by combining the results from the two methods, the clinician can conclude that the fetus has trisomy with a likelihood of (0.90)(0.95) / [(0.90)(0.95)+(1 - 0.90)(1 - 0.95)] = 99.42%. Even when the first and second methods are not orthogonal, i.e., there is a correlation between the two methods, the likelihoods can still be combined.

[0311] Examples of methods for analyzing multiple factors or variables are disclosed in U.S. Patent No. 8,024,128, issued September 20, 2011, U.S. Publication No. 2007 / 0027636, filed July 31, 2006, and U.S. Publication No. 2007 / 0178501, filed December 6, 2006, which are hereby incorporated by reference in their entirety.

[0312] In various embodiments, the combined probability of a particular hypothesis or diagnosis is greater than 80, 85, 90, 92, 94, 96, 98, 99, or 99.9%, or greater than some other threshold.

[0313] 8.12 Exemplary Methods for Phase-Determining Gene Data In some embodiments, the gene data is phase-determined using the methods described herein or any known method for phase-determining gene data (see, for example, PCT Publication WO2009 / 105531, filed February 9, 2009, PCT Publication WO2010 / 017214, filed August 4, 2009, U.S. Patent Publication 2013 / 0123120, filed November 21, 2012, U.S. Patent Publication 2011 / 0033862, filed October 7, 2010, U.S. Patent Publication 2011 / 0033862, filed August 19, 2010, U.S. Patent Publication 2011 / 0178719, filed February 3, 2011, U.S. Patent No. 8,515,679, filed March 17, 2008, U.S. Patent Publication 2007 / 0184467, filed November 22, 2006, U.S. Patent Publication 2008 / 0243398, filed March 17, 2008, and U.S. Application 61 / 994,791, filed May 16, 2014. Each of these documents is hereby incorporated by reference in its entirety).

[0314] In one embodiment, the individual's genetic data is phased using a computer program that estimates the most likely phase, for example, HapMap-based phasing, using a population based on haplotype frequencies. For example, a haploid dataset can be directly inferred from diploid data using statistical methods, taking advantage of known haplotype blocks in the general population (e.g., the populations generated for the public HapMap Project and the Perlegen Human Haplotype Project). Haplotype blocks are essentially a series of correlated alleles that occur repeatedly in various populations. Since these haplotype blocks are often common from ancient times, they may be used to predict haplotypes from diploid genotypes. Published and available algorithms for realizing this task include the incomplete pedigree method, the Bayesian method based on conjugate prior probability distributions, and the Bayesian method based on prior probability distributions from the genetic characteristics of the population. Some of these algorithms use hidden Markov models.

[0315] In one embodiment, the individual's genetic data is phase-determined using an algorithm that infers haplotypes from genotype data, such as an algorithm that uses decay of linkage disequilibrium with distance, order and spacing of genotype markers, imputation of missing data, estimation of recombination rates, or combinations thereof (see, e.g., Stephens and Scheet, “Accounting for Decay of Linkage Disequilibrium in Haplotype Inference and Missing-Data Imputation” Am. J. Hum. Genet. 76:449-462, 2005, which is incorporated herein by reference in its entirety). An exemplary program is v.2.1 or v2.1.1 of PHASE (available on the World Wide Web at stephenslab.uchicago.edu / software.html, which is incorporated herein by reference in its entirety).

[0316] In one embodiment, the individual's genetic data is phase-determined using an algorithm that infers haplotypes from genotype data, such as an algorithm that uses decay of linkage disequilibrium with distance, order and spacing of genotype markers, imputation of missing data, estimation of recombination rates, or combinations thereof (see, e.g., Stephens and Scheet, “Accounting for Decay of Linkage Disequilibrium in Haplotype Inference and Missing-Data Imputation” Am. J. Hum. Genet. 76:449-462, 2005, which is incorporated herein by reference in its entirety). An exemplary program is v.2.1 or v2.1.1 of PHASE (available on the World Wide Web at stephenslab.uchicago.edu / software.html, which is incorporated herein by reference in its entirety).

[0317] In one embodiment, the individual's genetic data is phase-determined using an algorithm that infers haplotypes from the population's genotype data, e.g., an algorithm that continuously varies the cluster constituents along the chromosome according to a hidden Markov model. This method is flexible and is possible for both the "block-like" pattern of linkage disequilibrium and the gradual decline of linkage disequilibrium at a distance (see, e.g., Scheet and Stephens, “A fast and flexible statistical model for large-scale population genotype data: applications to inferring missing genotypes and haplotypic phase.” Am J Hum Genet, 78:629-644, 2006. This document is incorporated herein by reference in its entirety). An exemplary program is fastPHASE (available on the World Wide Web at stephenslab.uchicago.edu / software.html. This program is incorporated herein by reference in its entirety).

[0318] In one embodiment, the genotype of an individual is determined by using a genotype imputation method, such as a method that uses one or more of the following reference datasets: the HapMap dataset, a control dataset genotyped on multiple SNP chips, and a high-density genotyped sample from the 1000 Genomes Project. Exemplary methods are flexible modeling frameworks that increase accuracy and combine information across multiple reference panels (see, e.g., Howie, Donnelly, and Marchini (2009) “A flexible and accurate genotype imputation method for the next generation of genome-wide association studies.” PLoS Genetics 5(6):e1000529, 2009, which is incorporated herein by reference in its entirety). Exemplary programs are IMPUTE or IMPUTE version 2 (also known as IMPUTE2) (available on the World Wide Web at mathgen.stats.ox.ac.uk / impute / impute_v2.html, which is incorporated herein by reference in its entirety).

[0319] In one embodiment, the genotype data of an individual is phased using an algorithm that infers haplotypes under a coalescent genetic model, such as an algorithm that infers haplotypes, such as the algorithm developed by Stephens in PHASE v2.1. Improvements to the main algorithms rely on the use of binary trees that represent the set of candidate haplotype sets for each individual. These binary tree representations (1) speed up the calculation of the posterior probability of haplotypes by avoiding the redundant operations created in PHASE v2.1, and (2) overcome the exponential aspect of the haplotype inference problem by performing a smart search of the most likely path (i.e., haplotype) in the binary tree (see, e.g., Delaneau, Coulonges and Zagury, “Shape-IT: new rapid and accurate algorithm for haplotype inference,” BMC Bioinformatics 9:540, 2008 doi:10.1186 / 1471-2105-9-540, which is incorporated herein by reference in its entirety). An exemplary program is SHAPEIT (available on the World Wide Web at mathgen.stats.ox.ac.uk / genetics_software / shapeit / shapeit.html, which is incorporated herein by reference in its entirety).

[0320] In one embodiment, an individual's genetic data is co-determined using an algorithm that infers haplotypes from a population's genotype data, such as an algorithm that uses haplotype-fragment frequencies to obtain an experimental-based probability for longer haplotypes. In some embodiments, the algorithm reconstructs the haplotype. Thereby, the haplotype comes to have maximum local coherence (see, e.g., Eronen, Geerts, and Toivonen, “HaploRec: Efficient and accurate large-scale reconstruction of haplotypes,” BMC Bioinformatics 7:542, 2006, which is incorporated herein by reference in its entirety). An exemplary program is HaploRec, such as version 2.3 of HaploRec (available on the World Wide Web at cs.helsinki.fi / group / genetics / haplotyping.html, which is incorporated herein by reference in its entirety).

[0321] In one embodiment, an individual's genetic data is co-determined using an algorithm that infers haplotypes from a population's genotype data, such as an algorithm that uses a partition-ligation strategy, and algorithms such as an expectation-maximization-based algorithm (see, e.g., Qin, Niu, and Liu, “Partition-Ligation-Expectation-Maximization Algorithm for Haplotype Inference with Single-Nucleotide Polymorphisms,” Am J Hum Genet. 71(5):1242-1247, 2002, which is incorporated herein by reference in its entirety). An exemplary program is PL-EM (available on the World Wide Web at people.fas.harvard.edu / ~junliu / plem / click.html, which is incorporated herein by reference in its entirety).

[0322] In one embodiment, the individual's genetic data is co-determined using an algorithm for inferring haplotypes from the population's genotype data, such as an algorithm that jointly co-determines genotypes into haplotypes and block partitions. In some embodiments, the expectation maximization algorithm is used (see, e.g., Kimmel and Shamir, “GERBIL: Genotype Resolution and Block Identification Using Likelihood,” Proceedings of the National Academy of Sciences of the United States of America (PNAS) 102:158 - 162, 2005, which is incorporated herein by reference in its entirety). An exemplary program is GERBIL, which is available as part of the GEVALT version 2 program (available on the World Wide Web at acgt.cs.tau.ac.il / gevalt / , which is incorporated herein by reference in its entirety).

[0323] In one embodiment, the individual's genetic data is co-determined using an algorithm for inferring haplotypes from the population's genotype data, such as the EM algorithm, to calculate the ML estimates of haplotype frequencies given unphased genotype measurements. The algorithm further omits measurements of some genotypes (e.g., due to PCR failure). The algorithm is also capable of multiple complementation of individual haplotypes (see, e.g., Clayton, D. (2002), “SNPHAP: A Program for Estimating Frequencies of Large Haplotypes of SNPs,” which is incorporated herein by reference in its entirety). An exemplary program is SNPHAP.

[0324] In one embodiment, the individual's genetic data is co-determined using an algorithm that infers haplotypes from the genotype data of a population, for example, an algorithm for haplotype inference based on statistical values of genotypes collected for SNP pair formation. Using this software, relatively accurate co-determination of a large number of long genomic sequences obtained, for example, from a DNA array can be performed. An exemplary program takes as input a genotype matrix and outputs a corresponding haplotype matrix (see, for example, Brinza and Zelikovsky, “2SNP: scalable phasing based on 2-SNP haplotypes,” Bioinformatics. 22(3):371-3, 2006. The entire disclosure of this document is incorporated herein by reference). An exemplary program is 2SNP (available on the World Wide Web at alla.cs.gsu.edu / ~software / 2SNP. The entire disclosure of this program is incorporated herein by reference).

[0325] In various embodiments, an individual's genetic data is phase determined using data regarding the probability that chromosomes cross over at different positions on a chromosome or chromosomal segment (e.g., using recombination data found in the HapMap database to generate a recombination risk score for any interval), and models the dependencies between polymorphic alleles on a chromosome or chromosomal segment. In some embodiments, the number of alleles at a polymorphic locus is calculated on a computer based on sequence data or SNP array data. In some embodiments, each of a number of hypotheses is generated (e.g., generated on a computer) in relation to various possible states of a chromosome or chromosomal segment (e.g., over-occurrence of the copy number of a first homologous chromosomal segment when compared to a second homologous chromosomal segment in the genome of one or more cells from an individual, duplication of the first homologous chromosomal segment, duplication of the second homologous chromosomal segment, or equal representation of the first and second homologous chromosomal segments). A model (e.g., a simultaneous distribution model) for the predicted number of alleles at polymorphic loci on the chromosome is constructed (e.g., constructed on a computer) for each hypothesis. The relative probability of each hypothesis is determined (e.g., determined on a computer) using the simultaneous distribution model and the number of alleles. And the hypothesis with the highest probability is selected. In some embodiments, the steps of constructing the simultaneous distribution model regarding the number of alleles and determining the relative probability of each hypothesis are performed using a method that does not require the use of a reference chromosome.

[0326] In some embodiments, a sample derived from an individual (e.g., a tumor biopsy, a blood sample, a plasma sample, a serum sample, or another sample likely to contain the primary or only cells, DNA, or RNA having the CNV of interest) is analyzed to determine the phase of one or more regions found or suspected to contain the CNV of interest (e.g., a deletion or duplication). In some embodiments, the sample has a high tumor fraction (e.g., 30, 40, 50, 60, 70, 80, 90, 95, 98, 99, or 100%). In some embodiments, the sample is a blood sample from a mother carrying a fetus. In some embodiments, the blood sample from a mother carrying a fetus contains circulating fetal cells and / or fetal DNA.

[0327] In some embodiments, the sample has a haplotype imbalance or any aneuploidy. In some embodiments, the sample contains any admixture of two types of DNA, the two types of DNA having two haplotypes in different ratios and sharing at least one haplotype. In some embodiments, at least 10, 100, 500, 1,000, 2,000, 3,000, 5,000, 8,000, or 10,000 polymorphic loci are analyzed and the phase of some or all of the alleles at the loci is determined. In some embodiments, the sample is derived from a cell or tissue that has been processed to be aneuploid, such as aneuploidy induced by long-term cell culture.

[0328] In some embodiments, most or all of the DNA or RNA in the sample has the CNV of interest. In some embodiments, the ratio of DNA or RNA from one or more target cells containing the CNV of interest to total DNA or total RNA in the sample is at least 80, 85, 90, 95, or 100%. For a sample with a deletion, for the cell (or DNA or RNA) with the deletion, only one haplotype is present. This first haplotype can be determined using standard methods for determining the identity of alleles present in the deleted region. In a sample containing only cells (or DNA or RNA) with deletions, only the signal from the first haplotype present in those cells is present. In a sample that also contains a small amount of cells (or DNA or RNA) without deletions (such as a small amount of non-cancerous cells), the weak signal from the second haplotype in these cells (or DNA or RNA) can be ignored. The second haplotype present in other cells, DNA, or RNA from an individual without the deletion can be determined by inference. For example, if the genotype of cells from an individual without the deletion is (AB, AB) and the phasing data for the individual indicates that the first haplotype is (A, A), it can be inferred that the other haplotype is (B, B).

[0329] For samples in which both cells (or DNA or RNA) with deletions and cells (or DNA or RNA) without deletions are present, the phase can be further determined. For example, the x-axis represents the linear position of individual loci along the chromosome, and a plot can be created where the y-axis represents the number of reads of the A allele as the ratio of the reads of the total (A + B) alleles. In some embodiments regarding deletions, the pattern includes two central bands representing SNPs for which the individual is heterozygous (the upper band represents AB from cells without deletions and A from cells with deletions, and the lower band represents AB from cells without deletions and B from cells with deletions). In some embodiments, as the proportion of cells, DNA, or RNA with deletions increases, the separation of these two bands also increases. Thus, the first haplotype can be determined using the identity of the A allele, and the second haplotype can be determined using the identity of the B allele.

[0330] For samples with duplications, for cells (or DNA or RNA) with duplications, there are extra copies of the haplotype. The haplotype in the duplicated region can be determined using standard methods for determining the identity of alleles present in increased amounts in the duplicated region, or the haplotype in non-duplicated regions can be determined using standard methods for determining the identity of alleles present in decreased amounts. Once one haplotype is determined, the other haplotype can be determined by inference.

[0331] For a sample in which both cells (or DNA or RNA) with duplications and cells (or DNA or RNA) without duplications are present, the phase can be further determined using a method similar to the method described above for deletions. For example, the x-axis can represent the linear position of individual loci along a chromosome, and a plot can be created where the y-axis represents the number of reads of the A allele as the ratio of the reads of the total (A + B) alleles. In some embodiments regarding deletions, the pattern includes two central bands representing SNPs for which the individual is heterozygous (the upper band represents AB from cells without duplications and AAB from cells with duplications, and the lower band represents AB from cells without duplications and ABB from cells with duplications). In some embodiments, as the proportion of cells, DNA, or RNA with duplications increases, the separation of these two bands also increases. Thus, the first haplotype can be determined using the identity of the A allele, and the second haplotype can be determined using the identity of the B allele. In some embodiments, for a sample (e.g., a tumor biopsy or plasma sample) from an individual known to have cancer, the phase of one or more CNV regions (e.g., the phase of at least 50, 60, 70, 80, 90, 95, or 100% of the polymorphic loci in the measured region) is determined, and subsequent samples from the same individual are analyzed to monitor the progression of the cancer (e.g., remission or recurrence of the cancer is monitored). In some embodiments, a sample with a high tumor proportion (e.g., a tumor biopsy or plasma sample from an individual with a high tumor burden) is used to obtain phase determination data, and this data is used to analyze subsequent samples with a low tumor proportion (e.g., a plasma sample from an individual who has received cancer treatment or is in remission).

[0332] In some embodiments, the genotype data of an individual is co-determined using two or more of the methods described herein. In some embodiments, a bioinformatics method (e.g., a method of estimating the most likely phase using population-based haplotype frequencies) and a molecular biology method (e.g., any of the molecular genotyping methods described herein for obtaining actual genotyping data rather than phase determination data estimated based on bioinformatics) are used. In some embodiments, genotyping data from other subjects (e.g., past subjects) is used to improve the population data. For example, the genotyping data from other subjects can be added to the population data to calculate the prior probability distribution of possible haplotypes for another subject. In some embodiments, the prior probability distribution of possible haplotypes for another subject is calculated using genotyping data from other subjects (e.g., past subjects).

[0333] In some embodiments, probabilistic data may be used. For example, due to the probabilistic nature of the occurrence of DNA molecules in a sample, as well as various amplification biases and measurement biases, the relative numbers of DNA molecules measured from two different loci, or from different alleles of a given locus, do not always represent the relative numbers of molecules in a mixture, or the relative numbers of molecules in an individual. When attempting to determine the genotype of a normal diploid individual at a given locus on an autosome by sequencing DNA from the plasma of an individual, it is expected that only one allele will be observed (homozygosity), or approximately equal numbers of two alleles will be observed (heterozygosity). If at that allele, 10 molecules of the A allele are observed and 2 molecules of the B allele are observed, it is not clear whether the individual is homozygous at that locus and the 2 molecules of the B allele are due to noise or contamination, or whether the individual is heterozygous and the low number of molecules of the B allele is due to random and statistical variation, amplification bias, contamination, or other various causes in the number of DNA molecules in the plasma. In this case, the probability that the individual is homozygous and the corresponding probability that the individual is heterozygous can be calculated and these probabilistic genotypes can be used for further calculations.

[0334] For a given allele ratio, note that the likelihood that the ratio closely represents the ratio of DNA molecules in an individual increases with the number of observed molecules. For example, when 100 molecules of A and 100 molecules of B are measured, the likelihood that the actual ratio is 50% is much higher than when 10 molecules of A and 10 molecules of B are measured. In one embodiment, Bayes' theory is combined with a detailed model of the data to determine the likelihood that a particular hypothesis is correct given the actual measurements. For example, when considering two hypotheses, namely the hypothesis corresponding to a trisomic individual and the hypothesis corresponding to a disomic individual, the probability that the disomic hypothesis is correct is much higher when 100 bialleles of each are observed compared to when 10 bialleles of each are observed. As the noise in the data increases due to bias, contamination, or other sources of noise, or as the number of actual measurements at a given locus decreases, the probability that the hypothesis of the maximum likelihood given the actual data is true decreases. In practice, it is possible to aggregate the probabilities at many loci to increase the confidence with which the maximum likelihood hypothesis can be determined to be the correct hypothesis. In some embodiments, the probabilities are simply aggregated without considering recombination. In some embodiments, the calculations take into account crossovers.

[0335] In one embodiment, probabilistically determined data is used for the determination of copy number variation. In some embodiments, the probabilistically determined data is, for example, data on haplotype block frequencies based on a population from a data source such as the HapMap database. In some embodiments, the probabilistically determined data is haplotype data obtained by molecular methods, for example, where individual segments of a chromosome are diluted to one molecule per reaction, but are co-determined by dilution where the identity of the haplotype is not fully determined due to probabilistic noise. In some embodiments, the probabilistically determined data is haplotype data obtained by molecular methods, where the identity of the haplotype may be known with high certainty.

[0336] Suppose a hypothetical case where a physician wishes to determine whether there are any cells in an individual's body with deletions in specific chromosomal segments by measuring plasma DNA from the individual. The physician can utilize the knowledge that if all the cells from which the plasma DNA originated are diploid and of the same genotype, for heterozygous loci, the relative numbers of DNA molecules observed for each of the two alleles fall into a single distribution that aggregates 50% A alleles and 50% B alleles. However, if a fraction of the cells from which the plasma DNA originated has a deletion in a specific chromosomal segment, for heterozygous loci, the relative numbers of DNA molecules observed for each of the two alleles are predicted to fall into two distributions. One distribution has more than 50% of the A alleles aggregated for loci where there is a deletion in the chromosomal segment containing the B allele, and the other distribution has less than 50% aggregated for loci where there is a deletion in the chromosomal segment containing the A allele. As the proportion of cells from which the plasma DNA originated increases, these two distributions move further apart from 50%.

[0337] In this example of the hypothesis, assume a physician who wants to determine whether an individual has a deletion in a chromosomal region in a part of the cells in the individual's body. The physician may collect blood from the individual into a vacutainer or other type of blood tube and centrifuge the blood to separate the plasma layer. The physician isolates DNA from the plasma and may concentrate the DNA of the targeted loci by targeted amplification or other amplification, locus capture techniques, size enrichment or other enrichment techniques as the case may be. The physician may perform the analysis by generating allele frequency data, for example, by measuring the number of alleles in an SNP set, that is, by using an assay such as qPCR, sequencing, microarray, or other techniques for measuring the amount of DNA in the sample to generate concentrated and / or amplified DNA. The inventors consider data analysis for a case where a physician uses a targeted amplification technique to amplify cell-free plasma DNA and then sequences the amplified DNA to obtain exemplary data with the following possibilities for 6 SNPs present on a chromosomal segment that is an indicator of cancer. In this case, the individual was heterozygous for the SNP:

[0338] SNP1: A allele in 460 reads, B allele in 540 reads (46% A)

[0339] SNP2: A allele in 530 reads, B allele in 470 reads (53% A)

[0340] SNP3: A allele in 40 reads, B allele in 60 reads (40% A)

[0341] SNP4: A allele in 46 reads, B allele in 54 reads (46% A)

[0342] SNP5: A allele in 520 reads, B allele in 480 reads (52% A)

[0343] SNP6: A allele in 200 reads, B allele in 200 reads (50% A)

[0344] From this series of data, it may be difficult to distinguish between cases where an individual is normal and all cells are diploid, or cases where an individual may have cancer and some of the cells have DNA that becomes cell-free DNA present in the plasma and have deletions or duplications in the chromosomes. For example, two hypotheses of maximum likelihood are that the individual has a deletion in this chromosomal segment, with a tumor proportion of 6%, and the deleted segment of the chromosome has genotypes of (A, B, A, A, B, B) or (A, B, A, A, B, A) for six SNPs. In the expression of the genotype of the individual regarding the SNP set, the first character within the parentheses corresponds to the genotype of the haplotype of SNP1, and the second character corresponds to SNP2.

[0345] Using a method for determining the haplotype of an individual in a chromosomal segment, for one of two chromosomes, if it is found that the haplotype was (A, B, A, A, B, B), it is consistent with the maximum likelihood hypothesis, and the calculated likelihood that the individual has a deletion in the segment and thus may have cancerous or precancerous cells increases significantly. On the other hand, if it is found that the individual has the haplotype (A, A, A, A, A, A, A), the likelihood that the individual has a deletion in the chromosomal segment decreases significantly, and presumably the likelihood of the non-deletion hypothesis becomes high (the actual likelihood value depends on other parameters such as measurement noise in the system, for example).

[0346] There are many methods for determining the haplotype of an individual, many of which are described separately herein. A partial list is shown here and is not intended to be exhaustive. One method is a biological method where individual DNA molecules are diluted in an arbitrary given reaction volume until there is approximately one molecule from each chromosomal region, and then the genotype is measured using methods such as sequencing. Another method is an information science-based method where population data regarding various haplotypes associated with frequencies can be used probabilistically. Another method is a method of measuring the diploid data of an individual together with one or more related individuals, where the related individuals are predicted to share a haplotype block with the individual and suggest a haplotype block. Another method is a method of collecting tissue samples containing high concentrations of deleted segments or duplicated segments and measuring the haplotype based on allelic imbalance. For example, using genotype measurements from tumor tissue samples with deletions, phasing data for the region can be determined, and then this data can be used to determine whether cancer regrew after resection.

[0347] In practice, typically more than 20 SNPs, more than 50 SNPs, more than 100 SNPs, more than 500 SNPs, more than 1,000 SNPs, or more than 5,000 SNPs are measured in a given chromosomal segment.

[0348] 9. Examples of mutations Examples of mutations associated with a disease or disorder such as cancer, or an increased risk of a disease or disorder such as cancer (e.g., a risk exceeding normal levels) include single nucleotide variants (SNVs), multiple nucleotide mutations, deletions (e.g., deletions in the region of 2 million to 30 million base pairs), duplications, or tandem repeats. In some embodiments, the mutation is in DNA such as cell-free DNA (cfDNA), cell-free mitochondrial DNA (cf mDNA), nuclear DNA (cf nDNA), cell DNA or cell-free DNA originating from mitochondrial DNA. In some embodiments, the mutation is in RNA such as cfRNA, cell RNA, cytoplasmic RNA, coding cytoplasmic RNA, non-coding cytoplasmic RNA, mRNA, miRNA, mitochondrial RNA, rRNA, or tRNA. In some embodiments, the mutation is present at a higher frequency in a subject having a disease or disorder (e.g., cancer) than in a subject not having the disease or disorder (e.g., cancer). In some embodiments, the mutation is an indicator of cancer, such as a causative mutation. In some embodiments, the mutation is a driver mutation that plays a causative role in a disease or disorder. In some embodiments, the mutation is not a causative mutation. For example, in some cancers, multiple mutations accumulate, some of which are not causative mutations. Even mutations that are not causative (e.g., mutations that are present at a higher frequency in a subject having a disease or disorder than in a subject not having the disease or disorder) may be useful for diagnosing the disease or disorder. In some embodiments, the mutation is loss of heterozygosity (LOH) at one or more microsatellites.

[0349] In some embodiments, a subject is screened for one or more polymorphisms or mutations known to be present in the subject (e.g., for their presence, changes in the amount of cells, DNA or RNA having these polymorphisms or mutations, or for cancer remission or recurrence testing). In some embodiments, a subject is screened for one or more polymorphisms or mutations known to be at risk in the subject (e.g., a subject with a relative having the polymorphism or the mutation). In some embodiments, a subject is screened for a panel of polymorphisms or mutations associated with a disease or disorder such as cancer (e.g., at least 5, 10, 50, 100, 200, 300, 500, 750, 1,000, 1,500, 2,000, or 5,000 polymorphisms or mutations).

[0350] Numerous coding variants associated with cancer are described in Abaan et al., “The Exomes of the NCI-60 Panel: A Genomic Resource for Cancer Biology and Systems Pharmacology”, Cancer Research, July 15, 2013, and are available worldwide on the World Wide Web at dtp.nci.nih.gov / branches / btb / characterizationNCI60.html. The foregoing is hereby incorporated by reference in its entirety. The NCI-60 human cancer cell line panel consists of 60 diverse cell lines representing cancers of the lung, colon, brain, ovary, breast, prostate, and kidney, as well as leukemia and melanoma. The genetic variations identified in these cell lines consist of two types: type I variants present in the normal population, and cancer-specific type II variants.

[0351] Exemplary polymorphisms or mutations (e.g., deletions or duplications) are in one or more of the following genes: TP53, PTEN, PIK3CA, APC, EGFR, NRAS, NF2, FBXW7, ERBBs, ATAD5, KRAS, BRAF, VEGF, EGFR, HER2, ALK, p53, BRCA, BRCA1, BRCA2, SETD2, LRP1B, PBRM, SPTA1, DNMT3A, ARID1A, GRIN2A, TRRAP, STAG2, EPHA3 / 5 / 7, POLE, SYNE1, C20orf80, CSMD1, CTNNB1, ERBB2.FBXW7, KIT, MUC4, ATM, CDH1, DDX11, DDX12, DSPP, EPPK1, FAM186A, GNAS, HRNR, KRTAP4-11, MAP2K4, MLL3, NRAS, RB1, SMAD4, TTN, ABCC9, ACVR1B, ADAM29, ADAMTS19, AGAP10, AKT1, AMBN, AMPD2, ANKRD30A, ANKRD40, APOBR, AR, BIRC6, BMP2, BRAT1, BTNL8, C12orf4, C1QTNF7, C20orf186, CAPRIN2, CBWD1, CCDC30, CCDC93, CD5L, CDC27, CDC42BPA, CDH9, CDKN2A, CHD8, CHEK2, CHRNA9, CIZ1, CLSPN, CNTN6, COL14A1, CREBBP, CROCC, CTSF, CYP1A2, DCLK1, DHDDS, DHX32, DKK2, DLEC1, DNAH14, DNAH5, DNAH9, DNASE1L3, DUSP16, DYNC2H1, ECT2, EFHB, RRN3P2, TRIM49B, TUBB8P5, EPHA7, ERBB3, ERCC6, FAM21A, FAM21C, FCGBP, FGFR2, FLG2, FLT1, FOLR2, FRYL, FSCB, GAB1, GABRA4, GABRP, GH2, GOLGA6L1, GPHB5, GPR32, GPX5, GTF3C3, HECW1, HIST1H3B, HLA-A, HRAS, HS3ST1, HS6ST1, HSPD1, IDH1, JAK2, KDM5B, KIAA0528, KRT15, KRT38, KRTAP21-1, KRTAP4-5, KRTAP4-7, KRTAP5-4, KRTAP5-5, LAMA4, LATS1, LMF1, LPAR4, LPPR4,LRRFIP1, LUM, LYST, MAP2K1, MARCH1, MARCO, MB21D2, MEGF10, MMP16, MORC1, MRE11A, MTMR3, MUC12, MUC17, MUC2, MUC20, NBPF10, NBPF20, NEK1, NFE2L2, NLRP4, NOTCH2, NRK, NUP93, OBSCN, OR11H1, OR2B11, OR2M4, OR4Q3, OR5D13, OR8I2, OXSM, PIK3R1, PPP2R5C, PRAME, PRF1, PRG4, PRPF19, PTH2, PTPRC, PTPRJ, RAC1, RAD50, RBM12, RGPD3, RGS22, ROR1, RP11-671M22.1, RP13-996F3.4, RP1L1, RSBN1L, RYR3, SAMD3, SCN3A, SEC31A, SF1, SF3B1, SLC25A2, SLC44A1, SLC4A11, SMAD2, SPTA1, ST6GAL2, STK11, SZT2, TAF1L, TAX1BP1, TBP, TGFBI, TIF1, TMEM14B, TMEM74, TPTE, TRAPPC8, TRPS1, TXNDC6, USP32, UTP20, VASN, VPS72, WASH3P, WWTR1, XPO1, ZFHX4, ZMIZ1, ZNF167, ZNF436, ZNF492, ZNF598, ZRSR2, ABL1, AKT2, AKT3, ARAF, ARFRP1, ARID2, ASXL1, ATR, ATRX, AURKA, AURKB, AXL, BAP1, BARD1, BCL2, BCL2L2, BCL6, BCOR, BCORL1, BLM, BRIP1, BTK, CARD11, CBFB, CBL, CCND1, CCND2, CCND3, CCNE1, CD79A, CD79B, CDC73, CDK12, CDK4, CDK6, CDK8, CDKN1B, CDKN2B, CDKN2C, CEBPA, CHEK1, CIC, CRKL, CRLF2, CSF1R, CTCF, CTNNA1, DAXX, DDR2, DOT1L, EMSY (C11orf30), EP300, EPHA3, EPHA5, EPHB1, ERBB4, ERG, ESR1, EZH2, FAM123B (WTX), FAM46C, FANCA, FANCC, FANCD2, FANCE, FANCF, FANCG, FANCL, FGF10, FGF14, FGF19, FGF23FGF3, FGF4, FGF6, FGFR1, FGFR2, FGFR3, FGFR4, FLT3, FLT4, FOXL2, GATA1, GATA2, GATA3, GID4 (C17orf39), GNA11, GNA13, GNAQ, GNAS, GPR124, GSK3B, HGF, IDH1, IDH2, IGF1R, IKBKE, IKZF1, IL7R, INHBA, IRF4, IRS2, JAK1, JAK3, JUN, KAT6A (MYST3), KDM5A, KDM5C, KDM6A, KDR, KEAP1, KLHL6, MAP2K2, MAP2K4, MAP3K1, MCL1, MDM2, MDM4, MED12, MEF2B, MEN1, MET, MITF, MLH1, MLL, MLL2, MPL, MSH2, MSH6, MTOR, MUTYH, MYC, MYCL1, MYCN, MYD88, NF1, NFKBIA, NKX2-1, NOTCH1, NPM1, NRAS, NTRK1, NTRK2, NTRK3, PAK3, PALB2, PAX5, PBRM1, PDGFRA, PDGFRB, PDK1, PIK3CG, PIK3R2, PPP2R1A, PRDM1, PRKAR1A, PRKDC, PTCH1, PTPN11, RAD51, RAF1, RARA, RET, RICTOR, RNF43, RPTOR, RUNX1, SMARCA4, SMARCB1, SMO, SOCS1, SOX10, SOX2, SPEN, SPOP, SRC, STAT4, SUFU, TET2, TGFBR2, TNFAIP3, TNFRSF14, TOP1, TP53, TSC1, TSC2, TSHR, VHL, WISP3, WT1, ZNF217, ZNF703, and combinations thereof (Su et al., J Mol Diagn 2011, 13:74-84; DOI:10.1016 / j.jmoldx.2010.11.010; and Abaan et al., “The Exomes of the NCI-60 Panel: A Genomic Resource for Cancer Biology and Systems Pharmacology”, Cancer Research, July 15, 2013, which is incorporated herein by reference in its entirety). In some embodiments, the duplication is a chromosome 1p (Chr1p) duplication associated with breast cancer. In some embodiments, one or more polymorphisms or mutations areFor example, it is in BRAF, such as the V600E mutation. In some embodiments, one or more polymorphisms or mutations are in K-ras. In some embodiments, there are combinations of one or more polymorphisms or mutations in K-ras and APC. In some embodiments, there are combinations of one or more polymorphisms or mutations in K-ras and p53. In some embodiments, there are combinations of one or more polymorphisms or mutations in APC and p53. In some embodiments, there are combinations of one or more polymorphisms or mutations in K-ras, APC, and p53. In some embodiments, there are combinations of one or more polymorphisms or mutations in K-ras and EGFR. Exemplary polymorphisms or mutations are in one or more of the following microRNAs: miR-15a, miR-16-1, miR-23a, miR-23b, miR-24-1, miR-24-2, miR-27a, miR-27b, miR-29b-2, miR-29c, miR-146, miR-155, miR-221, miR-222, and miR-223 (Calin et al., "A microRNA signature associated with prognosis and progression in chronic lymphocytic leukemia." N Engl J Med 353:1793-801, 2005. This document is incorporated herein by reference in its entirety).

[0352] In some embodiments, the deletion is at least 0.01 kb, 0.1 kb, 1 kb, 10 kb, 100 kb, 1 mb, 2 mb, 3 mb, 5 mb, 10 mb, 15 mb, 20 mb, 30 mb, or 40 mb. In some embodiments, the deletion is from 1 kb to 40 mb, for example, from 1 kb to 100 kb, from 100 kb to 1 mb, from 1 to 5 mb, from 5 to 10 mb, from 10 to 15 mb, from 15 to 20 mb, from 20 to 25 mb, from 25 to 30 mb, or from 30 to 40 mb.

[0353] In some embodiments, the duplication is a duplication of at least 0.01 kb, 0.1 kb, 1 kb, 10 kb, 100 kb, 1 mb, 2 mb, 3 mb, 5 mb, 10 mb, 15 mb, 20 mb, 30 mb, or 40 mb. In some embodiments, the duplication is a duplication of 1 kb or more to 40 mb or less, for example, 1 kb or more to 100 kb or less, 100 kb or more to 1 mb or less, 1 or more to 5 mb or less, 5 or more to 10 mb or less, 10 or more to 15 mb or less, 15 or more to 20 mb or less, 20 or more to 25 mb or less, 25 or more to 30 mb or less, or 30 or more to 40 mb or less.

[0354] In some embodiments, the tandem repeat is a repeat of 2 or more to 60 or less nucleotides, such as, for example, 2 or more to 6 or less, 7 or more to 10 or less, 10 or more to 20 or less, 20 or more to 30 or less, 30 or more to 40 or less, 40 or more to 50 or less, or 50 or more to 60 or less nucleotides. In some embodiments, the tandem repeat is a 2-nucleotide repeat (dinucleotide repeat). In some embodiments, the tandem repeat is a 3-nucleotide repeat (trinucleotide repeat).

[0355] In some embodiments, the polymorphism or mutation is predictive. Exemplary predictive mutations include, for example, K-ras mutations such as the K-ras mutation that is an indicator of postoperative disease recurrence in colorectal cancer (Ryan et al., "A prospective study of circulating mutant KRAS2 in the serum of patients with colorectal neoplasia: strong prognostic indicator in postoperative follow up," Gut 52:101-108, 2003; and Lecomte T et al., Detection of free-circulating tumor-associated DNA in plasma of colorectal cancer patients and its association with prognosis," Int J Cancer 100:542-548, 2002. Each of these references is incorporated herein by reference in its entirety).

[0356] In some embodiments, the polymorphism or mutation is associated with a change in response to a particular treatment (e.g., an increase or decrease in efficacy or side effects). An example is the K-ras mutation associated with a decrease in responsiveness to EGFR-based therapy in non-small cell lung cancer (Wang et al., "Potential clinical significance of a plasma-based KRAS mutation analysis in patients with advanced non-small cell lung cancer," Clin Canc Res 16:1324-1330, 2010. This reference is incorporated herein by reference in its entirety).

[0357] K-ras is an oncogene and is activated in many cancers. Exemplary K-ras mutations are mutations at codons 12, 13, and 61. K-ras cfDNA mutations have been identified in pancreatic cancer, lung cancer, colorectal cancer, bladder cancer, and gastric cancer (Fleischhacker & Schmidt “Circulating nucleic acids(CNAs)and caner - a survey,” Biochim Biophys Acta 1775:181-232,2007. This reference is incorporated herein by reference in its entirety).

[0358] p53 is a tumor suppressor and is mutated in many cancers, contributing to tumor progression (Levine & Oren “The first 30 years of p53:growing ever more complex.Nature Rev Cancer,”9:749-758,2009. This reference is incorporated herein by reference in its entirety). For example, many different codons, such as Ser249, may be mutated. p53 cfDNA mutations have been identified in breast cancer, lung cancer, ovarian cancer, bladder cancer, gastric cancer, pancreatic cancer, colorectal cancer, colon cancer and hepatocellular carcinoma (Fleischhacker & Schmidt “Circulating nucleic acids(CNAs)and caner - a survey,”Biochim Biophys Acta 1775:181-232,2007. This reference is incorporated herein by reference in its entirety).

[0359] BRAF is a cancer gene downstream of Ras. BRAF mutations have been identified in gliocytic neoplasms, melanoma, thyroid cancer, and lung cancer (Dias-Santagata et al., BRAF V600E mutations are common in pleomorphic xanthoastrocytoma: diagnostic and therapeutic implications. PLOS ONE 2011;6:e17948, 2011; Shinozaki et al., Utility of circulating B-RAF DNA mutation in serum for monitoring melanoma patients receiving biochemotherapy. Clin Canc Res 13:2068-2074, 2007; and Board et al., Detection of BRAF mutations in the tumor and serum of patients enrolled in the AZD6244 (ARRY-142886) advanced melanoma phase II study. Brit J Canc 2009;101:1724-1730. These references are hereby incorporated by reference in their entirety). The V600E mutation of BRAF occurs, for example, in melanoma and is more commonly seen in advanced stages. The V600E mutation has been detected in cfDNA.

[0360] EGFR contributes to cell proliferation and is dysregulated in many cancers (Downward J. Targeting RAS signalling pathways in cancer therapy. Nature Rev Cancer 3:11-22, 2003; and Levine & Oren “The first 30 years of p53: growing ever more complex. Nature Rev Cancer,” 9:749-758, 2009. The entire disclosures of these references are incorporated herein by reference). Exemplary EGFR mutations include mutations in exons 18-21 and have been identified in lung cancer patients. EGFR cfDNA mutations have been identified in lung cancer patients (Jia et al., “Prediction of epidermal growth factor receptor mutations in the plasma / pleural effusion to efficacy of gefitinib treatment in advanced non-small cell lung cancer,” J Canc Res Clin Oncol 2010;136:1341-1347, 2010. The entire disclosure of this reference is incorporated herein by reference).

[0361] Exemplary polymorphisms or mutations associated with breast cancer include loss of heterozygosity at microsatellites (Kohler et al., “Levels of plasma circulating cell free nuclear and mitochondrial DNA as potential biomarkers for breast tumors,” Mol Cancer 8:doi:10.1186 / 1476-4598-8-105,2009, which is incorporated herein by reference in its entirety), p53 mutations (e.g., mutations in exons 5-8) (Garcia et al., “Extracellular tumor DNA in plasma and overall survival in breast cancer patients,” Genes, Chromosomes & Cancer 45:692-701,2006, which is incorporated herein by reference in its entirety), HER2 (Sorensen et al., “Circulating HER2 DNA after trastuzumab treatment predicts survival and response in breast cancer,” Anticancer Res 30:2463-2468,2010, which is incorporated herein by reference in its entirety), and polymorphisms or mutations in PIK3CA, MED1, and GAS6 (Murtaza et al., ”Non-invasive analysis of acquired resistance to cancer therapy by sequencing of plasma DNA,” Nature 2013;doi:10.1038 / nature12065,2013, which is incorporated herein by reference in its entirety).

[0362] An increase in cfDNA levels and LOH are associated with a decrease in overall survival and disease-free survival. p53 mutations (exons 5-8) are associated with a decrease in overall survival. A decrease in circulating HER2 cfDNA levels is associated with a good response to HER2-targeted therapy in HER2-positive breast cancer patients. Activating mutations in PIK3CA, shortening of MED1, and splicing mutations in GAS6 result in treatment resistance.

[0363] Examples of polymorphisms or mutations associated with colorectal cancer include mutations in p53, APC, K-ras, and thymidylate synthase, as well as methylation of the p16 gene (Wang et al., “Molecular detection of APC, K-ras, and p53 mutations in the serum of colorectal cancer patients as circulating biomarkers,” World J Surg 28:721-726, 2004; Ryan et al., “A prospective study of circulating mutant KRAS2 in the serum of patients with colorectal neoplasia: strong prognostic indicator in postoperative follow up,” Gut 52:101-108, 2003; Lecomte et al., “Detection of free-circulating tumor-associated DNA in plasma of colorectal cancer patients and its association with prognosis,” Int J Cancer 100:542-548, 2002; Schwarzenbach et al., “Molecular analysis of the polymorphisms of thymidylate synthase on cell-free circulating DNA in blood of patients with advanced colorectal carcinoma,” Int J Cancer 127:881-888, 2009. Each of these references is incorporated herein by reference in its entirety). Postoperative detection of serum K-ras mutations is a strong predictor of disease recurrence. Detection of K-ras mutations and p16 gene methylation is associated with a decrease in survival rate and an increase in disease recurrence. Detection of mutations in K-ras, APC, and / or p53 is associated with recurrence and / or metastasis.Polymorphisms (including LOH, SNP, variable number tandem repeats, and deletions) of the thymidylate synthase (target of fluoropyrimidine-based chemotherapeutic agents) gene using cfDNA may be associated with treatment responsiveness.

[0364] Examples of polymorphisms or mutations associated with lung cancer (e.g., non-small cell lung cancer) include K-ras (e.g., mutations at codon 12) and EGFR mutations. Exemplary predictive mutations include EGFR mutations (deletion in exon 19 or mutation in exon 21) associated with increased overall survival and progression-free survival, and K-ras mutations (mutations at codons 12 and 13) are associated with decreased progression-free survival (Jian et al., “Prediction of epidermal growth factor receptor mutations in the plasma / pleural effusion to efficacy of gefitinib treatment in advanced non-small cell lung cancer,” J Canc Res Clin Oncol 136:1341-1347, 2010; Wang et al., “Potential clinical significance of a plasma-based KRAS mutation analysis in patients with advanced non-small cell lung cancer,” Clin Canc Res 16:1324-1330, 2010. Each of these references is incorporated herein by reference in its entirety). Examples of polymorphisms or mutations that serve as indicators of responsiveness to treatment include EGFR mutations (deletion in exon 19 or mutation in exon 21) that improve responsiveness to treatment, and K-ras mutations (codons 12 and 13) that decrease responsiveness to treatment. Resistance-conferring mutations have been identified in EFGR (Murtaza et al., “Non-invasive analysis of acquired resistance to cancer therapy by sequencing of plasma DNA,” Nature doi:10.1038 / nature12065, 2013. This reference is incorporated herein by reference in its entirety).

[0365] Examples of polymorphisms or mutations associated with melanoma (e.g., uveal melanoma) include polymorphisms or mutations in GNAQ, GNA11, BRAF, and p53. Exemplary mutations in GNAQ and GNA11 include mutations at R183 and Q209. The Q209 mutation in GNAQ or GNA11 is associated with metastasis to bone. The BRAF V600E mutation can be detected in patients with metastatic / advanced melanoma. BRAF V600E is an indicator of invasive melanoma. The presence of the BRAF V600E mutation after chemotherapy agents is associated with non-responsiveness to treatment.

[0366] Examples of polymorphisms or mutations associated with pancreatic cancer include polymorphisms or mutations in K-ras and p53 (e.g., p53 Ser249). p53 Ser249 is also associated with hepatitis B infection and hepatocellular carcinoma, as well as ovarian cancer, and non-Hodgkin lymphoma.

[0367] Even polymorphisms or mutations that are present at low frequency in a sample can be detected using the methods of the present invention. For example, a polymorphism or mutation present at a frequency of 1 in 1 million can be measured 10 times by performing 10 million sequence reads. If desired, the number of sequence reads can be varied according to the desired sensitivity level. In some embodiments, the sample is re-analyzed or another sample from the subject is analyzed using a greater number of sequence reads to improve sensitivity. For example, if no polymorphisms or mutations associated with cancer or associated with an increased risk of cancer are detected, or only a small number (e.g., 1, 2, 3, 4, or 5) are detected, the sample is re-analyzed or another sample is examined.

[0368] In some embodiments, multiple polymorphisms or mutations are required for cancer or metastatic cancer. In such cases, the ability to accurately diagnose cancer or metastatic cancer is improved by screening for the multiple polymorphisms or mutations. In some embodiments where the subject has a subset of the multiple polymorphisms or mutations required for cancer or metastatic cancer, the subject can later be re-screened to determine whether the subject has acquired additional mutations.

[0369] In some embodiments where multiple polymorphisms or mutations are required for cancer or metastatic cancer, the frequency of each of the polymorphisms or mutations can be compared to determine whether they occur at similar frequencies. For example, if two mutations (designated "A" and "B") are required for cancer, some cells have neither, some cells have A, some cells have B, and some cells have both A and B. If A and B are measured at similar frequencies, the subject is more likely to have some cells with both A and B. If A and B are observed at different frequencies, the subject is more likely to have different cell populations.

[0370] In some embodiments where multiple polymorphisms or mutations are required for cancer or metastatic cancer, the number or distinctiveness of such polymorphisms or mutations present in the subject can be used to predict the likelihood that the subject has a disease or disorder, or when. In some embodiments where polymorphisms or mutations tend to occur in a particular order, the subject may be periodically examined to determine whether the subject has acquired other polymorphisms or mutations.

[0371] In some embodiments, by determining the presence or absence of multiple polymorphisms or mutations (e.g., 2, 3, 4, 5, 8, 10, 12, 15, or more), the sensitivity and / or specificity of determining the presence or absence of a disease or disorder such as cancer, or the presence or absence of an increased risk of having a disease or disorder such as cancer, is increased.

[0372] In some embodiments, polymorphisms or mutations are detected directly. In some embodiments, polymorphisms or mutations are detected indirectly by detecting one or more sequences linked to the polymorphism or the mutation (e.g., polymorphic loci such as SNPs).

[0373] 9.1 Exemplary nucleic acid modifications In some embodiments, there are changes in the integrity of RNA or DNA (e.g., size changes in fragmented cfRNA or cfDNA, or changes in nucleosome composition) associated with a disease or disorder such as cancer, or an increased risk of a disease or disorder such as cancer. In some embodiments, there are changes in the methylation patterns of RNA or DNA (e.g., hypermethylation of tumor suppressor genes) associated with a disease or disorder such as cancer, or an increased risk of a disease or disorder such as cancer. For example, methylation of CpG islands in the promoter region of tumor suppressor genes has been suggested to induce local gene silencing. Aberrant methylation of the p16 tumor suppressor gene occurs in subjects with liver cancer, lung cancer, and breast cancer. Other frequently methylated tumor suppressor genes, including APC, Ras association domain family protein 1A (RASSF1A), glutathione S-transferase P1 (GSTP1), and DAPK, have been detected in various types of cancer such as oropharyngeal cancer, colorectal cancer, lung cancer, esophageal cancer, prostate cancer, bladder cancer, melanoma, and acute leukemia. Methylation of certain tumor suppressor genes such as p16 has been reported as an early event in carcinogenesis and is useful for early cancer screening.

[0374] In some embodiments, methylation patterns are determined using a bisulfite conversion method using methylation-sensitive restriction enzyme digestion or a non-bisulfite-based method (Hung et al., J Clin Pathol 62:308-313, 2009, which is incorporated herein by reference in its entirety). In the bisulfite conversion method, methylated cytosine remains as cytosine and unmethylated cytosine is converted to uracil. Methylation-sensitive restriction enzymes (e.g., BstUI) cut unmethylated DNA sequences at specific recognition sites (e.g., 5'-CG v CG-3 for BstUI), while methylated sequences remain intact. In some embodiments, intact methylated sequences are detected. In some embodiments, stem-loop primers are used to selectively amplify restriction enzyme-digested unmethylated fragments without co-amplifying enzymatically undigested methylated DNA.

[0375] 9.2 Examples of Changes in mRNA Splicing In some embodiments, changes in mRNA splicing are associated with a disease or disorder such as cancer, or an increased risk of a disease or disorder such as cancer. In some embodiments, the change in mRNA splicing is present in one or more of the following nucleic acids associated with cancer or an increased risk of cancer: DNMT3B, BRCA1, KLF6, Ron, or Gemin5. In some embodiments, the detected mRNA splice variants are associated with a disease or disorder such as cancer. In some embodiments, multiple mRNA splice variants are produced by healthy cells (e.g., non-cancerous cells, etc.), but a change in the relative amount of mRNA splice variants is associated with a disease or disorder such as cancer. In some embodiments, the change in mRNA splicing is due to a change in the mRNA sequence (e.g., a mutation in the splice site), a change in the level of splicing factors, a change in the amount of available splicing factors (e.g., a decrease in the amount of available splicing factors due to the binding of splicing factors to repeats), a change in splicing control, or the tumor microenvironment.

[0376] The splicing reaction is carried out by a multi - protein / mRNA complex called the spliceosome (Fackenthal1 and Godley, Disease Models & Mechanisms 1:37 - 42, 2008, doi:10.1242 / dmm.000331. This reference is hereby incorporated by reference in its entirety). The spliceosome recognizes intron - exon boundaries, removes intervening introns via two transesterification reactions, and results in the ligation of two adjacent exons. Since incorrect ligation can lead to a reduced likelihood of encoding a normal protein, the fidelity of this reaction must be maximized. For example, when the reading frame of the triplet codons that specify the identity and order of amino acids during translation is maintained by exon skipping, the alternatively spliced mRNA may specify a protein lacking important amino acid residues. More generally, exon skipping disrupts the translation reading frame and generates premature stop codons. These mRNAs are often degraded by at least 90% through a process known as nonsense - mediated mRNA decay, thereby reducing the likelihood that such defective messages will accumulate and produce truncated protein products. If mis - spliced mRNAs avoid this pathway, truncated, mutant, or unstable proteins are produced.

[0377] Alternative splicing is a means of expressing multiple or many different transcripts from the same genomic DNA and results from the incorporation of a subset of exons available for a particular protein. By excluding one or more exons, a domain of a particular protein is lost from the encoded protein, which may result in loss of function or gain of function of the m protein. Several of the following types of alternative splicing have been reported: exon skipping, alternative 5' or 3' splice sites, mutually exclusive exons, and, although very rare, intron retention. Some researchers have used bioinformatics methods to compare the amount of alternative splicing in cancer and normal cells and have determined that cancer exhibits lower levels of alternative splicing than normal cells. Furthermore, the distribution of the types of alternative splicing events was different between cancer and normal cells. Cancer cells showed less exon skipping than normal cells but more alternative 5' and 3' splice site selection and intron retention. When the exonization phenomenon (the use of sequences mainly used as introns in other tissues as exons) was verified, genes associated with exonization in cancer cells were preferentially associated with mRNA processing. This suggests a direct association between cancer cells and the generation of abnormal mRNA splice forms.

[0378] 9.3 Examples of DNA or RNA level changes In some embodiments, there is a change in the total amount or concentration of one or more types of DNA (e.g., cfDNA, cf mDNA, cf nDNA, cellular DNA, or mitochondrial DNA), or RNA (cfRNA, cellular RNA, cytoplasmic RNA, coding cytoplasmic RNA, non-coding cytoplasmic RNA, mRNA, miRNA, mitochondrial RNA, rRNA, or tRNA). In some embodiments, there is a change in the amount or concentration of one or more specific DNA (e.g., cfDNA, cf mDNA, cf nDNA, cellular DNA, or mitochondrial DNA) molecules, or RNA (cfRNA, cellular RNA, cytoplasmic RNA, coding cytoplasmic RNA, non-coding cytoplasmic RNA, mRNA, miRNA, mitochondrial RNA, rRNA, or tRNA) molecules. In some embodiments, one allele is expressed more than another allele at a locus of interest. Exemplary miRNAs are short 20-22 nucleotide RNA molecules that control gene expression. In some embodiments, there are changes in the transcriptome, such as changes in the identity or amount of one or more RNA molecules.

[0379] In some embodiments, an increase in the amount or concentration of cfDNA or cfRNA is associated with a disease or disorder such as cancer, or an increased risk of a disease or disorder such as cancer. In some embodiments, the total concentration of a type of DNA (e.g., cfDNA, cf mDNA, cf nDNA, cellular DNA, or mitochondrial DNA, etc.) or RNA (e.g., cfRNA, cellular RNA, cytoplasmic RNA, coding cytoplasmic RNA, non-coding cytoplasmic RNA, mRNA, miRNA, mitochondrial RNA, rRNA, or tRNA) is at least 2, 3, 4, 5, 6, 7, 8, 9, 10-fold, or more increased compared to the total concentration of that type of DNA or RNA in a healthy (e.g., non-cancerous) subject. In some embodiments, a total concentration of cfDNA of 75 ng / mL or more to 100 ng / mL or less, 100 ng / mL or more to 150 ng / mL or less, 150 ng / mL or more to 200 ng / mL or less, 200 ng / mL or more to 300 ng / mL or less, 300 ng / mL or more to 400 ng / mL or less, 400 ng / mL or more to 600 ng / mL or less, 600 ng / mL or more to 800 ng / mL or less, 800 ng / mL or more to 1,000 ng / mL or less, or a total concentration of cfDNA exceeding 100 ng, e.g., exceeding 200, 300, 400, 500, 600, 700, 800, 900, or 1,000 ng / mL, indicates cancer, indicates an increased risk of cancer, indicates an increased risk of a tumor that is more malignant than benign, indicates a decrease in cancer that has perhaps reached remission, or indicates a worsening prognosis of the cancer. In some embodiments, the amount of a type of DNA (e.g., cfDNA, cf mDNA, cf nDNA, cellular DNA, or mitochondrial DNA) or RNA (cfRNA, cellular RNA, cytoplasmic RNA, coding cytoplasmic RNA, non-coding cytoplasmic RNA, mRNA, miRNA, mitochondrial RNA, rRNA, or tRNA) having one or more polymorphisms / mutations (e.g., deletions or duplications) associated with a disease or disorder such as cancer, or an increased risk of a disease or disorder such as cancer, is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 14, 16, 18, 20, or 25% of the total amount of that type of DNA or RNA.In some embodiments, at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 14, 16, 18, 20, or 25% of the total amount of a certain type of DNA (e.g., cfDNA, cf mDNA, cf nDNA, cellular DNA, or mitochondrial DNA), or RNA (cfRNA, cellular RNA, cytoplasmic RNA, coding cytoplasmic RNA, non-coding cytoplasmic RNA, mRNA, miRNA, mitochondrial RNA, rRNA, or tRNA) has a specific polymorphism or mutation (e.g., deletion or duplication) associated with a disease or disorder such as cancer, or an increased risk of a disease or disorder such as cancer.

[0380] In some embodiments, the cfDNA is encapsulated. In some embodiments, the cfDNA is not encapsulated.

[0381] In some embodiments, the proportion of tumor DNA in total DNA (e.g., the proportion of tumor cfDNA in total cfDNA, or the proportion of tumor cfDNA having a specific mutation in total cfDNA) is determined. In some embodiments, the proportion of tumor DNA may be determined for multiple mutations, where the mutations can be single nucleotide variants, copy number variants, differential methylation, or combinations thereof. In some embodiments, the average tumor proportion calculated for one or a set of mutations having the highest calculated tumor proportion is considered the actual tumor proportion in the sample. In some embodiments, the average tumor proportion calculated for all of the mutations is considered the actual tumor proportion in the sample. In some embodiments, this tumor proportion is used to stage the cancer (since a higher tumor proportion may indicate a more advanced stage of the cancer). In some embodiments, since the proportion of tumor DNA in plasma may correlate with the size of the tumor, the tumor proportion is used for cancer sizing. In some embodiments, the tumor proportion is used to determine the size of the tumor fraction having one or more mutations, since there may be a correlation between the measured tumor proportion in a plasma sample and the size of tissue having a given mutant genotype. For example, the size of tissue having a given mutant genotype may correlate with the proportion of tumor DNA, and the proportion of tumor DNA can be calculated by focusing on the specific mutation.

[0382] 10. Exemplary Biochemical Methods and Compositions 10.1 Sample Examples In some embodiments of any part of the aspects of the present invention, the sample comprises cellular genetic material and / or extracellular genetic material derived from cells that are suspected of having deletions or duplications, such as cells suspected of being cancerous or circulating cells suspected of being of fetal origin. In some embodiments, the sample comprises any tissue or body fluid suspected of containing cells, DNA, or RNA having deletions or duplications. The genetic measurements used as part of these methods may be performed on any sample containing DNA or RNA, for example, but not limited to, samples such as tissue, blood, serum, plasma, urine, hair, tears, saliva, skin, nails, feces, bile, lymph, cervical mucus, semen, tumors, fetuses, or other cells or materials containing nucleic acids. The sample may contain or be used with DNA or RNA from any cell type, or from any cell type (e.g., cells from any organ or tissue suspected of being cancerous or neuronal). In some embodiments, the sample comprises nuclear DNA and / or mitochondrial DNA. In some embodiments, the sample is from any of the target individuals disclosed herein. In some embodiments, the target individual is a cancer patient.

[0383] Exemplary samples include samples containing cfDNA or cfRNA. In some embodiments, cfDNA is available for analysis without the need for a cell lysis step. Cell-free DNA may be obtained from various tissues, such as liquid tissues, for example, blood, plasma, lymph, ascites, or cerebrospinal fluid. In some examples, cfDNA is composed of DNA derived from fetal cells. In some examples, cfDNA is isolated from plasma isolated from whole blood that has been centrifuged to remove cellular material. CfDNA may be a mixture of DNA derived from target cells (e.g., cancer cells) and non-target cells (e.g., non-cancerous cells).

[0384] In some embodiments, the sample contains or is suspected of containing a mixture of DNA (or RNA), such as a mixture of DNA (or RNA) originating from cancer cells and DNA (or RNA) originating from non-cancerous (i.e., normal) cells. In some embodiments, at least 0.5, 1, 3, 5, 7, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 92, 94, 95, 96, 98, 99, or 100% of the cells in the sample are cancer cells. In some embodiments, at least 0.5, 1, 3, 5, 7, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 92, 94, 95, 96, 98, 99, or 100% of the DNA (e.g., cfDNA) or RNA (e.g., cfRNA) in the sample is derived from cancer cells. In various embodiments, the percentage of cells that are cancerous cells in the sample is from 0.5% or more to 99% or less, such as from 1% or more to 95% or less, from 5% or more to 95% or less, from 10% or more to 90% or less, from 5% or more to 70% or less, from 10% or more to 70% or less, from 20% or more to 90% or less, or from 20% or more to 70% or less. In some embodiments, the sample is enriched for cancer cells or for DNA or RNA derived from cancer cells. In some embodiments where the sample is enriched for cancer cells, at least 0.5, 1, 2, 3, 4, 5, 6, 7, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 92, 94, 95, 96, 98, 99, or 100% of the cells in the enriched sample are cancer cells. In some embodiments where the sample is enriched for DNA or RNA derived from cancer cells, at least 0.5, 1, 2, 3, 4, 5, 6, 7, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 92, 94, 95, 96, 98, 99, or 100% of the DNA or RNA in the enriched sample is derived from cancer cells.In some embodiments, cell sorting (e.g., fluorescence-activated cell sorting (FACS), etc.) is used to enrich cancer cells (see Barteneva et al., Biochim Biophys Acta., 1836(1):105-22, Aug 2013.doi:10.1016 / j.bbcan.2013.02.004.Epub 2013 Feb 24, and Ibrahim et al., Adv Biochem Eng Biotechnol. 106:19-39, 2007. Each of these is hereby incorporated by reference in its entirety).

[0385] In some embodiments, the sample is enriched for fetal cells. In some embodiments where the sample is enriched for fetal cells, at least 0.5, 1, 2, 3, 4, 5, 6, 7% or more of the cells in the enriched sample are fetal cells. In some embodiments, the percentage of cells that are fetal cells in the sample is from 0.5% or more to 100% or less, such as from 1% or more to 99% or less, from 5% or more to 95% or less, from 10% or more to 95% or less, from 10% or more to 95% or less, from 20% or more to 90% or less, or from 30% or more to 70% or less. In some embodiments, the sample is enriched for fetal DNA. In some embodiments where the sample is enriched for fetal DNA, at least 0.5, 1, 2, 3, 4, 5, 6, 7% or more of the DNA in the enriched sample is fetal DNA. In some embodiments, the percentage of DNA that is fetal DNA in the sample is from 0.5% or more to 100% or less, such as from 1% or more to 99% or less, from 5% or more to 95% or less, from 10% or more to 95% or less, from 10% or more to 95% or less, from 20% or more to 90% or less, or from 30% or more to 70% or less.

[0386] In some embodiments, the sample comprises one cell or DNA and / or RNA derived from one cell. In some embodiments, multiple individual cells (e.g., at least 5, 10, 20, 30, 40, or 50 cells derived from the same or different subjects) are analyzed in parallel. In some embodiments, cells from multiple samples derived from the same individual are combined, thereby reducing the workload compared to analyzing the samples separately. Combining multiple samples also enables simultaneous examination of multiple tissues for cancer (which can be used to provide a more thorough screening for cancer or to determine whether cancer has metastasized to other tissues).

[0387] In some embodiments, the sample contains one cell or a small number of cells, e.g., 2, 3, 5, 6, 7, 8, 9, or 10 cells. In some embodiments, the sample has 1 to 100, 100 to 500, or 500 to 1,000 cells. In some embodiments, the sample contains RNA and / or DNA in an amount of 1 picogram or more to 10 picograms or less, 10 picograms or more to 100 picograms or less, 100 picograms or more to 1 nanogram or less, 1 nanogram or more to 10 nanograms or less, 10 nanograms or more to 100 nanograms or less, or 100 nanograms or more to 1 microgram or less.

[0388] In some embodiments, the sample is paraffin-embedded. In some embodiments, the sample is preserved using a preservative such as formaldehyde and optionally paraffin-embedded, whereby the DNA is cross-linked and less DNA may be available for PCR. In some embodiments, the sample is a formaldehyde-fixed paraffin-embedded (FFPE) sample. In some embodiments, the sample is a fresh sample (e.g., a sample obtained within 1 or 2 days of analysis). In some embodiments, the sample is frozen prior to analysis. In some embodiments, the sample is a medical history sample.

[0389] These samples can be used in any of the methods of the present invention.

[0390] 10.2 Examples of sample preparation methods In some embodiments, the method includes isolating or purifying DNA and / or RNA. There are many standard methods known in the art for achieving this purpose. In some embodiments, the sample may be centrifuged and various layers may be separated. In some embodiments, DNA or RNA may be isolated using filtration. In some embodiments, the preparation of DNA or RNA may include amplification, separation, purification by chromatography, liquid separation, isolation, preferential concentration, preferential amplification, targeted amplification, or any of many other techniques known in the art or described herein. In some embodiments related to DNA isolation, RNase is used to degrade RNA. In some embodiments related to RNA isolation, DNase (e.g., DNase from Invitrogen, Carlsbad, California, USA) is used to degrade DNA. In some embodiments, RNA is isolated using an RNeasy mini kit (Qiagen) according to the manufacturer's protocol. In some embodiments, low RNA molecules are isolated using a mirVana PARIS kit (Ambion, Austin, Texas, USA) according to the manufacturer's protocol (Gu et al., J. Neurochem. 122:641-649, 2012. This reference is incorporated herein by reference in its entirety). The concentration and purity of RNA are optional and may be determined using a Nanovue (GE Healthcare, Piscataway, New Jersey, USA), and the integrity of RNA is optional and may be measured using a 2100 Bioanalyzer (Agilent Technologies, Santa Clara, California, USA) (Gu et al., J. Neurochem. 122:641-649, 2012. This reference is incorporated herein by reference in its entirety). In some embodiments, RNA is stabilized during storage using TRIZOL or RNAlater (Ambion).

[0391] In some embodiments, a universal-tagged adapter is added to create a library. Prior to ligation, the sample DNA may be blunt-ended and then one adenosine base is added to the 3-prime end. Prior to ligation, the DNA may be cleaved using a restriction enzyme or some other cleavage method. During ligation, ligation efficiency is enhanced by the 3-prime adenosine of the sample fragment and the complementary 3-prime tyrosine overhang of the adapter. In some embodiments, adapter ligation is performed using the ligation kit in the AGILENT SURESELECT kit. In some embodiments, the library is amplified using universal primers. In one embodiment, the amplified library is separated by size fractionation or by using a product such as AGENCOURT AMPURE beads or other similar methods. In some embodiments, target loci are amplified using PCR amplification. In some embodiments, the amplified DNA is sequenced (e.g., sequenced using an ILLUMINA IIGAX or HiSeq sequencer). In some embodiments, the amplified DNA is sequenced from each end of the amplified DNA to reduce sequencing errors. When sequenced from one end of the amplified DNA, if there is a sequence error at a particular base, the likelihood of a sequence error at the complementary base is low when sequenced from the opposite side of the amplified DNA (compared to sequencing the same end of the amplified DNA multiple times).

[0392] In some embodiments, whole genome application (WGA) is used to amplify nucleic acid samples. A number of methods are available for WGA: ligation-mediated PCR (LM-PCR), degenerate oligonucleotide primer PCR (DOP-PCR), and multiple displacement amplification (MDA). In LM-PCR, short DNA sequences called adapters are ligated to the blunt ends of DNA. These adapters contain universal amplification sequences that are used for DNA amplification by PCR. In DOP-PCR, random primers that further contain universal amplification sequences are used in the first round of annealing and PCR. Then, in the second PCR round, the sequences are further amplified with universal primer sequences. MDA uses phi-29 polymerase. This polymerase is a highly processive and non-specific enzyme that has been used for DNA replication and single cell analysis. In some embodiments, WGA is not performed.

[0393] In some embodiments, selective amplification or enrichment is used to amplify or enrich a target locus. In some embodiments, the amplification technique and / or selective enrichment technique may involve PCR, such as ligation-mediated PCR, fragment capture by hybridization, Molecular Inversion Probes, or other circularizing probes. In some embodiments, real-time quantitative PCR (RT-qPCR), digital PCR, or emulsion PCR, single allele base extension reaction followed by mass spectrometry is used (Hung et al., J Clin Pathol 62:308-313, 2009. This document is incorporated herein by reference in its entirety). In some embodiments, capture by hybridization using a hybrid capture probe is used to preferentially enrich DNA. In some embodiments, the method of amplification or selective enrichment involves the use of a probe, in which case when correctly hybridized to the target sequence, the 3-prime or 5-prime end of the nucleotide probe is separated from the polymorphic site of the polymorphic allele by a small number of nucleotides. This separation reduces the preferential amplification of one allele, called allelic bias. This is an improvement over methods involving the use of probes where the 3-prime or 5-prime end of the correctly hybridized probe is directly adjacent to or very close to the polymorphic site of the allele. In one embodiment, probes that may or surely contain the polymorphic site in the hybridization region are excluded. The polymorphic site at the hybridization site may cause non-uniform hybridization and may completely inhibit hybridization in some alleles. This results in preferential amplification of a particular allele. These embodiments are an improvement over other methods involving target amplification and / or selective enrichment in that they better retain the original allelic frequency of the sample at each polymorphic locus, regardless of whether the sample is a pure genomic sample from one individual or from a mixture of individuals.

[0394] In some embodiments, PCR (referred to as miniPCR) is used to generate very short amplicons (U.S. Patent Application No. 13 / 683,604, filed November 21, 2012; U.S. Patent Application Publication No. 2013 / 0123120; U.S. Patent Application No. 13 / 300,235, filed November 18, 2011; U.S. Patent Application Publication No. 2012 / 0270212, filed November 18, 2011; and U.S. Patent Application No. 61 / 994,791, filed May 16, 2014. The patent applications are incorporated herein by reference in their entirety). cfDNA (e.g., cancerous cfDNA released necrotically or apoptotically) is highly fragmented. For fetal cfDNA, the fragment size is approximately Gaussian, with an average of 160 bp, a standard deviation of 15 bp, a minimum size of about 100 bp, and a maximum size of about 220 bp. The polymorphic site of a particular target locus can occupy any position from the start position to the end position among the various fragments originating from that locus. Since cfDNA fragments are short, the likelihood that both primer sites are present is the likelihood of a fragment of length L that includes both the forward primer site and the reverse primer site, which is the ratio of the amplicon length to the fragment length. Under ideal conditions, assays with amplicons of 45, 50, 55, 60, 65, or 70 bp are successful in amplifying from 72%, 69%, 66%, 63%, 59%, or 56%, respectively, of the available template fragment molecules. In a particular related embodiment most preferred for cfDNA from a sample of an individual suspected of having cancer, the cfDNA is amplified using primers that yield a maximum amplicon length of 85, 80, 75, or 70 bp, and in a particularly preferred embodiment 75 bp, and have a melting point of 50-65°C, and in a particularly preferred embodiment 54-60.5°C. The amplicon length is the distance between the 5-prime ends of the forward and reverse priming sites. A shorter amplicon length than typically used by those skilled in the art can result in a more efficient measurement of the desired polymorphic locus by requiring only short sequence reads.In one embodiment, a substantial proportion of the amplicons are less than 100 bp, less than 90 bp, less than 80 bp, less than 70 bp, less than 65 bp, less than 60 bp, less than 55 bp, less than 50 bp, or less than 45 bp.

[0395] In some embodiments, the amplification is performed using direct multiplex PCR, continuous PCR, nested PCR, double nested PCR, one-and-a-half sided nested PCR, full nested PCR, one-sided full nested PCR, one-sided nested PCR, hemi-nested PCR, triple hemi-nested PCR, semi-nested PCR, one-sided semi-nested PCR, reverse semi-nested PCR, or one-sided PCR. These are described in U.S. Patent Application No. 13 / 683,604, filed on November 21, 2012, U.S. Patent Application Publication No. 2013 / 0123120, U.S. Patent Application No. 13 / 300,235, filed on November 18, 2011, U.S. Patent Application Publication No. 2012 / 0270212, and U.S. Patent Application No. 61 / 994,791, filed on May 16, 2014, and these patent applications are hereby incorporated by reference in their entirety. If desired, any of these methods can be used with miniPCR.

[0396] If desired, the extension step of the PCR amplification is limited in terms of time, and amplification of fragments longer than 200 nucleotides, 300 nucleotides, 400 nucleotides, 500 nucleotides, or 1,000 nucleotides may be reduced. This results in enrichment of fragmented DNA or short DNA (e.g., fetal DNA, or DNA from apoptotic or necrotic cancer cells), and may lead to improved test performance.

[0397] In some embodiments, multiplex PCR is used. In one embodiment, a method for amplifying target loci in a nucleic acid sample comprises: (i) contacting the nucleic acid sample with a primer library that simultaneously hybridizes to at least 100, 200, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 20,000, 25,000, 30,000, 40,000, 50,000, 75,000, or 100,000 different target loci to generate a reaction mixture; and (ii) subjecting the reaction mixture to primer extension reaction conditions (such as PCR conditions) to generate an amplification product comprising the target amplicon. In some embodiments, at least 50, 60, 70, 80, 90, 95, 96, 97, 98, 99, or 99.5% of the target loci are amplified. In various embodiments, less than 60, 50, 40, 30, 20, 10, 5, 4, 3, 2, 1, 0.5, 0.25, 0.1, or 0.05% of the amplification product is primer dimer. In some embodiments, the primer is in solution (e.g., dissolved in the liquid phase rather than being on a solid support). In some embodiments, the primer is in solution and not immobilized on a solid support. In some embodiments, the primer is not part of a microarray. In some embodiments, the primer does not include a molecular inversion probe (MIP).

[0398] In some embodiments, two or more (e.g., 3 or 4...

Claims

**Claim 1** A method for monitoring and detecting early recurrence or metastasis of tumors in cancer patients, said method comprising: a) selecting one or more patient-specific mutations based on somatic mutations identified in a tumor sample of a patient diagnosed with cancer; b) long-term collection of one or more blood samples from said patient after said patient has been treated with surgery, first-line chemotherapy, and / or adjuvant therapy; c) generating an amplicon set from cell DNA isolated from one or more circulating cells, presumed to be circulating tumor cells, obtained from a blood sample of said patient, said amplicon set being generated by multiplex amplification of genomic loci encompassing said patient-specific mutations associated with cancer; d) sequencing said amplicon set by next-generation sequencing and determining the origin of said one or more circulating cells based on the presence of said one or more patient-specific mutations in said amplicon set, wherein detection of one or more circulating tumor cells comprising said one or more patient-specific mutations indicates early recurrence or metastasis of said cancer. A method comprising the steps of: **Claim 2** The method according to claim 1, wherein said patient-specific mutations comprise single nucleotide variants (SNVs), copy number variants (CNVs), indels, and / or gene fusions associated with cancer. **Claim 3** The method according to claim 1, wherein said amplicon set is generated by multiplex amplification of genomic loci encompassing at least 8 patient-specific mutations associated with cancer. **Claim 4** The method according to claim 1, wherein said amplicon set is generated by multiplex amplification of genomic loci encompassing at least 16 patient-specific mutations associated with cancer. **Claim 5** The method according to claim 1, wherein the presence of at least 2 patient-specific mutations associated with cancer indicates that said one or more circulating cells are tumor cells. **Claim 6** The method according to claim 1, wherein the presence of at least 5 patient-specific mutations associated with cancer indicates that said one or more circulating cells are tumor cells. **Claim 7** The method according to claim 1, wherein detection of at least 2 circulating tumor cells comprising said one or more patient-specific mutations indicates early recurrence or metastasis of said cancer. **Claim 8** The method of claim 1, comprising dividing a plurality of circulating cells into individual reaction volumes, isolating cell DNA in each reaction volume, and attaching a sample barcode to the cell DNA isolated in each reaction volume.

9. The method of claim 1, wherein the cell DNA is isolated from one circulating cell presumed to be a tumor cell.

10. The method of claim 1, further comprising identifying additional patient-specific mutations associated with cancer from the genotypes of one or more circulating cells determined to be tumor cells.

11. The method of claim 1, wherein the patient has lung cancer.

12. The method of claim 1, wherein the patient has breast cancer.

13. The method of claim 1, wherein the patient has bladder cancer.

14. The method of claim 1, wherein the patient has colorectal cancer.

15. A method for determining the origin of a circulating cell presumed to be a donor cell, comprising: a) generating a first amplicon set from cell DNA isolated from one or more circulating cells presumed to be donor cells obtained from a blood sample of a transplant recipient, and a second amplicon set from cell-free DNA obtained from the plasma fraction of the blood sample, wherein the first amplicon set and the second amplicon set are obtained by performing a multiplex amplification reaction of a plurality of single nucleotide polymorphism (SNP) loci; b) sequencing the first amplicon set and the second amplicon set by next-generation sequencing; and c) determining the origin of the one or more circulating cells based on the sequence of the first amplicon set, wherein the sequence of the second amplicon set is used as a reference.

16. A method for monitoring and detecting early recurrence or metastasis of a tumor in a cancer patient, the method comprising: a) selecting one or more patient-specific mutations based on somatic mutations identified in a tumor sample of a patient diagnosed with cancer; b) longitudinally collecting one or more blood samples from the patient after the patient has been treated with surgery, first-line chemotherapy, and / or adjuvant therapy; c) generating a first amplicon set from cell-free DNA isolated from a blood sample of the patient, wherein the first amplicon set is generated by multiplex amplification of genomic loci comprising the patient-specific mutations associated with cancer; d) sequencing the first amplicon set by next-generation sequencing and detecting the presence of one or more patient-specific mutations in the amplicon set, wherein the presence of the one or more patient-specific mutations indicates early recurrence or metastasis of cancer; e) generating a second amplicon set from cell DNA isolated from one or more circulating cells obtained from a blood sample of the patient, wherein the second amplicon set is generated by multiplex amplification of genomic loci comprising the patient-specific mutations associated with cancer; f) sequencing the second amplicon set by next-generation sequencing and determining the origin of the one or more circulating cells based on the presence of one or more patient-specific mutations in the amplicon set, wherein the presence of the one or more patient-specific mutations indicates that the one or more circulating cells are tumor cells, and g) identifying one or more additional mutations associated with cancer from the sequences of the one or more circulating cells determined to be tumor cells. A method comprising.

17. The method according to claim 16, wherein the patient-specific mutations include single nucleotide variants (SNVs), copy number variants (CNVs), indels, and / or gene fusions associated with cancer.

18. The method according to claim 16, wherein the first and / or second amplicon sets are generated by multiplex amplification of genomic loci comprising at least 8 patient-specific mutations associated with cancer.

19. The method according to claim 16, wherein the first and / or second amplicon sets are generated by multiplex amplification of genomic loci comprising at least 16 patient-specific mutations associated with cancer.

20. The method according to claim 16, wherein in step (d), the presence of at least 2 or at least 5 patient-specific mutations associated with cancer indicates early recurrence or metastasis of cancer.

21. The method of claim 16, wherein in step (f), the presence of at least two or at least five patient-specific mutations associated with cancer indicates that the one or more circulating cells are tumor cells. **Claim 22** The method of claim 16, comprising dividing a plurality of circulating cells into individual reaction volumes, isolating cell DNA in each reaction volume, and attaching a sample barcode to the cell DNA isolated in each reaction volume. **Claim 23** The method of claim 16, wherein the cell DNA is isolated from one circulating cell suspected of being a tumor cell. **Claim 24** The method of claim 16, wherein the patient has lung cancer, breast cancer, bladder cancer, or colorectal cancer. **Claim 25** The method of claim 16, further comprising treating the patient based on the one or more additional mutations associated with cancer.