Determining Nucleic Acid Sequence Imbalances
Digital PCR with cutoff value determination and statistical methods addresses interference from maternal nucleic acids, improving the accuracy and efficiency of noninvasive prenatal diagnosis of fetal chromosomal aneuploidies by determining nucleic acid sequence imbalances.
Patent Information
- Application Number
- JP2023184105
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2007-07-23
- Filing Date
- 2023-10-26
- Publication Date
- 2025-10-29
- Estimated Expiration
- 2028-07-23
AI Technical Summary
Existing noninvasive prenatal diagnostic methods for fetal chromosomal aneuploidies, such as trisomy 21, face challenges due to interference from maternal nucleic acids in maternal plasma, limited fetal nucleic acid concentrations, and reliance on genetic polymorphisms, leading to suboptimal diagnostic accuracy and inefficiency.
A method using digital PCR to determine nucleic acid sequence imbalances by selecting cutoff values based on fetal sequence percentages and average concentrations, employing statistical methods like SPRT to minimize testing requirements and improve accuracy, independent of genetic polymorphisms.
Enhances diagnostic accuracy and efficiency by minimizing the amount of maternal nucleic acid interference, allowing for sensitive and specific detection of fetal chromosomal aneuploidies with reduced sample volume and cost.
Smart Images

Figure 0007761951000006 
Figure 0007761951000007 
Figure 0007761951000008
Abstract
Description
[Technical Field]
[0001] Priority claims This application claims priority from and is a regular application of U.S. Provisional Application No. 60 / 951,438 (Attorney Docket No. 016285-005200US), entitled "DETERMINING A NUCLEIC ACID SEQUENCE IMBALANCE," filed July 23, 2007, the entire contents of which are incorporated herein by reference for all purposes.
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS This application is also related to a concurrently filed regular application (Attorney Docket No. 016285-005220US) entitled "DIAGNOSING FETAL CHROMOSOMAL ANEUPLOIDY USING GEOMIC SEQUENCING," the entire contents of which are incorporated herein by reference for all purposes.
[0003] Technical field of the invention The present invention relates generally to diagnostic testing of genotypes and diseases by determining imbalances between different nucleic acid sequences, and more particularly to fetal Down's syndrome, other chromosomal aneuploidies, mutations, and genotypes through testing of maternal blood samples. The present invention also relates to cancer detection, transplant monitoring, and infectious disease monitoring. [Background technology]
[0004] Genetic diseases, cancer, and other conditions often cause or result in an imbalance between two corresponding chromosomes or alleles, or other nucleic acid sequences, where one gene is more or less than normal relative to the other. Normal ratios are usually exactly 50:50. Down syndrome (trisomy 21) involves an imbalance in the addition of chromosome 21.
[0005] Traditional prenatal diagnostic methods for fetal chromosomal aneuploidies, such as trisomy 21, involve invasive procedures, such as amniocentesis or chorionic villus sampling, to collect fetal samples, which carry the risk of fetal death. Noninvasive procedures, such as ultrasound and biochemical marker screening, have been used in at-risk pregnant women prior to definitive invasive diagnostic procedures. However, these screening methods typically measure epiphenomena associated with trisomy 21 rather than core chromosomal abnormalities, resulting in suboptimal diagnostic accuracy and significant influences by gestational age.
[0006] The discovery of circulating cell-free fetal DNA in maternal plasma in 1997 offered new possibilities for noninvasive prenatal diagnosis (Lo, YMD and Chiu, RWK 2007 Nat Rev Genet 8, 71-77). While this method was readily applied to the prenatal diagnosis of sex-linked genetic disorders (Costa, JM et al. 2002 N Engl J Med 346, 1502) and certain single-gene disorders (Lo, YMD et al. 1998 N Engl J Med 339, 1734-1738), its application to the prenatal detection of fetal chromosomal aneuploidies has presented considerable challenges (Lo, YMD and Chiu, RWK 2007, see above). First, fetal nucleic acids are usually present in maternal plasma, which contains a high background of maternal nucleic acids, which can interfere with the analysis of fetal nucleic acids (Lo, YMD et al. 1998 Am J Hum Genet 62, 768-775). Second, fetal nucleic acids circulate primarily in a cell-free form in maternal plasma, which makes it difficult to obtain information on gene dosage or chromosome dosage in the fetal genome.
[0007] Significant progress has been made in recent years to address these challenges (Benachi, A & Costa, JM 2007 Lancet 369, 440-442). One approach is to detect fetal-specific nucleic acids in maternal plasma, thereby overcoming the problem of maternal background disturbances (Lo, YMD and Chiu, RWK 2007, see above). The dosage of chromosome 21 has been estimated from the ratio of polymorphic alleles in placenta-derived DNA / RNA molecules. However, this method becomes less accurate as the content of the target nucleic acid in the sample decreases, and it can only be applied to fetuses that are heterozygous for the targeted polymorphism. If a single polymorphism is used, such fetuses represent only a small proportion of the population.
[0008] Dhallan et al. (Dhallan, R, et al. 2007, supra; Dhallan, R, et al. 2007 Lancet 369, 474-481) described an alternative method for increasing the proportion of circulating fetal DNA by adding formaldehyde to maternal plasma. The proportion of fetal chromosome 21 sequences in maternal plasma was determined by assessing the ratio of paternally inherited fetal-specific to non-fetal alleles for single nucleotide polymorphisms (SNPs) on chromosome 21. SNP ratios were calculated similarly for the reference chromosome. Fetal chromosome 21 imbalance was then inferred by detecting statistically significant differences between the SNP ratios of chromosome 21 and the reference chromosome, where significance was defined using a fixed p-value of 0.05 or less. To ensure high population coverage, more than 500 SNPs per chromosome were targeted. However, there is debate about the effectiveness of formaldehyde in increasing the proportion of fetal DNA (Chung, GTY, et al. 2005 Clin Chem 51, 655-658), and therefore the reproducibility of this method needs to be further evaluated. Similarly, because fetuses and mothers each provide information on different numbers of SNPs for each chromosome, the statistical power of comparing SNP ratios will vary from case to case (Lo, YMD & Chiu, RWK. 2007 Lancet 369, 1997). Furthermore, because these methods rely on the detection of genetic polymorphisms, they are limited to fetuses that are heterozygous for these polymorphisms.
[0009] Using polymerase chain reaction (PCR) and DNA quantification of the chromosome 21 locus and reference locus in amniotic cell cultures obtained from trisomy 21 and euploid fetuses, Zimmermann et al. (2002 Clin Chem 48, 362-363) were able to distinguish between the two groups of fetuses based on a 1.5-fold increase in the former chromosome 21 DNA sequence. Because a 2-fold difference in DNA template concentration corresponds to a difference of only one threshold cycle (Ct), discrimination of a 1.5-fold difference was the limit of conventional real-time PCR. To achieve a finer degree of quantitative discrimination, alternative methods are needed.
[0010] Digital PCR was developed to detect allele ratio imbalances in nucleic acid samples (Chang, HW et al. 2002 J Natl Cancer Inst 94, 1697-1703). Clinically, it has been shown to be useful in detecting loss of heterozygosity (LOH) in tumor DNA samples (Zhou, W. et al. 2002 Lancet 359, 219-225). To classify whether the experimental results suggest the presence of LOH in the sample, a previous study employed a sequential probability ratio test (SPRT) to analyze digital PCR results (El Karoui et al. 2006 Stat Med 25, 3124-3133). In the method used in the study, a fixed reference ratio of two alleles in 2 / 3 DNA was used as the cutoff value for determining LOH. Because the amount, proportion, and concentration of fetal nucleic acids in amniotic fluid varies, these methods are not suitable for detecting trisomy 21 using fetal nucleic acids in a background of maternal nucleic acids in amniotic fluid.
[0011] A noninvasive test for fetal trisomy 21 (and other imbalances) based on circulating fetal nucleic acid analysis, particularly one that does not rely on the use of genetic polymorphisms and / or fetal-specific markers, is desirable. It is also desirable to accurately determine cutoff values and enumerate sequences to reduce the number of data wells and / or the amount of maternal plasma nucleic acid molecules required for accuracy, thereby increasing efficiency and cost-effectiveness. It is also desirable for the noninvasive test to be highly sensitive and specific to minimize diagnostic errors.
[0012] Another application of detecting fetal DNA in amniotic fluid is the prenatal diagnosis of single-gene disorders such as beta-thalassemia. However, because fetal DNA represents only a small fraction of the DNA in amniotic fluid, this approach is likely to only detect mutations inherited by the fetus from the father and not present in the mother. Examples of such mutations include the 4-bp deletion in codon 41 / 42 of the beta-globin gene, which causes beta-thalassemia (Chiu RWK et al. 2002 Lancet, 360, 998-1000), and the Q890X mutation in the cystic fibrosis transmembrane conductance regulator gene, which causes cystic fibrosis (Gonzalez-Gonzalez et al. 2002 Prenat Diagn, 22, 946-8). However, because both beta-thalassemia and cystic fibrosis are autosomal recessive conditions, a fetus must inherit a mutation from each parent for the disease to develop. Therefore, detecting a paternally inherited mutation alone can only increase a fetus's risk of developing the disease by 25% to 50%. This is not ideal from a diagnostic perspective. That is, applying existing approaches to diagnosis can be considered a successful scenario when no paternally inherited mutation is detected in amniotic fluid, thereby ruling out the possibility that the fetus has a homozygous condition. However, diagnostically, this approach has the disadvantage that the results are based on not detecting the paternal mutation. Therefore, an approach that does not have these limitations and can determine a fetus's complete genotype (homozygous normal, homozygous mutant, or heterozygous) from amniotic fluid is highly desirable. Summary of the Invention [Means for solving the problem]
[0013] A brief overview Embodiments of the present invention provide methods, systems, and devices for determining whether a nucleic acid sequence imbalance (e.g., allelic imbalance, mutational imbalance, or chromosomal imbalance) is present in a biological sample. One or more cutoff values are selected to determine the imbalance, such as the ratio of the amounts of two sequences (or two sets of sequences).
[0014] In one embodiment, the cutoff value is determined, at least in part, based on the percentage of fetal (clinically relevant nucleic acid) sequences in a biological sample containing a background of maternal nucleic acid sequences, such as maternal plasma or serum or urine. In another embodiment, the cutoff value is determined based on the average concentration of the sequences in multiple reactions. In one aspect, the cutoff value is determined from the proportion of informative wells predicted to contain a particular nucleic acid sequence, where the proportion is determined based on the percentage and / or average concentration.
[0015] The cutoff value can be determined using a variety of methods, such as SPRT, false discovery, confidence intervals, receiver operating characteristics (ROC), etc. This strategy also minimizes the amount of testing required before a confident classification can be made, which is particularly important in the analysis of plasma nucleic acids, where template amounts are usually limited.
[0016] In one exemplary embodiment, 1. A method for determining the presence or absence of an imbalance in a nucleic acid sequence in a biological sample, the method comprising: Data is obtained from a plurality of reactions, wherein the data is (1) a first set of quantitative data indicating a first amount of a clinically relevant nucleic acid sequence; and (2) a second set of quantitative data indicating a second amount of a reference nucleic acid sequence that is not a clinically relevant nucleic acid sequence; Includes; determining a parameter from the two sets of data; deriving a first cutoff value from the average concentration of a reference nucleic acid sequence in each of a plurality of reactions, wherein the reference nucleic acid sequence is either a clinically relevant nucleic acid sequence or a reference nucleic acid sequence; comparing the parameter with the first cutoff value; and determining a classification of whether or not an imbalance exists in the nucleic acid sequence based on the comparison; This includes:
[0017] In another exemplary embodiment, 1. A method for determining the presence or absence of an imbalance in a nucleic acid sequence in a biological sample, the method comprising: Data is obtained from a plurality of reactions, wherein the data is (1) a first set of quantitative data indicating a first amount of a clinically relevant nucleic acid sequence; and (2) a second set of quantitative data indicating a second amount of a reference nucleic acid sequence that is not a clinically relevant nucleic acid sequence; wherein the clinically relevant nucleic acid sequence and the reference nucleic acid sequence are derived from a first type of cell and one or more second type of cell; determining a parameter from the two sets of data; deriving a first cutoff value from a first percentage derived from measuring the amount of nucleic acid sequences derived from a first type of cell in the biological sample; comparing the parameter with the cutoff value; and determining a classification of whether or not an imbalance exists in the nucleic acid sequence based on the comparison; This includes:
[0018] Other embodiments of the invention relate to systems and computer-readable media associated with the methods described herein.
[0019] A further understanding of the nature and advantages of the present invention may be obtained by reference to the following detailed description and accompanying drawings. [Brief explanation of the drawings]
[0020] [Figure 1] 1 is a flowchart illustrating a digital DNA experiment.
[0021] [Figure 2A] A digital RNA-SNP and RCD method according to one embodiment of the present invention is described.
[0022] [Figure 2B] Below is a list of chromosomal abnormalities commonly detected in cancer.
[0023] [Figure 3] 1 illustrates a graph with SPRT curves used to determine Down's syndrome, according to one embodiment of the present invention.
[0024] [Figure 4] 1 shows a method for determining a pathological condition using the percentage of fetal cells according to one embodiment of the present invention.
[0025] [Figure 5] 1 illustrates a method for determining a disease state using average concentrations according to one embodiment of the present invention.
[0026] [Figure 6] 1 shows a table of predicted digital RNA-SNP allele ratios and Pr for trisomy 21 samples across a range of template concentrations, expressed as mean reference template concentration per well (mr), according to one embodiment of the present invention.
[0027] [Figure 7] FIG. 1 shows a list of expected Pr for trisomy 21 samples at fractional fetal DNA concentrations of 10%, 25%, 50%, and 100% across a range of template concentrations expressed as the mean reference template concentration (mr) per well, according to one embodiment of the present invention.
[0028] [Figure 8] 1 shows a plot illustrating the degree of difference in SPRT curves when mr values are 0.1, 0.5, and 1.0 in digital RNA-SNP analysis according to one embodiment of the present invention.
[0029] [Figure 9A] 1 shows a table comparing the effectiveness of the old and new SPRT algorithms for classifying euploid and trisomy 21 cases in 96-well digital RNA-SNP analysis, according to one embodiment of the present invention.
[0030] [Figure 9B] 1 shows a table comparing the effectiveness of the old and new SPRT algorithms for classifying euploid and trisomy 21 cases in 384-well digital RNA-SNP analysis, according to one embodiment of the present invention.
[0031] [Figure 10] 1 is a table showing the percentage of correctly and incorrectly classified euploid or aneuploid, and unclassifiable fetuses at a given informative count, according to one embodiment of the present invention.
[0032] [Figure 11] 11 is a table 1100 showing a computer simulation of digital RCD analysis on a pure (100%) DNA sample according to one embodiment of the present invention.
[0033] [Figure 12] 12 is a table 1200 showing the results of a computer simulation of the accuracy of digital RCD analysis performed at mr=0.5 to classify samples from euploid or trisomy 21 fetuses with different fractional concentrations of fetal DNA, according to one embodiment of the present invention.
[0034] [Figure 13]Figure 13A shows a table 1300 of digital RNA-SNP analysis in placental tissue from euploid and trisomy 21 pregnancies, according to one embodiment of the present invention. Figure 13B shows a table 1350 of digital RNA-SNP analysis in maternal plasma tissue from euploid and trisomy 21 pregnancies, according to one embodiment of the present invention.
[0035] [Figure 14] 1 shows a plot illustrating the cutoff curve obtained in an RCD analysis, according to one embodiment of the present invention.
[0036] [Figure 15] Figure 15A shows a table of digital RNA-SNP analysis of placental tissue from euploid and trisomy 21 pregnancies, according to one embodiment of the present invention. Figure 15B shows a table of digital RNA-SNP data for 12 reaction panels from one maternal plasma sample, according to one embodiment of the present invention. Figure 15C shows a table of digital RNA-SNP analysis of maternal plasma from euploid and trisomy 21 pregnancies, according to one embodiment of the present invention.
[0037] [Figure 16] Figure 16A shows a table of digital RNA-SNP analysis of euploid and trisomy 18 placentas, according to one embodiment of the present invention. Figure 16B shows SPRT interpretation of digital RNA-SNP data of euploid and trisomy 18 placentas, according to one embodiment of the present invention.
[0038] [Figure 17] 1 shows a table of digital RCD analysis of 50% placental / maternal blood cell DNA mixtures from euploid and trisomy 21 pregnancies, according to one embodiment of the present invention.
[0039] [Figure 18] 1 shows an SPRT curve depicting the boundary determination of correct and incorrect classification according to one embodiment of the present invention.
[0040] [Figure 19]1 shows a table of digital RCD analysis of amniotic fluid from euploid and trisomy 21 pregnancies, according to one embodiment of the present invention.
[0041] [Figure 20] 1 shows a table of digital RNC analysis of placental DNA samples from euploid and trisomy 18 pregnancies (E=euploid; T18=trisomy 18) according to one embodiment of the present invention.
[0042] [Figure 21] FIG. 1 shows a table of multiplex digital RCD analysis of 50% placental / maternal blood cell DNA mixtures from euploid and trisomy 18 pregnancies (E=euploid; T21=trisomy 21; U=unclassifiable) according to one embodiment of the present invention.
[0043] [Figure 22A] 1 shows a table of multiplexed digital RCD analysis of a 50% euploid or trisomy 21 placental genomic DNA / 50% maternal buffy coat DNA mixture according to one embodiment of the present invention.
[0044] [Figure 22B] 1 shows a table of multiplexed digital RCD analysis of a 50% euploid or trisomy 21 placental genomic DNA / 50% maternal buffy coat DNA mixture according to one embodiment of the present invention.
[0045] [Figure 23] This shows the scenario when both parents carry the same mutation.
[0046] [Figure 24] Figure 24A shows a table of digital RMD analysis of female / male and male / male DNA mixtures, according to one embodiment of the present invention. Figure 24B shows a table of digital RMD analysis of a mixture of 25% female and 75% male DNA, according to one embodiment of the present invention.
[0047] [Figure 25]1 shows a table of digital RMD analysis of 15%-50% DNA mixtures simulating maternal plasma samples for HbE mutations, according to one embodiment of the present invention.
[0048] [Figure 26] Figure 26A shows a table of digital RMD analysis of a 5%-50% DNA mixture simulating maternal plasma samples for CD41 / 42 mutations, according to one embodiment of the present invention. Figure 26B shows a table of digital RMD analysis of a 20% DNA mixture simulating maternal plasma samples for CD41 / 42 mutations, according to one embodiment of the present invention.
[0049] [Figure 27] FIG. 1 shows a block diagram of an example of a computer device useful in the systems and methods according to one aspect of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0050] definition As used herein, the term "biological sample" refers to any sample obtained from a subject (eg, a human, such as a pregnant woman) and which contains one or more nucleic acid molecules of interest.
[0051] The term "nucleic acid" or "polynucleotide" refers to deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) and polymers thereof in either single- or double-stranded form. Unless specifically limited, the term includes nucleic acids containing known natural nucleotide analogs that have similar binding properties as the reference nucleic acid and are metabolized in a manner similar to naturally occurring nucleotides. Unless otherwise indicated, a particular nucleic acid sequence also implicitly includes conservatively modified variants thereof (e.g., degenerate codon substitutions), alleles, orthologs, SNPs, and complementary sequences, as well as the sequence explicitly indicated. Specifically, degenerate codon substitutions can be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and / or deoxyinosine residues (Batzer et al., Nucleic Acid Res. 19:5081 (1991); Ohtsuka et al., J. Biol. Chem. 260:2605-2608 (1985); and Rossolini et al., Mol. Cell. Probes 8:91-98 (1994)). The term nucleic acid is also used interchangeably with gene, cDNA, mRNA, small non-coding RNA, microRNA (miRNA), Piwi-interacting RNA, and short hairpin RNA (shRNA) encoded by a gene or locus.
[0052] The term "gene" refers to a DNA segment involved in producing a polypeptide chain. It may include regions preceding and following the coding region (leader and trailer), as well as intervening sequences between individual coding segments (exons).
[0053] As used herein, the term "reaction" refers to any process involving chemical, enzymatic, or physical action that indicates the presence or absence of a particular polynucleotide sequence of interest. An example of a "reaction" is an amplification reaction, such as the polymerase chain reaction (PCR). Another example of a "reaction" is a sequencing reaction, either by synthesis or ligation. An "informative reaction" refers to a reaction that indicates the presence of one or more particular polynucleotide sequences of interest, and in some cases, a reaction that indicates the presence of only one polynucleotide of interest. As used herein, the term "well" refers to a reaction in a predetermined location within a closed structure, such as a well-shaped vial, cell, or chamber in a PCR array.
[0054] As used herein, the term "clinically relevant nucleic acid sequence" can refer to a polynucleotide sequence corresponding to a fragment of a larger gene sequence whose potential imbalance is being tested, or to a polynucleotide sequence corresponding to the larger gene sequence itself. One example is the sequence of chromosome 21. Other examples include chromosomes 18, 13, X, and Y. Still other examples include mutant gene sequences, polymorphisms, or copy number variations that a fetus may have inherited from one or both of its parents. Still other examples include sequences that are mutated, deleted, or amplified in malignant tumors, such as sequences undergoing loss of heterozygosity or gene duplication. In some embodiments, multiple clinically relevant nucleic acid sequences, or equivalent multiple markers of clinically relevant nucleic acid sequences, can be used to provide data for detecting the imbalance. For example, data from five discontinuous sequences on chromosome 21 can be used in an additive method for determining possible chromosome 21 imbalances, effectively reducing the sample volume by 5-fold.
[0055] The term "reference nucleic acid sequence" as used herein refers to a nucleic acid sequence whose ratio to a clinically relevant nucleic acid sequence is known under normal circumstances, for example, a 1:1 ratio. In one example, the reference nucleic acid sequence and the clinically relevant nucleic acid sequence are two alleles from the same chromosome that can be clearly distinguished by heterozygosity. In another example, the reference nucleic acid sequence is one allele, and the other allele that is heterozygous to it is the clinically relevant nucleic acid sequence. Furthermore, some of each of the reference nucleic acid sequence and the clinically relevant nucleic acid sequence may be derived from different individuals.
[0056] As used herein, the term "reference nucleic acid sequence" refers to a nucleic acid sequence whose average concentration per reaction is known or determined to be equivalent.
[0057] As used herein, the term "overrepresented nucleic acid sequence" refers to a nucleic acid sequence that is more abundant than other sequences in a biological sample and that is contaminated with two sequences of interest (e.g., a clinically relevant sequence and a reference sequence).
[0058] As used herein, the term "based on" means "based at least in part on" and refers to a value (or result) used in determining one value that arises from the relationship between a method's inputs and the method's outputs. The term "derive" as used herein also refers to the relationship between a method's inputs and the method's outputs, such as occurs when the derivation is a formula calculation.
[0059] As used herein, the term "quantitative data" refers to data obtained from one or more reactions and providing one or more numerical values. For example, the number of wells that display a fluorescent marker for a particular sequence would be quantitative data.
[0060] As used herein, the term "parameter" refers to a numerical value that characterizes a set of quantitative data and / or a numerical relationship between quantitative data sets. For example, the ratio (or function of the ratio) between a first amount of a first nucleic acid sequence and a second amount of a second nucleic acid sequence is a parameter.
[0061] As used herein, the term "cutoff value" refers to a numerical value used to distinguish between two or more states (e.g., diseased and non-diseased) in the classification of a biological sample. For example, if a parameter is greater than the cutoff value, a first classification of the quantitative data is made (e.g., diseased); or if the parameter is less than the cutoff value, a different classification of the quantitative data is made (e.g., non-diseased).
[0062] As used herein, the term "imbalance" refers to any significant deviation from the reference amount, which is defined by at least one cutoff value, in the amount of clinically relevant nucleic acid sequence.For example, when the ratio of the reference amount is 3 / 5, and the measured ratio is 1:1, it can be said that imbalance occurs.
[0063] Detailed Description of the Invention The present invention provides methods, systems, and devices for determining whether the amount of a clinically relevant nucleic acid sequence present in a biological sample is increased or decreased (e.g., chromosomal or allelic imbalance) relative to other, non-clinically relevant nucleic acid sequences as compared to a reference (e.g., normal). One or more cutoff values are determined to determine whether an alteration (i.e., imbalance) exists (e.g., the ratio of the amounts of two sequences) relative to the reference amount. The detected alteration in the reference amount can be any deviation (upward or downward) between the clinically relevant nucleic acid sequence and other, non-clinically relevant sequences. Thus, the reference state can be any ratio or other quantity (e.g., other than a 1-1 correspondence), and the measured state indicating an alteration can be any ratio or other quantity that differs from the reference amount as determined by one or more cutoff values.
[0064] The clinically relevant nucleic acid sequences and the reference nucleic acid sequences can be derived from a first type of cell and one or more second type of cell. For example, fetal nucleic acid sequences from fetal / placental cells are present in a biological sample, such as maternal plasma, that contains a background of maternal nucleic acid sequences from maternal cells. In one embodiment, the cutoff value is determined at least in part based on the percentage of cells of the first type in the biological sample. It should be noted that the percentage of fetal sequences in the sample can be determined by any fetal-derived locus and is not limited to measuring clinically relevant nucleic acid sequences. In another embodiment, the cutoff value is determined at least in part based on the percentage of tumor sequences in a biological sample, such as plasma, serum, saliva, or urine, that contains a background of nucleic acid sequences from non-malignant cells in the body.
[0065] In yet another embodiment, the cutoff value is determined based on the average concentration of the sequence in multiple reactions. In one aspect, the cutoff value is determined from the proportion of informative wells that are predicted to contain a specific nucleic acid sequence, where the proportion is determined based on the percentage and / or average concentration. The cutoff value can be determined using various methods, such as SPRT, false positives, confidence intervals, receiver operating characteristics (ROC), etc. This approach also minimizes the amount of testing required before a confidence classification can be made. This is particularly important in the analysis of plasma nucleic acids, where the amount of template is usually limited. Although presented in relation to digital PCR, other methods can also be used.
[0066] In digital PCR, multiple PCR analyses are performed by diluting nucleic acids to the maximum extent possible so that most positive amplifications reflect signals from a single template molecule. Digital PCR therefore allows for the counting of individual template molecules. The percentage of positive amplifications among the total number of PCRs analyzed allows for estimation of the template concentration in the original or undiluted sample. This technique has been shown to be capable of detecting various genetic phenomena (Vogelstein, B. et al. 1999, supra) and has already been used to detect loss of heterozygosity in tumor samples (Zhou, W. et al. 2002, supra) and in cancer patient plasma (Chang, HW et al. 2002, supra). Because template quantification by digital PCR does not have a dose-dependent response relationship between reporter dye and nucleic acid concentration, its analytical precision theoretically exceeds that of real-time PCR. Therefore, digital PCR potentially allows for the discrimination of finer quantitative differences between target and reference loci.
[0067] To test this, we first investigated whether digital PCR can determine the allelic ratio of PLCA4 mRNA (Lo, YMD, et al. 2007 Nat Med 13, 218-223). This mRNA is a placental transcript derived from chromosome 21 and is present in maternal plasma, allowing for the clear distinction between fetal trisomy 21 and euploid fetuses. This approach is called the digital RNA-SNP method. We then investigated whether the increased accuracy of digital PCR would enable the detection of fetal chromosomal aneuploidy independent of genetic polymorphisms. We term this method digital relative chromosome dosage (RCD) analysis. The former approach is polymorphism-dependent but requires low quantitative discrimination, while the latter approach is polymorphism-independent but requires highly accurate quantitative discrimination.
[0068] I. Digital RNA-SNP A. Overview Digital PCR can detect the asymmetry of the abundance ratio of two alleles in DNA samples.For example, it is used to detect loss of heterozygosity (LOH) in tumor DNA samples.Suppose two alleles A and G exist in DNA samples.Suppose allele A can be lost in LOH cells.When LOH occurs in 50% of cells in tumor samples, the allele ratio G:A in the DNA samples can be 2:1.However, if LOH does not occur in tumor samples, the allele ratio G:A can be 1:1.
[0069] Figure 1 shows a flowchart 100 illustrating a digital PCR experiment. In step 110, a DNA sample is diluted and then dispensed into individual wells. Note that the inventors have determined that some nucleic acids are already highly diluted in the original sample. Thus, when the desired template is already present at the required concentration, dilution is not necessary. In previous studies, a DNA sample has been diluted to an average concentration of a specific "template DNA" of approximately 0.5 molecules per well, i.e., approximately one template for every two wells (see, e.g., Zhou et al., 2002). Note that the template DNA corresponds to either the A or G allele, and no rationale is given for this specific concentration.
[0070] In step 120, a PCR process is performed in each well to simultaneously detect the A and / or B alleles. In step 130, a marker (such as fluorescence) in each well is identified to determine whether the well contains either A or G, both, or neither. If LOH does not occur, the amounts of A and G alleles in the DNA sample will be the same (one copy per cell). Therefore, the frequency of wells positive for the A allele and the frequency of wells positive for the G allele will be the same. However, if LOH occurs in 50% or more of the tumor cells, the ratio of G and A alleles will be at least 2:1. According to conventional methods, this sample is simply assumed to be at least 50% cancerous. That is, the frequency of wells positive for the G allele is higher than the frequency of wells positive for the A allele. As a result, the number of wells positive for the G allele will be greater than the number of wells positive for the A allele.
[0071] In step 140, the number of wells that are positive for one of the alleles and negative for the other can be used to classify the results of the digital PCR. In the above example, the number of wells that are positive for allele A but negative for allele G, and the number of wells that are positive for allele G but negative for allele A, are counted. In one embodiment, the allele with fewer positive wells is considered the reference allele.
[0072] In step 150, the total number of informative wells is determined as the sum of the number of wells that are positive for either of the two alleles. In step 160, the proportion of informative wells (P r ) (a type of parameter) is calculated. P r = (number of wells positive for only the allele with more positive wells / total number of wells positive for only one allele (A or G)). In another embodiment, wells with one of the alleles separated by all wells with at least one allele can be used.
[0073] In step 170, P r The aim is to determine whether the value of α indicates allelic imbalance. This task is not trivial, as it requires accuracy and efficiency. One method for determining imbalance is the Bayesian likelihood method, the sequential probability ratio test (SPRT). SPRT is a method that allows two probabilistic hypotheses to be compared by accumulating data. In other words, it is a statistical method for classifying digital PCR results as indicating the presence or absence of allelic skew. This is advantageous for achieving a given statistical power and precision while minimizing the number of wells to be analyzed.
[0074] In an exemplary SPRT analysis, experimental results may be tested against a null hypothesis and an alternative hypothesis. The alternative hypothesis is accepted when there is skew in the allele ratios in the sample. The null hypothesis is accepted when there is no skew in the allele ratios in the sample. r can be compared to two cutoff values that accept either the null or alternative hypothesis. If neither hypothesis is accepted, the sample is marked as unclassifiable, meaning that the observed digital PCR results are insufficient to classify the sample with the desired statistical confidence.
[0075] The cutoff value for accepting the null or alternative hypotheses is typically a fixed P under the assumptions asserted in those hypotheses. r The null hypothesis is that the samples do not exhibit allele ratio distortion. Therefore, the frequency of wells positive for the A allele and the frequency of wells positive for the G allele are the same, and therefore, P r The expected value is 1 / 2. Under the alternative hypothesis, P r The value is expected to be 2 / 3 or halfway between 0.5 and 2 / 3, for example 0.585. Also, due to the limited number of experiments, an upper bound (0.585+3 / N) and a lower bound (0.585-3 / N) can be chosen.
[0076] B. Detection of Down Syndrome In one embodiment of the present invention, digital SNPs are used to detect fetal Down syndrome from maternal plasma. Using fetal / placental cell-specific markers, the allele ratio in chromosome 21 can be measured. For example, SPRT can be used to determine whether the observed degree of overrepresentation of the PLAC4 allele is statistically significant.
[0077] In one exemplary embodiment, digital RNA-SNP determines the imbalance in the ratio of polymorphic alleles of the A / G SNP located on the PLAC4 mRNA, which is transcribed from chromosome 21 and expressed by the placenta. In a heterozygous euploid fetus, the A and G alleles should be equally represented in the fetal genome (genomic ratio 1:1). In trisomy 21, trisomy 21 carries an additional copy of one of the SNP alleles in the fetal genome, resulting in a genomic ratio of 2:1. The purpose of digital PCR analysis is to determine whether the amounts of PLAC4 alleles in the analyzed sample are equal. Therefore, both the A and G alleles of PLAC4 are target templates. A real-time PCR assay was designed to amplify PLAC4 mRNA, and the two SNP alleles were distinguished using TaqMan fluorescent probes. Figure 2A shows a schematic diagram of the analysis process.
[0078] 2A illustrates a digital RNA-SNP method 200 according to an embodiment of the present invention. In step 210, a sample is prepared. In step 220, a nucleic acid sequence, such as PLAC4 mRNA, in an extracted RNA sample is quantified. In one embodiment, this step provides the practitioner with an idea of how much dilution is required for the target to reach the "domain" of digital PCR analysis.
[0079] In step 230, the sample is diluted. In step 240, the concentration of the diluted sample is measured. The concentration of the diluted sample is confirmed to be ∼1 template (i.e., reference or non-reference sequence, or either allele) per well. Some embodiments use the techniques described in Section IV for this determination. For example, the inventors dispensed the diluted sample into 96 wells for real-time PCR and confirmed that the desired dilution was achieved. Since the dilution concentration can remain unknown, embodiments are described below that omit this step.
[0080] In step 250, digital PCR is performed on each well of the array. For example, equally diluted samples are dispensed into 384 wells for real-time PCR analysis. From the PCR results, the amount of marker for each nucleic acid sequence and the number of informative wells are identified. An informative well is a well that is positive for either allele A or B, but not both. In step 260, the expected R r These steps are described in more detail below. The calculation involves determining a parameter from the values determined in step 250. For example, the actual average template concentration per well can be calculated.
[0081] In step 270, an SPRT or other likelihood ratio is performed to determine whether an imbalance exists. In the euploid case, we expect an equal number of A-positive and G-positive wells. However, when analyzing template molecules from a trisomy 21 fetus, the number of wells containing only one allele should be greater than the number of wells containing only one allele of the other. In short, allelic imbalance is expected for trisomy 21.
[0082] As mentioned above, SPRT is a Bayesian likelihood method that allows two probabilistic hypotheses to be compared as data accumulate. In digital PCR analysis for trisomy 21 detection, the alternative hypothesis is accepted when allelic balance exists (i.e., trisomy 21 is detected); the null hypothesis is accepted when allelic imbalance does not exist (i.e., trisomy 21 is not detected). The majority allele is called the potential overrepresented allele, and its proportion among all informative wells (P r ) can be calculated. SPRT is r is applied to determine whether it identifies a sufficient degree of allelic imbalance to be expected in trisomy 21 samples.
[0083] Operationally, SPRT can be applied and interpreted through the use of graphs with pairs of SPRT curves constructed to define probabilistic boundaries for accepting or rejecting either hypothesis. Figure 3 illustrates a graph with SPRT curves for determining Down's syndrome in one embodiment of the invention. These SPRT curves represent the ratio P of the number of informative wells positive for a given overrepresented allele to the total number of informative wells (x-axis), for which confidence intervals can be created. r As depicted in Figure 3, the upper curve sets the probability boundary for accepting the alternative hypothesis; the lower curve sets the probability boundary for accepting the null hypothesis.
[0084] Experimentally derived P r Values are the predicted P r The null hypothesis is then compared to the P value to reject or accept either hypothesis. If the null hypothesis is accepted, the pregnant woman from whom the samples were derived is classified as carrying a euploid fetus. If the alternative hypothesis is accepted, the pregnant woman from whom the samples were derived is classified as carrying a trisomy 21 fetus. Alternatively, if the P value for a given informative count is rWhen the statistical confidence in the disease classification was not reached, no hypothesis could be accepted. These cases were considered unclassifiable unless further data became available. If disease classification was not possible, additional 384-well plates could be run until sufficient data were accumulated to allow classification by SPRT.
[0085] Therefore, SPRT has the advantage over other statistical methods in that fewer tests are required to achieve a given level of confidence. In practical terms, SPRT can accept or reject any hypothesis once a desired amount of data has been accumulated, thereby minimizing unnecessary additional analysis. This property is particularly important in the analysis of nucleic acids, which are generally present at low concentrations in plasma, where the number of available template molecules is limited. In addition to strict classification, classification may also include percentage accuracy. For example, a classification obtained by comparison with a cutoff value indicates the probability that a sample is likely to have a specific percentage of imbalanced nucleic acid sequences, or equivalently, the determined heterogeneity is accurate to a given percentage.
[0086] A similar approach can be applied to determine the fetal genotype for mutations or genetic polymorphisms using fetal nucleic acids in maternal plasma or serum. Recall that a fetus inherits half of its genome from its mother. As an example, consider a particular locus with two alleles, A and B. If the mother has a heterozygous genotype AB, the fetus's genotype could theoretically be AA, BB, or AB. If the fetus's genotype is AB, i.e., identical to the mother's genotype, only nucleic acids of genotype AB (from the mother and fetus) will be present in the maternal plasma. Thus, nucleic acid or allele balance is confirmed in the maternal plasma. On the other hand, if the fetus's genotype is AA or BB, then allelic imbalance will occur in the maternal serum, with either the A or B allele overrepresented. This consideration is also applicable to disease-causing mutations (e.g., beta-thalassemia or spinal muscular atrophy), in which case A can be considered the wild-type allele and B can be considered the mutant allele.
[0087] II. Digital RCD The disadvantage of digital RNA-SNP is that it can only be applied when the analyzed SNP is heterozygous. An improvement would be that a non-invasive test for fetal trisomy 21 or other fetal chromosomal aneuploidies (e.g., trisomy 18, 13, and sex chromosome aneuploidies, etc.) based on circulating fetal nucleic acid analysis would ideally be independent of the use of genetic polymorphisms. Thus, in one embodiment, chromosome dosage is determined by digital PCR analysis of a non-polymorphic chromosome 21 locus relative to a locus located on a reference chromosome (referred to herein as chromosome 1). The ratio of chromosome 21 to chromosome 1 is 2:2 in the genome of a euploid fetus, but this changes in the case of trisomy 21. In digital PCR analysis for detecting trisomy 21, the two hypotheses to be compared can be the null hypothesis that there is no chromosomal imbalance (i.e., trisomy 21 is not detected) and the alternative hypothesis that there is a chromosomal imbalance (i.e., trisomy 21 is detected).
[0088] This approach can be generalized to other chromosomes, including other chromosomal aneuploidies, such as chromosome 18 in trisomy 18, chromosome 13 in trisomy 13, and chromosome X in Turner syndrome. In addition, other chromosomes unrelated to aneuploidy, apart from chromosome 1, can also be used as reference chromosomes. A similar approach can also be applied to cancer detection by analyzing changes in the proportion of chromosomal deletions that typically occur in cancers compared with the reference chromosome. Examples of the former include chromosome 5q in colorectal cancer, chromosome 3p in lung cancer, and chromosome 9p in nasopharyngeal carcinoma. Figure 2B lists several common cancer-associated chromosomal abnormalities that result in sequence imbalances.
[0089] Figure 2A also illustrates a digital RCD method 205 according to an embodiment of the present invention. In one embodiment of steps 220-230, extracted DNA is quantified, e.g., by Nanodrop technology, and diluted to a concentration where each well contains approximately one target template from either chromosome 21 or a normalizing chromosome (e.g., chromosome 1). In one embodiment of step 240, confirmation can be performed by analyzing the diluted DNA sample with an assay using chromosome 1 only in a 96-well format to determine whether less than 37% of wells are negative before proceeding with RCD analysis using both TaqMan probes in a 384-well plate. The significance of 37% is discussed in Section IV.
[0090] The testing of step 240 and the results of step 250 can be performed using a real-time PCR assay designed to amplify paralogous sequences present on both chromosomes, as identified by the paralogous sequence variations distinguished by the pair of TaqMan probes. In this context, an informative well is defined as one that is positive for either the chromosome 21 or chromosome 1 locus, but not both. In a euploid fetus, the number of informative wells positive for either locus should be approximately equal. In a trisomy 21 fetus, wells positive for chromosome 21 should be over-represented relative to those for chromosome 1. The exact percentage of over-representation is discussed in a later section.
[0091] III. Incorporation of fetal sequence percentage A disadvantage of embodiments of methods 200 and 205 described above is the need for a fetal-specific marker. Therefore, in one embodiment of the invention, a non-fetal-specific marker is used. To use such a non-fetal-specific marker, an embodiment of the invention measures the fractional concentration of fetal DNA in maternal plasma (i.e., the biological sample). This information is used to determine the fractional concentration of fetal DNA, as described below. r A more useful numerical value of is calculated.
[0092] Although the fractional percentage of fetal DNA in maternal plasma is small, a trisomy 21 fetus contributes an additional amount of chromosome 21 per genome equivalent (GE) of fetal DNA to maternal plasma. For example, a sample of plasma from a euploid pregnancy with 50 GE / ml of total DNA, of which 5 GE / ml is fetal (i.e., a fractional concentration of fetal DNA of 10%), should contain a total of 100 copies of chromosome 21 sequence per ml of sample plasma (90 copies from the sample + 10 copies from the fetus). In a trisomy 21 pregnancy, each fetal GE contains three copies of chromosome 21, resulting in a total of 105 copies of chromosome 21 sequence in maternal plasma (90 copies from the sample + 15 copies from the fetus). Thus, at a 10% fetal DNA concentration, the amount of chromosome 21 sequence in maternal plasma from a trisomic pregnancy should be 1.05 times that of a euploid pregnancy. Therefore, if analytical approaches are developed to determine small quantitative variations, a polymorphism-independent test for non-invasive prenatal diagnosis of fetal trisomy 21 may be achieved.
[0093] Therefore, the degree of overrepresentation may depend on the fractional concentration of fetal DNA in the DNA sample being analyzed. For example, when placental DNA is analyzed, the theoretical RCD ratio in the fetal genome should be 3:2, or 1.5-fold different. However, as noted above, when a maternal plasma sample containing 10% fetal DNA is analyzed, the theoretical RCD ratio may drop to 1.05-fold. Experimentally derived P r is calculated by dividing the number of wells that are positive for only the chromosome 21 locus by the total number of informative wells. r is the calculated P r and the theoretical RCD ratio are subjected to SPRT analysis.
[0094] FIG. 4 illustrates a method 400 for determining a pathological condition using the percentage of fetal nucleic acid in one embodiment of the present invention. In step 410, the differential percentage of fetal-derived molecules is measured. In one embodiment, the differential percentage is determined by measuring the amount of fetal-specific markers (e.g., Y chromosomes, genetic polymorphisms (e.g., SNPs), placental epigenetic signatures, etc.) relative to non-fetal-specific markers (i.e., gene sequences present in both the mother and the fetus). Accurate measurements can be made using real-time PCR, digital PCR, sequencing reactions (including massively parallel genome sequencing), or any other quantitative method. In one aspect, gene targeting, which can potentially result in allelic imbalances by itself, is preferably not used.
[0095] In step 420, digital PCR or other measurement methods are performed, which involves diluting the sample, dispensing the diluted sample into wells, and measuring the reaction in each well. In step 430, the PCR results are used to identify markers of different reference nucleic acid sequences (chromosomes or alleles). In step 440, the exact proportion of overrepresented sequences (P r In step 450, a cutoff value for determining a pathological condition is calculated using the percentage of fetal-derived molecules in the sample. In step 460, the accurate P r The presence or absence of imbalance is determined based on the cutoff value.
[0096] In one embodiment, the differential percentage of the reference nucleic acid sequence is incorporated into the digital RNA-SNP method. Thus, when investigating LOH due to cancer cells, this can be done when the cancer cells in the tumor sample are less than 50%. More accurate P rTumor samples with 50% or more cancer cells can also be used to obtain a genomic DNA fragment, which reduces the number of false positives that result in misdiagnosis. In another embodiment, the percentage of fetal nucleic acid is incorporated into a digital PCR method to determine whether the fetus has inherited a parental genetic mutation (such as those that cause cystic fibrosis, beta-thalassemia, or spinal muscular atrophy) or polymorphism from analysis of maternal plasma nucleic acid.
[0097] IV. Average concentration of incorporation per well Another drawback of the traditional method (see, for example, Zhou, W. et al. 2002, supra) is that the template concentration must be adjusted to one template per well. This can lead to errors if the exact concentration is difficult to determine. Furthermore, even if the exact concentration is adjusted to one template per well, the traditional method, i.e., the old algorithm, requires a higher expected P to accept the alternative hypothesis. r Values are allelic ratios and are independent of the average template concentration per well.
[0098] However, due to natural statistical variations in templates in diluted samples, there may not actually be one template per well. An embodiment of the present invention measures the concentration of at least one sequence and determines whether that measurement falls within a cutoff value, i.e., the expected P value. r In one embodiment, this calculation involves a statistical distribution to determine the probability of a well containing different nucleic acid sequences, which is used to calculate the expected P r is used to determine
[0099] In one embodiment, the average concentration is obtained from one reference nucleic acid sequence, which is, for example, a nucleic acid sequence that exists at a lower concentration in the DNA sample.If there is no imbalance, the concentrations of the two sequences in the sample should be the same, and one of them is considered as the reference allele.If the sample has, for example, LOH, the allele that is deleted in cancer cells can be considered as the reference allele.The average concentration of the reference allele is m rIn another embodiment, the sequence present at a higher concentration serves as the reference sequence.
[0100] A. Examples of using digital SNP / SPRT and digital PCR 5 shows a method 500 for determining a disease state using an average template concentration according to one embodiment of the present invention. In step 510, the amount of different sequences is measured. This can be done, for example, by counting markers in a digital PCR experiment, as described above. However, this method can also be done by other methods that do not include an amplification step or use fluorescent markers, but can use physical properties such as mass, specific optical properties, or other properties such as base pairing.
[0101] In step 520, the actual proportion of the overrepresented sequence is determined. This can be done by taking the number of wells that show only that sequence, as described above, and distinguishing by the number of informative wells. In step 530, the average concentration of at least one sequence (the reference sequence) is measured. In one embodiment, this reference sequence is an overrepresented sequence. In another embodiment, this reference sequence is an underrepresented sequence. The measurement can be made by counting the number of wells in a digital PCR experiment that are negative for the reference sequence. The relationship between the proportion of negative wells and the average template concentration is described by the Poisson distribution, as described in the next subsection.
[0102] In step 540, the expected amount of wells positive for different sequences is calculated, for example using a Poisson distribution. This expected amount can be the probability of a sequence per well, the average sequence per well, the number of wells containing that sequence, or any other suitable amount. In step 550, the expected P r is calculated from the predicted quantity. In step 560, a cutoff value is calculated based on the predicted P r , for example, calculated using SPRT. In step 570, the imbalance classification of the nucleic acid sequence is determined. Particular aspects of method 500 are now described.
[0103] Determining the expected amount of sequence Once the average concentration per well is known from step 530, the expected number of wells that will show that sequence is calculated in step 540. This quantity can be expressed as a %, a decimal value, or an integer value. For purposes of illustration, the average concentration (m r ) is 0.5, and the genotype of the fetus with trisomy 21 at PLAC4 SNP, rs8130833, is AGG. Therefore, the reference template is the A allele and the overrepresented template is the G allele.
[0104] In one embodiment, the Poisson distribution is taken to be the distribution of the A allele in the reaction mixture of a well of an assay procedure such as digital PCR. In other embodiments, other distribution functions are used, such as the binomial distribution.
[0105] The Poisson equation is: [ka] where n is the number of template molecules per well; P(n) is the probability of a template molecule in a particular well; and m is the average number of template molecules in a well in a particular digital PCR experiment.
[0106] Therefore, if allele A has an average concentration of 0.5, the probability that a well does not contain any molecules of allele A is: [ka] This becomes:
[0107] Therefore, the probability that any well contains at least one molecule of the A allele would be 1-0.6065=0.3935. Approximately 39% of the wells would therefore be expected to contain at least one molecule of the A allele.
[0108] For the non-reference nucleic acid sequence, the genomic ratio of A to G in each cell of a trisomy 21 fetus is 1:2. Assuming this A to G ratio is unchanged in the extracted RNA or DNA sample, the average concentration of the G allele per well is twice that of the A allele, i.e., 2 x 0.5 = 1.
[0109] Therefore, if the mean concentration of the G allele is 1, the probability that a well contains no molecules of the G allele is: [ka] This becomes:
[0110] Therefore, the probability that any well contains at least one molecule of the G allele would be 1-0.3679=0.6321. Approximately 63% of the wells would therefore be expected to contain at least one molecule of the G allele.
[0111] 2. Determining the percentage of overrepresented sequences After the expected amount is calculated, the percentage of over-represented nucleic acid sequences can be determined. Assuming that the alleles A and B loaded into the wells are independent, the probability of a well containing both alleles is 0.3935 x 0.6321 = 0.2487, so approximately 25% of the wells are expected to contain both alleles.
[0112] The proportion of wells containing the A allele but no G allele is the probability of containing at least one A allele minus the probability of containing both the A and B alleles: 0.3935-0.2487=0.1488. Similarly, the proportion of wells containing the G allele but no A allele is: 0.6321-0.2487=0.3834. Informative wells are defined as wells that are positive for either the A or B allele, but not both.
[0113] Therefore, the expected proportion of wells containing either the A or G allele in digital RNA-SNP is 0.1488 / 0.3834. In other words, there are 2.65 times as many wells that are only positive for the G allele as there are wells that are only positive for the A allele. This contrasts with the proportion of the fetal genome where the overrepresented allele is twice as likely as the other allele.
[0114] In the SPRT analysis, the proportion of informative wells positive for the overrepresented allele (P r ) is calculated and interpreted using the SPRT curve. In the example so far, the fraction of informative wells is: 0.1448 + 0.3834 = 0.5852. Therefore, m r 0.5, the expected P for trisomy 21 r is: 0.3834 / 0.5282=0.73.
[0115] Since the mean template concentration (m) is the key parameter in the Poisson equation, P r varies with m. Figure 6 shows the average reference template concentration per well (m r ) and P of trisomy 21 samples with different ranges of template concentrations. r Table 600 shows a table 600 listing the average reference template concentration (m r ) and the proportion of informative wells positive for the overrepresented allele (P r ), and the expected allelic ratios are shown.
[0116] Expected P r Values are the mean concentration of the reference allele per well (m r ) to accept the alternative hypothesis. r The expected value of m r The expected P for accepting the null hypothesis increases as r The value of m is fixed at 0.5, and samples with or without allelic imbalance are rAs increases, P r The difference in the amount of two different nucleic acid sequences can be determined on a case-by-case basis. Note that in other embodiments, the value for accepting the null hypothesis can be other than 0.5. This may be the case when the normal ratio is not 1:1, but for example, 5:3, and therefore, when the 5:3 ratio changes, an imbalance occurs. Therefore, the difference in the amount of two different nucleic acid sequences can be determined on a case-by-case basis.
[0117] However, conventional methods (see, e.g., Zhou, W. et al. 2002, supra) do not provide the expected P r Because we used P values, they were the P values for samples with LOH (accepting the alternative hypothesis). r The underestimation was due to the m r In other words, the higher the average concentration of the reference allele in the DNA sample, the less accurate the traditional method becomes. r A low estimate of can lead to inaccurate calculation of both the cutoff values for accepting the null and alternative hypotheses.
[0118] 3. Expected P r Calculation of cutoff value based on In embodiments using SPRT, the equations in El Karoui at al. (2006) for calculating the upper and lower limits of the SPRT curve boundary can be used. Furthermore, the level of statistical confidence preferred for accepting the null or alternative hypotheses can be varied by adjusting the threshold likelihood ratio of the equations. In this context, the threshold likelihood ratio is set to 8, as this value has been shown to provide good performance in distinguishing samples with and without allelic imbalance in cancer detection. Thus, in one embodiment, the equations for calculating the upper and lower limits of the SPRT curve boundary are: Upper boundary = [(ln8) / n-lnδ] / lnγ Lower boundary = [(ln1 / 8) / n-lnδ] / lnγ wherein: δ=-(1-θ1) / (1-θ0) γ=-(θ1(1-θ0) / θ0(1-θ1) θ0 = Proportion of informative wells containing the non-reference allele when the null hypothesis is true = 0.5 (see below) θ1 = the proportion of informative wells containing the non-reference (i.e., overrepresented) allele when the alternative hypothesis is true N = number of informative wells = number of wells that are positive for only one allele where ln is the natural logarithm, i.e., log e is the mathematical symbol for
[0119] To determine θ for accepting the null hypothesis, assume that the sample was obtained from a pregnant woman carrying a euploid fetus. Under this assumption, the expected number of wells positive for either template should be 1:1, and therefore the expected proportion of informative wells containing the non-reference allele should be 0.5.
[0120] To determine θ1 for accepting the alternative hypothesis, assume that the sample is obtained from a pregnant woman carrying a trisomy 21 fetus. The expected P for trisomy 21 obtained for digital RNA-SNP analysis is r The calculation of is detailed in table 600. Therefore, θ1 for digital RNA-SNP analysis refers to the data shown in the last column of table 600.
[0121] 4. Measurement of Average Concentration m r Measurement of m can be performed through a variety of mechanisms known or that may become known to those skilled in the art. r The value of m is determined during the experimental process of digital PCR analysis. r Since the relationship between the value of m and the total number of wells positive for the reference allele is governed by a distribution (e.g., Poisson distribution), r can be calculated from the number of wells positive for the reference allele, where m rThe formula used is: = -ln(1 - proportion of wells positive for the reference allele), where ln is the natural logarithm, i.e., log e This approach is based on the m r Provide a direct and accurate estimate of
[0122] This method can be used to achieve a desired concentration. For example, similar to step 240 of method 200, the extracted sample nucleic acid can be diluted to a specific concentration, such as one template molecule per reaction well. In one embodiment using a Poisson distribution, the expected proportion of wells that do not receive template is e -m where m is the average template molecule concentration per well. For example, at an average concentration of 1 template molecule per well, the expected fraction of wells that do not receive a template molecule is e -1 , or 0.37 (37%). The remaining 63% of wells may contain one or more template molecules. Typically, digital PCR is performed and the number of positive and informative wells is counted. The definition of an informative well and the method for interpreting digital PCR data are instrument-dependent.
[0123] In other embodiments, the average concentration m per well r is measured by other quantitative methods, such as quantitative real-time PCR, semi-quantitative competitive PCR, and real competitive PCR using mass spectrometry.
[0124] B. Digital RCD Digital PCR using average concentrations can be performed in a similar manner to the digital SNP method described above. The reference chromosome (non-chromosome 21) marker, the chromosome 21 marker, and the number of wells positive for both markers can be determined by digital PCR. The average concentration of the reference marker per well (m r ) is m in digital SNP analysis r As in the calculation of (1), it follows a Poisson distribution function and is calculated from the total number of wells that are negative for the reference marker, regardless of whether they are positive for the chromosome 21 marker.
[0125] SPRT analysis can then be used to classify plasma samples obtained from pregnant women carrying euploid or trisomy 21 fetuses. In this scenario, the expected ratio of wells positive for the reference marker and the chromosome 21 marker is 1:1, so the expected proportion of informative wells showing a positive signal for chromosome 21 is 0.5. When the fetus is trisomic for chromosome 21, the alternative hypothesis can be accepted. In this scenario, if the sample DNA is only of fetal origin, the average concentration of chromosome 21 in each well will be 0.5 times the average concentration of the reference marker (m r ) can be 3 / 2 times larger.
[0126] Although digital RCD can be used to determine chromosome dosage through the detection of fetal-specific markers, such as placental epigenetic signatures (Chim, SSC. et al. 2005 Proc Natl Acad Sci USA 102, 14753-14758), one embodiment of this digital RCD analysis uses non-fetal-specific markers. Therefore, when non-fetal-specific markers are used, an additional step is added to measure the percentage of fetal-derived molecules. Thus, the average concentration of chromosome 21 per well depends on the proportion of fetal DNA in the sample and, therefore, r It can be calculated using [(200% + fetal DNA percentage) / 200%].
[0127] To illustrate, we will once again use a specific example: Assume the average concentration of the reference chromosome, chromosome 1 () per well is 0.5, and assume that 50% of the DNA in the sample is of fetal origin and 50% of the DNA is of maternal origin.
[0128] Thus, using the Poisson distribution, the proportion of wells that do not contain any molecules at the chromosome 1 locus when the mean concentration per well is 0.5 is: [ka] This becomes:
[0129] Therefore, the probability that a well contains at least one molecule of the chromosome 1 locus is: 1-0.6065=0.3935. Therefore, approximately 39% of wells can be expected to contain at least one molecule of the locus.
[0130] In each cell of this trisomy 21 fetus, the genomic ratio of chromosome 21 to chromosome 1 is 3:2. The ratio between chromosome 21 and chromosome 1 in a DNA sample depends on the fractional fetal DNA concentration (fetal DNA%): 3 x fetal DNA% + 2(1 - fetal DNA%): 2 x fetal DNA% + 2 x (1 - fetal DNA%). Thus, in this case where the fractional fetal DNA concentration is 50%, the ratio is: (3 x 50% + 2 x 50%) / (2 x 50% + 2 x 50%) = 1.25. If the signal SNP method did not use fetal-specific markers, such calculations could also be used to calculate the average concentration of non-reference sequences.
[0131] Therefore, when the average concentration of chromosome 1 loci per well is 0.5, the average concentration of chromosome 21 per well is: 1.25 x 0.5 = 0.625. Therefore, the probability that no molecules of the chromosome 21 locus will enter the well when the average concentration per well is 0.625 is: [ka] This becomes:
[0132] Therefore, the probability that any well contains at least one chromosome 21 locus is: 1 - 0.5353 = 0.4647. Therefore, approximately 46% of wells are expected to contain at least one molecule of that locus. Assuming that each locus is independent of the other, the probability that a well contains both loci is: 0.3935 x 0.4647 = 0.1829. Therefore, approximately 18% of wells are expected to contain both loci.
[0133] The expected proportion of wells containing a chromosome 1 locus but no chromosome 21 locus is the number of wells containing at least one chromosome 1 locus minus the number of wells containing both loci: 0.3935-0.1829=0.2126. Similarly, the expected proportion of wells containing a chromosome 21 locus but no both loci is: 0.4647-0.1829=0.2818. Informative wells are defined as wells that are positive for either a chromosome 1 locus or a chromosome 21 locus, but not both.
[0134] Therefore, the expected ratio of chromosome 21 to chromosome 1 in digital RCD is 0.2818 / 0.2106 = 1.34. In other words, the proportion of wells that are positive for only the chromosome 21 locus is 1.34 times the proportion of wells that are positive for only the chromosome 1 locus. This contrasts with a ratio of 1.25 in the DNA sample.
[0135] The proportion of informative wells positive for the chromosome 21 locus in SPRT analysis (P r ) needs to be calculated and interpreted using the SPRT curve. In this case, the fraction of informative wells is: 0.2106 + 0.2818 = 0.4924. Therefore, m r 0.5 and the expected P for trisomy 21 with 50% fetal DNA r is: 0.2818 / 0.4924=0.57.
[0136] Since the mean template concentration (m) is the key parameter in the Poisson equation, P r varies with m. Figure 7 shows the average reference template concentration per well (m r ) in trisomy 21 samples with fractional fetal DNA concentrations of 10%, 25%, 50%, and 100% at a range of template concentrations expressed as r Expected P of Trisomy 21 Specimens for Digital RCD Analysis rThe calculation of θ is detailed in Table 700. Therefore, θ for digital RCD analysis of samples with various fetal DNA fractional concentrations is calculated using the corresponding expected P r The value can be obtained from the column that indicates the value.
[0137] C. Results 1. Different m r Comparison of A measure of the difference in the degree of allelic or chromosomal imbalance between theoretically derived (as in the fetal genome) and experimentally predicted imbalances, and the m in the latter case. r The calculations for determining the m value are shown in Tables 600 and 700. In the digital RNA-SNP analysis of trisomy 21 samples, r When m = 0.5, the ratio of wells containing only the overrepresented allele to wells containing only the reference allele, i.e., the digital RNA-SNP ratio, is 2.65 (Table 600). In the digital RCD analysis of a sample consisting of 100% fetal specimens, m r When m = 0.5, the ratio of wells that are positive for only the chromosome 1 locus to wells that are positive for only the chromosome 1 locus, i.e., the digital RCD ratio, is 0.63 / (1-0.63) = 1.7). As the fractional fetal DNA concentration decreases, m r The digital RCD ratio also decreases until it is the same (Table 700).
[0138] As shown in Tables 600 and 700, the degree of allele or chromosome overrepresentation is m r However, the percentage of informative wells increases with m r =0.5, and then m r As the number of wells increases, the efficiency gradually decreases. In fact, the decrease in the percentage of informative wells can be compensated for by increasing the total number of wells if the amount of template molecules in the sample is unlimited. However, adding wells increases the cost of reagents. Therefore, the optimal digital PCR efficiency is determined by a trade-off between template concentration and the total number of wells tested per sample.
[0139] 2. Examples of using SPRT curves As discussed above, the expected degree of allelic or chromosomal imbalance for a digital PCR experiment depends on the actual template concentration per reaction mixture (e.g., well). We use the template concentration based on the reference allele, i.e., the average reference template concentration per well (m r ) is written. As shown in the formula above, the expected P r can be used to plot the upper and lower SPTR curves. r is m r The plot of the SPRT curve depends essentially on the value of m r Therefore, in practice, the P obtained from a particular experiment may depend on the r To interpret the actual m of a digital PCR dataset, r It will be necessary to use a set of SPRT curves related to
[0140] FIG. 8 shows a schematic diagram of a digital RNA-SNP analysis system according to one embodiment of the present invention. r 8 shows a plot 800 illustrating the degree of difference in the SPRT curves for values of 0.1, 0.5, and 1.0. Each set of digital PCR data represents the exact m for a particular experiment. r The experimentally derived P values should be interpreted using specific curves related to the expected degree of allelic or chromosomal imbalance for the digital RNA-SNP and RCD approaches, which are different (2:1 for the former and 3:2 for the latter), necessitating different sets of SPRT curves for the two digital PCR systems. r is the corresponding m of the digital PCR reaction r This is in contrast to previously reported use of SPRT for molecular detection of LOH by digital PCR, which uses a fixed set of curves.
[0141] The actual method for interpreting digital PCR data using SPRT is explained below using a hypothetical digital RNA-SNP reaction. After digital RNA-SNP analysis in each case, the number of wells that are positive for only the A allele, the number of wells that are positive for only the G allele, or the number of wells that are positive for both alleles is counted. The reference allele is defined as the allele with the smaller number of positive wells. r The value of is calculated using the total number of wells negative for the reference allele, regardless of whether the other allele is positive, according to a Poisson probability density function. Data for our hypothetical example are shown below.
[0142] In a 96-well reaction, 20 wells are positive for only the A allele, 24 wells are positive for only the G allele, and 33 wells are positive for both. Since there are fewer A-positive wells than G-positive wells, allele A is considered the reference allele. The number of wells negative for the reference allele is 96-20-33=43. Therefore, using the Poisson equation, m r is calculated as: -ln(43 / 96) = 0.80. The experimentally determined P r is: 24 / (20+24)=0.55.
[0143] According to Table 600, m r = 0.8, the expected P for trisomy 21 samples r is 0.76. Therefore, in this case, θ1 is 0.76. The SPRT curve based on θ1=0.76 is r (0.55 in this case) can be used to interpret P r = 0.55 falls on the associated SPRT curve, the data point falls on the lower curve. Therefore, this case is classified as euploid. See Figure 3.
[0144] 3. Comparison with conventional methods FIG. 9A shows a table 900 comparing the effectiveness of the old and new SPRT algorithms for classifying euploid and trisomy 21 samples in a 96-well digital RNA-SNP analysis. FIG. 9B shows a table 950 comparing the effectiveness of the old and new SPRT algorithms for classifying euploid and trisomy 21 samples in a 384-well digital RNA-SNP analysis. The new algorithm uses m derived from digital PCR data. r This means selecting a specific SPRT curve for each reaction. The old algorithm means using a fixed set of SPRT curves for all digital PCR reactions. The effect of imprecise calculation of cutoff values on classification accuracy is demonstrated by the simulation analysis shown in Table 900.
[0145] As shown in Tables 900 and 950, the percentage of unclassifiable data is significantly lower in our approach compared to using a fixed set of SPRT curves in previous studies. r Using our approach with ρ = 0.5, 14% and 0% of trisomy 21 samples were unclassifiable in 96-well and 384-well digital RNA-SNP analysis, respectively, whereas using a fixed curve, 62% and 10% were unclassifiable, respectively (Table 900). Thus, our approach allows for disease classification with a smaller number of informative wells.
[0146] As shown in Table 900, the new algorithm r In the case where m is any value between 0.1 and 2.0, the sample can be classified more accurately depending on whether or not the allele ratio is distorted. For example, m r When performing 96-well digital RNA-SNP reactions with ρ = 1.0, the new algorithm classified samples with and without allele ratio skew with 88% and 92% accuracy, respectively, whereas using the old algorithm, the percentage of samples with and without allele ratio skew was only 19% and 36%, respectively.
[0147] Using this new algorithm, the accuracy of separation of samples with and without allele ratio distortion is m r As a result, the classification accuracy can be increased by m r m r As m increases beyond 2.0, the effect of improving classification accuracy in separating samples into two groups decreases because the percentage of informative wells decreases. On the other hand, when using the old algorithm, classification accuracy is m r decreases significantly as increases, because the deviation of the predicted P value from the true value increases.
[0148] Our experimental and simulation data demonstrate that digital RNA-SNP is an effective and accurate method for detecting trisomy 21. Because PLAC4 mRNA in maternal plasma is purely fetal, a single 384-well digital PCR experiment was sufficient for accurate classification in 12 of 13 maternal plasma samples (Table 1350 in Figure 13B). This homozygous, real-time digital PCR-based approach offers an alternative to mass spectrometry-based approaches for RNA-SNP analysis (Lo, YMD, et al. 2007 Nat Med, supra). Aside from placenta-specific mRNA transcripts, we also envision that other types of fetal-specific nucleic acid species in maternal serum could be used for digital PCR-based detection of fetal chromosomal aneuploidies. One example is fetal epigenetic markers (Chim, SSC et al. (2005) Proc Natl Acad Sci USA 102, 14753-14758; Chan, KCA et al. (2006) Clin Chem 52, 2211-2218), which have recently been used for non-invasive prenatal detection of trisomy 18 using the epigenetic allele ratio (EAR) approach (Tong, YK et al. (2006) Clin Chem 52, 2194-2202). This leads us to predict that digital EAR is a possible analytical technique.
[0149] Increased V.%, multiple markers, and PCR alternatives As noted above, applying embodiments of the present invention to DNA extracted from maternal plasma can be complicated when the fractional concentration of fetal DNA in maternal plasma is as low as approximately 3% between 11 and 17 weeks of gestation. However, as described herein, digital RCD enables aneuploidy detection even when aneuploid DNA is present as a minority population. The lower the fractional concentration of fetal DNA, such as in early pregnancy, the greater the number of informative counts required by digital RCD. The significance of this work is that we have provided a set of benchmark parameters, e.g., fractional fetal DNA and total template molecules, on which diagnostic experiments can be based, as summarized in Table 1200 of Figure 12. In our view, a total of 7680 responses at a fractional fetal DNA concentration of 25% is a particularly attractive set of benchmark parameters. These parameters allow classification of euploid and trisomy 21 samples with 97% accuracy at a fractional fetal DNA concentration of 25%, as shown in table 1200.
[0150] The number of plasma DNA molecules present per unit volume of maternal plasma is limited (Lo, YMD. et al. 1998 Am J Hum Genet 62, 768-7758). For example, in early pregnancy, the median maternal plasma concentration of an autosomal locus (β-globin) is 986 copies / ml, including both fetal and maternal loci (Lo, YMD. et al. 1998 Am J Hum Genet 62, 768-7758). To capture 7680 molecules, DNA extracted from approximately 8 ml of maternal plasma would be required. This plasma volume can be obtained from approximately 15 ml of maternal blood, which is within the limits of routine clinical practice. However, we envision that multiple sets of chr21 and reference chromosome targets can be combined for digital RCD analysis. For five pairs of chr21 and reference chromosome targets, just 1.6 mL of plasma material may be required to obtain the required number of template molecules for analysis. Multiplex single-molecule PCR can be performed. The robustness of this multiplex single-molecule analysis has already been demonstrated in single-molecule haplotyping (Ding, C. and Cantor, CR. 2003 Proc Natl Acad Sci USA 100, 7449-7453).
[0151] Alternatively, methods to selectively amplify fetal DNA in maternal plasma (Li, Y. et al. 2004 Clin Chem 50, 1002-1011) or suppress maternal DNA background (Dhallan, R et al. 2004 JAMA 291, 1114-1119), or both, can be employed to achieve a fractional fetal DNA concentration of 25%. In addition to physical methods such as amplifying fetal DNA and suppressing maternal DNA, it may also be possible to use molecular amplification methods, such as targeting fetal DNA molecules that exhibit specific DNA methylation patterns (Chim, SSC et al, 2005 Proc Natl Acad Sci USA 102, 14753-14758, Chan, KCA et al. 2006 Clin Chem 52, 2211-2218; Chiu, RWK et al. 2007 Am J Pathol 170, 941-950.).
[0152] In addition, currently, there are many alternative approaches to manually adjust digital real-time PCR analysis, which are used in research to carry out digital PCR. These alternative approaches include microfluidics digital PCR chips (Warren, L et al. 2006 Proc Natl Acad Sci USA 103, 17807-17812; Ottesen, EA et al. 2006 Science 314, 1464-1467), emulsion PCR (Dressman, D et al. 2003 Proc Natl Acad Sci USA 100, 8817-8822), and massively parallel genome sequencing using the Roche 454 platform, the Illumina Solexa platform, and Applied Biosystems' SOLiD™ system (Margulies, M. et al. 2005 Nature 437, 376-380). Regarding the latter, our approach can also be applied to massively parallel sequencing methods for single DNA molecules, which do not require an amplification step, such as Helicos True Single Molecule DNA Sequencing technology, Pacific Biosciences' Single Molecule Real-Time (SMRT™) technology, and nanopore sequencing (Soni GV and Meller A. 2007 Clin Chem 53, 1996-2001). Using these methods, digital RNA-SNP and digital RCD can be performed rapidly on large numbers of samples, thereby improving the clinical flexibility of the methods proposed for non-invasive prenatal diagnosis. [Example]
[0153] The following examples are offered by way of illustration, and are not intended to limit the claimed invention. I. Computer Simulation A computer simulation was performed to estimate the accuracy of trisomy 21 diagnosis using the SPRT approach. The computer simulation was performed using Microsoft Excel 2003 software (Microsoft Corp., USA) and SAS 9.1 for Windows software (SAS Institute Inc., NC, USA). The performance of digital PCR was evaluated using a reference template concentration (m r ), the number of informative counts and the estimated degree of allelic or chromosomal imbalance (P r ) and the interaction between them. Simulations were performed by varying each of these variables. Because the decision boundaries of the SPRT curves for digital RNA-SNP and digital RCD are different, the simulation analyses for these two methods were performed separately.
[0154] For each assumed digital PCR condition (i.e., fetal DNA fraction concentration, total number of wells), two rounds of simulation were performed. In the first round, a scenario was assumed in which test samples were obtained from pregnant women carrying euploid fetuses. In the second round, a scenario was assumed in which test samples were obtained from pregnant women carrying trisomy 21 fetuses. In each round, 5,000 fetuses were tested.
[0155] A. RNA-SNP In digital RNA-SNP, m r =0.1~m r A simulation was performed for a 384-well experiment where m = 2.0. r For this value, we assumed a scenario in which 5,000 euploid fetuses and 5,000 trisomy 21 fetuses were tested. To classify 10,000 fetuses, we used a given m r 10 is a table 1000 showing the percentage of fetuses that were correctly and incorrectly classified as euploid or aneuploid, as well as those that were unclassifiable, by a given informative count in accordance with an embodiment of the present invention. rWhen m is between 0.5 and 2.0, the accuracy of diagnosing both sex polyploids and aneuploids is 100%. r When = 0.5, only 57% and 88% of sex polyploid and trisomy 21 fetuses, respectively, were correctly classified after analysis of 384 wells.
[0156] These simulation data were generated by the steps described below.
[0157] In step 1, for each well, two random numbers representing the A and G alleles, respectively, were generated using a random (Poisson) function in the SAS program (www.sas.com / technologies / analytics / statistics / index.html). The random (Poisson) function generated positive integers starting from 0 (i.e., 0, 1, 2, 3, etc.), and the probability of each integer being generated was determined according to the probability of each integer according to the Poisson probability density function for a given mean value, which represents the average concentration of alleles per well. If the random number representing the A allele was greater than zero, the well was considered to be positive for the A allele, i.e., containing one or more molecules of the A allele. Similarly, if the random number representing the G allele was greater than zero, the well was considered to be positive for the G allele.
[0158] In the scenario of a pregnant woman carrying a euploid fetus, the same mean value was used to generate random numbers for the A allele and the G allele. For example, m r In analyses assuming digital RNA-SNP analysis with ρ = 0.5, the mean values for both the A and G alleles were set to 0.5. This means that the mean concentration of either allele is 0.5 molecules per well. Using the Poisson equation, the proportion of wells positive for either the A or B allele at a mean concentration of 0.5 can be the same, 0.3935. See Table 600.
[0159] m in pregnant women with trisomy 21 rAssuming digital RNA-SNP analysis with ρ = 0.5, the average concentration per well of the overrepresented allele can be expected to be twice the value of the reference allele, i.e., 1. Under this condition, the probability of a well being positive for the overrepresented allele was 0.6321. See Table 600.
[0160] After generating random numbers for digital PCR wells, the wells can be classified into one of the following states: a. Negative for both A and G alleles b. Both A and G alleles are positive c. Allele A is positive but G is negative d. Allele G is positive but A is negative
[0161] In step 2, step 1 was repeated until the desired number of wells (384 wells in this simulation) was reached. The number of wells positive for only the A allele or only the G allele was counted. The allele with fewer positive wells was considered the reference allele, and the allele with more positive wells was considered the potentially overrepresented allele. The number of informative wells was taken as the total number of wells positive for either allele but not both. The proportion of informative wells containing potentially overrepresented alleles (P r ) was calculated. In accordance with one embodiment of the present invention, the upper and lower boundaries on the SPRT curve for accepting the null or alternative hypotheses were calculated.
[0162] In step 3, 5,000 simulations were performed for each of two scenarios: a pregnant woman carrying a euploid or trisomy 21 fetus. Each simulation can be considered an independent biological sample obtained from the pregnant woman. In Table 1000, a correct classification of a euploid sample refers to a euploid sample for which the null hypothesis was accepted, and an incorrect classification of a euploid sample refers to a euploid sample for which the alternative hypothesis was accepted. Similarly, a trisomy 21 sample for which the alternative hypothesis was accepted is considered correctly classified, and a trisomy 21 sample for which the null hypothesis was accepted is considered incorrectly classified. In either group, samples for which neither the null nor the alternative hypothesis was accepted after a pre-specified total number of wells had been simulated were considered unclassifiable.
[0163] In step 4, m r was set in the range of 0.1 to 2.0, increasing by 0.1, and steps 1 to 3 were carried out.
[0164] B.RCD FIG. 11 shows m r 11 is a table 1100 showing computer simulations of digital RCD performed on pure (100%) fetal DNA samples with m ranging from 0.1 to 2.0. As the fractional fetal DNA concentration decreases, the degree of overrepresentation of chromosome 21 decreases, thereby requiring a larger number of informative wells to confirm disease classification. Therefore, the simulations are based on m r Further runs were performed at fetal DNA concentrations of 50%, 25%, and 10% with a range of total well numbers from 384 to 7680 wells at =0.5.
[0165] FIG. 12 illustrates a method for classifying samples from euploid or trisomy 21 fetuses with different fractional concentrations of fetal DNA in one embodiment of the present invention. rTable 1200 shows the results of a computer simulation of the accuracy of digital RCD analysis at ρ = 0.5. The efficiency of digital RCD improves as the sample's DNA fractional concentration increases. With a fetal DNA concentration of 25% and a total of 7680 PCR analyses, 97% of both euploid and aneuploid samples were classified successfully with no misclassifications. The remaining 3% of samples require further analysis to achieve classification.
[0166] The procedure for simulating digital RCD analysis was similar to that described for digital RNA-SNP analysis. The steps of this simulation are described below.
[0167] In step 1, two random numbers were generated under a Poisson probability density function to represent the reference locus, chromosome 1, and chromosome 21 loci. In subjects carrying euploid fetuses, the mean concentrations of chromosomes 1 and 21 were equal. In this simulation analysis, the mean template concentration per well for each locus was set to 0.5. In subjects carrying trisomy 21 fetuses, the m r was 0.5, but the average concentration of chromosome 21 per well depended on the fractional DNA concentration being tested in use, as shown in Table 700. The allocation of reference loci and / or chromosome 21 loci to a single well was determined by random numbers representing each locus generated according to a Poisson probability density function, at the appropriate average concentration of loci per well.
[0168] In step 2, step 1 was repeated until the desired number of wells was reached, for example, 384 wells in a 384-well plate experiment. The number of wells positive for only chromosome 1 and only chromosome 21 was counted. The number of informative wells was defined as the total number of wells positive for either chromosome but not both. The proportion of informative wells positive for chromosome 21 (P r) was calculated. As described above in the SPRT analysis section, in accordance with one embodiment of the present invention, upper and lower boundaries on the SPRT curve for accepting the null or alternative hypotheses were calculated.
[0169] In step 3, 5,000 simulations were performed for each of two scenarios: a pregnant woman carrying a euploid or trisomy 21 fetus. Each simulation can be considered an independent biological sample obtained from the pregnant woman. In Table 1100, a correct classification of a euploid sample refers to a euploid sample for which the null hypothesis was accepted, and an incorrect classification of a euploid sample refers to a euploid sample for which the alternative hypothesis was accepted. Similarly, a trisomy 21 sample for which the alternative hypothesis was accepted is considered correctly classified, and a trisomy 21 sample for which the null hypothesis was accepted is considered incorrectly classified. In either group, samples for which neither the null nor the alternative hypothesis was accepted after a pre-specified total number of wells had been simulated were considered unclassifiable.
[0170] In step 4, steps 1 to 3 were repeated for samples with fetal DNA concentrations of 10%, 25%, 50%, and 100% for a total number of wells ranging from 384 to 7680 wells.
[0171] II. Modified Methods for Detecting Trisomy 21 A.RNA-SNP for PLAC4 The rs8130833 SNP in PLAC4 on chromosome 21 (Lo, YMD et al. 2007 Nat Med 13, 218-223) was used to demonstrate the practical feasibility of digital RNA-SNP. Placental DNA and RNA from two euploid and two trisomy 21 heterozygous placentas were analyzed. These placental DNA samples were analyzed using the digital RNA-SNP protocol, but the reverse transcription step was omitted, essentially converting the procedure to digital DNA-SNP analysis. To balance the probability of correctly classifying specimens and the proportion of informative wells, the inventors diluted samples to contain one allele of any type per well and verified this by 96-well digital PCR analysis. This was confirmed by a 384-well digital RNA-SNP experiment. r and m r is calculated, and this m r The SPRT curve for was used for data interpretation.
[0172] FIG. 13A shows table 1300, which is the result of digital RNA-SNP analysis of placental tissue from pregnant women carrying euploid and trisomy 21 fetuses, in one embodiment of the present invention. Genotypes were determined by mass spectrometry assay. "Euploid" refers to experimentally obtained P r "T21" for trisomy 21 is assigned when experimentally obtained P r Each of these specimens was correctly classified in a single 384-well experiment using both DNA and RNA samples.
[0173] The inventors further tested plasma RNA collected from nine pregnant women carrying euploid fetuses and four pregnant women carrying trisomy 21 fetuses. Figure 13B shows Table 1350, which shows the results of digital RNA-SNP analysis of maternal plasma collected from pregnant women carrying euploid and trisomy 21 fetuses in one embodiment of the present invention. All samples were correctly classified. The initial results for one trisomy 21 sample (M2272P) fell within the unclassifiable region of the SPRT curve. Therefore, an additional 384-well experiment was performed. From the data collected from a total of 768 wells, new m r and P r is calculated, and this m r Classification was performed using a new set of SPRT curves selected based on the values, and the samples were correctly scored as aneuploid.
[0174] Our experimental and simulation data demonstrate that digital RNA-SNP is an effective and accurate method for detecting trisomy 21. Because PLAC4 mRNA in maternal plasma is purely fetal, only one 384-well digital PCR experiment is sufficient to correctly classify 12 out of 13 maternal plasma samples. Therefore, this homogeneous, real-time digital PCR-based approach offers an alternative to mass spectrometry-based approaches for RNA-SNP analysis. Apart from placenta-specific mRNA transcripts, we envision that other types of fetal-specific nucleic acid species in maternal serum could be used for digital PCR-based detection of fetal chromosomal aneuploidies. One example is the fetal epigenetic marker recently used for noninvasive prenatal detection of trisomy 18 using the epigenetic allele ratio (EAR) approach (Tong YK et al. 2006 Clin Chem, 52, 2194-2202). Therefore, the inventors anticipate that digital EAR may be a possible analysis technique.
[0175] B.RCD The actual feasibility of digital RCD for detecting trisomy 21 was also investigated using PCR assays targeting paralogous sequences on chromosomes 21 and 1. Paralogous loci were used here by way of example. Non-paralogous sequences on chromosome 21, as well as any other reference sequences, may also be used for RCD. Placental DNA samples from two euploid and two trisomy 21 placentas were diluted to a concentration of approximately one target template from either chromosome per well, which was confirmed by 96-well digital PCR analysis. Each confirmed sample was analyzed in a 384-well digital RCD experiment, and P r and m r In digital RCD, the paralog of chromosome 1 was used as the reference template. r The values were used to select the corresponding set of SPRT curves for data interpretation. As shown in Figure 14A, all placental samples were correctly classified.
[0176] To demonstrate that the digital RCD approach can be used to detect trisomy 21 DNA contaminated with excess euploid DNA, such as fetal DNA in maternal plasma, mixtures containing 50% and 25% trisomy 21 placental DNA in a background of euploid maternal blood cell DNA were analyzed. Placental DNA from 10 trisomy 21 specimens and 10 euploid specimens was each mixed with an equal amount of euploid blood cell DNA, resulting in 20 50% DNA mixtures. Figure 14B shows plot 1440 illustrating SPRT interpretation for RCD analysis of these 50% fetal DNA mixtures, in one embodiment of the invention. Similarly, placental DNA from 5 trisomy 21 specimens and 5 euploid specimens was each mixed with 3 times the amount of euploid blood cell DNA, resulting in 10 25% DNA mixtures. Figure 14C shows a plot 1470 illustrating the SPRT interpretation for the RCD analysis of these 25% fetal DNA mixtures. As shown in Figures 14B and 14C, all euploid and aneuploid DNA mixtures were correctly classified.
[0177] Each sample reached a point where it could be analyzed after many 384-well digital PCR analyses, as marked in Figures 14B and 14C. For a 50% DNA mixture, the number of 384-well plates required ranged from 1 to 5. For a 25% DNA mixture, the number of 384-well plates required ranged from 1 to 7. The cumulative percentage of correctly classified specimens after adding all 384 digital PCR analyses was comparable to that predicted by computer simulation, as shown in Table 1200.
[0178] III. Digital PCR Method A. Digital RNA-SNP All RNA samples were first reverse transcribed using ThermoScript reverse transcriptase (Invitrogen) with a gene-specific reverse transcription primer. The sequence of the reverse transcription primer was 5'-AGTATATAGAACCATGTTTAGGCCAGA-3' (Integrated DNA Technologies, Coralville, IA). Subsequently, reverse-transcribed RNA (i.e., cDNA) samples and DNA samples (e.g., placental DNA) for digital RNA-SNP were treated essentially the same way. Prior to digital DNA analysis, DNA and cDNA samples were first quantified using a real-time PCR assay for PLAC4. The primers used were 5'-CCGCTAGGGTGTCTTTTAAGC-3' and 5'-GTGTTGCAATACAAAATGAGTTTCT-3', and the fluorescent probe was 5'-(FAM)ATTGGAGCAAATTC(MGBNFQ)-3' (Applied Biosystems, Foster City, CA). FAM is 6-carboxyfluorescein, and MGBNFQ is a minor groove to which a non-fluorescent quencher binds.
[0179] A standard curve was prepared by serially diluting an HPLC-purified single-stranded synthetic DNA oligonucleotide encoding the amplicon. The sequence was 5'-CGCCGCTAGGGTGTCTTTTAAGCTATTGGAGCAAATTCAAATTTGGCTTAAAGAA AAAGAAACTCATTTTGTATTGCAACACCAGGAGTATCCCAAGGGACTCG-3'. Reactions were set up in a 25 μl reaction volume using 2X TaqMan Universal PCR Master Mix (Applied Biosystems). 400 nL of each primer and 80 nM of probe were used in each reaction. Reactions were initiated at 50°C for 5 minutes, followed by 95°C for 10 minutes and 45 cycles of 95°C for 15 seconds and 65°C for 1 minute, using an ABI PRISM 7900HT Sequence Detection System (Applied Biosystems). Serial dilutions of DNA or cDNA samples were then performed so that subsequent digital PCR amplifications could be performed with approximately one template molecule per well. At such concentrations, approximately 37% of the reaction wells were expected to be negative for amplification, which was confirmed first by 96-well digital real-time PCR analysis, followed by digital RNA-SNP analysis performed in 384-well plates using a set of non-intron-spanning primers (forward primer 5'-TTTGTATTGCAACACCATTTGG-3', gene-specific reverse primers previously described).
[0180] Two allele-specific TaqMan probes were designed, targeting each of the two alleles of the rs8130833 SNP on PLAC4. The sequences were 5'-(FAM)TCGTCGTCTAACTTG(MGBNFQ)-3' and 5'-(VIC)ATTCGTCATCTAACTTG(MGBNFQ) for the A and G alleles, respectively. Reactions were set up in a 5 μl reaction volume using 2X TaqMan Universal PCR Master Mix. Each reaction contained 1X TaqMan Universal PCR Master Mix, 572 nM of each primer, 107 nM of the G allele-specific probe, and 357 nM of the A allele-specific probe. Reactions were performed in an ABI PRISM 7900HT Sequence Detection System. Reactions began with 2 min at 50°C, followed by 10 min at 95°C, followed by 45 cycles of 15 s at 95°C and 1 min at 57°C. During the reaction, fluorescence data were collected using the "Absolute Quantification" application of SDS 2.2.2 software (Applied Biosystems). This software automatically calculated baseline and threshold values. The number of wells positive for either the A or G allele was recorded, and the results were used for SPRT analysis.
[0181] B. Digital RCD analysis All placental and maternal buffy coat DNA samples used in this study were first quantified using a NanoDrop spectrophotometer (NanoDrop Technology, Wilmington, DE). DNA concentrations were converted to copies / ml using a conversion of 6.6 pg / cell. DNA concentrations corresponding to approximately one template per well were determined by serial dilution of the DNA samples and confirmed by real-time PCR in a 96-well format, with approximately 37% of wells showing positive amplification. PCRs for the confirmation plate were prepared similarly, except that only probes for the reference chromosomes were added. In digital RCD analysis, paralogous loci on chromosomes 21 and 1 (Deutsch, S. et al. 2004 J Med Genet 41, 908-915) were first simultaneously amplified with the forward primer 5'-GTTGTTCTGCAAAAAACCTTCGA-3' and the reverse primer 5'-CTTGGCCAGAAATACTTCATTACCATAT-3'. Two chromosome-specific TaqMan probes were designed to target the chromosome 21 and 1 paralogs, with the sequences 5'-(FAM)TACCTCCATAATGAGTAA A(MGBNFQ)-3' and 5'-(VIC)CGTACCTCTGTAATGTGTAA(MGBNFQ)-3', respectively. Each reaction contained 1x TaqMan Universal PCR Master Mix (Applied Biosystems), 450 nM of each primer, and 125 nM of each probe. The total reaction volume was 5 μL / well. The reaction began with 2 min at 50°C, followed by 10 min at 95°C, followed by 50 cycles of 15 s at 95°C and 1 min at 60°C. All real-time PCR experiments were performed on an ABI PRISM 7900HT Sequence Detection System (Applied Biosystems), and fluorescence data were collected using the "Absolute Quantification" application in SDS 2.2.2 software (Applied Biosystems). Predefined reference values and manually entered thresholds were used.The number of wells positive for either chromosome 21 or 1 was recorded and submitted for SPRT analysis. More than one 384-well plate must be analyzed before SPRT classification is possible.
[0182] IV. Use of Microfluidics-Based Digital PCR A. Digital RNA-SNP This example demonstrates the performance of digital PCR analysis using microfluidics-based digital PCR. One method shown in this section uses the Fluidigm BioMark™ system, which is capable of performing over 9000 digital PCR reactions in a single reaction.
[0183] Placental tissue and maternal peripheral blood samples were obtained from pregnant women carrying euploid or trisomy 21 fetuses. Genotyping of the rs8130833 SNP in the PLAC4 gene was performed on placental DNA samples by primer extension followed by mass spectrometry. RNA was extracted from placental and maternal plasma samples.
[0184] All RNA samples were reverse transcribed with a gene-specific reverse transcription primer (5'-AGTATATAGAACCATGTTTAGGCCAGA-3') using ThermoScript reverse transcriptase (Invitrogen). Serial dilutions were performed on placental cDNA samples to achieve a concentration of approximately one template molecule per well for subsequent digital PCR amplification.
[0185] Digital PCR was performed using a BioMark System™ (Fluidigm) with a 12.765 digital array (Fluidigm). Each digital array consisted of 12 panels to accommodate 12 different sample assay mixtures. Each panel was further divided into 765 wells to accommodate 7-nL reactions per well. The rs8130833 SNP region on the PLAC4 gene was amplified using a forward primer (5'-TTTGTATTGCAACACCATTTGG-3') and the gene-specific reverse transcription primer. Two allele-specific TaqMan probes were designed to target each of the two alleles of the rs8130833 SNP. The sequences were 5'-(FAM)TCGTCGTCTAACTTG(MGBNFQ)-3' for the G allele and 5'-(VIC)ATTCGTCATCTAACTTG(MGBNFQ) for the A allele. Reactions for one array panel were prepared using TaqMan Universal PCR Master Mix in a 10 μL reaction volume. Each reaction contained 1x TaqMan Universal PCR Master Mix, 572 nM of each primer, 53.5 nM of the G allele-specific probe, 178.5 nM of the A allele-specific probe, and 3.5 μL of cDNA template. One reaction panel was used for each placental cDNA sample, while 12 panels were used for each maternal plasma sample. The sample assay mixture was injected into the digital array using a NanoFlex™ IFC controller (Fluidigm). Reactions were run in a BioMark™ System. The reaction began with 2 minutes at 50°C, followed by 10 minutes at 95°C, followed by 40 cycles of 15 seconds at 95°C and 1 minute at 57°C.
[0186] Placental RNA samples from one euploid and two T21 heterozygous placentas were analyzed in a 765-well reaction panel. In each sample, informative wells positive for either the A or G allele (but not both) were counted. The proportion of overrepresented alleles among all informative wells (P rThe average reference template concentration per well (m r ) is the SPRT curve that fits the experimentally obtained P r was used to determine whether the chromosomes indicated polyploid or T21 samples.
[0187] We further tested plasma RNA samples from four women carrying euploid fetuses and one woman carrying a trisomy 21 fetus. Each sample was analyzed in 12 765-well reaction panels per plasma RNA sample, i.e., 9,180 reactions. Figure 15B shows the number of informative wells in each of the 12 panels for this plasma RNA sample. As shown in the table, the template concentration in the plasma sample was so dilute that informative wells in any of the reaction panels were insufficient for SPRT classification. To classify this sample as euploid, informative wells from three reaction panels had to be combined (Figure 15C). Figure 15C demonstrates that all plasma specimens could be correctly classified using pooled data from 2 to 12 panels.
[0188] Compared to manually performing digital PCR, this microfluidics-based method is much faster and less labor-intensive: the entire process can be completed within two and a half hours.
[0189] Digital RNA-SNP analysis for prenatal detection of trisomy 18 In this example, we used a digital PCR-based allele discrimination assay for serpin protease inhibitor clade B (ovalbumin) member 2 (SERPINB2) mRNA, a placenta-expressed transcript on chromosome 18, to detect imbalances in the ratio of polymorphic alleles in trisomy 18 fetuses. DNA and RNA extraction from placental tissue samples was performed using a QIAamp DNA Mini Kit (Qiagen, Hilden, Germany) according to the manufacturer's instructions. Extracted RNA samples were treated with DNase I (Invitrogen) to eliminate genomic DNA contamination. Genotyping of the rs6098 SNP on the SERPINB2 gene was performed on placental tissue DNA samples using a homogeneous MassEXTEND (hME) assay using MassARRAY Compact (Sequenom, San Diego) as described above.
[0190] Reverse transcription of SERPINB2 transcripts was performed on placental tissue RNA samples using ThermoScript reverse transcriptase (Invitrogen) with the gene-specific primer 5'-CGCAGACTTCTCACCAAACA-3' (Integrated DNA Technologies, Coralville, IA). All cDNA samples were diluted to a concentration such that subsequent digital PCR amplification could be performed at an average concentration of one template molecule per reaction well. Digital PCR was prepared using TaqMan Universal PCR Master Mix (Applied Biosystems, Foster City, CA) and Biomark™ PCR Reagents (Fluidigm, San Francisco). The forward primer 5'-CTCAGCTCTGCAATCAATGC-3' (Integrated DNA Technologies) and the reverse primer (identical to the gene-specific primer for reverse transcription) were used at a concentration of 600 nM. The two TaqMan probes targeting the A or G allele of the rs6098 SNP on SERPINB2 were 5'-(FAM)CCACAGGGAATTATTT(MGBNFQ)-3' and 5'-(FAM)CCACAGGGGATTATTT(MGBNFQ)-3' (Applied Biosystems). FAM is 6-carboxyfluorescein, and MGBNFQ is a minor groove-binding non-fluorescent quencher. They were used at concentrations of 300 nM and 500 nM, respectively. Each sample-reagent mixture was dispensed into 756 reaction wells on a Biomark™ 12.765 Digital Array using a Nanoflex™ IFC Controller (Fluidigm). The array was then placed in a Biomark™ Real-time PCR System (Fluidigm) for thermal amplification and fluorescence detection. The reaction began with 2 min at 50°C, followed by 5 min at 95°C, followed by 45 cycles of 15 s at 95°C and 1 min at 59°C.After amplification, the number of informative wells (wells positive for only one of the A or G alleles) and the number of wells positive for both alleles were counted and subjected to sequential probability ratio test (SPRT) analysis.
[0191] In a heterozygous polyploid fetus, the A and G alleles should be equally present in the fetal genome (1:1), whereas in trisomy 18, an extra copy of one allele is present, resulting in a ratio of 2:1 in the fetal genome. A series of DPRT curves were generated to interpret the different samples. These curves represent the ratio P of the number of informative wells positive for the overrepresented allele to the total number of informative wells (x-axis) required for classification. r (y-axis) is plotted. For each sample, the experimentally derived P r is the expected P r The values were compared. Samples that fell above the upper curve were classified as trisomy 18, while samples that fell below the lower curve were classified as euploid. The area between the two curves is the unclassifiable region.
[0192] The feasibility of digital RNA-SNP analysis for detecting fetal trisomy 18 was demonstrated using the rs6098 SNP in the SERPINB2 gene. Placental tissue DNA samples from subjects carrying euploid and trisomy 18 fetuses were first genotyped by mass spectrometry to identify heterozygous specimens. Nine of the placental samples were found to be euploid and three were heterozygous for trisomy 18, which were then subjected to digital RNA-SNP analysis. In each sample, P r and m r is calculated, and this m r The SPRT curves corresponding to the values were used for disease classification. As shown in Figure 16A, all samples were correctly classified. P of trisomy 18 placentas r The values were above the unclassifiable region, whereas in the case of euploid placentas they were below this region.
[0193] m rSamples with SPRT curves based on Λ = 0.1, 0.2, and 0.3 are illustrated in Figure 16B. These data suggest that this digital RNA-SNP method is a useful diagnostic method for diagnosing pregnancies with trisomy 18. Samples with data points above the upper curve were classified as aneuploid, while samples with data points below the lower curve were classified as euploid.
[0194] C. Digital RCD analysis This example demonstrates the efficiency of digital PCR analysis using microfluidics-based digital PCR. One version of this approach is described here using the Fluidigm BioMark™ System, which is capable of performing over 9,000 digital PCR cycles in a single reaction.
[0195] Placental tissue, maternal blood cells, and amniotic fluid samples were collected from pregnant women carrying euploid or trisomy 21 (T21) fetuses. Placental DNA from 10 T21 and 10 euploid fetuses was mixed with an equal amount of euploid maternal blood cell DNA to create 20 50% DNA mixtures. To ensure accurate fetal fractions in the mixed samples, the extracted DNA was first quantified by measuring optical density (OD) at 260 nm. It was then digitally quantified using 12.765 Digital Arrays (Fluidigm) with a BioMark™ System (Fluidigm). The assays for quantifying the samples were performed similarly to those described below, except that probes for reference chromosomes were used.
[0196] The chromosome dosages in the 50% DNA mixture and amniotic fluid samples were determined by digital PCR analysis comparing non-polymorphic loci on chromosomes 1 and 21. 101-bp amplicons of a pair of paralogous loci on chromosomes 21 and 1 were first simultaneously amplified with the forward primer 5'-GTTGTTCTGCAAAAAACCTTCGA-3' and the reverse primer 5'-CTTGGCCAGAAATACTTCATTACCATAT-3'. Two chromosome-specific TaqMan probes were designed to distinguish the paralogs on chromosomes 21 and 1, with the sequences 5'-(FAM)TACCTCCATA ATGAGTAAA(MGBNFQ)-3' and 5'-(VIC)CGTACCTCTGTAATGTGTAA(MGBNFQ)-3', respectively. Paralogous loci are used here for illustrative purposes only. In other words, non-paralogous loci can also be used in such analyses.
[0197] To demonstrate the use of the digital RCD approach to detect trisomy 18 (T18), another assay was designed targeting paralogous sequences on chromosomes 21 and 18. 128-bp amplicons of the paralogous loci on chromosomes 21 and 18 were first simultaneously amplified with the forward primer 5'-GTACAGAAACCACAAACTGATCGG-3' and the reverse primer 5'-GTCCAGGCTGTGGGCCT-3'. Two chromosome-specific TaqMan probes were designed to distinguish the paralogs on chromosomes 21 and 18, with the sequences 5'-(FAM)AAGAGGCGAGGCAA(MGBNFQ)-3' and 5'-(VIC)AAGAGGACAGGCAAC(MGBNFQ)-3', respectively. The use of paralogous loci here is for illustrative purposes only. In other words, non-paralogous loci can also be used in such analyses.
[0198] All experiments were performed on a BioMark™ System (Fluidigm) using 12.765 Digital Arrays (Fluidigm). Reactions on one panel were set up in a 10 μl reaction volume using 2x TaqMan Universal PCR Master Mix (Applied Biosystems). Each reaction contained 1x TaqMan Universal PCR Master Mix, 900 nL of each primer, 125 nM of probe, and 3.5 μl of 50% placental / maternal blood cell DNA sample. The sample / assay mixture was injected into the digital array using a NanoFlex™ IFC controller (Fluidigm). Reactions were run on the BioMark™ System for detection. The reaction began with 2 minutes at 50°C, followed by 10 minutes at 95°C, followed by 40 cycles of 15 seconds at 95°C and 1 minute at 57°C.
[0199] Euploid and T21 50% placental / maternal blood cell DNA samples were analyzed on the digital array using the chr21 / chr1 assay. For each sample, the number of informative wells containing only one (but not both) marker positive was counted. The proportion of overrepresented markers among all informative wells (P = 0.01) was calculated. r The accurate mean reference template concentration per well (m r ) is the SPRT curve corresponding to the experimentally obtained P r This approach was used to determine whether the sample was euploid or T21. When unclassifiable samples remained, data from additional panels was accumulated until classification was possible. As shown in Figure 17, all 50% placental / maternal blood cell DNA samples were correctly classified using this approach, and their classification required a range of 1 to 4 panels. SPRT curves were also plotted to indicate the decision boundaries for correct classification, as shown in Figure 18.
[0200] The inventors further applied RCD analysis to amniotic fluid samples obtained from 23 pregnant women carrying euploid fetuses and 6 pregnant women carrying T21 fetuses. Each sample was analyzed in a single 765-well with the chr21 / chr1 assay. Figure 19 summarizes this SPRT classification. As shown in Figure 19, all 29 samples were correctly classified. Therefore, digital RCD is an alternative approach to detecting trisomy 21 using microsatellite (Levett LJ, et al. A large-scale evaluation of amnio-PCR for the rapid prenatal diagnosis of fetal trisomy. Ultrasound Obstet Gynecol 2001; 17: 115-8) or single nucleotide polymorphism (SNP) markers (Tsui NB, et al. Detection of trisomy 21 by quantitative mass spectrometric analysis of single-nucleotide polymorphisms. Clin Chem 2005; 51: 2358-62), or real-time non-digital PCR in various samples used for prenatal diagnosis, such as amniotic fluid and chorionic villus biopsy.
[0201] In an attempt to detect T18 specimens, we utilized the chr21 / chr18 assay in three euploid and five T18 placental DNA samples. The proportion of overrepresented markers in all informative wells (P r ) was calculated. Except for one T18 specimen that was misclassified as euploid, all other specimens were correctly classified. The results are summarized in Figure 20.
[0202] V. Use of Multiplexed Digital RCD Assays on a Mass Spectrometry Platform The number of plasma DNA molecules present per unit volume of maternal plasma is limited (Lo YMD. et al. 1998 Am J Hum Genet 62, 768-7758). For example, during early pregnancy, the median maternal plasma concentration of the autosomal locus, the β-globin gene, has been shown to be 986 copies / ml, including both maternal and fetal DNA (Lo YMD. et al. 1998 Am J Hum Genet 62, 768-7758). To obtain 7680 molecules, it may be necessary to extract DNA from approximately 8 ml of maternal plasma. This volume of plasma is equivalent to the amount obtained from approximately 15 ml of maternal blood, which is close to the limit of routine implementation. However, the inventors envision that a set of chr21 and reference chromosome targets can be combined for digital RCD analysis. For five pairs of chr21 and reference chromosome targets, only 1.6 ml of maternal plasma may be required to provide the number of template molecules required for analysis. Multiplex single-molecule PCR can be performed. The robustness of this multiplex single-molecule analysis has previously been demonstrated in single-molecule haplotyping (Ding, C. and Cantor, CR. 2003 Proc Natl Acad Sci USA 1W, 7449-7453).
[0203] In one example, placental tissue and maternal blood cell samples were obtained from pregnant women carrying euploid or trisomy 21 (T21) fetuses. Five euploid and five T21 placental DNA samples were mixed with equal amounts of maternal blood cell DNA to create 10 DNA mixtures, each mimicking a plasma sample containing 50% fetal DNA. To ensure accuracy of the fetal fraction in the mixed samples, the extracted DNA was first quantified by measuring optical density (OD) at 260 nm. They were then digitally quantified by real-time PCR in a 384-well format. The assay for quantifying the samples was performed similarly to the digital RCD analysis example described above.
[0204] The chromosome dosage in the 50% mixture was determined by digital PCR analysis, comparing non-polymorphic loci between chromosomes 1 and 21. This method is called digital relative chromosome dosage (RCD) analysis. A 121 bp (including 10 bp on each primer) amplicon of a pair of paralogous loci on chromosomes 21 and 1 was first co-amplified with the forward primer 5'-ACGTTGGATGGTTGTTCTGCAAAAAACCTTCGA-3' and the reverse primer 5'-ACGTTGGATGCTTGGCCAGAAATACTTCATTACCATAT-3'. An extension primer was designed to target the base difference between chromosomes 21 and 1, and its sequence was: It was 5'-CTCATCCTCACTTCGTACCTC-3'.
[0205] To demonstrate the utility of a multiplex digital PCR assay for detecting T21 specimens, another digital RCD assay was designed targeting paralogous sequences on chromosomes 21 and 18. A 148 bp (including 10 bp on each primer) amplicon of a pair of paralogous loci on chromosomes 21 and 18 was first simultaneously amplified with the forward primer 5'-ACGTTGGATGGTACAGAAACCACAAACTGATCGG-S' and the reverse primer 5'-ACGTTGGATGGTCC AGGCTGTGGGCCT-3'. An extension primer was designed targeting the base difference between chromosomes 21 and 18, and its sequence was: It was 5'-ACAAAAGGGGGAAGAGG-3'.
[0206] Multiplex digital RCD analysis was performed using a primer extension protocol. PCR reactions were prepared in a 5 μl reaction volume using the GeneAmp PCR Core Reagent Kit (Applied Biosystems). Each reaction contained 1x Buffer II, 2 mM MgCl2, 200 μM dNTP mix, 0.2 U of AmpliTaq Gold, 200 nM each of the four primers, and 50% DNA mix. The assay / sample mix was loaded into a 384-well PCR plate, and the reaction began with 2 min at 50°C, followed by 10 min at 95°C, followed by 40 cycles of 15 s at 95°C and 1 min at 57°C.
[0207] The PCR products were treated with shrimp alkaline phosphatase (SAP) to remove free dNTPs. The mixture was incubated at 37°C for 40 minutes, followed by 85°C for 5 minutes. A primer extension reaction was then performed. Briefly, 771 nM of extension primer for the chr21 / chr1 assay, 1.54 nM of extension primer for the chr21 / chr18 assay, 0.67 U of Thermosequenase (Sequenom), and 64 μM each of ddCTP, ddGTP, dATP, and dTTP in an extension cocktail were added to the SAP-treated PCR products. The reaction conditions were 94°C for 2 minutes, followed by 80 cycles of 94°C for 5 seconds, 50°C for 5 seconds, and 72°C for 5 seconds. For the final wash, 16 μl of water and 3 mg of Clean Resin (Sequenom) were added to the extension products. The mixture was mixed for 20–30 min on a rotator, followed by centrifugation at 351 g for 5 min. 15–25 nL of the final product was dispensed onto a SpectroCHIP (Sequenom) using a MassARRAY Nanodispenser S (Sequenom). Data acquisition from the SpectroCHIP was performed in a MassARRAY Analyzer Compact Mass Spectrometer (Sequenom). Mass data were transferred to MassARRAY Typer (Sequenom) software for analysis.
[0208] Five euploid and five T21 50% placental / maternal DNA samples were analyzed using the duplex RCD assay. For each sample, the number of informative wells from each assay, i.e., wells positive for only one of the chr21, chr1, or chr18 markers, was counted. The proportion of wells positive for the chr21 marker among all informative wells (P = 0.01) was calculated. r ) was calculated separately for each RCD assay. Then, P r A sequential probability ratio test (SPRT) was used to determine whether a β-ploid was indicative of a euploid or T21 sample, which reduced the number of wells required since each well was counted twice.
[0209] The Chr21 / chr1 assay is typically performed first. If any samples remain unclassifiable, the values obtained from the Chr21 / chr18 assay may be added for further calculations. Additional plates can be used for remaining unclassifiable samples until classification is possible. As shown in Figure 21, all 50% euploid mixed samples were correctly classified using a single 384-well plate. Some T21 samples required more than one plate for correct classification. If only one assay were used, more plates would have been needed to ensure the necessary number of informative wells for classification to be achieved. For example, sample N0230 could not be classified using either RCD assay alone. However, correct classification was achieved when the two assays were combined to obtain data. If this dual RCD assay had not been used, additional plates would have been required for analysis. The inventors intend to further reduce the number of wells by using a higher level of assay multiplexing.
[0210] In another example, the inventors developed a quadruple assay targeting four different amplicons on chromosome 21 and their corresponding paralogous partners located on autosomes other than chromosome 21. This quadruple assay was used in digital RCD analysis prior to SPRT typing of samples from euploid and trisomy 21 pregnancies. DNA extraction from placenta samples was performed using a QIAamp tissue kit (Qiagen, Hilden, Germany).
[0211] All placental and maternal buffy coat DNA samples used in this study were first quantified using a NanoDrop spectrophotometer (NanoDrop Technology, Wilmington, DE). DNA concentrations were converted to genome equivalents (GE) / μL using a conversion of 6.6 pg / cell. DNA concentrations corresponding to approximately one template per well were determined by serial dilution of the DNA samples. Under these conditions, the inventors would expect approximately 37% of wells to be negative for amplification. For multiplex digital RCD analysis, four sets of paralogous sequence targets were selected: Paralogous loci on chromosomes 21 and 1 were co-amplified with the forward primer 5′-ACGTTGGATGTTGATGAAGTCTCATCTCTACTTCG 3′ and the reverse primer 5′-ACGTTGGATGCAATAAGCTTGGCCAGAAATACT-3′, resulting in an 81-bp amplicon. Paralogous loci on chromosomes 21 and 7 were simultaneously amplified with the forward primer 5′-ACGTTGGATGGAATTTAAGCTAAATCAGCCTGAACTG-3′ and the reverse primer 5′-ACGTTGGATGGTTTCTCATAGTTCATCGTAGGCTTAT-3′, resulting in an 82-bp amplicon. Paralogous loci on chromosomes 21 and 2 were co-amplified with the forward primer 5′-ACGTTGGATGTCAGGCAGGGTTCTATGCAG-3′ and the reverse primer 5′-ACGTTGGATGAGGCGGCTTCCTGGCTCTT-3′, resulting in a 101-bp amplicon. Paralogous loci on chromosomes 21 and 6 were simultaneously amplified with the forward primer 5′-ACGTTGGATGGCTCGTCTCAGGCTCGTAGTT-3′ and the reverse primer 5′-ACGTTGGATGTTTCTTCGAGCCCTTCTTGG-3′, resulting in a 102-bp amplicon. Each reaction contained 10x Buffer II (Applied Biosystems), MgCl2, and 100 nM of each primer. The total reaction volume was 5 μL / well. The reaction began with 5 min at 95°C, followed by 45 cycles of 30 s at 95°C, 30 s at 62°C, and 30 s at 72°C, followed by a final 7 min at 72°C. All PCR amplifications were performed using the GeneAmp PCR Core Reagent Kit (Applied Biosystems). Free nucleotides were inactivated with shrimp alkaline phosphatase (SAP). Each reaction contained 10x SAP buffer (Sequenom) and SAP enzyme (Sequenom). 2 μl of SAP mixture was added to each PCR. The SAP reactions were incubated at 37°C for 40 min and 85°C for 5 min. After SAP treatment, primer extension reactions of the PCR products were performed using the iPLEX Gold kit (Sequenom). Paralogous sequence mismatches (PSMs) on chromosomes 21 and 1 were interrogated with the extension primer 5'-GTCTCATCTCTACTTCGTACCTC-3'. PSMs on chromosomes 21 and 7 were interrogated with the extension primer 5'-TTTTACGCTGTCCCCATTT-3'. PSMs on chromosomes 21 and 2 were interrogated with the extension primer 5'-GGTCTATGCAGGAGCCGAC-3'. PSMs on chromosomes 21 and 6 were interrogated with the extension primer 5'-TGGGCGCGGGAGCGGACTTCGCTGG-3'. Each reaction contained 10x PLEX buffer (Sequenom), iPLEX stop mix (Sequenom), iPLEX enzyme (Sequenom), and 343 nM of each extension primer, except that the extension primers for PSMs on chromosomes 21 and 6 were used at 1.3 μM. 2 μl of iPLEX mixture was added to 5 μl of PCR product. The iPLEX reaction was cycled according to a 200-cycle program. Briefly, the sample was first denatured at 94°C for 35 seconds, followed by annealing at 52°C for 5 seconds and extension at 80°C for 5 seconds.The annealing and extension cycles were repeated four more times, with five cycles each, followed by a 5-second denaturation step at 94°C, followed by another five-cycle annealing and extension loop. The five annealing and extension cycles and one denaturation step were repeated 39 times for a total of 40 cycles. A final extension step was performed at 72°C for 3 minutes. The iPLEX reaction products were diluted with 16 μl of water and desalted with 6 mg of resin for each PCR. The 384-well plates were centrifuged at 1600 g for 3 minutes and aliquoted onto a SpectroCHIP (Sequenom) for matrix-assisted laser desorption / ionization time-of-flight (MALDI-TOF) mass spectrometry analysis (Sequenom).
[0212] The number of wells positive for only one of either chromosome 21 or the reference chromosome in each of the four assays was recorded independently. Poisson-corrected counts of chromosome 21 and the reference chromosome were calculated for each assay. The sum of the Poisson-corrected counts of chromosome 21 and the reference chromosome across all four assays was calculated and considered the informative count for the quadruplicate assay. P r The value was calculated as the count for chromosome 21 in the quadruplicate assay divided by the sum of the counts for chromosome 21 and the reference chromosome in the quadruplicate assay. r The values were subjected to SPRT analysis. One or more 384-well plates were analyzed until SPRT classification was possible. Two 50% placental genomic DNA / 50% maternal buffy coat DNA mixtures and two 50% trisomy 21 placental genomic DNA / 50% maternal buffy coat DNA mixtures were analyzed.
[0213] The experimentally derived P r The value is the predicted P r value to test the null or alternative hypothesis. Alternatively, P r If the required level of statistical confidence for disease classification was not yet reached, neither the null nor the alternative hypothesis could be accepted. These specimens were considered unclassifiable until further data were obtained.
[0214] The results and SPRT classification of each sample are shown in Figures 22A and 22B. The two euploid samples required two and five 384-well digital RCD analyses, respectively, before SPRT classification was possible. Independent data from each component of the quadruple assays did not allow for SPRT classification of either sample. Both trisomy 21 samples were correctly classified using only a single 384-well digital RCD analysis. Similarly, independent data from each component of the quadruple assays did not allow for SPRT classification of either sample. However, the combined counts obtained from the quadruple assays allowed for correct SPRT classification. These data demonstrate that the use of multiplex digital RCD substantially increases the effective number of informative counts for a given number of digital PCR analyses compared to the use of a single-plex digital RCD assay.
[0215] VI. Use of Digital Epigenetic Relative Chromosome Dosage Here, we introduce an approach called digital epigenetic relative chromosome dosage (digital ERCD). In ERCD, epigenetic markers representing fetal-specific DNA methylation patterns or other epigenetic changes on the chromosome involved in chromosomal aneuploidy (e.g., chromosome 21 in trisomy 21) and a reference chromosome are subjected to digital PCR analysis. The ratio of the number of wells positive for the chromosome 21 epigenetic marker to the number of wells positive for the reference chromosome epigenetic marker in plasma DNA extracted from a pregnant woman carrying a normal fetus can provide a reference range. This ratio is expected to increase if the fetus has trisomy 21. It will be clear to those skilled in the art that one or more chromosome 21 markers and one or more reference chromosome markers can be used in this analysis.
[0216] An example of a gene on chromosome 21 that exhibits a fetal (placenta)-specific methylation pattern is the holocarboxylase synthetase (HLCS) gene. HLCS is hypermethylated in the placenta but hypomethylated in maternal blood cells. This is within the scope of U.S. Patent Application No. 11 / 784,499, which is incorporated herein by reference. An example of a gene on a reference chromosome that exhibits a fetal (placenta)-specific methylation pattern is RASSF1A on chromosome 3
[10] . RASSF1A is hypermethylated in the placenta but hypomethylated in maternal blood cells. See U.S. Patent Application No. 11 / 784,501, which is incorporated herein by reference.
[0217] To detect fetal trisomy 21 using digital PCR (DPCR) using maternal plasma cells to detect hypermethylated HLCS and hypermethylated RASSF1A, maternal peripheral blood cells are first collected. The blood is then centrifuged to collect plasma. DNA is then extracted from the plasma using techniques well known to those skilled in the art, such as the QIAamp Blood kit (Qiagen). The plasma DNA is then digested with one or more methylation-sensitive restriction enzymes, such as HpaII and BstUI. These methylation-sensitive restriction enzymes cleave unmethylated maternal genes but not hypermethylated fetal genes. The digested plasma DNA sample is then diluted to a level where an average of approximately 0.2 to 1 molecule of intact HLCS or RASSF1A sequences can be detected in the reaction well after restriction enzyme digestion. Two real-time PCR systems can be used to amplify the diluted DNA. One uses two primers covering the region where the HLCS gene should be cleaved by the restriction enzyme if the HLCS gene is not methylated, and one TaqMan probe specific to the HLCS gene. The other uses two primers and one probe for RASSF1A. An example of a primer / probe set for RASSF1A is described in Chan et al. (2006), Clin Chem 52, 2211-2218. The TaqMan probes targeting HLCS and RASSF1A may have different fluorescent reporters, such as FAM and VIC, respectively. Therefore, one 384-well plate may be sufficient to perform this digital PCR experiment. The number of wells scoring positive for only either HLCS or RASSF1A can be counted, and the ratio of these counts can be obtained. This HLCS:RASSF1A ratio is expected to be higher in maternal plasma from women carrying fetuses with trisomy 21 compared to maternal plasma from women carrying normal euploid fetuses, and the degree of overrepresentation may depend on the average reference template concentration per well in the digital PCR reaction.
[0218] Other methods can be used to score these results, such as counting the number of wells that are HLCS positive regardless of RASSF1A positivity; and conversely, counting the number of wells that are RASSF1A positive regardless of HLCS positivity. Furthermore, instead of calculating the ratio, either the sum or difference of HLCS and RASSF1A counts can be used to indicate the trisomy 21 status of the fetus.
[0219] Apart from performing digital PCR in plates, it is also clear to those skilled in the art that other digital PCR variants can be used, such as microfluidics chips, nanoliter PCR microplate systems, emulsion PCR, polony PCR and rolling circle amplification, primer extension, and mass spectrometry, etc. These digital PCR variants are listed for illustrative purposes and not for limiting purposes.
[0220] Apart from real-time PCR, it is clear to those skilled in the art that methods such as mass spectrometry can also be used to score the results of digital PCR.
[0221] Apart from the use of methylation-sensitive restriction enzymes to distinguish between fetal and maternal forms of HLCS and RASSF1A, it is clear to those skilled in the art that other methods for determining methylation status are also applicable, such as bisulfite modification, methylation-specific PCR, immunoprecipitation using anti-methylated cytosine antibodies, mass spectrometry, etc.
[0222] It will also be apparent to those skilled in the art that the approaches described in this and other embodiments of the present invention can also be used in other biological fluids in which fetal DNA may be found, including maternal urine, amniotic fluid, transcervical washings, chorion, maternal saliva, etc.
[0223] VII. Massively Parallel Genome Sequencing Using Emulsion PCR and Other Strategies Here, we describe another example in which digital readout of nucleic acid molecules can be used to detect fetal chromosomal aneuploidies, such as trisomy 21, in maternal plasma. Fetal chromosomal aneuploidies are caused by abnormal amounts of chromosomes or chromosomal regions. To minimize false diagnoses, noninvasive tests are desirable for high sensitivity and specificity. However, fetal DNA is present at absolutely low concentrations in maternal plasma and serum, accounting for only a small fraction of total DNA. Therefore, the number of digital PCR samples targeting specific loci cannot be increased indefinitely within the same sample. Therefore, analysis of multiple sets of specific target loci can be used to increase the amount of data that can be obtained from a single sample without increasing the number of digital PCR samples performed.
[0224] Thus, some embodiments enable non-invasive detection of fetal chromosomal aneuploidies by maximizing the amount of genetic information that can be inferred from the limited amount of fetal nucleic acid present in a biological sample containing maternal background nucleic acid. In one aspect, the amount of genetic information obtained is sufficient to make a correct diagnosis, but the cost and input amount of biological sample required is not too high.
[0225] Massively parallel sequencing, which can be achieved by the 454 platform (Roche) (Margulies, M. et al. 2005 Nature 437, 376-380), the Illumina Genome Analyzer (or Solexa platform), or the SOLiD system (Applied Biosystems), or Helicos True Single Molecule DNA sequencing technology (Harris TD et al. 2008 Science, 320, 106-109), Pacific Biosciences' single molecule, real-time (SMRT™) technology, and nanopore sequencing (Soni GV and Meller A. 2007 Clin Chem 53: 1996-2001), allows the sequencing of many nucleic acid molecules isolated from a sample in a highly multiplexed manner (Dear Brief Funct Genomic Proteomic 2003; 1: 397-416). Each of these platforms sequences single molecules of nucleic acid fragments that have not been clonally expanded or even amplified.
[0226] Since each reaction generates a large number of sequence data from each sample, on the order of 100,000 to 1 million, or even millions or billions, the resulting sequenced data forms a representative profile of the mixture of nucleic acid species in the original specimen. For example, the haplotype, transcriptome, and methylation profile of the sequenced data are similar to those of the original specimen (Brenner et al. Nat Biotech 2000;18:630-634; Taylor et al. Cancer Res 2007;67:8511-8518). To sample a large number of sequences from each specimen, the number of identical sequences, for example, those resulting from sequencing a nucleic acid pool with several-fold coverage or high redundancy, is also a good quantitative indication of the count of a specific nucleic acid species or locus in the original sample.
[0227] In one embodiment, random sequencing is performed on DNA fragments present in the plasma of a pregnant woman, thereby obtaining genomic sequences originating from either the mother or the fetus. Random sequencing involves sampling (sequencing) a random portion of the nucleic acid molecules present in a biological sample. Because the sequences are random, a different subset (fraction) of nucleic acid molecules (i.e., the genome) can be sequenced in each analysis. In some embodiments, this subset can work even when it varies from sample to sample and from analysis to analysis, even when the same sample is used. Examples of fractions are about 0.1%, 0.5%, or 1% of the genome. In other embodiments, the fraction is at least one of these values.
[0228] Bioinformatics techniques can then be used to place each of these DNA sequences into the human genome. Because such sequences are present in repetitive regions of the human genome or in regions that are subject to inter-individual variation, such as copy number variation, the proportion of such sequences may be excluded from subsequent analysis. The amount of the chromosome of interest and the amount of one or more other chromosomes can thus be determined.
[0229] In one embodiment, a parameter (e.g., fractional representation) of a chromosome potentially involved in chromosomal aneuploidy, e.g., chromosome 21, chromosome 18, or chromosome 13, can thus be calculated from the results of a bioinformatics procedure. The fractional representation can be obtained based on the abundance of the entire chromosome sequence (e.g., some measurement of all chromosomes, including the clinically important chromosome), or a specific subset of chromosomes (e.g., just one other than the one being tested).
[0230] In one embodiment, the parameters (e.g., fractional representation of clinically significant chromosomes) are then compared to reference ranges established for pregnancies involving normal (i.e., euploid) fetuses. In some variations of this procedure, the reference ranges (i.e., cutoff values) may be adjusted according to the fractional concentration (f) of fetal DNA in a particular maternal plasma sample. The value of f can be determined from a sequencing dataset, for example, using sequences mappable to the Y chromosome if the fetus is male. The value of f can also be determined by separate analysis, for example, using fetal epigenetic markers (Chan KCA et al. 2006 Clin Chem 52, 2211-8), or from analysis of single nucleotide polymorphisms.
[0231] In one embodiment, even when a pool of nucleic acids in a sample is sequenced with less than 100% gene coverage, the majority of each nucleic acid species is sequenced only once among the proportion of captured nucleic acid molecules. Similarly, the imbalance in the dosage of a particular locus or chromosome can be determined quantitatively. That is, the imbalance in the dosage of a locus or chromosome can be inferred from the percentage representation of that locus among other sequenced, mappable tags of the sample.
[0232] In one aspect of a massively parallel genome sequencing approach, representative data from all chromosomes can be produced simultaneously. The origin of specific fragments is not preselected. Sequencing is performed randomly, and then a database search is performed to confirm the origin of a specific fragment. Contrast this with the situation in which a specific fragment from chromosome 21 and another from chromosome 1 are amplified.
[0233] In one embodiment, the proportion of such sequences is obtained from the chromosome involved in the aneuploidy, e.g., chromosome 21 in this example. Additional sequences obtained by performing such sequencing may be derived from other chromosomes. By taking into account the relative size of chromosome 21 compared to other chromosomes, a normalized frequency of chromosome 21-specific sequences from such sequencing within the reference range can be obtained. If the fetus has trisomy 21, the normalized frequency of chromosome 21-derived sequences from such sequencing will increase, thus enabling the detection of trisomy 21. The degree of change in normalized frequency will depend on the fractional concentration of fetal nucleic acids in the analyzed sample.
[0234] In one embodiment, we used an Illumina genome analyzer for single-end sequencing of human genome and human plasma DNA samples. The Illumina genome analyzer sequences clonally expanded single DNA molecules captured on a solid surface called a flow cell. Each flow cell has eight lanes for sequencing eight individual specimens or pools of specimens. Each lane can generate approximately 200 Mb of sequence, representing only a portion of the 3 billion base pairs of sequence in the human genome. Each genomic DNA or plasma DNA sample was sequenced using one lane of the flow cell. The resulting short sequence tags were aligned with the human reference genome sequence, and chromosomal origins of replication were identified. The total number of individual sequenced tags aligned to each chromosome was tabulated and compared to the relative size of each chromosome expected from the reference human genome or a representative non-disease specimen. Chromosome gains or losses were then identified.
[0235] The described technique is only one example of the gene / chromosome dosage method described herein. Alternatively, paired-end sequencing can be performed. Instead of comparing the length of the sequenced fragments with that expected in the reference genome described by Campbell et al. (Nat Genet 2008;40:722-729), the number of aligned and sequenced tags is counted and sorted according to chromosomal location. Gain or loss of chromosomal regions or entire chromosomes was determined by comparing the number of tags with the expected chromosome size in the reference genome or that of a representative non-disease sample.
[0236] In another embodiment, the nucleic acid pool fraction to be sequenced in a single reaction is further selected before sequencing.For example, hybridization-type techniques, such as oligonucleotide arrays, can be used to initially select nucleic acid sequences from specific chromosomes, such as potentially aneuploid chromosomes and other chromosomes (single or multiple) that are not involved in the tested aneuploidy.In another example, a specific subgroup of nucleic acid sequences from a sample pool is selected or enriched before sequencing.For example, as discussed above, it has been reported that fetal DNA molecules in maternal plasma contain shorter fragments than maternal background DNA molecules (Chan et al. Clin Chem 2004;50:88-92).Therefore, one or more methods known to those skilled in the art can be used to fractionate nucleic acid sequences in a sample according to molecular size, for example, by gel electrophoresis, or size exclusion column, or by microfluidics. Alternatively, in the example of analyzing cell-free fetal DNA in maternal plasma, the fetal nucleic acid portion can be amplified by methods that suppress maternal background, such as by the addition of formaldehyde (Dhallan et al. JAMA 2004;291:1114-9). In one embodiment, a portion or subset of the pool of additionally selected nucleic acids is randomly sequenced.
[0237] Other single molecule sequencing methods, such as the Roche454 platform, the Applied Biosystems SOLiD platform, the Helicos True Single Molecule DNA sequencing technology, Pacific Biosciences' Single Molecule Real Time (SMRT®) technology, and nanopore sequencing methods, may be used in this application as well.
[0238] Examples of results and further discussion (e.g., sequencing and calculation parameters) can be found in the concurrently filed application, "DIAGNOSING FETAL CHROMOSOMAL ANEUPLOIDY USING GEOMIC SEQUENCING," (Attorney Docket No. 016285-005220US), which is incorporated by reference. The methods described herein for determining cutoff values can be applied when the reaction is a sequencing reaction, as described in this section.
[0239] The determination of the fractional concentration of fetal DNA in maternal plasma can also be performed separately from sequencing. For example, the Y chromosome DNA concentration can be determined in advance using real-time PCR, microfluidic PCR, or mass spectrometry. In fact, fetal DNA concentration can be determined using applicable loci in female fetuses other than the Y chromosome. For example, Chan et al. showed that fetal-derived methylated RASSF1A sequences are detected in the plasma of pregnant women in the background of maternal-derived unmethylated RASSF1A sequences (Chan et al., Clin Chem 2006;52:2211-8). Therefore, the fractional fetal DNA concentration can be determined by dividing the amount of methylated RASSF1A sequences by the total RASSF1A (methylated and unmethylated) sequences.
[0240] It is expected that maternal plasma is more preferable than maternal serum in carrying out the present invention, because DNA is released from maternal blood cells during blood clotting.Therefore, if serum is used, it is expected that the fractional concentration of fetal DNA will be lower in maternal plasma than in maternal serum.That is, if maternal serum is used, it is expected that more sequences will be required to diagnose fetal chromosomal aneuploidy compared with the plasma sample obtained from the same pregnant woman.
[0241] Yet another alternative method for determining the fractional concentration of fetal DNA can be through the quantification of polymorphic differences between pregnant women and fetuses (Dhallan R, et al. 2007 Lancet, 369, 474-481). An example of this method can target polymorphic sites where the pregnant woman is homozygous and the fetus is heterozygous. The amount of fetal-specific alleles can be compared with the amount of common alleles to determine the fractional concentration of fetal DNA.
[0242] In contrast to existing techniques for detecting chromosomal abnormalities, including comparative genomic hybridization, microarray comparative genomic hybridization, and quantitative real-time polymerase chain reaction, which detect and quantify one or more specific sequences, massively parallel sequencing does not rely on the detection or analysis of a predetermined or defined set of DNA sequences. Random, representative fragments of DNA molecules from a pool of specimens are sequenced. The number of distinct sequence tags aligned to various chromosomal regions was compared between specimens containing or not containing tumor DNA. Chromosomal abnormalities can be revealed by differences in the number (or percentage) of sequences aligned to a given chromosomal region in the specimens.
[0243] In another example, sequencing technology for cell-free DNA of plasma can be used to detect chromosomal abnormalities in plasma DNA for the detection of specific cancers.Different cancers have a set of typical chromosomal abnormalities.Changes (amplification and deletion) in multiple chromosomal regions can be used.Therefore, the proportion of sequences aligned to amplified regions can be increased, and the proportion of sequences aligned to reduced regions can be decreased.The percentage representation per chromosome can be compared with the size of each corresponding chromosome in a reference genome, which is expressed as the percentage of the genome representation of any given chromosome relative to the whole genome.Direct comparison or comparison with a reference chromosome can also be used.
[0244] VIII. Mutation Detection Fetal DNA in maternal plasma exists as a small population of fetal origin, accounting for an average of 3-6% of maternal plasma DNA. Therefore, prior art techniques have focused on detecting DNA targets that the fetus inherits from the father and that can be distinguished from the large maternal DNA background in maternal plasma. Examples of such previously detected targets include the SRF gene on the Y chromosome (Lo YMD et al. 1998 Am J Hum Genet, 62, 768-775) and the RHD gene if the mother is RhD-negative (Lo YMD et al. 1998 N Engl J Med, 339, 1734-1738).
[0245] Conventional strategies using maternal plasma to detect fetal mutations have been limited to autosomal dominant conditions in which the father is a carrier, and have excluded autosomal recessive conditions by direct mutation detection or by linkage analysis when the father and mother carry different mutations (Ding C. et al 2004 Proc Natl Acad Sci USA 101, 10762-10767). These conventional strategies have significant limitations. For example, when both the father and mother carry the same mutation, it becomes impossible to achieve meaningful prenatal diagnosis by direct mutation detection in maternal plasma.
[0246] Such a scenario is illustrated in Figure 23. In this scenario, there may be three possible fetal genotypes: NN, NM, and MM, where N represents the normal allele and M represents the mutant allele. Examples of mutant alleles include those responsible for cystic fibrosis, beta-thalassemia, alpha-thalassemia, sickle cell anemia, spinal muscular atrophy, and congenital adrenal hyperplasia. Other examples of such disorders can be found at Online Mendelian Inheritance in Man (OMIM) www.ncbi.nlm.nih.gov / sites / entrez?db=OMIM&itool=toolbar. In maternal plasma, most DNA is maternal and may be NM. Of the three fetal genotypes, there may not be a specific fetal allele that can be specifically detected in maternal plasma. Therefore, conventional strategies cannot be used in this case.
[0247] The embodiments described herein can address such scenarios. In a scenario where both the mother and fetus are NM, the N and M alleles are balanced. However, if the mother is NM and the fetus is NN, the N allele will be over-represented in maternal plasma, resulting in allelic imbalance. On the other hand, if the mother is NM and the fetus is MM, the M allele will be over-represented in maternal plasma, resulting in allelic imbalance. Thus, in detecting fetal mutations, the null hypothesis implies that there is no allelic imbalance when the fetal genotype is NM. The alternative hypothesis implies that there is allelic imbalance, and the fetal genotype is either NN or MM, depending on whether the N or M allele is over-represented.
[0248] The presence or absence of allelic imbalance can be determined using digital PCR as described herein. In a first scenario, suppose a particular volume of maternal plasma contains DNA released from 100 cells, 50 of which are maternal and 50 of which are fetal. The fractional concentration of fetal DNA in this volume of plasma is 50%. If the mother's genotype is NM, there will be 50 maternal N alleles and 50 M alleles. If the fetus's genotype is NM, there will be 50 maternal N alleles and 50 M alleles. Thus, there is no allelic imbalance between the N and M alleles, and they are present in 100 copies each. On the other hand, if the fetus's genotype is NN, there will be 100 fetal N alleles in this volume of plasma. Therefore, there will be 150 N alleles in total, and 50 M alleles. In other words, there is an allelic imbalance between N and M, with N being overrepresented relative to M by a ratio of 3:1.
[0249] Conversely, if the fetus's genotype is MM, there may be 100 fetal M alleles in this volume of plasma. Therefore, there may be a total of 150 M alleles and 50 N alleles. In other words, there is an allelic imbalance between N and M, with M overrepresented relative to N at a ratio of 3:1. Such allelic imbalance can be measured by digital PCR. This allele is considered the reference template. Similar to digital RNA-SNP and digital RCD analysis, the actual discrimination of alleles in a digital PCR experiment is governed by a Poisson probability density function. Thus, even if the theoretical degree of allelic imbalance in this scenario is 3:1, the expected degree of allelic imbalance may depend on the average template concentration per well during digital PCR analysis. Thus, the average reference template concentration per well (m r ) appropriate cutoff interpretation such as SPRT analysis should be used to classify specimens.
[0250] Furthermore, the degree of allelic imbalance to be measured depends on the fractional DNA concentration. In contrast to the above example, suppose a particular volume of maternal plasma contains DNA released from 100 cells, 90 of which are maternal and 10 of which are fetal. The fractional concentration of fetal DNA in that volume of plasma is then 10%. If the mother's genotype is NM, then there will be 90 maternal N alleles and 90 M alleles. If the fetus's genotype is NM, then there will be 10 fetal N alleles and 10 M alleles. Thus, in this case, there is no allelic imbalance between the N and M alleles, each with a total of 100 copies. On the other hand, if the fetus's genotype is NN, then there will be 20 fetal N alleles in this volume of plasma. This results in a total of 110 N alleles and 90 M alleles.
[0251] In other words, there is allelic imbalance between N and M, with N being overrepresented. Conversely, if the situation is reversed and the fetal genotype is MM, there will be 20 fetal M alleles in this volume of plasma. This results in 110 M alleles and 90 N alleles. In other words, there is allelic imbalance between N and M, with M being overrepresented. The theoretical allelic imbalance ratio for a 10% fetal DNA fraction is 110:90, which differs from the 3:1 ratio shown in the example above when 50% fetal DNA is present. Therefore, a cutoff interpretation, such as for SPRT analysis, appropriate for the fetal DNA fraction should be used to classify the sample.
[0252] Therefore, plasma DNA can be extracted and the amount of maternal and fetal DNA in the plasma sample can be quantified, for example, by conventional real-time PCR assays (Lo, et al. 1998 Am J Hum Genet 62, 768-775) or by other types of quantifiers known to those skilled in the art, such as SNP markers (Dhallan R et al. 2007 Lancet 369, 474-481) and fetal epigenetic markers (Chan KCA et al. 2006 Clin Chem 52, 2211-2218). The fetal DNA percentage can then be calculated. Then, during digital PCR analysis, the quantified plasma DNA sample is prepared (e.g., diluted or concentrated) so that each reaction well contains, on average, one molecule of template (which can be either the N or M allele). Digital PCR analysis can be performed using a pair of primers and two TaqMan probes, one specific for the N allele and the other for the N allele. The number of wells that are positive for only M and the number of wells that are positive for only N can be counted. The ratio of these wells can be used to determine whether there is evidence of allelic imbalance. Statistical evidence of allelic imbalance can be examined by methods well known to those of skill in the art, such as using SPRT. In one variation of this analysis, the number of wells that are positive for only M, or for both M and N, can be counted; and the number of wells that are positive for only N, or for both M and N, can be counted; and the ratio of these counts can be derived. Again, statistical evidence of allelic imbalance can be examined by methods well known to those of skill in the art, such as using SPRT.
[0253] Determination of fetal genetic mutation dosage, termed digital relative mutant dosage (RMD), was demonstrated using male / female (XX / XY) DNA mixtures. As depicted in Figure 24A, blood cell DNA from males or females was mixed with male DNA to prepare samples with fractional concentrations of 25% and 50% of the XX or XY genotype in an XY background.
[0254] Additionally, blood cell samples were obtained from 12 male and 12 female subjects. Female blood cell DNA (XX genotype) was mixed with a three-fold excess of male blood cell DNA (XY genotype), resulting in 12 DNA samples with 25% XX genotype DNA mixed in a 75% XY genotype background, as shown in Figure 24B.
[0255] The purpose of SPRT was to determine the minority genotype present in background DNA. In a DNA mixture containing 25% XX genotype DNA with a background of 75% XY genotype DNA, the minor allele could be Y, derived from 75% of the DNA. Because 25% of the DNA in the sample is from XX genotype, if there are 200 molecules of DNA in this sample, 150 molecules originate from an XY individual. Therefore, the expected number of Y alleles is 75. The number of X alleles derived from male DNA (XY genotype) is also 75. The number of X alleles derived from female DNA (XX genotype) is 50 (25 times 2). Therefore, the X:Y ratio is 125 / 75 = (1 + 25%) / (1 - 25%) = 5 / 3.
[0256] In the second part of this study, blood cell samples were obtained from male and female subjects carrying the HbE (G->A) and CD41 / 42 (CTTT / -) mutations in the beta globin gene, i.e., hemoglobin beta (HBB) gene. To mimic maternal plasma samples obtained from heterozygous (MN, where M = mutation and N = wild-type) mothers carrying male fetuses with all possible genotypes (MM, MN, or NN), blood cell DNA from men who were either homozygous for the wild-type allele (NN) or heterozygous for one of the two mutations (MN) was mixed with blood cell DNA samples from women heterozygous for the same mutation (MN). Thus, DNA mixtures with various male / female-differentiated DNA concentrations were prepared. Blood cell DNA from women homozygous for CD41 / 42 (MM) was also used to prepare the DNA mixtures. To confirm the actual proportion of male DNA used for SPRT typing, the fractional male DNA concentration of each DNA mixture was determined using the ZFY / X assay.
[0257] The digital ZFY / X assay was used to demonstrate SPRT and to determine the fractional male DNA concentration in DNA mixtures. The amounts of zinc finger protein sequences (ZFX and ZFY) on the X and Y chromosomes were determined by digital PCR analysis. First, 87-bp amplicons of the ZFX and ZFY loci were simultaneously amplified with the forward primer 5'-CAAGTGCTGGACTCAGATGTAACTG-3' and the reverse primer 5'-TGAAGTAATGTCAGAAGCTAAAACATCA-3'. Two chromosome-specific TaqMan probes were designed to distinguish the paralogues between the X and Y chromosomes: 5'-(VIC)TCTTTAGCACATTGCA(MGBNFQ)-3' and 5'-(FAM)TCTTTACCACACTGCAC(MGBNFQ)-3', respectively.
[0258] The amount of mutation in the DNA mixture was determined by digital PCR analysis of the normal allele versus the mutant allele. For the HbE mutation, the normal and mutant alleles were first simultaneously amplified with the forward primer 5'-GGGCAAGGTGAACGTGGAT-3' and the reverse primer 5'-CTATTGGTCTCCTTAAACCTGTCTTGTAA-3'. Two allele-specific TaqMan probes were designed to distinguish between the normal (G) and mutant (A) alleles: 5'-(VIC)TTGGTGGTGAGGCC(MGBNFQ)-3' and 5'-(FAM)TTGGTGGTAAGGCC(MGBNFQ)-3', respectively.
[0259] For the CD41 / 42 deletion mutation, 87-bp and 83-bp amplicons of the normal and mutant alleles were first coamplified using the forward primer 5'-TTTTCCCACCCTTAGGCTGC-3' and the reverse primer 5'-ACAGCATCAGGAGTGGACAGATC-3'. Two allele-specific TaqMan probes were designed to distinguish between the normal and mutant alleles (without and with deletion), with the sequences 5'-(VIC)CAGAGGTTCTTTGAGTCCT(MGBNFQ)-3' and 5'-(FAM)AGAGGTTGAGTCCTT(MGBNFQ)-3', respectively. The results for the HbE mutation are shown in Figures 26A and 26B.
[0260] These experiments were performed on a BioMark™ System (Fluidigm) using 12.765 Digital Arrays (Fluidigm). One panel reaction was prepared in a 10 μl reaction volume using 2x TaqMan Universal PCR Master Mix (Applied Biosystems). For the CD41 / 42 and ZFY / X assays, each reaction contained 1x TaqMan Universal PCR Master Mix, 900 nM of each primer, 125 nM of each probe, and 3.5 μL of 1 ng / μL DNA mixture. For the HbE assay, probes targeting the normal (G) and mutant (A) alleles were added at 250 nM and 125 nM, respectively. The sample / assay mixture was injected into the digital array using a NanoFlex™ IFC controller (Fluidigm). Reactions were run on the BioMark™ System for signal detection. The reaction began with 2 min at 50°C, followed by 10 min at 95°C, followed by 50 cycles of 15 s at 95°C and 1 min at 57°C (for ZFX / Y and CD41 / 42) or 56°C (for HbE). At least one reaction panel was used for each specimen, and if any samples remained unclassifiable, data were accumulated from additional panels until a decision could be made.
[0261] It may also be apparent to those skilled in the art that said digital PCR can be performed using methods well known to those skilled in the art, such as, for example, microfluidics chips, nanoliter PCR microplate systems, emulsion PCR, polony PCR, rolling circle amplification, primer extension and mass spectrometry.
[0262] IX. Cancer Example In one embodiment, the present invention can be implemented to classify samples according to the presence or absence of allele ratio distortion, which may occur in cancerous tumors. In one embodiment, the number of wells in each specimen with positive signals for only the A allele, only the G allele, and both alleles was determined by digital PCR. The reference sequence was defined as the allele with the fewer number of positive wells (in the unlikely event that the number of positive wells for both alleles was the same, both were used as reference alleles). The estimated mean concentration (m) of the reference allele per well was calculated. r ) was calculated using the sum of wells that were negative for the reference allele, regardless of whether the other allele was positive, according to a Poisson probability density function. We use a hypothetical example to illustrate this calculation.
[0263] In a 96-well reaction, 20 wells are positive for allele A, 24 wells are positive for allele G, and 28 wells are positive for both alleles. Allele A is considered the reference allele because there are fewer positive wells than the other alleles. The number of wells that are negative for the reference allele is 96-20-28=48. Therefore, m r is calculated using the Poisson distribution, which gives -ln(48 / 96) = 0.693.
[0264] In the context of LOH detection, the null hypothesis refers to samples in which the loss of one allele presumably leads to a loss of allele ratio distortion. Under this assumption, the expected ratio of the number of positive wells for the two alleles is 1:1, and therefore the proportion of informative wells (wells positive for only one allele) potentially containing the overrepresented allele is 0.5.
[0265] In the context of LOH detection, the alternative hypothesis refers to a sample in which the allele ratio is estimated to be distorted due to the loss of one allele in 50% of the cells in the sample. Because the ratio between the overrepresented allele and the reference allele is 2:1, the average concentration of the overrepresented allele per well may be twice the concentration of the reference allele. However, the number of wells positive for the overrepresented allele may not simply be twice the number of wells positive for the reference allele, but may follow a Poisson distribution.
[0266] Informative wells are defined as wells that are positive for either the A or G allele, but not both alleles. Calculation of the expected proportion of wells containing the overrepresented allele for samples with skewed allele ratios is similar to that shown in Table 600. In the example above, if LOH occurs in 50% of the tumor cells, the average concentration of G allele per well would be 2 times 0.693 = 1.386. If LOH occurs in more than 50% of the tumor cells, the average concentration of G allele per well would be calculated using the formula: 1 / [1-(proportion of LOH)] x m r can be followed.
[0267] The expected percentage of wells positive for the G allele is 1-C -1.386 = 0.75 (i.e., 75% or 72 wells). Assuming that wells testing positive for either the A or G allele are independent, 0.5 x 0.75 = 0.375 of those wells can be positive for both the A and G alleles. Therefore, 0.5 - 0.375 = 0.125 wells can be positive for only the A allele, and 0.75 - 0.375 = 0.375 wells can be positive for only the G allele. Thus, the proportion of informative wells can be 0.125 + 0.375 = 0.5. The expected proportion of informative wells carrying the G allele can be 0.375 / 0.5 = 0.75. This P r The predicted values of are used to construct an appropriate SPRT curve to measure the presence or absence of allelic ratio distortion (here, LOH) in the sample.
[0268] Furthermore, the actual proportion of informative wells carrying the non-reference allele, as determined experimentally by digital PCR analysis, was used to determine whether the null or alternative hypothesis was accepted, or whether further analysis with more wells might be required. r The decision boundary for was calculated based on a threshold likelihood ratio of 8, as this threshold likelihood ratio value has been shown to provide good performance in discriminating between samples with and without allelic imbalance in the context of cancer detection (Zhou, W, et al. (2001) Nat Biotechnol 19, 78-81; Zhou et al 2002, supra). In the above example, the number of informative wells is 20 + 24 = 44, and the experimentally obtained P r would be 24 / 44=0.5455. The decision boundary is <0.5879 to accept the null hypothesis and >0.6379 to accept the alternative hypothesis. Thus, the sample in this example can be classified as not having allele ratio distortion.
[0269] As a result, the inventors have outlined an approach for detecting sequence imbalances in a sample. In one embodiment, the invention can be used to noninvasively detect fetal chromosomal aneuploidies, such as trisomy 21, by analyzing fetal nucleic acids in maternal plasma. This approach can also be applied to other biomaterials containing fetal nucleic acids, including amniotic fluid, chorionic villus samples, maternal urine, endocervical samples, maternal saliva, and the like. First, the inventors demonstrated the use of the invention to determine allelic imbalances at SNPs in PLAC4 mRNA, which is placenta-specifically expressed by chromosome 21, in maternal plasma from pregnant women carrying trisomy 21 fetuses. Second, the inventors demonstrated that the invention can be used as a non-polymorphism-based method for noninvasive prenatal detection of trisomy 21 by relative chromosome dosage (RCD) analysis. Such a digital RCD-based approach involves directly assessing whether the total copy number of chromosome 21 in a sample containing fetal DNA is overrepresented relative to a reference chromosome. Without the need for complex equipment, digital PCR detection allows for the detection of trisomy 21 in samples containing 25% fetal DNA. The inventors applied the sequential probability ratio test (SPRT) to interpret the digital PCR data. Computer simulation analysis confirmed the high accuracy of this disease classification algorithm.
[0270] The inventors further outlined that this approach can be applied to the determination of other forms of nucleic acid sequence imbalances besides chromosomal abnormalities, such as the detection of fetal mutations or polymorphisms in maternal plasma, and the detection of local duplications and deletions in the genomes of malignant cells by analysis of tumor-derived nucleic acids in plasma.
[0271] Any software components or functions described in this application may be implemented using a suitable computer language, such as Java, C++, or Perl, using, for example, conventional or object-oriented techniques. The software code may be stored as a series of instructions or commands in a suitable medium, including a computer-readable medium for storage and / or transmission, such as random access memory (RAM), read-only memory (ROM), magnetic media such as a hard drive or floppy disk, or optical media such as a compact disc (CD) or DVD (digital versatile disc), flash memory, etc. The computer-readable medium may be any combination of such storage or transmission devices.
[0272] Such programs may also be encoded and transmitted using carrier signals adapted for transmission over wired, optical, and / or wireless networks conforming to various protocols, including the Internet. Thus, computer-readable media according to embodiments of the present invention may be created using data signals encoded with such programs. Such program-encoded computer-readable media may be packaged with a compatible device or provided separately from other devices (e.g., via Internet download). Any such computer-readable media may be provided on or within a single program product (e.g., a hard drive or an entire computer system), or may reside on or within different computer program products within a system or network. A computer system may include a monitor, printer, or other suitable display for providing a user with any of the results mentioned herein.
[0273] An example computer system is shown in FIG. 27. The subsystems shown in FIG. 27 are interconnected through a system bus 2775. Additional subsystems are shown, such as a printer 2774, a keyboard 2778, a fixed disk 2779, a monitor 2776 coupled to a display adapter 2782, and others. Peripherals and input / output (I / O) devices, coupled to an I / O controller 2771, may be connected to the computer system by any number of means known in the art, such as a serial port 2777. For example, the serial port 2777 or external interface 981 may be used to connect the computer system to a wide area network such as the Internet, a mouse input device, or a scanner. The interconnection via the system bus enables a central processor 2773 to communicate with the individual subsystems and to execute instructions from, and exchange information between, the system memory 2772 or fixed disk 2779. The system memory 2772 and / or fixed disk 2779 may embody computer-readable media.
[0274] The foregoing description of exemplary embodiments of the present invention is presented for purposes of illustration and description. The precise forms described are not intended to be exhaustive or to limit the invention, and many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described to best explain the principles of the invention and its practical applications so that others skilled in the art may optimally utilize the invention in various forms and with various modifications as suited to the particular use intended.
[0275] All applications, patents, and patent applications cited herein are hereby incorporated by reference in their entirety for all purposes.
Claims
1. 1. A method for determining the genotype at a first locus of a fetus of a pregnant woman, said fetus having a mother who is said pregnant woman, said mother being heterozygous at a first locus for a first allele and a second allele different from said first allele, said method comprising the steps of: obtaining data from a plurality of reactions comprising cell-free DNA from the maternal biological sample, wherein the biological sample comprises cell-free DNA from the woman and cell-free DNA from the fetus, wherein each reaction indicates the presence or absence of the first allele and the second allele from the woman or from the fetus, wherein the data comprises: (1) a first set of quantitative data indicating a first amount of a response positive for the presence of the first allele; and (2) a second set of quantitative data indicating a second amount of response positive for the presence of the second allele; Includes; determining a parameter from the two data sets, wherein the parameter provides a relative amount between the first amount and the second amount; and - comparing said parameters to one or more cut-off values to determine the genotype of said fetus, wherein said one or more cut-off values are determined based on a measured or estimated fractional concentration of fetal DNA in said biological sample; A method comprising:
2. 2. The method of claim 1, wherein the fetus is determined to have only the first allele at the first locus when the first amount is greater than the second amount by a degree specified by one of the one or more cutoff values.
3. 2. The method of claim 1, wherein the fetus is determined to have only the second allele at the first locus when the second amount is greater than the first amount by a degree specified by one of the one or more cutoff values.
4. 2. The method of claim 1, wherein the fetus is determined to be heterozygous for the first allele and the second allele when the first amount and the second amount are identified as statistically equal by the one or more cutoff values.
5. 2. The method of claim 1, wherein the first and second alleles are single nucleotide polymorphism alleles.
6. 10. The method of claim 1, wherein the first allele is a mutant allele.
7. 7. The method of claim 6, wherein the mutant allele causes cystic fibrosis, beta thalassemia, alpha thalassemia, sickle cell anemia, spinal muscular atrophy, or congenital adrenal hyperplasia.
8. 7. The method of claim 6, wherein the first allele comprises a deletion or an amplification.
9. 10. The method of claim 1, further comprising determining one or more cutoff values based on the measured fractional concentration of fetal DNA in the biological sample.
10. The method of claim 1 , wherein the reaction is a sequencing reaction or an amplification reaction.
11. 11. The method of claim 10, wherein the plurality of reactions are polymerase chain reactions, wherein the plurality of reactions contain an average of one first allele per reaction, and wherein the plurality of reactions contain an average of one second allele per reaction, and wherein the first amount is the number of reactions indicating the presence of the first allele, and wherein the second amount is the number of reactions indicating the presence of the second allele.
12. 2. The method of claim 1, wherein the biological sample is plasma, blood, urine, transcervical effluent, chorionic villi, or maternal saliva.
13. The method of claim 1 , wherein the parameter is the ratio or difference between the first amount and the second amount.
14. 10. The method of claim 1, further comprising determining that the mother is heterozygous at a first locus for the first allele and the second allele.
15. 10. The method of claim 1, wherein the fetus does not have aneuploidy.
16. A computer readable medium storing a plurality of instructions for controlling one or more processors to perform operations, the instructions designed to perform the method of any one of claims 1 to 15.
17. A computer system comprising one or more processors and a memory storing a plurality of instructions for controlling the processors designed to perform the method of any one of claims 1 to 15.
Citation Information
Patent Citations
Non-invasive prenatal diagnosis
JP2001513648A
A method for the detection of chromosomal aneuploidies
WO2006097049A1